Improving speech recognition of mobile devices
Abstract
Speech recognition in mobile processor-based devices may be improved by using location information. Location information may be derived from on-board hardware or from information provided remotely. The location information may assist in a variety of ways in improving speech recognition. For example, the ability to adapt to the local ambient conditions, including reverberation and noise characteristics, may be enhanced by location information. In some embodiments, pre-developed models or context information may be provided from a remote server for given locations.

Term
Term ended
Expired 10 June 2023, 3.3 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
27 claims: 3 independent, 24 dependent
- 1A method of adapting speech recognition processing comprising:obtaining information about the location of a mobile device;and using the obtained location information to adapt speech recognition processing;wherein obtaining information includes obtaining information about the location of nearby speakers being potential sources of interference.
- 11An article comprising a medium storing instructions that, if executed, enable a processor-based system to perform the steps of:obtaining information about the location of a mobile device;and using the obtained location information to adapt speech recognition processing;wherein obtaining information includes obtaining information about the location of nearby speakers being potential sources of interference.
- 21A speech recognition system comprising:a processor;a position detennining device coupled to said processor for obtaining information about the location of the system;and a storage coupled to said processor storing instructions that enable the processor to use the obtained location information to adapt speech recognition processing;wherein said system automatically obtains information about the location of nearby speakers being potential sources of interference.
Independent claims3
29 paragraphs in 4 sections, as filed
BACKGROUND
0001This invention relates generally to mobile processor-based systems that include speech recognition capabilities.
0002Mobile processor-based systems include devices such as handheld devices, personal digital assistants, digital cameras, laptop computers, data input devices, data collection devices, remote control units, voice recorders, and cellular telephones, to mention a few examples. Many of these devices may include speech recognition capabilities.
0003With speech recognition, the user may say words that may be converted to text. As another example, the spoken words may be received as commands that enable selection and operation of the processor-based system's capabilities.
0004In a number of cases, the ability of a given device to recognize speech or identify a speaker is relatively limited. A variety of ambient conditions may adversely affect the quality of the speech recognition or speaker identification. Because the ambient conditions may change unpredictably, the elimination of ambient effects is much more difficult with mobile speech recognition platforms.
0005Thus, there is a need for better ways to enable speech recognition with mobile processor-based systems.
0006One approach to this problem has been to use location information. European Patent Application <patcit id="pcit0001" dnum="EP1326232A"><text>EP-A-1326232</text></patcit> discloses a speech recognition system for mobile processor-based devices in which location information is used to select automatically a corresponding acoustic noise model for adaptation of speech recognition processing.
SUMMARY OF THE INVENTION
0007The present invention provides a method of adapting speech recognition processing of the kind, such as disclosed in <patcit id="pcit0002" dnum="EP1326232A"><text>EP-A-1326232</text></patcit>, comprising: obtaining information about the location of a mobile device; and using the obtained location information to adapt speech recognition processing. Notably, as set out in the appended claims, the obtaining of information includes obtaining information about nearby speakers.
0008The present invention also provides an article comprising a medium storing instructions that, if executed, enable a processor-based system to perform the steps of the method just mentioned.
0009The present invention further provides a speech recognition system of the kind, such as disclosed in <patcit id="pcit0003" dnum="EP1326232A"><text>EP-A-1326232</text></patcit>, comprising: a processor; a position determining device coupled to said processor; and a storage coupled to the processor storing instructions that enable the processor to use location information to adapt speech recognition processing. Notably, as set out in the appended claims, the system automatically obtains information about nearby speakers.
BRIEF DESCRIPTION OF THE DRAWINGS
0010<ul id="ul0001" list-style="none" compact="compact"><li><figref idref="f0001">FIG. 1</figref> is a schematic depiction of one embodiment of the present invention;</li><li><figref idref="f0002">FIG. 2</figref> is a flow chart useful with the embodiment shown in <figref idref="f0001">FIG. 1</figref> in accordance with one embodiment of the present invention; and</li><li><figref idref="f0003">FIG. 3</figref> is a flow chart useful with the embodiment shown in <figref idref="f0001">FIG. 1</figref> in accordance with one embodiment of the present invention.</li></ul>
DETAILED DESCRIPTION
0011Referring to <figref idref="f0001">FIG. 1</figref>, a speech enabled mobile processor-based system 14 may be any one of a variety of mobile processor-based systems that generally are battery powered. Examples of such devices include laptop computers, personal digital assistants, cellular telephones, digital cameras, data input devices, data collection devices, appliances, and voice recorders, to mention a few examples.
0012By incorporating a position detection capability within the device 14, the ability to recognize spoken words may be improved in the variety of enviromnents or ambient conditions. Thus, the device 14 may include a position detection or location-based services (LBS) client 26. Position detection may be accomplished using a variety of technologies such as global positioning satellites, hot-spot detection, cell detection, radio triangulation, or other techniques.
0013A variety of aspects of location may be used to improve speech recognition. The physical location of the system 14 may provide information about acoustic characteristics of the surrounding space. Those characteristics may include the size of the room, noise sources, such as ventilation ducts or exterior windows, and reverberation characteristics.
0014This data can be stored in a network infrastructure, such as a location-based services (LBS) server 12. For frequently visited locations, the characteristics may be stored in the system 14 data store 28 itself. The server 12 may be coupled to the system 14 through a wireless network 18 in one embodiment of the present invention.
0015Other aspects of location that may be leveraged to improve speech recognition include the physical location of nearby speakers who are using comparable systems 14. These speakers may be potential sources of interference and can be identified based on their proximity to the user of the system 14. In addition, the identity of nearby people who are carrying comparable systems 14 may be inferred by subscribing to their presence information or by ad hoc discovery peers. Also, the orientation of the system 14 may be determined and this may provide useful information for improving speech recognition.
0016The system 14 includes a speech context manager 24 that is coupled to the position detection/location-based services client 26, a speech recognizer 22, and a noise mitigating speech preprocessor 20.
0017When speech recognition is attempted by the system 14, the speech context manager 24 retrieves a current context from the server 12 in accordance with one embodiment of the present invention. Based on the size of the surrounding space, the context manager 24 adjusts the acoustic models of the recognizer 22 to account for reverberation.
0018This adjustment may be done in a variety of ways including using model adaptation, such as maximum likelihood linear regression to a known target. The target transformation may have been estimated in a previous encounter at that position or may be inferred from the reverberation time associated with the space. The adjustment may also be done by selecting from a set of previously trained acoustic models that match various acoustic spaces typically encountered by the user.
0019As another alternative, the context manager 24 may select from among feature extraction and noise reduction algorithms that are resistant to reverberation based on the size of the acoustic space. The acoustic models may also be modified to match the selected front-end noise reduction and feature extraction. Models may also be adapted based on the identity of nearby people, retrieving and loading speaker dependent acoustic models for each person, if available. Those models may be used for automatic transcription of hallway discussion in one embodiment of the present invention.
0020Another way that the adjustment may be done is by initializing and adapting a new acoustic model if the acoustic space has not been encountered previously. Once the location is adequately modeled, the system 14 may send the information to the server 12 to be stored in the remote data store 16 for future visitors to the same location.
0021As another example of adaptation, based on the identity of nearby speakers, the system 14 may assist the user to identify them as a transcription source. A transcription source is someone whose speech should be transcribed. A list of potential sources in the vicinity of the user may be presented to the user. The user may select the desired transcription sources from the list in one embodiment.
0022As still another example, based on the orientation of the system 10, the location of proximate people, and their designation as transcription sources, a microphone array controlled by preprocessor 20 may be configured to place nulls in the direction of the closest persons who are not transcription sources. Since that direction may not be highly accurate and is subj ect to abrupt change, this method may not supplant interferer tracking via a microphone array. However, it may provide a mechanism to place the nulls when the interferer is not speaking, thereby significantly improving performance when an interferer talker starts to speak.
0023Referring to <figref idref="f0002">Figure 2</figref>, in accordance with one embodiment of the present invention, the speech context manager 24 may be a processor-based device including both a processor and storage for storing instructions to be executed on the processor. Thus, the speech context manager 24 may be software or hardware. Initially, the speech context manager 24 retrieves a current context from the server 12, as indicated in block 30. Then the context manager 24 may determined the size of the surrounding space proximate to the device 14, as indicated in block 32. The device 14 may adjust the recognizer's acoustic models to account for local reverberation, as indicated in block 34.
0024Then feature extraction and noise reduction algorithms may be selected based on the understanding of the local environment, as indicated in block 36. In addition, the speaker-dependent acoustic models for nearby speakers may be retrieved and loaded, as indicated in block 38. These models may be retrieved, in one embodiment, from the server 12.
0025New acoustic models may be developed based on the position of the system 14 as detected by the position detection/LBS client 26, as indicated in block 40. The new model, linked to position coordinates, may be sent over the wireless network 18 to the server 12, as indicated in block 42, for potential future use. In some embodiments, models may be available from the server 12 and, in other situations, those models may be developed by a system 14 either on its own or in cooperation with the server 12 for immediate dynamic use.
0026As indicated in block 44, any speakers whose speech should be recognized may be identified. The microphone array preprocessor 20 may be configured, as indicated in block 46. Then speech recognition may be implemented, as indicated in block 48, having obtained the benefit of the location information.
0027Referring to <figref idref="f0003">Figure 3</figref>, the LBS server 12 may be implemented through software 50 in accordance with one embodiment of the present invention. The software 50 may be stored in an appropriate storage on the server 12. Initially, the server 12 receives a request for context information from a system 14, as determined in diamond 52. Once received, the server 12 obtains the location information from the system 14, as indicated in block 54. The location information may then be correlated to available models in the data storage 16, as indicated in block 56. Once an appropriate model is identified, the context may be transmitted to the device 14 over the wireless network, as indicated in block 58.
0028While the present invention has been described with respect to a limited number of embodiments, those skilled in the art will appreciate numerous modifications and variations therefrom.
0029The scope of the present invention is defined by the appended claims.
Contents4
3 sheets
Sheet 1 Sheet 2 Sheet 3
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| EP1326232A | Cites | European Patent Office (EPO) |
| UMING KO ET AL: "DSP for the third generation wireless communications" COMPUTER DESIGN, 1999. (ICCD '99). INTERNATIONAL CONFERENCE ON AUSTIN, TX, USA 10-13 OCT. 1999, LOS ALAMITOS, CA, USA,IEEE COMPUT. SOC, US, 10 October 1999 (1999-10-10), pages 516-520, XP010360534 ISBN: 0-7695-0406-X | Non-patent | – |
| DROPPO J, ET AL.: "Efficient on-line Acoustic Environment Estimation for FCDCN in a Continuous Speech Recognition System" INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING 2001, 2001, pages 209-212, XP002252982 Salt Lake City (USA) | Non-patent | – |
16 members in 9 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 176326 | United States of America | – | |
| 17632602 | United States of America | A | |
| 0318408 | United States of America | W |
Members16
| Document | Office | Kind | |
|---|---|---|---|
| US2003236099A1 | United States of America | A1 | |
| WO2004001719A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003245443A1 | Australia | A1 | |
| TW200412730A | Taiwan Province of China | A | |
| KR20050007429A | Republic of Korea | A | |
| EP1514259A1 | European Patent Office (EPO) | A1 | |
| TWI229984B | Taiwan Province of China | B | |
| CN1692407A | China | A | |
| US7224981B2 | United States of America | B2 | |
| KR20070065893A | Republic of Korea | A | |
| KR100830251B1 | Republic of Korea | B1 | |
| EP1514259B1This record | European Patent Office (EPO) | B1 | |
| AT465485T | Austria | T | |
| ATE465485T1 | Austria | T1 | |
| DE60332236D1 | Germany | D1 | |
| CN1692407B | China | B |
54 legal events, as 7 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Gb: european patent ceased through non-payment of renewal feeCeasedGBPC | GBPC | EP | |
| Application deemed withdrawn, or ip right lapsed, due to non-payment of renewal feeWithdrawnR119 | R119 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| Notification of lapseLapsedST | ST | FR | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent ceasedCeasedPL | PL | CH | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Discontinued in the netherlands as no translation has been filedVDEP | VDEP | NL | |
| Corresponds to:REF | REF | EP | |
| European patents granted designating irelandGrantedFG4D | FG4D | IE | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedFG4D | FG4D | GB | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Request for extension of the european patent (deleted)DAX | DAX | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAX | AX | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 1514259
- Application
- 37390838
Titles3
- German
- VERBESSERUNG DER SPRACHERKENNUNG VON MOBILGERÄTEN
- English
- IMPROVING SPEECH RECOGNITION OF MOBILE DEVICES
- French
- AMELIORATION DE LA RECONNAISSANCE DE LA PAROLE DE DISPOSITIFS MOBILES
Classification
- CPC, 5
- G10L15/20
- G10L15/30
- H04M2250/74
- G10L2015/228
- H04M1/72457
- IPC, 4
- G10L15 26
- G10L15 20
- G10L15 22
- G10L15 28
Designated states27
- Contracting states, 27
- Austria
- Belgium
- Bulgaria
- Switzerland
- Cyprus
- Czechia
- Germany
- Denmark
- Estonia
- Spain
- Finland
- France
- United Kingdom
- Greece
- Hungary
- Ireland
- Italy
- Liechtenstein
- Luxembourg
- Monaco
- Netherlands (Kingdom of the)
- Portugal
- Romania
- Sweden
and 3 moreShow fewer
- Slovenia
- Slovakia
- Türkiye