Data transmission over a coded voice channel
Abstract
A digital input symbol is transmitted to a receiver by determining one or more formant frequencies that correspond to the digital input symbol. In one embodiment, a pre-programmed addressable memory is used to map the set of possible digital input symbols onto a set of corresponding speech units, each comprising a superposition of one or more formant frequencies. A signal is then generated having the speech units. The signal is supplied for transmission over a voice channel. This may include supplying the signal to a voice coder prior to transmission. In another aspect of the invention, a forward error correction code (FEC) is determined for the digital input symbol, and the one or more speech units are modified as a function of the forward error correction code. In this way, the FEC may also be transmitted with the encoded input symbol. The modification may affect any of a number of attributes of the speech units, including a volume attribute and a pitch attribute.

Term
Term ended
Projected expiry passed 8 December 2018, 7.8 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
11 claims: 3 independent, 8 dependent
- 1PATENDINÕUDLUS 1. Meetod digitaalse sisendsümboli edastamiseks vastuvõtjale, mis koosneb järgmistest sammudest:5 ühe või enama digitaalsele sisendsümbolile vastava formantsageduse kindlaksmääramine;üht või enamat formantsagedust sisaldava signaali genereerimine;ning signaali väljastamine saatmiseks kõnekanali vahendusel. 10
- 2Meetod vastavalt nõudluspunktile 1, kus signaali väljastamise samm saatmiseks kõnekanali vahendusel hõlmab sammu signaali andmiseks kõnekoodrisse.
- 3Meetod vastavalt nõudluspunktile 1, kus ühe või enama digitaalsele sisendsümbolile vastava formantsageduse kindlaksmääramise samm koosneb järgmistest 15 sammudest:digitaalse sisendsümboli andmine adresseeritavate mäluvahendite aadresssisendporti, kusjuures adresseeritavatesse mäluvahenditesse on talletatud formantsageduste koodid aadressidele nii, et kui digitaalne sisendsümbol antakse adresseeritavate mäluvahendite aadress-sisendporti, tekib adresseeritavate mäluvahendite väljundporti 20 vastav formantsageduse kood;ning adresseeritava mälu väljundporti tekkiva vastava formantsageduse koodi kasutamine ühe või enama formantsageduse indikaatorina.
- 4Meetod vastavalt nõudluspunktile 3, kus vastav formantsageduse kood viitab 25 hulgale formantsagedustele, ning mille kohaselt üht või enamat formantsagedust sisaldava signaali genereerimise samm sisaldab hulga vastavas formantsageduse koodis ära näidatud formantsageduste genereerimist. 30 5. Meetod vastavalt nõudluspunktile 1, kus üht või enamat formantsagedust sisaldava signaali genereerimise samm sisaldab:edastuse veaparanduse koodkoodi määramist digitaalse sisendsümboli jaoks;üht või enamat formantsagedust sisaldava signaali genereerimine;ning signaali modifitseerimist edastuse veaparanduse koodkoodi funktsioonina. EE 200000341 Α 6. Meetod vastavalt nõudluspunktile 5, kus ühe või enama formantsageduse modifitseerimise samm hõlmab ühe või enama formantsageduse helitugevuse atribuudi modifitseerimist edastuse veaparanduse koodkoodi funktsioonina.
- 55 7. Meetod vastavalt nõudluspunktile 5, kus ühe või enama formantsageduse modifitseerimise samm hõlmab ühe või enama formantsageduse helikõrguse atribuudi modifitseerimist edastuse veaparanduse koodkoodi funktsioonina.
- 68. Meetod kõne ja digitaalsete sisendsümbolite edastamiseks vastuvõtjale, koosnedes 10 järgmistest sammudest:kõne edastamine vastuvõtjale kõnekanali vahenditega;edastusmoodi muutmine, genereerides selleks automaatselt ettemääratud formantsageduste järjendi ning edastades automaatselt genereeritud formantsagedused vastuvõtjale kõnekanali vahenditega;ning 15 digitaalsete sisendsümbolite sidumine vastavasse formantjärjendisse ning seejärel vastavat formantjärjendit esitava signaali edastamine vastuvõtjale kõnekanali vahenditega.
- 79. Meetod vastavalt nõudluspunktile 8, kus digitaalsete sisendsümbolite vastavasse formantjärjendisse sidumise samm sisaldab digitaalse sisendsümboli andmise sammu 20 adresseeritavate mäluvahendite aadress-sisendporti, kusjuures adresseeritavatesse mäluvahenditesse on talletatud formantsageduste koodid aadressidele nii, et kui digitaalne sisendsümbol antakse adresseeritavate mäluvahendite aadress-sisendporti, tekib adresseeritavate mäluvahendite väljundporti vastav formantsageduse kood. 25 10. Meetod vastavalt nõudluspunktile 8, sisaldades veel kõneedastuse moodi tagasipöördumise sammu, genereerides selleks teise automaatselt ettemääratud formantsageduste järjendi ning edastades vastuvõtjale kõnekanali vahenditega teise automaatselt genereeritud formantsageduste järjendi. 30 11. Meetod juhtsignaalide genereerimiseks kõnega juhitava automaatserveri juhtimiseks, koosnedes järgmistest sammudest:hääldatud käskluste muundamine esimeseks juhtsignaaliks;esimese käsklussignaali andmine kõnetuvastuse vahenditele;EE 200000341 Α kõnetuvastuse vahendite kasutamine esimesele käsklussignaalile vastava ühe või rohkema formantsageduse määramiseks, kusjuures üks või rohkem formantsagedusi moodustavad käskluse, mis on automaatserverile äratuntav;ning teise käsklussignaali genereerimine, mis sisaldab üht või rohkemat 5 formantsagedust. 12. Meetod digitaalse sisendsümboli vastuvõtuks, koosnedes järgmistest sammudest: üht või rohkemat formantsagedust sisaldava signaali, mis on eelnevalt määratud vastama digitaalsele sisendsümbolile ning on modifitseeritud eelnevalt edastatud 10 digitaalsete sisendsümbolite funktsioonina, vastuvõtt;vastuvõetud modifikatsiooni kindlaksmääramine eelnevalt edastatud digitaalsete sisendsümbolite funktsioonina;vastuvõetud modifikatsiooni kasutamine vastuvõetud signaalis pöördmodifikatsiooni sooritamiseks, seeläbi genereerides pöördmodifitseeritud signaali;15 digitaalse sisendsümboli kindlaksmääramine, mis vastab ühele või enamale formantsagedusele, mis sisalduvad pöördmodifitseeritud signaalis;ning vastuvõetud modifikatsiooni kasutades signaali genereerimine, mis on indikatiivne kindlaksmääratud digitaalse sisendsignaali õigsuse suhtes. 20 13. Meetod vastavalt nõudluspunktile 12, kus digitaalse sisendsümboli, mis vastab pöördmodifitseeritud signaalis sisalduvaile ühele või enamale formantsagedusele, kindlaksmääramise samm koosneb: eelnevalt digitaalsele sisendsümbolile vastama määratud ühe või enama formantsageduse detekteerimise sammust;ning 25 digitaalse sisendsümboli, mis vastab ühele või enamale detekteeritud formantsagedusele, kindlaksmääramise sammust. 14. Meetod vastavalt nõudluspunktile 12, kus digitaalse sisendsümboli, mis vastab pöördmodifitseeritud signaalis sisalduvale ühele või enamale formantsagedusele, 30 kindlaksmääramise samm sisaldab automaatse kõnetuvastuse tehnika kasutamist digitaalse sisendsümboli, mis vastab ühele või enamale pöördmodifitseeritud signaalis sisalduvale formantsagedusele, kindlaksmääramiseks. 15. Seade digitaalse sisendsümboli edastamiseks vastuvõtjale, mis sisaldab: EE 200000341 Α vahendeid ühe või enama digitaalsele sisendsümbolile vastava formantsageduse määramiseks;vahendeid üht või enamat formantsagedust sisaldava signaali genereerimiseks;ning vahendeid signaali väljastamiseks saatmiseks kõnekanali vahendusel. 16. Seade vastavalt nõudluspunktile 15, kus vahendid signaali väljastamiseks saatmiseks kõnekanali vahendusel sisaldavad vahendeid signaali andmiseks kõnekoodrisse. 17. Seade vastavalt nõudluspunktile 1, kus vahendid ühe või enama digitaalsele 10 sisendsümbolile vastava formantsageduse kindlaksmääramiseks sisaldavad: vahendeid digitaalse sisendsümboli andmiseks adresseeritavate mäluvahendite aadress-sisendporti, kusjuures adresseeritavatesse mäluvahenditesse on talletatud formantsageduste koodid aadressidele nii, et kui digitaalne sisendsümbol antakse adresseeritavate mäluvahendite aadress-sisendporti, tekib adresseeritavate mäluvahendite 15 väljundporti vastav formantsageduse kood;ning vahendeid adresseeritava mälu väljundporti tekkiva vastava formantsageduse koodi kasutamiseks ühe või enama formantsageduse indikaatorina. 18. Seade vastavalt nõudluspunktile 17, kus vastav formantsageduse kood viitab 20 hulgale formantsagedustele, ning mille kohaselt vahendid üht või enamat formantsagedust sisaldava signaali genereerimiseks sisaldavad hulga vastavas formantsageduse koodis ära näidatud formantsageduste genereerimise vahendeid. 25 19. Seade vastavalt nõudluspunktile 15, kus vahendid üht või enamat formantsagedust sisaldava signaali genereerimiseks sisaldavad: vahendeid edastuse veaparanduse koodi kindlaksmääramiseks digitaalse sisendsümboli jaoks;vahendeid üht või enamat formantsagedust sisaldava signaali genereerimiseks;ning 30 vahendeid signaali modifitseerimiseks edastuse veaparanduse koodi funktsioonina. 20. Seade vastavalt nõudluspunktile 19, kus vahendid ühe või enama formantsageduse modifitseerimiseks sisaldavad vahendeid ühe või enama formantsageduse helitugevuse atribuudi modifitseerimiseks edastuse veaparanduse kood funktsioonina. EE 200000341 Α 21. Seade vastavalt nõudluspunktile 19, kus vahendeid ühe või enama formantsageduse modifitseerimiseks hõlmavad vahendeid ühe või enama formantsageduse helikõrguse atribuudi modifitseerimiseks edastuse veaparanduse koodi funktsioonina. 5 22. Seade kõne ja digitaalsete sisendsümbolite edastamiseks vastuvõtjale, mis sisaldab: vahendeid kõne edastamiseks vastuvõtjale kõnekanali vahenditega;vahendeid edastusmoodi muutmiseks, genereerides selleks automaatselt ettemääratud formantsageduste järjendi ning edastades automaatselt genereeritud formantsagedused vastuvõtjale kõnekanali vahenditega;ning
- 810 vahendeid digitaalsete sisendsümbolite sidumiseks vastavasse formantjärjendisse ning seejärel vastavat formantjärjendit esitava signaali edastamiseks vastuvõtjale kõnekanali vahenditega. 23. Seade vastavalt nõudluspunktile 22, kus vahendid digitaalsete sisendsümbolite
- 915 vastavasse formantjärjendisse sidumiseks sisaldavad vahendeid digitaalse sisendsümboli andmiseks adresseeritavate mäluvahendite aadress-sisendporti, kusjuures adresseeritavatesse mäluvahenditesse on talletatud formantsageduste koodid aadressidele nii, et kui digitaalne sisendsümbol antakse adresseeritavate mäluvahendite aadress-sisendporti, tekib adresseeritavate mäluvahendite väljundporti vastav formantsageduse kood. 24. Meetod vastavalt nõudluspunktile 22, sisaldades veel vahendeid kõneedastuse moodi tagasipöördumiseks, genereerides teise automaatselt ettemääratud formantsageduste järjendi ning edastades vastuvõtjale kõnekanali vahenditega teise automaatselt genereeritud formantsageduste järjendi. 25. Seade juhtsignaalide genereerimiseks kõnega juhitava automaatserveri juhtimiseks, mis sisaldab:vahendeid hääldatud käskluste muundamiseks esimeseks käsklussignaaliks;kõnetuvastuse vahendeid, mis on seatud esimest käsklussignaali vastu võtma ühe 30 või rohkema formantsageduse kindlaksmääramiseks, mis vastavad esimesele käsklussignaalile, kusjuures üks või rohkem formantsagedusi moodustavad käskluse, mis on automaatserverile äratuntav;ning vahendeid teise käsklussignaali genereerimiseks, mis sisaldab üht või rohkemat formantsagedust. EE 200000341 Α 26. Seade digitaalse sisendsümboli vastuvõtuks, mis sisaldab: vahendeid signaali vastuvõtuks, mis sisaldab üht või rohkemat formantsagedust, mis on eelnevalt määratud vastama digitaalsele sisendsümbolile ning on modifitseeritud 5 eelnevalt edastatud digitaalsete sisendsümbolite funktsioonina;vahendeid vastuvõetud modifikatsiooni kindlaksmääramiseks eelnevalt edastatud digitaalsete sisendsümbolite funktsioonina;vahendeid vastuvõetud modifikatsiooni kasutamiseks vastuvõetud signaalis pöördmodifikatsiooni sooritamiseks, seeläbi genereerides pöördmodifitseeritud signaali;10 vahendeid digitaalse sisendsümboli kindlaksmääramiseks, mis vastab ühele või enamale formantsagedusele, mis sisalduvad pöördmodifitseeritud signaalis;ning vahendeid vastuvõetud modifikatsiooni kasutamiseks signaali genereerimisel, mis on indikatiivne kindlaksmääratud digitaalse sisendsignaali õigsuse suhtes. 15 27. Meetod vastavalt nõudluspunktile 26, kus vahendid digitaalse sisendsümboli, mis vastab pöördmodifitseeritud signaalis sisaldavaile ühele või enamale formantsagedusele, kindlaksmääramiseks sisaldavad: vahendeid eelnevalt digitaalsele sisendsümbolile vastama määratud ühe või enama formantsageduse detekteerimiseks;ning
- 1020 vahendeid digitaalse sisendsümboli, mis vastab ühele või enamale detekteeritud formantsagedusele, kindlaksmääramiseks. 28. Meetod vastavalt nõudluspunktile 26, kusjuures vahendid pöördmodifitseeritud signaalis sisalduvale ühele või enamale formatitsagedusele vastava digitaalse
- 1125 sisendsümboli kindlaksmääramiseks sisaldavad automaatse kõnetuvastuse vahendeid digitaalse sisendsümboli, mis vastab ühele või enamale pöördmodifitseeritud signaalis sisalduvale formantsagedusele, kindlaksmääramiseks.
Independent claims11
91 paragraphs in 4 sections, as filed
DATA TRANSFER IN CODED CODE
The present invention relates to a digital information communication technique, more particularly to a digital information communication technique in an encoded speech channel.
There is an increasing demand from customers for advanced telephone services, such as automated services, which can be accessed and controlled through control commands that are transmitted over a distance. As a result, techniques have been developed to provide access to services provided over communications networks. In the field of wireless communications, work in this area includes the development of Wireless Application Protocol (WAP), a layered communication protocol consisting of a network layer (e.g., transport and session layers) and an application environment consisting of micro-browser, scripting, phone enhancement and content formats. Part of WAP is Telephony Value Added Services (TelVAS), which is a secure way to access basic operating system-specific operating system and local subsystem functions of the subsystem, such as Call Control, Phonebook, Messaging, etc. through device-independent interface tools.
In fixed telephony networks, access to services offered on communications networks has been transformed into the use of Intelligent Networks, in which Service Access Points are network nodes that customers can access to use additional services. Access to network-based services that are independent of any traditional network operator is also becoming common. These nodes are implemented on the basis of service computers, which may be connected to independent computer networks (e.g. the Internet) and accessible from at least one communications network (e.g. a telephone network or a cellular network such as the European standard Global System for Mobile Communication (GSM)). A communications network (such as a public telephone network or a cellular network) is only used to provide access to these independent computer networks. In order to keep the services provided by network nodes more independent from traditional communication networks, access to the node through such a communication network may include both data (e.g., voice) and control signaling on the same channel (e.g., inband or same band signaling).
In the mobile communications system, operators generally provide a Short Message Service (Short Message Service)
Message Service (SMS) for sending short messages to a mobile terminal. Messages are controlled through a Short Message Service Center (SMS-C) server that:
EE 200000341 Α stores and transmits messages. The SMS service has several disadvantages in terms of exchanging control signals between the user terminal and the service node. For example, the SMS service does not give the sender any control over the delay or give any information about the status of the message. Moreover, the price level of SMS service varies significantly from operator to operator, and some operators keep prices at a level which makes the use of the service too expensive for many users. Another disadvantage is that different mobile operators provide interfaces other than the SMS-C interface servers outside the mobile network, which means that sending SMS messages is prevented from terminals belonging to other networks.
It is known how to establish separate voice and data connections between two terminals through multiple communication networks, one of which is a cellular telephone network. However, switching between the two modes is cumbersome and time consuming, which causes inconvenience to the user.
Although systems such as Internet Protocol (IP) can easily handle mixed voice and data, problems arise when a communication connection includes a cellular network, such as a GSM network. More specifically, in the latter case, the communication link includes a speech encoder optimized for human speech, and thus in-band modem signaling, for example, by using tone frequencies (e.g., Dual-Tone Multi-Frequency or DTMF), may result in slow data rates with increased risk of error. This is because the modem signal symbol is less predictable than the speech signal. Known methods of overcoming these difficulties carry the risk of becoming impractical for the user or, on the other hand, lead to technical solutions specific to each type of network involved. Moreover, speech encoders may in the future behave even more obstructively to DTMF signal throughput. Thus, a problem that requires an innovative solution is the same bandwidth signaling for communications over multiple communications networks where at least one of these networks includes speech coding.
PCT Publication No. W096 / 09708, Simultaneous Transmission of Speech and Data on a Mobile Communications System, by Hamalainen, describes the use of a voice channel in a wireless communication system for simultaneous voice and data transmission, and more particularly discloses a method and system for detecting non-speech periods of silence. , allowing data to be included in the transmitted frames. This publication further describes the acquisition of frames with information bits, allowing the separation of data and speech frames as viewed by the network. A characteristic of the described solution is that it depends on the radio interface protocol and in that
EE 200000341 Α voice and data separation facilities are integrated into the network. Such a solution is therefore unsuitable for solving the problem of simultaneous call and data transmission between the first user mobile terminal and the second service node when it is outside and independent of the communication network involved in the voice connection to the node connection.
From now on, the use of speech recognition methods to control user services by voice will be prevalent. A disadvantage of the known methods is the need to "train" the speech recognition system to learn specific vocabulary, language characteristics and even the speech characteristics of the person speaking.
EXECUTIVE SUMMARY
It is therefore an object of the present invention to provide a technique for matching voice data to be transmitted over an encoded voice channel in a radio interface of a cellular system (e.g. GSM) so that the radio interface can support signaling in the same frequency band as described above for terrestrial communication systems.
It is another object of the present invention to provide a common "language" for connecting to user service nodes that use a speech recognition system to control the interface.
According to one aspect of the invention, the foregoing and other objects are used in the art and apparatus for transmitting a digital input symbol to a transmitter. This is achieved by identifying one or more formant frequencies corresponding to the digital input symbol and generating a signal containing one or more formant frequencies. The signal can then be transmitted over a voice channel. Such a signal is more suitable for this purpose because it contains the formant frequencies for which the speech channel is adapted. For example, such a signal may be provided to a speech encoder which generates an encoded signal for transmission over a speech channel.
In another aspect of the invention, preprogrammed addressable memory is used to create links between input symbols and a plurality of corresponding formant frequencies. Specifically, the step of determining the formant frequencies corresponding to one or more digital input symbols comprises the steps of providing an address-input port of memory addressable memory means for a digital input symbol, wherein the addressable memory means itself has formant frequency code. Addressable
EE 200000341 Α the corresponding formant frequency code generated by the output of the storage means is then used to denote one or more formant frequencies specified.
In yet another aspect of the invention, the corresponding formant frequency code provides sequences of formant frequencies. The step of generating a signal containing one or more formant frequencies then comprises the step of generating a formant frequency sequence denoted by the corresponding formant frequency code.
In yet another aspect of the invention, the Forward Error Correction (FEC) code with formant frequencies is transmitted over the speech channel. More specifically, the transmission error correction code is assigned to a digital input symbol, and one or more formant frequencies are modified as a function of the transmission error correction code. A signal containing one or more modified formant frequencies is then generated for transmission in the speech channel. For example, the modification may affect one or more of the formant frequency volume or pitch attributes.
In yet another aspect of the invention, both speech and digital input symbols may be transmitted to a receiver. This includes transmitting the call to the recipient through a voice channel. If data is to be transmitted, the change to a data transmission mode is made automatically by generating a predetermined sequence of formant frequencies and automatically transmitting the generated formant frequencies to the receiver by voice channel means.
In another aspect of the invention, the reverse mode of voice transmission is accomplished automatically by generating a second imprint of a predetermined formant frequency and transmitting the automatically generated second formant frequency sequence to a receiver by means of a speech channel. Another sequence of predetermined formant frequencies is a mechanism that signals to the receiver the need to change the mode.
In another aspect of the invention, control signals may be generated to control a voice-controlled automatic server by converting said command into a first command signal, and provide a first command signal to speech recognition means. The speech recognition means are used to determine one or more formant frequencies corresponding to the first command signal, the one or more formant frequencies forming a command recognizable by the automatic server. A second command signal is then generated which consists of one or more formant frequencies. This feature allows almost all users to interface with an automated server because the user-spoken commands are "translated" to another set of formant frequencies to which the automated server is susceptible.
EE 200000341 Α
DESCRIPTION AND LIST OF DRAWINGS
The objects and the novelty of the invention will be understood by reading the following detailed description, taken in conjunction with the accompanying drawings, in which:
Figures ΙΑ, IB and IC are block diagrams illustrating exemplary embodiments of a communication device for transmitting input symbols in a voice channel in accordance with one aspect of the invention;
Figure 2 is a block diagram illustrating an exemplary embodiment of a device for receiving input symbols transmitted in a speech channel in accordance with one aspect of the invention;
FIG. 3 is a block diagram of an exemplary apparatus for modifying a speech apparatus A / A so as to encode transmission error correction code information according to one aspect of the invention;
Figure 4 is a block diagram of an exemplary apparatus for receiving FEC information in a receiver and detecting whether data symbols have been received without errors in accordance with one aspect of the invention;
FIG. 5A is a block diagram of a device for mapping each user command speech command to a standard formant frequency set for controlling an automatic server having voice recognition hardware, in accordance with one aspect of the invention;
FIG. 5B is a block diagram of a device for mapping each user command call command to a standard formant frequency set for controlling an automatic server equipped with speech recognition hardware, and for intermixing the generated formant frequencies with another user call in accordance with one aspect of the invention;
Figure 5C is a diagram depicting an exemplary output format that mixes automatically generated key words with the user's own speech; and Figures 6 and 7 illustrate exemplary implementations of transmitter and receiver hardware, respectively, allowing for switching modes for transmitting voice and data in the same voice channel.
DETAILED DESCRIPTION
The various embodiments of the invention will now be described with reference to the drawings, in which like components are identified by like numerals.
The invention includes methods and apparatus for using signaling in the same frequency band in combination with a speech encoder without suffering the disadvantages of tone signaling techniques described in the BACKGROUND section of this specification. This and other innovations have been acquired through techniques in the fields of speech synthesis and recognition. In one aspect of the invention, digital information is converted into formant sequences that can be easily transmitted by a transceiver's speech encoder. For example, it is well known in the art of human speech synthesis that formant is a resonant frequency of the vocal tract. Because formant signals have the same frequency as real human speech, such signals can be easily converted to a receiver using a conventional speech encoder (typically found in transceivers). This way you can avoid problems with other types of tones, such as DTMF.
The "alphabet" of formant frequencies (or combinations of formant frequencies) is predetermined to represent all possible values of digital information so that conversion for transmission involves mapping a given unit of digital information to a corresponding formant sequence and then transmitting a signal representing the corresponding formant quotient. The predefined "alphabet" is assumed to be easily distinguishable from the normal speech data stream. On the receiver side, the received formant frequency (or combination of formant frequencies) is converted back to the corresponding digital information by means of the reverse conversion process.
In another aspect of the invention, there is provided a technique for performing switching between a voice connection and a data connection on a common communication channel. A more detailed description of this and other aspects of the invention is provided below.
FIG. 1A is a block diagram of an exemplary device for communicating input symbols in a voice channel in accordance with one aspect of the invention. The symbol-formant frequency encoder 100 of the present invention provides equal length k input symbols 101 to an address circuit 103 that performs the translation (conversion) of the input symbol 101 into another bit structure required in a given embodiment. For example, the address chain 103 may be added to the base-tone input symbols 101. It is recognized that the address chain does not play a major role in the invention and may be omitted in some embodiments. To denote the address chain as an optional element, it is also represented by a dashed line in the figure. The output of the address circuit 103 is provided to the address (or data) input port of the formant generator 105. The formant generator 105 is a means for one 2<sup>k</sup> generating possible extended codes that provide predefined formant combinations.
A specific expanded code generated by the formant generator 105 is a function of the particular "address" (or input symbol) that is assigned to its input. Each of the extended codes is N-bit wide, each bit corresponding to one Je-bit width of the input symbols.
The formant generator 105 can be solved as one of many different methods.
For example, engineers of ordinary skill in the art are capable of designing hardware
EE 200000341 Α A logic circuit that translates between the desired input symbols 101 and the yV-wide expanded code. In another embodiment, also shown in Figure IB, this task is performed by preprogrammed memory 105 '. The memory 105 'may contain at least 2<sup>k</sup> a memory slot to be able to store an expanded code for every possible Λ-bit width input symbol. In this embodiment, each of the expanded codes is stored at a specific address in memory 105 'such that it can be provided with a memory output port at any time when the corresponding input symbol (or translated address, output from the address chain 103) is provided to the memory address input port. In this way, the memory 105 'is used as a device for mapping the input symbols to the corresponding expanded code. Although memory 105 'may be innovatively designed as a non-volatile memory device (e), it is not necessarily required.
As noted above, in some embodiments, each N-bit wide expanded code may represent a combination of formant frequencies. In one such embodiment shown in FIG. IC, the expanded code is in the form j formant combined with (e attached) to the N-bit expanded code form provided in the speech encoder.
In one embodiment, the number j may be the number of formants needed to generate a single phoneme. A phoneme is known to be the primary unit of language, meaning that it is the smallest difference in sound that distinguishes one word from another. English is typically described as consisting of either 44 or 45 phonemes.
In Figure 1C, the formant generator 105 "comprises two components:
a formant selector 107 and an adder circuit 109. The formant selector 107 may be constructed from a plurality of addressable memories that operate in parallel. The input symbol 101 (or translated address, the output of the address chain 103) is provided to the formant selector 107. Each j output test provides a formant, which is then provided, for example, to an adder circuit 109 which combines the j formant and generates an N-bit width expanded code.
FIG. 2 is a block diagram illustrating an exemplary embodiment of a device that receives symbols transmitted over a speech channel in the form of formants as described above. In the example device, the received formants (labeled "input" in the drawing) are stored in buffer memory 201, from which they are subsequently analyzed. Specifically, the pattern recognition module 203 examines patterns of received digital bits that represent transmitted formants (or combinations of formants) and determines which symbols correspond to those patterns. The output of the pattern recognition module 203 is an address for selecting a corresponding one or more data symbols (words) stored at different addresses of the addressable memory 205.
As noted above, each input symbol may be encoded as its respective formant, a combination of two or more formants. To further illustrate the invention, we introduce the term "speech unit", which is constructed to mean either a single formant that represents a symbol, or a combination of formants that together represent a symbol. Turning now to another aspect of the invention, when digital information is transmitted in a channel, it is a common method to use transmission error correction code (FEC) techniques to ensure the integrity of received data. The FEC technique typically involves adding additional information bits to the transmitted data, which additional bits may be used to detect and possibly correct errors in received data. In another aspect of the invention, the FEC technique may be applied to convey input symbols using contextual means for conveying additional FEC-related information (e.g., checksum bits) in a transmitted speech unit sequence. Such context dependency can be realized, for example, by modifying the default symbols stored in memory (e.g., memory 105 ') as described above. For example, suppose there is a stream of input symbols that are associated with a stream of speech units A, B, C. . . P, Q. The speech unit Q has at least one preceding sequence, called A, B, C. . . P. This immediately preceding portion may be associated with an address (A, B, C ... P) that may be used to address a modification of the speech unit to be added to the Q unit. The modification provides additional information consistent with the additional FEC-related information we described above. The type of modification may correspond to qualities such as volume, pitch, etc., in a normal voice connection. On the receiver side, the modification is detected and decoupled to determine the FEC-related information that was transmitted with the received symbol. This FEC-related information can be used to determine the correctness of the received string.
Let's take a string of l speech units, denoted by Ai, A<sub>2</sub>,. . . Am, Aj, and the A / modification designated M (A /; Ai, A<sub>2</sub>,. . . A / _i), and the modified speech unit so generated is denoted by [A / j. A block diagram of an exemplary speech unit A / modification unit will now be described with reference to Figure 3.<sub>2</sub>,. . . A buffer 301 is provided to store the Am, Ai. The modification calculation unit 303 has access to the buffer 303 and acquires the preceding speech units from the sequence IIa. As mentioned above, the calculation is preferably performed by extracting the address from speech units A<sub>b</sub> A<sub>2</sub>,. . . Am. The address thus formed is used to access the addressable memory (not shown), acquiring a modification that is assigned innovatively for each possible address.
The modification unit 305 receives the modification from the modification calculation unit 303. The modification unit 305 also accesses the buffer 301, thereby acquiring the speech unit A / to be modified. Modifier unit 305 then modifies speech unit A / according to the type of modification being performed (e.g., pitch and / or volume modification), and outputs the modified speech unit.
Figure 4 is a block diagram of an exemplary device for acquiring FEC information in a receiver and determining whether data symbols have been received without errors.
The exemplary device consists of two buffers: a first buffer 401 stores a string of received speech units containing a modified speech unit [A /]. The second buffer 403 stores the string of recently decoded speech units. If the FEC information is determined based on a plurality of, 7-1, speech symbols, the second buffer 403 should be capable of storing at least / -1 of the latest decoded speech units.
The second buffer 403 provides the latest decoded speech symbols to the modification calculation unit 405, which determines the expected modification. The modification calculation unit 405 may operate using the same principles as the modification calculation unit 303 described with reference to Figure 3.
The speech unit calculating unit 407 detects the received speech unit based on the expected modification value (given from the modification calculation unit 405) and the latest received modified speech unit [A /] (given from the first buffer 401). The received speech unit is separated by a rotation modification function (represented as [j '<sup>1</sup> received in the modified speech unit [A]). For example, if the modification is in the form of a waveform originally added to the speech unit Aj, the inverse modification may include separating the expected modification from the most recently received modified speech unit [A /].
The data symbol decoder 409 accepts the received speech unit and performs decoupling by extracting the corresponding data symbol provided at the first output.
To perform this decoupling, the data symbol decoder 409 (from the speech unit calculating unit 407) compares the received speech symbol with a stored "dictionary" (yocabulary) of predetermined speech symbols and determines which of the predetermined speech units is the most consistent. Most of the predefined speech units
EE 200000341 Α ίο matched address can then be used to identify, directly or indirectly, the decoded data symbol, which is then provided to the first output 415.
The second output 417 of the data symbol decoder 409 outputs a received modification calculation unit 411 to the most matched speech unit. The most matched predetermined speech unit is now received as the most identical transmitted speech unit. The received modification calculation unit 411 operates by detecting which modification was made so that the most recently modified speech unit [A /] could be generated by adding the most equally transmitted speech unit to it. For example, if the modification is in the form of a waveform attached to a speech unit, determining the received modification can be done by separating the most recently transmitted speech unit from the most recent received modified speech unit [A /]. The difference in this case is the actual modification received.
Both the actual received modification (from the received modification calculation unit 411) and the expected modification (given modification from the calculation unit 405) are provided to the error detector 413, which compares the two and generates an error signal when it detects discrepancies. The error signal may be used to determine the accuracy of the decoded data symbol generated at the first output 415 of the data symbol decoder 409. The comparison between the received modification and the expected modification can be made according to known algorithms for determining the "distance" between the two.
In another aspect of the invention, the transmission of symbols is intermittent, with intervals inserted between transmitted symbols. This makes it easier to determine the beginning and end of the received symbols and thus facilitates the use of known methods, such as pattern comparison, for character recognition. One of the disadvantages of intermittent transmission is that it reduces the transmission speed. By contrast, using continuous transmission eliminates this problem, but requires much more sophisticated technology to perform decoding on the receiving side of the transmission.
In yet another aspect of the invention, the formant frequencies representing the different speech units are selected only from the formant frequencies corresponding to the vocal (as opposed to the silent) voices. It can be used in combination with another aspect of the invention in which the strings of p symbols, which consist exclusively of audio speech units, are separated from silent speech units. It is innovative in that it facilitates the detection of one end of a sequence of speech units and the beginning of the next start of a sequence of speech units.
EE 200000341 Α
The invention, as described so far, can be innovatively applied to solve a number of communication problems. For example, it is known that users can be enabled to establish telephone connections to automated servers that can perform one of the countless services for the user. Such services may include, for example, providing information to the user (e.g., phonebook entries), or enabling the user to place an order made available by an automated service provider. Moreover, it is known that by using call recognition hardware on an automated server, it is possible for the user to issue call commands to control the automated server.
One problem with such a device, however, is that the audio recognition hardware is usually tuned to recognize speech articulated by a particular group of people (e.g., a group of people who speak a particular language). This means that anyone who speaks another language, or even the same language, but with different speech characteristics (such as country or region accent), may have difficulty in making their speech commands recognizable to an automated server. The invention can be applied to solve this problem.
An example of one such solution is shown in Figure 5A. Illustrated here is a similar symbol-formant frequency encoder 100 as used in Figure 1. To enable virtually any automated server to use voice recognition hardware, it is provided with a microphone 501 or other input device for receiving acoustic energy from a user and converting that energy into a corresponding signal.
The signal is provided to a speech recognizer 503 that is tuned to recognize a particular user's call. That is, speech recognition device 503 is configured to recognize a user's specific language and a particular accent, more specifically, speech recognition device 503 can be configured to recognize commands that a user may say when connecting to an automatic server (not shown).
The output of speech detector 503 is expected to be one of many predetermined symbols.
The symbols are provided to the input symbol-formant input of the encoder 100, which converts the received input symbols to the corresponding formant frequency superposition as fully described above. Specifically, the corresponding formant frequencies for which the automatic server (not shown) is tuned for recognition and response are selected. In this way, the speech of different, possibly even non-native, users is converted into a single "language" that is easily recognized on an automated server.
In one embodiment, the speech recognition device 503 is implemented in a mobile terminal used in a mobile communication system. Voice recognition systems are available that are possible
EE 200000341 Α integrate a personal mobile phone and tune it to the voice characteristics of a regular user of the phone, so there is no need to describe them in detail here.
In the much more complicated configuration of Figure 5B, the system is able to interleave a user's speech with predefined symbols as described above. In this embodiment, a buffer 505 is used to store an incoming call from a user. The analysis unit 507 includes a library of audio keywords. These keywords, which may also include synonyms, are read by the user to configure the speech terminals of the mobile terminal to understand these words. During the setup process, the user can browse through the memory that stores a number of commonly used keywords. Keywords can be displayed to the user and the user responds by saying the word. By pressing the button, the user can also indicate that there is a synonym.
When the user submits a request in the form of an audio query, the analysis unit 507 checks the user's speech in the buffer 505 and extracts the key words and transforms them into a standardized format as explained above. Those words that have not been recognized may be included unchanged in the words to be transmitted. The formatting unit 509 performs the task of generating an output format, which is the generated keywords mixed with the user's own speech. 5C below illustrates an exemplary format in which a predetermined header 515 signals that a stored symbol 517 is followed by a pause 513 inserted between a normal speech 511 and a header 515 to assist the receiver in performing the pattern recognition tool header 515 task.
The recipient side may have speech recognition circuits for interpreting additional words of the user. In this way, the key words are always identifiable at the service node, and the incoming call can be further analyzed on the receiving side for further identification to make sure that the correct message has been transmitted, thereby allowing the receiving party to receive as much information as possible from the received message.
The following exemplary embodiment of the invention is for informing a user of a mobile terminal who wants to start a conversation, e.g., a display or other analysis is performed before the user answers the call. According to this aspect of the invention, the terminal may, for example, exchange data with a service node even before the paging signal is heard. As noted above, such data exchange may, for example, be necessary to inform the recipient calling, to display such information, or to perform other analysis. Such a "business card" exchange usually requires changing the channel status to support the data connection. However, according to the present invention, the same speech channel can be used for both data exchange and subsequent voice communication.
EE 200000341 Α
In yet another exemplary embodiment of the invention, the channel can be used to transmit information, such as short messages or warning messages, to the terminal in the call frequency band. This gives the server the ability to immediately invite the user without any reference. This is possible both when there is no voice connection and when a conversation is in progress.
In another example application, the user may send short messages or orders to the server during a voice connection and even when no communication exists. The server may be the recipient of information or transmit it to its final destination on any carrier channel.
In another example application, the invention may be used to realize a distributed "blackboard" (<sup>i <</sup>whiteboard ') (an e display that can be modified by a user, for example, by drawing on it, and reproduced on another display terminal). For example, the user may have a shared whiteboard on the screen of his client device. The board is shared with another user connected through the server. During a call, both users can point to their screen, mark objects on the screen, or draw lines. These actions, which are reproduced on another user's display terminal, do not require a significant amount of bandwidth for transmission to the other party.
Since the same speech channel is used for both data and speech, there is an obvious need for a mechanism that allows for a particular mode (voice or data) to be controlled. Such a request is addressed in accordance with another aspect of the invention described with reference to Figures 6 and 7. The initiator of the fashion setting can be either a client or server having a terminal device (e.g., a fixed or mobile phone). The mode selection is made by the initiating party, turning off all speech devices. In this case, if the initiating party (ie the party initiating the fashion change) is a customer, this operation may involve switching off devices such as microphone and earpiece. If the initiating party is a server, the devices to be turned off may include any connected device or another client.
FIG. 6 illustrates an example of a client terminal having the ability to initiate a mode change in accordance with the invention. Here, the keypad 601 is connected to the control unit 603. By activating one or more buttons on the keypad 601, the user may cause the microphone 605 of the control unit 603 to be turned off. The control unit 603 then activates a speech processor 607 which generates a sequence of pre-programmed (e) predetermined symbols (e.g., in the form of a formant frequency sequence as explained above), denoted herein as "X". The predetermined symbols are fed to a speech encoder 609 which processes the sequence of symbols in a conventional manner and transmits to the transmitter 611.
Figure 7 illustrates an example of a device viewed from the other side of a connection.
A predetermined sequence of symbols, X, is received by antenna 701 and transmitted
EE 200000341 Α mobile call center gateway (GMSC) 705 to base station 703 and public network 707 for designated service node 709. Service node 709 includes a speech processing unit 711 that detects the transmitted symbol sequence X and informs the service node 709 of the response to the transmitted symbol node. the Y symbol of the confirmation symbols, and disconnects all loudspeakers or other connected parties.
The klint side detects a symbol imprint Y of the speech processing unit 607 and informs the control unit 603 of this recognition. In response, the control unit 603 stops sending the imprint X and sets the client data communication mode. The service node, after detecting the cessation of the symbol outline X, sets its own mode for data transmission. Subsequently, data exchange will take place as described above.
When it is necessary to stop data exchange and return to voice communication, this is easily accomplished by exchanging data sets with a predetermined setting that, if detected by the recipient, will cause the control unit 603 and service node 709 to change mode (including reconnecting all voice and disconnected during preparation for data exchange).
The apparatus and techniques described above allow data communication to take place in an encoded voice channel (e.g., a GSM channel or any cellular telephone system using digitally encoded communication) by means of signaling in the same frequency band. The apparatus and technology are independent of the type of speech encoder used and may be innovatively implemented to perform functions such as setting up signaling across multiple networks to transmit commands to a service computer. Another advantage of the innovative device and technique is the ability to mark the voice connection, thereby separating the header and body of the audio message. For example, the user may be provided with the ability to mark a message by pressing a button that causes the header to generate data symbols indicating a particular message recipient. As another example, the user may use this real-time tagging feature to add data that points to an address or web page where the receiver can find and acquire the relevant image.
In this aspect of the invention, predetermined markup is used to separate the header portion from the body (message body). Both the header part and the body part contain recorded sections of information.
Another application of such a marking procedure may be the ability to include additional information with audio messages that are left on the call partner.
Additional information may be, for example, an artist code, such as the identity of the caller or the name of the caller. As the caller processes messages, he or she may see additional information that may be helpful in selecting the message to listen to. Additional information may be provided in different ways:
1) The information may be stored in the calling party's terminal and sent in response to a request made by the called party's voicemail box (e.g., in the same way that the terminal's identity is sent when a fax is sent).
2) the information may be stored in the caller's terminal and sent in response to a caller's action, which may be to press a corresponding button on the terminal.
3) The information can be entered and transmitted manually by the caller either before or after recording the voice message.
The invention provides many innovations in communications, culminating in the fact that speech and data can be mixed in a manner that is transparent to the underlying communications system that connects the communications parties. This means that, for example, a message generated in accordance with the invention as a data message can be stored on the receiver side in a conventional voice mailbox. The identification symbol in a stored message can be used, for example, to define a message as a data message, whereby the recipient can receive the message, for example, in printed form.
The invention has been described with reference to a specific embodiment. However, it will be apparent to those skilled in the art that the invention can be embodied in other specific forms without substantially departing from the spirit of the invention. The embodiments shown are for illustrative purposes only and may not be construed as limiting. The scope of the invention is defined by the appended claims rather than the preceding claims, and any variations and equivalents within the scope of the claims are contemplated herein.
Contents4
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
9 members in 8 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 99077397 | United States of America | A | |
| 9802248 | Sweden | W | |
| 9802248 | – | – | – |
| 990773 | – | – | – |
| US19970990773 | – | – | – |
| WO1998SE02248 | – | – | – |
Members9
| Document | Office | Kind | |
|---|---|---|---|
| WO9931895A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU1896299A | Australia | A | |
| EP1040673A1 | European Patent Office (EPO) | A1 | |
| BR9813596A | Brazil | A | |
| CN1285118A | China | A | |
| US6208959B1 | United States of America | B1 | |
| EE200000341AThis record | Estonia | A | |
| US6385585B1 | United States of America | B1 | |
| MY120784A | Malaysia | A |
Numbers
- Publication, DOCDB
- 200000341
- Publication, EPODOC
- EE200000341
- Application
- 200000341
- Application, DOCDB
- P200000341
- Application, EPODOC
- EEP200000341
Titles2
- English
- Transfer data encoded konekanalis
- Estonian
- Andmete ülekanne kodeeritud kõnekanalis
Classification
- CPC, 2
- H04Q1/45
- G10L19/00
- IPC, 2
- G10L19 00
- H04Q1 45