Method for speech signals detection.
Abstract
In known detector circuits for speech signals, after bandpass filtering to obtain the fundamental speech frequency of the microphone signal, the speech envelope of the said signal is detected and the detection signal used for control of a control amplifier in such a way that this amplifier attenuates the microphone signal when the detection signal is absent and amplifies the microphone signal when the detection signal is present. A method of this kind results in a detection time of the order of magnitude of 200 ms. The new method is intended to permit a shorter detection time. This is achieved by the fact that the signals occurring at the output of the low-pass filter are checked for amplitude and for the duration of a defined amplitude and that a speech signal is detected if at least three sequential amplitudes have occurred in the range of the fundamental speech frequency. The detection of speech signals can be performed particularly in combination with a signal processor for controlling the amplification of microphone signals in an environment filled with interfering noise, amplification taking place only when speech signals are present; it can also be used in intercom or two-way communication systems to amplify the signal in the appropriate direction and attenuate it in the opposite direction when a speech signal is present. <IMAGE>

Term
Term ended
Projected expiry passed 20 February 2009, 17.6 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
20 claims: 1 independent, 19 dependent
- 1Verfahren zur Erkennung von Sprachsignalen, wobei diese zunächst einem Tiefpaß zugeführt werden, dessen Durchlaßbereich im Bereich der Sprachgrundfrequenz liegt, dadurch gekennzeichnet, daß die am Ausgang des Tiefpaßfilters auftretenden Signale auf Amplitude und Dauer einer bestimmten Amplitude überprüft werden und daß dann ein Sprachsignal erkannt wird, wenn mindestens drei aufeinanderfolgende Amplituden im Bereich der Sprachgrundfrequenz aufgetreten sind.
- 2Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß das Tiefpaßfilter eine obere Grenzfrequenz von höchstens 400 Hz aufweist.
- 3Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß die Dauer der Überprüfung einer Amplitude über ein Zeitfenster (ZF) erfolgt, dessen Länge kleiner ist, als die Hälfte der kürzesten Periode der Sprachgrundfrequenz.
- 4Verfahren nach einem der Ansprüche 1 bis 3, dadurch gekennzeichnet, daß sowohl positive als auch negative Amplituden überprüft werden.
- 5Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß die Überprüfung der folgenden Amplituden über ein Amplitudenfenster (AF) erfolgt, dessen Amplitudenbereich in Abhängigkeit von dem ersten erkannten Amplitudenhöchstwert festgelegt wird.
- 6Verfahren nach Anspruch 5, dadurch gekennzeichnet, daß das Amplitudenfenster einen Amplitudenbereich von + 20 bis -20 % des Amplitudenhöchstwertes aufweist.
- 7Verfahren nach einem der Ansprüche 1 oder 5, dadurch gekennzeichnet, daß der Zeitraum zwischen dem ersten erkannten Amplitudenhöchstwert und dem folgenden im Amplitudenfenster (AF) liegenden Amplitude innerhalb eines vorgegebenen Zeitrahmens gemessen wird.
- 8Verfahren nach Anspruch 7, dadurch gekennzeichnet, daß der Zeitrahmen (PF) zwischen 3 und 12,5 ms liegt.
- 9Verfahren nach Anspruch 7, dadurch gekennzeichnet, daß der dritte Amplitudenhöchstwert (A3) in einem Zeitfenster ( F) liegen muß, dessen Lage durch den Abstand zwischen dem ersten (A1) und dem zweiten (A2) Amplitudenhöchstwert bestimmt wird und innerhalb einer Toleranz von ± 7,5 % desselben liegt.
- 10Verfahren nach einem der Ansprüche 1 bis 9, dadurch gekennzeichnet, daß die erste Periode und die zweite Periode bzw. die erste Periode und die dritte Periode zur Bestimmung der Kreuz-Korrelationsgrade benutzt wird.
- 11Verfahren nach einem der Ansprüche 1 bis 10, dadurch gekennzeichnet, daß aus den gemessenen Zeiträumen der ersten und der zweiten bzw. der ersten und der dritten Periode der normierte mittlere quadratische Fehler ermittelt wird.
- 12Verfahren nach einem der Ansprüche 10 oder 11, dadurch gekennzeichnet, daß die ermittelten Werte mit Hilfe einer wählbaren Schwelle überprüft werden und daß bei Überschreiten der Schwelle durch einen ermittelten Wert ein Sprachsignal erkannt wird.
- 13Verfahren nach einem der Ansprüche 1 bis 12, dadurch gekennzeichnet, daß das analoge Sprachsignal einem Analog/Digital-Wandler zugeführt wird.
- 14Verfahren nach einem der Ansprüche 1 bis 13, dadurch gekennzeichnet, daß das digitalisierte Sprachsignal einem Signalprozessor (SP) zugeführt wird, welcher ein, das Vorliegen eines Sprachsignals kennzeichnendes Ausgangssignal liefert.
- 15Verfahren für eine Mikrofonverstärkerschaltung mit einem Regelverstärker nach einem der Ansprüche 1 bis 14 , dadurch gekennzeichnet, daß bei Vorliegen eines Sprachsignals das Eingangssignal des Regelverstärkers (RV) auf Veranlassung des Signalprozessors um einen vorgegebenen Wert verstärkt wird.
- 16Verfahren für eine Freisprecheinrichtung mit je einem Regelverstärker, für das Mikrofon- und das Lautsprechersignal nach einem der Ansprüche 1 bis 15, dadurch gekennzeichnet, daß bei Vorliegen eines Sprachsignals des Mikrofons (M) das Lautsprechersignal um einen vorgegebenen Wert durch den zugeordneten Regelverstärker (RV2) auf Veranlassung des Signalprozessors (SP) gedämpft wird.
- 17Verfahren nach einem der Ansprüche 14 bis 16, dadurch gekennzeichnet, daß durch den Signalprozessor (SP) das Mikrofonsignal um den Betrag der Erkennungszeit von Sprachsignalen verzögert wird.
- 18Verfahren für eine Gegensprecheinrichtung mit je einem in jeder der beiden Richtungen liegenden Regelverstärker nach einem der Ansprüche 1 bis 17, dadurch gekennzeichnet, daß durch den Signalprozessor (SP) bei Vorliegen eines Sprachsignals der betreffende Regelverstärker aufgesteuert und der andere Regelverstärker gedämpft wird.
- 19Verfahren nach einem der Ansprüche 1 bis 18, dadurch gekennzeichnet, daß das Steuersignal für den bzw. die Regelverstärker nach Ausbleiben eines Sprachsignals für eine bestimmte Zeit aufrechterhalten wird.
- 20Verfahren nach einem der Ansprüche 1 bis 19, dadurch gekennzeichnet, daß die Funktion der Regelverstärker (Rv1, Rv2) durch den Signalprozessor (SP) übernommen wird.
Independent claims20
33 paragraphs, as filed
0001The invention relates to a method for recognizing speech signals, wherein these are first fed to a low-pass filter, the pass band of which is in the range of the basic speech frequency.
0002In the field of electroacoustics, the recognition of voice signals is of great importance, since the presence of voice signals can be used as a criterion for increasing the gain. For example, for the acoustic decoupling of hands-free devices, the amplification of the transmit and receive signal is controlled as a function of the presence of a speech signal. The same applies to conference facilities.
0003It has already been proposed (P 37 34 446.3) to achieve noise compensation for a microphone by subjecting it to greater amplification in the presence of a speech signal, in order to achieve better intelligibility in the case of strong background noise. After a bandpass filtering for the basic speech frequency, the envelope of speech of the microphone signal is detected and the detection signal is fed to a timing element which has a certain response delay. The output signal of the timing element then serves to control a control amplifier which amplifies the microphone signal. The disadvantage of this method is the use of timers for processing the microphone signal, which means that there is a risk that initial syllables will be suppressed.
0004The object of the invention is to provide a method for recognizing speech signals, in which the presence of speech signals is recognized after a very short time, without this suppressing initial syllables.
0005This object is achieved in that the signals appearing at the output of the low-pass filter are checked for the amplitude and duration of a certain amplitude and in that a speech signal is recognized when at least three successive amplitudes have occurred within a predetermined time frame.
0006The signals are first checked for maximum amplitude values. As soon as a maximum amplitude value is determined, the period of time within which a further maximum amplitude value occurs is measured in order to be able to recognize speech signals in this way.
0007The invention is explained in more detail using an exemplary embodiment which is illustrated in the drawing.
0008It shows:<ul id="ul0001" list-style="none"><li>Fig. 1 shows the periods of a speech signal in connection with the evaluation criteria and</li><li>Fig. 2 shows the block diagram for an arrangement for performing the method.</li></ul>
0009The three amplitudes A1 to A3 shown in FIG. 1 are the amplitudes of a speech signal which are present at the output of a low-pass filter whose cut-off frequency is approximately 400 Hz. The signals supplied to the input of the low-pass filter are generated, for example, by a microphone and are composed of room noises and speech signals.
0010The method according to the invention for recognizing speech signals now essentially uses the frequency range of the fundamental speech frequency (80 to 333 Hz) for analysis. The most important feature for recognizing speech signals is the period of the vibrations of the speech signals, which, depending on the speaker, is in the range of 3 to 12.5 ms, depending on the speaker. This first feature is used to distinguish between speech and noise. For reliable detection of speech signals, the detection of zero crossings in the speech signal is not expedient, since in the event of a disturbance, for example due to noise, the number of zero crossings can increase so much that detection of speech is no longer possible in this way. The method according to the invention uses the maxima of the speech signal to recognize speech. If these are then within a predetermined amplitude time window, then a first criterion for the presence of speech signals is given. The choice of window parameters has a significant influence on the period detection.
1. Window size for a local amplitude maximum.
0011The window size is chosen such that it is smaller than half the smallest possible period of the basic speech frequency so that both positive and negative maximum values of the speech signal can be recognized. This is necessary because the speech signal is not symmetrical with respect to the dynamic range. The window size is therefore approximately 0.9 ms.
2nd Amplitude window
0012The amplitude tolerance of the maximum values is very small over a few periods in the case of an undisturbed speech signal, but can be increased significantly at high interference levels as a result of additive superimposition of the interference signal. The amplitude window is approximately plus minus 20% of the first maximum.
3rd Distance tolerance of the maximum values found
0013In the case of undisturbed speech, the distance between the maximum values of the signals is not constant, since the speech signal is frequency-modulated. A strictly periodic course of the excitation signal cannot be expected, the fluctuations in the basic speech frequency can therefore be considerable. However, the voiced, steady sounds have a quasi-periodic course. If the signal is disturbed (e.g. additively by noise), an additional shift of the signal maxima in the temporal direction can result. Investigations carried out have shown that the tolerance range for the detection of the signal maxima can be approximately 15%.
0014Under these boundary conditions, it can be assumed that even with an undisturbed speech signal never more than 10 periods of the signal meet the specified criteria, so that periodic, non-modulated interference signals, the frequency of which is in the range of the fundamental speech frequency, are distinguished from speech signals using the method according to the invention can.
0015As soon as a maximum value is recognized, its temporal position is saved. If the next maximum value that occurs does not meet the conditions described below, the data of the first maximum value is deleted and that of the next maximum value is stored in its place.
0016In the example of an amplitude sequence shown in FIG. 1, it is assumed that the three maximum values M1 to M3 shown meet all the conditions required for the recognition of speech signals. The amplitude A1 has been recognized as the maximum value, whereupon its duration t1 is stored as a period. At the time center of the amplitude A1 of the first maximum M1, the time window of the period PF begins, which is open between 3 and 12.5 ms. If the next amplitude A2 now falls within the time window of the period PF, since its time window ZF lies within the amplitude window AF, the duration of the amplitude A2 is identified as the second maximum by storing the value t2. If the amplitude A3 now lies within a time window F, which is determined by the period t2 ± 7.5%, the time value t3 of the third maximum M3 is also stored. It is also pointed out that the amplitude window AF is defined as a threshold as a function of the amplitude value of the first maximum M1.
0017A simple counting process for detecting the three successive amplitudes A1 to A3, which meet the conditions described above, can already be used to conclude that a speech signal is present, in which case it is not necessary to store the period durations t1 to t3. However, two methods can be used for a more precise determination of speech signals, which are described below.
0018If several periods of an oscillation in the speech fundamental frequency range were recognized, the degree of correlation between the individual periods is determined. Through a cross correlation between the successive signal sections of a period length, high values for the nominated cross correlation coefficient are achieved in the areas in which speech is present. However, if the detected period is only random maxima in the specified interval, the correlation analysis gives small values.<maths id="math0001" num=""><img file="EP0334023A2_D0001.tif" /></maths> To determine KKF (k. N<sub>p</sub>) the second or, in the case of detection of several periods, the third period is correlated with the first. If three periods are correlated, the smaller of the two values is used for the decision. This reduces the frequency of errors in the case of randomly detected periods, particularly in the event of interference by noise signals. If more periods are used for the detection, the detection speed slows down, however, no further improvement can be achieved because the values of KKF (k. N<sub>p</sub>) decrease significantly due to the amplitude and frequency modulation of the speech signal.
0019A further improvement in the decision can be achieved if, instead of evaluating the cross-correlation function for speech decision, the nominated mean square error between the recognized periods is used.<maths id="math0002" num=""><img file="EP0334023A2_D0002.tif" /></maths>
0020The use of this error criterion leads to similar results to the formation of the KKF (k. N<sub>p</sub>). However, there are differences when the speech signal is disturbed. When the KKF (k. N<sub>p</sub>) the distinction between speech and disorder on the basis of the correlation coefficient leads to wrong decisions more often than the formation 1-Δf². Both KKF (k. N<sub>p</sub>) and 1-Δf² can assume values in the range from 0 to 1. If the value of KKF (k. N<sub>p</sub>) or from 1-Δf² a value of 0.7, for example, the input signal is marked as speech. Studies have shown that the selection of the threshold is not critical, it can also be selected in the range from 0.3 to 0.9.
0021The key advantage of this method of speech detection is the recognition time. In the worst case, ie if the speaker has a basic voice frequency of 80 Hz and if three periods are detected, the detection time is 37.5 ms.
0022In the case of undisturbed signals, the analysis using the simplified method described at the outset gives approximately the same results as the evaluation method with cross-correlation or after determining the mean square error. The detection rate is on average 5% below the detection rate of the previously described method, but can also assume higher values depending on the noise situation. Differences to the above-mentioned procedure become clear when the speech sequence is disturbed. With the selected parameters, the period detection can, depending on the respective background noise, provide an increased number of wrong decisions for some background noise situations. In the case of interference caused by impulsive signals in particular, reflections of the interference signal, if they meet the criteria for the presence of speech, are recognized as speech and lead to incorrect decisions. The detection of sinusoidal interference in the range of the fundamental speech frequency is only possible on the basis of the duration and frequency constancy of this interference signal.
0023The selection of the method for speech detection to be used is essentially determined by the expected useful / interference power ratios and the interference noises. In the case of useful / interference power ratios of more than 12 dB, the simplified detection method can be used without arithmetic operations. However, all methods only result in a short signal delay in the range of the detection time (9 to 37 ms), so that initial syllables are not suppressed.
0024The method presented can be implemented, for example, with the aid of a signal processor SP (see FIG. 2). The analog signal of the microphone M is sampled and digitized via the analog / digital converter W1. The sample values obtained in this way can be used by the signal processor according to the method according to the invention for speech detection. If speech is recognized, the microphone signal can be amplified by the control amplifier RV1 by a fixed amount at the instigation of the signal processor SP.
0025Such an arrangement is suitable, for example, for microphones which are located in a room with a large amount of noise. The intelligibility of the speech signals is improved in this way.
0026In the application example shown in Fig. 2, a hands-free device is available, with this in the presence of a speech signal in the signal of the microphone M, the control amplifier RV2 is caused by the signal processor SP to attenuate the signal for the loudspeaker LS accordingly, in order in this way to to prevent acoustic feedback between loudspeaker LS and microphone M. Conversely, if there are voice signals for the loudspeaker LS, the control amplifier RV2 could be influenced at the instigation of the signal processor SP in such a way that it amplifies the input signal to achieve a better intelligibility of the loudspeaker signal LS.
0027The signal processor receives at its inputs SE and EE data words which represent the samples of the signals. Data words are also applied to the connected lines at the outputs SA and EA of the signal processor SP. To avoid the suppression of initial syllables, the input signals can be delayed by the signal processor SP by a time which is in the range of the recognition time (5-37ms). Likewise, the signal processor SP can generate a fall time for the control signals influencing the control amplifiers RV, which is of the order of magnitude of 200 to 900 ms and is used to bridge unvoiced sounds and short speech pauses between words and sentences. The low-pass filtering function with a cut-off frequency of 400 Hz can also be carried out by the signal processor SP.
0028Another application of the method according to the invention is also conceivable in the context of an intercom system, the other direction being attenuated accordingly in response to voice signals in one direction at the instigation of the signal processor.
0029The structure of a signal processor is not further discussed in the context of this description, but such signal processors are sold, for example, by Texas Instruments under the designation TMS 320 or by Fujitsu under the designation MB 8764. Such a signal processor is to be programmed in such a way that the described method steps run automatically. The analog / digital converters W1 and W4 serve to convert the analog signals into digital signals for signal processing in the signal processor SP, while the conversion of the digital signals occurring at the outputs SA and EA into analog signals by the digital / analog converters W2 and W3 takes place.
0030In contrast to the block diagram shown in FIG. 2, the control amplifiers RV1 and RV2 can also be dispensed with if the function of amplifying the signals is taken over by the signal processor SP itself, which can also be designed as a suitable microprocessor. It is also conceivable to carry out the method according to the invention by means of a corresponding, discretely constructed analog circuit arrangement or also a correspondingly designed customer circuit.
5 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO0070602A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO9700515A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO9213340A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| EP0120325A1 | Cites | European Patent Office (EPO) | Search report |
| FR2380612A1 | Cites | France | Search report |
| US3751602A | Cites | United States of America | Search report |
4 priority claims, no other members on record
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 3810068 | Germany | A | |
| 3810068 | Germany | – | |
| DE19883810068 | – | – | – |
| 3810068 | – | – | – |
9 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Application withdrawnWithdrawn18W | 18W | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: THE APPLICATION HAS BEEN WITHDRAWNSTAA | STAA | |
| First examination report despatched17Q | 17Q | |
| Request for examination filed17P | 17P | |
| Designated contracting statesAK | AK | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | |
| Designated contracting statesAK | AK | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI |
Numbers
- Publication
- 0334023
- Publication, DOCDB
- 0334023
- Publication, EPODOC
- EP0334023
- Application
- 89102876
- Application, DOCDB
- 89102876
- Application, EPODOC
- EP19890102876
Titles6
- German
- Verfahren zur Erkennung von Sprachsignalen.
- English
- Method for speech signals detection.
- French
- Procédé de détection de signaux de parole.
- German
- Verfahren zur Erkennung von Sprachsignalen
- English
- Method for speech signals detection
- French
- Procédé de détection de signaux de parole
Classification
- CPC, 1
- G10L25/78
- IPC, 1
- G10L25 78
Designated states12
- Contracting states, 12
- Austria
- Belgium
- Switzerland
- Germany
- Spain
- France
- United Kingdom
- Italy
- Liechtenstein
- Luxembourg
- Netherlands (Kingdom of the)
- Sweden