The manner and system of acoustic echo dampening in VoIP terminal
Abstract
A method of suppressing acoustic echo in a VoIP terminal using a digital adaptive filter to obtain an estimate of the acoustic echo, which is subtracted from the signal collected by the microphone, is characterized by the fact that the digital speech signal of a distant user, before being transformed into an analog form and fed to the loudspeaker (4), , is marked by adding an encoded digital tag to it, coming from the tag generator (14), and in the total signal received by the microphone (7), after converting it into digital form, a tag is detected and, depending on its detection or lack, the tuning process of the digital adaptive filter (9) is resumed or suspended. An acoustic echo suppression system in a VoIP terminal, including a digital adaptive filter with a control block, connected between the speech signal path of a distant user and the speech signal path of a close user, and a simultaneous speech detector, is characterized in that the simultaneous speech detector (11) includes a marker generator (l4 ), connected via the tag encoding block (15) to the tag recording block (16), connected between the speech decoder (2) and the digital-to-analog converter (3) in the remote user's speech signal reception path, with the tag generator (14) also connected to the tag decoding block (l7), connected to the output of the analog-to-digital converter (8 ) in the speech signal reception path of a close user (6), and the output of the tag decoding block (17) is connected through the decision block (18) to the control block (10) of the digital adaptive filter (9).

Term
1.4 yearsleft in the term
Expires 6 March 2028.
- Priority and filed
- Granted
- Today
- Expires
4 claims: 2 independent, 2 dependent
- 1Sposób tłumienia echa akustycznego w terminalu VoIP, w którym sygnał mowy docierający od odległego użytkownika przetwarza się za pomocą cyfrowego filtru adaptacyjnego w celu uzyskania estymaty echa akustycznego, którą odejmuje się od sygnału zebranego przez mikrofon, a otrzymany sygnał wykorzystuje się do strojenia filtru adaptacyjnego, przy czym proces strojenia filtru adaptacyjnego wstrzymuje się w czasie trwania mowy równoczesnej, znamienny tym, że cyfrowy sygnał mowy odległego użytkownika, przed przekształceniem na postać analogową i podaniem do głośnika (4), znakuje się poprzez dodanie do niego zakodowanego cyfrowego znacznika pochodzącego z generatora znacznika (14), a w całkowitym sygnale odebranym przez mikrofon (7), po przekształceniu go na postać cyfrową, przeprowadza się detekcję znacznika i w zależności od jego wykrycia lub braku, wznawia się bądź wstrzymuje proces strojenia cyfrowego filtru adaptacyjnego (9).
- 2Sposób według zastrz. 1, znamienny tym, że cyfrowy znacznik stanowi ustalona sekwencja bitów dobrana tak, że jest on tłumiony przez użyteczny sygnał mowy bliskiego użytkownika (6), a pozostaje w sygnale mowy odległego użytkownika zniekształconym w pętli akustycznego sprzężenia zwrotnego (5) między głośnikiem (4) i mikrofonem (7).
- 3Sposób według zastrz. 2, znamienny tym, że dokonuje się kodowania cyfrowego znacznika poprzez dodanie do sygnału określonej liczby kopii tego sygnału o różnych amplitudach i różnych wielkościach opóźnień.
- 4Układ tłumienia echa akustycznego w terminalu VoIP zawierający cyfrowy filtr adaptacyjny z blokiem sterującym włączony pomiędzy torem sygnału mowy odległego użytkownika i torem sygnału mowy bliskiego użytkownika oraz detektor mowy równoczesnej , znamienny tym, że detektor mowy równoczesnej (11) zawiera generator znacznika (14) połączony poprzez blok kodowania znacznika (15) z blokiem zapisu znacznika (16) włączonym pomiędzy dekoderem mowy (2) a przetwornikiem cyfrowo-analogowym (3) w torze odbioru sygnału mowy odległego użytkownika, przy czym generator PL 216 396 B1 znacznika (14) połączony jest także z blokiem dekodowania znacznika (17) dołączonym do wyjścia przetwornika analogowo-cyfrowego (8) w torze odbioru sygnału mowy bliskiego użytkownika (6), a wyjście bloku dekodowania znacznika (17) połączone jest poprzez blok decyzyjny (18) z blokiem sterującym (10) cyfrowego filtru adaptacyjnego (9).
Independent claims4
19 paragraphs, as filed
The present invention relates to a method and an acoustic echo suppression system in a VoIP terminal. The solution is intended for various types of client terminals of voice communication systems in the Internet, especially when the VoIP system user uses a loudspeaker during communication, not a handset.
The development of speech signal transmission technology using computer networks, which are defined by the broad concept of VoIP (Voice over IP) telephony, is the source of many new solutions, both in terms of software, as well as electronic methods and circuits ensuring not only effective transmission via this route speech signal, but also the best possible quality of this signal. Users of VoIP systems are recommended to use a special VoIP terminal during calls, resembling a classic telephone set, or a set of headphones and a microphone. In many cases and for a variety of reasons, users make use of loudspeakers. With such configurations, the quality of the conversation may be significantly reduced as a result of the so-called acoustic echo effect. It consists in that the distant speaker's speech signal emitted by the loudspeaker is picked up by a microphone which essentially transmits the near-speaker's speech signal. As a result, the speech signal returns to the sender, who hears his own voice delayed and distorted during the conversation, since the microphone of the client terminal receives and transmits not only a useful speech signal from a near speaker, but also a distorted echo signal from a distant speaker. In order to eliminate this phenomenon, various methods and signal processing systems are used, the purpose of which is to effectively prevent the signal from returning to the sender without introducing significant delays in communication between system clients, i.e. removing the echo signal from the input microphone signal and transmitting only the useful signal for transmission.
There are a number of methods and systems known, in which the elimination of the acoustic echo is carried out by means of an adaptive digital filter.
By processing the signal coming from the distant speaker by the adaptive filter, an estimate of the acoustic echo is obtained, which is then subtracted from the signal collected by the microphone. The result of this operation is used for tuning, i.e. filter adaptation. After the adaptation process is completed, the echo estimate obtained by the filter simulates the actual acoustic echo signal, which can be subtracted from the signal received by the microphone, which has the effect of suppressing the echo signal. The effectiveness and efficiency of solutions based on the use of an adaptive filter requires automatic suspension of the filter adaptation process while a microphone connected to the subscriber's terminal receives both a speech signal from a close speaker and an echo signal from a second speaker, in order to prevent the adaptive filter being detached and the signal being processed from being distorted. There are many different methods and systems of simultaneous speech detection of varying degrees of complexity and effectiveness.
There is known from US patent 4,894,820 a method of adaptive acoustic echo suppression, in which the speech detection simultaneously controlling the filter adaptation process is based on the use of a second adaptive filter to estimate the difference between the processed signal and the received signal on the basis of statistical parameters of the signal. In the solution known from the patent specification US 6608897 in the echo cancellation process, instead of direct simultaneous speech detection, a variable filter adaptation step is used, depending on the difference between the processed signal, after echo cancellation, and the received signal. A solution for simultaneous speech detection suitable for use in VoIP telephony is also known from the patent description US 6792107. In essence, it is based on the calculation of the correlation between the signal reaching the other interlocutor and the signal processed after echo cancellation. This correlation is a measure of the similarity of signals, so in the case of a low value of the correlation coefficient, simultaneous speech can be stated. Additionally, the dynamic detection threshold is determined on the basis of the signal energy analysis. Similarly, US Patent 6192126 proposes to detect simultaneous speech by analyzing the signal energy in several frequency ranges.
In the solution disclosed in US patent 4,894,820, the acoustic echo suppression system includes a digital adaptive filter and the so-called Geigel system for simultaneous speech detection, which is performed by comparing the amplitude or energy of the processed signal and the received signal. The echo cancellation circuit known from the international patent application published under No. WO 98/43368 has adaptive filter and double speech detector blocks connected in parallel, to which two nonlinear processors are connected also connected to a noise generator
The adjusting noise estimate block and the block of the adjusting noise estimator and power level. The system is also equipped with two tone switches. The acoustic echo suppression system disclosed in WO 98/51066 comprises, in addition to the adaptive filter between the near-user signal path (microphone signal) and the remote user signal path (signal to the loudspeaker), at least one additional microphone and a second adaptive filter connected thereto, as well as additional permanent filters.
Known solutions for the elimination of acoustic echo characterized by high efficiency are based on complex computational methods, which increases the time of simultaneous speech detection and introduces undesirable delays in data transmission. On the other hand, solutions that do not cause significant delays, i.e. ensure a short signal processing time, do not give satisfactory results in terms of elimination of distortions. Therefore, these solutions are not optimal for applications in VoIP client terminals that do not have high computing power.
An acoustic echo suppression method in a VoIP terminal according to the invention, wherein the speech signal from the remote user is processed with a digital adaptive filter to obtain an acoustic echo estimate which is subtracted from the signal collected by the microphone and the signal obtained is used to tune the adaptive filter. , wherein the tuning of the adaptive filter pauses during simultaneous speech is characterized by that the digital speech signal of the remote user, prior to being converted to analog form and delivered to the loudspeaker, is marked by adding thereto an encoded digital tag from the tag generator, and the tag is detected in the total signal received by the microphone after digital conversion, and Depending on its detection or absence, the tuning process of the digital adaptive filter is resumed or paused.
The digital tag is a fixed sequence of bits selected such that it is suppressed by the useful speech signal of a near user and remains in the distant user's speech signal distorted in the acoustic feedback loop between the loudspeaker and the microphone.
Preferably, the digital tag is encoded by adding to the signal a certain number of copies of this signal with different amplitudes and different delay amounts.
The acoustic echo suppression system of the VoIP terminal according to the invention, comprising a digital adaptive filter with a control block connected between the speech signal path of the distant user and the speech signal path of a near user and the simultaneous speech detector is characterized by: that the simultaneous speech detector comprises a tag generator connected via the tag coding block to a tag write block connected between the speech decoder and the digital to analog converter in the speech signal reception path of the remote user. The tag generator is also connected to a tag decoding block connected to the output of the analog-to-digital converter in the near-user speech signal reception path, and the output of the tag decoding block is connected via a decision block to the control block of the digital adaptive filter.
The solution according to the invention provides effective acoustic echo suppression without introducing significant delays in voice communication, which leads to a significant improvement in the quality of telecommunications services in voice communication systems using computer networks.
An exemplary embodiment of the invention is illustrated in a drawing showing a block diagram of a VoIP terminal.
The method of acoustic echo suppression in a VoIP terminal is that the speech signal coming from the remote user is processed with a digital adaptive filter 9 to obtain an acoustic echo estimate, which is subtracted from the signal collected by the microphone 7, and the received signal is used for tuning. digital adaptive filter 9. The tuning process of the digital adaptive filter 9 pauses during the simultaneous speech of the near and distant user, after it has been detected by the simultaneous speech detector 11. The detection of simultaneous speech is carried out as follows: the digital speech signal of a remote user, before being converted to analog form and fed to the loudspeaker 4, is marked by adding to it an encoded digital marker from the marker generator 14, and in the total signal received by the microphone 7, after converting it to digital form, detection of the digital tag is performed. The digital tag is a fixed sequence of bits selected so that it is suppressed by a useful speech signal from a near user 6 and remains in the distant user's speech signal distorted in the acoustic feedback loop 5 arising between the loudspeaker 4 and the microphone 7. To obtain a distortion resistant signal , the digital tag is encoded such that the sum of the specified tag becomes the tag
The number of copies of the original signal is scaled up and delayed by different amounts. This way of marking the signal corresponds to the method of marking the signal by hiding the echo. In the case of the absence of a digital marker in the signal received from the microphone 7, tuning of the digital adaptive filter 9 is suspended, and after its detection, the adaptation process of this filter is resumed.
The VoIP client terminal is connected to the telecommunications network 1. The terminal has at its input, in the speech signal receiving path of a remote user, a speech decoder block 2, the output of which is connected to a digital-to-analog converter 3 connected to a loudspeaker 4. The drawing shows schematically, in the form of a block, an acoustic feedback loop 5 that illustrates the distortion of the acoustic waves emitted by the loudspeaker 4, which, together with the useful speech signal from a user close by, reach the microphone 7. In the speech signal reception path of a close user 6, the microphone 7 is connected to an analog-to-digital converter 8, the output of which is connected to the summing node S, also connected to the output of the digital adaptive filter 9, whose input is connected to the output of the speech decoder 2 and the output of the control unit. 10. The input of the control block 10 is connected to the simultaneous speech detector block 11 and the output of the summing block S. The output of the summation block S is connected via a signal dynamics processor 12 and a speech coding block 13 to an input to the telecommunication network 1. The simultaneous speech detector 11 comprises a marker generator 14, a marker coding block 15, a marker write block 16, a marker decoding block 17 and a decision block 18 . The output of the tag generator 14 is connected via the tag coding block 15 to the tag recorder 16, which is connected between the output of the speech decoder 2, also connected to the tag coding block 15, and the input of the analog-to-digital converter 3 in the speech signal receiving path of the remote user. The output of the marker generator 14 is also connected, via a marker decoding block 17, to a decision block 18, the output of which is connected to the control block 10 of the digital adaptive filter 9. The output of the analog-to-digital converter 8 in the signal path is also connected to the marker decoding block 17. of the user's close speech 6.
The digital speech signal of the remote user received from the telecommunications network 1 is decoded in the speech decoder block 2, converted into an analog signal by a digital-to-analog converter 3 and transmitted to the loudspeaker 4. The loudspeaker 4 emits acoustic waves that are distorted in an acoustic feedback loop 5, mainly by multiple reflections in the vicinity of the terminal, causing reverberation. In a situation where the near user 6 of the terminal is silent at the moment, the microphone 7 collects distorted speech signal of the distant user from the surroundings. This signal is converted into digital form in the analog-to-digital converter block 8. In order that this signal, as unwanted, does not return to the remote user in the form of an acoustic echo, it is processed and suppressed in a system containing a digital adaptive filter 9 and a simultaneous speech detector 11 . The digital adaptive filter 9 computes an estimate of the acoustic echo which is subtracted at the summing node S from the signal from the output of the analog-to-digital converter 8 containing the echo. The result of this operation is used by the control unit 10 which controls the tuning of the digital adaptive filter 9. In the following steps, the coefficients of the digital adaptive filter 9 are modified by the control unit 10, and as a result, an exact echo estimate is calculated which, subtracted from the processed signal at the summing node S, makes it possible to obtain an acoustic echo-free signal. This process is effective as long as the adaptation of the digital adaptive filter 9 is stopped when the near user 6 starts talking to the microphone 7, and is resumed when the near user 6 becomes silent. Otherwise, the adaptive filter 9 will be de-tuned, with the result that the signal is significantly distorted. For detection of the simultaneous speech signal, the marking of the speech signal of the remote user, taken from the input of the terminal, is used after it has been processed by the speech decoder 2. The marker generator 14 produces a digital tag in the form of a predetermined sequence of bits selected in such a way as to allow later detection of the presence of this tag in a signal that has been distorted by the transmission of acoustic waves between the loudspeaker 4 and the microphone 7. The tag coding block 15 converts the digital tag so that it is the sum of a certain number of copies of the original signal, scaled up and delayed with respect to each other by different values, resulting in a signal covering a wide frequency range, immune to distortion. The thus obtained encoded digital mark is attenuated and added to the actual signal of the remote user in the mark record block 16. The signal with the digital mark stored therein is outputted, i.e. via the digital-to-analog converter 3 to the loudspeaker 4. The signal received by the microphone 7 is checked in the simultaneous speech detector block 11 for the presence of the mark. In the absence of simultaneous speech, only the acoustic echo signal and possibly noise and other disturbances are present in the signal received by the microphone 7, and thus the presence of a digital marker can be detected. On the other hand, in the case when a useful speech signal introduced by a close user 6 is also present in the analyzed signal, the digital marker contained in the echo signal is suppressed by this useful speech signal, which allows to determine the absence of the marker in the analyzed signal. The tag decoding block 17 performs signal preprocessing including, inter alia, normalization and synchronization with the encoded signal, followed by detection of the digital tag. The read tag is then compared with the tag obtained from the tag generator 14 previously inserted into the signal, whereupon the decision block 18 determines, based on the result of the attempt to read the tag, whether the tag is present in the analyzed signal and turns the control block 10 on or off, respectively. controls the adaptation of digital adaptive filter 9. Decision block 18 provides binary information: if the digital tag is not detected, it means stopping the tuning process of digital adaptive filter 9, while if digital tag is detected it means no simultaneous speech, so tuning of digital adaptive filter 9 should resume. Regardless of the detection result of the digital tag, the acoustic echo estimate obtained using the digital adaptive filter 9 is subtracted from the signal from the microphone 7, and the processed signal is subjected to residual echo suppression in the signal dynamics processor 12, after which the signal is encoded in a speech coding block 13 and transmitted to the telecommunications network 1.
Contrary to the known techniques and applications of signal marking, the present solution detects the mere presence of a digital marker, and its content is not read. The expected content of the tag is known, and only its presence in the analyzed signal is found in the detection process. For this reason, the marker signal is chosen so that its presence in the signal is detectable despite the distortion of the signal containing the marker introduced by the acoustic feedback loop 5 between the loudspeaker 4 and the microphone 7 and noise and other external disturbances, and at the same time that the speech signal of a close user 6 of the terminal caused the tag to be suppressed so that it was impossible to detect it, which allows for simultaneous speech to be confirmed.
2 sheets
Sheet 1 Sheet 2
6 members in 3 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 38461608 | Poland | A | |
| PL20080384616 | – | – | – |
Members6
| Document | Office | Kind | |
|---|---|---|---|
| WO2009110809A1 | World Intellectual Property Organization (WIPO) | A1 | |
| PL384616A1 | Poland | A1 | |
| US2011002458A1 | United States of America | A1 | |
| US8588404B2 | United States of America | B2 | |
| PL216396B1This record | Poland | B1 | |
| US2014133648A1 | United States of America | A1 |
Numbers
- Publication
- 216396
- Publication, DOCDB
- 216396
- Publication, EPODOC
- PL216396B
- Application
- 384616
- Application, DOCDB
- 38461608
- Application, EPODOC
- PL20080384616
Titles2
- English
- The manner and system of acoustic echo dampening in VoIP terminal
- Polish
- Sposób i układ tłumienia echa akustycznego w terminalu VoIP
Classification
- CPC, 2
- H04B3/234
- H04M9/082
- IPC, 2
- H04B3 23
- H04M9 08