Noise reduction method in a speech signal.
Abstract
Selon l'invention, le procédé consiste à: convertir cycliquement par transformée de Fourier des signaux numériques, résultant de conversions analogiques-numériques de signaux fournis par deux microphones éloignés d'une distance fixe et recevant ledit signal sonore, en deux séries de données discrètes, chaque donnée discrète desdites séries étant représentative de l'énergie et de la phase d'un canal de fréquence du spectre dudit signal sonore; déterminer l'angle d'arrivée dominant dudit signal sonore à partir des différences de phase entre les données discrètes correspondant aux mêmes canaux de fréquence desdites séries, ledit angle d'arrivée dominant correspondant à l'angle d'arrivée dudit signal de parole; obtenir un spectre instantané dudit signal sonore; mettre à jour un spectre de bruit en comparant, pour chaque canal de fréquence dudit spectre instantané, la valeur absolue de la différence entre ledit angle d'arrivée dominant et l'angle d'arrivée du canal de fréquence considéré avec une valeur de seuil de tolérance, les énergies desdits canaux de fréquence dudit spectre de bruit étant mises à jour à l'aide des énergies des canaux de fréquence dudit spectre instantané dont la valeur absolue de la différence entre ledit angle d'arrivée dominant et l'angle d'arrivée de ces canaux de fréquence est supérieure à ladite valeur de seuil de tolérance; soustraire ledit spectre de bruit mis à jour dudit spectre instantané pour obtenir un spectre dudit signal de parole.

Term
Term ended
Projected expiry passed 11 February 2013, 13.6 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
10 claims: 1 independent, 9 dependent
- 1Procédé de réduction de bruit acoustique compris dans un signal sonore reçu comprenant un signal de parole, du type consistant à soustraire les composantes spectrales de bruit dudit signal sonore reçu pour reconstituer le spectre dudit signal de parole, caractérisé en ce qu'il consiste à:- convertir (16,17) cycliquement par transformée de Fourier des signaux numériques, résultant de conversions analogiques-numériques (12,13) de signaux fournis par deux microphones (10,11) éloignés d'une distance fixe et recevant ledit signal sonore, en deux séries (S1,S2) de données discrètes, chaque donnée discrète desdites séries étant représentative de l'énergie et de la phase d'un canal de fréquence du spectre dudit signal sonore, lesdits canaux de fréquence étant adjacents de façon à être représentatifs dudit spectre dudit signal sonore reçu;- déterminer (18,19) l'angle d'arrivée dominant (ϑ max ) dudit signal sonore reçu à partir des différences de phase entre les données discrètes correspondant aux mêmes canaux de fréquence desdites séries (S1,S2), ledit angle d'arrivée dominant (ϑ max ) correspondant à l'angle d'arrivée dudit signal de parole;- obtenir un spectre instantané dudit signal sonore reçu correspondant à une desdites séries (S1,S2) de données discrètes ou obtenu par une combinaison (20,21) desdites séries (S1,S2) discrètes permettant d'obtenir une amplification dudit signal de parole par rapport audit bruit;- mettre à jour (22) un spectre de bruit (23) en comparant, pour chaque canal de fréquence dudit spectre instantané, la valeur absolue de la différence entre ledit angle d'arrivée dominant (ϑ max ) et l'angle d'arrivée (ϑ) du canal de fréquence considéré avec une valeur de seuil de tolérance (ϑ s ), ledit spectre de bruit (23) étant constitué des mêmes canaux de fréquence que ledit spectre instantané, les énergies desdits canaux de fréquence dudit spectre de bruit (23) étant mises à jour à l'aide des énergies des canaux de fréquence dudit spectre instantané dont la valeur absolue de la différence entre ledit angle d'arrivée dominant (ϑ max ) et l'angle d'arrivée (ϑ) de ces canaux de fréquence est supérieure à ladite valeur de seuil de tolérance (ϑ s );- soustraire (24) ledit spectre de bruit (23) mis à jour dudit spectre instantané pour obtenir un spectre de sortie constitué par ledit spectre dudit signal de parole.
- 2Procédé selon la revendication 1, caractérisé en ce qu'il consiste à corriger ledit spectre de bruit (23) mis à jour en fonction du résultat de ladite soustraction (24).
- 3Procédé selon la revendication 2, caractérisé en ce que ladite correction (24) dudit spectre de bruit (23) mis à jour consiste à compter, après ladite soustraction, le nombre de canaux de fréquence dont l'énergie est supérieure à une valeur (Sp) de seuil d'énergie, et à remplacer ledit spectre de bruit (23) mis à jour par la totalité dudit spectre instantané si le nombre desdits canaux de fréquence dont l'énergie est supérieure à ladite valeur (Sp) de seuil d'énergie est inférieur à une valeur numérique prédéterminée.
- 4Procédé selon la revendication 3, caractérisé en ce que les énergies des données discrètes du résultat de ladite soustraction supérieures à ladite valeur (Sp) de seuil d'énergie sont forcées à zéro avant que ledit résultat ne remplace ledit spectre de sortie.
- 5Procédé selon l'une des revendications 2 à 4, caractérisé en ce que ladite correction dudit spectre de bruit (23) mis à jour consiste à compter, après ladite soustraction, le nombre de canaux de fréquence dont l'énergie est supérieure à ladite valeur (Sp) de seuil d'énergie, et à remplacer les canaux dudit spectre de bruit (23) mis à jour par les canaux dudit spectre instantané ayant produit un résultat négatif après ladite soustraction (24) si le nombre desdits canaux de fréquence dont l'énergie est supérieure à ladite valeur (Sp) de seuil d'énergie est supérieur à ladite valeur numérique prédéterminée.
- 6Procédé selon l'une des revendications 1 à 5, caractérisé en ce que ladite détermination (18,19) dudit angle d'arrivée dominant (ϑ max ) est réalisée en sommant dans des cases mémoire, pour chacun desdits canaux de fréquence, des poids proportionnels à l'énergie desdits canaux de fréquence, chaque case mémoire correspondant à un intervalle d'angle d'arrivée, lesdits poids étant sommés dans les cases mémoire correspondant aux angles d'arrivée (ϑ), desdits canaux de fréquence, ledit angle d'arrivée dominant (ϑ max ) correspondant à l'angle d'arrivée (ϑ) affecté à la case mémoire dont le poids est le plus important.
- 7Procédé selon la revendication 6, caractérisé en ce que lesdits poids sont également proportionnels aux fréquences desdits canaux de fréquence.
- 8Procédé selon l'une des revendications 6 et 7, caractérisé en ce que lesdites sommations consistent à effectuer des moyennes glissantes.
- 9Procédé selon l'une des revendications 1 à 8, caractérisé en ce que ladite combinaison (20,21) desdites séries (S1,S2) discrètes permettant d'obtenir une amplification dudit signal de parole par rapport audit bruit consiste à:- mettre en phase (20) lesdites données discrètes d'une (S2) desdites séries avec celles de l'autre (S1) desdites séries à partir dudit angle d'arrivée dominant (ϑ max ) de façon à mettre en phase les données discrètes desdites séries (S1,S2) dont l'angle d'arrivée correspond audit angle d'arrivée dominant (ϑ max );- sommer (21) lesdites données discrètes desdites séries (S1,S2) en phase afin d'amplifier les données discrètes correspondant audit signal de parole par rapport aux données discrètes correspondant audit bruit acoustique.
- 10Procédé selon l'une des revendications 1 à 9, caractérisé en ce qu'il est appliqué pour le traitement d'un signal de parole dans un radiotéléphone.
Independent claims10
68 paragraphs, as filed
0001The field of the invention is that of methods for reducing the acoustic noise present in a speech signal.
0002In known manner, when taking sound in a noisy environment, it is advantageous to eliminate the ambient noise so that it is not taken into account by means of recording or retransmission of the sound signal. This latter scenario is found in particular in the field of mobile radiotelephony, where it is desirable not to transmit ambient noise, for example that of the engine of a vehicle in which a radiotelephone is used, to the recipient of the speech signal. .
0003The scientific article "Acoustic noise analysis and speech enhancement techniques for mobile radio applications" by DAL DEGAN and PRATI, Elsevier Science Publishers BV, Signal Processing 15 (Acoustic noise analysis and speech enhancement techniques for mobile radio applications) , 1988, p.43 to 56, describes and compares different techniques for processing noise in a speech signal taken from a motor vehicle.
0004According to this article, methods of signal processing are known which consist in carrying out an estimation of the ambient noise spectrum and in subtracting this noise spectrum from the spectrum of the measured signal coming from a microphone. This process based on the principle of noise cancellation by spectral subtraction is also described in the article "Suppression of acoustic noise in speech using spectral substraction" by SF BOLL, IEEE Trans. ASSP. Flight. ASSP-27, 1979, pages 113 to 120.
0005However, the main drawback of this method is that it is necessary to carry out frequent updates of the noise spectrum in order to take account of changes in ambient noise and that this update can only be done when the user do not speak, that is to say during periods of silence. Thus, in an environment where ambient noise often and significantly varies, especially in a motor vehicle, it is necessary to have many periods of silence to frequently update the ambient noise spectrum. However, there are not always sufficiently long periods of silence to update the noise spectrum and, when the periods of silence are too far apart, there is a degradation of the noise spectrum which can no longer take account of brief noises. The quality of the transmitted speech signals is thereby affected.
0006The present invention aims in particular to overcome these drawbacks.
0007More specifically, one of the objectives of the invention is to provide a method of processing sound signals making it possible to significantly attenuate the ambient noise and thus to increase the quality of a transmission of speech signals, this noise attenuation being performed from an updated noise spectrum without requiring a period of silence on the part of the speaker.
0008This objective, as well as others which will appear subsequently, is achieved by means of a method of reducing acoustic noise included in a received sound signal comprising a speech signal, this method being of the type consisting in subtracting the spectral components of noise. of said received sound signal to reconstruct the spectrum of said speech signal and consists of:<ul id="ul0001" list-style="dash"><li>convert cyclically by Fourier transform digital signals, resulting from analog-digital conversions of signals supplied by two microphones distant from a fixed distance and receiving said sound signal, into two series of discrete data, each discrete data of said series being representative of l energy and phase of a frequency channel of the spectrum of said sound signal, said frequency channels being adjacent so as to be representative of said spectrum of said received sound signal;</li><li>determining the dominant angle of arrival of said sound signal received from the phase differences between the discrete data corresponding to the same frequency channels of said series, said dominant angle of arrival corresponding to the angle of arrival of said speech signal;</li><li>obtaining an instantaneous spectrum of said received sound signal corresponding to one of said series of discrete data or obtained by a combination of said discrete series making it possible to obtain an amplification of said speech signal with respect to said noise;</li><li>update a noise spectrum by comparing, for each frequency channel of said instantaneous spectrum, the absolute value of the difference between said dominant angle of arrival and the angle of arrival of the frequency channel considered with a threshold value of tolerance, said noise spectrum consisting of the same frequency channels as said instantaneous spectrum, the energies of said frequency channels of said noise spectrum being updated using the energies of the frequency channels of said instantaneous spectrum whose absolute value of the difference between said dominant angle of arrival and the angle of arrival of these frequency channels is greater than said tolerance threshold value;</li><li>subtracting said updated noise spectrum from said instantaneous spectrum to obtain an output spectrum consisting of the spectrum of said speech signal.</li></ul>
0009The principle of this method is therefore based on the evaluation of a dominant angle of arrival corresponding to the position of the speaker with respect to the two microphones receiving the sound signal, in order to extract the speech signal from the noise signal by spectral subtraction.
0010Advantageously, the method also comprises a step of correcting said updated noise spectrum as a function of the result of said subtraction.
0011Preferably, said correction of said updated noise spectrum consists in counting, after said subtraction, the number of frequency channels whose energy is greater than an energy threshold value, and in replacing said updated noise spectrum by the whole of said instantaneous spectrum if the number of said frequency channels whose energy is greater than said energy threshold value is less than a predetermined digital value.
0012Advantageously, the energies of the discrete data of the result of said subtraction greater than said energy threshold value are forced to zero before said result constitutes said output spectrum.
0013This eliminates the high-amplitude noise carrier frequencies in the output spectrum.
0014Preferably and according to a complementary embodiment, said correction of said updated noise spectrum consists in counting, after said subtraction, the number of frequency channels whose energy is greater than said energy threshold value, and replacing the channels of said updated noise spectrum with the channels of said instantaneous spectrum having produced a negative result after said subtraction if the number of said frequency channels whose energy is greater than said energy threshold value is greater than said predetermined numerical value.
0015According to a preferred embodiment, said determination of said dominant angle of arrival is carried out by summing in memory boxes, for each of said frequency channels, weights proportional to the energy of said frequency channels, each memory box corresponding to an interval angle of arrival, said weights being summed in the memory boxes corresponding to the angles of arrival of said frequency channels, said dominant angle of arrival corresponding to the angle of arrival assigned to the memory box whose weight is the most important.
0016A dominant arrival angle is thus determined, corresponding to the position of the speaker with respect to the two microphones if this speaker emits a speech signal.
0017Advantageously, said weights are also proportional to the frequencies of said frequency channels and said summations consist in performing sliding averages.
0018According to a preferred embodiment, said combination of said discrete series making it possible to obtain an amplification of said speech signal with respect to said noise consists in:<ul id="ul0002" list-style="dash"><li>phasing said discrete data of one of said series with that of the other of said series from said dominant angle of arrival so as to phase the discrete data of said series whose angle of arrival corresponds to said angle of dominant arrival;</li><li>summing said discrete data from said series in phase in order to amplify the discrete data corresponding to said speech signal with respect to the discrete data corresponding to said acoustic noise.</li></ul>
0019The method of the invention is preferably applied to the processing of a speech signal in a radiotelephone.
0020Other characteristics and advantages of the invention will appear on reading the following description of a preferred embodiment of the method of the invention, given by way of illustration and not limitation, this method being implemented in a device, the block diagram of which is shown in the attached single figure.
0021The process of the present invention can be broken down into 7 successive steps which will each be detailed below:<ul id="ul0003" list-style="dash"><li>a first signal processing step consists in digitizing the signals supplied by two fixed microphones and in carrying out Fourier transforms on these digitized signals to obtain two series of discrete data, each discrete data of a series being representative of the energy and the phase of a given frequency channel in the spectrum of the sound signal received by the two microphones;</li><li>a second processing step consists in determining the angle of arrival of the signal picked up by the two fixed microphones from the phase differences existing between two identical frequency channels of the discrete series. Knowing this angle of arrival makes it possible to know the position of the speaker with respect to the two microphones;</li><li>a third processing step consists in recombining in phase the speech signals supplied by the two fixed microphones to increase the power of the speech signal compared to the noise;</li><li>a fourth processing step consists in updating the noise spectrum from the angle of arrival of the speech signal;</li><li>a fifth processing step consists in subtracting the noise spectrum from the instantaneous spectrum measured to obtain an output spectrum;</li><li>a sixth processing step consists in correcting the noise spectrum as a function of the result of the subtraction;</li><li>a seventh and last processing step consists in reconstructing the output signal to allow for example its emission (radiotelephone application).</li></ul>
0022The first step in signal processing consists in digitizing the analog signals supplied by two microphones and in applying a Fourier transform to these digital signals to obtain two series of digital signals.
0023The microphones, referenced 10 and 11 in the appended figure, form an acoustic antenna. They are both fixed and the device implementing the invention is therefore preferably applied to a hands-free system. The analog signals coming from microphones 10 and 11 are applied to analog-digital converters respectively denoted 12 and 13 also ensuring filtering of the signals (frequency band 300-3400 Hz). The digitized signals are then admitted into Hamming windows 14,15 in digital form and then converted into digital vectors by devices 16,17 for calculating fast Fourier transforms (FFT) of order N.
0024The analog-digital converters 12 and 13 each supply 256 digital values sequentially to an output buffer, not shown, comprising 512 memory locations. At an instant t, the buffer supplies 512 digital values to a Hamming window, these 512 digital values being constituted by the 256 values calculated at time t put following 256 other values calculated at time t-1. Each Hamming window 14 and 15 provides attenuation of the secondary lobes of the received signal to allow an increase in the resolution of the signals by the FFTs 16 and 17. Each of these windows 14 and 15 provides 512 digital values to one of the devices 16 and 17 The latter each provide a series of 512 frequency channels sharing the frequency band 0 to 8 kHz. Each frequency channel therefore measures a little more than 15 Hz. A clock H of timing is used in the devices for calculating FFT 14 and 15.
0025This clock can also be replaced by a device for counting the number of samples supplied by the FFTs 14 and 15, this counting device being reset when 512 treatments have been carried out.
0026The results S1 and S2 of these calculations therefore consist of a succession of vectors in digital form, each vector corresponding to a sampling of the spectrum of one of the input signals and being for example made up of two 16-bit words defining a complex number.
0027In reality, only 256 different vectors can be taken into account, since, in the case of a real signal (in the mathematical sense of the term), the modulus of the Fourier transform is an even function while the phase is odd. The remaining 256 vectors are not taken into account. At the output of FFT, each channel therefore provides 256 vectors each composed of two 16-bit words.
0028The second processing step consists in determining the angle of arrival of the signal picked up by the two fixed microphones from the phase differences existing between two identical frequency channels of the discrete series.
0029The vectors from the FFTs are successively supplied to a device 18 for calculating the phase shift between the signals from the microphones 10 and 11. The S1 and S2 series are respectively characterized by identical frequency channels, to each frequency channel of each input channel corresponding to a phase Φ1 and Φ2 and different modules if the sound signal arriving on the microphone 10 is not strictly identical to that arriving on the microphone 11 (phase shift of the signals due to the propagation delay).
0030The signals S1 and S2 therefore consist of series of vectors, each pair of vectors corresponding to a frequency channel, and are supplied to a device 18 for calculating the phase shift, frequency channel by frequency channel, existing between the signals S1 and S2.
0031The distance between the two microphones being known, by approximating the sound signal to a plane wave, we can find the angle of arrival of the sound signal in each frequency channel by the relation:<maths id="math0001"><math display="block"><mrow><mtext>sinϑ = </mtext><mfrac><mrow><mtext>v.δΦ</mtext></mrow><mrow><mtext>2.π.df</mtext></mrow></mfrac></mrow></math><img file="EP0557166A1_D0001.tif" /></maths> or:<ul id="ul0004" list-style="dash"><li>ϑ is the desired angle of arrival of the sound signal of the frequency channel considered;</li><li>v is the speed of sound;</li><li>δΦ is the phase difference between the two signals;</li><li>d is the distance between microphones 10 and 11;</li><li>f is the frequency in Hz corresponding to the frequency channel considered. The frequency f is for example the center frequency of this frequency channel.</li></ul>
0032The calculation device 18 therefore provides a calculation of the phase shift between the signals from the two microphones, channel by channel.
0033The device 18 supplies the angles of arrival ϑ calculated for the different frequency channels to a device 19 for finding the angle of arrival. The device 19 updates a histogram of the angles of arrival comprising m boxes covering the angles -90 to + 90 °. Each box therefore has a length of 180 / m degrees.
0034Updating the histogram consists, for each frequency channel, of adding, in the box corresponding to the angle of arrival calculated by the calculation device 18, a weight proportional to the frequency and proportional to the energy present in the channel considered (amplitude of the spectral component). The added weight is preferably proportional to the frequency because the determination of the angle is more reliable in the high frequencies which verify both a better approximation of the plane wave and a lower δf / f. The value stored in a box is in fact the result of a sliding average calculated according to the relation:<maths id="math0002"><math display="block"><mrow><mtext>c (n) = ac (n-1) + (1-a). weight (n)</mtext></mrow></math><img file="EP0557166A1_D0002.tif" /></maths> or:<ul id="ul0005" list-style="dash"><li>c (n) is the value contained in a cell of the histogram at time n;</li><li>a is a real less than 1 and close to 1;</li><li>weight (n) is the value of the weight at time n. This value is for example equal to the energy of the channel considered multiplied by the frequency of this channel.</li></ul>
0035If there is no speech signal in the sound signal received by the two microphones 10 and 11, the values stored in the different memory boxes decrease as new data arrive, so that ultimately, the weights of the various memory cells are substantially equal to each other, the noise being distributed uniformly in the various cells if it does not come from a localized source, such as for example the engine of a vehicle.
0036The histogram is therefore updated periodically, for example every 32 ms.
0037Means not shown also make it possible to block the updating of the histogram when a speech signal is received by the user.
0038When the update has been carried out on all the channels, the device 17 searches for the maximum of the histogram, that is to say the most significant weight box. The location of this box, that is to say the angle of arrival assigned to it, corresponds to the dominant angle of arrival. This dominant angle of arrival, hereinafter referred to as ϑ<sub>max</sub>, is the one under which the sound signals arrive with the most important energy.
0039When the audio signal includes speech, the dominant angle of arrival ϑ<sub>max</sub> corresponds to the position of the speaker in relation to the two microphones. In fact, the noise frequencies from non-localized sources are distributed almost uniformly in the different boxes of the histogram, while the speech frequencies from a localized source (coming from the speaker) will always accumulate in the same box, quickly showing a peak in the histogram of the box corresponding to the dominant angle of arrival ϑ<sub>max</sub>.
0040According to another embodiment, as many histograms are created as there are frequency channels and an average is made, on all the channels, of the different histograms, in order to detect the dominant angle of arrival. This mode of implementation however requires storage means of larger size and this is why it is preferable to calculate sliding averages for each memory cell.
0041Other embodiments making it possible to know the angle of arrival of the speech signal can also be implemented.
0042The second step of processing the signal of the method of the invention therefore makes it possible to know the angle of arrival ϑ<sub>max</sub> speech signals.
0043The third stage of signal processing consists in combining in phase the signals supplied by the two microphones. The purpose of this combination is to amplify the speech signal relative to the noise signal. This step involves the blocks 20 and 21 which are respectively means for re-phasing the two channels and means for adding the rephased channels.
0044The device 19 for finding the angle of arrival provides the means 20 for re-phasing the value ϑ<sub>max</sub> of the dominant angle of arrival. The means 20 calculate, for each frequency channel, the phase difference existing between the two input channels for the dominant angle of arrival ϑ<sub>max</sub> which is supplied to it by the device 19. This calculation is carried out on the basis of the preceding relation, where ϑ is replaced by ϑ<sub>max</sub>, that is to say:<maths id="math0003"><math display="block"><mrow><msub><mrow><mtext>δΦ = (2.π.dfsinϑ</mtext></mrow><mrow><mtext>max</mtext></mrow></msub><mtext>) / v</mtext></mrow></math><img file="EP0557166A1_D0003.tif" /></maths>
0045The phase difference obtained for each frequency channel is summed (or subtracted, depending on the calculation method of δΦ) from the phase of one of the two signals. In the embodiment shown, the phase difference is added (subtracted) to (from) S2. The re-phasing means 20 therefore make it possible to obtain a signal S2 whose frequency channels corresponding to the speech signal are in phase with those of the signal S1 (since these frequency channels carry the greatest energy making it possible to determine ϑ<sub>max</sub>).
0046Then, the means 21 for adding the rephased channels sum the signal S1 with the rephased signal S2. By adding the rephased signals of the two channels, the speech signal is summed coherently and a speech signal of large amplitude is therefore obtained. On the other hand, the noise signal is weakened compared to the speech signal thus obtained due to the spectral spread of the noise (the noise signal does not come from a localized source like the speech signal). The summation of the rephased signals thus constitutes an amplification of the speech signal with respect to the noise signal.
0047However, noise rejection is generally not sufficient since there is residual noise whose spectral components have the same angle of arrival as the speech signal. It is therefore necessary to carry out an additional processing step.
0048It should be noted that this third signal processing step is optional, one of the two signals, for example the signal S2, which can be used directly for the rest of the process. In this case, the signal from the summing means 21 is replaced by the signal S2.
0049It is also possible to use a larger number of fixed microphones. However, the use of digital signals corresponding to frequency components of signals picked up by several microphones complicates the algorithm for calculating the angle of arrival ϑ for each frequency channel and also the reshaping of these signals in order to allow for example their summation. It is also not interesting to use the signals corresponding to the sound signal sampled by a third microphone so that they replace those coming from the adder 21, because their angles of arrival will necessarily be different from that of the signals S1 and S2 and the dominant angle of arrival found will not be able to take account of these signals picked up by this third microphone.
0050The additional processing step constitutes the fourth step previously indicated and consists in updating the noise spectrum. This update of the noise spectrum is carried out on the one hand from the angle of arrival ϑ<sub>max</sub> recognized as being that of speech, and on the other hand from the instantaneous spectrum constituted by the series of digital data supplied by the means of addition 21.
0051For each frequency channel, a device 22 for updating the noise spectrum compares the angle of arrival ϑ calculated by the calculation device 18 with the angle of arrival ϑ<sub>max</sub> of the speech supplied to it by the device 19. The device 22 can for example compare the absolute value of the difference between ϑ<sub>max</sub> and ϑ, for each frequency channel, with a tolerance threshold ϑ<sub>s</sub>.
0052If the absolute value of the difference between these two angles is greater than the tolerance threshold ϑ<sub>s</sub>, the corresponding frequency channel is considered to belong to the noise spectrum. The energy present in this channel is then used to update the noise spectrum, for example by sliding average. This update can of course also consist in simply replacing part of the data of the noise spectrum by the corresponding data of the instantaneous spectrum. This noise spectrum is stored in a digital memory 23. The introduction of a tolerance threshold ϑ<sub>s</sub> allows to take into account small variations in the speaker's position relative to the two microphones and also calculation inaccuracies.
0053If the absolute value of the difference between the two angles ϑ<sub>max</sub> and ϑ is below the tolerance threshold ϑ<sub>s</sub>, the observed frequency channel is considered to belong to the speech spectrum and its energy is therefore not used to update the noise spectrum in memory 23.
0054The device 22 therefore makes it possible to update the noise spectrum by comparing the angle of arrival ϑ of each frequency channel with the dominant angle of arrival ϑ<sub>max</sub>. This angle ϑ<sub>max</sub> calculated therefore has the function of allowing a selection of the frequencies of the spectrum obtained by the FFTs. Of course, it is not essential to make an absolute value of the difference between the dominant angle of arrival and the angle of arrival of each frequency channel of the instantaneous spectrum. It is for example possible to define a range of angle of arrival, of width 2ϑ<sub>s</sub> and centered on ϑ<sub>max</sub>, and to check if the angle of arrival ϑ of each frequency channel is within this range.
0055It should be noted that the updating of the noise spectrum is carried out continuously, every 32 ms, whether or not there is a speech signal in the signal received by the two microphones 10,11. The method of the invention therefore differs from the aforementioned state of the art in that it is not necessary to have periods of silence to allow the updating of the noise spectrum, the determination of membership of each frequency channel with the noise spectrum or with the spectrum of the speech signal being carried out starting from the calculated dominant angle of arrival and the angle of arrival for the channel considered.
0056The fifth step is to subtract the noise spectrum from the measured instantaneous spectrum.
0057This step uses a device 24 for subtracting the noise spectrum from the instantaneous spectrum. The noise spectrum is read in the digital memory 23 and subtracted from the instantaneous spectrum coming from the device 22. If the amplification constituting the second step of the present method is not implemented, the instantaneous spectrum is constituted by the vectors of one of the two signals, for example those of signal S2.
0058This subtraction makes it possible to obtain an output spectrum consisting of a spectrum of the speech signal almost entirely devoid of noise spectral components. It is however possible to carry out additional processing of the spectrum obtained, making it possible in particular to correct the updated noise spectrum.
0059After subtraction, negative results are forced to zero. Two cases can then arise:<ul id="ul0006" list-style="dash"><li>the instantaneous spectrum did not contain a speech signal and the residual spectrum then contains only a small number of significant frequencies;</li><li>the instantaneous spectrum contained a speech signal and the residual spectrum then comprises a large number of energy-carrying frequencies essentially corresponding to the speech spectrum.</li></ul>
0060In order to know the content of the instantaneous spectrum, it suffices therefore to count the number of frequency channels for which the spectral power is greater than a threshold value Sp making it possible not to take into account the frequency channels carrying low energy. , which can therefore be eliminated. These frequency channels correspond either to residual noise or to speech frequency channels but which are of such low energy that it is not necessary to transmit them to their recipient (radiotelephone application).
0061Preferably, the value of the threshold Sp is not the same for each frequency channel and depends on the energy present in each of these frequency channels. It is possible, for example, to assign a first threshold value to the channels of the frequency band 0-2 kHz, and a second threshold value, for example equal to half of the first threshold value, to the channels of the frequency band 2 -4 kHz. This makes it possible to take into account the fact that the energies of the noise spectrum are greater at low frequencies than at high frequencies in a vehicle.
0062If the number of frequency channels having an energy greater than Sp is small (less than a digital threshold value), the frequency channels counted are considered to have residual noise frequencies. The entire instantaneous spectrum (that is to say the data presented at the input of the device 22) is then used to update the noise spectrum. This operation takes place in the digital memory 23 and constitutes the sixth step of signal processing. It consists more precisely in replacing the energies of the frequency channels of the updated noise spectrum by the energies of the corresponding frequency channels of the instantaneous spectrum. In addition, frequency channels whose energy exceeds the threshold Sp are forced to zero before replacing the frequency channels of the noise spectrum. This eliminates the high-amplitude noise carrier frequencies.
0063In an alternative embodiment, only the frequency channels of the instantaneous spectrum having energies higher than those of the corresponding frequency channels of the noise spectrum are used for replacement. Only the frequency channels of the instantaneous spectrum having a high energy are therefore taken into account.
0064If the spectrum obtained after subtraction corresponds to a speech signal (i.e. the number of frequency channels having after subtraction an energy greater than Sp is greater than the digital threshold value), only the energies of the frequency channels of the instantaneous spectrum corresponding to the frequency channels of the residual spectrum after subtraction exhibiting a negative result are used to correct the noise spectrum. Indeed, a negative result after subtraction means that the corresponding frequency channel of the updated noise spectrum has too much energy. This correction makes it possible to avoid that the residual noise spectrum (that is to say the updated spectrum) no longer consists of only a few frequency channels of high amplitude, which would make it particularly unpleasant and disturbing to the sound reproduction. .
0065Of course, the correction of the noise spectrum constituting this sixth processing step is optional and can be carried out in different ways from the moment when it is decided whether the spectrum obtained by subtraction should or should not be considered as a spectrum containing channels. speech frequency to be used, for example to transmit to a recipient.
0066The last and seventh processing step consists in constructing an analog output signal to allow for example its transmission. This step implements a device 25 for generating the output signal comprising a device 26 for fast Fourier inverse transform (FFT⁻¹) providing 512 samples of the speech signal. The FFT⁻¹ device is preceded by a device (not shown) making it possible to regenerate the 256 vectors received to obtain 512 vectors at the input of the FFT⁻¹ device. The device 26 is followed by a recovery device 27 allowing easy reconstruction of the output signal. The device 27 ensures recovery of the first 256 samples received with the last 256 samples that it has previously received (forming part of the previous processing). This overlap compensates at the output for the application of an input Hamming window. A digital-analog converter 28 makes it possible to obtain a little noisy sound signal ready to be sent to its recipient. It is also possible to record this signal, for example on a magnetic tape, or to make it undergo another treatment.
0067This seventh step may not be implemented in certain applications. For example, the method of the invention could be applied to voice recognition, and, in this case, the seventh processing step can be omitted, since the voice recognition devices exploit the spectral representation of a speech signal.
0068The method of the invention therefore makes it possible to significantly reduce the noise spectrum of a sound signal to provide a speech signal, without requiring periods of silence from the speaker to update the noise spectrum, since knowledge the angle of arrival of the signal is used to separate noise from speech.
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both ways
| Document | Relation | Office | Category | Cited during |
|---|---|---|---|---|
| US6522756B1 | Cited by | United States of America | – | Applicant |
| GB2516314A | Cited by | United Kingdom | – | Search report |
| US6766029B1 | Cited by | United States of America | – | Applicant |
| WO9952097A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search |
| AU749652B2 | Cited by | Australia | – | Search report |
| FR2761800A1 | Cited by | France | – | Search report |
| WO9904598A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search |
| WO9904598A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search |
| WO9952097A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search |
| GB2289593B | Cited by | United Kingdom | – | Search report |
| EP0802699A2 | Cited by | European Patent Office (EPO) | – | Search report |
| GB2516314B | Cited by | United Kingdom | – | Search report |
| EP0802699A3 | Cited by | European Patent Office (EPO) | – | Search report |
| US5546458A | Cited by | United States of America | – | Search report |
| GB2289593A | Cited by | United Kingdom | – | Search report |
| US112430A | Cites | United States of America | – | Examiner |
| US4112430A | Cites | United States of America | A | Search report |
| US4112430A | Cites | United States of America | A | Search report |
| US4333170A | Cites | United States of America | A | Search report |
| US4333170A | Cites | United States of America | A | Search report |
| US4653102A | Cites | United States of America | A | Search report |
| US4653102A | Cites | United States of America | A | Search report |
| ALTA FREQUENZA vol. 53, no. 3, 1 Mai 1984, MILANO IT pages 190 - 195 AUDISIO, PIRANI 'Noisy speech enhancement: a comparative analysis of three different techniques' | Non-patent | – | – | Search report |
| IEEE TRANSACTIONS ON ACOUSTICS,SPEECH AND SIGNAL PROCESSING vol. 27, no. 2, Avril 1979, NEW YORK US pages 113 - 120 BOLL 'Suppression of acoustic noise in speech using spectral subtraction' | Non-patent | – | – | Search report |
| SIGNAL PROCESSING. EUROPEAN JOURNAL DEVOTED TO THE METHODS AND APPLICATIONS OF SIGNAL PROCESSING vol. 15, no. 1, Juillet 1988, AMSTERDAM NL pages 43 - 56 DAL DEGAN, PRATI 'Acoustic noise analysis and speech enhancement techniques for mobile radio applications' | Non-patent | – | – | Search report |
17 members in 11 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 9201819 | France | A | |
| 9201819 | France | A | |
| 9201819 | France | – | |
| 9201819 | – | – | – |
| FR19920001819 | – | – | – |
Members17
| Document | Office | Kind | |
|---|---|---|---|
| AU3285493A | Australia | A | |
| FI930655A | Finland | A | |
| FI930655A7 | Finland | A7 | |
| FR2687496A1 | France | A1 | |
| EP0557166A1This record | European Patent Office (EPO) | A1 | |
| FR2687496B1 | France | B1 | |
| AU662199B2 | Australia | B2 | |
| NZ245850A | New Zealand | A | |
| US5539859A | United States of America | A | |
| EP0557166B1 | European Patent Office (EPO) | B1 | |
| DK0557166T3 | Denmark | T3 | |
| ATE159373T1 | Austria | T1 | |
| DE69314514D1 | Germany | D1 | |
| ES2107635T3 | Spain | T3 | |
| DE69314514T2 | Germany | T2 | |
| GR3025804T3 | Greece | T3 | |
| FI104526B | Finland | B |
59 legal events, as 6 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Notification of lapseLapsedST | ST | FR | |
| Gb: european patent ceased through non-payment of renewal feeCeasedGBPC | GBPC | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Announcement of lapse in spainLapsedFD2A | FD2A | ES | |
| Nl: lapsed or anulled due to non-payment of the annual feeLapsedNLV4 | NLV4 | EP | |
| Ep patent lapsedLapsedEBP | EBP | DK | |
| Se: european patent has lapsedLapsedEUG | EUG | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| European patent in force as of 2002-01-01IF02 | IF02 | GB | |
| Patent ceasedCeasedPL | PL | CH | |
| Be: lapsedLapsedBERE | BERE | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Definitive protectionFG2A | FG2A | ES | |
| Corresponds to:REF | REF | EP | |
| Gb: translation of ep patent filed (gb section 77(6)(a)/1977)GBT | GBT | EP | |
| Ep patent with danish claimsT3 | T3 | DK | |
| It: translation for a ep patent filedITF | ITF | EP | |
| It: translation for a ep patent filedITF | ITF | EP | |
| Designated contracting statesAK | AK | EP | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| New agentNV | NV | CH | |
| Corresponds to:REF | REF | EP | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOS IGRAGRAH | GRAH | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOS IGRAGRAH | GRAH | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Despatch of communication of intention to grantORIGINAL CODE: EPIDOS AGRAGRAG | GRAG | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 0557166
- Publication, DOCDB
- 0557166
- Publication, EPODOC
- EP0557166
- Application
- 93400346
- Application, DOCDB
- 93400346
- Application, EPODOC
- EP19930400346
Titles3
- German
- Rauchverminderungsverfahren in einem Sprachsignal
- English
- Noise reduction method in a speech signal
- French
- Procédé de réduction de bruit acoustique dans un signal de parole
Classification
- CPC, 3
- G10L21/0208
- G01S3/808
- H04R3/005
- IPC, 3
- G01S3 808
- G10L21 0208
- H04R3 00
Designated states1
- Contracting states, 1
- Sweden