Esimation of background noise in audio signals
15 claims: 4 independent, 11 dependent
- 1Zastrzeżenia patentowe 1. Sposób szacowania szumu tła w sygnale audio, który to sposób obejmuje:- uzyskanie (201) co najmniej jednego parametru powiązanego z wejściowym segmentem sygnału audio na podstawie: - zysku pierwszej predykcji liniowej obliczonego jako iloraz energii sygnału wejściowego i resztkowej energii sygnału z pierwszej predykcji liniowej dla segmentu sygnału audio;i - zysku drugiej predykcji liniowej obliczonego jako iloraz resztkowej energii sygnału z pierwszej predykcji liniowej i resztkowej energii sygnału z drugiej predykcji liniowej dla segmentu sygnału audio;- ustalenie (202) na podstawie tego co najmniej jednego parametru, czy segment sygnału audio zawiera przerwę wolną od mowy i muzyki;i: jeśli ustalono, że segment sygnału audio zawiera przerwę: - aktualizowanie (203) oszacowanie szumu tła na podstawie segmentu sygnału audio.
- 2Sposób według zastrz. 1, w którym pierwsza predykcja liniowa jest predykcją liniową 2. rzędu, a druga predykcja liniowa jest predykcją liniową 16. rzędu.
- 3Sposób według zastrz. 1 albo 2, w którym uzyskanie tego co najmniej jednego parametru obejmuje:ograniczenie zysków pierwszej i drugiej predykcji liniowej do przyjmowania wartości z określonego z góry przedziału.
- 4Sposób według dowolnego z zastrz. od 1 do 3, w którym uzyskanie tego co najmniej jednego parametru obejmuje:stworzenie co najmniej jednego oszacowania długookresowego każdego z zysków pierwszej i drugiej predykcji liniowej, w którym oszacowanie długookresowe jest dalej oparte na odpowiednich zyskach predykcji liniowej powiązanych z co najmniej jednym poprzedzającym segmentem sygnału audio. EP 3 309 784 B1
- 5Sposób według dowolnego z zastrz. od 1 do 4, w którym uzyskanie tego co najmniej jednego parametru obejmuje:ustalenie różnicy między jednym z zysków predykcji liniowej powiązanym z segmentem sygnału audio a oszacowaniem długookresowym wymienionego zysku predykcji liniowej.
- 6Sposób według dowolnego z zastrz. od 1 do 5, w którym uzyskanie tego co najmniej jednego parametru obejmuje:ustalenie różnicy między dwoma oszacowaniami długookresowymi powiązanymi z jednym z zysków predykcji liniowej.
- 7Sposób według dowolnego z zastrz. od 1 do 6, w którym uzyskanie tego co najmniej jednego parametru obejmuje filtrowanie dolnoprzepustowe zysku pierwszej i drugiej predykcji liniowej.
- 8Sposób według zastrz. 7, w którym współczynniki filtra, co najmniej jednego filtra dolnoprzepustowego, zależą od związku między zyskiem predykcji liniowej powiązanego z segmentem sygnału audio, a średnią odpowiedniego zysku predykcji liniowej otrzymaną na podstawie wielu poprzedzających segmentów sygnału audio.
- 9Sposób według dowolnego z poprzednich zastrzeżeń, w którym ustalenie, czy segment sygnału audio zawiera przerwę jest dalej oparte na mierze bliskości widmowej powiązanego z segmentem sygnału audio.
- 10Sposób według zastrz. 9, obejmujący ponadto uzyskanie miary bliskości widmowej na podstawie energii dla zbioru pasm częstotliwości segmentu sygnału aaudio i oszacowań szumu tła odpowiadających temu zbiorowi pasm częstotliwości.
- 11Sposób według zastrz. 10, w którym w okresie inicjalizacji wartość początkowa Emin jest stosowana jako oszacowania szumu tła na podstawie których otrzymywana jest miara bliskości widmowej.
- 12Aparatura (1100), do szacowania szumu tła w sygnale audio, obejmującym wiele segmentów sygnału audio, która to aparatura jest przystosowana do:- uzyskania co najmniej jednego parametru na podstawie: - zysku pierwszej predykcji liniowej obliczonego jako iloraz energii segmentu sygnału audio i resztkowej energii sygnału z pierwszej predykcji liniowej dla segmentu sygnału audio;i - zysku drugiej predykcji liniowej obliczonego jako iloraz resztkowej energii sygnału z pierwszej predykcji liniowej i resztkowej energii sygnału z drugiej predykcji liniowej dla segmentu sygnału audio;- ustalenia na podstawie tego co najmniej jednego parametru, czy segment sygnału audio zawiera przerwę wolną od mowy i muzyki;i: jeśli ustalono, że segment sygnału audio zawiera przerwę: - aktualizacji oszacowania szumu tła na podstawie segmentu sygnału audio.
- 13Aparatura według zastrz. 12, gdzie aparatura jest dalej przystosowana do wykonywania sposobu według dowolnego z zastrz. od 2 do 11.
- 14Kodek audio zawierający aparaturę według zastrz. 12 albo 13.
- 15Urządzenie komunikacyjne zawierające aparaturę według zastrz. 12 albo 13. EP 3 309 784 B1 ODNOŚNIKI CYTOWANE W OPISIE Poniższa lista odnośników cytowanych przez zgłaszającego ma na celu wyłącznie pomoc dla czytającego i nie stanowi części dokumentu patentu europejskiego. Pomimo, że dołożono największej staranności przy jej tworzeniu, nie można wykluczyć błędów lub przeoczeń i EUP nie ponosi żadnej odpowiedzialności w tym względzie. Dokumenty patentowe cytowane w opisie • WO 2011049514 A [0023] [0070] [0071] [0072] [0077] [0147] [0150] [0152] [0153] [0164] • WO 2011049515 A [0023] [0147] • WO 201109514 A [0148] [0167] • WO 201109515 A [0149] Literatura niepatentowa cytowana w opisie • M. JELINEK ;R. SALAMI. Noise reduction method for wideband speech coding. 12th European signal processing conference, 2004, 1959-1962 [0013] EP 3 309 784 Β1 EP 3 309 784 Β1 Fig. 2 Fig. 3 EP 3 309 784 Β1 Fig. 4 Energie podpasm Ncb(i) po wszystkich i=2...16 Fig. 5 EP 3 309 784 Β1 EP 3 309 784 Β1 Fig. 7 EP 3 309 784 Β1 Działanie cech E(0) i E(2) Numer ramki pjg g Działanie cech E(2) i E(16) *| 1200 1300 1400 1500 1600 1700 Numer ramki Fig. 9a EP 3 309 784 Β1 Działanie cech E(2) i E(16) 0.1 Etot ° θ_2_16 * Gd_2_16 _2------1-------1-------1--------1---- 200 1300 1400 1500 16001700 Numer ramki Fig.9b Działanie cech E(2) i E(16)3 1 ? 0.1 Et ° G_2_16 + Gmax 2 16 UJ 0 .4 1200 1300 1400 1500 16001700 Numer ramki Fig.9c EP 3 309 784 Β1 Fig. 10 EP 3 309 784 Β1 Fig. 11a Obwody przetwarzające 1101 Interfejs 1102 Estymator szumu tła 1100 Fig. 11b Fig. 11c EP 3 309 784 Β1 Estymator szumu tła 1200 Fig. 12 Fig. 13 EP 3 309 784 Β1 -209 Fig. A2 EP 3 309 784 B1 Estymator szumu tła Fig. A3 EP 3 309 784 B1 Estymator szumu tła Fig. A4 EP 3 309 784 Β1 EP 3 309 784 B1 Śledzenie zbyt wolne ooooooooooo 00 t Ό γ»Ί CM — —I CM I I (gp) luoizod Fig. A6 EP 3 309 784 B1 Śledzenie zbyt wolne Numer ramki EP 3 309 784 B1 O Śledzenie zbyt szybkie Numer ramki (gp) luoizoj EP 3 309 784 B1 O Odzysk muzyki (gp) luoizoj Fig. A9
Independent claims15
344 paragraphs in 1 section, as filed
Description
TECHNICAL FIELD [0001] Embodiments of the present invention relate to audio signal processing, and in particular to the estimation of background noise, e.g. to support decisions about audio activity.
BACKGROUND [0002] In communication systems using discontinuous broadcasting (DTX), it is important to find a balance between performance and non-degradation of quality. In such systems, an activity detector is used to detect active signals, e.g. speech or music, which are to be actively encoded, and segments with background signals, which can be replaced by natural noise (comfort noise) generated on the recipient side. If the activity detector is too efficient at detecting inactivity, it will introduce clipping of the active signal, which will then be seen as a subjective decrease in quality when the clipped active segment is replaced by natural noise. At the same time, DTX efficiency is reduced if the activity detector is not efficient enough and classifies background noise segments as active, and then actively encodes background noise instead of going into DTX mode with natural noise. In most cases, the clipping problem is considered worse.
[0003] Fig. 1.shows a block overview of a generalized soundactivity detector (SAD) or voice activity detector (VAD) that receives the sound signal as an input signal and produces an activity decision as output signal. The input signal is divided into data frames, i.e. audio segments with a length of e.g. 5-30 ms, depending on the implementation, and one activity decision per frame is generated as an output signal.
[0004] The main detector, shown in Fig. 1, makes the main decision labeled "prim". The main decision is essentially only a comparison of the features of the current frame with the background features estimated from previous input frames. The difference between the features of the current frame and background features greater than the threshold value causes the main decision about the activity. Add Delay Block hangover) is used to extend the validity of the main decision based on the old main decisions to create the final decision marked "flag". The reason for using the delay is mainly to reduce / eliminate the risk of cutting activity spikes in the middle and at the end. As indicated in the figure, the operation controller can adjust the threshold value (s) for the main detector and the length of the added delay according to the characteristics of the input signal. The background noise estimator block is used to estimate the background noise in the input signal. Background noise can also be referred to as "background" or "background feature".
[0005] The estimation of the background feature can be carried out according to two fundamentally different principles, either by applying the main decision, i.e. decision feedback or decision metric marked with a dotted line in Fig. 1, or by using some other characteristic of the input signal, i.e. without feedback. Combinations of these two strategies can also be used.
[0006] An example of a codec using decision feedback for background estimation is AMR-NB (Adaptive Multi-Rate Narrowband), and examples of codecs that do not use decision feedback are EVRC (Enhanced Variable Rate) Codec, improved codec with variable bit rate) and G.718.
[0007] There are a number of different signal features or characteristics that can be used, but one in common
The feature used in VAD is the frequency characteristics of the input signal. A commonly used type of frequency response is the energy of the subband frame, due to its low complexity and reliable operation at low signal-to-noise ratio (SNR). It is therefore assumed that the input signal is divided into different frequency subbands and the background level is estimated for each of the subbands. In this way, one of the features of the background noise is a vector with energy values for each subband. It is these values that characterize the background noise in the input signal in the frequency domain.
[0008] To achieve background noise tracking, the estimation of actual background noise can be updated in at least three different ways. One of these methods is to use Auto Regressive (AR) for frequency update support. Examples of such codecs are AMR-NB and G.718. Basically, for this type of update, the size of the update step is proportional to the observed difference between the current input signal and the current background estimate. Another way is to use multiplicative scaling of the current estimate with the limitation that the estimate can never be greater than the current input signal or less than a certain minimum value. This means that the estimate is increased every frame until it becomes larger than the current input signal. In this situation, the current input signal is used as an estimate. EVRC is an example of a codec that uses this technique to update the background estimate for a VAD function. It should be noted that EVRC uses different background estimates for VAD and noise suppression. It should also be noted that VAD can be used in contexts other than DTX. For example, in variable rate codecs such as EVRC, VAD can be used as part of the bit rate determination function.
[0009] A third method is to use the so-called minimum technique, in which the estimate is the minimum value during the sliding time window of earlier frames. This generally gives a minimum estimate that is scaled using a compensation factor to obtain and approximate the average estimate for stationary noise.
[0010] In cases of high SNR, when the level of the active signal is significantly higher than the background signal, deciding whether the input audio signal is active or inactive can be quite easy. However, separating active signals from inactive signals in low SNR cases, and especially when the background is non-stationary or even similar to an active signal in terms of characteristics, is very difficult.
[0011] The efficiency of VAD depends on the ability to track the background characteristics of the background noise estimator - especially when it comes to non-stationary backgrounds. With better tracking, you can make VAD more efficient without increasing the risk of speech clipping.
[0012] Although correlation is an important feature used to detect speech, mainly the vocal part of speech, there are also noise signals showing a high correlation. In these cases, the correlation noise will prevent the background noise estimates from being updated. The result is high activity, because both speech and background noise are encoded as active content. While in the case of high SNR (approximately> 20 dB) it would be possible to reduce the problem by detecting the gap based on energy, this is not reliable for an SNR range of 20 dB to 10 dB or possibly 5 dB. It is in this range that the solution described here makes the difference.
[0013] M. Jelinek and R. Salami in the article "Noise reduction method for wideband speech coding" [Noise reduction method for wideband speech coding] 2004, 12<sup>th</sup> European signal processing conference [12. European Conference on Signal Processing], pp. 1959-1962, teach a method of estimating background noise, in which the occurrence of pauses during which the said noise is estimated is determined on the basis of
EP 3 309 784 B1 ratio of the remainder of the 2nd order linear prediction and the residue of the 16th order linear prediction.
SUMMARY OF THE INVENTION [0014] It would be desirable to achieve improved estimation of background noise in audio signals. "Improved" may mean making better decisions about whether or not the sound signal contains active speech or music, and therefore more frequent estimation, e.g., updating the previous estimate, background noise in sound signal segments that actually do not contain active content, such as speech and / or music. An improved method of generating background noise estimation is provided herein, which may allow e.g. an audio activity detector to make more appropriate decisions.
[0015] For estimating background noise in audio signals, it is important to be able to find reliable features to identify the characteristics of the background noise signal also when the input signal contains an unknown mixture of active signals and background signals in which active signals may contain speech and / or music.
[0016] The inventor has realized that the residual energy-related features for different orders of the linear prediction model can be used to detect gaps in audio signals. These residual energies can be extracted, e.g., from linear prediction analysis common in speech codecs. Features can be filtered and combined to create a set of features or parameters that can be used to detect background noise, making the solution suitable for use in noise estimation. The solution described here is particularly effective in conditions where the SNR is in the range of 10 to 20 dB.
[0017] A further feature provided here is a measure of spectral proximity to the background that can be created, e.g., by using subband energy in the frequency domain, used e.g. in the SAD subband. The spectral proximity measure can also be used to decide whether an audio signal has a break or not. [0018] According to a first aspect, a method of estimating background noise is provided. The method includes obtaining at least one parameter associated with an audio signal segment, such as a frame or frame portion, based on the gain of the first linear prediction, calculated as the quotient of the input signal energy and the residual signal energy of the first linear prediction for the audio signal segment; and the gain of the second linear prediction calculated as the quotient of the residual signal energy of the first linear prediction and the residual signal energy of the second linear prediction for the segment of the audio signal. The method further includes determining whether the audio signal segment includes a gap based on the at least one parameter; and updating the background noise estimate based on the audio signal segment if the audio signal segment has been determined to include a gap.
[0019] According to a second aspect, apparatus for estimating background noise in an audio signal is provided. This apparatus is adapted to obtain at least one parameter based on the gain of the first linear prediction, calculated as the quotient of the energy of the audio signal segment and the residual signal energy of the first linear prediction for the audio signal segment; and the gain of the second linear prediction calculated as the quotient of the residual signal energy of the first linear prediction and the residual signal energy of the second linear prediction for the segment of the audio signal. The background noise estimator is further adapted to determine whether the audio signal segment includes a gap based on at least one parameter; and for updating the background noise estimate based on the audio signal segment if the audio signal segment has been determined to include a gap.
[0020] According to a third aspect, there is provided an audio codec that includes the apparatus of the second aspect.
[0021] According to a fourth aspect, there is provided a communication device that includes the apparatus of the second aspect.
BRIEF DESCRIPTION OF THE FIGURES [0022] The above and other objects, features and advantages of the technology disclosed herein will be apparent from the following more detailed description of the embodiments as shown in the accompanying drawings. The drawings are not necessarily to scale, instead the emphasis is on presenting the principles of the technology disclosed here.
Fig. 1 is a block diagram showing the activity detector and the delay determination logic.
Fig. 2 is a flowchart showing a method of estimating background noise according to one embodiment.
Fig. 3 is a block diagram showing the calculation of characteristics associated with residual energies for a linear prediction of 0 and 2 according to one embodiment.
Fig. 4 is a block diagram showing the calculation of characteristics associated with residual energies for the linear prediction of the 2nd and 16th order according to one embodiment.
Fig. 5 is a block diagram showing the calculation of features associated with the spectral proximity measure according to one embodiment.
Fig. 6 is a block diagram showing the background noise energy estimate of the subband.
Fig. 7 is a flowchart showing the logic of the decision to update the background from the solution described in Annex A.
Figures 8.-10. are diagrams showing the behavior of the various parameters presented here calculated for an audio signal containing two strings of speech.
Figures 11a-11c and 12.-13. are block diagrams showing different embodiments of the background noise estimator according to the embodiments.
Figures A2-A9 on the pages of the figures marked "Annex A" are related to Annex A and discussed in this Annex A with the number following the letter "A", i.e. 2-9.
DETAILED DESCRIPTION [0023] The solution disclosed herein relates to estimating background noise in audio signals. In the generalized activity detector shown in Fig. 1, the background noise estimation function is performed by a block designated "background noise estimator". Some embodiments of the solution described herein may be seen in combination with the solutions disclosed previously in WO2011 / 049514, WO2011 / 049515, as well as in Annex A (Appendix A). The solution disclosed here will be compared with the embodiments of these solutions previously disclosed. Although the solutions disclosed in WO2011 / 049514, WO2011 / 049515 and Annex A are good solutions, the solution presented here has advantages over these solutions. For example, the solution presented here is even more suitable in terms of tracking background noise.
[0024] The efficiency of VAD depends on the ability to track the background characteristics of the background noise estimator - especially when it comes to non-stationary backgrounds. With better tracking, you can make VAD more efficient without increasing the risk of speech clipping.
[0025] One problem with current noise estimation methods is that to achieve good background noise tracking at low SNR, a reliable gap detector is needed. For a speech-only input, you can use the tempo to find speech gaps
Syllabic or that a person cannot speak all the time. Such solutions could mean that after a sufficiently long period of time without performing background updates, the requirements for gap detection are "relaxed" so that detection of a speech break becomes more likely. This allows you to respond to rapid changes in the characteristics or level of noise. Some examples of such noise recovery logic are: 1) As statements (speech) contain segments with high correlation, it is usually safe to assume that after a sufficient number of non-correlated frames there will be a speech break. 2) When the Signal to Noise Ratio (SNR)> 0, speech energy is greater than background noise, so if the frame energy is close to the minimum energy for a long time, e.g. 1-5 seconds, it's also safe to assume it's a speech break. While earlier techniques work well for a speech-only input, they are insufficient when music is considered an active input. There may be long segments with little correlation in music that are still music. Further, energy dynamics in music can also cause erroneous detection of a gap, which can lead to unwanted, erroneous updates of the background noise estimate.
[0026] Ideally, an inverse function of the activity detector or something that would be called a "break detector" would be needed to control the noise estimation. This would ensure that an update of the background noise characteristics would only be made if there is no active signal in the current frame. However, as indicated above, determining whether the audio signal segment contains an active signal or not is not an easy task.
[0027] Traditionally, when it was known that the active signal was a speech signal, the activity detector was called the Voice Activity Detector (VAD). The term "VAD" is often used for activity detectors also when the input signal can contain music. However, in modern codecs it is also common to specify an activity detector as an audio activity detector ( Sound Activity Detector, SAD) when music is also to be detected as an active signal. [0028] The background noise estimator shown in Fig. 1 uses feedback from the main detector and / or delay block to locate inactive segments of the audio signal. When developing the technology described herein, it was desirable to remove or at least reduce the dependence on such feedback. Therefore, for the background estimation disclosed herein, the inventor considered it important to be able to find reliable characteristics for identifying the characteristics of the background signals when only an input signal that is an unknown mixture of the active signal and the background signal is available. The creator also realized that it cannot be assumed that the input signal starts with a noise segment, or even that the input signal is speech mixed with noise, as it may be that the active signal is music.
[0029] One aspect is that even if the current frame may have the same energy level as the current noise estimate, the frequency characteristics may be quite different, making noise estimation updates using the current frame becoming undesirable. To prevent updates in these cases, you can apply the background noise update to the proximity feature.
[0030] Furthermore, during initialization, it is desirable to allow the estimation to be started as soon as possible avoiding incorrect decisions, as it could lead to SAD clipping if the background noise update were performed using active content. The use of an initialization-dependent version of the proximity feature during initialization may at least partially solve this problem.
[0031] The solution described herein relates to a method of estimating background noise, especially a method of detecting interruptions in an audio signal that works well in difficult SNR situations. This solution will be described below with reference to Figs. 2.-5.
[0032] In the field of speech coding, so-called linear prediction is commonly used to analyze the spectral shape of an input signal. This analysis is usually carried out twice per frame and for increased time accuracy, the results are then interpolated so that a filter is generated for each 5-millisecond input signal block.
[0033] Linear prediction is a mathematical operation in which future values of discontinuous signal over time are estimated as a linear function of previous samples. In digital signal processing, linear prediction is often called linearpredictive coding (LPC) and can therefore be considered a subset of filter theory. In the linear prediction in the speech encoder, the linear prediction filter A (z) is applied to the speech input signal. A (z) is a filter of all zeros, which when applied to the input signal removes redundancy that can be modeled using the filter A (z) from the input signal. Therefore, the output signal from the filter has less energy than the input signal when the filter successfully models some aspects or aspects of the input signal. This output signal is designated as "residual", "residual energy" or "residual signal". Such linear prediction filters, alternatively referred to as residual filters, may have a different model row with a different number of filter coefficients. For example, for correct speech modeling, a linear prediction filter with model order 16 may be required. Thus, a linear prediction filter A (z) with model order 16 may be used in the speech encoder.
[0034] The inventor has realized that features associated with linear prediction can be used to detect interruptions in audio signals in the SNR range from 20 dB to 10 dB or possibly 5 dB. According to the embodiments of the solution described herein, a relationship between residual energies for different orders of models for the audio signal is used to detect gaps in the audio signal. The compound used is the quotient of residual energy of the lower order model and the higher order model. The residual energy quotient can be referred to as "linear prediction gain" because it is an indicator of how much signal energy the linear prediction filter has managed to model, i.e. remove, between a model of one order and a model of another order.
[0035] The residual energy will depend on the order of the M model of the linear prediction filter A (z). A popular way of calculating the filter coefficients for a linear prediction filter is the Levinson-Durbin algorithm. This algorithm is recursive and when creating a prediction filter of A (from) order M as a "by-product" it will also give residual energies of lower order models. According to embodiments of the invention, this can be used. [0036] Fig. 2. shows an example of a general way of estimating background noise in a sound signal. This method can be performed by a background noise estimator. The method includes obtaining at least one parameter associated with an audio signal segment, such as a frame or frame portion, based on the gain of the first linear prediction, calculated as the quotient of the residual signal from zero-order linear prediction and the residual signal from linear prediction 2. row for the audio signal segment; and the gain of the second linear prediction calculated as the quotient of the residual signal from the 2nd order linear prediction and the residual signal from the 16th order linear prediction for the audio signal segment.
[0037] The method further includes determining 202 whether the audio signal segment includes a gap, i.e. it does not contain active content such as speech and music, based on at least the obtained parameter; and updating 203 estimating background noise based on the audio segment when the audio segment includes a gap. That is, the method includes updating the background noise estimate when a gap has been detected in the audio signal segment based on at least the obtained at least one parameter.
[0038] Linear prediction gains could be described as the first linear prediction gain associated with the transition
From zero order linear prediction to 2nd order linear prediction for a segment of an audio signal; and the gain of the second linear prediction associated with the transition from the 2nd order linear prediction to the 16th order linear prediction for the audio signal segment. Furthermore, obtaining this at least one parameter could alternatively be described as fixing, calculating, deriving or creating. Residual energies associated with linear predictions of model order 0, 2 and 16 can be obtained, obtained or taken from parts of the encoder in which linear prediction is performed as part of normal coding, i.e. these residual energies can somehow be provided by these parts. Thus, the computational complexity of the solution described herein can be reduced compared to when residual energies must be derived specifically for background noise estimation.
[0039] At least one parameter obtained based on the features of linear prediction can provide a level independent input signal analysis that improves the decision on whether or not to perform background noise update. This solution is particularly useful in the SNR range from 10 to 20 dB, where energy-based SADs have limited efficiency due to the normal dynamic range of speech signals.
[0040] Among others, the variables E (0), ... E (m), ... E (M) denote the residual energies for the order rows from 0 to M M + 1 of the Am (z) filters. Note that E (0) simply means the energy of the input signal. The analysis of the audio signal according to the solution described here provides several new features or parameters by analyzing the linear prediction gain calculated as the quotient of the residual signal from the zero-order linear prediction and the residual signal from the linear prediction 2. order and linear prediction gain calculated as the quotient of the residual signal from the 2nd order linear prediction and the residual signal from the 16th order linear prediction. That is, the gain of linear prediction for the transition from zero order linear prediction to 2nd order linear prediction means the same as "residual energy" E (0) (for the zero order model) divided by residual energy E (2) (for model 2. row). Corresponding linear prediction gain for transition from linear prediction 2. order for 16th order linear prediction means the same as residual energy E (2) (for the 2nd order model) divided by residual energy E (16) (for the 16th order model). Examples of parameters and setting parameters based on prediction gains will be described in more detail below. At least one parameter obtained according to the general embodiment described above may form part of the decision criterion used to assess whether to update the background noise estimate or not.
[0041] To improve the long-term stability of the at least one parameter or characteristic, a limited version of the prediction gain can be calculated. That is, obtaining at least one parameter may include limiting the gain of linear prediction associated with the transition from zero order linear prediction to 2nd order linear prediction and from 2nd order linear prediction to 16th order linear prediction, to take a predetermined value range. For example, linear prediction gains may be limited to taking values from 0 to 8, as shown e.g. in equals 1 and eq. 6 below.
[0042] Obtaining this at least one parameter may further comprise creating at least one long-term estimate of each of the gains of the first and second linear prediction, e.g. by low-pass filtering. Such at least one long-term estimate would then be further associated based on the respective linear prediction gains with at least one preceding segment of the audio signal. More than one long-term estimate could be created, where e.g. the first and second long-term estimates associated with linear prediction gain would react differently to changes in the audio signal. For example, the first long-term estimate may respond to changes faster than the second long-term estimate. Such a first long-term estimate may alternatively be referred to as a short-term estimate.
[0043] Obtaining this at least one parameter may further include determining a difference, such as the Gd_0_2 absolute equation (equation 3) described below, between one of the linear prediction gains associated with the audio signal segment and the long-term estimate of said gain linear prediction. Alternatively or additionally, a difference could be established between two long-term estimates, as in 9 below. The term "determination" could alternatively be converted to calculation, creation or derivation.
[0044] Obtaining this at least one parameter may, as indicated above, include low-pass filtering of linear prediction gains, and thus deriving long-term estimates, some of which can alternatively be referred to as short-term estimates, depending on how many segments are included in the estimation. The filter coefficients of at least one low-pass filter may depend on the relationship between the associated linear prediction gain, e.g. only, with the current segment of the audio signal and the average, defined as e.g. long-term average or long-term estimation, of the corresponding prediction gain obtained from the many preceding audio signal segments. This can be done to create, for example, estimates of long-term prediction gains. Low-pass filtering can be performed in two or more stages, where each stage can give a parameter, i.e. an estimate, used to make decisions regarding the occurrence of a break in the segment of the audio signal. For example, different long-term estimates (such as G1_0_2 (equiv. 2) and Gad_0_2 (equiv. 4), and / or G1_2_16 (equiv. 7), G2_2_16 (equiv. 8) and Gad_2_16 (equiv. 10) described below), reflecting changes in the audio signal in various ways, can be analyzed or compared to detect a break in the current segment of the audio signal.
[0045] Determining 202 whether the audio signal segment includes a gap or not may be further based on the spectral proximity measure associated with the audio signal segment. A measure of spectral proximity will indicate how close the "per band frequency" energy level is to the current background noise estimate, e.g. The initial value or estimate that is the result of a previous update performed prior to analyzing the current audio segment is the energy level "per frequency band" of the audio segment currently being processed. An example of determining or deriving a spectral proximity measure is given below in equiv. 12 and eq. 13. A measure of spectral proximity can be used to prevent noise updates based on low-energy frames with a large difference in frequency characteristics compared to the current background estimate. For example, the average energy across frequency bands could be equally small for the current signal segment and the current background noise estimate, but a measure of spectral proximity would reveal whether energy is distributed differently between the frequency bands. Such a difference in energy distribution could suggest that the current signal segment, e.g., frame, may be low activity content and updating the background noise estimation based on the frame, e.g., would prevent future frames with similar content from being detected. As the SNR subband is most sensitive to energy increases, the use of even low activity content can lead to a large update of the background estimate if this particular frequency range does not exist in the background noise, just like part of the high frequency speech compared to the car noise low frequency. After such an update, speech detection will be more difficult.
[0046] As already suggested above, a spectral proximity measure can be derived, derived or calculated based on energy for a set of frequency bands (alternatively referred to as subbands) of the currently analyzed audio signal segment and current background noise estimates corresponding to this set of frequency bands. This will also be exemplified and described in more detail
EP 3 309 784 B1 below; this is also shown in figure 5.
[0047] As indicated above, a spectral proximity measure can be derived, obtained or calculated by comparing the current energy level per frequency band of the currently processed audio signal segment with the energy level per frequency band of the current background noise estimation. However, at the beginning, i.e. during the first period or first number of frames at the beginning of the analysis of the audio signal, there may be no reliable estimate of the background noise, e.g. because the background noise estimate has not yet been reliably updated. Therefore, an initialization period may be used to determine the spectral proximity value. During such an initialization period, energy levels per frequency band of the current audio signal segment will instead be compared to an initial background estimate, which may be e.g. a configurable constant value. In the following examples below, the initial background noise estimate is set to the example value Emin = 0.0035. After the initialization period, the procedure may switch to normal operation and compare the current energy level per frequency band of the currently processed audio signal segment with the energy level per frequency band of the current background noise estimation. The length of the initialization period can be configured, e.g. based on simulations or tests indicating the time it will take before a reliable and / or satisfactory estimation of background noise is provided, for example. In the example below, a comparison with the initial background noise estimate (instead of the "real" estimate based on the current audio signal) is made during the first 150 frames.
[0048] The at least one parameter may be a parameter given, for example, in the code below, designated NEW_POS_BG, and / or one or more of the many parameters described below, leading to the creation of a decision criterion or a decision criterion component for detecting gaps. In other words, at least one parameter or feature obtained on the basis of linear prediction gain may be one or more of the parameters described below, and may include one or more of the parameters described below and / or may be based on one or more parameters described below. Features or parameters associated with residual energies E (0) and E (2) [0049] Fig. 3. shows an overview block diagram of deriving the features or parameters associated with E (0) and E (2) according to one embodiment. As can be seen in Figure 3, the prediction gain is first calculated as E (0) / E (2). The limited version of the profit prediction is calculated as
G_0_2 = max (0, min (8, E (0) / E (2))) (equals 1) where E (0) is the energy of the input signal and E (2) is the residual energy after the 2nd order linear prediction . Expression in equiv 1 limits the prediction gain to between 0 and 8. In normal cases, the prediction gain should be greater than zero, but anomalies may occur for values close to zero, and therefore a "greater than zero" limit (0 <) may be useful. The reason for limiting the prediction gain to a maximum of 8 is that for the purposes of the solution described here, it is enough to know that the prediction gain is about 8 or is greater than 8, which indicates a significant linear prediction gain. It should be noted that when there is no difference between residual energy for two different rows of models, the gain of linear prediction will be 1, which indicates that the higher order model filter is not much more successful in modeling the audio signal than the lower order model filter. In addition, if the G_0_2 prediction gain had too large values in subsequent expressions, this could threaten the stability of the derived parameters. Note that 8 is only an example value that has been chosen for a specific embodiment. The parameter G_0_2 could alternatively be marked e.g. epsP_0_2 or gLP_0_2 · [0050] The limited prediction gain is then filtered in two stages to create estimates
EP 3 309 784 B1 long-term profit. The first low-pass filtering, i.e. deriving the first long-term feature or parameter, is performed as:
G1_0_2 = 0.85 G1_0_2 + 0.15 G_0_2, (equiv. 2) [0051] Where the second "G1_0_2" in the expression should be read as the value from the previous segment of the audio signal. This parameter will usually be either 0 or 8, depending on the type of background noise in the input signal when there is an input segment containing only the background. The G1_0_2 parameter could alternatively be marked e.g. epsP_0_2_lp or gLP_0_2. You can then create or calculate another feature or parameter using the difference between the first long-term feature G1_0_2 and the limited prediction gain frame by frame G_0_2, according to:
Gd_0_2 = abs (G1_0_2-G_0_2) (equiv. 3) [0052] This will give an indication of the prediction gain of the current frame as a comparison with the long-term prediction gain estimate. The Gd_0_2 parameter could alternatively be marked e.g. epsP_0_2_ad or gad_0_2. In Fig. 3, this difference is used to create a second long-term estimate or feature of Gad_0_2. This is done by using a filter that uses different filter coefficients depending on whether the long-term difference is larger or smaller than the currently estimated average difference according to:
Gad02 = (1 -a) Gad_0_2 + a Gd_0_2 (equals 4) where if Gd_0_2 <Gad_0_2, then a = 0.1, otherwise a = 0.2 [0053] Where the second "Gad_0_2" in the expression should be read as value from the previous segment of the sound signal.
The Gad_0_2 parameter could alternatively be marked e.g. Glp_0_2, epsP_0_2_ad_lp or g<sub>ad 0 2</sub>. To prevent occasional large frame differences from being masked by filtering, you can derive another parameter that is not shown in the figure. That is, the second long-term feature of Gad_0_2 can be combined with a frame difference to prevent such masking. This parameter can be derived assuming the maximum version of the Gd_0_2 frame and the long-term version of Gad_0_2 of the prediction gain as:
Gmax_0_2 = max (Gad_0_2, Gd_0_2) (equ. 5) [0054] The Gmax_0_2 parameter could alternatively be marked, e.g. epsP_0_2_ad_lp_max or g<sub>m</sub>ax_o_2.
Features or parameters associated with residual energies E (2) and E (16) [0055] Fig. 4 is an overview block diagram of deriving features or parameters associated with E (2) and E (16) according to one embodiment. As can be seen in Figure 4, the prediction gain is first calculated as E (2) / E (16). Features or parameters created using the difference or relationship between 2nd order residual energy and residual energy 16. orders are derived in a slightly different way than described above in relation to the relationship between the zero order residual energy and the second order residual energy.
[0056] Here, too, the prediction gain is calculated as
G_2_16 = max (0, min (8, E (2) / E (16))) (equiv. 6) where E (2) means residual energy after the 2nd order linear prediction, and E (16) means residual energy after 16th order linear prediction. The G_2_16 parameter could alternatively be marked e.g. epsP_2_16 or gLP_2_16. This limited prediction gain is then used to create two long-term estimates of that gain: one in which the filter coefficient is different, depending on whether the long-term estimate is to be increased or not as shown in:
G1_2_16 = (1-a) G1_2_16 + a G_2_16 (equiv. 7) where if G_2_16> G1_2_16, then a = 0.2, otherwise a = 0.03
[0057] Parameter G1_2_16 could alternatively be marked eg epsP_2_16_lp or g<sub>LP 216</sub>.
[0058] The second long-term estimate uses a constant filter coefficient, as per:
G2_2_16 = (1-b) G2_2_16 + b G_2_16, where b = 0.02 (equiv. 8) [0059] Parameter G2_2_16 could alternatively be marked, e.g. epsP_2_16_lp2 or g<sub>LP2 0 2</sub>.
[0060] For most types of background signals, both G1_2_16 and G2_2_16 will be close to 0, but will have different responses to content where 16th order linear prediction is needed, which is typical for speech and other active content. The first long-term estimate, G1_2_16, will usually be larger than the second long-term estimate G2_2_16. This difference between long-term features is measured according to: Gd_2_16 = G1_2_16 - G2_2_16 (equiv 9) [0061] The Gd_2_16 parameter could alternatively be marked as epsP_2_16_dlp or gad_2_16.
[0062] Gd_2_16 can then be used as an input to the filter that creates the third long-term feature according to:
Gad_2_16 = (1-c) Gad_2_16 + c Gd_2_16 (equiv. 10) where if Gd_2_16 <Gad_2_16, then c = 0.02, otherwise c = 0.05 [0063] This filter uses different filter coefficients depending on this whether the third long-term signal is to be increased or not. Alternatively, the Gad_2_16 parameter can be marked e.g. epsP_2_16_dlp_lp2 or g<sub>ad 216</sub>. Also here, the long-term signal Gad_2_16 can be combined with the filter input of Gd_2_16 to prevent masking by filtering sporadic large input signals for the current frame. The final parameter is therefore the maximum of the frame or segment and the long-term version of the feature
Gmax_2_16 = max (Gad_2_16, Gd_2_16) (equals 11) [0064] The parameter Gmax_2_16 could alternatively be marked, e.g. epsP_2_16_dlp_max or g<sub>It has</sub>x_0_2.
Proximity / Spectral Difference Measure [0065] The spectral proximity measure uses the frequency analysis of the current input frame or input segment in which the subband energy is calculated and compared with the subband background estimate. The spectral proximity parameter or feature can be used in combination with the parameter described above associated with the linear prediction gain e.g. to ensure that the current segment or frame is relatively close or at least not far from the previous background estimate.
[0066] Fig. 5. shows a block diagram for calculating spectral proximity or difference measure. During the initialization period, e.g. the first 150 frames, a comparison is made with the constant corresponding to the initial background estimate. After initialization, it goes to normal operation and compares it with the background estimate. It should be noted that although spectral analysis produces subband energies for 20 subbands, the nonstaB calculation only uses subbands I = 2, ... 16, because speech energy is mainly located in these bands. nonstaB reflects non-stationarity here.
[0067] Thus, during initialization, nonstaB is calculated using Emin, here set to Emin = 0.0035, as:
nonstaB = sum (abs (log (Ecb (i) +1) -log (Emin + 1))) (equiv. 12) where summation is performed after i = 2 ... 16.
This is done to reduce the impact of decision errors in estimating background noise during initialization. After the initialization period, the calculation is performed using the current background noise estimate of the corresponding subband, according to:
nonstaB = sum (abs (log (Ecb (i) +1) -log (Ncb (i) +1))) (equiv. 13) where summation is performed after i = 2 ... 16
EP 3 309 784 Β1 [0068] The addition of a constant 1 to the energy of each subband before the logarithm reduces the sensitivity to spectral difference for low energy frames. The nonstaB parameter could alternatively be marked, e.g. non_staB or nonstate.
[0069] A block diagram showing an embodiment of the background noise estimator is shown in Fig. 6. The embodiment in Fig. 6 includes an input signal framing block 601 in which the input audio signal is divided into frames or segments of appropriate length, e.g. 5-30 ms. The embodiment further includes a feature extraction block 602 in which features, also referred to as parameters, are calculated for each frame or segment of the input signal. The embodiment further includes an update decision logic block 603 to determine whether or not the background estimation can be updated based on the signal in the current frame, i.e. whether the signal segment does not contain active content such as speech and music. The embodiment further includes a background updater 604 for updating the background noise estimate when the logic of the update decision indicates that its execution is correct. In the embodiment shown, the background noise estimate can be output to a subband, i.e. for a number of frequency bands.
[0070] The solution described herein can be used to improve the prior background noise estimation solution described in Annex A as well as in WO2011 / 049514. The solution described here will be described below in the context of this solution described earlier. Code examples from the code of the embodiment of the background noise estimator will be given. The following describes the details of the actual embodiment for an embodiment of the invention in a G.718 based encoder. This embodiment uses many of the energy features described in the solution in Annex A and WO2011 / 049514. For further details than below, please refer to Annex A and WO2011 / 049514.
[0071] The following energy features are defined in WO2011 / 049514:
Etat;
Etotllp;
Etot_v_h;
totalNoise;
s ign_dyn_lp;
[0072] The following correlation features are defined in WO2011 / 049514:
AEN;
ha.rrri_cor_cnt
a.ct_pred cor_est [0073] The following features are defined in the solution given in Annex A:
Etot_v_h;
lt_cor_est = 0. Olf * cor_est -I- 0.9 9f * lt_cor_est;
lt_tn_track = 0.03f * (Etot - totalNoise <10) -I- 0.97f * lt_tn_track; lt_tn_dist = 0.03f * (Etot - totalNoise) -l- 0.97f * lt_tn_dist;
lt_Ellp_dist = 0.03f * (Etot-Etot_l_lp) -I- 0.97f * lt_Ellp_dist; harm_cor_cnt low_tn_track_cnt [0074] The noise updating logic of the solution given in Annex A is shown in Fig. 7. The noise estimator improvements in Annex A related to the solution described herein are mainly associated with part 701 in which the features are calculated; part 702 where decisions are made on a break based on
Various parameters; and further with section 703 in which, based on whether or not a break has been detected, various actions are taken. In addition, improvements may have an effect on updating the background noise estimation 704, which may be updated e.g. when a gap has been detected based on new features that would not have been detected before the implementation of the solution described herein. In the exemplary embodiment described herein, the new features introduced herein are calculated as follows, starting from non_staB, which is determined using the energy of the subbands of the current frame enr [i], which corresponds to Ecb (i) above and in Fig. 6, and the current noise estimation backgrounds bckr [i], which corresponds to Ncb (i) above and in Figure 6. The first part of the first section of the code below is associated with a special initial procedure for the first 150 frames of the audio signal, before deriving the proper background estimation.
Γ calculate non-stationarity feature relative background (spectral closeness feature non_staB 7 if (ini_frame <150) {
/ * During init don't include updates 7 if (i> = 2 && i <= 16) {
non_staB + = (float) fabs (log (en r [i] + 1 .Of) log (E_MIN + 1.0f));
} }
else {
Γ After init compare with background estimate 7 if (i> = 2 && i <= 16) non_staB + = (float) fabs (log (en r [i] + 1 .Of) log (bckr [i] +1 .Of )):
} }
if (norr_staB> = 128) {
non_staB = 32767.0 / 256.0f;
} [0075] The code sections below show how new features are calculated for the residual linear prediction energy, i.e. for linear prediction gain. Residual energies are here called epsP [m] (see E (m) used earlier).
EP 3 309 784 B1 ____ _ _ ____ _ _ _ * * Linear prediction efficiency 0 to 2 order * (linear prediction gain going from 0<sup>lh</sup> it's 2<sup>nd</sup> order model of linear prediction filter) _ __ __________ _ - * / epsP_0_2 = max (0, min (8, epsP [0] / epsP [2]));
epsP_0_2_lp = 0.15f * epsP_0_2 + (1.0f-0.15f) * st-> epsP_0_2_lp;
epsP_0_2_ad = (float) fabs (epsP_0_2 - epsP_0_2_lp);
ii (epsP_0_2_ad <epsP_0_2_ad_l p) {
epsP_0_2_ad_lp = 0.1f * epsP_0_2_ad + (1.Of - 0.1f) * epsP0 2ad_lp;
} else {
epsP_0_2_ad_lp = 0.2f * epsP_0_2_ad + (1.Of - 0.2f) * epsP_0_2_ad_lp;
} epsP_0_2_ad_lp_max = max (epsP_0_2_ad, st-> epsP_0_2_ad_lp);
y * - - - - - - - - - - - - _ _ _ * * Linear predition efficiency 2 to 16 order * (linear prediction gain going from 2 ^ to 16th order model of linear prediction filter) * __________________________________________________ * 1 epsP_2_16 = max (0, min (8, epsP [2] / epsP [16]));
if (epsP_2_16> epsP_2_16_lp) {
epsP_2_16_lp = 0.2f 'epsP_2_16 + (1.0f-0.2f)' epsP_2_16_lp;
} else {
epsP_2_16_lp = 0.03f * epsP_2_16 + (1.0f-0.03f) * epsP_2_16_lp;
} epsP_2_16_l p2 = 0.02f * epsP_2_16 + (1.0f-0.02f) * epsP_2_16_lp2;
epsP_2_16_dlp = epsP_2_16_lp-epsP_2_16_lp2;
if (epsP_2_16_dl p <epsP_2_16_dlp_lp2) {
epsP_2_16_dl p_l p2 = 0.02f * epsP_2_16_dl p + (1.0f-0.02f) * epsP_2_16_dlp_lp2;
} else {
epsP_2_16_dl pj p2 = 0.05f * epsP_2_16_dl p + (1.0f-0.05f) * epsP_2_16_dlp_lp2;
} epsP_2_16_dlp_max = max (epsP_2_16_dlp, epsP_2_16_dlp_lp2);
[0076] The code below illustrates the creation of combined metrics, thresholds and flags used for the actual update decision, i.e. determining whether to update the background noise estimate or not. At least some of the parameters associated with linear prediction gains and / or spectral proximity are shown in bold.
Comb_ahc_epsP = max (max (act_pred, lt_haco_ev), epsP_2_16_dlp); comb_hcm_epsP = max (max (lt_haco_ev, epsP_2_16_dlp_max), epsP_0_2_ad_lp_max);
haco_ev_max = max (st_harm_cor_cnt == 0,> lt_haco_ev);
EtotJJp_thr = st-> Etot_IJp + (1.51 + 1.5f * (EtotJp <50.01)) * Etot_v_h2;
enr_bgd = Etot <EtotJ_lp_thr;
cns_bgd = (epsP_0_2> 7.95f) && (non_sta <1e3f);
lp_bgd = epsP_2_16dlp_max <0.10f;
nsjnask = non_sta <1e5f;
lt_haco_mask = lt_haco_ev <0.5f;
bg_haco_mask = haco_ev_max <0.4f;
SD_1 = ((epsP_0_2_ad> 0.5f) && (epsP_0_2> 7.95f));
bg_bgd3 = enr_bgd || ((cns_bgd || lp_bgd) && ns_mask && lt_haco_mask && SD_1 == 0);
PD_1 = (ep $ P_2_16_dlp_max <0.1 Of);
PD_2 = (ep $ P_0_2_ad_lp_max <0.10f);
PD_3 = (comb_ahc_epsP <0.851);
PD_4 = cotnb.ahc ^ epsP <0.15f;
PD_5 = comb_hcm_epsP <0.30f;
BG_1 = ((SD_1 == 0) || (Etot <Etot_l_lp_thr)) && bg_haco_mask && (act_pred <0.85Ϊ) && (Etotjp <50. Of);
PAU = (aEn == 0) || ((Etot <55.0f) && (SD_1 == 0) && ((PD_3 && (PD_11 | PD_2)) || (PD_41 | PD_5)));
NEW_POS_BG = (PAU | BG_1) & bg_bgd3;
Γ Original silence detector works in most cases 7 aE_bgd = aEn == 0;
Γ When the signal dynamics is high and the energy is close to the background estimate 7 sd1 _bgd = ($ t-> sign_dyn_lp> 15) && (Etot - st-> Etot_l_lp) <2 * st-> Etotvh2 && st-> harm_cor_cnt > 20;
i * init conditions steadily dropping acLpred and / or lt_haco_ev * 1 tnjni = inrframe <150 && harm_cor_cnt> 5 &&
((st-> act_pred <0.59f && st-> lt_haco_ev <0.23f) || st-> act_pred <0.38f || st-> lt_haco_ev <0.15f | (non_staB <50.011 | aE.bgd);
Γ Energy close to the background estimate serves as a mask for other background detectors 7 bg_bgd2 = Etot <EtotJJp_thr || tn_ini;
[0077] Since it is important not to update the background noise estimate when the current frame or segment contains active content, several conditions are evaluated to decide if the update should be performed. The main decision-making stage in the noise updating logic is whether or not to update it, and this decision is made by evaluating the logical expression, which is highlighted below. The new parameter NEW_POS_BG (new in relation to the solution in Annex A and WO2011 / 049514) is the gap detector, obtained on the basis of linear prediction gain at the transition from zero to 2nd and 2nd and 16th order of the linear prediction filter model, and tn_ini is obtained at based on a characteristic associated with spectral proximity. Here comes the decision logic using the new features according to
EP 3 309 784 B1 embodiment.
updt_step = O.Of;
if ((bg bgd2 && (aE bgd II sd 1 bod II It tn track> 0.90f II NEW POS BG)) II tn ini) {
if (((act_pred <0.85f) && aE_bgd &&
(lt_Ellp_dist <10 || sd 1_bgd) && lt_tn_dist <40 &&
((Etot-totalNoise) <10.0f)) || (st-> first_noise_updt == 0 && st-> harm_cor_cnt> 80 && aE_bgd && st-> lt_aEn_zero> 0.5f) || (tn_ini && (aE_bgd || non_staB <10.01 | st-> harm_cor_cnt> 80)) {
updt_step = 1,0f;
st-> first_noise_updt = 1;
for (i = 0; i <NB_BANDS; i ++) {
st-> bckr [i] = tmpN [i];
} }
eise ii (((st-> act_pred <0.801) && (aE_bgd || PAU) && st-> lt_haco_ev <0.10f) || ((st-> act_pred <0.701) && (aE_bgd || non_staB <17.0f) && PAU && st-> lt_haco_ev <0.15f) || (st-> harm_cor_cnt> 80 && st-> totalNoise> 5.01 && Etot <max (1.0f, EtotJJp <sup>+ 1</sup> -5f * st-> Etot_v_h2)) || (st-> harm_cor_cnt> 50 && st-> first_noise_updt> 30 && aE_bgd && st-> lt_aEn_zero> 0.5t) || tn_ini)
{updt_step = 0.1f;
if (! aE_bgd &&
st-> harm_cor_cnt <50 &&
(st-> act_pred> 0.6f || (! tn_ini && EtotJJp - st-> totalNoise <10.0f && non_staB> 8.Of))) {
updt_step = 0.01f;
} if (updt_step> O.Of) {
st-> first noise_updt = 1;
for (i = O; i <NB_BANDS; i ++) {
st-> bckr [i] = st-> bckr [i] + updt_step * (tmpN [i] -st-> bckr [i]);
} }
} eise if (aE_bgd || st-> harm_cor_cnt> 100) {
(st-> first_noise_updt) + = 1;
} }
eise {
/ * If in musie lower bekr to drop further 7 if (st-> low tn track ent> 300 && st-> lt_haco_ev> 0.9f && st-> totalNoise> O.Of) {
UPdt. step = -0.02f;
for (i = 0; i <NB_BANDS; i ++) {
if (st-> bckr [i]> 2 * E_MIN) {
st-> bckr [i] = 0.98f * st-> bckr [i];
} }
} )
st-> lt_aEn_zero = 0.2f * (st-> aEn == 0) + (1-0.2f) * st-> lt_aEn_zero;
[0078] As noted earlier, linear prediction features provide level-independent input signal analysis that improves the decision to update background noise, which is particularly useful in the SNR range from 10 to 20 dB, where SAD based on energy has limited efficiency due to the normal dynamic range of speech signals. [0079] Background proximity features also improve the estimation of background noise as they can be used in both initialization and normal operation. During initialization, they can enable fast initialization for (lower level) background noise with mainly low frequency content common to car noise. These features can also be used to prevent noise updates using low-energy frames with a large difference in frequency characteristics compared to the current background estimate, suggesting that the current frame may have low active content and the update could prevent future frames with similar content from being detected.
[0080] Figures 8.-10. show how individual parameters or metrics behave in the case of speech against the background of car noise with a 10 dB SNR. In Figures 8.-10. the dots represent the energy of each frame. In Figs. 8 and 9a-c, energy has been divided by 10 to make it easier to compare for features based on G_0_2 and G_2_16. The graphs correspond to the sound signal containing two statements, where the approximate positions of the first statement are frames 1310-1420, and the second one - frames 1500-1610.
[0081] Fig. 8. shows the frame energy (/ 10) (dot "·") and the features G_0_2 (circle "o") and Gmax_0_2 (plus "+") for car noise speech with a 10dB SNR. It should be noted that G_0_2 is 8 during car noise, because there is some correlation in the signal that can be modeled using linear prediction with the order of model 2. During the speech, the Gmax_0_2 feature reaches over 1.5 (in this case), and after a speech spike it decreases to 0. In a specific implementation of decision logic, Gmax_0_2 must be less than 0.1 to allow noise updates using this feature.
[0082] Fig. 9a shows the frame energy (/ 10) (dot "·") and features G_2_16 (circle "o"), G1_2_16 (cross "x"), G2_2_16 (plus "+"). Fig. 9b shows the frame energy (/ 10) (dot "·") and features G_2_16 (circle "o"), Gd_2_16 (cross "x"), and Gad_2_16 (plus "+"). Fig. 9c shows the energy of the frame (/ 10) (dot "·") and features G_2_16 (circle "o") and Gmax_2_16 (plus "+"). The graphs shown in Figs. 9a-c also relate to speech with car noise of 10dB SNR. Features are presented in three graphs to facilitate the observation of individual parameters. It should be noted that G_2_16 (circle "o") is slightly more than 1 during car noise (ie outside of statements), which indicates that the gain from the higher order model is small for this type of noise. During the statement, the Gmax_2_16 (plus "+" in Fig. 9c) increases and then decreases back to 0. In a specific implementation of decision logic, the Gmax_2_16 feature must also be less than 0.1 to allow noise updates. This is not the case in this particular sample of the sound signal.
[0083] Fig. 10 shows the frame energy (dot "·") (this time without dividing by 10) and the nonstaB (plus "+") feature for car noise speech with a 10dB SNR. The nonstaB feature is in the range of 0-10 during segments containing only noise, and for the speech it becomes much larger (because the frequency characteristics are different for speech). However, it should be noted that even during speech there are frames in which the nonstaB feature falls to the range of 0-10. For these frames, you may be able to update the background noise, and thus better track the background noise.
[0084] The solution disclosed herein also relates to a background noise estimator made in hardware and / or in software. Background noise estimator, Figures 11a-11c
[0085] One embodiment of the background noise estimator is generally shown in Fig. 11a. A module or object adapted to estimate background noise in audio signals containing e.g. speech and / or music is defined as the background noise estimator. Encoder 1100 is adapted to perform at least one method corresponding to the methods described above with reference to e.g. Figures 2 and 7. Encoder 1100 is associated with the same technical features, objects and advantages as the previously described embodiments of the method. The background noise estimator will be described briefly to avoid unnecessary repetition.
[0086] The background noise estimator can be made and / or described as follows:
The background noise estimator 1100 is adapted to estimate the background noise of the audio signal. The background noise estimator 1100 includes processing circuits, i.e. processing means 1101, and communication interface 1102. The processing circuits 1101 are adapted to cause the encoder 1100 to obtain, e.g. determine or calculate, at least one parameter, e.g. NEW_POS_BG, based on the gain of the first linear prediction calculated as the quotient of the residual signal from the zero-order linear prediction and the residual signal from the second-order linear prediction for the audio signal segment; and the gain of the second linear prediction calculated as the quotient of the residual signal from the 2nd order linear prediction and the residual signal from the 16th order linear prediction for the audio signal segment.
[0087] The processing circuits 1101 are further adapted to cause the background noise estimator to determine whether the audio signal segment includes a gap, i.e. it does not contain active content such as speech and music, based on the at least one parameter. Processing circuits 1101 are further adapted to cause the background noise estimator to update the background noise estimation based on the audio signal segment when the audio signal segment includes a gap.
[0088] Communication interface 1102, which may also be designated as e.g. an input / output (I / O or I / O) interface, includes an interface for sending data to and receiving data from other objects or modules. For example, you can get residual signals associated with the ninth, second and 16th order of the linear prediction model, e.g., receive through the I / O interface from an audio signal encoder performing linear predictive coding.
[0089] Processing circuits 1101 could, as shown in Fig. 11b, include processing means such as processor 1103, e.g., CPU, and memory 1104 for storing or holding instructions. The memory would then contain instructions, e.g., in the form of a computer program 1105, which when processing means 1103 would cause the encoder 1100 to perform the operations described above.
[0090] An alternative embodiment of the processing circuits 1101 is shown in Fig. 11c. The processing circuits comprise here an acquiring or retaining assembly or module 1106 adapted to cause the background noise estimator 1100 to obtain, e.g., determine or calculate, at least one parameter, e.g. NEW_P0S_BG, based on the gain of the first linear prediction calculated as the quotient of the residual signal from zero-order linear prediction and the residual signal from linear prediction 2. row for the audio signal segment; and the gain of the second linear prediction calculated as the quotient of the residual signal from the 2nd order linear prediction and the residual signal from the 16th order linear prediction for the audio signal segment. The processing circuits further include a retention assembly or module 1107 adapted to cause the background noise estimator 1100 to determine whether the audio signal segment has a gap, i.e. does not contain active content, such as speech and music, based on at least one parameter. The processing circuits 1101 further include an updating or estimating assembly or module 1110 adapted to cause the background noise estimator to update the background noise estimation based on the audio segment when the audio segment includes a gap.
[0091] Processing circuits 1101 could contain more assemblies, such as a filter assembly or module adapted to cause the background noise estimator to low pass filter the linear prediction gain, thereby creating one or more estimates of long-term linear prediction gain operations, just like low-pass filtering can otherwise be done e.g. by the retainer assembly or module 1107.
[0092] The embodiments of the background noise estimator described above could be adapted to the various embodiments of the method described herein, such as limiting and lowpass filtering of linear prediction gain; determining the difference between linear prediction gains and long-term estimates, and between long-term estimates; and / or obtaining and applying a spectral proximity measure, etc.
[0093] It can be assumed that the background noise estimator 1100 includes a further functionality for performing the background noise estimation, such as e.g. the function given in Annex A.
[0094] Fig. 12 shows a background noise estimator 1200 according to an embodiment. The background noise estimator 1200 includes an input assembly e.g. for receiving residual energy for the zero, 2nd and 16th order of the model. The background noise estimator further includes a processor and memory that contains commands executable by that processor, whereby the background noise estimator is capable of performing the method according to one embodiment described herein.
[0095] Accordingly, the background noise estimator may include, as shown in Fig. 13, an I / O assembly 1301, a calculator 1302 for calculating the first two sets of characteristics from residual energy for the zero, 2nd and 16th order models and a frequency analyzer 1303 for calculating the spectral proximity feature.
[0096] The background noise estimator as described above may be included e.g. in a VAD or SAD, encoder and / or a decoder, i.e. a codec, and / or in a device such as a communication device. The communication device may be a user equipment (UE) in the form of a mobile phone, video camera, sound recorder, tablet, desktop, laptop, television receiver or home server / home network gateway / home access point / home router. The communication device in some embodiments may be a telecommunications network device adapted to encode and / or transcode audio signals. Examples of such telecommunications network devices are servers such as media servers, application servers, routers, gateways and radio base stations. The communication device can also be adapted to be placed, i.e. being built into a vehicle such as a ship, drone, plane and road vehicle such as a passenger car, bus or truck. Such a built-in device would normally belong to the vehicle telematics unit or vehicle information and entertainment system.
[0097] The steps, functions, procedures, modules, assemblies and / or blocks described herein can be implemented in hardware using any conventional technology, such as isolated circuit technology or integrated circuit technology, including both universal and specialized electronic circuits.
[0098] Specific examples include one or more suitably adapted digital signal processors and other known electronic circuits, e.g., isolated logic gates connected to one another to perform a specialized function, or Application Specific Integrated Circuits (ASICs).
[0099] Alternatively, at least part of the steps, functions, procedures, modules, assemblies and / or blocks described above may be implemented in software, such as in a computer program intended for execution by appropriate processing systems comprising one or more processors. Before using and / or when using computer programs on network nodes, the software could
It may be carried on a carrier such as an electronic signal, an optical signal, a radio signal or a computer readable medium.
[0100] The activity networks depicted herein can be considered as computer activity networks when executed by one or more processors. The appropriate apparatus can be defined as a group of functional modules, where each stage performed by the processor corresponds to a functional module. In this case, function modules are implemented as a computer program running on the processor.
[0101] Examples of processing circuits include, but are not limited to, one or more microprocessors, one or more Digital Signal Processors (DSP), one or more Central Processing Units (CPUs) and / or any suitable programmable logic circuit such as one or more directly programmable gate arrays ( Field Programmable Gate Array, FPGA), or one or more Programmable Logic Controller (PLC). That is, the above-described assemblies or modules in circuits at various nodes could be implemented by a combination of analog and digital circuits, and / or one or more processors configured using software and / or firmware, e.g., stored in memory. One or more of these processors, as well as other digital equipment, may be contained in one specialized integrated circuit (ASIC), or several processors and different digital equipment may be distributed among several separate components, whether housed in separate housings or connected into a one-system system (called system-on-a-chip, SoC).
[0102] It should also be understood that it may be possible to reuse the general processing capabilities of any conventional device or assembly in which the proposed technology is implemented. It may also be possible to reuse existing software, e.g. by reprogramming existing software or adding new software components.
[0103] The above described embodiments are given only as examples and it should be understood that the proposed technology is not limited to them. Those skilled in the art will understand that various modifications, combinations and changes can be made to the embodiments without departing from the present scope. In particular, different parts of the solutions in individual embodiments can be combined in other configurations when technically possible.
[0104] The use of the word "comprises / includes" or "containing / including" should be interpreted as non-limiting, i.e. meaning "consists of at least".
[0105] It should also be noted that in some alternative embodiments, the functions / actions noted in the blocks may occur outside the order noted in the flow networks. For example, the two blocks shown as sequential may actually be executed substantially simultaneously or may sometimes be executed in reverse order, depending on the respective functions / activities. Furthermore, the functionality of a given block of block diagrams and / or flow networks may be divided into a plurality of blocks and / or the functionality of two or more blocks of block diagrams and / or flow networks may be at least partially integrated. Finally, you can add / insert other blocks between the presented blocks and / or blocks / actions can be omitted without departing from the scope of creative ideas.
[0106] It should be understood that the selection of interacting assemblies, as well as the names of these assemblies in the present disclosure are exemplary only, and the nodes capable of performing any of the methods described above can be configured in a variety of alternative ways to allow the suggested procedures to be performed.
[0107] It should also be noted that the assemblies described in the present disclosure are to be considered logical objects and not necessarily separate physical objects.
[0108] Reference to a given element in the singular is not intended as meaning "one and only one" unless it is indicated so clearly, but rather as "one or more." Furthermore, it is not necessary for the device or method to deal with any problem that is sought by the technology disclosed herein, for that device or method to be covered.
[0109] In some cases, detailed descriptions of well-known devices, circuits and methods are not omitted so as not to obscure the description of the disclosed technology with unnecessary details. All of the statements given herein that set out the principles, aspects and embodiments of the disclosed technology, as well as specific examples thereof, are intended to include both structural and functional equivalents. In addition, it was intended that such counterparts would include both known counterparts and counterparts developed in the future, e.g., any developed components performing the same function, regardless of construction.
Appendix [0110] Provided is a method of estimating background noise in an audio signal for a background noise estimator, in which method the audio signal includes multiple segments of the audio signal, and the method includes:
- obtaining (201) at least one parameter associated with one segment of the sound signal, based on:
- the gain of the first linear prediction calculated as the quotient of the residual signal (E (0)) from the zero-order linear prediction and the residual signal (E (2)) from the second-order linear prediction for the audio signal segment; and
- second linear prediction gain calculated as the quotient of the residual signal (E (2)) from the 2nd order linear prediction and the residual signal (E (16)) from the 16th order linear prediction for the audio signal segment;
- determining (202) whether the audio signal segment includes a gap, i.e. does not contain active content such as speech and music, based on at least the obtained parameter; and:
when the audio segment contains a gap:
- updating (203) an estimate of the background noise based on the segment of the audio signal.
[0111] Obtaining this at least one parameter may include limiting the gains of the first and second linear prediction to take values from the predetermined range.
[0112] Obtaining this at least one parameter may include creating at least one estimation of the long-term first and second linear prediction gain, e.g. by low-pass filtering, wherein the long-term estimation is further based on the respective linear prediction gains associated with at least one preceding sound signal segment.
[0113] Obtaining this at least one parameter may include determining a difference between one of the linear prediction gains associated with the audio signal segment and the long-term estimate of said linear prediction gain and / or the difference between two different long-term estimates associated with the linear prediction gain.
[0114] Getting this at least one parameter may include low pass filtering of the first gain and the second linear prediction gain.
[0115] The filter coefficients of the at least one low-pass filter may depend on the relationship between the linear prediction gain associated with the audio signal segment and the average of the corresponding gain
A linear prediction obtained based on a plurality of preceding audio signal segments.
[0116] Determining whether the audio signal segment includes a gap may be further based on a spectral proximity measure associated with the audio signal segment.
[0117] The method may further comprise obtaining a spectral proximity measure based on energy for a set of frequency bands of the audio signal segment and background noise estimates corresponding to this set of frequency bands. During the initialization period, the initial value, Emin, can be used as estimates of background noise from which a measure of spectral proximity is obtained.
[0118] Further provided is a background noise estimator (1100) for estimating background noise in an audio signal comprising a plurality of segments of the audio signal, which background noise estimator is adapted to:
- obtaining at least one parameter based on:
- the gain of the first linear prediction calculated as the quotient of the residual signal from the zero-order linear prediction and the residual signal from the second-order linear prediction for the audio signal segment;
- second linear prediction gain calculated as the quotient of the residual signal from the 2nd order linear prediction and the residual signal from the 16th order linear prediction for the audio signal segment;
- determining whether the audio segment contains a break, i.e. it does not contain active content such as speech and music, based on at least one parameter; and when the audio segment contains a gap:
- background noise estimation update based on the audio signal segment.
[0119] The background noise estimator according to claim The method of claim 10, wherein obtaining at least one parameter includes limiting the gains of the first and second linear predictions to take values from the predetermined range.
[0120] Obtaining this at least one parameter in the background noise estimator may include: creating at least one estimation of the long-term gain of the first and the second linear prediction gain, e.g. by low-pass filtering, in which the long-term estimation is further based on the respective associated linear prediction gains with at least one preceding audio signal segment.
[0121] Obtaining this at least one parameter in the background noise estimator may include determining the difference between one of the linear prediction gains associated with the audio signal segment and the long-term estimate of said linear prediction gain and / or the difference between two different long-term estimates associated with the linear prediction gain.
[0122] Obtaining this at least one parameter in the background noise estimator may include low pass filtering of the first gain and the second linear prediction gain.
[0123] The filter coefficients of the at least one lowpass filter in the background noise estimator may depend on the relationship between the linear prediction gain associated with the audio signal segment and the average of the corresponding linear prediction gain obtained from the many preceding audio signal segments.
[0124] The background noise estimator may be further adapted to base whether the audio signal segment includes a gap, based on the spectral proximity measure associated with the audio signal segment.
[0125] The background noise estimator may be adapted to obtain a spectral proximity measure based on energy for a set of frequency bands of the audio signal segment and background noise estimates
EP 3 309 784 B1 corresponding to this set of frequency bands.
[0126] The background noise estimator may be adapted to be used during the initialization period initial value, Emin, as background noise estimates from which a spectral proximity measure is obtained.
[0127] A sound activity detector, SAD, further comprising a background noise estimator as described above is provided.
[0128] A codec further comprising a background noise estimator as described above is provided.
[0129] A wireless device further comprising a background noise estimator as described above is provided. [0130] A network node further comprising a background noise estimator as described above is provided.
[0131] Further provided is a computer program containing instructions that, when executed by at least one processor, cause the at least one processor to perform the method as described above. Also provided is a medium containing this computer program, wherein the medium is one of an electronic signal, an optical signal, a radio signal or a computer readable memory medium.
ANNEX A [0132] References to the figures in the text below are references to Figs. A2-A9, also "Fig. 2 "below corresponds to Fig. A2 in the drawing.
[0133] Fig. 2. is a flowchart showing an embodiment of a method for estimating background noise according to the technology proposed herein. It was intended that the method be performed by a background noise estimator, which may be part of SAD. The background noise estimator and SAD may be further included in an audio encoder, which in turn may be contained in a wireless device or network node. In the case of the background noise estimator described, the downward noise estimation adjustment is not limited. A possible new estimate of the subband noise is calculated for each frame, regardless of whether the frame is background or active content; if the new value is less than the current value, it is used directly because it will most likely come from the background frame. The subsequent noise estimation logic is the second stage in which a decision is made as to whether the estimation of subband noise can be increased, and if so, how much. The increase is based on a previously calculated possible new subband noise estimate. Basically, this logic makes a decision about whether the current frame is a background frame, and if it is not certain, it may allow for a smaller increase compared to what was originally estimated.
[0134] The method shown in Fig. 2 includes: when the energy level of the audio segment is greater than by a threshold value of 202: 1 than the long-term minimum energy level, lt_min, or when the energy level of the audio segment is less than by the threshold value 202: 2 from lt_min, but no 204: 1 gap was detected in the audio segment:
- a reduction 206 of the current background noise estimate, when it is determined 203: 2 that the audio signal segment contains music, and the current background noise estimate exceeds the minimum value 205: 1, designated "T" in Fig. 2, and further exemplified by e.g. as 2 * E_MIN in the code below.
[0135] By performing the above and providing a background noise estimate for SAD, SAD obtains the ability to perform a more appropriate detection of audio activity. In addition, recovery from erroneous updates of background noise estimation is enabled.
[0136] The energy level of the audio signal segment used in the method described above can alternatively be determined as e.g. current frame energy, Etot or signal segment energy, or a frame which can be calculated by adding the subband energy for the current signal segment.
[0137] Another energy characteristic used in the above method, i.e. the long-term minimum energy level, lt_min, is an estimate established on a plurality of preceding audio signal segments or frames. It_min could alternatively be marked e.g. with Etot_l_lp. One basic way to derive lt_min would be to use a minimum value from the energy history of the current frame from a number of past frames. If the value calculated as: "current frame energy - long-term minimum estimation" is less than the threshold value, e.g. THR1, it is stated here that the energy of the current frame is close to long-term minimum energy or is close to long-term minimal energy. That is, when (Etot - lt_min) <THR1, it can be determined 202 that the energy of the current frame, Etot, is close to the long-term minimum energy lt_min. In the event that (Etot - lt_min) = THR1, you can refer to one of two decisions, 202: 1 or 202: 2, depending on the implementation. Reference 202: 1 in Fig. 2 indicates the decision that the current frame energy is not close to lt_min, while 202: 2 indicates the decision that the current frame energy is close to lt_min. Other designations in the form XXX: Y in Fig. 2 indicate the respective decisions. The lt_min feature will be described below.
[0138] It can be assumed that the minimum value that the current background noise estimate is to exceed to be reduced is zero or a small positive number. For example, as will be shown, for example, in the code below, it may be required that the current total energy of the background estimate, which can be labeled "totalNoise" and set to e.g. 10-Iog10Zbackr [i], exceeds the zero minimum value so that its reduction is taken into consideration. Alternatively or additionally, each item in the backr [i] vector including subband background estimates can be compared with the minimum value, E_MIN, for the reduction to be made. In the code example below, E_MIN is a small positive number.
[0139] It should be noted that according to the preferred embodiment of the solution suggested herein, the decision whether the energy level of the audio signal segment is greater than lt_min more than a threshold value is based solely on the information derived from the input audio signal, i.e. on feedback from the audio activity detector's decision.
[0140] Determining whether the current frame includes a gap or not can be done in various ways based on one or more criteria. The gap criterion can also be defined as a gap detector. A single gap detector or a combination of different gap detectors can be used. In the case of a combination of break detectors, each can be used to detect breaks in various conditions. One indicator that the current frame may contain a gap or inactivity is that the correlation trait for this frame is small and that a number of preceding frames also had low correlation features. If the current energy is close to the long-term minimum energy and an interval is detected, the background noise can be updated according to the current input signal, as shown in Fig. 2. An interval can be considered detected, in addition to the fact that the energy level of the audio signal segment is greater than lt_min less than the threshold value, it was determined that the predetermined number of successive preceding audio signal segments does not contain the active signal and / or the dynamics of the audio signal exceeds the threshold. This is also shown below in the code example below.
[0141] Reducing the background noise estimation 206 allows dealing with situations where the background noise estimation has become "too high", i.e. relative to the actual background noise. This can also be expressed, for example, as the background noise estimate differs from the actual background noise. Too much background noise estimation can lead to SAD making inappropriate decisions in which it was determined that the current segment of the signal is inactive even though it contained active speech or music. The reason that the background noise estimate becomes too large is because of an incorrect or unnecessary update of the background noise in music,
Where the noise estimation confused the music with the background and allowed to increase the noise estimation. The disclosed method allows correcting such an erroneously updated background noise estimate, e.g., when it is determined that the next input signal frame contains music. This correction is made by forcibly reducing the background noise estimate in which the noise estimate is scaled down even if the energy of the current input signal segment is greater than the current background noise estimate, e.g. in the subband. It should be noted that the background noise estimation logic described above is used to control the increase of background subband energy. Subband energy reduction is always allowed when the subband energy of the current frame is less than the background noise estimate. This function is not shown directly in Fig. 2. Such a reduction usually has a fixed setting for the step size. However, increasing the background noise estimate should only be allowed in conjunction with the decision logic of the method described above. When an interval is detected, the correlation and energy features can also be used to decide 207 how large the correction step size is to increase the background estimate before performing the actual background noise update.
[0142] As mentioned earlier, distinguishing some segments of music from background noise can be difficult because they are very similar to noise. Therefore, the noise update logic may accidentally allow increased subband energy estimates, even when the input signal was an active signal. This can cause problems because the noise estimate may become larger than it should be.
[0143] In the background noise estimators of the prior art, the subband energy estimates could be reduced only when the input subband energy fell below the current noise estimation. However, since distinguishing some segments of music from background noise can be difficult, because they are very similar to noise, the developers realized that a recovery strategy for music is needed. In the embodiments described herein, such recovery can be performed by forcibly reducing the noise estimate when the input signal returns to a music-like characteristic. This means that the energy described above and the 202: 1,204: 1 gap logic prevent the increase in noise estimation, test 203 is performed, whether the input signal is suspected to be music, and if so 203: 2, the subband energies are reduced 206 by a small the amount in each frame until the noise estimates reach the lowest level 205: 2.
[0144] The background noise estimator, as described above, may be included or implemented in a VAD or SAD, and / or in an encoder and / or decoder, where the encoder and / or decoder can be implemented in a user device, such as a mobile phone, laptop, tablet, etc. The background noise estimator could be further included in a network node, such as a media network gateway, e.g. as part of a codec.
[0145] Fig. 5. is a block diagram schematically showing an embodiment of the background noise estimator according to one embodiment. The input signal framing block 51 first divides the input signal into frames of the appropriate length, e.g. 5-30 ms. For each frame, the feature extractor 52 calculates at least the following features from the input signal: 1) the feature extractor analyzes the frame in the frequency domain and the energy for the subband set is calculated. These subbands are the same subbands that were used in the background estimation. 2) the feature extractor further analyzes the time domain frame and calculates the correlation marked e.g. cor_est and / or lt_cor_est, used to determine if the frame contains active content or not. 3) the feature extractor further uses the total energy of the current frame, marked e.g. Etot, for updating traits for current energy history and earlier input frames, such as long-term minimum energy, lt_min. The correlation and energy features are then passed to the update decision logic block 53.
[0146] Here, the decision logic is implemented according to the solution disclosed in the logic block
EP 3 309 784 decyzji1 update decision 53, where correlation and energy features are used to create a decision about whether the current frame's energy is close to long-term minimum energy or not; whether the current frame is part of the break (not an active signal) or not; and whether the current frame is part of the music or not. The solution to the embodiments described here is related to how these features and decisions are used to update the background noise estimate in a robust manner.
[0147] In the following, some details of the embodiments of the disclosed solution will be described. The following implementation details are taken from an embodiment in the G.718 based encoder. This embodiment uses some of the features described in WO2011 / 049514 and WO2011 / 049515.
[0148] The following features are defined in the modified G.718 described in WO2011 / 09514
<td>etot</td><td>total energy for the current input frame</td>
<td>EtotJ EtotJJp totalNoise</td><td>traces the minimum energy envelope smoothed version of the minimum energy envelope EtotJ current total energy of the background estimate</td>
<td>bckr [i] tmpN [i] aEn</td><td>vector with subband background estimates pre-calculated new possible background estimate background detector using many features (counter)</td>
<td>harm_cor_cnt act_pred cor [i]</td><td>counts the frames from the last frame from a harmonic event or correlation activity forecast only for the features of the input frame vector with correlation estimates for i = 0 end of the current frame, i = 1 start of the current frame, i = 2 end of the previous frame</td>
[0149] The following features are defined in the modified G.718 described in WO2011 / 09515
<td>Etot_h sign_dynjp</td><td>tracks the maximum energy envelope, smoothed dynamics of the input signal</td>
[0150] Also the feature Etot_v_h is defined in WO2011 / 049514, but in this embodiment it has been modified and is now implemented as follows:
Etot_v = (float) fabs (* Etot_last - Etot);
if (Etot_v <7.0f) / * note that no VAD flag or similar is used here * / {* Etot_v_h - = O.Of;
if (Etot_v> * Etot_v_h) {if ((* Etot_v - * Etot_v_h)> 0.2f) {
* Etot_v_h = * Etot_v_h + 0.2f;
} else {
* Etot_v_h = Etot_v; }}} [0151] Etot_v measures the absolute energy variation between frames, i.e. the absolute value of the instantaneous energy variation between frames. In the above examples, it is determined that the energy variation between the two frames is "small" when the difference between the energy of the last frame and the energy of the current frame is less than 7 units. This is used as an indicator that the current frame (and previous frame) may be part of the gap, i.e. contain only background noise. However, such a small variability could also be found, for example, in the middle of a speech fragment. The Etotjast variable is the energy level of the previous frame.
[0152] The above steps described in the code may be performed as part of the "Calculation and update of correlation and energy" steps in the flowchart of Fig. 2, i.e. as part of step 201. In the embodiment of WO2011 / 049514 VAD to determine if the current audio segment contains background noise or not. The developers have realized that dependency on feedback can cause problems. In the solution described here, the decision whether to update the background noise estimate or not depends on the VAD (or SAD) decision.
[0153] Furthermore, in the solution described herein, the following features, which are not part of the embodiment of WO2011 / 049514, can be calculated / updated as part of the same steps, ie the steps "Calculation / update of correlation and energy" shown in Fig. 2. These features are also used in the decision logic of whether to update the background estimate or not.
[0154] For a more accurate estimate of background noise, a number of features are defined below. For example, new features related to cor_est and lt_cor_est have been defined. The cor_est feature is a correlation estimation in the current frame and cor_est is also used to obtain lt_cor_est, which is a smoothed long-term correlation estimation.
cor_est = (cor [0] + cor [1] + cor [2]) / 3.0f;
st-> lt_cor_est = 0.01f * cor_est + 0., 9f * st-> lt_cor_est;
[0155] As defined above, cor [i] is a vector containing correlation estimates and cor [0] is the end of the current frame, cor [1] is the beginning of the current frame, and cor [2] is the end of the previous frame.
[0156] In addition, a new feature, lt_tn_track, is calculated that provides a long-term estimate of how often the background estimates are close to the current energy of the frame. When the energy of the current frame is close enough to the current background estimate, this is recorded by the condition that it signals (1/0) whether the background is close or not. This signal is used to create the long-term measure lt_tn_track.
st-> lt_ln_track = 0.03f * (Etot - st-> totalNoise <10) + 0.97f * st-> lt_ln_track;
[0157] In this example, 0.03 is added when the energy of the current frame is close to the background noise estimate, otherwise the remaining term is only 0.97 of the previous value. In this example, "close" is defined as the difference between the current frame's energy, Etot, and the background noise estimate, totalNoise, less than 10 units. Other definitions of the term "similar" are also possible.
[0158] Furthermore, the distance between the current background estimation, Etot, and the energy of the current frame, totalNoise, is used to determine the feature lt_tn_dist, which gives a long-term estimate of this distance. A similar feature, lt_Ellp_dist, is created for the distance between the long-term minimum energy Etot_l_lp and the energy of the current frame, Etot.
st-> lt_lnJist = 0.03f * (Etot - st-> total Noise) + 0.97f * st-> lt_ln_dist; st-> lt_Ellp_dist = 0.03f * (Etot - st-> Etot_l_lp) + 0.97f * st-> lt_Ellp_dist;
[0159] The above feature harm_cor_cnt is used to count the number of frames since the last frame having a correlation or harmonic event, i.e. from a frame that meets certain activity-related criteria. That is, when the condition harm_cor_cnt == 0, this suggests that the current frame is probably the active frame, because it presents a correlation or harmonic event. This is used to create a smoothed long-term estimate, lt_haco_ev, how often such events occur. In this case, the update is not symmetrical, i.e. different time constants are used if the estimate is increased or decreased, as can be seen below.
EP 3 309 784 Β1 if (st-> ha.rrri_cor_cnt == 0) / * when probably active * / (
st-> lt_haco_ev = 0.03f + 0.97f * st-> lt_haco_ev; / * increase long term estiiriate * /} else (
st-> lt_haco_ev = 0.99f * st-> lt_haco_ev; / Mecreaae long term estimate * /} [0160] The low value of the lt_ln_track feature introduced above indicates that the input frame energy was not close to the background energy for some frames. This is because lt_ln_track is reduced for each frame where the energy of the current frame is not close to the background energy estimate. It_ln_track is only increased when the energy of the current frame is close to the background energy estimate as shown above. To get a better estimate of how long it did "tracking", ie, the frame energy was far from the background estimate, a counter, low_tn_track_cnt, is created for the number of frames with this lack of tracking:
if (st-> lt_tn_tr-ack <0.05f) / * when lt_tn_track is Iow * / (
st-> low_tn_track_cnt ++; / * add 1 to counter * / else {
st-> low_tn_track_cnt = 0; / * reset counter * / [0161] In the above examples, "small" is defined as a value below 0.05. This should be seen as an example value that can be chosen as another value.
[0162] For the "Making decision on interruption and music" step shown in Fig. 2, the following three code expressions are used to generate a break detection, also referred to as background detection. Other break detection criteria can be added to other embodiments and implementations. The actual music decision is made in the code using correlation and energy features.
1:
bg_bgd = Etot <EtotJJp + 0.6f * st-> Etot_v_h;
bg_bgd will be "1" or "true" when Etot is close to estimating background noise. bg_bgd serves as a mask for other background detectors. That is, if bg_bgd is not "true", then the background detectors 2 and 3 below need not be estimated. Etot_v_h means an estimation of the noise variation that could alternatively be determined by Nvar. Etot_v_h is derived from the total energy of the input signal (in the log domain) using Etot_v, which measures the absolute energy variation between frames. Note that the Etot_v_h feature is limited only to increasing the maximum of a small fixed value, e.g. 0.2 for each frame. EtotJJp means the smoothed version of the EtotJ minimum energy envelope
2:
aE_bgd = st-> aEn == 0;
When aEn is zero, aE_bgd gets "1" or "true". aEn is a counter whose value is increased when it is determined that the active signal is present in the current frame, and decreased when it is determined that the current frame does not contain the active signal. aEn cannot be increased above a certain number, e.g. 6, or reduced to less than zero. After a certain number of consecutive frames, e.g. 6, without an active signal, aEn will be zero.
3:
sd1 _bgd = (st-> sign_dynjp> 15) && (Etot - st-> Etot_l_lp) <st-> Etot_v_h && st-> harm_cor_cnt> 20; [0163] Here, sd1_bgd will be "1" or "true" when three different conditions are true: Signal dynamics sign_dyn_lp is high, in this example greater than 15; The energy of the current frame is close to
EP 3 309 784 Β1 background estimation; i: A number of frames without correlation or harmonic events have passed, in this example 20 frames.
[0164] The bg_bgd function is being a flag for detecting that the energy of the current frame is close to long-term minimum energy. The last two, aE_bgd and sd1_bgd, indicate the detection of a gap or background in different conditions. aE_bgd is the more general detector of the two, while sd1_bgd mainly detects speech breaks at high SNR.
The new decision logic according to an embodiment of the technology disclosed herein is constructed as follows in the code below. Decision logic includes masking condition bg_bgd and two break detectors aE_bgd and sd1_bgd. There could also be a third gap detector that would evaluate long-term statistics for how well totalNoise tracks the minimum energy estimate. The evaluated condition, when the first line is true, is the decision logic about how large the size of the updt_step stage should be and the actual update of the noise estimation is the assignment of the value to "st-> bckr [i] = -". It should be noted that tmpN [i] is a previously calculated potentially new noise level, calculated according to the solution described in WO2011 / 049514. The decision logic below applies to part 209 of Fig. 2, which is partially indicated in conjunction with the code below.
if (bg_bgd && (aE_bgd II sdl_bgd II st-> lt_tn_track> 0.90f)) / * if 202: 2 and 204: 2) * / ł
if ((st-> act_pred <0.85f II (aE_bgd && st-> lt_haco_ev <0.05f)) &&
(st-> lt_Ellp_dist <10 II sdl_bgd) && st-> it_tn_dist <40 &&
((Etot - st-> totalNoise) <15.Of II st-> lt_haco_ev <O.lOf)) / * 207 * / {st-> first_noise_updt = 1;
for (i = 0; i <NB_BANDS; i ++) {
st-> bckr [i] = tmpN [i) / * 208 * /)} else if (aE_bgd && st-> lt_haco_ev <0.15f) (
updt stop-C<sup>1</sup>. if;
if (st-> act_pred> 0.85f) {
updt_step = 0.Olf / * 207 * /)) if (updt_step> O.Of) {
st-> first_noise_updt = 1;
fox [i = 0; and <NB_BANDS; i ++) (
st-> bckr [i] = st-> bckr [i] + updt_step * (tmpN [i] -st-> bckr [i]); / * 208 * /}}) else (
(st-> first_noise_updt) + = 1;
}) else {/ * If in musie lower bekr to drop further * // * if 203: 2 and 205: 1 * /
If (st-> low_tn_track_cnt> 300 S & st-> lt_haco_ev> O.Of & Ł st-> totalNoise> O.Of) (For (i = 0; i <NB_BMJDS; i ++) (If (st-> bckr [i] > 2 * E_MIN {
St-> bckr [i] = 0.98f * st-> bckr [i]; / * 206 * /)))
Else {
(st-> first_noise_updt) + = 1;
}} [0165] The code segment in the last block of code starting with "/ * If in musie ... * /" contains forced
Scaling down the background estimate, used when it is suspected that the current input signal is music. The decision is made as a function: the long period of poor background noise tracking compared to the minimum energy estimation and the frequent occurrence of harmonic or correlation events, and, as the last condition, "totalNoise> 0" is a check that the current total energy of the background estimation is greater than zero, suggesting that a reduction in background estimate may be considered. In addition, it is determined whether "bckr [i]> 2 * E_MIN", where E_MIN is a small positive number. This is a check of each item in the vector containing background estimates of the background subbands, so that the item must exceed E_MIN to be reduced (in the example - by multiplying by 0.98). These checks are performed to avoid reducing the background estimate to too low values.
[0166] Embodiments improve background noise estimation, which enables improved SAD / VAD operation to obtain a highly effective DTX solution and to avoid degradation of speech or music quality caused by clipping.
[0167] By removing the feedback feedback of the decision described in WO2011 / 09514 from Etot_v_h, there is a better separation between noise estimation and SAD. This is of benefit as the noise estimation is not changed if / when the SAD function / adjustment is changed. That is, determining the background noise estimate becomes independent of the SAD function. Also, adjusting the noise estimation logic becomes easier because there is no effect of secondary effects from SAD when background estimates are changed.
29 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29
70 members in 19 offices
Priority claims13
| Document | Office | Kind | Date |
|---|---|---|---|
| 201462030121 | United States of America | P | |
| 201462030121 | United States of America | P | |
| 15739357 | European Patent Office (EPO) | A | |
| 15739357 | European Patent Office (EPO) | A | |
| 17202308 | European Patent Office (EPO) | A | |
| 2015050770 | Sweden | W | |
| 2015050770 | Sweden | W | |
| 172023087 | – | – | – |
| 201462030121P | – | – | – |
| EP20150739357 | – | – | – |
| EP20170202308 | – | – | – |
| US201462030121P | – | – | – |
| WO2015SE50770 | – | – | – |
Members70
| Document | Office | Kind | |
|---|---|---|---|
| CA2956531A1 | Canada | A1 | |
| WO2016018186A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20170026545A | Republic of Korea | A | |
| US2017069331A1 | United States of America | A1 | |
| CN106575511A | China | A | |
| MX2017000805A | Mexico | A | |
| PH12017500031A1 | Philippines | A1 | |
| EP3175458A1 | European Patent Office (EPO) | A1 | |
| JP2017515138A | Japan | A | |
| JP6208377B2 | Japan | B2 | |
| EP3175458B1 | European Patent Office (EPO) | B1 | |
| US9870780B2 | United States of America | B2 | |
| BR112017001643A2 | Brazil | A2 | |
| JP2018041083A | Japan | A | |
| EP3309784A1 | European Patent Office (EPO) | A1 | |
| ES2664348T3 | Spain | T3 | |
| US2018158465A1 | United States of America | A1 | |
| HUE037050T2 | Hungary | T2 | |
| RU2017106163A | Russian Federation | A | |
| RU2017106163A3 | Russian Federation | A3 | |
| NZ728080A | New Zealand | A | |
| RU2665916C2 | Russian Federation | C2 | |
| KR101895391B1 | Republic of Korea | B1 | |
| KR20180100452A | Republic of Korea | A | |
| RU2018129139A | Russian Federation | A | |
| ZA201903140A0 | South Africa | A0 | |
| MX365694B | Mexico | B | |
| US10347265B2 | United States of America | B2 | |
| MX2019005799A | Mexico | A | |
| KR102012325B1 | Republic of Korea | B1 | |
| KR20190097321A | Republic of Korea | A | |
| US2019267017A1 | United States of America | A1 | |
| EP3309784B1 | European Patent Office (EPO) | B1 | |
| ZA201708141B | South Africa | B | |
| JP6600337B2 | Japan | B2 | |
| PT3309784T | Portugal | T | |
| EP3582221A1 | European Patent Office (EPO) | A1 | |
| RU2018129139A3 | Russian Federation | A3 | |
| RU2713852C2 | Russian Federation | C2 | |
| JP2020024435A | Japan | A | |
| PL3309784T3This record | Poland | T3 | |
| CA2956531C | Canada | C | |
| ES2758517T3 | Spain | T3 | |
| ZA201903140B | South Africa | B | |
| MY178131A | Malaysia | A | |
| JP6788086B2 | Japan | B2 | |
| BR112017001643B1 | Brazil | B1 | |
| CN106575511B | China | B | |
| EP3582221B1 | European Patent Office (EPO) | B1 | |
| NZ743390A | New Zealand | A | |
| DK3582221T3 | Denmark | T3 | |
| CN112927724A | China | A | |
| CN112927725A | China | A | |
| KR102267986B1 | Republic of Korea | B1 | |
| RU2020100879A | Russian Federation | A | |
| PL3582221T3 | Poland | T3 | |
| MX385944B | Mexico | B | |
| US11114105B2 | United States of America | B2 | |
| RU2020100879A3 | Russian Federation | A3 | |
| ES2869141T3 | Spain | T3 | |
| RU2760346C2 | Russian Federation | C2 | |
| US2021366496A1 | United States of America | A1 | |
| MX2021010373A | Mexico | A | |
| US11636865B2 | United States of America | B2 | |
| US2023215447A1 | United States of America | A1 | |
| CN112927724B | China | B | |
| CN112927725B | China | B | |
| MX385944B | Mexico | B | |
| US12347446B2 | United States of America | B2 | |
| US2025285630A1 | United States of America | A1 |
Numbers
- Publication
- 3309784
- Publication, DOCDB
- 3309784
- Publication, EPODOC
- PL3309784T
- Application
- 17202308
- Application, DOCDB
- 17202308
- Application, EPODOC
- PL20170202308T
Titles2
- English
- ESIMATION OF BACKGROUND NOISE IN AUDIO SIGNALS
- Polish
- Szacowanie szumu tła w sygnałach audio
Classification
- CPC, 10
- G10L25/78
- G10L19/012
- G10L19/0208
- G10L25/03
- G10L21/0216
- G10L25/12
- G10L19/02
- G10L21/0324
- G10L19/04
- G10L21/0388
- IPC, 1
- G10L25 78
