Comfort noise addition for modeling background noise at low bit-rates.
Abstract
The invention provides a decoder being configured for processing an encoded audio bitstream (BS), wherein the decoder (1 ) comprises: a bitstream decoder (2) configured to derive a decoded audio signal (DS) from the bitstream (BS), wherein the decoded audio signal (DS) comprises at least one decoded frame; a noise estimation device (3) configured to produce a noise estimation signal (NE) containing an estimation of the level and/or the spectral shape of a noise (N) in the decoded audio signal (DS); a comfort noise generating device (4) configured to derive a comfort noise signal (CN) from the noise estimation signal (NE); and a combiner (5) configured to combine the decoded frame of the decoded audio signal (DS) and the comfort noise signal (CN) in order to obtain an audio output signal (OS).

Term
7.2 yearsleft in the term
Expires 19 December 2033.
- Priority
- Filed
- Granted
- Today
- Expires
27 claims: 18 independent, 9 dependent
- 1CLAIMS REIVINDICACIONES 1. Un decodificador que está configurado para procesar un flujo de bits de audio codificado (BS), donde el decodificador (1) comprende:one. A decoder that is configured to process an encoded audio bit stream (BS), where the decoder (1) comprises: a bit stream decoder (2) configured to derive a decoded audio signal (DS) from the bit stream (BS), where the decoded audio signal (DS) comprises at least one decoded frame;un decodificador de flujos de bits (2) configurado para derivar una señal de audio decodificada (DS) del flujo de bits (BS), donde la señal de audio decodificada (DS) comprende por lo menos una trama decodificada;a noise estimation device (3) configured to produce a noise estimation signal (NE) containing an estimate of the level and / or spectral form of a noise (N) of the decoded audio signal (DS);un dispositivo de estimación de ruido (3) configurado para producir una señal de estimación de ruido (NE) que contiene una estimación del nivel y/o la forma espectral de un ruido (N) de la señal de audio decodificada (DS);a comfort noise generating device (4) configured to derive a comfort noise signal (CN) from the noise estimation signal (NE) and a combiner (5) configured to combine the decoded frame of the decoded audio signal (DS) and the comfort noise signal (CN) to obtain an output audio signal (OS), so that the decoded frame of the output audio signal (OS) comprises artificial noise corresponding to the noise (N) contained in the decoded audio signal (DS). un dispositivo generador de ruido de confort (4) configurado para derivar una señal de ruido de confort (CN) de la señal de estimación de ruido (NE) y un combinador (5) configurado para combinar la trama decodificada de la señal de audio decodificada (DS) y la señal de ruido de confort (CN) para obtener una señal de audio de salida (OS), de modo que la trama decodificada de la señal de audio de salida (OS) comprenda ruido artificial correspondiente al ruido (N) contenido en la señal de audio decodificada (DS).
- 2El decodificador de acuerdo con la reivindicación precedente, donde la trama decodificada es una trama activa. two. The decoder according to the preceding claim, wherein the decoded frame is an active frame.
- 3The decoder according to one of the preceding claims, wherein the decoded frame is an active frame. 3. El decodificador de acuerdo con una de las reivindicaciones anteriores, en el cual la trama decodificada es una trama activa.
- 4El decodiflcador de acuerdo con una de las reivindicaciones anteriores, en el cual el dispositivo de estimación de ruido (3) comprende un dispositivo de análisis espectral (6) configurado para crear una señal de análisis (AS) que contiene el nivel y la forma espectral del ruido (N) en la señal de audio decodificada (DS) y un dispositivo para producir estimaciones de ruido (7) configurado para producir la señal de estimación de ruido (NE) sobre la base de la señal de análisis (AS). Four. The decoder according to one of the preceding claims, wherein the noise estimation device (3) comprises a spectral analysis device (6) configured to create an analysis signal (AS) containing the level and spectral form of the noise (N) in the decoded audio signal (DS) and a device for producing noise estimates (7) configured to produce the noise estimation signal (NE) based on the analysis signal (AS).
- 5The decoder according to one of the preceding claims, wherein the comfort noise generating device (4) comprises a noise generator (8) configured to create a comfort noise signal in the frequency domain (FD) over the base of the noise estimation signal (NE) and a spectral synthesizer (9) configured to create the comfort noise signal (CN) based on the comfort noise signal in the frequency domain (FD). 5. El decodificador de acuerdo con una de las reivindicaciones anteriores, donde el dispositivo generador de ruido de confort (4) comprende un generador de ruido (8) configurado para crear una señal de ruido de confort en el dominio de la frecuencia (FD) sobre la base de la señal de estimación de ruido (NE) y un sintetizador espectral (9) configurado para crear la señal de ruido de confort (CN) sobre la base de la señal de ruido de confort en el dominio de la frecuencia (FD).
- 6The decoder according to one of the preceding claims, wherein the decoder (1) comprises a switching device (10) configured to switch the decoder alternately to a first mode of operation or a second mode of operation, where in the first mode The comfort noise signal (CN) is fed to the combiner (5), while the comfort noise signal (CN) is not fed to the combiner (5) in the second mode of operation. 6. El decodificador de acuerdo con una de las reivindicaciones anteriores, donde el decodificador (1) comprende un dispositivo conmutador (10) configurado para conmutar el decodificador en forma alternada a un primer modo de operación o a un segundo modo de operación, donde en el primer modo de operación la señal de ruido de confort (CN) es alimentada al combinador (5), en tanto que la señal de ruido de confort (CN) no es alimentada al combinador (5) en el segundo modo de operación.
- 7The decoder according to the preceding claim, wherein the decoder (1) comprises a control device (11) configured to control the switching device (10) automatically, wherein the control device (11) comprises a noise detector ( 12) and configured to control the switching device (11) depending on a signal-to-noise ratio of the decoded audio signal (DS) where, under conditions of low signal-to-noise ratio, The decoder (1) is switched to the first mode of operation and in conditions of high signal to noise ratio to the second mode of operation. 7. El decodificador de acuerdo con la reivindicación precedente, donde el decodificador (1) comprende un dispositivo de control (11) configurado para controlar el dispositivo conmutador (10) en forma automática, donde el dispositivo de control (11) comprende un detector de ruido (12) y configurado para controlar el dispositivo conmutador (11) dependiendo de una relación señal a ruido de la señal de audio decodificada (DS) donde, en condiciones de baja relación señal a ruido, el decodificador (1) se conmuta al primer modo de operación y en condiciones de alta relación señal a ruido al segundo modo de operación.
- 8The decoder according to the preceding claim, wherein the control device (11) comprises a complementary information receiver (13) configured to receive complementary information contained in the bit stream (BS), which corresponds to the signal to decoded audio signal (DS) noise, and configured to create a noise detection (ND) signal, where the noise detector (12) switches the switching device (11) depending on the noise detection signal (ND). 8. El decodificador de acuerdo con la reivindicación precedente, en el cual el dispositivo de control (11) comprende un receptor de información complementaria (13) configurado para recibir información complementaria contenida en el flujo de bits (BS), que corresponde a la relación señal a ruido de la señal de audio decodificada (DS), y configurado para crear una señal de detección de ruido (ND), donde el detector de ruido (12) conmuta el dispositivo conmutador (11) dependiendo de la señal de detección de ruido (ND).
- 9The decoder according to the preceding claim, wherein the complementary information corresponding to the signal-to-noise ratio of the decoded audio signal (DS) consists of at least one dedicated bit in the bit stream (BS). 9. El decodificador de acuerdo con la reivindicación precedente, en el cual la información complementaria que corresponde a la relación señal a ruido de la señal de audio decodificada (DS) consiste en por lo menos un bit dedicado en el flujo de bits (BS).
- 12The decoder according to one of the preceding claims, wherein the bit stream comprises active frames and inactive frames, wherein the decoder (1) comprises a complementary information receiver (17) configured to discriminate between active frames and frames. inactive based on complementary information contained in the bit stream (BS) that indicates whether the current frame is active or inactive. 12. El decodificador de acuerdo con una de las reivindicaciones anteriores, en el cual el flujo de bits comprende tramas activas y tramas inactivas, donde el decodificador (1) comprende un receptor de información complementaria (17) configurado para discriminar entre las tramas activas y las tramas inactivas sobre la base de información complementaria contenida en el flujo de bits (BS) que indica si la trama actual es activa o inactiva.
- 13The decoder according to the preceding claim, wherein the Supplementary Information indicating whether the current frame is active or inactive consists of at least one dedicated bit in the bit stream (BS). 13. El decodificador de acuerdo con la reivindicación precedente, en el cual la Información complementaria que indica si la trama actual es activa o inactiva consiste en por lo menos un bit dedicado en el flujo de bits (BS).
- 16The decoder according to one of the preceding claims, wherein the comfort noise generating device (4) is configured to create the comfort noise signal (CN) on the basis of a target level signal of comfort noise (TNL). 16. El decodificador de acuerdo con una de las reivindicaciones anteriores, en el cual el dispositivo generador de ruido de confort (4) está configurado para crear la señal de ruido de confort (CN) sobre la base de una señal de nivel objetivo de ruido de confort (TNL).
- 17The decoder according to the preceding claim, wherein the signal of comfort noise target level (TNL) is adjusted depending on a bit rate of the bit stream (BS). 17. El decodificador de acuerdo con la reivindicación precedente, en el cual la señal de nivel objetivo de ruido de confort (TNL) es ajustada dependiendo de una tasa de bits del flujo de bits (BS).
- 20El decodificador de acuerdo con una de las reivindicaciones anteriores, donde el decodificador (1) comprende un decodificador adicional de flujos de bits, donde el decodificador de flujos de bits (2) y el decodificador adicional de flujos de bits son de diferentes tipos, donde el decodificador (1) comprende un conmutador configurado para alimentar la señal decodificada (DS) procedente del decodificador de flujos de bits (2) o la señal decodificada procedente del decodificador adicional de flujos de bits al dispositivo de estimación de ruido (3) y al combinador (5). twenty. The decoder according to one of the preceding claims, wherein the decoder (1) comprises an additional bit stream decoder, wherein the bit stream decoder (2) and the additional bit stream decoder are of different types, wherein the decoder (1) comprises a switch configured to feed the decoded signal (DS) from the bitstream decoder (2) or the decoded signal from the additional bitstream decoder to the noise estimation device (3) and to the combiner (5).
- 21Un codificador que está configurado para producir un flujo de bits de audio (BS), donde el codificador (18) comprende:twenty-one. An encoder that is configured to produce an audio bit stream (BS), where the encoder (18) comprises: a bit stream encoder (20) configured to produce an encoded audio signal (ES) corresponding to an input audio signal (IS) and to derive the bit stream (BS) of the encoded audio signal (ES) );un codificador de flujos de bits (20) configurado para producir una señal de audio codificada (ES) que corresponde a una señal de audio de entrada (IS) y para derivar el flujo de bits (BS) de la señal de audio codificada (ES);a signal analyzer (30) consisting of a signal-to-noise ratio estimator (33) configured to determine the signal-to-noise ratio of the input audio signal (IS) based on the desired signal energy (WS) of the input audio signal (IS) determined by an energy estimator of the desired signal (31) and based on an energy of a noise (N) of the input audio signal (IS) determined by an estimator of noise energy (32);un analizador de señales (30) que consta de un estimador de relación señal a ruido (33) configurado para determinar la relación señal a ruido de la señal de audio de entrada (IS) sobre la base de la energía de señal deseada (WS) de la señal de audio de entrada (IS) determinada por un estimador de energía de la señal deseada (31) y en base a una energía de un ruido (N) de la señal de audio de entrada (IS) determinada por un estimador de energía de ruido (32);a noise reduction device (27, 28) configured to produce an audio signal with noise reduction (TS) and a switching device (35) configured to feed, depending on the signal to noise ratio determined from the audio signal of input (IS), either the input audio signal (IS) or the noise reduction audio signal (TS), to the bitstream encoder (20) in order to encode the respective signal (IS, TS ), where the bit stream encoder (20) is configured to transmit complementary information (NF), which indicates whether the input audio signal (IS) or the noise reduction audio signal (TS) is being encoded within the bit stream (BS). un dispositivo de reducción de ruido (27, 28) configurado para producir una señal de audio con reducción de ruido (TS) y un dispositivo conmutador (35) configurado para alimentar, dependiendo de la relación señal a ruido determinada de la señal de audio de entrada (IS), ya sea la señal de audio de entrada (IS) o la señal de audio con reducción de ruido (TS), al codificador de flujos de bits (20) con el fin de codificar la señal respectiva (IS, TS), donde el codificador de flujos de bits (20) está configurado para transmitir una información complementaria (NF), que indica si se está codificando la señal de audio de entrada (IS) o la señal de audio con reducción de ruido (TS) dentro del flujo de bits (BS).
- 232. 3. A method for decoding an audio bit stream (BS), where the method comprises:23. Un método para decodificar un flujo de bits de audio (BS), donde el método comprende: derivar una señal de audio decodificada (DS) del flujo de bits (BS), en donde la señal de audio decodificada (DS) comprende por lo menos una trama decodificada;deriving a decoded audio signal (DS) from the bit stream (BS), wherein the decoded audio signal (DS) comprises at least one decoded frame;produce a noise estimation signal (NE) containing an estimate of the level and / or spectral form of a noise (N) of the decoded audio signal (DS);producir una señal de estimación de ruido (NE) que contiene una estimación del nivel y/o la forma espectral de un ruido (N) de la señal de audio decodificada (DS);derivar una señal de ruido de confort (CN) de la señal de estimación de ruido (NE);y combinar la trama decodificada de la señal de audio decodificada (DS) y la señal de ruido de confort (CN) para obtener una señal de audio de salida (OS), de tal manera que la trama decodificada en la señal de audio de salida (OS) comprenda ruido artificial correspondiente al ruido (N) contenido en la señal de audio decodificada (DS). derive a comfort noise signal (CN) from the noise estimate signal (NE);and combining the decoded frame of the decoded audio signal (DS) and the comfort noise signal (CN) to obtain an output audio signal (OS), such that the frame decoded into the output audio signal (OS) comprises artificial noise corresponding to the noise (N) contained in the decoded audio signal (DS).
- 24A method of encoding audio signals to produce an audio bit stream (BS), where the method comprises:24. Un método de codificación de señales de audio para producir un flujo de bits de audio (BS), donde el método comprende: determinar la relación señal a ruido de una señal de audio de entrada (IS) sobre la base de una energía determinada de señal deseada (WS) de la señal de audio de entrada (IS) y una energía determinada de un ruido (N) de la señal de audio de entrada (IS);determine the signal-to-noise ratio of an input audio signal (IS) on the basis of a certain desired signal energy (WS) of the input audio signal (IS) and a determined energy of a noise (N) of the input audio signal (IS);produce an audio signal with noise reduction (TS);producir una señal de audio con reducción de ruido (TS);produce an encoded audio signal (ES) corresponding to the input audio signal (IS), where, depending on the signal-to-noise ratio of the input audio signal (IS), the audio signal is encoded input (IS) or audio signal with noise reduction (TS);producir una señal de audio codificada (ES) que corresponde a la señal de audio de entrada (IS), en donde, dependiendo de la relación señal a ruido determinada de la señal de audio de entrada (IS), se codifica la señal de audio de entrada (IS) o la señal de audio con reducción de ruido (TS);derivar el flujo de bits (BS) de la señal de audio codificada (ES);y transmitir una información complementaria (NF), que indica si se está codificando la señal de audio de entrada (IS) o la señal de audio con reducción de ruido (TS) dentro del flujo de bits (BS). derive the bit stream (BS) from the encoded audio signal (ES);and transmit complementary information (NF), which indicates whether the input audio signal (IS) or the noise reduction audio signal (TS) is being encoded within the bit stream (BS).
- 27A computer readable medium for signal coding of 27. Un medio legible por computadora para codificación de señales de 5 audio to produce an audio bit stream (BS), comprising the method according to claim 24. 5 audio para producir un flujo de bits de audio (BS), que comprende el método de acuerdo con la reivindicación 24.
Independent claims18
170 paragraphs, as filed
The present invention relates to the processing of audio signals and, in particular, to the encoding of loud voice and the addition of control noise to the audio signals.
Usually, comfort noise generators are used in the discontinuous transmission (DTX) of audio signals, in particular audio signals with voice content. In that mode, the audio signal is classified, first, into active and inactive frames by a voice activity detector (VAD). An example of VAD can be found in [1], Based on the result of the VAD, only active voice frames are encoded and transmitted at the nominal bit rate. During prolonged pauses, when only background noise is present, the bit rate is reduced or set to zero and the background noise is coded episodically and parametrically. This is how the average bit rate is significantly reduced. The noise is generated during inactive frames on the decoder side by means of a comfort noise generator (CNG). For example, the AMR-WB [2] and ITU G.718 [1] voice encoders have the possibility of working in both cases in DTX mode.
The voice coding and, especially the loud voice at low bit rates, is prone to alterations. Voice encoders are usually based on a voice production model that is no longer supported in the presence of background noise. In that case, the encoding loses efficiency and the quality of the decoded audio signal is reduced. Moreover, certain features of voice coding can be especially annoying when handling loud voice. Indeed, at low bit rates, the coarse quantification of the coding parameters produces some fluctuation over time, and the fluctuations are perceptually annoying with the voice coding on fixed background noise.
Noise reduction is a well known technique to intensify speech intelligibility and improve communication in the presence of background noise. It has also been adopted in voice coding. For example, the G.718 encoder uses noise reduction to deduce certain coding parameters such as voice tone. It also has the possibility of encoding the intensified signal instead of the original signal. Then the voice is more predominant compared to the noise level in the decoded signal. However, it usually sounds more degraded or less natural, since noise reduction could distort voice components and cause audible musical noise alterations in addition to encoding alterations.
The objective of the present invention is to present improved concepts for the processing of audio signals. The object of the present invention is obtained by means of a decoder according to claim 1, an encoder according to claim 18, a system according to claim 19, a method according to claim 20 or
21, a bit stream according to claim 22 and a computer program according to claim 15.
In one aspect, the invention features a decoder that is configured to process an encoded audio bit stream, where the decoder comprises:
a bit stream decoder configured to derive a decoded audio signal from the bit stream, where the decoded audio signal comprises at least one decoded frame;
a noise estimation device configured to produce a noise estimation signal containing an estimate of the level and / or spectral form of a noise in the decoded audio signal;
a comfort noise generating device configured to derive a comfort noise signal from the noise estimation signal and a combiner configured to combine the decoded frame of the decoded audio signal and the comfort noise signal to obtain a signal of audio output
The bit stream decoder can be a device or a computer program capable of decoding an audio bit stream, which is a digital data stream that contains audio information. The decoding process gives rise to a digital decoded audio signal, which can be fed to an A / D converter to produce an analog audio signal, which can then be fed to a speaker to produce a sound signal.
The decoded audio signal is divided into so-called frames, where each of these frames contains audio information regarding a certain time interval. Such frames can be classified into active frames and inactive frames, where an active frame is a frame that contains desired components of the audio information, such as voice or music, while an inactive frame is a frame that does not contain any component. Desired audio information. Inactive frames usually appear during pauses, where there is no presence of any desired component, such as music or voice. Therefore, inactive frames usually contain only background noise.
In the discontinuous transmission (DTX) of audio signals, only the active frames of the decoded audio signal are obtained by decoding the bit stream, since during inactive frames, the encoder does not transmit the audio signal within the bit stream.
In the non-discontinuous (non-DTX) transmission of audio signals, the active frames and also the inactive frames are obtained by decoding the bit stream.
Frames that are obtained by decoding the bit stream by the bit stream decoder are called decoded frames.
The noise estimation device is configured to produce a noise estimation signal that contains an estimate of the level and / or spectral form of a noise included in the decoded audio signal. Moreover, the comfort noise generating device is configured to derive a comfort noise signal from the noise estimation signal. The noise estimation signal may be a signal that contains information regarding the characteristics of the noise contained in the audio signal decoded in parametric form. The signal of comfort noise is an artificial audio signal, which corresponds to the noise contained in the decoded audio signal. These features allow comfort noise to sound like real background noise without requiring any additional information regarding the background noise contained in the bit stream.
The combiner is configured to combine the decoded frame of the decoded audio signal and the comfort noise signal to obtain an output audio signal. As a result, the output audio signal comprises decoded frames, which comprise artificial noise. The artificial noise contained in the decoded frames makes it possible to mask the alterations of the output audio signal especially when the bit stream is transmitted at low bit rates. It smoothes the fluctuations usually observed and at the same time masks the predominant coding alterations.
Unlike the prior art, the present invention applies the principle of adding artificial comfort noise to decoded frames. The concept of the invention can be applied to both DTX and non-DTX modes.
The invention discloses a method for intensifying the quality of the encoded and transmitted loud voice at low bit rates. At low bit rates, the coding of the loud voice, that is to say voice recorded with background noise, is usually not as efficient as the coding of the clean voice. Decoded synthesis is generally prone to alterations. The two types of origins, noise and voice, cannot be efficiently encoded by an encoding scheme that is based on a single origin model. The present invention offers a concept for modeling and synthesizing background noise on the decoder side and requires little or no complementary information. This is obtained by estimating the level and spectral shape of the background noise on the decoder side and artificially generating a comfort noise. The generated noise is combined with the decoded audio signal and allows the coding artifacts to be masked.
In addition, the concept can be combined with a noise reduction scheme applied on the encoder side. Noise reduction intensifies the level of the signal-to-noise ratio (SNR) and improves the efficiency of subsequent audio coding. The amount of missing noise in the decoded audio signal is then compensated by the comfort noise on the decoder side. However, it usually sounds more degraded or less natural, since noise reduction could distort audio components and cause distortions of musical noise in addition to coding alterations. One aspect of the present invention is to mask such unpleasant distortions by adding a comfort noise from the decoder side. When a noise reduction scheme is used, the addition of comfort noise does not impair the SNR. Moreover, comfort noise conceals a large part of the annoying musical noise typical of noise reduction techniques.
In a preferred embodiment of the invention the decoded frame is an active frame. This feature extends the principle of adding comfort noise to decoded active frames.
In a preferred embodiment of the invention the decoded frame is an active frame. This feature extends the principle of adding comfort noise to decoded idle frames.
In a preferred embodiment of the invention the noise estimation device comprises a spectral analysis device configured to create an analysis signal containing the level and spectral form of the noise present in the decoded audio signal and a device for producing Noise estimates configured to produce the noise estimate signal based on the analysis signal.
In a preferred embodiment of the invention, the comfort noise generating device comprises a noise generator configured to create a comfort noise signal in the frequency domain based on the noise estimation signal and a spectral synthesizer. configured to create the comfort noise signal based on the comfort noise signal in the frequency domain.
In a preferred embodiment of the invention, the decoder comprises a switching device configured to switch the decoder alternately to a first mode of operation or a second mode of operation where, in the first mode of operation, the comfort noise signal it is fed to the combiner, while the comfort noise signal is not fed to the combiner in the second mode of operation. These characteristics allow the use of artificial comfort noise to be abandoned in situations where it is not necessary.
In a preferred embodiment of the invention, the decoder comprises a control device configured to control the switching device automatically, wherein the control device comprises a noise detector configured to control the switching device depending on a signal-to-noise ratio. of the decoded audio signal, where in low signal-to-noise conditions the decoder is switched to the first mode of operation and in conditions of high signal-to-noise ratio to the second mode of operation. Under these characteristics comfort noise can be activated only in loud voice situations, that is, not in clean voice situations or clean music. In order to discriminate between conditions of low signal-to-noise ratio and conditions of high signal-to-noise ratio, a threshold for the signal-to-noise ratio can be defined and used.
In a preferred embodiment of the invention, the control device comprises a complementary information receiver configured to receive complementary information contained in the bit stream, which corresponds to the signal-to-noise ratio of the decoded audio signal, and configured to create a noise detection signal, where the noise detector controls the switching device depending on the noise detection signal. These features allow the switching device to be controlled based on an analysis of the signal performed by an external device that produces and / or processes the received bit stream. The external device can be especially an encoder to produce the bit stream.
In a preferred embodiment of the invention, the complementary information corresponding to the signal-to-noise ratio of the decoded audio signal consists of at least one dedicated bit of the bit stream. In general, a dedicated bit is a bit that contains, alone or together with other dedicated bits, defined information. In this context, the dedicated bit can indicate whether the signal-to-noise ratio is higher or lower than a predefined threshold.
In a preferred embodiment of the invention, the control device comprises a desired signal energy estimator configured to determine the energy of a desired signal of the decoded audio signal, a noise energy estimator configured to determine the energy of a noise of the decoded audio signal and a signal to noise ratio estimator configured to determine the signal to noise ratio of the decoded audio signal on the basis of the energy of the desired signal and based on the noise energy, where the switching device is switched depending on the signal to noise ratio determined by the control device. In this case, no additional information is necessary in the bit stream. Since the energy of the desired signal is generally higher than the noise energy of the decoded signal, the total energy of the decoded audio signal, which includes the desired signal energy as well as the noise energy, provides Approximate estimation of the desired signal energy of the decoded audio signal. For this reason, the signal-to-noise ratio can be calculated by approximation by dividing the total energy of the decoded audio signal by the noise energy of the decoded signal.
In a preferred embodiment of the invention, the bit stream contains active frames and inactive frames, where the control device is configured to determine the desired signal energy of the decoded audio signal during the active frames and to determine the energy of the noise of the decoded audio signal during inactive frames. In this way, great precision can be obtained by estimating the signal to noise ratio in a simple way.
In a preferred embodiment of the invention, the bit stream contains active frames and inactive frames, where the decoder comprises a complementary information receiver configured to discriminate between active frames and inactive frames based on complementary information contained in the flow. bit, which indicates whether the current frame is active or inactive. Thanks to this feature, active frames or inactive frames can be identified respectively, without calculation effort.
In a preferred embodiment of the invention, the complementary information indicating whether the current frame is active or inactive consists of at least one dedicated bit in the bit stream.
In a preferred embodiment of the invention the control device is configured to determine the desired signal energy of the decoded audio signal on the basis of the analysis signal. In this case the analysis signal, which usually must be computed to perform the noise estimation, can be reused, so that complexity can be reduced.
In a preferred embodiment of the invention the control device is configured to determine the noise energy of the decoded audio signal on the basis of the noise estimation signal. In that embodiment the noise estimation signal, which generally has to be computed in order to generate comfort noise, can be reused, whereby complexity can be further reduced.
In a preferred embodiment of the invention, the comfort noise generating device is configured to create the comfort noise signal based on a comfort level target level signal. The level of added comfort noise should be limited to preserve intelligibility and quality. This can be achieved by scaling comfort noise using a target noise signal that indicates a predetermined target noise level.
In a preferred embodiment of the invention the target comfort noise level signal is adjusted depending on a bit rate of the bit stream. In general, the decoded audio signal exhibits a higher signal-to-noise ratio than the original input signal, especially at low bit rates where coding alterations are more severe. This attenuation of the noise level in voice coding comes from a paradigm of origin modeling that hopes to have the voice as input. Otherwise, source modeling coding is not entirely appropriate and cannot reproduce all the energy of non-voice components. Therefore, the comfort level target level signal can be adjusted depending on the bit rate to approximately compensate for the noise attenuation inherently introduced by the coding process.
In a preferred embodiment of the invention, the comfort noise target level signal is adjusted depending on a noise attenuation level caused by a noise reduction method applied to the bit stream. Thanks to these characteristics, the attenuation of the noise caused by a reduction of the noise module in an encoder can be compensated.
In a preferred embodiment of the invention the energy of the comfort noise signal in the domain of the random noise frequency w (/ c) is adjusted depending on the target level signal of comfort noise, which indicates a level comfort noise target g<sub>go</sub>, for each frequency fc as follows: = max {(^<sub>tar</sub> - 1) É<sub>n</sub>(/ c); o], where É<sub>n</sub>(k) refers to an estimate of the noise energy of the decoded audio signal at the frequency k, provided by the device to produce noise estimates. Under these characteristics, the intelligibility and quality of the output signal can be intensified.
In a preferred embodiment of the invention the decoder comprises an additional bit stream decoder, where the bit stream decoder and the additional bit stream decoder are of different types, wherein the decoder comprises a switch configured to feed the decoded signal from the bitstream decoder or the decoded signal from the additional bitstream decoder to the noise estimation device and the combiner. As the addition of comfort noise is effected when the bit stream decoder is used, as well as when the additional bit stream decoder is used, transition alterations can be minimized by switching between the stream decoder. bits and the additional bit stream decoder. For example, the bitstream decoder can be a linear predicted bitstream decoder (ACELP), while the additional bitstream decoder can be a core bitstream decoder based on transformed (TCX).
The invention also features an encoder for processing audio signals that is configured to produce an audio bit stream, wherein the encoder comprises:
a bitstream encoder configured to produce an encoded audio signal corresponding to an input audio signal and to derive the bitstream of the encoded audio signal;
a signal analyzer consisting of a signal-to-noise ratio estimator configured to determine the signal-to-noise ratio of the input audio signal based on the desired signal energy of the audio signal determined by an energy estimator of the desired signal and based on the energy of a noise of the input audio signal determined by a noise energy estimator;
a noise reduction device configured to produce an audio signal with noise reduction and a switching device configured to feed, depending on the determined signal to noise ratio of the input audio signal, either the input audio signal or the noise reduction audio signal to the bit stream encoder for the purpose of encoding the respective signal, where the bit stream encoder is configured to transmit complementary information, which indicates whether the input audio signal or the noise-reduced audio signal has been encoded within the bit stream.
The bitstream encoder may be a device or a computer program capable of encoding an audio signal, which is a digital data signal that contains audio information. The coding process gives rise to a digital bit stream, which can be transmitted by a digital data link to a decoder located at a remote point.
The input audio signal is directly encoded by the bit stream encoder. The bit stream encoder can be a voice encoder or a low delay scheme switching between an ACELP voice encoder and an audio encoder based on TCX transforms. The bitstream encoder is responsible for encoding the input audio signal and generating the bitstream necessary to decode the audio signal. In parallel, the input signal is analyzed by a module called signal analyzer. In a preferred embodiment the signal analysis is equal to that used in G.718. It consists of a spectral analysis device followed by the device to produce noise estimates. The spectra of the original signal and the estimated noise are sent as input to the noise reduction module. Noise reduction attenuates the level of background noise in the frequency domain. The amount of reduction is given by the intended level of attenuation. The intensified signal in the time domain (audio signal with noise reduction) is generated after spectral synthesis. The signal is used to deduce some characteristics, such as tone stability, which is then used by the VAD to discriminate between active and inactive frames. The result of the classification can also be used by the coding module. In the preferred embodiment, a specific coding mode is used to handle inactive frames. In this way, the decoder can say the VAD flag of the bit stream without the need for a dedicated bit.
To avoid unnecessary distortions in situations without noise (clean voice or clean music), noise reduction is applied only in case of loud voice and otherwise omitted. Discrimination between signals with and without noise is obtained by estimating the long-term energy of both the noise and the desired signal (voice or music). Long-term energy is computed by first-order autoregressive filtering of frame input energy (during active frames) or by using the noise estimation module output (during inactive frames). In this way an estimate of the signal to noise ratio can be calculated, which is defined as the ratio of the long-term energy of the voice or music during the long-term energy of the noise. If the signal to noise ratio is below a predetermined threshold, the frame is considered noisy, otherwise it is classified as a clean voice. Since the bitstream encoder is configured to transmit complementary information within the bitstream, which indicates whether the input audio signal or the noise-reduced audio signal is being encoded, the decoder can adjust the target level signal. Comfort noise automatically to encoder operating mode.
In the preferred embodiment of the invention, during active frames, only the long-term voice / music energy estimate is updated. During inactive frames, only the noise energy estimate is updated.
The invention further presents a system comprising a decoder for processing audio signals and an encoder for processing audio signals, where the decoder is designed in accordance with the claimed invention and / or the encoder is designed in accordance with the invention. claimed.
In another aspect the invention presents a method for decoding an audio bit stream, where the method comprises:
deriving a decoded audio signal from the bit stream, where the decoded audio signal comprises at least one decoded frame;
produce a noise estimation signal containing an estimate of the level and / or spectral form of a noise in the decoded audio signal;
derive a comfort noise signal from the noise estimation signal and combine the decoded frame of the decoded audio signal and the comfort noise signal to obtain an output audio signal.
The invention further presents a method for encoding audio signals to produce an audio bit stream, where the method comprises:
determining the signal-to-noise ratio of an input audio signal on the basis of a certain desired signal energy of the input audio signal and a determined energy of a noise of the input audio signal;
produce an audio signal with noise reduction;
produce an encoded audio signal corresponding to the input audio signal, where, depending on the determined signal-to-noise ratio of the input audio signal, the input audio signal or the audio signal with reduction of noise;
derive the bit stream of the encoded audio signal and transmit complementary information, which indicates whether the input audio signal or the audio signal with noise reduction within the bit stream is being encoded.
The invention also presents a bit stream produced in accordance with the method set forth above. The claimed bit stream contains complementary information, which indicates whether the input audio signal or the noise reduction audio signal is being encoded.
In a further aspect the invention discloses a computer program for practicing the methods of the invention when running on a computer or a processor.
The preferred embodiments of the invention are described below with respect to the accompanying drawings, in which:
Fig. 1 illustrates a first embodiment of a decoder according to the invention;
Fig. 2 illustrates a second embodiment of a decoder according to the invention;
Fig. 3 illustrates an encoder according to the prior art;
Fig. 4 illustrates a first embodiment of an encoder according to the invention;
Fig. 5 illustrates a second embodiment of an encoder according to the invention and
Fig. 6 illustrates an embodiment of a bit stream frame format according to the invention.
Fig. 1 illustrates a first embodiment of a decoder 1 according to the invention. Decoder 1 is configured to process a bit stream of encoded audio BS, where decoder 1 comprises:
a bit stream decoder 2 configured to derive a decoded audio signal DS from the bit stream BS, where the decoded audio signal DS comprises at least one decoded frame;
a noise estimation device 3 configured to produce a noise estimation signal NE containing an estimate of the level and / or spectral form of a noise N present in the decoded audio signal DS;
a comfort noise generating device 4 configured to derive a comfort noise audio signal CN from the noise estimate signal NE and a combiner 5 configured to combine the decoded frame of the decoded audio signal DS and the noise signal of comfort CN to obtain an audio signal of output OS.
The bit stream decoder 2 may be a device or a computer program capable of decoding an audio bit stream BS, which is a digital data stream that contains audio information. The decoding process gives rise to a digital decoded DS audio signal, which can be fed to an A / D converter to produce an analog audio signal, which can then be fed to a speaker to produce a sound signal.
The decoded audio signal DS comprises the so-called frames, where each of these frames contains audio information concerning a particular moment. Said frames can be classified into active frames and inactive frames, where an active frame is a frame containing desired components WS of the audio information, which is also referred to as the desired WS signal, such as voice or music, in so much so that an inactive frame is a frame that does not contain any desired component of the audio information. Inactive frames usually appear during pauses, where there is no presence of any desired component, such as music or voice. Therefore, inactive frames usually contain only background noise N.
The noise estimation device 3 is configured to produce a noise estimate signal NE containing an estimate of the level and / or spectral form of a noise included in the decoded audio signal DS. In addition, the comfort noise generating device 4 is configured to derive a comfort noise audio signal CN from the noise estimate signal NE. The noise estimation signal NE may be a signal that contains information regarding the characteristics of the noise N contained in the decoded audio signal DS in parametric form. The comfort noise signal CN is an artificial audio signal, which corresponds to the noise N contained in the decoded audio signal DS. These characteristics allow the comfort noise CN to sound like the actual background noise N without requiring any additional information regarding the background noise N contained in the bit stream BS.
The combiner 5 is configured to combine the decoded frame of the decoded audio signal DS and the comfort noise signal CN to obtain an output audio signal OS. As a result, the output audio signal OS comprises decoded frames, which comprise artificial noise CN. The artificial noise CN contained in the decoded frames makes it possible to mask the disturbances of the output audio signal OS especially when transmitting the bit stream BS at low bit rates.
Unlike the prior art, the present invention applies the principle of adding artificial comfort noise CN to decoded active or non-active frames. The concept of the invention can be applied to both DTX and non-DTX modes.
The invention presents a method for intensifying the quality of the encoded and transmitted loud voice at low bit rates. At low bit rates, the coding of the loud voice, that is the voice recorded with background noise N, is generally not as efficient as the clean voice coding WS. Decoded synthesis is usually prone to alterations. The two types of sources, the N noise and the WS voice, cannot be efficiently coded by an encoding scheme that is based on a single source model. The present invention offers a concept for modeling and synthesizing background noise N on the decoder side and requires little or no complementary information. This is obtained by estimating the level and spectral shape of the background noise N on the decoder side and artificially generating a comfort noise CN. The generated noise CN is combined with the decoded audio signal DS and allows masking the alterations by encoding during the decoded frames.
In addition, the concept can be combined with a noise reduction scheme applied on the encoder side. Noise reduction intensifies the level of the signal-to-noise ratio (SNR) and improves the efficiency of subsequent audio coding. The amount of missing noise N is then compensated for in the decoded audio signal DS by the comfort noise CN on the decoder side.
However, however, it usually sounds more degraded or less natural, since noise reduction could distort voice components and cause audible musical noise alterations in addition to encoding alterations. An aspect of the present invention consists in masking said unpleasant distortions by adding a comfort noise CN on the decoder side. When a noise reduction scheme is used, the addition of comfort noise does not impair the SNR. Moreover, comfort noise conceals a large part of the annoying musical noise typical of noise reduction techniques.
In a preferred embodiment of the invention the decoded frame is an active frame. This feature extends the principle of adding comfort noise to decoded active frames.
In a preferred embodiment of the invention the decoded frame is an active frame. This feature extends the principle of adding comfort noise to decoded idle frames.
In a preferred embodiment of the invention the noise estimation device 3 comprises a spectral analysis device 6 configured to create an analysis signal AS containing the level and spectral form of the noise present in the decoded audio signal DS and a device for producing noise estimates 7 configured to produce the noise estimation signal NE based on the analysis signal AS.
In a preferred embodiment of the invention, the comfort noise generating device comprises 4 a noise generator 8 configured to create a comfort noise signal in the frequency domain FD based on the noise estimation signal NE and a spectral synthesizer 9 configured to create the comfort noise signal CN based on the comfort noise signal in the frequency domain FD.
In a preferred embodiment of the invention the decoder 1 comprises a switching device 10 configured to switch the decoder 1 alternately to a first mode of operation or a second mode of operation where, in the first mode of operation, the signal of comfort noise CN is fed to the combiner, while the comfort noise signal CN is not fed to combiner 5 in the second mode of operation. These features allow the use of artificial comfort noise CN to be abandoned in situations where it is not necessary.
In a preferred embodiment of the invention, the decoder 1 comprises a control device 11 configured to control the switching device 10 automatically, where the control device 10 comprises a noise detector 12 configured to control the switching device 10 depending of a signal-to-noise ratio of the decoded audio signal DS where, under conditions of low signal-to-noise ratio, The decoder is switched to the first mode of operation and in conditions of high signal to noise ratio to the second mode of operation. By virtue of these characteristics, the use of CN comfort noise can be activated only in loud voice situations, that is, not in clean voice situations or clean music. To discriminate between conditions of low signal-to-noise ratio and conditions of high signal-to-noise ratio, a threshold for the signal-to-noise ratio can be defined and used.
In a preferred embodiment of the invention, the control device 11 comprises a complementary information receiver 13 configured to receive complementary information contained in the bit stream BS, which corresponds to the signal-to-noise ratio of the decoded audio signal DS , and configured to create a noise detection signal ND, where the noise detector 12 switches the switching device 11 depending on the noise detection signal ND. These features allow the switching device 10 to be controlled on the basis of an analysis of the signal executed by an external device that produces and / or processes the received bit stream BS. The external device can be, in particular, an encoder that produces the bit stream BS.
In a preferred embodiment of the invention, the complementary information corresponding to the signal-to-noise ratio of the decoded audio signal DS consists of at least one dedicated bit in the bit stream BS. In general, a dedicated bit is a bit that contains, alone or together with other dedicated bits, defined information. In this context, the dedicated bit can indicate whether the signal-to-noise ratio is higher or less than a predefined threshold.
In a preferred embodiment of the invention, the comfort noise generating device 4 is configured to create the comfort noise signal CN based on a target comfort noise level signal TNL. The added CN comfort noise level must be limited to preserve intelligibility and quality. This can be achieved by scaling comfort noise CN using a TNL target noise signal indicating a predetermined target noise level.
In a preferred embodiment of the invention, the TNL comfort noise target level signal is adjusted depending on a bit rate of the bit stream BS. In general, the decoded audio signal DS exhibits a higher signal-to-noise ratio than the original input signal, especially at low bit rates where the coding alterations are more severe. This attenuation of the noise level in voice coding comes from the origin modeling paradigm that the voice considers to be input. Otherwise, source modeling coding is not entirely appropriate and cannot reproduce all the energy of non-voice components. Therefore, the TNL comfort noise target level signal can be adjusted depending on the bit rate to approximately compensate for the noise attenuation inherently introduced by the coding process.
In a preferred embodiment of the invention the target comfort noise level signal TNL is adjusted depending on a noise attenuation level caused by a noise reduction method applied to the bit stream BS. Thanks to these characteristics, the noise attenuation caused by a noise reduction module in an encoder can be compensated.
In a preferred embodiment of the invention, the energy of the comfort noise signal in the domain of the frequency FD of the random noise w (fc) is adjusted depending on the target level signal of comfort noise TNL, which indicates a target comfort noise level £<sub>tar</sub>, for each frequency as follows: E<sub>w</sub>(/ c) = max {(^<sub>tar</sub> -1) F<sub>n</sub>(/ c); 0}, where E<sub>n</sub>(k) refers to an estimate of the noise energy N of the decoded audio signal DS at the frequency k, emitted by the device to produce noise estimates 7. Under these characteristics the intelligibility and quality of the output signal OS.
Fig. 2 illustrates a second embodiment of a decoder 1 according to the invention. The second embodiment of the decoder 1 is based on the decoder 1 of the first embodiment. Only the differences with respect to the first embodiment are explained and explained below.
In a preferred embodiment of the invention, the control device comprises a desired signal energy estimator 14 configured to determine the energy of the desired signal WS of the decoded audio signal DS, a noise energy estimator 15 configured to determine the energy of a noise N of the decoded audio signal DS and a signal to noise ratio estimator 16 configured to determine the signal to noise ratio of the decoded audio signal DS on the based on the energy of the desired signal WS and on the basis of the energy of the noise N, where the switching device 10 is switched depending on the signal to noise ratio determined by the control device 11. In this case, no additional information is necessary in the bit stream with respect to the signal-to-noise ratio. Therefore, the receiver of complementary information 13 in accordance with the first embodiment is also not necessary.
In a preferred embodiment of the invention, the bit stream BS contains active frames and inactive frames, where the control device 11 is configured to determine the desired signal energy WS of the decoded audio signal DS during the active frames and to determine the noise energy N of the decoded audio signal DS during the inactive frames. In this way, great precision can be obtained easily by estimating the signal-to-noise ratio.
In a preferred embodiment of the invention, the bit stream BS contains active frames and inactive frames, where the decoder 1 comprises a complementary information receiver 17 configured to discriminate between active frames and inactive frames based on complementary information. contained in the bit stream that indicates whether the current frame is active or inactive. Thanks to this feature, active frames or active frames can be identified, respectively, without calculation effort.
In the preferred embodiment of the invention, the complementary information receiver 17 may be configured to control and a switch 17<sup>to</sup>, which alternately feeds an output signal OW of the energy estimator of the desired signal 14 or an output signal ON of the noise energy estimator 15 to the signal to noise ratio estimator 16, where the output signal OW of an energy estimator of the desired signal 14 is fed to the signal-to-noise ratio estimator 16 during the active frames and where the output signal ON of the noise energy estimate estimate of 15 is fed to the estimator Signal to noise ratio 16 during inactive frames. By virtue of these characteristics, the signal-to-noise ratio can be calculated easily and accurately.
In a preferred embodiment of the invention, the control device 11 is configured to determine the desired signal energy of the decoded audio signal based on the AS analysis signal. In this case, the AS analysis signal, which usually must be computed for noise estimation, can be reused, so complexity can be reduced.
In a preferred embodiment of the invention, the control device 11 is configured to determine the noise energy N of the decoded audio signal DS based on the noise estimate signal NE. In that embodiment, the NE noise estimation signal can be reused, which generally must be computed for the generation of comfort noise, so that complexity can be further reduced.
In a preferred embodiment of the invention, the decoder 1 comprises an additional bitstream decoder (not illustrated in the figures), where the bitstream decoder 2 and the additional bitstream decoder are of different types, wherein the decoder 1 comprises a switch (not illustrated in the figures) configured to feed either the decoded signal DS from the bitstream decoder 2 or the decoded signal from the additional bitstream decoder to the device for estimating noise 3 and combiner 5. As the addition of comfort noise is performed when the bit stream decoder 2 is used, as well as when the additional bit stream decoder is used, transitional disturbances can be minimized by switching between the stream decoder bit 2 and the additional bit stream decoder. For example, the bit stream decoder 2 may be a linear predicted bit stream decoder excited by algebraic codes (ACELP), while the additional bit stream decoder may be a core based bit stream decoder in transformed (TCX.
The decoder 1 of the invention is described in Figures 1 and 2, where the addition of comfort noise is carried out blindly in the frequency domain. To have a comfort noise CN that resembles the actual background noise N, a noise estimation device 3 is used in the decoder 1 to determine the level and spectral shape of the background noise N, without requiring information complementary some.
The comfort noise generating device 4 comfort noise can only be activated in loud voice situations, that is, not in clean voice situations or clean music. The discrimination can be based on the detection made in the encoder. In this case, the decision must be transmitted using a dedicated bit. In a preferred embodiment, on the contrary, a device is applied to produce noise estimates 7 that is similar to the noise estimation device used in the encoder. It consists in estimating a long-term signal-to-noise ratio by adapting separately the long-term estimates of the noise energy N or the energy of the desired WS signal, such as voice and / or music, depending on the decision of the VAD The latter can be deduced directly from the index of ACELP and TCX modes. Indeed, TCX and ACELP can be executed in a specific mode called TCXNA and ACELP-NA, respectively, when the signal is from non-active voice / music frames, that is frames with background noise only. All ACELP and TCX modes refer to active frames. Therefore, the presence of a dedicated VAD bit in the bit stream can be avoided.
The level of added comfort noise must be limited to preserve intelligibility and quality. Thus, comfort noise is scaled to reach a predetermined target noise level. Yes #<sub>tar</sub> denotes the target level of noise amplification, the energy E is adjusted<sub>w</sub>of the random noise w (k) corresponding to each frequency k as follows:
Ew (fc) = max {(g<sub>tar</sub> - 1) É<sub>n</sub>(fc); O}, where É<sub>n</sub>(k) refers to an estimate of the noise energy present in the decoded audio output at the frequency k, emitted by the noise estimation module.
In general, the decoded audio signal DS exhibits a higher signal-to-noise ratio than the original input signal, especially at low bit rates where the coding alterations are more severe. This attenuation of the noise level in voice coding comes from a paradigm of origin modeling that estimates having the voice as input. Otherwise, source modeling coding is not entirely appropriate and cannot reproduce all the energy of non-voice components. Therefore, with reference to the first aspect of the invention, using the encoder illustrated in Figure 3, adjust the target comfort noise level ^<sub>tar</sub> depending on the bit rate to approximately compensate for the noise attenuation inherently introduced by the coding process.
As for the second aspect of the invention using the encoder illustrated in Figures 4 and 5, the target level of comfort noise #<sub>tar</sub> you must also take into account the noise attenuation caused by the noise reduction module included in the encoder.
Moreover, the addition of comfort noise described herein allows smoothing of the transitional alterations between one type of coding (e.g.) and the other (e.g., TCX) by uniformly adding a comfort noise in all the plots
Fig. 3 illustrates an encoder according to the prior art that can be used in combination with the decoders illustrated in Figures 1 and 2.
The input signal IS is directly encoded by the bit stream encoder 20. The bit stream encoder 20 can be a voice encoder or a low delay scheme switching between an ACELP voice encoder and an audio based encoder in TCX transforms. The bit stream encoder 20 comprises a signal encoder 21 to encode the IS signal and a bit stream producer 22 to generate the bit stream BS necessary to produce the decoded signal DS in the decoder 1. In parallel, the signal input IS is analyzed by a module called signal analyzer 23, which comprises a noise estimation device 24. In the preferred embodiment the noise estimation device 24 is equal to that used in G.718. It consists of a spectral analysis device 25 followed by a device for producing noise estimates 26. The SI spectrum of the original IS signal and the NI spectrum of the estimated noise are input to the noise reduction module 27. The noise reduction module 27 attenuates the background noise level in the intensified signal in the FS frequency domain. The amount of reduction is given by the attenuation level signal TAS. The intensified signal in the time domain (audio signal with noise reduction) TS is generated once the spectral synthesis is performed by the spectral synthesis device 28. The TS signal is used to deduce some characteristics, such as the stability of the tone, which is then used by the signal activity detector 29 to discriminate between active and inactive frames. The result of the classification can be used in turn by the encoder module 18. In a preferred embodiment, a specific coding mode is used to handle the inactive frames. In this way, decoder 1 can deduce the signal activity flag (VAD flag) from the bit stream without the need of a dedicated bit.
Fig. 4 illustrates a first embodiment of an encoder 18 according to the invention. The encoder 18 illustrated in Figure 4 is based on the encoder 18 set forth in Figure 3.
The encoder 18 set forth in Figure 4 is configured to produce an audio bit stream BS, where the encoder 18 comprises:
a bitstream encoder 20 configured to produce an encoded audio signal ES corresponding to an input audio signal IS and to derive the bitstream BS from the encoded audio signal ES;
a signal analyzer 19 consisting of a signal-to-noise ratio estimator 33 configured to determine the signal-to-noise ratio of the input audio signal IS based on the desired signal energy WS of the input audio signal IS determined by an energy estimator of the desired signal 31 and based on the energy of a noise N of the input audio signal IS determined by the noise energy estimator 32;
a noise reduction device 27, 28 configured to produce an audio signal with noise reduction TS and a switching device 35 configured to feed, depending on the signal to noise ratio determined from the input audio signal IS, the signal from IS input audio or noise reduction audio signal TS to bitstream encoder 20 in order to encode the respective signal IS, TS, where the bit stream encoder 20 is configured to transmit complementary information within the bit stream, which indicates whether the input audio signal IS or the audio signal with noise reduction TS is being encoded.
The bit stream encoder 20 may be a device or a computer program capable of encoding an audio signal, which is a digital data signal that contains audio information. The coding process gives rise to a digital bit stream, which can be transmitted through a digital data link to an existing decoder in a remote site.
The part of the encoder according to an embodiment of the invention is set forth in Figure 4. The main difference compared to Figure 3 lies in the fact that it encodes the noise reduction output, that is, the signal TS intensified. To avoid unnecessary distortions in situations without noise (clean voice or clean music), noise reduction is applied only in the case of loud voice and otherwise omitted. The discrimination between signals with and without noise is obtained by estimating the long-term energy of the desired signal WS (voice or music) performed by the energy estimator of the desired signal 31 and by estimating the long-term energy of the noise N made by the noise energy estimator 32. For this purpose the energy estimator of the desired signal 31 receives the signal of the SI spectrum corresponding to the input signal IS provided by the spectral analysis device 25. In turn, the noise energy estimator receives the estimation signal of NI noise corresponding to the input signal IS provided by the device to produce noise estimates 26. During active, long-term voice / music frames
WE. During inactive frames, only the NE noise energy estimate is updated. Long-term energy is computed by first-order self-regressive filtering of frame input energy (during active frames) or by using the noise estimation module output (during inactive frames). In this way the signal-to-noise ratio estimator 33 can calculate an estimate of the signal-to-noise ratio SR, which contains the ratio of the long-term energy of the voice or music WS during the long-term energy of the noise N. The signal-to-noise ratio RS is fed to a noise detector 34 which determines whether the current frame contains a loud audio signal or a clean audio signal. If the signal-to-noise ratio RS is less than a predetermined threshold, the frame is considered to be loud, otherwise it is classified as a clean voice.
The result of the classification is output as an output of the NF noise flag signal, which is used to control the switch 35. Moreover, the NF noise flag signal is fed to the bit stream encoder 20. The bit stream encoder 20 is configured to produce and transmit complementary information based on the noise flag signal NF within the bit stream, which indicates whether the input audio signal IS or the signal is being encoded. TS noise reduction audio. By decoding this flag, a decoder can adjust the target noise level automatically without classifying the decoded DS signal as loud or clean.
Fig. 5 illustrates a second embodiment of an encoder 18 according to the invention. The encoder 18 illustrated in Figure 5 is based on the encoder set forth in Figure 4. The additional features are explained below. In Figure 4 the signal analyzer 30 comprises a signal activity detector 36 that receives the spectrum signal SI corresponding to the input signal IS and the noise estimation signal NI. The signal activity detector 36 is configured to discriminate between active frames and inactive frames based on these two signals. The signal activity detector produces a signal activity signal SA which is transmitted, on the one hand, to the bit stream encoder 20 in order to adapt the bit stream BS to the signal activity and, on the other hand, it it is used to switch a switch 37 that is configured to alternately feed the desired signal energy signal WE or the noise energy signal EN to the signal-to-noise ratio estimator 33.
Fig. 6 illustrates an embodiment of a frame format FF of the bit stream BS according to the invention. The frame according to the FF frame format comprises a vector of SV signals consisting of a plurality of bits that are located at positions 0 to n. In the n + 1 position there is a bit consisting of an AF activity flag that indicates whether the frame is in an active or inactive frame. Likewise, the bit of position n + 2 is provided, which is an NF noise flag that indicates whether the frame contains a noisy signal or a clean signal. In the n + 3 position there is a bit that is a filling bit PB.
In a preferred embodiment of the invention, the complementary information indicating whether the current frame is active or inactive consists of at least one dedicated bit in the bit stream.
In summary, it can be said that, in one aspect of the invention, the original signal is encoded in decoder 1. It is encoded before adding it to an artificially generated comfort noise CN. The comfort noise generating device 4 does not require or only requires a small amount of complementary information. In a first embodiment, the comfort noise generating device 4 does not require complementary information and all processing is carried out blindly. In the preferred embodiment, the comfort noise generating device 4 needs to retrieve the VAD information (result of the classification of active and inactive frames) of the bit stream BS, which may already be present in the bit stream, and use it For other purposes. In a third embodiment, the comfort noise generating device 4 requires the encoder 18 a loud voice flag to discriminate between clean voice and loud voice. You can imagine any type of information encoded in a parametric way that can help direct the comfort noise generating device 4.
In another aspect of the invention, noise reduction is first applied to the original IS signal and an intensified signal TS is sent to the bit stream encoder 20, encoded and transmitted. At the end of the decoding, a comfort noise CN artificially generated is added to the decoded (intensified) DS signal. The target level of attenuation used for noise reduction in the encoder is a fixed value shared with the CNG module in the decoder. Therefore, it is not necessary to explicitly convey the objective level of attenuation.
While some aspects have been described in the context of an apparatus, it is obvious that these aspects also represent a description of the corresponding method, in which a block or device corresponds to a method step or a characteristic of a method step. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or item or a characteristic of a corresponding apparatus. Some or all of the steps of the method can be executed by means of (or using) a hardware device, such as a microprocessor, a computer program or an electronic circuit. In some embodiments, any one or more of the most important steps of the method may be executed by that type of apparatus.
Depending on certain implementation requirements, embodiments of the invention may be implemented in hardware or software. The implementation can be done using a digital storage medium, for example a soft disk, a DVD, a Blue-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, which is stored in the same electronically readable control signals, which cooperate (or have the capacity to cooperate) with a programmable computing system in such a way that the respective method is executed. Therefore, the digital storage medium can be readable by a computer.
Some embodiments according to the invention comprise a data transporter comprising electronically readable control signals, capable of cooperating with a programmable computing system such that one of the methods described herein is executed.
In general, the embodiments of the present invention can be implemented in the form of a computer program product with a program code, where the program code fulfills the function of executing one of the methods when the computer program is executed on a computer. The program code can be stored, for example, in a carrier readable by a machine.
Other embodiments include the computer program for executing one of the methods described herein, stored in a carrier readable by a machine.
In other words, an embodiment of the method of the invention consists, therefore, of a computer program consisting of a program code for performing one of the methods described herein when the computer program is executed on a computer.
Another of the embodiments of the methods of the invention consists, therefore, of a data carrier (or digital storage medium, or computer readable medium) comprising, recorded therein, the computer program for executing one of The methods described here. The data carrier, the digital storage medium or the recorded medium are generally tangible and non-transient.
Another embodiment of the method of the invention is, therefore, a data flow or signal sequence representing the computer program for executing one of the methods described herein. The data flow or the signal sequence may be configured, for example, to be transferred through a data communication connection, for example over the Internet.
Another embodiment comprises a processing means, for example a computer, a programmable logic device, configured or adapted to execute one of the methods described herein.
Another embodiment comprises a computer in which the computer program has been installed to execute one of the methods described herein.
Another embodiment according to the invention comprises an apparatus or system configured to transfer (for example electronically or optically) a computer program to implement one of the methods described herein in a receiver. The receiver can be, for example, a computer, a mobile device, a memory device and the like. The apparatus or system may comprise, for example, a file server to transfer the computer program to the receiver.
In some embodiments, a programmable logic device (for example an array of programmable doors in the field) can be used to execute some or all of the functionalities of the methods described herein. In some embodiments, an array of programmable doors in the field can cooperate with a microprocessor to execute one of the methods described herein. Generally, the methods are preferably executed by any hardware apparatus.
The embodiments described above are merely illustrative of the principles of the present invention. It is understood that the modifications and variations of the provisions and details described herein must be apparent to persons skilled in the art. Therefore, it is only intended to limit the scope of the following patent claims and not to the specific details presented by way of description and explanation of the embodiments presented herein.
Reference signs:
decoder bitstream decoder noise estimation device comfort noise generating device combiner spectral analysis device device to produce noise estimates noise generator spectral synthesizer switching device control device noise detector complementary information receiver signal energy estimator desired noise energy estimator signal to noise ratio estimator information receiver complementary
17<sup>to</sup> switch encoder signal analyzer bitstream encoder signal encoder bitstream producer signal analyzer noise estimation device spectral analysis device device to produce noise estimates noise reduction module spectral analysis device signal activity detector signal analyzer desired signal energy estimator noise energy estimator signal to noise ratio estimator noise switch signal activity detector switch
BS encoded audio bit stream
DS decoded audio signal
NE noise estimation signal
No noise
CN comfort noise signal
OS audio output signal
AS signal analysis
FD comfort noise signal in the frequency domain
ND noise detection signal
TNL comfort noise target level
IS input signal
ES coded signal
OW output signal of the desired signal energy estimator
ΟΝ noise energy estimator output signal
IF spectrum signal corresponding to the input signal
NI noise estimation signal corresponding to the input signal
TAS target attenuation signal
FS signal intensified in the frequency domain
TS audio signal with noise reduction
AD activity detector signal
WE desired signal energy signal
IN noise power signal
RS signal to signal to noise ratio
NF noise flag
SA signal activity signal
FF frame format
SV vector sign
AF activity flag
NF noise flag signal
PB fill bit
References:
[1] Recommendation ITU-T G.718: “Frame error robust narrow-band and wideband embedded variable bit-rate coding of speech and audio fro m 8-32 kbit / s [2] 3GPP TS 26.190“ Adaptive Multi-Rate wideband speech transcoding, ”3GPP
Technical Specification
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
42 members in 20 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 201261740883 | United States of America | P | |
| 201261740883 | United States of America | P | |
| 61740883 | United States of America | – | |
| 2013077527 | European Patent Office (EPO) | W | |
| 2013077527 | European Patent Office (EPO) | W | |
| US201261740883P | – | – | – |
| WO2013EP77527 | – | – | – |
Members42
| Document | Office | Kind | |
|---|---|---|---|
| CA2895391A1 | Canada | A1 | |
| CA2948015A1 | Canada | A1 | |
| WO2014096280A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201432671A | Taiwan Province of China | A | |
| AU2013366552A1 | Australia | A1 | |
| AR094279A1 | Argentina | A1 | |
| SG11201504899XA | Singapore | A | |
| KR20150107751A | Republic of Korea | A | |
| EP2936486A1 | European Patent Office (EPO) | A1 | |
| US2015364144A1 | United States of America | A1 | |
| CN105210148A | China | A | |
| JP2016500453A | Japan | A | |
| MX2015007854A | Mexico | A | |
| ZA201505191B | South Africa | B | |
| TWI553629B | Taiwan Province of China | B | |
| HK1217244A1 | Hong Kong, China | A1 | |
| KR101692659B1 | Republic of Korea | B1 | |
| KR20170001751A | Republic of Korea | A | |
| RU2015129782A | Russian Federation | A | |
| AU2013366552B2 | Australia | B2 | |
| RU2633107C2 | Russian Federation | C2 | |
| CA2948015C | Canada | C | |
| JP6335190B2 | Japan | B2 | |
| JP2018084834A | Japan | A | |
| BR112015014217A2 | Brazil | A2 | |
| EP2936486B1 | European Patent Office (EPO) | B1 | |
| PT2936486T | Portugal | T | |
| ES2688021T3 | Spain | T3 | |
| US2018342253A1 | United States of America | A1 | |
| US10147432B2 | United States of America | B2 | |
| PL2936486T3 | Poland | T3 | |
| US10339941B2 | United States of America | B2 | |
| MX366279BThis record | Mexico | B | |
| CA2895391C | Canada | C | |
| US2020013417A1 | United States of America | A1 | |
| CN111145767A | China | A | |
| CN105210148B | China | B | |
| US10789963B2 | United States of America | B2 | |
| KR102167541B1 | Republic of Korea | B1 | |
| MY178710A | Malaysia | A | |
| JP6849619B2 | Japan | B2 | |
| JP2021092816A | Japan | A |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 366279
- Publication, DOCDB
- 366279
- Publication, EPODOC
- MX366279
- Application
- 20150007854
- Application, DOCDB
- 2015007854
- Application, EPODOC
- MX20150007854
Titles2
- Spanish
- ADICION DE RUIDO DE CONFORT PARA MODELAR EL RUIDO DE FONDO A BAJAS TASAS DE BITS.
- English
- ADDITION OF COMFORT NOISE TO MODEL THE BACKGROUND NOISE AT LOW BIT RATES.
Classification
- CPC, 1
- G10L19/012
- IPC, 2
- G10L19 012
- G10L19 00