Apparatus and method for improved signal fade out in different domains during error concealment.
Abstract
Se proporciona un aparato para decodificar una señal de audio. El aparato comprende una interfaz receptora (110), en donde la interfaz receptora (110) está configurada para recibir un primer cuadro que comprende una primera porción de la señal de audio de la señal de audio, y en donde la interfaz receptora (110) está configurada para recibir un segundo cuadro que comprende una segunda porción de la señal de audio de la señal de audio. Mas aun, el aparato comprende una unidad de rastreo del nivel de ruido (130), en donde la unidad de trazado del nivel de ruido (130) está configurada para determinar información del nivel de ruido que depende de por lo menos una de la primera porción de señal de audio y la segunda porción de señal de audio, en donde la información del nivel de ruido se presentada en el dominio de trazado. En forma adicional, el aparato comprende una primera unidad de reconstrucción (140) para reconstruir, en un primer dominio de reconstrucción, una tercera porción de la señal de audio de la señal de audio que depende de la información del nivel de ruido, si un tercer cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho tercer cuadro se recibe a través de la interfaz receptora (110) pero está corrupto, en donde el primer dominio de reconstrucción es diferente de o igual al dominio de rastreo. Mas aun, el aparato comprende una unidad de transformación (121) para transformar la información del nivel de ruido del dominio de trazado a un segundo dominio de reconstrucción, si un cuarto cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho cuarto cuadro se recibe a través de la interfaz receptora (110) pero está corrupto, en donde el segundo dominio de reconstrucción es diferente del dominio de trazado, y en donde el segundo dominio de reconstrucción es diferente del primer dominio de reconstrucción. Mas aun, el aparato comprende una segunda unidad de reconstrucción (141) para reconstruir, en el segundo dominio de reconstrucción, una cuarta porción de señal de audio de la señal de audio que depende de la información del nivel de ruido que sea representada en el segundo dominio de reconstrucción, si la dicho cuarto cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho cuarto cuadro se recibe a través de la interfaz receptora (110) pero está corrupto.

Term
7.7 yearsleft in the term
Expires 23 June 2034.
- Priority
- Filed
- Granted
- Today
- Expires
24 claims: 9 independent, 15 dependent
- 1REIVINDICACIONES Habiendo así especialmente descripto y determinado la naturaleza de la presente invención y la forma como la misma ha de ser llevada a la práctica, se declara reivindicar como de propiedad y derecho exclusivo:1. Un aparato para decodificar una señal de audio, que comprende: una interfaz receptora (110), en donde la interfaz receptora (110) está configurada para recibir un primer cuadro que comprende una primera porción de la señal de audio de la señal de audio, y en donde la interfaz receptora (110) está configurada para recibir un segundo cuadro que comprende una segunda porción de la señal de audio de la señal de audio, una unidad de rastreo del nivel de ruido (130), en donde la unidad de trazado del nivel de ruido (130) está configurada para determinar información del nivel de ruido que depende de por lo menos una de la primera porción de señal de audio y la segunda porción de la señal de audio, en donde la información del nivel de ruido se representa en el dominio de trazado, una primera unidad de reconstrucción (140) para reconstruir, en un primer dominio de reconstrucción, una tercera porción de la señal de audio de la señal de audio que depende de la información del nivel de ruido, si un tercer 170 cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho tercer cuadro se recibe a través de la interfaz receptora (110) pero está corrupto, en donde el primer dominio de reconstrucción es diferente de o igual al dominio de rastreo, una unidad de transformación (121) para transformar la información del nivel de ruido del dominio de trazado a un segundo dominio de reconstrucción, si un cuarto cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho cuarto cuadro se recibe a través de la interfaz receptora (110) pero está corrupto, en donde el segundo dominio de reconstrucción es diferente del dominio de trazado, y en donde el segundo dominio de reconstrucción es diferente del primer dominio de reconstrucción, y una segunda unidad de reconstrucción (141) para reconstruir, en el segundo dominio de reconstrucción, una cuarta porción de la señal de audio de la señal de audio que depende de la información del nivel de ruido que se representa en el segundo dominio de reconstrucción, si dicho cuarto cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho cuarto cuadro se recibe a través de la interfaz receptora (110) pero está corrupto.
- 2Un aparato de acuerdo con la reivindicación 1, 171 en donde el dominio de trazado es un dominio de tiempo, un dominio espectral, un dominio FFT, un dominio MDCT, o un dominio de excitación, en donde el primer dominio de reconstrucción es el dominio de tiempo, el dominio espectral, el dominio FFT, el dominio MDCT, o el dominio de excitación, y en donde el segundo dominio de reconstrucción es el dominio de tiempo, el dominio espectral, el dominio FFT, el dominio MDCT, o el dominio de excitación, pero no el mismo dominio que el primer dominio de reconstrucción.
- 3Un aparato de acuerdo con la reivindicación 2, en donde el dominio de trazado es el dominio FFT, en donde el primer dominio de reconstrucción es el dominio de tiempo, y en donde el segundo dominio de reconstrucción es el dominio de excitación.
- 4Un aparato de acuerdo con la reivindicación 2, en donde el dominio de trazado es el dominio de tiempo, en donde el primer dominio de reconstrucción es el dominio de tiempo, y en donde el segundo dominio de reconstrucción es el dominio de excitación.
- 5Un aparato de acuerdo con una de las reivindicaciones que anteceden, en donde dicha primera porción de la señal de audio se representa en el primer dominio de entrada,, y en donde la 172 segunda porción de la señal de audio que se representa es un segundo dominio de entrada, en donde la unidad de transformación (121) es una segunda unidad de transformación (121), en donde el aparato comprende además una primera unidad de transformación (120) para transformar la segunda porción de la señal de audio o un valor o señal derivado de la segunda porción de la señal de audio del segundo dominio de entrada al dominio de trazado para obtener una segunda información de la porción de señal, en donde la unidad de trazado del nivel de ruido está configurada para recibir una primera información de la porción de señal que se representa en el dominio de trazado, en donde la primera información de la porción de señal depende de la primera porción de la señal de audio, en donde la unidad de trazado del nivel de ruido (130) está configurada para recibir la segunda porción de señal que se representa en el dominio de trazado, y en donde la unidad de trazado del nivel de ruido (130) está configurada para determinar la información del nivel de ruido que depende de la primera información de la porción de señal que se representa en el dominio de trazado y que depende de la información de la segunda porción de señal que se representa en el dominio de trazado. 173
- 6Un aparato de acuerdo con la reivindicación 5, en donde el primer dominio de entrada es un dominio de excitación, y en donde el segundo dominio de entrada es un dominio MDCT.
- 7Un aparato de acuerdo con la reivindicación 5, en donde el primer dominio de entrada es un dominio MDCT, y en donde el segundo dominio de entrada es el dominio MDCT.
- 8Un aparato de acuerdo con una de las reivindicaciones que anteceden, en donde la primera unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio al conducir un primer desvanecimiento a un espectro tipo ruido, en donde la segunda unidad de reconstrucción (141) está configurada para reconstruir la cuarta porción de la señal de audio al conducir un segundo desvanecimiento a un espectro del tipo ruido y/o un segundo desvanecimiento de una ganancia LTP, y en donde la primera unidad de reconstrucción (140) y la segunda unidad de reconstrucción (141) se configuran para conducir el primer desvanecimiento y el segundo desvanecimiento a un espectro del tipo ruido y/o un segundo desvanecimiento de una ganancia LTP con la misma velocidad de desvanecimiento.
- 9Un aparato de acuerdo con una de las reivindicaciones 5 a 8, 174 en donde el aparato comprende además una primera unidad de agregación (150) para determinar un primer valor agregado que depende de la primera porción de la señal de audio, en donde el aparato comprende además una segunda unidad de agregación (160) para determinar, que depende de la segunda porción de la señal de audio, un segundo valor agregado como el valor derivado de la segunda porción de la señal de audio, en donde la unidad de trazado del nivel de ruido (130) está configurada para recibir el primer valor agregado como la primera información de la porción de señal que se representa en el dominio de trazado, en donde la unidad de trazado del nivel de ruido (130) está configurada para recibir el segundo valor agregado como la información de la segunda porción de señal que se representa en el dominio de trazado, y en donde la unidad de trazado del nivel de ruido (130) está configurada para determinar la información del nivel de ruido que depende de el primer valor agregado que se representa en el dominio de trazado y que depende de el segundo valor agregado que se representa en el dominio de trazado.
- 10Un aparato de acuerdo con la reivindicación 9, en donde la primera unidad de agregación (150) está configurada para determinar el primer valor agregado de modo tal que el primer valor agregado indica un cuadrado de 175 promedio de raíz de la primera porción de la señal de audio o de una señal derivada de la primera porción de la señal de audio, y en donde la segunda unidad de agregación (160) está configurada para determinar el segundo valor agregado de modo tal que el segundo valor agregado indica un cuadrado de promedio de raíz de la segunda porción de la señal de audio o de una señal derivada de la segunda porción de la señal de audio.
- 11Un aparato de acuerdo con una de las reivindicaciones 8 a 10, en donde la primera unidad de transformación (120) está configurada para transformar el valor derivado de la segunda porción de la señal de audio del segundo dominio de entrada al dominio de trazado por aplicación de un valor de ganancia en el valor derivado de la segunda porción de la señal de audio.
- 12Un aparato de acuerdo con la reivindicación 11, en donde el valor de ganancia indica una ganancia introducida por Síntesis de codificación predictiva lineal, o en donde el valor de ganancia indica una ganancia introducida por Síntesis de codificación predictiva lineal y des-énfasis.
- 13Un aparato de acuerdo con una de las reivindicaciones que anteceden, en donde la unidad de trazado del nivel de ruido (130) está configurada para determinar la información 176 del nivel de ruido por aplicación de un método de estadística mínima.
- 14Un aparato de acuerdo con una de las reivindicaciones que anteceden, en donde la unidad de trazado del nivel de ruido (130) está configurada para determinar un nivel de ruido de control como la información del nivel de ruido, y en donde la unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio que depende de la información del nivel de ruido, si dicho tercer cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho tercer cuadro se recibe a través de la interfaz receptora (110) pero está corrupto.
- 15Un aparato de acuerdo con la reivindicación 13, en donde la unidad de trazado del nivel de ruido (130) está configurada para determinar un nivel de ruido de control como la información del nivel de ruido que deriva de un espectro del nivel de ruido, en donde dicho espectro del nivel de ruido se obtiene por aplicación del método de estadística mínima, y en donde la unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio que depende de una pluralidad de Coeficientes Predictivos Lineales, si dicho tercer cuadro de la 177 pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho tercer cuadro se recibe a través de la interfaz receptora (110) pero está corrupto.
- 16Un aparato de acuerdo con una de las reivindicaciones que anteceden, en donde la primera unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio que depende de la información del nivel de ruido y que depende de la primera o la segunda porción de la señal de audio, si dicho tercer cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (140) o si dicho tercer cuadro se recibe a través de la interfaz receptora (140) pero está corrupto.
- 17Un aparato de acuerdo con la reivindicación 16, en donde la primera unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio por medio de la atenuación o amplificación de una señal derivada de la primera o la segunda porción de la señal de audio.
- 18Un aparato de acuerdo con una de las reivindicaciones que anteceden, en donde la segunda unidad de reconstrucción (141) está configurada para reconstruir la cuarta porción de la señal de audio que depende de la información del nivel de ruido y que depende de la segunda porción de la señal de audio.
- 19Un aparato de acuerdo con la reivindicación 18, en donde la segunda unidad de reconstrucción (141) está configurada 178 para reconstruir la cuarta porción de la señal de audio por medio de la atenuación o amplificación de una señal derivada de la primera o la segunda porción de la señal de audio.
- 20Un aparato de acuerdo con una de las reivindicaciones que anteceden, en donde el aparato comprende además una unidad de predicción a largo plazo (170) que comprende un búfer de retardo (180), en donde la unidad de predicción a largo plazo (170) está configurada para generar una señal procesada que depende de la primera o la segunda porción de la señal de audio, que depende de una entrada del búfer de retardo que se almacena en el búfer de retardo (180) y que depende de una ganancia de predicción a largo plazo, y en donde la unidad de predicción a largo plazo (170) está configurada para desvanecer la ganancia de predicción a largo plazo hacia cero, si dicho tercer cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho tercer cuadro se recibe a través de la interfaz receptora (110) pero está corrupto.
- 21Un aparato de acuerdo con la reivindicación 20, en donde la unidad de predicción a largo plazo (170) está configurada para desvanecer la ganancia de predicción a largo plazo hacia cero, en donde una velocidad con la cual la ganancia de 179 predicción a largo plazo se desvanece hacia cero depende de un factor de desvanecimiento.
- 22Un aparato de acuerdo con la reivindicación 20 o 21, en donde la unidad de predicción a largo plazo (170) está configurada para actualizar la entrada del búfer de retardo (180) por almacenamiento de la señal procesada generada en el búfer de retardo (180), si dicho tercer cuadro de la pluralidad de cuadros no se recibe a través de la interfaz receptora (110) o si dicho tercer cuadro se recibe a través de la interfaz receptora (110) pero está corrupto.
- 23Un método para decodificar una señal de audio, que comprende:recibir un primer, cuadro que comprende una primera porción de la señal de audio de la señal de audio, y recibir un segundo cuadro que comprende una segunda porción de la señal de audio de la señal de audio, determinar la información del nivel de ruido que depende de por lo menos una de la primera porción de señal de audio y la segunda porción de la señal de audio, en donde la información del nivel de ruido se representa en el dominio de trazado, reconstruir, en un primer dominio de reconstrucción, una tercera porción de la señal de audio de la señal de audio que depende de la información del nivel de ruido, si un tercer cuadro de la pluralidad de cuadros no se recibe o si dicho 180 tercer cuadro se recibe pero está corrupto, en donde el primer dominio de reconstrucción es diferente de o igual al dominio de rastreo, transformar la información del nivel de ruido del dominio de trazado a un segundo dominio de reconstrucción, si un cuarto cuadro de la pluralidad de cuadros no se recibe o si dicho cuarto cuadro se recibe pero está corrupto, en donde el segundo dominio de reconstrucción es diferente del dominio de trazado, y en donde el segundo dominio de reconstrucción es diferente del primer dominio de reconstrucción, y reconstruir, en el segundo dominio de reconstrucción, una cuarta porción de la señal de audio de la señal de audio que depende de la información del nivel de ruido que se representa en el segundo dominio de reconstrucción, si dicho cuarto cuadro de la pluralidad de cuadros no se recibe o si dicho cuarto cuadro se recibe pero está corrupto.
- 24Un programa para computadoras para implementar el método de la reivindicación 23 cuando se ejecuta en una computadora o un procesador de señal.. 181
Independent claims24
980 paragraphs in 8 sections, as filed
(54) Title: APPARATUS AND METHOD FOR IMPROVED SIGNAL FADING IN DIFFERENT DOMAINS DURING HIDING OF ERRORS.
(54) Title: APPARATUS AND METHOD FOR IMPROVED SIGNAL FADE OUT IN DIFFERENT DOMAINS DURING ERROR CONCEALMENT.
(57) Summary
An apparatus is provided for decoding an audio signal. The apparatus comprises a receiving interface (110), wherein the receiving interface (110) is configured to receive a first frame comprising a first portion of the audio signal of the audio signal, and wherein the receiving interface (110) it is configured to receive a second frame comprising a second portion of the audio signal of the audio signal. Furthermore, the apparatus comprises a noise level tracing unit (130), wherein the noise level tracing unit (130) is configured to determine noise level information that depends on at least one of the first audio signal portion and the second audio signal portion, wherein the noise level information is presented in the plotting domain. Additionally, the apparatus comprises a first reconstruction unit (140) for reconstructing, in a first reconstruction domain, a third portion of the audio signal from the audio signal that depends on the noise level information, if a third frame of the plurality of frames is not received through the receiving interface (110) or if said third frame is received through the receiving interface (110) but is corrupted, where the first reconstruction domain is different from or equal to the crawl domain. Furthermore, the apparatus comprises a transformation unit (121) to transform the noise level information from the plotting domain to a second reconstruction domain, if a fourth frame of the plurality of frames is not received through the receiving interface. (110) or if said fourth frame is received through the receiving interface (110) but is corrupt, where the second reconstruction domain is different from the plotting domain, and wherein the second reconstruction domain is different from the first reconstruction domain. Furthermore, the apparatus comprises a second reconstruction unit (141) to reconstruct, in the second reconstruction domain, a fourth audio signal portion of the audio signal that depends on the noise level information that is represented in the second reconstruction domain, if the said fourth frame of the plurality of frames is not received through the receiving interface (110) or if said fourth frame is received through the receiving interface (110) but is corrupted.
(57) Abstract
An apparatus for decoding an audio signal ¡s provided. The apparatus comprises a receiving interface (110), wherein the receiving interface (110) is configured to receive a first trame, comprising a first audio signal portion of the audio signal, and wherein the receiving interface (110) is configured. to receive a second trame comprising a second audio signal portion of the audio signal. Moreover, the apparatus comprises a noise level tracing unit (130), wherein the noise level tracing unit (130) ¡s configured to determine noise level ¡nformation depending on at least one of the first audio signal portion and the second audio signal portion, wherein the noise level ¡nformation ¡s represented ¡na tracing domain. Furthermore, the apparatus comprises a first reconstruction unit (140) for reconstructing, ¡na first reconstruction domain, a third audio signal portion of the audio signal depending on the noise level ¡nformation, ¡fa third trame of the plurality of trames ¡s not received by the receiving ¡nterface (110) or ¡f said third trame ¡s received by the receiving ¡nterface (110) but ¡s corrupted, wherein the first reconstruction domain ¡s different from or equal to the tracing domain. Moreover, the apparatus comprises a transform unit (121) for transforming the noise level ¡nformation from the tracing domain to a second reconstruction domain, ¡fa fourth trame of the plurality of trames ¡s not received by the receiving ¡nterface (110) or ¡F said fourth trame ¡s received by the receiving interface (110) but ¡s corrupted, wherein the second reconstruction domain ¡s different from the tracing domain, and wherein the second reconstruction domain is different from the first reconstruction domain. Moreover, the apparatus comprises a second reconstruction unit (141) for reconstructing, ¡n the second reconstruction domain, a fourth audio signal portion of the audio signal depending on the noise level ¡nformation being represented ¡n the second reconstruction domain, ¡f said fourth trame of the plurality of trames ¡s not received by the receiving ¡nterface (110) or ¡f said fourth trame ¡s received by the receiving ¡nterface (110) but ¡s corrupted.
APPARATUS AND METHOD FOR IMPROVED SIGNAL FADING IN DIFFERENT DOMAINS DURING ERROR HIDING Description
The present invention relates to audio signal encoding, processing and decoding, and in particular to an apparatus and method for improved signal fading for switched audio encoding systems during error concealment.
In the text that follows, the prior art is described with respect to audio and speech codes that fade during packet loss concealment (PLC). Explanations regarding the prior art begin with the G-series ITU-T codes (G.718, G.719, G.722, G.722.1, G.729. G.729.1), they are followed by the 3GPP code (AMR, AMR-WB, AMR-WB +) and an IETF code (OPUS), and conclude with two MPEG code (HE-AAC, HILN) (ITU = International Telecommunication Union; 3GPP = 3rd Generation Partnership Project; AMR = Adaptive Multi-Rate; WB = Wideband; IETF = Internet Engineering Task Force). Later, the technique
<td>previous</td><td>related with</td><td>he</td><td>level plot</td><td>noise</td><td>of</td>
<td>background is</td><td>analyze, followed</td><td>of</td><td colspan="2">a summary that provides</td><td>a</td>
<td colspan="2">generalized view.</td><td></td><td></td><td></td><td></td>
<td>How</td><td>first measure,</td><td>I know</td><td>consider G.718.</td><td>G.718 is</td><td>a</td>
wideband and narrowband voice codec, which supports
DTX / CNG (DTX = Digital Theater Systems; CNG = Comfort Noise
Generation). Since the particular embodiments relate to a low delay code, the low delay version mode will be described in more detail here.
If ACELP (Layer 1) (ACELP = Algebrare Code Excited Linear Prediction) is considered, the ITU-T recommends for G.718 [ITU08a, section 7.11] a flexible fading in the linear predictive domain to control the fading rate. In general, concealment follows this principle:
According to G.718, in the case of frame corrections, the concealment strategy can be summarized as a convergence of the signal energy and the spectral envelope for the estimated parameters of the background noise. The periodicity of the signal converges to zero. The speed of convergence is dependent on the parameters of the last correctly received frame and the number of consecutive frames erased, and is controlled by an attenuation factor, a. The attenuation factor a, additionally depends on the stability, Θ, of the LP filter (LP = Linear Prediction) for NON-VOICE frames. In general, convergence is slow if the last good frame received is in a stable segment and fast if the frame is in a transition segment.
The attenuation factor a depends on the voice signal class, which is derived from the signal classification described in [ITU08a, section 6.8.1.3.1 and 7.11.1.1]. The stability factor Θ is computed based on the measurement of distance between the filters of the adjacent ISF (Immittance Spectral Frequency) [ITU08a, section 7.1.2.4.2].
Table 1 shows the calculation scheme for a:
<td>last good picture received</td><td>Number of successive frames removed</td><td>I heard</td>
<td>ARTIFICIAL START</td><td></td><td> 0, 6</td>
<td>START UP, WITH VOICE</td><td> £ 3</td><td>O t — 1</td>
<td></td><td> > 3</td><td> 0,4</td>
<td>TRANSITION WITH VOICE</td><td></td><td> 0,4</td>
<td>VOICELESS TRANSITION</td><td></td><td> 0,8</td>
<td>WITHOUT VOICE</td><td> = 1</td><td>0.2 * Θ + 0.8</td>
<td></td><td> = 2</td><td> 0,6</td>
<td></td><td> > 2</td><td> 0,4</td>
Table 1: Values of the attenuation factor a, the value Θ is a stability factor computed from the distance measurement between the adjacent LP filters. [ITU08a, section 7.1.2.4.2].
Furthermore, G.718 provides a fading method that is intended to modify the spectral envelope. The general idea is to converge the latest ISF parameters towards a flexible ISP average vector. As a first measure, an average ISF vector is calculated from the last 3 known ISF vectors. The average ISF vector is then averaged again with an offline trained long-term ISF vector (which is a constant vector) [ITU08a, section 7.11.1.2].
Furthermore, G.718 provides a fading method to control long-term behavior and thus interaction with background noise, where the pitch excitation energy (and thus the periodicity of the excitation) is convergent. at 0, while the random excitation energy is convergent to the CNG excitation energy [ITU08a, section 7.11.1.6]. The innovation gain attenuation is calculated as
4<sup>1]</sup> = «4<sup>01</sup> + (<sup>1</sup> - <sub>(1)</sub> where g ^ is the innovation gain at the beginning of the next table, g '<sup>01</sup> is the innovation gain at the beginning of the current frame, g<sub>n</sub> is the gain of the excitation used during the generation of the comfort noise and the attenuation factor a.
Similar to periodic excitation attenuation, gain attenuates linearly across the frame on a sample-by-sample basis starting with, g [<sup>0] </sup>, and reaches gp<sup>1</sup> at the beginning of the next frame.
<td>The</td><td>Fig. 2 defines</td><td>the</td><td colspan="2">structure of</td><td>decoder</td><td>of</td>
<td>G.718.</td><td>In particular,</td><td>the</td><td>Fig.</td><td>2 illustrates</td><td>a structure</td><td>of</td>
<td colspan="3">level decoder</td><td>tall</td><td>by G.718</td><td>for PLC,</td><td>than</td>
characterizes a high pass filter.
Through the previously described method of G.718, the innovative gain g<sub>s</sub> converges with the gain used during comfort noise generation g<sub>n</sub> during long bursts of packet loss. As described in [ITU08a, section 6.12.3], the comfort noise gain g<sub>n </sub>is given as the square root of the energy E. The conditions of the E update are not described in detail. If the reference implementation is followed (floating point C code, stat_noise_uv_mod.c), E is obtained as follows:
yes (unvoiced_vad == 0) {yes (unv_cnt> 20) {ftmp = lp_gainc * lp_gainc;
lp_ener = 0.7f * lp_ener + 0.3f * ftmp;
} else {unv_cnt ++;
} }
else {unv_cnt = 0;
} where unvoiced_vad retains detection of voice activity, where unv_cnt refers to the number of frames without voice in a row, where lp_gainc refers to the fixed codec book low pass increases, and where lp_ener refers to the low pass CNG energy estimate E, it is initialized with 0.
Additionally, G.718 provides a high-pass filter, introduced in the signal path of the no-voice drive, if the signal of the last good frame was classified different from NO VOICE, see Fig. 2, see also [ITU08a , section 7.11.1.6]. This filter has a low storage characteristic with a frequency response at DC that is around 5 dB less than at the Nyquist frequency.
Furthermore, G.718 proposes a decoupled LTP feedback loop (LTP = Long-Term Prediction): While during normal operation the feedback loop for the flexible codebook is updated in the form of sub-frames ([ITU08a, section 7.1.2.1.4]) based on full excitation. During cloaking this feedback loop is updated in the form of frames (see [ITU08a, sections 7.11.1.4, 7.11.2.4, 7.11.1.6, 7.11.2.6; dec_GV_exc@dec_gen_voic.c and syn_bfi_post@syn_bfi_pre_post.c]) based on excitement with voice only. With this approach, the flexible codebook is not contaminated with noise originating from the randomly chosen innovation drive.
With respect to G.718 transform encoded enhancement layers (3-5), during masking, the decoder behaves with respect to high layer decoding similar to normal operation, so that the MDCT spectrum is fixed at zero. No special fading behavior is applied during concealment.
With respect to CNG, in G.718, the synthesis of CNG is done in the following order. As a first measure, the parameters of a comfort noise frame are decoded. Then a frame of comfort noise is synthesized. Thereafter the tone buffer is reset. Then the synthesis for the FER (Frame Error Recovery) classification is saved. Thereafter, spectrum de-emphasis is conducted. Then low-frequency post-filtering is conducted. Then the CNG variables are updated.
In the case of concealment, the exact same thing is done, except that the CNG parameters are not decoded from the bitstream. This means that the parameters are not updated during frame loss, but the decoded parameters from the last good SID frame (Silence Insertion Descriptor) are used.
<td>Now,</td><td>I know</td><td>consider G.719. G</td><td>.719, which</td><td>I know</td><td>based on</td><td>a</td>
<td>Siren 22,</td><td>it is</td><td colspan="2">a full audio codec</td><td>bas</td><td>adored in</td><td>the</td>
<td colspan="2">transformation</td><td colspan="3">The ITU-T recommends for</td><td>G.719</td><td>a</td>
<td colspan="3">fading with repeat</td><td>squared</td><td>in</td><td colspan="2">The Dominion</td>
<td>spectral</td><td colspan="2">[ITU08b, section 8.6].</td><td>Agree</td><td>with</td><td>G.719,</td><td>a</td>
<td>mechanism</td><td>of</td><td>concealment by</td><td>elimination</td><td>of</td><td>picture</td><td>I know</td>
<td>incorporates</td><td>in</td><td>the decoder.</td><td colspan="2">When a painting</td><td colspan="2">is received</td>
correctly, the reconstructed transformation coefficients are stored in a buffer. If the decoder receives the report that a frame has been lost or a frame is corrupted, the reconstructed transform coefficients on the most recently received frame are scaled down by a factor of 0.5 and then used as the coefficients of reconstructed transformation for frame in progress. The decoder proceeds by its transformation in the time domain and the window-overlap-aggregate operation is performed.
In the text that follows, G.722 is described.
a coding system from 50 to 7000
Hz using sub-band flexible differential pulse code modulation (SB-ADPCM) within a transfer rate of up to kbit / s.
The signal is separated into an upper and a lower sub-band, with the use of QMF analysis (QMF
Quadrature Mirror Filter
The two resulting bands are encoded by ADPCM (ADPCM =
Adaptive Differential Pulse Code Modulation (Flexible Differential Pulse Code Modulation).
For G.722, a high complexity algorithm for packet loss concealment is specified in Appendix III [ITU06a] and a low complexity algorithm for packet loss concealment is specified in Appendix IV [ITU07]. G.722 - Appendix III ([ITU06a, section
III. 5]) proposes a gradual muting, starting after 20 ms of frame loss, completing after 60 ms of frame loss. Furthermore, G.722 - Appendix IV proposes a fading technique that applies a magnification factor to each sample that is computed and adapted sample by sample [ITU07, section
IV. 6.1.2.7].
In G.722, the squelch process takes place in the sub-band domain immediately before QMF synthesis and as the last step of the PLC module. The calculation of the squelch factor is done with the use of class information from the signal classifier which is also part of the PLC module. The distinction is made between TRANSIENT, UV_TRANSITION and other classes. Additionally, a distinction is made between single 10 ms frame losses and other cases (multiple 10 ms frame losses and single / multiple 20 ms frame losses).
This is illustrated through Fig. 3. In particular, Fig. 3 presents a. scenario, where the G.722 fading factor, depends on the class information and where 80 samples are equivalent to 10 ms.
According to G.722, the PLC module creates the signal for the missing frame and some additional signal (10 ms) that supposedly crossfades with the next good frame. Squelching for this additional signal follows the same rules. In G.722 high-band concealment, crossfade does not take place.
In the text that follows, G.722.1 is considered. G.722.1, which is based on a Siren 7, is a wideband audio codec based transformation with a super wideband extension mode, referred to as G.722.1CG 722.1C itself is based a Siren 14. The ITU-T recommends for G.722.1 a frame repeat with post squelch [ITU05, section 4.7]. If the decoder receives information, through an external signaling mechanism not defined in this recommendation, that a frame has been lost or corrupted, it repeats the coefficients of the decoded MLD of the previous frame (Modulated Lapped Transform). It proceeds by its transformation in the time domain, and performs the superposition and aggregate operation with the decoded information from the following table. If the previous frame was also lost or corrupted, then the decoder sets all current frame MLT coefficients to zero.
Now, it is considered G.729. G.729 is an audio data compression algorithm for speech that compresses digital speech into packets of 10 milliseconds in length. This is officially described as 8 kbit / s Speech Coding with the use of Code Excited Linear Prediction Speech Coding (CS-ACELP) [ITU12].
As stated in [CPK08], G.729 recommends fading in the LP domain. The PLC algorithm employed in the G.729 standard reconstructs the speech signal for the current frame based on the previously received speech information. In other words, the PLC algorithm replaces the missing excitation with an equivalent characteristic of a previously received frame, through the excitation energy finally gradually decaying, the increases of the fixed and flexible codebooks are attenuated by a constant factor .
The attenuated fixed code log is given by:
with m being the sub-frame index.
The flexible codebook gain is based on an attenuated version of the previous flexible codebook gain:
= 0.9 ·, joined by g ™ <0.9
Nam in Park et al. suggest for G.729, a signal amplitude control using prediction by means of linear regression [CPK08, PKJ + 11]. It is cited to drive packet loss and its linear regression as a kernel technique. Linear regression is based on the linear model as<sub>9</sub>i = a + bi <sub>(2)</sub> where g, is the newly predicted current amplitude, a and b are coefficients for the first-order linear function, and i is the index in the table. In order to find the optimized coefficients a * and b *, the sum of the squared prediction error is minimized:
<sup>and</sup> = Σ j = i-4. <sub>(3)</sub> ε is the square error, g¿ is the original passed j-th amplitude. To minimize this error, simply the derivative related to a and b is set to zero. Using the optimized parameters a * and b *, an estimate of each g * is denoted by g * = a * + b * i<sub>(4)</sub>
Fig. 4 shows the amplitude prediction, in particular the g * amplitude prediction, by using linear regression.
To obtain the amplitude A¡ of the lost packet i, a proportion σ, (5) <7i =
9i-i is multiplied with a scale factor Si:
A'i = Si * σι (6) where the scale factor Si depends on the number of consecutive hidden frames 2 (i):
<td rowspan="2"></td><td>í io, 0.9.</td><td>if Z (i) = 1,2 if Z (0 = 3.4</td>
<td>I 0-8, (one</td><td>if l (i) = 5.6 'else</td>
(7)
In [PKJ + 11], a slightly different magnification is proposed.
According to G.729, thereafter A- will be smoothed to avoid individual fading at the edges of the frame. The final smoothed amplitude ^ (n) is multiplied by the excitation, obtained from the previous PLC components.
In the text that follows, G.729.1 is considered. G.729.1 is an embedded variable rate encoder based on G.729: A scalable wideband encoder bitstream from 8-32 kbit / s interoperable with G.729 [ITU06b].
According to G.729.1, as in G.718 (see above), a flexible fading is proposed, which depends on the stability of the signal characteristics ([ITU06b, section
7.6.1]). During the bumping, the signal usually attenuates based on an attenuation factor a which depends on the class parameters of the last good frame received and the number of consecutive erased frames. The attenuation factor a is additionally dependent on stability. LP filter for VOICELESS frames. In general, fading is slow if the last good frame received is in a stable segment and fast if the frame is in a transition segment.
Additionally, the attenuation factor a depends on the average tone gain per subframe g<sub>p</sub> ([ITU06b, eq. 163, 164]):
= 0.14°’ + 0.2,/),<sup>11</sup> + + 0.4<sub>S</sub><<sup>3</sup>'where g ^ is the tone gain in sub-frame i.
Table 2 shows the calculation scheme of a, where β = with 0.85> β> 0.98
During the hiding process, a is used in the following hiding tools:
<td>last good picture received</td><td>Number of successive frames eliminated</td><td>to</td>
<td>WITH VOICE</td><td> 1</td><td>β</td>
<td></td><td> 2,3</td><td>g<sub>P</sub></td>
<td></td><td> > 3</td><td> 0,4</td>
<td>START</td><td> 1</td><td>0.8 β</td>
<td></td><td> 2,3 > 3</td><td> 8<sub>P</sub> 0,4</td>
<td>ARTIFICIAL START</td><td> 1</td><td>0.6 β</td>
<td></td><td> 2,3</td><td> 8<sub>P</sub></td>
<td></td><td> > 3</td><td> 0,4</td>
<td>TRANSITION WITH VOICE</td><td> < 2</td><td> 0,8</td>
<td></td><td> > 2</td><td> 0,2</td>
<td>VOICELESS TRANSITION</td><td></td><td> 0,88</td>
<td>WITHOUT VOICE</td><td> 1</td><td> 0,95</td>
<td></td><td> 2.3</td><td>0.6 Θ + 0.4</td>
<td></td><td> > 3</td><td> 0,4</td>
Table 2: Values of the attenuation factor a, the value Θ is a stability factor computed from the measurement of
<td>distance between</td><td>the</td><td>filters</td><td>by LP</td><td>adjacent.</td><td>[ITU06b,</td>
<td>section 7.6.1].</td><td></td><td></td><td></td><td></td><td></td>
<td>Agree</td><td>with</td><td>G. 729.1,</td><td>with</td><td>about</td><td>the re</td>
<td>synchronization of</td><td>pulse</td><td>glottal,</td><td>as</td><td colspan="2">the last pulse of the</td>
The excitation of the above frame is used for the construction of the periodic part, its magnification is approximately correct at the beginning of the hidden frame and can be set to 1. The gain is then linearly attenuated throughout the entire frame on a sample basis by sample to achieve the value of a at the end of the table. The energy evolution of voiced segments is extrapolated using the excitation increase values per tone of each sub-frame of the last good frame. In general, if these increases are greater than 1, the signal energy is increasing, if they are greater than 1, the energy is decreasing, a is thus set at β = ^ ξ<sub>ρ</sub> as previously described, see [ITU06b, eq. 163, 164]. The value of β is fixed between 0.98 and 0.85 to avoid strong increases and decreases in energy, see [ITU06b, section 7.6.4].
With regard to the construction of the random part of the excitation, according to G. 729.1, at the beginning of a removed block, the innovation gain g, is initialized by using the innovation excitation increases of each sub -Frame of the last good frame:
g<sub>s</sub> = 0.1^<sup>or)</sup> + 0.2 yes?<sup>(1)</sup> + 0.3ir<sup>(2)</sup> + 0.4<sub>5</sub><sup>(3)</sup> where gr<sup>(0)</sup>, g<sup>(1></sup>, gr<sup>(2)</sup> and g<sup>(3)</sup> are the increases of the fixed code log, or innovation, of the four sub-frames of the last frame received correctly. The attenuation of the innovation gain is done as:
<img file="MX2015018024A_D0001.tif" />
<* - g (0) s where is the innovation gain at the beginning of the next frame, is the innovation gain at the beginning of the current frame, and ori is as defined in Table 2 above. Similar to the periodic excitation attenuation, the gain is thereby attenuated linearly throughout the frame on a sample-by-sample basis beginning with and approaching the value that could be achieved at the beginning of the next frame.
According to G.729.1, if the last good frame is NO VOICE, only the innovation drive is used and is further attenuated by a factor of 0.8. In this case, the past drive buffer is updated with the innovation drive since no periodic part of the drive is available, see [ITU06b, section 7.6.6].
In the text that follows, it is considered AMR. 3GPP AMR [3GP12b] is a speech codec that uses the ACELP algorithm. AMR is capable of encoding speech with a sampling rate of 8000 samples / s and a transfer rate between 4.75 and 12.2 kbit / s and supports signaling of description frames of silence (DTX / CNG).
In AMR, during error concealment (see [3GP12a]), a distinction is made between frames that are error prone ('bit errors) and frames, which are completely lost (no data at all).
For ACELP concealment, AMR introduces a state machine that estimates the channel quality: The higher the value of the state counter, the worse the channel quality. The system boots to state 0. Every time a bad frame is detected, the status counter is incremented by one and saturates when it reaches 6. Every time a good voice frame is detected, the status counter is reset to zero, except when the state is 6, where the state counter is set to 5. The flow of control of the state machine can be described by the following C code (BFI is a bad frame indicator, State is a state variable):
if (BFI! = 0) {
State = State + 1;
} or if (State == 6) {
State = 5;
) O well {
State = 0;
} if (State> 6) {
State = 6;
}
In addition to this state machine, in AMR, the indicators of a bad frame of the current and previous frames are monitored (prevBFI).
Three different combinations are possible:
The first of the three combinations is BFI = 0, prevBFI = 0, Status = 0: No error detected in previous received or received voice frame. The received speech parameters are used in normal mode in speech synthesis. The voice parameters of the current frame are saved.
The second of the three combinations is BFI = 0, prevBFI = 1, Status = 0, or 5: No error detected in received voice frame, but previous received voice frame was bad.
The LTP gain and fixed codebook gain are limited below the values used for the last good subframe received:
9p -
9p>
9p <g<sub>p</sub>(-l) 9p> fl'pC<sup>-</sup>'!) (10)
<td>where g<sub>p</sub></td><td>= increase in</td><td>Decoded LTP</td><td>current gr<sub>p</sub>(-l</td>
<td>LTP gain</td><td>used for the</td><td>last sub-frame</td><td>Okay</td>
<td>(BFI = 0)</td><td>, and</td><td></td><td></td>
<td></td><td>n)<sup>9c</sup>'</td><td>9c <g<sub>c</sub>{- ^ 9c> 9c {- ^</td><td> (11)</td>
where g<sub>c</sub> = current decoded gain of the fixed codebook, and g<sub>c</sub>(-l) = fixed codebook increase used for last good subframe (BFI = 0).
The rest of the received speech parameters are normally used in speech synthesis. The voice parameters of the current frame are saved.
The third of the three combinations is BFI = 1, prevBFI = 0 or 1, Status = 1 ... 6: An error is detected in the received speech box and the substitution and muting procedure is started. The LTP gain and fixed codebook rise are replaced by attenuated values from the previous sub-tables:
9ρ -
Mat) (? Ρ (-1), g<sub>p</sub>(—1) <median 5><sub>p</sub>(-l) ...., g<sub>p</sub>(-b))
P (state) -median 5 (, 9<sub>Ρ</sub>(-1), · · ·, .9<sub>P</sub>(~ S)) 9p (~ 1)> median 5 (.9p (~ 1) · ··, 9p (~ ty) (12) where g<sub>p</sub> indicates current decoded LTP gain and g<sub>p</sub>(-l),. . g<sub>p</sub>(-n) indicate uses of the gain of LTPs for the last n sub-frames and median5 () indicates a median operation of 5 points and
P (state) = attenuation factor, where (P (l) = 0.98, P (2) = 0.98, P (3) = 0.8, P (4) = 0.3,
P (5) = 0.2, P (6) = 0.2) and state - state number, and
9c = <P '(state) 9 c (<sup>—</sup>1) ·
C (state) median 5 (g<sub>c</sub>(—1), .... g<sub>c</sub>[—5))
9c {~ 1) <· median 5 (gc (~ 1) · · ·. 9c {—5))
9cÁ-1)> median 5 (.% (-!), .... <7<sub>c</sub>(-5)) (13) where g<sub>c</sub> indicates the current decoded gain of the fixed codebook and g<sub>c</sub>(-l), ..., g<sub>c</sub> (-n) indicate the gain of the fixed code log used for the last n sub-frames and median5 () indicates a 5-point median operation and C (state) = attenuation factor, where (C (l) = 0, 98, C (2) = 0.98, C (3) = 0.98, C (4) = 0.98, C (5) = 0.98, C (6) = 0.7) and state = state number.
In AMR, the LTP-lag values (LTP = Long-Term Prediction) are replaced by the past value of 4<sup>to</sup> sub-table of the previous table (12.2 mode) or slightly modified values based on the last value received correctly (all other modes).
According to AMR, the innovation pulses from the fixed code log received from the erroneous frame are used in the state in which they were received when corrupted data is received. In the case where no data was received, random fixed code log indexes should be used.
With regard to CNG in AMR, according to [3GP12a, section 6.4], each SID frame lost first is replaced by using the SID information from valid SID frames received before and the procedure for valid SID frames is applied. For later lost SID frames, an attenuation technique is applied to the comfort noise that will gradually reduce the output level. Therefore it is controlled if the last SID update was more than 50 frames (= 1 s) before, if yes, the output will be muted (level attenuation at a rate of -6/8 dB per frame [3GP12d, dtx_dec[-lex.europa.eu@sp_dec.c] which produces 37.5 dB per second). It should be noted that the CNG applied fading is performed in the LP domain.
In the text that follows, it is considered AMR-WB. Adaptive Multirate - WB [ITU03, 3GP09c] is a voice codec, ACELP, based on AMR (see section 1.8). It uses parametric bandwidth extension and also supports DTX / CNG. In the description of the [3GP12g] standard there are exemplary concealment solutions given which are the same as for AMR [3GP12a] with minor deviations. Therefore, only the differences with AMR are described here. For the standard description, see the description above.
With respect to ACELP, in AMR-WB, ACELP fading is performed based on the reference source code [3GP12c] by modifying the tone gain g<sub>p</sub> (for the above AMR referred to as LTP gain) and by modifying the g-code gain<sub>c</sub>.
In the case of the lost frame, the tone gain g<sub>p </sub>for the first sub-frame it is the same as in the last good frame, except that it is limited between 0.95 and 0.5. For the second, third, and subsequent subframes, the tone gain g<sub>p</sub> it is reduced by a factor of 0.95 and limited again.
AMR-WB proposes that in a hidden box, g<sub>c</sub> is based on the last g<sub>c</sub>:
c, current - 9c. past * (1.4 (/ ρ. past (14) c 9 c, current * 9 c (15)
<img file="MX2015018024A_D0002.tif" />
<img file="MX2015018024A_D0003.tif" />
____<sup>βη & Γ</sup>1ηον ____ subframe _size
1.0 (16) í'iiet í<sub>nm</sub>>
sub-table size-1 code '[í
<img file="MX2015018024A_D0004.tif" />
(17)
To hide the LTP lags, in AMR-WB, the history of the last five good LTP lags and LTP increases are used to find the best method to update, in the event of a lost frame. In the case where the frame is received with bit errors, a prediction is made, whether the received LTP lag can be used or not [3GP12g].
Regarding CNG, in AMR-WB, if the last frame successfully received was a SID frame and a frame is classified as lost, it will be replaced with the information from the last valid SID frame and the procedure for valid SID frames should be applied.
For later lost SID frames, AMR-WB proposes to apply a comfort noise attenuation technique that will gradually reduce the output level. Therefore it is controlled if the last SID update was more than 50 frames (= 1 s) before, if yes, the output will be muted (level attenuation at the rate of -3/8 dB per frame [3GP12f, dtx_dec {}@dtx.c] which produces 18.75 dB per second). It should be noted that the fading applied to CNG is done in the LP domain.
Now, it is considered AMR-WB +. Adaptive Multirate - WB + [3GP09a] is a switched code using ACELP and TCX (TCX = Transform Coded Excitation).
Transformation)) as central codes. It uses parametric bandwidth extension and also supports DTX / CNG.
In AMR-WB +, a mode extrapolation logic is applied to extrapolate the modes from lost frames within a distorted subframe. This extrapolation of modes is based on the fact that there is redundancy in the definition of mode indicators. The decision logic (given in [3GP09a, figure 18]) proposed by AMR-WB + is the following:
A vector mode, (m_i, mo, mi, m<sub>2</sub>, m<sub>3</sub>), is defined, where m_i indicates the mode of the last frame of the previous super frame and m<sub>0</sub>my m<sub>2</sub>, m<sub>3</sub> indicate the modes of the frames in the current super frame, (decoded from the bit stream), where m<sub>k</sub> = -1, 0, 1, 2 or 3 (-1: lost, 0: ACELP, 1: TCX20, 2: TCX40, 3: TCX80), and where the number of lost frames nloss can be between 0 and 4.
If m_i = 3 and two of the mode indicators in tables 0-3 are equal to three, all indicators will be set to three because it is then ensured that a TCX80 table was indicated within the super table.
If only one indicator in frames 0 - 3 is three (and the number of lost frames nloss is three), the mode will be fixed at (1, 1, 1, 1), because then 3/4 of the TCX80's weighted spectrum is lose and the overall TCX profit is very likely to be lost.
If the mode is indicative (x, 2, -1, x, x) or (x, -l,
2, x, x), it will extrapolate to (x, 2, 2, x, x), indicating a TCX40 frame. If the mode indicates (x, x, x, 2, -1) or (x, x, -l, 2) it will extrapolate to (x, x, x, 2, 2), which also indicates a TCX40 frame. It should be noted that (x, [0, 1], 2, 2, [0, 1]) are invalid configurations.
After that, for each frame that is lost (mode = -1), the mode is set to ACELP (mode = 0) if the preceding frame was ACELP and the mode is set to TCX20 (mode = 1) for all others cases.
With respect to ACELP, according to AMR-WB +, if a dropped frame mode produces m<sub>k</sub> = 0 after mode extrapolation, the same method as in [3GP12g] is applied for this table (see above).
In AMR-WB +, depending on the number of lost frames and the extrapolated mode, the following TCX-related concealment methods are distinguished (TCX = Transform Coded Excitation):
If an entire frame is lost, then ACELP-like concealment is applied: The last excitation is repeated and the hidden ISF coefficients (slightly shifted towards their flexible average) are used to synthesize the time domain signal. Additionally, a fading factor of 0.7 per frame (20 ms) [3GP09b, dec_tcx.c] is multiplied in the linear predictive domain, immediately before the LPC (Linear Predictive Coding) synthesis.
If the last mode was
TCX80 as well as sub-frame extrapolated mode (partially lost) is
TCX80 (nloss [1, 2], mode (3, 3, 3, concealment is extrapolation from last frame extrapolation of interest here details, see modification [3GP09a, previous:
current:
performed in the FFT domain, which uses amplitude and phase, having received correctly. The phase information method is not considered at all (none and therefore [3GP09a, relation with the strategy of both is not described. For further section 6.5.1.2.4]. Regarding the amplitude of AMR- WB +, the method performed section 6.5.1.2.3]:
I know
I know
The spectral computes:
The following steps compute the spectrum of compute the spectrum of magnitude of magnitude of the box box energy increase difference of no lost between the previous box coefficients and current ones. _ / V increase - U and%<sub>ldA [k] 2</sub>
The amplitude of the missing spectral coefficients is extrapolated with the use of:
if {^ [k] lost) A [fc] = increase.old Λ [Á ']
In any other case of a missing frame with mk = [2, 3], the TCX objective (inverse FFT of the decoded spectrum plus noise induction (using a noise level decoded from the bit stream)) is synthesized with the use of all available info (including global TCX increase). No fading is applied in this case.
With respect to CNG in AMR-WB +, the same method is used as in AMR-WB (see above).
In the text that follows, it is considered OPUS. OPUS [IET12] incorporates two-code technology: the voice-oriented SILK (known as the Skype codec) and low-latency CELT (CELT = Constrained-Energy Lapped Transform). Opus can be optimally tuned between high and low transfer rates, and internally switches between a linear prediction codec at lower transfer rates (SILK) and a transform code at higher transfer rates (CELT) as well as well as a hybrid for a short overlay.
Regarding compression and decompression of SILK audio data, in OPUS, there are several parameters that are attenuated during concealment in the SILK decoder routine. The gain of the LTP parameter is attenuated by multiplying all the LPC coefficients with either 0.99, 0.95, or 0.90 per frame, depending on the number of consecutive lost frames, where excitation is built with the use of the last tone cycle from the excitation of the previous frame. The tone lag parameter increases very slowly during consecutive losses. For simple losses it remains constant when compared to the last frame. Furthermore, the excitation increase parameter attenuates exponentially with 0.99<sup>/ ftS / c</sup><sup>z</sup> per frame, such that the drive-up parameter is 0.99 for the first drive-up parameter, such that the drive-up parameter is 0.992 for the second drive-up parameter, and so on. Excitation is generated with the use of a random number generator that generates variable overflow weighted noise. Additionally, the LPC coefficients that are extrapolated / averaged based on the last set of correctly received coefficients. After generating the attenuated excitation vector, the hidden LPC coefficients are used in OPUS to synthesize the time domain output signal.
Now, in the context of OPUS, it is considered CELT. CELT is a transform-based codec. CELT concealment features a tone-based PLC method, which is applied for five consecutive missed frames. Starting with Table 6, a noise masking method is applied, generating background noise, the characteristic of which is supposed to sound like the preceding background noise.
Fig. 5 illustrates a CELT burst loss behavior. In particular, Fig. 5 presents a spectrogram (x-axis: time; y-axis: frequency) of a voice segment hidden with CELT. The green light box indicates the first 5 lost frames consecutively, where the tone-based PLC method is applied. Beyond that, noise is shown as concealment. It should be noted that the switching is done instantaneously, it does not pass smoothly.
With respect to pitch-based concealment, in OPUS, pitch-based concealment consists of finding the periodicity in the decoded signal by auto-correlation and repetition of the window waveform (in the excitation domain with the use of LPC analysis and synthesis) with the use of tone compensation (tone lag). The wavelength of the window is superimposed in such a way that the cancellation with time domain bias with the previous frame and the next frame is preserved [IET12]. Additionally a fade factor is divided is applied by the following code:
opus_val32 El = l, E2 = l;
int period;
if (tone_index <= MAX_PERIOD / 2) {period = tone_index;
either {period = MAX_PERIOD / 2;
The + = exc [MAX_PERIOD- period + i] 'k exc [MAX PERIODperiod + i];
E2 + = exc [MAX_PERIOD-2 * period + i] exc [MAX PERIOD2 * period + i]}
yes (He> E2)
El = E2;
} fall = sqrt (E1 / E2)); attenuation = fall;
In this code, exc contains the drive signal up to samples of MAX_PERIOD before loss.
The excitation signal is later multiplied with attenuation, then synthesized and generated by LPC synthesis.
The fading algorithm for the time domain method can be summarized as follows:
Find the synchronous energy of the tone of the last tone cycle before the loss.
Find the synchronous energy of the tone of the second last tone cycle before the loss.
If the energy is increasing, limit it to remain constant: attenuation = 1
If the energy is decreasing, continue with the same attenuation during concealment.
Regarding noise concealment, according to OPUS, for the 6<sup>cough</sup> and following consecutive lost frames a noise substitution method in the MDCT domain is performed, in order to stimulate comfortable background noise.
Regarding background level and shape noise tracking, in OPUS, background noise estimation is performed as follows: After MDCT analysis, square root of MDCT by energies is calculated by frequency band, where grouping of the MDCT containers follow the bark scale according to [IET12, Table 55]. Then the square root of the energies becomes the domain log<sub>2 </sub>by:
bandLogE [i] = · log<sub>and</sub>(bandE [i] - eMeans [¿]) fpara í = 0 ... 21 where e is the 'Euler number, bandE is the square root of the MDCT and eMeans is a vector of constants (necessary to maintain the average result to zero, which produces an enhanced encoding gain).
In OPUS, the background noise is loaded on the decoder side like this [IET12, amp2Log2 and log2Amp @ quant_bands.c]:
backgroundLogE [i] = min {backgroundLogE [i \ +8 · 0.001, bandLogE \ i]} for i = 0 ... 21 (19)
The minimum energy tracked is basically determined by the square root of the energy that is per current frame, but the gain from one frame to the next is limited by 0.05 dB.
Regarding the application of the background level and shape noise, according to OPUS, if the noise is applied as PLC, backgroundLogE as drift in the last good frame is used and it is converted back to the linear domain:
bandE [i] = <sub>and</sub>(log ^) <b<sub>to</sub>ckgroundLogE [i]<sub>+ e</sub>M<sub>ea</sub>n<sub>S</sub>[i])) <sub>=</sub> (). . . 21 (20) where e is the Euler number and eMeans is the same vector of constants as for the linear transformation to log.
The current masking procedure is to fill the MDCT frame with weighted noise produced by a random number generator, and increase this weighted noise in such a way that it matches the energy of bandE. Subsequently, the reverse MDCT is applied which produces a time domain signal. After the addition of overlap and de-emphasis (as in regular decoding ) it is discarded.
In the text that follows, it is considered MPEG-4 HE-AAC (MPEG = Moving Picture Experts Group; HE-AAC = High Efficiency Advanced Audio Coding). High Efficiency Advanced Audio Coding consists of a transform-based audio codec (AAC), supplemented by a parametric bandwidth extension (SBR).
With respect to AAC (AAC = Advanced Audio Coding), the DAB consortium specifies for AAC in DAB +, a fading to zero in the frequency domain [EBU10, section Al.2] (DAB = Digital Audio Broadcastingj. The fading behavior, For example, the attenuation ramp can be fixed or can be adjusted by the user. The spectral coefficients of the last AU (AU = Access Unit) are attenuated by a factor corresponding to the fading characteristics and then it is passed to frequency mapping in time. Depending on the attenuation ramp, the concealment changes to silence after a number of consecutive invalid AUs, which means that the entire spectrum will be fixed at 0.
The DRM consortium (DRM = Digital Rights Management) specifies for AAC and DRM a fading in the frequency domain [EBU12, section 5.3.3]. Concealment works on the spectral data immediately before the final frequency to time conversion. If multiple frames are corrupted, concealment first implements a fading based on slightly modified spectral values from the last valid frame. Furthermore, similar to DAB +, the fading behavior, eg. , the attenuation ramp, can be fixed or can be adjusted by the user. The spectral coefficients in the last frame are attenuated by a factor corresponding to the fading characteristics and then passed to frequency mapping in time. Depending on the attenuation ramp, the concealment changes to silence after a number of consecutive invalid frames, which means the entire spectrum will be fixed at 0.
3GPP introduces for AAC in Enhanced aacPlus fade in. DRM-like frequency domain [3GP12e, section 5.1] · Concealment works on spectral data immediately before final frequency conversion to time. If multiple frames are corrupted, concealment first implements a fading based on slightly changed spectral values since the last good frame. A complete fade unfolds in 5 frames. The spectral coefficients since the last good frame are copied and attenuated by a factor of:
fadeOutFac = 2- <<sup>nPadeOwíFrowe</sup>/<sup>2</sup>) with nFadeOutFrame as the frame counter since the last good frame.
After five fading frames the concealment changes to silent, which means that the full spectrum will be fixed at 0.
Lauber
Sperschneider presents for
AAC a frame fade of the MDCT spectrum, based on energy extrapolation [LS01, section 4.4].
The energy shapes of a preceding spectrum could be used to extrapolate the shape of an estimated spectrum. Energy extrapolation can be done independent of concealment techniques as a post-concealment class.
With regard to AAC, the energy calculation is done on a scale factor per base with the goal of being close to the critical bands of the human auditory system. Individual power values are lowered on a frame-by-frame basis in order to smoothly reduce the volume, eg to fade the signal. This becomes necessary since the probability, that the estimated values represent the current signal, decreases rapidly with the passage of time.
For the generation of the spectrum to fade, they suggest frame repetition or noise substitution [LS01, sections 3.2 and 3.3].
Quackenbusch and Driesen suggest for AAC an exponential frame fade to zero [QD03]. A repetition 'of an adjacent group of time / frequency coefficients is proposed, where each repetition has exponentially increasing attenuation, thus gradually fading to silence in case of prolonged outputs.
With regard
Spectral
Band
Replication
Spectral Band)) in MPEG-4 HE-AAC,
3GPP suggests for
SBR on Enhanced aacPlus that stores the data in decoded envelope and, in the event of a loss of a frame, reuses the stored energies of the transmitted envelope data and reduces it by a constant 3 dB ratio for each hidden frame. The result is fed into the normal decoding process where it is used by the wrapper adjuster to calculate the magnification, used to adjust the high patched bands created by the HF generator. Then decoding takes place
SBR as usual. Furthermore, the delta encoded noise floor and sine level values are detected. Since no difference from the previous information remains available, the decoded noise floor and sine levels remain proportional to the energy of the generated HF signal [3GP12e, section 5.2].
The DRM consortium specifies for SBR in conjunction with AAC the same technique as 3GPP [EBU12, section 5.6.3.1]. Furthermore, the DAB consortium specifies for SBR in DAB + the same technique as 3GPP [EBU10, section A2].
In the text that follows, MPEG-4 CELP and MPEG4 HVXC (HVXC = Harmonic Vector Excitation Coding) are considered. The DRM consortium specific to SBR in conjunction with CELP and
HVXC [EBU12, the minimum requirement concealment for SBR for speech codecs is to apply a predetermined set of data values, as long as a frame of
Corrupt SBR. These values produce a spectral static high band envelope at a low relative playback level, exhibiting movement toward higher frequencies.
The goal is simply to ensure that no potentially loud, misbehaving audio pops reach the listener's ears, by inserting comfort noise (as opposed to strict muting). This is not actually a real fading but rather a jump to a certain level of strategy in order to insert some kind of comfort noise.
Later, an alternative is mentioned [EBU12, section 5.6.3.2] that reuses the last correctly coded data and slowly fades the (L) levels towards 0, analogously to the AAC + SBR case.
Now, it is considered HILN MPEG-4 (HILN = Harmonic and Individual Lines plus Noise). Meine et al. introduce a fading for the HILN MPEG-4 [ISO09] parametric codec in a parametric domain [MEP01]. For continuous harmonic components, a good default behavior for replacing corrupted differentially coded parameters is to maintain a frequency constant, to reduce the amplitude by an attenuation factor (e.g. -6 dB), and to allow the spectral envelope converge toward that low-pass average characteristic. An alternative for the spectral envelope would remain unchanged. With regard to amplitudes and spectral envelopes, noise components can be treated in the same way as harmonic components.
In the text that follows, the plotting of the background noise level in the prior art is considered. Rangachari and Loizou [RL06] provide a good overview of various methods and discuss some of their limitations. The methods for plotting the background noise level with eg. , minimal trace procedure [RL06] [Coh03] [SFB00] [Dob95], based on VAD (VAD = voice activity detection); Raiman filtering [Gan05] [BJH06], sub-space decompositions [BP06] [HJH08]; Soft Decision [SS98] [MPC89] [HE95], and minimum statistics.
The minimum statistics method was chosen to be used within the scope for USAC-2, (USAC - Unified Speech and Audio Coding) and is defined in greater detail below.
The noise energy spectral density estimation based on minimum statistics and optimal smoothing [MarOl] introduces a noise estimator, which is capable of working independently of the signal that is background noise or active voice. In contrast to other methods, the minimum statistics algorithm does not use any explicit thresholds to distinguish between voice activity and voice pause and is therefore more closely related to decisions.
Density during
The smooth, voiced methods
<td>soft than with</td><td colspan="2">methods</td><td colspan="2">detection of</td><td>exercise</td>
<td>traditional.</td><td>Similary</td><td>to</td><td colspan="2">the methods of</td><td>decision</td>
<td>also can</td><td colspan="2">to update</td><td>the PSD i</td><td>(Power</td><td>Spectral</td>
<td colspan="2">(Spectral Density of</td><td colspan="2">Energy)) of</td><td>noise</td><td>Dear</td>
<td>the activity of</td><td>voice.</td><td></td><td></td><td></td><td></td>
<td colspan="2">statistical method</td><td></td><td>minimum is</td><td>based</td><td>in two</td>
Noise are usually observations, that is, that speech and speech are statistically independent and that the energy of a noisy speech signal frequently decays to the noise energy level. Therefore, it is possible to derive an estimate of PSD (PSD = power spectral density (Spectral Density of
Accurate noise energy)) by tracking the noisy signal PSD minimum. Since the minimum is less than (or in other cases equal to) the average value, the minimum trace method requires compensation for bias.
The bias is a function of the variance of the smoothed signal PSD and as such depends on the smoothing parameter of the PSD estimator. In contrast to previous work on minimal tracking, which uses a constant smoothing parameter and a constant correction for minimal bias, a frequency and time dependent PSD smoothing is used, which also requires time and frequency dependent bias compensation. .
Using the minimum trace provides an estimate of the noise power. However, there are no downsides. Smoothing with a fixed smoothing parameter widens the speech activity peaks of the smoothed PSD estimate. This will lead to inaccurate noise estimates as the slip window for minimum search would slip on wide peaks. Thus, smoothing parameters close to one cannot be used, and as a consequence the noise estimate will have a relatively large variance. Furthermore, the noise estimate leans towards lower values. Additionally, in the case of increasing noise power, the minimum trace is left behind.
The low complexity MMSE-based noise PSD plot [HHJ10] introduces a background noise PSD method that uses an MMSE search used on a DFT (Discrete Fourier Transform) spectrum. The algorithm consists of these processing steps:
-The maximum probability estimator is computed based on the noise PSD of the previous table.
-The minimum mean square estimator is computed.
-The maximum probability estimator is estimated using the decision-oriented method [EM84].
The inverse bias factor is computed assuming that the noise and speech DFT coefficients are Gaussian distributed.
The spectral density of the estimated noise power is smoothed.
Also, there is a net security method in order to avoid algorithm jam.
Non-fixed noise plotting based on data-driven recursive noise power estimation [EH08] presents a method for estimating the spectral variance of noise from voice signals contaminated by highly non-fixed noise sources. This method also uses smoothing in the frequency / time orientation.
A low complexity noise estimation algorithm based on noise power estimation smoothing and estimation bias correction [Yu09] enhances the method presented in [EH08]. The main difference is that the spectral gain function for estimating noise power is found by an iterative data driven method.
The statistical methods for the enhancement of the noisy voice [Mar03] combine the minimum statistical method given in [MarOl] by modification of soft decision augmentation [MCA99], by an a-priori SNR estimate [MCA99], by a [ MC99] limiting of adaptation gain and by a spectral amplitude estimator of log MMSE [EM85].
Fading is of particular interest for a plurality of audio and voice codees, in particular, AMR (see [3GP12b]) (including ACELP and CNG), AMR-WB (see [3GP09c]) (including ACELP and CNG), AMR -WB + (see [3GP09a]) (including ACELP, TCX and CNG), G.718 (see [ITU08a]), G.719 (see [ITU08b]), G.722 (see [ITU07]),
G.722.1 (see [ITU05]), G.729 (see [ITU12, CPK08, PKJ + 11]), MPEG-4 HE-AAC / Enhanced aacPlus (see [EBU10, EBU12, 3GP12e, LS01, QD03]) ( including AAC and SBR), HILN MPEG-4 (see [ISO09, MEP01]) and OPUS (see [IET12]) (including SILK and CELT).
Depending on the codec, fading is performed in different domains:
For codecs using LPC, fading is done in the linear predictive domain (also known as the excitation domain). This applies to the truth for codes that are based on ACELP, eg, AMR, AMR-WB, the ACELP kernel of AMR-WB +, G.718, G.729, G.729.1, the SILK kernel in OPUS; codecs that further process the excitation signal with the use of a frequency-time transformation, eg. , the TCX core in AMR-WB +, the CELT core in OPUS; and for the generation of comfort noise schemes (CNG), which operate in the linear predictive domain, eg, CNG in AMR, CNG in AMR-WB, CNG in AMR-WB +.
For codes that directly transform the time signal into the frequency domain, fading is performed in the spectral domain / sub-band. This is true for codecs that are based on MDCT or a similar transformation, such as AAC in MPEG-4 HE-AAC, G.719, G.722 (sub-band domain) and G.722.1.
For parametric codes, the fading is applied in the parametric domain. This actually happens for HILN MPEG-4.
With regard to the fading rate and the fading curve, a fading is commonly realized by applying an attenuation factor, which is applied to the signal representation in the appropriate domain. The size of the fader factor controls the fade rate and the fade curve. In most cases the attenuation factor is applied frame by frame, but one application per sample is also used, see eg G.718 and G.722.
The attenuation factor for a certain signal segment could be provided in two ways, absolute, and relative.
In the case where an attenuation factor is given in absolute form, the reference level is always the one of the last frame received. Absolute attenuation factors usually start with a value close to 1 for the signal segment immediately after the last good frame and then degrade faster or slower towards 0. The fading curve directly depends on these factors. This is the case, for example, for the concealment described in Appendix IV of G.722 (see in particular [ITU07, Figure IV.7]), where the possible fading curves are linear or gradually linear. If the magnification factor g (n) is considered, while g (0) represents the gain factor of the last good frame, an absolute attenuation factor at<sub>ajbs</sub>(n), the magnification factor of any subsequent lost frames can be derived as g (n) = a<sub>abs</sub>(n) g (Q) (21)
In the case where an attenuation factor is relatively provided, the reference level is that of the previous table. This has disadvantages in the case of a recursive hiding procedure, eg. , if the already attenuated signal is further processed and attenuated again.
If an attenuation factor is applied recursively, then this could be a fixed value independent of the number of consecutive lost frames, eg 0.5 for G.719 (see above); a fixed value relative to the number of frames lost consecutively, eg, as proposed for G.729 in [CPK08]: 1.0 for the first two frames, 0.9 for the next two frames, 0.8 for frames 5 and 6, and 0 for all subsequent tables (see above); or a value that is relative to the number of consecutive lost frames and that depends on the signal characteristics, eg. , a faster fading for an unstable signal and a slower fading for a stable signal, eg G.718 (see previous section and [ITU08a, table 44]);
Assuming a relative fading factor of 0 ceri (n ') <1, while n is the amount of frame lost (n> 1); the magnification factor of any subsequent frame can be derived as g (n) = a<sub>re</sub>i (n) · g (n - 1) (22) </ (η) = (Π «(0] · δ (θ) \ m = l / (23) íK<sup>n</sup>) =®”<sub>and</sub>r # (°) (24) which produces an exponential fading.
Regarding the fading procedure, usually, the attenuation factor is specified, but in some application standards (DRM, DAB +) the latter is left to the manufacturer.
If different signal parts fade out separately, different attenuation factors could be applied, eg to fade out tonal components with a certain speed and components like noise with another speed (eg AMR, SILK).
Usually a certain magnification is applied to the whole picture. When fading is in the spectral domain, it is the only possible mode. However, if the fading is done in the time domain or linear predictive domain, a more granular fading is possible. Such more granular fading is applied in G.718, where individual magnification factors are derived for each sample by linear interpolation between the magnification factor of the last frame and the magnification factor of the current frame.
For codes with a variable frame length, a constant, relative dimming factor leads to a different fading rate depending on the frame length. This is the case, for example, for AAC, where the frame duration depends on the sample collection speed.
To adopt the fading curve applied to the time shape of the last received signal, the fading factors (static) could be further adjusted. Such additional dynamic adjustment can, for example, be applied for AMR where the median of the five previous magnification factors is taken into account (see [3GP12b] and section 1.8.1). Before performing any damping, the current gain is set to the median, if the median is less than the last boost, otherwise the last boost is used. Furthermore, such additional dynamic adjustment, eg, applies to G729, where the amplitude is predicted using linear regression of the previous magnification factors (see [CPK08, PKJ + 11] and section 1.6). In this case, the resulting magnification factor for the first, hidden frames could exceed the magnification factor for the last received frame.
Regarding the weighted level of the fading, with the exception of G.718 and CELT, the weighted level is 0 for all the codes analyzed, including the comfort noise generation of those codes (CNG).
In G.718, tone drive fading (representing tonal components) and random drive fading (representing components such as noise) is performed separately. While the pitch boost factor fades to zero, the innovation boost factor fades to the CNG drive energy.
Assuming relative attenuation factors are given, this leads - based on formula (23) - to the following absolute attenuation factor:
g (n) = a<sub>re</sub>i (n) g (n - 1) + (1- a<sub>re</sub>i (n)) g<sub>n (25)</sub> with g<sub>n</sub> as the gain of the excitation used during the generation of the comfort noise. This formula corresponds to formula (23), when g<sub>n</sub> = 0.
G.718 does not perform fading in the case of DTX / CNG.
In CELT there is no fade towards the weighted level, but after 5 frames of tonal concealment (including a fade) the level is instantly switched to
<td>the weighted level</td><td>in the 6th</td><td>lost painting</td><td>in</td><td>shape</td>
<td>consecutive. Level</td><td>drift by</td><td>portions with the</td><td>use</td><td>of the</td>
<td>formula (19).</td><td></td><td></td><td></td><td></td>
<td>With respect to</td><td>the shape</td><td colspan="2">spectral weighted</td><td>of the</td>
<td colspan="2">fading, all</td><td>based code</td><td>in</td><td>the</td>
Pure transformation analyzed (AAC, G.719, G.722, G.722.1) as well as SBR simply extend the spectral shape of the last good frame during fading.
Various speech codes fade spectral shape down to average with the use of LPC synthesis. The average could be static (AMR) or flexible (AMR-WB, AMR-WB +, G.718), while the latter is derived from a static average and a short-term average (derived by averaging the last sets of n LP coefficients) (LP = Linear Prediction).
All CNG modules in the AMR, AMR-WB, AMR-WB +, G.718 discussed codecs extend the spectral shape of the last good frame during fading.
With regard to tracking the background noise level, there are five different methods known from the literature:
Based on Voice Activity Detector: based on SNR / VAD, but very difficult to tune and difficult to use for low SNR language.
Soft decision scheme: The soft decision method takes into account the probability of speech presence [SS98] [MPC89] [HE95].
Minimum statistic: The minimum of the PSD is plotted by holding a certain amount of values over time in a buffer, in this way it is possible to find the minimum noise of past samples [MarOl] [HHJ10] [EH08] [Yu09].
Kalman filtering: The algorithm uses a series of measurements observed over time, containing noise (random variations), and produces noise PSD estimates that tend to be more accurate based on a single single measurement. The Kalman filter operates recursively on noisy input data streams to produce a statistically optimal estimate of system state [Gan05] [BJH06].
Sub-space decomposition: This approach attempts to decompose a noise signal into a clean speech signal and a noise part, using for example the KLT (Karhunen-Loéve transformation, also known as principal component analysis) and / or the DFT (Discrete Time Fourier Transform). Then the auto vectors / auto values can be plotted with the use of an arbitrary smoothing algorithm [BP06] [HJH08].
The aim of the present invention is to provide improved concepts for audio coding systems. The object of the present invention is solved by means of an apparatus according to claim 1, by a method according to claim 23 and by a computer program according to claim 24.
Furthermore, an apparatus is provided for decoding an audio signal. The apparatus comprises a receiving interface, wherein the receiving interface is configured to receive a first frame comprising a first audio signal portion
<td>of the signal</td><td>audio, and where</td><td>the</td><td>receiving interface</td><td>I know</td>
<td>configure for</td><td>receive a second</td><td colspan="2">box comprising</td><td>a</td>
<td>second portion</td><td>of the audio signal</td><td>of</td><td>the audio signal.</td><td></td>
<td>Furthermore, the</td><td colspan="2">apparatus comprises a</td><td>trace unit</td><td>of the</td>
<td>Noise level,</td><td>where the unit</td><td>of</td><td>level tracking</td><td>of</td>
Noise is set to determine noise level information depending on at least one of the first portion of the audio signal and the second portion of the audio signal (this means: depending on the first portion of the audio signal and / or the second portion of the audio signal), where the noise level information is represented in a plotting domain.
Additionally, the apparatus comprises a first reconstruction unit for the reconstruction of, in a first reconstruction domain, a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received through the receiving interface or if said third frame is received through the receiving interface but is corrupted, where the first reconstruction domain is different from or equal to the plotting domain.
Furthermore, the apparatus comprises a transformation unit for transforming the noise level information from the plotting domain to a second reconstruction domain, if a fourth frame of the plurality of frames is not received via the receiving interface or if said fourth frame is received through the receiving interface but is corrupted, where the second reconstruction domain is different from the plot domain, and where the second reconstruction domain is different from the first reconstruction domain, and
Additionally, the apparatus comprises a second reconstruction unit for the reconstruction of, in the second reconstruction domain, a fourth audio signal portion of the audio signal depending on the noise level information is represented in the second domain reconstruction, if said fourth frame of the plurality of frames is not received through the receiving interface or if said fourth frame is received through the receiving interface but is corrupted.
According to some embodiments, the plotting domain can, eg, be where the plotting domain is a time domain, a spectral domain, an FFT domain, an MDCT domain, or an excitation domain. The first reconstruction domain can, eg. , being the time domain, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain. The second reconstruction domain can, eg, be the time domain, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain.
In one embodiment, the plotting domain may, eg, be the FFT domain, the first reconstruction domain may, eg, be the time domain, and the second reconstruction domain may, eg, be the domain of excitement.
In another embodiment, the plotting domain may, eg, be the time domain, the first reconstruction domain may, eg, be the time domain, and the second reconstruction domain may, eg. . , be the domain of excitement.
According to one embodiment, said first portion of the audio signal can eg. , be represented in a first input domain, and said second portion of the audio signal may, eg, be represented in a second input domain. The transformation unit can, for example, be a second transformation unit. The apparatus may, for example, further comprise a first transformation unit for transforming the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from the second input domain to the plotting domain. to obtain a second information from the signal portion. The noise level tracking unit can, for example, be configured to receive a first information of the signal portion is represented in the plotting domain, where the first information of the signal portion depends on the first signal portion. audio, where the noise level tracking unit is configured to receive the second portion of the signal represented in the trace domain, and wherein the noise level tracking unit is configured to determine the noise level information depending on the first information of the signal portion is represented in the plotting domain and depending on the second portion of the signal information is displayed. represents in the plotting domain.
According to one embodiment, the first input domain may, eg, be the excitation domain, and the second input domain may, eg. , be the MDCT domain.
<td>In</td><td colspan="2">Another way</td><td>of</td><td colspan="2">realization, the</td><td>first</td><td colspan="2">domain</td><td>of</td>
<td>entry</td><td>may,</td><td>by</td><td>ex. ,</td><td>be the</td><td>domain</td><td>MDCT, and</td><td>in</td><td>where</td><td>he</td>
<td>second</td><td>domain</td><td>of</td><td colspan="3">input can, by</td><td>eg, be</td><td>he</td><td colspan="2">domain</td>
<td>MDCT.</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>Of</td><td>agreement</td><td>with</td><td>a</td><td>shape</td><td>of real</td><td>ization,</td><td>the</td><td colspan="2">first</td>
Reconstruction unit can, eg, be configured to reconstruct the third audio signal portion by driving a first fading into a noise-like spectrum. The second reconstruction unit can, for example, be configured to reconstruct the fourth portion of the audio signal by driving a second fading to a noise-like spectrum and / or a second fading to an LTP gain. Furthermore, the first reconstruction unit and the second reconstruction unit can eg be configured to drive the first fading and the second fading to a noise-like spectrum and / or a second fading of an LTP gain with the same speed. of fading.
In one embodiment, the apparatus can, eg. , further comprising a first aggregation unit for determining a first aggregated value depending on the first audio signal portion. Furthermore, the apparatus may further, eg, comprise a second aggregation unit for determining, depending on the second portion of the audio signal, a second aggregated value as the derived value of the second portion of the audio signal. The noise level tracking unit can, for example, be configured to receive the first added value since the first information of the signal portion is represented in the plotting domain, where the noise level tracking unit can e.g. be configured to receive the second added value since the second portion of the signal information is represented in the plotting domain, and where the unit, Noise level tracking is configured to determine noise level information depending on the first added value is represented in the plotting domain and depending on the second added value it is represented in the plotting domain.
According to one embodiment, the first aggregation unit may, for example, be configured to determine the first aggregated value such that the first aggregated value indicates a root average square of the first audio signal portion or of a signal derived from the first audio signal portion. The second aggregation unit is configured to determine the second aggregate value such that the second aggregate value indicates an average root square of the second portion of the audio signal or of a signal derived from the second portion of the audio signal. .
In one embodiment, the first transformation unit can, eg. , be configured to transform the derived value of the second portion of the audio signal from the second input domain to the trace domain by applying a gain value on the derived value of the second portion of the audio signal.
According to one embodiment, the gain value may, eg, indicate a gain introduced by linear predictive coding synthesis, or wherein the gain value indicates a gain introduced by linear predictive coding synthesis and de-emphasis. .
In one embodiment, the noise level tracking unit can eg be configured to determine the noise level information by applying a minimum statistics method.
According to one embodiment, the noise level tracking unit can eg be configured to determine a comfort noise level as the noise level information. The rebuilding unit can eg. , be configured to reconstruct the third portion of the audio signal depending on the noise level information, if said third frame of the plurality of frames is not received through the receiving interface or if said third frame is received through the interface receiver but is corrupt.
In one embodiment, the noise level tracking unit can eg. , be configured to determine a comfort noise level as the noise level information derived from a noise level spectrum, wherein said noise level spectrum is obtained by applying the minimum statistics method. The reconstruction unit may, for example, be configured to reconstruct the third audio signal portion depending on a plurality of Linear Predictive Coefficients, if said third frame of the plurality of frames is not received via the receiving interface or if said third frame is received through the receiving interface but is corrupt.
According to one embodiment, the first reconstruction unit can, eg. , be configured to reconstruct the third audio signal portion depending on the noise level information and depending on the first audio signal portion, if said third frame of the plurality of frames is not received by means of the receiving interface or if said third frame is received through the receiving interface but is corrupted.
In one embodiment, the first reconstruction unit can, eg, be configured to reconstruct the third audio signal portion by attenuating or amplifying the first audio signal portion.
According to one embodiment, the second reconstruction unit can eg be configured to reconstruct the fourth portion of the audio signal depending on the noise level information and depending on the second portion of the audio signal.
In one embodiment, the second reconstruction unit can, eg. , be configured to reconstruct the fourth portion of the audio signal by attenuating or amplifying the second portion of the audio signal.
According to one embodiment, the apparatus can eg. , further comprising a long-term prediction unit comprising a delay buffer, wherein the long-term prediction unit can, for example, be configured to generate a processed signal depending on the first or second portion of the signal from audio, depending on an input to the delay buffer stored in the delay storage device and depending on a long-term prediction gain, and wherein the long-term prediction unit is configured to fade the long-term prediction gain toward zero, if said third frame of the plurality of frames is not received via the receiving interface or if said third frame is received at through the receiving interface but it is corrupt.
In one embodiment, the long-term prediction unit can, eg. , be configured to fade the long-term prediction gain toward zero, where a rate at which the long-term prediction gain fades to zero depends on a fading factor.
In one embodiment, the long-term prediction unit can, for example, be configured to update the input of the delay storage device by storing the processed signal generated in the delay storage device, if said third frame of the plurality of frames is not received through the receiving interface or if said third frame is received through the receiving interface but is corrupted.
Furthermore, a method for decoding an audio signal is provided. The method comprises:
Receiving a first frame comprises a first audio signal portion of the audio signal, and receiving a second frame comprising a second audio signal portion of the audio signal.
Determine the noise level information depending on at least one of the first portion of the audio signal and the second portion of the audio signal, where the information of the noise level is represented in a plotting domain.
Reconstruct, in a first reconstruction domain, a third portion of the audio signal of the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received or if said third frame is received but it is corrupt, where the first rebuild domain is different from or equal to the plot domain.
Transform the noise level information from the plotting domain into a second reconstruction domain, if a fourth frame of the plurality of frames is not received or if said fourth frame is received but is corrupted, where the second reconstruction domain is different from the plotting domain, and where the second reconstruction domain is different from the first domain of
<td>reconstruction. AND:</td><td></td><td></td><td></td><td></td>
<td>Rebuild, in</td><td>he</td><td>second</td><td>domain</td><td>of</td>
<td>reconstruction, a fourth</td><td>portion</td><td>signal of</td><td>audio of</td><td>the</td>
<td>audio signal depending</td><td>of the</td><td>information</td><td>Of the level</td><td>of</td>
Noise is represented in the second reconstruction domain, if said fourth frame of the plurality of frames is not received or if said fourth frame is received but corrupted.
Furthermore, a computer program for the implementation of the method described above when running on a computer or signal processor is provided.
Furthermore, an apparatus is provided for decoding an audio signal.
The apparatus comprises a receiver interface. The receiving interface is configured to receive a plurality of frames, wherein the receiving interface is configured to receive a first frame of the plurality of frames, said first frame comprises a first audio signal portion of the audio signal, said first portion of the audio signal is represented in a first domain, and where the receiving interface is configured to receive a second frame of the plurality of frames, said second frame comprises a second portion of the audio signal of the audio signal.
Furthermore, the apparatus comprises a transformation unit for transforming the second portion of the audio signal or a value or a signal derived from the second portion of the audio signal from a second domain to a tracing domain to obtain a second information. of the signal portion, where the second domain is different from the first domain, where the tracing domain is different from the second domain, and where the plotting domain is the same as or different from the first domain.
Additionally, the apparatus comprises a noise level tracking unit, wherein the noise level tracking unit is configured to receive a first information from the signal portion that is represented in the trace domain, wherein the first Signal portion information depends on the first portion of the audio signal. The noise level tracking unit is configured to receive the second portion of the signal that is represented in the plotting domain, and where the noise level tracking unit is configured to determine the noise level information depending on the First information of the signal portion is represented in the plotting domain and depending on the second portion of the signal information it is represented in the plotting domain.
Furthermore, the apparatus comprises a reconstruction unit for reconstructing a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received by means of of the receiving interface but it is corrupt.
An audio signal can, for example, be a voice signal, or a music signal, or a signal comprising voice and music, etc.
The statement that the first signal portion information depends on the first audio signal portion means that the first signal portion information is either the first audio signal portion, or that the first audio signal portion signal has been obtained / generated depending on the first audio signal portion or in some other way dependent on the first audio signal portion. For example, the first audio signal portion may have been transformed from one domain to another domain to obtain the first information from the signal portion.
Similarly, a statement that the second portion of the signal information depends on a second portion of the audio signal means that the second portion of the signal information is either the second portion of the audio signal, or that the second portion of the signal information has been obtained / generated depending on the second portion of the audio signal or in some other way dependent on the second portion of the audio signal. For example, the second portion of the audio signal may have been transformed from one domain to another domain to obtain the second information from the signal portion.
In one embodiment, the first audio signal portion can, eg, be represented in a time domain as the first domain. Furthermore, the transformation unit can, eg. , be configured to transform the second portion of the audio signal or the value derived from the second portion of the audio signal from an excitation domain that is the second domain in the time domain is the plotting domain. Additionally, the noise level tracking unit can eg be configured to receive the first information of the signal portion represented in the time domain as the trace domain. Furthermore, the noise level tracking unit can eg. , set to receive the second signal portion is represented in the time domain as the plotting domain.
According to one embodiment, the first audio signal portion can, for example, be represented in an excitation domain as the first domain. Furthermore, the transform unit can, for example, be configured to transform the second portion of the audio signal or the value derived from the second portion of the audio signal from a time domain which is the second domain to excitation domain which is the plotting domain. Additionally, the noise level tracking unit can eg. , set to receive the first information of the signal portion is represented in the excitation domain as the trace domain. Furthermore, the noise level tracking unit can eg. , set to receive the second signal portion is represented in the excitation domain as the trace domain.
In one embodiment, the first audio signal portion can, eg. , represented in an excitation domain as the first domain, wherein the noise level tracking unit may, e.g., be configured to receive the first information from the signal portion, wherein said first information from the signal portion is represented in the FFT domain, which is the tracing domain, and wherein said first information of the signal portion depends on said first portion of the audio signal is represented in the excitation domain, wherein the transform unit can, e.g., be configured to transform the second portion of the audio signal or the derived value of the second portion of the audio signal from a time domain which is the second domain to a FFT domain which is the plotting domain, and where the noise level tracking unit can, eg, be configured to receive the second portion of the audio signal is represented in the FFT domain.
In one embodiment, the apparatus may, eg, further comprise a first aggregation unit for determining a first aggregated value depending on the first audio signal portion. Furthermore, the apparatus may, eg, further comprise a second aggregation unit for determining, depending on the second portion of the audio signal, a second aggregated value as the derived value of the second portion of the audio signal. Additionally, the noise level tracking unit can eg. , be configured to receive the first added value since the first information of the signal portion is represented in the plotting domain, where the noise level tracking unit can, e.g., be configured to receive the second added value already that the second portion of the signal information is represented in the plotting domain, and where the noise level tracking unit can, e.g., be configured to determine the noise level information depending on the first added value is represented in the plotting domain and depending on the second added value it is represented in the plotting domain.
According to one embodiment, the first aggregation unit may, for example, be configured to determine the first aggregated value such that the first aggregated value indicates a root average square of the first audio signal portion or of a signal derived from the first audio signal portion. Furthermore, the second aggregation unit can, for example, be configured to determine the second aggregated value such that the second aggregated value indicates an average root square of the second portion of the audio signal or of a signal derived from the second portion of the audio signal.
In one embodiment, the transformation unit can, eg,. be configured to transform the derived value of the second portion of the audio signal from the second domain to the plotting domain by applying a gain value on the derived value of the second portion of the audio signal.
According to the embodiments, the gain value may, eg, indicate a gain entered by linear predictive coding synthesis, or the gain value may, eg, indicate a gain entered by linear predictive encoding synthesis. and de-emphasis.
In one embodiment, the noise level tracking unit can eg be configured to determine the noise level information by applying a minimum statistics method.
According to one embodiment, the noise level tracking unit can eg. , set to determine a comfort noise level as the noise level information. The rebuilding unit can eg. , be configured to reconstruct the third portion of the audio signal depending on the noise level information, if said third frame of the plurality of frames is not received through the receiving interface or if said third frame is received through the interface receiver but is corrupt.
In one embodiment, the noise level tracking unit can, for example, be configured to determine a comfort noise level as the noise level information derived from a noise level spectrum, wherein said spectrum noise level is obtained by applying the minimum statistics method. The rebuilding unit can eg. , be configured to reconstruct the third portion of the audio signal depending on a plurality of Linear Predictive Coefficients, if said third frame of the plurality of frames is not received through the receiving interface or if said third frame is received through the interface receiver but is corrupt.
According to another embodiment, the noise level tracking unit can eg. , be configured to determine a plurality of Predictive Coefficients
Linear indicating a comfort noise level as the noise level information, and the reconstruction unit can eg be configured to reconstruct the third portion of the audio signal depending on the plurality of Linear Predictive Coefficients.
In one embodiment, the noise level tracking unit is configured to determine a plurality of FFT coefficients indicating a comfort noise level as the noise level information, and the first reconstruction unit is configured to reconstruct the third portion of the audio signal depending on a comfort noise level derived from said FFT coefficients, if said third frame of the plurality of frames is not received through the receiving interface or if said third frame is received through the receiving interface but is corrupted.
In one embodiment, the reconstruction unit can, eg. , be configured to reconstruct the third audio signal portion depending on the noise level information and depending on the first audio signal portion, if said third frame of the plurality of frames is not received by means of the receiving interface or if said third frame is received through the receiving interface but is corrupted.
According to one embodiment, the reconstruction unit can eg. , configured to reconstruct the third portion of the audio signal by attenuating or amplifying a signal derived from the first or second portion of the audio signal.
In one embodiment, the apparatus may, eg, further comprise a long-term prediction unit comprising a delay buffer. Furthermore, the long-term prediction unit can, for example, be configured to generate a processed signal depending on the first or second portion of the audio signal, depending on an input to the delay buffer stored in the storage device. delay and depending on a long-term prediction gain. Additionally, the long-term prediction unit may, eg, be configured to fade the long-term prediction gain toward zero, if said third frame of the plurality of frames is not received via the receiving interface or if said third frame is received through the receiving interface but it is corrupted.
According to one embodiment, the long-term prediction unit may, for example, be configured to fade the long-term prediction gain toward zero, where a speed with which the long-term prediction gain is reduced. fading to zero depends on a fading factor.
In one embodiment, the long-term prediction unit can, eg. , be configured to update the input of the delay storage device by storing the processed signal generated in the delay storage device, if said third frame of the plurality of frames is not received by means of the receiving interface or 'if said third frame is received through receiving interface but corrupted.
According to one embodiment, the transformation unit can eg. , be a first transformation unit, and the rebuilding unit is a first rebuilding unit. The apparatus further comprises a second transformation unit and a second reconstruction unit. The second transformation unit can eg. , be configured to transform the noise level information from the plotting domain to the second domain, if a fourth frame of the plurality of frames is not received through the receiving interface or if said fourth frame is received through the receiving interface but it is corrupt. Furthermore, the second reconstruction unit can eg. , be configured to reconstruct a fourth audio signal portion of the audio signal depending on the noise level information represented in the second domain if said fourth frame of the plurality of frames is not received via the receiving interface or if said fourth frame is received through the receiving interface but is corrupted.
In one embodiment, the second reconstruction unit may, eg, be configured to reconstruct the fourth portion of the audio signal depending on the noise level information and depending on the second portion of the audio signal.
According to one embodiment, the second reconstruction unit may, eg, be configured to reconstruct the fourth portion of the audio signal by attenuating or amplifying a signal derived from the first or second portion of the audio signal. .
Furthermore, a method for decoding an audio signal is provided.
The method comprises:
Receiving a first frame of a plurality of frames, said first frame comprises a first audio signal portion of the audio signal, said first portion of the audio signal is represented in a first domain.
Receiving a second frame of the plurality of frames, said second frame comprises a second portion of the audio signal of the audio signal.
Transform the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from a second domain to a plotting domain to obtain a second information from the signal portion, where the second domain is different from the first domain, where the plotting domain is different from the second domain, and where the plotting domain is the same as or different from the first domain.
Determine the noise level information depending on the information of the first signal portion, it is represented in the plotting domain, and depending on the second portion of the signal information it is represented in the plotting domain, where the first information of the signal portion depends on the first signal portion of the
<td colspan="2">Audio. AND:</td><td colspan="2">Rebuild</td><td>a</td><td>third</td><td>signal portion</td><td>of</td>
<td>Audio</td><td>of</td><td>the signal</td><td>of</td><td>Audio</td><td colspan="2">depending on the information</td><td>of the</td>
<td>level</td><td>of</td><td>noise is</td><td colspan="3">represents in the</td><td>plotting domain, yes</td><td>a</td>
third frame of the plurality of frames is not received from whether said third frame is received but corrupted.
Additionally, a computer program for implementing the method described above when running on a computer or signal processor is provided.
Some of the embodiments of the present invention provide a time-varying smoothing parameter such that the tracking capabilities of the smoothed periogram and its variance are better balanced, to develop an algorithm for skew compensation, and to speed up tracing noise in general.
The embodiments of the present invention are based on the finding that with respect to fading, the following parameters are of interest: The fading domain;
the fading rate, or, more generally, the fading curve; the weighted level of fading; the weighted spectral shape of the fading; and / or the background noise level plot.
In this context, the embodiments are based on the finding that the prior art has significant disadvantages.
An apparatus and method for improved signal fading for auto switched coding systems during error concealment is provided.
<td colspan="4">Furthermore, a computer program</td><td>for</td><td>the</td>
<td>implementation of</td><td>method described</td><td colspan="2">previously</td><td>when</td><td>I know</td>
<td>run in a</td><td>computer or</td><td>processor</td><td>of</td><td>signal</td><td>I know</td>
<td>provides.</td><td></td><td></td><td></td><td></td><td></td>
<td>The forms</td><td>of realization</td><td>they observe</td><td>a</td><td>level</td><td>of</td>
<td>fading to</td><td>comfort noise.</td><td>Agree</td><td>with</td><td colspan="2">the forms</td>
Common comfort noise a level of performance is observed, which traces in the excitation domain.
The level of the weighted comfort noise during burst packet loss will be the same, regardless of the center encoder (ACELP / TCX) in use, and will always be up to date. It is not known prior art, where a common noise level tracking is necessary. The embodiments provide fading of a switched codec to a signal such as comfort noise during burst packet losses.
Furthermore, the embodiments observe that the overall complexity will be less when compared to having two independent noise level tracking modules, since the functions (PROM) and memory can be shared.
In the embodiments, the level tap in the excitation domain (when compared to the level tap in the time domain) provides more minima during active talk, as some of the voice information is covered by the coefficients of LP.
In the case of ACELP, according to the embodiments, the level shunting takes place in the excitation domain. In the case of TCX, in the embodiments, the level drifts in the time domain, and the gain from the LPC synthesis and de-emphasis is applied as a correction factor in order to model the energy level. in the excitation domain. Plotting the level in the excitation domain, eg before the FDNS, would also be possible in theory, but level compensation between the excitation domain of TCX and the excitation domain of ACELP is considered quite complex.
No prior art incorporates such a common pool level tracking across different domains. The prior techniques do not have such common comfort noise level tracking, eg, in the drive domain, in a codec switched system. Thus, the embodiments are advantageous over the prior art, as regards the prior art, the comfort noise level sought during burst packet losses may be different, depending on the preceding encoding mode. (ACELP / TCX), where the level was tracked; As in the prior art, the tracing that is separated for each encoding mode will cause unnecessary excess and additional computing complexity; and as in the prior art, no comfort noise level to date would be available in any of the cores due to the recent switching by this core.
According to some embodiments, the level plot is driven in the excitation domain, but fading by TCX is driven in the time domain. Due to the fading in the time domain, TDAC failures are avoided, which would produce aliases. This becomes of particular interest when the components of the tone signal are hidden. Furthermore, level conversion between the excitation domain of ACELP and the spectral domain of MDCT is avoided and thus, eg, computing resources are saved. Since in switching between the excitation domain and the time domain, a level adjustment is required between the excitation domain and the time domain. This is solved by the derivation of the gain that would be introduced by the LPC synthesis and pre-emphasis and to use this gain as a correction factor to convert the level between the two domains.
In contrast, the prior techniques do not drive level rationing in the excitation domain and TCX fading in the time domain. With regard to the state of the art transformation-based codes, the attenuation factor is applied both in the excitation domain (for masking methods such as ACELP / time domain, see [3GP09a]) or in the frequency domain (for methods domain names such as frame repetition or noise replacement, see [LS01]). A disadvantage of the prior art method of applying the attenuation factor in the frequency domain is that such aliasing will occur in the time domain overlap aggregate region. This will be the case for adjacent frames to which different attenuation factors are applied, because the fading procedure causes the TDAC (time domain alias concealment) to fail. This is particularly relevant when the components of the tone signal are hidden. The aforementioned embodiments are thus advantageous over the prior art.
The embodiments compensate for the high-pass filter influence on the gain of LPC synthesis. According to the embodiments, to compensate for the unwanted gain change of the LPC analysis and the emphasis produced by driving high-pass filters without speech, a correlation factor is derived. This correlation factor takes this unwanted gain change into account and modifies the weighted comfort noise level in the excitation domain such that the correct weighted level is achieved in the time domain.
In contrast, the prior art, eg, G.718 [ITU08a], introduces a high-pass filter into the signal path of the no-voice drive, as shown in Fig. 2, if the last good frame signal it was not classified as NO VOICE. By means of this, the techniques of the prior art produce unwanted side effects, since the gain of the subsequent LPC synthesis depends on the signal characteristics, which are altered by means of this high-pass filter. Since the background level is plotted and applied in the excitation domain, the algorithm relies on the gain of LPC synthesis, which in turn again depends on the characteristics of the excitation signal. In other words: Modification of the signal characteristics of the drive due to high-pass filtering, as driven by the prior art, could lead to a modified (usually reduced) gain of LPC synthesis. This leads to the wrong output level even though the drive level is correct.
The embodiments overcome these disadvantages of the prior art.
In particular, the embodiments observe a flexible spectral shape of comfort noise. In contrast to G.718, by plotting the spectral shape of the background noise, and by applying (fading to) this shape during burst packet losses, the noise characteristic of the preceding background noise will match, leading to a pleasant noise characteristic of comfort noise. This avoids obstructive spectral shape mismatches that can be introduced by using a spectral envelope that has been derived from offline training and / or the spectral shape of the last received frames.
Furthermore, an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal is provided. The apparatus comprises a receiving interface for receiving one or more frames, a coefficient generator, and a signal reconstructor. The coefficient generator is configured to determine, if a current frame of the one or more frames is received through the receiving interface and if the current frame received by the receiving interface is not corrupted, one or more of the first signal coefficients of audio, comprised by the current frame, wherein said one or more of the first audio signal coefficients indicate a characteristic of the encoded audio signal, and one or more noise coefficients indicating a background noise of the encoded audio signal. Furthermore, the coefficient generator is configured to generate one or more second coefficients of the audio signal, depending on the one or more of the first audio signal coefficients and depending on the one or more noise coefficients, if the current frame does not received via the receiving interface or if the current frame received by the receiving interface is corrupted. The audio signal reconstructor is configured to reconstruct a first portion of the reconstructed audio signal depending on the one or more of the first audio signal coefficients, whether the current frame is received through the receiving interface, and whether the frame current received by the receiving interface is not corrupted. Furthermore, the audio signal reconstructor is configured to reconstruct a second portion of the reconstructed audio signal depending on one or more second coefficients of the audio signal, if the current frame is not received via the receiving interface or if the current frame received by the receiving interface is corrupt.
In some of the embodiments, the one or more of the first audio signal coefficients may, eg. , be one or more coefficients of the linear predictive filter of the encoded audio signal. In some of the embodiments, the one or more of the first audio signal coefficients may, eg, be one or more coefficients of the linear predictive filter of the encoded audio signal.
According to one embodiment, the one or more noise coefficients may, eg, be one or more coefficients of the linear predictive filter indicating the background noise of the encoded audio signal. In one embodiment, the one or more coefficients of the linear predictive filter may, eg, represent a spectral shape of the background noise.
In one embodiment, the coefficient generator may, eg, be configured to determine the one or more second portions of the audio signal such that the one or more second portions of the audio signal are one or more coefficients. of the linear predictive filter of the reconstructed audio signal, or such that the one or more of the first audio signal coefficients are one or more spectral pairs mimicking the reconstructed audio signal.
According to one embodiment, the coefficient generator can, eg, be configured to generate the one or more second coefficients of the audio signal by applying the formula:
/ current [^] - θ: · 4 ~ (1 - cv) · ptmean [í] where f<sub>C</sub>urrent [i] indicates one of the one or more second coefficients of the audio signal, where fiastli] indicates one of the one or more of the first coefficients of the audio signal, where pt<sub>mean</sub>[i] is one of the one or more noise coefficients, where a is a real number with 0 c and 1, and where i is an index. In one embodiment, 0 <a <1.
According to one embodiment, fase [t] indicates a coefficient of the linear predictive filter of the encoded audio signal, and where fcurrenttil indicates a coefficient of the linear predictive filter of the reconstructed audio signal.
In one embodiment, pt<sub>mean</sub>[í] can, eg,
<td>indicate</td><td>he</td><td>noise</td><td>of</td><td>sign background</td><td>1 audio</td><td colspan="2">encoded.</td><td></td>
<td>In</td><td>a</td><td>shape</td><td>of</td><td>realization, the</td><td>generator</td><td colspan="3">coefficients</td>
<td>may,</td><td>by</td><td>ex. ,</td><td colspan="2">set up for</td><td>decide</td><td>, if he</td><td colspan="2">picture</td>
<td>current</td><td>of the</td><td>one</td><td>or</td><td>more pictures are</td><td>receive</td><td>through</td><td>of</td><td>the</td>
<td colspan="4">receiving interface</td><td>and if the painting</td><td colspan="2">current received</td><td>by</td><td>the</td>
Receiver interface is not corrupted, the one or more noise coefficients by determining a noise spectrum of the encoded audio signal.
According to one embodiment, the coefficient generator can, for example, be configured to determine coefficients of LPC representing background noise by using a method of minimal statistics on the signal spectrum to determine a spectrum of noise from background and by calculating the LPC coefficients representing the shape of the background noise from the background noise spectrum.
Furthermore, a method of decoding an encoded audio signal to obtain a reconstructed audio signal is provided. The method comprises:
Receive one or more pictures.
Determine, if a current frame of the one or more frames is received and if the current frame being received is not corrupted, one or more of the first audio signal coefficients, comprised by the current frame, where said one or more of the first audio signal coefficients indicate a characteristic of the encoded audio signal, and one or more noise coefficients that indicate a background noise of the encoded audio signal.
Generate one or more second coefficients of the audio signal, depending on the one or more of the first coefficients of the audio signal and depending on the one or more noise coefficients, if the current frame is not received or if the current frame is being reception is corrupt.
Reconstructing a first portion of the reconstructed audio signal depending on the one or more of the first audio signal coefficients, whether the current frame is received and whether the current frame being received is not corrupted. AND:
Reconstruct a second portion of the reconstructed audio signal depending on the one or more second coefficients of the audio signal, if the current frame is not received or if the current frame being received is corrupted.
Furthermore, a computer program for the implementation of the method described above when running on a computer or signal processor is provided.
Having common means of plotting and applying the comfort noise spectral shape during fading has several advantages. Tracing and applying the spectral shape in such a way that it can be performed similarly for both kernel codes allows a simple common method. CELT teaches only the banding of energies in the spectral domain and the banding of the spectral shape in the spectral domain, which is not possible for the CELP core.
In contrast, in the prior art, the spectral shape of the comfort noise introduced during blast losses is fully static, or partially static or partially flexible to the short-term average of the spectral shape (as noted in G.718 [ITU08a] ), and will usually not match the background noise in the signal before packet loss. This mismatch of comfort noise characteristics could be disturbing. According to the prior art, a form of noise can be employed in an off-line (static) trained way that may be pleasant sound for particular signals, but less pleasant for others, e.g. car noise sounds totally different than the noise from the office.
Furthermore, in the prior art, a short-term averaging of the spectral shape of the previously received frames may be employed that could bring the signal characteristics closer to the previously received signal, but not necessarily due to the characteristics background noise. In the prior art, the tracing of the spectral shape by bands in the spectral domain (as observed in CELT [IET12]) is not applicable for a codec 'switched with the use not only of a core based on an MDCT domain ( TCX) but also a kernel based on an ACELP. The aforementioned embodiments are thus advantageous over the prior art.
Furthermore, an apparatus is provided for decoding an encoded audio signal to obtain a reconstructed audio signal. The apparatus comprises a receiving interface for receiving one or more frames comprising information on a plurality of audio signal samples of a signal from the audio spectrum of the encoded audio signal, and a processor for generating the reconstructed audio signal. The processor is configured to generate the reconstructed audio signal by fading a modified spectrum to a weighted spectrum, if a current frame is not received through the receiving interface or if the current frame is received through the receiving interface but is corrupted, wherein the modified spectrum comprises a plurality of modified signal samples, wherein, for each of the modified signal samples of the modified spectrum, an absolute value of said modified signal sample is equal to an absolute value of one of the audio signal samples of the spectrum of the audio signal. Furthermore, the processor is configured not to fade the modified spectrum to the weighted spectrum, if the current frame of the one or more frames is received through the receiving interface and if the current frame received by the receiving interface is not corrupted.
According to one embodiment, the weighted spectrum may, eg, be a noise-type spectrum.
In one embodiment, the noise-like spectrum may, eg, represent weighted noise.
According to one embodiment, the noise-like spectrum may, for example, be shaped.
In one embodiment, the noise-like spectrum shape can, eg. , depend on a signal from the audio spectrum of a previously received signal.
According to one embodiment, the noise-like spectrum may, eg, be shaped depending on the shape of the spectrum of the audio signal.
In one embodiment, the processor may, eg, employ a tilt factor to shape the noise-like spectrum.
According to one embodiment, the processor can, for example, use the formula shaped_noise [i] = noise * power (tilt_factor, i / N) where N indicates the number of samples, where i is an index, where 0 <= i <N, with tilt_factor> 0, and where power is a function of power.
power (x, y) indicates x<sup>and</sup> i
power (tilt_factor, i / N) indicates tilt_factor<sup>N</sup>
If the tilt_factor is less than 1 this means attenuation with increasing i. If the tilt_factor is greater than 1 it means amplification with increasing i.
According to another embodiment, the processor can, for example, use the formula shaped_noise [i] = noise * (1 + i / (Nl) * (tilt_factor-l)) where N indicates the number of samples, where i is an index, where 0 <= · i <N, with tilt_factor> 0.
If the tilt_factor is less than 1 this means attenuation with increasing i. If the tilt_factor is greater than 1 it means amplification with increasing i.
According to one embodiment, the processor can eg be configured to generate the modified spectrum, by changing a sign of one or more of the audio signal samples from the audio signal spectrum, if the current frame not received through the receiving interface or if the current frame received by the receiving interface is corrupted.
In one embodiment, each of the audio signal samples from the audio signal spectrum can, eg. , represented by a real number but not by an imaginary number.
According to one embodiment, the audio signal samples from the audio signal spectrum may, eg, be represented in a Modified Individual Cosine Transformation Domain.
In another embodiment, the audio signal samples from the audio signal spectrum may, eg, be represented in a Modified Individual Sine Transformation Domain.
According to one embodiment, the processor may, eg, be configured to generate the modified spectrum by employing a random sign function that randomly or pseudo-randomly produces both a first and a second value.
In one embodiment, the processor can, eg. , set to fade the modified spectrum to the weighted spectrum by subsequently reducing an attenuation factor.
According to one embodiment, the processor can, eg. , set to fade the modified spectrum to the weighted spectrum by subsequently increasing an attenuation factor.
In one embodiment, if the current frame is not received via the receiving interface or if the current frame received by the receiving interface is corrupted, the processor can eg '. , be configured to generate the reconstructed audio signal using the formula:
x [i] = (l-cum_damping) * noise [i] + cum_damping * random_sign () * x_old [i] where i is an index, where x [i] indicates a sample of the reconstructed audio signal, in where cum_damping is an attenuation factor, where x_old [i] indicates one of the audio signal samples from the audio signal spectrum of the encoded audio signal, where random_sign () returns 1 or -1, and in where noise is a random vector indicating the weighted spectrum.
In one embodiment, said random vector noise can, for example, be increased such that its mean square is similar to the mean square of the spectrum of the encoded audio signal comprised of one of the frames received last by the receiving interface. .
According to a general embodiment, the processor can, for example, be configured to generate the reconstructed audio signal, by employing a random vector that is increased such that its mean square is similar to the mean square of the spectrum of the encoded audio signal comprised of one of the frames received last by the receiving interface.
Furthermore, a method of decoding an encoded audio signal to obtain a reconstructed audio signal is provided. The method comprises:
Receiving one or more frames comprising information on a plurality of audio signal samples from an audio signal spectrum of the encoded audio signal. AND:
Generate the reconstructed audio signal.
Generating the reconstructed audio signal is driven by fading a modified spectrum to a weighted spectrum, if a current frame is not received or if the current frame is received but corrupted, wherein the modified spectrum comprises a plurality of modified signal samples , where, for each of the modified signal samples of the modified spectrum, an absolute value of said modified signal sample is equal to an absolute value of one of the audio signal samples of the spectrum of the audio signal. The modified spectrum does not fade to a noise weighted spectrum, if the current frame of the one or more frames is received and if the current frame being received is not corrupted.
Furthermore, a computer program for the implementation of the method described above when running on a computer or signal processor is provided.
The embodiments observe a weighted noise-to-noise MDCT fading spectrum prior to FDNS Application (FDNS = Frequency Domain Noise Substitution).
According to the prior art, in ACELP-based codes, the innovative codebook is replaced with a random vector (eg, with noise). In the embodiments, the ACELP method, which consists of replacing the innovative codebook with a random vector (eg, with noise) is adopted to the structure of the TCX decoder. Here, the equivalent of the innovative codebook is the MDCT spectrum usually received within the bitstream and fed into the FDNS.
The classic MDCT masking method would be to simply repeat this spectrum as is or apply some scrambling process, which basically lengthens the spectral shape of the last frame received [LS01]. This has the disadvantage that the short-term spectral shape is prolonged, which often leads to a repetitive metallic sound, which is not like background noise, and thus cannot be used as comfort noise.
The use of the proposed method of short-term spectral form is performed by the use of FDNS and TCX LTP, the spectral form over a long cycle is performed only by FDNS. The shape by the FDNS fades from the short-term spectral shape to the traced long-term spectral shape of the background noise, and the TCX LTP fades to zero.
The fading of the FDNS coefficients to the plotted background noise coefficients leads to a smooth transition between the last good spectral envelope and the background spectral envelope that should be sought in the long cycle, in order to achieve background noise. nice in the case of long burst frame losses.
In contrast, according to the prior art, for transform-based codecs, noise concealment is driven by frame repetition or noise substitution in the frequency domain [LS01]. In the prior art, noise substitution is usually done by scrambling the sign of spectral vessels. If scrambling of the TCX sign (frequency domain) is used in the prior art during concealment, the coefficients of the last received MDCT are reused and each sign is randomized before the spectrum is inversely transformed into the time domain. . The disadvantage of this prior art procedure is that for consecutively lost frames the same spectrum is used over and over again, only with different sign randomizations and overall attenuation. When looking for the spectral envelope over time in a coarse time grid, it can be seen that the envelope is approximately constant during the loss of consecutive frames, because the band energies remain relatively constant with each other within a frame and they fade only globally. In the coding system used, according to the prior art, the spectral values are processed with the use of FDNS, in order to restore the original spectrum. This means, that if one wishes to fade the MDCT spectrum to a certain spectral envelope (with the use of FDNS coefficients, e.g. describing the current background noise), the result is not dependent on the FDNS coefficients, but also dependent previously decoded spectrum that is interfered with in the form of a sign scramble. The aforementioned embodiments overcome these disadvantages of the prior art.
The embodiments are based on the finding that it is necessary to fade the spectrum used for weighted sign-to-noise scrambling interference before feeding it into the FDNS process. Otherwise the spectrum generated will never match the envelope used for the FDNS process.
In the embodiments, the same fading rate is used to fade the LTP gain as for the weighted noise fading.
Furthermore, there is provided an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal. The apparatus comprises a receiving interface for receiving a plurality of frames, a delay buffer for storing audio signal samples from the decoded audio signal, a sample selector for selecting a plurality of selected audio signal samples from the audio signal samples. audio signal stored in the delay storage device, and a sample processor for processing the selected audio signal samples to obtain reconstructed audio signal samples from the reconstructed audio signal. The sample selector is configured to select, whether a current frame is received through the receiving interface and whether the current frame received by the receiving interface is not corrupted, the plurality of selected audio signal samples from the audio signal samples. stored in the delay storage device depending on a pitch lag information comprised by the current frame. Furthermore, the sample selector is configured to select, if the current frame is not received via the receiving interface or if the current frame received by the receiving interface is corrupted, the plurality of selected samples of audio signal from the samples of audio signal stored in the delay storage device depending on a lag per tone information comprised by another frame previously received by the receiving interface.
According to one embodiment, the sample processor can, eg. , set to get the reconstructed audio signal samples, if the current frame is received through the receiving interface, and if the current frame received by the receiving interface is not corrupted, by re-increasing the selected audio signal samples depending on of the gain information comprised by the current frame. Furthermore, the sample selector can eg be configured to get the reconstructed audio signal samples, if the current frame is not received via the receiving interface or if the current frame received by the receiving interface is corrupted, by increasing the selected audio signal samples again depending on the gain information comprised by said other frame previously received by the receiving interface.
In one embodiment, the sample processor can, for example, be configured to obtain the reconstructed audio signal samples, if the current frame is received through the receiving interface and if the current frame received by the receiving interface is not. is corrupted, by multiplying the selected audio signal samples and a value depending on the gain information comprised by the current frame. Furthermore, the sample selector is configured to obtain the reconstructed audio signal samples, if the current frame is not received via the receiving interface or if the current frame received by the receiving interface is corrupted, by multiplying the selected samples of audio signal and a value depending on the gain information comprised by said other frame previously received by the receiving interface.
According to one embodiment, the sample processor can, eg, be configured to store the reconstructed audio signal samples in the delay storage device.
In one embodiment, the sample processor may, eg, be configured to store the reconstructed audio signal samples in the delay storage device before receiving an additional frame through the receiver interface.
According to one embodiment, the sample processor can, eg. , be configured to store the reconstructed audio signal samples in the delay storage device after receiving an additional frame through the receiving interface.
In one embodiment, the sample processor can, for example, be configured to re-boost selected samples of audio signal depending on the gain information to obtain newly-boosted audio signal samples and by combining the Audio signal samples again augmented with input audio signal samples to obtain the processed audio signal samples.
According to one embodiment, the sample processor may, for example, be configured to store the processed audio signal samples, indicating the combination of the newly augmented audio signal samples and the audio signal samples. input, in the delay storage device, and do not store the newly augmented audio signal samples in the delay storage device, if the current frame is received through the receiving interface and if the current frame received by the receiving interface is not corrupted. Furthermore, the sample processor is configured to store the newly augmented audio signal samples in the delay storage device and not to store the processed audio signal samples in the delay storage device, if the current frame is not stored. received via the receiving interface or if the current frame received by the receiving interface is corrupt.
According to another embodiment, the sample processor can eg be configured to store the processed audio signal samples in the delay storage device, if the current frame is not received via the receiving interface. or if the current frame received by the receiving interface is corrupt.
<td>In</td><td>a</td><td>shape</td><td>of realization,</td><td colspan="2">the selector</td><td>of samples</td>
<td>may,</td><td>by</td><td>eg</td><td>set up for</td><td>get</td><td>the</td><td>samples of</td>
<td>signal</td><td>of</td><td>Audio</td><td>rebuilt to</td><td>return</td><td>to</td><td>increase the</td>
selected samples of audio signal depending on a modified gain, where the modified gain is defined according to the formula:
gain = gain_past * damping;
where gain is' the modified gain, where the sample selector can eg. , set to set gain_past to increase after the gain has been calculated, and where articulation is an actual value.
According to one embodiment, the sample selector can, eg, be configured to calculate the modified gain.
In one embodiment, joint can, eg, be defined according to: 0 joint 1.
According to one embodiment, the modified gain may, eg, be set to zero, if at least a predefined number of frames has not been received through the receiving interface since a frame was last received. once over the receiving interface.
100
Furthermore, a method of decoding an encoded audio signal to obtain a reconstructed audio signal is provided. The method comprises:
Receive a plurality of frames.
Store audio signal samples of the decoded audio signal.
Selecting a plurality of selected audio signal samples from the audio signal samples stored in the delay storage device. AND:
Process selected audio signal samples to obtain reconstructed audio signal samples from the reconstructed audio signal.
If a current frame is received and the current frame being received is not corrupted, the step of selecting the plurality of selected audio signal samples from the audio signal samples stored in the delay storage device is conducted depending of a lag information per tone comprised by the current frame. Furthermore, if the current frame is not received or if the current frame being received is corrupted, the step of selecting the plurality of selected audio signal samples from the audio signal samples stored in the delay storage device it is driven depending on a lag information per tone understood
101 by another frame previously received by the receiving interface.
Furthermore, a computer program is provided for the implementation of the method described above when run on a computer or signal processor.
The embodiments employ TCX LTP (TXC LTP = Transform Coded Excitation Long-Term Prediction.) During normal operation, the TCX LTP memory is
<td>update with</td><td>the</td><td>synthesized signal,</td><td>than</td><td>contains</td><td>noise</td><td>and</td>
<td colspan="2">tonal components</td><td>rebuilt.</td><td></td><td></td><td></td><td></td>
<td>Instead</td><td>of</td><td>disable the</td><td>TCX</td><td colspan="2">LTP during</td><td>he</td>
<td>concealment,</td><td>its</td><td colspan="2">normal functioning</td><td>may</td><td colspan="2">continue</td>
during concealment with the parameters received in the last good frame. This preserves the spectral shape of the signal, particularly those tonal components that are modeled by the LTP filter.
Furthermore, the embodiments decouple the feedback loop from the TCX LTP. A simple continuation of the normal TCX LTP operation introduces additional noise, since with each update step the additional randomly generated noise from the LTP drive is introduced. The tonal components are therefore distorted more and more over time by the added noise.
102
To overcome this, only the updated TCX LTP buffer can be fed back (without adding noise), in order not to contaminate the tonal information with unwanted random noise.
Additionally, according to the embodiments, the LTP TCX gain fades to zero.
These embodiments are based on the finding that TCX LTP continuation helps preserve signal characteristics in the short term, but has long-term disadvantages: The signal reproduced during concealment will include the voice / tonal information that was present before loss. Especially for the clean voice or voice over background noise, it is extremely unlikely that a tone or harmonica will drop very slowly over a very long time. When continuing the TCX LTP operation during cloaking, particularly if the LTP memory update is decoupled (only tonal components are fed back and not the part encrypted with the sign), the voice / tonal information will be present in the signal hidden for complete loss, attenuated only by the general level of fading to comfort noise. Furthermore, it is impossible to achieve the comfort noise envelope during packet burst losses, if TCX LTP is applied during burst loss without attenuating with step
103 of the time, because the signal then always incorporates the voice information of the LTP.
Therefore, the LTP TCX gain fades towards zero, such that the tonal components represented by the LTP will fade to zero, at the same time the signal fades to the level and shape of the background signal, and so that the fading reaches the desired spectral background envelope (comfort noise) without incorporating unwanted tonal components.
In l<sup>ace</sup> In embodiments, the same fading rate is used for LTP gain fading as for weighted noise fading.
In contrast, in the prior art, it is not known transform codec that uses LTP during concealment.
For MPEG-4 LTP [ISO09] there is no masking method in the prior art. Another prior art MDCT-based codec that makes use of an LTP is CELT, but this codec uses ACELP-like concealment for the first five frames, and background noise is generated for all subsequent frames, which makes no use of of the LTP. A disadvantage of the prior art of not using TCX LTP is, that all the tonal components modeled with the LTP abruptly disappear. Furthermore, in prior art ACELP-based codes, the LTP operation is prolonged during concealment, and the gain of the codelog
104 flexible fades to zero. With regard to the operation of the feedback loop, the prior art employs two methods, either full drive, eg. , the sum of innovative excitation and adaptation, is feedback (AMR-WB); or only the updated flexible drive, eg, the parts of the tonal signal, is fed back (G.718). The above mentioned embodiments overcome the disadvantages of the prior art.
In the text that follows, the embodiments of the present invention are described in greater detail with reference to the figures, in which:
Fig. 1A illustrates an apparatus for decoding an audio signal according to one embodiment,
Fig. IB illustrates an apparatus for decoding an audio signal according to another embodiment,
Fig. 1C illustrates an apparatus for decoding an audio signal according to another embodiment, wherein the apparatus further comprises a first and a second aggregation unit,
Fig. ID illustrates an apparatus for. decoding an audio signal according to a further embodiment, wherein the apparatus further comprises a long-term prediction unit comprising a delay buffer,
105
Fig. 2 illustrates the structure of the decoder
G.718,
Fig. 3 presents a scenario, where the G.722 fading factor depends on the class information,
Fig. 4 shows a method for predicting amplitude with the use of linear regression,
Fig. 5 illustrates the Retained Energy Overlapping Transformation (CELT) burst loss behavior,
Fig. 6 shows the background noise level tracking according to an embodiment in the decoder for an error-free mode of operation,
Fig. 7 illustrates LPC synthesis gain derivation and de-emphasis according to one embodiment,
Fig. 8 presents a comfort noise level application during packet loss according to one embodiment,
Fig. 9 illustrates advanced high pass gain compensation during ACELP concealment according to one embodiment,
Fig. 10 presents the decoupling of the LTP feedback loop during concealment according to one embodiment,
106
<td>The</td><td>Fig. 11</td><td>illustrates</td><td>an apparatus for</td><td>decode</td><td>a</td>
<td>signal</td><td>audio</td><td>encoded</td><td>To get one</td><td colspan="2">audio signal</td>
<td colspan="2">rebuilt from</td><td>agree with</td><td colspan="2">an embodiment,</td><td></td>
<td>The</td><td>Fig. 12</td><td>shows</td><td>an apparatus for</td><td>decode</td><td>a</td>
<td>signal</td><td>audio</td><td>encoded</td><td>To get one</td><td colspan="2">audio signal</td>
<td colspan="2">rebuilt from</td><td>agree with</td><td colspan="2">another embodiment, and</td><td></td>
<td>The</td><td>Fig. 13</td><td>illustrates</td><td>an apparatus for</td><td>decode</td><td>a</td>
<td>signal</td><td>audio</td><td>encoded</td><td>To get one</td><td colspan="2">audio signal</td>
<td colspan="2">rebuilt in</td><td colspan="3">a further embodiment, and</td><td></td>
<td>The</td><td>Fig. 14</td><td>illustrates</td><td>an apparatus for</td><td>decode</td><td>a</td>
<td>signal</td><td>audio</td><td>encoded</td><td>To get one</td><td colspan="2">audio signal</td>
rebuilt another embodiment.
Fig. Is illustrated by an apparatus for decoding an audio signal in accordance with one embodiment.
The apparatus comprises a receiving interface 110. The receiving interface is configured to receive a plurality of frames, wherein the receiving interface 110 is configured to receive a first frame of the plurality of frames, said first frame comprising a first audio signal portion of the audio signal, said first portion of the audio signal is represented in a first domain. Furthermore, the receiving interface 110 is configured to receive a second frame of the plurality of frames, said second frame comprising a second portion of the audio signal of the audio signal.
107
Furthermore, the apparatus comprises a transformation unit 120 for transforming the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from a second domain to a plotting domain to obtain a second signal portion information, where the second domain is different from the first domain, where the tracing domain is different from the second domain, and where the plotting domain is the same as or different from the first domain.
Additionally, the apparatus comprises a noise level tracking unit 130, wherein the noise level tracking unit is configured to receive a first information of the signal portion represented in the trace domain, wherein the first signal portion information depends on the first audio signal portion, wherein the noise level tracking unit is configured to receive the second signal portion is represented in the plotting domain, and wherein the noise level tracking unit is configured to determine the noise level information depending on the first information of the signal portion is represented in the plotting domain and depending on the second portion of the signal information is displayed. represents in the plotting domain.
Furthermore, the apparatus comprises a reconstruction unit for the reconstruction of a third portion
108 of the audio signal of the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received by the receiving interface but is corrupted.
With respect to the first and / or the second portion of the audio signal, for example, the first and / or the second portion of the audio signal can, for example, be fed into one or more processing units (no sample) to generate one or more speaker signals for one or more speakers, such that the received sound information comprised by the first and / or second portion of the audio signal can be replayed.
Furthermore, however, the first and second audio signals are also used for masking, eg in the case that subsequent frames do not reach the receiver or in the case that subsequent frames are in error.
Inter alia, the present invention is based on the finding that level noise tracing must be conducted in a common domain, referred to herein as a tracing domain. The tracing domain can, for example, be an excitation domain, for example the domain in which the signal is represented by LPCs (LPC = Linear Predictive Coefficient) or by ISPs (ISP = Immittance Spectral Pair) as described in AMR-WB and AMR-WB + (see [3GP12a], [3GP12b],
109 [3GP09a], [3GP09b], [3GP09c]). Plotting the noise level in a single domain has inter alia the advantage that aliasing effects are avoided when the signal switches between a first representation in a first domain and a second representation in a second domain (for example, when the representation change from ACELP to TCX or vice versa).
With respect to the transform unit 120, what is transformed is either the second portion of the audio signal itself, or a signal derived from the second portion of the audio signal (e.g., the second portion of the audio signal has been processed to obtain the derived signal), or a value derived from the second portion of the audio signal (eg, the second portion of the audio signal has been processed to obtain the derived value).
With respect to the first audio signal portion, in some of the embodiments, the first audio signal portion can be processed and / or transformed to the plotting domain.
In other embodiments, however, the first audio signal portion may already be represented in the plotting domain.
In some of the embodiments, the first information in the signal portion is identical to the first audio signal portion. In other embodiments, the first information in the signal portion is, eg, a
110 Added value depending on the first portion of the audio signal.
Now, as a first measure, the fading to a comfort noise level is considered in greater detail.
The fading method described can eg. , be implemented in a low-delay version of xHE-AAC [NMR + 12] (xHE-AAC = Extended High Efficiency AAC), which is able to optimally switch between ACELP (voice ) and MDCT (music / noise) on a per frame basis.
With respect to common level tracing in a tracing domain, for example an excitation domain, to apply a soft fading to an appropriate comfort noise level during packet loss, such comfort noise level must be identified during the process. normal decoding. It can, for example, be assumed that the noise level similar to background noise is more comfortable. In this way, the background noise level can be constantly derived and updated during normal decoding.
The present invention is based on the finding that when a core codec (eg ACELP and TCX) has been switched, considering a common background noise level independent of the chosen core encoder is particularly appropriate.
111
Fig. 6 presents background noise level tracking according to a preferred embodiment at the decoder during error free mode of operation, eg during normal decoding.
The plotting itself can, for example, be done using the least statistic method (see [MarOl]).
The plotted background noise level can eg be considered as the noise level information mentioned above.
For example, the noise estimate of the minimum statistic presented in the document: Rainer Martín, Noise power spectral density estimation based on optimal smoothing and minimal statistics, IEEE Transactions on Speech and Audio Processing 9 (2001), no. 5, 504-512 [MarOl] can be used to track the background noise level.
Correspondingly, in some of the embodiments, the noise level tracking unit 130 is configured to determine the noise level information by applying a minimum statistical method, eg, by employing the estimation of the noise level. minimum noise statistic of [MarOl].
Later, some considerations and details of this plotting method are described.
Regarding the level plot, the background is assumed to be similar to noise. Therefore it is preferred to perform the
112 level plotting in the excitation domain to avoid plotting the front tonal components that are picked up by the LPC. For example, ACELP noise filling can also employ the background noise level in the excitation domain. With excitation domain plotting, only a single background noise level plot can serve two purposes, saving computational complexity. In a preferred embodiment, the screening is done in the excitation domain of ACELP.
Fig. 7 illustrates LPC synthesis gain derivation and de-emphasis in accordance with one embodiment.
With regard to level shunting, level shunting can, for example, be conducted either in the time domain or in the excitation domain, or in any other appropriate domain. If the domains for level derivation and level tracing differ, a gain compensation may, for example, be necessary.
In the preferred embodiment, level shunting for ACELP is in the excitation domain. Therefore, no compensation for gain is required.
For TCX, a gain compensation may, for example, be necessary to adjust the level derived with the ACELP excitation domain.
In the preferred embodiment, level shunting for TCX occurs in the time domain. I know
113 found for this approach a manageable gain trade-off: The gain introduced by LPC synthesis and de-emphasis is derived as shown in Fig. 7 and the derived level is divided by this gain.
Alternatively, the level tap for
TCX could be performed in the excitation domain of
TCX.
However, the gain compensation between the excitation domain of TCX and the excitation domain of ACELP was considered too complicated.
Thus, returning to Fig. 1A, in some of the embodiments, the first audio signal portion is represented in a time domain as the first domain.
The transformation unit 120 is configured to transform the second portion of the audio signal or the value derived from the second portion of the audio signal to an excitation domain that is the second domain in the time domain that is the domain of tracing. In such embodiments, the noise level tracking unit 130 is configured to receive the first information of the signal portion that is represented in the time domain as the trace domain. Furthermore, the noise level tracking unit 130 is configured to receive the second signal portion that is represented in the time domain as the trace domain.
114
In other embodiments, the first audio signal portion is represented in an excitation domain as the first domain. The transformation unit 120 is configured to transform the second portion of the audio signal or the derived value of the second portion of the audio signal from a time domain that is the second domain to the excitation domain that is the domain plotting. In such embodiments, the noise level tracking unit 130 is configured to receive the first information from the signal portion that is represented in the excitation domain as the tracing domain. Furthermore, the noise level tracking unit 130 is configured to receive the second signal portion that is represented in the drive domain as the tracing domain.
In one embodiment, the first audio signal portion can, for example, be represented in an excitation domain as the first domain, where the noise level tracking unit
130 can, for example, be configured to receive the first information of the signal portion, wherein said first information of the signal portion is represented in the FFT domain which is the plotting domain, and wherein said first information of the signal portion depends on said first portion of the audio signal is represented in the excitation domain, where the transformation unit 120 can, for example, be configured
115 to transform the second portion of the audio signal or the derived value of the second portion of the audio signal from a time domain that is the second domain to an FFT domain that is the plotting domain, and where the noise level tracking unit 130 can eg. , set up to receive the second portion of the audio signal is represented in the FFT domain.
Fig. IB illustrates an apparatus according to another embodiment. In Fig. IB, the transformation unit 120 of Fig. 1A is a first transformation unit 120, and the reconstruction unit 140 of Fig. 1A is a first reconstruction unit 140. The apparatus further comprises a second unit transformation unit 121 and a second reconstruction unit 141.
The second transformation unit 121 is configured to transform the noise level information from the plotting domain to the second domain, if a fourth frame of the plurality of frames is not received via the receiving interface or if said fourth frame it is received through the receiving interface but it is corrupted.
Furthermore, the second reconstruction unit 141 is configured to reconstruct a fourth audio signal portion of the audio signal depending on the noise level information is represented in the second domain if said fourth frame of the plurality of frames is not shown. receive for
116 middle of the receiving interface or if said fourth frame is received through the receiving interface but is corrupted.
Fig. 1C illustrates an apparatus for decoding an audio signal according to another embodiment. The apparatus further comprises a first aggregation unit 150 for determining a first aggregated value depending on the first audio signal portion. Furthermore, the apparatus of Fig. 1C further comprises a second aggregation unit 160 for determining a second aggregated value as the derived value of the second portion of the audio signal depending on the second portion of the audio signal. In the embodiment of Fig. 1C, the noise level tracking unit 130 is configured to receive a first added value since the first information of the signal portion is represented in the plotting domain, where the noise level tracking unit 130 is configured to receive the second added value since the second portion of the signal information is represented in the plotting domain. The noise level tracking unit 130 is configured to determine the noise level information depending on the first added value is represented in the plotting domain and depending on the second added value it is represented in the plotting domain.
In one embodiment, the first aggregation unit 150 is configured to determine the first value
117 aggregated such that the first aggregated value indicates a root average square of the first audio signal portion or of a signal derived from the first audio signal portion. Furthermore, the second aggregation unit 160 is configured to determine the second aggregated value such that the second aggregated value indicates an average root square of the second portion of the audio signal or of a signal derived from the second portion of the audio signal. the audio signal.
Fig. 6 illustrates an apparatus for decoding an audio signal according to a further embodiment.
In Fig. 6, the background level tracing unit 630 implements a noise level tracing unit 130 in accordance with Fig. 1A.
Furthermore, in Fig. 6, the RMS 650 unit (RMS = root mean square) is a first aggregation unit and the RMS 660 unit is a second aggregation unit.
According to some embodiments, the (first) transform unit 120 of Fig. 1A, Fig. IB and Fig. 1C is configured to transform the derived value of the second portion of the audio signal from the second domain to the plotting domain by applying a gain value (x) to the derived value of the second portion of the audio signal, e.g. by dividing the derived value of the second portion of the audio signal by a
118 gain value (x). In other embodiments, a gain value can, eg, be multiplied.
In some of the embodiments, the gain value (x) may, eg, indicate a gain entered by linear predictive coding synthesis, or the gain value (x) may, eg, indicate an entered gain. by synthesis of linear predictive coding and de-emphasis.
In Fig. 6, unit 622 provides the value (x) which indicates the gain introduced by linear predictive coding synthesis and de-emphasis. Unit 622 then divides the value, provided by second aggregation unit 660, which is a value derived from the second portion of the audio signal, by the provided gain value (x) (e.g., either dividing by x, or by multiplying the value 1 / x). Thus, the unit 620 of Fig. 6 comprising units 621 and 622 implements the first transformation unit of Fig. 1A, Fig. IB or Fig. 1C.
The apparatus of Fig. 6 receives a first frame with a first audio signal portion which is a voice drive and / or a non-voice drive and is represented in the plotting domain, in Fig. 6 an LPC domain ( ACELP). The first portion of the audio signal is placed in a 671 LPC and De-emphasis synthesis unit for processing to obtain a production of the first portion of the audio signal with
119 time domain. Furthermore, the first audio signal portion is placed in the RMS module 650 to obtain a first value indicating a root mean square of the first audio signal portion. This first value (first RMS value) is represented in the plotting domain. The first RMS value is represented in the plotting domain, then placed in the noise level tracking unit 630.
Furthermore, the apparatus of Fig. 6 receives a second frame with a second portion of the audio signal comprising an MDCT spectrum and represented in an MDCT domain. The noise filling is conducted by means of a noise filling module 681, the frequency domain noise shaping is driven by a frequency domain noise shaping module 682, the transformation in the Time is driven by an iMDCT / OLA 683 module (OLA = aggregate with overlap) and the long-term forecast by a long-term forecast unit 684. The long-term prediction unit may, eg, comprise a delay buffer (not shown in Fig. 6).
The signal that is derived from the second portion of the audio signal is then placed in the RMS 660 module to obtain a second value indicating that a root-average square of the signal that is derived from the second portion of the audio signal is obtained. . This second value (the second RMS value)
120 it is still represented in the time domain. Unit 620 then transforms the second RMS value from the time domain to the domain. traced, here, the domain of LPC (ACELP). The second RMS value is plotted in the plotting domain, then placed in the noise level tracking unit 630.
In embodiments, level plotting is conducted in the excitation domain, but fading by TCX is conducted in the time domain.
While the background noise level is plotted during normal decoding, it can, for example, be used during packet loss as an indicator of an appropriate comfort noise level, with which the last received signal fades smoothly. level by level.
Deriving the trace level and applying the fading level are generally independent of each other and could be done in different domains. In the preferred embodiment, the level application is performed in the same domains as the level derivation, which leads to the same benefits as for ACELP, no gain compensation is needed, and that for TCX, the compensation for gain inverse to that of the level tap (see Fig. 6) is needed, and therefore the same gain tap can be used, as illustrated through Fig. 7.
121
In the text that follows, the compensation of a high-pass filter influence on the gain of LPC synthesis is described according to the embodiments.
Fig. 8 defines this approach. In particular, Fig. 8 illustrates a comfort noise level application during packet loss.
In Fig. 8, the high-pass gain filter unit 643, multiplication unit 644, fading unit 645, high-pass filter unit 646, fading unit 647, and combining unit 648 together form a first unit of reconstruction.
Furthermore, in Fig. 8, the background level provision unit 631 provides the noise level information. For example, the background level provision unit 631 can be implemented in the same way as the background level plotting unit 630 of Fig. 6.
Additionally, in Fig. 8, the synthesis of the
LPC Gain and De-emphasis unit 649 and multiplication unit 641 together for a second transformation unit 640.
Furthermore, in the
Fig. 8, the fading unit
642 represents a second rebuild unit.
In the embodiment of Fig. 8, the voiced and non-voiced arousal fade separately: The voice arousal fades to zero, but the non-voice arousal fades to zero.
122 fades towards the comfort noise level. Fig. 8 additionally presents a high-pass filter, which is introduced into the signal chain of the no-voice drive to suppress the low-frequency components for all cases except when the signal was classified as no-voice.
In order to model the influence of the high-pass filter, the level after LPC synthesis and de-emphasis is computed once with and once without the high-pass filter. The ratio of those two levels is then derived and used to alter the applied background level.
This is illustrated through Fig. 9. In particular, Fig. 9 presents advanced high pass gain compensation during ACELP concealment in accordance with one embodiment.
Instead of the current drive signal only a single pulse is used as input for this computation. This allows for reduced complexity as the impulse response decays rapidly and thus the RMS tapping can be done in a shorter time frame. In practice, only one sub-frame is used instead of the full frame.
According to one embodiment, the noise level tracking unit 130 is configured to determine a comfort noise level as the noise level information. The reconstruction unit 140 is configured to reconstruct the third audio signal portion depending on
123 of the noise level information, if said third frame of the plurality of frames is not received via the receiver interface 110 or if said third frame is received via the receiver interface 110 but is corrupted.
According to one embodiment, the noise level tracking unit 130 is configured to determine a comfort noise level as the noise level information. Reconstruction unit 140 is configured to reconstruct the third audio signal portion depending on the noise level information, if said third frame of the plurality of frames is not received via receiver interface 110 or if said third frame is receives through receiving interface 110 but is corrupted.
In one embodiment, the noise level tracking unit 130 is configured to determine a comfort noise level as the noise level information derived from a noise level spectrum, wherein said noise level spectrum It is obtained by applying the minimum statistics method. The reconstruction unit 140 is configured to reconstruct the third portion of the audio signal depending on a plurality of Linear Predictive Coefficients, if said third frame of the plurality of frames is not received via the receiver interface 110 or if said third frame is receives through receiving interface 110 but is corrupted.
124
In one embodiment, the (first and / or second) reconstruction unit 140, 141 may, eg, be configured to reconstruct the third audio signal portion depending on the noise level information and depending on the first portion of audio signal, if said third (fourth) frame of the plurality of frames is not received via receiver interface 110 or if said third (fourth) frame is received via receiver interface 110 but is corrupted.
According to one embodiment, the (first and / or second) reconstruction unit 140, 141 may, eg, be configured to reconstruct the third (or fourth) portion of the audio signal by attenuating or amplifying the first audio signal portion.
Fig. 14 illustrates an apparatus for decoding an audio signal. The apparatus comprises a receiver interface 110, wherein the receiver interface 110 is configured to receive a first frame comprising a first audio signal portion of the audio signal, and wherein the receiver interface 110 is configured to receive a second frame that comprises a second portion of the audio signal of the audio signal.
Furthermore, the apparatus comprises a noise level tracking unit 130, wherein the noise level tracking unit 130 is configured to determine the noise level information depending on at least one of the first
125 portion of the audio signal and the second portion of the audio signal (this means: depending on the first portion of the audio signal and / or the second portion of the audio signal), where the noise level information is represented in a plot domain.
Additionally, the apparatus comprises a first reconstruction unit 140 for the reconstruction of, in a first reconstruction domain, a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received through the receiving interface 110 or if said third frame is received through the receiving interface 110 but is corrupted, where the first reconstruction domain is different from or equal to the plotting domain.
Furthermore, the apparatus comprises a transformation unit 121 for transforming the noise level information from the plotting domain into a second reconstruction domain, if a fourth frame of the plurality of frames is not received via the receiving interface. 110 or if said fourth frame is received through receiving interface 110 but is corrupted, where the second reconstruction domain is different from the plotting domain, and where the second reconstruction domain is different from the first reconstruction domain, and
126
Additionally, the apparatus comprises a second reconstruction unit 141 for the reconstruction of, in the second reconstruction domain, a fourth audio signal portion of the audio signal depending on the noise level information is represented in the second reconstruction domain, if said fourth frame of the plurality of frames is not received via receiver interface 110 or if said fourth frame is received via receiver interface 110 but is corrupted.
According to some embodiments, the plotting domain can, eg, be where the plotting domain is a time domain, a spectral domain, an FFT domain, an MDCT domain, or an excitation domain. The first reconstruction domain can, eg, be the time domain, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain. The second reconstruction domain can, eg. , being the time domain, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain.
In one embodiment, the plotting domain may, eg, be the FFT domain, the first reconstruction domain may, eg, be the time domain, and the second reconstruction domain may, eg, be the domain of excitement.
127
In another embodiment, the plotting domain can, eg. , be the time domain, the first reconstruction domain may, eg, be the time domain, and the second reconstruction domain may, eg, be the excitation domain.
According to one embodiment, said first portion of the audio signal may, eg, be represented in a first input domain, and said second portion of the audio signal may, eg. , rendered in a second input domain. The transformation unit can, for example, be a second transformation unit. The device can, eg. , further comprising a first transformation unit for transforming the second portion of the audio signal or a value or a signal derived from the second portion of the audio signal from the second input domain to the plotting domain to obtain a second information from the signal portion. The noise level tracking unit can, for example, be configured to receive a first information of the signal portion is represented in the plotting domain, where the first information of the signal portion depends on the first signal portion. where the noise level tracking unit is configured to receive the second portion of the signal is represented in the plotting domain, and where the noise level tracking unit is configured to determine the
128 Noise level information depending on the first information of the signal portion is represented in the plotting domain and depending on the second portion of the signal information is represented in the plotting domain.
According to one embodiment, the first input domain can, eg, be the excitation domain, and the second input domain can, eg, be the MDCT domain.
In another embodiment, the first input domain may, eg, be the MDCT domain, and where the second input domain may, eg, be the MDCT domain.
If, for example, a signal is represented in a time domain, it can, for example, be represented by time domain samples of the signal. Or, for example, if a signal is represented in a spectral domain, it can, for example, be represented by spectral samples from a spectrum of the signal.
In one embodiment, the plotting domain can, eg. , be the FFT domain, the first reconstruction domain may, eg, be the time domain, and the second reconstruction domain may, eg, be the excitation domain.
In another embodiment, the plotting domain can, eg. , be the time domain, the first domain
129 Reconstruction can, eg, be the time domain, and the second reconstruction domain can, eg. , be the domain of excitement.
In some of the embodiments, the units illustrated in Fig. 14 can, for example, be configured as described for Fig. 1A, IB, 1C and ID.
With regard to particular embodiments, in, for example, a low ranking mode, an apparatus according to one embodiment may, for example, receive ACELP frames as input, which are represented in an excitation domain, and which are then transformed into a time domain through LPC synthesis. Furthermore, in the low-rating mode, the apparatus according to one embodiment may, for example, receive TCX frames as input, which are represented in an MDCT domain, and then transformed into a time domain by means of of a reverse MDCT.
The trace is then driven into an FFT Domain, where the FFT signal is derived from the time domain signal by conducting an FFT (Fast Fourier Transform). The plotting can, for example, be conducted by conducting a minimal statistics method, separated from all spectral lines to obtain a spectrum of the comfort noise.
The concealment is then conducted by conducting the level tap based on the comfort noise spectrum.
130
The level derivation is conducted on the basis of the comfort noise spectrum. Level conversion in the time domain is conducted for FD TCX PLC. A fading in the time domain is driven. The level tap in the excitation domain is driven for ACELP PLC and for TD TCX PLC (as
ACELP). Then a fading is conducted in the excitation domain.
The following list summarizes the above:
low index:
•entry:
or acelp (excitation domain -> time domain, via LPC synthesis) or tcx (mdct domain -> time domain, via reverse MDCT) • tracing:
or fft domain, which is derived from time domain by means of FFT or minimal statistics, separated from all spectral lines -> comfort noise spectrum • concealment:
o the level derivation based on the comfort noise spectrum or the level conversion in time domain for
131
-FD TCX PLC
-> fading in the time domain or level conversion in the excitation domain for
ACELP PLC
TD TCX PLC (As ACELP)
-> fading in the excitation domain
In, for example, a high-speed mode, it can, for example, receive TCX frames as input, which are represented in the MDCT domain, and which are then transformed into the time domain by means of an inverse MDCT.
The plot can then be conducted in the time domain. Tracing can, for example, be conducted by conducting a minimum statistical method based on the energy level to obtain a comfort noise level.
For concealment, for FD TCX PLC, the level can be used as is and only time domain fading can be conducted. For TD TCX PLC (Like ACELP), level conversion in the excitation domain and fading in the excitation domain is driven.
The following list summarizes the above:
high index;
•entry:
or tcx (mdct domain -> time domain, via reverse MDCT)
132 • layout:
o time domain o minimal statistic on energy level
-> comfort noise level • concealment:
or use level as is
FD TCX PLC
-> time domain fading or level conversion in excitation domain for • TD TCX PLC (As ACELP)
-> fading in the excitation domain
The FFT domain and the MDCT domain are both spectral domains, while the excitation domain is some kind of time domain.
In accordance with one embodiment, the first reconstruction unit 140 may, eg, be configured to reconstruct the third audio signal portion by driving a first fading into a noise-like spectrum. The second reconstruction unit 141 may, eg, be configured to reconstruct the fourth portion of the audio signal by driving a second fading to a noise-like spectrum and / or a second fading to an LTP gain. Furthermore, the first rebuild unit 140 and the second rebuild unit 141
133 you can, eg. , configured to drive the first fading and the second fading to a noise-like spectrum and / or a second fading of an LTP gain with the same fading rate.
Now, the flexible spectral shape allocation of comfort noise is considered.
In order to achieve the flexible allocation to comfort noise during burst packet loss, as a first step, the finding of appropriate LPC coefficients which represent the background noise can be conducted. These LPC coefficients can be derived during active speech using a minimal statistical method to find the background noise spectrum and then calculate the LPC coefficients from it using an arbitrary algorithm for LPC derivation. known from the bibliography. Some embodiments, for example, can directly convert the background noise spectrum into a representation that can be used directly for FDNS in the MDCT domain.
Comfort noise fading can be performed in the ISF domain (also applicable in the LSF domain; LSF line spectral frequency);
/ curreníf ^ J - + (1 “Ci) · ptmean [¿] i - 0 ... 16 (26) when setting pt<sub>piss</sub>Comfort noise is described at appropriate LP coefficients.
134
With respect to the previously described flexible spectral shape mapping of comfort noise, a more general embodiment is illustrated through Fig. 11.
Fig. 11 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in accordance with one embodiment.
The apparatus comprises a receiver interface 1110 for receiving one or more frames, a coefficient generator 1120, and a signal reconstructor 1130.
The coefficient generator 1120 is configured to determine, if a current frame of the one or more frames is received through the receiving interface 1110 and if the current frame received by the receiving interface 1110 is not corrupted / erroneous, one or more of the first audio signal coefficients, comprised by the current frame, wherein said one or more of the first audio signal coefficients indicate a characteristic of the encoded audio signal, and one or more noise coefficients indicating a background noise of the encoded audio signal. Furthermore, the coefficient generator 1120 is configured to generate one or more second coefficients of the audio signal, depending on the one or more of the first audio signal coefficients and depending on the one or more noise coefficients, if the current frame not received via receiving interface
135
1110 or if the current frame received by the receiving interface 1110 is corrupted / erroneous.
The audio signal reconstructor 1130 is configured to reconstruct a first portion of the reconstructed audio signal depending on the one or more of the first audio signal coefficients, if the current frame is received through the receiver interface 1110 and if the current frame received by the receiving interface 1110 is not corrupted. Furthermore, the audio signal reconstructor 1130 is configured to reconstruct a second portion of the reconstructed audio signal depending on the one or more second coefficients of the audio signal, if the current frame is not received via the receiving interface. 1110 or if the current frame received by the receiving interface 1110 is corrupted.
The determination of a background noise is well known in the art (see, for example, [MarOl]: Rainer Martín, Noise power spectral density estimation based on optimal smoothing and minimal statistics, IEEE Transactions on Speech and Audio Processing 9 (2001 ), no. 5, 504-512), and in one embodiment, the apparatus proceeds accordingly.
In some of the embodiments, the one or more of the first audio signal coefficients may, eg, be one or more coefficients of the linear predictive filter of the encoded audio signal. In some of the embodiments, the one or more of the first coefficients of
136 The audio signal may, for example, be one or more coefficients of the linear predictive filter of the encoded audio signal.
It is well known in the art how to reconstruct an audio signal, eg a speech signal, from linear predictive filter coefficients or from mimic spectral pairs (see, for example, [3GP09c]: Speech codee speech processing functions; adaptive multi-rate - wide band (AMRWB) speech codee; transcoding functions, 3GPP TS 26.190, 3rd Generation Partnership Project, 2009), and in one embodiment, the reconstructor signal proceeds accordingly.
According to one embodiment, the one or more noise coefficients may, eg, be one or more coefficients of the linear predictive filter indicating the background noise of the encoded audio signal. In one embodiment, the one or more coefficients of the linear predictive filter may, eg, represent a spectral shape of the background noise.
In one embodiment, the coefficient generator 1120 may, eg. , be configured to determine the one or more second portions of the audio signal such that the one or more second portions of the audio signal are one or more coefficients of the linear predictive filter of the reconstructed audio signal, or so that the one or more of the first audio signal coefficients are one or more spectral pairs mimicking the reconstructed audio signal.
137
According to one embodiment, the coefficient generator 1120 may, eg. , be configured to generate the one or more second coefficients of the audio signal by applying the formula:
/ current | /] <sup>=</sup> ® '/ Zast ^ J 4 * (1 λ)' ptmean [í] where fcurrent [1] indicates one of the one or more second coefficients of the audio signal, where fz<sub>ace</sub>t [í] indicates one of the one or more of the first audio signal coefficients, where pt<sub>mean</sub>[i] is one of the one or more noise coefficients, where a is a real number with 0 <a <1, and where i is an index.
According to one embodiment, -fiastC-i] indicates a coefficient of the linear predictive filter of the encoded audio signal, and where fcurrenti] indicates a coefficient of the linear predictive filter of the reconstructed audio signal.
In one embodiment, pt<sub>I</sub>an [i] can, eg. , be a coefficient of the linear predictive filter indicating the background noise of the encoded audio signal.
According to one embodiment, the coefficient generator 1120 can, eg. , set to generate at least 10 second coefficients of the audio signal as the one or more second coefficients of the audio signal.
In one embodiment, the coefficient generator 1120 may, eg. , be configured to determine, if the
138 current frame of the one or more frames is received through the receiving interface 1110 and if the current frame received by the receiving interface 1110 is not corrupted, the one or more noise coefficients by determining a noise spectrum of the signal encoded audio.
In the text that follows, the fading of the MDCT spectrum to Noise weighted is considered prior to FDNS Application.
Instead of randomly modifying the sign of an MDCT (sign scrambling interference) basket, the entire spectrum is filled with weighted noise, which is shaped using the FDNS. To avoid an instantaneous change in spectrum characteristics, a crossfade is applied between sign scrambling interference and noise filling. Cross fading can be seen as follows:
for (i = 0; i <L_frame; i ++) {if (old_x [i]! = 0) {x [i] = (1 - 'cum_damping) * noise [i] + cum_damping * random_sign () * x_old [i ];
} }
where:
cum_damping is the (absolute) damping factor - it reduces from frame to frame, starting from 1 and decreasing towards 0
139 x_old is the spectrum of the last frame received random_sign returns to 1 or -1 noise contains a random vector (weighted noise) that is increased such that its root mean square (RMS) is similar to the last good spectrum.
The term random_sign () * old_x [i] characterizes the sign scrambling interference process to Randomize the phases and thus avoid harmonic repetitions.
Subsequently, another normalization of the energy level could be performed after the crossfade to ensure that the summed energy does not drift during the correlation of the two vectors.
In accordance with the embodiments, the first reconstruction unit 140 may, eg, be configured to reconstruct the third audio signal portion depending on the noise level information and depending on the first audio signal portion. In a particular embodiment, the first reconstruction unit 140 may, eg, be configured to reconstruct the third audio signal portion by attenuating or amplifying the first audio signal portion.
In some of the embodiments, the second reconstruction unit 141 may, eg, be configured to reconstruct the fourth portion of the audio signal.
140 depending on the noise level information and depending on the second portion of the audio signal. In a particular embodiment, the second reconstruction unit 141 can, eg, be configured to reconstruct the fourth portion of the audio signal by attenuating or amplifying the second portion of the audio signal.
With respect to the previously described fading of the MDCT spectrum to noise weighted before FDNS Application, a more general embodiment is illustrated through Fig. 12.
Fig. 12 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in accordance with one embodiment.
The apparatus comprises a receiver interface 1210 for receiving one or more frames comprising information on a plurality of audio signal samples from a spectrum of the audio signal from the encoded audio signal, and a processor 1220 for generating the audio signal. reconstructed.
The processor 1220 is configured to generate the reconstructed audio signal by fading a spectrum modified to a weighted spectrum, if a current frame is not received via the receiver interface 1210 or if the current frame is received via the receiver interface. 1210 but is corrupt, where the modified spectrum comprises a
141 plurality of modified signal samples, wherein, for each of the modified signal samples of the modified spectrum, an absolute value of said modified signal sample is equal to an absolute value of one of the audio signal samples of the spectrum of the audio signal.
Furthermore, the processor 1220 is configured not to fade the modified spectrum to the weighted spectrum, if the current frame of the one or more frames is received through the receiving interface 1210 and if the current frame received by the receiving interface 1210 is not corrupted. .
According to one embodiment, the weighted spectrum is a noise-like spectrum.
In one embodiment, the noise-like spectrum represents weighted noise.
According to one embodiment, the noise-like spectrum receives its shape.
In one embodiment, the shape of the noise-like spectrum depends on a spectrum of the audio signal from a previously received signal.
According to one embodiment, the noise-like spectrum is shaped depending on the spectrum shape of the audio signal.
In one embodiment, processor 1220 employs a skew factor to shape the noise-like spectrum.
142
According to one embodiment, the processor
1220 use the formula shaped_noise [i] = noise * power (tilt_factor, i / N)
<td>in</td><td>where</td><td>N indicates the</td><td>quantity</td><td>of</td><td>samples,</td>
<td>in</td><td>where</td><td>i is an indi</td><td>EC,</td><td></td><td></td>
<td>in</td><td>where</td><td>0 <= i <N,</td><td>with tilt</td><td colspan="2">factor> 0,</td>
<td>in</td><td>where</td><td>power is a</td><td>function</td><td>of</td><td>power.</td>
<td>Yes</td><td>he</td><td>tilt factor</td><td colspan="2">is minor</td><td>what 1 does this mean</td>
attenuation with increasing i. If the tilt_factor is greater than 1 it means amplification. With increasing i.
According to another embodiment, processor 1220 may employ the formula shaped_noise [i] = noise * (1 + i / (Nl) * (tilt_factor-l)) where N indicates the number of samples, where i is an index, where 0 <= i <N, with tilt_factor> 0.
According to one embodiment, the processor 1220 is configured to generate the modified spectrum, by changing a sign of one or more of the audio signal samples from the audio signal spectrum, if the current frame is not received by middle of the receiving interface 1210 or if the current frame received by the receiving interface 1210 is corrupted.
143
In one embodiment, each of the audio signal samples in the audio signal spectrum is represented by a real number but not by an imaginary number.
According to one embodiment, the audio signal samples from the spectrum of the audio signal are represented in a Domain of
Individual Modified.
In another embodiment, audio from the spectrum of the
Cosine Transformation the audio signal samples are represented in a
Sine Transformation Domain
Individual Modified.
According to one embodiment, the processor
1220 is configured to generate the modified spectrum by employing a random sign function that randomly or pseudo-randomly produces both a first and a second value.
In one embodiment, processor 1220 is configured to fade the modified spectrum to the weighted spectrum by subsequently reducing an attenuation factor.
According to one embodiment, processor 1220 is configured to fade the modified spectrum to the weighted spectrum by subsequently increasing an attenuation factor.
In one embodiment, if the current frame is not received via the receiving interface 1210 or if the frame
144 current received by receiving interface 1210 is corrupted, processor 1220 is configured to generate the reconstructed audio signal using the formula:
x [i] = (l-cum_damping) * noise [i] + cum_damping * random_sign () * x_old [i] where i is an index, where x [i] indicates a sample of the reconstructed audio signal, in where cum_damping is an attenuation factor, where x_old [i] indicates one of the audio signal samples from the audio signal spectrum of the encoded audio signal, where random_sign () returns to -1, and where noise is a random vector indicating the weighted spectrum.
Some embodiments continue a TCX LTP operation. In those embodiments, the TCX LTP operation is continued during concealment with the LTP parameters (LTP lag and LTP gain) derived from the last good frame.
LTP operations can be summarized as:
Feed the LTP delay buffer based on previously derived output.
Based on LTP lag: choose the appropriate signal portion from the LTP delay buffer that is used as the LTP contribution to shape the current signal.
Increase this LTP contribution again with the use of LTP gain.
145
Add this newly increased contribution of LTP to the LTP input signal to generate the LTP output signal.
Different methods could be considered with respect to time, when LTP delay buffer update is done:
As the first LTP operation on frame n using the output from last frame n-1. This updates the LTP delay buffer in frame n to be used during LTP processing in frame n.
As the last LTP operation on frame n using the output of current frame n. This updates the LTP delay buffer in frame n to be used during LTP processing in frame n + 1.
In the text that follows, we consider decoupling the feedback loop from the TCX LTP.
Decoupling of the TCX LTP feedback loop prevents the introduction of additional noise (resulting from noise substitution applied to the LPT input signal) during each feedback loop of the LTP decoder when in stealth mode .
Fig. 10 illustrates this decoupling. In particular, Fig. 10 shows the decoupling of the LTP feedback loop during concealment (bfi = l).
146
FIG. 10 illustrates a delay buffer 1020, a sample selector 1030, and a sample processor 1040 (sample processor 1040 is indicated by the dotted line).
Regarding timing, when updating the LTP 1020 delay buffer, some embodiments proceed as follows:
For normal operation: To update the LTP delay buffer 1020 as the first LTP operation might be preferred, as the summed output signal is usually stored persistently. With this approach, an exclusive buffer can be omitted.
For decoupled operation: To update the LTP delay buffer 1020 as the last LTP operation might be preferred, since the contribution of LTP to the signal is usually stored temporarily. With this approach, the transient contribution of the LTP signal is maintained. In an implementation way, this contribution from the LTP buffer could be made persistent.
Assuming the latter approach is used in any case (normal operation and stealth), the embodiments, may eg implement the following:
During normal operation: The LTP decoder time domain signal output after
147 this addition to the LTP input signal is used to feed the LTP delay buffer.
During concealment: The time domain signal output from the LTP decoder prior to its incorporation into the LTP input signal is used to feed the LTP delay buffer.
Some embodiments fade the LTP TCX gain toward zero. In such an embodiment, the LTP TCX gain can eg fade towards zero with a certain flexible signal fading factor. This can, for example, be done iteratively, for example, according to the following pseudo-code:
gain = gain_past * damping;
[· · ·] Gain_past = gain;
where:
gain is the gain of the TCX LTP decoder in the current frame;
gain_past is the TCX LTP decoder gain applied in the previous frame;
damping is the (relative) fade factor.
Fig. ID illustrates an apparatus according to in a further embodiment, wherein the apparatus further comprises a long-term prediction unit 170 comprising a delay buffer 180. The long-term prediction unit 170 is configured to generate a
148 processed signal depending on the second portion of the audio signal, depending on an input to the delay buffer stored in the delay storage device 180 and depending on a long-term prediction gain. Furthermore, the long-term prediction unit is configured to fade the long-term prediction gain toward zero, if said third frame of the plurality of frames is not received via the receiving interface 110 or if said third frame is received. through the receiving interface 110 but it is corrupted.
In other embodiments (not shown), the long-term prediction unit can, e.g., be configured to generate a processed signal depending on the first audio signal portion, depending on an input to the delay buffer stored in the delay storage device and depending on a long-term prediction gain.
In Fig. ID, the first reconstruction unit 140 may, eg, generate the third audio signal portion additionally depending on the processed signal.
In one embodiment, the long-term prediction unit 170 may, eg, be configured to fade the long-term prediction gain toward zero, where a speed at which the long-term prediction gain
149 Term fades to zero depends on a fading factor.
Alternatively or additionally, the long-term prediction unit 170 may, eg, be configured to update the input delay storage device 180 by storing the generated processed signal in the delay storage device 180 if said third frame of the plurality of frames is not received via the receiver interface 110 or if said third frame is received via the receiver interface 110 but is corrupted.
With respect to the above-described use of TCX LTP, a more general embodiment is illustrated through Fig. 13.
Fig. 13 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal.
The apparatus comprises a receiver interface 1310 for receiving a plurality of frames, a delay buffer 1320 for storing audio signal samples of the decoded audio signal, a sample selector 1330 for selecting a plurality of selected audio signal samples from the audio signal samples stored in the delay storage device 1320, and a sample processor 1340 to process the samples
150
<td>selected from</td><td>signal</td><td>audio</td><td>pair.</td><td colspan="3">to get samples</td><td>of</td>
<td>audio signal</td><td colspan="2">rebuilt</td><td>of</td><td colspan="2">the sign of</td><td colspan="2">Audio</td>
<td>reconstructed.</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>The selector</td><td>of</td><td>samples</td><td> 1330</td><td>I know</td><td>configure</td><td colspan="2">for</td>
<td colspan="3">select, if a current frame</td><td>I know</td><td>receives</td><td>through</td><td>of</td><td>the</td>
<td>receiving interface</td><td> 1310</td><td colspan="2">and if the painting</td><td>current</td><td>received</td><td>by</td><td>the</td>
receiver interface 1310 is not corrupted, the plurality of selected audio signal samples from the audio signal samples stored in delay storage device 1320 depending on a tone lag information comprised by the current frame. Furthermore, the sample selector 1330 is configured to select, whether the current frame is not received via the receiving interface 1310 or whether the current frame received by the receiving interface 1310 is corrupted, the plurality of selected audio signal samples from the audio signal samples stored in the delay storage device 1320 depending on a lag per tone information comprised by another frame previously received by the receiver interface 1310.
According to one embodiment, the sample processor 1340 may, eg, be configured to obtain the reconstructed audio signal samples, if the current frame is received through the receiver interface 1310 and if the current frame received by the receiving interface 1310 is not
151 corrupted, by re-boosting selected samples of audio signal depending on the gain information comprised by the current frame. Furthermore, the sample selector 1330 can, eg, be configured to obtain the reconstructed audio signal samples, if the current frame is not received via the receiving interface 1310 or if the current frame received by the receiving interface 1310 is corrupted, by re-increasing the selected audio signal samples depending on the gain information comprised by said other frame previously received by the receiving interface 1310.
In one embodiment, the sample processor 1340 may, eg. , set to obtain the reconstructed audio signal samples, if the current frame is received through the receiving interface 1310 and if the current frame received by the receiving interface 1310 is not corrupted, by multiplying the selected signal samples from audio and a value depending on the information of the gain understood, by the current frame. Furthermore, the sample selector 1330 is configured to obtain the reconstructed audio signal samples, if the current frame is not received via the receiver interface 1310 or if the current frame received by the receiver interface 1310 is corrupted, by the multiplication of the selected audio signal samples and a value depending on the information of
152 the gain comprised by said other frame previously received by the receiving interface 1310.
According to one embodiment, the sample processor 1340 can, eg, be configured to store the reconstructed audio signal samples in the delay storage device 1320.
In one embodiment, the sample processor 1340 can, eg, be configured to store the reconstructed audio signal samples in the delay storage device 1320 before receiving an additional frame through the receiver interface 1310.
According to one embodiment, the sample processor 1340 may, eg, be configured to store the reconstructed audio signal samples in the delay storage device 1320 after receiving an additional frame through the receiver interface. 1310.
In one embodiment, the sample processor 1340 may, eg. , set to re-boost selected audio signal samples depending on the gain information to obtain newly augmented audio signal samples and by combining the newly augmented audio signal samples with input audio signal samples to obtain the processed audio signal samples.
153
According to one embodiment, the sample processor 1340 can, eg. , be configured to store the processed audio signal samples, indicating the combination of the newly augmented audio signal samples and the input audio signal samples, in the delay storage device 1320, and do not store the samples audio signal levels again increased in delay storage device 1320, if the current frame is received through the receiving interface 1310 and if the current frame received by the receiving interface 1310 is not corrupted. Furthermore, the sample processor 1340 is configured to store the newly augmented audio signal samples in the delay storage device 1320 and not to store the processed audio signal samples in the delay storage device 1320, if the frame The current frame is not received via the receiving interface 1310 or if the current frame received by the receiving interface 1310 is corrupt.
According to another embodiment, the sample processor 1340 may, eg, be configured to store the processed audio signal samples in the delay storage device 1320, if the current frame is not received via the interface. receiver 1310 or if the current frame received by receiver interface 1310 is corrupted.
154
<td>In</td><td>a</td><td>shape</td><td>of realization</td><td>, he</td><td>selector</td><td>of</td><td>samples</td><td> 1330</td>
<td>may,</td><td>by</td><td>ex. ,</td><td>set up</td><td>for</td><td>get</td><td>the</td><td colspan="2">samples of</td>
<td>signal</td><td>of</td><td>Audio</td><td>rebuilt</td><td>to the</td><td>return</td><td>to</td><td>increase</td><td>the</td>
selected samples of audio signal depending on a modified gain, where the modified gain is defined according to the formula:
gain = gain_past * damping;
where gain is the modified gain, where the sample selector 1330 can eg. , set to set gain_past to increase after the gain has been calculated, and where joint is a real number.
In accordance with one embodiment, the sample selector 1330 may, eg, be configured to calculate the modified gain.
In one embodiment, articulation may, eg, be defined according to: 0 <articulation <1.
According to one embodiment, the modified gain may, eg, be set to zero, if at least a predefined number of frames has not been received through the receiving interface 1310 since a frame was last received. once over the receiving interface 1310.
In the text that follows, the fading rate is considered. There are several concealment modules that apply a certain kind of fading. While the speed of this fading could be chosen in the form
155 different across those modules, it is beneficial to use the same fade rate for all cloaking modules for a kernel (ACELP or TCX). For example:
For ACELP, the same fading rate should be used, in particular, for the flexible codebook (by altering the gain), and / or for the innovative codebook signal (by altering the gain).
Also, for TCX, the same fade rate should be used, in particular, for the time domain signal, and / or for LTP gain (fades to zero), and / or for LPC weighing (fades to zero). to one), and / or for LP coefficients (fading to background spectral shape), and / or for weighted noise fading.
It might also be preferable to use the same fading rate for ACELP and TCX, but due to the different nature of the cores it could also be chosen to use different fading rates.
This fading rate could be static, but is preferably flexible to signal characteristics. For example, the fading rate may, eg, depend on the LPT stability factor (TCX) and / or on a rating, and / or on a number of consecutive lost frames.
156
The rate of fading can, for example, be determined depending on the attenuation factor, which could occur absolutely or relatively, and which could also change with the passage of time during a certain fading.
In the embodiments, the same fading rate is used to fade the LTP gain as for the weighted noise fading.
An apparatus, method, and computer program have been provided for generating a comfort noise signal as described above.
Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a characteristic of a method step. Analogously, the aspects described in the context of a method step also represent a description of a corresponding block or element or characteristic of a corresponding apparatus.
The decomposed signal of the invention can be stored on a digital storage medium or it can be transmitted on a transmission medium such as a wireless transmission medium or a cable transmission medium such as
Internet.
157
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or software. The implementation can be carried out with the use of a digital storage medium, for example a floppy disk, a DVD, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, which have control signals capable of read electronically in them, which cooperate (or are capable of cooperating) with a computer system programmed in such a way that the respective method is performed.
Some embodiments according to the invention comprise a non-transient data vehicle having control signals read by electronic means, which are capable of cooperating with a programmable system for computers, such that one of the described methods is performed in this document.
In general, the embodiments of the present invention can be implemented as a computer program product with a program code, the program code is operative to perform one of the methods when the computer program product runs on a computer . The program code can for example be stored in a vehicle that is read through the machine.
Other embodiments include computer program to perform one of the methods listed above.
158 described in this document, stored in a vehicle that can be read on the machine.
In other words, an embodiment of the method of the invention is therefore a computer program that has a program code to perform one of the methods described in this document, when the computer program is run on a computer. .
A further embodiment of the method of the invention is, therefore, a data vehicle (or a digital storage medium, or a computer reading medium) comprising, registered in it, the computer program to perform a of the methods described in this document.
A further embodiment of the method of the invention is, therefore, a data stream, a sequence of signals representing the computer program to perform one of the methods described in this document. The data throughput or the signal sequence can, for example, be configured to be transferred over a data communication connection, for example over the internet.
A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured for or adapted to perform one of the methods described herein.
159
<td>Form</td><td>of</td><td>realization</td><td>additional</td><td>understands</td><td>a</td>
<td>computer that</td><td colspan="2">has installed</td><td>in her</td><td>the program</td><td>for</td>
<td colspan="2">computers for</td><td>perform one</td><td>of the</td><td>methods that</td><td>I know</td>
described in this document.
In some of the embodiments, a programmable logic device (eg, a field-programmable opening arrangement) may be used to perform some or all of the functionality of the methods described herein. In some of the embodiments, a field programmable aperture arrangement may cooperate with a microprocessor in order to perform one of the methods described herein. In general, the methods are preferably performed with any hardware apparatus.
The embodiments described above are merely illustrative for the principles of the present invention. It is understood that modifications and variations to the provisions and details described in this document will be obvious to others of skill in the art. It is therefore understood that they are limited only by the scope of the patent pending claims and not by the specific details presented by way of description and explanation of the embodiments herein.
160 [3GP09a] [3GP09b] [3GP09c] [3GP12a] [3GP12b] [3GP12d]
References
3GPP; Technical Specification Group Services and System Aspects, Extended adaptive multi-rate wideband (AMR-WB +) codee, 3GPP TS 26.290, 3rd Generation Partnership Project, 2009.
Extended adaptive multi-rate - wideband (AMR-WB +) codee; floating-point ANSI-C code, 3GPP TS 26.304, 3rd Generation Partnership Project, 2009.
Speech codee speech processing functions; adaptive multi-rate - wideband (AMRWB) speech codee; transcoding functions, 3GPP TS 26.190, 3rd Generation Partnership Project, 2009.
Adaptive multi-rate (AMR) speech codee; error concealment of lost frames (release 11), 3GPP TS 26.091, 3rd Generation Partnership Project, Sep 2012.
Adaptive multi-rate (AMR) speech codee; transcoding functions (release 11), 3GPP TS 26.090, 3rd Generation Partnership Project, Sep 2012. [3GP12c], ANSI-C code for the adaptive multi-rate wideband (AMR-WB) speech codee, 3GPP TS 26.173, 3rd Generation Partnership Project , Sep 2012.
ANSI-C code for the floating-point adaptive multirate (AMR) speech codee (releasell), 3GPP TS
161 [3GP12e] [3GP12f] [3GP12g] [BJH06] [BP06] [Coh03]
26.104, 3rd Generation Partnership Project, Sep 2012.
General audio codee audio processing functions; Enhanced aacPlus general audio codee; additional decoder tools (release 11), 3GPP TS 26.402, 3rd Generation Partnership Project, Sep 2012.
Speech codee speech processing functions; adaptive multi-rate - wideband (amr-wb) speech codee; ansi-c code, 3GPP TS 26.204, 3rd Generation Partnership Project, 2012.
Speech codee speech processing functions; adaptive multi-rate - wideband (AMR-WB) speech codee; error concealment of erroneous or lost frames, 3GPP TS 26.191, 3rd Generation Partnership Project, Sep 2012.
I. Batina, J. Jensen, and R. Heusdens, Noise power spectrum estimation for speech enhancement using an autoregressive model for speech power spectrum dynamics, in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. 3 (2006), 1064-1067.
A. Borowicz and A. Petrovsky, Minima controlled noise estimation for klt-based speech enhancement, CD-ROM, 2006, Italy, Florence.
I. Cohén, Noise spectrum estimation in adverse environments: Improved minima controlled recursive
162 averaging, IEEE Trans. Speech Audio Process. 11 (2003), no. 5, 466-475.
[CPK08] Choong Sang Cho, Nam In Park, and Hong Kook Kim, A packet loss concealment algorithm robust to burst packet loss for cell-type speech coders, Tech. Report, Korea Enectronics Technology Institute, Gwang Institute of Science and Technology, 2008 , The 23rd International Technical Conference on Circuits / Systems, Computers and Communications (ITC-CSCC 2008).
[Dob95] G. Doblinger, Computa ti onally efficient speech enhancement by minimal spectral tracking in subbands, in Proc. Eurospeech (1995), 1513-1516.
[EBU10] EBU / ETSI JTC Broadcast, Digital audio broadcasting (DAB); transport of advanced audio coding (AAC) audio, ETSI TS 102 563, European Broadcasting Union, May 2010.
[EBU12] Digital radio mondiale (DRM); system specification, ETSI ES 201 980, ETSI, Jun 2012.
[EH08] Jan S. Erkelens and Richards Heusdens, Tracking of Nonstationary Noise Based on Data-Driven Recursive Noise Power Estimation, Audio, Speech, and Language Processing, IEEE Transactions on 16 (2008), no. 6,
1112 -1123.
163 [ΕΜ84] Y. Ephraim and D. Malah, Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator, IEEE Trans. Acoustics, Speech and Signal Processing 32 (1984), no. 6, 1109-1121.
[EM85] Speech enhancement using a minimum mean-square
<td></td><td>log-spectral error Acoustics, Speech 443-445.</td><td>amplitude and Signal</td><td>estimator, IEEE Processing 33</td><td>Trans. (1985),</td>
<td>[Gan05]</td><td>S. Gannot, Speech</td><td colspan="2">enhancement: Application</td><td>of the</td>
kalman filter in the estimate-maximize (em framework), Springer, 2005.
[HE95] HG Hirsch and C. Ehrlicher, Noise estimation techniques for robust speech recognition, Proc. IEEE Int. Conf. Acoustics, Speech, Signal Processing, no. pp. 153-156, IEEE, 1995.
[HHJ10] Richard C. Hendriks, Richard Heusdens, and Jesper Jensen, MMSE based noise PSD tracking with low complexity, Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on, Mar 2010, pp. 4266 -4269.
[HJH08] Richard C. Hendriks, Jesper Jensen, and Richard Heusdens, Noise tracking using dft domain subspace decompositions, IEEE Trans. Audio, Speech, Lang. Process. 16 (2008), no. 3, 541-553.
164 [ΙΕΤ12] [ISO09] [ITU03] [ITU05] [ITU06a] [ITU06b]
IETF, Definition of the Opus Audio Codee, Tech. Report RFC 6716, Internet Engineering Task Force, Sep 2012.
ISO / IEC JTC1 / SC29 / WG11, Information technology coding of audio-visual objeets - part 3: Audio, ISO / IEC IS 14496-3, International Organization for Standardization, 2009.
ITU-T, Wideband coding of speech at around 16 kbit / s using adaptive multi-rate wideband (amr-wb), Recommendation ITU-T G.722.2, Telecommunication Standardization Sector of ITU, Jul 2003.
Low-complexity coding at 24 and 32 kbit / s for hands-free operation in systems with low frame loss, Recommendation ITU-T G.722.1,
Telecommunication Standardization Sector of ITU, May 2005.
G. 722 Appendix III: A high-complexity algorithm for packet loss concealment for G. 722, ITU-T Recommendation, ITU-T, Nov 2006.
G. 729.1: G.729-based embedded variable bit-rate coder: An 8-32 kbit / s scalable wideband coder bitstream interoperable with g. 729, Recommendation ITU-T G.729.1, Telecommunication Standardization Sector of ITU, 'May 2006.
165 [ITU07]
G. 722 Appendix
IV: A low-complexity algorithm for packet loss concealment with G. 722,
ITU-T
Recommendation,
ITU-T, Aug 2007.
[ITU08a]
G. 718: Frame error robust narrow-band and wideband embedded variable bit-rate coding of speech and audio from 8-32 kbit / s, Recommendation ITU-T G.718,
Telecommunication Standardization
Sector of ITU,
Jun 2008.
[ITUO8b]
G. 719: Low-complexity, full-band audio coding for high-quality, conversational applications, [ITU12] [LS01]
Recommendation
ITU-T G.719,
Telecommunication
Standardization
G. 729: Coding
Sector of ITU, Jun of speech at
2008 .
kbit / s using conjugate-structure algebraic-code-excited prediction (cs-acelp), Recommendation ITU-T linear
G.729,
Telecommunication Standardization Sector of
June 2012.
Pierre Lauber and Ralph Sperschneider,
AND YOU,
Error concealment for compressed digital audio,
Audio
Engineering Society Convention 111, no.
5460, Sep
2001.
[MarOl] Rainer Martín, Noise power spectral density estimation based on optimal smoothing and minimum statistics, IEEE Transactions on Speech and Audio
Processing 9 (2001), no. 5, 504-512.
166
<td>[Mar03]</td><td>Statistical methods for the enhancement of noisy speech, International Workshop on Acoustic Echo and Noise Control (IWAENC2003), Technical University of Braunschweig, Sep 2003.</td>
<td>[MC99]</td><td>R. Martín and R. Cox, New speech enhancement techniques for low bit rate speech coding, in Proc. IEEE Workshop on Speech Coding (1999), 165-167.</td>
<td>[MCA99]</td><td>D. Malah, RV Cox, and AJ Accardi, Tracking speech-presence uncertainty to improve speech enhancement in nonstationary noise environments, Proc. IEEE Int. Conf. On Acoustics Speech and Signal Processing (1999), 789-792.</td>
[ΜΕΡ01] Nikolaus Meine, Bernd Edler, and Heiko Purnhagen,
Error protection and concealment for HILN MPEG-4 parametric audio coding, Audio Engineering Society
Convention 110, no. 5300, May 2001.
[MPC89] Y. Mahieux, J.-P. Petit, and A. Charbonnier,
Transform coding of audio signáis using correlation between successive transform blocks, Acoustics,
Speech, and Signal Processing, 1989. ICASSP-89., 1989 International Conference on, 1989, pp. 20212024 vol.3.
[NMR + 12]
Max Neuendorf, Markus Multrus
Nikolaus Rettelbach,
Guillaume Fuchs,
Julien Robilliard, Jérémie
Lecomte, Stephan Wilde, Stefan Bayer, Sascha Disch,
167 [PKJ + 11] [QD03] [RL06] [SFB00]
Christian Helmrich, Roch Lefebvre, Philippe Gournay, Bruno Bessette, Jimmy Lapierre, Kristopfer Kjórling, Heiko Purnhagen, Lars Villemoes, Werner Oomen, Erik Schuijers, Kei Kikuiri, Toru Chinen, Takeshi Norimatsu, Chong Kok Seng, Eunmi Oh, Miyolerung Quackenbush, and Berndhard Grill, MPEG Unified Speech and Audio Coding - The ISO / MPEG Standard for High-Efficiency Audio Coding of all Content Types, Convention Paper 8654, AES, April 2012, Presented at the 132nd Convention Budapest, Hungary.
Nam In Park, Hong Kook Kim, Min A Jung, Seong Ro Lee, and Seung Ho Choi, Burst packet loss concealment using multiple codebooks and comfort noise for celp-type speech coders in wireless sensor networks, Sensors 11 (2011), 5323- 5336.
Schuyler Quackenbush and Peter F. Driessen, Error mitigation in MPEG-4 audio packet communication systems, Audio Engineering Society Convention 115, no. 5981, Oct 2003.
S. Rangachari and PC Loizou, A noise-estimation algorithm for highly non-stationary environments, Speech Commun. 48 (2006), 220-231.
V. Stahl, A. Fischer, and R. Bippus, Quantile based noise estimation for spectral subtraction and
168 wiener filtering, in Proc. IEEE Int. Conf. Acoust., Speech and Signal Process. (2000), 1875-1878.
[SS98] J. Sohn and W. Sung, A voice activity detector employing soft decision based noise spectrum adaptation, Proc. IEEE Int. Conf. Acoustics, Speech, Signal Processing, no. pp. 365-368, IEEE, 1998.
<td>[Yu09]</td><td colspan="2">Rongshan Yu, A</td><td>low-complexity</td><td>noise estimation</td>
<td></td><td>algorithm</td><td>based</td><td>on smoothing</td><td>of noise power</td>
<td></td><td>estimation</td><td>and</td><td>estimation</td><td>bias correction,</td>
<td></td><td>Acoustics,</td><td>Speech</td><td>and Signal</td><td>Processing, 2009.</td>
<td></td><td>ICASSP 2009</td><td>. IEEE</td><td>International</td><td>Conference on, Apr</td>
2009, pp. 4421-4424.
169
Contents8
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
178 members in 20 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 13173154 | European Patent Office (EPO) | A | |
| 13173154 | European Patent Office (EPO) | A | |
| 131731549 | European Patent Office (EPO) | – | |
| 14166998 | European Patent Office (EPO) | A | |
| 14166998 | European Patent Office (EPO) | A | |
| 141669986 | European Patent Office (EPO) | – | |
| 2014063177 | European Patent Office (EPO) | W | |
| 2014063177 | European Patent Office (EPO) | W | |
| 131731549 | – | – | – |
| 141669986 | – | – | – |
| EP20130173154 | – | – | – |
| EP20140166998 | – | – | – |
| PCTEP2014063177 | – | – | – |
| WO2014EP63177 | – | – | – |
Members178
| Document | Office | Kind | |
|---|---|---|---|
| CA2913578A1 | Canada | A1 | |
| CA2914869A1 | Canada | A1 | |
| CA2914895A1 | Canada | A1 | |
| CA2915014A1 | Canada | A1 | |
| CA2916150A1 | Canada | A1 | |
| WO2014202784A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202786A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202788A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202789A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202790A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201508736A | Taiwan Province of China | A | |
| TW201508737A | Taiwan Province of China | A | |
| TW201508738A | Taiwan Province of China | A | |
| TW201508739A | Taiwan Province of China | A | |
| TW201508740A | Taiwan Province of China | A | |
| AR096692A1 | Argentina | A1 | |
| AR096693A1 | Argentina | A1 | |
| AR096695A1 | Argentina | A1 | |
| AR096696A1 | Argentina | A1 | |
| AR096698A1 | Argentina | A1 | |
| SG11201510352YA | Singapore | A | |
| SG11201510353RA | Singapore | A | |
| SG11201510508QA | Singapore | A | |
| SG11201510510PA | Singapore | A | |
| SG11201510519RA | Singapore | A | |
| AU2014283123A1 | Australia | A1 | |
| AU2014283194A1 | Australia | A1 | |
| AU2014283124A1 | Australia | A1 | |
| AU2014283196A1 | Australia | A1 | |
| AU2014283198A1 | Australia | A1 | |
| CN105340007A | China | A | |
| CN105359209A | China | A | |
| CN105359210A | China | A | |
| KR20160021295A | Republic of Korea | A | |
| KR20160022363A | Republic of Korea | A | |
| KR20160022364A | Republic of Korea | A | |
| KR20160022365A | Republic of Korea | A | |
| CN105378831A | China | A | |
| KR20160022886A | Republic of Korea | A | |
| CN105431903A | China | A | |
| MX2015016892A | Mexico | A | |
| MX2015017126A | Mexico | A | |
| US2016104487A1 | United States of America | A1 | |
| US2016104488A1 | United States of America | A1 | |
| US2016104489A1 | United States of America | A1 | |
| US2016104497A1 | United States of America | A1 | |
| US2016111095A1 | United States of America | A1 | |
| EP3011557A1 | European Patent Office (EPO) | A1 | |
| EP3011558A1 | European Patent Office (EPO) | A1 | |
| EP3011559A1 | European Patent Office (EPO) | A1 | |
| EP3011561A1 | European Patent Office (EPO) | A1 | |
| EP3011563A1 | European Patent Office (EPO) | A1 | |
| MX2015018024AThis record | Mexico | A | |
| JP2016522453A | Japan | A | |
| JP2016523381A | Japan | A | |
| JP2016526704A | Japan | A | |
| JP2016527541A | Japan | A | |
| MX2015017261A | Mexico | A | |
| TWI553631B | Taiwan Province of China | B | |
| JP2016532143A | Japan | A | |
| AU2014283123B2 | Australia | B2 | |
| AU2014283124B2 | Australia | B2 | |
| AU2014283194B2 | Australia | B2 | |
| AU2014283196B2 | Australia | B2 | |
| AU2014283198B2 | Australia | B2 | |
| TWI564884B | Taiwan Province of China | B | |
| TWI569262B | Taiwan Province of China | B | |
| TWI575513B | Taiwan Province of China | B | |
| MX347233B | Mexico | B | |
| EP3011557B1 | European Patent Office (EPO) | B1 | |
| EP3011561B1 | European Patent Office (EPO) | B1 | |
| TWI587290B | Taiwan Province of China | B | |
| RU2016101469A | Russian Federation | A | |
| BR112015031177A2 | Brazil | A2 | |
| BR112015031178A2 | Brazil | A2 | |
| BR112015031180A2 | Brazil | A2 | |
| BR112015031343A2 | Brazil | A2 | |
| BR112015031606A2 | Brazil | A2 | |
| PT3011557T | Portugal | T | |
| PT3011561T | Portugal | T | |
| EP3011558B1 | European Patent Office (EPO) | B1 | |
| EP3011559B1 | European Patent Office (EPO) | B1 | |
| RU2016101521A | Russian Federation | A | |
| RU2016101600A | Russian Federation | A | |
| RU2016101604A | Russian Federation | A | |
| RU2016101605A | Russian Federation | A | |
| HK1224009A1 | Hong Kong, China | A1 | |
| HK1224076A1 | Hong Kong, China | A1 | |
| HK1224423A1 | Hong Kong, China | A1 | |
| HK1224424A1 | Hong Kong, China | A1 | |
| HK1224425A1 | Hong Kong, China | A1 | |
| JP6190052B2 | Japan | B2 | |
| JP6196375B2 | Japan | B2 | |
| JP6201043B2 | Japan | B2 | |
| ES2635027T3 | Spain | T3 | |
| ES2635555T3 | Spain | T3 | |
| PT3011558T | Portugal | T | |
| MX351363B | Mexico | B | |
| KR101785227B1 | Republic of Korea | B1 | |
| JP6214071B2 | Japan | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 2015018024
- Publication, EPODOC
- MX2015018024
- Application
- 2015018024
- Application, DOCDB
- 2015018024
- Application, EPODOC
- MX20150018024
Titles
- Spanish
- APARATO Y MÉTODO PARA DESVANECIMIENTO DE SEÑAL MEJORADO EN DIFERENTES DOMINIOS DURANTE OCULTAMIENTO DE ERRORES.
Classification
- CPC, 14
- G10L19/005
- G10L19/002
- G10L19/0212
- G10L19/09
- G10L19/012
- G10L19/083
- H03M7/30
- G10L19/22
- G10L19/06
- G10L19/12
- G10L2019/0002
- G10L2019/0011
- G10L2019/0016
- G10L19/07
- IPC, 3
- G10L19 005
- G10L19 09
- G10L25 90