Apparatus and method for improved signal fade out for switched audio coding systems during error concealment.
Abstract
An apparatus for decoding an audio signal is presented. The apparatus comprises a reception interface (110), where the reception interface (110) is configured to receive a first frame comprising a first portion of the audio signal of the audio signal, and where the reception interface (110 ) is configured to receive a second frame comprising a second portion of the audio signal of the audio signal. Furthermore, the apparatus comprises a noise level tracking unit (130), wherein the noise level tracking unit (130) is configured to determine the noise level information depending on at least one of the first portion of the audio signal and the second portion of the audio signal, where the noise level information is represented in a trace domain. In addition, the apparatus comprises a first reconstruction unit (140) for reconstructing, in a first reconstruction domain, a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received by the reception interface (110) or if said third frame is received by the reception interface (110) but is corrupted, where the first reconstruction domain is different from or equal to the crawl domain. Furthermore, the apparatus comprises a transformation unit (121) for transforming the noise level information from the tracking domain to a second reconstruction domain, if a fourth frame of the plurality of frames is not received by the reception interface ( 110) or if said fourth frame is received by the reception interface (110) but is corrupted, where the second reconstruction domain is different from the tracking domain, and where the second reconstruction domain is different from the first reconstruction domain. In addition, the apparatus comprises a second reconstruction unit (141) to reconstruct, in the second reconstruction domain, a fourth portion of the audio signal from the audio signal depending on the noise level information that is represented in the second reconstruction domain, if said fourth frame of the plurality of frames is not received by the reception interface (110) or if said fourth frame is received by the reception interface (110) but is corrupted.

Term
7.7 yearsleft in the term
Expires 23 June 2034.
- Priority
- Filed
- Granted
- Today
- Expires
22 claims: 10 independent, 12 dependent
- 1REIVINDICACIONES 1. Un aparato para decodificar una señal de audio, que comprende:una interfaz de recepción (110) para recibir una pluralidad de tramas, donde la interfaz de recepción (110) está configurada para recibir una primera trama de la pluralidad de tramas, donde dicha primera trama comprende una primera porción de la señal de audio de la señal de audio, donde dicha primera porción de la señal de audio está representada en un primer dominio, y donde la interfaz de recepción (110) está configurada para recibir una segunda trama de la pluralidad de tramas, donde dicha segunda trama comprende una segunda porción de la señal de audio de la señal de audio, una unidad de transformación (120) para transformar la segunda porción de la señal de audio o un valor o señal derivada de la segunda porción de la señal de audio de un segundo dominio a un dominio de rastreo para obtener una información de la segunda porción de la señal, donde el segundo dominio es diferente del primer dominio, donde el dominio de rastreo es diferente del segundo dominio, y donde el dominio de rastreo es 152 igual ejrsirtvTO kzxjcamo rí— os m «omídad _ ,. „ . . , wwmiAL o diferente del primer dominio, una unidad de rastreo de nivel de ruido (130), donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir una información de la primera porción de la señal que está representada en el dominio de rastreo, donde la información de la primera porción de la señal depende de la primera porción de la señal de audio, donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir la segunda porción de la señal que está representada en el dominio de rastreo, y donde la unidad de rastreo de nivel de ruido (130) está configurada para determinar la información de nivel de ruido dependiendo de la información de la primera porción de la señal que está representada en el dominio de rastreo y dependiendo de la información de la segunda porción de la señal que está representada en el dominio de rastreo, donde la información de nivel de ruido está representada en el dominio de rastreo y una unidad de reconstrucción (140) para reconstruir una tercera porción de la señal de audio de la señal de audio dependiendo de la información de nivel de ruido, si una tercera trama de la pluralidad de tramas no es recibida por - IMPIAS tNSrtnrtOMWCAMo W t a fikwcoa o r~lmmsswuM. * J»' la interfaz de recepción (110) o si dicha tercera trama es recibida por la interfaz de recepción (110) pero está corrupta.
- 2Un aparato de acuerdo con la reivindicación 1, en el cual la primera porción de la señal de audio está representada en un dominio del tiempo como primer dominio, donde la unidad de transformación (120) está configurada para transformar la segunda porción de la señal de audio o el valor derivado de la segunda porción de la señal de audio de un dominio de la excitación que es el segundo dominio al dominio del tiempo que es el dominio de rastreo, donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir la información de la primera porción de la señal que está representada en el dominio:del tiempo como dominio de rastreo y donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir la segunda porción de la señal gue está representada en el dominio del tiempo como dominio de rastreo.
- 3Un aparato de acuerdo con la reivindicación 1, 154 CNSmVTO MEXICANO J ge u norato C^aSL^ IMWJSntlAL en el cusí la primera porción de- la señal de audio •«MeiinBWeMWwr-wiEwweEaerwWMK^rs está representada en el dominio de la excitación como primer dominio, donde la unidad de transformación (120) está configurada para transformar la segunda porción de la señal de audio o el valor derivado de la segunda porción de la señal de audio de un dominio del tiempo que es el segundo dominio al dominio de la excitación que es el dominio de rastreo, donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir la información de la primera porción de la señal que está representada en el dominio de la excitación como dominio de rastreo y donde la unidad de rastreo de nivel de ruido (.130) está configurada para recibir la segunda porción de la señal que está representada en el dominio de la excitación como dominio de rastreo.
- 4Un aparato de acuerdo con la reivindicación 1, en el cual la primera porción de la señal de audio está representada en el dominio de la excitación como primer dominio, donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir la información de la primera 155 IMPI^ Qtnnuro MAJOOf*', αε la tkw**o OwmmtMí porción de la señal, donde dicha información de^^Ta* - primera porción de la señal está representada en el dominio de FFT, que es el dominio de rastreo, y donde dicha información de la primera porción de la señal depende de dicha primera porción de la señal de audio que está representada en el dominio de la excitación, donde la unidad de transformación (120) está configurada para transformar la segunda porción de la señal de audio o el valor derivado de la segunda porción de la señal de audio de un dominio del tiempo que es el segundo dominio a un dominio de FFT que es el dominio de rastreo y donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir la segunda porción de la señal de audio que está representada en el dominio de FFT.
- 5Un aparato de acuerdo con una de las reivindicaciones anteriores, donde el aparato comprende asimismo una primera unidad de agregación (150) para determinar un primer valor agregado dependiendo de la primera porción de la señal de audio, donde el aparato comprende asimismo una segunda unidad de agregación (160) para determinar, dependiendo de la segunda porción de la señal de audio, un segundo valor 156 IMPI INSTITUTO MEXICANO n LA PROPIEDAD agregado como valor derivado de la segunda señal de audio, donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir el primer valor agregado como 5 información de la primera porción de la señal que está representada en el dominio de rastreo, donde la unidad de rastreo de nivel de ruido (130) está configurada para recibir el segundo valor agregado como información de la segunda porción de la señal que está representada en el dominio de 10 rastreo, y donde la unidad de rastreo de nivel de ruido (130) está configurada para determinar la información de nivel de ruido dependiendo del primer valor agregado que está representada en el dominio de rastreo y dependiendo del segundo valor agregado que está representada en el dominio de 15 rastreo.
- 6Un aparato de acuerdo con la reivindicación 5, en el cual la primera unidad de agregación (150) está configurada para determinar el primer valor agregado de tal manera que el primer valor agregado indique una raíz media 20 cuadrática de la primera porción de la señal de audio o de una señal derivada de la primera porción de la señal de audio 157 donde la segunda unidad de IMPI IHSTuUTO mexicano cela «ohedao agregación (160 configurada para determinar el segundo valor agregado dé 1 Ldf 1 · manera que el segundo valor agregado indique una raíz media cuadrática de la segunda porción de la señal de audio o de una señal derivada de la segunda porción de la señal de audio.
- 7Un aparato de acuerdo con una de las reivindicaciones anteriores, donde la unidad de transformación (120) está configurada para transformar el valor derivado de la segunda porción de la señal de audio del segundo dominio al dominio de rastreo mediante la aplicación de un valor de ganancia al valor derivado de la segunda porción de la señal de audio.
- 8Un aparato de acuerdo con la reivindicación 7, en el cual el valor de ganancia indica una ganancia introducida por síntesis de Codificación por Predicción Lineal o donde el valor de ganancia indica una ganancia introducida por síntesis y desénfasis de Codificación por Predicción Lineal.
- 9Un aparato de acuerdo con una de las reivindicaciones anteriores, en el cual la unidad de rastreo de nivel de ruido (130) está configurada para determinar la 158 información de nivel de _ j ____, · _ . Ί _ _ _, INDUSTRIAL· ruido mediante la aplicación IMPI INSTITUTO MEXICANO Ot LA EROUWAO estrategia de estadísticas mínimas.
- 10Un aparato de acuerdo con una de las reivindicaciones anteriores, en el cual la unidad de rastreo de nivel de ruido (130) está configurada para determinar un nivel de ruido de confort como información de nivel de ruido y donde la unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio dependiendo de la información de nivel de ruido, si dicha tercera trama de la pluralidad de tramas no es recibida por la interfaz de recepción (110) o si dicha tercera trama es recibida por la interfaz de recepción (110) pero está corrupta.
- 11Un aparato de acuerdo con la reivindicación 9, en el cual la unidad de rastreo de nivel de ruido (130) está configurada para determinar un nivel de ruido de confort como información de nivel de ruido derivado de un espectro de nivel de ruido, donde dicho espectro de nivel de ruido se obtiene mediante la aplicación de la estrategia de estadísticas mínimas y donde la unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal 159 , , . , , , , INSTITUTO MEXICANO JA de audio dependiendo de una pluralidad d i^^^^^ae predicción lineal, si dicha tercera trama Hp la pluralidad de tramas no es recibida por la interfaz de recepción (110) o si dicha tercera trama es recibida por la interfaz de recepción (110) pero está corrupta.
- 12Un aparato de acuerdo con una de las reivindicaciones 1 a 9, en el cual la unidad de rastreo de nivel de ruido (130) está configurada para determinar una pluralidad de Coeficientes de FFT que indican un nivel de ruido de confort como información de nivel de ruido y donde la primera unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio dependiendo de a nivel de ruido de confort derivado de dichos coeficientes de FFT, si dicha tercera trama de la pluralidad de tramas no es recibida por la interfaz de recepción (140) o si dicha tercera trama es recibida por la interfaz de recepción (110) pero está corrupta.
- 13Un aparato de acuerdo con una de las reivindicaciones anteriores, en el cual la unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio dependiendo de la 160 IMPI MROVVa.'MEMKMIIO WFJW.*»3^»ej CS-eR*©·' información de nivel de ruido y dependiendo de la primera o segunda porción de la señal de audio, si dicha tercera trama de la pluralidad de tramas no es recibida por la interfaz de recepción (110) o si dicha tercera trama es recibida por la interfaz de recepción (110) pero está corrupta.
- 14Un aparato de acuerdo con la reivindicación 13, en el cual la unidad de reconstrucción (140) está configurada para reconstruir la tercera porción de la señal de audio mediante la atenuación o amplificación de una señal derivada de la primera porción de la señal de audio o la segunda porción de la señal.
- 15Un aparato de acuerdo con una de . las reivindicaciones anteriores, donde el aparato comprende asimismo una unidad de predicción a largo plazo (170) que comprende un búfer de retardo (180), donde la unidad de predicción a largo plazo (170) está configurada para generar una señal procesada dependiendo de la primera o segunda porción de la señal de audio, dependiendo de un ingreso del búfer de retardo (180) que está almacenado en el búfer de retardo (180) y dependiendo de una ganancia de predicción a largo plazo y 161 IMPI ΙΝίΤΠυΤΟ MEXICANO DE LAtEOMOAO iHWimiAL donde la unidad de predicción a largo plazo (170) está configurada para desvanecer la ganancia de predicción a largo plazo hacia cero, si dicha tercera trama de la pluralidad de tramas no es recibida por la interfaz de 5 recepción (110) o si dicha tercera trama es recibida por la interfaz de recepción (110) pero está corrupta.
- 16Un aparato de acuerdo con la reivindicación 15, en el cual la unidad de predicción a largo plazo (170) está configurada para desvanecer la ganancia de predicción a 10 largo plazo hacia cero, donde la velocidad a la cual se desvanece la ganancia de predicción a largo plazo hacia cero depende de un factor de desvanecimiento.
- 17Un aparato de acuerdo con la reivindicación 15 o 16, en el cual la unidad de predicción a largo plazo (170) 15 está configurada para actualizar la entrada del búfer de retardo (180) almacenando la señal procesada generada en el búfer de retardo (180), si dicha tercera trama de la pluralidad de tramas no es recibida por la interfaz de recepción (110) o si dicha tercera trama es recibida por la 20 interfaz de recepción (110) pero está corrupta.
- 18Un aparato de acuerdo con una de las reivindicaciones anteriores, 62 ΙΜΡΙ^ι ίΝβτττϊ/ηα mmucaki νζϋαβτΤί·, jl OtlA FW3RW.O en el cual la unidad de transformación (120) es una primera unidad de transformación (120), donde la unidad de reconstrucción (140) es una primera unidad de reconstrucción (140), donde el aparato comprende asimismo una segunda unidad de transformación (121) y una segunda unidad de reconstrucción (141), donde la segunda unidad de transformación (121) está configurada para transformar la información de nivel de ruido del dominio de rastreo al segundo dominio, si una cuarta trama de la pluralidad de tramas no es recibida por la interfaz de recepción (110) o si dicha cuarta trama es recibida por la interfaz de recepción (110) pero está corrupta y donde la segunda unidad de reconstrucción (141) está configurada para reconstruir una cuarta porción de la señal de audio de la señal de audio dependiendo de la información de nivel de ruido que está representada en el segundo dominio si dicha cuarta trama de la pluralidad de tramas no es recibida por la interfaz de recepción (110) o si dicha cuarta trama es recibida por la interfaz de recepción (110) pero está corrupta. 163 IMPI
- 19INSTrnWO MiWCANO . OtW.WOKIOW . WftiCTIMJL Un aparato de acuerdo ’^eon la reivindicación 18, en el cual la segunda unidad de reconstrucción (141) está configurada para reconstruir la cuarta porción de la señal de audio dependiendo de la información de nivel de ruido y dependiendo de la segunda porción de la señal de audio.
- 20Un aparato de acuerdo con la reivindicación 19, en el cual la segunda unidad de reconstrucción (141) está configurada para reconstruir la cuarta porción de la señal de audio mediante la atenuación o amplificación de la segunda porción de la señal de audio.
- 21Un método para decodificar una señal de audio, que comprende:recibir una primera trama de una pluralidad de tramas, donde dicha primera trama comprende una primera porción de la señal de audio de la señal de audio, donde dicha primera porción de la señal de audio está representada en un primer dominio, recibir una segunda trama de la pluralidad de tramas, donde dicha segunda trama comprende una segunda porción de la señal de audio de la señal de audio, transformar la segunda porción de la señal de audio o un valor o señal derivada de la segunda porción de la señal de audio de un segundo dominio a un dominio de rastreo para obtener una información de la segunda 164 donde el segundo dominio es diferente donde el dominio de rastreo es diferente del segundo dominio, y donde el dominio de rastreo es igual o diferente del primer dominio, determinar la información de nivel de ruido dependiendo de la información de la primera porción de la señal, gue está representada en el dominio de rastreo, y dependiendo de la información de la segunda porción de la señal que está representada en el dominio de rastreo, donde la información de la primera porción de la señal depende de la primera porción de la señal de audio, reconstruir una tercera porción de la señal de audio de la señal de audio dependiendo de la información de nivel de ruido, si una tercera trama de la pluralidad de tramas no es recibida o si dicha tercera trama es recibida pro está corrupta.
- 22Un medio leíble por computadora para decodificar una señal de audio, que comprende el método de la reivindicación 21. 165
Independent claims22
1,037 paragraphs in 128 sections, as filed
(54) Title: APPARATUS AND METHOD FOR IMPROVED SIGNAL FADING FOR AUDIO CODING SYSTEMS SWITCHED DURING HIDING OF ERRORS.
(54) Title: APPARATUS AND METHOD FOR IMPROVED SIGNAL FADE OUT FOR SWITCHED AUDIO CODING SYSTEMS DURING ERROR CONCEALMENT.
(57) Summary
An apparatus for decoding an audio signal is presented. The apparatus comprises a reception interface (110), where the reception interface (110) is configured to receive a first frame comprising a first portion of the audio signal of the audio signal, and where the reception interface (110 ) is configured to receive a second frame comprising a second portion of the audio signal of the audio signal. Furthermore, the apparatus comprises a noise level tracking unit (130), wherein the noise level tracking unit (130) is configured to determine the noise level information depending on at least one of the first portion of the audio signal and the second portion of the audio signal, where the noise level information is represented in a trace domain. In addition, the apparatus comprises a first reconstruction unit (140) for reconstructing, in a first reconstruction domain, a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received by the reception interface (110) or if said third frame is received by the reception interface (110) but is corrupted, where the first reconstruction domain is different from or equal to the crawl domain. Furthermore, the apparatus comprises a transformation unit (121) for transforming the noise level information from the tracking domain to a second reconstruction domain, if a fourth frame of the plurality of frames is not received by the reception interface ( 110) or if said fourth frame is received by the reception interface (110) but is corrupted, where the second reconstruction domain is different from the tracking domain, and where the second reconstruction domain is different from the first reconstruction domain. In addition, the apparatus comprises a second reconstruction unit (141) to reconstruct, in the second reconstruction domain, a fourth portion of the audio signal from the audio signal depending on the noise level information that is represented in the second reconstruction domain, if said fourth frame of the plurality of frames is not received by the reception interface (110) or if said fourth frame is received by the reception interface (110) but is corrupted.
(57) Abstract
An apparatus for decoding an audio signal ¡s provided. The apparatus comprises a receiving interface (110), wherein the receiving interface (110) is configured to receive a first trame comprising a first audio signal portion of the audio signal, and wherein the receiving interface (110) is configured to receive a second trame comprising a second audio signal portion of the audio signal. Moreover, the apparatus comprises a noise level tracing unit (130), wherein the noise level tracing unit (130) is configured to determine noise level Information depending on at least one of the first audio signal portion and the second audio signal portion, wherein the noise level Information is represented in a tracing domain. Furthermore, the apparatus comprises a first reconstruction unit (140) for reconstructing, in a first reconstruction domain, a third audio signal portion of the audio signal depending on the noise level Information, if a third trame of the plurality of trames is not received by the receiving interface (110) or if said third trame is received by the receiving interface (110) but is corrupted, wherein the first reconstruction domain is different from or equal to the tracing domain. Moreover, the apparatus comprises a transform unit (121) for transforming the noise level Information from the tracing domain to a second reconstruction domain, if a fourth trame of the plurality of trames is not received by the receiving interface (110) or if said fourth trame is received by the receiving interface (110) but is corrupted, wherein the second reconstruction domain is different from the tracing domain, and wherein the second reconstruction domain is different from the first reconstruction domain. Furthermore, the apparatus comprises a second reconstruction unit (141) for reconstructing, in the second reconstruction domain, a fourth audio signal portion of the audio signal depending on the noise level Information being represented in the second reconstruction domain, if said fourth trame of the plurality of trames is not received by the receiving interface (110) or if said fourth trame is received by the receiving interface (110) but is corrupted.
Headlines)
Home:
Denomination:
Classification lnventor («s):
Number:
Mexican Institute of Industrial Property
PATENT TITLE NO. 347233
<img file="MX347233B_D0001.tif" />
FRAUNHOFER-GESELLSCHAFT ZUR FÓRDERUNG DER ANGEWANDTEN FORSCHUNG EV
Hansastrasse 27C, 80686, Munich, GERMANY
APPARATUS AND METHOD FOR IMPROVED SIGNAL FADING FOR AUDIO ENCODING SYSTEMS SWITCHED DURING HIDING OF ERRORS.
Int.CI.8: G10L19 / 005; G10L19 / 09; G10L25 / 90
MICHAEL SCHNABEL; GORAN MARKOVIC; RALPH SPERSCHNEIDER;
JÉRÉMIE LECOMTE; CHRISTIAN HELMRICH
REQUEST
International filing date:
MX / a / 2015/016638 of June 2014
PRIORITY
Country:
Date:
Number:
EP EP June 2013 May 5 2014
13173154.9
14166998.6
Validity: Twenty years
Expiration Date: June 23, 2034
The reference patent is granted based on articles 1, 2, section V, 6, section M, and 68 of the Industrial Property Law.
In accordance with article 23 of the Industrial Property Law, this patent is valid for twenty years, non-extendable, counted from the date of presentation of the international application and will be subject to the payment of the fee to keep the rights in force. . . .
Whoever signs this title does so based on the provisions of articles 6 ° sections III and 7 ° bis 2 of the Industrial Property Law (Official Gazette of the Federation (DOF) 06/27/1991, amended on 02 / 08/1994, 10/25/1996, 12/26/1997, 05/17/1999, 01/26/2004, 06/16/2005, 01/25/2006, 05/06/2009. 06/01/2010 , 06/18/2010, 06/28/2010, 01/27/2012 and 04/09/2012). Articles 1, 3<sup>or</sup> fraction V subsection a), 4 'and 12 · fractions I and III of the Regulations of the Mexican Institute of Industrial Property (DOF, 12/14/1999, amended on 07/01/2002, 07/15/2004, 07/28 / 2004 and 7/09/2007); Articles 1, 3, 4, 5 section V subsection a), 16 sections I and III and 30 of the Organic Statute of the Mexican Institute of Industrial Property (DOF 12/27/1999, amended on 10/10/2002, 07/29/2004, 08/04/2004 and 09/13/2007); 1st, 3rd and 6th subsections a) of the Agreement that delegates powers to the Deputy General Directors, Coordinator, Divisional Directors, Heads of Regional Offices, Divisional Deputy Directors, Departmental Coordinators and other subordinates of the Mexican Institute of Industrial Property. (DOF 12/15/1999, amended on 02/04/2000, 07/29/2004, 08/04/2004 and 09/13/2007).
<img file="MX347233B_D0002.tif" />
Issue Date: April 19, 2017
TO DIVISIONAL DIRECTOR OF PATENTS
<img file="MX347233B_D0003.tif" />
NAHANNY CANAL REYES
No 550 P Santa Mari
MX / 2017/33819
INSTITUTO MEXICANO Jñ
OF PROPERTY V »= aseSU £ ¿F INDUSTRIAL
APPARATUS AND METHOD FOR IMPROVED SIGNAL FADING FOR AUDIO CODING SYSTEMS SWITCHED DURING
HIDING ERRORS
Description
The present invention relates to the encoding, processing and decoding of audio signals and, in particular, to an apparatus and method for improved signal fading for switched audio coding systems during error concealment.
The following describes the state of the art regarding the fading of voice and audio codes during packet loss concealment (PLC). The explanations regarding the state of the art with the ITU-T codes of the G series (G.718, G.719, G.722, G.722.1, G.729. G.729.1), precede the 3GPP code (AMR, AMR-WB, AMR-WB +) and an IETF codec (OPUS), and conclude with two MPEG code (HE-AAC, HILN) (ITU = International Telecommunication Union ( International Telecommunications Union); 3GPP = 3rd Generation Partnership Project (Sodieties Project of 3<sup>to</sup> Generation; AMR = Adaptive Multi-Rate; WB = Broadband; IETF = Internet Engineering Task Forcé (Internet Engineering Task Force). The state of the art with respect to the noise level trace is then discussed, followed by a summary that provides an overview.
First, G.718 is considered. G.718 is a narrowband and wideband speech codec that supports DTX / CNG (DTX = Digital Theater Systems
IMPI
MEXICAN INSTITUTE>
DE LA PtOREDAD ΐ \ Γ ~ ϋΓ {€> (Digital Theater Systems); CNG = Noise generation of<sup>w</sup>5lW6rt (Comfort Noise Generation)). As the embodiments relate, eh pfTICCflar, to low-delay coding, the low-delay mode is described here in more detail.
Considering ACELP (Layer 1) (ACELP = Algébrale Code Excited Linear Prediction), the ITU-T recommends for G.718 [ITU08a, section 7.11] an adaptive fading in the linear prediction domain to control the fading speed. In general, concealment follows this principle:
According to G.718, in case of frame erasures, the concealment strategy can be summarized as convergence of the signal energy and the spectral envelope with the estimated parameters of the background noise. The periodicity of the signal is converged to zero. The speed of convergence depends on the parameters of the last correctly received frame and the number of consecutive erased frames, and is controlled by an attenuation factor, a. The attenuation factor a, also depends on the stability, Θ, of the LP filter (LP = Linear Prediction, Linear Prediction) corresponding to NO VOICE frames. In general, convergence is slow if the last correct frame received is on a stable segment and it is fast if the frame is on a transition segment.
The attenuation factor a depends on the voice signal class, which is derived by the signal classification described in [ITU08a, section 6.8.1.3.1 and 7.11.1.1]. The stability factor Θ is computed on the basis of a measure of
I - f T ~ 1
IMPIOUS"
INSTITUTO MEXICANO Λ '*'!
DE LA PROHEDAO the distance between filters ISF (Immittance Spectral Freéfü ^ ñi ^, Frequency
Spectral Immittance) [ITU08a, section 7.1.2.4.2].
Table 1 shows the calculation scheme for a:
<td>Last good plot received</td><td>Number of successive frames erased</td><td>to</td>
<td>ARTIFICIAL START</td><td></td><td> 0,6</td>
<td>HOME, WITH VOICE</td><td> <3</td><td> 1,0</td>
<td></td><td> >3</td><td> 0,4</td>
<td>TRANSITION WITH VOICE</td><td></td><td> 0,4</td>
<td>VOICELESS TRANSITION</td><td></td><td> 0,8</td>
<td>WITHOUT VOICE</td><td> = 1</td><td> 0,2 0 + 0,8</td>
<td></td><td> = 2</td><td> 0,6</td>
<td></td><td> >2</td><td> 0,4</td>
Table 1: Values of the attenuation factor a, the value Θ is a stability factor computed from the measurement of the distance between the adjacent LP filters.
[ITU08a, section 7.1.2.4.2],
Furthermore, G.718 introduces a fading method to modify the spectral envelope. The general idea is that the latest ISF parameters converge towards an adaptive ISF mean vector. First, an average ISF vector is calculated from the last 3 known ISF vectors. Then ___ is averaged
ΙΜΡΙ ^^
MEXICAN INSTITUTE> f4
I OWNED once again the average ISF vector with an offline long-term learning latex vector (which is a constant vector) [ITU08a, section 7.11.1.2].
Furthermore, G.718 offers a fading method to control long-term behavior and therefore interaction with background noise, where the tonal excitation energy (and hence the excitation periodicity) converges to 0, whereas the random excitation energy converges on the CNG excitation energy [ITU08a, section 7.11.1.6]. The innovative gain attenuation is calculated as follows
4<sup>11</sup> = + (<sup>1</sup> ” <sup>to</sup>) 9n <sub>(1</sub>) where g ^<sup>1]</sup> is the innovative gain at the beginning of the next frame, is the innovative gain at the beginning of the current frame, g „is the gain of the excitation used during the generation of comfort noise and the attenuation factor a.
As with the periodic attenuation of the excitation, the gain is attenuated in a linear way during the whole frame sample by sample from, and reaches g ^<sup>11</sup> at the beginning of the next frame.
Fig. 2 outlines the structure of the G.718 decoder. In particular, Fig. 2 illustrates a G.718 high-level decoder structure for PLC, featuring a high-pass filter.
IMPIAS
INSTITUTO MEXICANO Xi, t> E LA PROPERTY ”£ Γ - ÍÍJ
¡. ^ INDUSTRIAL —-_ ^ 8 ^
Using the previously described technique of G.718, fa-gain ^ innnwadnra converges towards the gain used during the generation of comfort noise g<sub>n </sub>corresponding to long packet loss trains. As described in [ITU08a, section 6.12.3], the comfort noise gain g<sub>n</sub> It is expressed in terms of the square root of energy E. The conditions of the E update are not described in detail. Following the reference implementation (floating-point code C, stat_noise_uv_mod. c), E is derived as follows: if (unvoiced vad == 0) {if (unv_cnt> 20) {ftmp = lp_gainc * lp_gainc;
lp_ener = 0.7f * lp_ener + 0.3f * ftmp; } else {unv_cnt ++; }} else {unv_cnt = 0;
}
IMPI
<img file="MX347233B_D0004.tif" />
INSTITUTO MEXICANO PE LA ntOHEDAD INDUSTRIAL where unvoiced_vad maintains voice activity detection, where unv_cnt maintains the number of consecutive frames without voice, where lp_gainc maintains the fixed codebook low-pass gains and where lp_ener maintains the energy estimate calculation of low-pass CNG E, is initialized with 0.
In addition, G.718 presents a high-pass filter, introduced in the path of the excitation signal without voice, if the signal of the last correct frame had a different classification of NO VOICE, see Fig. 2, see also [ITU08a, section 7.11.1.6]. This filter has a low-end characteristic with a DC frequency response that is around 5 dB lower than the Nyquist frequency.
Furthermore, G.718 proposes a decoupled LTP feedback loop (LTP = Long-Term Prediction): Although during normal operation the feedback loop corresponding to the adaptive codebook is updated by units less than the frames ([ITU08a, section 7.1.2.1.4]) based on full excitation. During concealment this feedback loop is updated frame by frame (see [ITU08a, sections 7.11.1.4, 7.11.2.4, 7.11.1.6, 7.11.2.6; dec_GV_exc @ dec_gen_voic. C and syn_bf i_post @ syn_bf i_pre_post. C]) based only in excitement with voice. With this strategy, the adaptive codebook is not “contaminated” with noise as it originates from a randomly chosen innovation drive.
<img file="MX347233B_D0005.tif" />
IMPI
MEXICAN INSTITUTE
OF INDUSTRIAL MOBILITY
As regards the enhancement layers encoded by
-<sup>111</sup>^ -----— G.718 transforms (3-5), during concealment, the decoder behaves with respect to upper layer decoding in a similar way to normal operation, only the spectrum of MDCT at zero. No fading behavior is applied during hiding.
Regarding CNG, in G.718, CNG synthesis is performed in the following order. First, the parameters of a comfort noise frame are decoded. A comfort noise pattern is then synthesized.
The tone buffer is then reset. Subsequently, the synthesis corresponding to the FER classification (Frame Error Recovery) is saved. Next, the de-emphasis of the spectrum is carried out. This is followed by low-frequency post-filtering. Later, the CNG variables are updated.
In the case of concealment, the exact same thing is done, except that the CNG parameters of the bitstream are not decoded. This means that the parameters are not updated during loss of frames, but the decoded parameters of the last correct SID frame (Silence Insertion Descriptor) are used.
It is now considered G.719. G.719, which is based on Siren 22, is a transform-based full-band audio codec. The ITU-T recommends a frame repetition fading in the spectral domain for G.719 [ITU08b, section 8.6]. According to G.719, a frame erasure concealment mechanism is incorporated into the decoder. When a plot is
INSTITUTO MEXICANO DE ΙΑ FttóHBDAU INDUSTRIAL received correctly, the reconstructed transform coefficients are stored in a buffer. If the decoder is informed that a frame has been lost or the frame is corrupted, the reconstructed transform coefficients on the last received frame are scaled downward by a factor of 0.5 and then used as reconstructed transform coefficients for the current frame. The decoder continues by transforming them to the time domain and executing the windowing and overlapping and adding operation.
Hereinafter, G.722 is described. G.722 is a 50-7000 Hz coding system using subband adaptive pulse code differential modulation (SB-ADPCM) within a bit rate of up to 64 kbit / s. The signal is divided into an upper subband and a lower subband using a QMF analysis (QMF = Quadrature Mirror Filter). The two bands obtained are encoded by ADPCM (ADPCM = Adaptive Differential Pulse Code Modulation).
For G.722, a high complexity algorithm for packet loss concealment is specified in Annex III [ITU06a] and a low complexity algorithm for packet loss concealment is specified in Annex IV [ITU07]. G.722 - Annex III ([ITU06a, section III.5]) proposes a squelch performed gradually, starting with 20ms of frame loss, completing after 60ms of loss of frames. Furthermore, G.722 - Annex IV proposes a fading technique that applies to each sample a
<img file="MX347233B_D0006.tif" />
<sup>9</sup> IMPI rNSTFTUTC MEXICANO DE LA PROPERTY INDUSTRIAL gain factor that is computed and adapted sample by sample [ITU07, section IV.6.1.2.7].
In G.722, the muting process takes place in the subband domain immediately before QMF synthesis and as the last step of the PLC module. The squelch factor calculation is done using class information from the signal classifier which is also part of the PLC module. The distinction is made between TRANSIENT, UV_TRANSITION and other classes. In addition, a distinction is made between individual 10-ms frame losses and other cases (multiple 10-ms frame losses and single / multiple 20-ms frame losses).
This is illustrated in Fig. 3. In particular, Fig. 3 illustrates a situation in which the fading factor of G.722 depends on the class information and where 80 samples are equivalent to 10 ms.
According to G.722, the PLC module creates the signal corresponding to the missing frame and some additional signal (10ms) that is supposed to be cross-faded with the next correct frame. Mute for this additional signal follows the same rules. In G.722 high-band concealment, cross-fading does not take place.
Hereinafter, it is considered G.722.1. G.722.1, which is based on Siren 7, is a transform-based wideband audio codec with a super-wideband extension mode, referred to as G.722.1CG 722.1 C itself is based on Siren 14. The ITU-T recommends for G.722.1 a repetition of frames with subsequent muting [ITU05, section 4.7]. If he
<img file="MX347233B_D0007.tif" />
IMPI
MIXICAN INSTITUTE
DE LA FROMEDAO INDUSTRIAL decoder is informed, by means of an external signaling mechanism not defined in this recommendation, that a frame has been lost or corrupted, it repeats the MLT coefficients (from English, Modulated Lapped Transform) decoded from The previous frame It proceeds by transforming them to the time domain and executing the overlap and sum operation with the decoded information of the previous and next frames. If the previous frame had also been lost or corrupted, then the decoder sets all the MLT coefficients of the current frames to zero.
G.729 is considered below. G.729 is an audio data compression algorithm for speech that compresses digital speech into packets lasting 10 milliseconds. It is officially described as 8 kbit / s speech coding using code-excited linear prediction (CS-ACELP) speech coding [ITU12],
As reported in [CPK08], G.729 recommends fading in the LP domain. The PLC algorithm employed in the G.729 standard reconstructs the speech signal corresponding to the current frame based on the previously received speech information. In other words, the PLC algorithm replaces the missing drive with an equivalent characteristic from the previously received frame, although the drive energy gradually decays ultimately, the fixed and adaptive codebook gains are attenuated by a constant factor.
Fixed codebook attenuated gain is dated by:
<img file="MX347233B_D0008.tif" />
<4<sup>m)</sup> = 0.98 ·
IMPI
MU1CANO INSTITUTE OF INDUSTRIAL PROPERTY where m is the subplot index.
The adaptive codebook gain is based on a dimmed version of the adaptive codebook gain above:
*4” - <sup>11</sup> limited by u<sup>1</sup>™* - 0 0
Nam de Park et al. suggests for G.729, a signal amplitude control that uses prediction by means of linear regression [CPK08, PKJ + 11], It aims to lower the packet loss rate and uses linear regression as a central technique. Linear regression is based on the linear model, namely<sub>10</sub> g<sup>F</sup>i = a + bi <sub>(2)</sub> where g 'is the newly predicted current amplitude, a and b are coefficients corresponding to the first order linear function and i is the index of the frame. To find the optimized coefficients a * and b *, the sum of the squared prediction error is minimized:
i — 1 «= Σ (¾” .½)<sup>2</sup> (3) ε is the quadratic error, g<sub>}</sub> is the y-<sup>to</sup> original past amplitude. To minimize this error, simply adjust the derivative with respect to a and b to zero. Using the optimized parameters a * and b *, an estimate of each g- is indicated by = a * + b * i (4)
<img file="MX347233B_D0009.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PRCPiíTY
Fig. 4 illustrates the amplitude prediction, in particular the prediction of the mII m ίβιιμ i ua¿Mr.-amplitude g *, by using linear regression.
To obtain the amplitude A- of the lost packet i, multiply a relation σ, * _ 9i (Ti - ^ -<sup>1</sup> (5) by a scale factor Si:
Ai <sup>=</sup> * <Ti (θ) where the scale factor Si depends on the number of consecutive hidden frames / (/):
('1.0, if 1 (7) = 1.2
0.0. if / (/) = 3.4
0.8, if / (/) = 5.6
0, otherwise
In [PKJ + 11], a slightly different scaling is proposed.
According to G.729, later, A- is smoothed to prevent discrete fading at the edges of the frames. The final smoothed amplitude A, (n) is multiplied by the drive, obtained from the above PLC components.
G.729.1 is explained hereinafter. G.729.1 is a built-in variable bit rate encoder based on G.729: A scalable 8-32 kbit / s wideband encoder that interfaces with G.729 [ITU06b].
Mexican Institute OF PROPERTY
In accordance with G.729.1, as in G.718 (see above, P ^^ é proposes adaptive fading that depends on the stability of the signal ([ITU06b, section 7.6.1]). During concealment, usually the signal is attenuated based on an attenuation factor a which depends on the parameters of the class of the last good frame received and the number of consecutive erased frames. The attenuation factor a also depends on the stability of the LP filter for NON-VOICE frames. In general, the attenuation is slow if the last good frame received is in a stable segment and it is fast if the frame is in a transition segment.
In addition, the attenuation factor a depends on the average tone gain per subframe g<sub>p</sub> ([ITU06b, eq. 163, 164]):
9p = 0.14 °) + 0.2 ^<sup>1</sup>) + 0.3 ^) + 0.4 ^) (8) where g ^ is the pitch gain of subframe i.
| Table 2 illustrates the calculation scheme of a, where i = where IIX »> d>
ΙΜΡΙ ^ 5
MEXICAN INSTiniTO ¡During the concealment process, a is used in the following ^ «^ ferrw ^ © sYES concealment:
<td>last good plot received</td><td>Number of successive frames erased</td><td>to</td>
<td>WITH VOICE</td><td> 1</td><td>β</td>
<td></td><td> 2,3</td><td>S p</td>
<td></td><td> >3</td><td> 0,4</td>
<td>START</td><td> 1</td><td> 0,8 £</td>
<td></td><td> 2,3</td><td>Sp</td>
<td></td><td> > 3</td><td> 0,4</td>
<td>ARTIFICIAL START</td><td> 1</td><td> 0,6^</td>
<td></td><td> 2,3</td><td>Sp</td>
<td></td><td> > 3</td><td> 0,4</td>
<td>TRANSITION WITH VOICE</td><td> <2</td><td> 0,8</td>
<td></td><td> >2</td><td> 0,2</td>
<td>VOICELESS TRANSITION</td><td></td><td> 0,88</td>
<td>WITHOUT VOICE</td><td> 1</td><td> 0,95</td>
<td></td><td> 2.3</td><td> 0,6 0 + 0,4</td>
<td></td><td> > 3</td><td> 0,4</td>
IMPIOS Mexican institute 'DELAfROflEOAD t INDUSTRIAL ^ »W ·
Table 2: Attenuation factor values a. ^ He value Θ is a stability factor computed from a distance measure between adjacent LP filters. [ITU06b, section 7.6.1].
According to G.729.1, regarding glottic pulse resynchronization, as the last pulse of the previous frame excitation is used for the construction of the periodic part, its gain is approximately correct at the beginning of the hidden frame and is You can set it to 1. Then the gain is linearly attenuated throughout the frame sample by sample to obtain the value of a at the end of the frame. The evolution of the energy of the voiced segments is extrapolated using the tonal excitation gain values of each subframe of the last correct frame. In general, if these gains are greater than 1, the signal energy is increasing; if they are less than 1, the energy is decreasing. Therefore a β = ^^ is set as described above; see [ITU06b, ec. 163, 164]. The value of β is cut between 0.98 and 0.85 to avoid strong increases and decreases in energy; see [ITU06b, section 7.6.4],
Regarding the construction of the random part of the excitation, according to G.729.1, at the beginning of an erased block, the innovation gain g is initialized<sub>s</sub> using the innovation excitation gains of each subframe of the last correct frame:
<sub>9s</sub> = ().1/°) + 0.2(/<sup>1</sup>} + 0.3</<sup>2)</sup> + 0.4^<sup>3</sup>)
IMPI where g ^, g ^ and g ^ are the book gains of the four subframes of the last frame correctly.
<img file="MX347233B_D0010.tif" />
profit for innovation is realized as follows:
= λ · where g ^ is the innovation gain at the beginning of the next frame, g ^ is the innovation gain at the beginning of the current frame, and it is as previously defined in Table 2. Similar to the attenuation of With periodic excitation, the gain is thus attenuated in a linear manner throughout the frame sample by sample starting from g ^ and up to the value of g® that would be obtained at the beginning of the next frame.
According to G.729.1, if the last correct frame is NO VOICE, then only innovation drive is used and attenuated in turn by a factor of 0.8. In this case, the previous excitation buffer is updated with innovation excitation, since no periodic part of the excitation is available; see [ITU06b, section 7.6.6].
Hereinafter it is considered AMR. 3GPP AMR [3GP12b] is a speech codec that uses the ACELP algorithm. AMR is suitable for encoding speech with a sample rate of 8000 samples / s and a bit rate between 4.75 and 12.2 kbit / s and supports signaling of silence descriptor frames (DTX / CNG).
In AMR, during error concealment (see<sup>D</sup>[^^^), match between error-prone frames (bit errors) and traTTraTpertfWarpüT ^^ (absolutely no data).
In the case of ACELP concealment, AMR introduces a state machine that estimates the channel quality: The higher the value of the status counter, the worse the channel quality. The system boots to state 0. Each time a failed frame is detected, the status counter is incremented by one and saturates upon reaching 6. Each time a successful voice frame is detected, the status counter is reset to zero, except when the status is 6, when the status counter is set to 5. The flow of control of the state machine can be described by the following C code (BFI if it is a frame failed indicator, State is a state variable): if (BFI! = 0) {
State = State + 1; } else if (State == 6) {
State = 5;
else
State if (State> 6) {
State = 6;
IMPI ^
Mexican Institute of Industrial Property
<img file="MX347233B_D0011.tif" />
atrau * ·:
In addition to this state machine, in AMR, the failed frame flags of the current and previous frames (prevBFI) are checked.
There are three different possible combinations:
The first of the three combinations is BFI = 0, prevBFI = 0, State = 0: No error is detected in the received voice frame or in the previous one received. The received speech parameters are used in the normal way in speech synthesis. The current frame of voice parameters is saved.
The second of the three combinations is BFI = 0, prevBFI = 1, State = 0 or 5: No error is detected in the received voice frame, but the previous received voice frame was failed. The LTP gain and the fixed codebook gain are limited to less than the values used for the last correct subframe received:
(9p · 9p <9p {
1), 9p> 9pk 1) (10) where g<sub>p</sub> = current decoded LTP gain, g<sub>p</sub>(-1) = LTP gain used for the last correct subframe (BFI = 0), and
<img file="MX347233B_D0012.tif" />
9c <gc {-V)
9c> Pe (“l) (11)
IMPI • «τπντο Mexican M LA FROFIEGAD INDUSTRIAL
<img file="MX347233B_D0013.tif" />
where g<sub>c</sub> = current decoded fixed codebook gain, yg<sub>c</sub>(-1) = fixed codebook gain used for last correct subframe (BFI = 0).
The rest of the received speech parameters are normally used in speech synthesis. The current frame of voice parameters is saved.
The third of the three combinations is BFI = 1, prevBFI = 0 or 1, State = 1 ... 6: An error is detected in the received voice frame and the substitution and muting procedure is started. The LTP gain and the fixed codebook gain are replaced by attenuated values from the previous subframes;
<} p = <sub><</sub> Pástate) -Pp (-l), P (state) m. & I, iari5 (g<sub>p</sub>(— 1).
9<sub>P</sub>(-ty <median5 (g<sub>p</sub>(—L), ..., ^ (- 5)) g<sub>p</sub>(-l)> median. ^^. ., ^ (- 5)) (12) where g<sub>p</sub> indicates current decoded LTP gain and g<sub>p</sub>(-1), . . . , g<sub>p</sub>(-n) indicate the LTP gains used for the last n subframes and median5 () indicates a median operation of 5 points and
P (state) = attenuation factor, where (P (1) = 0.98, P (2) = 0.98, P (3) = 0.8, P (4) = 0.3, P (5 ) = 0.2, P (6) = 0.2) and
State = state number and
9c =
C (state) g<sub>c</sub>(—1), C (state) · median5 (g<sub>c</sub>(- l) ,.
9c (- ^)
IMPIAS
INSTITUTO MEXICANO '*, D £ LA PROPERTY g<sub>c</sub>(—1) <
g<sub>c</sub>{- 1)> Tnedian & (g<sub>c</sub>(—1),.
(13) where g<sub>c</sub> indicates current fixed codebook gain and g<sub>c</sub>(-1) ..... g<sub>c</sub> (-n) indicate the fixed codebook gains corresponding to the last n subframes and median5 () indicates a 5-point median operation and C (state) = attenuation factor, where (C (1) = 0.98 , C (2) = 0.98, C (3) = 0.98, C (4) = 0.98,
C (5) = 0.98, C (6) = 0.7) and State = state number.
In AMR, the LTP delay values (LTP = Long-Term Prediction) are replaced by the previous value of the 4<sup>to </sup>subframe of the previous frame (mode 12.2) or slightly modified values based on the last correctly received value (all other modes).
According to AMR, the fixed codebook innovation pulses received from the erroneous frame are used in the state that they were received upon receiving the corrupted data. In the case of not receiving any data, the random indexes of fixed code books should be used.
As regards AMR CNG, according to [3GP12a, section 6.4], every first lost SID frame is replaced using the SID information of previously received valid SID frames and the procedure for valid frames is applied. As for the subsequent lost SID frames, a comfort noise attenuation technique is applied which gradually has to reduce the output level. So it is checked if the last SID update occurred
IMPI
MEXICAN INSTITUTE νή <sub>F</sub> I GAVE OWNERSHIP more than 50 frames (= 1 5) before; if so, it is Wfltídeanirsíwa (attenuation level of -6/8 dB per frame [3GP12d, dtx_dec {} @sp_dec. cj which gives 37.5 dB per second). Note that the fading applied to CNG runs in the LP domain.
Hereinafter AMR-WB is taken into account. Adaptive Multirate - WB [ITU03, 3GP09c] is a voice codec, ACELP, based on AMR (see section 1.8). It uses parametric bandwidth extension and also supports DTX / CNG. Illustrative concealment solutions are given in the description of the [3GP12g] standard which are the same as for AMR [3GP12a] with minute deviations. Therefore, only differences to AMR are described here. For the description of the standard, see the description above.
Regarding ACELP, in AMR-WB, ACELP fading is executed on the basis of reference code [3GP12c] by modifying the tone gain g<sub>p</sub> (from previous AMR called LTP gain) and modifying the g-code gain<sub>c</sub>.
In the case of frame loss, the tone gain g<sub>p</sub> corresponding to the first subframe is the same as the last correct frame, except that it is limited between 0.95 and 0.5. Regarding the second, third and subsequent frames, the tone gain g<sub>p</sub> it is reduced by a factor of 0.95 and, again, it is limited.
AMR-WB proposes that in a hidden frame, g<sub>c</sub> is based on the latest g<sub>c</sub>:
(14) tic = ffc, current; * ff<sub>Ciuov</sub>
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY • (15)
<img file="MX347233B_D0014.tif" />
1.0
I, .............., í
V tarnaño_subplot
.............—
II (16) (17)
To hide the LTP delays, in AMR-WB, the history of the last five correct LTP delays and LTP gains is used to find the best method to update, in case of frame loss. In case of receiving the 10 bit error frame a prediction is executed as to whether the received LTP delay is usable or not [3GP12g].
As regards CNG, in AMR-WB, if the last correctly received frame was a SID frame and a frame is classified as lost, it should be replaced by the last valid SID frame information and the procedure for applying valid SID frames.
For subsequent lost SID frames, AMR-WB proposes the application of a comfort noise attenuation technique that gradually reduces the output level. Therefore, it is verified that the last SID update has been more than 50 frames (= 1 5 ·) before, if so, the
IMPI
INSTITUTO MEXICANO Λ1 ι ·, ζ · ιι. ., -,<sub>Λ lr</sub><. M LA PROPERTY output (attenuation level at -3/8 dB per frame [3GP12f, dtxj ^ féW which gives 18.75 dB per second). Note that the fade <Wentü apllcadu d CNO is<sup>1</sup> * performed in the domain of the LP.
It is now considered AMR-WB +. Adaptive Multispeed - WB + [3GP09a] is a switched codec using ACELP and TCX (TCX = Transform Coded Excitation, Transform Coded Excitation) as core codecs. It uses parametric bandwidth extension and also supports DTX / CNG.
In AMR-WB +, a mode extrapolation logic is applied to extrapolate the modes from the lost frames within a distorted superframe. This mode extrapolation is based on the fact that there is redundancy in the definition of mode indicators. The decision logic (exposed in [3GP09a, figure 18]) proposed by AMR-WB + is the following:
- A vector mode is defined, (mi, mo, mi, m<sub>2</sub>, m<sub>3</sub>), where mi indicates the mode of the last frame of the previous superframe and mo, m ^ m<sub>2</sub>, m<sub>3</sub> indicate the frame modes of the current superframe (decoded from the bitstream), where mk = -1, 0, 1, 2 or 3 (-1: lost, 0: ACELP, 1: TCX20, 2: TCX40, 3 : TCX80), and where the number of lost frames nloss can be between 0 and 4.
- If mi = 3 and two of the mode indicators in frames 0-3 are equal to three, all indicators are set to three because then it is certain that a TCX80 frame was indicated within the super frame.
WICKED
MWUCANQ INSTITUTE • OF THE PROPERTY
- If only one indicator of frames 0 - 3 is three * (and the numSró<sup>L</sup> detrarTTas losses nloss is three), the mode is set to (1, 1, 1,1), since then ^ 3/4 of the TCX80 target spectrum is lost and it is very likely that the overall TCX gain will be lost.
- If the mode is indicating (x, 2, -1, x, x) or (x, -1, 2, x, x), it is extrapolated to (x, 2, 2, x, x), indicating a TCX40 frame. If the mode indicates (x, x, x, 2, -1) or (x, x, —1, 2) it is extrapolated to (x, x, x, 2, 2), also indicating a TCX40 frame. It should be noted that (x, [0, 1], 2, 2, [0, 1]) are invalid configurations.
- After that, for each frame that is lost (mode = -1), the mode is set to ACELP (mode = 0) if the preceding frame was ACELP and the mode is set to TCX20 (mode = 1) in all other cases.
Regarding ACELP, according to AMR-WB +, if a lost frame mode results in mk = 0 after mode extrapolation, the same strategy as in [3GP12g] is applied to this frame (see above ).
In AMR-WB +, depending on the number of lost frames and the extrapolated mode, the following masking techniques related to TCX (TCX = Transform Encoded Excitation) are distinguished:
- In case of missing a complete frame, then a concealment similar to ACELP is applied: The last excitation is repeated and the hidden coefficients of ISF (slightly shifted towards its adaptive mean) are used to synthesize the signal in the time domain. In addition, a fading factor of 0.7 is multiplied per frame (20ms) [3GP09b, -> tai.HM * «WNB>
IMPI
INSTITUTO MEXICANO '& * í »£ j? ¡Íí * Λ dec_tcx. c] in the domain of linear prediction, the synthesis of LPC LPC (Linear Prediction Coding). ,, - .........-—
- If the last mode was TCX80 and also the extrapolated mode of the superframe (partially lost) is TCX80 (nloss = [1, 2], mode = (3, 3,
3, 3, 3)), concealment is performed in the FFT domain, using phase and amplitude extrapolation, taking into account the last correctly received frame. The phase information extrapolation strategy is not of interest here (not related to the fading strategy) and therefore is not described. For more details, see
[3GP09a, section 6.5.1.2.4], Regarding the amplitude modification of
AMR-WB +, the technique performed for TCX hiding consists of the following steps [3GP09a, section 6.5.1.2.3]:
- The magnitude spectrum of the previous frame is computed:
oldA [k] = oldX [k]
- The magnitude spectrum of the current frame is computed:
A [k] = Χ [λ ·]
The difference in gain of the energy of spectral coefficients not lost between the previous frame and the current one is computed:
/<sup>-</sup> gain ^^^^
MEXICAN INSTITUTE
The amplitude of the missing spectral coefficients is ^ Be using:
if (lost) [k]) A m 'gain, ant A f A';
- In all other cases of lost frame with mk = [2, 3], the TCX target (inverse FFT of the decoded spectrum most filled with noise (using a decoded noise level of the bitstream)) is synthesized using all the information available (including global TCX gain). No fading is applied in this case.
As regards CNG in AMR-WB +, the same technique is used as in AMR-WB (see above).
Hereinafter, it is considered OPUS. OPUS [IET12] incorporates two-code technology: the voice-oriented SILK (known as Skype codec) and low-latency CELT (CELT = Power Restricted Overlap Transform). Opus can be seamlessly tuned between high and low bit rates and internally switches between a lower bit rate linear prediction codec (SILK) and a higher bit rate transform codec (CELT), as well as also a hybrid for brief overlaps.
When it comes to compression and decompression of SILK audio data, in OPUS, there are several parameters that are attenuated during concealment in the SILK decoder routine. The LTP gain parameter is attenuated by multiplying all the LPC coefficients by 0.99, 0.95 or 0.90 per frame,
WICKED
MEXICAN INSTITUTE β.
depending on the number of consecutive lost frames, 3on ^<sup>P</sup>^ exa «SgÍ ^« accumulates using the last tonal cycle obtained from the excitation of the previous frame. The pitch delay parameter increases very slowly during consecutive losses. In the case of individual losses it remains constant compared to the last frame. Furthermore, the excitation gain parameter attenuates exponentially with o.99<sup>/or</sup>^<sup>cw</sup> per frame, so the drive gain parameter is 0.99 in the case of the first drive gain parameter, so the drive gain parameter is 0.992 for the second drive gain parameter, and so on . Excitation is generated using a random number generator that generates variable overflow white noise. In addition, the LPC coefficients are extrapolated / averaged on the basis of the last correctly received series of coefficients. After generating the attenuated excitation vector, the LPC coefficients hidden in OPUS are used to synthesize the output signal in the time domain.
CELT is now considered in the context of OPUS. CELT is a transform-based codee. CELT concealment includes a tone-based PLC strategy, which is applied for up to five consecutively lost frames. Starting from frame 6, a noise-like masking technique is applied, which generates background noise, a characteristic that is supposed to sound preceding the background noise.
IMPI ^
MEXICAN INSTITUTE
Say THE PROPERTY
Fig. 5 illustrates the loss behavior in rapha ^ iHfe CÉtrr In particular, Fig. 5 illustrates a spectrogram (x-axis: time; y-axis: frequency) of a voice segment hidden by CELT. The light gray box indicates the first 5 consecutively lost frames, where a pitch-based PLC strategy is applied. Beyond that, noise-like concealment is illustrated. It should be noted that the switching is done instantaneously, the transition is not smooth.
Regarding tone-based concealment, in OPUS, tone-based concealment consists of finding the periodicity in the decoded signal by autocorrelation and repeating the windowed waveform (in the excitation domain using analysis and synthesis of LPC) using pitch shift (pitch delay). The windowed waveform overlaps in such a way as to preserve the cancellation of overlap in the time domain with the previous frame and the next frame [IET12]. Furthermore, a fading factor is derived and applied using the following code: opus val32 El = l, E2 = l; int period;
if (pitch_index <= MAX_PERIOD / 2) {period = pitch_index;
else {period = MAX_PERIOD / 2;
} for (i = 0; i <period; i ++)
<img file="MX347233B_D0015.tif" />
The + = exc [MAX_PERIOD- period + i] period + i];
E2 + = exc [MAX_PERI0D-2 * period + i] * exc [MAX_PERIOD2 * period + i]; } if (El> E2) {
El = E2;
} decay = sqrt (E1 / E2));
attenuation = decay;
In this code, exc contains the excitation signal until MAX_PERIOD samples before loss.
The excitation signal is then multiplied with attenuation, then synthesized and sent by LPC synthesis.
The fading algorithm for the time domain strategy can be summarized as follows:
- Find the synchronous energy of the tone of the last tonal cycle before the loss.
- Find the synchronous tone energy of the second last tonal cycle before the loss.
<img file="MX347233B_D0016.tif" />
IMPI ^
INSTITUTO MEXICANO DE LA PROriEDAC __
If energy is increasing, limit it to keep it going:
attenuation = 1
- If the energy is reducing, it continues with the same attenuation during concealment.
Regarding noise concealment, according to OPUS, for the 6<sup>to</sup> and the next consecutive lost frames a noise substitution strategy is executed in the MDCT domain in order to simulate background comfort noise.
Regarding the background noise level and shape tracking, in OPUS, the background noise calculation is executed as follows: After the MDCT analysis, the square root of the MDCT band energies is calculated by frequency band, where the grouping of the MDCT boxes follows the barks scale according to [IET12, Table 55], then the square root of the energies is transformed to the log domain<sub>2</sub> according:
bandLogE [i] = logz (e) · log<sub>and</sub>(bandE [i \ - eMeans \ i \) for i = 0 ... 21 (18) where e is the Euler number, bandE is the square root of the MDCT band and eMeans is a vector of constants (necessary for obtain the mean zero result, which results in increased encoding gain).
In OPUS, the background noise on the decoder side is entered as follows [IET12, amp2Log2 and log2Amp @ quant bands.c]:
(19) backgroundLogE [i]
IMPI ^ 3
MEXICAN INSTITUTE
OF THE PROPERTY = min (backgroundLogE [i] + 8 · () for i = 0 ... ¿í
The minimum energy tracked is basically determined by the square root of the energy of the current frame band, although the increase from one frame to the next is limited to 0.05 dB.
Regarding the application of the level and shape of the background noise, according to OPUS, if the noise-type PLC is applied, backgroundLogE derived in the last correct frame is used and it is converted back to the linear domain:
for (20) where is the Euler number and eMeans is the same vector of constants as in the case of the “linear to log” transform.
The present screening procedure consists of filling the MDCT frame with white noise produced by a random number generator and scaling this white noise so that it coincides, in the band, with the energy of bandE. The inverse MDCT is then applied which gives rise to a signal in the time domain. After overlapping and addition and de-emphasis (as in normal decoding) it is output.
Hereinafter referred to as MPEG-4 HE-AAC (MPEG = Moving Picture
Experts Group, Group of Experts on Moving Images; HE-AAC = High
IMPI ^^ mbxicano institute
DB LA MOHEDAL ·
Efficiency Advanced Coding (Audió'WAIta Efficiency Advanced Coding).
Advanced High Efficiency Audio Coding — Cóháiste<sup>1</sup> in a transform-based codeé (AAC), supplemented by a parametric bandwidth extension (SBR).
As regards AAC (AAC = Advanced Audio Coding), the DAB consortium specifies for AAC in DAB +, a fade to zero in the frequency domain [EBU10, section A1.2] (DAB = Digital Audio Broadcasting, ( Digital Audio Transmission) / The fading behavior, eg the attenuation ramp, could be fixed or user adjustable. The spectral coefficients corresponding to the last AU (AU =
Access Unit) are attenuated by a factor corresponding to the fading characteristics and then passed to frequency mapping in time. Depending on the attenuation ramp, the concealment switches to mute after a number of consecutive invalid AUs, which means that the entire spectrum is set to 0.
The DRM consortium (DRM = Digital Rights Management) specifies for the DRM CAA a fading in the frequency domain [EBU12, section 5.3.3]. Concealment acts on the spectral data just before the final frequency-to-time conversion. If multiple frames are corrupted, concealment first implements a fading based on slightly modified spectral values relative to the last valid frame. Furthermore, as with DAB +, the fading behavior, e.g. the attenuation ramp, could be
IMPI
MEXICAN INSTITUTE
INDUSTRIAL PROPERTY • Fixed or user adjustable. The spectral coefficients of the last frame are attenuated by a factor corresponding to the fading characteristics and then passed to frequency mapping in time. Depending on the attenuation ramp, the concealment switches to mute after a number of consecutive invalid frames, which means that the entire spectrum is set to 0.
3GPP introduces DRM-like frequency domain fading for AAC in aacPlus [3GP12e, section 5.1]. Concealment acts on the spectral data just before the final frequency-to-time conversion. If multiple frames are corrupted, concealment first implements a fading based on slightly modified spectral values of the last good frame. A complete fade takes 5 frames. The spectral coefficients of the last correct frame are copied and attenuated by a factor of:
fadeOutFac = 2 - (nFadeOutFrame / 2) with nFadeOutFrame as the counter of frames since the last good frame.
After five frames of squelch the concealment switches to squelch, which means that the entire spectrum is set to 0.
Lauber and Sperschneider introduce a raster fading of the MDCT spectrum for AAC, based on energy extrapolation [LS01, section 4.4]. The energy forms of a previous spectrum could be used to extrapolate the shape of an estimated spectrum. Energy extrapolation is
IMPI
INSTITUTO MEXICANO can independently execute ocL ^ afflféntc ^ 3Sra ^ na kind of post-concealment techniques. ---------।
For AAC, the energy calculation is performed based on a band of scale factors to choose the critical bands of the human auditory system. Individual power values are lowered frame by frame to reduce volume smoothly, eg to fade the signal. This becomes essential since the probability that the estimated values represent the current signal decreases rapidly with time.
For the generation of the spectrum to be transmitted, they suggest the repetition of frames or the substitution of noise [LS01, sections 3.2 and 3.3].
Quackenbusch and Driesen suggest for AAC an exponential fading per frame to zero [QD03]. A repetition of an adjacent series of time / frequency coefficients is proposed, where each repetition has increasing attenuation, thus gradually fading to mute in the case of extended suspensions.
Regarding SBR (SBR = Spectral Band Replication in MPEG-4 HE-AAC, 3GPP suggests for SBR in Enhanced aacPlus to buffer the decoded envelope data and, in case of loss of frame, reuse the buffered energies of the transmitted envelope data and reduce it by a constant 3 dB ratio for each hidden frame. The result is fed into a normal decoding process where it is used by the envelope adjuster to calculate the gains, used to adjust the patched high bands generated by the encoder.
<img file="MX347233B_D0017.tif" />
<sup>3</sup> IMPI
MEXICAN INSTITUTE
OF THE KKOP1EDAD
INDUSTRIAL
HF. SBR decoding then takes place as normal. Furthermore, the floor of «« • ΟΜΜννΜΜΜΜΜΜΜνΜΝΜΜΜΚΒΜΒαΜΤΜηβΓαΒ »?
delta modulation encoded noise and sine level values are being erased. Since there are no differences with the previous information, the decoded noise floor and sine levels remain proportional to the energy of the generated HF signal [3GP12e, section 5.2].
The DRM consortium specifies for SBR in combination with AAC the same technique as 3GPP [EBU12, section 5.6.3.1]. Furthermore, the DAB consortium specifies for SBR in DAB + the same technique as 3GPP [EBU10, section A2],
Hereinafter, MPEG-4 CELP and MPEG-4 HVXC (HVXC = Harmonium Vector Excitation Coding)) are considered. The DRM consortium specifies for SBR, in connection with CELP and HVXC [EBU12, section 5.6.3.2] that the minimum concealment requirement for SBR corresponding to voice codes is to apply a predetermined series of data values, provided that it has been detected a corrupt SBR frame. Those values give a static high band spectral envelope at a relatively low reproduction level, exhibiting a discharge towards the higher frequencies. The goal is simply to ensure that no unpleasant, potentially loud audio bursts reach the listener's ears, through the insertion of "comfort noise" (as opposed to strict muting). This is not actually a true fading but rather a jump to a certain energy level to insert some kind of comfort noise.
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
An alternative is mentioned below [EBU12, 5.t ^^ íjue reuses the last correctly decoded data and slowly fades the levels (L) towards 0, analogously to the case of AAC + SBR.
It is now considered MPEG-4 HILN (HILN = Harmonics and Individual Lines plus Noise). Meine et al. introduce a fade out for the MPEG-4 HILN [ISO09] parametric codec in a parametric domain [MEP01]. In the case of harmonic components, a good default replacement behavior for corrupted differentially encoded parameters is to keep the frequency constant to reduce the amplitude by an attenuation factor (e.g. -6 dB), and to allow that the spectral envelope converges towards that of the averaged low-pass characteristic. An alternative to the spectral envelope would be to keep it unchanged. With respect to amplitudes and spectral envelopes, noise components can be treated in the same way as harmonic components.
Hereinafter, background noise level tracking of the prior art is discussed. Rangachari and Loizou [RL06] present a good review of various methods and explain some of their limitations. Methods for tracking the background noise level are, for example, the minimal trace procedure [RL06] [Coh03] [SFB00] [Dob95], based on VAD (VAD = voice activity detection); Kalman filtering [Gan05] [BJH06], subspace decompositions [BP06] [HJH08]; Soft Decision [SS98] [MPC89] [HE95] and minimal statistics.
<sup>37</sup> IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL ^ * L —The minimum statistics strategy was chosen within the scope for USAC-2, (USAC = Unified Speech and Audio Coding, Unified Voice and Audio Coding)) and is described in more detail below.
The noise power spectral density estimate based on optimal smoothing and minimum statistics [Mar01] introduces a noise estimator capable of operating regardless of whether the signal is active speech or background noise. Unlike other methods, the minimum statistics algorithm does not use any explicit threshold to distinguish between voice activity and voice pause, and is therefore more closely related to soft decision methods than to traditional voice detection methods. human voice activity. Similar to the soft decision methods, you can also update the PSD (Power Spectral Density) of estimated noise.
The minimal statistics method is based on two observations, that is, speech and noise are usually statistically independent, and the power of a noisy speech signal drops to the noise power level. Therefore, it is possible to derive an accurate PSD (PSD = power spectral density) estimate of the noise by tracking the PSD minimum of the noisy signal. Since the minimum is less than (or in some cases equal to) the average value, the minimum tracking method requires compensation for bias (“bias”).
The deviation is a function of the variance of the PSD of the smoothed signal and, therefore, depends on the smoothing parameter of the PSD estimator. Unlike the Mexican institute Ρ »LA PROPERTY INDUSTRIAL previous work on minimum tracking; Using a constant smoothing parameter and a constant minimum deviation correction, a frequency and time-dependent PSD smoothing is used, which also requires frequency and time-dependent deviation compensation.
Using minimum tracking gives a rough estimate of the noise power. However, it has some disadvantages. Smoothing with a fixed smoothing parameter widens the speech activity peaks of the smoothed PSD estimate. This leads to inaccurate noise estimates, since the moving window for the search for the minimum could slide towards wider peaks. Consequently, smoothing parameters close to one cannot be used, and consequently the noise estimate must have a relatively larger variance. Furthermore, the noise estimate is biased towards the lower values. In addition, in case of increasing the noise power, the tracking of the minimum is delayed.
Low complexity MMSE-based noise PSD tracking [HHJ10] introduces a background noise PSD technique that uses an MMSE search used on a DFT spectrum (Discrete Fourier Transform). . The algorithm consists of these processing steps:
- The maximum probability estimator is computed on the basis of the noise PSD of the previous frame.
- The mean square error estimator is computed.
IMPI
MEXICAN INSTITUTE „....__. _ ..___. . _ ...... .PELAfRQHtDAD
- The probability estimator n ^ wmeLUtHñ3B3SM6 decision-directed technique [EM84] is estimated. · ———
- The inverse deviation is computed assuming that the voice and noise DFT coefficients have a Gaussian distribution.
- The estimated noise power spectral density is smoothed.
There is also a safety net strategy applied to avoid a complete algorithm crash.
Non-stationary noise tracking based on data-based recursive noise power estimation [EH08] introduces a method for estimating the noise spectral variance of voice signals contaminated by highly non-stationary noise sources. This method also uses smoothing in the time / frequency direction.
A low complexity noise estimation algorithm based on smoothing the noise power estimate and correcting deviations in the estimate [Yu09] powers the strategy introduced in [EH08]. The main difference is that the spectral gain function corresponding to the noise power estimate is found by an iterative method driven by data.
The statistical methods for the potentiation of the noisy voice [Mar03] combine the minimum statistics strategy given in [Mar01] by modifying the gains with soft decision [MCA99], by an estimate of the previous SNR [MCA99], by a limitation of adaptive gain [MC99] and by an MMSE logarithmic spectral amplitude estimator [EM85].
WICKED
MEXICAN INSTITUTE
OI THE PROPERTY
Fade-out is of particular interest from voice and audio codecs, in particular, AMR (see [3GPI2b]) (including ACELP and ~ CNG), AMR-WB (see [3GP09c]) (including ACELP and CNG), AMR -WB + (see [3GP09a]) (including ACELP, TCX and CNG), G.718 (see [ITU08a]), G.719 (see [ITU08b]), G.722 (see [ITU07]), G.722.1 (see [ITU05]), G.729 (see [ITU12, CPK08, PKJ + 11]), MPEG-4 HE-AAC / aacPlus Enhanced (see [EBU10, EBU12, 3GP12e, LS01, QD03]) (including AAC and SBR), MPEG-4 HILN (see [ISO09, MEP01]) and OPUS (see [IET12]) (including SILK and CELT).
Depending on the codec, the fade runs in different domains:
For codecs using LPC, fading is in the linear prediction domain (also known as the excitation domain). This is so in the case of ACELP based codes, eg AMR, AMR-WB, the ACELP core of AMR-WB +, G.718, G.729, G.729.1, the SILK core of OPUS; the codecs that further process the drive signal using a time-frequency transformation, eg, the TCX core of AMR-WB +, the CELT core of OPUS; and in the case of comfort noise generation (CNG) schemes, which operate in the domain of linear prediction, eg, CNG in AMR, CNG in AMR-WB, CNG in AMR-WB +.
In the case of codes that directly transform the time signal in the frequency domain, fading is done in the spectral / subband domain. This is so in the case of codes that are based on
IMPI ^^
INSTITUTO MEXICANO V ^ i-SSST; From THE PRGPltDAb
MDCT or a similar transformation, such as ÁÁC in Ι \ Λϊ ^ 8 ^ HETWÚ?
G.719, G.722 (subband domain) and G.722.1. ''
In the case of parametric codes, the fading is applied in the parametric domain. This is so in the case of MPEG-4 HILN.
Regarding fading rate and fading curve, a fade is commonly obtained by applying an attenuation factor, which is applied to the representation of the signal in the appropriate domain. The size of the fade factor controls the fading rate and fading curve. In most cases the attenuation factor is applied per frame, although one application per sample is also used; see, eg, G.718 and G.722.
The attenuation factor for a certain segment of the signal can exist in two ways, absolute and relative.
In the case where an attenuation factor is presented in absolute form, the reference level is always that of the last frame received. The absolute attenuation factors usually start with a value close to 1 corresponding to the segment of the signal immediately after the last correct frame and then degrade more or less rapidly towards 0. The fading curve depends directly on these factors. This is the case, for example, of the concealment described in Annex IV of G.722 (see, in particular, [ITU07, Figure IV.7]), where the possible fading curves are linear or gradually linear. Considering a gain factor g (n), while g (0)
<img file="MX347233B_D0018.tif" />
IMPI «NSTITUTO MEXICANO DE LA PROPERTY INDUSTRIAL represents the gain factor of the last correct frame, an absolute attenuation factor at<sub>to</sub>í><sub>s</sub>(n), the gain factor of any subsequent lost frames can be derived as follows g (n) = a<sub>to</sub>bs (n) g (Q) (21)
In the case where an attenuation factor appears in relative form, the reference level is from the previous frame. This offers advantages in the case of a recursive masking procedure, eg if the already attenuated signal is reprocessed and attenuated once more.
If an attenuation factor is recursively applied, then this could be a fixed value independent of the number of consecutively lost frames, eg 0.5 in the case of G.719 (see above); a fixed value with respect to the number of consecutively lost frames as proposed, eg for the case of G.729 in [CPK08]: 1.0 for the first two frames, 0.9 for the next two frames, 0 , 8 for frames 5 and 6, and 0 for all subsequent frames (see above); or a value that is relative to the number of consecutively lost frames and that depends on the characteristics of the signal, eg faster fading for an unstable signal and a slower fading for a stable signal, eg G. 718 (see previous section and [ITU08a, table 44]);
<img file="MX347233B_D0019.tif" />
TíSTTTVTO MEXICANO J.
. OF THE PROPERTY __
Assuming a fading factor 0 ^^
1, while n is the number of the lost frame (η> I'J, WRTCtü! 'Of gain of any subsequent frame can be derived in the following way g (n) = a<sub>re</sub>i (n) -g (n- 1) g (n) (<sup>n</sup> \
Π s (°) m = l / (23) g (n) = o $<sub>he</sub> · Í / (0) (24) resulting in an exponential fading.
As regards the fade-out procedure, the attenuation factor is usually specified, although in some application standards (DRM, DAB +) the latter is left to the manufacturer.
If different parts of the signal are faded separately, different attenuation factors can be applied, eg to fade tonal components at one speed and noise-type components at another speed (eg, AMR, SILK).
Usually a certain gain is applied to the entire frame. When the fading is done in the spectral domain, this is the only possible way. However, if the fading is in the time domain or
<img file="MX347233B_D0020.tif" />
in the domain of linear prediction, more granular fading is possible. Such more granular fading is applied in G.718, where individual gain factors for each sample are derived by linear interpolation between the gain factor of the last frame and the gain factor of the current frame.
As for the codes with variable frame duration, a constant relative attenuation factor leads to a different fading rate which depends on the duration of the frame. This is the case, for example, of AAC, where the duration of the frame depends on the sampling rate.
To adapt the fading curve applied to the time shape of the last received signal, the fading factors (static) could be further adjusted. This additional dynamic adjustment is applied, for example, in the case of AMR where the median of the five previous gain factors is taken into account (see [3GP12b] and section 1.8.1). Before performing any attenuation, the current gain is adjusted to the median, if the median is less than the last gain; otherwise the last gain is used. Furthermore, such additional dynamic adjustment applies, eg, in the case of G729, where the amplitude is predicted using linear regression of the previous gain factors (see [CPK08, PKJ + 11] and section 1.6). In this case, the gain factor obtained corresponding to the first hidden frames could exceed the gain factor of the last received frame.
Regarding the target level of fading, except for
G.718 and CELT, the target level is 0 for all analyzed codecs, including comfort noise generation (CNG) from those codecs.
IMPI
INSTITUTO MEXICANO DB LA PROPERTY INDUSTRIAL
In G.718, the pitch drive fading (representing the tonal components) and the random drive fading (representing the noise-like components) are performed separately. Although the tonal gain factor fades to zero, the innovation gain factor fades to the driving energy of CNG.
Assuming you are given the relative attenuation factors, this leads based on formula (23) - to the following absolute attenuation factor:
ÍZ (n) = arelan) g (n -1) + (1 - a<sub>r</sub>the (n)) g<sub>n</sub> where g<sub>n</sub> is the gain of the excitation used during the generation of comfort noise. This formula corresponds to formula (23), when g<sub>n</sub> = 0.
G.718 does not perform fading in the case of DTX / CNG.
In CELT there is no fading towards the target level, but after 5 frames of pitch squelch (including fading) it is instantly switched to the target level at 6 o'clock.<sup>to</sup> consecutively lost plot. The level is derived by bands using formula (19).
Regarding the target spectral shape of the fading, all analyzed pure transform-based codes (AAC, G.719, G.722, G.722.1) as well as SBR simply extend the spectral shape of the last correct frame for the fading.
<img file="MX347233B_D0021.tif" />
IMPI
MEXICAN INSTITUTE
OF THE PHOHKMD
INDUSTRIAL
Various voice codes fade the spectral shape to a mean • eewswweewwwwwraeaweeiíafcL'Wi using LPC synthesis. This mean could be static (AMR) or adaptive (AMR-WB, AMR-WB +, G.718), while the latter is derived from a static mean and a short-term mean (which is derived by averaging the last n series of LP coefficients) (LP = Linear Prediction).
All the CNG modules of the described codes AMR, AMR-WB, AMRWB +, G.718 extend the spectral shape of the last correct frame during fading.
When it comes to tracking the background noise level, there are five different strategies reported in the literature:
- Voice Activity Detector: based on SNRA / AD, although very difficult to tune and difficult to use for low SNR voice.
- Soft decision scheme: The soft decision strategy takes into account the probability of voice presence [SS98] [MPC89]
[HE95],
- Minimum statistics: The minimum of PSD is tracked, maintaining a certain amount of values over time in a buffer or buffer, thus allowing the finding of the minimum noise of the previous samples [Mar01] [HHJ10] [EH08] [Yu09],
- Kalman filtering: The algorithm uses a series of measurements observed over time, with noise content (random variations), and produces estimates of the noise PSD that tend to be more accurate than those based on a single measurement. Filter
<img file="MX347233B_D0022.tif" />
Kalman recursively operates on noisy input data streams to produce a statistically optimal estimate of the state of the system [Gan05] [BJH06].
- Subspace Decomposition: This strategy attempts to decompose a noise-like signal into a clean voice signal and a noise part, using, for example, the KLT (Karhunen-Loéve transform, Karhunen-Loéve transform, also known as analysis of principal components) and / or the DFT (Discrete Time Fourier Transform, Discrete Time Fourier Transform). Then the eigenvectors / eigenvalues can be traced using an arbitrary smoothing algorithm [BP06] [HJH08],
The object of the present invention is to provide improved concepts for audio coding systems. The object of the present invention is achieved by an apparatus according to claim 1, a method according to claim 21 and a computer program according to claim 22.
An apparatus for decoding an audio signal is presented.
The apparatus comprises a reception interface. The reception interface is configured to receive a plurality of frames, where the reception interface is configured to receive a first frame of the plurality of frames, where said first frame comprises a first portion of the audio signal of the audio signal, where said first portion of the audio signal is represented in a first domain, and where the receiving interface is
IMPI
Mexican Wsthuto DE LA PROPERTY INDUSTRIAL configured to receive a second frame of the plurality of frames, where said second frame comprises a second portion of the audio signal of the audio signal.
Furthermore, the apparatus comprises a transformation unit for transforming the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from a second domain to a tracking domain to obtain information of the second portion of the signal, where the second domain is different from the first domain, where the tracking domain is different from the second domain, and where the tracking domain is the same or different from the first domain.
In addition, the apparatus comprises a noise level tracking unit, where the noise level tracking unit is configured to receive information from the first portion of the signal that is represented in the tracking domain, where the information from the first portion of the signal depends on the first portion of the audio signal. The noise level tracking unit is configured to receive the second portion of the signal that is represented in the tracking domain, and where the noise level tracking unit is configured to determine noise level information depending on the information of the first portion of the signal that is represented in the tracking domain and depending on the information of the second portion of the signal that is represented in the tracking domain.
Furthermore, the apparatus comprises a reconstruction unit for reconstructing a third portion of the audio signal from the audio signal.
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL depending on the noise level information, if a third frame of the plurality of frames is not received by the receiving interface but is corrupted.
An audio signal can be, for example, a voice signal or a music signal, or a signal comprising voice and music, etc.
The statement that the information in the first portion of the signal depends on the first portion of the audio signal means that the information in the first portion of the signal is the first portion of the audio signal, or has been obtained / generated the information of the first portion of the signal depending on the first portion of the audio signal or that in some other way depends on the first portion of the audio signal. For example, the first portion of the audio signal may have been transformed from one domain to another domain to obtain the information from the first portion of the signal.
Similarly, the statement that the information in the second portion of the signal depends on a second portion of the audio signal means that the information in the second portion of the signal is the second portion of the audio signal, or that the information of the second portion of the signal has been obtained / generated depending on the second portion of the audio signal or in some other way dependent on the second portion of the audio signal. For example, the second portion of the audio signal may have been transformed from one domain to another domain to obtain information from the second portion of the signal.
In one embodiment, the first portion of the audio signal may be represented, eg, in a time domain as the first domain. Furthermore, the transformation unit may be configured, e.g., to
IΜ ΡI ·.
INSTITUTO MÍJUCa ^ '. · Λ
OF THE EROP1ΕΟ *:
transform the second portion of the audio signal or the value ^ Wf ^ do ^ d ^ féF * ^ second portion of the audio signal of a domain of the second domain, tfn'e'esngl ', to the time domain that is the tracking domain. In addition, the noise level tracking unit may be configured, eg, to receive the information from the first portion of the signal that is represented in the time domain as the tracking domain. Furthermore, the noise level tracking unit may be configured, eg, to receive the second portion of the signal that is represented in the time domain as the tracking domain.
According to one embodiment, the first portion of the audio signal may be represented, eg, in the excitation domain as * first domain. Furthermore, the transform unit may be configured, eg, to transform the second portion of the audio signal or the value derived from the second portion of the audio signal from a time domain which is the second domain to the domain. of arousal that is the tracking domain. In addition, the noise level tracking unit may be configured, eg, to receive the information from the first portion of the signal that is represented in the drive domain as the tracking domain. Furthermore, the noise level tracking unit may be configured, eg, to receive the second portion of the signal that is represented in the drive domain as the tracking domain.
IMPI
MEXICAN INOTITC
In one embodiment, the first portion of the will be represented, eg, in the domain of the domain, as the primary domain. where the noise level tracking unit may be configured, e.g., to receive the information from the first portion of the signal, where said information from the first portion of the signal is represented in the FFT domain, which is the tracking domain, and where said information of the first portion of the signal depends on said first portion of the audio signal that is represented in the excitation domain, where the transformation unit can be configured, e.g. to transform the second portion of the audio signal or the derived value of the second portion of the audio signal from a time domain that is the second domain to an FFT domain that is the tracking domain, and where the noise level tracking unit may be configured, eg, to receive the second portion of the audio signal that is represented in the FFT domain.
In one embodiment, the apparatus may further comprise, eg, a first aggregation unit for determining a first aggregate value that depends on the first portion of the audio signal. Furthermore, the apparatus may further comprise eg a second aggregation unit for determining, depending on the second portion of the audio signal, a second aggregated value as a derived value of the second portion of the audio signal. In addition, the noise level tracking unit may be configured, e.g., to receive the first added value as information from the first portion of the signal that is represented in the tracking domain, where the tracking unit
IMPI
MEXICAN INSTITUTE> ·,
......,, „,. OF THE NOISE LEVEL PROPERTY can be set, for example, for redlt ^ Wsegefl ^ tfálor added as information of the second portion of the signal that is represented in the trace domain, and where the noise level trace unit can be configured, e.g. to determine noise level information depending on the first aggregate value that is represented in the tracking domain and depending on the second aggregate value that is represented in the tracking domain.
According to one embodiment, the first aggregation unit may be configured, eg, to determine the first aggregated value in such a way that the first aggregated value indicates a root mean square of the first portion of the audio signal or of a signal derived from the first portion of the audio signal. Furthermore, the second aggregation unit may be configured, e.g., to terminate the second aggregate value in such a way that the second aggregate value indicates a root mean square of the second portion of the audio signal or of a derived signal. of the second portion of the audio signal.
In one embodiment, the transform unit may be configured, eg, to transform the derived value of the second portion of the audio signal from the second domain to the tracking domain by applying a gain value to the derived value. of the second portion of the audio signal.
According to these embodiments, the gain value may indicate, for example, a gain introduced by the encoding synthesis.
MEXICAN INSTITUTE
OF PROPERTY C y,, INDUSTRIAL> ^ ** 88 ** »^ ·
Linear Prediction, or the gain value can indicate eg a gain introduced by the synthesis and de-emphasis of Linear Prediction Coding.
In one embodiment, the noise level tracking unit may be configured, eg, to determine noise level information by applying a minimal statistics strategy.
According to one embodiment, the noise level tracking unit may be configured, eg, to determine a comfort noise level as noise level information. The reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal depending on the noise level information, if said third frame of the plurality of frames is not received by the receiving interface or if said third frame is received by the receiving interface but it is corrupted.
In one embodiment, the noise level tracking unit may be configured, eg, to determine a comfort noise level as noise level information derived from a noise level spectrum, where said level spectrum noise is obtained by applying the minimum statistics strategy. The reconstruction unit may be configured, e.g., to reconstruct the third portion of the audio signal depending on a plurality of linear prediction coefficients, if said third frame of the plurality of frames is not received by the receiving interface or if said third frame is received by the receiving interface but is corrupt.
According to another embodiment, the noise level tracking unit may be configured, e.g., to determine a plurality of
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL linear prediction coefficients indicating a noise level to conform to noise level information, and the reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal depending on the plurality of linear prediction coefficients.
In one embodiment, the noise level tracking unit is configured to determine a plurality of FFT Coefficients that indicate a comfort noise level as noise level information, and the first reconstruction unit is configured to reconstruct the third portion of the audio signal depending on a comfort noise level derived from said FFT coefficients, if said third frame of the plurality of frames is not received by the reception interface or if said third frame is received by the reception interface but is corrupted.
In one embodiment, the reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal depending on the noise level information and depending on the first portion of the audio signal, if said. third frame of the plurality of frames is not received by the receiving interface or if said third frame is received by the receiving interface but is corrupted.
According to one embodiment, the reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal by attenuating or amplifying a signal derived from the first or second portion of the audio signal. .
MEXICAN INSTITUTE
OF THE PROPERTY
In one embodiment, the apparatus can buy a long-term prediction unit comprising a delay buffer. Furthermore, the long-term prediction unit may be configured, e.g., to generate a processed signal depending on the first or second portion of the audio signal, depending on an input from the delay buffer that is buffered. delay and depending on a long-term prediction gain. In addition, the long-term prediction unit may be configured, e.g., to fade the long-term prediction gain towards zero, if said third frame of the plurality of frames is not received by the receiving interface or if said third frame is received by the receiving interface but is corrupted.
According to one embodiment, the long-term prediction unit may be configured, eg, to fade the long-term prediction gain toward zero, where the rate at which the long-term prediction gain fades to zero depends on a fading factor.
In one embodiment, the long-term prediction unit may be configured, eg, to update the delay buffer input by storing the generated processed signal in the delay buffer, if said third frame of the plurality of frames does not is received by the receiving interface or if said third frame is received by the receiving interface but is corrupted.
According to one embodiment, the transformation unit can be, eg, a first transformation unit, and the reconstruction unit is a first reconstruction unit. The apparatus comprises
<img file="MX347233B_D0023.tif" />
<sup>56</sup> IMPI Mexican iNsTmrro
OF THE PRCMtDAO
INDUSTRIAL also a second transformation unit and a second reconstruction unit. The second transformation unit may be configured, eg, to transform noise level information from the tracking domain to the second domain, if a fourth frame of the plurality of frames is not received by the receiving interface or if said fourth frame is received by the receiving interface but is corrupted. Furthermore, the second reconstruction unit may be configured, eg, to reconstruct a fourth portion of the audio signal from the audio signal depending on the noise level information that is represented in the second domain if said fourth frame of the plurality of frames is not received by the receiving interface or if said fourth frame is received by the receiving interface but is corrupted.
In one embodiment, the second reconstruction unit may be configured, eg, to reconstruct the fourth portion of the audio signal depending on the noise level information and depending on the second portion of the audio signal.
According to one embodiment, the second reconstruction unit may be configured, eg, to reconstruct the fourth portion of the audio signal by attenuating or amplifying a signal derived from the first or second portion of the audio signal. Audio.
A method for decoding an audio signal is also presented.
The method comprises:
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
Receiving a first frame of a plurality of frames, wherein said first frame comprises a first portion of the audio signal of the audio signal, wherein said first portion of the audio signal is represented in a first domain.
Receiving a second frame of the plurality of frames, wherein said second frame comprises a second portion of the audio signal of the audio signal.
Transform the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from a second domain to a tracking domain to obtain information from the second portion of the signal, where the second domain is different from the first domain, where the crawl domain is different from the second domain, and where the crawl domain is the same or different from the first domain.
Determine noise level information depending on the information of the first portion of the signal, which is represented in the tracking domain, and depending on the information of the second portion of the signal that is represented in the tracking domain, where the Information of the first portion of the signal depends on the first portion of the audio signal Y:
- Reconstructing a third portion of the audio signal from the audio signal depending on the noise level information that is represented in the tracking domain, if a third frame of the plurality
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL of frames is not received or if said third frame is received but is corrupt.
In addition, a computer program is presented to implement the method described above when running on a computer or signal processor.
Some of the embodiments of the present invention disclose a time-varying smoothing parameter such that the tracking capabilities of the smoothed priodogram and its variance are better balanced, to develop an algorithm to compensate for deviations and to accelerate noise tracking in general.
The embodiments of the present invention are based on the finding that, with respect to fading, the following parameters are of interest: The domain of fading; the fading rate, or more generally, the fading curve; the target level of fading; the target spectral shape of the fading and / or background noise level tracking. In this context, the embodiments are based on the finding that the prior art has significant disadvantages.
An apparatus and method for improved signal fading for switched audio coding systems during error concealment is disclosed.
Furthermore, a computer program is presented to implement the above-described method when run on a computer or signal processor.
IMPI
INSTITUTO MEXICANO »t THE PROPERTY. INDUSTRIAL
The embodiments perform a fading at the comfort noise level. According to these embodiments, a common tracking of the comfort noise level is performed in the drive domain. The target comfort noise level during burst packet loss is the same regardless of the core encoder (ACELP / TCX) in use, and is always updated. There is no known prior art, where common noise level tracking is necessary. The embodiments feature the fading of a codec switched to a comfort noise type signal during bursty packet losses.
Furthermore, the embodiments achieve that the overall complexity is lower compared to the presence of two independent noise level tracking modules, since both the functions (PROM) and the memory can be shared.
In embodiments, the level drift in the excitation domain (compared to the level drift in the time domain) produces more minima during active speech, since some of the voice information is covered by the coefficients of LP.
In the case of ACELP, according to these embodiments, the level shifting takes place in the domain of excitation. In the case of TCX, in the embodiments, the level is derived in the time domain, and the gain from LPC synthesis and de-emphasis is applied as a correction factor to model the energy level in the excitation domain. It would also be theoretically possible to trace the level in the excitation domain, e.g.
IMPIAS
INSTITUTO MEXICANO zl
OF THE PROPERTY Ά &
. rr ^. «·. ,. >. , INDUSTRIAL before FDNS, although the level of compensation between the TCX excitation domain and the ACELP excitation domain is considered to be quite complex.
There is no prior art that incorporates such common background level tracking across different domains. The prior art techniques do not include that common comfort noise level tracking, eg, in the drive domain, in a switched codec system. Accordingly, the embodiments are advantageous over the prior art in that, for prior art techniques, the comfort noise level that is targeted during burst packet losses may be different, depending on the mode of operation. preceding coding (ACELP / TCX) where the level would be tracked, since in the prior art, separate tracing for each coding mode causes unnecessary costs and additional computing complexity, and since in the prior art an updated comfort noise level could or could be available in any of the cores due to recent switching to this core.
According to some embodiments, level tracking is carried out in the excitation domain, but TCX fading is carried out in the time domain. By time domain fading, TDAC failures, which would cause overlap, are avoided. This is of particular interest when tonal components of the signal are hidden. Furthermore, the level of conversion between the excitation domain of ACELP and the spectral domain of MDCT is avoided and thus, eg, computing resources are saved. Due to the switching between the domain of the
<img file="MX347233B_D0024.tif" />
IMPI
INSTITUTO MEXICANO PE LA INDUSTRIAL PROPERTY excitation and the time domain, a level of adjustment is required between the excitation domain and the time domain. This is solved by deriving the gain that would be introduced by LPC synthesis and pre-emphasis and using this gain as a correction factor to convert the level between the two domains.
In contrast, the prior techniques do not perform level tracking in the excitation domain and TCX fading in the time domain. For state-of-the-art transform-based codes, the attenuation factor is applied in the excitation domain (by masking strategies in the time domain / ACELP type, see [3GP09a]), or either in the frequency domain (for strategies in the frequency domain such as frame repetition or noise substitution, see [LS01]). A disadvantage of the prior art strategy for applying the attenuation factor in the frequency domain is that overlap occurs in the region of overlap and sum in the time domain. This is so in the case of adjacent frames to which different attenuation factors are applied, since the fading procedure causes the failure of the TDAC (time domain overlap cancellation). This is especially relevant when tonal components of the signal are hidden. Consequently, the aforementioned embodiments are advantageous over the prior art.
The embodiments compensate for the influence of the high pass filter on the gain of LPC synthesis. According to these forms of realization, for t INSTITUTO MEXICANO DE LA FRORSDAD • INDUSTRIAL. INDUSTRIAL To compensate for the unintended gain change of the LPC Analysis and Emphasis caused by high-pass filtered unvoiced excitation, a correction factor is derived. This correction factor takes this unintended gain change into account and modifies the target comfort noise level in the excitation domain in such a way that the correct target level is obtained in the time domain.
In contrast, the prior art, eg G.718 [ITUOSa], introduces a high-pass filter into the signal path of the speechless drive, as illustrated in Fig. 2, if the signal from the last correct frame was not classified as NO VOICE. Thus, prior art strategies cause detrimental side effects, since the gain of subsequent LPC synthesis depends on the characteristics of the signal, which are altered by virtue of this high-pass filter. Since the background level in the excitation domain is tracked and applied, the algorithm is based on the gain of the LPC synthesis, which in turn depends once again on the characteristics of the excitation signal. In other words: The modification of the characteristics of the excitation signal due to the high pass filtering that was carried out in the prior art, could lead to a modified (usually reduced) gain of the LPC synthesis. This results in the wrong output level even though the drive level is correct.
The embodiments overcome these disadvantages of the prior art.
In particular, the embodiments obtain an adaptive spectral shape of the comfort noise. Unlike G.718, by tracking the spectral shape
INSTITUI · MEXICANO DE LA PROPERTY industrial background noise, and by applying (fading to) this shape during burst packet losses, the noise characteristic is equal to the previous background noise, thus leading to a noise characteristic nice comfort noise. This avoids annoying spectral shape discrepancies that could be introduced by the use of an offline learning derived spectral envelope and / or the spectral shape of the last received frame.
Furthermore, an apparatus for decoding an audio signal is disclosed. The apparatus comprises a reception interface, where the reception interface is configured to receive a first frame comprising a first portion of the audio signal of the audio signal, and where the reception interface is configured to receive a second frame that it comprises a second portion of the audio signal of the audio signal.
Furthermore, the apparatus comprises a noise level tracking unit, wherein the noise level tracking unit is configured to determine the noise level information depending on at least one of the first portion of the audio signal and the second portion of the audio signal (this means: depending on the first portion of the audio signal and / or the second portion of the audio signal), where the noise level information is represented in a tracking domain.
Furthermore, the apparatus comprises a first reconstruction unit to reconstruct, in a first reconstruction domain, a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received by the interface <sup>04</sup> IMPI
MEXICAN INSTITUTE ΛΤ
OF THE PROPERTY
INDUSTRIAL reception or if said third frame is received by the reception interface but is corrupted, where the first reconstruction domain is different or equal to the tracking domain.
Furthermore, the apparatus comprises a transformation unit for transforming the noise level information from the tracking domain to a second reconstruction domain, if a fourth frame of the plurality of frames is not received by the reception interface or if said fourth frame is received by the receiving interface but it is corrupted, where the second reconstruction domain is different from the tracking domain and where the second reconstruction domain is different from the first reconstruction domain, and
In addition, the apparatus comprises a second reconstruction unit to reconstruct, in the second reconstruction domain, a fourth portion of the audio signal from the audio signal depending on the noise level information that is represented in the second domain of reconstruction, if said fourth frame of the plurality of frames is not received by the reception interface or if said fourth frame is received by the reception interface but is corrupted.
According to some embodiments, the tracking domain can be, eg, where the tracking domain is a time domain, a spectral domain, an FFT domain, an MDCT domain, or an excitation domain. . The first reconstruction domain can be, eg, the time domain, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain. The second reconstruction domain can be, for example, the domain of »» · - Μ ι · Ι i I -Γ — Τ I »
INSTITUTO MEXICANO 31
FROM HOOP TO AGE
INDUSTRIAL.<sup>1</sup>¾ ^ time, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain.
In one embodiment, the tracking domain can be, eg, the FFT domain, the first reconstruction domain can be, eg, the time domain, and the second reconstruction domain can be, eg. , the domain of arousal.
In another embodiment, the tracking domain can be, e.g., the time domain, the first reconstruction domain can be, e.g., the time domain, and the second reconstruction domain can be, e.g. ., the domain of arousal.
According to one embodiment, said first portion of the audio signal may be represented, eg, in a first input domain and said second portion of the audio signal, may be represented, eg, in a second. input domain. The transformation unit can be eg a second transformation unit. The apparatus may further comprise, eg, a first transformation unit for transforming the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from the second input domain to the tracking domain to obtain information from the second portion of the signal. The noise level tracking unit can be configured, eg, to receive information from the first portion of the signal that is represented in the tracking domain, where the information from the first portion of the signal depends on the first portion of the audio signal, where the noise level tracking unit is set to 'IMPI ^ w
MEXICAN INSTITUTE
OF PROPERTY VV ^ 3iÍfJ £ & ___... ___________ ·, _ |. ~. ,,. industrial, 'SCLÍ-í? *. receive the second portion of the signal that is represented in the tracking domain, and where the noise level tracking unit is configured to determine the noise level information depending on the information of the first portion of the signal that is represented in the tracking domain and depending on the information of the second portion of the signal that is represented in the tracking domain.
According to one embodiment, the first input domain can be, eg, the excitation domain, and the second input domain can be, eg, the MDCT domain.
In another embodiment, the first input domain can be, eg, the MDCT domain, and where the second input domain can be, eg, the MDCT domain.
According to one embodiment, the first reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal by performing a first fading to a noise-like spectrum. The second reconstruction unit may be configured, eg, to reconstruct the fourth portion of the audio signal by executing a second fade to a noise-like spectrum and / or a second fade to a LTP gain. Furthermore, the first reconstruction unit and the second reconstruction unit may be configured, eg, to perform the first fading and the second fading to a noise-like spectrum and / or a second fading of a LTP gain. at the same fade rate.
«Wnyro MEXICAN
In one embodiment, the apparatus can compress ^ ngíi ^ a ^? ^^, a first aggregation unit for ΉΓπτιίη-1Γ lili μι mu volor ogragnrin depending on the first portion of the audio signal. Furthermore, the apparatus may also comprise eg a second aggregation unit for determining, depending on the second portion of the audio signal, a second aggregated value as a derived value of the second portion of the audio signal. The noise level tracking unit may be configured, e.g., to receive the first added value as information from the first portion of the signal that is represented in the tracking domain, where the noise level tracking unit can be configured, e.g., to receive the second added value as information from the second portion of the signal that is represented in the tracking domain, and where the noise level tracking unit is configured to determine the noise level information depending on the first added value that is represented in the tracking domain and depending on the second added value that is represented in the tracking domain.
According to one embodiment, the first aggregation unit may be configured, eg, to determine the first aggregated value in such a way that the first aggregated value indicates a root mean square of the first portion of the audio signal or of a signal derived from the first portion of the audio signal. The second aggregation unit is configured to determine the second aggregate value in such a way that the second aggregate value indicates a root mean square of the second portion of the audio signal or of a signal derived from the second portion of the audio signal.
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
In one embodiment, the first transform unit may be configured, eg, to transform the derived value of the second portion of the audio signal from the second input domain to the tracking domain by applying a gain value. to the derived value of the second portion of the audio signal.
According to one embodiment, the gain value may indicate, eg, a gain introduced by the synthesis of Linear Prediction Coding, or where the gain value indicates a gain introduced by the synthesis and de-emphasis of the Coding by Linear Prediction.
In one embodiment, the noise level tracking unit may be configured, eg, to determine noise level information by applying a minimal statistics strategy.
According to one embodiment, the noise level tracking unit may be configured, eg, to determine a comfort noise level as noise level information. The reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal depending on the noise level information, if said third frame of the plurality of frames is not received by the receiving interface or if said third frame is received by the receiving interface but it is corrupted.
In one embodiment, the noise level tracking unit may be configured, eg, to determine a comfort noise level as noise level information derived from a noise level spectrum, where said level spectrum noise is obtained by applying the
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL strategy of minimal statistics. The reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal depending on a plurality of linear prediction coefficients, if said third frame of the plurality of frames is not received by the receiving interface or if said third frame is received by the receiving interface but is corrupt.
According to one embodiment, the first reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal depending on the noise level information and depending on the first portion of the audio signal. , if said third frame of the plurality of frames is not received by the reception interface or if said third frame is received by the reception interface but is corrupted.
In one embodiment, the first reconstruction unit may be configured, eg, to reconstruct the third portion of the audio signal by attenuating or amplifying the first portion of the audio signal.
According to one embodiment, the second reconstruction unit may be configured, eg, to reconstruct the fourth portion of the audio signal depending on the noise level information and depending on the second portion of the audio signal.
In one embodiment, the second reconstruction unit may be configured, eg, to reconstruct the fourth portion of the audio signal by attenuating or amplifying the second portion of the audio signal.
IMPIAS
INSTITUTO MEXICANO 1Λ
OF THE PROPERTY
According to one embodiment, the apparatus further, eg, a long-term prediction unit that has a delay, where the long-term prediction unit may be configured, eg, to generate a processed signal depending on of the first or second portion of the audio signal, depending on a delay buffer input that is stored in the delay buffer and depending on a long-term prediction gain, and where the long-term prediction unit is configured to fade the long-term prediction gain towards zero, if said third frame of the plurality of frames is not received by the receiving interface or if said third frame is received by the interface reception but is corrupt.
In one embodiment, the long-term prediction unit may be configured, eg, to fade the long-term prediction gain toward zero, where the rate at which the long-term prediction gain fades to zero depends on a fading factor.
In one embodiment, the long-term prediction unit may be configured, eg, to update the delay buffer input by storing the generated processed signal in the delay buffer, if said third frame of the plurality of frames does not is received by the receiving interface or if said third frame is received by the receiving interface but is corrupted.
Furthermore, a method for decoding an audio signal is presented. The method comprises:
IMPI iNsrmrrc Mexican
OF THE PROPERTY
Γ, ... _ _ .____.___ __ _1 · industhi ^ l.
Receiving a first frame comprising a first portion of the audio signal of the audio signal, and receiving a second frame comprising a second portion of the audio signal of the audio signal.
Determine the noise level information depending on at least one of the first portion of the audio signal and the second portion of the audio signal, where the noise level information is represented in a tracking domain.
Reconstruct, in a first reconstruction domain, a third portion of the audio signal of the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received or if said third frame is received but it is corrupt, where the first rebuild domain is different from or equal to the crawl domain.
Transform the noise level information from the trace domain to a second reconstruction domain, if a fourth frame of the plurality of frames is not received or if said fourth frame is received but is corrupted, where the second reconstruction domain is different from the trace domain, and where the second reconstruction domain is different from the first reconstruction domain Y:
Reconstruct, in the second reconstruction domain, a fourth portion of the audio signal of the audio signal depending on the noise level information that is represented in the second reconstruction domain, if said fourth frame of the plurality of frames does not is received or if said fourth frame is received but is corrupted.
<img file="MX347233B_D0025.tif" />
IMPI
INSTITUTO MEXICANO DB LA PROPftoAD industrial
Furthermore, a computer program is disclosed to implement the above-described method when run on a computer or signal processor.
Furthermore, an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal is presented. The apparatus comprises a reception interface for receiving one or more frames, a coefficient generator, and a signal reconstructor. The coefficient generator is configured to determine, if a current frame of said one or more frames is received by the receiving interface and if the current frame being received by the receiving interface is not corrupted, one or more first coefficients of the audio signal, which are contained in the current frame, where said one or more first coefficients of the audio signal indicate a characteristic of the encoded audio signal, and one or more noise coefficients indicating a background noise of the encoded audio signal. Furthermore, the coefficient generator is configured to generate one or more second coefficients of the audio signal, depending on said one or more first coefficients of the audio signal and depending on said one or more noise coefficients, if the frame current frame is not received by the receiving interface or if the current frame being received by the receiving interface is corrupted. The audio signal reconstructor is configured to reconstruct a first portion of the reconstructed audio signal depending on said one or more first coefficients of the audio signal, if the current frame is received by the receiving interface and if the current frame that is being received by the receiving interface is not corrupted. Plus
ΙΜΡΙΟ ^
MEXICAN INSTITUTE
DE PROPtí DAD still, the audio signal reconstructor is configured '^ Yá ^ ecofisfroTr a second portion of the reconstructed audio signal depending on said one or more second coefficients of the audio signal, if the current frame is not received by the receiving interface or if the current frame being received by the receiving interface is corrupted.
In some embodiments, said one or more first coefficients of the audio signal may be, eg, one or more linear prediction filter coefficients of the encoded audio signal. In some embodiments, said one or more first coefficients of the audio signal may be, eg, one or more linear prediction filter coefficients of the encoded audio signal.
According to one embodiment, said one or more noise coefficients may be, eg, one or more linear prediction filter coefficients indicating the background noise of the encoded audio signal. In one embodiment, said one or more linear prediction filter coefficients may represent, eg, a spectral shape of the background noise.
In one embodiment, the coefficient generator may be configured, eg, to determine said one or more second portions of the audio signal such that said one or more second portions of the audio signal are one or more linear prediction filter coefficients of the reconstructed audio signal, or such that said one or more first coefficients of the audio signal are one or more immittance spectral pairs of the reconstructed audio signal.
IMPI ^^ 3
INSTITUTO MEXICANO JÁ
D £ LA PROPtEDAC »
According to one embodiment, the generator ^ coefficients can be configured, for example, to generate said one or more second coefficients of the audio signal by applying the formula:
/ currentf ^] - 'flast [*] 4 ~ (1 <sup>—</sup> d) · ptmean [¿] where /<sub>ΜηβΜ</sub>[ΐ] indicates one of said one or more second coefficients of the audio signal, where fi<sub>ast</sub>[í] indicates one of said one or more first coefficients of the audio signal, where pt<sub>mean</sub>[i \ is one of said one or more noise coefficients, where a is a real number where 0 sas 1, and where i is an index. In one embodiment, 0 <a <1.
According to one embodiment, fias<sub>t</sub>[í \ indicates a linear prediction filter coefficient of the encoded audio signal, and where fcurr<sub>in</sub>t [í \ indicates a linear prediction filter coefficient of the reconstructed audio signal.
In one embodiment, pt<sub>mean</sub>[í \ can indicate eg the background noise of the encoded audio signal.
In one embodiment, the coefficient generator may be configured, eg, to determine whether the current frame of said one or more frames is received by the receiving interface and whether the current frame being received by the interface reception is not corrupted, said one or more noise coefficients by determining a noise spectrum of the encoded audio signal.
MEXICAN INSTITUTE
DS INDUSTRIAL PROPERTY
According to one embodiment, the coefficient generator may be configured, eg, to determine LPC coefficients representing background noise using a strategy of minimal statistics on the signal spectrum to determine a spectrum of the background noise. and calculating the LPC coefficients representing the shape of the background noise from the spectrum of the background noise.
A method for decoding an encoded audio signal to obtain a reconstructed audio signal is further disclosed. The method comprises:
- Receive one or more frames.
- Determine, if a current frame of said one or more frames is received and if the current frame being received is not corrupted, one or more first coefficients of the audio signal, which are contained in the current frame, where said one One or more first coefficients of the audio signal indicate a characteristic of the encoded audio signal, and one or more noise coefficients that indicate a background noise of the encoded audio signal.
- Generate one or more second coefficients of the audio signal, depending on said one or more first coefficients of the audio signal and depending on said one or more noise coefficients, if the current frame is not received or if the current frame is being received is corrupt.
- Reconstruct a first portion of the reconstructed audio signal depending on said one or more first coefficients of the audio signal.
MEXICAN INSTITUTE Ά
DELA PROner> AI>
audio, if the current frame is received and if the acfuaPquesS ^^ frame is receiving it is not corrupted and: <sup>:</sup> ”
- Reconstructing a second portion of the reconstructed audio signal depending on said one or more second coefficients of the audio signal, if the current frame is not received or if the current frame that is being received is corrupted.
Furthermore, a computer program is disclosed to implement the above-described method when run on a computer or signal processor.
Having common means of tracking and applying the comfort noise spectral shape during fading offers several advantages. By tracking and applying the spectral shape in such a way that it can be done in a similar way in both core codes, a single common strategy results. CELT only reports banding of energies in the spectral domain and banding of the spectral shape in the spectral domain, which is not possible for the CELP core.
In contrast, in the prior art, the spectral shape of comfort noise introduced during burst losses is fully static or partially static and partially adaptive to the short-term mean of the spectral shape (as implemented in G.718 [ ITU08a]), and generally does not match the background noise of the signal before packet loss. These discrepancies in comfort noise characteristics could be annoying. According to the prior art, a form of background noise may be employed
ΙΜΡΙ ^^
INSTITUTO Mexicano DE LA PROPIEDAD INDUSTRIAL obtained offline (static) that may sound pleasant in the case of certain signals, but less pleasant in others, e.g. car noise sounds are completely different from office noise.
Furthermore, in the prior art, a short-term averaging of the spectral shape of previously received frames may be employed that could bring the characteristics of the signal closer to those of the previously received signal, although not necessarily characteristics of background noise. In the prior art, band tracking of spectral shape in the spectral domain (as performed in CELT [IET12]) is not applicable to a switched codec that uses not only a domain-based core of MDCT (TCX) but also a kernel based on ACELP. Accordingly, the above-discussed embodiments are advantageous over the prior art.
Furthermore, an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal is presented. The apparatus comprises a reception interface for receiving one or more frames comprising information on a plurality of samples of audio signals from a spectrum of the audio signal from the encoded audio signal, and a processor for generating the audio signal reconstructed. The processor is configured to generate the reconstructed audio signal by fading a modified spectrum to a target spectrum, if a current frame is not received by the receiving interface or if the current frame is received by the receiving interface but is corrupted. , where the modified spectrum comprises a plurality of modified signal samples, where, for each of the modified signal samples of the modified spectrum, an absolute value of said sample of the modified signal is equal to an absolute value of one of the samples of the audio signal from the spectrum of the audio signal. Furthermore, the processor is configured not to fade the modified spectrum to the white noise spectrum, if the current frame of said one or more frames are received by the receiving interface and if the current frame being received by the receiving interface it is not corrupt.
According to one embodiment, the target spectrum may be, eg, a noise-type spectrum.
In one embodiment, the noise spectrum can represent, eg, white noise.
According to one embodiment, the noise spectrum can be, eg, patterned.
In one embodiment, the shape of the noise spectrum may depend, eg, on a spectrum of the audio signal of a previously received signal.
According to one embodiment, the noise spectrum can be modeled, eg depending on the shape of the spectrum of the audio signal.
In one embodiment, the processor may employ, eg, a tilt factor to model the noise spectrum.
According to one embodiment, the processor may employ eg the formula shaped_noise [i] = noise * power (tilt_factor, i / N)
IMPI ^ ►MEXICANISTITUTE
DS INDUSTRIAL PROPERTY where N indicates the number of samples, where i is an index, where o <= i <n, with tilt_factor> 0, and where power is a power function.
power (x, y) indicates x<sup>AND</sup> power (tilt_factor, i / N) indicates tilt_factor<sup>N</sup>
If the tilt_factor is less than 1 this means attenuation with increasing i. If the tilt_f actor is greater than 1, this means amplification with increasing i.
According to another embodiment, the processor can use, for example, the formula shaped_noise [i] = noise * (1 + i / (Nl) * (tilt_factor-l)) where N indicates the number of samples, where i is an index, where 0 <= i <N, with tilt_factor> 0.
If the tilt factor is less than 1 this means attenuation with increasing i. If the tilt_factor is greater than 1, this means amplification with increasing i.
According to one embodiment, the processor may be configured, eg, to generate the modified spectrum, by changing the sign of one or more of the samples of the audio signal from the spectrum of the audio signal, if the current frame is not received by the receiving interface or if the current frame being received by the receiving interface is corrupted.
In one embodiment, each of the audio signal samples from the audio signal spectrum may be represented, eg, by a real number and not an imaginary number.
MEXICAN WSTITVrO OF INDUSTRIAL PROPERTY
According to one embodiment, the audio signal samples from the audio signal spectrum may be represented, eg, in a Domain of the Modified Discrete Cosine Transform.
In another embodiment, the audio signal samples from the audio signal spectrum may be represented, eg, in a Modified Discrete Sine Transform Domain.
According to one embodiment, the processor may be configured, eg, to generate the modified spectrum by employing a random sign function that randomly or pseudo-randomly outputs a first or a second value.
In one embodiment, the processor may be configured, eg, to fade the modified spectrum toward the target spectrum by subsequently reducing an attenuation factor.
In accordance with one embodiment, the processor may be configured, eg, to fade the modified spectrum towards the target spectrum by subsequently increasing an attenuation factor.
In one embodiment, if the current frame is not received by the receive interface or if the current frame that is being received by the receive interface is corrupted, the processor may be configured, e.g., to generate the signal signal. reconstructed audio using the formula:
x [i] = (l-cum_damping) * noise [i] + cum_damping * random_sign () * x_old [i]
<img file="MX347233B_D0026.tif" />
IMPI Mexican institute DE LA PROPIEDAD INDUSTRIAL where i is an index, where x [i] indicates a sample of the reconstructed audio signal mi ιι ^ ρκτΜ ea, where cum_damping is an attenuation factor, where x_old [i] indicates one of the samples of the audio signal from the audio signal spectrum of the encoded audio signal, where random_sign () returns 1 or -1, and where noise is a random vector indicating the target spectrum.
In one embodiment, said random noise vector can be scaled, eg, in such a way that its root mean square is similar to the mean square of the spectrum of the encoded audio signal that is comprised of one of the last received frames. term by the receiving interface.
According to a general embodiment, the processor may be configured, eg, to generate the reconstructed audio signal, by employing a random vector that is scaled such that its root mean square is similar to the mean square. of the spectrum of the encoded audio signal that is comprised of one of the frames ultimately received by the receiving interface.
Furthermore, a method of decoding an encoded audio signal to obtain a reconstructed audio signal is disclosed. The method comprises:
- Receive one or more frames that comprise information about a plurality of samples of audio signals from a spectrum of the audio signal of the encoded audio signal and:
- Generate the reconstructed audio signal.
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
The generation of the reconstructed audio signal is carried out by fading a modified spectrum to a target spectrum, if a current frame is not received or if the current frame is received but is corrupted, where the modified spectrum comprises a plurality of modified signal samples, where, for each of the modified signal samples of the modified spectrum, an absolute value of said sample of the modified signal is equal to an absolute value of one of the samples of the audio signal from the spectrum of the audio signal. The modified spectrum does not fade into a white noise spectrum, if the current frame of said one or more frames is received and if the current frame being received is not corrupted.
Furthermore, a computer program is disclosed to implement the above-described method when run on a computer or signal processor.
The embodiments obtain an MDCT spectrum faded to white noise prior to FDNS Application (FDNS = Frequency Domain Noise Substitution. Frequency Domain Noise Substitution)).
According to the prior art, in ACELP-based codecs, the innovative codebook is replaced with a random vector (eg, with noise). In embodiments, the ACELP technique, which consists of replacing the innovative codebook with a random vector (eg, by noise), is adapted to the structure of the TCX decoder. In this case, the innovative codebook equivalent is the MDCT spectrum usually received within the bit stream and fed to the FDNS.
IMPI Mexican mstitutc DE LA PROPERTY INDUSTRIAL
The classical MDCT masking strategy would simply repeat the spectrum as is or apply a certain scrambling process, which basically lengthens the spectral shape of the last received frame [LS01]. This has the disadvantage that the spectral shape is prolonged in the short term, often giving rise to a repetitive metallic sound that does not resemble background noise and therefore cannot be used as comfort noise.
Using the proposed method, short-term spectral modeling is performed by FDNS and TCX LTP, long-term spectral modeling is performed only by FDNS. Modeling by the FDNS fades from a short-term spectral shape to the long-term tracked spectral shape of the background noise, and the TCX LTP fades to zero. The fading of the FDNS coefficients to the tracked background noise coefficients leads to a smooth transition between the last correct spectral envelope and the background spectral envelope that would be the goal in the long run, to get a nice background noise in case frame loss on long bursts.
On the contrary, according to the current state of the art, in the case of transform-based codes, noise concealment is carried out by means of the repetition of frames or substitution of noise in the frequency domain [LS01] . In the prior art, noise substitution is usually carried out by sign transposition in the spectral boxes. If in the prior art TCX (frequency domain) sign transposition is used during concealment, the last received MDCT coefficients are reused and each sign is randomized before inverse transforming the MEXICAN RSTTTVTO
FROM THE PROPERTY ^^ ¿33 ^ 0 ^ INDUSTRIAL spectrum to the time domain. The drawback of this procedure is that, for consecutive lost frames, the same spectrum is used over and over again, only with different sign scrambling and global attenuation. When the spectral envelope over time is considered in a coarse time grid, it can be seen that the envelope is approximately constant during the loss of consecutive frames, since the band energies remain constant with each other within a frame and only attenuate globally. In the coding system employed according to the prior art, spectral values are processed using FDNS to restore the original spectrum. This means that if you want to fade the MDCT spectrum to a certain spectral envelope (using FDNS coefficients, e.g. describing the current background noise), the result does not depend only on the FDNS coefficients, but also depends of the spectrum decoded with transposed sign. The aforementioned embodiments overcome these disadvantages of the prior art.
The embodiments are based on the finding that it is necessary to fade the spectrum used for transposition from sign to white noise, before feeding it to FDNS processing. Otherwise the emitted spectrum never matches the target envelope used for FDNS processing.
In embodiments, the same feed rate is used for LTP gain fading as for white noise fading.
IMPI ^ ag ^ wsrmiTO Mexican
OF THE PROPERTY
In addition, an apparatus is presented for decoding a encoded audio signal to obtain a reconstructed audio signal. The comp7erd * apparatus has a reception interface for receiving a plurality of frames, a delay buffer for storing samples of audio signals of the decoded audio signal, a sample selector for selecting a plurality of samples of the selected audio signal of the audio signal samples that are stored in the delay buffer, and a sample processor for processing the selected audio signal samples to obtain samples of the reconstructed audio signal from the reconstructed audio signal. The sample selector is configured to select, if a current frame is received by the reception interface and if the current frame being received by the reception interface is not corrupted, the plurality of samples of the audio signal selected from the Audio signal samples that are stored in the delay buffer depending on a pitch delay information that are contained in the current frame. Furthermore, the sample selector is configured to select, if the current frame is not received by the receiving interface or if the current frame being received by the receiving interface is corrupted, the plurality of samples of the audio signal. selected from the audio signal samples that are stored in the delay buffer depending on a tone delay information that is contained in another frame previously received by the receiving interface.
υατπυτο MEXICAN DE LA PRQP! L »AD
According to one embodiment, the '^ rsample processor may be configured, e.g. to sample the reconstructed audio signal, if the current frame is received by the receiving interface and if the current frame being received by the receiving interface is not corrupted, by rescaling the selected audio signal samples depending on the gain information that are contained in the current frame. Furthermore, the sample selector may be configured, e.g., to obtain samples from the reconstructed audio signal, if the current frame is not received by the receiving interface or if the current frame is being received by the interface. reception is corrupted, by rescaling the selected audio signal samples depending on the gain information that is included in said additional frame previously received by the reception interface.
In one embodiment, the sample processor may be configured, e.g., to sample the reconstructed audio signal, if the current frame is received by the receiving interface and if the current frame is being received by the receiving interface is not corrupted, multiplying the selected audio signal samples and a value that depends on the gain information that are contained in the current frame. In addition, the sample selector is configured to obtain the samples of the reconstructed audio signal, if the current frame is not received by the reception interface or if the current frame that is being received by the reception interface is corrupted, multiplying the selected audio signal samples and a value
<img file="MX347233B_D0027.tif" />
IMPI «Γπτντο Mexican
DE LA MOHEDA or «DWTIUAL which depends on the gain information that is included in said additional frame previously received by the reception interface.
According to one embodiment, the sample processor may be configured, eg, to store the samples of the reconstructed audio signal in the delay buffer.
In one embodiment, the sample processor may be configured, eg, to store samples of the reconstructed audio signal in the delay buffer prior to receiving another frame by the receiving interface.
According to one embodiment, the sample processor may be configured, eg, to store samples of the reconstructed audio signal in the delay buffer upon receipt of another frame by the receive interface.
In one embodiment, the sample processor may be configured, eg, to rescale the selected audio signal samples depending on the gain information to obtain rescaled audio signal samples and combining the signal samples. samples rescaled with samples of the input audio signal to obtain the samples of the processed audio signal.
According to one embodiment, the sample processor may be configured, eg, to store the processed audio signal samples, indicating the combination of the rescaled audio signal samples and the audio signal samples. input audio, in the delay buffer,
IMPI ^ tNSTTTUTO MUICANO
Say THE «OniDAÍi
INDUSTRIAL and not to store the rescaled audio signal samples in the delay buffer, if the current frame is received by the receiving interface and if the current frame that is being received by the receiving interface is not corrupted. Furthermore, the sample processor is configured to store the rescaled audio signal samples in the delay buffer and not to store the processed audio signal samples in the delay buffer, if the current frame is not received by the receiving interface or if the current frame being received by the receiving interface is corrupted.
According to another embodiment, the sample processor may be configured, e.g., to store the samples of the processed audio signal in the delay buffer, if the current frame is not received by the receiving interface or if the current frame being received by the receiving interface is corrupted.
In one embodiment, the sample selector may be configured, eg, to obtain the samples of the reconstructed audio signal by rescaling the selected audio signal samples depending on a modified gain, where the modified gain is defined according to the formula:
gain = gain_past * damping;
Mexican ηβπτυτο
OF THE NTHEDAD + ΐ *
INDUSTRIAL where gain is the modified gain, where the sample selector w — W — m — M'iniWMrn'L r · can be set, eg, to set gain past to gain after gain y is calculated, and where damping is a real value.
According to one embodiment, the sample selector may be configured, eg, to calculate the modified gain.
In one embodiment, damping can be defined, eg, according to: 0 damping ú 1.
According to one embodiment, the modified gain gain may be, eg, set to zero, if at least a predetermined number of frames has not been received by the receiving interface since the last frame that was received by the reception interface.
Furthermore, a method for decoding an encoded audio signal to obtain a reconstructed audio signal is presented. The method comprises:
- Receive a plurality of frames.
- Store samples of audio signals from the decoded audio signal.
- Select a plurality of samples of the audio signal selected from the audio signal samples that are stored in the delay buffer and:
<img file="MX347233B_D0028.tif" />
IMPI
MEXICAN INSTITUTE
OF THE INDUSTRIAL TOOP1CDAO
Process selected audio signal samples
Μ · * φΜ · ΜΜΜ · η · ν>
to obtain samples of the reconstructed audio signal from the reconstructed audio signal.
If a current frame is received and if the current frame being received is not corrupted, the step of selecting the plurality of samples of the audio signal selected from the samples of the audio signal that are stored in the delay buffer depending on a pitch delay information that is contained in the current frame. Furthermore, if the current frame is not received or if the current frame being received is corrupted, the step of selecting the plurality of samples of the audio signal selected from the audio signal samples that are buffered Delay is carried out depending on a tone delay information that is contained in another frame previously received by the receiving interface.
Also disclosed is a computer program for implementing the above-described method when run on a computer or signal processor.
The embodiments employ TCX LTP (TXC LTP = Long Term Prediction Transform Coded Excitation). During normal operation, the TCX LTP memory is updated with the synthesized signal, which contains noise and reconstructed tonal components.
Instead of disabling the TCX LTP during cloaking, its normal operation can be continued during cloaking with the parameters received in the last good frame. This preserves the spectral shape of the
INSTITUTO MEXICANO DE LA PROPIEDAD signal, especially the tonal components that are modeled by ehffiro de
LTP.
Furthermore, the embodiments decouple the TCX LTP feedback loop. A simple continuation of the normal TCX LTP operation introduces additional noise, since with each update step more noise generated randomly from the LTP drive is introduced. Thus, the tonal components become increasingly distorted over time by the added noise.
To overcome this, the unique updated TCX LTP buffer can be re-fed (without adding noise) so as not to contaminate the tonal information with unwanted random noise.
In addition, according to these embodiments, the TCX LTP gain fades to zero.
These embodiments are based on the finding that continuation of TCX LTP contributes to preserving signal characteristics in the short term, although it has long-term drawbacks. The signal reproduced during concealment typically includes the voice / tonal information that was present before the loss. Especially in the case of the clean voice or voice over background noise, it is very unlikely that the pitch or harmonic will decay very slowly over a long period of time. Continuing the TCX LTP operation during masking, especially if the LTP memory update is decoupled (only the tonal components are fed back and not the transposed signed part), the voice / tonal information is still present in the hidden signal throughout the loss, which is attenuated only by the fading
IMPI
MEXICAN INSTITUTE
FROM INDUSTRIAL PROPERTY, “'TOTAL MUSIRIAL TO COMFORT NOISE LEVEL. Furthermore, it is impossible to achieve the comfort noise envelope during burst packet losses if TCX LTP is applied during burst losses without attenuating in time, since then the signal always incorporates the voice information of the LTP. .
Therefore, the TCX LTP gain fades to zero, such that the tonal components represented by the LTP fade to zero, at the same time that the signal fades to the level and shape of the background signal, and such that the fading reaches the spectral background envelope (comfort noise) without incorporating unwanted tonal components.
In embodiments, the same fading rate is used to fade the LTP gain as for the white noise fading.
In contrast, in the prior art, no transform coding is known to use LTP during concealment. In the case of MPEG-4 LTP [ISO09] there are no masking strategies in the prior art. Another prior art MDCT-based codec that makes use of an LTP is CELT, although this codec uses ACELP-like concealment in the first five frames, and background noise is generated for all subsequent frames, which does not make use of the LTP. One disadvantage of the prior art of not using TCX LTP is that all tonal components that are modeled with LTP abruptly disappear. Furthermore, in prior art ACELP-based codes, the LTP operation is prolonged during concealment and the adaptive codebook gain fades towards zero. Regarding the operation of the feedback loop, the prior art employs two
<img file="MX347233B_D0029.tif" />
strategies, total arousal is fed back, eg, the sum of innovative and adaptive arousal (AMR-WB); or only the updated adaptive drive, eg, the tonal parts of the signal (G.718), is fed back. The above-cited embodiments overcome the disadvantages of the prior art.
Hereinafter, embodiments of the present invention are described in more detail with reference to the figures, in which:
Fig. 1a illustrates an apparatus for decoding an audio signal according to one embodiment,
Fig. 1b illustrates an apparatus for decoding an audio signal according to another embodiment,
Fig. 1c illustrates an apparatus for decoding an audio signal according to another embodiment, wherein the apparatus also comprises a first and a second aggregation unit,
Fig. 1d illustrates an apparatus for decoding an audio signal according to another embodiment, wherein the apparatus further comprises a long-term prediction unit comprising a delay buffer,
Fig. 2 illustrates the structure of the G.718 decoder,
Fig. 3 illustrates a situation where the G.722 fading factor depends on class information,
Fig. 4 shows a strategy for the prediction of the amplitude that uses linear regression, <sup>94</sup> IMPI ^^
MEXICAN INSTITUTE *
FROM THE NDMDAn V—
INDUSTRIAL ~ *
Fig. 5 illustrates the burst loss behavior of the Energy Constrained Overlap Transform (CELT),
Fig. 6 illustrates a background noise level tracking according to an embodiment in the decoder during an error-free mode of operation,
Fig. 7 illustrates the derivation of the LPC Synthesis gain and de-emphasis according to one embodiment,
Fig. 8 illustrates the application of the comfort noise level during packet loss according to one embodiment,
Fig. 9 illustrates advanced high pass gain compensation during concealment by ACELP in accordance with one embodiment,
Fig. 10 illustrates LTP feedback loop decoupling during concealment in accordance with one embodiment,
Fig. 11 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in accordance with one embodiment,
Fig. 12 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in accordance with another embodiment and
Fig. 13 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in another embodiment and
Mexican IMPI
Fig. 14 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in another embodiment.
Fig. 1a illustrates an apparatus for decoding an audio signal in accordance with one embodiment.
The apparatus comprises a reception interface 110. The reception interface is configured to receive a plurality of frames, where the reception interface 110 is configured to receive a first frame of the plurality of frames, where said first frame comprises a first portion of the audio signal of the audio signal, wherein said first portion of the audio signal is represented in a first domain. Furthermore, the reception interface 110 is configured to receive a second frame of the plurality of frames, wherein said second frame comprises a second portion of the audio signal of the audio signal.
Furthermore, the apparatus comprises a transformation unit 120 for transforming the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from a second domain to a tracking domain to obtain a tracking information. the second portion of the signal, where the second domain is different from the first domain, where the tracking domain is different from the second domain, and where the tracking domain is the same or different from the first domain.
In addition, the apparatus comprises a noise level tracking unit 130, where the noise level tracking unit is configured to receive information from the first portion of the signal that is represented in
<img file="MX347233B_D0030.tif" />
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD industrial the tracking domain, where the information of the first portion of the signal depends on the first portion of the audio signal, where the noise level tracking unit is configured to receive the second portion of the signal that is represented in the trace domain, and wherein the noise level tracking unit is configured to determine the noise level information depending on the information of the first portion of the signal that is represented in the tracking domain and depending on the information of the second portion of the signal that is represented in the trace domain.
Furthermore, the apparatus comprises a reconstruction unit for reconstructing a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received by the interface of reception but it is corrupt.
As regards the first and / or the second portion of the audio signal, for example, the first and / or the second portion of the audio signal may, for example, be fed to one or more processing units (not illustrated) to generate one or more speaker signals for one or more speakers, so that the received sound information comprised by the first and / or the second portion of the audio signal can be reproduced.
Furthermore, however, the first and second portions of the audio signal are also used for concealment, eg, in case subsequent frames do not reach the receiver or in case subsequent frames are in error.
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
The present invention is based, among other things, on the finding that noise level tracking must be carried out in a common domain referred to herein as "tracking domain". The tracking domain can be, for example, an excitation domain, for example the domain in which the signal is represented by LPCs (LPC = Linear Predictive Coefficient) or by ISPs (ISP = Immittance Spectral Pair) as described in AMR-WB and AMR-WB + (see [3GP12a], [3GP12b], [3GP09a], [3GP09b], [3GP09c]). Tracking the noise level in a single domain has the advantage, among others, that overlapping effects are avoided when the signal switches between a first representation in a first domain and a second representation in a second domain (for example, when the representation of the signal switches from ACELP to TCX or vice versa).
As far as transformation unit 120 is concerned, it is the second portion of the audio signal itself that is transformed, or a signal derived from the second portion of the audio signal (e.g., the second portion of the audio signal has been processed to obtain the derived signal), or a value derived from the second portion of the audio signal (eg, the second portion of the audio signal has been processed to obtain the derived value).
Regarding the first portion of the audio signal, in some embodiments, the first portion of the audio signal can be processed and / or transformed to the tracking domain.
In other embodiments, however, the first portion of the audio signal may already be represented in the tracking domain.
IMPI ^ a
Χ5ΤΓΠ ΓΓΟ Μ RXICA NO
GIVE THE INDUSTRIAL PROPERTY AND
In some embodiments, the information in the first portion of the signal is identical to the first portion of the audio signal. In other embodiments, the information in the first portion of the signal is, eg, an aggregate value that depends on the first portion of the audio signal.
Now we consider in more detail, first, fading at a comfort noise level.
The fading strategy described can be implemented, for example, in a low-delay version of xHE-AAC [NMR + 12] (xHE-AAC = Extended High Efficiency AAC), which can seamlessly switch between ACELP (voice) and MDCT (music / noise) encoding frame by frame.
When it comes to tracking a common level in a tracking domain, for example an excitation domain, in order to apply a soft fading to an appropriate comfort noise level during packet loss, it is necessary to identify that level of comfort noise during normal decoding process. It can be assumed, for example, that a noise level similar to background noise is very pleasant. Consequently, the background noise level can be derived and constantly updated during normal decoding.
The present invention is based on the finding that when a switched core codec (eg, ACELP and TCX) is available, a common background noise level independent of the chosen core encoder is considered to be particularly suitable.
• MEXICAN INSTRUMENT OF INDUSTRIAL PROPERTY
Fig. 6 illustrates a background noise level tracking according to a preferred embodiment at the decoder during error-free mode of operation, eg, during normal decoding.
Tracing itself can be done, eg, using the minimal statistics strategy (see [Mar01]).
This tracked background noise level can be considered eg as the above-mentioned noise level information.
For example, the minimum statistical noise estimation presented in the document: ““ Rainer Martin, Noise power spectral density estimation based on optimal smoothing and minimum statistics, IEEE Transactions on Speech and Audio Processing 9 (2001), no. 5, 504-512 ”[Mar01] for background noise level tracking.
Correspondingly, in some embodiments, the noise level tracking unit 130 is configured to determine the noise level information by applying a minimal statistics strategy, e.g., by using the estimation of noise by minimum statistics of [Mar01].
Below are some considerations and details of this tracking strategy.
As far as level tracking is concerned, the background is assumed to be noise. Therefore, it is preferable to run level tracking in the excitation domain to avoid tracking foreground tonal components extracted by the LPC. For example, ACELP's Noise Fill can also use the level of 'r lili LII MI jltKtT100
IMPI iNSTmn-o Mexican DI LA PROPIEDAD
<img file="MX347233B_D0031.tif" />
...... ......, INDUSTRIAL.
background noise in the excitation domain. With excitation domain tracking, a single background noise level scan can serve two purposes, avoiding computational complexity. In a preferred embodiment, the screening is performed in the excitation domain of ACELP.
Fig. 7 illustrates the derivation of the Synthesis gain and de-emphasis of
LPC according to one embodiment.
As regards the level derivation, the level derivation can be carried out, for example, in the time domain or in the excitation domain or in any other suitable domain. If the domains for level derivation and level tracking differ, eg gain compensation may be required.
In the preferred embodiment, the derivation of the level corresponding to ACELP is done in the domain of excitation. Hence, profit compensation is not necessary.
In the case of TCX, eg gain compensation may be necessary to adjust the derived level to the ACELP drive domain.
In the preferred embodiment, the tapping of the level in TCX occurs in the time domain. A controllable gain tradeoff has been found for this strategy. The gain introduced by the LPC synthesis and de-emphasis is derived as illustrated in Fig. 7 and the derived level divided by this gain.
On the other hand, the derivation of the level corresponding to TCX could be performed in the excitation domain of TCX. However, it was considered that
INSTITUTO MEXICANO DE LA PROPERTY INDUSTRIAL Gain trade-off between the excitation domain of TCX and the excitation domain of ACELP was too complicated.
Therefore, returning to Fig. 1a, in some embodiments, the first portion of the audio signal is represented in a time domain as the first domain. The transformation unit 120 is configured to transform the second portion of the audio signal or the value derived from the second portion of the audio signal from an excitation domain that is the second domain to the time domain that is the domain of tracking. In those embodiments, the noise level tracking unit 130 is configured to receive information from the first portion of the signal that is represented in the time domain as the tracking domain. Furthermore, the noise level tracking unit 130 is configured to receive the second portion of the signal that is represented in the time domain as the tracking domain.
In other embodiments, the first portion of the audio signal is represented in the drive domain as the first domain. The transformation unit 120 is configured to transform the second portion of the audio signal or the derived value of the second portion of the audio signal from a time domain that is the second domain to the excitation domain that is the domain of tracking. In those embodiments, the noise level tracking unit 130 is configured to receive information from the first portion of the signal that is represented in the drive domain as the tracking domain. Furthermore, the noise level tracking unit 130 is
<img file="MX347233B_D0032.tif" />
configured to receive the second portion of the signal that is represented in
Μ · Μ »« * υ · ΙΧ14 the excitation domain as a tracking domain.
In one embodiment, the first portion of the audio signal may be represented, eg, in the drive domain as the first domain, where the noise level tracking unit 130 may be configured, eg, to receive the information of the first portion of the signal, where said information of the first portion of the signal is represented in the FFT domain, which is the tracking domain, and where said information of the first portion of the signal depends on said first portion of the audio signal that is represented in the drive domain, where the transformation unit 120 may be configured, eg, to transform the second portion from the audio signal or the value derived from the second portion of the audio signal from a time domain that is the second domain to an FFT domain that is the tracking domain, and where the noise level tracking unit 130 may be configured, eg, to receive the second portion of the audio signal that is represented in the FFT domain.
Fig. 1b illustrates an apparatus according to another embodiment. In Fig. 1b, the transformation unit 120 of Fig. 1a is a first transformation unit 120, and the reconstruction unit 140 of Fig. 1a is a first reconstruction unit 140. The apparatus also comprises a second unit transformation 121 and a second reconstruction unit 141.
103
IMPI
MEXICAN INSTITUTE
The second transformation unit 121 is ^ wrr ^ j
<img file="MX347233B_D0033.tif" />
transforming the noise level information of the scan domlHTO to the second, domain, if a fourth frame of the plurality of frames is not received by the reception interface or if said fourth frame is received by the reception interface but is corrupted.
Furthermore, the second reconstruction unit 141 is configured to reconstruct a fourth portion of the audio signal from the audio signal depending on the noise level information that is represented in the second domain if said fourth frame of the plurality of frames it is not received by the receiving interface or if said fourth frame is received by the receiving interface but is corrupted.
Fig. 1c illustrates an apparatus for decoding an audio signal according to another embodiment. The apparatus also comprises a first aggregation unit 150 for determining a first aggregated value depending on the first portion of the audio signal. Furthermore, the apparatus of Fig. 1c also comprises a second aggregation unit 160 for determining a second aggregated value as a derived value of the second portion of the audio signal depending on the second portion of the audio signal. In the embodiment of Fig. 1c, the noise level tracking unit 130 is configured to receive first added value as information from the first portion of the signal that is represented in the tracking domain, where the noise level tracking unit 130 is configured to receive the second value added as information of the second portion of the signal that is
<img file="MX347233B_D0034.tif" />
<sup>104</sup> IMPI
MEXICAN INSTITUTE
OF THE INDUSTRIAL PROPERTY represented in the tracking domain. The noise level tracking unit 130 is configured to determine the noise level information depending on the first added value that is represented in the tracking domain and depending on the second added value that is represented in the tracking domain.
In one embodiment, the first aggregation unit 150 is configured to determine the first aggregated value such that the first aggregated value indicates a root mean square of the first portion of the audio signal or of a signal derived from the first portion of the audio signal. Furthermore, the second aggregation unit 160 is configured to determine the second aggregate value such that the second aggregate value indicates a root mean square of the second portion of the audio signal or of a signal derived from the second portion of the audio signal. audio signal.
Fig. 6 illustrates an apparatus for decoding an audio signal according to another embodiment.
In Fig. 6, the background level tracking unit 630 implements a noise level tracking unit 130 in accordance with Fig. 1a.
Furthermore, in Fig. 6, the RMS unit 650 (RMS = root mean square) is a first aggregation unit and the RMS unit 660 is a second aggregation unit.
According to some embodiments, the (first) transform unit 120 of Fig. 1a, Fig. 1b and Fig. 1c is configured to transform the derived value of the second portion of the audio signal from the —— III lili
105
INSTTH / TO MEXICAN second domain to the tracking domain using the apífcaw <^^^ gain (x) to the value derived from the second portion of your audio signal Id, t> ef ^ ·, dividing the value derived from the second portion of the audio signal by a gain value (x). In other embodiments, eg a gain value can be multiplied.
In some embodiments, the gain value (x) may indicate, eg, a gain entered by Linear Prediction Coding synthesis or the gain value (x) may indicate, eg, a gain entered by the synthesis and de-emphasis of Linear Prediction Coding.
In Fig. 6, unit 622 gives the value (x) that indicates the gain introduced by the synthesis and offset of Linear Prediction Coding. Unit 622 then divides the value provided by the second aggregation unit 660, which is a value derived from the second portion of the audio signal, by the provided gain value (x) (i.e., either dividing by x, or multiplying the value 1 / x). Consequently, unit 620 of Fig. 6 comprising units 621 and 622 implements the first transformation unit of Fig. 1a, Fig. 1b or Fig. 1c.
The apparatus of Fig. 6 receives a first frame with a first portion of the audio signal which is a voice drive and / or a non-voice drive and which is represented in the tracking domain, in Fig. 6 a domain of LPC (ACELP). The first portion of the audio signal is fed to an LPC 671 synthesis and de-emphasis unit for processing to output the first portion of the audio signal in the time domain. Plus
IMPI
Mexican XSTmrro
Still from INDUSTRIAL PROPERTY, the first portion of the audio signal is fed to the RMS 650 module to obtain a first value that indicates a root mean square of the first portion of the audio signal. This first value (first RMS value) is represented in the trace domain. The first RMS value, which is represented in the tracking domain, is then fed to the noise level tracking unit 630.
Furthermore, the apparatus of Fig. 6 receives a second frame with a second portion of the audio signal comprising an MDCT spectrum and which is represented in an MDCT domain. Noise fill is performed by 681 noise fill module, frequency domain noise shaping is performed by 682 frequency domain noise shaping module, time domain transformation is performed by an ¡MDCT / OLA module 683 (OLA = overlap and sum) and the long-term prediction is executed by a long-term prediction unit 684. The long-term prediction unit may comprise eg a delay buffer (not illustrated in Fig. 6).
The signal derived from the second portion of the audio signal is then fed to the RMS 660 module to obtain a second value that indicates the obtaining of a root mean square of that signal derived from the second portion of the audio signal. This second value (second RMS value) is still represented in the time domain. Next, unit 620 transforms the second RMS value from the time domain to the tracking domain, in this case the LPC domain (ACELP). The second is then fed <sup>, υ</sup>'ΙΜΡΙ ^ ι
OF THE PROPERTY
INDUSTRIAL RMS value, which is represented in the tracking domain, to the noise level tracking unit 630.
In embodiments, level tracking is performed in the excitation domain, although TCX fading is performed in the time domain.
While the background noise level is tracked during normal decoding, it can be used eg during packet losses as an indicator of an appropriate comfort noise level, at which the last received signal per level gradually fades.
The derivation of the level to trace and the application of the fading by levels are, in general, independent of each other and could be executed in different domains. In the preferred embodiment, the level application is performed in the same domains as the level derivation, giving rise to the same benefits as in the case of ACELP, no gain compensation is necessary and, in the TCX case, Gain compensation inverse to that required for the level tap is needed (see Fig. 6) and thus the same gain tap can be used, as illustrated in Fig. 7.
Compensation for an effect of the high pass filter on the gain of LPC synthesis according to these embodiments is described hereinafter.
Fig. 8 summarizes this strategy. In particular, Fig. 8 illustrates the application of the comfort noise level during packet losses.
<img file="MX347233B_D0035.tif" />
<sup>108</sup> IMPI
MEXICAN INSTITUTE
OBLA PtOPIEJAO
INDUSTRIAL
In Fig. 8, the high-pass gain filter unit 643, the multiplier unit 644, the fading unit 645, the high-pass filter unit 646, the fading unit 647, and the combination unit 648 together , form a first reconstruction unit.
Furthermore, in Fig. 8, the background level provision unit 631 provides the noise level information. For example, the background level provision unit 631 can be implemented in the same manner as the background level tracking unit 630 of Fig. 6.
In addition, in FIG. 8, the LPC De-emphasis and Synthesis Gain unit 649 and the multiplication unit 641 together form a second transformation unit 640.
Furthermore, in Fig. 8, the fading unit 642 represents a second reconstruction unit.
In the embodiment of Fig. 8, the voice and non-voice excitement fades separately: The voice excitement fades to zero, but the non-voice excitement fades toward the comfort noise level. Fig. 8 further illustrates a high pass filter, which is introduced into the signal chain of the no-voice drive to suppress the low-frequency components for all cases, except when the signal is classified as no-voice.
Regarding the modeling of the influence of the high pass filter, the level after synthesis and de-emphasis of LPC is computed once with and once without the high pass filter. The relationship of these two levels is then derived and used to modify the applied background level.
109
IMPI ^ a
MEXICAN INSTITUTE
M LA FRONEDAO VWaffljnrfy
INDUSTRIAL
This is illustrated by Fig. 9. In particular, Fig. 9 illustrates advanced high-pass gain compensation during ACELP concealment in accordance with one embodiment.
Instead of the current drive signal, only a single pulse is used as input for this computation. This results in reduced complexity since the impulse response decays rapidly and therefore the RMS tap can be executed in a shorter time frame. In practice, only one subframe is used instead of the entire frame.
According to one embodiment, the noise level tracking unit 130 is configured to determine a comfort noise level as noise level information. The reconstruction unit 140 is configured to reconstruct the third portion of the audio signal depending on the noise level information, if said third frame of the plurality of frames is not received by the reception interface 110 or if said third frame is received by the receiving interface 110 but is corrupted.
According to one embodiment, the noise level tracking unit 130 is configured to determine a comfort noise level as noise level information. The reconstruction unit 140 is configured to reconstruct the third portion of the audio signal depending on the noise level information, if said third frame of the plurality of frames is not received by the reception interface 110 or if said third frame is received by the receiving interface 110 but is corrupted.
<img file="MX347233B_D0036.tif" />
110
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
In one embodiment, the noise level tracking unit 130 is configured to determine a comfort noise level as noise level information derived from a noise level spectrum, where said noise level spectrum is obtained by the application of the minimum statistics strategy. The reconstruction unit 140 is configured to reconstruct the third portion of the audio signal depending on a plurality of linear prediction coefficients, if said third frame of the plurality of frames is not received by the reception interface 110 or if said third frame it is received by the reception interface 110 but is corrupted.
In one embodiment, the (first and / or second) reconstruction unit 140, 141 may be configured, e.g., to reconstruct the third portion of the audio signal depending on the noise level information and depending on the first portion of the audio signal, if said third (fourth) frame of the plurality of frames is not received by the reception interface 110 or if said third (fourth) frame is received by the reception interface 110 but is corrupted.
According to one embodiment, the (first and / or second) reconstruction unit 140, 141 may be configured, eg, to reconstruct the third (or fourth) portion of the audio signal by attenuating or amplifying the first portion of the audio signal.
Fig. 14 illustrates an apparatus for decoding an audio signal. The apparatus comprises a reception interface 110, where the reception interface 110 is configured to receive a first frame comprising a first portion of
111 , IMPI ^ * INSTITUTO MEXICANO ---
<img file="MX347233B_D0037.tif" />
PROPERTY v industrial _ the audio signal of the audio signal, and where the receiving interface 110 is configured to receive a second frame comprising a second portion of the audio signal of the audio signal.
Furthermore, the apparatus comprises a noise level tracking unit 130, wherein the noise level tracking unit 130 is configured to determine noise level information depending on at least one of the first portion of the signal signal. audio and the second portion of the audio signal (this means: depending on the first portion of the audio signal and / or the second portion of the audio signal), where the noise level information is represented in a tracking domain.
In addition, the apparatus comprises a first reconstruction unit 140 for reconstructing, in a first reconstruction domain, a third portion of the audio signal from the audio signal depending on the noise level information, if a third frame of the plurality of frames is not received by reception interface 110 or if said third frame is received by reception interface 110 but is corrupted, where the first reconstruction domain is different from or equal to the crawl domain.
Furthermore, the apparatus comprises a transformation unit 121 for transforming the noise level information from the tracking domain to a second reconstruction domain, if a fourth frame of the plurality of frames is not received by the reception interface 110 or if said fourth frame is received by the reception interface 110 but it is corrupted, where the second domain of <sup>112</sup> ΙΜΡΙ ^ λ íwnmrro MEXICAN • DB LA PROPERTY, INDUSTRIAL reconstruction is different from the trace domain, and where the second reconstruction domain is different from the first reconstruction domain and
In addition, the apparatus comprises a second reconstruction unit 141 to reconstruct, in the second reconstruction domain, a fourth portion of the audio signal from the audio signal depending on the noise level information that is represented in the second domain. reconstruction, if said fourth frame of the plurality of frames is not received by the reception interface 110 or if said fourth frame is received by the reception interface 110 but is corrupted.
According to some embodiments, the tracking domain can be, eg, where the tracking domain is a time domain, a spectral domain, an FFT domain, an MDCT domain, or an excitation domain. . The first reconstruction domain can be, eg, the time domain, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain. The second reconstruction domain can be, eg, the time domain, the spectral domain, the FFT domain, the MDCT domain, or the excitation domain.
In one embodiment, the tracking domain can be, e.g., the FFT domain, the first reconstruction domain can be, e.g., the time domain, and the second reconstruction domain can be, e.g. ., the domain of arousal.
In another embodiment, the tracking domain may be, eg, the time domain, the first reconstruction domain may be, eg, the
<img file="MX347233B_D0038.tif" />
time domain and the second reconstruction domain can be, eg, the excitation domain.
According to one embodiment, said first portion of the audio signal may be represented, eg, in a first input domain and said second portion of the audio signal, may be represented, eg, in a second. input domain. The transformation unit can be eg a second transformation unit. The apparatus may further comprise, eg, a first transformation unit for transforming the second portion of the audio signal or a value or signal derived from the second portion of the audio signal from the second input domain to the tracking domain to obtain information from the second portion of the signal. The noise level tracking unit can be configured, eg, to receive information from the first portion of the signal that is represented in the tracking domain, where the information from the first portion of the signal depends on the first portion of the audio signal, where the noise level tracking unit is configured to receive the second portion of the signal that is represented in the tracking domain, and wherein the noise level tracking unit is configured to determine the noise level information depending on the information of the first portion of the signal that is represented in the tracking domain and depending on the information of the second portion of the signal that is represented in the trace domain.
114
INSTITUTO MEXICANO DI LA MOHEDAL · INDUSTRIAL
I INDUSTRIAL
According to one embodiment, the first input domain can be, eg, the excitation domain, and the second input domain can be, eg, the MDCT domain.
In another embodiment, the first input domain can be, eg, the MDCT domain, and where the second input domain can be, eg, the MDCT domain.
If, for example, a signal is represented in a time domain, it may be represented, eg, by samples of the signal in the time domain. Or, for example, if a signal is represented in a spectral domain, it may be represented, eg, by spectral samples of a spectrum of the signal.
In one embodiment, the tracking domain can be, eg, the FFT domain, the first reconstruction domain can be, eg, the time domain, and the second reconstruction domain can be, eg. , the domain of arousal.
In another embodiment, the tracking domain can be, e.g., the time domain, the first reconstruction domain can be, e.g., the time domain, and the second reconstruction domain can be, e.g. ., the domain of arousal.
In some embodiments, the units illustrated in Fig. 14 may be configured, for example, in accordance with what is described in Figs. 1a, 1b, 1c and 1d.
115
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
With regard to specific embodiments, in, for example, a low-rate mode, an apparatus according to an embodiment may receive, for example, ACELP frames as input, which are represented in the drive domain and which are then transformed to a time domain by LPC synthesis. Furthermore, in the low-rate mode, the apparatus according to one embodiment can receive, for example, TCX frames as input, which are represented in an MDCT domain, and which are then transformed into a time domain. using a reverse MDCT.
Tracing is then carried out in an FFT domain, where the FFT signal is derived from the signal in the time domain by executing an FFT (Fast Fourier Transform). Tracing can be carried out, for example, by running a minimum statistics strategy, separately for all spectral lines to obtain a comfort noise spectrum.
The concealment is then carried out by executing the level tap based on the comfort noise spectrum. The level derivation is performed on the basis of the comfort noise spectrum. Level conversion to time domain is performed for FD TCX PLC. A fade is executed in the time domain. A level tap to the excitation domain is performed in the case of ACELP PLC and TD TCX PLC (similar to ACELP). This is followed by fading in the excitation domain.
<img file="MX347233B_D0039.tif" />
116
The following list summarizes this:
low rate:
IMPI
INSTITUTO MEXICANO DE LA ERONEDAD INDUSTRIAL • entry:
or acelp (excitation domain -> time domain, by LPC synthesis) or tcx (MDCT domain -> time domain, by reverse MDCT) • screening:
o fft domain, derived from time domain by FFT or minimum statistics, separately for all spectral lines -> comfort noise spectrum • concealment:
o level derivation based on the comfort noise spectrum o level conversion to time domain for
FD TCX PLC
-> fading in time domain or level conversion to excitation domain for
ACELP PLC
<img file="MX347233B_D0040.tif" />
117 • TD TCX PLC (similar to
IMPI
MEXICAN INSTITUTE
OF THE INDUSTRIAL FROFITY
ACELP)
-> fading in the arousal domain
In, for example, a high rate mode, it can receive, for example, TCX frames as input, which are represented in the MDCT domain and which are then transformed to the time domain by inverse MDCT.
The time domain tracking is then carried out. Tracing can be done, for example, by running a minimum statistics strategy based on energy level to obtain a comfort noise level.
For concealment, in the case of FD TCX PLC, the level can be used as is and only a fade can be performed in the time domain. In the case of TD TCX PLC (similar to ACELP), level conversion is performed to the excitation domain and fading to the excitation domain.
The following list summarizes this:
high rate:
• entry:
or tcx (MDCT domain -> time domain, by reverse MDCT) • trace:
o time domain o minimal stats on energy level ->
comfort noise level rtW4EaUKU9
118 concealment
MEXICAN INSTITUTE
US LA TRCNLDAD INDUSTRIAL or use the level as is <sup>1</sup><sup>1</sup>
FD TCX PLC
-> time domain fading or level conversion to excitation domain for
TD TCX PLC (similar to ACELP)
-> fading in the arousal domain
Both the FFT domain and the MDCT domain are spectral domains, while the excitation domain is some kind of time domain.
According to one embodiment, the first reconstruction unit 140 may be configured, eg, to reconstruct the third portion of the audio signal by executing a first fading to a noise-like spectrum. The second reconstruction unit 141 may be configured, eg, to reconstruct the fourth portion of the audio signal by executing a second fade to a noise-like spectrum and / or a second fade to a LTP gain. Furthermore, the first reconstruction unit 140 and the second reconstruction unit 141 may be configured, eg, to perform the first fading and the second fading to a noise-like spectrum and / or a second
119 fading from a LTP gain to * fading.
Adaptive spectral modeling of comfort noise is now considered.
To obtain the adaptive modeling of comfort noise during burst packet loss, a first step can be carried out, of finding the appropriate LPC coefficients that represent the background noise. These LPC coefficients can be derived during active speech using a minimal statistics strategy to find the spectrum of the background noise and then calculate the LPC coefficients from it using an arbitrary algorithm for LPC derivation disclosed in the literature. Some embodiments, for example, can directly convert the background noise spectrum into a representation that can be used directly for FDNS in the MDCT domain.
Comfort noise fading can be done in the ISF domain (also applicable in the LSF domain; LSF, Line spectral frequency):
Jcurrent [i] = «* flast [i] + (1 - a) · pfmeanH i = 0 ... 16 (26) adjusting pr<sub>mean</sub> to the appropriate LP coefficients that describe the comfort noise.
mexican institute
OF THE PROPERTY t_S — .LJn «@r
INDUSTRIAL
Regarding the above-described adaptive spectral modeling of comfort noise, a more general embodiment is illustrated in Fig. 11.
Fig. 11 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in accordance with one embodiment.
The apparatus comprises a reception interface 1110 for receiving one or more frames, a coefficient generator 1120 and a signal reconstructor 1130.
The coefficient generator 1120 is configured to determine, if a current frame of said one or more frames is received by the reception interface 1110 and if the current frame being received by the reception interface 1110 is not corrupted / erroneous, one or more first coefficients of the audio signal, which are contained in the current frame, where said one or more first coefficients of the audio signal indicate a characteristic of the encoded audio signal, and one or more noise coefficients indicating a background noise of the encoded audio signal. Furthermore, the coefficient generator 1120 is configured to generate one or more second coefficients of the audio signal, depending on said one or more first coefficients of the audio signal and depending on said one or more noise coefficients, if the frame The current frame is not received by the receiving interface 1110 or if the current frame being received by the receiving interface 1110 is corrupted / erroneous.
The audio signal reconstructor 1130 is configured to reconstruct a first portion of the reconstructed audio signal depending on said one
121
INSTITUTO MEXICANO DE LA FROFIÉDAD or more first coefficients of the audio signal, if the a¿ftüST ^ b frame is recorded by the reception interface 1110 and if the current frame that is being received by<sup>-</sup> the<sup></sup>Receive interface 1110 is not corrupted. Furthermore, the audio signal reconstructor 1130 is configured to reconstruct a second portion of the reconstructed audio signal depending on said one or more second coefficients of the audio signal, if the current frame is not received by the reception interface 1110. or if the current frame being received by the receiving interface 1110 is corrupted.
The determination of background noise is well known in the art (see, for example, [MarOlj: Rainer Martin, Noise power spectral density estimation based on optimal smoothing and minimal statistics, IEEE Transactions on Speech and Audio Processing 9 (2001), no. 5, 504-512), and in one embodiment, the apparatus proceeds accordingly.
In some embodiments, said one or more first coefficients of the audio signal may be, eg, one or more linear prediction filter coefficients of the encoded audio signal. In some embodiments, said one or more first coefficients of the audio signal may be, eg, one or more linear prediction filter coefficients of the encoded audio signal.
It is well known in the art how to reconstruct an audio signal, eg a speech signal, from linear prediction filter coefficients or immittance spectral pairs (see, for example, [3GP09c]: Speech codee speech Processing functions; adaptive multi-rate - wideband (AMRWB) speech codee; transcoding
122
IMPI
<img file="MX347233B_D0041.tif" />
INSTITUTO MEXIO m> PE LA fErtPIEDAO __ functions, 3GPP TS 26.190, 3rd Generation Partnership Project, in one embodiment, the signal reconstructor proceeds accordingly.
According to one embodiment, said one or more noise coefficients may be, eg, one or more linear prediction filter coefficients indicating the background noise of the encoded audio signal. In one embodiment, said one or more linear prediction filter coefficients may represent, eg, a spectral shape of the background noise.
In one embodiment, the coefficient generator 1120 may be configured, eg, to determine said one or more second portions of the audio signal such that said one or more second portions of the audio signal are one or more more linear prediction filter coefficients of the reconstructed audio signal, or such that said one or more first coefficients of the audio signal are one or more spectral pairs of immittance of the reconstructed audio signal.
According to one embodiment, the coefficient generator 1120 may be configured, eg, to generate said one or more second coefficients of the audio signal by applying the formula:
fcurrent [? '] - θ ”JÍast [*] + (1 θ)' Pernean [q in which fcu<sub>rrent</sub>[í] indicates one of said one or more second coefficients of the audio signal, where f<sub>ast</sub>\ í \ indicates one of said one or more first
IMPI
<img file="MX347233B_D0042.tif" />
MEXICAN INSTITUTE OF THE FHOMEDaD
123 coefficients of the audio signal, where pt<sub>mean</sub>[í \
INDUSTRIAL _______ is one of said one or more noise coefficients, where a is a real number where 0 <a <1, and where i is an index.
According to one embodiment, f<sub>ast</sub>[í \ indicates a linear prediction filter coefficient of the encoded audio signal, and where fcumnt [í \ indicates a linear prediction filter coefficient of the reconstructed audio signal.
In one embodiment, pt<sub>mean</sub>[t] may be, eg, a linear prediction filter coefficient indicating the background noise of the encoded audio signal.
According to one embodiment, the coefficient generator 1120 may be configured, eg, to generate at least 10 second coefficients of the audio signal as said one or more second coefficients of the audio signal.
In one embodiment, the coefficient generator 1120 may be configured, eg, to determine whether the current frame of said one or more frames is received by the receive interface 1110 and whether the current frame being received by the receiving interface 1110 is not corrupted, said one or more noise coefficients by determining a noise spectrum of the encoded audio signal.
Hereinafter considered the fading of the spectrum from MDCT to White Noise prior to FDNS Application.
Instead of randomly modifying the sign of an MDCT (sign transpose) box, the entire spectrum is filled with white noise, which is
<img file="MX347233B_D0043.tif" />
<img file="MX347233B_D0044.tif" />
models using the FDNS. To avoid an instantaneous change of the. characteristics of the spectrum, a crossfade is applied between sign transposition and noise fill. Cross fading can be done as follows:
for (i = 0; i <L_frame; i ++) {if (old_x [i]! = 0) {x [i] = (1 - cum_damping) * noise [i] + cum_damping * random_sign () * x_old [i] ;
} }
where:
cum_damping is the attenuation factor (absolute) - it decreases from frame to frame, starting at 1 and decreasing towards 0 x old is the spectrum of the last frame received random_sign returns 1 or -1 noise contains a random vector (white noise) that it is scaled in such a way that its root mean square (RMS) is similar to the last correct spectrum.
ΙΜΤΠΤυΤΟ MEXICAN INDUSTRIAL PROPERTY
The term random_sign () * old_x [i] characterizes the sign transposition process to randomize the phases and thus avoid harmonic repetitions.
Then another normalization of the energy level could be performed after the crossfade to ensure that the summation of energy does not drift due to the correlation of the two vectors.
In accordance with these embodiments, the first reconstruction unit 140 may be configured, eg, to reconstruct the third portion of the audio signal depending on the noise level information and depending on the first portion of the audio signal. Audio. In a specific embodiment, the first reconstruction unit 140 may be configured, eg, to reconstruct the third portion of the audio signal by attenuating or amplifying the first portion of the audio signal.
In some embodiments, the second reconstruction unit 141 may be configured, eg, to reconstruct the fourth portion of the audio signal depending on the noise level information and depending on the second portion of the audio signal. In a specific embodiment, the second reconstruction unit 141 may be configured, eg, to reconstruct the fourth portion of the audio signal by attenuating or amplifying the second portion of the audio signal.
126
IMPI ^ R mSTíTUTQ MEXICAN
OF THE VWbmJMAf PROPERTY
Regarding the fading previously described in MDCT to white noise prior to the application of FONS. ^^ fTITE ^ TT ^ elus ^ more general embodiment.
Fig. 12 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal in accordance with one embodiment.
The apparatus comprises a reception interface 1210 for receiving one or more frames comprising information on a plurality of samples of audio signals from a spectrum of the audio signal of the encoded audio signal and a processor 1220 for generating the audio signal. reconstructed audio.
The processor 1220 is configured to generate the reconstructed audio signal by fading a modified spectrum to a target spectrum, if a current frame is not received by the receive interface 1210 or if the current frame is received by the receive interface 1210. but is corrupted, where the modified spectrum comprises a plurality of modified signal samples, where, for each of the modified signal samples of the modified spectrum, an absolute value of said sample of the modified signal is equal to an absolute value of one of the samples of the audio signal from the spectrum of the audio signal.
Furthermore, the processor 1220 is configured not to fade the modified spectrum to the white noise spectrum, if the current frame of said one or more frames is received by the receiving interface 1210 and if the current frame being received by the interface 1210 reception is not corrupted.
127
<img file="MX347233B_D0045.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
According to one embodiment, the target spectrum is a noise-type spectrum.
In one embodiment, the noise spectrum represents white noise.
According to one embodiment, the noise spectrum is modeled.
In one embodiment, the shape of the noise spectrum depends on a spectrum of the audio signal of a previously received signal.
According to one embodiment, the noise spectrum is modeled depending on the spectrum shape of the audio signal.
In one embodiment, processor 1220 uses a tilt factor to model the noise spectrum.
According to one embodiment, processor 1220 uses the formula shaped_noise [i] = noise * power (tilt_factor, i / N) where n indicates the number of samples, where i is an index, where 0 <= i <N , with tilt factor> 0, where power is a function of power.
If the tilt_factor is less than 1 this means attenuation with increasing i.
If the tílt_factor is greater than 1 it means amplification with increment of i.
In accordance with another embodiment, processor 1220 may employ the formula
128 mexican institute
FROM THE PROPERTY shaped_noise [i] = noise * (1 + i / (Nl) - * (til ^^ actor ^ T) where N indicates the number of samples, 'where i is an index, where 0 <= i <n , where tilt factor> 0.
According to one embodiment, the processor 1220 is configured to generate the modified spectrum, by changing the sign of one or more of the samples of the audio signal from the spectrum of the audio signal, if the current frame is not received by receive interface 1210 or if the current frame being received by receive interface 1210 is corrupted.
In one embodiment, each of the audio signal samples in the audio signal spectrum is represented by a real number and not an imaginary number.
According to one embodiment, the audio signal samples from the audio signal spectrum are represented in a Domain of the Modified Discrete Cosine Transform.
In another embodiment, the audio signal samples from the audio signal spectrum are represented in a Domain of the Modified Discrete Sine Transform.
In accordance with one embodiment, processor 1220 is configured to generate the modified spectrum by employing a random sign function that randomly or pseudo-randomly outputs a first or second value.
129
IMPIAS
MEXICAN INSTITUTE
OF LAFROFIEDAD
In one embodiment, processor 1220 is<sup>N</sup>^ ft £ ursR1fFpí ^ fade the modified spectrum towards the spectr5 ~ ^ üj5tlvóniisawena— subsequent reduction of an attenuation factor.
In accordance with one embodiment, processor 1220 is configured to fade the modified spectrum toward the target spectrum by subsequently increasing an attenuation factor.
In one embodiment, if the current frame is not received by the receive interface 1210 or if the current frame that is being received by the receive interface 1210 is corrupted, the processor 1220 is configured to generate the reconstructed audio signal. by using the formula:
x [i] = (l-cum_damping) * noise [i] + cum_damping * random_sign () * x_old [i] where i is an index, where x [i] indicates a sample of the reconstructed audio signal, where cum_damping is an attenuation factor, where x_old [i] indicates one of the audio signal samples from the audio signal spectrum of the encoded audio signal, where random_sign () returns 1 or -1, and where noise is a vector random indicating the target spectrum.
Some embodiments continue a TCX LTP operation. In those embodiments, the TCX LTP operation continues during concealment with the LTP parameters (LTP delay and LTP gain) derived from the last good frame.
LTP operations can be summarized as follows:
<img file="MX347233B_D0046.tif" />
IMPI «ffim / TO MEXICANO kiawwhoac
INDUSTRIAL
130
LTP Delay Buffer Feed based on previously derived output.
- Based on LTP delay: choice of the appropriate portion of the signal from the LTP delay buffer that is used as LTP's contribution to shaping the current signal.
- Rescaling of this LTP contribution using the LTP gain.
- Sum of this rescaled LTP contribution to the LTP input signal to generate the LTP output signal.
Different strategies could be considered with respect to timing, when executing the LTP delay buffer update:
As the first LTP operation in frame n using the output of the last frame n-1. This updates the LTP delay buffer in frame n to be used during LTP processing in frame n.
As the last LTP operation in frame n using the output of current frame n. This updates the LTP delay buffer in frame n to be used during LTP processing in frame n + 1.
Hereinafter, the decoupling of the TCX LTP feedback loop is described.
Decoupling the TCX LTP feedback loop prevents the introduction of additional noise (produced by noise substitution applied to the LPT input signal) during each feedback loop of the LTP decoder when in cloaking mode.
<sup>131</sup> .
IMPIAS. '^ «MEXICAN TITUTE
Fig. 10 illustrates this decoupling. In particular,<sup>D £</sup>l4 ™ ^ * ¿1í! ^^ LTP feedback loop decoupling during ^^ adiating ^ l ^^
FIG. 10 illustrates a delay buffer 1020, a sample selector 1030, and a sample processor 1040 (the sample processor 1040 is indicated by the dashed line).
In terms of timing, when the LTP 1020 delay buffer update is executed, some embodiments proceed as follows:
- For normal operation: It might be preferable to update the LTP 1020 delay buffer as the first LTP operation, since the summed output signal is usually stored persistently. With this strategy, a specialized buffer can be bypassed.
- For decoupled operation: It might be preferable to update the LTP 1020 delay buffer as the last operation, since LTP's contribution to the signal is usually only temporarily stored. With this strategy, the signal from the transient contribution of LTP is preserved. Implementation wise, this LTP contribution buffer could be kept constant.
Assuming the latter strategy is used in any case (normal operation and hiding), implementations can implement, for example, the following:
- During normal operation: The LTP decoder time-domain output signal is used after its addition to the LTP input signal to feed the LTP delay buffer.
132
<img file="MX347233B_D0047.tif" />
IMPI
MEXICAN INSTrTUTE
FROM AGE
IN0WHUa<sub>L</sub>
During concealment: The time domain output signal from the LTP decoder is used prior to its addition to the LTP input signal to feed the LTP delay buffer.
Some embodiments fade the TCX LTP gain to zero. In those embodiments, the TCX LTP gain can be faded, eg, to zero with a certain signal adaptive fading factor. This can be done, for example, iteratively, for example, according to the following pseudo code:
gain = gain_past * damping;
[· · ·] Gain_past = gain;
where:
gain is the gain of the TCX LTP decoder applied in the current frame;
gain past is the gain of the TCX LTP decoder applied in the previous frame;
damping is the (relative) fade factor.
Fig. 1d illustrates an apparatus according to another embodiment, wherein the apparatus further comprises a long-term prediction unit 170 comprising a delay buffer 180. The long-term prediction unit 170 is configured to generate a signal processed depending on the second
<img file="MX347233B_D0048.tif" />
portion of the audio signal, depending on a delay buffer input that is stored in delay buffer 180 and depending on a long-term prediction gain. Furthermore, the long-term prediction unit is configured to fade the long-term prediction gain towards zero, if said third frame of the plurality of frames is not received by the reception interface 110 or if said third frame is received by the receiving interface 110 but is corrupt.
In other embodiments (not illustrated), the long-term prediction unit may be configured, eg, to generate a processed signal depending on the first portion of the audio signal, depending on an input from the delay buffer. which is stored in the delay buffer and depending on a long-term prediction gain.
In Fig. 1d, the first reconstruction unit 140 may generate, eg, the third portion of the audio signal in addition depending on the processed signal.
In one embodiment, the long-term prediction unit 170 may be configured, eg, to fade the long-term prediction gain toward zero, where the rate at which the long-term prediction gain fades to zero. zero depends on a fading factor.
On the other hand, or in addition, the long-term prediction unit 170 may be configured, eg, to update the input of the delay buffer 180 by storing the processed signal generated in the delay buffer 180 if said third frame of the plurality of frames is not received by the reception interface ^ **** x *. ^ · *. 1 II w
<img file="MX347233B_D0049.tif" />
110 or if said third frame is received by the reception interface 110 but is corrupted.
Regarding the above-described use of TCX LTP, a more general embodiment is illustrated in Fig. 13.
Fig. 13 illustrates an apparatus for decoding an encoded audio signal to obtain a reconstructed audio signal.
The apparatus comprises a reception interface 1310 for receiving a plurality of frames, a delay buffer 1320 for storing audio signal samples of the decoded audio signal, a sample selector 1330 for selecting a plurality of samples of the audio signal selected from the audio signal samples that are stored in the delay buffer 1320 and a sample processor 1340 for processing the samples of the selected audio signal for obtain samples of the reconstructed audio signal from the reconstructed audio signal.
The sample selector 1330 is configured to select, if a current frame is received by the reception interface 1310 and if the current frame being received by the reception interface 1310 is not corrupted, the plurality of samples of the audio signal selected from the audio signal samples that are stored in the delay buffer 1320 depending on a pitch delay information that are contained in the current frame. Furthermore, the sample selector 1330 is configured to select, if the current frame is not received by the receive interface 1310 or if the current frame that is being received by the receive interface 1310 is corrupted, the
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL plurality of samples of the audio signal selected from the audio signal samples that are stored in the delay buffer 1320 depending on a tone delay information that is contained in another frame previously received by the interface reception 1310.
According to one embodiment, the sample processor 1340 may be configured, e.g., to obtain the samples of the reconstructed audio signal, if the current frame is received by the receive interface 1310 and if the current frame that is being received by the receiving interface 1310 is not corrupted, by rescaling the selected audio signal samples depending on the gain information that is contained in the current frame. Furthermore, the sample selector 1330 may be configured, e.g., to sample the reconstructed audio signal, if the current frame is not received by the receive interface 1310 or if the current frame is being received by the receiving interface 1310 is corrupted, by rescaling the selected audio signal samples depending on the gain information that is included in said additional frame previously received by the receiving interface 1310.
In one embodiment, the sample processor 1340 may be configured, e.g., to sample the reconstructed audio signal, if the current frame is received by the receive interface 1310 and if the current frame being received by the reception interface 1310 is not corrupted, multiplying the selected audio signal samples and a value that depends on the gain information that are contained in the current frame.
>36 <sub>:</sub> IMPI »?
* iNSTnvnoMsxiONG
;. l OF THE nomOAO Q ^ mwLiPGg i *<sub>and</sub> ENWUSTRÍAt \ ·.
Furthermore, the sample selector 1330 is configured to sample the reconstructed audio signal, if the current frame is not received by the receive interface 1310 or if the current frame that is being received by the receive interface 1310 is corrupted, multiplying the selected audio signal samples and a value that depends on the gain information that is included in said additional frame previously received by the receiving interface 1310.
According to one embodiment, the sample processor 1340 may be configured, eg, to store the samples of the reconstructed audio signal in the delay buffer 1320.
In one embodiment, sample processor 1340 may be configured, eg, to store samples of the reconstructed audio signal in delay buffer 1320 prior to receiving another frame by receiving interface 1310.
According to one embodiment, sample processor 1340 may be configured, eg, to store samples of the reconstructed audio signal in delay buffer 1320 upon receipt of another frame by receive interface 1310.
In one embodiment, the sample processor 1340 may be configured, eg, to rescale the selected audio signal samples depending on the gain information to obtain rescaled audio signal samples and combining the samples of the audio signal. audio signal; IMPI ^ a · MEXICAN USTITUTE
i. i DSLAFROfiEtiAo ______. J __________.__. _ -. . _<sub>x</sub> , INDUSTRIAL rescaled with samples of the ^ input audio signal to obtain the samples of the processed audio signal.
In accordance with one embodiment, the sample processor 1340 may be configured, eg, to store the processed audio signal samples, indicating the combination of the rescaled audio signal samples and the samples of the audio signal. input audio signal, in delay buffer 1320 and not to store the rescaled audio signal samples in delay buffer 1320, if the current frame is received by the receive interface 1310 and if the current frame that is being received by the receive interface 1310 is not corrupted. Furthermore, the sample processor 1340 is configured to store the rescaled audio signal samples in the delay buffer.
1320 and not to store the processed audio signal samples in the delay buffer 1320, if the current frame is not received by the receive interface 1310 or if the current frame that is being received by the receive interface 1310 is corrupted.
According to another embodiment, the sample processor 1340 may be configured, eg, to store the samples of the processed audio signal in the delay buffer 1320, if the current frame is not received by the receiving interface. 1310 or if the current frame being received by the receiving interface 1310 is corrupted.
In one embodiment, the sample selector 1330 may be configured, eg, to obtain the samples from the reconstructed audio signal by rescaling the samples from the. selected audio signal
138 <sub>;</sub> ΪΜΡΙ® ^ 'Mexican INSTrnrro
CE LA PRCFtEPAD depending on a modified gain, where the gain rffóflTffóhdaSeiféme according to the formula:
gain = gain_past * damping;
where gain is the modified gain, where the 1330 sample selector can be set, eg to set gain_past to gain once gain y has been calculated and where damping is a real number.
According to one embodiment, the sample selector 1330 may be configured, eg, to calculate the modified gain.
In one embodiment, damping can be defined, eg, according to: 0 <damping <1.
According to one embodiment, the modified gain gain can be set, eg, to zero, if at least a predetermined number of frames has not been received by the receiving interface 1310 since the last frame that was received by the receiving interface 1310.
Hereinafter the fading rate is considered. There are several concealment modules that apply a certain type of fading. Although the rate of this fading could be chosen differently between those modules, it is convenient to use the same rate of fading for all concealment modules corresponding to a kernel (ACELP or TCX). For example:
- 'niax — ej— i *.
<img file="MX347233B_D0050.tif" />
In the case of ACELP, the same fading rate should be used, in particular, for the adaptive codebook (by altering the gain), and / or for the innovative codebook signal (by altering the gain).
Furthermore, in the case of TCX, the same fading rate should be used, in particular, for the time-domain signal, and / or for the LTP gain (zero fading), and / or for the weighting of LPC (fade to one, and / or for LP coefficients (fade to spectral shape) and / or for cross fade to white noise.
It might also be preferable to use the same fade rate for ACELP and TCX, although due to the different nature of the cores, different fade rates could also be chosen.
This fading rate could be static, although it is preferably adaptive to the characteristics of the signal. For example, the fading rate may depend, eg, on the LPC stability factor (TCX) and / or a classification and / or a number of consecutively lost frames.
The fading speed can be determined, eg depending on the attenuation factor, which could be given in absolute or relative form, and which could also change with time during a fading.
In embodiments, the same fading rate is used for the LTP gain fading as for the white noise fading.
<img file="MX347233B_D0051.tif" />
An apparatus, method and computer program for generating a comfort noise signal in accordance with the above has been disclosed.
Although some aspects have been described in the context of an apparatus, it is obvious that these aspects also represent a description of the corresponding method, in which a block or device corresponds to a method step or a characteristic of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or item or of a characteristic of a corresponding apparatus.
The decomposed signal of the present invention can be stored on a digital medium or it can be transmitted by a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or software. The implementation can be carried out using a digital storage medium, for example a floppy disk, a DVD, a Blue-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, which has stored in the same electronically readable control signals, which cooperate (or have the ability to cooperate) with a programmable computer system in such a way that the respective method is executed.
IMPI «srnu-na Mexican INDUSTRIAL PROPERTY
Some embodiments according to the invention comprise a data carrier comprising electronically readable control signals, capable of cooperating with a programmable computing system in such a way that one of the methods described herein is executed.
In general, embodiments of the present invention can be implemented as a computer program product with a program code, where the program code performs the function of executing one of the methods when the computer program is executed on a computer. The program code can be stored, for example, on a machine-readable carrier.
Other embodiments comprise the computer program to execute one of the methods described herein, stored on a machine-readable carrier.
In other words, an embodiment of the method of the invention therefore consists of a computer program consisting of a program code to perform one of the methods described herein when the computer program is executed on a computer.
Another form of embodiment of the methods of the invention consists, therefore, in a data carrier (or digital storage medium, or computer-readable medium) comprising, recorded therein, the computer program to execute one of the methods described here.
Another embodiment of the method of the invention is, therefore, a data stream or a signal sequence representing the program of
IMPIAS
MEXICAN WSTrrUTO
OF THE PROPERTY Γ ^ ί ^ Π / computation to execute one of the methods described here'.Tffiüjo a & wrtos or the signal sequence can be configured, for example, páT9 SST * transferred through a data communication connection, for example on the Internet.
Another embodiment comprises a processing means, for example a computer, a programmable logic device, configured or adapted to execute one of the methods described herein.
Another embodiment comprises a computer in which the computer program has been installed to execute one of the methods described here.
In some embodiments, a programmable logic device (eg, an array of field-programmable gates) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field-programmable gate array may cooperate with a microprocessor to execute one of the methods described herein. In general, the methods are preferably executed by any hardware apparatus.
The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations to the provisions and details described herein are to be apparent to persons trained in the art. Therefore, it is only intended to limit the scope of the following patent claims
143
IMPÍ
MEXICAN INSTITUTE
OF INDUSTRIAL PROPERTY
<img file="MX347233B_D0052.tif" />
and not to the specific details presented by way of description and explanation of the embodiments presented herein.
<img file="MX347233B_D0053.tif" />
144
IMPI INSTITUTO MEXICANO _ DE LA PROPERTY
Industrial reference
[3GP09a] 3GPP; Technical Specification Group Seulceü 91 id Sybltíiu 'Aspecly, · - Extended adaptive multi-rate - wideband (AMR-WB +) codee, 3GPP TS 26.290, 3rd Generation Partnership Project, 2009.
[3GP09b] Extended adaptive multi-rate - wideband (AMR-WB +) codee; floating-point ANSI-C code, 3GPP TS 26.304, 3rd Generation Partnership Project, 2009.
[3GP09c] Speech codee speech Processing functions; adaptive multi-rate - wideband (AMRWB) speech codee; transcoding functions, 3GPP TS 26.190, 3rd Generation Partnership Project, 2009.
[3GP12a] Adaptive multi-rate (AMR) speech codee; error concealment of lost frames (release 11), 3GPP TS 26.091, 3rd Generation Partnership Project, Sep 2012.
[3GP12b] Adaptive multi-rate (AMR) speech codee; transcoding functions (release 11), 3GPP TS 26.090, 3rd Generation Partnership Project, Sep 2012. [3GP12c], ANSI-C code for the adaptive multi-rate - wideband (AMR-WB) speech codee, 3GPP TS 26.173, 3rd Generation Partnership Project, Sep 2012.
[3GP12d] ANSI-C code for the floating-point adaptive multi-rate (AMR) speech codee (releasell), 3GPP TS 26.104, 3rd Generation Partnership Project, Sep 2012.
[3GP12e] General audio codee audio Processing functions; Enhanced aacPlus general audio codee; additional decoder tools (release 11), 3GPP TS 26.402, 3rd Generation Partnership Project, Sep 2012.
IMPI ^^ ustttvto Mexican Ja
OF PROPERTY nBW
INDUSTRIAL
[3GP12f] Speech codee speech Processing functions; adaptive multi-rate - wideband (amr-wb) speech codee; ansi-c code, 3GPP TS 26.04, ¿rd Generation Partnership Project, 2012.
[3GP12g] Speech codee speech processing functions; adaptive multi-rate - wideband (AMR-WB) speech codee; error concealment of erroneous or lost frames, 3GPP TS 26.191, 3rd Generation Partnership Project, Sep 2012.
[BJH06] I. Batina, J. Jensen, and R. Heusdens, Noise power spectrum estimation for speech enhancement using an autoregressive model for speech power spectrum dynamics, in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. 3 (2006), 1064-1067.
[BP06] A. Borowicz and A. Petrovsky, Minimal controlled noise estimation for kltbased speech enhancement, CD-ROM, 2006, Italy, Florence.
[Coh03] I. Cohen, Noise spectrum estimation in adverse environments: Improved minimal controlled recursive averaging, IEEE Trans. Speech Audio Process. 11 (2003), no. 5, 466-475.
[CPK08] Choong Sang Cho, Nam In Park, and Hong Kook Kim, A packet loss concealment algorithm robust to burst packet loss for cell-type speech coders, Tech, report, Korea Enectronics Technology Institute, Gwang Institute of Science and Technology, 2008 , The 23rd International Technical Conference on Circuits / Systems, Computers and Communications (ITCCSCC 2008).
[Dob95] G. Doblinger, Computationally efficient speech enhancement by minimal spectral trackingin subbands, in Proc. Eurospeech (1995), 1513-1516.
146
<img file="MX347233B_D0054.tif" />
IMPI
ENSTRUTO MEXICANO OBLA FRGMETM OR
[EBU10]
[EBU12]
[ΕΗ08]
[ΕΜ84]
[ΕΜ85]
[Gan05]
[HE95]
[HHJ10] -, -<sub>Λ</sub> INDUSTRIAL _____
EBU / ETSI JTC Broadcast, Digital audio broadcastlng (DAB); transport of advanced audio coding (AAC) audio, ETt> l TS 102 563, European
Broadcasting Union, May 2010.
Digital radio mondiale (DRM); system specification, ETSI ES 201 980, ETSI,
Jun 2012.
Jan S. Erkelens and Richards Heusdens, Tracking of Nonstationary Noise Basad on Data-Driven Recursive Noise Power Estimation, Audio, Speech, and Language Processing, IEEE Transactions on 16 (2008), no. 6, 1112-1123.
Y. Ephraim and D. Malah, Speech enhancement using a minimum meansquare error short-time spectral amplitude estimator, IEEE Trans. Acoustics, Speech and Signal Processing 32 (1984), no. 6, 1109-1121.
Speech enhancement using a minimum mean-square error log-spectral amplitude estimator, IEEE Trans. Acoustics, Speech and Signal Processing 33 (1985), 443-445.
S. Gannot, Speech enhancement: Application of the kalman fiiter in the estimate-maximize (em framework), Springer, 2005.
HG Hirsch and C. Ehrlicher, Noise estimation techniques for robust speech recognition, Proc. IEEE Int. Conf. Acoustics, Speech, Signal Processing, no. pp. 153-156, IEEE, 1995.
Richard C. Hendriks, Richard Heusdens, and Jesper Jensen, MMSE based noise PSD tracking with low complexity, Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on, Mar 2010, pp. 4266 -4269.
147
IMPI ^
<td>[HJH08]</td><td>MEXICAN ITEMSrnVTO tx? ua fbomtdao Richard C. Hendriks, Jesper Jensen, and Richard Heus3ens * '1Vo / se? Rac ^) g</td>
<td>[ΙΕΤ12]</td><td>using dfí domain subspace decompositions, TEEE'TTShs. Audio, S'péech, Lang. Process. 16 (2008), no. 3, 541-553. IETF, Definition of the Opus Audio Codee, Tech. Report RFC 6716, Internet</td>
<td>[ISO09]</td><td>Engineering Task Force, Sep 2012. ISO / IEC JTC1 / SC29 / WG11, Informaron technology - coding of audio-visual objeets - part 3: Audio, ISO / IEC IS 14496-3, International Organization for Standardization, 2009.</td>
<td>[ITU03]</td><td>ITU-T, Wideband coding of speech at around 16 kbit / s using adaptive multirate wideband (amr-wb), Recommendation ITU-T G.722.2, Telecommunication Standardization Sector of ITU, Jul 2003.</td>
<td>[ITU05]</td><td>Low-complexity coding at 24 and 32 kbit / s for hands-free operation in systems with low frame loss, Recommendation ITU-T G.722.1,</td>
<td>[ITU06a]</td><td>Telecommunication Standardization Sector of ITU, May 2005. G. 722 Appendix III: A high-complexity algorithm for packet loss concealment for G.722, ITU-T Recommendation, ITU-T, Nov 2006.</td>
<td>[ITU06b]</td><td>G.729.1: G.729-based embedded variable bit-rate coder: An 8-32 kbit / s scalable wideband coder bitstream interoperable with g. 729, Recommendation ITU-T G.729.1, Telecommunication Standardization</td>
<td>[ITU07]</td><td>Sector of ITU, May 2006. G. 722 Appendix IV: A low-complexity algorithm for packet loss concealment with G.722, ITU-T Recommendation, ITU-T, Aug 2007.</td>
<td>[ITU08a]</td><td>G.718: Frame error robust narrow-band and wideband embedded variable bit-rate coding of speech and audio from 8-32 kbit / s, Recommendation ITUT G.718, Telecommunication Standardization Sector of ITU, Jun 2008.</td>
<sup>148</sup> IMPI ^ ®
MEXICAN INSTITUTE
Ot THE PROPERTY ιμχζγγχιαι
[ITU08b] G.719: Low-complexity, full-band audio coding for high-quality, conversational applications, Recommendation ITU-T G.719, Telecommunication Standardization Sector of ITU, Jun 2008.
[ITU 12] G.729: Coding of speech at 8 kbit / s using conjugate-structure algebraiccode-excited linear prediction (cs-acelp), Recommendation ITU-T G.729, Telecommunication Standardization Sector of ITU, June 2012.
[LS01] Pierre Lauber and Ralph Sperschneider, Error concealment for compressed digital audio, Audio Engineering Society Convention 111, no. 5460, Sep 2001.
<td>[Mar01]</td><td>Rainer Martin, Noise power spectral density estimation based on optimal smoothing and minimal statistics, IEEE Transactions on Speech and Audio Processing 9 (2001), no. 5, 504-512.</td>
<td>[Mar03]</td><td>Statistical methods for the enhancement of noisy speech, International Workshop on Acoustic Echo and Noise Control (IWAENC2003), Technical University of Braunschweig, Sep 2003.</td>
<td>[MC99]</td><td>R. Martin and R. Cox, New speech enhancement techniques for low bit rate speech coding, in Proc. IEEE Workshop on Speech Coding (1999), 165167.</td>
<td>[MCA99]</td><td>D. Malah, RV Cox, and AJ Accardi, Tracking speech-presence uncertainty to improve speech enhancement in nonstationary noise environments, Proc. IEEE Int. Conf. On Acoustics Speech and Signal Processing (1999), 789-792.</td>
<td>[MEP01]</td><td>Nikolaus Meine, Bernd Edler, and Heiko Purnhagen, Error protection and concealment for HILN MPEG-4 parametric audio coding, Audio Engineering Society Convention 110, no. 5300, May 2001.</td>
149
IMPI Msrrtvro mmscamo ¿e Ια ns? MW
[MPC89] Y. Mahieux, J.-P. Petit, and A. Charbonnier, TransfS ^^ clingoraudio signáis using correlation between successive HdHSMin blóólts, ^ coustic3 ~ Speech, and Signal Processing, 1989. ICASSP-89., 1989 International Conference on, 1989, pp. 2021-2024 vol.3.
[NMR + 12] Max Neuendorf, Markus Multrus, Nikolaus Rettelbach, Guillaume Fuchs, Julien Robilliard, Jérémie Lecomte, Stephan Wilde, Stefan Bayer, Sascha Disch, Christian Helmrich, Roch Lefebvre, Philippe Gournay, Bruno Bessette, Jimmy Lapierre, Kristopfer Kjorling, Heiko Purnhagen, Lars Villemoes, WernerOomen, Erik Schuijers, Kei Kikuiri, Toru Chinen, Takeshi Norimatsu, Chong Kok Seng, Eunmi Oh, Miyoung Kim, Schuyler Quackenbush, and Berndhard Grill, MPEG Unified Speech and Audio Coding - The ISO / MPEG Standard for High-Efficlency Audio Coding of all Content Types, Convention Paper 8654, AES, April 2012, Presented at the 132nd Convention Budapest, Hungary.
[PKJ + 11] Nam In Park, Hong Kook Kim, Min A Jung, Seong Ro Lee, and Seung Ho Choi, Burst packet loss concealment using multiple codebooks and comfort noise for celp-type speech coders in wireless sensor networks, Sensors 11 ( 2011), 5323-5336.
[QD03] Schuyler Quackenbush and Peter F. Driessen, Error mitigation in MPEG-4 audio packet communication systems, Audio Engineering Society Convention 115, no. 5981, Oct2003.
[RL06] S. Rangachari and PC Loizou, A noise-estimation algorithm for highly non-stationary environments, Speech Commun. 48 (2006), 220-231.
<sup>150</sup> ΙΜΡΙ ^
ÜNSTrrVTO MEXICO, • S? » LA MOMEO * OC »», „¿í_. WKWmiAl
[SFB00] V. Stahl, A. Fischer, and R. Bippus, Quantile based noise estimation for spectral subtraction and wiener filtering, in Proc. IEEE Int. Conf. Acoust., Speech and Signal Process. (2000), 1875-1878.
[SS98] J. Sohn and W. Sung, A voice activity detector employing soft decision based noise spectrum adaptation, Proc. IEEE Int. Conf. Acoustics, Speech, Signal Processing, no. pp. 365-368, IEEE, 1998.
[Yu09] Rongshan Yu, A low-complexity noise estimation algorithm based on smoothing of noise power estimation and estimation bias correction, Acoustics, Speech and Signal Processing, 2009. ICASSP 2009. IEEE International Conference on, Apr 2009, pp. 4421-4424.
<img file="MX347233B_D0055.tif" />
151
Contents128
76 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76
178 members in 20 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 13173154 | European Patent Office (EPO) | A | |
| 13173154 | European Patent Office (EPO) | A | |
| 131731549 | European Patent Office (EPO) | – | |
| 14166998 | European Patent Office (EPO) | A | |
| 14166998 | European Patent Office (EPO) | A | |
| 141669986 | European Patent Office (EPO) | – | |
| 2014063171 | European Patent Office (EPO) | W | |
| 2014063171 | European Patent Office (EPO) | W | |
| 131731549 | – | – | – |
| 141669986 | – | – | – |
| EP20130173154 | – | – | – |
| EP20140166998 | – | – | – |
| PCTEP2014063171 | – | – | – |
| WO2014EP63171 | – | – | – |
Members178
| Document | Office | Kind | |
|---|---|---|---|
| CA2913578A1 | Canada | A1 | |
| CA2914869A1 | Canada | A1 | |
| CA2914895A1 | Canada | A1 | |
| CA2915014A1 | Canada | A1 | |
| CA2916150A1 | Canada | A1 | |
| WO2014202784A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202786A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202788A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202789A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014202790A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201508736A | Taiwan Province of China | A | |
| TW201508737A | Taiwan Province of China | A | |
| TW201508738A | Taiwan Province of China | A | |
| TW201508739A | Taiwan Province of China | A | |
| TW201508740A | Taiwan Province of China | A | |
| AR096692A1 | Argentina | A1 | |
| AR096693A1 | Argentina | A1 | |
| AR096695A1 | Argentina | A1 | |
| AR096696A1 | Argentina | A1 | |
| AR096698A1 | Argentina | A1 | |
| SG11201510352YA | Singapore | A | |
| SG11201510353RA | Singapore | A | |
| SG11201510508QA | Singapore | A | |
| SG11201510510PA | Singapore | A | |
| SG11201510519RA | Singapore | A | |
| AU2014283123A1 | Australia | A1 | |
| AU2014283194A1 | Australia | A1 | |
| AU2014283124A1 | Australia | A1 | |
| AU2014283196A1 | Australia | A1 | |
| AU2014283198A1 | Australia | A1 | |
| CN105340007A | China | A | |
| CN105359209A | China | A | |
| CN105359210A | China | A | |
| KR20160021295A | Republic of Korea | A | |
| KR20160022363A | Republic of Korea | A | |
| KR20160022364A | Republic of Korea | A | |
| KR20160022365A | Republic of Korea | A | |
| CN105378831A | China | A | |
| KR20160022886A | Republic of Korea | A | |
| CN105431903A | China | A | |
| MX2015016892A | Mexico | A | |
| MX2015017126A | Mexico | A | |
| US2016104487A1 | United States of America | A1 | |
| US2016104488A1 | United States of America | A1 | |
| US2016104489A1 | United States of America | A1 | |
| US2016104497A1 | United States of America | A1 | |
| US2016111095A1 | United States of America | A1 | |
| EP3011557A1 | European Patent Office (EPO) | A1 | |
| EP3011558A1 | European Patent Office (EPO) | A1 | |
| EP3011559A1 | European Patent Office (EPO) | A1 | |
| EP3011561A1 | European Patent Office (EPO) | A1 | |
| EP3011563A1 | European Patent Office (EPO) | A1 | |
| MX2015018024A | Mexico | A | |
| JP2016522453A | Japan | A | |
| JP2016523381A | Japan | A | |
| JP2016526704A | Japan | A | |
| JP2016527541A | Japan | A | |
| MX2015017261A | Mexico | A | |
| TWI553631B | Taiwan Province of China | B | |
| JP2016532143A | Japan | A | |
| AU2014283123B2 | Australia | B2 | |
| AU2014283124B2 | Australia | B2 | |
| AU2014283194B2 | Australia | B2 | |
| AU2014283196B2 | Australia | B2 | |
| AU2014283198B2 | Australia | B2 | |
| TWI564884B | Taiwan Province of China | B | |
| TWI569262B | Taiwan Province of China | B | |
| TWI575513B | Taiwan Province of China | B | |
| MX347233BThis record | Mexico | B | |
| EP3011557B1 | European Patent Office (EPO) | B1 | |
| EP3011561B1 | European Patent Office (EPO) | B1 | |
| TWI587290B | Taiwan Province of China | B | |
| RU2016101469A | Russian Federation | A | |
| BR112015031177A2 | Brazil | A2 | |
| BR112015031178A2 | Brazil | A2 | |
| BR112015031180A2 | Brazil | A2 | |
| BR112015031343A2 | Brazil | A2 | |
| BR112015031606A2 | Brazil | A2 | |
| PT3011557T | Portugal | T | |
| PT3011561T | Portugal | T | |
| EP3011558B1 | European Patent Office (EPO) | B1 | |
| EP3011559B1 | European Patent Office (EPO) | B1 | |
| RU2016101521A | Russian Federation | A | |
| RU2016101600A | Russian Federation | A | |
| RU2016101604A | Russian Federation | A | |
| RU2016101605A | Russian Federation | A | |
| HK1224009A1 | Hong Kong, China | A1 | |
| HK1224076A1 | Hong Kong, China | A1 | |
| HK1224423A1 | Hong Kong, China | A1 | |
| HK1224424A1 | Hong Kong, China | A1 | |
| HK1224425A1 | Hong Kong, China | A1 | |
| JP6190052B2 | Japan | B2 | |
| JP6196375B2 | Japan | B2 | |
| JP6201043B2 | Japan | B2 | |
| ES2635027T3 | Spain | T3 | |
| ES2635555T3 | Spain | T3 | |
| PT3011558T | Portugal | T | |
| MX351363B | Mexico | B | |
| KR101785227B1 | Republic of Korea | B1 | |
| JP6214071B2 | Japan | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 347233
- Publication, DOCDB
- 347233
- Publication, EPODOC
- MX347233
- Application
- 2015016638
- Application, DOCDB
- 2015016638
- Application, EPODOC
- MX20150016638
Titles2
- Spanish
- APARATO Y MÉTODO PARA DESVANECIMIENTO MEJORADO DE SEÑAL PARA SISTEMAS DE CODIFICACIÓN DE AUDIO CONMUTADOS DURANTE EL OCULTAMIENTO DE ERRORES.
- English
- APPARATUS AND METHOD FOR IMPROVED SIGNAL FADING FOR AUDIO ENCODING SYSTEMS SWITCHED DURING HIDING OF ERRORS.
Classification
- CPC, 14
- G10L19/005
- G10L19/002
- G10L19/0212
- G10L19/09
- G10L19/012
- G10L19/083
- H03M7/30
- G10L19/22
- G10L19/06
- G10L19/12
- G10L2019/0002
- G10L2019/0011
- G10L2019/0016
- G10L19/07
- IPC, 3
- G10L19 005
- G10L19 09
- G10L25 90