Audio decoder and method for providing a decoded audio information using an error concealment based on a time domain excitation signal.
Abstract
An audio decoder (100; 300) for providing a decoded audio information (112; 312) on the basis of an encoded audio information (110; 310) comprises an error concealment (130; 380; 500) configured to provide an error concealment audio information (132; 382; 512) for concealing a loss of an audio frame following an audio frame encoded in a frequency domain representation (322) using a time domain excitation signal (532).

Term
8.1 yearsleft in the term
Expires 27 October 2034.
- Priority
- Filed
- Granted
- Today
- Expires
29 claims: 7 independent, 22 dependent
- 1CLAIMS REIVINDICACIONES 1. Un decodificador de audio (100; 300) para proveer una información de audio decodificada (112; 312) sobre la base de una información de audio codificada (110; 310), el decodificador de audio comprende:one. An audio decoder (100;300) to provide decoded audio information (112;312) based on encoded audio information (110;310), the audio decoder comprises: an error concealment (130;380;500) configured to provide error concealment audio information (132;382;512) for concealing a loss of an audio frame after an audio frame encoded in a representation frequency domain (322) using a time domain excitation signal (532);un ocultamiento de error (130;380;500) configurado para proveer una información de audio de ocultamiento de error (132;382;512) para el ocultamiento de una pérdida de una trama de audio luego de una trama de audio codificada en una representación de dominio de frecuencia (322) usando una señal de excitación de dominio de tiempo (532);el decodificador de audio caracterizado porque: the audio decoder characterized in that: el ocultamiento de error (130;380;500) está configurado para combinar una señal de excitación de dominio de tiempo extrapolada (552) y una señal de ruido (562), a fin de obtener una señal de entrada (572) para una síntesis de codificación predictiva lineal (LPC) (580);y en donde el ocultamiento de error está configurado para realizar la síntesis de codificación predictiva lineal (LPC), en donde la síntesis de codificación predictiva lineal (LPC) está configurada para filtrar la señal de entrada (572) de la síntesis de codificación predictiva lineal (LPC) de the error concealment (130;380;500) is configured to combine an extrapolated time domain drive signal (552) and a noise signal (562), to obtain an input signal (572) for synthesis linear predictive coding (LPC) (580);and where error concealment is configured to perform linear predictive coding synthesis (LPC), where linear predictive coding synthesis (LPC) is configured to filter input signal (572) from linear predictive coding synthesis (LPC) of 161 161 IMPI O IMPI O INSTITUTO MEXICANO MEXICAN INSTITUTE DE LA MOHEDAL · OnINDUSTRIAL according to linear prediction coding parameters, in order to obtain the audio information of error concealment (132;382;512), where the error concealment (130;380;500) is DÉ LA MOHEDAL· OnINDUSTRIAL acuerdo con parámetros de codificación de predicción lineal, a fin de obtener la información de audio de ocultamiento de error (132;382;512), en donde el ocultamiento de error (130;380;500) está 5 configured for the high pass filter of the noise signal (562) that is combined with the extrapolated time domain excitation signal (552). 5 configurado para el filtro paso alto de la señal de ruido (562) que se combina con la señal de excitación de dominio de tiempo extrapolada (552).
- 14The audio decoder (100;300) according to claims 1 to 13, wherein the error concealment (130;380;500) is configured to compute a gain of the extrapolated time domain drive signal (552) , which is used to obtain the input signal (572) for linear predictive coding (LPC) synthesis (580), using a time domain correlation that is performed based on a time domain representation (122 ;372;378;510) of the audio frame encoded in the frequency domain representation (322) preceding the missing audio frame, where 14. El decodificador de audio (100;300) de acuerdo con las reivindicaciones 1 a 13, en donde el ocultamiento de error (130;380;500) está configurado para computar una ganancia de la señal de excitación de dominio de tiempo extrapolada (552), que se usa para obtener la señal de entrada (572) para la síntesis de codificación predictiva lineal (LPC) (580), usando una correlación en el dominio de tiempo que se realiza sobre la base de una representación de dominio de tiempo (122;372;378;510) de la trama de audio codificada en la representación de dominio de frecuencia (322) que precede la trama de audio perdida, en donde se 167 167 INSTlfÜTO MEXICANO,, DE LA MONEDAD establishes a correlation delay that agfféWde-hie ^ 'a height information obtained on the bá SS' “δδ '...... time domain excitation signal (532), or using a correlation in the excitation domain. INStlfÜTO MEXICANO , , DE LA MONEDAD establece una demora de correlación que agfféWde-hie^'una información de altura obtenida sobre la bá SS'“δδ'......ráseñal de excitación de dominio de tiempo (532), o usando una correlación en el dominio de excitación.
- 17The audio decoder (100;300) according to 17. El decodificador de audio (100;300) de acuerdo con 168 MEXICAN tax claim one of claims 1 to 16, whereWildebeest^’¥IAlerror occuTt! (130;380;500) is configured to mo3'rTÍcaT '”a time domain excitation signal (532) obtained on the basis of one or more audio frames preceding a frame 168 ííIsfiTUTO MEXICANO una de las reivindicaciones 1 a 16, en dond¿Nu^’¥IAlocuTt^!ííiento de error (130;380;500) está configurado para mo3'rTÍcaT'”úña señal de excitación de dominio de tiempo (532) obtenida sobre la base de una o más tramas de audio que preceden una trama 5 audio files, in order to obtain the audio information of the error concealment (132;382;512). 5 de audio perdida, a fin de obtener la información de audio de ocultamiento de error (132;382;512).
- 19The audio decoder (100;300) according to 19. El decodificador de audio (100;300) de acuerdo con 15 una de las reivindicaciones 17 o 18, en donde el ocultamiento de error (132;380;500) está configurado para modificar la señal de excitación de dominio de tiempo (532) obtenida sobre la base de una o más tramas de audio que preceden una trama de audio perdida o una o más de sus copias, de modo de fifteen one of claims 17 or 18, wherein the error concealment (132;380;500) is configured to modify the time domain driving signal (532) obtained on the basis of one or more audio frames preceding a lost audio frame or one or more of its copies, so
- 2020 reducir un componente periódico de la información de audio de ocultamiento de error (132;382;512) en función del tiempo. twenty reduce a periodic component of the error concealment audio information (132;382;512) as a function of time. 20. El decodificador de audio (100;300) de acuerdo con una de las reivindicaciones 17 a 19, en donde el ocultamiento twenty. The audio decoder (100;300) according to one of claims 17 to 19, wherein the concealment 169 169 IMPI de error (132;380;500) esta conf iguradoD£^g(^R1^s señal de excitación de dominio de t i empo_ f 532) chToniri? pobre la base de una o más tramas de audio que preceden la trama de audio perdida, o una o más de sus copias, de modo de modificar la señal de excitación de dominio de tiempo. IMPI error (132;380;500) is configuredD £^ g (^ R1^ s signal of excitation of domain of you empo_ f 532) chToniri? poor the basis of one or more audio frames preceding the missing audio frame, or one or more of its copies, so as to modify the time domain drive signal.
- 26The audio decoder (100;300) according to 26. El decodificador de audio (100;300) de acuerdo con 172 172 IMPI IMPI INSTITUTO MEXICANO DE LA fkOflEÜAD una de las reivindicaciones 1 a 25, en donde eTDuocultamTefito de error (130;380;500) está configurado para proveer’ la información de audio de ocultamiento de error (132;382;512) para un tiempo que es mayor que una duración temporal de una o más tramas de audio perdidas. MEXICAN INSTITUTE OF THE fkOflEÜAD one of claims 1 to 25, where eTDuerror concealment (130;380;500) is configured to provide the error concealment audio information (132;382;512) for a time that is greater than a time duration of one or more missing audio frames.
- 2930. A computer-readable storage medium to provide decoded audio information about the 30. Un medio de almacenamiento legible por computadora para proveer una información de audio decodificada sobre la 174 174 IMPI IMPI INSTITUTO MEXICANO DS LA PROPIEDAD base de una información de audio codificada, quéIÜLólólrttp método de acuerdo con la reivindicación 297 INSTITUTO MEXICANO DS LA PROPIEDAD based on coded audio information, whatIÜLorlorlrttp method according to claim 297 175 175 IMPI IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIA! MEXICAN INSTITUTE OF INDUSTRY PROPERTY!
Independent claims7
910 paragraphs in 109 sections, as filed
(54) Title: AUDIO DECODER AND METHOD FOR PROVIDING DECODED AUDIO INFORMATION USING AN ERROR HIDING BASED ON A TIME DOMAIN EXCITATION SIGNAL.
(54) Title: AUDIO DECODER AND METHOD FOR PROVIDING A DECODED AUDIO INFORMATION USING AN ERROR CONCEALMENT BASED ON A TIME DOMAIN EXCITATION SIGNAL.
(57) Summary
An audio decoder (100; 300) to provide decoded audio information (112; 312) based on encoded audio information (110; 310) comprises an error concealment (130; 380; 500) configured to provide an error concealment audio information (132; 382; 512) for concealing a loss of an audio frame after an audio frame encoded in a frequency domain representation (322), using a time domain drive signal (532).
(57) Abstract
An audio decoder (100; 300) for providing a decoded audio Information (112; 312) on the basis of an encoded audio Information (110; 310) comprises an error concealment (130; 380; 500) configured to provide an error concealment audio Information (132; 382; 512) for concealing a loss of an audio frame following an audio frame encoded in afrequency domain representation (322) using a time domain excitation signal (532).
Yes.
<img file="MX356334B_D0001.tif" />
PATENT TITLE No. 356334
Headlines):
Home
FRAUNHOFER-GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Hansastrasse 27c, 80686, Munich, GERMANY
Denomination:
Classification:
AUDIO DECODER AND METHOD FOR PROVIDING USANQEKUN DECODED AUDIO INFORMATION HIDDEN ERROR ON THE BASIS OF A SEMDE φ (|> Ο | 0 <φΕ TIME DOMAIN.
CIP: G1ÓL19 / 0¿5 ^ GT0L19 / 02; G1X19 / £) ¿jfl 025/90
9 / oéi
CPC:, «ft) OLI 9/005;
G10L19 / 08; G
JÉR <WED LECOMJÉ - ^ _
10L19 / 09; ÍWL19 / 125; G10L19 / 0212;
Inventor (s):
r
--i
Z '/
Country:
"•Goes
<img file="MX356334B_D0002.tif" />
end:
: ifc (I ^ FMB ^ ficinÉ la fro | ji5at || p8fetrial.
KyHAEL '^ CHNABEL; GRZEGORZ
Number; .
MX / a / 2016 / fi05535
Validity: VeMfe ^ ftos Date of Veiteihiienti Date of Ex |
Reference patent
Pursuant to the date of
Who subscribes to this title is (Official Gazette of the Federation (I 25/01/2006, 06/05 / 2009,06 / 01/2010, í Regulation of the Mexican Institute of articles 1, 3, 4, 5 ° fraction V subsection a) .1 12/27/1999, amended on 10/10/2002, 07/29/2004 Deputy Generals, Coordinator, Divisioi Departmental Directors and other subordinates of the Mexican Institute on 08/04/2004 and 13 / 09/2007).
<sup>r</sup>fi
Nttniéro:
EP1ÍÍ91133
ΕΡΐ4φ824 Industrial id.
isí / n extendable, counted to lot ^ rights.
Industrial Property Law
1999, 01/26/2004, 06/16/2005, smooth a), 4th and 12 'sections I and III of Ί 07/05/2004, 07/28/2004 and 09/07/2007); Mexican Industrial Property (DOF agreement that delegates powers to the Directors, Divisional Deputy Directors, Coordinators (DOF 12/15/1999, amended on 02/04/2000, 07/29/2004,
This letter is signed with an advanced electronic signature (FIEL), based on articles 7 BIS 2 of the Industrial Property Law; 3 of its Regulations, and 1 fraction III, 2 fraction V, 26 BIS and 26 TER of the Agreement establishing the guidelines for the use of the Electronic Payment and Services Portal (PASE) of the Mexican Institute of Industrial Property, in the procedures indicated.
DIVISIONAL DIRECTOR OF PATENTS NAHANNY CANAL REYES
<img file="MX356334B_D0003.tif" />
Original string:
NAHANNY MARISOL CANAL REYES | 00001000000403252793 | Tax Administration Service | 1695 || MX / 2018/43317 | MX / a / 2016/005535 | Patent title PCT | 1220 | RRGO | Page (s) | mTz1 + zty6qGvJx0qv0Lhci6j1 + U =
Digital stamp:
AGIQgFywfGtHqyAH6auDfvuQZ781bBxTDCTY7t9w9NU5eqexlvBoSkCoMoghj3aKvgk1¡Wu05grobefXecFyjncVoN OX6wxkJKG9wDJ9ue e2RRcwg2QbX9t9Vj44VzkGCeFNgK + + + KuT8CcYYofgpP7KJbisVmMzpxMHvRC JVzr1Xhre7awr t1zJ¡MflcHckgtGC4NMqWWZr + E + + iX + jdGqWqmzG2FNmX8UtWROjRDu zYv6nMr // AyoeqUxaotDgSLQOJxHJvV5TZZ
UM9Bp / X + Y7PQHXaJqUexng2xK1swjs3xSs / UotMKW42zLfBWKThn8MTnwlSXbPheGrCPNpOA ==
Arenal No. 550. Floor 1, Pueblo Santa Mana Tepepan, Xochimilco. 15020. Mexico City> 55) 53340700 www gob.mx/impi
<img file="MX356334B_D0004.tif" />
356331
IMPI
MEXICAN INSTITUTE OE LA MONEDAD INDUSTRIAL
<img file="MX356334B_D0005.tif" />
AUDIO DECODER AND Mffwnn papa ppovf.ER A DECODED AUDIO INFORMATION USING AN ERROR HIDING BASED ON A DOMAIN EXCITATION SIGNAL
WEATHER
Technical field.
Embodiments according to the invention create audio decoders to provide decoded audio information based on encoded audio information.
Some embodiments in accordance with the invention create methods for providing decoded audio information based on encoded audio information.
Some embodiments according to the invention create computer programs for carrying out one of said methods.
Some embodiments according to the invention relate to a time domain concealment for a transform domain [encoder-decoder] code.
Background of the invention.
JL1VL i I 'f% ¡^
MEXICANC INSTITUTE
OF THE PROPERTY V ^ e * ®¡3 $ -JÍ
INDUSTRIAL> ¿^ 2 ^
In recent years, there has been a growing demand for digital streaming and storage of audio content.
However, audio content is often transmitted over unreliable channels, which carries the risk that data units (eg packets) comprising one or more audio frames (eg in the form of a encoded representation, such as an encoded time domain representation or a frequency encoded domain representation) are lost. In some situations, it may be possible to require a repeat (forwarding) of lost audio frames (or of data units, such as packets, comprising one or more lost audio frames). However, this will typically produce a substantial delay, and, therefore, will require extensive buffering of audio frames. In other cases, it is almost impossible to require a repeat of missed audio frames.
In order to obtain good, or at least acceptable, audio quality in the event that the audio frames are lost without providing extensive temporary storage (which would consume a large amount of memory, and which would also substantially degrade the capacities real-time audio encoding), it is desirable to have concepts
IMPI
<img file="MX356334B_D0006.tif" />
to handle a loss of one or more audio frames. In particular, it is desirable to have concepts that produce good audio quality, or at least acceptable audio quality, even in the event that the audio frames are lost.
In the past, some error concealment concepts have been developed, which can be used in different audio coding concepts.
In the following, a conventional audio encoding concept will be described.
In the 3gpp TS26.290 standard, a transformed encoded excitation decoding (TCX decoding) with error concealment is explained. In the following, some explanations will be provided, based on the Signal Synthesis and TCX Mode Decoding section in reference [1].
A TCX decoder according to the Standard
International 3gpp TS 26.290 is shown in Figs. 7A-7B and
8, where Figs. 7A-7B and 8 show block diagrams of the transformed encoded excitation decoder (TCX). However, Fig. 7A-7B shows those functional blocks that are relevant for TCX decoding in
<img file="MX356334B_D0007.tif" />
IMPL imctit »rr <> υίγιη ^ ', a normal operation, or in a case of partial packet loss. In contrast, Fig. 8 shows the relevant processing of TCX decoding in the case of TCX-256 packet erase concealment.
In other words, Figs. 7A-7B and 8 show a block diagram of the TCX decoder that includes the following cases:
Case 1 (Fig. 8): Concealment of packet deletion in TCX-256 when TCX frame length is 256 samples and the related packet is lost, ie BFI_TCX = (1);
and
Case 2 (Fig. 7A-7B): normal TCX decoding, possibly with partial packet loss.
In the following, some explanations will be provided in relation to Figs. 7A-7B and 8.
As mentioned, Fig. 7A-7B shows a block diagram of a TCX decoder that performs TCX decoding in normal operation, or, in the case of partial packet loss. The TCX 700 decoder according to Fig. 7A-7B receives specific parameters from
TCX 710 and provides, on its basis, the decoded audio information 712, 714.
<img file="MX356334B_D0008.tif" />
<sup>5</sup> IMPI •• WTITUT MUICANO »» »W» tDAD «•« ΪΠΤΗΛΙ
Audio decoder 700 comprises a DEMUX TCX 720 demultiplexer, which is configured to receive TCX 710 specific parameters and BFI_TCX information. The demultiplexer 720 separates the specific parameters of TCX 710, and provides encoded excitation information 722, encoded noise fill information 724, and encoded global gain information 726. Audio decoder 700 comprises an excitation decoder 730, which is configured to receive encoded excitation information 722, encoded noise filler information 724, and encoded global gain information 726, as well as some additional information (for example , a flag_bit_rate bit rate flag, a BFI_TCX information and a TCX frame length information. The excitation decoder 730 provides, on its basis, a time domain excitation signal 728 (also denoted by x). Excitation decoder 730 comprises an excitation information processor 732, which demultiplexes encoded excitation information 722 and decodes algebraic quantization parameters. The excitation information processor 732 provides an intermediate excitation signal 734, which is typically located in a
<img file="MX356334B_D0009.tif" />
IMPI representation of the frequency domain, and which is designated by Y. The excitation encoder 730 further comprises a noise injector 736, which is configured to inject noise into unquantized subbands in order to derive an excitation signal. noise filled 738 of intermediate drive signal 734. Noise filled drive signal 738 is typically in the frequency domain, and is designated Z. Noise injector 736 receives noise intensity information 742 from a noise fill level decoder 740. The drive decoder further comprises an adaptive low-frequency de-emphasis 744, which is configured to perform a low-frequency de-emphasis operation based on the noise-filled drive signal 738, so as to obtain a processed drive signal 746, which is still in the frequency domain, and is designated by X '. The excitation decoder 730 further comprises a frequency domain to time domain transformer 748, which is configured to receive the processed excitation signal 746 and to provide, on its basis, a time domain excitation signal 750, which is associated with a certain portion of time represented by a set of frequency domain excitation parameters (for
<img file="MX356334B_D0010.tif" />
IMPI example, of processed excitation signal 746). The drive decoder 730 further comprises a scaler 752, which is configured to scale the time domain drive signal 7 50 to obtain a scaled time domain drive signal 754. Scaler 752 receives global gain information 756 from global gain decoder 758, where, in response, global gain decoder 758 receives encoded global gain information 726. Excitation decoder 730 further comprises an overlay synthesis and addition 760, which receives the scaled time domain driving signals 754 associated with a plurality of time slices. Overlap and add synthesis 760 performs an overlay and add operation (which may include a window operation) based on the scaled time domain drive signals
754, so as to obtain a temporarily combined time domain driving signal 728 for a longer period in time (longer than the periods in time for which the individual time domain driving signals 750 are provided, 754).
The audio decoder 700 further comprises a linear predictive coding synthesis (LPC, according to
ΙΜΡΪ
MEXICAN INSTITUTE
PROPERTY, INDUSTRIAL - »,?» r ^. · 770, which receives the time domain excitation signal 728 provided by the overlay and add synthesis 7 60 and one or more linear predictive coding coefficients (LPC) that define a filter function Predictive Coding Synthesis (LPC) 772. The linear predictive coding synthesis (LPC) 770, for example, can comprise a first filter 774, which, for example, can synthesize the time domain excitation signal 728, in order to obtain the decoded audio signal 712 . Optionally, the linear predictive coding synthesis (LPC) 770 may further comprise a second synthesis filter 772 which is configured to synthesize the output signal of the first filter 774 using another synthesis filter function, so as to obtain the signal decoded audio 714.
In the following, TCX encoding will be described in the case of a TCX-256 packet erase concealment. The
Fig. 8 shows a block diagram of the TCX decoder, in this case.
The packet drop concealment 800 receives a height 810 information, which is further designated tcx_height, and is derived from a previously decoded TCX frame. For example, height information
IMPI-
<img file="MX356334B_D0011.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL ___
810 can be obtained using a dominant height estimator
747 from the processed excitation signal 746 in the excitation decoder 730 (during normal decoding). Furthermore, the concealment of package deletion
800 receives linear predictive coding (LPC) parameters
812, which may represent a linear predictive coding synthesis (LPC) filter function. The linear predictive coding (LPC) parameters 812, for example, may be identical to the linear predictive coding (LPC) parameters 772. Accordingly, packet deletion concealment 800 can be configured to provide, based on height information 810 and linear predictive coding (LPC) parameters 812, an error concealment signal 814, which can be considered information audio concealment error. The packet erase concealment 800 comprises an excitation buffer 820, which, for example, may temporarily store a previous excitation. The Excitation Buffer
820, for example, can make use of the adaptive codebook ACELP [adaptive codebook-excited linear prediction], and can provide an excitation signal 822. Packet erase concealment 800 may comprise additionally a
<img file="MX356334B_D0012.tif" />
IMPI
MEXICAN INSTITUTE • j = · Ί J_ or Λ X r-. z, PROPERTY first filter 824, a filter function queiN ^ gireide rse as shown in Fig. 8. Therefore<sup>1</sup>,··<sup>1</sup> the first 'fITT 824 can filter the drive signal 822 based on the linear predictive coding (LPC) parameters 812, so as to obtain a filtered version 826 of the drive signal 822. Packet erasure concealment further comprises an amplitude limiter 828, which can limit an amplitude of the filtered excitation signal 826 based on target information or rms level information<sub>wsyn</sub>. Furthermore, the packet erase concealment 800 may comprise a second filter 832, which may be configured to receive the limited amplitude filtered excitation signal 830 from the amplitude limiter 822 and to provide, on its basis, the concealment signal of mistake
814. A filter function of the second filter 832, for example, can be defined as shown in Fig. 8.
In the following, some details regarding decoding and error concealment will be described.
In Case 1 (concealment of package deletion in
TCX-256), no information is available for decoding the 256-sample TCX frame. The synthesis of TCX is found by processing the past excitation delayed by T, where T = height_tcx is a delay of
<img file="MX356334B_D0013.tif" />
height estimated in the TCX frame previously a nonlinear filter approximately equivalent to
<img file="MX356334B_D0014.tif" />
use a non-linear filter in 1 / A (z) luqar to avoid clicks in synthesis. This filter breaks down into 3 steps.
Step 1: filtration by:
Α (ζ! Γ) 1 Á (z) 1-qz '<sup>1</sup> to map the delayed excitation by T in the TCX target domain;
Step 2: the application of a limiter (the magnitude is limited to ± rms<sub>wsyn</sub>)
Step 3: filtering by:
A (z! Χ) to find the synthesis. Note that the OVLP_TCX buffer is set to zero, in this case.
Decoding of algebraic VQ parameters.
In Case 2, 'TCX decoding involves decoding the algebraic VQ parameters that describe each quantized block B'<sub>k</sub> of the X 'scaled spectrum, where
<img file="MX356334B_D0015.tif" />
X 'is as described in Stage 2 of Section 5.3.5.7 of 3gpp TS 26.290. Remember that X 'has dimension N, where N = 288, 576 and 1152 for TCX-256, 512 and 1024, respectively, and that each block B'<sub>k</sub> it has dimension 8. The number K of blocks B '<sub>k</sub> it is therefore 36, 72 and 144 for TCX-256, 512 and 1024, respectively. The alqebraic VQ parameters for each block B '<sub>k</sub> Step 5 in Section 5.3.5.7 are described.
For each block B '<sub>k</sub> , three groups of binary indexes are sent by the encoder:
a) the codebook index n<sub>kr</sub> transmitted in unary code as described in Step 5 of Section
5.3.5.7;
b) the series of a selected grid point c in a so-called base codebook, indicating the permutation to be applied to a specific leader (see
<td>Step 5</td><td>of</td><td>the section</td><td> 5.3.5.7</td><td>) to get a</td><td>point</td><td>of</td>
<td>grating</td><td>c;</td><td></td><td></td><td></td><td></td><td></td>
<td>c)</td><td>and.</td><td colspan="3">if the quantized block # '* (a</td><td>point</td><td>of</td>
<td>grating)</td><td>not</td><td>was presented</td><td>at</td><td colspan="2">base codebook, the</td><td> 8</td>
<td>indices</td><td>of the</td><td>vector of</td><td>index</td><td>extension</td><td>Voronoi</td><td>Jt</td>
calculated in sub-step VI of Step 5 in Section; From Voronoi extension indices, a vector of z extension can be computed as in reference [1] of 3gpp TS
IMPI * 3 Mexican wawro V% _ '. · - </>!
Be la prupiídao \\ cj_ <sub>w</sub> * ·; · '' PP
26,290. The number of bits in each component '^ el'Vector of index k is provided by the order of extension ~ r ~,' which 'can be obtained from the unary code value of index n<sub>k</sub> . The scale factor M of the Voronoi extension is provided by M = 2<sup>r</sup>.
Then, from the scale factor M, the
Voronoi extension vector z (one grid point in REg} and grid point c in the base codebook (also, one grid point in REg), each quantized scaled block B '<sub>k</sub> can be computed as:
B '<sub>k</sub> = Me + z
When there is no Voronoi extension (i.e. n<sub>k</sub><5, M = 1 and z = 0), the base codebook is either the Qo, Qi, Q3 codebook or the 3gpp TS reference [1]
26,290. So no bits are required to transmit vector k. Otherwise, when using the extension
Voronoi because B '<sub>k</sub> is big enough then just Q<sub>3</sub> o Qh of reference [1] is used as a 20 base code book. Q selection<sub>3</sub> o is implicit in codebook index value n<sub>k</sub>, as described in Step 5 of Section 5.3.5.7.
<img file="MX356334B_D0016.tif" />
Estimation of the dominant height value.
The estimation of the dominant height is made from ”” moHd such that the next frame to be decoded can be properly extrapolated if it corresponds to TCX-256, and if the related packet is lost. This estimate is supported by the assumption that the peak of maximum magnitude in the spectrum of the TCX objective corresponds to the dominant height.
The search for the maximum M is restricted to a frequency lower than Fs / 64 kHz
M = maXi = i ..<sub>N</sub>/ 32 (X'2i) <sup>2</sup>+ (X '21 + 1)<sup>2</sup> and the minimum index 1 <í<sub>max</sub><N / 32 so that (X '<sub>2</sub>i) <sup>2</sup>+ (X'2i + i)<sup>2</sup> = M is also found. The dominant height is then estimated from the number of samples as r<sub>is</sub>t =
N / magnet (this value may not be an integer). Remember that the dominant height is calculated for packet drop concealment in TCX-256. In order to avoid temporary storage problems (the excitation buffer is limited to
256 samples), if T<sub>is</sub>t> 256 samples, height_tcx is set to 256; otherwise if T<sub>is</sub>t ^ 256, multiple height period is avoided in 256 samples by setting height_tcx in height tcx = max {LnT<sub>is</sub>tJ n integer> 0 and nT<sub>sst</sub><
256}
IMPI
MEXICAN INSTITUTE Ví — Ss-sS ^ S 2 *
OF THE PROPERTY :
INDUSTRIAL where LJ denotes rounding to the nearest integer toward
In the following, some additional conventional concepts will be briefly described.
In ISO_IEC_DIS_23003-3 (reference [3]), a TCX decoding employing MDCT [Modified Discrete Cosine Transform] is explained in the context of Unified Voice and Audio Codee.
In the state of the art of AAC [Advanced Audio Coding] (confer, for example, reference [4]), only one interpolation mode is described. According to reference [4], the AAC core decoder includes a hide function that increases the decoder delay by one frame.
In European Patent EP 1207519 Bl (reference [5]), the provision of a speech decoder and error compensation method capable of achieving further improvement for decoded speech in a frame in which an error is detected is described. According to the patent, a speech encoding parameter includes information so that it expresses features of each speech · short (frame) segment. The voice encoder adaptively calculates the delay parameters and the gain parameters used for the
<img file="MX356334B_D0017.tif" />
INSTITUTO MEXICANu DELA PROPIEDAD,. ,, INDUSTRIAL voice decoding according to the mode information. Furthermore, the voice decoder adaptively controls the adaptive drive gain ratio and the set drive gain according to the mode information. Furthermore, the concept according to the patent comprises adaptive control of adaptive excitation gain parameters and fixed excitation gain parameters used for speech decoding according to decoded gain parameter values in a normal decoding unit. in which no error is detected, immediately after a decoding unit whose encoded data is detected with an error.
In view of the prior art, there is a need to find a further improvement in error concealment, which provides a better auditory impression.
3. Synthesis of the invention.
An embodiment in accordance with the invention creates an audio decoder to provide decoded audio information based on encoded audio information. The audio decoder comprises an error concealment configured to provide error concealment audio information for the
<img file="MX356334B_D0018.tif" />
concealment of an audio frame loss (or more than one frame loss) after an audio frame encoded in a frequency domain representation, using a time domain drive signal.
This embodiment according to the invention is supported by the finding that improved error concealment can be obtained by providing the error concealment audio information on the basis of a time domain drive signal, even if the audio frame preceding a lost audio frame is encoded in a frequency domain representation. In other words, it has been recognized that a quality of an error concealment is typically better if the error concealment is performed on the basis of a time domain excitation signal, when compared to an error concealment performed in a domain. frequency so that it is worth switching to a time domain error concealment, using a time domain drive signal, even if the audio content preceding the missing audio frame is encoded in the frequency domain (ie, in a frequency domain representation). This is true, for example, for a monophonic signal and mostly for voice.
<img file="MX356334B_D0019.tif" />
IMPI <* srmjTQ Mexican bE u property, industrial invention allows to obtain if the audio frame is not
Therefore, the present good error concealment even precedes the missing audio frame is encoded in the frequency domain (ie, in a frequency domain representation).
In a preferred embodiment, the frequency domain representation comprises an encoded representation of a plurality of spectral values and an encoded representation of a plurality of factor factors.
<td>scale for</td><td>the scale</td><td>of</td><td>the values</td><td>spectral, or</td><td>the</td>
<td>decoder</td><td>audio</td><td>this</td><td>configured</td><td>to derive</td><td>a</td>
<td>plurality of</td><td>factors</td><td>of</td><td>scale for</td><td>the scale of</td><td>the</td>
spectral values from a coded representation of linear predictive coding (LPC) parameters. This could be done using FDNS (Frequency Domain Noise Form). However, it has been found that it is desirable to derive the time domain excitation signal (which can serve as an excitation for a linear predictive coding (LPC) synthesis) even if the audio frame preceding the missing audio frame is originally encoded in the frequency domain representation comprising substantially different information (i.e., an encoded representation of a plurality of values
<img file="MX356334B_D0020.tif" />
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY . ...
INDUSTRIAL spectral in a coded representation of a plurality of scale factors for the scale of the spectral values). For example, in the case of TCX, we do not send scale factors (from an encoder to a decoder), but linear predictive coding (LPC), and then, in the decoder, we transform linear predictive coding (LPC) into a representation factor factor for the bins of the Modified Discrete Cosine Transform (MDCT). In other words, in the case of TCX, we send the linear predictive coding coefficient (LPC), and then in the decoder we transform those linear predictive coding coefficients (LPC) into a scale factor representation for TCX in USAC or in AMRWB + where there is no scale factor.
In a preferred embodiment, the audio decoder comprises a frequency domain decoder core configured for scaling based on scale factors, to a plurality of values
<td>spectral</td><td>derivatives</td><td>of</td><td>the representation</td><td>of</td><td>domain</td><td>of</td>
<td>frequency.</td><td>In this</td><td>case,</td><td>concealment</td><td>of</td><td colspan="2">error is</td>
<td>configured</td><td colspan="2">to provide</td><td>information</td><td>of</td><td>Audio</td><td>of</td>
<td>concealment</td><td>of mistake</td><td>for</td><td>hiding</td><td>a</td><td>lost</td><td>of</td>
an audio frame after an audio frame encoded in
<img file="MX356334B_D0021.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL ___ the frequency domain representation comprising a plurality of scale factors encoded using a time domain excitation signal derived from the frequency domain representation. This embodiment according to the invention is supported by the finding that the derivation of the time domain excitation signal from the above mentioned frequency domain representation typically provides a better error concealment result compared to an error concealment performed directly in the frequency domain. For example, the excitation signal is created based on the synthesis of the previous frame; so it doesn't really matter if the previous frame is a frequency domain frame (MDCT (Modified Discrete Cosine Transform), FFT (Fast Fourier Transform ...) or a time domain frame However, particular advantages can be seen if the previous frame was a frequency domain. Furthermore, it should be noted that particularly good results are achieved, for example, for monaural signal such as voice. As another example, the scale factors could be transmitted as linear predictive coding coefficients (LPC), for example, using
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX356334B_D0022.tif" />
a polynomial representation that is then converted to scale factors on the decoder side.
In a preferred embodiment, the audio decoder comprises a frequency domain decoder core configured to derive a time domain audio signal representation from the frequency domain representation without the use of an excitation signal time domain as an intermediate quantity for the audio frame encoded in the frequency domain representation. In other words, the use of a time domain drive signal for error concealment has been found to be convenient even if the audio frame preceding the missing audio frame is encoded in a real frequency mode that does not use no time domain excitation signal as an intermediate quantity (and, consequently, not supported by a linear predictive coding synthesis (LPC)).
In a preferred embodiment, error concealment is configured to obtain the time domain drive signal based on the audio frame encoded in the frequency domain representation preceding a lost audio frame. In this case, the
INSTITUTO MEXICANO MI y *
OF THE PROPERTY
INDUSTRIAL error concealment is configured — prnvppr, the error concealment audio information for hiding the lost audio frame using said time domain drive signal. In other words, it has been recognized that the time domain drive signal, which is used for error concealment, should derive from the audio frame encoded in the frequency domain representation preceding the missing audio frame, since this time domain driving signal derived from the audio frame encoded in the frequency domain representation preceding the missing audio frame provides a good representation of an audio content of the audio frame preceding the lost frame. lost audio, so that error concealment can be performed with moderate effort and good accuracy.
In a preferred embodiment, error concealment is configured to perform predictive coding analysis. linear (LPC) based on the audio frame encoded in the frequency domain representation preceding the missing audio frame, in order to obtain a set of linear prediction encoding parameters and the time domain drive signal representing an audio content of the plot of
<img file="MX356334B_D0023.tif" />
IMPI iHsmvTTj μβικανο OF THE PROFIF.UAIt audio encoded in the representation of <sup>mt</sup>What is the frequency preceding the lost audio frame? weather, even if the audio frame preceding the missing audio frame is encoded in a frequency domain representation (which does not contain any linear prediction encoding parameters and no representation of a time domain drive signal), because good quality error concealment audio information can be obtained for many input audio signals on the basis of said time domain driving signal. Alternatively, error concealment can be configured to perform linear predictive coding (LPC) analysis based on the audio frame encoded in the frequency domain representation preceding the missing audio frame, in order to obtain the signal time domain drive representing an audio content of the audio frame encoded in the frequency domain representation preceding the missing audio frame. Also, alternatively, the decoder
<img file="MX356334B_D0024.tif" />
Audio IMPI may be configured to obtain a set of linear prediction encoding parameters using an estimate of linear prediction encoding parameters, or the audio decoder may be configured to obtain a set of linear prediction encoding parameters based on of a set of scale factors using a transform. In other words, linear predictive coding (LPC) parameters can be obtained using linear predictive coding (LPC) parameter estimation. This could be done either by windows / autocorr / levinson durbin based on the audio frame encoded in the frequency domain representation or by transforming from the scale factor prior directly to the linear predictive encoding representation (LPC ).
In a preferred embodiment, error concealment is configured to obtain height (or delay) information describing a height of the audio frame encoded in the frequency domain preceding the missing audio frame, and to provide the Audio hiding error information according to the height information. By considering height information, error concealment audio information can be achieved
<img file="MX356334B_D0025.tif" />
(which is typically a blackout audio signal of - N · -, -, - T, --- ** 'error covering the time duration of at least one missing audio frame) suits the audio content well real.
In a preferred embodiment, error concealment is configured to obtain height information based on the time domain drive signal derived from the encoded audio frame in the frequency domain representation preceding the frame of lost audio. It has been found that a derivation of the height information from the time domain excitation signal carries high accuracy. Furthermore, it has been found that it is convenient if the height information is well suited to the time domain drive signal, since the height information is used for a modification of the time domain drive signal. By deriving the height information from the time domain excitation signal, such a close relationship can be achieved.
In a preferred embodiment, error concealment is configured to evaluate a cross correlation of the time domain excitation signal to determine approximate height information. Furthermore, error concealment can be configured to refine
<img file="MX356334B_D0026.tif" />
the approximate height information or closed loop search around a height determined by the approximate height information. Consequently, highly accurate height information can be achieved with moderate computational effort.
In a preferred embodiment, the audio decoder error concealment may be configured to obtain height information based on lateral information of the encoded audio information.
In a preferred embodiment, error concealment may be configured to obtain height information based on available height information for a previously decoded audio frame.
In a preferred embodiment, error concealment is configured to obtain height information based on a height search performed on a time domain signal or a residual signal.
In other words, the height can be transmitted as lateral information or it could also come from the previous frame if there is, for example, LTP. The height information could also be transmitted in the bit stream if it is available in the encoder. It could optionally be done
<img file="MX356334B_D0027.tif" />
IMPI search for height above the ~ ™ domiñio 'sign of<sup>-</sup>Time directly, or on the residual, which usually provides better results on the residual (time domain excitation signal).
In a preferred embodiment, the error concealment is configured to copy a high cycle of the time domain drive signal derived from the encoded audio frame in the frequency domain representation preceding the missing audio frame by one once or multiple times, in order to obtain an excitation signal for an error concealment audio signal synthesis. By copying the time domain excitation signal once or multiple times, the deterministic (i.e. substantially periodic) component of the audio error concealment information can be achieved with good accuracy, and is a good continuation of the deterministic (eg, substantially periodic) component of the audio content of the audio frame preceding the missing audio frame.
In a preferred embodiment, the error concealment is configured to filter in low pass the height cycle of the time domain drive signal derived from the frequency domain representation of the
IMPI
<img file="MX356334B_D0028.tif" />
audio frame encoded in the frequency domain representation preceding the lost audio frame using a sample rate dependent filter, the bandwidth of which depends on a sampling rate of the audio frame encoded in a domain representation of frequency. Accordingly, the time domain driving signal can be adapted for an available audio bandwidth, which produces a good auditory impression of the error concealment audio information. For example, the low pass is preferred only over the first lost frame, and preferably, furthermore, the low pass only if the signal is not 100% stable. However, it should be noted that the optional low-pass filtration ·, and can be performed only on the first height cycle. For example, the filter may depend on the sampling rate, so that the cutoff frequency is independent of the bandwidth.
In a preferred embodiment, the error concealment is configured to predict a height at one end of a lost frame in order to adapt the time domain drive signal, or one or more of its copies, to the predicted height . Consequently, the expected height changes during the lost audio frame can be considered. Consequently, failures in a
IMPI
MEXICAN INSTITUTE 'f¿ * OF PROPERTY
<img file="MX356334B_D0029.tif" />
transition between audio information from ocWE ^ iemno error and audio information from one frame to
<img file="MX356334B_D0030.tif" />
decoded after one or more missing audio frames (or at least reduced, since it is only a predicted frame, not the real one). For example, adaptation ranges from last good height to predicted height. This is done by means of pulse resynchronization [7].
In a preferred embodiment, error concealment is configured to combine an extrapolated time domain excitation signal and a noise signal, to obtain an input signal for linear predictive coding (LPC) synthesis. In this case, error concealment is configured to perform linear predictive coding synthesis (LPC), where linear predictive coding synthesis (LPC) is configured to filter the input signal from linear predictive coding synthesis (LPC) according to linear prediction coding parameters, in order to obtain the error hiding audio information. Consequently, both a deterministic (eg, approximately periodic) component of the audio content and a noise-like component of the audio content can be considered. Therefore, it is achieved that the information of
IMPI
<img file="MX356334B_D0031.tif" />
error concealment audio comprise a natural auditory impression.
In a preferred embodiment, the error concealment is configured to compute an extrapolated time domain excitation signal gain, which is used to obtain the input signal for linear predictive coding (LPC) synthesis, using a correlation in the time domain that is performed based on a time domain representation of the coded audio frame in the frequency domain that precedes the missing audio frame, where a correlation delay dependent on a height information obtained based on the time domain excitation signal is established.
In other words, an intensity of a periodic component is determined within the audio frame preceding the missing audio frame, and this determined intensity of the periodic component is used to obtain the error concealment audio information. However, the aforementioned computation of the intensity of the periodic component has been found to provide particularly good results, since the real time domain audio signal of the audio frame preceding the missing audio frame is considered. Alternatively, a
IMPI
<img file="MX356334B_D0032.tif" />
correlation in the excitation domain or directly in the time domain in order to obtain the height information. However, there are also different possibilities, depending on the embodiment used. In one embodiment, the height information could be just the height obtained from the last frame ltp, or the height that is transmitted as lateral or calculated information.
In a preferred embodiment, error concealment is configured for the high-pass filter of the noise signal that is combined with the extrapolated time domain drive signal. High-pass filtering of the noise signal (which is typically input into linear predictive coding synthesis (LPC)) has been found to achieve a natural auditory impression. For example, the high pass characteristic may change with the amount of frame lost, after a certain amount of frame loss there can no longer be high pass. The high pass characteristic may also depend on the sampling rate at which the decoder is running. For example, the high pass depends on the sampling rate, and the filter characteristic can change as a function of time (on consecutive frame loss). The high pass feature can also optionally change over consecutive frame loss, from
<img file="MX356334B_D0033.tif" />
IMPI mode such that after a certain amount of frame loss, there is no longer any filtering, just to get the full band shape noise so as to get a good comfort noise close to the background noise.
In a preferred embodiment, error concealment is configured to selectively change the spectral shape of the noise signal (562) using the pre-emphasis filter where the noise signal is combined with the extrapolated time domain drive signal if the audio frame encoded in a frequency domain representation that precedes the missing audio frame is a speech audio frame or comprises a start. It has been found that the auditory impression of the error concealment audio information can be improved by this concept. For example, in some cases, it is better to decrease the profits and the form, and somewhere, it is better to increase them.
In a preferred embodiment, error concealment is configured to compute a gain of the noise signal according to a correlation in the time domain, which is performed based on a time domain representation of the audio encoded in the frequency domain representation that precedes the missing audio frame. Such determination of the
IMPI
<img file="MX356334B_D0034.tif" />
noise signal gain ~ p * ^ * e ^ «__ ^<sub>></sub>gs_ultados particularly accurate, since the real-time domain audio signal associated with the audio frame preceding the missing audio frame can be considered. Using this concept, it is possible to obtain a hidden plot energy close to the previous good plot energy. For example, the gain for the noise signal can be generated by measuring the energy of the result: input signal drive - drive based on generated height.
In a preferred embodiment, error concealment is configured to modify a time domain drive signal obtained based on one or more audio frames preceding a missing audio frame, in order to obtain the audio information of error concealment. It has been found that modification of the time domain drive signal enables adaptation of the time domain drive signal to a desired time evolution. For example, modification of the time domain excitation signal allows the outgoing fading of the deterministic (eg, substantially periodic) component of the audio content into the error concealment audio information. Furthermore, the
<img file="MX356334B_D0035.tif" />
Modification of the time domain excitation signal also allows adapting the time domain excitation signal to a height variation (estimated or expected). This allows adjustment of the characteristics of the audio error concealment information as a function of time.
In a preferred embodiment, error concealment is configured to use one or more modified copies of the time domain drive signal obtained on the basis of one or more audio frames preceding a missing audio frame, in order to get the error concealment information. Modified copies of the time domain excitation signal can be obtained with moderate effort, and the modification can be made using a simple algorithm. Accordingly, the desired characteristics of the error concealment audio information can be achieved with moderate effort.
In a preferred embodiment, error concealment is configured to modify the obtained time domain drive signal based on one or more audio frames preceding a missing audio frame, or one or more of its copies, in order to reduce a periodic component of the error hiding audio information as a function of time. In consecuense,
IMPÍ
<img file="MX356334B_D0036.tif" />
the correlation between the audio content of the audio frame preceding the missing audio frame and the audio content of one or more missing audio frames can be considered to decrease as a function of time. Furthermore, causing an unnatural auditory impression can be avoided by long preservation of a periodic component of the error concealment audio information.
In a preferred embodiment, error concealment is configured to scale the obtained time domain drive signal based on one or more audio frames preceding the missing audio frame, or one or more of its copies, in order to modify the time domain excitation signal. It has been found that the scaling operation can be performed with little effort, where the scaled time domain drive signal typically provides good error concealment audio information.
In a preferred embodiment, the error concealment is configured to gradually reduce a gain applied to scale the time domain drive signal obtained on the basis of one or more audio frames preceding a missing audio frame, or a or more of your copies. Therefore, a
<img file="MX356334B_D0037.tif" />
IMPI
OF THE C 'INDUSTRIAL PROPERTY>
component fade out — oh j not within error concealment audio information.
In a preferred embodiment, error concealment is configured to adjust a rate used to gradually reduce a gain applied to scale the time domain drive signal obtained based on one or more audio frames preceding a frame of lost audio, or one or more of its copies, according to one or more parameters of one or more audio frames that precede the lost audio frame, and / or according to a number of consecutive lost audio frames. Accordingly, it is possible to adjust the rate at which the deterministic component (eg, at least approximately periodic) is fade out in the error concealment audio information. The outgoing fade rate can be tailored to specific characteristics of the audio content, which can typically be observed from one or more parameters of one or more audio frames preceding the missing audio frame. Alternatively, or in addition, the number of consecutive missing audio frames can be considered when determining the rate used for the outgoing fade of the deterministic component (by
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX356334B_D0038.tif" />
example, at least approximately periodic) of the error concealment audio information, which helps to tailor the error concealment to the specific situation. For example, the gain of the tonal part and the gain of the noisy part may fade out separately. The gain for the tonal part may converge to zero after a certain amount of frame loss, while the noise gain may converge to the determined gain to achieve a certain comfort noise.
In a preferred embodiment, error concealment is configured to adjust the rate used to gradually reduce an applied gain to scale the time domain drive signal obtained based on one or more audio frames preceding a frame of lost audio, or one or more of its copies, according to a length of a period of time domain excitation signal height, such that a time domain excitation signal input in a linear predictive coding synthesis (LPC) is faded out faster for signals that have a shorter height period length compared to signals that have a longer length of the height period.
<img file="MX356334B_D0039.tif" />
Accordingly, signals having a shorter height period length can be prevented from being repeated too often at high intensity, as this will typically result in unnatural auditory impression. Consequently, an overall quality of the error concealment audio information can be improved.
In a preferred embodiment, error concealment is configured to adjust the rate used to gradually reduce an applied gain to scale the time domain drive signal obtained based on one or more audio frames preceding a frame of lost audio, or one or more of its copies, based on a result of a height analysis or a height prediction, such that a deterministic component of the time domain excitation signal input in a linear predictive coding (LPC) synthesis fades out more rapidly for signals that have a greater change in height per unit time compared with signals that have a smaller height change per unit time, and / or such that a deterministic component of the time domain excitation signal input in a linear predictive coding (LPC) synthesis fades out faster for signals for the
IMPI
<img file="MX356334B_D0040.tif" />
which a height prediction fails compared to signals for which the height prediction is successful. Consequently, the outward fading can be done more quickly for signals in which there is a high height uncertainty compared to signals for which there is a lower height uncertainty. However, by outgoing fade of a deterministic component more quickly for signals that comprise a comparatively large height uncertainty, audible failures can be avoided, or at least substantially reduced.
In a preferred embodiment, the error concealment is configured for the time domain excitation signal time scale obtained on the basis of one or more audio frames preceding a missing audio frame, or one or more of their copies, based on a predicted height-for-time of one or more missing audio frames. Accordingly, the time domain excitation signal can be adapted to a variable height, such that the error concealment audio information comprises a more natural auditory impression.
<img file="MX356334B_D0041.tif" />
IMPI
In a preferred embodiment, the error concealment is configured to provide the error concealment audio information for a time that is greater than a time duration of one or more lost audio frames. Therefore, it is possible to perform an overlay and add operation based on the error concealment audio information, which helps to reduce crash failures.
In a preferred embodiment, the error concealment is configured to superimpose and add the error concealment audio information and a time domain representation of one or more appropriately received audio frames after one or more lost audio frames. Consequently, it is possible to avoid (or at least reduce) blocking failures.
In a preferred embodiment, the error concealment is configured to derive the error concealment audio information based on at least three partially overlapping windows or frames preceding a missing audio frame or a lost window. Consequently, error concealment audio information can be obtained with good accuracy, even for modes
<img file="MX356334B_D0042.tif" />
IMPI encoding in which more than two frames (or windows) are overlapped (where such overlap can help reduce delay).
Another embodiment in accordance with the invention creates a method of providing decoded audio information on the basis of audio information.
<td>encoded.</td><td>The</td><td>method comprises</td><td>the</td><td>provision</td><td>of</td><td>a</td>
<td>information</td><td>of</td><td colspan="2">blackout audio</td><td>of mistake</td><td>for</td><td>the</td>
<td>concealment</td><td>of</td><td>a loss of a</td><td>plot</td><td>audio</td><td>then</td><td>of</td>
<td>a plot</td><td>of</td><td>audio encoded in</td><td>a</td><td colspan="2">representation</td><td>of</td>
<td>domain of</td><td colspan="2">frequency using a</td><td>signal</td><td colspan="2">of excitement</td><td>of</td>
time domain. This method is based on the same considerations as the aforementioned audio decoder.
Still another embodiment according to the invention creates a computer program for carrying out said method when the computer program is run on a computer.
Another embodiment in accordance with the invention creates an audio decoder to provide decoded audio information based on encoded audio information. The audio decoder comprises an error concealment configured to provide a
<img file="MX356334B_D0043.tif" />
IMPI
MEXICAN INSTITUTE W LA «OnttiAD INDUSTRIA!
<td>information</td><td>of</td><td>Audio</td><td>concealment</td><td>of</td><td>error</td><td>for</td><td>the</td>
<td>concealment</td><td>of</td><td>a</td><td>loss of a</td><td>plot</td><td>of</td><td>Audio.</td><td>The</td>
<td>concealment</td><td>of</td><td>error</td><td>is configured</td><td>for</td><td colspan="2">Modify</td><td>a</td>
time domain drive signal obtained on the basis of one or more audio frames preceding a lost audio frame, in order to obtain the error concealment audio information.
This embodiment according to the invention is based on the idea that error concealment with good audio quality can be obtained on the basis of a time domain drive signal, where a modification of the time domain excitation signal obtained on the basis of one or more audio frames preceding a missing audio frame allows an adaptation of the error concealment audio information to expected (or predicted) changes of the audio content during the lost frame. Accordingly, failures, and in particular, unnatural auditory impression, which would be caused by unchanged use of the time domain excitation signal can be avoided. Accordingly, an improved provision of error concealment audio information is achieved such that lost audio frames can be concealed with improved results.
<img file="MX356334B_D0044.tif" />
IMPI
In a preferred embodiment, the error concealment is configured to use one or more modified copies of the obtained time domain excitation signal for one or more audio frames preceding a lost audio frame, in order to obtain the error concealment information. By using one or more modified copies of the obtained time domain excitation signal for one or more audio frames preceding a missing audio frame, good quality of error concealment audio information can be achieved with little effort computational.
In a preferred embodiment, error concealment is configured to modify the obtained time domain drive signal for one or more audio frames preceding a missing audio frame, or one or more of its copies, in order to reduce a periodic component of the audio masking error information as a function of time. By reducing the periodic component of the error hiding audio information as a function of time, artificially long preservation of a deterministic sound (eg roughly periodic) can be avoided, helping to make the sound of the audio information natural of error concealment.
IMPI
<img file="MX356334B_D0045.tif" />
In a preferred embodiment, the error nniltamianha— is configured to scale the time domain drive signal obtained on the basis of one or more audio frames preceding the missing audio frame, or one or more of its copies , in order to modify the time domain excitation signal. Scaling the time domain drive signal is a particularly efficient way to vary the error concealment audio information as a function of time.
In a preferred embodiment, error concealment is configured to gradually reduce a gain applied to scale the time domain drive signal obtained for one or more audio frames preceding a missing audio frame, or one or more of your copies. It has been found that the gradual reduction of the gain applied to scale the time domain excitation signal obtained for one or more audio frames preceding a missing audio frame, or one or more of its copies, allows obtaining a signal of time domain excitation for the provision of error concealment audio information such that the deterministic components (eg, components at least approximately periodic) are faded out. For example, there may not be
<img file="MX356334B_D0046.tif" />
just a profit. For example, you could have a gain for the tonal part (also referred to as the roughly periodic part), and a gain for the noise part. Both excitations (or excitation components) can be attenuated separately with different velocity factor, and then the resulting two excitations (or excitation components) can be combined before being fed into linear predictive coding (LPC) for synthesis. In the case of not having any estimate of background noise, the outgoing fading factors for the noise and for the tonal part may be similar, and so, you could have only one outgoing fading application on the results of the two excitations, multiplied with their own profit and combined with each other.
Therefore, the error concealment audio information can be prevented from comprising a temporarily extended deterministic (eg, at least approximately periodic) audio component, which would typically provide an unnatural auditory impression.
In a preferred embodiment, error concealment is configured to adjust a rate used to gradually reduce a gain applied to scale the time domain drive signal.
IMPI
<img file="MX356334B_D0047.tif" />
obtained for one or more audio frames..qne -.-, f (X £££ deii ·. a lost audio frame, or one or more of its copies, according to one or more parameters of one or more frames audio frames preceding the missing audio frame, and / or according to a number of consecutive missing audio frames. Therefore, the outgoing fade rate of the deterministic component (eg, at least approximately periodic) in the error concealment audio information can be tailored to the specific situation, with moderate computational effort. Because the time domain drive signal used for the provision of the error concealment audio information is typically a scaled version (scaled using the gain mentioned above) of the time domain drive signal obtained for a or more audio frames preceding the missing audio frame, A variation of said gain (used to derive the time domain driving signal for the provision of the error concealment audio information) is a simple, yet effective method of tailoring the error concealment audio information to the needs specific. However, the speed of the outgoing fade is also controllable with very little effort.
<img file="MX356334B_D0048.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX356334B_D0049.tif" />
In a preferred embodiment, "the error concealment is configured to adjust the rate used to gradually reduce a gain applied to scale the time domain drive signal obtained on the basis of one or more audio frames preceding a frame of lost audio, or one or more of its copies, according to a length of a period of time domain excitation signal height, such that a time domain excitation signal input in a linear predictive coding (LPC) synthesis is faded out faster for signals that have a shorter height period length compared to signals that have a longest length of the height period. Consequently, fade-out is performed faster for signals that have a shorter height period length, preventing a height period from being copied too many times (which would typically achieve an unnatural auditory impression) .
In a preferred embodiment, error concealment is configured to adjust the rate used to gradually reduce a gain applied to scale the time domain drive signal obtained for one or more audio frames preceding an audio frame.
<img file="MX356334B_D0050.tif" />
s
1.
Lost ívlPI, or one or more of your copies, give <sup>1</sup> according to a result of a height analysis or a height prediction, such that a deterministic component of a time domain excitation signal input in a linear predictive coding (LPC) synthesis is faded out more quickly for signals that have a greater change in height per unit time, compared to signals that have a smaller change in height per unit time, and / or such that a deterministic component of a time domain excitation signal input in a linear predictive coding (LPC) synthesis is faded out more rapidly for signals for which a comparison height prediction fails with signals for which height prediction is successful. Consequently, a deterministic component (eg, at least roughly periodic) is faded out more quickly for signals for which there is greater height uncertainty (where a greater change in height per unit time, or even, a failure of the height prediction indicates a comparatively large height uncertainty). Consequently, failures, which would arise from the provision of audio information from
IMPI
<img file="MX356334B_D0051.tif" />
highly deterministic error concealment pp nr. =. vituauión in which the actual height is uncertain.
In a preferred embodiment, error concealment is configured for the time domain excitation signal time scale obtained for (or based on) one or more audio frames preceding a missing audio frame, or one or more of its copies, according to a prediction of a height for the time of the one or more missing audio frames. Accordingly, the time domain excitation signal, which is used for the provision of the error concealment audio information, is modified (compared to the time domain excitation signal obtained for (or on the basis of ) one or more audio frames preceding a lost audio frame, such that the height of the time domain drive signal follows the requirements of a time period of the lost audio frame. Accordingly, the auditory impression can be improved, which can be achieved by the error concealment audio information.
In a preferred embodiment, the error concealment is configured to obtain a time domain drive signal, which has been used for decoding one or more preceding audio frames
<img file="MX356334B_D0052.tif" />
INSTITUTO MEXICANO OE LA PROPIEDAD the lost audio plot, and for the mcdiliuuciúu — dU '<sup>,</sup>d'ÍTTra 'time domain drive signal, which has been used for decoding one or more audio frames preceding the missing audio frame, in order to obtain a modified time domain drive signal. In this case, the time domain concealment is configured to provide the error concealment audio information based on the modified time domain audio signal. Accordingly, it is possible to reuse a time domain drive signal, which has already been used to decode one or more audio frames preceding the missing audio frame. Consequently, very little computational effort can be maintained if the time domain drive signal has already been acquired for decoding one or more audio frames preceding the missing audio frame.
In a preferred embodiment, error concealment is configured to obtain height information, which has been used for decoding one or more audio frames preceding the missing audio frame. In this case, the error concealment is further configured to provide the error concealment audio information according to that information.
IMFl
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL -------- height. Consequently, the previously used height information can be reused, which avoids a computational effort for a new computation of the height information. Therefore, error concealment is particularly computationally efficient. For example, in the case of ACELP, we have 4 height delays and frame gains. We can use the last two frames to be able to predict the height at the end of the frame that we have to hide.
Next, we compare with the previously described frequency domain codec where only one or two heights are derived per frame (we can have more than two, although this would add a lot of complexity for a not very large gain in quality). In the case of a switching codec that is, for example, ACELP - FD - loss, then we have much better height accuracy since the height is transmitted in the bitstream and based on the original input signal (not in the decoded one, as it is done in the decoder). In the case of high bit rate, for example, we can also send a height and gain delay information, or LTP information, per frequency domain encoded frame.
IMPI
<img file="MX356334B_D0053.tif" />
In a preferred embodiment, the error concealment of the audio decoder may be configured to obtain height information on the basis of lateral information of encoded audio information.
In a preferred embodiment, error concealment may be configured to obtain height information based on available height information for a previously decoded audio frame.
In a preferred embodiment, error concealment is configured to obtain height information based on a height search performed on a time domain signal or a residual signal.
In other words, the height can be transmitted as lateral information or it could also come from the previous frame if there is LTP, for example. The height information could also be transmitted in the bit stream if it is available in the encoder. We can optionally do the height search in the time domain signal directly or in the residual, which usually provides better results on the residual (time domain excitation signal).
IMPI
<img file="MX356334B_D0054.tif" />
In a preferred embodiment, error concealment is configured to obtain a set of linear prediction coefficients, which have been used to decode one or more audio frames preceding the missing audio frame. In this case, the error concealment is configured to provide the error concealment audio information according to said set of linear prediction coefficients. Consequently, the efficiency of error concealment is increased by reusing previously generated (or previously decoded) information, for example, the previously used set of linear prediction coefficients.
Consequently, unnecessary high computational complexity is avoided.
In a preferred embodiment, error concealment is configured to extrapolate a new set of linear prediction coefficients based on the set of linear prediction coefficients, which have been used to decode one or more audio frames preceding the frame. lost audio. In this case, error concealment is configured to use the new set of linear prediction coefficients to provide the error concealment information. When deriving
IMPI
<img file="MX356334B_D0055.tif" />
the new set of · ρΐ ·· ι · gü-ociún. — üaoAir coefficients used to provide audio error concealment information, from a set of linear prediction coefficients previously used using extrapolation, a full recalculation can be avoided of linear prediction coefficients, which helps keep computational effort reasonably low. Furthermore, by extrapolating based on the previously used set of linear prediction coefficients, it can be guaranteed that the new set of linear prediction coefficients is at least similar to the previously used set of linear prediction coefficients, which helps to avoid discontinuities when providing error concealment information. For example, after a certain amount of frame loss, we tend to estimate the shape of linear noise linear predictive coding (LPC). The speed of this convergence, for example, may depend on the signal characteristic.
In a preferred embodiment, error concealment is configured to obtain information about an intensity of a deterministic signal component in one or more audio frames preceding a missing audio frame. In this case the error concealment is
IMPI
INSTITUTO nexican; OF THE PROPERTY
INDUSTRIAL
<img file="MX356334B_D0056.tif" />
configured to compare information<sup>1</sup>Lower-intensity of a deterministic signal component in one or more audio frames preceding a missing audio frame with a threshold value, in order to decide whether to input a deterministic component of a time domain drive signal in a linear predictive coding (LPC) synthesis (synthesis based on the linear prediction coefficient), or whether to enter only a noise component of a time domain excitation signal in linear predictive coding (LPC) synthesis. Accordingly, it is possible to omit the provision of a deterministic (eg, at least approximately periodic) component of the error concealment audio information in the event that there is only a small contribution of deterministic signal within one or more frames than they precede the missing audio frame. This has been found to help obtain a good hearing impression.
In a preferred embodiment, the error concealment is configured to obtain height information describing a height of the audio frame preceding the missing audio frame, and to provide the error concealment audio information in accordance with the height information. Therefore, it is possible to adapt
IMPI
<img file="MX356334B_D0057.tif" />
the height of the hidden information Gl'i'lU ύΰ — ΈΓίΐυι · - · »—height of the audio frame preceding the missing audio frame. Accordingly, discontinuities are avoided, and a natural auditory impression can be achieved.
In a preferred embodiment, error concealment is configured to obtain height information based on the time domain drive signal associated with the audio frame preceding the missing audio frame. The height information obtained on the basis of the time domain excitation signal has been found to be particularly reliable, and furthermore very well suited to the processing of the time domain excitation signal.
In a preferred embodiment, error concealment is configured to evaluate a cross correlation of the time domain drive signal (or, alternatively, a time domain audio signal), in order to determine a approximate height, and refine the approximate height information using a closed-loop search around a height determined (or described) by the approximate height information. It has been found that this concept allows obtaining very precise height information with moderate effort. <sup>57</sup> IMPI ^
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL computational. In other words, in some codees, we do the height search directly on the time domain signal, while in some others we do the height search on the time domain excitation signal.
In a preferred embodiment, the error concealment is configured to obtain the height information for the provision of the error concealment audio information on the basis of previously computed height information, which was used for decoding a or more audio frames preceding the missing audio frame, and based on an evaluation of a cross-correlation of the time domain excitation signal, which is modified to obtain a modified time domain drive signal for the provision of the error concealment audio information. Consideration of both previously computed height information and height information obtained on the basis of the time domain excitation signal (using cross correlation) has been found to improve the reliability of the height information, and in Consequently, it helps to avoid failures and / or discontinuities.
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX356334B_D0058.tif" />
In a preferred embodiment, the orjiliamianifl. error is configured to select a peak of the cross correlation, from a plurality of peaks of the cross correlation, such as a peak representing a height according to the previously computed height information, such that a peak representing a height that is closest to the height represented by the previously computed height information. Therefore, possible cross-correlation ambiguities can be overcome, which, for example, can produce multiple peaks. The previously computed height information is thus used to select the appropriate peak of the cross correlation, which helps to substantially increase reliability. On the other hand, the real-time domain excitation signal is primarily considered for height determination, which provides good accuracy (which is substantially better than an accuracy obtained based on only previously computed height information. ).
In a preferred embodiment, the audio decoder error concealment may be configured to obtain height information based on lateral information of the encoded audio information.
<img file="MX356334B_D0059.tif" />
IMPI
MEXICAN INSTITUTE Of INDUSTRIAL property
In a preferred embodiment '/ -ol ide error may be configured to obtain height information based on available height information for a previously decoded audio frame.
In a preferred embodiment, error concealment is configured to obtain information about
<td>height above</td><td>base of</td><td>a</td><td>search</td><td>of</td><td>height realized</td>
<td>on a sign</td><td colspan="2">Of domain</td><td>of time</td><td> 0</td><td>on a sign</td>
<td>residual.</td><td></td><td></td><td></td><td></td><td></td>
<td colspan="2">In other words, the</td><td colspan="2">height can</td><td>to be</td><td>transmitted as</td>
Lateral information, or it could also come from the previous frame, if there is LTP, for example. The height information could also be transmitted in the bit stream if it is available in the encoder. We can optionally do the height search on the time domain signal directly, or on the residual, which usually provides better results on the residual (time domain excitation signal).
In a preferred embodiment, error concealment is configured to copy one cycle high of the time domain drive signal associated with the audio frame preceding the lost audio frame once or multiple times, in order to get an excitation signal
<img file="MX356334B_D0060.tif" />
<img file="MX356334B_D0061.tif" />
IMPI
INSTITUTO MEXICANO DE LA PROPÍKÜAD INDUSTRIAL (or at least one of its deterministic components) for a synthesis of the audio information of error concealment. By copying the height cycle of the time domain drive signal associated with the audio frame preceding the missing audio frame once or multiple times, and modifying those one or more copies using a comparatively simple modification algorithm, the excitation signal (or at least its deterministic components) for the synthesis of the error concealment audio information can be obtained with little computational effort. However, reuse of the time domain drive signal associated with the audio frame preceding the missing audio frame (by copying that time domain drive signal) avoids audible discontinuities.
In a preferred embodiment, error concealment is configured for the low-pass filter of the signal height cycle. of domain time excitation associated with the audio frame preceding the missing audio frame using a sample rate dependent filter, the bandwidth of which depends on a sample rate of the audio frame encoded in a domain representation of frequency. Therefore, the excitation signal of
IMPI
<img file="MX356334B_D0062.tif" />
Time domain adapts to a signal bandwidth of the audio decoder, resulting in good playback of the audio content. For details and optional improvements, reference is made, for example, to the above explanations.
For example, low pass is preferred for only the first
<td>lost plot,</td><td colspan="3">and preferably,</td><td colspan="3">in addition, we do the step</td>
<td>low alone if</td><td>the signal</td><td>It is not</td><td>without</td><td>voice. Without</td><td colspan="2">However, you should</td>
<td>be noted that</td><td colspan="2">the filtration</td><td>of</td><td>low pass</td><td>is</td><td>optional.</td>
<td>Further,</td><td>the filter</td><td>can</td><td>to be</td><td>dependent</td><td>of</td><td>the rate of</td>
<td>sampling of</td><td>so</td><td>than</td><td>the</td><td>frequency</td><td>of</td><td>court is</td>
independent of bandwidth.
In a preferred embodiment, error concealment is configured to predict a height at one end of a lost frame. In this case, the error concealment is configured to adapt the time domain excitation signal, or one or more of its copies, to the predicted height. By modifying the time domain excitation signal, such that the time domain excitation signal that is actually used for the provision of the error concealment audio information is modified with respect to the domain excitation signal of time associated with an audio frame preceding the audio frame
<img file="MX356334B_D0063.tif" />
so that error adapts to evolution
By elemolo, the
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL lost, expected (or predicted) changes in height during the lost audio frame, concealment audio information from either the actual (or at least expected or predicted) evolution of the audio content .
adaptation ranges from the last good height to that predicted. This is done by means of pulse resynchronization [7].
In a preferred embodiment, error concealment is configured to combine an extrapolated time domain excitation signal and a noise signal, to obtain an input signal for linear predictive coding (LPC) synthesis. In this case, error concealment is configured to perform linear predictive coding synthesis (LPC), where linear predictive coding synthesis (LPC) is configured to filter the input signal from linear predictive coding synthesis (LPC) according to linear prediction coding parameters, in order to obtain the error hiding audio information. By combining the extrapolated time domain excitation signal (which is typically a modified version of the derived time domain excitation signal for one or
<img file="MX356334B_D0064.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY plus audio frames preceding the missing audio frame) and a noise signal, can be considered both deterministic components (eg approximately periodic) and noise components of audio content, in error concealment . Therefore, the error concealment audio information can be achieved to provide an auditory impression that is similar to the auditory impression provided by the frames preceding the missing frame.
Furthermore, by combining a time domain excitation signal and a noise signal, in order to obtain the input signal for linear predictive coding (LPC) synthesis (which can be considered a combined time domain excitation signal) , it is possible to vary a percentage of the deterministic component of the input audio signal for linear predictive coding (LPC) synthesis, while maintaining an energy (of the input signal of the linear predictive coding synthesis (LPC), or even, of the output signal of the linear predictive coding synthesis (LPC)). Accordingly, it is possible to vary the characteristics of the error concealment audio information (eg, the tonality characteristics), without substantially changing an energy or volume of the error concealment audio signal, such that
IMPI
<img file="MX356334B_D0065.tif" />
that it is possible to modify the excitation signal in time domain without causing unacceptable audible distortions.
An embodiment according to the invention creates a method for providing decoded audio information on the basis of encoded audio information. The method comprises providing error concealment audio information for concealing a loss of an audio frame. The provision of the error concealment audio information comprises modifying a time domain drive signal obtained on the basis of one or more audio frames preceding a lost audio frame, in order to obtain the audio information. of error concealment.
This method is based on the same considerations as the audio decoder described above.
A further embodiment according to the invention creates a computer program for carrying out said method, when the computer program is executed on a computer.
Brief description of the Figures.
/..Yh,
<img file="MX356334B_D0066.tif" />
inStíT¿) 1st Mexican OF THE FkOFIECMí IÑDUSTMAL
The embodiments of the present invention will be described below with reference to the attached figures, in which:
Fig. 1 shows a schematic block diagram of an audio decoder, in accordance with an embodiment of the invention;
Fig. 2 shows a schematic block diagram of an audio decoder, in accordance with another embodiment of the present invention;
Fig. 3 shows a schematic block diagram of an audio decoder, in accordance with another embodiment of the present invention;
Figs. 4A-4B show a schematic block diagram of an audio decoder, in accordance with another embodiment of the present invention;
Fig. 5 shows a schematic block diagram of a time domain concealment for a transform encoder;
Fig. 6 shows a schematic block diagram of a time domain concealment for a switching codee;
<img file="MX356334B_D0067.tif" />
ΙΜΡΪ
Figs. 7A-7B show a diagram of .ds blocks of a TCX decoder that performs a TCX decoding in normal operation or in the case of partial packet loss;
Fig. 8 shows a schematic block diagram of a TCX decoder that performs a TCX decoding in the case of TCX-256 packet erase concealment;
Fig. 9 shows a flowchart of a method of providing decoded audio information based on encoded audio information, in accordance with an embodiment of the present invention; and
Fig. 10 shows a flow chart of a method for providing decoded audio information based on encoded audio information, in accordance with another embodiment of the present invention;
Fig. 11 shows a schematic block diagram of an audio decoder, in accordance with another embodiment of the present invention.
Detailed description of the embodiments.
one. Audio decoder according to Fig. 1.
Fig. 1 shows a schematic block diagram of an audio decoder 100, in accordance with an embodiment of the present invention. The decoder
IMPI
<img file="MX356334B_D0068.tif" />
Audio 100 receives encoded audio information 110, which, for example, may comprise an encoded audio frame in a frequency domain representation. The encoded audio information, for example, can be received via an unreliable channel, so that frame loss occurs from time to time. The audio decoder 100 further provides, on the basis of the encoded audio information 110, the decoded audio information 112.
Audio decoder 100 may comprise decoding / processing 120, which provides the decoded audio information based on the encoded audio information in the absence of frame loss.
Audio decoder 100 additionally comprises error concealment 130, which provides error concealment audio information. Error concealment 130 is configured to provide error concealment audio information 132 for concealing a loss of an audio frame after an audio frame encoded in the frequency domain representation, using an excitation signal of time domain.
In other words, decoding / processing 120 can provide decoded audio information 122 for
<img file="MX356334B_D0069.tif" />
<img file="MX356334B_D0070.tif" />
audio frames that are encoded in the form of a frequency domain representation, that is, in the form of an encoded representation, the encoded values of which describe intensities at different frequency bins.
In other words, decoding / processing 120, for example, may comprise a frequency domain audio decoder, which derives a set of spectral values from encoded audio information 110 and performs a frequency domain to time domain transform. , to thereby derive a time domain representation that constitutes the decoded audio information
122, or which forms the basis for the provision of the decoded audio information 122 in the event of further post-processing.
However, error concealment 130 does not perform error concealment in the frequency domain, but instead uses a time domain drive signal, which, for example, can serve to drive a synthesis filter, for example, a linear predictive coding synthesis (LPC) filter, which provides a time domain representation of an audio signal (eg, audio error concealment information) based on
IMPI
<img file="MX356334B_D0071.tif" />
of the time domain excitation signal, and further, based on linear predictive coding filter coefficients (LPC) (linear prediction coding filter coefficients).
Accordingly, error concealment 130 provides error concealment audio information 132, which, for example, may be a time domain audio signal, for lost audio frames, where the time domain excitation signal used by error concealment 130 may be underpinned by one or more appropriately received previous audio frames (preceding the lost audio frame), they are encoded in the form of, or can be derived from, a frequency domain representation. In conclusion, the audio decoder 100 can perform error concealment (i.e. provide error concealment audio information 132), which reduces degradation of an audio quality due to the loss of an audio frame on the encoded audio information base, where at least some audio frames are encoded in a frequency domain representation. It has been found that performing error concealment using a time domain excitation signal, even if one frame after a frame of
<img file="MX356334B_D0072.tif" />
MEXICAN INSTITUTE '^ C<sup>fcí £</sup>lSÉp <* '<sub>fi </sub>OF THE PROPERTY
INDUSTRIAL ^ & a ---->
encoded audio in the appropriately received frequency domain representation is missing, carries improved audio quality compared to error concealment that is performed in the frequency domain (for example, using a frequency domain representation of the encoded audio in the frequency domain representation preceding the missing audio frame).
This is because a smooth transition can be achieved between the decoded audio information associated with the audio frame preceding the appropriately received missing audio frame, and the error concealment audio information associated with the lost audio frame, using a time domain excitation signal, since signal synthesis, which is usually performed on the basis of the time domain excitation signal, helps to avoid discontinuities. Therefore, a good (or at least acceptable) auditory impression can be achieved, using audio decoder 100, even if an audio frame following an audio frame encoded in the appropriately received frequency domain representation is lost. . For example, the time domain approach produces an improvement over the monaural signal, such as voice, since it is closer than it is in the case of the
<img file="MX356334B_D0073.tif" />
MEXICAN INSTITUTE voice codec concealment. The use of linear predictive coding (LPC) helps avoid discontinuities, and provides a better shape for the frames.
Furthermore, it should be noted that the audio decoder 100 can be supplemented by any of the features and functionalities described below, either individually, or taken in combination.
2. Audio decoder according to Fig. 2.
Fig. 2 shows a schematic block diagram of an audio decoder 200 in accordance with an embodiment of the present invention. Audio decoder 200 is configured to receive encoded audio information 210 and to provide, on its basis, decoded audio information 220. The encoded audio information 210, for example, can take the form of a sequence of audio frames encoded in a time domain representation, encoded in a frequency domain representation, or encoded in both a time domain representation and in a frequency domain representation. In other words, all the frames of the encoded audio information 210 may be encoded in a frequency domain representation, or all the frames of the audio information
<img file="MX356334B_D0074.tif" />
<img file="MX356334B_D0075.tif" />
encoded 210 may be encoded in uiHa-_time domain representation (eg, in the form of an encoded time domain excitation signal and encoded signal synthesis parameters, eg, linear predictive encoding (LPC) parameters) . Alternatively, some frames of the encoded audio information may be encoded in a frequency domain representation, and some other frames of the encoded audio information may be encoded in a time domain representation, for example, if the audio decoder
200 It is a switching audio decoder that can switch between different decoding modes. The decoded audio information 220, for example, can be a time domain representation of one or more audio channels.
Audio decoder 200 may typically comprise decoding / processing 220, which, for example, may provide decoded audio information 232 for audio frames that are appropriately received. In other words, the decoding / processing 230 can perform a frequency domain decoding (for example, an AAC [Advanced Audio Coding] type decoding, or the like) based on one or more
<img file="MX356334B_D0076.tif" />
IMPI encoded audio frames, encoded in a frequency domain representation. Alternatively, or in addition, decoding / processing 230 may be configured to perform decoding in the time domain (or decoding in the linear prediction domain) based on one or more encoded audio frames, encoded in a representation time domain (or, in other words, in a linear prediction domain representation), for example, a TCX-excited linear prediction decoding (TCX = transformed encoded excitation) or a decoding of
ACELP (Adaptive Codebook Excited Linear Prediction Decoding). Optionally, the decoding / processing 230 can be configured to switch between different decoding modes.
Audio decoder 200 further comprises error concealment 240, which is configured to provide error concealment audio information 242 for one or more lost audio frames. Error concealment 240 is configured to provide error concealment audio information 242 for concealment of a loss of one audio frame (or even a loss of multiple audio frames). Concealment of error 240 is
INSTltUTO MKWCAf * '<sup>z</sup>
OF THE INDUSTRIAL FMOPIEDAD -►.Vir-- 'configured to modify a time domain excitation signal obtained on the basis of one or more audio frames preceding a lost audio frame, in order to obtain the audio information of concealment of error 242. In other words, error concealment 240 can obtain (or derive) a time domain drive signal for (or based on) one or more encoded audio frames preceding a lost audio frame, and can modify said time domain drive signal, which is obtained for (or on the basis of) one or more appropriately received audio frames preceding a missing audio frame, so as to obtain (by means of the modification) a time domain excitation signal which is used to provide the error concealment audio information 242. In other words, the modified time domain excitation signal can be used as an input (or as a component of an input) for a synthesis (eg linear predictive coding synthesis (LPC)) of the audio information of error concealment associated with missing audio frame '(or even multiple missing audio frames). By providing the audio concealment information of error 242 based on the time domain excitation signal obtained based on a
<img file="MX356334B_D0077.tif" />
ICηΠΌ MEXICANO
OF THE PROPERTY <sup>AND</sup> > ~~ INDUSTRIAL or more appropriately received audio frames preceding the missing audio frame, audible discontinuities can be avoided. On the other hand, by modifying the derived time domain excitation signal for (or from) one or more audio frames preceding the missing audio frame, and by providing the error concealment audio information on the basis of the modified time domain excitation signal, it is possible to consider the variation of the characteristics of the audio content (for example, a change in pitch), and in addition it is possible to avoid unnatural auditory impression (for example, by outgoing fading of a deterministic signal component (eg, at least approximately periodic). Therefore, the error concealment audio information 242 can be made to comprise some similarity to the decoded audio information 232 obtained on the basis of appropriately decoded audio frames preceding the missing audio frame, and may be achieved even though error concealment audio information 242 comprises somewhat different audio content when compared to decoded audio information
232 associated with the audio frame preceding the lost audio frame by some modification of the signal
IMPI
<img file="MX356334B_D0078.tif" />
IÑSTt + Üto MEXICANO time domain excitation. The time domain excitation modification 'Nd ^^ a used bitch' Id pj_'uu iuióii of · the error concealment audio information (associated with the missing audio frame), for example, may comprise an amplitude scale or a time scale. However, other types of modifications are possible (or even a combination of an amplitude scale and a time scale), where preferably a certain degree of relationship should remain between the obtained time domain excitation signal (such as a input information) by error concealment and modified time domain excitation signal.
In conclusion, the audio decoder 200 allows the provision of the error concealment audio information 242, such that the error concealment audio information provides a good auditory impression, even in the event that one or more frames audio is lost.
Error concealment is performed on the basis of a time domain drive signal, where a variation of the signal characteristics of the audio content during the lost audio frame is considered by modifying the drive domain domain signal. time ϊ Μ Ρ1 obtained on the basis of one or more audio frames preceding a lost audio frame.
Furthermore, it should be noted that the audio decoder 200 can be supplemented by any of the features and functionality described in this application, either individually or in combination.
3. Audio decoder according to Fig. 3.
Fig. 3 shows a schematic block diagram of an audio decoder 300, in accordance with another embodiment of the present invention.
Audio decoder 300 is configured to receive encoded audio information 310 and to provide, on its basis, decoded audio information
312. Audio decoder 300 comprises a bitstream parser 320, which may further be designated as a bitstream warper or bitstream parser. The bitstream analyzer 320 receives the encoded audio information 310 and provides, on its basis, a frequency domain representation 322 and possibly additional control information 324. The frequency domain representation
IMPI
MfcXICAN INSTITUTE '; OF THE PROPERTY
INDUSTRIAL
<img file="MX356334B_D0079.tif" />
322, for example, may comprise coded spectral values 326, coded scale factors 328, and optionally additional side information 330 which, for example, can control specific processing steps, for example noise fill, intermediate processing or further processing. The audio decoder 300 further comprises a spectral value decoding 340 which is configured to receive the encoded spectral values 326, and to provide, on its basis, a set of decoded spectral values 342. The audio decoder 300 may further comprise a scale factor decoding 350, which may be configured to receive the encoded scale factors 328 and to provide, on its basis, a set of decoded scale factors 352.
As an alternative to scale factor decoding, a linear predictive encoding (LPC) conversion to scale factor 354 can be used, for example, in the case where the encoded audio information comprises encoded linear predictive encoding (LPC) information , instead of a scale factor information. However, in some encoding modes (for example, in the decoder TCX encoding mode
IMPI
<img file="MX356334B_D0080.tif" />
MEXICAN INSTITUTE,, OF THE PROPERTY,
Audio USAC or in the audio decoder% l<sup>s</sup>™ E'VS use a set of linear predictive coefficient ^ of '' cOdlf-icagrión · (LPC) to derive a set of scale factors from the audio decoder side. This functionality can be achieved by converting linear predictive coding (LPC) to scale factor 354.
The audio decoder 300 may further comprise a scaler 360, which may be configured to apply the scaled factor set 352 to the set of spectral values 342, so as to obtain a set of scaled decoded spectral values 362. For example, a first frequency band comprising multiple decoded spectral values 342 can be scaled using a first scale factor, and a second frequency band comprising multiple decoded spectral values 342 can be scaled using a second scale factor. Therefore, the set of scaled decoded spectral values, 362, is obtained. Audio decoder 300 may further comprise optional processing 366, which may apply some processing to scaled decoded spectral values 362. For example, optional processing .366 may comprise noise padding or some other operation.
<img file="MX356334B_D0081.tif" />
The audio decoder 300 further comprises a frequency domain to time domain transform 370, which is configured to receive the scaled decoded spectral values 362, or a processed version 368 thereof, and to provide an associated time domain representation 372 with a set of scaled decoded spectral values 362. For example, the frequency domain to time domain transform 370 may provide a representation of time domain 372, which is associated with a frame or subframe of the audio content. For example, the frequency domain to time domain transform can receive a set of Modified Discrete Cosine Transform (MDCT) coefficients (which can be considered scaled decoded spectral values) and provide, on its basis, a block of domain samples of time, which can form the time domain representation 372.
The audio decoder 300 may optionally comprise a postprocessing 376, which can receive the time domain representation 372 and somewhat modify the time domain representation 372 so as to obtain a postprocessed version 378 of the time domain representation 372.
<img file="MX356334B_D0082.tif" />
frequency domain to time domain 370 and that, for example, may provide error concealment audio information 382 for one or more lost audio frames. In other words, if an audio frame is lost, such that, for example, 326 encoded spectral values are not available for that audio frame (or audio subframe), error concealment 380 may provide the audio information error concealment based on the time domain representation 372 associated with one or more audio frames preceding the missing audio frame. The error concealment audio information can typically be a time domain representation of an audio content.
It should be noted that error concealment 380, for example, can perform the functionality of error concealment 130 described above. Furthermore, error concealment 380, for example, may comprise the functionality of error concealment 500 described with reference to Fig. 5. However, generally speaking, error concealment 380 may comprise any of
IMPIOS,
IÑSTlTÜTO MEXICANO \ y ^, 3f! Sr ^! „GIVE THE PROPERTY the features and functionalities that are <sup>IND</sup>áescrW ^ rPconcerning error concealment in this document.
Regarding error concealment, it should be noted that error concealment does not occur at the same time as frame decoding. For example, if frame n is good, then, we do a normal decoding, and in the end, we save some variable that will help if we have to hide the next frame, then, if n + 1 is lost, we call the hide function providing the variable that comes from the previous good plot. Also, we will update some variables to help with the next frame loss or recovery for the next good frame.
Audio decoder 300 further comprises a signal combination 390, which is configured to receive time domain representation 372 (or postprocessed time domain representation 378 in the event of postprocessing 376). Furthermore, the signal combination -390 can receive the error concealment audio information 382, which is typically also a time domain representation of an error concealment audio signal provided for a lost audio frame. Combination of signals 390, for example, can combine time domain representations
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX356334B_D0083.tif" />
associated with subsequent audio frames. In the event that there are subsequent appropriately decoded audio frames, signal combination 390 can combine (eg, overlay and add) time domain representations associated with subsequent appropriately decoded audio frames. However, if an audio frame is lost, signal combination 390 can combine (eg, overlay and add) the time domain representation associated with the appropriately decoded audio frame preceding the lost audio frame, and the error concealment audio information associated with the missing audio frame, so as to obtain a smooth transition between the appropriately received audio frame and the lost audio frame. Similarly, signal combination 390 may be configured to combine (eg, overlay and add) the error concealment audio information associated with the missing audio frame and the time domain representation associated with another audio frame appropriately decoded after the lost audio frame (or other error concealment audio information associated with another lost audio frame, in the event that multiple consecutive audio frames are lost).
<img file="MX356334B_D0084.tif" />
Accordingly, the as signal combination 390 can provide a decoded audio information 312, so as to provide the time domain representation 372, or a post-processed version 378 thereof, for appropriately decoded audio frames, and so such that the error 382 masking audio information is provided for missing audio frames, where an overlap and add operation is usually performed between the audio information (regardless of whether it is provided by a frequency domain to time domain transform
370 or by hiding error 380) of subsequent audio frames. Because some codees have some aliasing on the overlay and add part that needs to be canceled, optionally we can create some artificial aliaslng over the half of the frame we created to perform the overlay add.
It should be noted that the functionality of the audio decoder 300 is similar to the functionality of the audio decoder 100 according to Fig. 1, where additional details are shown in Fig. 3. Furthermore, it should be noted that the audio decoder 300 according to Fig. 3 can be supplemented by any of the features and functionalities described herein.
<img file="MX356334B_D0085.tif" />
IMPI request. In particular, error concealment 380 can be supplemented by any of the features and functionality described in this application regarding error concealment.
Four. Audio decoder 400 according to Fig. 4.
Fig. 4 shows an audio decoder 400 in accordance with another embodiment of the present invention. Audio decoder 400 is configured to receive encoded audio information and to provide, on its basis, decoded audio information
412. Audio decoder 400, for example, may be configured to receive encoded audio information
410, where different audio frames are encoded using different encoding modes. For example, the audio decoder 400 can be considered a multi-mode audio decoder or a switching audio decoder. For example, some of the audio frames can be encoded using a frequency domain representation, where the encoded audio information comprises an encoded representation of spectral values (eg, FFT (Fast Fourier Transform) values or MDCT (Transform) values). cosine
ΙΜΡϊ Mexican iñítitUto? ·>
IÑíflTÜTO MEXICANO V ^ * üs? X7 ^ 3 <Λ tí THE PROPERTY V modified discrete)) and scale factors qiid<sup>NIJ</sup>W ^ etñíPPa-rt 'a scale of different frequency bands — Alill llÍ¿s; -' ía— encoded audio information 410 may further comprise a time domain representation of audio frames, or a linear prediction domain representation multiple audio frames. The linear prediction coding domain representation (also briefly referred to as linear predictive coding (LPC) representation), for example, may comprise an encoded representation of an excitation signal, and an encoded representation of linear predictive coding parameters (LPC ) (linear prediction encoding parameters), where the linear prediction encoding parameters describe, for example, a linear prediction coding synthesis filter, which is used to reconstruct an audio signal based on the time domain excitation signal.
In the following, some details of the audio decoder 400 will be described.
Audio decoder 400 comprises a bitstream analyzer 420 that, for example, can analyze encoded audio information 410 and extract, from encoded audio information 410, a representation of
<img file="MX356334B_D0086.tif" />
frequency domain 422, comprising, eg
encoded spectral values, encoded e sc'aTa factors, and optionally additional lateral information. The bitstream analyzer 420 may further be configured to extract a linear prediction coding domain representation 424, which, for example, may comprise coded excitation 426 and coded linear prediction coefficients 428 (which may also be considered prediction parameters linear encoded). Furthermore, the bitstream analyzer can optionally extract lateral information, which can be used to control additional processing steps, from the encoded audio information.
Audio decoder 400 comprises a frequency domain encoding path 430, which, for example, may be substantially identical to the encoding path of audio decoder 300 according to Fig. 3. In other words, the frequency domain encoding path 430 may comprise a spectral value decoding 340, a scale factor decoding 350, a scaler 360, an optional processing 366, a frequency domain to time domain transform.
370, optional 376 postprocessing, and
<img file="MX356334B_D0087.tif" />
η
MEXICAN INSTITUTE OF PROPERTY error 380, as previously described co ^ re'Éere Fig. 3.
Audio decoder 400 may further comprise a linear prediction domain decoding path 440 (which may further be considered a time domain decoding path, since linear predictive coding (LPC) synthesis is performed in the time domain ).
The linear prediction domain decoding path comprises an excitation decoding 450, which receives the encoded excitation 426 provided by the bitstream analyzer 420 and provides, on its basis, a decoded excitation 452 (which may take the form of a decoded time domain excitation signal). For example, drive decoding 450 may receive encoded transform coded drive information, and may provide, on its basis, a decoded time domain drive signal. Therefore, drive decoding 450, for example, can perform functionality that is performed by drive decoder 730 described with reference to Fig. 7A-7B. However, alternatively or in addition, drive decoding 450 may receive an adaptive codebook (ACELP) driven linear prediction drive.
<img file="MX356334B_D0088.tif" />
IMPI
INSTITUTO MEXICANO Di LA PROPIEDAD encoded, and can provide the 452 time excitation signal decoded on the basis of said encoded ACELP excitation information.
It should be noted that there are different options for excitation decoding. Reference is made, for example, to the relevant Standards and publications that define the concepts of codebook-excited Linear Prediction (CELP), the concepts of adaptive codebook-excited Linear Prediction (ACELP), the Modifications to the Coding Excited Linear Prediction (CELP) Coding Concepts and Coding Concepts of
Adaptive codebook-excited linear prediction (ACELP) and the concept of transformed encoded excitation coding (TCX).
The linear prediction domain decoding pathway
440 optionally comprises a processing 454 in which a processed time domain drive signal 456 is derived from the time domain drive signal 452.
The linear prediction domain decoding path 440 further comprises a linear prediction coefficient decoding 460, which is configured to receive encoded linear prediction coefficients and to provide,
IMPÍg
INSTmrro w ^ Zi on its basis, predicted coefficients «« 5r6tt decoded 462. Decoding »Utí tí u tí f 1Cfents ^^ linear prediction 460 can use different representations of a linear prediction coefficient as information of 5 input 428 , and can provide different representations of the decoded linear prediction coefficients as the output information 462. For details, reference is made to different Standards documents in which a coding and / or decoding of linear prediction coefficients 10 is described.
The linear prediction domain decoding pathway
440 optionally it comprises a 464 processing, which can process the decoded linear prediction coefficients and provide a processed version of these 466.
The linear prediction domain decoding pathway
440 it further comprises a linear predictive coding synthesis (LPC) 470, which is configured to receive the decoded excitation 452, or its processed version 456, and the decoded linear prediction coefficients 462, or its processed version 466, and to provide a signal of decoded time domain audio 472. For example, Linear Predictive Coding Synthesis (LPC) 470 may be configured to apply filtering, which is defined by
<img file="MX356334B_D0089.tif" />
IMPI
INSTITUTO MEXICANO DE LA RRORJEDAD linear prediction coefficients decodifidStlS ^ My processed version 466), to the excitation signal<sup>1</sup>'<sup>1</sup> of 'decoded time dominid 452, or its processed version, so that the decoded time domain audio signal
472 it is obtained by filtering (synthesis filtering) of the 452 (or 456) time domain excitation signal. Linear prediction domain encoding pathway 440 may optionally comprise postprocessing 474, which can be used to refine or adjust the characteristics of the decoded time domain audio signal 472.
The linear prediction domain decoding pathway
440 it further comprises an error concealment 480, which is configured to receive the decoded linear prediction coefficients 462 (or its processed version 466) and the decoded time domain drive signal 452 (or its processed version 456). The error concealment 480 may optionally receive additional information, for example height information. Error concealment 480 may accordingly provide error concealment audio information, which may be presented in the form of a time domain audio signal, in the event that a frame (or subframe) of the information encoded audio
IMPI irarrifUto mexicana - “'.'. Vi-. · Λ
410 miss out. Therefore, the occultation of the error concealment 482 audio information may be provided in such a way that the characteristics of the 482 error concealment audio information are substantially adapted to the characteristics of a last audio frame. appropriately decoded that precedes the lost audio frame. It should be understood that error concealment
480 it may comprise any of the features and functionality described with respect to 240 error concealment. Also, it should be noted that 480 error concealment may further comprise any of the features and functionality described with respect to error concealment. time domain of Fig. 6.
The audio decoder 400 further comprises a signal combiner (or signal combination 490), which is configured to receive the decoded time domain audio signal 372 (or its post-processed version 378), the error concealment audio information 382 provided by concealment of error 380, the decoded time domain audio signal 472 (or its post-processed version 476) and the 482 error concealment audio information provided by the 480 error concealment. The signal combiner · 490 may be configured to combine these
I .Μ ΡI Mexican mrtnvro y> e? ^ V ^ $ OF PROPERTY ^ 5-. J'ALCjp signals 372 (or 378), 382, 472 (or 476) and 482 in order to obtain the decoded audio information 412. Particular? an overlay and add operation can be applied by means of signal combiner 490. Accordingly, signal combiner 490 can provide smooth transitions between subsequent audio frames for which the time domain audio signal is provided by means of different entities (eg, by different encoding pathways 430, 440). However, the signal combiner
490 it can further provide smooth transitions if the time domain audio signal is provided by the same entity (eg frequency domain to time domain transform 370, or linear predictive coding synthesis (LPC) 470) for subsequent frames. Because some codees have some aliasing over the overlay and add part that needs to be canceled, optionally, we can create some artificial aliasing over the half of the frame that we created to perform the overlay add. In other words, an artificial time domain aliasing offset (TDAC [Time Domain Aliasing Effect Cancellation]) can optionally be used.
iAsVíTOtomexicani.)
FROM THE PRO! »Tí.D, \ L>? í>, i. /
In addition, the signal combiner 4 90 '^ prdvéer flat transitions to and from frames for Tas ^ cuaTés provides error concealment audio information (which is also typically a time domain audio signal).
In summary, audio decoder 400 allows decoding of audio frames that are encoded in the frequency domain, and audio frames that are encoded in the linear prediction domain. In particular, it is possible to switch between the use of the frequency domain coding path and the use of the linear prediction domain coding path according to the characteristics of the. signal (for example, using signaling information provided by an audio encoder). Different types of error concealment can be used for the provision of error concealment audio information, in the event of a frame loss, according to whether a last appropriately decoded audio frame was encoded in the frequency domain (or , equivalently, in a frequency domain representation), or in the time domain (or equivalently, in a time domain representation, or, equivalently, in a linear prediction domain, or, λ t4V. * aritr
<img file="MX356334B_D0090.tif" />
representation df £?<sup>uS</sup>™ ¿fq
IMPI equivalently, in a linear prediction). ———
5. Time domain concealment according to the
Fig. 5.
Fig. 5 shows a schematic block diagram of an error concealment in accordance with an embodiment of the present invention. The error concealment according to Fig. 5 is designated in its entirety as 500.
The error concealment 500 is configured to receive a time domain audio signal 510 and to provide, on its basis, an error concealment audio information 512, which, for example, may take the form of an audio signal time domain.
It should be noted that error concealment 500 may, for example, take the place of error concealment 130, such that error concealment audio information 512 may correspond to error concealment audio information 132. Still further , it should be noted that error concealment 500 may take the place of error concealment 380, such that the time domain audio signal 510 may correspond to the signal of
IMPI ,,,,. . ,. MEXICAN puppet 'time domain audio 372 (or to - .....
time domain 37 8), and so gn »i =» infnrmar.ión ds error concealment audio 512 may correspond to error concealment audio information 382.
Error concealment 500 comprises a pre-emphasis 520, which may be considered optional. The pre-emphasis receives the time domain audio signal and provides, on its basis, a pre-emphasized time domain audio signal 522.
The error concealment 500 further comprises a linear predictive coding (LPC) analysis 530, which is configured to receive the time domain audio signal 510, or its pre-emphasized version 522, and to obtain linear predictive coding information ( LPC) 532, which may comprise a set of linear predictive coding (LPC) parameters 532. For example, the linear predictive coding (LPC) information may comprise a set of linear predictive coding (LPC) filter coefficients (or a representation of these) and a time domain drive signal (which is matched for an drive of a linear predictive coding synthesis filter (LPC) configured according to the linear predictive coding filter (LPC) coefficients, in order to reconstruct, at least in form
IMPI
MEXICAN INSTITUTE
PROPERTY £ *** α & * & 9τ approximate, the input signal of the analysis d¿<sup>ND</sup>c% ^ linear predictive stress (LPC)). -<sup>1</sup> '<sup>1</sup> ’
Error concealment 500 further comprises a height search 540, which is configured to obtain height information 542, for example, based on a previously decoded audio frame.
The error concealment 500 further comprises an extrapolation 550, which may be configured to obtain an extrapolated time domain excitation signal based on the result of the linear predictive coding (LPC) analysis (for example, based on the signal of time domain excitation determined by linear predictive coding (LPC) analysis, and possibly based on the height search result.
The 500 error concealment further comprises a
<td>noise generation</td><td> 560,</td><td>than</td><td>provides</td><td>a noise signal</td><td> 562.</td>
<td>Concealment</td><td>of</td><td>error</td><td> 500</td><td>also includes</td><td>a</td>
<td colspan="2">combiner / fader</td><td> 570,</td><td>than</td><td>is configured</td><td>for</td>
<td>receive the signal</td><td>of</td><td colspan="2">excitement</td><td colspan="2">time domain</td>
extrapolated 552 and noise signal 562, and to provide, on its basis, a combined time domain drive signal 572. Combiner / fader 570 may be
IMPI
MÜUCANn INSTITUTE DC INDUSTRIAL PROPERTY
<img file="MX356334B_D0091.tif" />
configured to combine the extrapolated time domain excitation signal 552 and the noise signal 562, where a fading can be performed, such that a relative contribution of the extrapolated time domain excitation signal 552 (determining a deterministic component of the input signal of the linear predictive coding synthesis (LPC)) decreases as a function of time, while a relative contribution of the noise signal 562 increases as a function of time. However, different functionality of the combiner / fader is also possible. In addition, reference is made to the description below.
The error concealment 500 further comprises a 580 linear predictive coding (LPC) synthesis, which receives the combined time domain drive signal
572 and that it provides a 582 time domain audio signal on its base. For example, linear predictive coding (LPC) synthesis may further receive linear predictive coding (LPC) filter coefficients that describe a linear predictive coding (LPC) shape filter, which is applied to the domain domain excitation signal. combined time 572, in order to derive the 582 time domain audio signal. Linear predictive coding synthesis
IMPI
<img file="MX356334B_D0092.tif" />
(LPC) 580 may, for example, use linear predictive coding coefficients (LPC) obtained on the basis of one or more previously decoded audio frames (eg, provided by linear predictive coding (LPC) analysis 530).
The · error concealment 500 also includes in emphasis 584, which can be considered optional. De-emphasis 584 can provide a de-emphasized error concealment time domain audio signal 586.
The error concealment 500 further optionally comprises an overlay and add 590, which performs an overlay and add operation of the time domain audio signals associated with subsequent frames (or subframes). However, it should be noted that overlay and addition 590 should be considered optional, since error concealment may otherwise use a combination of signals that is already provided in the audio decoder environment. For example, the overlay and add 590 can be replaced by the combination of signals 390 in the audio decoder 300 in some embodiments.
In the following, some additional details regarding the 500 error concealment will be described.
100
IMPI
<img file="MX356334B_D0093.tif" />
The 500 error concealment according to Fig. 5 covers the context of a transform domain codec as
AAC_LC or AAC_ELD. In other words, the error concealment 500 is well suited for use in said transform domain codec (and, in particular, in said transform domain audio decoder). In the case of a transform codec only (for example, in the absence of a linear prediction domain decoding path), an output signal from a last frame is used as a starting point. For example, a time domain audio signal
372 it can be used as a starting point for error concealment. Preferably, no drive signal is available, only an output time domain signal from (one or more) previous frames (eg, time domain audio signal 372).
In the following, the subunits and functionalities of error concealment 500 will be described in more detail.
5.1. Linear predictive coding analysis (LPC).
In the embodiment according to Fig. 5, all concealment is performed in the excitation domain in order to obtain a smoother transition between consecutive frames. Therefore, it is necessary to first find (or,
101 'N.'T mjT (' '-'i <sup>! l</sup> i
Ot UZ> KK., »(Tf.A [· * · '» ·., ^ Ί <sub>z</sub><sup>l</sup>M> d> 1 k I Al “'n. <sup>J</sup> · * __ more generally, obtain) an appropriate set of linear predictive coding (LPC) parameters. In the embodiment according to Fig. 5, a linear predictive coding (LPC) analysis 530 is performed on the past pre-emphasized time domain signal 522. The linear predictive coding (LPC) parameters (or linear predictive coding (LPC) filter coefficients) are used to perform the linear predictive coding (LPC) analysis of the past synthesis signal (for example, based on the time domain audio signal 510, or based on the pre-emphasized time domain audio signal 522) in order to obtain an drive signal (eg, a time domain drive signal).
5.2. Height search.
There are different approaches to obtaining the height, since they are used to achieve the construction of the new signal (for example, the error hiding audio information).
In the context of the codec using an LTP filter (long-term prediction filter), as a long-term prediction filter of the
102
<img file="MX356334B_D0094.tif" />
IMPI
INSTITUTO MEXICANO Dt LA PROPIEDAD INDUSTRIAL advanced audio coding [AAC-LTP], if the last frame was advanced audio coding (AAC) with long-term prediction (LTP), we use this latest long-term prediction height delay (LTP). ) received and the corresponding gain for the generation of the harmonic part. In this case, the gain is used to decide whether to build the harmonic part in the signal or not. For example, if the long-term prediction gain (LTP) is greater than 0.6 (or any other predetermined value), then the long-term prediction information (LTP) is used to construct the harmonic part.
If there is no height information available from the previous frame, then there are, for example, two solutions, which will be described below.
For example, it is possible to perform a height search on the encoder and transmit the height delay and gain in the bit stream. This is similar to long-term prediction (LTP), although there is no filtering application (also, no long-term prediction (LTP) filtering on the clean channel).
Alternatively, it is possible to perform a height search on the decoder. Adaptive multi-speed broadband height search (AMR-WB, conforming to
103
IMPI
<img file="MX356334B_D0095.tif" />
MEXICAN INSTITUTE
PROPERTY _ in the case of the encoded transformed excitation ™ (TCX) is performed in the £ T domain of the fast Fourier transform (FFT). In Extra Low Delay (ELD), for example, if the Modified Discrete Cosine Transform (MDCT) domain was used, then the phases will be lost. Therefore, the height search is preferably performed directly in the excitation domain. This provides better results than performing the height search in the synthesis domain. The search for height in the excitation domain is first performed with an open circuit by means of a normalized cross correlation. Next, optionally, we refine the height search by performing a closed-circuit search around the open-circuit height, with a certain delta. Due to the limitations of the extra low delay (ELD) window, an erroneous height could be found, and consequently, we also verify that the height found is correct, or else we discard it.
In conclusion, the height of the last appropriately decoded audio frame preceding the missing audio frame can be considered when providing the error concealment audio information. In some cases, there is a
104
IMPI
<img file="MX356334B_D0096.tif" />
height information available from the decoding of the previous frame (ie the last frame preceding the lost audio frame). In this case, this height can be reused (possibly with some extrapolation and a consideration of a height change as a function of time). In addition, we can optionally reuse the height of more than one frame from the past, in order to try to extrapolate the height we need at the end of our hidden frame.
In addition, if there is information (eg, designated as long-term prediction gain) available, describing an intensity (or relative intensity) of a deterministic signal component (eg, at least approximately periodic), this value may be used to decide whether a deterministic (or harmonic) component should be included in the error concealment audio information. In other words, by comparing said value (eg LTP gain) with a predetermined threshold value, it can be decided whether a time domain drive signal derived from a previously decoded audio frame should be considered for the provision of the information audio hiding error or not.
If there is no height information available from the previous frame (or, more precisely, from the frame decoding
105
<img file="MX356334B_D0097.tif" />
IMPI prior), there are different options. The height information could be transmitted from an audio encoder to an audio decoder, which would simplify the audio decoder while creating a bit rate overhead. Alternatively, the height information may be determined in the audio decoder, for example, in the drive domain, ie based on a time domain drive signal. For example, the time domain excitation signal derived from an appropriately decoded previous audio frame may be evaluated to identify the height information to be used for the provision of the error concealment audio information.
5.3. Extrapolation of the excitation or creation of the harmonic part.
The excitation (for example, the time domain excitation signal) obtained from the previous frame (or only computed for the lost frame or already saved in the previous lost frame for multiple frame loss) is used for the construction of the harmonic part (also designated as deterministic component or approximately periodic component) in the excitation (for example, in the signal of
106
<img file="MX356334B_D0098.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY i- ·· * 'input of the synthesis of linear predictive coding (LPC)) by copying the last height cycle as many times as necessary to obtain a plot and a half. In order to save complexity, we can also create one and a half frames only for the first loss frame and then change the processing for subsequent frame loss to half the frame, and create only one frame for each one. Next, we always have access to half of an overlay pattern.
In the case of the first frame lost after a good frame (i.e. a properly decoded frame), the first height cycle (for example, of the time domain drive signal obtained based on the last appropriately decoded audio frame preceding the missing audio frame) is the low pass filter with a filter dependent on the sampling rate (since the extra low delay (ELD) covers a really wide sampling rate combination - ranging from AAC-ELD core to AACELD with SBR or AAC-ELD dual SBR rate).
The height in a voice signal is almost always changing.
Therefore, the concealment presented above tends to create some problems (or at least distortions) in the recovery, since the height at the end of the hidden signal
107
IMPI
MiatCANf INSTITUTE. W IA »« OR £ BAI) inovstrial
<img file="MX356334B_D0099.tif" />
(that is, at the end of the error concealment information) often does not match the height of the first good frame. Therefore, optionally, in some embodiments, it is a matter of predicting the height at the end of the hidden frame so as to match the height at the beginning of the retrieval frame. For example, the height at the end of a missing frame (which is considered a hidden frame) is predicted, where the goal of the prediction is to set the height at the end of the lost frame (hidden frame) in order to approximate the height at start of first appropriately decoded frame after one or more lost frames (whose first appropriately decoded frame is also called recovery frame). This could be done during frame loss or during the first good frame (ie, during the first appropriately received frame). For even better results, some conventional tools can be optionally reused and adapted, such as pulse and height prediction resynchronization. For details, reference is made, for example, to reference [6] and [7].
108
<img file="MX356334B_D0100.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
If a long-term prediction (LTP) is used in a frequency domain codec, it is possible to use the delay as the starting information about the height. However, in some embodiments, it is further desired to have better granularity in order to better track the height contour. Therefore, it is preferred to perform a height search at the beginning and end of the last good (properly decoded) frame. In order to adapt the signal to the moving height, it is desirable to use a pulse resynchronization, which is presented in the state of the art.
5.4. Height gain.
In some embodiments, applying a gain to the previously obtained excitation is preferred in order to achieve the desired level. The height gain (for example, the gain of the deterministic component of the time domain drive signal, i.e. the gain applied to a time domain drive signal derived from a previously decoded audio frame, in order to of obtaining the input signal of the synthesis of linear predictive coding (LPC)), can, for example, be obtained by performing a
109
IMPI
<img file="MX356334B_D0101.tif" />
normalized correlation in the time domain at the end of the last good (eg appropriately decoded) frame. The length of the map can be equivalent to the length of two subframes, or it can be adaptively changed. The delay is equivalent to the height delay used to create the harmonic part.
We can also optionally perform the gain calculation only on the first lost frame and then only apply an out fade (reduced gain) for the next consecutive frame loss.
The height gain will determine the amount of hue (or the number - of deterministic, at least approximately periodic signal components) to be created. However, it is desirable to add some shaped noise so as not to have just an artificial tone. If we get very low gain in height, then we build a signal that consists only of shaped noise.
As a conclusion, in some cases, the time domain drive signal obtained, for example, based on a previously decoded audio frame, is scaled according to gain (for example, in order to obtain the input signal for linear predictive coding (LPC) analysis). Therefore, because the
110
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX356334B_D0102.tif" />
dctorminn domain excitation signal nn deterministic signal component (at least approximately periodic), the gain can determine a relative intensity of said deterministic signal components (at least approximately periodic) in the error concealment audio information. Furthermore, the error concealment audio information may be supported by noise, which is further formed by linear predictive coding synthesis (LPC), such that a total power of the error concealment audio information is adapted, at least to some degree, to an appropriately decoded audio frame preceding the missing audio frame and, ideally, in addition to an appropriately decoded audio frame after the one or more missing audio frames.
5.5. Creation of the noise part.
An innovation is created by a random noise generator. Optionally, this noise is additionally high-pass filtered and optionally pre-emphasized for voice and start frames. As for the low pass of the harmonic part, this filter (for example, the high pass filter) is dependent on the sampling rate. This noise (which is
111
<img file="MX356334B_D0103.tif" />
ΙΜΡΪ
INSTITUTO MEXICANO Ut LA PKOWti 'AO INDUSTRIAL provided, for example, by a noise generation 560) will be formed by linear predictive coding (LPC) (for example, by the synthesis of linear predictive coding (LPC) 580) to get the most as close to background noise as possible. The high pass feature is also optionally changed over consecutive frame loss, so that over a certain amount of a frame loss, there is no more filtering, just to get the full band-shaped noise to achieve a comfort noise close to background noise.
An innovation gain (which, for example, can determine a noise gain 562 in the outgoing fade / blend 570, i.e. a gain using noise signal 562 that is included in input signal 572 of the synthesis of linear predictive coding (LPC)) is, for example, calculated by removing the previously computed contribution of height (if any) (for example, a scaled version, scaled using height gain, of the obtained time domain excitation signal based on the last appropriately decoded audio frame preceding the missing audio frame) and performing a correlation at the end of the last good frame. Refering to
112
IMPI
<img file="MX356334B_D0104.tif" />
height gain, this could be done optionally 1 blanket snln on the first frame lost, and then the outgoing fade, although in this case the outgoing fade could go either 0, resulting in a completed mute, or a estimated noise level present in the background. The length of the correlation is, for example, equivalent to the length of two subframes, and the delay is equivalent to the height delay used to create the harmonic part.
Optionally, this gain is further multiplied by (l-height gain) in order to apply as much gain on the noise as to achieve the gain loss if the height gain is not one. Optionally, this gain is further multiplied by a noise factor. This noise factor comes, for example, from the previous valid frame (for example, from the last appropriately decoded audio frame preceding the missing audio frame).
5.6. Outgoing fading.
The 'fade out' is mostly used for multiple frame loss. However, the outgoing fade can also be used in the event that only a single audio frame is lost.
113
IMPI
MEXICAN INSTITUTE
FROM THE PROPERTY 'Λ ^ -νΗί?' · JJV
In the case of a multiple loss d ^^ tMama<sup>5</sup>?<sup>1</sup>-—Pds predictive coding parameters lifie'ai '' (L'PCT “TKr ~ Süir recalculated. Either the last computed is maintained, or the linear predictive coding (LPC) concealment is performed by convergence to In this case, the periodicity of the signal converges to zero. For example, the time domain driving signal 502 obtained on the basis of one or more audio frames preceding a missing audio frame still uses a gain that is gradually reduced as a function of time, while noise signal 562 stays constant or scaled with a gain that is gradually increasing as a function of time, such that the relative weight of the time domain driving signal 552 is reduced as a function of time compared to the relative weight of the noise signal 562. Accordingly, the input signal 572 of the predictive coding synthesis Linear (LPC) 580 is becoming increasingly noise-like. Therefore, the periodicity (or, more precisely, the deterministic component, or at least approximately periodic component of the output signal 582 of linear predictive coding synthesis (LPC) 580) is reduced as a function of time.
114
IMPI
MEXICAN INSTITUTE OF PROPERTY <sub>τ</sub> ί ·,,,, INDUSTRIAL The speed of convergence according to which the periodicity of signal 572, and / or the periodicity ~ of signal 582, converges to 0, depends on the parameters of the last correctly received frame ( or appropriately decoded) and / or the number of consecutive erased frames, and is controlled by an attenuation factor, a. The factor, a, is additionally dependent on the stability of the LP filter. Optionally, it is possible to alter the factor a in relation to the height length. If the height (for example, a periodic length associated with the height) is really long, then we keep it normal, but if the height is really short, it is usually necessary to copy the same part of the past excitation a number of times.
This will quickly sound too artificial, and therefore the faster outgoing fade of this signal is preferred.
Also, optionally, if available, we can consider the height prediction output. If a height is predicted, this means that the height was already changing in the previous frame, and then, the more frames we lose, the further we are from the truth. Therefore, it is preferred to somewhat accelerate the fade in protrusion of the tonal part, in this case.
<img file="MX356334B_D0105.tif" />
115
<img file="MX356334B_D0106.tif" />
MEXICAN INSTITUTE OF PROPERTY _. ....,. INDUSTRIAL - -a »
If the height prediction fails because the height changes too much, this means that either the "height values are not really reliable, or the signal is really unpredictable. Therefore, again, it is preferred to perform outgoing fading faster (eg, outgoing fading of time domain drive signal 552 obtained on the basis of one or more appropriately decoded audio frames preceding one or more lost audio frames).
5.7. Linear predictive coding synthesis (LPC).
In order to return to the time domain, it is preferred to perform a linear predictive coding synthesis (LPC)
580 over to the sum of the two excitations (tonal part and noisy part), followed by a de-emphasis. In other words, it is preferred to perform 580 linear predictive coding (LPC) synthesis on the basis of a heavy combination of a 552 time domain excitation signal obtained on the basis of one or more appropriately decoded audio frames preceding the missing audio frame (tonal part) and noise signal 562 (noisy part). As mentioned above, the 552 time domain excitation signal can be modified compared to the
116
<img file="MX356334B_D0107.tif" />
IMPI /
MEXICAN INSTITUTE \
PROPERTY time domain excitation signal 532 see linear predictive coding analysis (LPC) 5T0 ~ 7a ^ ema'έΓ of linear predictive coding coefficients (LPC) describing a filter characteristic of linear predictive coding synthesis (LPC) used for the synthesis of linear predictive coding (LPC) 580). For example, time domain drive signal 552 may be a time scaled copy of time domain drive signal 532 obtained by linear predictive coding (LPC) analysis 530, where the time scale can be used to adapt the height of the 552 time domain driving signal to a desired height.
5.8. Overlap and addition.
In the case of a transform codec only, in order to obtain the best overlay and addition, we create an artificial signal for half a frame more than the hidden frame, and we can create artificial aliasing on it. However, different concepts of overlap and addition may apply.
In the context of advanced audio encoding (AAC) or regular transformed encoded excitation (TCX),
117
IMPIOUS
INSTITUTO MEXICANO applies an overlay and addition between 'extra that comes from concealment and the first good first frame (could be half or less, for smaller delay windows like AAC-LD).
In the special case of extra low delay (ELD) for the first missed frame, it is preferred to run the analysis three times in order to obtain the appropriate contribution from the three windows, and then for the first concealment frame, and all subsequent , the analysis is run once more. Next, an extra low delay (ELD) synthesis is performed, to return to the time domain with all the appropriate memory for the next frame in the
Modified Discrete Cosine Transform (MDCT).
In conclusion, input signal 572 of linear predictive coding synthesis (LPC) 580 (and / or time domain drive signal 552) can be provided for a time duration that is greater than a duration of an audio frame lost. Accordingly, the output signal 582 of the linear predictive coding synthesis (LPC) 580 can further be provided for a period of time that is greater than a lost audio frame. Accordingly, an overlay and addition can be performed between the error concealment audio information (which is consequently
118
<img file="MX356334B_D0108.tif" />
,,,,, IÑSflfUTO MEXICANO obtained for a period of time more λια »« * «ραι>
tRial temporal extension of the audio frame. Lost) and Ί-decoded audio information provided for an appropriately decoded audio frame after one or more lost audio frames.
In summary, error concealment 500 is well suited to the case where audio frames are encoded in the frequency domain. Even when the audio frames are encoded in the frequency domain, the provision of the error concealment audio information is performed on the basis of a time domain drive signal. Different modifications are applied to the obtained time domain excitation signal based on one or more appropriately decoded audio frames preceding a missing audio frame. For example, the time domain excitation signal provided by linear predictive coding analysis (LPC) 530 adapts to changes in height, for example, using a time scale. Furthermore, the time domain excitation signal provided by linear predictive coding analysis (LPC) 530 is further modified by a scale (application of a gain), where an outward fading of the deterministic (or tonal, or otherwise less about
119
<img file="MX356334B_D0109.tif" />
INSTITUTO MEXICANO. $ Periodical) can be performed by the escalacfói ^ 'w.val ^ SéÓ ^' r '57 0, so that the input signal · - &? £ —de —-! · © —-sí- nb-e-si-sde linear predictive coding (LPC) 580 comprises both a component derived from the time domain excitation signal obtained by linear predictive coding (LPC) analysis and a noise component that is based on the noise signal 562. The deterministic component of input signal 572 of linear predictive coding synthesis (LPC) 580 however, is usually modified (eg, time scale and / or amplitude scale) with respect to the excitation domain signal of time provided by the linear predictive coding (LPC) analysis 530.
<td>In</td><td>consequence the signal</td><td>of</td><td>excitement</td><td>Of domain</td><td>of</td>
<td>15 time</td><td>can be adapted to</td><td colspan="2">needs,</td><td>and it is avoided</td><td>a</td>
<td colspan="2">unnatural auditory impression.</td><td></td><td></td><td></td><td></td>
<td> 6.</td><td>Domain concealment</td><td>of</td><td>time of</td><td>agree with</td><td>the</td>
<sup>Fi</sup>g <sup>6</sup>·
Fig. 6 shows a schematic block diagram of a time domain concealment that can be used for a switching codec. For example, the time domain concealment 600 according to Fig. 6 can, for example,
120
IMPI
INSTITUTO MEXICANO,. 'U take the place of concealment of error 240, concealment of error 480. ..................
Furthermore, it should be noted that the embodiment according to Fig. 6 covers the context (which can be used within the context) of a switching codec using combined time and frequency domains, such as USAC [Unified Voice Coding and audio] (MPEG-D / MPEG-H) or EVS (3GPP). In other words, time domain concealment 600 can be used in audio decoders in which there is a switch between a frequency domain decoding and a time decoding (or, equivalently, a decoding based on linear prediction coefficients ).
However, it should be noted that the error concealment 600 according to Fig. 6 can also be used in audio decoders that merely perform decoding in the time domain (or equivalently, in the linear prediction coefficient domain).
In the case of a switched codec (and even in the case of a codec that merely performs decoding in the linear prediction coefficient domain), we usually already have the excitation signal (for example, the excitation signal of the domain of time) that comes from a plot
121
<img file="MX356334B_D0110.tif" />
MEXICAN INSTITUTE OF PROPERTY,. INDUSTRIAL —____ previous (for example, an appropriately decoded audio frame preceding a lost audio frame). Otherwise (for example, if the time domain excitation signal is not available), it is possible to act as explained in the embodiment according to Fig. 5, i.e. perform a linear predictive coding analysis (LPC). If the previous frame was of type Linear prediction excited by adaptive codebook (ACEL), we also already have the height information of the subframes in the last frame. If the last frame was TCX (Transformed Coded Excitation) with LTP (Long-Term Prediction), then we also have the delay information that comes from the long-term prediction. And if the last frame was in the frequency domain without long-term prediction (LTP), then the height search is preferably performed directly in the excitation domain (for example, based on a domain excitation signal of time provided by a linear predictive coding (LPC) analysis).
If the decoder already uses some predictive encoding parameters line al (LPC) in the time domain, we reuse them and extrapolate a new set of predictive linear encoding parameters (LPC). The
122 instíVótó ^ eXícanc S<sup>m</sup>.
, -, „, 0 £ t \ QA \ M # 3'AI> i, · ^ extrapolation of the coding parameters 108 ^ 6 ^^^^^ 6 ^ va linear (LPC) is supported in the pass cod: ffíTdtr ^^ TÍ ~ * j linear (LPC), for example, the mean of the last three frames and (optionally), the form of linear predictive coding (LPC) derived during DTX noise estimation if DTX (discontinuous transmission) exists in the codec.
All concealment is performed in the excitation domain in order to obtain a smoother transition between consecutive frames.
In the following, error concealment 600 will be described in more detail in accordance with Fig. 6.
The error concealment 600 receives a past excitation 610 and a past height information 640. Furthermore, the error concealment 600 provides an error concealment audio information 612.
It should be noted that the past excitation 610 received by the error concealment 600 may, for example, correspond to output 532 of the linear predictive coding (LPC) analysis 530. Furthermore, the past height information 640 may, for example, correspond to output information 542 of height search 540.
123
<img file="MX356334B_D0111.tif" />
Error concealment 600 additionally comprises extrapolation 650, which may correspond to extrapolation 550, such that reference is made to the description above.
Furthermore, the error concealment comprises a noise generator 660, which may correspond to the noise generator 560, such that reference is made to the above description.
Extrapolation 650 provides an extrapolated time domain drive signal 652, which may correspond to extrapolated time domain drive signal 552.
<td>The noise generator</td><td> 660</td><td colspan="2">provides a signal</td><td>noise 662,</td><td>than</td>
<td>corresponds to the signal</td><td>of</td><td>noise 562.</td><td></td><td></td><td></td>
<td>Concealment</td><td>of</td><td>error 600</td><td>also</td><td>understands</td><td>a</td>
<td colspan="2">combiner / fader</td><td>670 which</td><td>receives</td><td>the signal</td><td>of</td>
time domain excitation extrapolated 652 and noise signal 662 and provides, on its basis, an input signal 672 for a linear predictive coding synthesis (LPC) 680, where the linear predictive coding synthesis (LPC) 680 can correspond to Linear Predictive Coding (LPC) Synthesis 580, so that the above explanations also apply. Linear Predictive Coding Synthesis (LPC) 680 provides a signal of
124 time domain audio 682
IMPIOUS that can correspond ¡MEXICAN INSTITUTE JE LA I'ROEIELÍAJ INDUSTRIAL
<img file="MX356334B_D0112.tif" />
time domain audio signal 582. Error concealment further comprises (optionally) a de-emphasis 684, which may correspond to de-emphasis 584 and which provides a de-emphasized error concealment time domain audio signal 686. Error concealment 600 optionally comprises an overlay and addition 690, which may correspond to the overlay and addition 590. However, the explanations regarding superposition and addition 590, superposition and addition also apply.
690. In other words, the overlay and addition 690 can also be replaced by the overlay and override of the audio decoder, such that the linear predictive coding synthesis (LPC) output signal 682 or the. Output 686 of the de-emphasis can be considered the error concealment audio information.
In conclusion, error concealment 600 differs substantially from error concealment 500, in that error concealment 600 directly derives past excitation information 610 and past height information 640 from one or more previously decoded audio frames , without the need to carry out an analysis of
125
IMPIé Mexican institute and you
OF THE PROPERTY ·*
INDUSTRIAL linear predictive coding (LPC) and / or height analysis. However, it should be noted that error concealment 600 may optionally comprise linear predictive coding (LPC) analysis and / or height analysis (height search).
In the following, some features of error concealment 600 will be described in more detail. However, it should be noted that specific details should be considered exemplary, rather than essential features.
6.1. Height search height passed.
There are different approaches to obtain the height to be used in the construction of the new signal.
In the context of encoding using the long-term prediction filter (LTPE) e, as a long-term prediction filter of advanced audio encoding [AAC-LTP], if the last frame (preceding the missing frame) was advanced audio encoding (AAC) with long-term prediction (LTP), we have the height information that comes from the last long-term prediction height delay (LTP) and the corresponding gain. In this case, we use the gain in order to decide if we want to build the harmonic part in the signal or
126
<img file="MX356334B_D0113.tif" />
INjflf UTü MEXICANO OE LA INDUSTRIAL property no. For example, if the long-term prediction gain (LTP) is greater than 0.6, then we use the long-term prediction (LTP) information to construct the harmonic part.
If we don't have any height information available from the previous frame, then there are, for example, two additional solutions.
One solution is to perform a height search on the encoder and transmit the height delay and gain in the bitstream. This is similar to long-term prediction (LTP), although we do not apply any filtering (either, no long-term prediction filtering on the clean channel).
Another solution is to perform a height search on the decoder. The adaptive multi-rate Broadband (AMR-WB) height search in the case of transformed encoded excitation (TCX) is performed in the Fast Fourier transform (FFT) domain. In transformed encoded excitation (TCX), for example, we use the domain of the modified discrete cosine transform (MDCT), so we lose the phases. Therefore, the height search is performed directly in the excitation domain (for example, based on the signal of
127
MEXICAN INSTITUTE OF PROPERTY, ρΑ, ί *
INDUSTRIAL - * time domain excitation used as the input for linear predictive coding synthesis (LPC), or used to derive the input for linear predictive coding synthesis (LPC)), in a preferred embodiment. This usually provides better results than performing the height search in the synthesis domain (eg, based on a fully decoded time domain audio signal).
The height search in the excitation domain (for example, on the basis of the time domain excitation signal) is first performed with an open circuit by means of a normalized cross correlation. Optionally, the height search can then be refined by performing a closed-circuit search around the open-circuit height with a certain delta.
In preferred implementations, we do not simply consider a maximum correlation value. If we have height information from a previous frame that is not error-prone, then we select the height that corresponds to that of the five highest values in the normalized cross-correlation domain, although the closest to the height of the previous frame. Then, it is also verified that
ΙΜΡΪ
<img file="MX356334B_D0114.tif" />
128 the maximum found is not a maximum -rrnea ttSBTcTo 'to the window constraint.
In conclusion, there are different concepts for determining height, where it is computationally efficient to consider a past height (i.e. height associated with a previously decoded audio frame).
Alternatively, the height information can be transmitted from an audio encoder to an audio decoder. As another alternative, a height search can be performed on the audio decoder side, where the height determination is preferably performed on the basis of the time domain drive signal (ie in the drive domain).
A two stage height search comprising an open circuit search and a closed circuit search can be performed in order to obtain particularly reliable and accurate height information. Alternatively, or in addition, height information from a previously decoded audio frame may be used to ensure that the height search provides a reliable result.
129
6.2.
Extrapolation of
<img file="MX356334B_D0115.tif" />
MEXICAN INSTITUTE OF PROPERTY,. , INDUSTRIAL, --excitation or creation of the harmonic part.
The excitation (eg in the form of a time domain excitation signal) obtained from the previous frame (or only computed for the lost frame or already saved in the previous lost frame for multiple frame loss) is used to construct the harmonic part in the excitation (eg, the extrapolated time domain excitation signal 662) by copying the last height cycle (eg, a portion of the time domain excitation signal 610, whose temporal duration is equal to a period period of the height) as many times as necessary to obtain, for example, one and a half of the plot (lost).
In order to obtain even better results, it is optionally possible to reuse some tools known from the state of the art and adapt them. For details, reference is made, for example, to references [6] and [7].
The height in a voice signal has been found to be almost always changing. Therefore, the concealment presented above has been found to tend to create some problems in recovery, since the height at the end of the
130
<img file="MX356334B_D0116.tif" />
MEXICAN INSTITUTE ~ η n, OWNERSHIP OstoS & L® * / hidden signal often does not match the '^^ Ptur ^^ hr the first good plot. So, opciottailTimife,<sup>Ί</sup> 5'5trat ^ -OS predict the height at the end of the hidden frame, in order to match the height at the beginning of the recovery frame. This functionality will be performed, for example, by extrapolation 650.
If long-term prediction (LTP) is used in Transformed Coded Excitation (TCX), delay can be used as the initial information about height. However, it is desirable to have better granularity in order to better track the height contour. Therefore, a height search is optionally performed at the beginning and end of the last good frame. In order to adapt the signal to the moving height, a pulse resynchronization can be used, which is presented in the state of the art.
In conclusion, extrapolation (eg, of the time domain drive signal associated with, or obtained on the basis of, a last appropriately decoded audio frame preceding the lost frame) may comprise a copying of a time slice of said time domain excitation signal associated with a previous audio frame, where the copied time portion may be modified according to a computation or a
131
<img file="MX356334B_D0117.tif" />
I NSTlturo mexicana i de lA PROPERTY '
INDUSTRIAL estimate, of a change in height (expected) during the lost audio frame. Different concepts can be obtained for determining the height change.
6.3. Height gain.
In the embodiment according to Fig. 6, a gain is applied on the previously obtained excitation in order to reach a desired level. The height gain is obtained, for example, by performing a normalized correlation in the time domain at the end of the last good frame. For example, the length of the correlation can be equivalent to the length of two subframes, and the delay can be equivalent to the height delay used for the creation of the harmonic part (for example, for copying the excitation signal of time domain). It has been found that doing the gain calculation in the time domain provides a much more reliable gain than doing it in the excitation domain. The linear predictive coding (LPC) changes in each frame, and then the application of a gain, calculated on the previous frame, on an excitation signal that will be processed by another set of linear predictive coding (LPC), will not provide the energy expected in the time domain.
132
<img file="MX356334B_D0118.tif" />
MEXICAN INSTITUTE INDUSTRIAL PROPERTY
The height gain determines the amount of hue that will be created, although some shaping noise will also be added so as not to have just an artificial tone. If a very low height gain is obtained, then a signal consisting of just shaped noise can be constructed.
In conclusion, a gain that is applied to scale the time domain drive signal obtained on the basis of the previous frame (or a time domain drive signal that is obtained for a previously decoded frame, or which is associated with the previously decoded frame) is adjusted so as to determine a value of a tonal component (or deterministic, or at least approximately periodic) within the 680 linear predictive coding synthesis (LPC) input signal, and consequently within the error concealment audio information. Said gain can be determined on the basis of a correlation, which is applied to the time domain audio signal obtained by a decoding of the previously decoded frame (where said time domain audio signal can be obtained using a synthesis of linear predictive coding (LPC) that is performed in the course of decoding).
133
<img file="MX356334B_D0119.tif" />
6.4. Creation of the noise part.
An innovation is created by means of a 660 random noise generator. This noise is additionally high-pass filtered and optionally pre-emphasized for start and voice frames. High-pass filtering and pre-emphasis, which can be done selectively for speech and start frames, are not explicitly shown in Fig. 6, although they can be done, for example, within the noise generator
660 or inside the combiner / fader 670.
Noise will be formed (eg, after combining with the time domain excitation signal
652 obtained by extrapolation 650) by linear predictive coding (LPC) in order to obtain as close as possible to the background noise.
For example, the innovation gain can be calculated by removing the previously computed height contribution (if any) and correlating at the end of the last good plot. The length of the correlation can be equivalent to the length of two subframes, and the delay can be equivalent to the height delay used to create the harmonic part.
134
<img file="MX356334B_D0120.tif" />
Optionally, this gain can also be _______ multiplied by (1-height gain) in order to apply as much gain on noise to achieve energy loss if the height gain is not one. Optionally, this gain is further multiplied by a noise factor.
This noise factor can come from a previous valid frame.
In conclusion, a noise component of the error concealment audio information is obtained by the noise formation provided by the noise generator 660 using the 680 linear predictive coding synthesis (LPC) (and possibly the de-emphasis 684) .
In addition, additional high pass filtration and / or pre-emphasis may be applied. Linear Predictive Coding Synthesis (LPC) 680 input contribution noise contribution 672 (further designated innovation gain) can be computed based on the last appropriately decoded audio frame preceding the audio frame missing, where a deterministic (or at least roughly periodic) component can be removed from the audio frame preceding the missing audio frame, and where a correlation can then be performed to determine the intensity (or gain) of the noise component
135
ΙΜ1
<img file="MX356334B_D0121.tif" />
i. * 4 i + Mexican iVtrro within the time domain signal deca ^^^ ea audio frame preceding the frame of
Optionally, certain additional modifications may be applied to the gain of the noise component.
6.5. Outgoing fading.
Outbound fading is mostly used for multiple frame losses. However, fade-out can also be used in the event that only a single audio frame is lost.
In the case of multiple frame loss, the linear predictive coding (LPC) parameters are not recalculated.
Either the last computation is maintained, or a linear predictive coding (LPC) concealment is performed as explained above.
A periodicity of the signal converges to zero. The rate of convergence depends on the parameters of the last correctly received (or correctly decoded) frame and the number of consecutive erased (or lost) frames 20, and is controlled by an attenuation factor, I hear.
The factor, I heard, also depends on the stability of the linear prediction filter (LP). Optionally, the factor oi can be altered relative to the height length. For example,
136
IMPI
MEXICAN INSTITUTE
FROM PROPERTY Lj'Y 'If the height is really long, then I heard it can stay normal, but if the height is really short, it may be convenient (or necessary) to copy the same part of past excitation a number of times. Because this has been found to quickly sound too artificial, the signal therefore fades out more quickly.
Also optionally, it is possible to consider the height prediction output. If a height is predicted, this means that the height was already changing in the previous frame, and so the more frames that are lost, the further we are from the truth. Therefore, it is desirable to somewhat accelerate the fade in protrusion of the tonal part, in this case.
If the height prediction fails because the height changes too much, this means that either the height values are not really reliable, or the signal is really unpredictable. Therefore, again, we should perform the outgoing fade faster.
In conclusion, the contribution of the extrapolated time domain excitation signal 652 to the input signal 672 of the linear predictive coding synthesis (LPC) 680 is usually reduced as a function of time. This
137
<img file="MX356334B_D0122.tif" />
INDUSTRIAL rrofiedaíj Z · 'e gain, can be achieved, for example, by reducing a valueé that is applied to the extrapolated time domain excitation signal 652, as a function of time. The rate used to gradually reduce the gain applied to scale the 552 time domain drive signal obtained on the basis of one or more audio frames preceding a missing audio frame (or one or more of its copies) is adjusted from according to one or more parameters of one or more audio frames (and / or according to a number of consecutive lost audio frames). In particular, the length of height and / or the rate at which height changes as a function of time, and / or the question of whether a height prediction fails or succeeds, can be used to adjust said speed.
6.6. Linear predictive coding synthesis (LPC).
Ά In order to return to the time domain, a linear predictive coding synthesis (LPC) 680 is performed on the sum in general (or generally, the heavy combination) of the two excitations (tonal part 652 and noisy part 662), followed by de-emphasis 684.
In other words, the result of the heavy combination (fading) of the domain excitation signal of
138
<img file="MX356334B_D0123.tif" />
time extrapolated 652 and noise signal 662 forms a combined time domain excitation signal, which is input into the 680 linear predictive coding (LPC) synthesis, which, for example, can perform synthesis filtering based on said combined time domain excitation signal 672 according to linear predictive coding coefficients (LPC) describing the synthesis filter.
6.7. Overlap and addition.
Because the mode of the next arriving frame (for example, adaptive codebook-excited linear prediction (ACELP), is not known during concealment,
Transformed encoded excitation (TCX) or frequency domain (FD)), it is preferred to prepare different overlays in advance. In order to achieve the best overlay and addition if the next frame is in a transform domain (TCX or FD), an artificial signal (for example, error concealment audio information) can, for example, be created for the half a frame more than the hidden (lost) frame. Furthermore, artificial aliasing can be created on it (where artificial aliasing
139
IMPI
INSTITUTO MEXICANO DC THE PROPERTY
INDUSTRIAL
<img file="MX356334B_D0124.tif" />
it can, for example, be adapted to the superposition and addition of inverse modified discrete cosine transform (MDCT)).
In order to obtain a good overlap and addition without discontinuity with the future plot in the time domain (ACE1P [Linear Prediction Excited by Adaptive Codebook]), we do as above, but without aliasing, so that we can apply long superposition and addition, or if we want to use a square window, the zero input response (ZIR) at the end of the synthesis buffer is computed.
In conclusion, in a switching audio decoder (which can, for example, switch between an adaptive codebook-excited linear prediction decoding (ACE1P), a transformed encoded excitation decoding (TCX), and a frequency domain decoding (FD decoding)), an overlay and addition can be performed between the error concealment audio information that is mainly provided for a lost audio frame, but in addition, for a certain portion of time after the lost audio frame, and the decoded audio information provided for the first appropriately decoded audio frame after a sequence of one or more lost audio frames. For the purpose of
140
API
<img file="MX356334B_D0125.tif" />
.LB
MEXICAN INSTITUTE f * ~
GIVE THE EROPISDAD,. . ,,,, ... INDUSTRIAL · * - to obtain an appropriate superposition and addition, even, for decoding modes that carry a time domain aliasing in a transition between subsequent audio frames, an aliasing cancellation information can be provided ( for example, designated artificial aliasing). Accordingly, an overlap and addition between the error concealment audio information and the time domain audio information obtained on the basis of the first appropriately decoded audio frame after a lost audio frame, achieves aliasing cancellation. .
If the first appropriately decoded audio frame after the sequence of one or more missing audio frames is encoded in adaptive codebook-excited linear prediction mode (ACELP), specific overlay information may be computed, which may be supported by a zero input response (ZIR) from a linear predictive coding filter (LPC).
In conclusion, the 600 error concealment is well suited for use in a switching audio codec. However, error concealment 600 can also be used in an audio codec that merely decodes content from
141
<img file="MX356334B_D0126.tif" />
ΪΜΡΙ audio encoded in a transformed encoded excitation mode (TCX) or in an adaptive codebook excited linear prediction mode (ACELP).
6.8. Conclusion.
It should be noted that particularly good error concealment is achieved by the above-mentioned concept, for extrapolation of a time domain excitation signal, the combination of the extrapolation result with a noise signal using a fade (eg a cross fade), and for performing a linear predictive coding (LPC) synthesis based on a cross fade result.
7. Audio decoder according to Fig. 11.
Fig. 11 shows a schematic block diagram of an audio decoder 1100, in accordance with an embodiment of the present invention.
It should be noted that the audio decoder 1100 may be part of a switching audio decoder. For example, audio decoder 1100 can replace linear prediction domain decoding path 440 in audio decoder 400.
142
IMPI iñstítütü Mexican Ót lA PfombAD
The audio decoder 1100 is configured to receive encoded audio information 1110 and to provide, on its basis, decoded audio information
1112. The encoded audio information 1110 may, for example, correspond to the encoded audio information
410, and the decoded audio information 1112 may, for example, correspond to the decoded audio information
412.
The audio decoder 1100 comprises a bitstream analyzer 1120, which is configured to extract an encoded representation 1122 from a set of spectral coefficients and an encoded representation of linear prediction encoding coefficients 1124 from the encoded audio information 1110. Without However, the bitstream analyzer 1120 can optionally extract additional information from the encoded audio information 1110.
The audio decoder 1100 further comprises a spectral value decoding 1130, which is configured to provide a set of decoded spectral values 1132 based on the encoded spectral coefficients 1122. Any concept of
143 irsWuto Mexican JR
OF THE hlOHtíMa 'Λ. ¡»» ¿3SE-Λ »
INDUSTRIAL decoding for decoding known spectral coefficients.
Audio decoder 1100 further comprises a linear prediction encoding coefficient for scale factor conversion 1140, which is configured to provide a set of scale factors 1142 based on the encoded representation 1124 of linear prediction encoding coefficients. . For example, the linear prediction coding coefficient for 1142 scale factor conversion can perform functionality that is described in the USAC [Unified Voice and Audio Coding] standard. For example, the coded representation 1124 of the linear prediction coding coefficients may comprise a polynomial representation, which is decoded and converted to a set of scale factors by the linear prediction coding coefficient for the scale factor conversion 1142.
Audio decoder 1100 further comprises scaling 1150, which is configured to apply scale factors 1142 to decoded spectral values 1132, so as to obtain scaled decoded spectral values 1152. Furthermore, audio decoder 1100
144
IMPI
<img file="MX356334B_D0127.tif" />
optionally comprises a processing 1160, which, for example, may correspond to the processing 366 described above, where the processed scaled decoded spectral values 1162 are obtained by the optional processing 1160. The audio decoder 1100 further comprises a frequency domain to time domain transform 1170, which is configured to receive scaled decoded spectral values 1152 (which may correspond to scaled decoded spectral values 362), or processed scaled decoded spectral values 1162 (which may correspond to the processed scaled decoded spectral values 368) and provide, on their basis, a time domain representation 1172, which may correspond to the time domain representation 372 described above. Audio decoder 1100 further comprises a first optional postprocessing 1174, and a second optional postprocessing 1178, which, for example, may correspond, at least in part, to the aforementioned optional postprocessing 376. Accordingly, the audio decoder 1110 (optionally) obtains a post-processed version 1179 of the time domain audio representation 1172.
145
MEXICAN INSTITUTE _ _. . , _, PROPERTY V¿ * a¡ »¡wS5LJ3
The 1100 audio decoder plus 1180 error concealment block qvre— (¿srS — όδΐϊΠ ^ 'ύϊ'3' <3ο to receive the 1172 time domain audio representation, or a post-processed version of it, and the coefficients linear prediction encoding (either in encoded form or in decoded form) and provides, on its basis, 1182 error concealment audio information.
The 1180 error concealment block is configured to provide the error concealment audio information
1182 for hiding a loss of an audio frame after an audio frame encoded in a frequency domain representation using a time domain drive signal, and therefore is similar to 380 error concealment and concealment error 480, and in addition to the 500 error concealment and 600 error concealment.
However, error concealment block 1180 comprises linear predictive coding analysis (LPC)
1184, which is substantially identical to linear predictive coding (LPC) analysis 530. However, linear predictive coding (LPC) analysis 1184 can optionally use linear predictive coding (LPC) coefficients 1124 to facilitate analysis ( in
146
IMPI
MHKtCANO INSTITUTE • T HE INDUSTRIAL MONEDAD
<img file="MX356334B_D0128.tif" />
Comparison with Linear Predictive Coding Analysis (LPC) 530). Linear Predictive Coding Analysis (LPC) 1134 provides a time domain 1186 drive signal, which is substantially identical to the 532 time domain drive signal (and in addition to the 610 time domain drive signal ). Furthermore, the error concealment block 1180 comprises an error concealment.
1188, which, for example, can perform the functionality of blocks 540, 550, 560, 570, 580, 584 of error concealment 500, or which, for example, can perform the functionality of blocks 640, 650, 660, 670, 680, 684 from error concealment 600. However, error concealment block 1180 differs slightly from error concealment 500, and in addition, from error concealment 600. For example, the 1180 error concealment block (comprising linear predictive coding analysis (LPC)
1184) differs from error concealment 500 in terms that linear predictive coding (LPC) coefficients (used for linear predictive coding (LPC) synthesis 580) are not determined by linear predictive coding (LPC) analysis 530, although they are (optionally) received from the bit stream.
Also, the block gives error 1188 concealment, which
147
<img file="MX356334B_D0129.tif" />
comprises linear predictive coding analysis (LPC)
1184 differs from error concealment 600 in terms that past excitation 610 is obtained by linear predictive coding (LPC) analysis 1184, rather than being directly available.
The audio decoder 1100 further comprises a signal combination 1190, which is configured to receive the time domain audio representation 1172, or a post-processed version thereof, and furthermore the error concealment audio information 1182 (naturally, for subsequent audio frames), and combines said signals, preferably using an overlay and add operation, so as to obtain the decoded audio information 1112.
For further details, reference is made to the explanations above.
8. Method according to Fig. 9.
Fig. 9 shows a flow chart of a method for providing decoded audio information based on encoded audio information. The method 900 according to Fig. 9 comprises the provision of 910 an error concealment audio information for the
148
IMPI
MEXICAN INSTITUTE OF PROPERTY concealment of a loss of a plot of auáio '<sup>TO THE</sup>lue§v — ae
<img file="MX356334B_D0130.tif" />
an audio frame encoded in a 3Ξ * frequency domain representation using a time domain drive signal. The 900 method according to Fig. 9 is based on the same considerations as the audio decoder according to Fig. 1. Furthermore, it should be noted that the 900 method can be supplemented by any of the features and functionalities. described in this application, either individually, or in combination.
9. Method according to Fig. 10.
Fig. 10 shows a flowchart of a method of providing decoded audio information based on encoded audio information. Method 1000 comprises providing 1010 an error concealment audio information for concealing a loss of an audio frame, where a time domain drive signal obtained for (or based on) one or more frames The audio that precede a lost audio frame is modified in order to obtain the audio error concealment information.
149
IMPI
<img file="MX356334B_D0131.tif" />
The method 1000 according to Fig. 10 is based on the same considerations as the above mentioned audio decoder according to Fig. 2.
Furthermore, it should be noted that the method according to Fig. 10 can be supplemented by any of the features and functionality described in this application, either individually, or in combination.
10. Additional remarks.
In the embodiments described above, multiple frame losses can be handled in different ways. For example, if two or more frames are lost, the periodic part of the time domain excitation signal for the second lost frame may be derived from (or equal to) a copy of the tonal part of the domain excitation signal. of time associated with the first lost frame. Alternatively, the time domain excitation signal for the lost second frame can be supported by linear predictive coding (LPC) analysis of the lost previous frame synthesis signal. For example, in a code, linear predictive coding (LPC) can be changeable with each frame lost; then, re-running the analysis makes sense for each missing frame.
150
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX356334B_D0132.tif" />
eleven. Implementation alternatives.
Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or to a feature of a method step. Similarly, the aspects described in the context of a method step further represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by said apparatus.
In accordance with certain implementation requirements, the embodiments of the invention can be implemented in hardware or software. Implementation can be done using a digital storage medium, for example, a floppy disk, a DVD (digital versatile disc), a Blu-Ray, a CD (compact disc), a ROM (memory alone
151
IMPI
<img file="MX356334B_D0133.tif" />
read, a PROM (programmable read-only memory), an EPROM (programmable read-only erase memory), an EEPROM (memory programmable read-only electronic erase, or a FLASH memory, which has electronically readable control signals stored there, cooperating (or capable of cooperating) with a programmable computer system in such a way as to carry out the respective method. Therefore, the digital storage medium can be computer readable.
Some embodiments in accordance with the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, in order to carry out one of the methods described herein. request.
In general, the embodiments of the present invention can be implemented as a computer program product with a program code, where the program code is operative to carry out one of the methods when the program product is executed. computer into a computer. The program code can
152
IMPIí
MEXICAN INSTITUTE OF PROPERTY η,, INDUSTRIAL B .., to be stored, for example, in a machine-readable carrier.
Other embodiments comprise the computer program for carrying out one of the methods described in the present application, stored in a machine-readable carrier.
In other words, an embodiment of the method of the invention, therefore, is a computer program having a program code for performing one of the methods described in the present application, when the program is run from computer to computer.
A further embodiment of the method of the invention is therefore a data carrier (or a digital storage medium, or a computer readable medium) comprising, there recorded, the computer program for carrying out one of the methods described in the present application. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transient.
A further embodiment of the method of the invention is therefore a data stream or a sequence of signals representing the computer program to carry out one of the methods that is
<img file="MX356334B_D0134.tif" />
153 βτιτυτυ mlxicano / 2 · “'* --- - - o
INDUSTRIAL <sup>Λ</sup> described in the present application. The industrial signal sequence, for example, may be configured to be transferred via a data communication connection, eg via the Internet.
A further embodiment comprises a processing means, for example a computer, or a programmable logic device configured or adapted to carry out one of the methods described in the present application.
A further embodiment comprises a computer having the computer program installed there to carry out one of the methods described in the present application.
A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for carrying out one of the methods described in this application, to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may comprise, for example, a file server for transferring the computer program to the receiver.
154
In some embodiments<sup>1</sup>
<img file="MX356334B_D0135.tif" />
a
INLHJM RIAL programmable logic device (for example, an array ...... of field programmable gates) to perform some or all of the functionalities of the methods described in the present application. In some embodiments, an array of field programmable gates can cooperate with a microprocessor in order to carry out one of the methods described in the present application.
In general, the methods are preferably carried out by any physical support apparatus.
The apparatus described in the present application can be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
The methods described in this application can be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
The embodiments described above are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the provisions and details described in the present application will be apparent to those skilled in the art.
155 technique. Therefore, it is only for the scope of imminent, and not for the mode of description and realization of the present
IMPI SiL-a
MEXICAN INSTITUTE f OF THE PROPERTY intends<sup>INC</sup>d§<sup>ESTUARY</sup>lim ¥ taTiíón the patent claims specific details presented in explanation of the application forms.
12. Conclusions.
In conclusion, although some concealment for transform domain codes has been described in the field, the embodiments according to the invention outperform conventional codes (or decoders). The embodiments according to the invention use a domain change for concealment (frequency domain to time domain or excitation). Accordingly, the embodiments according to the invention create high quality voice concealment for transform domain decoders.
The transform encoding mode is similar to that in USAC (confer, for example, reference [3]).
It uses the Modified Discrete Cosine Transform (MDCT) as a transform, and spectral noise formation is achieved by applying the spectral envelope of
156
IMPIOS
INSTITUTO MEXICANO heavy linear predictive coding (LPC) <sup>M</sup>e ^ c> ^ MA ^ omÍfiS ^ Se frequency (also known as FDNS, form of rni'do — frequency domain). In other words, the embodiments according to the invention can be used in an audio decoder, which uses the decoding concepts that are described in the USAC standard. However, the concept of error concealment disclosed in this application may also be used in an audio decoder that is of the AAC (Advanced Audio Coding) type, or in any encoding (or decoder) of the AAC family.
The concept according to the present invention applies to a switched codec such as USAC, as well as a pure frequency domain codec. In both cases, concealment is done in either the time domain or the excitation domain.
In the following, some advantages and features of time domain concealment (or excitation domain concealment) will be described.
Conventional transformed encoded excitation (TCX) concealment, as described, for example, with reference to Figs. 7A-7B and 8, also called noise substitution, is not suitable for voice type signals or even for tonal signals. The forms of
157
IMPI
MEXICAN INSTITUTE V .......
embodiment according to the invention concealment for a j-ranyfnrmgda domain codec ??. .
applies in the time domain (or in the excitation domain of a linear prediction coding decoder). It is similar to an ACELP-type concealment (adaptive codebook-excited linear prediction), and increases the quality of the concealment. Height information has been found to be convenient (or even required, in some cases) for an ACELP-type concealment. Therefore, the embodiments according to the present invention are configured to find reliable height values for the previous frame encoded in the frequency domain.
Different parts and details have been explained above, for example, on the basis of the embodiments according to Figs. 5 and 6.
In conclusion, the embodiments according to the invention create an error concealment that overcomes conventional solutions.
158
Bibliography
IMPI ^> 5 iRtflTllTK) MEXICAN ty- ~
Dt THE PROPERTY V VK INDUSTRIAL [1] 3GPP, Audio codee processing functions; Extended Adaptive Multi-Rate - Wideband (AMR-WB +) code; Transcoding functions, 2009, 3GPP TS 26.290.
[2] MDCT-BASED CODER FOR HIGHLY ADAPTIVE SPEECH AND
AUDIO CODING; Guillaume Fuchs & al .; EUSIPCO 2009.
[3] ISO_IEC_DIS_23003-3_ (E); Information technology
<td>MPEG audio coding.</td><td>technologies -</td><td>Part 3:</td><td>Unified</td><td>speech</td><td>and audio</td>
<td> [4]</td><td>3GPP, General</td><td>Audio</td><td>Elbow</td><td>Audio</td><td>Processing</td>
<td>15 functions;</td><td>Enhanced aacPlus</td><td colspan="2">general audio</td><td>elbow</td><td>Additional</td>
<td>decoder</td><td>tools, 2009, 3GPP TS 26.402.</td>
<td> [5]</td><td>Audio decoder and coding error compensating</td>
<td>method,</td><td>2000, EP 1207519 B1</td>
<td> [6]</td><td>Apparatus and method for improved concealment of</td>
the adaptive codebook in ACELP-like concealment employing improved pitch lag estimation, 2014, PCT / EP2014 / 062589
IMPI
<img file="MX356334B_D0136.tif" />
159
I Jk Ai. I j · * '· ..— „_r <j mexican it tA PROPERTY CJaa INDUSTRIAL [7] Apparatus and method for improved concealment of the adaptive codebook in ACELP-like concealment employing improved press resynchronization, 2014, PCT / EP2014 / 062578
160
IMPI
IHStfWTO MEXICANO W THE PROPERTY
INDUSTRIAL
Contents109
148 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111 Sheet 112 Sheet 113 Sheet 114 Sheet 115 Sheet 116 Sheet 117 Sheet 118 Sheet 119 Sheet 120 Sheet 121 Sheet 122 Sheet 123 Sheet 124 Sheet 125 Sheet 126 Sheet 127 Sheet 128 Sheet 129 Sheet 130 Sheet 131 Sheet 132 Sheet 133 Sheet 134 Sheet 135 Sheet 136 Sheet 137 Sheet 138 Sheet 139 Sheet 140 Sheet 141 Sheet 142 Sheet 143 Sheet 144 Sheet 145 Sheet 146 Sheet 147 Sheet 148
190 members in 21 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 13191133 | European Patent Office (EPO) | A | |
| 13191133 | European Patent Office (EPO) | A | |
| 13191133 | European Patent Office (EPO) | – | |
| 14178824 | European Patent Office (EPO) | A | |
| 14178824 | European Patent Office (EPO) | A | |
| 14178824 | European Patent Office (EPO) | – | |
| 2014073035 | European Patent Office (EPO) | W | |
| 2014073035 | European Patent Office (EPO) | W | |
| EP13191133 | – | – | – |
| EP14178824 | – | – | – |
| EP20130191133 | – | – | – |
| EP20140178824 | – | – | – |
| PCTEP2014073035 | – | – | – |
| WO2014EP73035 | – | – | – |
Members190
| Document | Office | Kind | |
|---|---|---|---|
| CA2928974A1 | Canada | A1 | |
| CA2929012A1 | Canada | A1 | |
| CA2984017A1 | Canada | A1 | |
| CA2984030A1 | Canada | A1 | |
| CA2984042A1 | Canada | A1 | |
| CA2984050A1 | Canada | A1 | |
| CA2984066A1 | Canada | A1 | |
| CA2984532A1 | Canada | A1 | |
| CA2984535A1 | Canada | A1 | |
| CA2984562A1 | Canada | A1 | |
| CA2984573A1 | Canada | A1 | |
| WO2015063044A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015063045A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201521016A | Taiwan Province of China | A | |
| TW201523584A | Taiwan Province of China | A | |
| AR098257A1 | Argentina | A1 | |
| AR098258A1 | Argentina | A1 | |
| SG11201603425UA | Singapore | A | |
| SG11201603429SA | Singapore | A | |
| AU2014343905A1 | Australia | A1 | |
| AU2014343904A1 | Australia | A1 | |
| KR20160079056A | Republic of Korea | A | |
| KR20160079849A | Republic of Korea | A | |
| MX2016005535A | Mexico | A | |
| CN105765651A | China | A | |
| CN105793924A | China | A | |
| MX2016005542A | Mexico | A | |
| US2016240203A1 | United States of America | A1 | |
| US2016247506A1 | United States of America | A1 | |
| EP3063759A1 | European Patent Office (EPO) | A1 | |
| EP3063760A1 | European Patent Office (EPO) | A1 | |
| JP2016535867A | Japan | A | |
| JP2016539360A | Japan | A | |
| SG10201609146YA | Singapore | A | |
| SG10201609186UA | Singapore | A | |
| SG10201609218XA | Singapore | A | |
| SG10201609234QA | Singapore | A | |
| SG10201609235UA | Singapore | A | |
| US2016379645A1 | United States of America | A1 | |
| US2016379646A1 | United States of America | A1 | |
| US2016379647A1 | United States of America | A1 | |
| US2016379648A1 | United States of America | A1 | |
| US2016379649A1 | United States of America | A1 | |
| US2016379650A1 | United States of America | A1 | |
| US2016379651A1 | United States of America | A1 | |
| US2016379652A1 | United States of America | A1 | |
| US2016379657A1 | United States of America | A1 | |
| TWI569261B | Taiwan Province of China | B | |
| TWI571864B | Taiwan Province of China | B | |
| BR112016009805A2 | Brazil | A2 | |
| BR112016009819A2 | Brazil | A2 | |
| KR20170117615A | Republic of Korea | A | |
| KR20170117616A | Republic of Korea | A | |
| KR20170117617A | Republic of Korea | A | |
| KR20170118246A | Republic of Korea | A | |
| KR20170118247A | Republic of Korea | A | |
| AU2017251669A1 | Australia | A1 | |
| AU2017251670A1 | Australia | A1 | |
| AU2017251671A1 | Australia | A1 | |
| ZA201603528B | South Africa | B | |
| AU2014343905B2 | Australia | B2 | |
| RU2016121148A | Russian Federation | A | |
| RU2016121172A | Russian Federation | A | |
| AU2017265032A1 | Australia | A1 | |
| AU2017265038A1 | Australia | A1 | |
| EP3063760B1 | European Patent Office (EPO) | B1 | |
| AU2014343904B2 | Australia | B2 | |
| AU2017265060A1 | Australia | A1 | |
| AU2017265062A1 | Australia | A1 | |
| EP3063759B1 | European Patent Office (EPO) | B1 | |
| SG10201709061WA | Singapore | A | |
| SG10201709062UA | Singapore | A | |
| EP3285254A1 | European Patent Office (EPO) | A1 | |
| EP3285255A1 | European Patent Office (EPO) | A1 | |
| EP3285256A1 | European Patent Office (EPO) | A1 | |
| EP3288026A1 | European Patent Office (EPO) | A1 | |
| KR20180023063A | Republic of Korea | A | |
| KR20180026551A | Republic of Korea | A | |
| KR20180026552A | Republic of Korea | A | |
| ES2659838T3 | Spain | T3 | |
| TR201802808T4 | Türkiye | T4 | |
| PT3063759T | Portugal | T | |
| PT3063760T | Portugal | T | |
| ES2661732T3 | Spain | T3 | |
| JP6306175B2 | Japan | B2 | |
| JP6306177B2 | Japan | B2 | |
| US2018114533A1 | United States of America | A1 | |
| KR101854296B1 | Republic of Korea | B1 | |
| MX356036B | Mexico | B | |
| MX356334BThis record | Mexico | B | |
| PL3063760T3 | Poland | T3 | |
| KR101854297B1 | Republic of Korea | B1 | |
| EP3336839A1 | European Patent Office (EPO) | A1 | |
| EP3336840A1 | European Patent Office (EPO) | A1 | |
| EP3336841A1 | European Patent Office (EPO) | A1 | |
| PL3063759T3 | Poland | T3 | |
| EP3355305A1 | European Patent Office (EPO) | A1 | |
| EP3355306A1 | European Patent Office (EPO) | A1 | |
| RU2667029C2 | Russian Federation | C2 | |
| AU2017265032B2 | Australia | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 356334
- Publication, DOCDB
- 356334
- Publication, EPODOC
- MX356334
- Application
- 2016005535
- Application, DOCDB
- 2016005535
- Application, EPODOC
- MX20160005535
Titles
- Spanish
- DECODIFICADOR DE AUDIO Y METODO PARA PROVEER UNA INFORMACION DE AUDIO DECODIFICADA USANDO UN OCULTAMIENTO DE ERROR SOBRE LA BASE DE UNA SEÑAL DE EXCITACION DE DOMINIO DE TIEMPO.
Classification
- CPC, 10
- G10L19/005
- G10L19/02
- G10L19/08
- G10L19/008
- G10L19/0212
- G10L19/04
- G10L19/09
- G10L19/125
- G10L25/90
- G10L2019/0011
- IPC, 4
- G10L19 005
- G10L19 02
- G10L19 08
- G10L25 90