Apparatus and method for encoding and decoding an encoded audio signal using temporal noise/patch shaping.
Abstract
An apparatus for decoding an encoded audio signal, comprises: a spectral domain audio decoder (602) for generating a first decoded representation of a first set of first spectral portions being spectral prediction residual values; a frequency regenerator (604) for generating a reconstructed second spectral portion using a first spectral portion of the first set of first spectral portions, wherein the reconstructed second spectral portion additionally comprises spectral prediction residual values; and an inverse prediction filter (606) for performing an inverse prediction over frequency using the spectral residual values for the first set of first spectral portions and the reconstructed second spectral portion using prediction filter information (607) included in the encoded audio signal.

Term
7.8 yearsleft in the term
Expires 15 July 2034.
- Priority
- Filed
- Granted
- Today
- Expires
6 claims: 2 independent, 4 dependent
- 1REIVINDICACIONES IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL 1. Un aparato para decodificar codificada, el cual comprende:un decodificador de audio de dominio espectral (602) para generar una primera representación decodificada de un primer conjunto de primeras porciones espectrales que son los valores residuales de predicción espectral;un regenerador de frecuencia (604) para generar una segunda porción espectral reconstruida utilizando una primera porción espectral del· primer conjunto de primeras porciones espectrales, en el que la segunda porción espectral reconstruida y el primer conjunto de primeras porciones espectrales comprenden valores residuales de predicción espectral;y un filtro de predicción inversa (606, 616, 626) para llevar a cabo una predicción inversa sobre la frecuencia utilizando los valores residuales de predicción espectral para el primer conjunto de primeras porciones espectrales y la segunda porción espectral reconstruida utilizando la información del filtro de predicción (607) incluida en la señal de audio codificada.
- 2Un aparato de acuerdo con la reivindicación 1, una señal de audio 127 que comprende además un modelador INSTITUTO MEXICANO DF LA PROHEDAD INDUSTRIAL _ de la envolvente espectral (614) para modelar una envolvente espectral de una señal de entrada o una señal de salida del filtro de predicción inversa (606).
- 3Un aparato de acuerdo con la reivindicación 2, en el que la señal de audio codificada comprende información de la envolvente espectral para la segunda porción espectral, la información de la envolvente espectral tiene una segunda resolución espectral, la segunda resolución espectral es más baja que una primera resolución espectral asociada con la primera representación decodificada, en el que el modelador de la envolvente espectral (624) está configurado para aplicar una operación de modelado de la envolvente espectral en la salida del filtro de predicción inversa (622), en el que la información del filtro (607) se ha determinado utilizando una señal de audio antes del filtrado de predicción, o en el que el modelador de la envolvente espectral (614) está configurado para aplicar una operación de modelado de la envolvente espectral en la entrada del filtro de predicción inversa (616), cuando la información del filtro de predicción (607) se ha determinado utilizando una señal de audio luego de un filtrado de predicción en un codificador. 128
- 4Un aparato de reivindicaciones anteriores, acuerdo IMPI INSTITUTO MEXICANO Τ4Γ***ϋ·ί3?**ΰ DE LA PROPIEDAD con IN[ W Tte^-^as reivindicaciones anteriores, en el que el decodificador de audio en el dominio espectral (602) está configurado para generar la primera representación decodificada de manera que la primera representación decodificada tiene una frecuencia de Nyquist igual a una tasa de muestreo de una señal de dominio temporal generada por la conversión de tiempo-frecuencia de una salida del filtro de predicción inversa (606). 7. Un aparato de acuerdo con una de las reivindicaciones anteriores, IMPI INSTITUT· MtXICANO •E LA FR FliDAB en el que el decodificador de audio en^eT espectral (602) está configurado de modo que una úrecuencia máxima representada por un valor espectral para la frecuencia máxima en la primera representación decodificada es igual a una frecuencia máxima incluida en una representación de tiempo generada por la conversión de frecuencia-tiempo de una reivindicaciones anteriores, en el que la primera representación decodificada de un primer conjunto de primeras porciones espectrales comprende valores espectrales reales, en el que el aparato comprende además un estimador (724) para estimar valores imaginarios para el primer conjunto de primeras porciones espectrales a partir del primer conjunto de valor real de primeras porciones espectrales, y en el que el filtro de predicción inversa (728) es un filtro complejo de predicción inversa definido por la información del filtro de predicción de valor complejo (714), y en el que el aparato comprende además un convertidor de frecuencia-tiempo (730) configurado para llevar a cabo una conversión de un 130 espectro de valor complejo en una temporal. IMPI INSTITUTO MSJDCANC DE LA FROFUOAD señal de aíftft'ÉP'AHe io 9. Un aparato de acuerdo con una de las reivindicaciones anteriores, en el que el filtro de predicción inversa (606, 728) está configurado para aplicar una pluralidad de subfiltros, en el que un límite de frecuencia de cada subfiltro coincide con un límite de frecuencia de una banda de reconstrucción que coincide con un recuadro de frecuencia. espectral; un filtro de predicción (704) para llevar a cabo una predicción sobre la frecuencia en la representación espectral para generar valores residuales espectrales, el filtro de predicción está definido por la información del filtro derivada de la señal de audio; un codificador de audio (708) para codificar un primer conjunto de primeras porciones espectrales de los valores residuales espectrales para obtener un primer conjunto 131 IMPI INSTITUTO mejicano OE LA FROFieOAD codificado de primeros valores espectrales qué UST *£i< primera resolución espectral; un decodificador paramétrico (706) para codificar paramétricamente un segundo conjunto de segundas porciones espectrales de los valores residuales espectrales o de valores de la representación espectral con una segunda resolución espectral que es más baja que la primera resolución espectral; y una interfaz de salida (710) para emitir una señal codificada que comprende el segundo conjunto codificado, el primer conjunto codificado y la información del filtro (714). 11. Un aparato de acuerdo con la reivindicación 10, en el que el convertidor de tiempo-espectro está configurado para llevar a cabo una transformada coseno discreta modificada, y en el que los valores residuales espectrales son valores residuales espectrales de la transformada coseno discreta modificada. 12. Un aparato de acuerdo con una de las reivindicaciones 10 y 11, en el que el filtro de predicción (704) comprende una calculadora de información del filtro, la calculadora de información del filtro está configurada para utilizar valores espectrales de una representación espectral para calcular la 132 información del filtro y en el que el filtróle ''^predicción IMPI INSTITUTO MEXICANO DE LA PROPIEDAD está configurado para calcular los valores residuales la de espectrales utilizando valores espectrales representación espectral, en el que los valores espectrales para calcular la información del filtro y los valores espectrales introducidos en el filtro de predicción se obtienen de la misma señal de audio. 13. Un aparato de acuerdo con una las reivindicaciones 10 a 12, en el que el filtro de predicción comprende una calculadora de filtros para calcular la información del filtro utilizando valores espectrales de una frecuencia de inicio de TNS a una frecuencia de fin de TNS, en el que la frecuencia de inicio de TNS es inferior a 4 kHz y la frecuencia de fin de TNS es superior a 9 kHz. 14. Un aparato de acuerdo con una de las reivindicaciones 10 a 13, que comprende además un analizador (102, 706) para determinar el primer conjunto de primeras porciones espectrales para ser codificadas por el codificador de audio (708), el analizador utiliza una frecuencia de relleno de espacios, en el que las porciones espectrales por debajo de la frecuencia de inicio de relleno de espacios son primeras porciones espectrales, y 133 IMPI INSTITUTO MEXICANO en el que la frecuencia de fin de la frecuencia de relleno de espacios. 15. Un aparato de acuerdo con una las reivindicaciones 10 a 14, en el que el convertidor de tiempo-frecuencia (702) está configurado para proporcionar una representación espectral compleja, en el que el filtro de predicción está configurado para llevar a cabo una predicción sobre la frecuencia con la representación espectral de valor complejo, y en el que la información del filtro (714) está configurada para definir un filtro complejo de predicción inversa. 15. Un método para decodificar una señal de audio codificada, el cual comprende:generar (502) una primera representación decodificada de un primer conjunto de primeras porciones espectrales que son los valores residuales de predicción espectral;regenerar (604) una segunda porción espectral reconstruida utilizando una primera porción espectral del primer conjunto de primeras porciones espectrales, en el que la segunda porción espectral reconstruida y el primer 134 τ -w r -r^ IMPI INSTITUTO MEXICANO DE IA EROME DA O . , . [NDUSTRIAL conjunto de primeras porciones espectrales comprenden residuales de predicción espectral;y llevar a cabo (606, 616, 626) una predicción inversa sobre la frecuencia utilizando los valores residuales de predicción espectral para el primer conjunto de primeras porciones espectrales y la segunda porción espectral reconstruida utilizando la información del filtro de predicción (607) incluida en la señal de audio codificada. que comprende además un modelador de la envolvente espectral (614) para modelar una envolvente espectral de una señal de entrada o una señal de salida del filtro de predicción inversa (606). 17. El método de acuerdo con la reivindicación 16, en el que la señal de audio codificada comprende información de la envolvente espectral para la segunda porción espectral, la información de la envolvente espectral tiene una segunda resolución espectral, la segunda resolución espectral es más baja que una primera resolución espectral asociada con la primera representación decodificada, en el que la regeneración (604) comprende un modelado de la envolvente espectral (624) que comprende una operación de modelado de la envolvente espectral en una salida del paso que lleva a cabo (606, 616, 626) una predicción inversa sobre 135 IMPI» INSTITUTO MEXICANO »t LA FROFIEDAD la frecuencia, en el que la información del filtY’^'tfeO/f—?e ha determinado utilizando una señal de áüd'ió áhfes cTeT filtrado de predicción, o en el que la regeneración (604) comprende un modelado de 5 la envolvente espectral (624) que comprende aplicar una operación de modelado de la envolvente espectral en una entrada del paso que lleva a cabo (606, 616, 626) una predicción inversa sobre la frecuencia, en el que la información del filtro de predicción (607) se ha determinado en la representación espectral para generar valores residuales espectrales, el filtro de predicción está definido por la información del filtro derivada de la señal de audio;20 codificar (708) un primer conjunto de primeras porciones espectrales de los valores residuales espectrales para obtener un primer conjunto codificado de primeros valores espectrales que tienen una primera resolución espectral;136 IMPI INSTITUTO MEXICANO DI LA MONEDAD codificar paramétricamente (706) un segúnüfeusTsanjoSwyMre segundas porciones espectrales de los ^-g-τ ¡, i ljrsoldnrl oc espectrales o de valores de la representación espectral con una segunda resolución espectral que es más baja que la
- 55 primera resolución espectral;y emitir (710) una señal codificada que comprende el segundo conjunto codificado, el primer conjunto codificado y la información del filtro (714). 19. Un medio leíble por computadora para decodificar
- 610 una señal de audio, que comprende el método de la reivindicación 16. 20. Un medio leíble por computadora para codificar una señal de audio, que comprende el método de la reivindicación 18 . 137
Independent claims6
817 paragraphs in 154 sections, as filed
(54) Title: APPARATUS AND METHOD FOR CODING AND DECODING A CODED AUDIO SIGNAL USING TEMPORARY / PATCH NOISE MODELING.
(54) Title: APPARATUS AND METHOD FOR ENCODING AND DECODING AN ENCODED AUDIO SIGNAL USING TEMPORAL NOISE / PATCH SHAPING.
(57) Summary
An apparatus for decoding an encoded audio signal, comprising: a spectral domain audio decoder (602) for generating a first decoded representation of a first set of first spectral portions that are spectral prediction residuals; a frequency regenerator (604) for generating a second reconstructed spectral portion using a first spectral portion of the first set of first spectral portions, wherein the second reconstructed spectral portion additionally comprises spectral prediction residual values; and an inverse prediction filter (606) to perform an inverse prediction on frequency using the residual spectral values for the first set of first spectral portions and the second reconstructed spectral portion using the included prediction filter information (607) on the encoded audio signal.
(57) Abstract
An apparatus for decoding an encoded audio signal, comprises: a spectral domain audio decoder (602) for generating a first decoded representation of a first set of first spectral portions being spectral prediction residual values; a frequency regenerator (604) for generating a reconstructed second spectral portion using a first spectral portion of the first set of first spectral portions, where the reconstructed second spectral portion additionally comprises spectral prediction residual values; and an inverse prediction filter (606) for performing an inverse prediction over frequency using the spectral residual values for the first set of first spectral portions and the reconstructed second spectral portion using prediction filter Information (607) included in the encoded audio signal.
<img file="MX340575B_D0001.tif" />
PATENT TITLE NO. 340575 _SE_
SteeiMfc I heard KSNOMIA
Headlines):
Home:
Denomination:
Classification:
Inventor (s):
Mexican Institute of Industrial Property
<img file="MX340575B_D0002.tif" />
FRAUNHOFER-GESELLSCHAFT ZUR FÚRDERUNG DER ANGEWANDTEN FORSCHUNG EV
Hansastrasse 27C, 80686, Munich, GERMANY
APPARATUS AND METHOD FOR CODING AND DECODING A CODED AUDIO SIGNAL USING TEMPORARY / PATCH NOISE MODELING lnt.CI.8: G10L19 / 02; G10L19 / 028; G10L19 / 03; G10L21 / 0388
SASCHA DISCH; FREDERIK NAGEL; RALF GEIGER; BALAJI NAGENDRAN THOSHKAHNA; KONSTANTIN SCHMIDT; STEFAN BAYER; CHRISTIAN
NEUKAM; BERND EDLER; CHRISTIAN HELMRICH
REQUEST
MX / a / 2015/004022
International filing date:
July 2014
PRIORITY
<td>Country:</td><td>Date:</td><td>Number:</td>
<td>EP</td><td>July 22, 2013</td><td> 13177353.3</td>
<td>EP</td><td>July 22, 2013</td><td> 13177350.9</td>
<td>EP</td><td>July 22, 2013</td><td> 13177348.3</td>
<td>EP</td><td>July 22, 2013</td><td> 13177346.7</td>
<td>EP</td><td>October 18, 2013</td><td> 13189358.8</td>
Validity: Twenty years<sub>:</sub>
Expiration Date: July 15, 2034
The reference patent is granted based on articles 1, 2, section V. 6, section III, and 50 of the Industrial Property Law.
Oe in accordance with article 23 tle the Accounting Law from the right of presentation of rights.
Industrial Property, this patent has a non-extendable term of twenty years, of the international application and will be subject to the payment of the fee to keep the
Whoever subscribes the presartte title to does so based on Ό provided by read articles 6 fractions lll and 7 bis bis 2 of the Industrial Property Law (Official Mario de la Federación (DOF) 27 / 08M9W, ietoeiAtfcf 08/02/1994 , 10/25/1996, 12/26/19 S07, 05/17/1999, ¢ 6/01/2004 06/16/2005, 01/25/2006, 06/05/2009 / 06/01/2010, 18 / 06/2010, 06/28/2010, 01/27/2012 and 04/09/2012), articles 1 *. 3rd section V Wdfee e, 4 "and 45 * wagones i yn« á defWmwfito deFtnewt »Merfeane« fe te Pwpiéded MMM {00 4 «« «» », amended on 07/01/2002, 07/15/2004, 28 / 07/2004 and 7/09/2007); Articles 1, 3, 4, 5, section V, Section a), 16 sections I and III and 30 of the Organic Statute of the Mexican Institute of Industrial Property (DOF) 12/27/1999, amended on 10/10/2002, 07/29/2004, 08/04/2004 and 09/13/2007); 1, 3 and 5 Clause a) of the Agreement that delegates powers to the Deputy Directors General, Coordinator, Divisional Directors, Holders of the Regional Offices, Divisional Deputy Directors, Departmental Coordinators and other subordinates of the Mexican Institute of Industrial Property. (DOF 12/15/1999, amended on 02/04/2000, 07/29/2004, 08/04/2004 and 09/13/2007).
<img file="MX340575B_D0003.tif" />
Arenal No. 550. Floor 1.
People Santa María Tepepan. Xochimilco. CP 16C2C,
Mexico City
Ί. (55) 53 34 07 00 www.impi flob.rnx
Issue Date: July 13, 2016
DIVISIONAL DIRECTOR OF PATENTS
<img file="MX340575B_D0004.tif" />
NAHANNY CANAL REYES
<img file="MX340575B_D0005.tif" />
MX / 2015/54902
<img file="MX340575B_D0006.tif" />
<img file="MX340575B_D0007.tif" />
<img file="MX340575B_D0008.tif" />
<img file="MX340575B_D0009.tif" />
MEXICAN INSTITUTE OF THE INDUSTRIAL FROHT.DAD
DEVICE AND METHOD FOR CODING AND DECODING A SIGNAL OF
CODED AUDIO USING TEMPORARY / SOUND NOISE MODELING
PATCH
DESCRIPTION
The present invention relates to the coding / decoding of the audlo and, in particular, to the encoding of the audlo using Intelligent Space Fill (IGF). The encoding of the audlo is the domain of signal compression that deals with the exploitation of the redundancy and irrelevance of audio signals using psychoacoustic knowledge. Currently, audio codes generally need around 60 kbps / channel for transparent perceptual encoding of almost any type of audio signal. Newer codes aim to reduce the encoding bit rate by taking advantage of spectral similarities in the signal and using techniques such as bandwidth extension (BWE). A bandwidth extension scheme (BWE) uses a low bit rate parameter set to represent the high frequency components (HF) of an audio signal. The high frequency spectrum (HF) is populated with the low frequency ruler spectral content (LF) and the spectral shape, slope, and time continuity are adjusted to maintain the timbre and color of the original signal. These bandwidth extension (BWE) methods allow audio codes to retain good quality at even low bit rates of around 24 kbps / channel.
.rjk js. MEXICAN INSTITUTE
SAY INDUSTRIAL PROPERTY
<img file="MX340575B_D0010.tif" />
The storage or transmission of audio signals is often subject to strict bit rate limitations. In the past, encoders have been forced to dramatically reduce the transmitted audio bandwidth when only a very low bit rate was available.
Modern audio codes are now capable of encoding broadband signals using Bandwidth Extension Methods (BWE) [1]. These algorithms are based on a parametric representation of the high-frequency (HF) content - which is generated from the wave-encoded low-frequency (LF) portion of the decoded signal by transposing it to the high-spectral region. frequency (HF) (interconnect) and the application of parameter-based post-processing. In bandwidth extension schemes (BWE), the reconstruction of the high frequency spectral region (HF) above a so-called determined crossover frequency is often based on spectral interconnection. In general, the high frequency region (HF) consists of multiple adjacent connections and each of these connections is derived from bandpass regions (BP) of the low frequency spectrum (LF) below the determined crossover frequency. Current state of the art systems efficiently perform interconnection within a representation of filter banks, for example Quadrature Mirror Filter Bank (QMF), copying a set of coefficients of adjacent sub-bands from a region of origin to the region of destination.
Another technique found in current audio codes that increases compression efficiency and thus allows for audio bandwidth
IMPI
INSTITUTO MEXICANO I heard the extended industrial source at low bit rates is the parameter-based synthetic replacement of appropriate parts of audio spectra. For example, parts of the noise-like signal of the original audio signal can be replaced without substantial loss of subjective quality by artificial noise generated in the decoder and scaled by side information parameters. An example is the Perceptual Noise Substitution tool (PNS) contained in the Advanced Audio Coding MPEG-4 (AAC) [5].
Another provision that also allows for extended audio bandwidth at low bit rates is the noise filler technique contained in the Unified MPEG-D Audio and Voice Coding System (USAC) [7] . The spectral spaces (zeros) that are deduced by the dead zone of the quantizer due to too thick quantization, are subsequently filled with artificial noise in the decoder and scaled by post-processing based on parameters.
Another state-of-the-art system is called Precise Spectral Replacement (ASR) [2-4]. In addition to waveform coding, Precise Spectral Replacement (ASR) employs a dedicated signal synthesis stage that perceptually restores significant sinusoidal portions of the signal to the decoder. Also, a system described in [5] relies on sinusoidal modeling in the high frequency region (HF) of a waveform encoder to allow the extended audio bandwidth to have adequate perceptual quality at bit rates low. Everybody
<img file="MX340575B_D0011.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0012.tif" />
these methods involve transforming the data into a second domain apart from the Modified Discrete Cosine Transform (MDCT) and also quite complex analysis / synthesis steps for the conservation of high-frequency sinusoidal (HF) components.
Fig. 13a illustrates a schematic diagram of an audio encoder for a bandwidth extension technology such as that used in Advanced High Efficiency Audio Coding (HE-AAC). English). An audio signal on line 1300 is input to a filter system comprising a low pass 1302 and a high pass 1304. The signal emitted by the high pass filter 1304 is input to a parameter extractor / encoder 1306. Parameter extractor / encoder 1306 is configured to calculate and encode parameters such as, for example, a spectral envelope parameter, a noise addition parameter, a missing harmonics parameter, or a reverse filter parameter. These extracted parameters are input to a bit stream multiplexer 1308. The low-pass output signal is input to a processor that generally comprises the functionality of a downstream sampler 1310 and a central encoder 1312. Low-pass 1302 restricts bandwidth to be encoded at significantly less bandwidth than produced on the original input audio signal on line 1300. This provides a significant encoding gain due to the fact that all the functionalities that occur in the central encoder only have to operate on a signal with low bandwidth. When, for example, the signal bandwidth of
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0013.tif" />
audio on line 1300 is 20 kHz and when low pass filter 1302 has an exemplary bandwidth of 4 kHz, in order to fulfill the sampling theorem, it is theoretically sufficient that the signal subsequent to the downstream sampler have a sampling rate of 8 kHz, which is a substantial reduction in the sampling rate required for the 1300 audio signal that has to be at least 40 kHz.
Fig. 13b illustrates a schematic diagram of a respective bandwidth extension decoder. The decoder comprises a 1320 bit stream multiplexer (sic). The bitstream demultiplexer 1320 extracts an input signal for a central decoder 1322 and an input signal for a parameter decoder 1324. An output signal from the central decoder has, in the example above, a sampling rate of 8 kHz and therefore a bandwidth of 4 kHz whereas, for a complete reconstruction of bandwidth, the output signal of a 1330 high frequency rebuilder must be 20 kHz which requires a sampling rate of at least 40 kHz. In order to make this possible, a decoder processor that has the functionality of an upstream sampler 1325 and a filter bank 1326 is required. The high frequency reconstructor 1330 then receives the low frequency signal analyzed by frequency emitted by the bank. of filters 1326 and reconstructs the frequency range defined by the high-pass filter 1304 of Fig. 13a using the parametric representation of the high-frequency band. The 1330 High Frequency Rebuilder has various functionalities such as regeneration of the upper frequency range using the range
IMPI
MEXICAN INSTITUTE ΠΕ THE PROPERTY i
<img file="MX340575B_D0014.tif" />
source in the low-frequency range, a wrap-around setting ^ triant noise addition functionality and a functionality to introduce missing harmonics in the upper frequency range and, if applied and calculated in the encoder of Fig. 13a, a reverse filtering operation in order to account for the fact that the upper frequency range is usually not as tonal as the lower frequency range. In Advanced High Efficiency Audio Coding (HE-AAC), the missing harmonics are re-synthesized on the decoder side and placed exactly in the middle of a reconstruction band. Therefore, all missing harmonic lines that have been determined in a certain reconstruction band are not placed at the frequency values where they were located in the original signal. Instead, these missing harmonic lines are placed at frequencies in the center of the given band. Therefore, when a missing harmonic line in the original signal was placed very close to the limit of the reconstruction band in the original signal 15, the error in the frequency introduced by placing this missing harmonic line in the reconstructed signal in the The center of the band is close to 50% of the individual reconstruction band, for which parameters have been generated and transmitted.
Furthermore, despite the fact that typical central audio encoders operate in the spectral domain, the central decoder nevertheless generates a time domain signal which is then converted back to a spectral domain by the filter bank functionality. 1326. This introduces additional processing delays, may introduce faults due to processing in
IMPI
MEXICAN INSTITUTE BE LA FROFISBAL
INDUSTRIAL
<img file="MX340575B_D0015.tif" />
tandem transformation first of the spectral domain into the frequency domain and again transformation into generally a different frequency domain and of course this also requires a substantial amount of computational complexity and therefore electrical power, which basically represents a problem when applying bandwidth extension technology on mobile devices, such as mobile phones, tablets or laptops etc.
Current audio codes perform low-bit-rate audio encoding using Bandwidth Extension (BWE) as an integral part of the encoding scheme. However, bandwidth extension (BWE) techniques are limited to replacing only high-frequency (HF) content. They also do not allow waveform encoding of perceptually important content above a certain crossover frequency. Therefore, contemporary audio codecs either lose high-frequency detail (HF) or timbre when bandwidth extension (BWE) is implemented, since the exact alignment of the signal's tonal harmonics is not account in most systems.
Another disadvantage of the current state of the art bandwidth extension systems (BWE) is the need to transform the audio signal into a new domain for the implementation of the BWE (for example, transformation of the Cosine Transform). Discrete Modified (MDCT) to the Quadrature Mirror Filters (QMF) domain, which generates complications of
Γ ί
MEXICAN INSTITUTE 5 '·' <Λί
PROPERTY C ¿X
INDUSTRIAL 'timing, additional computational complexity, and increased memory requirements.
In particular, if a bandwidth extension system is implemented in a filter bank or the domain of the time-frequency transform, there is only a limited possibility to control the temporal form of the bandwidth extension signal. Generally, temporal granularity is limited by the jump size used between adjacent windows of the transform. This can generate unwanted pre or post echoes in the spectral range of the bandwidth extension. In order to increase temporal granularity, shorter bandwidth extension frames or shorter hop sizes can be used, although this causes a bit rate overload due to the fact that, during a certain period of time, it must transmit a larger number of parameters, generally a certain set of parameters for each time frame. Otherwise, if the individual time frames become too large, then pre and post echoes are generated particularly for the transition portions of an audio signal.
An object of the present invention is to provide an improved encoding / decoding concept.
This objective is achieved by an apparatus for decoding an encoded audio signal according to claim 1, an apparatus for encoding an audio signal according to claim 10, a decoding method
<img file="MX340575B_D0016.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0017.tif" />
according to claim 16, a coding method according to claim 18 or a computer program according to claim 19.
The present invention is based on the discovery that an improvement in quality and reduction in bit rate specifically for signals comprising transitional portions, as is very often the case with audio signals, is obtained by combining Modeling technology. Temporal Noise (TNS) or Temporal Box Modeling (TTS) with high-frequency reconstruction. Temporal Noise Modeling (TNS) / Temporal Box Modeling (TTS) processing on the encoder side Implemented by a frequency prediction reconstructs the time envelope of the audio signal. Depending on the implementation, that is, when the temporal noise shaping filter is determined within a frequency range not only encompassing the source frequency range but also the destination frequency range to be reconstructed in a regeneration decoder frequency, the time envelope is applied not only to the center audio signal up to a gap fill start frequency, but the time envelope is also applied to the spectral ranges of reconstructed second spectral portions. Therefore, the pre-echoes or post-echoes that would occur without time-frame modeling are reduced or eliminated. This is accomplished by applying an inverse frequency prediction not only within the center frequency range up to a certain gap fill start frequency but also within a frequency range above the center frequency range. To this end, the
<img file="MX340575B_D0018.tif" />
Frequency regeneration or generation of frequency frames are carried out on the decoder side before applying a frequency prediction. However, the frequency prediction can be applied either before or after the spectral envelope modeling depending on whether the calculation of power information has been carried out on the residual post-filter spectral values or on the spectral values (complete ) before modeling the envelope.
The processing of Temporal Box Modeling (TTS) on one or more frequency boxes additionally establishes a correlation continuity between the origin range and the reconstruction range or in two adjacent reconstruction ranges or frequency boxes.
In one implementation it is preferred to use complex Temporal Noise Modeling (TNS) / Temporal Box Modeling (TTS) filtering. This prevents (temporary) overlap failures of a critically sampled real representation, such as Modified Discrete Cosine Transform (MDCT). A complex Temporal Noise Modeling (TNS) filter can be computed on the encoder side by applying not only a modified discrete cosine transform but also a modified discrete sinusoidal transform to further obtain a complex modified transform. However, only the values of the modified discrete cosine transform, that is, the actual part of the complex transform, are transmitted. However, on the decoder side, it is possible to estimate the imaginary part of the transform using the MDCT spectra from previous or subsequent tables so that, in the
<img file="MX340575B_D0019.tif" />
<img file="MX340575B_D0020.tif" />
INSTITUTO MEXICANO E »LA PROREDAf INDUSTRIAL side of the decoder, the complex filter can be applied again in the inverse prediction on the frequency and, specifically, the prediction on the limit between the origin range and the reconstruction range and also on the limit between the adjacent frequency boxes of the frequency within the reconstruction range.
A further aspect is based on the discovery that the problems related to bandwidth extension separation on the one hand and central coding on the other hand can be addressed and overcome by carrying out the bandwidth extension on the same spectral domain in which the central decoder operates. Therefore, a full rate central decoder is provided that encodes and decodes the entire range of the audio signal. This does not create the need for a downstream sampler on the encoder side and an upstream sampler on the decoder side. Instead, all processing is done across the entire sampling rate or across the entire bandwidth domain. In order to obtain a high coding gain, the audio signal is analyzed in order to find a first set of first spectral portions that has to be encoded with a high resolution, where this first set of first spectral portions can include , in one embodiment, tonal portions of the audio signal. On the other hand, the non-tonal or noisy components in the audio signal that constitute a second set of second spectral portions are parametrically encoded with low spectral resolution. So the encoded audio signal only requires that the first set of first spectral portions
IMí r 1
MEXICAN INSTITUTE -> 1
FROM OWN <sub>D</sub> <%, '* <sup>T</sup>¡~\ -. 3
INDUSTRIAL 'is coded so as to preserve the waveform with high spectral resolution and, additionally, that the second set of second spectral portions is parametrically coded with low resolution using frequency frames obtained from the first set. On the decoder side, the center decoder, which is a full-band decoder, reconstructs the first set of first spectral portions in a way that preserves the waveform, i.e., ignoring whether there is additional frequency feedback. However, the spectrum thus generated has a lot of spectral spaces. These spaces are subsequently filled with the invention's Intelligent Space Fill (IGF) technology using a frequency regeneration that applies parametric data on the one hand, and that uses a source spectral range, i.e. first spectral portions reconstructed by the full-rate audio decoder, on the other hand.
In other embodiments, the spectral portions, which are reconstructed by noise padding only rather than bandwidth replication or frequency box padding, constitute a third set of third spectral portions. Due to the fact that the concept of encoding operates in a single domain for central encoding / decoding on the one hand, and frequency regeneration on the other hand, Intelligent Space Fill (IGF) is not only limited to filling a range of higher frequency but can fill lower frequency ranges by either noise filling without
IΜ ΡI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY frequency regeneration or by frequency regeneration using a frequency box in a different frequency range.
Likewise, it is insisted that an information on spectral powers, an information on individual powers or an information on individual power, an information on a conservation power or an information of conservation power, an information on a box power or an information of box power, or a missing power information or a missing power information may comprise not only a power value but also an amplitude value (eg absolute), a level value or any other value, from which a value can be obtained of final power. Therefore, the information on a power may comprise, for example, the power value itself, and / or a value of a level and / or of an amplitude and / or of an absolute amplitude.
An additional aspect is based on the discovery that the correlation situation is not only important for the origin range, but is also important for the destination range. Furthermore, the present invention recognizes the fact that different correlation situations can occur in the origin range and in the destination range. When considering, for example, a voice signal with high frequency noise, it may happen that the low frequency band that comprises the voice signal with a small number of higher harmonics is highly correlated in the left channel and the right channel, when the speaker is placed in the middle. The high frequency portion, however, can be strongly correlated due to the fact that there could be a noise from
<img file="MX340575B_D0021.tif" />
IMPI
INSTITUTO MEXICANO DF LA PROPIEDAD INDUSTRIAL different high frequency noise on the left side compared to other high frequency noise or no high frequency noise on the right side. Therefore, when a simple space fill operation is performed that ignores this situation, then the high frequency portion would be correlated as well, and this could lead to severe spatial segregation failures in the reconstructed signal. In order to address this question, the parametric data for a reconstruction band or, in general, for the second set of second spectral portions that has to be reconstructed using a first set of first spectral portions, are computed to identify either a first or a second representation of two channels, different for the second spectral portion or, in other words, for the reconstruction band. Therefore, on the encoder side, a two-channel identification is calculated for the second spectral portions, that is, for the portions for which the power information for the reconstruction bands is additionally calculated. Next, a frequency regenerator on the decoder side regenerates a second spectral portion based on a first portion of the first set of first spectral portions, i.e. the origin range and parametric data for the second portion such as Information of the spectral envelope power or any other data of the spectral envelope and additionally depends on the identification of two channels for the second portion, i.e. for this reconstruction band under reconsideration.
<img file="MX340575B_D0022.tif" />
IMPI
MEXICAN INSTITUTE Dt THE PROPERTY
INDUSTRIAL
<img file="MX340575B_D0023.tif" />
The identification of two channels is preferably transmitted as a label for each reconstruction band and this data is transmitted from an encoder to a decoder and the decoder then decodes the core signal preferably indicated by tags calculated for the core bands. Then, in one implementation, the center signal is stored in both stereo representations (for example, left / right and middle / lateral) and for the filling of intelligent space filling frequency (IGF) boxes, the representation is chosen of source boxes to adapt the representation of target boxes indicated by the two-channel identification labels for the intelligent filling of spaces or the reconstruction bands, that is, for the target range.
It is emphasized that this procedure not only works for stereo signals, that is, for a left channel and the right channel, but also works for multi-channel signals. In the case of multichannel signals, several different pairs of channels can be processed, for example as a Left channel and a right channel as a first pair, a left surround channel and a right surround channel as the second pair and a central channel and a LFE channel as the third pair. Other pair formations can be determined for the upper output channel formats such as 7.1, 11.1 and so on.
An additional aspect is based on the discovery that certain alterations in audio quality can be solved by applying an adaptive frequency frame filling scheme of the signal. For
IMPIg
MEXICAN INSTITUTE OF THE industrial rXOFIEDAP
<img file="MX340575B_D0024.tif" />
For this purpose, an analysis is carried out on the encoder side in order to find the best adaptation originating line candidate for a certain destination region. An adaptation information that identifies for a destination region a certain source region together with optionally some additional information is generated and transmitted as lateral information to the decoder. The decoder then applies a frequency frame fill operation using the matching information. For this purpose, the decoder reads the adaptation information from the transmitted data stream or data file and accesses the identified origin region for a given reconstruction band and, if indicated in the adaptation information, leads to In addition, perform some processing of this source screed data to generate raw spectral data for the reconstruction band. Then, this result of the frequency frame filling operation, that is, the raw spectral data for the reconstruction band, is modeled using information of the spectral envelope in order to finally obtain a reconstruction band that includes the first spectral portions as well as tonal portions. These tonal portions, however, are not generated by the adaptive frame fill scheme, but these first spectral portions are output by the audio decoder or center decoder directly.
The adaptive spectral box selection scheme can operate with low granularity. In this implementation, a source region is subdivided into generally overlapping source regions and the target region or
IMPI
MEXICAN INSTITUTE VSjSít. * FROM THE PROniDAI? CV.JTSj. /
INDUSTRIAL reconstruction bands are given by frequency target regions that do not overlap. Then, the similarities between each source region and each destination region are determined on the encoder side and the best match pair of a source region and the destination region is identified by the match information and, on the decoder, the source region identified in the adaptation information is used for the generation of raw spectral data for the reconstruction band.
In order to obtain higher granularity, each region of origin is allowed to change in order to obtain a certain delay when the similarities are maximum. This delay can be as fine as a frequency range and allows even better adaptation between a source region and the destination region.
Also, in addition to only identifying a best match pair, this correlation delay can also be transmitted within the match information and, additionally, a sign can even be transmitted. When the signal is determined to be negative on the encoder side, then a corresponding sign tag is also transmitted within the adaptation information and, on the decoder side, the region spectral values of the source region are multiplied by -Γ or, in a complex representation, they are rotated 180 degrees.
A further implementation of the present invention applies a box bleaching operation. Spectrum bleaching removes raw spectral envelope information and emphasizes fine spectral structure
INSTITUTO MEXICANO OT LA PROPIEDAD
MDUSTRIAL
<img file="MX340575B_D0025.tif" />
which is of primary interest for assessing box similarity. Therefore, a frequency box on the one hand and / or the source signal on the other hand is bleached before calculating a cross-correlation measure. When only the frame is bleached using a predefined procedure, a bleach tag is transmitted indicating to the decoder that the same predefined bleaching process will be applied to the frequency frame within IGF.
As for the box selection, it is preferred to use the correlation delay to spectrally shift the regenerated spectrum by an integer number of transformation intervals. Depending on the underlying transformation, the spectral shift may require addition corrections. In the case of odd delays, the box is further modulated through multiplication by an alternative time sequence of -1/1 to compensate for the inverse frequency representation of any other band within the
Transformed Discrete Modified Cosine (MDCT). Also, the sign of the correlation result is applied when the frequency box is generated.
Also, it is preferred to use trim and frame stabilization to ensure that failures created by rapid change of source regions to the same rebuild region or destination region are avoided. For this, a similarity analysis is carried out between the different identified origin regions and when an origin box is similar to other origin boxes with a similarity above a threshold, then this origin box can be removed from the set of Potential source boxes, as it is highly correlated with other source boxes. Also, as a type of
IMPI
INSTITUTO MEXICANO OE LA PROPERTY INDUSTRIAL stabilization of selection of boxes, it is preferred to keep the order of the previous box if none of the boxes of origin in the current box correlates (better than a certain threshold) with the target boxes in the current box.
The audio encoding system efficiently encodes arbitrary audio signals in a wide range of bit rates. Therefore, for high bit rates, the system of the invention converges with transparency, for low bit rates perceptual discomfort is minimized. Therefore, the main bit rate available is used to wavecode only the most perceptually relevant signal structure in the encoder, and the resulting spectral spaces are filled in the decoder with the content of the signal that roughly approximates the original spectrum. A very limited bit budget is consumed to control the so-called Intelligent Space Fill (IGF) based on parameters by dedicated side information transmitted from the encoder to the decoder.
Hereinafter, preferred embodiments of the present invention will be described with reference to the accompanying drawings, in which:
Fig. 1a illustrates an apparatus for encoding an audio signal.
Fig. 1b illustrates a decoder for decoding an encoded audio signal that matches the encoder of Fig. 1 a.
Fig. 2a illustrates a preferred implementation of the decoder.
Fig. 2b illustrates a preferred implementation of the encoder.
<img file="MX340575B_D0026.tif" />
<img file="MX340575B_D0027.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL MONEDAD
Fig. 3a illustrates a schematic representation of a spectrum generated by the spectral domain decoder of Fig. 1b.
Fig. 3b illustrates a table indicating the relationship between scale adjustment factors for scale factor bands and powers for reconstruction bands and noise fill information for a noise fill band.
Fig. 4a illustrates the functionality of the spectral domain encoder to apply the selection of spectral portions to the first and second set of spectral portions.
Fig. 4b illustrates an implementation of the functionality of Fig. 4a.
Fig. 5a illustrates a functionality of a Transform encoder
Modified Discrete Cosine (MDCT).
Fig. 5b illustrates a decoder functionality with an MDCT technology.
Fig. 5c illustrates an implementation of the frequency regenerator.
Fig. 6a illustrates an audio encoder with temporal noise modeling / temporal box modeling functionality.
Fig. 6b illustrates a decoder with temporal noise modeling / temporal frame modeling technology.
Fig. 6c illustrates additional temporal noise modeling / temporal box modeling functionality with a different order of the spectral prediction filter and the spectral modeler.
<img file="MX340575B_D0028.tif" />
<img file="MX340575B_D0029.tif" />
Fig. 7a illustrates an implementation of Time Box Modeling (TTS) functionality.
Fig. 7b illustrates an implementation of the decoder that is adapted to the implementation of the encoder of Fig. 7a.
Fig. 7c illustrates a spectrogram of an original signal and an extended signal without Temporal Box Modeling (TTS).
Fig. 7d illustrates a frequency representation illustrating the correspondence between the intelligent gap fill frequencies and the box or time modeling powers .
Fig. 7e illustrates a spectrogram of an original signal and an extended signal without Temporal Box Modeling (TTS).
Fig. 8a illustrates a two-channel decoder with frequency feedback.
Fig. 8b illustrates a table illustrating different combinations of the 15 representations and the origin / destination ranges.
Fig. 8c illustrates a flowchart illustrating the functionality of the two-channel decoder with frequency feedback of Fig. 8a.
Fig. 8d illustrates a more detailed implementation of the decoder of Fig. 8a.
Fig. 8e illustrates an implementation of an encoder for two-channel processing that will be decoded by the decoder of Fig. 8a.
MEXICAN INSTITUTE j0
DF. THE PROPERTY
INDUSTRIAL
Fig. 9a illustrates a decoder with frequency regeneration technology using power values for the regeneration frequency range.
Fig. 9b illustrates a more detailed implementation of the 5-frequency regenerator of Fig. 9a.
Fig. 9c illustrates a schematic illustrating the functionality of Fig. 9b.
Fig. 9d illustrates a further implementation of the decoder of the
Fig. 9a.
Fig. 10a illustrates a block diagram of an encoder that matches the decoder of Fig. 9a.
Fig. 10b illustrates a block diagram to illustrate additional functionality of the parameter calculator of Fig. 10a.
Fig. 10c illustrates a block diagram to illustrate additional functionality of the parameter calculator of Fig. 10a.
Fig. 10d illustrates a block diagram illustrating additional functionality of the parameter calculator of Fig. 10a.
Fig. 11a illustrates an additional decoder that has a source range specific identification for a spectral frame filling operation in the decoder.
Fig. 11b illustrates the additional functionality of the frequency regenerator of Fig. 11a.
Fig. 11c illustrates an encoder used to cooperate with the decoder of Fig. 11a.
<img file="MX340575B_D0030.tif" />
MEXICAN INSTITUTE OF INBUSTRIAL PROPERTY
Fig. 11d illustrates a block diagram of an implementation of the parameter calculator of Fig. 11c.
Fig. 12a and 12b illustrate frequency schemes to illustrate a source range and a target range.
Fig. 12c illustrates an example mapping scheme of two signals.
Fig. 13a illustrates a prior art encoder with bandwidth extension.
Fig. 13b illustrates a prior art decoder with bandwidth extension.
Fig. 1a illustrates an apparatus for encoding an audio signal 99. The audio signal 99 is input to a time spectrum converter 100 to convert an audio signal having a sampling rate into a spectral representation 101 emitted by the time spectrum converter. Spectrum 101 is fed into a spectral analyzer 102 to analyze spectral representation 101. Spectral analyzer 101 is configured to determine a first set of first spectral portions 103 to be encoded with a first spectral resolution and a different second set of second spectral portions 105 to be encoded with a second spectral resolution. The second spectral resolution is smaller than the first spectral resolution. The second set of second spectral portions 105 is input to a parameter calculator or parametric encoder 104 to calculate the spectral envelope information having the second spectral resolution. Also, an audio encoder is provided
<img file="MX340575B_D0031.tif" />
Spectral domain 106 to generate a first coded representation 107 of the first set of first spectral portions having the first spectral resolution. Furthermore, the parameter calculator / parametric encoder 104 is configured to generate a second encoded representation 109 of the second set of second spectral portions. The first encoded representation 107 and the second encoded representation 109 are input to a bit stream multiplexer or bit stream former 108 and block 108 finally outputs the encoded audio signal for transmission or storage in a storage device.
Generally, a first spectral portion such as 306 in Fig. 3a will be surrounded by two second spectral portions such as 307a, 307b. This does not apply in Advanced High Efficiency Audlo Coding (HE AAC), where the frequency range of the central encoder is band-limited.
Fig. 1b illustrates a decoder that matches the encoder of Fig. 1a. The first encoded representation 107 is input to a spectral domain audio decoder 112 to generate a first decoded representation of a first set of first spectral portions, where the decoded representation has a first spectral resolution. Furthermore, the second encoded representation 109 is input to a parametric decoder 114 to generate a second decoded representation of a second set of second spectral portions having a second spectral resolution that is lower than the first spectral resolution.
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
The decoder further comprises a frequency regenerator 116 for regenerating a reconstructed second spectral portion having the first spectral resolution using a first spectral portion. The frequency regenerator 116 carries out a box filling operation, that is, it uses a box or a portion of the first set of first spectral portions and copies this first set of first spectral portions into the reconstruction range or reconstruction band that the second spectral portion has and generally performs spectral envelope modeling or other operation indicated by the second decoded representation emitted by the parametric decoder 114, that is, using the information on the second set of second spectral portions. The first decoded set of first spectral portions and the second reconstructed set of spectral portions indicated at the output of the frequency regenerator 116 on line 117 is inserted into a spectrum-time converter 118 configured to convert the first decoded representation and the second spectral portion reconstructed into a time representation 119, where the time representation has a certain high sampling rate.
Fig. 2b illustrates an implementation of the encoder of Fig. 1a. An audio input signal 99 is input to an analysis filter bank 220 corresponding to the time spectrum converter 100 of Fig. 1a. Next, the temporal noise modeling operation is performed in the temporal noise modeling block (TNS) 222. Therefore, the input to the spectral analyzer 102 in Fig. 1a corresponding to a masking
<img file="MX340575B_D0032.tif" />
<img file="MX340575B_D0033.tif" />
From the industrial tonal PROmi> A!> Block 226 of Fig. 2b they can be full spectral values, when the temporal noise modeling / temporal box modeling operation is not applied or they can be spectral residual values, when the Temporary Noise Modeling (TNS) operation as illustrated in Fig. 2b, block 222. For two-channel signals or multichannel signals, a joint channel 228 encoding may further be carried out, whereby the spectral domain encoder 106 of Fig. 1a may comprise the joint channel encoding block 228. Also, provides an entropy encoder 232 for performing lossless data compression which is also a portion of the spectral domain encoder 106 of Fig. 1a.
Spectral analyzer / tonal masking 226 separates the output of the temporal noise modeling block (TNS) 222 in the center band and the tonal components corresponding to the first set of first spectral portions 103 and the residual components corresponding to the second set of second spectral portions 105 of Fig. 1a. Block 224 denoted as the Intelligent Space Fill Parameter Extraction (IGF) encoding corresponds to the parametric encoder 104 of FIG. 1a and the bitstream multiplexer 230 corresponds to the bitstream multiplexer 108 of FIG. 1st.
Preferably, the analysis filter bank 222 is implemented as an MDCT (modified discrete cosine transform filter bank) and the MDCT is used to transform signal 99 into a frequency domain.
IMPI
INSTITUTO MiXICA no DE LA EROME DAD INDUSTRIAL
<img file="MX340575B_D0034.tif" />
temporal with the modified discrete cosine transform acting as the frequency analysis tool.
Spectral analyzer 226 preferably applies a hue masking. This tonality masking estimation step is used to separate the tonal components from the noise-like components in the signal. This allows the central encoder 228 to encode all the tonal components with a psychoacoustic module. The tonality masking estimation stage can be implemented in many different ways and is preferably implemented similarly in functionality to the sinusoidal track estimation stage used in noise and sinusoidal modeling for speech / audio coding [8 , 9] or an audio encoder based on the HILN model described in [10]. An implementation is preferably used that is easy to implement without the need to maintain birth-death paths, but any other hue or noise detector can also be used.
The Intelligent Space Fill Module (IGF) calculates the similarity that exists between a source region and a destination region. The destination region will be represented by the spectrum of the origin region. The similarity between the source and destination regions is measured using a cross-correlation approach. The destination region is divided into nTar frequency boxes that do not overlap. For each frame in the destination region, nSrc creates a frame of origin from a fixed start frequency. These
IMPIAS
MEXICAN INSTITUTE
OF INDUSTRIAL PROPERTY
U £ -; _. ·. -Λ origin boxes overlap by a factor between 0 and 1, where 0 means
0% overlap and 1 means 100% overlap. Each of these source boxes is mapped to the target box at various delays to find the source box that best matches the target box. The best adaptation box number is stored in tfZe / Vum [idx_tar], the delay in which it best correlates with the target is stored in xcorr_laglidx_tar] [idx_src] and the correlation sign is stored in xcorr_s¿5n [tdx_tar ] [tdx_rrc]. In case the correlation is very negative, the origin box must be multiplied by -1 before the box filling process in the decoder. The Intelligent Gap Fill Module (IGF) also takes care not to overwrite tonal components in the spectrum as the tonal components are preserved using hue masking. A band power parameter is used to store the power of the target region to allow accurate spectrum reconstruction.
This method has certain advantages compared to classical SBR [1] where the harmonic grid of a multi-tone signal is preserved by the central encoder while only the spaces between the sinusoids are filled with the best modeled adaptive noise of the region of origin. Another advantage of this system compared to ASR (Precise Spectral Replacement) [2-4] is the absence of a signal synthesis step that creates significant portions of the signal in the decoder. Instead, this task is assumed by the central encoder, which allows the conservation of the components
IM. pi
MEXICAN INSTITUTE OF THE ΓϋΟΠϊΠΑΌ
INDUSTRIAL major spectrum. Another advantage of the proposed system is the continuous scalability offered by the features. Only the use of tileNum [idx_tar] and xcorrjag = O, for each frame is called raw granularity adaptation and can be used for low bit rates while the use of the xcorrjag variable for each frame allows better adaptation of the target spectra and of origin.
In addition, a frame choice stabilization technique is proposed that eliminates frequency domain faults such as trill and musical noise.
For stereo channel pairs, additional set stereo processing is applied. This is necessary since, for a given target range, the signal can be a highly correlated, panned sound source. In case the chosen origin regions for this particular region are not well correlated, even though the powers are adapted to the target regions, the spatial image may suffer due to the uncorrelated origin regions. The encoder analyzes each power band in the target region, usually by performing a cross-correlation of the spectral values, and if a certain threshold is exceeded, it establishes a joint label for this power band. In the decoder, the power bands of the left and right channels are treated individually if this joint stereo label is not set. In case the stereo set tag is set, both powers and interconnect are done in the stereo set domain. Information
IMPI
MEXICAN INSTITUTE £> S LA PMOWE> AD INDUSTRIAL
<img file="MX340575B_D0035.tif" />
Stereo ensemble for Intelligent Gap Fill Regions (IGF) is reported similarly to Stereo Joint information for central encoding, which includes a label indicating in the case of the prediction whether the direction of the prediction is down to residual mix or vice versa.
Powers can be calculated from the powers transmitted in the L / R domain.
mtdNrg [k] = leftNrg [k] + rightNrg [t¿l;
sldeNrg [k] = leftNrg [k] - rightNrglk];
where k is the frequency index in the domain of the transform.
Another solution is to calculate and transmit the powers directly into the set stereo domain for the bands where the set stereo is active, so there is no need for additional power transformation on the decoder side.
Source boxes are always created according to the Matrix
Middle / Side:
midTile [k] = 0.5 · (leftTile [k] + righ (Tile [k]) sideTile [k \ = 0.5 · (leftTile [k] - rightTile [k])
Power adjustment:
midTile [k] = médTíte [fc] * mtdNrg [k] ·, SídeTííeífel = s £ deF0e [fe] * sideNrg [k];
Joint stereo transformation -> LR:
If no additional prediction parameter is coded:
IMPI
INSTITUTO MRXICANO DF LA rROPltDAP INDUSTRIAL leftTile [k] = midTile [k] + sideTiltfk] rightTile [k] = midTile [k] - sideTile [k]
If an additional prediction parameter is coded and if the indicated direction is middle to side:
sideTile [k] = sideTile [k] - predictionCoeff midTile [k] leftTile [k] = midTile [k] + sideTiie [k] rightTile [k] = midTile [k \ - sideTile \ k \
If the direction indicated is from the side to the middle:
2θ midTile \ [k] = midTile [k \ - predictionCoeff sideTile [k \ leftTile [k] = midTilei [k] - sideTile [k \ rightTile [k] = midTile \ [k] + sideTile [k]
This processing ensures that from the boxes used to regenerate highly correlated destination regions and panned destination regions, the resulting left and right channels continue to represent a correlated and panned sound source even if the source regions are uncorrelated, preserving the stereo image for those regions.
In other words, stereo joint tags are transmitted in the bitstream indicating whether L / R or M / S will be used as an example for general stereo joint encoding. In the decoder, first, the center signal is decoded as indicated by the stereo joint labels for the center bands. Second, the core signal is stored in both the L / R and M / S representations. For Smart Space Fill (IGF) box fill, the source box rendering is selected to
<img file="MX340575B_D0036.tif" />
<img file="MX340575B_D0037.tif" />
<sup>32</sup> IMPI
INSTITUT MEJeCAN <?
OF THE PROPERTY
INDUSTRIAL adjust the rendering of destination boxes as indicated by the joint stereo information for the IGF bands.
Temporal noise modeling (TNS) is a standard technique and is part of Advanced Audio Coding (AAC) [11-13]. Temporal noise modeling (TNS) can be considered as an extension of the basic scheme of a perceptual encoder, by inserting an optional processing step between the filter bank and the quantization stage. The main task of the Temporal Noise Modeling Module (TNS) is to hide the quantization noise produced in the temporal masking region of transition signals and thus produce a more efficient coding scheme. First, temporal noise modeling (TNS) calculates a set of prediction coefficients using direct prediction in the transform domain, for example, the Modified Discrete Cosine Transform (MDCT). These coefficients are then used to flatten the time envelope of the signal. As quantization affects the temporal noise modeling (TNS) filtered spectrum, quantization noise is also temporarily flat. By applying reverse filtering of temporal noise modeling (TNS) on the decoder side, the quantization noise is modeled according to the temporal envelope of the TNS filter, and therefore the quantization noise is masked by the transient.
Space Smart Fill (IGF) is based on a representation of MDCT. Preferably, for efficient coding, long blocks of approximately 20 ms blocks should be used. If the signal inside said
IMPI
MEXICAN INSTITUTE OF THE INDUSTRIAL RRORISPAD
<img file="MX340575B_D0038.tif" />
long block contains transients, in the smart space filling (IGF) spectral bands audible pre- and post-echoes occur due to box filling. Fig. 7c shows a typical pre-echo effect before transient onset due to Intelligent Space Fill (IGF). On the left side the spectrogram of the original signal is shown and on the right side the spectrogram of the extended bandwidth signal without temporal noise modeling (TNS) filtering is shown.
This pre-echo effect is reduced by using TNS in the context of Intelligent Space Fill (IGF). In this instance, the TNS is used as a time frame modeling tool (TTS) since spectral regeneration at the decoder is performed on the residual TNS signal. The required TTS prediction coefficients are calculated and applied using the full spectrum on the encoder side as usual. The start and end frequencies of Temporal Noise Modeling (TNS) / Temporal Box Modeling (TTS) are not affected by the Intelligent Fill Space (IGF) Start frequency ^<sub>eFsrarr</sub>of the IGF tool. Compared to the prior art time noise modeling (TNS), the end frequency of time frame modeling (TTS) increases to the end frequency of the Intelligent Space Fill Tool (IGF), which is greater than /}<sub>CFlterc</sub>. On the decoder side, the TNS / TTS coefficients are applied over the entire spectrum again, ie the central spectrum plus the regenerated spectrum plus the tonal components of the tonality map (see Fig. 7e). The application of
IMPI Mexican Institute Dt THE PROPERTY
INDUSTRIAL
<img file="MX340575B_D0039.tif" />
Temporal Box Modeling (TTS) is required to form the temporal envelope of the regenerated spectrum to adapt to the original signal envelope again. Therefore, the illustrated pre-echoes are reduced. Additionally, it still models the quantization noise in the signal below<sup>5</sup> fiGFstart as usual in temporal noise modeling (TNS).
In prior art decoders, the spectral interconnection in an audio signal alters the spectral correlation at the interconnection limits, and thus affects the time envelope of the audio signal by introducing scatter. Therefore, another advantage of applying the Intelligent Space Fill (IGF) box fill on the residual signal is that, after the modeling filter is applied, the box boundaries are perfectly correlated, resulting in a more faithful temporal reproduction of the signal.
In an encoder of the invention, the spectrum that has been subjected to TNS / TTS filtering, hue masking processing, and intelligent space fill parameter (IGF) estimation lacks any signal above the frequency of IGF onset except tonal components. This spread spectrum is now encoded by the central encoder using the principles of arithmetic coding and predictive coding. These encoded components together with the signaling bits form the audio bitstream.
Fig. 2a illustrates the corresponding implementation of the decoder. The bitstream in Fig. 2a corresponding to the encoded audio signal is
IMPIAS
MIXICANO INSTITUT
3S PROPERTY V ^ wTSfC / ^ industrial ^ ¿τβτ ^ ε ^ introduces in the demultiplexer / decoder that it would be connected, with respect to Fig. 1b, to blocks 112 and 114. The bitstream demultiplexer separates the signal from input audio in the first coded representation 107 of Fig.
1b and the second coded representation 109 of Fig. 1b. The first encoded representation having the first set of first spectral portions is inserted into the channel joint decoding block 204 corresponding to the spectral domain decoder 112 of Fig. 1b. The second coded representation is inserted into the parametric decoder 114 which is not shown in Fig. 2a and then inserted into the filler block
Space Intelligent (IGF) 202 corresponding to the frequency regenerator 116 of Fig. 1b. The first set of first spectral portions necessary for frequency regeneration is input to IGF block 202 via line 203. Also, after joint decoding of channels 204, specific core decoding is applied in the masking block. tonal
206 such that the output of the tonal masking 206 corresponds to the output of the spectral domain decoder 112. Next, the combiner 208 performs a merge, ie a frame construction where the output of the combiner 208 now has the spectrum of full range, but still in the filtered temporal noise modeling (TNS) / temporal box modeling (TTS) domain. Subsequently, in block 210, a TNS / TTS Inverse operation is carried out using TNS / TTS Filter Information provided through line 109, i.e. the TTS Side Information is preferably Included in the first encoded representation generated by the
<img file="MX340575B_D0040.tif" />
spectral domain encoder 106 which may be, for example, a direct forward audio encoding (AAC) central encoder or unified voice and audio encoding (USAC), or may also be included in the second encoded representation. At the output of block 210 a full spectrum is provided up to the maximum frequency which is the full range frequency defined by the sampling rate of the original input signal. A spectrum / time conversion is then carried out in the synthesis filter bank 212 to finally obtain the audio output signal.
Fig. 3a illustrates a schematic representation of the spectrum. The spectrum is subdivided into SCB scale factor bands where there are seven scale factor bands SCB1 to SCB7 in the illustrated example of Fig. 3a. Scale factor bands can be advanced audio coding scale factor (AAC) bands that are defined in the AAC standard and have increasing bandwidth up to higher frequencies as illustrated in Fig.
3a schematically. It is preferred to carry out Intelligent Space Fill (IGF) not from the beginning of the spectrum, ie at low frequencies, but to start IGF operation at an IGF start frequency illustrated at 309. Therefore, the band Center frequency ranges from the lowest frequency to the IGF start frequency. Above the IGF start frequency, spectrum analysis is applied to separate the high-resolution spectral components 304, 305, 306, 307 (the first set of first spectral portions) from low-resolution components represented by the second set of second spectral portions. Fig. 3a illustrates a spectrum that
A VAT T 1 f
INSTITUTE ΜΜΚ7ΛΗΟ Ϊ t »f LA FÍ OHEB-AD \ rwewrwAt introduces, by way of example, the spectral domain encoder 106 or the
<img file="MX340575B_D0041.tif" />
channel set encoder 228, i.e. the core encoder operates over the entire range but encodes a significant number of zero spectral values, i.e. these zero spectral values are quantized to zero or set to zero before or after quantization. quantification. Regardless, the core encoder operates over the full range, i.e. as if the spectrum were as illustrated, i.e. the core decoder does not necessarily have to be aware of any intelligent space padding from the second set of second spectral portions with lower spectral resolution.
Preferably, the high resolution is defined by a line coding of spectral lines, such as lines of the modified direct cosine transform (MDCT), while the second resolution or low resolution is defined, for example, by calculating only a single spectral value per scale factor band, where a scale factor band spans multiple frequency lines. Therefore, the second low resolution, with respect to its spectral resolution, much lower than the first or high resolution defined by line coding is generally applied by the central encoder such as an advanced advanced encoding (AAC) central encoder. or unified voice and audio coding (USAC).
As for the scale adjustment factor or power calculation, the situation is illustrated in Fig. 3b. Due to the fact that the encoder is a central encoder and due to the fact that there may, but not necessarily, be components of the first set of spectral portions in each band, the
IMPI
MEXICAN INSTITUTE OE THE PROPERTY
INDUSTRIAL
<img file="MX340575B_D0042.tif" />
Center Encoder calculates a scale adjustment factor for each band not only in the center range below the Intelligent Space Fill (IGF) start frequency 309, but also above the IGF start frequency up to the maximum frequency <sup>is</sup> less than or equal to half the sampling frequency, i.e. f<sub>s</sub>/<sub>2</sub>. Therefore, the coded tonal portions 302, 304, 305, 306, 307 of Fig. 3a and, in this embodiment together with the scaling factors SCB1 to SCB7 correspond to the high resolution spectral data. Low resolution spectral data are calculated from the IGF start frequency and correspond to the values of
Ei, E power information<sub>2</sub>, E<sub>3</sub>, E<sub>4</sub>, which are transmitted together with the scaling factors SF4 to SF7.
In particular, when the central encoder is in a low bit rate condition, the additional noise fill operation can also be applied in the central band, that is, a frequency lower than the start frequency of the Intelligent Space Fill (IGF). ), that is, in the scale factor bands SCB1 to SCB3. In the noise fill, there are several adjacent spectral lines that have been zero rated. On the decoder side, these zero-qualified spectral values are re-synthesized and the re-synthesized spectral values are adjusted in magnitude using a noise fill power such as NF<sub>2</sub> Illustrated at 308 in Fig. 3b. The noise fill power, which can be given in absolute or relative terms particularly with respect to the scaling factor as in unified speech encoding
IMPI
INSTITUTO MDCICANO OS LA PROREDAD industrial
<img file="MX340575B_D0043.tif" />
and audio (USAC) corresponds to the power of the set of spectral values quantized to zero. These noise fill spectral lines can also be considered a third set of third spectral portions that are regenerated by simple noise fill synthesis without any intelligent space filling (IGF) operation based on frequency regeneration using frequency frames. of other frequencies for the reconstruction of frequency boxes using spectral values of an origin range and the power information Ei, E<sub>2</sub>, E3, E<sub>4</sub>.
Preferably, the bands for which the power information is calculated coincide with the scale factor bands. In other embodiments, a grouping of Power Information values is applied such that, for example, for scale factor bands 4 and 5, only a single power information value is transmitted, but even in this embodiment , the limits of the clustered reconstruction bands coincide with the limits of the scale factor bands. If different band separations are applied, then new calculations or timing calculations can be applied, and this may make sense depending on the particular application.
Preferably, the spectral domain encoder 106 of Fig. 1a is a psychoacoustically activated encoder as illustrated in Fig. 4a. Generally, as illustrated for example in the MPEG2 / 4 AAC or MPEG1 / 2, Layer 3 standard, the audio signal to encode after being transformed into the spectral range (401 in Fig. 4a) is sent to a
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY scale factor calculator 400. The scale adjustment factor calculator is controlled by a psychoacoustic model that additionally receives the audio signal to quantify or receive, as in the MPEG1 / 2 Layer 3 or MPEG standard. AAC, a complex spectral representation of the audio signal. The psychoacoustic model calculates, for each band of scale factor, a scale factor representing the psychoacoustic threshold. Additionally, the scaling factors are then adjusted, by the well-known cooperation of the internal and external iteration loops or by any other suitable encoding method so that certain bit rate conditions are met.
Next, the spectral values to quantize on the one hand, and the scaling factors calculated on the other hand, are input to a quantizer processor 404. In the simple audio encoding operation, the spectral values to quantize are weighted by the factors. Scale adjustment and weighted spectral values are then entered into a fixed quantizer that generally has compression functionality for higher amplitude ranges. So, at the output of the quantizer processor there are quantization indices that are then sent to an entropy encoder that generally has specific and highly efficient encoding for a set of zero quantization indices for adjacent frequency values or, as it is also called in the technique, a run of zero values.
In the audio encoder of Fig. 1a, however, the quantizer processor generally receives information about the second spectral portions of the spectral analyzer. Therefore, the quantizer processor 404
<img file="MX340575B_D0044.tif" />
IMPI
NSTmrro Mexican OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0045.tif" />
ensures that, at the output of quantizer processor 404, the second spectral portions identified by spectral analyzer 102 are zero or have a representation recognized by an encoder or decoder as a zero representation that can be encoded very efficiently, specifically when there are zero value runs in the spectrum.
Fig. 4b illustrates an implementation of the quantizer processor. The spectral values of the modified discrete cosine transform (MDCT) can be entered into a zero-set block 410. Subsequently, the second spectral portions are already set to zero before weighting by the scale adjustment factors in the block 412. In a further implementation, block 410 is not provided, but the zero-set cooperation is carried out in block 418 after weighting block 412. Even in another implementation, the zero-set operation can also be carried out in a zero-set block 422 subsequent to quantization at quantizer block 420. In this implementation, blocks 410 and 418 would not be present. In general at least one of blocks 410, 418, 422 is provided depending on the specific implementation.
Then, at the output of block 422, a quantized spectrum corresponding to what is illustrated in Fig. 3a is obtained. This quantized spectrum is then input to an entropy encoder such as 232 in Fig. 2b which may be a Huffman encoder or an arithmetic encoder as defined, for example, in the Unified Voice and Audio Coding (USAC) standard.
<img file="MX340575B_D0046.tif" />
INDUSTRIAL
The zero-set blocks 410, 418, 422, which are provided alternately with each other or in parallel, are controlled by the spectral analyzer 424. The spectral analyzer preferably comprises any implementation of a well-known hue detector or comprises any different type of operational detector. to separate a spectrum into components to encode with a high resolution and components to encode with a low resolution. Other of these algorithms implemented in the spectral analyzer can be a voice activity detector, a noise detector, a voice detector or any other detector that determines, based on the spectral information or associated metadata, the resolution requirements for different spectral portions.
Fig. 5a illustrates a preferred implementation of the time spectrum converter 100 of Fig. 1a as implemented, for example, in advanced audio encoding (AAC) or unified speech and audio encoding (USAC). Time spectrum converter 100 comprises a window splitter 502 controlled by a transient detector 504. When transient detector 504 detects a transient, it then signals an exchange from long windows to short windows to the window splitter. The window splitter 502 then computes for the overlapping blocks, window-divided boxes, where each window-divided box typically has two N values, such as the 2048 values. A transformation is then carried out within a block transformer 506, and generally this block transformer further provides a deletion so that it performs a
<img file="MX340575B_D0047.tif" />
INDUSTRIAL elimination / transform combined to obtain a spectral table with N values such as ios spectral values of the modified discrete cosine transform (MDCT). Therefore, for a long window operation, the box at the input of block 506 comprises two N values, such as
2048 values and a spectral box then has 1024 values. However, an exchange is then performed on the short blocks, i.e. when eight short blocks are performed where each short block has 1/8 window-divided time domain values compared to a long window and each spectral block has 1/8 spectral values compared to a long block. Therefore, when this elimination is combined with a 50% window splitter overlap operation, the spectrum is a critically sampled version of the 99 time domain audio signal.
Reference is then made to Fig. 5b which illustrates a specific implementation of the frequency regenerator 116 and the spectrum time converter 118 of Fig. 1b, or the combined operation of blocks 208, 212 of Fig. 2a . In Fig. 5b a specific reconstruction band is illustrated such as the scale factor band of Fig. 3a. The first spectral portion in this reconstruction band, that is, the first spectral portion
306 of Fig. 3a is entered into the frame builder / regulator block 510.
Likewise, a second reconstructed spectral portion for scale factor band 6 is also input to frame builder / regulator 510. In addition, power information such as E<sub>3</sub> of Fig. 3b for a band of
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY scale factor 6 is also entered in block 510. The second reconstructed spectral portion in the reconstruction band has already been generated by the filling of frequency boxes using an origin range and the reconstruction band then corresponds to the target rank. In this instance, a power adjustment of the frame is carried out to finally obtain the complete reconstructed frame that has the N values, such as those obtained at the output of the combiner 208 in Fig. 2a. Then in block 512 an inverse block transform / interpolation is performed to obtain 248 time domain values for the 124 spectral values, for example at the input of block 512. Next, at block 514, a window split synthesis operation is performed which is again controlled by a long window / short window indication transmitted as side information in the encoded audio signal. Then, at block 516, an overlap / add operation is performed with a previous time frame.
Preferably, the modified discrete cosine transform (MDCT) applies a 50% overlap so that, for each new time frame of 2N values, the time domain N values are finally output. A 50% overlap is preferred due to the fact that providing critical sampling and continuous crossover from one frame to the next frame due to the overlap / add operation of block 516.
As illustrated at 301 in Fig. 3a, a noise fill operation can be additionally applied, not only below the intelligent space fill start frequency (IGF) but also above the frequency of
<img file="MX340575B_D0048.tif" />
IMPI
INSTITUT · MEXICAN
OF THE INDUSTRIAL ERORITY
<img file="MX340575B_D0049.tif" />
start of IGF as for the contemplated reconstruction band coinciding with the scale factor band 6 of Fig. 3a. Then the noise fill spectral values can also be entered into the frame builder / slider 510 and the adjustment of the noise fill spectral values can also be applied within this block or the noise fill spectral values can already be entered can be adjusted using noise fill power before being entered into the frame builder / slider
510.
Preferably, an IGF operation, that is, a frequency frame fill operation using spectral values from other portions, can be applied over the entire spectrum. Therefore, a spectral box fill operation can not only be applied in the high band above an intelligent gap fill start frequency (IGF) but can also be applied in the low band. In addition, noise padding without frequency frame padding can also be applied not only below the Smart Space Fill Start (IGF) frequency but also above the IGF start frequency. However, it has been found that the high quality and high efficiency of audio encoding can be obtained when the noise fill operation is limited to the frequency range below the IGF start frequency and when the fill operation Frame rate is limited to the frequency range above the IGF start frequency, as illustrated in Fig. 3a.
<img file="MX340575B_D0050.tif" />
Preferably, the target boxes (TT) (which have frequencies greater than the IGF start frequency) are subject to the full rate encoder scale factor band limits. The origin boxes (ST), from which information is obtained, that is, for frequencies below the IGF start frequency are not subject to the limits of the scale factor band. The size of the source boxes (ST) must correspond to the size of the associated target box (TT). This is demonstrated using the following example. TT [0] is 10 MDCT Intervals in length. This corresponds exactly to the length of two subsequent SBCs (such as 4 + 6). Thus, all possible origin boxes (ST) that must be correlated with TT [0], also have a length of 10 intervals. A second target box TT [1] that is adjacent to TT [0] is 15 intervals 1 long (SCB is 7 + 8 long). So the origin box (ST) for the above has a length of 15 intervals instead of 10 intervals as for TT [0].
In case a target box (TT) cannot be found for a source box (ST) with the length of the target box (for example, when the length of the TT is greater than the available source range), then a correlation is not calculated and the source range is copied a number of times into this TT (copying is performed one after the other so that a frequency line for the lowest frequency of the second copy follows immediately - at frequency - the frequency line for the highest frequency of the first copy), until the destination box (TT) is completely filled.
<img file="MX340575B_D0051.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
Reference is made below to Fig. 5c which illustrates a further preferred embodiment of the frequency regenerator 116 of Fig. 1b or the Intelligent Space Fill Block (IGF) 202 of Fig. 2a. Block 522 is a frequency frame generator that not only receives an ID from the destination band but also receives an ID from the source band. By way of example, it has been determined on the encoder side that the scale factor band 3 in Fig. 3a is very suitable for the reconstruction of the scale factor band 7. Therefore, the ID of the source band would be 2 and the ID of the destination band would be 7. Based on this information, the box generator of Frequency 522 applies a harmonic frame fill or copy operation or any other frame fill operation to generate the second raw portion of the 523 spectral components. The second raw portion of the spectral components has a frequency resolution identical to the frequency resolution included in the first set of first spectral portions.
Then, the first spectral portion of the reconstruction strip such as 307 in Fig. 3a is input to a frame builder 524 and the second raw portion 523 is also input to the frame builder 524. The reconstructed frame is then adjusted by the 526 regulator using a gain factor for the reconstruction band calculated by the 528 gain factor calculator. It is important to note, however, that the first spectral portion in the box is not affected by regulator 526, but only the second raw portion for the reconstruction box is affected
IMPI
MBX1CANO INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0052.tif" />
by regulator 526. To this end, the gain factor calculator 528 analyzes the source band or the second raw portion 523 and further analyzes the first spectral portion in the reconstruction band to finally find the correct gain factor 527 of so that the power of the adjusted frame emitted by the 526 regulator has the power E<sub>4</sub> when a scale factor 7 band is contemplated.
In this context, it is very important to evaluate the precision of the high frequency reconstruction of the present invention compared to advanced high efficiency audio coding (HE-AAC). This is explained with respect to the scale factor band 7 in Fig. 3a. It is assumed that a prior art encoder illustrated in Fig. 13a would detect the spectral portion 307 to be encoded at high resolution as a missing harmonic. Then, the power of this spectral component would be transmitted along with spectral envelope information for the reconstruction band such as the scale factor 7 band to the decoder. The decoder would then recreate the missing harmonic. However, the spectral value, at which the missing harmonic 307 would be reconstructed by the prior art decoder of Fig. 13b would be in the middle of band 7 at a frequency indicated by the reconstruction frequency 390. Therefore, the present invention avoids a frequency error 391 that would be introduced by the prior art decoder of Fig. 13d.
In one implementation, the spectral parser is also implemented to calculate similarities between first and second spectral portions
IM F
INSTITUTO MRXICAi DF LA FROMEP spectral portions and to determine, on the basis of calculated * simlliTQ9es, for a second spectral portion in a reconstruction range a first spectral portion that matches the second spectral portion as much as possible. So, in this variable source range / target range implementation, the parametric encoder will also input adaptation information indicating an adaptation source range for each destination range in the second encoded representation. On the decoder side, this information could then be used by a frequency frame generator 522 in Fig. 5c illustrating a generation of a second raw portion 523 based on an ID of the source band and an ID of the destination band.
Also, as illustrated in Fig. 3a, the spectral analyzer is configured to analyze the spectral representation up to a maximum analysis frequency that is only a small amount below half the sampling frequency and is preferably at least a quarter of the sampling rate or generally higher.
According to the illustration, the encoder operates without sampling reduction and the decoder operates without operating without upsampling. In other words, the spectral domain audio encoder is configured to generate a spectral representation that has a Nyquist frequency defined by the sampling rate of the originally input audio signal.
Also, as illustrated in Fig. 3a, the spectral analyzer is configured to analyze the spectral representation that starts with a
IMPI
INSTITUTO MEXICANO OE LA PROPIEDAD INDUSTRIAL space filling frequency and ending with a maximum frequency represented by a maximum frequency included in the spectral representation, where a spectral portion that extends from a minimum frequency to the space filling start frequency belongs to the first set of spectral portions and where another spectral portion such as 304, 305, 306, 307 which has frequency values above the space fill frequency, is additionally included in the first set of first spectral portions.
As explained above, the spectral domain audio decoder 112 is configured such that a maximum frequency represented by a spectral value in the first decoded representation is equal to a maximum frequency included in the time representation having the sampling rate, where the spectral value for the maximum frequency in the first set of first spectral portions is zero or different from zero. Anyway, for this maximum frequency in the first set of spectral components there is a scale adjustment factor for the scale factor band, which is generated and transmitted regardless of whether all the spectral values in this scale factor band are set to zero or not, as described in the context of Figs. 3a and 3b.
Therefore, the invention is advantageous over other parametric techniques for increasing compression efficiency, for example, noise substitution and noise filler (these techniques are exclusively for efficient representation of local noise type signal content) , so the
<img file="MX340575B_D0053.tif" />
IMPI
MEXICAN INSTITUTE • E INDUSTRIAL TROPIÍDAD
<img file="MX340575B_D0054.tif" />
Invention allows accurate frequency reproduction of tonal components. To date, no state-of-the-art method addresses the efficient parametric representation of arbitrary signal content by filling in spectral spaces without the restriction of a priori fixed division in the low band (LF) and in the high band ( HF).
Embodiments of the system of the invention enhance current state of the art approaches and thus provide high compression efficiency, no or only minor perceptual discomfort, and full audio bandwidth, even at low rates. bit.
The system in general consists of:
Full-band center encoding Intelligent gap filling (frame fill or noise fill) Core-scattered tonal parts, selected by tonal masking Stereo pair coding for the full band, including frame fill
TNS in the spectral bleaching box in the smart gap fill (IGF) range
A first step towards a more efficient system is to eliminate the need to transform spectral data into a second domain of
ΙΜΡΪ
MEXICAN INSTITUTE D5 THE INDUSTRIAL PROPERTY transformed different from the domain of the central encoder. Since most audio codes such as Advanced Audio Coding (AAC) use Modified Discrete Cosine Transform (MDCT) as the basic transform, it is also useful to perform Bandwidth Extension (BWE). ) in the domain of the MDCT. A second requirement for the BWE system would be the need to preserve the tonal grid by which even high frequency (HF) tonal components are preserved and therefore the quality of the encoded audio is superior to existing systems. To take into account both requirements mentioned above for a bandwidth extension scheme (BWE), a new system called Intelligent Fill of Spaces (IGF) is proposed. Fig. 2b shows the block diagram of the proposed system on the encoder side and Fig. 2a shows the system on the decoder side.
Fig. 6a illustrates an apparatus for decoding an encoded audio signal in another implementation of the present invention. The decoding apparatus comprises a spectral domain audio decoder 602 to generate a first decoded representation of a first set of spectral portions and as the frequency regenerator 604 connected downstream of the spectral domain audio decoder 602 to generate a second spectral portion reconstructed using a first spectral portion of the first set of first spectral portions. As illustrated in 603, the spectral values in the first spectral portion and in the second spectral portion are residual spectral prediction values. In order to transform these values
<img file="MX340575B_D0055.tif" />
IMPI
ΤΟυΤΟ MSXICAN DR LA tNDUSTtIAL MORTALITY
<img file="MX340575B_D0056.tif" />
Spectral prediction residuals in a full spectral representation is provided by a 606 spectral prediction filter. This inverse prediction filter is configured to perform an inverse prediction on frequency using the spectral residuals for the first set of the first frequency and the second reconstructed spectral portions. The spectral inverse prediction filter 606 is configured by the filter information included in the encoded audio signal. Fig. 6b illustrates a more detailed implementation of the embodiment of Fig. 6a. The spectral prediction residuals 603 are entered into a frequency box generator 612 that generates raw spectral values for a reconstruction band or for a certain second frequency slice and these raw data which now have the same resolution as the first High resolution spectral representation are entered into the 614 Spectral Modeler. The spectral modeler now models the spectrum using envelope information transmitted in the bitstream, and the spectrally modeled data is then applied to the spectral prediction filter 616, finally generating a complete spectral value table using the filter information 607 transmitted from the encoder. to the decoder through the bit stream.
In Fig. 6b, it is assumed that, on the encoder side, the calculation of the filter information transmitted through the bit stream and used through line 607 is performed subsequent to the calculation of the filter information. the envelope. Therefore, in other words, an encoder that matches the
IMPI
MEXICAN INSTITUTE Of. INDUSTRIAL PROPERTY
<img file="MX340575B_D0057.tif" />
decoder of Fig. 6b would first compute the spectral residuals and then compute the envelope information with the spectral residuals as illustrated, for example, in Fig. 7a. However, the other implementation is also useful for certain implementations, when the envelope information is computed before performing Time Noise Modeling (TNS) or Time Box Modeling (TTS) filtering on the encoder side. . The spectral prediction filter 622 is then applied before spectral modeling is performed at block 624. Therefore, in other words, the (full) spectral values are generated before the spectral modeling operation 624 is applied.
Preferably a TNS filter or a complex value TTS filter is calculated. This is illustrated in Fig. 7a. The original audio signal is input into a complex modified discrete cosine transform (MDCT) block 702. The computation of the TTS filter and TTS filtering are then performed in the complex domain. Next, in block 706 the lateral information of the intelligent space filling (IGF) is calculated and any other operation such as spectral analysis for encoding, etc. is also calculated. Subsequently, the first set of first spectral portion generated by block 706 is encoded with a psychoacoustic model activated encoder illustrated in 708 to obtain the first set of first spectral portions indicated at X (k) in Fig. 7a and all of this data is sent to the bitstream multiplexer 710.
On the decoder side, the encoded data is input to a demultiplexer 720 to separate the IGF side information on the one hand, the
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0058.tif" />
TTS lateral information on the other hand and the coded representation of the first set of first spectral portions.
Then block 724 is used to calculate a complex spectrum from one or more actual value spectra. Next, both actual value spectra and complex spectra are input to block 726 to generate reconstructed frequency values in the second set of second spectral portions for a reconstruction band. Then, in the full-obtained full band frame and frame fill the time frame modeling reverse operation (TTS) 728 is performed and, on the decoder side, a complex MDCT final reverse operation is performed on the block 730. Therefore, the use of complex TNS filter information enables it to be generated automatically, when applied not only within the center band or within the frame bands separately but also over the center / box limits or over the limits frame / frame, a frame boundary processing that finally re-introduces a spectral correlation between the frames. This spectral correlation over the box boundaries is not obtained by generating only frequency boxes and performing an adjustment of the spectral envelope on these raw data from the frequency boxes.
Fig. 7c illustrates a comparison of an original signal (left panel) and an extended signal without Temporal Box Modeling (TTS). It can be seen that there are strong faults illustrated by the enlarged portions in the upper frequency range illustrated at 750. This, however, does not occur in Fig.
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0059.tif" />
7e when the same spectral portion at 750 is compared to the component related to faults 750 in Fig. 7c.
The embodiments or the audio coding system of the invention use the main bit rate available to waveform encode only the most perceptually relevant signal structure in the encoder, and the resulting spectral spaces are filled in. the decoder with the content of the signal that roughly approximates the original spectrum. A very limited bit budget is consumed to control the so-called Intelligent Space Fill (IGF) based on parameters by
Dedicated side information transmitted from the encoder to the decoder.
The storage or transmission of audio signals is often subject to strict bit rate limitations. In the past, encoders have been forced to dramatically reduce the transmitted audio bandwidth when only a very low bit rate was available. Modern audio codes are now capable of encoding broadband signals using bandwidth extension (BWE) methods such as Spectral Bandwidth Replication (SBR) [1]. These algorithms are based on a parametric representation of the high-frequency (HF) content that is generated from the wave-encoded low-frequency (LF) portion of the decoded signal by transposition to the high-frequency spectral ruler. (HF) (interconnect) and the application of parameter-based post-processing. In bandwidth extension schemes (BWE), the reconstruction of the high frequency spectral region (HF) above a
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0060.tif" />
So-called determined crossover frequency is often based on spectral interconnection. In general, the high frequency region (HF) consists of multiple adjacent connections and each of these connections is derived from bandpass regions (BP) of the low frequency spectrum (LF) below the determined crossover frequency. Current state of the art systems efficiently perform interconnection within a representation of filter banks by copying a set of coefficients from adjacent subbands from a source region to the destination region.
If a BWE system is implemented in a filter bank or the time-frequency transform domain, there is only a limited possibility to control the temporal shape of the bandwidth extension signal.
Generally, temporal granularity is limited by the jump size used between adjacent windows of the transform. This can lead to unwanted pre or post echoes in the bandwidth extension spectral range (BWE).
In perceptual audio coding, it is known that the time envelope shape of an audio signal can be restored using spectral filtering techniques such as Temporal Envelope Modeling (TNS) [14]. However, the TNS filter known from the current state of the art is a true value filter in the true value spectra. Such actual value filter on actual value spectra can be seriously affected by overlap failures, especially if the underlying actual transform is a Modified Discrete Cosine Transform (MDCT).
IMPI
MSXtCANO INSTITUTE • f. THE INDUSTRIAL PROPERTY
<img file="MX340575B_D0061.tif" />
Box modeling of the time envelope applies complex filtering on complex value spectra, such as those obtained, for example, through a Complex Modified Discrete Cosine Transform (CMDCT). In this way, overlap failures are avoided.
Temporal box modeling consists of • estimation of complex filter coefficients and application of a flattening filter on the spectrum of the original signal in the encoder • transmission of filter coefficients in lateral information • application of a modeling filter on the reconstructed spectrum filling boxes in the decoder
The invention extends the state-of-the-art technique known from audio transform coding, specifically Temporal Noise Modeling (TNS) by linear prediction along the frequency direction, for use in a form modified in the context of bandwidth extension.
Likewise, the invention's bandwidth extension algorithm is based on Intelligent Space Fill (IGF), but employs an oversampled Complex Value Transform (CMDCT) as opposed to the standard Intelligent Space Fill (IGF) configuration. which is based on a representation of the modified discrete cosine transform (MDCT)
<img file="MX340575B_D0062.tif" />
critically sampled of actual value of a signal. The CMDCT can be seen as the combination of the MDCT coefficients in the real part and the modified discrete sinusoidal transform (MDST) coefficients in the imaginary part of each complex value spectral coefficient.
Although the new approach is described in the context of IGF, the processing of the invention can be used in combination with any bandwidth extension (BWE) method that is based on a representation of filter banks of the audio signal.
In this new context, linear prediction along the frequency direction is not used as temporal noise modeling, but rather as a temporal box modeling (TTS) technique. The name change is justified by the fact that the box-filled signal components are modeled temporarily by TTS compared to the quantization noise modeling by TNS in the perceptual transform codes of the current state of the art.
Fig. 7a shows a block diagram of a BWE encoder using IGF and the new TTS approach.
Therefore, the basic coding scheme works as follows:
calculate the CMDCT of a time domain signal x (n) to obtain the frequency domain signal X (k) calculate the complex value TTS filter
<img file="MX340575B_D0063.tif" />
INSTITUTO MEXICANO DT LA PROPIEDAD INDUSTRIAL obtain the lateral information for the BWE and eliminate the spectral information that has to be replicated by the decoder apply the quantification using the psychoacoustic module (PAM, for its acronym in English) store / transmit the data, only transmit the coefficients of
Actual Value MDCT
Fig. 7b shows the corresponding decoder. This mainly reverses the steps performed in the encoder.
In this instance, the basic decoding scheme works as follows:
estimate the coefficients of the modified discrete sinusoidal transform (MDST) from the values of the modified discrete cosine transform (MDCT) (this processing adds a block decoder delay) and combine the coefficients of the MDCT and the MDST into coefficients of the Complex Value Discrete Modified Cosine Transform (CMDCT) Carry out frame filling with further processing Apply Time Frame Modeling Filtering (TTS) inverse with the transmitted coefficients of the TTS filter calculate the inverse CMDCT
Alternatively, it should be noted that the order of TTS synthesis and IGF post-processing can also be reversed at the decoder if the
JL JL w A JL JL
MEXICAN INSTITUTE DS THE INDUSTRIAL FRCPÍÍDAD
<img file="MX340575B_D0064.tif" />
TTS analysis and estimation of IGF parameters are systematically reversed in the encoder.
For efficient transform encoding, so-called long blocks of approximately 20 ms should preferably be used to achieve reasonable transform gain. If the signal within said long block contains transients, audible pre- and post-echoes occur in the reconstructed spectral bands due to the filling of frames. Fig. 7c shows typical pre-and post-echo effects that alter transients due to Intelligent Space Fill (IGF). In the left panel of Fig. 7c the spectrogram of the original signal is shown and in the right panel the spectrogram of the box fill signal is shown without the filtering of the temporal box modeling (TTS) of the invention. In this example, the IGF start frequency f<sub>l6Fssart</sub> or fspiit between the center band and the box-filled band is selected as /<sub>s</sub>/4. In the right panel of Fig. 7c different pre and post echoes are observed around the transients, especially prominent at the upper spectral end of the replicated frequency region.
The main task of the TTS module is to restrict these unwanted signal components in close proximity around a transient and thereby hide them in the temporal region governed by the temporal masking effect of human perception. Therefore, the required prediction coefficients of TTS are calculated and applied using direct prediction in the CMDCT domain.
IMPI
INSTITUTO MSXICANO DR LA PROPIEDAD INDUSTRIAL
<img file="MX340575B_D0065.tif" />
In an embodiment that combines TTS and IGF in a codec, it is important to align certain TTS parameters and IGF parameters so that an IGF box is fully filtered by a TTS filter (flattening or modeling filter) or not. Therefore, all TTSstart [..] or TTSstop [..] frequencies will not fall within an IGF box, but rather will be aligned with frequencies f<sub>IGF</sub> respective. Fig. 7d shows an example of TTS and IGF operating areas for a set of three filter filters.
TTS.
The end frequency of time frame modeling (TTS) is set to the end frequency of the Intelligent Space Fill Tool (IGF), which is greater than f<sub>lGFstart</sub>. If time frame modeling (TTS) uses more than one filter, you have to make sure that the crossover frequency between two TTS filters has to match the split frequency of the smart space fill (IGF). Otherwise, a TTS sub-filter will exceed the limit f¡<sub>GFttan</sub> This will lead to unwanted failures such as over-modeling.
In the implementation variant shown in Fig. 7a and Fig. 7b, special care must be taken that in that decoder, the IGF powers are correctly adjusted. This is especially so if, in the course of TTS and IGF processing, different TTS filters that have different prediction gains are applied to the source region (such as a flattening filter) and the target spectral region (such as a modeling that is not the exact counterpart of said flattening filter) of a box of the
ΙΜΡΙ
INSTITUTO MIXICANO SM LA INDUSTRIAL FROFIBDAD
<img file="MX340575B_D0066.tif" />
IGF. In this case, the ratio of the prediction gain of the two applied TTS filters is no longer equal to one, and therefore a power adjustment must be applied for this ratio.
In the alternate implementation variant, the IGF and TTS post-processing order is reversed. In the decoder, this means that the power adjustment by IGF post-processing is calculated after TTS filtering and is thus the final processing step before the synthesis transform. Therefore, regardless of the different TTS filter gains applied to a frame during encoding, the final power is always correctly adjusted by IGF processing.
On the decoder side, the TTS filter coefficients are applied across the entire spectrum again, ie the core spectrum spread over the regenerated spectrum. Time frame modeling (TTS) application is necessary to form the time envelope of the regenerated spectrum to adapt to the original signal envelope again. Therefore, the illustrated preecos are reduced. Additionally, it still temporarily models the quantization noise in the signal below / j<sub>CFítarí</sub> as usual in the prior art temporal noise modeling (TNS).
In the prior art encoders, the spectral interconnection of an audio signal (for example, Spectral Bandwidth Replication (SBR)) alters the spectral correlation at the interconnection limits and therefore affects the time envelope of the audio signal by introducing dispersion. Thus,
ΙΜΡΪ
Another advantage of applying Smart Space Fill (IGF) box fill to the residual signal is that, after applying the Time Box Modeling Filter (TTS), the box boundaries are They correlate perfectly, resulting in a more faithful temporal reproduction of the signal.
The result of the corresponding processed signal is shown in Fig.
7e. In comparison, the unfiltered version (Fig. 7c, right panel) the filtered TTS signal shows a good reduction of unwanted pre- and post-echoes (Fig. 7e, right panel).
Also, according to the description, Fig. 7a illustrates an encoder that matches the decoder of Fig. 7b or the decoder of Fig. 6a. Basically, an apparatus for encoding an audio signal comprises a time spectrum converter such as 702 for converting an audio signal to a spectral representation. The spectral representation can be a real value spectral representation or, as illustrated in block 702, a complex value spectral representation. In addition, a prediction filter such as 704 is provided to perform a prediction on the frequency to generate spectral residual values, where the prediction filter 704 is defined by the prediction filter information obtained from the audio signal and sent to a bitstream multiplexer 710, as illustrated at 714 in Fig.
7a. Also, an audio encoder such as psychoacoustically activated audio encoder 704 is provided. The audio encoder is configured to encode a first set of first spectral portions of the spectral residual values to obtain a first encoded set of first
<img file="MX340575B_D0067.tif" />
i ivi ri
MEXICAN INSTITUTE
OF OsoseraSiM PROPERTY
INDUSTRIAL spectral values. Additionally, a parametric encoder such as that illustrated at 706 in Fig. 7a is provided to encode a second set of second spectral portions. Preferably, the first set of first spectral portions is encoded with higher spectral resolution compared to the second set of second spectral portions.
Lastly, as illustrated in Fig. 7a, an output interface is provided to output the coded signal comprising the second parametrically coded set of second spectral portions, the first coded set of first spectral portions, and the Filter Information illustrated as Lateral Box Modeling Information (TTS) at 714 in Fig. 7a.
Preferably, the prediction filter 704 comprises a filter information calculator configured to use the spectral values of the spectral representation to calculate the filter information. Also, the prediction filter is configured to calculate the spectral residual values using the same spectral values of the spectral representation used to calculate the filter information.
Preferably, the TTS filter 704 is configured in the same manner known to prior art audio encoders that apply the Temporal Noise Modeling Tool (TNS) in accordance with the Advanced Audio Coding Standard (AAC).
Subsequently, an additional application using two-channel decoding is discussed in the context of Figs. 8a to 8e. In addition, reference is made
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0068.tif" />
to the description of the corresponding elements in the context of Figs. 2a, 2b (joint channel encoding 228 and joint channel decoding
204).
Fig. 8a illustrates an audio decoder for generating a two channel decoded signal. The audio decoder comprises four 802 audio decoders to decode a two channel encoded signal to obtain a first set of first spectral portions and additionally a parametric decoder 804 to provide parametric data for a second set of second spectral portions and, additionally, a two-channel identification that identifies, either a first or a second different representation of two channels for the second spectral portions. Additionally, a frequency regenerator 806 is provided to regenerate a second spectral portion based on a first spectral portion of the first set of first spectral portions and parametric data for the second portion and the identification of two channels for the second portion. Fig. 8b illustrates different combinations for the representations of two channels in the source range and in the destination range. The source range can be in the first two-channel representation, and the destination range can also be in the first two-channel representation. Alternatively, the source range can be in the first two-channel representation and the destination range can be in the second two-channel representation. Also, the source range can be in the second representation of two channels and the destination range can be in the second representation of
IMPI
INSTITUTO l * eX! CANO DE LA PROPIEDAD,, INDUSTRIAL
<img file="MX340575B_D0069.tif" />
first representation of two channels as indicated in the third column of Fig. 8b. Finally, both the source range and the destination range can be in the second representation of two channels. In one embodiment, the first two-channel representation is a separate two-channel representation 5, wherein the two channels of the two-channel signal are represented individually. So, the second two-channel representation is a joint representation where the two channels of the two-channel representation are represented together, that is, when further processing or rendering transformation is necessary to recalculate a representation of two separate channels that is required for output to the corresponding speakers.
In an implementation, the first two-channel representation may be a left / right (L / R) representation and the second two-channel representation is a joint stereo representation. However, 15 other two-channel representations in addition to left / right or M / S or stereo prediction can be applied and used for the present invention.
Fig. 8c illustrates a flow chart for the operations carried out by the audio decoder of Fig. 8a. In a step 812, the audio decoder 802 performs a decoding of the source range. The origin range 20 may comprise, with respect to Fig. 3a, scale factor bands SCB1 to SCB3. Also, there may be a two-channel identification for each scale factor band and scale factor band 1 may be, for example, in the first representation (such as L / R) and the third scale factor band
-Λ .at
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0070.tif" />
it may be in the second two-channel representation such as M / S or downmix / residual prediction. Therefore, step 812 can result in different representations for different bands. Then, in step 814, the frequency regenerator 806 is configured to select a source range for a frequency regeneration. At step 816, frequency regenerator 806 then checks the representation of the source range, and at block 818, frequency regenerator 806 compares the two-channel representation of the source range with the two-channel representation of the destination range. If both representations are identical, the frequency regenerator 806 provides a separate regeneration frequency for each channel of the two-channel signal. When, however, both representations detected in block 818 are not identical, then signal flow 824 is taken and block 822 computes the other two-channel representation of the source range and uses this other calculated two-channel representation to regeneration of the target rank. Therefore, the decoder in Fig. 8a makes it possible to regenerate a destination range that is indicated to have the second two-channel identification using a source range that is in the first two-channel representation. Naturally, the present invention allows to regenerate, in addition, a destination range using an origin range that has the same identification of two channels. And, additionally, the present invention allows to regenerate a destination range that has a two-channel identification indicating a joint representation of two channels and then transform this representation into a representation.
INSTITUTE A «XICANO
OF PROPERTY C «« 3ÍLMr
INDUSTRIAL channels separately, required for storage or transmission to the corresponding speakers for the two-channel signal.
It is emphasized that the two channels in the two-channel representation can be two stereo channels, such as the left channel and the right channel. However, the signal can also be a multi-channel signal that has, for example, five channels and one subwoofer channel, or that has even more channels. Then, a pairwise two-channel processing described in the context of Fig. 8a to 8e can be carried out when the pairs can be, for example, a left channel and a right channel, a left surround channel and a right surround channel and a center channel and an LFE (subwoofer) channel. Any other pair formation can be used to represent, for example, six input channels by three two-channel processing procedures.
Fig. 8d illustrates a block diagram of a decoder of the invention corresponding to Fig. 8a. A source range or central decoder 830 may correspond to the audlo 802 decoder. The other blocks 832, 834, 836, 838, 840, 842 and 846 may be parts of the frequency regenerator 806 of Fig. 8a. In particular, block 832 is a representation transformer for transforming representations of the source range into individual bands so that, at the output of block 832, a complete set of the source range is present in the first representation on one side and in the second representation of two channels on the other hand. These two representations
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0071.tif" />
Full range of origin can be stored in storage 834 for both representations of the origin range.
Then, block 836 applies a generation of frequency frames using, as input, an ID of the source range and, in addition, using a two-channel ID as input for the destination range. Based on the two-channel ID for the destination range, the frequency frame generator accesses storage 834 and receives the two-channel representation of the source range that matches the two-channel ID for the entered destination range in the frequency box generator at 835. Therefore, when the two-channel ID for the target range indicates stereo co-processing, then frequency frame generator 836 accesses storage 834 in order to obtain the co-stereo representation of the source range indicated by the ID of the source range 833.
The frequency frame generator 836 performs this operation for each destination range and the output of the frequency frame generator is such that each channel of the channel representation identified by the two channel identification is present. An envelope regulator 838 then performs an envelope adjustment. The envelope adjustment is carried out in the two-channel domain identified by the two-channel identification. For this purpose, envelope adjustment parameters are required and these parameters are transmitted from the encoder to the decoder in the same two-channel representation described. When the identification of two channels in the target range to be processed by the envelope regulator has a
IMPI
MEXICAN INSTmjTO • E LA «OPISDA · INDUSTRIAL
<img file="MX340575B_D0072.tif" />
Two-channel identification indicating a two-channel representation of the envelope data for this target range, then a parameter transformer 840 transforms the envelope parameters into the required two-channel representation. When, for example, the identification of two channels for a band indicates the stereo joint coding and when the parameters for this destination range have been transmitted as parameters of the L / R envelope, then the parameter transformer computes the joint parameters of the stereo envelope from the parameters of the L / R envelope described so that the correct parametric representation is used for adjusting the spectral envelope of a target range.
In another preferred embodiment, the envelope parameters are already transmitted as stereo set parameters when using the set stereo in a target band.
When the input to the envelope controller 838 is assumed to be a set of target ranges having different representations of two channels, then the output of the envelope controller 838 is also a set of target ranges to different representations of two channels. channels.
When a destination range has a joint representation such as M / S, then this destination range is processed by a representation transformer 842 to calculate the separate representation required for storage or transmission to the speakers. However, when a
MEXICAN INSTITUTE! C¡¡ * ^ g5é ^ DF THE PROPERTY QexsñdaU
INDUSTRIAL * target range already has a separate representation 844 signal flow is taken and 842 representation transformer is avoided. At the output of block 842, a two-channel spectral representation is obtained which is a separate two-channel representation that can then be further processed as indicated by block 846, where this additional processing may be, for example, a conversion of frequency / time or any other processing required.
Preferably, the second spectral portions correspond to the frequency bands, and the identification of two channels is provided as an array of labels corresponding to the table in Fig. 8b, where there is a label for each frequency band. Then, the parametric decoder is configured to check whether the tag has been set or not and to control the frequency regenerator 106 according to a tag to use, either a first representation or a second representation of the first spectral portion.
In one embodiment, only the reconstruction range that starts with the intelligent gap fill (IGF) start frequency 309 of Fig. 3a has two-channel identifications for different reconstruction bands. In another embodiment, this also applies for the frequency range below the IGF 309 start frequency.
In a further embodiment, the source band identification and the destination band identification can be adaptively determined by a similarity analysis. However, the processing of two
ΙΜΡΪ
MEXICAN INSTITUTE INDUSTR INDUSTRIAL PROPERTY Invention channels can also be applied when there is a fixed association
<img file="MX340575B_D0073.tif" />
from an origin range to a destination range. A source range can be used to recreate, with respect to frequency, a wider target range, either by a harmonic frequency frame fill operation or a copy frequency frame fill operation using two or more processing-like frequency frame fill operations for multiple known interconnects from high-efficiency Advanced Audio Coding (AAC) processing.
Fig. 8e illustrates an audio encoder for encoding a two channel audio signal. The encoder comprises an 860 time spectrum converter to convert the two-channel audio signal into a spectral representation. Also, a 866 spectral analyzer to convert the two-channel audio channel audio signal into a spectral representation. In addition, a spectral analyzer 866 is provided to carry out an analysis to determine the spectral portions that will be encoded with a high resolution, that is, to discover the first set of first spectral portions and to additionally discover the second set of second portions. spectral.
Additionally, a two-channel analyzer 864 is provided to analyze the second set of second spectral portions to determine a two-channel identification that identifies a first two-channel representation or a second two-channel representation.
IMPI (NSTITUTO MEXICANO ΐζϊ ^ βΜβκβΓ GIVE THE PROPIEPAC IJ · —-iT ^
INDUSTRIAL
Depending on the result of the two-channel analyzer, a band in the second spectral representation is parameterized using either the first two-channel representation or the second two-channel representation, and this is done using an 868 parameter encoder. The frequency range ie the frequency band below the start frequency of the intelligent space fill (IGF) 309 of Fig. 3a is encoded by a central encoder 870. The output from blocks 868 and 870 is input to an output interface 872. As noted above, the two-channel analyzer provides two-channel identification for each band, either above the IGF start frequency or for the entire frequency range, and this two-channel identification is also sent to the 872 Output Interface so that this data is also included in an 873 encoded signal emitted by the 872 output interface.
It is also preferred that the audio encoder comprises a strip transformer 862. Based on the decision of the two channel analyzer 862, the output signal of the time spectrum converter 862 is transformed into a representation indicated by the two channel analyzer. channels and in particular by the two-channel ID 835. Therefore, an output of the band transformer 862 is a set of frequency bands where each frequency band may be in the first two-channel representation or the second different two-channel representation. When the present invention is applied in full band, that is to say when both ranges, the origin range and the reconstruction range, are processed by the band transformer, the
IMPI λ MEXICAN INSTITUTE OF INDUSTRIAL POTOPTY
<img file="MX340575B_D0074.tif" />
860 spectral analyzer can analyze this representation. Alternatively, however, the 860 spectral analyzer can also analyze the signal output by the time spectrum converter indicated by the 861 control line. Therefore, the spectral analyzer 860 can apply the preferred hue analysis to the output of the 862 band transformer or the output of the 860 time spectrum converter before being processed by the 862 band transformer. Likewise, the spectral analyzer can apply the identification of the best adaptation source range for a certain target range, either in the result of the 862 band transformer or in the result of the 860 time spectrum converter.
Later reference is made to Figs. 9a to 9d to illustrate a preferred calculation of the power information values already discussed in the context of Fig. 3a and Fig. 3b.
Modern state-of-the-art audio encoders apply 15 different techniques to minimize the amount of data that represents a given audio signal. Audio encoders such as unified speech and audio encoding (USAC) [1] apply a time-to-frequency transformation such as the modified discrete cosine transform (MDCT) to obtain a spectral representation of a given audio signal.
These MDCT coefficients are quantified taking advantage of the psychoacoustic aspects of the human auditory system. If the available bit rate is reduced, the quantization becomes thicker by entering a large number of spectral values reduced to zero that generate audible faults on the side of the
IMPI
MEXICAN INSTITUTE OF PROPERTY
<img file="MX340575B_D0075.tif" />
INDUSTRIAL - 7 ^ decoder. In order to improve the quality of perception, the decoders of the state of the art fill these spectral parts reduced to zero with random noise. The Intelligent Gap Fill Method (IGF) collects frames from the remaining nonzero signal to fill those gaps in the spectrum. It is crucial for the perceptual quality of the decoded audio signal that the spectral envelope and power distribution of the spectral coefficients be preserved. The power adjustment method presented in the present invention uses the transmitted lateral information to reconstruct the MDCT spectral envelope of the audio signal.
Within spectral bandwidth replication (eSBR) [15] the audio signal is subsampled by at least a factor of two and the high-frequency portion of the spectrum is completely reduced to zero [1,17]. This removed part is replaced by parametric techniques, eSBR, on the decoder side. ESBR involves the use of an additional transform, the Quadrature Mirror Filter (QMF) transform, which is used to replace the empty high-frequency portion and to resample the audio signal [17]. This adds computational complexity and memory consumption to an audio encoder.
The USAC encoder [15] offers the possibility of filling spectral gaps (spectral lines reduced to zero) with random noise but it has the following drawbacks: random noise cannot preserve the temporal fine structure of a transient signal and the harmonic structure of a tonal signal.
IMPI
INSTITUTO MSXICANO OI LA PIONBOA ·
INDUSTRIAL
<img file="MX340575B_D0076.tif" />
The area where the eSBR operates on the decoder side was completely eliminated by the encoder [1]. Therefore, the eSBR tends to eliminate tonal lines in the high frequency region or distort the harmonic structures of the original signal. As the frequency resolution of quadrature mirror filters (QMF) of the spectral bandwidth replication (eSBR) is very low and the reinsertion of sinusoidal components is only possible in the coarse resolution of the underlying filter bank, the regeneration of components tonal values in the eSBR in the replicated frequency range have very little precision.
The eSBR uses techniques to adjust the powers of the interconnected areas, adjusting the spectral envelope [1]. This technique uses the transmitted power values on a QMF frequency time grid to reshape the spectral envelope. This state of the art does not deal with partially removed spectra and due to the high temporal resolution it tends to require a relatively large amount of bits to transmit appropriate power values or to apply coarse quantization to the power values.
The IGF method does not require an additional transformation, as it uses the prior art MDCT transformation that is calculated as described in [15],
The power adjustment method presented in the present invention uses the lateral information generated by the encoder to reconstruct the envelope
IMPI
MEXICAN INSTITUTE OF. INDUSTRIAL PROPERTY
<img file="MX340575B_D0077.tif" />
spectral of the audio signal. This lateral information is generated by the encoder as indicated below:
a) Apply a window-split modified discrete cosine transform (MDCT) to the input audio signal [16, section 4.6], optionally calculate a window-split modified discrete sinusoidal transform (MDST), or estimate a window-split MDST to from the calculated MDCT.
b) Apply Temporal Noise Modeling (TNS) / Temporal Box Modeling (TTS) on MDCT coefficients [15, section 7.8]
c) Calculate the average power for each band of the MDCT scale factor above the intelligent space filling start frequency (IGF) (f<sub>lGFscart</sub>) up to the end frequency of IGF (/<sub>ÍCmep</sub>)
d) Quantify the mean power values fieman AND ficFstop are parameters given by the user.
The values calculated in step c) and d) are lossless encoded and transmitted as bitstream side information to the decoder.
The decoder receives the transmitted values and uses them to adjust the spectral envelope.
a) Dequantify the transmitted values of the MDCT
b) Apply the noise filler of the prior art unified voice and audio coding (USAC) if indicated
IMPI
c)
MEXICAN INSTITUTE Μ LA rWONEDAP
INDUSTRIAL
<img file="MX340575B_D0078.tif" />
Apply the Smart Space Fill (IGF) box fill
d) Quantize the transmitted power values
e) Adjust the spectral envelope by scale factor band
f) Apply TNS / TTS if indicated
Let R * be the MDCT transform, the actual value spectral representation of an audio signal divided into windows of window length 2N. This transformation is described in [16]. The encoder optionally applies TNS in f.
[16, 4.6.2] describes a partition of x in scale factor bands.
Scale factor bands are a set of a set of indices and are indicated in this text with scb.
The limits of each scb<sub>k</sub>wíth k = Q, l, 2, ... max_sfb are defined by an array swb_offset (16, 4.6.2), where sw & _o // set [fe] and swt_o / fset [fc +1] -1 15 define the first and last index for the lowest and highest spectral coefficient line contained in scb<sub>k</sub>. The scale factor band is indicated as follows:
scb<sub>faith</sub>: = {swb_offset [k], 1 + swb_offset [k], 2 + swb_offset [k], ..., swb_offset [k + 1] -1}
If the IGF tool is used by the encoder, the user defines an IGF start frequency and an IGF end frequency. These two values
IMPI
MEXICAN INSTITUTE PE LA «OPISIMO
INDUSTRIAL
<img file="MX340575B_D0079.tif" />
they are mapped to the best fit scale factor band index igfStartSfb and igfStopSfb. Both are sent in the bitstream to the decoder.
[16] describes a long-block and short-block transformation. For long blocks, only one set of spectral coefficients together with a set of scale factors is transmitted to the decoder. For short blocks, eight short windows are calculated with eight different sets of spectral coefficients. To save the bit rate, the scale factors of said eight short block windows are grouped by the encoder.
In the case of Intelligent Space Fill (IGF), the method presented in this invention uses prior art scale factor bands to group spectral values that are transmitted to the decoder:
X<sup>2</sup>
Where k = igfStartSfb ,! + igfStartSfb, 2 + igfStartSfb,
To quantify, igfEndSfb is calculated.
B<sub>k</sub> = nlNT (4log<sub>2</sub>(AND<sub>k</sub>))
All É values<sub>k</sub> they are transmitted to the decoder.
The encoder is supposed to decide to group the scale factor sets num_window_group.
IMPIOS
MEXICAN INSTITUTE • E LA rrOIMEDAD
INDUSTRIAL
Indicated with w is this grouping-partition of the set {0,1,2, .., 7} which are the indices of the eight short windows. w<sub>2</sub> indicates the ith subset of
U7, where l indicates the index of the window group, 0 <l <num_wtndow_group. For the short block calculation, the user defined that the IGF start / end frequency is assigned to appropriate scale factor bands. However, for simplicity reasons it is also indicated for short blocks k = igfStartSfb ,! + igfStartSfb, 2 + igfStartSfb, igfEndSfb.
The IGF power calculation uses the grouping information to group the Se-
<img file="MX340575B_D0080.tif" />
to quantify, calculate = nINT (4log<sub>2</sub>(AND<sub>or</sub>)·)
All É values<sub>kl</sub> they are transmitted to the decoder.
The coding formulas mentioned above operate using only MDCT coefficients of real value δ. To obtain a more stable power distribution in the IGF range, i.e. to reduce temporal amplitude fluctuations, an alternative method can be used to calculate the values
AND<sub>k</sub>:
Let x<sub>r</sub> e R<sup>N</sup> let the MDCT transform be the actual value spectral representation of an audio signal divided into windows of window length 2N, and
ΙΜΡΙ
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0081.tif" />
x in<sup>N</sup> the spectral representation of the actual value MDST transform of the same portion of the audio signal. The spectral representation of the modified discrete sinusoidal transform (MDST) x, could be calculated or estimated exactly from c: = ($<sub>Γ</sub>Λ) € C<sup>N</sup> indicates the complex spectral representation of the windowed audio signal, which has x<sub>r</sub> as its real part and x, as its imaginary part. The encoder optionally applies temporal noise modeling (TNS) on x<sub>r</sub> and jq.
In this instance, the signal strength in the IGF range can be measured with
<img file="MX340575B_D0082.tif" />
The real and complex value powers of the reconstruction band, i.e. the box to be used on the decoder side in the reconstruction of the IGF scb range<sub>k</sub>, is calculated with:
'«« K
<img file="MX340575B_D0083.tif" />
where tr<sub>k</sub> is a set of indexes - the associated origin box range, based on scb<sub>k</sub>. In the two previous formulas, instead of the set
<img file="MX340575B_D0084.tif" />
i
IMPI
MEXICAN INSTITUTE OF THE FROPIEOAD
<img file="MX340575B_D0085.tif" />
of scb indices<sub>k</sub> you could use the scb set<sub>k</sub> (defined later in this text) to create tr<sub>k</sub> to achieve more accurate values E<sub>t</sub> and E<sub>r</sub>.
Calculate <sup>AND</sup>okay
Etk if Ε *> 0, otherwise f<sub>k</sub> = 0.
With:
<sup>AND</sup>k => / fkErk now a more stable version of E is calculated<sub>k</sub>, since a calculation of E<sub>k</sub> with 10 the MDCT values are only affected by the fact that the MDCT values do not obey Parseval's theorem and therefore do not reflect the full power information of the spectral values. AND<sub>k</sub> it is calculated as indicated above.
As noted above, for cut blocks it is assumed that the encoder decides to group the scale factor sets num_window_group.
As above, w¡ indicates the f-th subset of «?, Where l indicates the index of the window group, 0 <l <num_window ^ group.
Once again the alternative version described above could be applied to calculate a more stable version of Ejj. With the definitions of c: = € C<sup>M</sup> , x<sub>r</sub>eB<sup>N</sup> which is the modified discrete cosine transform
<img file="MX340575B_D0086.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY (MDCT) yx¡ eR<sup>N</sup> which is the audio signal divided into windows of length 2N of the modified discrete sinusoidal transform (MDST), calculate
EoW = iv<sup>-1</sup> i | wj | L Iscbkl lew¡
<img file="MX340575B_D0087.tif" />
ic scbfc
Calculate analogously
<img file="MX340575B_D0088.tif" />
and continue with the factor, <sup>Epki </sup>»Ια = ΐ— <sup>h</sup>tkJ which is used to adjust E<sub>rkJ</sub> previously calculated:
Ευ is calculated as indicated above.
The procedure that not only uses the power of the reconstruction band, whether derived from the complex reconstruction band or the MDCT values, but also uses power information from the source range provides an improved power reconstruction.
Specifically, the parameter calculator 1006 is configured to calculate the power information for the reconstruction band using information on the power of the reconstruction band and, furthermore, using the information on a power of a source range to be used for the reconstruction of the reconstruction band.
<img file="MX340575B_D0089.tif" />
Also, parameter calculator 1006 is configured to calculate power information (E<sub>or</sub>k) in the reconstruction band of a complex spectrum of the original signal, to calculate additional power information (E,<sub>k</sub>) in a source range of a real value part of the complex spectrum of the original signal to be used to reconstruct the reconstruction band, and where the parameter calculator is configured to calculate the power information for the reconstruction band using the power information (E<sub>OR</sub>k) and the additional power information (E<sub>r</sub>k).
Furthermore, the parameter calculator 1006 is configured to determine a first power information (E<sub>okay</sub>) in a scale factor band to be reconstructed from a complex spectrum of the original signal, to determine a second power information (E<sub>t</sub>k) in a range of origin of the complex spectrum of the original signal to be used to reconstruct the scale factor band to be reconstructed, to determine a third power information (E<sub>r</sub>k) in a range of origin of a real value part of the complex spectrum of the original signal to be used to reconstruct the scale factor band to be reconstructed, to determine a weighting information based on a relationship between at minus two of the first power information, the second power information, and the third power information, and to weight one of the first power information and the third power information using the weight information to obtain a weighted power information and to use the
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX340575B_D0090.tif" />
weighted power information as the power information for the reconstruction band.
Examples for calculations are presented below although many other examples may be left to those of skill in the art in light of the above general principle:
A) f_k = E_ok / E_tk;
E_k - sqrt (f_k * E_rk);
B) f_k = E_tk / E_ok;
E_k = sqrt ((1 / f_k) * E_rk);
C) f_k = E_rk / E_tk;
E_k = sqrt (f_k * E_ok)
D) f_k = E_tk / E_rk;
E_k = sqrt ((1 / f_k) * E_ok)
All of these examples confirm that although only actual MDCT values are processed on the decoder side, the actual calculation is - due to overlap and addition - of the implicit time domain overlap cancellation procedure using complex numbers. However, in particular, the determination 918 of the power information of
IMPI
MEXICAN INSTITUTE OF LA FROPIEOAU
INDUSTRIAL
<img file="MX340575B_D0091.tif" />
Inset of additional spectral portions 922, 923 of reconstruction band 920, for different frequency values of the first spectral portion 921 having frequencies in reconstruction band 920, is based on actual values of the MDCT. Therefore, the power information transmitted to the decoder will generally be less than the Power Information E<sub>OR</sub>k over the reconstruction band of the complex spectrum of the original signal. For example, for case C above, this means that the factor f_k (weight information) will be less than 1.
On the decoder side, if the Intelligent Space Fill Tool (IGF) is pointed to ON, the transmitted values B<sub>k</sub> are obtained from the bit stream and will be quantized with
<img file="MX340575B_D0092.tif" />
for all faith = ÍgfStartSfb ,! + tgfStartSfb, 2 + iggStartSfb, igfEndSfb.
A decoder quantizes the transmitted values of the MDCT ax R * 15 and calculates the remaining holding power:
where k is in the range defined above.
We indicate that scb<sub>k</sub> = {£ | í e scb<sub>k</sub>Ax<sub>i</sub> = 0}. This set contains all the indices for the scb scale factor band<sub>k</sub> they have been zeroed by the encoder.
IMPIAS
INSTITUTE Μ EM CANO DE LA PROPIEDAD
INDUSTRIAL
The Intelligent Gap Fill Subband (IGF) method (not described in the present invention) is used to fill spectral gaps that result from coarse quantization of MDCT spectral values on the encoder side using non-zero values of the transmitted MDCT, x will additionally contain the values that replace all previous values reduced to zero. The power of the box is calculated by:
tea<sub>fc</sub>: = £ x?
where k is in the range defined above.
The missing power in the reconstruction band is calculated by:
£ m<sub>fc</sub> - | E<sub>fc</sub><sup>2</sup> - I know<sub>k</sub> g =
And the gain factor for the adjustment is obtained by:
less
- YES (mE<sub>k</sub> > 0 Λ tE<sub>k</sub> > 0) íE<sub>faith</sub> on the contrary
With:
g '= min (^ 10)
The spectral envelope setting using the gain factor is:
= 9% for all ie scb<sub>k</sub> and faith is in the range defined above.
IMPI
MEXICAN INSTITUTE OF INDTTTIjlAL PROPERTY
<img file="MX340575B_D0093.tif" />
This reshapes the spectral envelope of x to the shape ^ dTIa<sup>1</sup> envelope shape of the original spectral envelope x.
In principle, with the short window sequence, all the calculations defined above remain the same, but the grouping of scale factor bands is taken into account. It is indicated as E<sub>k!</sub> the grouped and quantized power values, obtained from the bit stream. Calculate
<img file="MX340575B_D0094.tif" />
/ «Wj}> i
The / index describes the window index of the short block sequence.
Calculate
O Λ pE<sub>kl</sub> > O) or otherwise
With g '= 10)
Apply <sup>χ</sup>μ ·
9<sup>> x</sup>
IMPI
MBXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0095.tif" />
for all ie scb<sub>kl</sub>.
For low bit rate applications, pairwise grouping of E values is possible<sub>k</sub> without losing too much precision. This method applies 5 only with long blocks:
«?
where k = igfStartSfb ,! + IgfStartSfb, 4 + igfStartSfb, ...<sub>t</sub>igfEndSfb,
Again, after dequantization, all £ values<sub>k</sub>^<sub>1</sub>they are transmitted to the decoder.
Fig. 9a illustrates an apparatus for decoding an encoded audio signal comprising an encoded representation of a first set of first spectral portions and an encoded representation of parametric data indicating the spectral powers for a second set of second spectral portions. The first set of first spectral portions is indicated in
901a in Fig. 9a, and the coded representation of the parametric data is
Indicated at 901b in Fig. 9a. An audlo 900 decoder is provided to decode the encoded representation 901a of the first set of first spectral portions to obtain a first decoded set of first spectral portions 904 and to decode the encoded representation of the parametric data to obtain decoded parametric data 902 for the second set of second spectral portions indicating the powers
IMPI
MEXICAN INSTITUTE OF LA PROPIBDA ·
INDUSTRIAL
<img file="MX340575B_D0096.tif" />
for the reconstruction bands, where the second spectral portions are located in the reconstruction bands. Furthermore, a frequency regenerator 906 is provided to reconstruct spectral values of a reconstruction band comprising a second spectral portion.
Frequency regenerator 906 uses a first spectral portion of the first set of first spectral portions and individual power information for the reconstruction band, wherein the reconstruction band comprises a first spectral portion and the second spectral portion. The frequency regenerator 906 comprises a calculator 912 for determining a holding power information comprising a cumulative power of the first spectral portion having frequencies in the band of the reconstruction. Likewise, the frequency regenerator 906 comprises a calculator 918 for determining frame power information from other spectral portions of the reconstruction band and for frequency values that are different from the first spectral portion, where these frequency values have frequencies. in the reconstruction band, wherein the other spectral portions must be generated by frequency regeneration using a different first spectral portion from the first spectral portion in the reconstruction band.
Frequency regenerator 906 further comprises a calculator 914 for a missing power in the rebuild band, and calculator 914 operates using the individual power for the rebuild band and the holding power generated by block 912. In addition, the regenerator
IMPI
INSTITUT · MEXICANO Dt LA PIOPISDA »HDUSTUAl
<img file="MX340575B_D0097.tif" />
906 The frequency module comprises a spectral envelope regulator 916 for adjusting the additional spectral portions in the reconstruction band based on the missing power information and the frame power information generated by block 918.
Referring to Fig. 9c, there is illustrated a certain reconstruction band 920. The reconstruction band comprises a first spectral portion in the reconstruction band such as the first spectral portion 306 in Fig. 3a schematically illustrated at 921. Also, the rest of the spectral values in the reconstruction band 920 must be generated using an origin region, for example, of the scale factor band 1, 2, 3 below the start frequency of the smart space filling 309 of Fig. 3a. Frequency regenerator 906 is configured to generate raw spectral values for second spectral portions 922 and 923. A gain factor g is then calculated as illustrated in Fig. 9c in order to finally adjust the raw spectral values in the frequency bands 922, 923 in order to obtain the second reconstructed and adjusted spectral portions in the reconstruction band 920, which now have the same spectral resolution, i.e. the same line distance as the first spectral portion 921. It is important to understand that the first spectral portion in the reconstruction band illustrated at 921 in Fig. 9c is decoded by audio decoder 900 and is not influenced by the envelope adjustment carried out by block 916 of Fig. 9b. Instead, the first spectral portion in the reconstruction band indicated at 921 is left as is, since this first spectral portion is
<img file="MX340575B_D0098.tif" />
broadcast by the full-bandwidth or full-rate audio decoder 900 over line 904.
Next, a specific example will be analyzed with real numbers. The remaining holding power calculated by block 912, for example, is five power units and this power is the power of the four spectral lines indicated by way of example in the first spectral portion 921.
Likewise, the value of the power E3 for the reconstruction band corresponding to the scale factor band 6 of Fig. 3b or Fig. 3a is equal to 10 units. It is important to note that the power value comprises not only the power of the spectral portions 922, 923, but also the total power of the reconstruction band 920 calculated on the encoder side, i.e. before carrying out the analysis spectral, using, for example, tonality masking. Therefore, the ten power units encompass the first and second spectral portions in the reconstruction band. So, the power of the source range data for blocks 922, 923 or the target range raw data for block 922, 923 is assumed to be equal to eight power units. Therefore, a missing power of five units is calculated.
A gain factor of 0.79 is calculated based on the missing power divided by the box power tEk. Then, the raw spectral lines for the second spectral portions 922, 923 are multiplied by the calculated gain factor. Thus, only the spectral values are adjusted for the second spectral portions 922, 923 and the lines
IMPI
MEXICAN INSTITUTE OF THE FROPIBDAP industrial
<img file="MX340575B_D0099.tif" />
Spectral for the first spectral portion 921 are not influenced by this envelope adjustment. After multiplication of the raw spectral values for the second spectral portions 922, 923 a complete reconstruction band has been calculated consisting of the first spectral portions in the reconstruction band, and consisting of spectral lines in the second spectral portions 922, 923 on reconstruction band 920.
Preferably, the source range for generating the raw spectral data in bands 922, 923 is, with respect to frequency, below the smart gap fill (IGF) start frequency 309 and reconstruction band 920 is above the IGF 309 Start frequency.
Furthermore, it is preferred that the limits of the reconstruction band coincide with the limits of the scale factor band. Therefore, a reconstruction band has, in one embodiment, the size of the respective scale factor bands of the central audio decoder or is dimensioned such that, when power pair formation is applied, a power value for a reconstruction band provides the power of two or more integers of scale factor bands. Therefore, when power accumulation is assumed to be carried out for scale factor band 4, scale factor band 5 and scale factor band, then the lower frequency limit of the reconstruction 920 is equal to the lower limit of the scale factor band 4 and the upper power limit of the reconstruction band 920 coincides with the upper limit of the scale factor band 6.
IMPI
MEXICAN INSTITUTE OFLAPRC
<img file="MX340575B_D0100.tif" />
Fig. 9d is now described in order to show the additional functionalities of the decoder of Fig. 9a. The audio decoder 900 receives the quantized spectral values corresponding to the first spectral portions of the first set of spectral portions and, additionally, the scale factors for the scale factor bands, as illustrated in Fig. 3b are provided to a 940 reverse scale adjustment block. The inverse scale adjustment block 940 provides all the first sets of first spectral portions below the IGF start frequency 309 of Fig. 3a and, additionally, the first spectral portions above the IGF start frequency, that is, the first spectral portions 304, 305, 306, 307 of Fig. 3a that are all located in a reconstruction band illustrated at 941 in Fig. 9d. Furthermore, the first spectral portions in the source band for the filling of frequency boxes in the reconstruction band are provided to the envelope regulator / calculator 942 and this block also receives the power information for the reconstruction band. provided as parametric side information of the encoded audio signal illustrated at 943 in Fig. 9d. Then the envelope regulator / calculator 942 provides the functionalities of Fig. 9b and 9c and finally outputs the adjusted spectral values for the second spectral portions in the reconstruction band. These adjusted spectral values 922, 923 for the second spectral portions in the reconstruction band and the first spectral portions 921 in the reconstruction band indicated on line 941 in Fig. 9d
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PRORITY
<img file="MX340575B_D0101.tif" />
together they represent the complete spectral representation of the reconstruction band.
Later reference is made to Figs. 10a to 10b to explain the preferred embodiments of an audio encoder that encodes an audio signal to provide or generate an encoded audio signal. The encoder comprises a time / spectrum converter 1002 that powers a spectral analyzer 1004, and the spectral analyzer 1004 is connected to a parameter calculator 1006 on the one hand and to an audio encoder 1008 on the other hand. Audio encoder 1008 provides the encoded representation of a first set of first spectral portions and does not span the second set of second spectral portions. Furthermore, parameter calculator 1006 provides power information for a reconstruction band spanning the first and second spectral portions . Also, audio encoder 1008 is configured to generate a first encoded representation of the first set of first spectral portions having the first spectral resolution, where audio encoder 1008 provides scale adjustment factors for all bands in the spectral representation. generated by block 1002. Additionally, as illustrated in Fig. 3b, the encoder provides power information for at least the reconstruction bands located, with respect to frequency, above the IGF 309 start frequency as illustrated in Fig. 3a. Therefore, in order for the reconstruction bands to preferably coincide with the scale factor bands or with groups of scale factor bands,
IMPI
NSTITUTO MEXICANO D € THE PROPERTY
INDUSTRIAL
<img file="MX340575B_D0102.tif" />
they provide two values, that is, the corresponding scale adjustment factor of the audio encoder 1008 and, additionally, the information of the power emitted by the parameter calculator 1006.
Preferably, the audio encoder has scale factor bands with different frequency bandwidths, that is, with a different number of spectral values. Therefore, the parametric calculator comprises a normalizer 1012 to normalize the powers for the different bandwidth with respect to the bandwidth of the specific reconstruction band. To this end, normalizer 1012 receives, as inputs, a power in the band and a number of spectral values in the band, and normalizer 1012 then outputs a normalized power per reconstruction band / scale factor band.
In addition, the parametric calculator 1006a of Fig. 10a comprises a power value calculator that receives control information from the audio or core encoder 1008 as illustrated on line 1007 in Fig. 10a.
This control information may comprise information about the long / short blocks used by the audio encoder and / or grouping information. Therefore, while the information on long / short blocks and the grouping information on short windows refer to a temporal grouping, the grouping information can also refer to a spectral grouping, that is, the grouping of two factor bands. Scale in a single reconstruction band. Therefore, the power value calculator 1014 outputs a single power value for each band.
IMPI Mexican institute Vcfy-; ;
OF PROPERTY V?<sup>31</sup>·^
INDUSTRIAL ------- * grouped covering a first and a second spectral portion when only the spectral portions have been grouped.
Fig. 10d illustrates a further embodiment for the implementation of spectral grouping. For this purpose, block 1016 is configured to calculate power values for two adjacent bands. Next, at block 1018, the power values for the adjacent bands are compared, and when the power values are not as different or less different than defined, for example, by a threshold, then a single value is generated. (normalized) for both bands as indicated in the block
1020. As illustrated on line 1019, block 1018 can be skipped. Also, the generation of a unique value for two or more bands that is carried out in block 1020 can be controlled by an encoder bit rate control.
1024. Therefore, when the bit rate is to be reduced, the 1024 bit rate encoded control controls block 1020 to generate a single normalized value for two or more bands, even when the comparison in block 1018 would not have been allowed for group the power information values.
In case the audio encoder performs the grouping of two or more short windows, this grouping is also applied for the power information. When the core encoder performs a grouping of two or more short blocks, then for these two or more blocks, only a single set of scaling factors is calculated and transmitted. On the decoder side, the audio decoder then applies the same set of scaling factors to both grouped windows.
<img file="MX340575B_D0103.tif" />
<img file="MX340575B_D0104.tif" />
Regarding the calculation of the power information, the spectral values ... in the reconstruction band accumulate over two or more short windows. In other words, this means that the spectral values in a given reconstruction band for a short block and for the subsequent short block are accumulated and only a single power information value is transmitted for this reconstruction band that spans two short blocks. . Next, on the decoder side, the envelope adjustment described in Fig. 9a to 9d is not carried out individually for each short block, but is carried out together for the set of grouped short windows.
Then the corresponding normalization is applied again so that, although any grouping in the frequency or temporal grouping has been carried out, the normalization easily allows that, for the calculation of the power value information on the decoder side, only know the value of the power information on the one hand and the number of spectral lines in the reconstruction band or in the set of grouped reconstruction bands.
In prior art bandwidth extension (BWE) schemes, the reconstruction of the high frequency spectral region (HF) above a so-called determined crossover frequency is often based on spectral interconnection. In general, the high frequency region (HF) consists of multiple adjacent connections and each of these connections is derived from bandpass regions (BP) of the low frequency spectrum (LF) below
<img file="MX340575B_D0105.tif" />
100
<img file="MX340575B_D0106.tif" />
of the determined crossover frequency. Within a filter bank representation of the signal, such systems copy a set of coefficients from adjacent subbands of the low frequency spectrum (LF) in the destination region. The limits of the selected sets usually depend on the system and do not depend on the signal. For some signal contents, this selection of static interconnect can cause unpleasant ringing and coloration of the reconstructed signal.
Other approaches transfer the low frequency (LF) signal to the high frequency (HF) signal adaptive single sideband modulation (SSB). These approaches are of high computational complexity compared to [1] since they operate at a high sampling rate in samples of time domain. Furthermore, the interconnection can become unstable, especially for non-tonal signals (for example, voiceless) and, therefore, the adaptive interconnection of the state of the art can introduce alterations in the signal.
The approach of the invention is called Intelligent Fill of Spaces (IGF) and, in its preferred configuration, it is applied in a bandwidth extension system (BWE) based on a temporal frequency transform such as, for example, the Transformed Discrete Cosine Modified (MDCT). However, the teachings of the invention are generally applicable, for example, analogously within a Quadrature Mirror Filter Bank (QMF) based system.
101
JL J.VJL J. JL
INSTITUTO MEXICANO »S LA FROPtlDAD INDUSTRIAL
<img file="MX340575B_D0107.tif" />
An advantage of the MDCT-based IGF configuration is the seamless integration into MDCT-based audio encoders, for example MPEG Advanced Audio Coding (AAC). Sharing the same transform for waveform audio encoding and for BWE significantly reduces the overall computational complexity for the audio codec.
Furthermore, the invention provides a solution to the inherent stability problems found in adaptive interconnection schemes of the state of the art.
The proposed system is based on the observation that for some signals, an unguided interconnection selection can generate timbre changes and colorations in the signal. If a signal that is tonal in the source spectral region (SSR) but is noise-like in the destination spectral region (STR), interconnecting the noise-like STR by the tonal SSR can generate an unnatural timbre. The timbre of the signal can also change as the tonal structure of the signal could be misaligned or even destroyed by the interconnection process.
The proposed IGF system performs intelligent box selection using cross-correlation as a measure of similarity between a
SSR in particular and a specific STR. The cross correlation of two signals provides a measure of the similarity of those signals and also the maximum correlation delay and its sign. Therefore, the correlation-based box selection approach can also be used to fit with
<img file="MX340575B_D0108.tif" />
102
<img file="MX340575B_D0109.tif" />
the spectral displacement of the copied spectrum is as accurate as possible as close to the original spectral structure as possible.
The fundamental contribution of the proposed system is the choice of an appropriate measure of similarity, and also techniques to stabilize the process of selecting boxes. The proposed technique provides an optimal balance between instantaneous signal adaptation and, at the same time, temporal stability. The provision of temporal stability is especially important for signals that have little SSR and STR similarity, and therefore exhibit low cross-correlation values or when using similarity measures 10 that are ambiguous. In such cases, stabilization prevents the heavy-random behavior of adaptive box selection.
For example, a class of signals that often poses problems for state-of-the-art bandwidth extension is characterized by a different concentration of power in arbitrary spectral regions, as shown in Fig. 12a (a the left). Although there are methods available to adjust the spectral envelope and tonality of the reconstructed spectrum in the target region, for some signals, these methods are not able to conserve timbre well as shown in Fig. 12a (on the right). In the example illustrated in Fig. 12a, the magnitude of the spectrum in the destination region of the original signal above a so-called crossover frequency f<sub>xover </sub>(Figure 12a, left) decreases almost linearly. On the contrary, in the
103
IMPI
INSTITUTO MEXICANO DE LA FROFISDAD INDUSTRIAL reconstructed spectrum (Fig. 12a, on the right) there is a different set of slopes and peaks that is perceived as a failure of timbre coloration.
An important step in the new approach is to define a set of boxes between which the choice based on subsequent similarity can take place. First, the boundaries of the boxes, both in the region of origin and the region of destination, have to be defined with each other. Therefore, the destination region between the IGF start frequency of the central encoder f,<sub>GFsíart</sub> and a higher frequency available f<sub>lGFatop</sub> is divided into an arbitrary integer number nTar of boxes, each of which has a predefined individual size. So for each tar [idx_tar] target box a set of source boxes of equal size src [tdx_src] is generated. Therefore, the basic degree of freedom of the IGF system is determined. The total number of source boxes nSrc is determined by the bandwidth of the source region, ^<sup>w</sup>src <sup>=</sup> (ftGFsean ~ fiGPrmn) where f<sub>1GFmin</sub> is the lowest frequency available for frame selection so that an integer nSrc of source frames fits in bw ^. The minimum number of source boxes is 0.
To further increase the degree of freedom for selection and fit, the source boxes can be defined to overlap each other by an overlap factor between 0 and 1, where 0 means no overlap and 1 means
<img file="MX340575B_D0110.tif" />
104
IMPI
MEXICAN INSTITUTE OF PROPERTY HDUSTWAL
<img file="MX340575B_D0111.tif" />
100% overlap. The case of 100% overlap implies that only one or no source box is available.
Fig. 12b shows an example of the box boundaries of a box set. In this case, all the target boxes are mapped to each of the source boxes. In this example, the source boxes overlap by 50%.
For a target box, the cross correlation is calculated with multiple source boxes at delays of up to xcorr_maxLag intervals. For a given target box tdxjtar and a source box idxsrc, xcorr_vaZ [idxjwj | [idx_sre] provides the maximum value of the absolute cross-correlation between the boxes, while xcorrja £ [idx_tar] [idx_src] provides the delay at which this maximum occurs and xcorr_si5n [¿dx_tarj [úíx_src] provides the sign of the cross correlation in xcorr_lag [idxjar] [idxjrrc].
The xcorr_lag parameter is used to control the proximity of the match between the source frame and the destination frame. This parameter results in fault reduction and helps to better preserve the timbre and color of the signal.
In some cases it may happen that the size of a specific destination box is larger than the size of the available source boxes. In this case, the available source box is repeated as often as necessary to completely fill in the specific destination box. Still
105
ΙΜΡΙ
MKXCANO INSTITUTE OF INDUSTRIAL HtOFlITY
<img file="MX340575B_D0112.tif" />
it is possible to perform cross-correlation between the large target box and the smallest source box in order to obtain the best position of the source box in the target box in terms of the xcorrlag cross-correlation delay and the sign xcorr_s¡gn.
The cross correlation of the raw spectral frames and the original signal may not be the most suitable measure of similarity applied to audio spectra with a strong formant structure. Spectrum bleaching removes the information from the raw envelope and therefore emphasizes the fine spectral structure that is of primary interest for assessing frame similarity. Bleaching also aids in easy modeling of the easy envelope of the target spectral ruler (STR) in the decoder for regions processed by IGF. Therefore, optionally, the frame and the origin signal are bleached before calculating the cross correlation.
In other configurations, only the frame is whitewashed using a predefined procedure. A transmitted bleach tag indicates to the decoder that the same predefined bleach process will be applied to the frequency box within the Intelligent Space Fill (IGF).
To whiten the signal, an estimate of the spectral envelope is first calculated. The spectrum of the MDCT is then divided by the spectral envelope. The spectral envelope estimate can be estimated on the MDCT spectrum, the MDCT spectrum powers, the MDCT-based complex power spectrum estimates, or the
106
M RXICAN INSTITUTE OF PROPERTY
INDUSTRIAL power spectrum. The signal at which the envelope is estimated will be called the base signal hereafter.
Envelopes calculated on the basis of complex power spectrum estimates based on the MDCT or power spectrum as the base signal have the advantage of not having time fluctuation in the tonal components.
IF the base signal is in a power domain, the spectrum of the MDCT has to be divided by the square root of the envelope to properly bleach the signal.
There are different methods to calculate the envelope:
· Transforming the base signal with a discrete cosine transform (DCT), retaining only the lowest DCT coefficients (setting the highest to zero) and then calculating an inverse DCT • calculating a spectral envelope of a set of Prediction Coefficients Linear (LPC) calculated on the time domain audio box • by filtering the base signal with a low-pass filter Preferably, the last approach is chosen. For applications that require low computational complexity, some simplification can be carried out for the whitening of a spectrum from the MDCT: First, the envelope is calculated using a moving average. This only needs two processor cycles per MDCT Interval. Then, in order to avoid calculating the division and the square root, the spectral envelope is approached by 2 ", in
107
INSTITUTO MSXICANC Ot LA RROPtF.DAD
INDUSTRIAL
<img file="MX340575B_D0113.tif" />
where n is the integer logarithm of the envelope. In this domain, the square root operation is simply converted to an offset operation, and furthermore, division by the envelope can be performed by another offset operation.
After calculating the correlation of each source box with each target box, for all nTar target boxes, the source box with the highest correlation is selected to replace it. To better match the original spectral structure, the correlation delay is used to modulate the spectrum replicated by an integer number of intervals in the transform. In case of odd delays, the box is further modulated through multiplication by an alternative time sequence of -1/1 to compensate for the inverse frequency representation of any other band within the Modified Discrete Cosine Transform (MDCT).
Figure 12c shows an example of a correlation between a source box and a target box. In this example, the correlation delay is 5, so the origin box has to be modulated by 5 Intervals towards the highest frequency intervals in the bandwidth extension algorithm (BWE) copying stage. Additionally, the sign in the box has to be inverted since the maximum correlation value is negative and additional modulation as described above represents the odd delay.
Therefore, the total amount of lateral information to transmit from the encoder to the decoder could consist of the following data:
108
IMPI
MEXICAN INSTITUTE DF LA PRtFIEDAD INDI ISTR1AL
<img file="MX340575B_D0114.tif" />
t¡leNum [nrar]: index of the selected source box by destination box
<td></td><td> •</td><td>tileSign [nrar]:</td><td>target box sign</td>
<td> 5</td><td>• destination</td><td>tileMod [nrar]:</td><td>correlation delay per box of</td>
Box trimming and stabilization are an important step in Intelligent Space Fill (IGF). Their need and advantages are explained with an example, assuming a stationary tonal audio signal such as a stable pitch pitch note. Logic determines that fewer faults are entered if, for a given destination region, you always select the origin squares from the same origin region through the boxes. Although the signal is assumed to be stationary, this condition would not apply well in each frame since the similarity measure (eg correlation) from another region of similar origin could still dominate the similarity result (eg , cross correlation). This causes tileNum [nTar] between adjacent frames to hesitate between two or three very similar options. This can be the source of an annoying musical noise type failure.
In order to eliminate such faults, the set of source boxes will be trimmed so that the remaining elements of the source set are maximally dissimilar. This is accomplished through a set of source boxes
109
IMPI
INSTITUTO MEXlCANU £> E LA PROFIP.DAP INDUSTRIAL
S = {Yes, S<sub>2</sub>, ... S<sub>n</sub>} ----— as follows. For any box of origin s¡, we correlate it with all other boxes of origin, finding the best correlation between s¡ and Sj and storing it in a matrix S<sub>x</sub>. Here, S<sub>x</sub>[i] [j] contains the absolute maximum value of cross-correlation between s¡ and Sj. Adding the matrix S<sub>x</sub> Along the columns, it provides the sum of the cross correlations of a box of origin s, with all other boxes of origin T.
T [i] = S<sub>x</sub>[¡] [1] + S<sub>x</sub>[i] [2] ... + S<sub>x</sub>[i] [n]
Here, T represents a good measure of similarity between an origin box and other origin boxes. If, for any box of origin i, threshold T>
the source box i can be removed from the set of potential sources, as it is highly correlated with other sources. The box with the lowest correlation of the set of boxes that meets the condition in equation 1 is chosen as a representative box for this subset. In this way, we ensure that the origin boxes are maximally dissimilar from each other.
The box trimming method also involves a memory of the set of cropped boxes used in the previous table. The boxes that were active in the previous table are retained in the following table if there are also alternative candidates for the cut.
Let S3 boxes, s<sub>4</sub>one s<sub>5</sub> are active from the boxes {yes, s<sub>2</sub>.s<sub>5</sub>} in box k, then in box k + 1, even if boxes Yes, s<sub>3</sub> ys<sub>2</sub>
<img file="MX340575B_D0115.tif" />
110
ΙΜΡΪ
MRXICANO INSTITUTE OF INDUSTRIAL PROPERTY compete to be cut where<sub>3</sub> is most closely correlated with the others, s<sub>3</sub> it is retained since it was a useful origin box in the previous table and, therefore, its retention in the set of origin boxes is beneficial to reinforce the temporal continuity in the selection of boxes. This method is preferably applied if the cross correlation between origin i and destination j, represented as T<sub>x</sub>[i] [j] is high
An additional method of stabilizing boxes is to retain the order of the boxes in the previous k-1 box if none of the source boxes in the current k-box correlate well with the target boxes.
This can happen if the cross correlation between origin i and destination j, represented as T<sub>x</sub>[i] [j] is very low for all i, j
For example, yes
T<sub>x</sub>[i] [¡] <0.6 then a provisional threshold is used, then 15 tileNum [nTar]<sub>k</sub> = tileNumlnTar] ^ for all nTar's in this table k.
The above two techniques greatly reduce failures that occur from the rapid change of fixed frame numbers across the frames. Another additional benefit of this frame cropping and stabilization is that no extra information needs to be sent to the decoder and no change in decoder architecture is needed. This frame trimming proposal is an elegant way to reduce power musical noise in the form of faults or excessive noise in the spectral regions of the frames.
<img file="MX340575B_D0116.tif" />
111
IMPI
MEXICAN INSTITUTE >> '»F. LA FROFIEUAD O = * «X¿¿4. ·<sup>r M</sup> '
INDUSTRIAL
Fig. 11a illustrates an audio decoder to decode an encoded audio signal. The first audio decoder comprises an (central) audio decoder 1102 to generate a first decoded representation of a first set of first spectral portions, wherein the decoded representation has a first spectral resolution.
i
Also, the audio decoder comprises a parametric decoder 1104 to generate a second decoded representation of a second set of second spectral portions having a second spectral resolution that is lower than the first spectral resolution. Furthermore, a frequency regenerator 1106 is provided which receives, as a first input 1101, the first decoded spectral portions and as a second input in 1103 the parametric information that includes, for each target frequency box or target reconstruction band, information range of origin. The frequency regenerator of 1106 then applies frequency regeneration using spectral values from the source range identified by the adaptation information in order to generate the spectral data for the target range. Next, the first spectral portions 1101 and the output of the frequency regenerator 1107 are both fed into a spectrum-time converter 1108 to finally generate the decoded audio signal.
Preferably, the audio decoder 1102 is a spectral domain audio decoder, although the audio decoder can also be
112
IMPI
MEXICAN INSTITUTE OF IJk PROHEDAD
INDUSTRIAL implement like any other audio decoder such as, for example, a time domain or parametric audio decoder.
As indicated in Fig. 11b, the frequency regenerator 1106 may comprise the functionalities of block 1120 illustrating a frame modulator - source range selector for odd delays, a bleached filter 1122, when a bleach tag 1123 is provided, and, additionally, a spectral envelope with tuning functionality implemented as illustrated in block 1128 using raw spectral data generated by either block 1120 or 1122 or the cooperation of both blocks. However, the frequency regenerator 1106 may comprise a switch 1124 reactive to a received bleach tag 1123. When the bleach tag is affixed, the output range selector / frame modulator output for odd lags is input to the bleach filter 1122. However, the bleach tag 1123 is not affixed for a certain rebuild band, whereby a bypass line 1126 is then activated such that the output of block 1120 is provided to the spectral envelope adjustment block 1128 without any bleaching.
There may be more than one bleach level (1123) signaled in the bit stream and these levels may be flagged. In case there are three levels indicated by box, the levels will be coded as follows:
bit = readBit (1); yes (bit == 1) {se »
MEXICAN INSTITUTE *? OFÍA FRORFOA O INDUSTRIAL
<img file="MX340575B_D0117.tif" />
113 for (t¡le_index = O..nT) / * the same levels as the last frame * / whitening_level [tile_index] = whitening_level_prev_frame [tile_index];
} or {/ 'first box: * / tile_index = 0; bit = readBit (1); if (bit == 1) {whitening_level [tile_index] = MID_WHITENING;
} or {bit = readBit (1);
if (bit == 1) {whitening_level [tile_index] = STRONG_WHITENING;
} or {whitening_level [tile_index] = OFF; / * no bleaching * /}
} / 'remaining boxes: * / bit = readBit (1);
su (bit == 1) {/ * the flattening levels for the remaining squares are the same as for the first one. * / / * No other Intervals have to be read * /
114
ΙΜΡΙ63
XXICAN INSTITUTE
OF THE PROPERTY ü> *
INDUSTRIAL for (tile_index = 1 ..nT) whiten¡ng_level [t¡le_¡ndex] = wh¡tening_level [O];
} or {/ * b¡ts read for the remaining boxes as for the first box * / for (t¡le_index = 1 ..nT) {bit = readB¡t (1); if (bit == 1) {wh¡ten¡ng_level [tile_¡ndex] = MID_WHITENING;
} or {bit = readBit (1);
if (bit == 1) {whitening_level [tile_index] = STRONG_WHITENING;
} or {whitening_level [tile_index] = OFF; / * no bleaching * /}
} }
} }
MID_WHITENING and STRONG_WHITENING refer to different bleach filters (1122) that may differ in how the envelope is calculated (as described above).
115
INSTITUI MtXICANO DE LA FROFIFDA »
INDUSTRIAL
<img file="MX340575B_D0118.tif" />
The decoder-side frequency regenerator can be controlled by a source range ID 1121 when only a raw spectral frame selection scheme is applied. However, when a precise timing spectral box selection scheme is applied, then a source range delay 1119 is further provided. Likewise, as long as the correlation calculation provides a negative result, then additionally a correlation sign can also be applied to block 1120 so that each of the page data spectral lines is multiplied by -Γ to represent the negative sign.
Therefore, the present Invention as described in Fig. 11a, 11b ensures optimum audio quality is obtained due to the fact that the best matching source range for a given destination or destination range is calculated in the encoder side and applied on the decoder side.
FIG. 11c is an audio encoder determined to encode an audio signal comprising a time-spectrum converter 1130, a downstream spectral analyzer 1132, and additionally a parameter calculator 1134 and a core encoder 1136. The core encoder 1136 outputs coded source ranges, and parameter calculator 1134 outputs adaptation information for target ranges.
The encoded source ranges are transmitted to a decoder along with adaptation information for the destination ranges so that the decoder illustrated in Fig. 11a is in the position to carry out a frequency regeneration.
116
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
Parameter calculator 1134 is configured to calculate similarities between first spectral portions and second spectral portions and to determine, based on the similarities calculated for a second spectral portion, a matching first spectral portion that matches the second spectral portion. Preferably, the adaptation results for different ranges of origin and ranges of destination, as illustrated in Figs. 12a, 12b for determining a selected adaptation pair comprise the second spectral portion, and the parameter calculator is configured to provide this adaptation information that identifies the adaptation pair in an encoded audio signal. Preferably, the parameter calculator 1134 of the present invention is configured to use predefined target rulers in the second set of second spectral portions or predefined origin rulers in the first set of first spectral portions as illustrated, for example, in the Fig. 12b. Preferably, the predefined target regions do not overlap or the predefined source regions overlap. When the predefined origin regions are a subset of the first set of first spectral portions below a space-filling ice frequency 309 of Fig. 3a, and preferably, the predefined destination region encompassing a spectral ruler
Lower coincides, with its Lower frequency limit, with the space fill start frequency so that any target range is above the space fill start frequency and the source ranges are below the Space fill start frequency.
<img file="MX340575B_D0119.tif" />
117
MEXICAN INSTITUTE
OF THE PROPERTY ^ iijT
INDUSTRIAL
As explained previously, a fine granularity is obtained by comparing a destination region with an origin screed without any delay in the origin screed and the same origin screed, but with a certain delay. These delays are applied in the cross-correlation calculator 1140 of Fig. 11d and the selection of adaptation pairs is finally carried out by the box selector 1144.
In addition, it is preferred to perform source range and / or target range bleaching as illustrated in block 1142. This block 1142 then provides a bleach tag to the bit stream that is used to control the decoder side switch 1123 of Fig. 11b. Likewise, if the cross-correlation calculator 1140 provides a negative result, then this negative result is also signaled to a decoder. Therefore, in a preferred embodiment, the box selector issues a source range ID for a destination range, a delay, a sign, and block 1142 further provides a bleach tag.
Also, the parameter calculator 1134 is configured to perform a source frame clipping 1146 by reducing the number of potential source ranges whereby a Source Interconnect is removed from a set of potential source boxes based on a threshold of similarity. Therefore, when two source boxes are more similar or equal to a similarity threshold, then one of these two source boxes is removed from the set of potential sources, and the removed source box is no longer used for further processing. and, specifically, it cannot be selected by the
118
IMPI INDUSTRIAL PROPERTY iNSTmrro mbxicanq box picker 1144 or not used for calculation of cross-correlation between different source ranges and target ranges as performed in block 1140.
Different implementations have been described with respect to different 5 figures. Figs. 1a-5c refer to a full bandwidth or full rate encoder / decoder scheme. Figs. 6a-7e refer to an encoder / decoder scheme with Temporal Noise Modeling Processing (TNS) or Temporal Box Model (TTS). Figs. 8a-8e refer to an encoder / decoder scheme with two-channel specific processing. Figs. 9a-10d refer to a specific calculation of Power Information and the application, and Figs. 11a-12c refer to a specific mode of box selection.
All these different aspects can be of inventive use and are independent of each other but, additionally, they can also be applied together as basically illustrated in Fig. 2a and 2b. However, the specific two-channel processing can also be applied to an encoder / decoder scheme illustrated in Fig. 13a and 13b, and the same applies to the TNS / TTS processing, the calculation of envelope power information and the application in the reconstruction band or the identification of the adaptation origin range and the corresponding application on the side of the decoder. On the other hand, the full rate aspect can be applied with or without TNS / TTS processing, with or without two-channel processing, with or without adaptation source range identification, or with other types of
<img file="MX340575B_D0120.tif" />
119
IMPI
INSTITUTO MEXICANO Dt LA FROFÍIDAD INDUSTRIAL power calculations for the representation of the spectral envelope. Therefore, it is evident that the characteristics of one of these individual aspects can be applied in other aspects as well.
Although some aspects have been described in the context of an apparatus 5 for encoding or decoding, it is evident that these aspects also represent a description of the corresponding method, where a block or device corresponds to a step of the method or to a characteristic of a step of the method. method. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or element or feature of a respective apparatus. Some or all of the steps in the method may be carried out by (or with) a hardware apparatus such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or more of most of the important steps of the method can be carried out by said apparatus.
Depending on certain implementation requirements, the embodiments of the invention can be implemented in hardware or in software. The implementation can be carried out using a non-transient storage medium such as a digital storage medium, for example a floppy disk, a Hard Drive (HDD), a DVD, a Blu-Ray, a CD, a ROM, a PROM memory, and an EPROM memory, an EEPROM memory or a FLASH memory, which have electronic read control signals stored in them, cooperating (or capable of cooperating) with a programmable computer system such that the respective method is carried out.
<img file="MX340575B_D0121.tif" />
120
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0122.tif" />
Therefore, the digital storage medium can be computer readable.
Some embodiments in accordance with the invention comprise a data carrier having electronic read control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
In general, the embodiments of the present invention can be implemented as a computer program product with a program code, the program code of which is operative to carry out one of the methods when the computer program product is run in a computer. The program code can be stored, for example, on a computer-readable carrier.
Other embodiments comprise computer programs for carrying out one of the methods described herein, stored on a computer-readable carrier.
In other words, one embodiment of the method of the Invention is, therefore, a computer program that has a program code to carry out one of the methods described herein, when the Computer program is run on a computer. .
Therefore, another embodiment of the method of the invention is a data carrier (or a digital storage medium, or a computer readable medium) comprising, recorded thereon, the computer program for carrying out one of the methods described herein. The data carrier,
121
IMPI
MEXICAN INSTITUTE Dt LA PtOP (K) AD INtUSTlIAl
<img file="MX340575B_D0123.tif" />
the digital storage medium or the recorded medium are generally tangible and / or non-transient.
Therefore, a further embodiment of the invention is a data stream or signal sequence representing the computer program for carrying out one of the methods described herein. The data stream or signal sequence, for example, may be configured to be transferred over a data communication connection, for example, over the Internet.
A further embodiment comprises a processing means, for example, a computer or a programmable logic device configured or adapted to carry out one of the methods described in the present invention.
Another embodiment comprises a computer having the computer program installed therein to carry out one of the methods described herein.
Another embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program to carry out one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may comprise, for example, a file server for transferring the computer program to the receiver.
122
ΙΜΡΙ
MEXICAN INSTITUTE OF IA PROriSDAD
INDUSTRIAL
<img file="MX340575B_D0124.tif" />
In some embodiments, a programmable logic device (eg, a field programmable gate array) can be used to perform some or all of the functionality of the methods described in the present invention. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods will preferably be carried out by any hardware apparatus.
The above described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to other experts in the field. It is the intention, therefore, that the invention be limited only by the scope of the Imminent patent claims and not by the specific details presented by way of description and explanation of the embodiments of the present.
123
IMPI
MEXICAN INSTITUTE IM THE INDUSTRIAL PROPERTY
<img file="MX340575B_D0125.tif" />
List of References [1] Dietz, L. Liljeryd, K. Kjórling and O. Kunz, “Spectral Band Replication, a novel approach in audio coding,” in 112th AES Convention, Munich, May
2002.
[2] Ferreira, D. Sinha, “Accurate Spectral Replacement”, Audio
Engineering Society Convention, Barcelona, Spain 2005.
[3] D. Sinha, A. Ferreiral and E. Harinarayanan, “A Novel Integrated Audio Bandwidth Extension Toolkit (ABET)”, Audio Engineering Society
Convention, Paris, France 2006.
[4] R. Annadana, E. Harinarayanan, A. Ferreira and D. Sinha, “New
Results in Low Bit Rate Speech Coding and Bandwidth Extension ”, Audio Engineering Society Convention, San Francisco, USA. 2006.
[5] T. Zernicki, M. Bartkowiak, “Audio bandwidth extension by frequency scaling of sinusoidal partías”, Audio Engineering Society Convention, San
Francisco, USA 2008.
[6] J. Herre, D. Schulz, Extending the MPEG-4 AAC Codee by Perceptual Noise Substitution, 104th AES Convention, Amsterdam, 1998, Preprint
4720.
[7] M. Neuendorf, M. Multrus, N. Rettelbach, et al., MPEG Unified
Speech and Audio Coding-The ISO / MPEG Standard for High-Efficiency Audio
Coding of all Contení Types, 132nd AES Convention, Budapest, Hungary, April
2012.
124
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX340575B_D0126.tif" />
[8] McAulay, Robert J., Quatieri, Thomas F. "ftpaarh Anaiysis / Synthesis Based on a Sinusoidal Representation". IEEE Transactions on Acoustics, Speech, And Signal Processing, Vol 34 (4), August 1986.
[9] Smith, JO, Serra, X. “PARSHL: An analysis / synthesis program for 5 non-harmonic sounds based on a sinusoidal representation”, Proceedings of the
International Computer Music Conference, 1987.
[10] Purnhagen, H .; Meine, Nikolaus, HILN-the MPEG-4 parametric audio coding tools, Circuits and Systems, 2000. Proceedings. ISCAS 2000 Geneva. The 2000 IEEE International Symposium on, vol.3, no., Pp. 201, 204 vol.3,
2000 [11] International Standard ISO / IEC 13818-3, Generic Coding of Moving
Pictures and Associated Audio: Audio, Geneva, 1998.
[12] M. Bosi, K. Brandenburg, S. Quackenbush, L. Fielder, K. Akagiri, H. Fuchs, M. Dietz, J. Herre, G. Davidson, Oikawa: MPEG-2 Advanced Audio
Coding, 101st AES Convention, Los Angeles 1996 [13] J. Herre, “Temporal Noise Shaping, Quantization and Coding methods in Perceptual Audio Coding: A Tutorial introduction, 17th AES International Conference on High Quality Audio Coding, August 1999 [14] J. Herre, "Temporal Noise Shaping, Quantization and Coding 20 methods in Perceptual Audio Coding: A Tutorial Introduction", 17th AES
International Conference on High Quality Audio Coding, August 1999 [15] International Standard ISO / IEC 23001-3: 2010, Unified speech and audio coding Audio, Geneva, 2010.
Sfjvjrjff ·. YES:>:!
125
IMPI
MEXICAN INSTITUTE OF LA PROPIBDA · ΙΝΓΗ ISTRIAL
<img file="MX340575B_D0127.tif" />
[16] International Standard ISO / IEC 14496-3: 2005, Information technology - Coding of audio-visual objects - Part 3: Audio, Geneva, 2005.
[17] P. Ekstrand, “Bandwidth Extension of Audio Signáis by Spectral Band Replication, in Proceedings of 1st IEEE Benelux Workshop on MPCA, Leuven,
November 2002 [18] F. Nagel, S. Disch, S. Wilde, A continuous modulated single sideband bandwidth extension, ICASSP International Conference on Acoustics, Speech and Signal Processing, Dallas, Texas (USA), April 2010
126
Contents154
156 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111 Sheet 112 Sheet 113 Sheet 114 Sheet 115 Sheet 116 Sheet 117 Sheet 118 Sheet 119 Sheet 120 Sheet 121 Sheet 122 Sheet 123 Sheet 124 Sheet 125 Sheet 126 Sheet 127 Sheet 128 Sheet 129 Sheet 130 Sheet 131 Sheet 132 Sheet 133 Sheet 134 Sheet 135 Sheet 136 Sheet 137 Sheet 138 Sheet 139 Sheet 140 Sheet 141 Sheet 142 Sheet 143 Sheet 144 Sheet 145 Sheet 146 Sheet 147 Sheet 148 Sheet 149 Sheet 150 Sheet 151 Sheet 152 Sheet 153 Sheet 154 Sheet 155 Sheet 156
298 members in 22 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 131773467 | European Patent Office (EPO) | – | |
| 131773483 | European Patent Office (EPO) | – | |
| 131773509 | European Patent Office (EPO) | – | |
| 131773533 | European Patent Office (EPO) | – | |
| 13177346 | European Patent Office (EPO) | A | |
| 13177348 | European Patent Office (EPO) | A | |
| 13177350 | European Patent Office (EPO) | A | |
| 13177353 | European Patent Office (EPO) | A | |
| 131893588 | European Patent Office (EPO) | – | |
| 13189358 | European Patent Office (EPO) | A | |
| 2014065123 | European Patent Office (EPO) | W |
Members298
| Document | Office | Kind | |
|---|---|---|---|
| EP2830054A1 | European Patent Office (EPO) | A1 | |
| EP2830056A1 | European Patent Office (EPO) | A1 | |
| EP2830059A1 | European Patent Office (EPO) | A1 | |
| EP2830061A1 | European Patent Office (EPO) | A1 | |
| EP2830063A1 | European Patent Office (EPO) | A1 | |
| EP2830064A1 | European Patent Office (EPO) | A1 | |
| EP2830065A1 | European Patent Office (EPO) | A1 | |
| CA2886505A1 | Canada | A1 | |
| CA2918524A1 | Canada | A1 | |
| CA2918701A1 | Canada | A1 | |
| CA2918804A1 | Canada | A1 | |
| CA2918807A1 | Canada | A1 | |
| CA2918810A1 | Canada | A1 | |
| CA2918835A1 | Canada | A1 | |
| CA2973841A1 | Canada | A1 | |
| WO2015010947A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010948A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010949A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010950A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010952A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010953A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010954A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201513098A | Taiwan Province of China | A | |
| AU2014295302A1 | Australia | A1 | |
| TW201514974A | Taiwan Province of China | A | |
| TW201517019A | Taiwan Province of China | A | |
| TW201517023A | Taiwan Province of China | A | |
| TW201517024A | Taiwan Province of China | A | |
| SG11201502691QA | Singapore | A | |
| KR20150060752A | Republic of Korea | A | |
| TW201523589A | Taiwan Province of China | A | |
| TW201523590A | Taiwan Province of China | A | |
| EP2883227A1 | European Patent Office (EPO) | A1 | |
| MX2015004022A | Mexico | A | |
| CN104769671A | China | A | |
| US2015287417A1 | United States of America | A1 | |
| JP2015535620A | Japan | A | |
| AR096985A1 | Argentina | A1 | |
| AR096988A1 | Argentina | A1 | |
| AR096989A1 | Argentina | A1 | |
| AR096990A1 | Argentina | A1 | |
| AR096991A1 | Argentina | A1 | |
| AR096992A1 | Argentina | A1 | |
| AR096993A1 | Argentina | A1 | |
| SG11201600401RA | Singapore | A | |
| SG11201600422SA | Singapore | A | |
| SG11201600464WA | Singapore | A | |
| SG11201600494UA | Singapore | A | |
| SG11201600496XA | Singapore | A | |
| SG11201600506VA | Singapore | A | |
| KR20160024924A | Republic of Korea | A | |
| AU2014295295A1 | Australia | A1 | |
| AU2014295296A1 | Australia | A1 | |
| AU2014295297A1 | Australia | A1 | |
| AU2014295298A1 | Australia | A1 | |
| AU2014295300A1 | Australia | A1 | |
| AU2014295301A1 | Australia | A1 | |
| KR20160030193A | Republic of Korea | A | |
| CN105453175A | China | A | |
| CN105453176A | China | A | |
| KR20160034975A | Republic of Korea | A | |
| KR20160041940A | Republic of Korea | A | |
| CN105518776A | China | A | |
| CN105518777A | China | A | |
| KR20160042890A | Republic of Korea | A | |
| MX2016000940A | Mexico | A | |
| KR20160046804A | Republic of Korea | A | |
| CN105556603A | China | A | |
| MX2016000857A | Mexico | A | |
| MX2016000924A | Mexico | A | |
| CN105580075A | China | A | |
| EP3017448A1 | European Patent Office (EPO) | A1 | |
| US2016133265A1 | United States of America | A1 | |
| US2016140973A1 | United States of America | A1 | |
| US2016140979A1 | United States of America | A1 | |
| US2016140980A1 | United States of America | A1 | |
| US2016140981A1 | United States of America | A1 | |
| HK1211378A | Hong Kong, China | A | |
| HK1211378A1 | Hong Kong, China | A1 | |
| EP3025328A1 | European Patent Office (EPO) | A1 | |
| EP3025337A1 | European Patent Office (EPO) | A1 | |
| EP3025340A1 | European Patent Office (EPO) | A1 | |
| EP3025343A1 | European Patent Office (EPO) | A1 | |
| EP3025344A1 | European Patent Office (EPO) | A1 | |
| MX2016000854A | Mexico | A | |
| AU2014295302B2 | Australia | B2 | |
| MX2016000935A | Mexico | A | |
| MX2016000943A | Mexico | A | |
| TWI541797B | Taiwan Province of China | B | |
| MX340575BThis record | Mexico | B | |
| US2016210974A1 | United States of America | A1 | |
| TWI545558B | Taiwan Province of China | B | |
| TWI545560B | Taiwan Province of China | B | |
| TWI545561B | Taiwan Province of China | B | |
| EP2883227B1 | European Patent Office (EPO) | B1 | |
| JP2016525713A | Japan | A | |
| JP2016527556A | Japan | A | |
| JP2016527557A | Japan | A | |
| TWI549121B | Taiwan Province of China | B | |
| JP2016529545A | Japan | A |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 340575
- Application
- 4022
Titles2
- Spanish
- APARATO Y METODO PARA CODIFICAR Y DECODIFICAR UNA SEÑAL DE AUDIO CODIFICADA UTILIZANDO MODELADO DE RUIDO TEMPORAL/DE PARCHE
- English
- APPARATUS AND METHOD FOR ENCODING AND DECODING AN ENCODED AUDIO SIGNAL USING TEMPORAL NOISE/PATCH SHAPING.
Classification
- CPC, 19
- G10L21/0388
- G10L19/02
- G10L19/03
- G10L19/008
- G10L19/025
- G10L19/028
- G10L21/038
- G10L19/0204
- G10L19/0212
- G10L19/022
- G10L19/032
- G10L19/06
- G10L19/18
- H03M7/30
- G10L19/0208
- H04S1/007
- G10L25/18
- G10L25/21
- G10L25/06
- IPC, 4
- G10L19 03
- G10L19 02
- G10L19 028
- G10L21 0388