Audio encoder and decoder for encoding and decoding audio samples.
Abstract
An audio encoder (100) for encoding audio samples, comprising a first time domain aliasing introducing encoder (110) for encoding audio samples in a first encoding domain, the first time domain aliasing introducing encoder (110) having a first framing rule, a start window and a stop window. The audio encoder (100) further comprises a second encoder (120) for encoding samples in a second encoding domain, the second encoder (120) having a different second framing rule. The audio encoder (100) further comprises a controller (130) switching from the first encoder (110) to the second encoder (120) in response to characteristic of the audio samples, and for modifying the second framing rule in response to switching from the first encoder (110) to the second encoder (120) or for modifying the start window or the stop window of the first encoder (110), wherein the second framing rule remains unmodified.

Term
2.8 yearsleft in the term
Expires 26 June 2029.
- Priority
- Filed
- Granted
- Today
- Expires
33 claims: 23 independent, 10 dependent
- 1Reivindicaciones 1. Un codificador de audio (100) para codificar muestras de audio, que comprende:5 un primer codificador introductor de aliasing en dominio de tiempo (110) para codificar muestras de audio en un primer dominio de codificación, el primer codificador introductor de aliasing en dominio de tiempo (110) que tiene una primer regla de tramado, una ventana de ¡nielo y una ventana de detención;10 un segundo codificador (120) para codificar muestras en un segundo dominio de codificación, el segundo codificador (120) que tiene un número de muestras de audio de tamaño de trama predeterminado, y un número de muestras de audio de periodo de puesta a punto de codificación, y el segundo codificador (120) que tiene una segunda regla de tramado 15 diferente, una trama del segundo codificador (120) que es una representación codificada de un número de muestras de audio posteriores en el tiempo, el número es igual al número de muestras de audio de tamaño de trama predeterminado;y un controlador (130) para cambiar del primer codificador (110) al 20 segundo codificador (120) en respuesta a una característica de las muestras de audio, y para modificar la segunda regla de tramado en respuesta al cambio del primer codificador (110) al segundo codificador (120) o para modificar la ventana de inicio o la ventana de detención del primer codificador (110), en donde la segunda regla de tramado permanece sin modificaciones.
- 2El codificador de audio (100) de la reivindicación 1, en donde 5 el primer codificador introductor de aliasing en dominio de tiempo (110) comprende un transformador de dominio de frecuencia para transformar una primera trama de muestras de audio posteriores al dominio de frecuencia. 10
- 3El codificador de audio (100) de la reivindicación 2, en donde se adapta el primer codificador introductor de aliasing en dominio de tiempo (110) para la ponderar la última trama con la ventana de inicio cuando una trama posterior es codificada por el segundo codificador (120) y/o para ponderar la primera trama con la ventana de detención cuando una trama 15 precedente va a ser codificada por el segundo codificador (120).
- 4El codificador de audio (100) de una de las reivindicaciones 2 o 3, en donde se adapta el transformador de dominio de frecuencia para transformar la primer trama al dominio de frecuencia en base a una 20 transformada del coseno discreta modificada (MDCT) y en donde se adapta el primer codificador introductor de aliasing en dominio de tiempo (110) para adaptar un tamaño de MDCT a las ventanas de inicio y/o detención y/o inicio y/o detención modificadas.
- 5El codificador de audio (100) de una de las reivindicaciones 1 a 4, en donde se adapta el primer codificador introductor de aliasing en dominio de tiempo (110) para utilizar una ventana de inicio y/o detención 5 que tiene una parte de aliasing y/o parte libre de aliasing.
- 6El codificador de audio (100) de una de las reivindicaciones de 1 a 5, en donde se adapta el primer codificador introductor de aliasing en dominio de tiempo (110) para utilizar una ventana de inicio y/o ventana de 10 detención que tiene una parte libre de aliasing como una parte de flanco de •subida de la ventana cuando la trama precedente es codificada por el segundo codificador (120) y en una parte de flanco de bajada cuando la trama posterior es codificada por el segundo codificador (120). 15
- 7El codificador de audio (100) de una de las reivindicaciones 5 o 6, en donde se adapta el controlador (130) para iniciar el segundo codificador (120), de manera que la primer trama de una secuencia de tramas del segundo codificador (120) comprende una representación codificada de una muestra procesada en la parte libre de aliasing 20 precedente del primer codificador (110).
- 8El codificador de audio (100) de una de las reivindicaciones 5 o 6, en donde se adapta el controlador (130) para iniciar el segundo codificador (120), de manera que el número de muestras de audio de periodo de puesta a punto de codificación se superpone con la parte libre de aliasing de la ventana de inicio del primer codificador introductor de aliasing en dominio en tiempo (110) y la trama posterior del segundo 5 codificador (120) se superpone con la parte de aliasing de la ventana de detención.
- 9El codificador de audio (100) de una de las reivindicaciones 5 a 7, en donde se adapta el controlador (130) para iniciar el segundo
- 1010 codificador (120), de manera que el periodo de puesta a punto de codificación se superpone con la parte de aliasing de la ventana de inicio. 10. El codificador de audio (100) de una de las reivindicaciones de 1 a 9, en donde además se adapta el controlador (130) para cambiar del 15 segundo codificador (120) al primer codificador (110) en respuesta a una característica diferente de las muestras de audio y para modificar la segunda regla de tramado en respuesta al cambio del segundo codificador (120) al primer codificador (110) o 20 para modificar la ventana de inicio o la ventana de detención del primer codificador (110), en donde la segunda regla de tramado permanece sin modificaciones.
- 11El codificador de audio de la reivindicación 10, en donde se adapta el controlador (130) para iniciar el primer codificador de aliasing en dominio de tiempo (110), de manera que la parte de aliasing de la ventana de detención se superpone con la trama del segundo codificador (120).
- 12El codificador de audio (100) de la reivindicación 11, en donde se adapta el controlador (130) para iniciar el primer codificador introductor de aliasing en dominio de tiempo (110), de manera que la parte libre de aliasing de la ventana de detención se superpone con una trama del 10 segundo codificador (120).
- 13El codificador de audio (100) de una de las reivindicaciones 1 a 12, en donde el primer codificador de aliasing en dominio de tiempo (110) comprende un codificador AAC de acuerdo con la Codificación Genérica de 15 Imágenes de Movimiento y Audio Asociado:Codificación de Audio Avanzado, Estándar Internacional 13818-7, ISO/IEC JTC1/SC29/WG11 MPEG, 1997.
- 14El codificador de audio (100) de una de las reivindicaciones 1 20 a 13, en donde el segundo codificador comprende un codificador AMR o AMR-WB+ de acuerdo con el Proyecto de Asociación de Tercera Generación (3GPP), especificación técnica (TS), 26.290, versión 6.3.0 a partir de junio de 2005.
- 15El codificador de audio de la reivindicación 14, en donde se adapta el controlador para modificar la regla de tramado AMR, de manera que la primera supertrama AMR comprende cinco tramas AMR. 5
- 16Un método para codificar las tramas de audio, que comprende los pasos de:codificar las muestras de audio en un primer dominio de codificación usando una primer regla de tramado, una ventana de inicio y una ventana de detención;10 codificar muestras de audio en un segundo dominio de codificación usando un número de muestras de audio de tamaño de trama predeterminado y un número de muestras de audio de periodo de puesta a punto de codificación y usar una segunda regla de tramado diferente, la trama del segundo dominio de codificación que es una representación 15 codificada de un número de muestras de audio posteriores en el tiempo, el número que es igual al número de muestras de audio de tamaño de trama predeterminado;cambiar del primer dominio de codificación al segundo dominio de codificación;y 20 modificar la segunda regla de tramado en respuesta al cambio del primer al segundo dominio de codificación modificar la ventana de inicio o la ventana de detención del primer dominio de codificación, en donde la segunda regla de tramado permanece sin modificaciones. 5
- 17Un programa informático que tiene un código de programa para llevar a cabo el método de la reivindicación 16, cuando el código de programa funciona en una computadora o procesador.
- 18Un decodificador de audio (150) para decodificar tramas 10 codificadas de muestras de audio, que comprende:un primer decodificador introductor de aliasing en dominio de tiempo (160) para decódificar muestras de audio en un primer dominio de decodificación, el decodificador introductor de aliasing en dominio de tiempo (160) que tiene una primera regla de tramado, una ventana de inicio y una 15 ventana de detención;un segundo decodificador (170) para decodificar muestras de audio en un segundo dominio de decodificación y el segundo decodificador (170) que tiene un número de muestras de audio de tamaño de trama predeterminado y un número de muestras dé audio de periodo de puesta a 20 punto de codificación, el segundo decodificador (170) que tiene un segunda regla de tramado diferente, una trama del segundo codificador (170) que es representación codificada de un número de muestras de audio posteriores en el tiempo, el número que es igual al número de muestras de audio de tamaño de trama predeterminado;y uh controlador (180) para cambiar del primer decodificador (160) al segundo decodificador (170) en base a una indicación en la trama 5 codificada de las muestras de audio, en donde se adapta el controlador (180) para, modificar la segunda regla de tramado en respuesta al cambio del primer decodificador (160) al segundo decodificador (170) o para modificar la ventana de inicio o la ventana de detención del primer decodificador (160), en donde la segunda regla de tramado permanece sin 10 modificaciones.
- 19El decodificador de audio (150) de la reivindicación 19, en donde el primer decodificador (160) comprende un transformador de dominio de tiempo para transformar una primera trama de muestras de 15 audio decodificadas al dominio de tiempo.
- 20El decodificador de audio (150) de una de las reivindicaciones 18 o 19, en donde se adapta el primer decodificador (160) para ponderar la última trama decodificada con la ventana de inicio cuando la trama posterior 2 0 es decodificada por el segundo decodificador (170) y/o para ponderar la primera trama decodificada con la ventana de detención cuando la trama precedente va a ser decodificada por el segundo decodificador (170).
- 21El decodlflcador de audio (150) de una de las reivindicaciones 19 o 20, en donde se adapta el transformador de dominio de tiempo para transformar la primer trama al dominio de tiempo en base a la MDCT Inversa (IMDCT) y en donde se adapta el primer decodificador introductor 5 de aliasing en dominio de tiempo (160) para adaptar un tamaño de IMDCT a las ventanas de ¡nielo y/ o detención o Inicio y/o detención modificadas.
- 22El decodlflcador de audio (150) de una de las reivindicaciones 18 a 21, en donde se adapta el primer decodificador introductor de aliasing 10 en dominio de tiempo (160) para utilizar una ventana de inicio y/o ventana de detención que tiene una parte de aliasing y un parte libre de aliasing.
- 23El decodificador de audio (150) de una de las reivindicaciones 18 a 22, en donde se adapta el primer decodlflcador Introductor de allaslng 15 en dominio de tiempo (110) para utilizar una ventana de ¡nielo y/o una ventana de detención que tiene una parte libre de allaslng en una parte de flanco de subida de la ventana cuando la trama precedente es decodificada por el segundo decodlflcador (170) y una parte de flanco de bajada cuando la trama posterior es codificada por el segundo decodificador (170).
- 24El decodlflcador de audio (150) de acuerdo con una de las reivindicaciones 22 o 23, en donde se adapta el controlador (180) para Iniciar el segundo decodiflcador (170), de manera que la primera trama de la secuencia de tramas del segundo decodiflcador (170) comprende una representación codificada de una muestra procesada en la parte libre de aliasing precedente del primer codificador (160). 5
- 25El decodificador de audio (150) de una de las reivindicaciones 22 a 24, en donde se adapta el controlador (180) para Iniciar el segundo decodiflcador (170), de manera que el número de muestras de audio del periodo de puesta punto de codificación se superpone con la parte libre de aliasing de la ventana de inicio del primer decodificador Introductor de 10 aliasing en dominio de tiempo (160) y la trama posterior del segundo decodificador (170) se superpone con la parte de aliasing de la ventana de detención.
- 26El decodificador de audio (150) de una de las reivindicaciones 15 22 a 24, en donde se adapta el controlador (180) para Iniciar el segundo decodiflcador (170), de manera que el periodo de puesta a punto de codificación se superpone con la parte de aliasing de la ventana de detención. 20
- 27El decodiflcador (150) de una de las reivindicaciones 18 a 26, en donde además se adapta el controlador (180) para cambiar del segundo decodificador (170) al primer decodificador (160) en respuesta a una indicación de las muestras de audio y para modificar la segunda regla de tramado en respuesta al cambio del segundo decodificador (170) al primer decodificador (160) o para modificar la ventana de inicio o la ventana de detención del 5 primer decodificador (160), en donde la segunda regla de tramado permanece sin modificaciones.
- 28El decodificador de audio (150) de la reivindicación 27, en donde se adapta el controlador (180) para iniciar el primer decodificador 10 introductor de aliasing en dominio de tiempo (160), de manera que la parte de aliasing de la ventana de detención se superpone con una trama del segundo decodificador (170).
- 29El decodificador de audio (150) de una de las reivindicaciones 15 18 a 28, en donde se adapta el controlador (180) para aplicar un desvanecimiento cruzado entre tramas consecutivas de muestras de audio decodificadas de diferentes decodificadores.
- 30El decodificador de audio (150) de una de las reivindicaciones 20 18 a 29, en donde se adapta el controlador (180) para determinar un aliasing en un parte de aliasing de la ventana de inicio o detención de una trama decodificada del segundo decodlficador (170) y para reducir el aliasing en la parte de aliasing en base al aliasing determinado.
- 31El decodlficador de audio (150) de una de las reivindicaciones 18 a 30, en donde se adapta el controlador (180) para descartar el periodo de puesta a punto de codificación de las muestras de audio del segundo 5 decodificador (170).
- 32Un método para decodificar tramas codificadas de muestras de audio, que comprende los pasos de decodificar muestras de audio en un primer dominio de 10 decodificación, el primer dominio de decodificación que introduce aliasing en tiempo y que tiene una primera regla de tramado, una ventana de inicio y una ventana de detención;decodificar muestras de audio en un segundo dominio de decodificación, el segundo dominio de decodificación que tiene un número 15 de muestras de audio de tamaño de trama predeterminado y un número de muestras de audio de periodo de puesta a punto de codificación, el segundo . dominio de decodificación que tiene una regla de tramado diferente, una trama del segundo dominio de decodificación que es una representación decodificada de un número de muestras de audio posteriores en el tiempo, 20 el número que es igual al número de muestras de audio de tamaño de trama predeterminado;y cambiar del primer dominio de decodificación al segundo dominio de decodificación en base a una indicación de la trama codificada de muestras de audio;modificar la segunda regla de tramado en respuesta al cambio del 5 primer dominio de codificación al segundo dominio de codificación o modificar la ventana de inicio y/o la ventana de detención del primer dominio de decodificación, en donde la segunda regla de tramado permanece sin modificaciones.
- 33Un programa informático que tiene un código de programa para llevar a cabo el método de la reivindicación 32, cuando el código de programa funciona en una computadora o procesador.
Independent claims33
227 paragraphs in 2 sections, as filed
(54) Title: AUDIO ENCODER AND DECODER TO ENCODE AND DECODE AUDIO SAMPLES. (54) Title: AUDIO ENCODER AND DECODER FOR ENCODING AND DECODING AUDIO SAMPLES.
(57) Summary
An audio encoder (100) for encoding audio samples, comprising a first time domain aliasing introducer encoder (110) for decoding audio samples in a first encoding domain, the first time domain aliasing introducer encoder (110) having a first screening rule, a start window, and a stop window. The audio encoder 100 further comprises a second encoder 120 for encoding samples in a second encoding domain, the second encoder 120 having a number of audio samples of predetermined frame size, and a number of samples. encoding timing period audio, the second encoder (120) having a different second screening rule, a frame of the second encoder (120) which is an encoded representation of a number of audio samples later in time, the number being equal to the number of audio samples of predetermined frame size. The audio encoder (100) further comprises a controller (130) that switches from the first encoder (110) to the second encoder (120) in response to a characteristic of the audio samples, and to modify the second screening rule in response to the changing from the first encoder (110) to the second encoder (120) or to modify the start window or the stop window of the first encoder (110), where the second screening rule remains unchanged.
(57) Abstract
An audio encoder (100) for encoding audio samples, comprising a first time domain aliasing introducing encoder (110) for encoding audio samples in a first encoding domain, the first time domain aliasing introducing encoder (110) having a first framing rule, a start window and a stop window. The audio encoder (100) further comprises a second encoder (120) for encoding samples in a second encoding domain, the second encoder (120) having a different second framing rule. The audio encoder (100) further comprises a controller (130) switching from the first encoder (110) to the second encoder (120) in response to characteristic of the audio samples, and for modifying the second framing rule in response to switching from the first encoder (110) to the second encoder (120) or for modifying the start window or the stop window of the first encoder (110), where the second framing rule remains unmodified.
AUDIO ENCODER AND DECODER TO CODE AND
DECODE AUDIO SAMPLES
Descriptive memory
The present invention falls within the field of audio encoding in different encoding domains, for example in the time domain and transformation domain.
In the context of low bit rate audio and speech coding technology, different coding techniques have traditionally been used to achieve low bit rate coding of such signals with the best possible subjective quality at a bit rate Dadaist. General music / sound signal encoders seek to optimize subjective quality by giving a spectral (and temporal) shape of quantization error according to a masking threshold curve that is estimated from the input signal by the model perceptual (“perceptual audio encoding”). On the other hand, low-bit-rate speech coding has been shown to work efficiently when based on a human speech production model, that is, using Linear Prediction Coding (LPC). to model resonant effects of the human vocal tract along with efficient encoding of the residual excitation signal.
As a consequence of these two different approaches, general audio encoders, such as MPEG-1 Layer 3 (MPEG = Moving Pictures Expert Group), or MPEG-2/4 Advanced Audio Coding (AAC), generally , they do not work as well for speech signals at very low very low data rates as dedicated LPC-based speech encoders due to the lack of exploitation of a speech source model. In contrast, LPC-based speech encoders generally do not achieve convincing results when applied to general music signals due to their inability to flexibly shape the spectral envelope of encoding distortion according to a masking threshold curve. . Below, we describe concepts that combine the advantages of LPC-based encoding and perceptual audio encoding in a single framework, and therefore describe a unified audio encoding that is efficient for both general audio and audio signals. speaks.
Traditionally, perceptual audio encoders use a bank filter-based approach to efficiently encode audio signals and shape quantization distortion based on an estimate of the masking curve.
Fig. 16a shows a basic block diagram of a monophonic perceptual coding system. An analysis filter bank 1600 is used to delineate the time domain samples into subsampled spectral components. Depending on the number of spectral components, also
0 The system is referred to as a subband encoder (small number of subbands, for example, 32) or a transformer encoder (large number of frequency lines, for example, 512). A perceptual ("psychoacoustic") model 1602 is used to estimate the real-time dependent masking threshold. The spectral components ("subband" or "frequency domain") are quantized and encoded 1604 so that the quantization noise is hidden under the actually transmitted signal and is not noticeable after decoding. This is accomplished by varying the quantification granularity of the spectral values over time and frequency.
The entropy encoded and quantized spectral coefficients or subband values are, in addition to supplemental information, input into a bit stream formatter 1606, which provides an encoded audio signal that is suitable for transmission or storage. The output bit stream of block 1606 can be transmitted over the internet or can be stored in any machine-readable data driver.
On the decoder side, a decoder input interface 1610 receives the encoded bit string. Block 1610 separates the quantized and entropy encoded spectral / subband values from the supplemental information. The encoded spectral values are entered into an entropic decoder as a Huffman decoder, which is positioned between 1610 and 1620. The outputs of this entropy decoder are quantized spectral values. The quantized spectral values are entered into a quantizer, which performs an "inverse" quantization as indicated at 1620 in Fig. 16a. The departure of the block 1620 is to a synthesis bank filter 1622, which performs a synthesis filter including a frequency / time transformation and typically an “aliasing” operation (foreign signal generation - effect produced by the distortion generated in the digitization of an audio signal when the sampling frequency is insufficient) of time domain as an overlay or aggregate and / or a complementary synthesis window operation to finally obtain the output audio signal.
Traditionally, efficient speech coding has been based on Linear Prediction Coding (LPC) to model the resonant effects of the human vocal tract along with efficient coding of the residual excitation signal. Both the LPC and excitation parameters are transmitted from the encoder to the decoder. This principle is illustrated in Figs. 17a and 17b.
Fig. 17a indicates the coding side of a coding / decoding system based on linear prediction coding. The speech input is the input to an LPC 1701 parser, which provides, at its output, LPC filter coefficients. Based on these LPC filter coefficients, an LPC filter 1703 is adjusted. The LPC filter outputs a spectrally bleached audio signal, which is also called a "prediction error signal". This spectrally bleached audio signal is input to a residual / drive encoder 1705, which generates drive parameters. Therefore, speech input is encoded in excitation parameters, on the one hand, and LPC coefficients, on the other hand.
On the decoder side illustrated in Fig, 17b, the excitation parameters are entered into the excitation decoder 1707, which generates an excitation signal, which can be input into an LPC synthesis filter. The LPC synthesis filter is adjusted using the transmitted LPC filter coefficients.
Therefore, the LPC 1709 synthesis filter generates a reconstructed or synthesized speech output signal.
Over time, many methods have been proposed for efficient and perceptually convincing representation of the residual (excitation) signal, such as Multi-Pulse Excitation (MPE), Regular Pulse Excitation (RPE, and Line Excited Linear Prediction (CELP).
Linear Prediction Coding attempts to produce an estimate of the current sample value of a sequence based on the observation of a certain number of past values as a linear combination of the past observations. In order to reduce redundancies in the input signal, the LPC encoder filter "blanks" the input signal into its spectral envelope, that is, it is a model of the Inverse of the signal spectral envelope. In contrast, the decoder LPC synthesis filter is a model of the spectral envelope of the signal. Specifically, the well-known linear autoregressive predictive analysis (ARj is known to model the signal spectral envelope using an all-pole approximation.
Typically, narrowband speech encoders (i.e. 8kHz sample rate speech encoders) employ an LPC filter with an order between 8 and 12. Due to the nature of the LPC filter, a uniform frequency resolution is effective across the entire frequency range. This does not correspond to a perceptual frequency scale.
In order to combine the strengths of traditional LPC / CELP based encoding (better quality for speech signals) and traditional bank filter based perceptual audio encoding approach (better for music), combined encoding has been proposed between these architectures. In the encoder
AMR-WB + (AMR-WB = Adaptive Multi-speed Broadband) B. Bessette, R. Lefebvre, R. Salami, UNIVERSAL SPEECH / AUDIO CODING USING HYBRID ACELP / TCX TECHNIQUES, ”Proc. IEEE ICASSP 2005, pp. 301-304, 2005 two alternative coding nuclei operate on a residual signal from the LPC. One is based on ACELP (ACELP = Linear Prediction by Excitation with Code
Algebraic) and is therefore extremely efficient for encoding speech signals. The other encoding core is TCX-based (TCX = Transformer Encoded Excitation), i.e. a bank filter-based encoding approach that resembles traditional audio encoding techniques in order to achieve good quality for musical signs.
Depending on the characteristics of the input signals, one of the two coding modes is selected for a short period of time to transmit the residual LPC signal. In this way, 80ms frames can be divided into 40ms or 20ms subframes in which a decision is made between the two coding modes.
0 The AMR-WB + (AMR-WB + = Extended Adaptive Multi-rate codec), cf.
3GPP (3GPP = Third Generation Partnership Project) technical specification number 26,290, version 6.3.0, June 2005, can change between two essentially different modes ACELP and TCX. ACELP mode a signal
Ί time domain is encoded by the excitation of algebraic code. In TCX mode a Fourier transform (FFT = Fast Fourier transform) is used and the spectral values of the LPC weighted signal (from which LPC excitation can be derived) are encoded based on vector quantization.
The decision, which modes to use, can be made by testing and decoding the two options and comparing the segmental signal-to-noise ratios (SNR = Signal-to-Noise Ratio).
This case is also called a closed-loop decision, since there is a closed control loop, evaluating the coding performance or efficiencies, respectively, and then choosing the one with the best SNR.
It is known that for windowed speech and audio encoding applications a windowless block transformer is not possible. Therefore, for TCX mode the signal is divided by window with low overlay window with 1/8 overlay. This overlap region is required for the purposes of fading from a previous block or frame while merging into the next, for example to suppress artifacts (in this context refers to conversion errors) due to noise from uncorrelated quantization in consecutive audio frames. In this way, the overhead compared to the noncritical sample is kept reasonably low and the decoding required for the closed-loop decision reconstructs at least 7/8 of the samples in the current frame.
The AMR-WB + introduces 1/8 of overhead in TCX mode, that is, the number of spectral values to be encoded is 1/8 greater than the number of input samples. This provides the disadvantage of increased data overhead. Also, the frequency response of the corresponding bandpass filters is not advantageous due to the deep overlap region of
1/8 of the consecutive frames.
In order to further detail a code overload and consecutive frame overlay, Fig. 18 illustrates a definition of window parameters. The window shown in Fig. 18 it has a rising flank part on the left side, which is called the "L" and is also called the left overlay region, a central region named with the "1", which is also called the 1-region or override part ( bypass), and a descending flank part, which is called the "R" and is also called the right overlay region. In addition, Fig. 18 shows an arrow indicating the "perfect reconstruction PR" region within a frame. Furthermore, Fig. 18 shows an arrow indicating the length of the transformation core, which is called the "T".
Fig. 19 shows a graph of an AMR-WB + window sequence and at the end a table of a window parameter according to Fig. 18. The window sequence shown at the top of Fig. 19 is ACELP, TCX20 (for a 20ms frame length), TCX20, TCX40 (for a 40ms frame length), TCX80 (for an 80ms frame length),
TCX20, TCX20, ACELP, ACELP.
From the window sequence one can see the various overlapping regions, which overlap exactly 1/8 of the central part M. The table at the bottom of Fig. 19 also shows that the transformation length "T" is always 1 / 8 larger than the region of perfectly reconstructed new samples "PR". Also, it should be noted that this is not only the case for ACELP to TCX transitions, but also for transitions from TCXx to TCXx (where "x" indicates TCX frames of arbitrary length). Therefore, an overload of 1/8 is introduced in each block, that is, the critical sample is never reached.
When switching from TCX to ACELP the window samples are discarded from the FFT-TCX frame in the overlap region, as indicated, for example, at the top of Fig. 19 by the region marked 1900. When changes the zero input response (ZIR = zero input response) from ACELP to TCX, which is also indicated by the dotted line 1910 at the top of Fig. 19, it is removed in the encoder before splitting into windows and added to the decoder for retrieval. When switching from TCX to TCX frames the window split samples are used for cross fading. Since TCX frames can be quantized differently, the quantization error or quantization noise between consecutive frames can be different and / or independent. Therefore, when switching from one frame to the other without cross fading, noticeable artifacts can occur, and consequently cross fading is necessary to acquire a certain quality.
From the table at the bottom of Fig. 19 it can be seen that the cross fade region grows with increasing weft length. Fig. 20 provides another table with illustrations of the different windows for the possible transitions in AMR-WB +. When transitioning from TCX to ACELP, overlap samples may be discarded. When transitioning from ACELP to TCX, the zero input response from the ACELP can be removed in the encoder and added to the decoder for recovery.
Next, the audio coding will be explained, which uses time domain (TD = Time Domain) and frequency domain (FD = Frequency Domain) encoding. Also, between the two coding domains, the change can be used. In Fig. 21, a timeline is shown during which a first frame 2101 is encoded by an FD encoder followed by another frame 2103, which is encoded by a TD encoder and overlaps in region 2102 with the first frame 2101. The frame Time domain encoded 2103 is followed by a frame 2105, which is encoded in the frequency domain again and overlaps in region 2104 with the preceding frame 2103. Overlap regions 2102 and 2104 occur whenever the coding domain is changed.
The purpose of these overlapping regions is to smooth the transitions. However, the overlapping regions may still be susceptible to loss of coding efficiency and artifacts. Therefore, the overlapping rulers or transitions are generally chosen as a compromise between some overload of the transmitted information, that is, coding efficiency and the quality of the transition, that is, the audio quality of the decoded signal. . In order to establish this compromise, care must be taken when handling transitions and designing transition windows 2111, 2113 and 2115 as indicated in Fig. 21.
Conventional concepts regarding the manipulation of transitions between the coding modes of the frequency domain and the time domain are, for example, the use of cross fade windows, that is, introducing an overload as large as the overlap region. A cross fade window is used, the fading of the preceding frame and intensification of the next frame simultaneously. This approach, due to overload, introduces deficiencies in decoding efficiency, since whenever a transition is made, the signal does not continue to be critically sampled. Critically sampled overlapping transformations are described, for example, in J. Princen, A. Bradley, "Analysis / Synthesis Filter Bank Design Based on Time Domain Aliasing Cancellation", IEEE Trans. ASSP, ASSP-34 (5): 1153-1161, 1986, and are used, for example, in AAC (AAC = Advanced Audio Coding), cf. Generic Coding of Motion Images and Associated Audio: Audio Coding
Advanced, International Standard 13818-7, ISO / IEC JTC1 / SC29 / WG11 MPEG,
1997.
Also, cross-fade transitions without aliasing are described in Fielder, Louis D., Todd, Craig C., “The Design of a Video Friendly Audio Coding System for Distribution Applications,” Brief No. 17-008, La
AES International Conference 17: High Quality Audio Coding (August 1999) and in Fielder, Louis D., Davidson, Grant A., “Audio Coding Tools for Digital Television Distribution”, Preprint Number 5104, Convention 108 of the AES (January 2000).
WO 2008/071353 describes a concept for switching between a time domain and frequency domain encoder. The concept could be applied to any codec based on a time domain / frequency domain change. For example, the concept could be applied to time domain encoding according to the ACELP mode of the AMR-WB + codec and AAC as an example of the frequency domain codec. Fig. 22 shows a block diagram of a conventional encoder that uses a frequency domain decoder in the upper branch and a time domain decoder in the lower branch. The frequency decoding portion is exemplified by an AAC decoder, which comprises a re-quantization block 2202 and a modified discrete cosine transform block 2204. In AAC the modified discrete cosine transform (MDCT = Modified Discrete Cosine Transform) is used as a transformation between the time domain and the frequency domain. In Fig. 22 the time domain decoding path is exemplified as an AMR-WB + 2206 decoder followed by a MDCT 2208 block, for the purpose of combining the result of decoder 2206 with the result of requalifier 2202 in the domain of frequency.
This allows a combination in the frequency domain, while the overlap and aggregate stage, not shown in Fig. 22, can be used after the reverse MDCT 2204, in order to combine the adjacent cross-fade blocks, regardless of whether they have been encoded in the time domain or the frequency domain.
In another conventional approach that is described in W02008 / 071353 is to avoid MDCT 2208 in Fig. 22, ie DCT-IV and IDCT-IV in the case of time domain decoding, another approach of the called time domain aliasing cancellation (TDAC = Cancellation of Aliasing of
Time domain). This is shown in Fig. 23. Fig. 23 shows another decoder having a frequency domain decoder exemplified as an AAC decoder comprising a re-quantization block 2302 and an IMDCT block 2304. The time domain path is again exemplified by an AMR-WB + 2306 decoder and TDAC 2308 block. The decoder shown in Fig. 2. 3 allows a combination of decoded blocks in the time domain, i.e. after IMDCT 2304, since the TDAC 2308 introduces the necessary time aliasing for the appropriate combination, i.e. for the cancellation of time aliasing, directly in the time domain. To save some calculations instead of using MDCT on each first and last superframe, that is, on each 1024 sample, s, of each AMR-WB + segment, TDAC can only be used in overlap zones or regions on 128 samples. Aliasing in normal time domain introduced by AAC processing can be maintained, while aliasing in inverse time domain in AMR-WB + parts is introduced.
0 Cross fade windows not subject to aliasing have the. disadvantage that they are not efficient in coding, because they generate coded coefficients not sampled critically, and add an information overload to encode. The introduction of TDA (TDA = Aliasing in Domain of
Time) in the time domain decoder, for example, in WO 2008/071353, reduces this overhead, but could only be applied as the time frames of the two encoders merge with each other. Otherwise, the coding efficiency is reduced again. Furthermore, the
TDA on the decoder side can be problematic, especially at the starting point of a time domain encoder. After a potential reset, a time domain encoder or decoder will generally produce a burst of quantization noise due to memory voiding of the time domain encoder or decoder using, for example, LPC (LPC = Encoding Linear Prediction). The decoder will then take some time before being in a permanent or stable state and providing more uniform quantization noise over time. This burst error is not advantageous as it is generally audible.
Therefore, it is an object of the present invention to provide an improved concept for changing audio encoding in multiple domains.
The objective is achieved by an encoder according to claim 1, and methods for encoding according to claim 16, an audio decoder according to claim 18 and a method for audio decoding according to claim 32.
It is a finding of the present invention that an improved change in an audio encoding concept using time domain and frequency domain encoding can be achieved, when the dithering of the corresponding encoding domains is adapted or cross fade windows are used modified. In one embodiment, for example, AMR-WB + can be used as a time domain codec and AAC can be used as an example of a frequency domain codec, a more efficient switch between the two codes can be achieved by embodiments, either by adaptation dithering of the AMR-WB + portion or by using start or stop windows for the respective AAC encoding portion.
It is another finding of the invention that TDAC can be applied in the decoder and cross fade windows can be used without aliasing.
Embodiments of the present invention can provide the advantage that an information overload, introduced in an overlapping transition, can be reduced, while maintaining moderate to moderate cross fading regions which ensures the quality of cross fading. Embodiments of the present invention will be detailed using the attached figures, in which
Fig. 1a shows an embodiment of an audio encoder;
Fig. 1 b shows an embodiment of an audio decoder;
Figs. 2a-2j show equations for the MDCT / IMDCT;
Fig. 3 shows an embodiment using modified screening;
0 Fig. 4a shows a quasi-periodic signal in the time domain;
Fig. 4b shows a sound signal in the frequency domain;
Fig. 5a shows a sound like signal in a time domain;
Fig. 5b shows a non-sound signal in the frequency domain;
Fig. 6 shows a CELP synthesis analysis;
Fig. 7 illustrates an example of an LPC analysis step in one embodiment;
Fig. 8a shows an embodiment with a modified stop window 5;
Fig. 8b shows an embodiment with a modified start-stop window;
Fig. 9 shows a beginning window;
Fig. 10 shows a more advanced window;
Fig. 11 shows an embodiment of a modified stop window;
Fig. 12 illustrates an embodiment with different zones or regions of overlap;
Fig. 13 illustrates an embodiment of a modified start window;
Fig. 14 shows an embodiment of an aliasing-free modified stop window applied to an encoder;
Fig. 15 shows an aliasing-free modified stop window applied on the decoder;
Fig. 16 illustrates examples of conventional encoder and decoder;
0 Figs. 17a, 17b illustrate LPCs for sound and non-sound signals;
Fig. 18 illustrates a prior art cross fade window;
Fig. 19 illustrates a prior art sequence of AMRWB + windows;
Fig. 20 illustrates windows used to transmit on AMR-WB + between
ACELP and TCX;
Fig. 21 shows an example of a sequence of consecutive audio frames in different coding domains;
Fig. 22 illustrates the conventional approach to audio decoding in different domains; and
Fig. 23 illustrates an example of time domain aliasing cancellation.
Fig. 1a shows an audio encoder 100 for encoding audio samples. Audio encoder 100 comprises a first time domain aliasing introducer encoder 110 for encoding audio samples in a first encoding domain, the first time domain aliasing introducer encoder 110 has a first screening rule, a window of start and a stop window. Also, the audio encoder 100 comprises a second encoder 120 for encoding audio samples in the second encoding domain. The second encoder 120 has a number of audio samples of predetermined frame size and a number of audio samples of encoding timing period. The encoding tuning period may be true or predetermined, it may be dependent on the audio samples, an audio sample frame or a sequence of audio signals. The second encoder 120 has a different second screening rule. A frame of the second encoder 120 is an encoded representation of a number of audio samples later in time, the number being equal to the number of audio samples of predetermined frame size.
The audio encoder 100 further comprises a controller 130 for switching from the first time domain aliasing introducer encoder 110 to the second encoder. 120 in response to a characteristic of the audio samples and to modify the second screening rule in response to a change from the first time domain aliasing introducer encoder 110 to the second encoder 120 or to modify the start or stop window of the first introducer aliasing encoder in time domain 110, where the second screening rule remains unchanged.
In embodiments controller 130 may be adapted to determine the characteristic of the audio samples based on the input audio samples or based on the output of the first time domain aliasing introducer encoder 110 or the second encoder 120. This it is indicated by the dotted line in Fig. 1a, through which input audio samples can be provided by controller 130. More details of the change decision will be described below. .
In embodiments controller 130 can control the first time domain aliasing introducer encoder 20 and the second encoder 120 in one way, which both encode the audio samples in parallel, and controller 130 makes the change decision based to the respective result, carry out the modifications before the change. In other embodiments the controller
130 You can analyze the characteristics of the audio samples and decide which encoding branch to use, but turning off the other branch. In such an embodiment the coding tuning period of the second encoder 120 becomes relevant, as before the change, the coding tuning period has to be taken into consideration, which will be detailed below.
In embodiments, the first time domain allaslng introducer 110 may comprise a frequency domain transformer to transform the first frame of subsequent audio samples to frequency domain. The first time domain allaslng Introducer encoder 110 can be adapted to weight the first frame encoded with the Start window, when the subsequent frame is encoded by the second encoder 120 and can further be adapted to weight the first frame encoded with the start window stop when a preceding frame is to be encoded by a second encoder 120.
It should be noted that different notations can be used, the first Allaslng Introducer encoder in time domain 110 applies a Start window or a Stop window. Here, and for the rest it is assumed that an ice window is applied before the change to the second encoder 120 and when it is changed again from the second encoder 120 to the first introducer encoder in time domain 120 the stop window is applied in a first time domain allasing introducer encoder 110. Without loss of generality, the expression could be used vice versa in reference to the second encoder 120. In order to avoid confusion, here the expressions "ice" and "stop" refer to windows applied to the first encoder 110, when the second encoder 120 starts or after it is stopped.
In embodiments the frequency domain transformer as used in the first aliasing encoder in time domain 110 can be adapted to transform the first frame in the frequency domain based on a MDCT and the first aliasing introducer encoder in domain of Time 110 can be tailored to fit a MDCT size to the modified start and stop or start and stop windows. The details for the MDCT and its size will be established below.
In embodiments, the first time domain aliasing introducer encoder 110 may accordingly be adapted to use a start and / or stop window having an aliasing free part, i.e., within the window there is a part, without aliasing in time domain. Also, the first time domain aliasing introducer encoder 110 can be adapted to use a start window and / or a stop window having an aliasing free part in a rising edge part of the window, when the preceding frame it is encoded by the second encoder 120, that is, the first time domain aliasing introducer encoder 110 uses a stop window, which has a leading edge portion that is free from aliasing. Accordingly, the first time domain aliasing introducer encoder 110 can be adapted to use a window having a falling edge portion that is free from aliasing, when a subsequent frame is encoded by the second encoder 120, i.e. using a stop window with a falling flank part, which is free from aliasing.
In embodiments, controller 130 may be adapted to start second encoder 120 such that a first frame in a frame sequence of second encoder 120 comprises an encoded representation of the samples processed in the preceding aliasing free portion of the first aliasing encoder in time domain 110. In other words, the output of the first time domain aliasing introducer encoder 110 and the second time encoder 120 can be coordinated by controller 130 such that the aliasing free portion of the encoded audio samples of the first domain aliasing introducer encoder Time 110 overlaps with the output of audio samples encoded by the second encoder 120. Controller 130 can further be adapted for cross fading, ie fading of one encoder while intensifying the other encoder.
Controller 130 can be adapted to start second encoder 120 such that the number of samples of the encoding tuning period overlaps with the aliasing free part of the start window of the first time domain aliasing introducer encoder 110 and a subsequent frame of the second encoder 120 overlaps the aliasing part of the stop window. In other words, controller 130 can coordinate second encoder 120 such that unaliased audio samples from the encoding tuning period are available from first encoder 110, and when only aliasing audio samples are available from the first introducer aliasing encoder in time domain 110, the tuning period of the second encoder 120 has ended and encoded the audio samples are available at the output of the second encoder 120 in a regular manner.
Controller 130 may further be adapted to start second encoder 120 such that the encoding tuning period overlaps with the aliasing portion of the start window. In this embodiment, during the overlay part, the aliasing audio samples are available from the output of the first time domain aliasing introducer encoder 110, and to the output of the second encoder 120 coded audio samples of the tuning period , which may have increased quantization noise, may be available. Controller 130 can still be adapted for cross fading between two suboptimal encoded audio streams during the overlay period.
In other embodiments, controller 130 may further be adapted to change the first encoder 110 in response to a different characteristic of the audio samples and to modify the second screening rule in response to the change of the first time domain aliasing introducer encoder. 110 to the second encoder 120 or to modify the start sale or stop window of the first encoder, where the second screening rule remains unchanged. In other words, controller 130 can be adapted to switch back and forth between the two audio encoders.
In other embodiments controller 130 may be adapted to start the first time domain aliasing introducer encoder 110 such that the aliasing free portion of the stop window overlaps the frame of the second encoder 120. In other words, in embodiments The controller can be adapted for cross fading between the outputs of the two encoders. In some embodiments, the output of the second encoder fades, while only the suboptimal encoded output, i.e., the aliasing audio samples from the first time domain aliasing introducer encoder 110 are boosted. In other embodiments, controller 130 may be adapted for cross fading between a frame of the second encoder 120 and frames without aliasing of the first encoder 110.
In embodiments, the first time domain aliasing introducer encoder 110 may comprise an AAC encoder in accordance with the Generic Motion Image Coding and Associated Audio: Audio Coding
Advanced, International Standard 13818-7, ISO / IEC JTC1 / SC29 / WG11 MPEG,
1997.
In embodiments, the second encoder 120 may comprise an AMR-WB + encoder in accordance with 3GPP (3GPP = Third Generation Partnership Project), Technical Specification 26.290, Version 6.3.0 as of June 2005 "Codec Processing Function of Audio; Extended Adaptive Multi-speed Broadband Codec; Transcoding Functions ”, broadcast 6.
Controller 130 can be adapted to modify the AMR or AMR-WB + screening rule such that a first AMR superframe comprises five AMR frames, where according to the aforementioned technical specification, a superframe comprises four regular AMR frames, compare Fig. 4, Table 10 on page 18 and Fig. 5 on page 20 of the aforementioned Technical Specification. As will be detailed later, controller 130 can be adapted to add an extra frame to an AMR superstar. It should be noted that in embodiments the superframe can be modified by a frame attached to the beginning or end of any superframe, i.e., the screening rules can also be attached to the end of a superframe.
Fig. 1b shows an embodiment of an audio decoder 150 for.
decode encoded frames of audio samples. The audio decoder
150 it comprises a first introducer aliasing decoder in 15 time domain 160 to decode audio samples in a first decoding domain. The first time domain aliasing introducer encoder
160 it has a first screening rule, a start window, and a stop window. Audio decoder 150, furthermore, comprises a second decoder 170 for decoding audio samples in a second decoding domain. The second decoder 170 has a number of predetermined frame size audio samples and a number of encoding tuning period audio samples. Also, the second decoder 170 has a different second screening rule. A frame of the second decoder 170 may correspond to a decoded representation of a number of audio samples later in time, where the number equals the number of audio samples of predetermined frame size.
The audio decoder 150 further comprises a controller 180 for changing from the first time domain aliasing input decoder 160 to the second decoder 170 based on an indication in the coded frame of the audio samples, wherein controller 180 is adapted to modify the second screening rule in response to changing the first time domain introducer decoder 160 to the second decoder 170 or to modify the start or stop window of the first decoder 160, wherein the second screening rule remains unchanged.
According to the above description, for example, in the AAC encoder and decoder, the start and stop windows apply to both the encoder and the decoder. In accordance with the above description of the audio encoder 100, the audio decoder 150 provides the corresponding decoding components. The change indication of controller 180 can be provided in terms of a bit, a flag, or any supplemental information along with the encoded frames.
In certain embodiments, the first decoder 160 may comprise a
0 time domain transformer for transforming a first frame of decoded audio samples to the time domain. The first time domain aliasing introducer decoder 160 may be adapted to weight the first decoded frame with the start window when a subsequent frame is decoded by the second decoder 170 and / or to weight the first decoded frame with the stop window when a preceding frame is to be decoded by the second decoder 170. The time domain transformer can be adapted to transform the first frame to the time domain based on an inverse MDCT (IMDCT = inverse MDCT) and / or the first time domain aliasing introducer decoder 160 can be adapted to adapt a size of IMDCT to the modified start and / or stop or start and / or stop windows. IMDCT sizes will be detailed later.
In certain embodiments, the first time domain aliasing input decoder 160 can be adapted by using a start window and / or a stop window that are aliasing free or contain an aliasing free portion. The first time domain aliasing introducer decoder 160 can be further adapted by using a stop window containing an aliasing free part in the rising part of the window when the preceding frame has been decoded by the second decoder 170 and / or the first time domain aliasing input decoder 160 may have a start window that has an aliasing free part on the falling edge when the frame posterior is decoded by the second
0 decoder 170.
Concerning the previously described embodiments of audio encoder 100, controller 180 may be adapted to initiate second decoder 170 such that the first frame of a frame sequence of second decoder 170 comprises a decoded representation of a processed sample in the preceding aliasing free part of the first decoder 160. Controller 180 may be adapted to start second decoder 170 such that the number of encoding tuning period audio samples overlaps with the aliasing free part of the start window of the first domain aliasing introducer decoder in Time 160 and a subsequent frame of the second decoder 170 overlaps with the aliasing part of the stop window.
In other embodiments, controller 180 may be adapted to start 10 second decoder 170 such that the encoding tuning period overlaps with the aliasing portion of the start window.
In other embodiments, controller 180 may be further adapted to change from second decoder 170 to first decoder 160 in response to an indication of the encoded audio samples and to modify the second screening rule in response to changing from second decoder 170 to first decoder 160 or to modify the start window or the stop window of the first decoder 160, where the second screening rule remains unchanged. The indication may be provided in terms of a flag, a bit or any supplementary information together with the encoded frames.
0 In certain embodiments, controller 180 may be adapted to start the first time domain aliasing introducer decoder 160 such that the aliasing portion of the stop window overlaps a frame of the second decoder 170.
Controller 180 can be adapted to apply cross-fade between consecutive frames of decoded audio samples from different decoders. Also, controller 180 can be adapted to determine aliasing in an aliasing part of the start window or stop window of a decoded frame of the second decoder 170 and controller 180 can be adapted to reduce aliasing in the aliasing part based on the aliasing determined.
In certain embodiments, controller 180 may be further adapted by ruling out the encoding tuning period of the audio samples from the second decoder 170.
Next, the modified discrete cosine transform (MDCT = Modified Discrete Cosine Transform) will be explained in more detail and the IMDCT will be described. The MDCT will be explained in more detail with the help of the equations illustrated in Figures 2a-2j. The modified discrete cosine transform is a Fourier-related transform based on the type IV discrete cosine transform (DCT-IV = Cosine Transform
Discrete Type IV), with the additional property of being overlapped, that is, it is designed to be carried out in consecutive blocks of a larger data set, where subsequent blocks are superimposed so that, for
0 For example, the last half of a block matches the first half of the next block. This overlay, in addition to the power compression qualities of DCT, makes the MDCT especially attractive for signal compression applications, as it helps prevent artifacts from going outside the block boundary. Therefore, an MDCT in MP3 (MP3 = MPEG2 / 4 layer 3), AC-3 (AC-3 = Audio Codec 3 by Dolby), Ogg Vorbis, and AAC (AAC = Advanced Audio Coding) is used to audio compression, for example.
MDCT was proposed by Princen, Johnson, and Bradley in 1987, 5 subsequent to previous work (1986) by Princen and Bradley to develop the underlying MDCT principle of time domain aliasing cancellation (TDAC). English), described below. Too<sup>1</sup> there is an analogous transform, the MDST (MDST = Modified DST, DST = Discrete Sinus Transform), based on the discrete sine transform, as well as other, rarely used, forms of the MDCT based on types of DCT combinations or DCT combinations / DST, which can also be used in embodiments by the time domain aliasing introducer transformer.
In MP3, MDCT is not applied to the audio signal directly, but to an output of a 32-band polyphase quad filter filter (PQF = Polyphase Quadrature Filter). The output of this MDCT is postprocessed by an alias reduction formula to reduce the typical, filter bank aliasing
PQF. Such a combination of a bench filter with an MDCT is called a hybrid filter bank or a subband MDCT. AAC, on the other hand, normally uses a pure MDCT; only the MPEG-4 AAC-SSR variant (rarely used) (for
Sony) uses a four-band PQF bank followed by an MDCT. ATRAC (ATRAC = Adaptive Transformed Audio Coding) uses stacked quadrature mirror filters (QMF) followed by an MDCT.
As an overlapping transform, the MDCT is a bit unusual compared to the other Fourier-related transforms in that it has half the outputs as the inputs (instead of the same number). In particular, it is a linear function F; R<sup>2N</sup> -> R<sup>N</sup>, where R denotes the set of real numbers. The real numbers 2N xo ..... X2N-1 are transformed into the real numbers of N Xo, ....
Xn-i according to the formula in Fig. 2a.
The normalization coefficient against this transform, here unity, is an arbitrary convention and differs between treatments. Only the product of the MDCT and IMDCT normalizations, below, is constrained.
The Reverse of MDCT is known as the IMDCT 'Since there are different numbers of inputs and outputs, in principle it may seem that the MDCT should not be invertible. However, perfect invertibility is achieved by adding overlapping IMDCTs to subsequent overlapping blocks, causing errors to cancel and original data to be recovered; this technique is known as time domain aliasing cancellation (TDAC).
The IMDCT transforms real numbers N Xo ..... Xn-i into real numbers 2N i ..... y2N-i according to the formula in Fig. 2b. As for DCT-IV, an orthogonal transform, the inverse has the same shape as the direct transform.
0 In the case of MDCT divided by windows with the usual window normalization (see below), the normalization coefficient against IMDCT must be multiplied by 2, that is, 2 / N is taken.
Despite the direct application of the MDCT formula it would require O (N<sup>2</sup>), it is possible to compute the same thing with only complexity O (N log N) by recursive factoring of computation, as in fast Fourier transform (FFT). One can also compute MDCTs through other transforms, typically a DFT (FFT) or a DCT, combined with O (N) pre and post processing steps. Also, as described below, any algorithm for DCT-IV immediately provides a method for computing MDCT and IMDCT of equal size.
In typical signal compression applications, the properties of the transform are likewise improved using a window function w<sub>n</sub> (n = 0, ..., 2N-1) that multiplies with x<sub>n</sub> yy<sub>n</sub> in the MDCT and IMDCT formulas above in order to avoid discontinuities in the n = 0 and 2N limits by making the function go smoothly from zero to those points. That is, the information is divided by windows before the MDCT and after the IMDCT. In principle, x and y could have different window functions, and the window function could also change from one block to the next, especially for the case where data blocks of different sizes are combined, but for simplicity the common case of Identical window functions for blocks of equal size are considered first.
The transform remains invertible, that is, TDAC works for a symmetric window w<sub>n</sub> = w<sub>2N</sub>-in, provided that w meets the PrincenBradley condition according to Fig. 2c.
The various different window functions are common, an example is given in Fig. 2d for MP3 and MPEG-2 AAC, and in Fig. 2e for Vorbis. AC-3 uses a derived window (KBD = Kaiser-Bessel Derived), and MPEG-4 AAC can also use a KBD window.
Note that the windows applied to the MDCT are different from the windows used for other types of signal analysis, since they must meet the Princen-Bradley condition. One of the reasons for this difference is that the MDCT windows are applied twice, for the MDCT (analysis filter) and the IMDCT (synthesis filter).
As can be seen from the inspection of the definitions, for equal N the MDCT is essentially equivalent to DCT-IV, where the input is changed to N / 2 and two N blocks of data are transformed immediately. By examining this equivalence more carefully, important properties like TDAC can be easily derived.
In order to define the precise relationship to DCT-IV, one must realize that DCT-IV corresponds to alternating even / odd limit conditions, it is even at its left limit (approximately n = -1 / 2), odd in its right limit (approximately n = N-1/2), and so on (instead of periodic limits as for a DFT). This follows from the identities given in Fig. 2f. For the
0 So, if your inputs with an order x of length N, imagine extending this order to (x, -xr, —x, xr, ...) and so on you can imagine, where xr denotes x in a reverse order.
Consider an MDCT with 2N inputs and N outputs, where the inputs can be divided into four blocks (a, b, c, d) each of size N / 2. If these are changed to N / 2 (from the term + N / 2 in the MDCT definition), then (b, c, d) extends past the end of the N DCT-IV inputs, therefore they must “bend ”Again in accordance with the limit conditions described above.
Therefore, the MDCT of inputs 2N (a, b, c, d) is exactly equivalent to a DCT-IV of inputs N: (-c<sub>R</sub>-d, ab<sub>R</sub>), where R denotes inversion as before. In this way, any algorithm for computing DCT10 IV can be trivially applied to the MDCT.
Similarly, the IMDCT formula, as mentioned above, is precisely 1/2 of the DCT-IV (which is its own inverse), where the output is changed to N / 2 and extended (through the conditions limits) to a length of 2N. The inverse of DCT-IV would simply require returning the entries (-c<sub>R</sub>-d, ab<sub>R</sub>) of the above. When this is changed and extended through the boundary conditions, one gets the result shown in Fig. 2g. Half of the IMDCT outputs are therefore redundant.
One can understand how TDAC works. Suppose one computes the MDCT of the subsequent, 50% overlay, block 2N (c, d, e, f). The. IMDCT will then render, analogous to the above: (cd<sub>R</sub>, dc<sub>R</sub>, e + f<sub>R</sub>, e<sub>R</sub>+ f) / 2. When this is added with the previous IMDCT result in the overlapping half, the reverse terms cancel and one simply gets (c, d), retrieving the original data.
The Origin of the phrase "cancellation of aliasing in time domain" is now clear. Using input data that extends beyond the limits of the DCT-IV logic results in the data being aliased in exactly the same way that frequencies beyond the Nyquist frequency are subject to aliasing with lower frequencies , except that this aliasing occurs in the time domain instead of the frequency domain. Therefore, cd combinations<sub>R</sub> and following, which have precisely the correct signs for combinations to cancel when added.
For odd N (which is rarely used in practice), N / 2 is not an integer 10 so MDCT is not simply a change permutation of a DCT-IV. In this case, the additional change per half sample means that the MDCT / IMDCT becomes equivalent to DCT-lll / ll, and the analysis is analogous to the previous one.
The TDAC property above was tested for the common MDCT which shows that adding the IMDCTs to subsequent blocks in their overlapping half retrieves the original data. Deriving this inverse property for the windowed MDCT is a bit more complicated.
Recall from the above that when (a, b, c, d) and (c, d, e, f) are subject to MDCT, IMDCT, and added to their overlapping half, we obtain (c + d<sub>R</sub>, c<sub>R</sub> + d) / 2 + (c - d<sub>R</sub>, d - c<sub>R</sub>) / 2 = (c, d), the original data.
0 Now, multiplying the MDCT inputs and the IMDCT outputs by a window function of length 2N is assumed. As before, we assume a symmetric window function, which is, therefore, of the form (w, z, z<sub>R</sub>, w<sub>R</sub>), where w and z are vectors of length-N / 2 and R denotes inverse as before. Then the Princen-Bradley condition can be written w<sup>2</sup> + z<sup>2</sup>r = (1,1, ...), with the multiplications and sums performed per element, or equivalently u £ +? = (1,1, ...) reversing w and z.
Therefore, instead of subjecting MDCT (a, b, c, d), MDCT (wa, zb, z<sub>R</sub>c, w<sub>R</sub>d) it is submitted to MDCT with all the multiplications carried out by element. When this is subjected to IMDCT and multiplied again (per item) by the window function, the results of half of last N are shown in the
Fig. 2h.
Note that multiplication by <sup>1</sup>/> is no longer present, because the IMDCT normalization differs by a factor of 2 in the window case. Similarly, the window-divided MDCT and IMDCT of (c, d, e, f) yields, in their first half N according to Fig. 2¡. When these two halves are added together, the results of Fig. 2j are obtained, recovering the original data.
Next, an embodiment will be detailed in which encoder-side controller 130 and decoder-side controller 180, respectively, will modify the second screening rule in response to the change from the first encoding domain to the second encoding domain. In the embodiment, a smooth transition into a changed encoder, ie, switching between AMR-WB + and AAC encoding. In order to have a smooth transition, some overlap is used, that is, a short segment of a signal or a number of audio samples, to which both encoding modes are applied. In other words, in the following description, an embodiment is provided, wherein the first time domain allaslng encoder 110 and the first time domain allaslng decoder 160 correspond to AAC encoding and decoding. The second encoder 120 and decoder 170 correspond to AMR-WB + in ACELP mode. The embodiment corresponds to an option of the respective controllers 130 and 180 in which the AMR-WB + dithering is modified, that is, the second dithering rule.
Fig. 3 shows a timeline showing a number of windows and frames. In Fig. 3, a regular AAC window 301 is followed by an AAC Home window 302. In AAC, the AAG ice window 302 is used between long and short frames. For the purposes of illustrating the AAC screening legacy, ie, the first screening rule of the time domain aliasing introducer encoder 110 and decoder 160, a sequence of short AAC windows 303 is also shown in Fig. 3. The AAC 303 short window sequence ends with an AAC 304 stop window, which prints an AAC long window sequence. In accordance with the description above, it is assumed in the present embodiment that the second encoder 120, decoder
170, respectively, use the ACELP mode of the AMR-WB +. The AMR-WB + uses equal size frames of which a sequence 320 is shown in Fig. 3. Fig. 3 shows a sequence of pre-filter frames of different types according to ACELP in AMR-WB +. Prior to changing AAC to ACELP, controller 130 or 180 modifies the ACELP frame so that the first superframe 320 consists of five frames instead of four. Therefore, ACE 314 data is available on the decoder, while decoded AAC data is also available. Therefore, the first part can also be discarded in the decoder, since this refers to the encoding tuning period of the second encoder 120, the second decoder 170, respectively. In general, in other embodiments the AMR-WB + superframe can be extended by the frame attachment also at the end of the superframe.
Fig. 3 shows two mode transitions, ie from AAC to AMR-WB + and from AMR-WB + to AAC. In one embodiment, typical AAC codec start / stop windows 302 and 304 are used and the frame length of the AMRWB + codec is increased to overlap with the fade portion of the AAC codec start / stop window, i.e. the second frame rule is modified. According to Fig. 3, the transitions from AAC to AMR-WB +, ie
0 the first time aliasing input encoder 110 to the second encoder 120 or the first time aliasing input decoder 160 to the second decoder 170, respectively, is handled by maintaining the AAC frame and time domain frame extension in transition to cover overlay effects. The AMR-WB + superframe in transition, i.e. the first 320 superframe in Fig. 3, use five frames instead of four; the fifth covers the overlay. This introduces a data overhead, however, the embodiment provides the advantage of ensuring a smooth transition between AAC and AMR-WB + modes.
As mentioned above, controller 130 can be adapted to make the switch between the two coding domains based on the characteristic of the audio samples where different analysis or options are conceivable. For example, controller 130 can change the encoding mode based on a stationary or transient fraction of the signal.
Another option would be for the change to be made based on whether the audio samples correspond to a louder or non-voiced speech signal. In order to provide a detailed embodiment for determining the characteristics of the audio samples, in the following, an embodiment of the controller 130, which changes based on the voice similarity of the signal.
By way of example, reference is made to Figs. 4a and 4b, 5a and 5b, respectively. The quasi-periodic signal segments or pulse signal portions and the noise signal segments or signal portions are treated by way of example. Generally, the 130, 180 controllers can be adapted
0 to decide based on different criteria, such as seasonality, transience, spectral whiteness, etc. An exemplary criterion is given below as part of one embodiment. Specifically, a voiced speech is illustrated in Fig. 4a in the time domain and in Fig. 4b in the frequency domain and is explained as an example for a quasi-periodic pulse signal portion and a non-beep speech segment as an example for a noise signal portion is explained in connection with Figs. 5a and 5b.
Generally speaking, speech can be classified as voiced, non-voiced, or mixed. Sound 5 is quasi-periodic in the time domain and harmonically structured in the frequency domain, while non-sound speech is random and broadband. Furthermore, the energy of the sound segments is generally greater than the energy of the non-sound segments. The short-term spectrum of sound speech is characterized by its fine and formative structure. The fine harmonic structure is a consequence of the quasi-periodicity of speech and can be attributed to vibrating vocal cords. The formant structure, which is also called the spectral envelope, is due to the interaction of the source and the vocal tracts. The vocal tracts consist of the pharynx and the oral cavity. The shape of the spectral envelope that "fits" the short-term spectrum of sound speech is associated with the transfer characteristics of the vocal tract and spectral tilt (6 dB / octave) due to the glottis pulse.
The spectral envelope is characterized by a set of peaks, which are called formants. Formants are the resonant modes of the vocal tract. For the average vocal tract there are 3 to 5 formants below 5 kHz. The amplitudes and locations of the first three formants generally occur below 3 kHz are quite important, both in the synthesis of speech and perception. Higher formants are also important for broadband and non-sounding speech representations. The properties of speech are related to the physical speech production systems as below. Excitation of the vocal tract with quasi-periodic glottic air pulses by vibration of the vocal cords produces sound speech. The frequency of periodic pulses is referred to as the fundamental or pitch frequency. Forcing air through a constriction in the vocal tract produces non-voiced speech. Nasal sounds are due to acoustic coupling of the nasal tract to the vocal tract, and occlusive sounds are reduced by abruptly reducing air pressure, which was built behind the closure of the tract.
Therefore, an audio signal noise potion can be a stationary portion in the time domain as illustrated in Fig. 5a or a stationary portion in the frequency domain, which is different from the quasi-imposed portion. -periodic as illustrated in the example in Fig. 4a, due to the fact that the stationary portion in the time domain does not show permanently repeating pulses. As will be described later, however, the differentiation between the noise portions and the quasi-periodic pulse portions can also be observed after LPC for the purposes of signal excitation. LPC is a method that models the vocal tract and excitation of the vocal tracts. When the frequency domain of the signal is considered, the impulse signals show a prominent appearance of the individual formants, i.e., prominent peaks in Fig. 4b, while the stationary spectrum has a fairly wide spectrum as illustrated in Fig. 5b, or in the case of harmonic signals, a fairly continuous noise floor that has prominent peaks representing specific tones occurring, for example , in the music signal, but which do not have the regular distance from each other as the impulse signal in Fig. 4b.
Furthermore, quasi-periodic pulse portions and noise portions 5 may occur over time, i.e. this means that one portion of the time audio signal is noisy and another portion of the time audio signal is quasi-periodic. , that is, tonal. Alternatively or additionally, the characteristic of a signal may be different in different frequency bands. Therefore, the determination, if an audio signal is noisy or tonal, can be carried out by a selection frequency so that a certain frequency band or several certain frequency bands are considered noisy and other frequency bands are considered tonal. . In this case, a certain time portion of the audio signal may include tonal components and noisy components.
Next, the CELP analysis encoder will be analyzed synthetically with respect to Fig. 6. Details of a CELP encoder can also be found in “Speech Coding: A tutorial review”, Andreas Spanias, Proceedings of IEEE, Vol. 84, no. 10, October 1994, pp. 1541-1582. The CELP encoder as illustrated in Fig. 6 includes a long-term prediction component 60 and a short-term prediction component 62. In addition, a codebook is used.
0 used as noted at 64. A perceptual weighting filter W (z) is implemented at 66, and an error minimization driver is provided at 68. s (n) is the input audio signal in time domain . After being perceptually weighted, the weighted signal is input to a subtractor 69, which calculates the error between the weighted synthesis signal at block output 66 and the actual weighted signal Sw (n).
Generally, the short-term prediction A (z) is calculated by an LPC analysis step that will be explained later. Depending on this information, the long-term prediction A<sub>L</sub>(z) includes the long-term prediction gain b and delay T (also known as pitch gain and pitch delay). The CELP algorithm then encodes the residual signal obtained after the short and long term predictions using a codebook of, for example, Gaussian sequences. The ACELP algorithm, where "A" stands for "algebraic" has a specific code book designed algebraically.
The codebook can contain more or fewer vectors each vector has a length according to a number of samples. A gain factor g scales the code vector and the won encoded samples are filtered by the long-term synthesis filter and the short-term prediction synthesis filter. The "optimal" code vector is selected such that the perceptually weighted average square error is minimized. The search process in CELP is evident from the synthesis analysis scheme illustrated in Fig. 6. It should be noted that Fig. 6 only illustrates an example of a CELP analysis by synthesis and that the embodiments should not be limited to structure shown in Fig. 6.
In CELP, the long-term predictor is generally implemented as an adaptive codebook containing the pre-excitation signal. The long-term prediction delay and gain are represented by an adaptive codebook index and gain, which are also selected by minimizing the average square weighted error. In this case the drive signal consists of the addition of two scaled gain vectors, one from an adaptive codebook and the other from a fixed codebook. The perceptual weighting filter in AMR-WB + is based on the LPC filter, therefore the perceptually weighted signal is in the form of an LPC domain signal. In the transformation domain encoder used in AMR-WB +, the transformation is applied to the weighted signal. In the decoder, the drive signal can be obtained by filtering the decoded weighted signal through a filter consisting of the inverse of synthesis and weighting filters.
The functionality of an embodiment of a predictive coding analysis step 12 will be further analyzed in accordance with the embodiment shown in Figs. 7, using LPC analysis and LPC synthesis on controllers 130,180 in the corresponding embodiments.
Fig. 7 illustrates a more detailed implementation of an embodiment of an LPC analysis block. The audio signal is input to a filter determination block, which determines the filter information A (z), that is, the information on the coefficients for the synthesis filter. The information is quantized and the output is produced as the short-term prediction information required by the decoder. On a 786 subtractor, a current signal sample is entered and a prediction value for the current sample is subtracted such that for the sample, the prediction error signal is generated on line 784. Note that the prediction can also be called an excitation signal or an excitation frame (usually after encoding).
Fig. 8a shows another window time sequence that is accomplished with another embodiment. In the embodiment considered below, the AMR-WB + codec corresponds to the second encoder 120 and the AAC codec corresponds to the first time domain aliasing introducer encoder 110. The following embodiment maintains the AMR-WB + codec frame, that is, the second dither rule remains unchanged, but the window split is changed in the transition from AMR-WB + codec to AAC codec, start windows are manipulated / AAC codec stop. In other words, the window division of the AAC codec will be longer in the transition.
Figs. 8a and 8b illustrate this embodiment. Both Figures show a sequence of conventional AAC windows 801 in which, in Fig. 8a a new modified stop window 802 is introduced and in Fig. 8b, a new stop / start window 803. Regarding ACELP , similar screening is shown as described with respect to the embodiment used in Fig. 3. In the embodiment resulting in the window sequence as shown in Figs. 8a and 8b, it is assumed that normal AAC codec dithering is not maintained, that is, start, stop, or modified start / stop windows are used. The first window that appears in Figs. 8a is for the transition from AMR-WB + to AAC, where the AAC codec will use a long 802 stop window. Another window will be described with the help of Fig. 8b, showing the transition from AMR-WB + to AAC when the AAC codec will use a short window, using a long AAC window for this transition, as indicated in Fig. 8b. Fig. 8a shows that the first ACELP superframe 820 comprises four frames, that is, it is in accordance with conventional ACELP screening, that is, the second screening rule. In order to maintain the ACELP screening rule, that is, the second screening rule remains unchanged, the modified windows 802 and 803 are used as indicated in Figs. 8a and 8b.
Therefore, some details regarding window splitting will generally be introduced below.
Fig. 9 shows a general rectangular window, in which the window sequence information may comprise a first zero part, in which the window masks samples, a second bypass part, in which the samples of a frame, that is, an input time domain frame or an overlapping time domain frame can pass unchanged, and a third part zero, which again masks samples at the end of the frame. In other words, window functions can be applied, which deletes a number of samples from a frame in the first part zero, passes through the samples in the second bypass part, and then deletes samples at the end of the frame in a third part zero. In this context, deletion may also refer to attaching a sequence of zeros to the beginning and / or end of the bypass part of the window. The second bypass part can be such that the window function simply has a value of 1, i.e. samples pass
0 unchanged, i.e. the window function changes through the plot samples.
Fig. 10 shows another embodiment of a window sequence or window function, wherein the window sequence further comprises a leading edge portion between the first zero portion and the second bypass portion and a falling edge portion between the second part of bypass and the third part zero. The rising flank part can also be considered as a stepping part and the falling flank can be considered as the fading part. In embodiments, the second bypass part may comprise a sequence of ones by not modifying samples of an excitation frame in any way.
Returning to the embodiment shown in Fig. 8a, the modified stop window, as used in the embodiment that transits between the
AMR-WB + and AAC, when transiting from AMR-WB + to AAC is shown in more detail in Fig. 11. Fig. 11 shows ACELP frames 1101, 1102, 1103 and 1104. The modified stop window 802 is then used to transition to AAC, i.e. to the first time domain aliasing introducer encoder 110, decoder 160, respectively. According to the details of the MDCT set forth above, the window already starts in the middle of frame 1102, with a first zero portion of 512 samples. This part is followed by the rising edge part of the window, which extends through 128 samples followed by the second bypass part which, in this embodiment, extends to 576, that is, 512 samples after the leading edge part to which the first zero part is folded, followed by 64 more samples of the second bypass part, resulting from the third zero part at the end of the window extended through 64 samples. The falling edge portion of that window results in 1024 samples, which should overlap with the next window.
The embodiment can also be described using a pseudo code, which is exemplified by:
/ * Attack based block change * /
Yes (there is an attack) {nextwindowSequence = SHORT_WINDOW;
} or {nextwindowSequence = LONG_WINDOW;
} / * Block change based on ACELP Change Decision * / if (the next frame is AMR) {nextwindowSequence = SHORT_WINDOW;
} / * Block change based on ACELP change decision for STOP_WINDOW_1152 * / if (the real name is AMR && next frame is not AMR) {nextwindowSequence = STOP_WINDOW_1152;
/ }
/ 'Block change for STOPSTART_WINDOW_1152 * / si (nextwindowSequence == SHORT_WINDOW) {si (windowSequence == STOP_WINDOW_1152) {windowSequence = STOPSTART_WINDOW_1152;
} }
Returning to the embodiment shown in Fig. 11, there is a time aliasing doubling section within the rising edge portion of the window, which spans 128 samples. Since this section overlaps with the last ACELP 1104 frame, the output of the ACELP 1104 frame can be used to cancel aliasing in time on the leading edge part. Aliasing cancellation can be done in either the time domain or the frequency domain, in line with the examples described above. In other words, the output of the last ACELP frame can be transformed into the frequency domain and can then be overlaid with the rising edge portion of the modified stop window 802. Alternatively, TDA or TDAC can be applied to the last ACELP frame before overlapping with the rising edge portion of the modified 802 stop window.
The previously described embodiment reduces the overhead generated in the transitions. It also eliminates the need for modifications to be made to the time domain code dithering, i.e. the second dither rule.
Likewise, it also adapts the frequency domain encoder, that is, the time domain aliasing introducer encoder 110 (AAC), which is generally more flexible in terms of bit distribution and number of coefficients to transmit than a domain encoder of time, that is, the second encoder 120. ·
Later, another embodiment will be described, which provides an aliasing-free cross fading when the switch occurs between the first time domain aliasing introducer encoder than 110 and the second encoder 120, decoders 160 and 170, respectively. This embodiment provides the advantage that noise due to TDAC is avoided, especially at low bit rates, in case of start or restart procedures. The advantage is achieved by an embodiment having a modified AAC start window with no time aliasing on the right hand side or on the falling side of the window. The modified start window constitutes an asymmetric window, that is, the right side or the falling edge part of the window ends before the MDCT bend point point. Consequently, the window is free of time aliasing. At the same time, the overlap region can be reduced by embodiments to 64 samples instead of 128 samples.
In certain embodiments, audio encoder 100 or audio decoder 150 may take a certain time before entering a permanent or stable state. In other words, during the start period of the time domain encoder, that is, the second encoder 120 and also the decoder 170, a certain amount of time is needed in order to start, for example, the coefficients of an LPC. In order to smooth out the error in the case of resetting, in certain embodiments, the left part of an AMR-WB + input signal can be divided into windows with a short sine window at encoder 120, for example, which has a long of 64 samples. Also, the left part of the synthesis signal can be divided into windows with the same signal in the second decoder 170. In this way, the square sine window can be applied in a similar way to AAC, by applying the square sine to the right side of its home window.
By using this window splitting, in one embodiment, the transition from AAC to AMR-WB + can be accomplished without time aliasing and can be accomplished by a cross fade sine window such as, for example, 64 samples. Fig. 12 shows a timeline that exemplifies a transition from AAC to AMR-WB + and back to AAC. Fig. 12 shows an AAC 1201 start window followed by the AMR-WB + 1203 part that overlaps with the AAC 1201 window and with the ruler 1202, which spans 64 samples. The AMR-WB + part is followed by an AAC 1205 stop window, which overlaps with 128 samples.
According to Fig. 12, the embodiment applies to the respective aliasing free window in the transition from AAC to AMR-WB +.
Fig. 13 shows the modified Start window, as applied when transitioning from AAC to AMR-WB + on both sides at encoder 100 and decoder 150, encoder 110 and decoder 160, respectively.
The window that appears in Fig. 13 shows that the first part zero is not present. The window starts directly with the leading edge part, which extends through 1024 samples, that is, the doubling axis is in the middle of the 1024 Interval shown in Fig. 13. The axial symmetry appears then on the right side of Interval 1024. As can be seen in Fig. 13, the third part zero extends to 512 samples, that is, there is no allasing in the right-hand part of the entire window, that is, the bypass part extends from the center towards the beginning of Sample Interval 64. It can also be seen that the descending flank part extends through 64 samples, which provides the advantage that the cross section is narrow. Sample Interval 64 is used for cross fading, however no aliasing is present in this interval. Thus, only a low overload is introduced.
Embodiments with the modified windows described above can avoid encoding too much Overload Information, that is, encoding some samples twice. In accordance with the description above, similarly designed windows can optionally be applied to the transition from AMR-WB + to AAC according to an embodiment where again modifying the AAC window, also reducing the overlap to samples.
Thus, the modified stop window elongates 2304 samples in one embodiment and is used at one point 1152 MDCT. The Left part of the window can be made time-free by starting the fade after the MDCT dub axis. In other words, by making the first part zero larger than a quarter the entire size of MDTC. The complementary square sine window is then applied to the last 64 decoded samples from the AMR-WB + segment. These two cross fade windows allow a smooth transition from AMR-WB + to AAC by limiting the transmitted information of the overhead.
Fig. 14 illustrates the window for the transition from AMR-WB + to AAC such as 5 would have been applied on encoder side 100 in one embodiment. It can be seen that the dubbing axis is posterior to. 576 samples, that is, the first part zero is spread across 576 samples. These consequences on the left side of the entire window are aliasing free. Cross fade starts in the second quarter of the window, that is after 576 samples or, in other words, just beyond the dub axis. The cross fade section, that is, the rising flank part of the window can shrink to 64 samples according to Fig. 14.
Fig. 15 shows the window for the transition from AMR-WB + to ACC applied to the side of decoder 150 in one embodiment. The window is similar to the window described in Fig. 14, in such a way that the application of both windows through encoded samples and then decoded again results in a square sine window.
The following pseudo code describes an embodiment of a boot window selection procedure, when the change from AAC to AMR-WB + occurs.
These embodiments can also be described by using a pseudo code such as, for example:
/ * Fit to an Allowed Window Sequence * / s¡ (nextw¡ndowSequence == SHORT_WINDOW) {si (w¡ndow¿equence == LONG_WINDOW) {si (the actual frame is not AMR && next frame is AMR) {windowSequence = START_WINDOW_AMR;
<sup>5 } </sup>or {windowSequence = START_WINDOW;
} }
Embodiments as described above reduce information overload by using small overlapping regions in consecutive windows during transition. What is more, these embodiments provide the advantage that these small overlapping regions are still sufficient to smooth out blocking artifacts, i.e. to have smooth cross fade. Also, it reduces the impact of the error explosion due to the start of the time domain encoder, that is, the second encoder 120, decoder 170, respectively, by initiating with a fade input.
Summarizing the embodiments of the present invention provides the advantage that smoothed cross regions can be performed in a multi-mode audio encoding concept at high encoding efficiency, i.e. transition windows introduce only low overhead in terms of additional information to be transmitted. Even realizations allow the use of multiple-mode encoders, while adapting the frame or division of. window from one mode to another.
Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or apparatus corresponds to a step of a method or to a characteristic of a step of a method. . Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or element or feature of a corresponding apparatus.
The inventive encoded audio signal can be stored on a digital storage medium, or it can be transmitted on a transmission medium such as a wireless or wired transmission medium, such as the internet.
Subject to certain implementation requirements, the embodiments of this invention can be implemented in hardware or software. The implementation can be done through the use of a digital storage medium, for example a floppy disk, a DVD, a CD, a ROM, a
PROM, EPROM, EEPROM, or FLASH memory, with electronically readable control signals stored therein, that cooperate (or may cooperate) with a programmable computing system such that the method
0 respective is made.
Some embodiments in accordance with the invention comprise a data transporter having electronically readable control signals, capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
In general, the embodiments of the present invention can be implemented as a computer program product with a program code, which is operative to perform one of the methods when the computer program product operates on a computer. The program code can, for example, be saved on a machine-readable carrier.
Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier.
In other words, an embodiment of the inventive method thus constitutes a computer program that has a program code to perform one of the methods described herein, when the computer program operates on the computer.
Another embodiment of the methods of the invention therefore constitutes a data carrier (or a digital storage medium, or a computer readable medium) comprising, stored there, the computer program to perform one of the methods described in the Present.
Another embodiment of the methods of the invention therefore constitutes a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or signal sequence can, for example, be configured to be transferred via a data communication connection, for example, via the internet.
Another embodiment comprises a processing means, eg, a computer, or a programmable logic device, configured or adapted to perform one of the methods described herein.
Another embodiment comprises a computer that has the computer program installed to perform one of the methods described herein.
In some embodiments, a programmable logic device (eg, a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, a programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. In general, the methods are preferably performed by a hardware apparatus.
The embodiments described above simply illustrate the principles of the present invention. It is understood that the modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. Therefore, an attempt is made to find a limit only in the scope of the claims of the present patent and not in the specific details presented by means of a description and
0 explanation of the achievements found in the present.
Contents2
25 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25
200 members in 22 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 7985608 | United States of America | P | |
| 10382508 | United States of America | P | |
| 2009004651 | European Patent Office (EPO) | W |
Members200
| Document | Office | Kind | |
|---|---|---|---|
| EP2144171A1 | European Patent Office (EPO) | A1 | |
| EP2144230A1 | European Patent Office (EPO) | A1 | |
| AU2009267394A1 | Australia | A1 | |
| AU2009267466A1 | Australia | A1 | |
| AU2009267467A1 | Australia | A1 | |
| AU2009267555A1 | Australia | A1 | |
| CA2729878A1 | Canada | A1 | |
| CA2730195A1 | Canada | A1 | |
| CA2730204A1 | Canada | A1 | |
| CA2730315A1 | Canada | A1 | |
| CA2871372A1 | Canada | A1 | |
| CA2871498A1 | Canada | A1 | |
| WO2010003491A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2010003563A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2010003564A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2010003663A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201007705A | Taiwan Province of China | A | |
| TW201009815A | Taiwan Province of China | A | |
| TW201011738A | Taiwan Province of China | A | |
| TW201011739A | Taiwan Province of China | A | |
| AU2009301358A1 | Australia | A1 | |
| CA2739736A1 | Canada | A1 | |
| WO2010040522A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AR072421A1 | Argentina | A1 | |
| AR072424A1 | Argentina | A1 | |
| WO2010040522A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AR072556A1 | Argentina | A1 | |
| AR072738A1 | Argentina | A1 | |
| EP2301023A1 | European Patent Office (EPO) | A1 | |
| IL210331A0 | Israel | A0 | |
| IL210331D0 | Israel | D0 | |
| IL210332A0 | Israel | A0 | |
| IL210332D0 | Israel | D0 | |
| MX2011000362A | Mexico | A | |
| KR20110036906A | Republic of Korea | A | |
| EP2311032A1 | European Patent Office (EPO) | A1 | |
| EP2311034A1 | European Patent Office (EPO) | A1 | |
| WO2010003563A8 | World Intellectual Property Organization (WIPO) | A8 | |
| KR20110043592A | Republic of Korea | A | |
| MX2011000366AThis record | Mexico | A | |
| MX2011003824A | Mexico | A | |
| AR076060A1 | Argentina | A1 | |
| KR20110052622A | Republic of Korea | A | |
| MX2011000375A | Mexico | A | |
| KR20110055545A | Republic of Korea | A | |
| AU2009301358A8 | Australia | A8 | |
| CN102089758A | China | A | |
| CN102089811A | China | A | |
| CN102105930A | China | A | |
| CN102113051A | China | A | |
| KR20110081291A | Republic of Korea | A | |
| US2011173008A1 | United States of America | A1 | |
| US2011173010A1 | United States of America | A1 | |
| US2011173011A1 | United States of America | A1 | |
| EP2345030A2 | European Patent Office (EPO) | A2 | |
| MX2011000369A | Mexico | A | |
| US2011202354A1 | United States of America | A1 | |
| CN102177426A | China | A | |
| ZA201009163B | South Africa | B | |
| US2011238425A1 | United States of America | A1 | |
| ZA201009257B | South Africa | B | |
| ZA201100089B | South Africa | B | |
| ZA201100090B | South Africa | B | |
| JP2011527444A | Japan | A | |
| JP2011527453A | Japan | A | |
| JP2011527454A | Japan | A | |
| JP2011527459A | Japan | A | |
| TW201142827A | Taiwan Province of China | A | |
| CO6351832A2 | Colombia | A2 | |
| CO6351833A2 | Colombia | A2 | |
| CO6351837A2 | Colombia | A2 | |
| ZA201102537B | South Africa | B | |
| CO6362072A2 | Colombia | A2 | |
| AU2009267467B2 | Australia | B2 | |
| JP2012505423A | Japan | A | |
| HK1155552A | Hong Kong, China | A | |
| HK1155552A1 | Hong Kong, China | A1 | |
| HK1156142A | Hong Kong, China | A | |
| HK1156142A1 | Hong Kong, China | A1 | |
| HK1157489A | Hong Kong, China | A | |
| HK1157489A1 | Hong Kong, China | A1 | |
| RU2010154747A | Russian Federation | A | |
| HK1158333A | Hong Kong, China | A | |
| HK1158333A1 | Hong Kong, China | A1 | |
| RU2011102422A | Russian Federation | A | |
| RU2011104003A | Russian Federation | A | |
| RU2011104004A | Russian Federation | A | |
| CN102105930B | China | B | |
| AU2009267394B2 | Australia | B2 | |
| RU2011117699A | Russian Federation | A | |
| KR101224559B1 | Republic of Korea | B1 | |
| KR101227729B1 | Republic of Korea | B1 | |
| AU2013200679A1 | Australia | A1 | |
| AU2013200680A1 | Australia | A1 | |
| CN102089811B | China | B | |
| US2013096930A1 | United States of America | A1 | |
| AU2009267466B2 | Australia | B2 | |
| US8447620B2 | United States of America | B2 | |
| RU2485606C2 | Russian Federation | C2 | |
| KR20130069833A | Republic of Korea | A |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG | |
| Transfer or rightsGB | GB | |
| Transfer or rightsGB | GB |
Numbers
- Application
- 2011000366
Titles2
- English
- AUDIO ENCODER AND DECODER FOR ENCODING AND DECODING AUDIO SAMPLES.
- Spanish
- CODIFICADOR Y DECODIFICADOR DE AUDIO PARA CODIFICAR Y DECODIFICAR MUESTRAS DE AUDIO.
Classification
- CPC, 7
- G10L19/00
- G10L19/20
- G10L19/18
- G10L19/022
- G10L19/02
- G10L19/04
- G10L19/12
- IPC, 2
- G10L19 02
- G10L19 14