Apparatus and method for downsampling an audio signal
Abstract
An apparatus for processing an audio signal to generate an extended bandwidth signal that has a high frequency part and a low frequency part that uses parametric data for the high frequency part, the parametric data relating to frequency bands of the high frequency part comprise a patch edge calculator (2302) to calculate a patch edge such that the patch edge coincides with a frequency band edge of the bands of frequency. The apparatus further comprises a patch (2312) to generate a patched signal using the audio signal (2300) and the patch edge.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
16 claims: 9 independent, 7 dependent
- 1Habiendo asi especialmente descrito y determinado Ia naturaleza de Ia presente invención y la forma còrno la misma ha de ser llevada a la practica se déclara reivindicar corno de propiedad y derecho exclusivo:5 1. Un aparato para procesar una serial de audio para generar una serial extendida en ancho de banda que tiene una parte de alta frecuencia (102) y una parte de baja frecuencia (104) usando datos paramétricos (2302) para la parte de alta frecuencia (102), los datos paramétricos que se relacionan con bandas de frecuencia (100, 101) de la parte de alta frecuencia (102), que comprende: 10 un calculador de borde de patching (2302) es decir de parche, para calcular un borde de parche (1001c, 1002c, 1002d, 1003c, 1003b) tal que el borde de parche coincide con un borde de banda de frecuencia de las bandas de frecuencia (101, 100);y un parcheador (2312) para generar una serial parcheada usando la serial 15 de audio (2300) y el borde de parche (1001c, 1002c, 1002b, 1003c, 1003b).
- 2Un aparato de acuerdo con la reivindìcación 1, en el cual el calculador de borde de parche (2302) està configurado para usar un borde de parche bianco (1001b, 1002a, 1002b, 1003a) que no coincide con un borde de banda de frecuencia de una banda de frecuencia (101), y 20 en el cual el calculador de borde de parche (2302) està configurado para poner el borde de parche diferente del borde de parche bianco.
- 3Un aparato de acuerdo con una cualquiera de las reivindicaciones 1 o 2, en el cual el calculador de borde de parche (2302) està configurado para calcular bordes de parche para très diferentes factores de transposición tal que cada 25 borde de parche coincida con un borde de banda de frecuencia (100, 101) de las bandas de frecuencia de la parte de alta frecuencia, y — 70 — en el cual el parcheador (2312) esta configurado para generar la senal parcheada usando los très factores de transposición diferentes (2308) de modo que un borde entre parches adyacentes coincide con un borde entre dos bandas de frecuencia adyacentes (100, 101).
- 4Un aparato de acuerdo con una cualquiera de las reivindicaciones precedentes, en el cual el calculador de borde de parche )2302) esta configurado para calcular el borde de parche corno un borde de frecuencia (k) en un rango de frecuencia de sintesis correspondiente a la parte de alta frecuencia (102), y en donde el parcheador (2312) esta configurado para seleccionar una porción de frecuencia de la parte de banda baja (104) usando un factor de transposición y el borde de parche.
- 5Un aparato de acuerdo con una cualquiera de las reivindicaciones precedentes, que ademâs comprende:un reconstructor de alta frecuencia (1030, 2510) para ajustar la senal parcheada (2509) usando los datos paramétricos (2302), estando configurado el reconstructor de alta frecuencia para calcular, para una banda de frecuencia o un grupo de bandas de frecuencia, un factor de ganancia a ser usado para ponderar la correspondiente banda de frecuencia o grupos de bandas de frecuencia de la senal parcheada (2509).
- 6Un aparato de acuerdo con una cualquiera de las reivindicaciones precedentes, en el cual el calculador de borde de parche (2302) està configurado para :calcular (2520) una tabla de frecuencia que define las bandas de frecuencia de la parte de alta frecuencia (102) usando los datos paramétricos o datos adicionales de entrada de configuración;determinar (2522) un borde de parche de sintesis bianco usando por lo menos un factor de transposición, — 71 — buscar (2524) en la tabla de frecuencia, una banda de frecuencia coincidente;y seleccionar (2525, 2527) la banda de frecuencia coincidente corno el borde de parche. 5
- 7Un aparato de acuerdo con la reivindicación 6, en el cual el calculador de borde de parche està configurado para buscar, en la tabla de frecuencia, una banda de frecuencia coincidente que tiene un borde coincidente que coincide con el borde de la frecuencia bianco dentro de un rango de coincidencia predeterminado, o para buscar la banda de frecuencia que tiene un bode de 10 banda de frecuencia que es el mâs cercano al borde de frecuencia bianco.
- 8Un aparato de acuerdo con la reivindicación 7, en el cual el rango de coincidencia predeterminado se fija a un valor mäs pequefio o igual que cinco bandas QMF o 40 bandejas de frecuencia de la parte de alta frecuencia (102).
- 9Un aparato de acuerdo con una cualquiera de las reivindicaciones 15 precedentes, en el cual los datos paramétricos comprenden un valor de datos de envolvente espectral, en donde para cada banda de frecuencia se da un valor de datos de envolvente espectral separado, en donde el aparato ademâs comprende un reconstructor de alta frecuencia (2510, 1030) para ajuste de envolvente espectral de cada banda de la serial parcheada usando el valor de datos de 20 envolvente espectral para esta banda.
- 10Un aparato de acuerdo con una cualquiera de las reivindicaciones precedentes, en el cual el calculador de borde de parche (2302) està configurado para buscar el borde mâs alto en la tabla de frecuencia, que no excede un limite de ancho de banda de una serial regenerada de alta frecuencia por un factor de 25 transposición, y para usar el borde mâs alto hallado, corno el borde de parche.
- 11Un aparato de acuerdo con la reivindicación 10, en el cual el calculador de borde de parche (2302) està configurado para recibir, para cada factor de — 72 — transposición de la pluralidad de diferentes factores de transposición, un borde de parche bianco diferente.
- 12Un aparato de acuerdo con una cualquiera de las reivindicaciones precedentes, que ademâs comprende una herramienta de software de limitador 5 (2505, 2510) para calcular bandas de limitador usadas para limitar valores de ganancia para ajustar las senales parcheadas, el aparato que ademâs comprende un calculador de banda de limitador configurado para fijar un borde de limitador de modo que por lo menos un borde de parche determinado por el calculador de borde de parche (2302) sea fijado corno un borde de limitador también. 10
- 13Un aparato de acuerdo con la reivindicación 12, en el cual el calculador de banda de limitador (2505) està configurado para calcular ademâs bordes de limitador de modo que otros bordes de limitador coincidan con bordes de banda de frecuencia de las bandas de frecuencia de la parte de alta frecuencia (102).
- 14Un aparato de acuerdo con cualquiera de las reivindicaciones
- 1515 precedentes, en el cual el parcheador (2312) està configurado para generar mùltiples parches usando diferentes factores de transposición (2308), en el cual el calculador de borde de parche (2302) està configurado para calcular los bordes de parche de cada parche de los mùltiples parches de modo que los bordes de parche coinciden con diferentes bordes de banda de frecuencia 20 de las bandas de frecuencia de la parte de alta frecuencia (102), en donde el aparato ademàs comprende un ajustador de envolvente (2510) para ajustar una envolvente de la parte de alta frecuencia (102) después de parchear o para ajustar la parte de alta frecuencia antes de parchear usando factores de escala incluidos en los datos paramétricos dados para bandas de 25 factor de escala. 15. Un mètodo de procesar una senal de audio para generar una senal extendida en ancho de banda que tiene una parte de alta frecuencia (102) y una — 73 parte de baja frecuencia (104) usando datos paramétricos (2302) para la parte de alta frecuencia (102), los datos paramétricos que se relacionan con bandas de frecuencia (100, 101) de la parte de alta frecuencia (102), que comprende:calcular (2302) un borde de parche (1001c, 1002c, 1002d, 1003c, 1003b) 5 tal que el borde de parche coincide con un borde de banda de frecuencia de las bandas de frecuencia (101, 100);y generar (2312) una senal parcheada usando la serial de audio (2300) y el borde de parche (1001c, 1002c, 1002b, 1003c, 1003b).
- 16Un programa de computadora que tiene un código de programa para 10 ejecutar cuando corre en una computadora, el mètodo de la reivindicación 15. p.de:FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER
Independent claims16
348 paragraphs in 9 sections, as filed
The present invention relates to audio source coding systems which make use of a harmonic transposition method for high frequency reconstruction (HFR), and digital effect processors, for example, the so-called exciters, where the generation of distortion Harmonica adds brightness to the processed serial, and to time extenders, where the duration of a serial is extended while maintaining the spectral content of the originai.
BACKGROUND OF THE INVENTION
In PCT WO 98/57436 the concept of transposition was established as a method to recreate a high frequency band from a lower frequency band of an audio signal. A substantial saving in the amount of bits transmitted can be obtained using this concept for audio coding. In an HFR-based audio coding system, a low bandwidth signal is processed by one encoder per waveform core and the higher frequencies are regenerated using transposition and additional side information of very low amount of transmitted bits that Describe the white spectral shape of the decoder side. For low amounts of transmitted bits, where the bandwidth of the signal coded by a core (core coded) is narrow, it becomes increasingly important to recreate a high band with perceptually pleasing characteristics. The transposition - 2 -
<img file="AR080476A1_D0001.tif" />
PCT defined harmonic WO 98/57436 works very well for complex musical material in a situation with low transition frequency. The principle of a harmonic transposition is that a sinusoid with frequency ω is mapped to a sinusoid with frequencyΤω where T> \ is an integer that defines the order of transposition. In contrast to this, an HFR method based on simple sideband modulation (SSB) maps a sinusoid with frequency ω to a sinusoid with frequency ω + Δω where Δω is a fixed frequency shift. Given a serial core with low bandwidth, an artifact that sounds dissonant from the SSB transposition may result.
To achieve the best possible audio quality, high-quality harmonic HFR methods of the current state of the art employ complex modulated filter banks, for example, a Short Time Fourier Transformation (STFT), with high frequency resolution and a high degree of oversampling to achieve the required audio quality. Fine resolution is necessary to avoid unwanted intermodulation distortion that appears from nonlinear processing of sinus sums. With sufficiently high frequency resolution, that is, narrow subbands, high quality methods aim to have a maximum of one sinusoid in each subband. A high degree of over-sampling in time is needed to avoid a type of aliasing distortion, and a certain degree of over-sampling in the frequency is needed to avoid pre-echoes for serial with transient component. The obvious disadvantage is that computational complexity can be made high.
Subsonic block-based harmonic transposition is another HFR method used to suppress intermodulation products, in which case a filter bank with coarse frequency resolution and a lower degree of oversampling is used, for example, a multichannel QMF bank. In this method, a time block of complex subband samples is processed by a - 3 -
<img file="AR080476A1_D0002.tif" />
Common phase modifier while the superposition of several modified samples forms an output subband sample. This has the net effect of suppressing intermodulation products that would otherwise appear when the input subband signal consists of several sinusoids. Transposition based on block-based subband processing has much less computational complexity than high quality transposition media and achieves almost the same quality for many signals. However, the complexity is still much higher than for trivial SSB-based HFR methods, since a plurality of analysis filter banks are required, each processing signals of different transposition orders T, in a typical application of HFR to synthesize the required bandwidth. Additionally, a common approach is to adapt the sampling rate of the input signals to adjust analysis filter banks of a constant size, although the filter banks process signals of different transposition orders. It is also common to apply bandpass filters to the input signals to obtain output signals processed from different transposition orders, with spectral densities that do not overlap.
The storage and transmission of audio signals are often subject to strict restrictions on the number of bits transmitted. In the past, encoders were forced to drastically reduce the transmitted audio bandwidth when only a very low amount of transmitted bits was available. Modern encoders — audio decoders today are capable of encoding broadband signals using bandwidth extension (BWE) methods [1-12]. These algorithms are based on a parametric representation of the high frequency (HF) content, which is generated from the low frequency (LF) part of the decoded signal by transposition within the spectral region of HF (patching, is - 4 -
<img file="AR080476A1_D0003.tif" />
say the “patches” of audio) and application of a subsequent processing governed by parameters. The LF part is encoded with any audio or voice encoder. For example, the bandwidth extension methods described in [1-4] are supported by simple sideband modulation (SSB), which is often referred to as the term copy-up method, to generate multiple patching sectors , ie the patches of HF.
Recently, a new algorithm has been presented, which uses a bank of phase Vocoders [15-17] for the generation of the different patches [13] (see Figure 20). This method has been developed to avoid the auditory harshness that is frequently observed in signals subjected to SSB bandwidth extension. Despite being beneficial for many tonal signals, this method called harmonic bandwidth extension (HBE) is prone to quality degradation of the transient components contained in the audio signal [14], since the conservation of the vertical coherence on subbands in the standard phase Vocoder algorithm and, likewise, the recalculation of the phases has to be performed on time blocks of a transformation or, alternatively from a bank of filters. Therefore, there is a need for a special treatment for serial parts that contain transient components.
However, computational complexity is a serious matter, because the BWE algorithm is performed on the decoder side of an encoder-decoder chain. The methods of the current state of the art, especially the HBE based on phase Vocoder comes at the cost of a greatly increased computational complexity compared to the SSB-based methods.
As detailed above, existing bandwidth extension schemes apply only a patch method on a serial block given to - 5 -
<img file="AR080476A1_D0004.tif" />
At the same time, either SSB-based patching [1-4] or HBE Vocoder-based patching [15-17]. Additionally, modern audio encoders [19-20] offer the possibility of switching the patch method globally over a time block base between alternative patching schemes.
The SSB copy-up patch introduces unwanted roughness into the audio serial, but is computationally simple and retains the time envelope of transient components. In encoders — audio decoders that use HBE patching, the playback quality of the transient component is often below optimal. Thus, computational complexity is significantly increased over the SSB copy-up method of very simple computational complexity.
In terms of complexity reduction, sampling rates are of particular importance. This is due to the fact that a high sampling rate means high complexity and a low sampling rate generally means low complexity due to the small number of operations required. On the one hand, however, the situation in bandwidth extension applications is particularly such that the sampling rate of the coder output serial per core will typically be so low that this sampling rate is too low for a serial of full bandwidth In other words, when the sampling rate of the decoder output serial is, for example, 1 or 2.5 times the maximum frequency of the encoder output serial per core, then an extension of bandwidth, per example, by a factor 2, It means that an upward sampling operation is required so that the bandwidth extended serial sampling rate is so high that the sampling can cover the additionally generated high frequency components.
<img file="AR080476A1_D0005.tif" />
Additionally, filter banks such as analysis filter banks and synthesis filter banks are responsible for a considerable amount of processing operations. Therefore, the size of the filter banks, that is, if the filter bank is a 32-channel filter bank, a 64-channel filter bank or even a filter bank with a larger number of channels, will significantly influence in the complexity of the audio processing algorithm. In general, one could say that a high number of filter bank channels requires more processing operations and, therefore, more complexity than a small number of filter bank channels. In view of this, in bandwidth extension applications and also in other audio processing applications, where different sampling rates are a subject, such as in Vocoder type applications or any other audio effect application, there are a specific interdependence between complexity and sampling rate or audio bandwidth, which means that sampling operations up or subband filtering can drastically improve complexity without specifically influencing audio quality in a good way when inappropriate algorithms or software tools are chosen for specific operations.
In the context of bandwidth extension, parametric data sets are used to perform a spectral envelope adjustment and to perform other manipulations to a signal generated by a patching operation, that is, by an operation that takes some data from the range source, this is from the low band portion of the extended bandwidth serial that is available at the input of the bandwidth extension processor and then maps this data to a high frequency range. Spectral envelope adjustment can take place before actually mapping the low band serial to
<img file="AR080476A1_D0006.tif" />
high frequency range or subsequently having mapped the range of the source to the high frequency range.
Typically, parametric data sets are provided with a certain frequency resolution, that is, parametric data refers to frequency bands of the high frequency part. On the other hand, patching from the low band to the high band, that is, which source ranges are used to obtain which bianco or high frequency ranges, is an independent operation of the resolution, in which the parametric data sets They are given with respect to frequency. The fact that the transmitted parametric data is, in a sense, independent of what is actually used as a patching algorithm, is an important feature, since this allows great flexibility on the decoder side, that is, when it comes to the Bandwidth extension processor implementation. Different patching algorithms can be used here but one and the same spectral envelope setting must be performed. In other words, the high frequency reconstruction processor or the spectral envelope adjustment processor in a bandwidth extension application does not need to have information about the patching algorithm applied to perform the spectral envelope adjustment.
A disadvantage of this procedure, however, is that poor alignment can occur between the frequency bands for which the parametric data sets are provided on the one hand, and the spectral edges of one patch on the other. Particularly in situations where the spectral energy changes greatly in the vicinity of a patch edge, artifacts may appear, specifically in this region, which degrade the quality of the signal extended in bandwidth.
SYNTHESIS OF THE INVENTION - 8 -
<img file="AR080476A1_D0007.tif" />
It is an objective of the present invention to provide an improved concept of audio processing that allows for good audio quality.
This objective is achieved with an apparatus for processing an audio serial according to claim 1, a method for processing a high frequency audio serial according to claim 15 or a computer program according to claim 16.
Embodiments of the present invention relate to an apparatus for processing an audio serial to generate an extended serial bandwidth having a high frequency portion and a low frequency portion, where parametric data is used for the high frequency portion, and where the parametric data is related to frequency bands of the high frequency part. The apparatus comprises a patch edge calculator for calculating a patch edge such that the patch edge coincides with a frequency band edge of the frequency bands. The apparatus also comprises a patch to generate a serial patch using the audio serial and the calculated patch edge. In one embodiment, the patch edge calculator is configured to calculate the patch edge as a frequency edge in a synthetic frequency range corresponding to the high frequency portion. In this context, the patch is configured to select a frequency portion of the low band portion using a transposition factor and the patch edge. In another embodiment, the patch edge calculator is configured to calculate the patch edge using a white patch edge that does not match a frequency band edge of the frequency band. Then, the patch edge calculator is set to adjust the patch edge different from the white patch edge to obtain alignment. Particularly in the context of a plurality of patches that use different transposition factors, the patch edge calculator is set to - 9 -
<img file="AR080476A1_D0008.tif" />
calculate patch edges, for example, for three different transposition factors such that each patch edge coincides with a frequency band edge of the frequency bands of the high frequency part. The patch is then configured to generate the patch signal using three different transposition factors such that the border between two adjacent patches coincides with an edge between two adjacent frequency bands to which the parametric data is related.
The present invention is particularly useful because of the avoidance of artifacts that appear along the edges of badly aligned patches on the one hand, and on the other, the frequency bands for parametric data. On the other hand, due to the perfect alignment, even strongly changing serials or signals that have strongly changing portions in the patch edge region, they are subjected to bandwidth extension with good quality.
Also, the present invention is advantageous in that it nevertheless allows high flexibility due to the fact that the encoder does not have to deal with a patching algorithm to be applied on the decoder side. The independence between patching on one side and spectral envelope adjustment is maintained, that is, using the parametric data generated by the bandwidth extension encoder on the other, and allows the application of different patching algorithms or even a combination of different patching algorithms. This is possible since the patch edge alignment ensures that at the end the patch data on one side and the parametric data sets on the other, match each other with respect to the frequency bands, which are also called bands of scale factor.
Depending on the calculated patch edges, which can be related, for example, to the white range, that is, the high frequency stop of the extended serial in finally extended bandwidth, the - 10 -
<img file="AR080476A1_D0009.tif" />
<img file="AR080476A1_D0010.tif" />
Corresponding source ranges to determine the patch source data form the low band portion of the audio serial. It turns out that only a certain (small) bandwidth of the low band portion of the audio serial is required due to the fact that in some embodiments harmonic transposition factors are applied. Therefore, in order to efficiently extract this portion of the low band audio serial, a specific analysis filter bank structure that relies on cascading individual filter banks is used.
Such embodiments are based on a specific cascade location of analysis and / or synthesis filter banks to obtain low complexity re-sampling without sacrificing audio quality. In one embodiment, an apparatus for processing an input audio serial comprises a bank of synthetic filters to synthesize an intermediate audio serial from the input audio serial, where the input audio serial is represented by a plurality of first subband serials generated by a bank of analysis filters placed in the processing direction before the synthesis filter bank, where a number of filter bank channels of the synthesis filter bank is smaller than a number of channels of the analysis filter bank. The intermediate serial is also processed by an additional analysis filter bank to generate a plurality of second subband serial from the intermediate audio serial, wherein the additional analysis filter bank has a number of channels that is different from the number of channels of the synthesis filter bank so that a sampling rate of a serial subband of the plurality of subband serials is different from a rate of Sampling of a first serial subband of the plurality of first subband serials generated by the bank of analysis filters.
The cascade of a bank of synthesis filters and a bank of additional analysis filters connected subsequently provides a conversion rate of - 11 -
<img file="AR080476A1_D0011.tif" />
sampling and additionally a modulation of the bandwidth portion of the original input audio serial that has been entered into the synthesis filter bank to a baseband. This intermediate time serial, which has now been extracted from the original input audio serial which may be, for example, the output serial of a decoder per core of a bandwidth extension scheme, is now preferably represented as a sampled serial critically modulated to the baseband, and this representation was found, that is, the re-sampled output serial, when it is being processed by an additional analysis filter bank to obtain a subband representation, it allows a low complexity processing of additional processing operations that may or may not occur, and which may be, for example, processing operations related to taies bandwidth extension as non-linear subband operations followed by high frequency reconstruction processing and by a fusion of subbands in the final synthesis filter bank.
The present application provides different aspects of devices, methods or computer programs for processing audio serials in the context of bandwidth extension and in the context of other audio applications, which are not related to bandwidth extension. The features of the individual aspects subsequently described and claimed can be partially or totally combined, but they can also be used separately from each other, since the individual aspects already provide advantages with respect to perceptual quality, computational complexity and processor resources. / memory when implemented in a computer or microprocessor system.
Embodiments provide a method to reduce the computational complexity of a harmonic HFR method based on subband block by - 12 -
<img file="AR080476A1_D0012.tif" />
means of efficient filtering and conversion of sampling rate of the input signals to the stages of HFR filter bank analysis. In addition, it can be shown that the passband filters applied to the input signals are obsolete in a transposition media based on a subband block.
The present embodiments help reduce the computational complexity of the subband block-based harmonic transposition by efficiently implementing several subband block-based transposition orders within the framework of a single pair of bank analysis and synthesis filters. Depending on the perceptual quality compromise solution based on computational complexity, only a subset of orders or all transposition orders can be performed together within a pair of filter banks. Also, a combined transposition scheme where only certain transposition orders are calculated directly while the remaining bandwidth is filtered by replicating available transposition orders, that is, previously calculated (for example, 2nd order) and / or the width of band coded by nucleus. In this case, patching can be carried out using any conceivable combination of source ranges available for replication.
Additionally, there are embodiments that provide a method for improving both high quality harmonic HFR methods as well as subband block based harmonic HFR methods by means of spectral alignment of HFR tools. In particular, greater performance is achieved by aligning the spectral edges of the signals generated by HFR to spectral edges of the envelope adjustment frequency table. In addition, the spectral edges of the limiting tool are aligned by the same principle to the spectral edges of the signals generated by HFR.
— 13 —
<img file="AR080476A1_D0013.tif" />
Other embodiments are configured to improve the perceptual quality of transient components while reducing computational complexity, for example, by applying a patching scheme that applies a mixed patch consisting of harmonic patching and copy-up patching.
In specific embodiments, the individual filter banks of the cascade filter bank structure are quadrature mirror filter banks (QMF), which are based on a low-pass prototype filter or modulated window using a set of modulation frequencies that define the center frequencies of the filter bank channels. Preferably, all the prototype window or filter functions depend on each other in such a way that the filters of the filter banks are different sizes (filter bank channels) also depend on each other. Preferably, the larger filter bank is a cascading structure of filter banks comprising, in embodiments, a first analysis filter bank, a subsequent connected filter bank, an additional analysis filter bank and in some subsequent state. Processing, a bank of final synthesis filters, has a window function or prototype filter response that has a certain number of window or prototype filter function coefficients. The smaller filter banks are all sub-sampled versions of this window function, which means that the window functions for the other filter banks are sub-sampled versions of the large window function. For example, if a filter bank has half the size of the large filter bank, then the window function has half the number of coefficients, and the coefficients of smaller size filter banks are derived by subsampling. In this situation, subsampling means that, for example, for the smallest filter bank that is half the size, every second filter coefficient is taken. However, when there are other relationships - 14 -
<img file="AR080476A1_D0014.tif" />
between filter bank sizes that are not integer values, then, a certain type of interpolation of window coefficients is performed so that in the end the window of the smallest filter bank again is a subsampled version of the window of the largest filter bank.
Embodiments of the present invention are particularly useful in situations where only a portion of the input audio serial is required for further processing, and this situation appears particularly in the context of harmonic bandwidth extension. In this context, Vocoder-type processing operations are particularly preferred.
It is an advantage in some embodiments that the embodiments provide less complexity for a means of transposition of QMF through efficient time and frequency domain operations and better audio quality for harmonic spectral band replication based on QMF and DFT using spectral alignment.
Some embodiments relate to audio source encoding systems that employ, for example, a harmonic transposition method based on subband block for high frequency reconstruction (HFR), and digital effect processors, for example, the so-called exciters, where Harmonic distortion generation adds brightness to the processed signal, and to time extenders, where the duration of a signal is extended while maintaining the original spectral content. Embodiments provide a method to reduce the computational complexity of a subsonic block-based harmonic HFR method by means of efficient filtering and conversion of sampling rate of the input signals before the HFR filter bank analysis stages. In addition, there are embodiments that show that conventional bandpass filters applied to the input signals are obsolete in a subband block based HFR system. Additionally, there are - 15 -
<img file="AR080476A1_D0015.tif" />
embodiments that provide a method to improve both high quality harmonic HFR methods as well as subband block harmonic HFR methods by means of spectral alignment of HFR tools. In particular, there are embodiments that show that greater performance is achieved by aligning the spectral edges of the serials generated by HFR to spectral edges of the envelope adjustment frequency table. In addition, the spectral edges of the limiting tool are aligned by the same principle to the spectral edges of the serials generated by HFR.
BRIEF DESCRIPTION OF THE DRAWINGS
The present invention will now be described by way of illustrative examples, without limiting the scope of the invention, with reference to the accompanying drawings, in which:
Figure 1 illustrates the operation of a block-based transposition medium using transposition orders 1, 3 and 4 in an improved HFR decoder frame;
Figure 2 illustrates the operation of the non-line subband stretch units of Figure 1;
Figure 3 illustrates an efficient implementation of the block-based transposition medium of Figure 1, where the re-samplers and bandpass filters that precede the HFR analysis filter side are implemented using multi-rate time domain samplers and QMF based bandpass filters;
Figure 4 illustrates an example of building blocks for efficient implementation of a multi-rate time domain resampler of Figure 3;
<img file="AR080476A1_D0016.tif" />
Figures 5a-5f illustrate the effect of an exemplary serial processed by the different blocks of Figure 4 for a transposition order of 2;
Figure 6 illustrates an efficient implementation of the block-based transposition medium of Figure 1, where the re-samplers and bandpass filters that precede the HFR analysis filter side are replaced by small banks of sub-sampled synthesis filters that they operate on selected subbands of a 32-band analysis filter bank;
Figure 7 illustrates the effect of an exemplary serial processed by a subsampled synthesis filter bank of Figure 6 for a transposition order of 2;
Figures 8a-8f illustrate the implementation blocks of a multi-rate time domain sampling rate reducer of two a factor 2;
Figures 9a-9e illustrate the implementation blocks of a sampling rate reducer or multi-rate efficient time domain, of two a 3/2 factor;
Figure 10 illustrates the alignment of spectral edges of the serials of the HFR transposition medium to the edges of the envelope adjustment frequency bands in an enhanced HFR encoder;
Figure 11 illustrates a scenario where artifacts emerge due to spectral edges of the serial misaligned HFR transposition media;
Figure 12 illustrates a scenario where the artifacts of Figure 11 are avoided as a result of spectral edges of the aligned HFR transposition media serials;
— 17 —
<img file="AR080476A1_D0017.tif" />
<img file="AR080476A1_D0018.tif" />
Figure 13 illustrates the adaptation of spectral edges in the limiting tool to the spectral edges of the signals of the HFR transposition medium;
Figure 14 illustrates the principle of harmonic transposition based on subband block;
Figure 15 illustrates an example scenario for the application of subband block-based transposition using several transposition orders in an encoder — enhanced HFR audio decoder;
Figure 16 illustrates an example scenario of the prior art for the operation of a multi-order subband block based transposition by applying a separate filter filter bank for each transposition order;
Figure 17 illustrates an inventive example scenario for the efficient operation of a multi-order subband block-based transposition applying a single bank of 64-band QMF analysis filters;
Figure 18 illustrates another example for forming a subband signal processing;
Figure 19 illustrates a single sideband modulation (SSB) patch Figure 20 illustrates a harmonic bandwidth extension (HBE) patch Figure 21 illustrates a mixed patch, where the first patch is generated by spreading frequency and the second patch is generated by SSB copy-up of a low frequency portion;
Figure 22 illustrates an alternative mixed patch using the first HBE patch for an SSB copy-up operation to generate a second patch;
— 18
<img file="AR080476A1_D0019.tif" />
Figure 23 illustrates an overview of an apparatus for processing an audio signal using spectral band alignment in accordance with one embodiment;
Figure 24a illustrates a preferred implementation of the patch edge calculator of Figure 23.
Figure 24b illustrates another overview of a sequence of steps performed by embodiments of the invention;
Figure 25a illustrates a block diagram illustrating more details of the patch edge calculator and more details about the spectral envelope setting in the context of patch edge alignment;
Figure 25b illustrates a logical diagram for the procedure indicated in Figure 24a as a pseudo code;
Figure 26 illustrates an overview of the framework in the context of bandwidth extension processing; and Figure 27 illustrates a preferred implementation of a subband signal processing delivered by the additional analysis filter bank of Figure 23.
DESCRIPTION OF PREFERRED EMBODIMENTS
The embodiments described below are merely illustrative and may provide a lower complexity of a QMF transposition medium by efficient operations in the time and frequency domain, and improved audio quality of both, harmonic SBR based QMF and DFT, by spectral alignment. It is understood that the possible modifications and variations of the provisions and details described herein will be apparent to those skilled in the art. Therefore, it is the intention that the invention be limited only by the scope of the following claims of - 19 -
<img file="AR080476A1_D0020.tif" />
<img file="AR080476A1_D0021.tif" />
patent and not for the specific details presented by the description and explanation of the embodiments herein.
Figure 23 illustrates an embodiment of an apparatus for processing an audio serial 2300 to generate an extended serial bandwidth having a high frequency part and a low frequency part, using parametric data for the high frequency part, where Parametric data is related to frequency bands of the high frequency part. The apparatus comprises a patch edge calculator 2302 for calculating a patch edge preferably using a bianco patch edge 2304 that does not match a frequency band edge of the frequency band. Information 2306 on the frequency bands of the high frequency part can be taken, for example, from an encoded data transmission suitable for bandwidth extension. In a further embodiment, the patch edge calculation not only calculates a single patch edge for a single patch but also calculates several patch edges for a several different patches belonging to different transposition factors, where information on transposition factors they are provided to patch edge calculator 2302 as indicated in 2308. The patch edge calculator is configured to calculate the patch edges so that a patch edge matches a frequency band edge of the frequency bands. Preferably, when the patch edge calculator receives information 2304 about a white patch edge, then the patch edge calculator is configured to set the patch edge different from the white patch edge to obtain alignment. The patch edge calculator delivers the calculated patch edges, which are different from the bianco patch edges, on line 2310 to a patch 2312. Patch 2312 generates a patched serial or several patched serials at output 2314 using the low band audio serial - 20 -
<img file="AR080476A1_D0022.tif" />
2300 and patch edges 2310, and in the embodiments where multiple transpositions are performed, using transposition factors on line 2308.
The table in Figure 23 illustrates a numerical example to illustrate the basic concept. For example, when it is assumed that the low-band audio signal has a low frequency portion that extends from 0 to 4 kHz (that is, the source range does not really start at 0 Hz but near 0, such as at 20 Hz) It is also the intention of the user to carry out a 4 kHz serial bandwidth extension to an extended serial bandwidth of 16 kHz. Additionally, the user has indicated that the user wishes to make a bandwidth extension using three harmonic patches with transposition factors 2, 3 and 4. Then, the white edges of the patches can be set to a first patch that extends from 4 to 8 kHz, a second patch that extends from 8 to 12 kHz and a third patch that extends from 12 to 16 kHz. Thus, the patch edges are 8, 12 and 16 when it is assumed that the first patch edge that matches the maximum or transition frequency of the low frequency band serial is not changed. However, changing this edge of the first patch is also within the embodiments of the present invention if required. The white edges would correspond to a source range of 2 to 4 kHz for the transposition factor 1, 2.66 to 4 kHz for the transposition factor of 3, and 3 to 4 kHz for the transposition factor of 4. Specifically, the Source range is calculated by dividing the edges bianco by the transposition factor actually used.
For the example of Figure 23 it is assumed that the edges 8, 12, 16 do not match the frequency band edges of the frequency bands to which the parametric input data are related. For den, the patch edge calculator calculates aligned patch edges and does not apply - 21 -
<img file="AR080476A1_D0023.tif" />
immediately the edges bianco. This may result in an upper patch edge of 7.7 kHz for the first patch, an upper edge of 11.9 kHz for the second patch and 15.8 kHz for the upper edge for the third patch. Then, using the transposition factor again for the individual patch, certain adjusted source ranges are calculated and used for patching, which are indicated in Figure 23 in exemplary form.
Although it has been expressed that the source ranges are changed along with the bianco ranges, for other implementations one could also manipulate the transposition factor and maintain the source range or the edges bianco or for other applications one could even change the source range and the transposition factor to finally up to adjusted patch edges which coincide with edges of frequency band of the frequency bands to which the parametric bandwidth extension data describing the spectral envelope of the portion of the high band of the originai serial.
Figure 14 illustrates the principle of transposition based on subband block. The serial of the input time domain is fed to a bank of analysis filters 1410 which provides a multitude of serial subbands of complex value. These are fed to the subband processing unit 1402. The multitude of complex value output subbands is fed to the synthesis filter bank 1403, which in turn delivers the serial of the modified time domain. Subband processing unit 1402 performs subband processing operations based on a nonlinear block such that the modified time domain serial is a transposed version of the input serial corresponding to a transposition order T> \. The notion of a block-based subband processing is defined as comprising nonlinear operations on blocks of more than a subband sample at a - 22
<img file="AR080476A1_D0024.tif" />
time, where the subsequent blocks are windows and added in superposition to generate the output subband signals.
The filter banks 1401 and 1403 can be any complex exponential modulated type such as QMF horn or a windowed DFT. They can be stacked evenly or oddly in modulation and can be defined from a wide range of prototype filters or windows. It is important to know the ratio δ / ^ / δ /) of the next two filter bank parameters, measured in physical units.
• & f<sub>TO</sub> : the frequency subband spacing of the analysis filter bank 1401;
• \ f<sub>s</sub> : the frequency subband spacing of the synthesis filter bank 1403.
For the configuration of subband processing 1402 it is necessary to find the correspondence between subband source and bianco indices. It is noted that a physical frequency input sinusoid Ω will result in a major contribution occurring in input subbands with index "~ Ω / δ /<sub>4</sub> . An output sinusoid of the desired transposed physical frequency 7 · Ω will result from feeding the synthesis subband with index m * T-Ω / Hf<sub>s</sub>. Therefore, the appropriate values of the subband source index of I subband processing split a bianco subband index die m must obey
A4
A4
<img file="AR080476A1_D0025.tif" />
(D
Figure 15 illustrates an example scenario for the application of subband block-based transposition using several transposition orders in an encoder — enhanced HFR audio decoder. A series of time bits transmitted is received in the decoder per core 1501, which provides a decoded core signal of low bandwidth to -
<img file="AR080476A1_D0026.tif" />
a sampling frequency fs. The low frequency is re-sampled to the 2fs output sampling frequency by means of a complex modulated QMF analysis band 1502 followed by a 64-band QMF synthesis bank (reverse QMF) 1505. The two filter banks 1502 and 1505 have the same physical resolution parameters & f<sub>s</sub> = kf<sub>TO</sub> and the processing unit of HFR 1504 simply passes the unmodified lower subbands corresponding to the low bandwidth core signal. The high frequency content of the output signal is obtained by feeding the upper subbands of the 64-band QMF synthesis bank 1505 with the output bands of the multiple transponder unit 1503, subjected to spectral modeling and modification performed by the processing unit of HFR 1504. Multiple transponder 1503 takes as input the decoded core serial and delivers a multitude of subband signals which represent the OMF band analysis of 64 of an overlay or combination of several transposed serial components. The objective is that if the HFR processing is skipped, each component corresponds to an entire physical transposition of the core series, (7 = 2,3, ...).
Figure 16 illustrates an example scenario of the prior art for the operation of a multi-order subband block based transposition
1603 applying a bank of analysis filters separated by each transposition order. Here, three transposition orders 7 = 2,3,4 have to be produced and supplied in the domain of a 64-band QMF that operates at the 2fs output sampling rate. Fusion unit 1604 simply selects and combines the relevant subbands of each transposition factor branch into a single multitude of QMF subbands to be fed into the HFR processing unit.
— 24 —
<img file="AR080476A1_D0027.tif" />
Let me first consider case T = 2. The objective is specifically that the processing chain of a 64-band QMF analysis 1602-2, a sub-band processing unit 1603-2 and a 64-band QMF synthesis 1505 resurface in a transposition physics of 7 = 2. Identifying these three blocks with 1401, 1402 and 1403 of Figure 14, one finds that and Ef<sub>s</sub>INf<sub>TO</sub> = 2 such that (1) results in the specification for 1603-2 that the correspondence between source subbands «and bianco m is given by n = m.
For case 7 = 3, the exemplary system includes a sampling rate converter 1601-3 which reduces the input sampling rate by a factor 3/2 from fs to 2fs / 3. The objective is specifically that the processing chain of a 64-band QMF analysis 1602-3, the subband processing unit 1603-3 and a synthesis of 64-band QMF 1505 resurface in a physical transposition of 7 = 3. Identifying these three blocks with 1401, 1402 and 1403 of Figure 14, one finds due to re-sampling, that Nf<sub>s</sub>/ Nf<sub>TO</sub> = 3, such that (1) provides the specification for 1603—3, where the correspondence between source subbands «and bianco m is again given by n = m.
For case 7 = 4, the exemplary system includes a sampling rate converter 1601-4 which reduces the input sampling rate by a factor of two from fs to fs / 2. The objective is specifically that the processing chain of a 64-band QMF analysis 1602-4, the sub-band processing unit 1603-4 and a 64-band QMF synthesis 1505 resurface in a physical transposition of 7 = 4. Identifying these three blocks with 1401, 1402 and 1403 of Figure 14, one finds due to re-sampling, which bf<sub>s</sub>ltrf<sub>TO</sub> = 4, such that (1) provides the specification for 1603-4, where the correspondence between source subbands «and bianco m is also given by - 25 -
<img file="AR080476A1_D0028.tif" />
Figure 17 illustrates an inventive example scenario for the efficient operation of a multi-order subband block based transposition applying a single bank of 64-band QMF analysis filters. In fact, the use of three separate QMF analysis banks and two sample rate converters in Figure 16 results in quite high computational complexity, as well as some disadvantages of implementation by table-based processing due to rate conversion Sample 1601—3. Current embodiments would replace both branches 1601—3 -> 1602—3 - * 1603—3 and 1601—4 -> 1602—4 - »1603—4 with subband processing 1703—3 and 1703—4, respectively, while branch 1602—2 -> 1603—2 remains unchanged compared to Figure 16. The three transposition orders will now have to be performed in a filter bank domain with reference to Figure 14, where Nf<sub>s</sub>/ Nf<sub>TO</sub> = 2. For case T = 3, the specification for 1703—3 given by (1) is that the correspondence between source subbands “and bianco west” given by n ~ 2m / 3. For case 7 = 4, the specifications for 1703-4 given by (1) are that the correspondence between source subbands "and bianco west" given by n ~ 2m. To further reduce complexity, some transposition orders can be generated by copying already calculated transposition orders or decoder output per core.
Figure 1 illustrates the operation of a subband block based transposition medium using transposition orders of 2, 3, and 4 in an improved HFR decoder framework, such as SBR [ISO / IEC 14496-3: 2009, Information technology - Coding of audio - visual objects - Part 3: Audio]. The series of bits in time (bitstream) is decoded to the time domain by decoding by kernel 101 and is passed to the module of - 26 -
<img file="AR080476A1_D0029.tif" />
HFR 103, which generates a high frequency signal from the serial baseband core. After generation, the serial generated by the HFR is dynamically adjusted to match the original serial as closely as possible by means of transmitted side information. This adjustment is made by the HFR 105 processor on subband serials, obtained from one or several QMF analysis banks. A typical scenario is where the decoder per core works on a serial of the time domain sampled at half the frequency of the input and output serials, this is the HFR decoder module will effectively re-sample the core serial at twice the sampling frequency This sampling rate conversion is usually obtained by the first step of filtering the serial encoder by core through a 32-band QMF analysis bank, 102. The subbands below the so-called transition frequency, this is the lower subset of the 32 subbands that contain all the coder serial energy per core, are combined with the set of subbands that carry the serial generated in the HFR. Usually, the number of subbands thus combined is 64, which, after filtering through the synthesis bank QMF 106, results in a serial encoder per core of the converted sampling rate combined with the output of the HFR module.
In the transposition medium based on subband block of the module
HFR 103, three transposition orders T = 2, 3 and 4 have to be produced and supplied, in the domain of a 64-band QMF that operates at an output sampling rate of 2fs. The serial of the entry time domain is filtered passband in blocks 103-12, 103-13 and 103-14. This is done to make the output serials processed by the different transposition orders, to have spectral contents that do not overlap. The sampling rate of the serials is reduced further (103-23, 103-24) to adapt -
<img file="AR080476A1_D0030.tif" />
the sampling rate of the input signals to match the analysis filter banks of a constant size (in this case 64). It can be noted that the increase in the sampling rate, from fs to 2fs, can be explained by the fact that the sample rate converters use factors to reduce the sampling rate of 772 instead of T, and with the The latter would be serial subbands transposed with the same sampling rate as the input signal. Signals with a reduced sampling rate are fed to separate HFR analysis filter banks (103-32, 103-33 and 103-34), one for each transposition order, which provide a multitude of sub-band signals of complex values . These are fed to the non-linear subband extender units (103—42,103—43 and 103—44). The multitude of output subbands of complex value is fed to the Merge / Combine module 104 together with the output of the subsampled analysis bank 102. The Merge / Combine unit simply merges the subbands from the 102-core analysis filter bank and each stretch factor branch into a single multitude of QMF subbands to be fed into the HFR 105 processing unit.
When the signal spectra of different transposition orders are adjusted so as not to overlap, that is, the spectrum of the 7th transposition order signal must begin when the spectrum of the serial order Γ— 1 ends, the transposed signals need to be of character passband. Hence the traditional bandpass filters 103—12—103—14 of Figure 1. However, through a simple exclusive selection between subbands available through the Cast / Merge 104 unit, the separate band pass filters are redundant and can be avoided. On the other hand, the inherent passband characteristic provided by the QMF bank is exploited by feeding the different contributions from the transposition medium branches independently to different - 28 -
<img file="AR080476A1_D0031.tif" />
subband channels in 104. It is also enough to apply time stretching only to bands that are combined in 104.
Figure 2 illustrates the operation of a non-linear subband stretch unit. The block extractor 201 samples a finite frame of samples of the complex value input signal. The table is defined by an entry pointer position. This frame undergoes non-linear processing at 202 and is subsequently windowed through a finite length window 203. The resulting samples are added to the samples in the overlapping unit — and — sum 204 where the output frame position is defined by an exit pointer position. The input pointer is increased by a fixed magnitude and the output pointer is increased by a stretch factor subbanding times the same magnitude. An iteration of this chain of operations will produce an output signal with the duration that is the stretch factor subbands times the duration of the input subband signal, up to the length of the synthesis window.
While the SSB transposition medium used by SBR [ISO / IEC 14496—3: 2009, Information Technology - Information technology encoding (Visual technology) - Part 3: Audio] typically takes advantage of all The baseband, excluding the first subband, to generate the high band serial, a harmonic transposition medium generally uses a smaller part of the encoder spectrum per core. The magnitude used, the so-called source range, depends on the order of transposition, the bandwidth extension factor, and the rules applied for the combined result, for example, if allowed or not, the spectral superposition of the signals generated at from different transposition orders. As a consequence, only a limited part of the output spectrum - 29 -
<img file="AR080476A1_D0032.tif" />
of the harmonic transposition medium for a given transposition order, it will actually be used by the HFR 105 processing module.
Figure 18 illustrates another embodiment of an exemplary processing implementation for processing a simple serial subband. The single subband serial has been subjected to any type of decimation either before or after being filtered by a bank of analysis filters not shown in Figure 18. Therefore, the time length of the single subband serial It is shorter than the length of time before decimation. The single serial subband is entered into block extractor 1800, which can be identical to block extractor 201, but which can also be implemented in a different way. Block extractor 1800 of Figure 18 operates using a sample / block advance value called for the example, e. The sample / block feed value can be variable or can be fixed and illustrated in Figure 18 as a float in block extractor box 1800. At the output of block extractor 1800 there is a plurality of extracted blocks. These blocks have high overlap, since the sample / block feed value e is significantly smaller than the block length of the block extractor. An example is that the block extractor extracts blocks from 12 samples. The first block comprises samples 0 to 11, the second block comprises samples 1 to 12, the third block comprises samples 2 to 2 to 13, and so on. In this embodiment, the sample / block feed value e is equal to 1, and there is an overlap of 11 times.
The individual blocks are entered in a window means 1802 to window the blocks using a window function for each block. Additionally, a phase calculator 1804 is provided, which calculates a phase for each block. The phase calculator 1804 may use the individual block before the sale or subsequent to the sale. Then a value of - 30 - is calculated
<img file="AR080476A1_D0033.tif" />
pxky phase adjustment is entered to a 1806 phase adjuster. The phase adjuster applies the adjustment value to each sample of the block. Also, the k factor is equal to the bandwidth extension factor. When, for example, the bandwidth extension must be obtained by a factor of 1, then the phase p calculated for a block extracted by block extractor 1800 is multiplied by factor 2 and the adjustment value applied to each sample of the block in phase adjuster 1806 is p multiplied by 2. This is an exemplary value / rule. Alternatively, the phase corrected by synthesis is k * p, p + (k — 1) * p. Thus, in this example, the correction factor is 2 if it is multiplied, or 1 * p if it is added. Other values / rules can be applied to calculate the phase correction value.
In one embodiment, the simple subband serial is a complex subband signal, and the phase of a block can be calculated by a plurality of different ways. One way is to turn the sample in the middle or around the middle of the block and calculate the phase of this complex sample. It is also possible to calculate the phase for each sample.
Although Figure 18 illustrates the way in which a subsequent phase adjuster operates to the window means, these two blocks can also be exchanged, so that the phase adjustment is made to the blocks extracted by the extractor of block and a subsequent window operation is performed. As both operations, that is, window and phase adjustment are multiplications of real values or complex values, these two operations can be summarized in a single operation using a complex multiplication factor which, in itself, is the product of a factor of phase adjustment multiplication and a window factor.
The phase-adjusted blocks are entered into an overlap / sum block and amplitude correction 1808, where the windowed and phase-adjusted blocks are superimposed — added. However, of - 31 -
<img file="AR080476A1_D0034.tif" />
Importantly, the sample / block advance value is block 1808 is different from the value used in block extractor 1800. Particularly, the sample / block advance value in block 1808 is greater than the value e used in the block. block 1800, so that a time stretch of the serial delivered by block 1808 is obtained. Thus, the processed subband serial delivered by block 1808 has a length that is longer than the subband serial entered in block 1800. When the bandwidth extension of two is to be obtained, then the sample / block advance value is used, which is twice the corresponding value in block 1800. This results in a time stretch by a factor of two. However, when other time stretch factors are needed, then other sample / block advance values can be used so that the output of block 1808 has a required length of time.
To address the overlap problem, an amplitude correction is preferably performed to address the problem of different overlays in block 1800 and 1808. However, this amplitude correction could also be introduced in the window / phase adjuster multiplication factor. , but amplitude correction can also be performed subsequent to overlay / processing.
In the example above with a block length of 12 and a sample / block advance value in the block extractor of one, the sample / block advance value for the overlap / sum block 1808 would be equal to two, when bandwidth extension is done by a factor of two. This would still result in an overlap of five blocks. When a bandwidth extension has to be made by a factor of three, then the sample / block advance value used by block 1808 would be equal to three, and the overlap would fall to a three-fold overlap. When a - 32 - is to be carried out
<img file="AR080476A1_D0035.tif" />
bandwidth extension by four, then the overlap / sum block 1808 would have to use a sample advance / block value of four, which would still result in an overlap of more than two blocks.
Large computational savings can be achieved by restricting the input signals to the transposition medium branches that only contain the source range, and this at a sampling rate adapted to each transposition order. The basic block scheme of such a system for a block-based HFR generator is illustrated in Figure 3. The encoder signal per input core is processed by dedicated sample rate reducers that precede the HFR analysis filter banks.
The essential effect of each sampling rate reducer is to filter out the source range signal and supply it to the analysis filter bank at the lowest possible sampling rate. Here, lowest possible refers to the lowest sampling rate that is still suitable for downstream processing, not necessarily the lowest sampling rate that avoids aliasing after decimation. The conversion of sampling rate can be obtained in several ways. Without limiting the scope of the invention, two examples will be given: the first shows the re-sampling performed by processing in the multi-rate time domain, and the second illustrates the re-sampling achieved by means of QMF subband processing.
Figure 4 shows an example of the blocks in a multi-time domain sampling rate reducer for a transposition order of 2. The input serial, which has a bandwidth B Hz, and a frequency of sampling s<sub>s</sub>, is modulated by a complex exponential (401) to run in frequency the beginning of the source range at DC frequency according to x<sub>m</sub> (n) = x (n) · exp f -ϊ2π /<sub>5</sub> — — 33 —
<img file="AR080476A1_D0036.tif" />
Examples of an input signal and the spectrum after modulation are shown in Figures 5 (a) and (b). The modulated serial is interpolated (402) and filtered through a complex low-pass filter with bandpass limits 0 and S / 2 Hz (403). The spectra after the respective steps are shown in Figures 5 (c) and (d). The filtered serial is subsequently decimated (404) and the real part of the serial is computed (405). The results of these steps are shown in Figures 5 (e) and (f). In this particular example, when T-2, B = 0, Q (on a standardized scale, that is, faith = 2), P<sub>2</sub> Corno 24 is chosen to safely cover the source range. The sampling rate reduction factor gives
327 _64 8
P<sub>2</sub><sup>_</sup> 24 <sup>_</sup> 3 'where the fraction has been reduced by the common factor 8. From there, the interpolation factor is 3 (as seen in Figure 5 (c)) and the decimated factor is 8. Using the identities of Noble [Systems Multiritmo and Filter Banks (“Multirate Systems And Filter Banks) by PP Vaidyanathan, 1993, Prentice
Hall, Englewood Cliffsj, the decimator can be moved all the way to the left, and the interpolator all the way to the right in Figure 4. In this way, modulation and filtering are done at the lowest sampling rate possible and the computational complexity is further diminished.
Another approach is to use the subband outputs of the QMF bank for analysis of
32 sub-sampled bands 102 already present in the SBR HFR method. The subbands covering the source ranges for the different branches of the transposition medium are synthesized to the time domain by small subsampled QMF banks that precede the HFR analysis filter banks. This type of HFR system is illustrated in Figure 6. The small QMF banks are obtained by subsampling the original 64-band QMF bank, where the prototype filter coefficients are found by line interpolation of the original prototype filter. Following the notations in Figure 6, the QMF bank of - 34 -
<img file="AR080476A1_D0037.tif" />
synthesis that precedes the branch of the 2nd order transposition medium has Q<sub>2</sub>= 12 bands (subbands with zero-based indices from 8 to 19 in the 32-band QMF). To avoid aliasing in the synthesis process, the first (index 8) and the last (index 19) bands are set to zero. The resulting spectral output is shown in Figure 7. Note that the block filter bank of the block-based transposition medium has 2Q<sub>2</sub>= 24 bands, that is, the same number of bands as in the example based on a sampling rate reducer in the multi-rate time domain (Figure 3).
The system detailed in Figure 1 can be seen as a simplified special case 10 of the re-sampling detailed in Figures 3 and 4. To simplify the arrangement, the modulators are omitted. In addition, all HFR analysis filtering is obtained using 64-band analysis filter banks. From there, P<sub>2</sub> = P<sub>3</sub> = P<sub>4 </sub>= 64 of Figure 3, and the sampling rate reduction factors are 1, 1.5 and 2 for the branches of the transposition medium of 2nd, 3rd and 4th order, respectively.
A block diagram of a factor 2 sampling rate reducer is shown in Figure 8 (a). The low pass filter now of actual value can be written // (z) = 5 (z) / ^ (z ), where /? (z) is the non-recursive part (FIR) and J (z) is the recursive part (IIR). However, for efficient implementation, using the
Noble identities to reduce computational complexity, it is beneficial to design a filter where all poles have multiplicity 2 (double poles) as A (z<sup>2</sup>). Therefore, the filter can be factored as shown in Figure 8 (b). Using the Identity of Noble 1, the recursive part can be moved beyond the decimated means as in Figure 8 (c). The non-recursive filter B (z) can be implemented using standard 2-component polyphase decomposition according to
Λ '. 1 N./2
B (z) = £ b (n) z ~<sup>n</sup> = Σz-'E, (z<sup>2</sup>), where E, (z) = £ 0 (2 · «+ / jz 'n = 0 / = 0 n = 0 -
<img file="AR080476A1_D0038.tif" />
Thus, the sampling rate reducer can be structured as in Figure 8 (d). After using the Noble 1 Identity, the FIR part is computed at the lowest possible sampling rate as shown in Figure 8 (e). From Figure 8 (e) it is easy to see that the FIR operation (delay, decimation means and polyphase components) can be seen as a window-sum operation using a two sample entry block. For two input samples, a new output sample will be produced, effectively resulting in a reduction in the sampling rate of a factor 2.
A block diagram of the sampling rate reduction of factor 1.5 = 3/2 is shown in Figure 9 (a). The low pass filter of reai value can be rewritten H (z) = B (z) lA (z), where B (z) is the non-recursive part (FIR) and A (z) is the recursive part (IIR). As before, for efficient implementation, using Noble Identities to decrease computational complexity, it is beneficial to design a filter where all the poles have multiplicity 2 (double poles) or multiplicity 3 (tripole poles) as A (z<sup>2</sup>) or A (z<sup>3</sup>) respectively. Here, double poles are chosen since the dissertation algorithm for the low pass filter is more efficient, although the recursive part actually gives 1.5 times more complex to implement compared to the triple pole approach. Therefore, the filter can be factored as shown in Figure 9 (b). Using the Identity of Noble 2, the recursive part can be moved to the front of the interpolation medium as in Figure 9 (c). The non-recursive filter B (z) can be implemented using component 2-3 = 6standard standard decomposition
N. 5 N- 6
B (z) = £ b (n) z ~ = £ <sub>z</sub>-'E, (z<sup>b</sup> ), where E, (z) = £ b (6 · n + /) z 'n = 0 / = 0 n = 0
Thus, the sampling rate reducer can be structured as in Figure 9 (d). After using both Noble Identity 1 and 2, the FIR part is computed at the lowest possible sampling rate as shown in - 36 -
<img file="AR080476A1_D0039.tif" />
Figure 9 (e). From Figure 9 (e) it is easy to see that the even index output samples are computed using the group of three lower polyphase filters (Ε<sub>ϋ</sub>(ζ), E<sub>2</sub>(z), E<sub>4</sub>(z)) while odd index samples are computed from the upper group (Æj (z), E<sub>3</sub>(z), E / z)). The opération of each group (delay chain, decimation means and polyphase components) can be seen as a window operation — summing up using an input step of three samples. The window coefficients used in the upper group are the odd index coefficients, while the lower group uses the even index coefficients of the original filter B (z). From there, for a group of three input samples, two new output samples will be produced, effectively resulting in a reduction of the sampling rate of a factor of 1.5.
The signal in the time domain of the decoder per core (101 in Figure 1) can also be sub-sampled using a smaller sub-sampled synthesis transformation in the decoder per core. The use of a smaller synthesis transformation still offers a reduction in computational complexity. Depending on the transition frequency, that is, the bandwidth of the encoder signal per core, the ratio of the synthesis transformation size and the nominal size Q (Q <1), results in an encoder output signal per core that has a Qfs sampling rate. To process the encoder serial by sub-sampled core in the examples detailed in this application, all the analysis filter banks of Figure 1 (102, 103-32, 103-33 and 103-34) need to be set to scale by the Q factor, as well as the sample rate reducers (301-2, 301-3 and 301-T) of Figure 3, the decimation element 404 of Figure 4, and the analysis filter bank 601 of Figure 6. Obviously, Q has to be selected so that all filter bank sizes are integers.
— 37 —
<img file="AR080476A1_D0040.tif" />
Figure 10 illustrates the alignment of the spectral edges of the HFR transposition medium signals to the spectral edges of the envelope adjustment frequency table in an improved HFR encoder, such as SBR [ISO / IEC 14496-3: 2009, Information technology - Coding of audio — visual objects - Part 3: Audio], Figure 10 (a) shows a schematic graph of the frequency bands that make up the envelope adjustment table, the so-called scale factor bands, covering the frequency range from the transition frequency k<sub>x</sub> at the stop frequency k<sub>s</sub>. The scale factor bands constitute the frequency grid used in the enhanced HFR encoder when adjusting the energy level of the regenerated high band over frequency, that is, the frequency envelope. To adjust the envelope, the signal energy is averaged over a time / frequency block restricted by the selected scale factor band edges and time edges.
Specifically, Figure 10 illustrates in the upper portion, a division into frequency bands 100, and it turns out from Figure 10 that the frequency bands increase with the frequency, where the horizontal axis corresponds to the frequency and has the notation of Figure 10, filter bank channels k, where the filter bank can be implemented as a QMF filter bank such as a 64-channel filter bank or it can be implemented via a digital Fourier transformation, where k corresponds to a certain frequency tray of the DFT application. Therefore, a frequency tray of a DFT application and a filter bank channel of a QMF application indicate the same in the context of this description. Therefore, the parametric data is given for the high frequency part 102 in the frequency trays 100 or frequency bands. The low frequency part of the signal extended in width - 38 -
<img file="AR080476A1_D0041.tif" />
The band is finally indicated at 104. The intermediate illustration in Figure 10 illustrates the patch ranges for a first patch 1001, a second patch 1002 and a third patch 1003. Each patch extends between two patch edges, where there is an edge of lower patch 1001a and an upper patch edge 1001b, pair the first patch. The upper edge of the first patch indicated in 1001b corresponds to the lower edge of the second patch which is indicated in 1002a. Therefore, reference numbers 1001b and 1002a really refer to one and the same frequency. An upper patch edge 1002b of the second patch, again, corresponds to a lower patch edge
1003a of the third patch, and the third patch also has an upper patch edge 1003b. It is preferred that there are no holes between individual patches, but this is not a final requirement. It is visible in Figure 10 that the patch edges 1001b, 1002b do not coincide with corresponding edges of the frequency bands 100 but rather are within certain frequency bands 101. The bottom line in Figure 10 illustrates different patches with aligned edges 1001c, where the alignment of the upper edge 1001c of the first patch automatically means the alignment of the lower edge 1002c of the second patch and vice versa. Additionally, it is indicated that the upper edge of the second patch 1002d is now aligned with the frequency edge or lower frequency band 101 in the first line of Figure 10 which, therefore, automatically the lower edge of the third patch indicated in 1003c , is also aligned.
In the embodiment of Figure 10 it is shown that the aligned edges are aligned to the lower frequency edge of the coincident frequency band 101, but the alignment could also be made in a different direction, that is, that the patch edge 1001c, 1002c is aligned to the upper frequency edge of the band 101 instead of the lower frequency edge of the - 39 -
<img file="AR080476A1_D0042.tif" />
same. Depending on the actual implementation, one of those possibilities can be applied and can even be a mixture of both possibilities for different patches.
If the signals generated by different transposition orders are not aligned to the scale factor bands, as illustrated in Figure 10 (b), artifacts may appear if the spectral energy changes dramatically in the vicinity of a transposition band edge , since the envelope adjustment process will maintain the spectral structure within a scale factor band. Thus, the invention adapts the frequency edges of the transposed signals to the edges of the scale factor bands as shown in Figure 10 (c). In the figure, the upper edge of the signals generated by transposition orders of 2 and 3 (7 = 2, 3) are diminished or small amount, compared to Figure 10 (b), to align the frequency edges of the bands of transposition to existing scale factor band edges.
A realistic scenario that shows potential artifacts when non-aligned edges are used, is represented in Figure 11. Figure 11 (a) again shows the scale factor band edges. Figure 11 (b) shows the HFR-generated signals not adjusted for transposition orders 7 = 2, 3 and 4 together with the coreband serial decoded by kernel. Figure 11 (c) shows the serial set in envelope when a flat white envelope is assumed. Blocks with squared areas represent scale factor bands with high intra-band energy variations, which can cause anomalies in the output serial.
Figure 12 illustrates the scenario of Figure 11, but this time using aligned edges. Figure 12 (a) shows the scale factor band edges, Figure 12 (b) represents the serials generated by unadjusted HFRs, of transposition orders 7 = 2, 3 and 4 together with the baseband serial - 40 -
<img file="AR080476A1_D0043.tif" />
decoded by nucleus and, in line with Figure 11 (c), Figure 12 (c) shows the envelope-adjusted signal if a flat white envelope is assumed. As seen from this figure, there are no scale factor bands with high intra-band energy variations due to poor alignment of transposed serial bands and scale factor bands, and therefore, the potential artifacts are diminished.
Figure 25a illustrates a general view of an implementation of patch edge calculator 2302 and the patch and the location of those elements within the bandwidth extension scenario in accordance with a preferred embodiment. Specifically, an input interface 2500 is provided, which receives the low band data 2300 and the parametric data 2302. The parametric data may be bandwidth extension data, such as what is known from ISO / IEC 14496-3: 2009, which is incorporated herein by reference in its entirety, and particularly with respect to the section related to extension of bandwidth, which is section 4.6.18 SBR tool (SBR tool). Of particular relevance in section 4.6.18 is section 4.6.18.3.2 Frequency band tables, and in particular the calculation of some frequency tables fmasten fïabieHighi fïabieLow, fïabieNoise and fïabieLim · In particular, the section
4.6.18.3.2.1 of the Standard defines the calculation of the master frequency band tables, and section 4.6.18.3.2.2 defines the calculation of the frequency band tables derived from the master frequency band table, and in particular delivery is calculated fîabieHigh, frabieLow and fïabieNoise · Section 4.6.18.3.2.3 defines the calculation of the limiting frequency band table.
The frabieLow low resolution frequency table is for low resolution parametric data and the fîabieHigh high resolution frequency table is for high resolution parametric data, which are both - 41 -
<img file="AR080476A1_D0044.tif" />
possible in the context of the MPEG-4 SBR software tool, as discussed in the aforementioned Standard and if the parametric data is low-resolution parametric data or high-resolution parametric data, it depends on the encoder implementation. The input interface 2500 determines whether the parametric data is low or high resolution data and provides this information to the frequency table calculator 2501. The frequency table calculator then calculates the master table or generally derives a high resolution table 2502 and a low resolution table 2503 and provides the same to the patch edge calculator core 2504, which additionally comprises or cooperates with, a calculator of limiting band 2505. Elements 2504 and
2505 they generate 2506 aligned synthesis patch edges and corresponding limiting band edges related to the synthesis range. This information
2506 it is provided with a source band calculator 2507, which calculates the source range of the low band audio signal for a certain patch so that together with the corresponding transposition factors, the aligned synthetic patch edges 2506 are obtained after of patching using, for example, a harmonic transposition medium 2508 as a patch.
In particular, harmonic transposition means 2508 can execute different patching algorithms such as the DFT-based patching algorithm or a QMF-based patching algorithm. The harmonic transposition means 2508 can be implemented to perform a Vocoder type processing which is described in the context of Figures 26 and 27 for the realization of harmonic transposition means based on QMF, but other media operations can also be used. of transposition such as a transposition medium based on DFT for the purpose of generating a high frequency portion in a Vocoder type structure. For the DFT-based transposition medium, the source band calculator calculates frequency windows for the low range - 42 -
<img file="AR080476A1_D0045.tif" />
frequency. For the QMF-based implementation, source band calculator 2507 calculates the required QMF bands of the source range for each patch. The source range is defined by the low band audio data 2300, which is typically provided in encoded form and sent by the input interface 2500 to a decoder per core 2509. Core decoder 2509 feeds its output data to a bank of 2510 analysis filters, which can be a QMF implementation or a DFT implementation. In the QMF implementation, the analysis filter bank 2510 can have 32 filter bank channels, and these 32 filter bank channels define the maximum source range, and the harmonic transposition medium 2508 then selects, from these 32 bands , the current bands that make up the adjusted source range as defined by source band calculator 2507 to, for example, meet the adjusted source range data in the table in Figure 23, provided that the frequency values in the table of Figure 23 are converted to subband indexes of a synthesis filter bank. A similar procedure can be performed for the DFT-based transposition medium, which receives for each patch a certain window for the low frequency range and this window is then sent to the DFT block 2510 to select the source range according to the edges of adjusted or aligned synthesis patch calculated by block 2504.
The transposed signal 2509 delivered by the transposition means 2508 is sent to an envelope adjuster and gain limiter 2510, which receives as input the high resolution table 2502 and the low resolution table 2503, the adjusted limiting bands 2511 and, of course , parametric data 2302. The high band adjusted by envelope on the line 2512 is then entered into a synthesis filter bank 2514, which additionally receives the low band typically in the form of output by the 2509 core decoder.
— 43 —
<img file="AR080476A1_D0046.tif" />
Both contributions are merged through the synthesis filter bank 2514 to finally obtain the reconstructed high frequency signal in line 2515.
It is darò that the fusion of the high band and the low band can be done in a different way, such as performing a fusion in the time domain instead of in the frequency domain. It is also true that the order of fusion can be changed, regardless of the implementation of the fusion and the envelope setting, that is, so that the envelope setting of a certain frequency range can be made subsequent to the fusion or , alternatively, before the merger, where the last case is illustrated in Figure 25a. Furthermore, it is detailed that the envelope adjustment can even be performed before transposition in the transposition means 2508, so that the order of the transposition means 2508 and the envelope adjuster 2510 may also be different from what is illustrated in the Figure 25a as an embodiment.
As already detailed in the context of block 2508, in the embodiments a harmonic transposition medium based on DFT or a harmonic transposition medium based on QMF can be applied. Both algorithms rely on phase spread of Vocoder frequency. The serial time domain encoder per core is extended in bandwidth using a modified phase Vocoder structure. The bandwidth extension is performed by time stretching followed by decimation, that is, transposition, using several transposition factors (t = 2, 3, 4) in a common analysis / synthesis transformation stage. The output signal of the transposition medium will have a sampling rate twice that of the input signal, which means that for a transposition factor of two, the signal will be stretched over time but not decimated, efficiently producing a serial of the same duration as the input serial but having twice the sampling frequency. The combined system can be interpreted as très -
<img file="AR080476A1_D0047.tif" />
parallel transposition means using transposition factors of 2, 3 and 4, respectively, where the decimated factors are 1; 1.5 and 2. To reduce complexity, the transposition means of factor 3 and 4 (third and fourth order transposition means) are integrated into the factor 2 transposition means (second order transposition means) by means of interpolation as discussed subsequently in the context of Figure 27.
For each frame, a nominal tamarium transformation schedule of a transposition medium is determined, depending on an over-sampling in the serial-adaptive frequency domain that can be applied to improve the transient component response or that can be turned off. . This value is indicated in Figure 24a as FFTSizeSyn. Then, blocks of sold input samples are transformed, where for block extraction a block advance value or analysis step value of a much smaller number of samples is executed, to have a significant block overlap. The extracted blocks are transformed to the frequency domain by means of a DFT depending on the serial control over-sampling of the serial frequency domain-adaptable. The phases of the DFT coefficients of complex values are modified according to the three transposition factors used. For second order transposition, the phases are duplicated, for third and fourth order transpositions, the phases are triplicate, quadruplicate or interpolated from two consecutive DFT coefficients. The modified coefficients are subsequently transformed back to the time domain by means of a DFT, are sold and combined by means of superimposing — adding using an output step different from the input step. Then, using algorithm illustrated in Figure 24a, the patch edges are calculated and written in the xOverBin array. Then the patch edges are used for - 45 -
<img file="AR080476A1_D0048.tif" />
calculate transformation windows in the time domain for the application of the DFT transposition medium. For the QMF transposition medium, source range channel numbers are calculated based on the patch edges calculated in the synthesis range. Preferably, this is happening before transposition as control information is needed to generate the transposed spectrum.
Subsequently, the heavy code indicated in Figure 24a is discussed in relation to the flow chart of Figure 25b illustrating a preferred implementation of the patch edge calculator. In step 2520 a frequency table is calculated based on the input data such as a high or low resolution table. From there, block 2520 corresponds to block 2501 of Figure 25a. Then, in step 2522 a bianco synthesis patch edge is determined based on the transposition factor. In particular, the bianco synthesis patch edge corresponds to the result of the multiplication of the patch value of Figure 24a and fTabieLow (O), where fTabieLow (O) indicates the first channel or tray of the bandwidth extension range, that is, the first band above the transition frequency, below which the 2300 input audio data is given with high resolution. In step 2524 it is verified whether the bianco synthesis patch edge matches an entry in the low resolution table within an alignment range. In particular, an alignment range of 3 is preferred, as indicated, for example, in 2525 in Figure 24a. However, other ranges are also useful, such as ranks smaller than or equal to 5. If in step 2524 it is determined that the blank coincides with a low resolution table entry, then this matching entry is taken as the new patch edge instead of the white patch edge. However, if it is determined that there is no entry within the alignment range, step 2526 is applied, in which the same search is made with the registration table - 46 -
<img file="AR080476A1_D0049.tif" />
resolution also indicated in 2527 in Figure 24a. If in step 2526 it is determined that there is a table entry within the alignment range, then the matching entry is taken as a new patch edge instead of the bianco synthesis patch edge. However, if in step 2526 it is determined that even in the high resolution table there is no value within the alignment range, then step 2528 is applied in which the edge of bianco synthesis is used without any alignment. This is also indicated in Figure 24a in 2529. Therefore, step 2528 can be seen as a position in case of a fall so that in any case it is guaranteed that the bandwidth extension decoder does not remain in a loop, but rather reaches a solution in any case even if There is a very specific and problematic selection of frequency tables and bianco ranges.
With respect to the pseudo code of Figure 24a, it is detailed that lines of code 2531 execute some processing to ensure that all variables are in a useful range. Likewise, the verification about whether the bianco coincides with an entry in the low resolution table within an alignment range, is performed as the calculation of a difference (lines 2525, 2527) between the bianco synthesis patch edge calculated by the product indicated near block 2522 in Figure 25b and indicated on lines 2525, 2527 and a current table entry defined by the sfbL parameter for line 2525 or sfbH for line 2527 (sfb = scale factor band). Naturally, other verification operations can also be executed.
Also, it is not necessarily the case that a match is sought within an alignment range when the alignment range is predetermined. Instead, you can search the table to find the best matching table entry, that is, the table entry that is -
<img file="AR080476A1_D0050.tif" />
closer to the value of the bianco frequency regardless of whether the difference between those two is small or high.
Other implementations refer to a search in the table, such as fîabieLow or fiabieHigh for the highest edge that does not exceed the bandwidth limits (fundamental) of the signal generated by HFR for a transposition factor T. Then this is used higher edge as the frequency limit of the signal generated by HFR of the transposition factor T. In this implementation the calculation of bianco indicated near box 2522 in Figure 25b is not required.
Figure 13 illustrates the adaptation of the limiter band edges of
HFT, as described, for example, in SBR [ISO / IEC 14496—3: 2009, Information Technology - Information technology encoding (Visual technology) - Part 3: Audio] Harmonic patches in an enhanced HFR encoder. The limiter operates on frequency bands that have a much thicker resolution than the scale factor bands, but the principle of operation is very similar. In the limiter the average gain value is calculated for home one of the limiter bands. Individual gain values, that is, envelope gain values calculated for each of the scale factor bands, are not allowed to exceed the average gain value of the limiter by more than a certain multiplicative factor. The purpose of the limited is to suppress large variations of scale factor band gains within each of the bands of the limiter. While the adaptation of the bands generated by the means of transposition to the scale factor bands ensures small variations of the intra-band energy within the scale factor band, the adaptation of the limiter band edges to the edges of transposition media band, according to the present invention, handles the differences of - 48 -
<img file="AR080476A1_D0051.tif" />
larger scale energy between processed bands of the transposition medium. Figure 13 (a) shows the frequency limits of the signals generated by HFR of transposition orders T = 2, 3 and 4. The energy levels of the different transposed signals can be substantially different. Figure 13 (b) shows the frequency bands of the limiter which are typically of constant width on a logarithmic frequency scale. The frequency band edges of the transposition medium are added as constant limiter edges and the remaining limiter edges are re-calculated to keep the logarithmic relationships as close as possible, as illustrated, for example, in Figure 13 (c ).
Other embodiments employ a mixed patch scheme which is shown in Figure 21, where the mixed patch method is executed within a time block. For complete coverage of the different regions of the HF spectrum, a BWE comprises several patches. In HBE, higher patches require high transposition factors within phase Vocoders, which particularly deteriorates the perceptual quality of the transient components.
Embodiments thus generate the highest order patches that occupy the upper spectral regions preferably by computationally efficient SSB copy-up patching and the lower order patches that cover the median spectral regions, for which conservation of the harmonic structure is preferably desired, by patching HBE. The individual mixture of patching methods may be static over time or, preferably, may be signaled in the series of bits over time.
For the copy-up operation, the low frequency information can be used as shown in Figure 21. Alternatively, the patch data that was generated using HBR methods can be used as illustrated in -
<img file="AR080476A1_D0052.tif" />
Figure 21. The latter leads to a less dense tonal structure for higher patches. In addition to these two examples, any other combination of copy-up and HBE can be devised.
The advantages of the proposed concepts are • Better perceptual quality of transient components • Reduced computational complexity
Figure 26 illustrates a preferred processing chain for the purpose of bandwidth extension, where different processing operations can be performed within the non-line subband processing indicated in blocks 1020a, 1020b. In one implementation, the band-selective signal processing in the processed time domain such as the extended bandwidth signal is executed in the time domain rather than in the subband domain, which exists before the filter bank of synthesis 2311.
Figure 26 illustrates an apparatus for generating an extended audio signal in bandwidth from a low band input signal 1000 according to another embodiment. The apparatus comprises a bank of analysis filters 1010, a non-line mode subband processor - subband 1020a, 1020b, a envelope adjuster subsequently connected 1030 or, in general, a high frequency reconstruction processor operating on reconstruction parameters high frequency, for example, as input into parameter line 1040. The envelope adjuster, or as is generally expressed, the high frequency reconstruction processor, processes individual subband signals for each subband channel and sends processed subband signals for each subband channel into a 1050 synthesis filter bank. Synthesis filter bank 1050 receives, in its lower channel input signals, a subband representation of the decoder signal per low band core. Depending on the implementation, the low band can also -
<img file="AR080476A1_D0053.tif" />
be derived from the outputs of the analysis filter bank 1010 of Figure 26. The transposed subband signals are fed into higher filter bank channels of the synthesis filter bank to perform high frequency reconstruction.
The filter bank 1050 finally delivers a serial output of transposition media which comprises bandwidth extensions by transposition factors 2, 3 and 4, and the serial delivered by block 1050 is no longer limited in bandwidth to the transition frequency, that is, at the highest frequency of the serial coder per core corresponding to the lowest frequency of the serial components generated by SBR or HFR. The analysis filter bank 1010 of Figure 26 corresponds to the analysis filter bank 2510 and the synthesis filter bank 1050 may correspond to the synthesis filter bank 2514 of Figure 25a. In particular, as discussed in the context of Figure 27, the source band calculation illustrated in block 2507 in the figure
25a is performed within the non-linear subband processing 1020a, 1020b, using the aligned synthetic patch edges and the limiter band edges calculated by blocks 2504 and 2505.
With respect to the limiter frequency band tables, it should be noted that the limiter frequency band tables can be constructed to have either a limited band over the entire reconstruction range, or approximately 1.2 ; 2 or 3 bands per octave, serialized by a bit string element in time bsjimi-ter_bands as defined in ISO / IEC 14496-3: 2009, 4.6.18.3.2.3. The band table may comprise additional bands corresponding to the high frequency generator patches. The table may contain indexes of synthetic filter bank subbands, where the number of elements is equal to the number of bands plus one. When harmonic transposition is active, it is ensured that the calculator of - 51 -
<img file="AR080476A1_D0054.tif" />
Limiter band introduces limiter band edges that match the patch edges defined by patch edge calculator 2504. Additionally, the remaining limiter band edges are then calculated between those limiter band edges fixedly set for the edges. of patch.
In the embodiment of Figure 26, the filter bank performs a twice sampled overlay and has a certain subband analysis spacing 1060. The filter bank 1050 has a synthetic subband spacing 1070 which is, in this embodiment, double of the sub-band spacing of analysis which results in a transposition contribution as will then be discussed in the context of Figure 27.
Figure 27 illustrates a detailed implementation of a preferred embodiment of a non-linear subband processor 1020a of Figure 26. The circuit illustrated in Figure 27 receives as input a single serial subband 1080, which is processed in three branches. The upper branch 110a is for a transposition by a transposition factor of 2. The middle branch of Figure 27 indicated in 11 Ob is for a transposition by a transposition factor of 3, and the lower branch of Figure 27 is for a transposition by a transposition factor 4, and is indicated by the number of reference 11 Oc. However, the reai transposition obtained by each processing element of Figure 27 is only 1 (that is, without transposition) per branch 110a. The real transposition obtained by the processing element illustrated in Figure 27 by the middle branch 11 Ob is equal to 1.5 and the real transposition for the lower branch 11 Oc is equal to 2. This is indicated by the numbers in square brackets. to the left of Figure 27, where transposition factors T are indicated. Transpositions of 1.5 and 2 represent a first transposition contribution obtained by having a decimation operation on branches 110b, - 52 -
<img file="AR080476A1_D0055.tif" />
110c and a time stretch using the overlay processor — add. The second contribution, that is, duplication of the transposition, is obtained by the synthesis filter bank 105, which has a subband spacing of synthesis 1070 which is twice the subband spacing of the analysis filter bank. Therefore, as the synthesis filter bank has twice the synthesis subband spacing, the decimation functionality does not take place on branch 110a.
However, branch 11 Ob has a decimated functionality to obtain a transposition by 1.5. Due to the fact that the synthesis filter bank has twice the physical subband spacing than the analysis filter bank, a transposition factor of 3 is obtained as indicated in Figure 27 to the left of the block extractor for second branch 110b.
Similarly, the third branch has a decimation functionality corresponding to a transposition factor of 2, and the final contribution of different subband spacing in the analysis filter bank and the synthesis filter bank finally corresponds to a transposition factor of 4 of the third branch 110c.
In particular, each branch has a block extractor 120a, 120b, 120c and each of these block extractors can be similar to the block extractor
1800 of Figure 18. Also, each branch has a phase calculator 122a,
122b and 122c, and the phase calculator may be similar to the phase calculator 1804 of Figure 18. Also, each branch has a phase adjuster 124a, 122b and 122c, and the phase adjuster may be similar to the phase adjuster 1806 of Figure 18. Also, each branch has a window element 126a, 120b,
120c, where each of these window elements can be similar to window element 1802 of Figure 18. In case focus, window elements 126a, 126b, 126c can also be configured to apply a
<img file="AR080476A1_D0056.tif" />
rectangular window along with some filled with zeros. The transposed or patch signals of each branch 110a, 110b, 110c of the embodiment of Figure 11 are entered in adder 128, which adds the contribution from each branch to the real subband signal to finally obtain what is called blocks of transposes at the output of adder 128. An overlay procedure is then performed — sum in the superposition medium — sum 130, and the superposition medium — sum 130 may be similar to the superposition block — sum 1808 in Figure 18. The superposition means — sum applies a superposition feed value — sum d 2 e, where e is the superposition value — feed or step value of block extractors 120a, 120b, 120c, and the superposition means — sum 130 delivers the transposed signal, which in this embodiment of Figure 27, is a subband signal output for channel k, this is for the subband channel currently observed. The processing illustrated in Figure 27 is performed for each analysis subband or for a certain group of analysis subbands and, as illustrated in Figure 26, transposed subband signals are input into the synthesis filter bank 105 after being processed by block 103 to finally obtain the output signal of transposition means illustrated in Figure 26 at the output of the output block 105.
In one embodiment, block extractor 120a of the first branch of transposition medium 110a extracts 10 subband samples and subsequently converts these 10 QMF samples to polar coordinates. This output, generated by the phase adjuster 124a, is then sent to the window element 126a, which extends the output by zeros for the first and last value of the block, where this operation is equivalent to a window (synthesis) with a rectangular window. of length 10. Block extractor 120a of branch 110a does not decimate. Therefore, samples taken using - 54 -
<img file="AR080476A1_D0057.tif" />
Block extractors are mapped into a block extracted in the same sample spacing as where they were extracted.
However, this is different for branches 110b and 110c. Block extractor 120b preferably extracts a block of 8 subband samples and distributes these 8 subband samples of the extracted block in a different subband sample spacing. The subband sample entries not whole for the extracted block are obtained by interpolation, and the QMF samples thus obtained together with the interpolated samples are converted to polar coordinates and processed by the phase adjuster. Then, again, the window is made in the window element 126b to extend the block output by the phase adjuster 124b by means of zeros for the first two samples and the last two samples, the operation of which is equivalent to a window (synthesis) with a rectangular window of length 8.
The block extractor 120c is configured to extract a block with a time extension of 6 subband samples and decimates a decimated factor 2, converts the QMF samples into polar coordinates and again performs an operation on the adjuster phase 124b, and the output again is extended by zeros, although now for the first three subband samples and for the last three subband samples. This operation is equivalent to a window (synthesis) with a rectangular sale of length 6.
The transposition outputs of each branch are then added to form the combined QMF output by means of adder 128, and the combined QMF outputs are finally superimposed using overlay — sum in block 130, where the overlap advance — sum or value of step is twice that the step value of block extractors 120a, 120b, 120c as discussed above.
— 55 —
<img file="AR080476A1_D0058.tif" />
Figure 27 further illustrates the functionality performed by source band calculator 2507 of Figure 25a, where reference number 108 is considered to illustrate the subband analysis signals available for a patch, that is, the signals indicated at 1080 of the Figure 26, which are delivered by the analysis filter bank 1010 of Figure 26. The selection of the correct subband of the subband analysis signals or, in the other embodiment related to the DFT transposition medium, the application of the correct analysis frequency window is realized by block extractors 120a, 120b, 120c. To this end, the patch edges indicating the first subband signal, the last subband signal and the intermediate subband signals for each patch are provided to the block extractor for each transposition branch. The first branch finally resulting in a transposition factor of 7 = 2 receives, with its block extractor 120a all subband indexes between xOverQmf (0) and xOverQmf (1), and block extractor 120a then extracts a block of the analysis sub-band thus selected. It should be noted that the patch edges are given as a channel index of the synthesis range indicated by k, and the analysis bands are indicated by n with respect to their subband channels. Therefore, since n is calculated by dividing 2k by T, the channel numbers of the analysis band n, therefore, are equal to the channel numbers of the synthesis range due to the double frequency spacing of the synthesis filter bank as discussed in the context of Figure 26. This is indicated above block 120a for the first block extractor 120a or, generally, for the first branch of transposition means 110a. Then, for the second patch branch 110b, the block extractor receives all channel indexes of synthesis range between xOverQmf (1) and xOverQmf (2). In particular, the source range channel indices, from which the block extractor has to extract blocks for further -
<img file="AR080476A1_D0059.tif" />
processing, are calculated from the synthesis range channel indices given by the determined patch edges by multiplying k with the 2/3 factor. Then, the entire part of this calculation is taken as the analysis channel number n, from which the block extractor then extracts the block to be further processed by elements 124b, 126b.
For the third branch 110c, the block extractor 120c, again receives the patch edges and performs a block extraction of the subbands corresponding to synthetic bands defined by xOverQmf (2) to xOverQmf (3). The analysis numbers n are calculated by 2 multiplied by k, and this is the calculation rule for calculating the analysis channel numbers from the synthesis channel numbers. In this context, it should be noted that xOverQmf corresponds to xOverBin of Figure 24a, although Figure 24a corresponds to the DFT-based patch, while xOverQmf corresponds to the QMF-based patch. The calculation rules for determining xOverQmf (i) is determined in the same way as illustrated in Figure 24a, but the fftSizeSyn / 128 factor is not required to calculate xOverQmf.
The procedure for determining the patch edges to calculate the analysis ranges for the embodiment of Figure 27, is also illustrated in Figure 24. In the first step 2600 the patch edges for the patches corresponding to the transposition factors are calculated. 2, 3, 4 and, optionally, even more, as discussed in the context of Figures 24a or Figures 25a. The source range frequency window for the DFT patch or the source range subbands for the QMF patch is then calculated using the equations discussed in the context of blocks 120a, 120b, 120c, which are also illustrated to the right of block 2602. A patch is then made by calculating the transposed signal and mapping the transposed serial at high frequencies as indicated in block 2604, and
<img file="AR080476A1_D0060.tif" />
—
<img file="AR080476A1_D0061.tif" />
the calculation of the transposed serial is illustrated in particular in the procedure of Figure 27, where the transposed serial delivered by the overlay means - sum of block 130 corresponds to the result of the patching generated by the procedure of block 2604 of Figure 24.
One embodiment comprises a method for decoding an audio signal using harmonic transposition based on a subband block, which comprises filtering a decoded signal by nucleus through a bank of M-band analysis filters to obtain a subband signal set; synthesize a subset of said subband signals by means of synthetic filter banks with reduced sampling rate, to obtain source range signals with reduced sampling rate.
One embodiment relates to a method for aligning the spectral band edges of signals generated by HFR to spectral edges used in a parametric process.
One embodiment relates to a method for aligning spectral edges of the signals generated by HFR to spectral edges of the envelope adjustment frequency table comprising: the search for the highest edge in the envelope adjustment frequency table that does not exceed the fundamental bandwidth limits of the signal generated by HFR of transposition factor T;
and use the highest edge found as the frequency limit of the signal generated by HFR of transposition factor T.
One embodiment relates to a method for aligning the spectral edges of the limiter software tool to the spectral edges of the signals generated by HFR, comprising: adding the frequency edges of the signals generated by HFR to the edge table used when creating the frequency band edges used by the limiter software tool; and - 58 -
<img file="AR080476A1_D0062.tif" />
force the limiter to use the frequency edges added as constant edges and adjust the remaining edges accordingly.
One embodiment relates to the combined transposition of an audio serial comprising several integer transposition orders in a low resolution filter bank domain where the transposition operation is performed on time blocks of the subband signals.
Another embodiment relates to combined transposition, where transposition orders greater than 2 are packaged in a transposition environment of order 2
Another embodiment relates to combined transposition, where transposition orders greater than 3 are packaged in a transposition environment of order 3, while transposition orders less than 4 are performed separately.
Another embodiment relates to combined transposition, where transposition orders (for example, transposition orders greater than 2) are created by replication of previously calculated transposition orders (that is, especially lower orders) including the bandwidth encoded by kernel. Any conceivable combination of available transposition orders and core bandwidths is possible without restrictions.
One embodiment relates to reduction of computational complexity due to the small number of analysis filter banks that are required for transposition.
An embodiment relates to an apparatus for generating an extended bandwidth signal from an input audio serial, comprising: a patch to patch an input audio signal to obtain a first patched signal and a second patched serial , the second signal having a different patch frequency compared to the first one - 59 -
<img file="AR080476A1_D0063.tif" />
patched serial, where the first patched serial is generated using a first patching algorithm, and the second patched serial is generated using a second patching algorithm; and a combiner to combine the first patched signal and the second patched signal to obtain the extended bandwidth signal.
Another embodiment relates to this apparatus, in which the first patch algorithm is a harmonic patch algorithm, and the second patch algorithm is a non-harmonic patch algorithm.
Another embodiment relates to a prediction apparatus, in which the first patch frequency is less than the second patch frequency or vice versa.
Another embodiment relates to a prediction apparatus, in which the input signal comprises patch information; and in which the patch is configured to be controlled by the patch information extracted from the input serial to vary the first patch algorithm or the second patch algorithm in accordance with the patch information.
Another embodiment relates to a prediction apparatus, in which the patch is operative to patch subsequent blocks of audio signal samples, and in which the patch is configured to apply the first patching algorithm and the second patching algorithm to the Same block of audio samples.
Another embodiment relates to a prediction apparatus, in which the patcher comprises, in arbitrary orders, a decimation means controlled by a bandwidth extension factor, a filter bank, and an extender for a bank subband signal. of filters.
Another embodiment relates to a prediction apparatus , in which the extender comprises a block extractor for extracting a number of blocks of - 60
<img file="AR080476A1_D0064.tif" />
overlap in accordance with an extraction advance value; a phase adjuster or window element for adjusting subband sampling values in each block based on a window function or phase correction; and an overlay means — sum to perform overlay processing — sum of blocks sold and adjusted in phase using an overlap advance value greater than the extraction advance value.
Another embodiment relates to an apparatus for bandwidth extension of an audio serial comprising: a bank of filters to filter the audio serial to obtain serial subbands with reduced sampling rate; a plurality of different subband processors to process different subband serials in different ways, the subband processors performing different serial subband time stretching operations using different stretching factors; and a fusion means for merging processed subbands delivered by the plurality of different subband processors to obtain an extended serial audio bandwidth.
Another embodiment relates to an apparatus for reducing the sampling rate of an audio serial, which comprises; a modulator; an interpolator that uses an interpolation factor; a complex low pass filter; and a decimation means that uses a decimated factor, where the decimated factor is higher than the interpolation factor.
An embodiment relates to an apparatus for reducing the sampling rate of an audio serial, comprising: a first bank of filters to generate a plurality of subband serials from the audio serial, wherein a sampling rate of the serial subband is smaller than a sampling rate of the audio serial; at least one synthesis filter bank followed by an analysis filter bank to perform a sample rate conversion, the synthesis filter bank having a different number of channels of -
<img file="AR080476A1_D0065.tif" />
a number of channels of the analysis filter bank; a time stretching processor to process the signal with converted sampling rate; and a combiner to combine the time stretched signal and a low band signal or a serial time stretched signal.
Another embodiment relates to an apparatus for reducing the sampling rate of an audio signal by a factor of reduction of the non-integer sampling rate, which comprises; a digital filter; an interpolator that has an interpolation factor; a poly-phase element that has odd and even taps and a decimation means that has a decimation factor that is greater than the interpolation factor, the decimated factor and the interpolation factor selected such that a quotient of interpolation factor and decimated factor is not integer.
One embodiment refers to an apparatus for processing an audio signal, comprising: a decoder per core having a synthesis transformation size that is smaller than a nominal transformation size by a factor, so that the output signal it is generated by the decoder per core that has a smaller sampling rate than a nominal sampling rate corresponding to the nominal transformation size; and a post-processor that has one or more filter banks, one or more time extenders and a fusion medium, wherein a number of filter bank channels of the one or more filter banks is reduced compared to a number according to what is determined by the nominal transformation size.
Another embodiment relates to an apparatus for processing a low band signal, comprising: a patch generator for generating multiple patches using the low band audio signal; an envelope adjuster for adjusting a serial envelope using scale factors given by adjacent scale factor bands having scale factor band edges, wherein - 62
<img file="AR080476A1_D0066.tif" />
The patch generator is configured to perform multiple patches, so that an edge between adjacent patches matches a bode between adjacent scale factor bands in the frequency scale.
One embodiment relates to an apparatus for processing a low band audio signal, comprising: a patch generator for generating multiple patches using the low band audio serial; and an envelope adjuster limiter to limit envelope adjustment values for a serial limiting on adjacent limiter bands that have limiter band edges, where the patch generator is configured to perform multiple patches, so that one edge between adjacent patches matches a bode between adjacent limiter bands on a frequency scale.
The processing of the invention is useful for improving encoders — audio decoders that rely on a bandwidth extension scheme. Especially if optimal perceptual quality at a quantity of transmitted bits is very important and, at the same time, the processing power is a limited resource.
Most of the featured applications are audio decoders that are frequently implemented in portable devices and thus work on a battery power source.
The inventive encoded audio serial can be stored in a digital storage medium or it can be transmitted through a transmission medium such as a wireless transmission medium or a physical transmission medium such as the Internet.
Depending on certain implementation requirements, embodiments of the invention may be implemented in hardware or software. The implementation can be carried out using a digitai storage medium, for example a floppy disk, a DVD, a CD, a ROM, a - 63 -
<img file="AR080476A1_D0067.tif" />
EPROM, an EEPROM or a FLASH memory, which have electronically readable control signals stored in them, which cooperate (or are able to cooperate) with a programmable computing system so that the respective method is executed.
Some embodiments according to the invention comprise a data carrier that has electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is executed.
Generally, embodiments of the present invention can be implemented as a computer program with a program code, the program code being operative to execute one of the methods when the computer program product runs on a computer. The program code can be stored, for example, on a carrier readable by a machine.
Other embodiments comprise the computer program for executing one of the methods described herein, stored in a carrier readable by a machine.
In other words, an embodiment of the inventive method is, therefore, a computer program that a program code to execute one of the methods described herein, when the computer program runs on a computer.
A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded therein, the computer program for executing one of the methods described herein.
A further embodiment of the inventive method is, therefore, a data transmission or a sequence of serials representing the program of - 64
<img file="AR080476A1_D0068.tif" />
computer to execute one of the methods described herein. The data transmission or the signal sequence can be configured, for example, to be transferred via a data communication connection, for example, via the Internet.
A further embodiment comprises a processing means, for example, a computer, or a programmable logic device, configured for or adapted to execute one of the methods described herein.
A further embodiment comprises a computer that has the computer program installed therein to execute one of the methods described herein.
In some embodiments, a programmable logic device (for example a field programmable composite arrangement) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, the field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by some hardware apparatus.
The embodiments described above are purely illustrative for the principles of the present invention. It is understood that the possible modifications and variations of the provisions and details described herein will be apparent to those skilled in the art. Therefore, it is the intention that the invention be limited only by the scope of the following patent claims and not by the specific details presented by the description and explanation of the embodiments herein.
— 65 —
<img file="AR080476A1_D0069.tif" />
LITERATURE:
[1] M. Dietz, L. Liljeryd, K. Kjörling and O. Kunz, Spectral Band Replication, a novel approach to audio coding (“Spectral Band Replication, a novel approach in audio coding”) at the 112th Convention AES,
Munich, May 2002.
[2] S. Meitzer, R. Böhm and F. Henn, Encoders — SBR enhanced audio decoders for digital broadcasting such as Digital Radio Mondiale (DRM) (“SBR enhanced audio coding for digital broadcasting such as“ Digital Radio Mondiale ” (DRM), ”) at the 112th AES Convention, Munich, May
2002.
[3] T. Ziegler, A. Ehret, P. Ekstrand and M. Lutzky, Improving mp3 with SBR: Features and Capabilities of the new mp3PRO Algorithm (“Enhancing mp3 with SBR: Features and Capabilities of the new mp3PRO Algorithm,”) at the 112th AES Convention, Munich, May 2002.
[4] International Standard ISO / IEC 14496-3: 2001 / FPDAM 1 “Bandwidth Extension ISO / IEC, 2002. (International Standard ISO / IEC 14496— 3: 2001 / FPDAM 1,“ Bandwidth Extension ”ISO / IEC , 2002.) Voice bandwidth extension method and apparatus Vasu lyengar et al ..
[5] Larsen, RM Aarts, and M. Danessis. Efficient high-frequency bandwidth extension of music and speech at the 112th AES convention, Munich, Germany, May 2002.
[6] RM Aarts, E. Larsen, and O. Ouweltjes. A unified approach to low and high frequency bandwidth extension, in The 115th AES Convention, New York, USA, October 2003.
— 66 —
<img file="AR080476A1_D0070.tif" />
<img file="AR080476A1_D0071.tif" />
[7] K. Käyhkö. A Robust Broadband Enhancement for Narrowband Voice Signal (A Robust Wideband Enhancement for Narrowband Speech Signal). Research report, University of Technology of Helsinki, Acoustics and Audio Processing Laboratory (Research Report,
Helsinki University of Technology, Laboratory of Acoustics and Audio Signal Processing), 2001.
[8] E. Larsen and RM Aarts. Audio Bandwidth Extension - Application to psychoacoustics, Signal Processing and Audio Bandwidth Extension - Application to Psychoacoustics, Serial Processing and Speaker Design
Loudspeaker Design). John Wiley & Sons, Ltd, 2004.
[9] Larsen, RM Aarts, and M. Danessis. Efficient high-frequency bandwidth extension of music and speech at the 112th AES convention, Munich, Germany, May 2002.
[10] J. Makhoul. Spectral Voice Analysis using Linear Prediction (Spectral Analysis of Speech by Linear Prédiction). IEEE Transactions on Audio and Electroacoustics, AU — 21 (3), June 1973.
[11] United States Patent Application No. 08 / 951.029, Ohmori, et al. Audio bandwidth extension system and method (12) US Patent No. 6895375, Malah, D & Cox, RV: Bandwidth Bandwidth Extension System narrow (System for bandwidth extension of Narrow — band speech).
[13] Frederik Nagel, Sascha Disch, A harmonic bandwidth extension method for audio encoder (“A harmonie bandwidth extension method for audio codées”), ICASSP International Congress - 67 -
<img file="AR080476A1_D0072.tif" />
on acoustics, voice and signal processing (International Conference on Acoustics, Speech and Signal Processing), IEEE CNF, Taipei, Taiwan, April 2009 [14] Frederik Nagel, Sascha Disch, Nikolaus Rettelbach, An extension method of band driven by phase Vocoder with a new treatment of transient components for audio codes. (A phase Vocoder driven bandwidth extension method with novel transient handling for audio codées, ”) 126th AES Convention, Munich, Germany, May 2009.
[15] M. Puckette. Synchronized phase vocoder. IEEE ASSP Congress on Senal Processing Applications in Audio and Acoustics. (Phase— locked Vocoder. IEEE ASSP Conference on Applications of Signal Processing to Audio and Acoustics), Mohonk 1995., A. Röbel,: Detection and preservation of transient components in the phase Vocoder. (Transient detection and preservation in the phase Vocoder,) citeseer.ist.psu.edu/679246.html [16] Laroche L., Dolson M .: Improved modification of the audio phase Vocoder time scale (“Improved phase Vocoder timescale modification of audio), IEEE Trans, on voice and audio processing (IEEE Trans. Speech and Audio Processing), vol. 7, no. 3, pp. 323-332, [17] US Patent 6,559,884, Laroche, J. & Dolson, M .: Phase tone shift — Vocoder (Phase — Vocoder pitch — shifting ”) [18] Herre, J .; Faller, C .; Ertel, C .; Hilpert, J .; Hölzer, A .; Spenger, C, MP3 Surround: Efficient and Compatible Coding of Multi-Channel Audio Signals (MP3 Surround: Efficient and Compatible Coding of Multi-Channel Audio), 116th Congress of the Society of Audio Engineers, May 2004 (116th Conv. Aud. Eng. Soc., May 2004) [19] Neuendorf, Max; Gournay, Philippe; Multrus, Markus; Lecomte, Jérémie; Kiss you, Bruno; Geiger, Ralf; Bayer, Stefan; Fuchs, Guillaume; Hilpert,
<img file="AR080476A1_D0073.tif" />
— 68 —
<img file="AR080476A1_D0074.tif" />
Johannes; Rettelbach, Nikolaus; Salami, Redwan; Schüller, Gerald; Lefebvre, Roch; Grill, Bernhard: Unified Speech and Audio Coding Scheme for High Quality at Lowbitrates, ICASSP 2009 (International Congress on Acoustic, Voice and Audio Processing) signal), April 19-24, 2009, Taipei, Taiwan [20] Bayer, Stefan; Kiss you, Bruno; Fuchs, Guillaume; Geiger, Ralf; Gournay, Philippe; Grill, Bernhard; Hilpert, Johannes; Lecomte, Jérémie; Lefebvre, Roch; Multrus, Markus; Nagel, Frederik; Neuendorf, Max; Rettelbach, Nikolaus;
Robilliard, Julien; Salami, Redwan; Schüller, Gerald: A Novel Scheme for Unified Voice and Audio Coding with a Low Number of Transmitted Bits (A Novel Scheme for Low Bitrate Unified Speech and Audio Coding), at the 126th AES convention, Munich, Germany, May 7 of 2009.
Contents9
98 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98
81 members in 19 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 31212710 | United States of America | P | |
| 2011053313 | European Patent Office (EPO) | W |
Members81
| Document | Office | Kind | |
|---|---|---|---|
| CA2792450A1 | Canada | A1 | |
| CA2792452A1 | Canada | A1 | |
| WO2011110499A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2011110500A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201207841A | Taiwan Province of China | A | |
| TW201207842A | Taiwan Province of China | A | |
| AR080476A1This record | Argentina | A1 | |
| AR080477A1 | Argentina | A1 | |
| MX2012010415A | Mexico | A | |
| AU2011226211A1 | Australia | A1 | |
| AU2011226212A1 | Australia | A1 | |
| SG183967A1 | Singapore | A1 | |
| MX2012010416A | Mexico | A | |
| KR20120131206A | Republic of Korea | A | |
| KR20120139784A | Republic of Korea | A | |
| EP2545548A1 | European Patent Office (EPO) | A1 | |
| EP2545553A1 | European Patent Office (EPO) | A1 | |
| CN102939628A | China | A | |
| US2013051571A1 | United States of America | A1 | |
| CN103038819A | China | A | |
| US2013090933A1 | United States of America | A1 | |
| JP2013521538A | Japan | A | |
| JP2013525824A | Japan | A | |
| HK1181180A | Hong Kong, China | A | |
| HK1181180A1 | Hong Kong, China | A1 | |
| AU2011226211B2 | Australia | B2 | |
| AU2011226212B2 | Australia | B2 | |
| RU2012142732A | Russian Federation | A | |
| JP5523589B2 | Japan | B2 | |
| TWI444991B | Taiwan Province of China | B | |
| TWI446337B | Taiwan Province of China | B | |
| EP2545553B1 | European Patent Office (EPO) | B1 | |
| KR101414736B1 | Republic of Korea | B1 | |
| KR101425154B1 | Republic of Korea | B1 | |
| JP5588025B2 | Japan | B2 | |
| ES2522171T3 | Spain | T3 | |
| PL2545553T3 | Poland | T3 | |
| CN103038819B | China | B | |
| CN102939628B | China | B | |
| MY154204A | Malaysia | A | |
| US9305557B2 | United States of America | B2 | |
| CA2792450C | Canada | C | |
| RU2586846C2 | Russian Federation | C2 | |
| US2017194011A1 | United States of America | A1 | |
| US9792915B2 | United States of America | B2 | |
| CA2792452C | Canada | C | |
| US10032458B2 | United States of America | B2 | |
| US2018366130A1 | United States of America | A1 | |
| EP3570278A1 | European Patent Office (EPO) | A1 | |
| US2020279571A1 | United States of America | A1 | |
| US10770079B2 | United States of America | B2 | |
| BR112012022740A2 | Brazil | A2 | |
| BR112012022574A2 | Brazil | A2 | |
| BR112012022740B1 | Brazil | B1 | |
| BR122021019078B1 | Brazil | B1 | |
| BR112012022574B1 | Brazil | B1 | |
| BR122021014305B1 | Brazil | B1 | |
| BR122021019082B1 | Brazil | B1 | |
| BR122021014312B1 | Brazil | B1 | |
| EP3570278B1 | European Patent Office (EPO) | B1 | |
| US11495236B2 | United States of America | B2 | |
| ES2935637T3 | Spain | T3 | |
| US2023074883A1 | United States of America | A1 | |
| EP4148729A1 | European Patent Office (EPO) | A1 | |
| PL3570278T3 | Poland | T3 | |
| US11894002B2 | United States of America | B2 | |
| US2024135939A1 | United States of America | A1 | |
| EP4475124A2 | European Patent Office (EPO) | A2 | |
| EP4475124A3 | European Patent Office (EPO) | A3 | |
| EP4148729B1 | European Patent Office (EPO) | B1 | |
| EP4148729C0 | European Patent Office (EPO) | C0 | |
| ES3010370T3 | Spain | T3 | |
| US12308036B2 | United States of America | B2 | |
| PL4148729T3 | Poland | T3 | |
| HUE070311T2 | Hungary | T2 | |
| EP4475124B1 | European Patent Office (EPO) | B1 | |
| EP4475124C0 | European Patent Office (EPO) | C0 | |
| EP4661004A2 | European Patent Office (EPO) | A2 | |
| EP4661004A3 | European Patent Office (EPO) | A3 | |
| ES3058766T3 | Spain | T3 | |
| PL4475124T3 | Poland | T3 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant, registrationFG | FG |
Numbers
- Application
- 110100723
Titles2
- English
- APPLIANCE AND METHOD FOR PROCESSING AN AUDIO SIGNAL USING PATCHING EDGE ALIGNMENT
- Spanish
- APARATO Y METODO PARA PROCESAR UNA SENAL DE AUDIO USANDO ALINEACION DE BORDE DE PATCHING
Classification
- CPC, 5
- G10L19/0204
- G10L21/0232
- G10L19/008
- G10L21/038
- G10L21/04
- IPC, 3
- G10L19 02
- G10L21 038
- G10L21 04