Bandwidth extension method, bandwidth extension apparatus, program, integrated circuit, and audio decoding apparatus
Abstract
This record has no abstract on file.
Term
4.7 yearsto projected expiry
Projected expiry 6 June 2031, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
8 claims: 4 independent, 4 dependent
- 1Zastrzeżenia patentowe 1. Sposób rozszerzania pasma częstotliwości w celu generowania pełnopasmowego sygnału audio z sygnału audio w paśmie niskiej częstotliwości, przy czym wymieniony sposób obejmuje:pierwszy krok (S21) transformacyjny polegający na przekształcaniu sygnału w paśmie niskiej częstotliwości do dziedziny zespołu kwadraturowych filtrów lustrzanych (QMF), aby generować pierwsze widmo QMF niskiej częstotliwości;krok (S22) generowania wstawki harmonicznej niskiego rzędu poprzez rozciąganie w czasie sygnału w paśmie niskiej częstotliwości w dziedzinie QMF;krok (S23) generowania wysokiej częstotliwości obejmujący (i) generowanie sygnałów, które są przesunięte pod względem wysokości tonu, poprzez zastosowanie różnych współczynników przesunięcia do wstawki harmonicznej niskiego rzędu, oraz (ii) generowanie widma QMF wysokiej częstotliwości z sygnałów;krok (S24) modyfikacji widma polegający na modyfikowaniu widma QMF wysokiej częstotliwości, aby spełnić warunki energii wysokiej częstotliwości i tonalności;oraz krok (S25) generowania pełnopasmowego widma polegający na generowaniu pełnopasmowego sygnału poprzez łączenie zmodyfikowanego widma QMF wysokiej częstotliwości z pierwszym widmem QMF niskiej częstotliwości.
- 2Sposób rozszerzania pasma częstotliwości według zastrzeżenia 1, w którym ten krok generowania wstawki harmonicznej niskiego rzędu obejmuje:drugi krok transformacyjny polegający na przekształcaniu sygnału w paśmie niskiej częstotliwości na drugie widmo QMF niskiej częstotliwości, gdzie to drugie widmo QMF niskiej częstotliwości ma wyższą rozdzielczość niż pierwsze widmo QMF niskiej częstotliwości.
- 3Sposób rozszerzania pasma częstotliwości według jednego z zastrzeżeń 1 i 2, w którym wymieniony krok generowania wysokiej częstotliwości obejmuje:krok generowania wstawki polegający na przetwarzaniu pasmowym wstawki harmonicznej niskiego rzędu, aby generować przetworzone pasmowo wstawki;krok generowania wysokiego rzędu polegający na odwzorowywaniu każdej z przetworzonych pasmowo wstawek do wysokiej częstotliwości, aby generować wstawki harmoniczne wysokiego rzędu;oraz krok sumowania polegający na sumowaniu wstawek harmonicznych wysokiego rzędu ze wstawkami harmonicznymi niskiego rzędu.
- 4Sposób rozszerzania pasma częstotliwości według zastrzeżenia 3, w którym wymieniony krok generowania wysokiego rzędu obejmuje:krok rozdzielania polegający na dzieleniu każdego podpasma QMF w każdej z przetworzonych pasmowo wstawek na wiele podpasm;krok odwzorowywania polegający na odwzorowywaniu podpasm do podpasm QMF wysokiej częstotliwości;oraz krok łączenia polegający na łączeniu wyników odwzorowywania podpasm.
- 5Sposób rozszerzania pasma częstotliwości według zastrzeżenia 4, w którym wymieniony krok odwzorowywania obejmuje:krok dzielenia polegający na dzieleniu podpasm każdego z podpasm QMF na zaporową część pasma i przepustową część pasma;krok obliczania częstotliwości polegający na obliczaniu transponowanych środkowych częstotliwości podpasm na przepustowej części pasma ze współczynnikiem zależnym od rzędu wstawki;pierwszy krok odwzorowywania polegający na odwzorowywaniu podpasm na przepustową część pasma do podpasm QMF wysokiej częstotliwości zgodnie z częstotliwościami środkowymi;oraz drugi krok odwzorowywania polegający na odwzorowywaniu podpasm na zaporowej części pasma do podpasm QMF wysokiej częstotliwości zgodnie z podpasmami przepustowej części pasma.
- 6Urządzenie do rozszerzania pasma częstotliwości, które wytwarza pełnopasmowy sygnał audio z sygnału audio w paśmie niskiej częstotliwości, przy czym wymienione urządzenie do rozszerzania pasma częstotliwości zawiera:pierwszą jednostkę przekształcającą (1503) skonfigurowaną do przekształcania sygnału w paśmie niskiej częstotliwości do dziedziny zespołu kwadraturowych filtrów lustrzanych (QMF), aby generować pierwsze widmo QMF niskiej częstotliwości;jednostkę (1504) generującą wstawki harmoniczne niskiego rzędu skonfigurowaną do generowania wstawki harmonicznej niskiego rzędu poprzez rozciąganie w czasie sygnału w paśmie niskiej częstotliwości w dziedzinie QMF;jednostkę (1506) generowania wysokiej częstotliwości skonfigurowaną do (i) generowania sygnałów, które są przesunięte pod względem wysokości tonu, poprzez stosowanie różnych współczynników przesunięcia do wstawki harmonicznej niskiego rzędu, oraz (ii) do generowania widma QMF wysokiej częstotliwości z sygnałów;jednostkę (1507) modyfikacji widma skonfigurowaną do modyfikowania widma QMF wysokiej częstotliwości, aby spełnić warunki energii wysokiej częstotliwości i tonalności;oraz pełnopasmową jednostkę generującą (1512) skonfigurowaną do generowania pełnopasmowego sygnału poprzez łączenie zmodyfikowanego widma QMF wysokiej częstotliwości z pierwszym widmem QMF niskiej częstotliwości.
- 7Układ scalony wytwarzający pełnopasmowy sygnał audio z sygnału audio w paśmie niskiej częstotliwości, zawierający:pierwszą jednostkę przekształcającą skonfigurowaną do przekształcania sygnału w paśmie niskiej częstotliwości do dziedziny zespołu kwadraturowych filtrów lustrzanych (QMF), aby generować pierwsze widmo QMF niskiej częstotliwości;jednostkę generującą wstawki harmoniczne niskiego rzędu skonfigurowaną do generowania wstawki harmonicznej niskiego rzędu poprzez rozciąganie w czasie sygnału w paśmie niskiej częstotliwości w dziedzinie QMF;jednostkę generowania wysokiej częstotliwości skonfigurowaną do (i) generowania sygnałów, które są przesunięte pod względem wysokości tonu, poprzez stosowanie różnych współczynników przesunięcia do wstawki harmonicznej niskiego rzędu, oraz (ii) do generowania widma QMF wysokiej częstotliwości z sygnałów;jednostkę modyfikacji widma skonfigurowaną do modyfikowania widma QMF wysokiej częstotliwości, aby spełnić warunki energii wysokiej częstotliwości i tonalności;oraz pełnopasmową jednostkę generującą skonfigurowaną do generowania pełnopasmowego sygnału poprzez łączenie zmodyfikowanego widma QMF wysokiej częstotliwości z pierwszym widmem QMF niskiej częstotliwości.
- 8Urządzenie dekodujące sygnał audio zawierające:jednostkę rozdzielającą skonfigurowaną do oddzielania kodowanego sygnału w paśmie niskiej częstotliwości od zakodowanej informacji;jednostkę dekodującą skonfigurowaną do dekodowania kodowanego sygnału w paśmie niskiej częstotliwości;jednostkę przekształcającą skonfigurowaną do przekształcania sygnału w paśmie niskiej częstotliwości generowanego w drodze dekodowania przez wymienioną jednostkę dekodującą, do dziedziny zespołu kwadraturowych filtrów lustrzanych (QMF), aby generować widmo QMF niskiej częstotliwości;jednostkę generującą wstawki harmoniczne niskiego rzędu skonfigurowaną do generowania wstawki harmonicznej niskiego rzędu poprzez rozciąganie w czasie sygnału w paśmie niskiej częstotliwości w dziedzinie QMF;jednostkę generowania wysokiej częstotliwości skonfigurowaną do (i) generowania sygnałów, które są przesunięte pod względem wysokości tonu, poprzez stosowanie różnych współczynników przesunięcia do wstawki harmonicznej niskiego rzędu, oraz (ii) do generowania widma QMF wysokiej częstotliwości z sygnałów;jednostkę modyfikacji widma skonfigurowaną do modyfikowania widma QMF wysokiej częstotliwości, aby spełnić warunki energii wysokiej częstotliwości i tonalności;pełnopasmową jednostkę generującą skonfigurowaną do generowania pełnopasmowego sygnału poprzez łączenie zmodyfikowanego widma QMF wysokiej częstotliwości z widmem QMF niskiej częstotliwości;oraz jednostkę przekształcania odwrotnego skonfigurowaną do przekształcania pełnopasmowego sygnału, z sygnału dziedziny zespołu kwadraturowych filtrów lustrzanych (QMF) na sygnał dziedziny czasu. Panasonic Intellectual Property Corporation of America Pełnomocnik: EP 2 581 905 B1 Rysunek parametr HF CG Cl ο Αϊ Ο ω Ν ω 2 ΕΞ CG £ ο “Ρ ο ΑΤ ω ρ CG ra c co . PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 anie bloków audio i dodawanie ich na za PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 1100 1200 1300 1400 1500 1600 1700 lfloo 1900 2000 częstotliwość PL-PAT-2012-572 EP 2 581 905 B1 na bazie QMF x 4 PL-PAT-2012-572 EP 2 581 905 B1 S22 Generowanie wstawki harmonicznej niskiego rzędu poprzez rozciąganie w czasie w dziedzinie QMF 823 Przesuwanie wysokości tonu wstawki harmonicznej niskiego rzędu na podstawie współczynnika przesunięcia, aby generować widmo QMF wysokieji częstotliwości FIG. 12 Koniec Przekształcanie QMF sygnału w paśmie niskiej częstotliwości, aby generować pierwsze widmo QMF niskiej częstotliwości Ξ24 Modyfikowanie widma QMF wysokieji częstotliwości 325 Łączenie widma QMF wysokiej częstotliwości z pierwszym widmem QMF niskiej częstotliwości PL-PAT-2012-572 EP 2 581 905 B1 χ Τ/2 PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572 EP 2 581 905 B1 (θΡ) Epni!|dujE częstotliwość normalizowana ix ™ rad/próbka) PL-PAT-2012-572 EP 2 581 905 B1 PL-PAT-2012-572
Independent claims8
230 paragraphs in 2 sections, as filed
Technical field] [0001] The present invention relates to a method of extending the frequency band to extend the frequency band of an audio signal.
[Background of the invention] [0002] Audio frequency bandwidth (BWE) technique is usually used in modern audio codecs to efficiently encode a wideband audio signal at a low data rate. Its principle is to use the parametric representation of the original high frequency (HF) content to synthesize the HF approximation from the low frequency (LF) data.
[0003] ZHOU HUAN AND IN: "Core Experiment on the eSBR module of USAC", 90. MPEG MEETING, ISO / IEC JTC1 / SC29 / WG11, MPEG2009 / M16933, October 2009, discloses a method of extending the frequency band to generate full bandwidth signal from a low frequency signal.
[0004] FIG. 1 is a diagram showing such an audio codec based on BWE technology. In his encoder, the broadband audio signal is first split (101 & 103) into the LF part and the HF part; its LF portion is encoded (104) in a manner that maintains the shape of a wave; during this time, the relationship between its LF and HF parts is analyzed (102) (usually in the frequency domain) and described by the HF parameter set. Thanks to the parametric description of the HF parts, multiplexed (105) waveform data and HF parameters can be sent to the decoder at a low transfer rate.
[0005] In the decoder, the LF portion is first decoded (107). To approximate the original HF portion, the decoded LF portion is converted (108) to the frequency domain, and the resulting LF spectrum is modified (109) to generate the HF spectrum, with some decoded HF parameters being the leading elements. The HF spectrum is further purified (110) by post-processing, also taking some decoded HF parameters as the leading. The purified spectrum is processed (111) into the time domain and combined (112) with the delayed portion of LF. As a result, the final reproduced broadband audio signal is output.
[0006] It should be noted that in BWE technology one important step is to generate the HF spectrum from the LF spectrum (109). There are several ways to implement it, such as copying the LF part to the location of the HF, non-linear processing or increasing the sampling frequency (so-called upsampling).
[0007] The best-known audio codec using such BWE technology is MPEG-4 HE-AAC, where BWE technology is referred to as SBR (spectral band replication) or SBR technology, where the HF part is generated by simple copying the LF part in the QMF representation to the spectral location of the HF.
[0008] This spectral copying operation, also referred to as "patching", is simple and proves to be effective in most cases. However, at very low transfer rates (e.g. <20kbits / s mono), where only small parts of the LF bandwidth are feasible, this SBR technology can lead to unwanted sensations of acoustic artifacts, such as roughness and unpleasant timbre (see e.g. Patent (NPL) 1).
[0009] Therefore, to avoid such artifacts resulting from mirroring or copying operations occurring in the coding scenario at a low bit rate, the standard SBR technology is improved and expanded with the following major changes (see e.g. NPL 2):
(1) modifying the insertion algorithm from the copy model to the phase vocoder-driven insertion model, (2) increasing the adaptive time resolution for the post-processing parameters.
[0010] As a result of the first modification ((1), see above), by extending the LF spectrum by a number of integer factors, harmonic continuity in HF is generally ensured. In particular, no unwanted roughness sensations due to rumbling phenomena may appear at the border between low frequency and high frequency, and between different parts of high frequency (see NPL 1).
[0011] And the second modification ((2), see above) facilitates the purified HF spectrum to be more adaptive with respect to signal fluctuations in replicated frequency bands.
[0012] Since the new insertion method maintains a harmonic relationship, it is referred to as harmonic bandwidth extension (HBE). The advantages of the known HBE method over the standard SBR method have also been experimentally confirmed for audio coding at a low bit rate (see e.g. NPL 1).
[0013] It should be noted that both of the above modifications only affect the HF spectrum generator (109), while the other processes in HBE are identical to those in SBR.
[0014] Fig. 2 is a schematic diagram of the HF spectrum generator in known HBE. Note that the HF spectrum generator includes the TF transformation module 108 and the HF reproduction module 109. Assuming the LF part of the signal, let's assume that its HF spectrum consists of (T-1) HF harmonic inserts (each insert creation process creates one HF insert), from the 2nd order (the lowest frequency HF insert) to the T-order ( highest frequency HF insert). In known HBE, all these HF inserts are generated independently in parallel from phase vocoders. As shown in Fig. 2, (T-1) phase vocoders (201 ~ 203) with different stretch ratios (from 2 to k) are used to stretch the input part LF. The stretched output portions, of various lengths, are filtered in bandpass filters (204 ~ 206) and resampled (207 ~ 209) to generate HF inserts by converting time-extension to frequency-extension. By setting a stretch ratio twice that of the resampling factor, the HF inserts maintain the harmonic signal structure and have double the length of the LF part. Then all HF inserts are aligned in terms of delay (210 ~ 212) to compensate for potential different delays in the resampling operation. In the last step, all delay-tuned HF inserts are added together and transformed (213) to the QMF domain to generate the HF spectrum.
[0016] Observing the above HF spectrum generator, it is found that it requires a large amount of calculations. The number of calculations results mainly from the time-stretching operation, carried out by means of the STFT (Short Time Fourier Transform) and Inverse Short Time Fourier Transform (ISTFT) transformations used in phase vocoders and the subsequent QMF operation, applied to the stretched part HF.
[0017] The general introduction to phase vocoder and QMF transformation is described below.
[0018] The phase vocoder is a well-known technique that uses frequency-domain transformations to introduce a stretching effect over time. This means modifying the signal's development over time, while its local spectral parameters remain unchanged. Its basic principles are described below.
[0019] Figs. 3A and 3B show the basic principle of stretching over time performed in a phase vocoder.
[0020] Divide the audio into overlapping blocks and redeploy these blocks, the distance in time (so-called hop size) (the time interval between successive blocks) is not the same at the input and output as shown in Fig. 3A. Thus, the input distance Ra is smaller than the output distance Rs, as a result of which the original signal is extended in the degree r shown in equation 1 below.
[0021] [Calculation 1] r = - (equation 1)
R<sub>s</sub> [0022] As shown in Fig. 3B, the rearranged blocks overlap in a coherent manner that requires frequency domain transformation. Usually, input blocks are converted to frequency, and after proper phase modification, new blocks are converted back to output blocks.
[0023] According to the above principle, the most classic phase vocoders adopt the STF Fourier transformation as the frequency domain transformation and introduce an explicit string of analysis, modification and re-synthesis for time extension.
[0024] QMFs transform time-domain representations into combined time-frequency representations (and vice versa), which is typically used in parameter-based coding schemes, such as SBR (spectral band replication), parametric coding stereo (PS) and spatial audio coding (spatial audio coding, SAC) etc. A characteristic feature of these filter assemblies is that signals with complex domain values (frequency subbands) undergo effective oversampling with a factor of two. This enables post-processing of subband domain signals without introducing aliasing signal distortions.
[0025] More specifically, assuming the discontinuous real-time signal x (n) with a QMF analysis unit, the signals sk (n) with complex values of the subband domain are obtained from Equation 2 below.
[0026] [Calculation 2] p<sub>k</sub>(n) Φ Σ χ (Μ η - 0p (0e<sup>7</sup>M 1 = 0 (fc + 0.5) (Z + zz) (equation 2) [0027] In equation 2, p (n) represents the L-1 impulse response of the prototype low-pass filter, α represents the phase parameter, M represents the number bands, while k is the subband indicator at k = 0, 1, ..., M-1).
[0028] It should be noted that, like STFT, the QMF transformation is also a combined time-frequency transformation. This means that it provides both the frequency content of the signal and the frequency change of the content over time, with the frequency content representing, respectively, the frequency subband and time representing the time window.
[0029] Fig. 4 is a diagram showing the course of QMF analysis and synthesis.
[0030] In detail, as shown in Fig. 4, the actual audio input data is divided into successive overlapping blocks of length L and distance in time equal to M (Fig. 4 (a)), and the QMF analysis process transforms each block for one time window consisting of M complex subband signals. In this way, L time-domain input samples are converted to L complex QMFs, consisting of L / M time windows and M subbands (Fig. 4 (b)). Each time window, in combination with the previous (L / M-1) time windows, is synthesized in the QMF synthesis process to reproduce M real-time domain samples (Fig. 4 (c)) with almost perfect reproduction.
[List of cited documents] [Non-patent literature] [0031] [NPL 1] Frederik Nagel and Sascha Disch, 'A harmonic bandwidth extension method for audio codecs', IEEE Int. Conf. On Acoustics, Speech and Signal Proc., 2009 [ NPL 2] Max Neuendorf, et al., 'A novel scheme for low bitrate unified speech and audio coding MPEG RM0', in 126th AES Convention, Munich, Germany, May 2009.
[Summary of the invention] [Technical problem] [0032] The problem associated with known HBE technology is the large amount of calculations. The traditional phase vocoder used in HBE to stretch the signal has more calculations, because subsequent FFT and IFFT are used, i.e. subsequent FFT (fast Fourier transforms) and IFFT (inverted fast Fourier transforms); and the next QMF transformation increases the number of calculations by applying it to the time-extended signal. In addition, in general, attempts to reduce the number of calculations lead to a potential quality degradation problem.
[0033] Thus, the present invention as defined in the claims has been developed taking into account the above problem, and its purpose is to provide a method for expanding the frequency band which can reduce the amount of calculations in the process of extending the frequency band as well as eliminating deterioration in the extended frequency band.
[Solution to the problem] [0034] To achieve the above goal, the bandwidth extension method in an aspect of the present invention is a bandwidth extension method to generate a full bandwidth audio signal from an low frequency band audio signal, the method comprising: a first transformational step consisting in converting a low-frequency signal in the field of the QMF (quadrature mirror filter) to generate the first low-frequency QMF spectrum; the step of generating low order harmonic inserts by generating the low order harmonic insert by stretching the signal in the low frequency band in the QMF domain; the high frequency generation step of (i) generating signals that are shifted in terms of pitch by applying different shift coefficients to the low order harmonic insert, and (ii) generating high frequency QMF spectrum from the signals; the spectrum modification step of modifying the high frequency QMF spectrum to meet high frequency energy and tonality conditions; the step of generating the full bandwidth of generating a full bandwidth signal by combining the modified high frequency QMF spectrum with the first low frequency QMF spectrum.
[0035] Accordingly, the high frequency QMF spectrum is generated by stretching in time and shifting the pitch of the signal in the QMF domain. It is therefore possible to avoid the classic complex processing (repeated FFT and IFFT transformations followed by QMF transformation), to generate high frequency QMF spectrum, and thus the amount of calculations can be reduced. In addition, because the pitch-shifted signals are generated by using mutually shifted different shift factors instead of just one shift factor, and the high-frequency QMF spectrum is generated from these signals, the high-quality QMF spectrum degradation can be eliminated. In addition, because the high frequency QMF spectrum is generated from a lower order harmonic insert, the quality degradation of the high frequency QMF spectrum can further be eliminated.
[0036] It should be noted that in the frequency band expansion method, the pitch offset also works in the QMF domain. This is to split the LF QMF subband on the low order insert into multiple subbands for higher frequency resolution, and then map these subbands to the high QMF subband to generate in and out the high order insert.
[0037] Furthermore, the step of generating the low order harmonic insert includes: a second transformation step of converting the low frequency band signal into a second low frequency QMF spectrum; a step of processing the second low frequency QMF spectrum; a stretching step of stretching the band-processed second low frequency QMF spectrum in a time dimension.
[0038] Furthermore, the second low frequency QMF spectrum has a higher frequency resolution than the first low frequency QMF spectrum.
[0039] In addition, the high frequency generation step includes: an insert generation step of processing low order harmonic band to generate band inserts; the high order generation step of mapping each of the high frequency band elements to generate high order harmonic inserts; and the summation step of summing the high order harmonics with the lower order harmonics.
[0040] In addition, the high order generation step includes: a separation step of separating each QMF subband in each of the band-processed inserts into multiple subbands; step of mapping the subbands to the high frequency QMF subbands; and a combining step of combining the subband mapping results.
[0041] In addition, the mapping step includes: a dividing step of dividing the subbands of each of the QMF subbands into the barrier portion of the band and the bandwidth portion of the band; a frequency calculation step of calculating the transposed center subband frequencies for a portion of the bandwidth with a factor depending on the insert order; a first mapping step of mapping the subbands on a portion of the bandwidth to the high frequency QMF subbands according to the center frequencies; and a second step of mapping the subbands on the portion of the bandwidth to the high frequency QMF subbands according to the subbands of the portion of the bandwidth.
[0042] It should be noted that in the method of extending the frequency band according to the present invention, the process operations (steps) described previously can be combined in any way.
[0043] Such a method of extending the frequency band as the method of the present invention is a low-computing HBE technology that uses a low-computing HF spectrum generator that introduces the largest amount of calculations into HBE. To reduce the number of calculations, a new QMF based phase vocoder is used which performs stretching in time in the QMF domain with a small amount of calculations. In addition, to avoid possible quality problems associated with the solution, a new pitch shift algorithm is used that generates high order harmonic inserts from low order inserts in the QMF domain.
[0044] It is an object of the present invention to design a QMF based insert in which both time stretching and frequency extension can be done in the QMF domain to go further, developing a low-computing HBE technology driven by a QMF-based phase vocoder .
[0045] It should be noted that the present invention can be implemented not only as such a frequency band extension method but also as a frequency band extension device and an integrated circuit that extends the frequency band of the audio signal using the frequency band extension method.
[Advantageous Effects of the Invention] [0046] The frequency band expansion method in the present invention designs a new harmonic frequency band extension (HBE) technology. The essence of the technology is to perform both stretching in time and shifting the pitch in the QMF domain instead of the traditional FFT and time domain respectively. Compared to known HBE technology, the method of extending the frequency band in the present invention can provide good sound quality and significantly reduces the amount of calculations.
[Brief description of the drawings] [0047] [FIG. 1] Fig. 1 is a diagram of the audio codec using ordinary BWE technology.
[FIG. 2] Fig. 2 is a diagram showing a spectrum generator with preserved HF harmonic structure.
[FIG. 3A] Fig. 3A is a diagram showing the principle of stretching over time by changing the position of audio blocks.
[FIG. 3B] Fig. 3B is a diagram showing the principle of stretching over time by changing the position of audio blocks.
[FIG. 4] Fig. 4 is a diagram showing QMF analysis and synthesis diagram.
[FIG. [Fig. 5] Fig. 5 is a block diagram of a method for extending the frequency band in the first example.
[FIG. [Fig. 6] Fig. 6 is a diagram showing the HF spectrum generator in the first example.
[FIG. [Fig. 7] Fig. 7 is a diagram showing an audio decoder in the first example.
[FIG. [Fig. 8] Fig. 8 is a diagram showing a diagram of a signal time change based on the QMF transformation in the first example.
[FIG. [Fig. 9] Fig. 9 is a diagram showing the method of stretching over time in the QMF domain in the first example.
[FIG. [Fig. 10] Fig. 10 is a diagram showing a comparison of stretching results for a sinusoidal audio signal with different stretching ratios.
[FIG. 11] The figure shows a diagram of the phenomenon of energy shift and dispersion in the HBE diagram. [FIG. [Fig. 12] Fig. 12 is a block diagram of a method of extending a frequency band in the form of the present invention.
[FIG. [Fig. 13] Fig. 13 is a diagram showing an HF spectrum generator in the form of the present invention.
[FIG. [Fig. 14] Fig. 14 is a diagram showing an audio decoder in the form of the present invention.
[FIG. [Fig. 15] Fig. 15 is a diagram showing a method of extending the frequency band in the QMF domain in the form of the present invention.
[FIG. [Fig. 16] Fig. 16 shows the distribution of subband spectra in an embodiment of the present invention.
[FIG. [Fig. 17] Fig. 17 is a diagram showing the relationship between the bandwidth component and the barrier band component for a sinusoidal signal in the complex domain of QMF in the embodiment of the present invention.
[Description of the form] [0048] The following forms are merely illustrative of the principles of the individual inventive steps. It should be understood that modifications to the details described herein will be clear to those skilled in the art.
(First example) [0049] Hereinafter, the HBE scheme (harmonic frequency band expansion method) and decoder (audio decoder or audio decoding device) using this method will be described.
[0050] Fig. 5 is a block diagram of a method for extending a frequency band.
[0051] This method of expanding the frequency band is a method of expanding the frequency band to generate a full-range signal from a signal in the low frequency band, which method comprises: (quadrature mirror filter) to generate the first low frequency QMF spectrum; a pitch shift step consisting in generating tone shifted signals by applying different shift coefficients to a signal in the low frequency band; the high frequency generation step of generating the high frequency QMF spectrum by stretching the pitch shifted signals in the QMF domain; the spectrum modification step of modifying the high frequency QMF spectrum to meet high frequency energy and tonality conditions; and the full-band generation step of generating the full-band signal by combining the modified high frequency QMF spectrum with the first low frequency QMF spectrum.
[0052] It should be noted that the first transformation step (S11) is performed by the TF transformation unit 1406 described below, the pitch shift step (S12) is performed by the sampling units 504 to 506 and the resampling unit 1403 described in continue. In addition, the high frequency generation step (S13) is performed by QF transformation units 507 to 509, phase vocoders 510 to 512, QMF transformation unit 404, and time stretching unit 1405, described later. Furthermore, the step (S15) for generating the full band signal is performed by the adding unit 1410 described below.
[0053] In addition, the high frequency generation step includes: a second transformation step of converting the shifted pitch signals into the QMF domain to generate QMF spectra; the step of generating a harmonic insert consisting in extending the QMF spectra in time dimension with different stretching factors to generate harmonic inserts; setting step consisting of setting harmonic inserts over time; and the summation step consisting of summing the harmonics set in time.
[0054] It should be noted that the second transformation step is performed by the QMF transforming units 507 to 509 and the QMF transforming unit 1404, while the harmonic insertion generation step is performed by phase vocoders 510 to 512 and the 1405 stretching unit. Furthermore, the setting step is performed by the delay units 513 to 515 described later, and the adding step is carried out by the adding unit 516 described below.
[0055] In the HBE scheme in the present embodiment, the HF spectrum generator in HBE technology is designed with pitch-shifting processes in the time domain followed by vocoder-controlled stretching processes in time in the QMF domain.
[0056] Fig. 6 is a schematic of the HF spectrum generator used in the HBE scheme. The HF spectrum generator includes: bandpass units 501, 502, ..., and 503; sampling units 504, 505, ..., and 506; QMF conversion units 507, 508, ..., and 509; phase vocoders 510, 511, ..., and 512; units 513, 514, ..., and 515 controlling the delay and adding unit 516.
[0057] A given input signal in the LF frequency band is first band-filtered (501 ~ 503) and sampled (504 ~ 506) to generate its HF portions of the frequency band. These HF portions of the frequency band are converted (507 ~ 509) in the QMF domain, and the resulting QMF output signals are stretched in time (510 ~ 512) with stretch ratios twice that of the resampling ratios. The extended HF spectra are set for delay (513 ~ 515) to compensate for potential different delays from the resampling process and summed (516) to generate the final HF spectrum. It should be noted that each of the above reference numbers 501 to 516 in parentheses indicates the component of the HF spectrum generator.
[0058] Comparing the scheme with known schemes (Figure 2), it can be seen that the main differences are 1) using more QMF transformations; and 2) the time-stretch operation is performed in the QMF domain and not in the FFT domain. A detailed description of time stretching in the QMF domain will be described in more detail below.
[0059] Fig. 7 is a schematic diagram of a decoder using an HF spectrum generator. The decoder (audio decoding device) includes a demultiplexer 1401, a decoding unit 1402, a sampling unit 1403 in time, a transformer unit QMF 1404 and a stretching unit 1405. Note that the demultiplexer 1401 corresponds to a separation unit that separates the encoded low-frequency signal from the encoded information (bit stream). Furthermore, the TF inverse transformation unit 1409 corresponds to the inverse transformation unit that converts the full-range signal from the time domain signal (QMF) of the quadrature mirror filter assembly to the time domain signal.
[0060] In the decoder, the bit stream is first demultiplexed (1401), then the LF part of the signal is decoded (1402). To approximate the original HF portion, the decoded LF portion (low frequency signal) is sampled (1403) in the time domain to generate the HF portion, and the resulting HF portion is converted (1404) to the QMF domain, the resulting HF QMF spectrum extends ( 1405) in the time direction, the extended HF spectrum is further purified (1408) by post-processing, based on some decoded HF parameters. At the same time, the decoded portion of LF is also converted (1406) to the QMF domain. Finally, the purified HF spectrum is combined (1410) with the delayed (1407) LF spectrum to produce a full-spectrum QMF spectrum. The resulting full-band QMF spectrum transforms (1409) back to the time domain to output a wide-band audio signal. It should be noted that each of the above reference numerals 1401 to 1410 in parentheses denotes a component element of the decoder.
The method of stretching over time [0061] The HBE process is, in the case of an audio signal, its time-extended signal that can be generated by QMF transformation, phase manipulation and inverse QMF transformation. More specifically, the harmonic insert generation step includes: a step of calculating the amplitude and phase of the QMF spectrum among the QMF spectra; a phase manipulation step of manipulating the phase to create a new phase; and the step of generating the QMF coefficient by combining the amplitude with the new phase to generate a new set of QMF coefficients. It should be noted that each of the steps, namely the calculation step, the phase manipulation step and the step of generating the QMF coefficient is carried out by the module 702 described below.
[0062] Fig. 8 is a diagram showing a QMF-based stretching process over time performed by the QMF 1404 conversion unit and the 1405 stretching unit over time. First, the audio signal is converted into a set of QMF coefficients, say X (m, n), by way of analytical transformation QMF (701). These QMFs are modified in module 702.
At the same time, for each of the QMF coefficients, its amplitude r and phase a are calculated, say X (m, n) = r (m, n) ^ exp (j ^ a (m, n)). Phases a (m, n) are modified (manipulated) to a<sup>~</sup>(m, n). Modified phases a<sup>~</sup> and the original amplitudes r form a new set of QMF coefficients. For example, a new set of QMF coefficients is shown in Equation 3 below.
[0063] [Calculations 3] (equation 3) [0064] Finally, this new set of QMF coefficients is converted (703) into a new audio signal corresponding to the original audio signal with a modified time scale.
[0065] The QMF based time stretching algorithm in the HBE scheme mimics the STFT based stretching algorithm: 1) in the modification step, the instantaneous frequency concept is used to modify the phases; 2) to reduce the number of calculations, overlapping - addition is performed in the QMF domain using the additivity property of the QMF transform.
[0066] The time-stretching algorithm in the HBE scheme is further described in detail. [0067] Assuming that there are 2L real-time signals in the time domain, x (n), intended for stretching with a stretch coefficient s, after the QMF analysis stage, there are 2L complex QMF coefficients, composed of 2L / M time windows and M subbands .
[0068] It should be noted that, as in the STFT-based stretching method, the transformed QMF coefficients are optionally subjected to window analysis prior to phase manipulation.
In the present invention this can be done either in the time domain or in the QMF domain. [0069] In the time domain, the time domain signal can naturally be windowed as in Equation 4 below.
[0070] [Calculations 4] (equation 4) [0071] The symbol mod (.) In equation 4 indicates a modulation operation.
[0072] In the field of QMF, an equivalent operation can be performed by:
1) transforming the analytical window h (n) (length L) into the QMF domain to generate H (v, k) with L / M time windows and M subbands;
2) simplifying the QMF representation of the window as shown in equation 5 below. [0073] [Calculations 5]
ml
H<sub>0</sub>(v) = H (v, k) k = 0 (equation 5) [0074] In this case, v = 0, ..., L / M-1.
[0075] 3) Performing window analysis in the QMF domain with X (m, k) = X (m, k / H<sub>0</sub>(w), where w = mod (m, L / M) (it should be noted that mod (.) means a modulation operation).
[0076] In addition, in the HBE scheme, in the phase manipulation step, a new phase is created based on the original phase of the entire set of QMF coefficients. More specifically, as a detailed implementation of stretching in time, phase manipulation is performed on the basis of a QMF block.
[0077] Fig. 9 is a schematic of methods of stretching over time in the QMF domain.
[0078] These original QMFs can be considered as L + 1 overlapping QMF blocks with a measure of offset amount of 1 time window and a L / M time window block length as shown in (a) in Fig. 9.
[0079] To ensure no phase jump phenomenon, each original QMF block is modified to generate a new QMF block with modified phases, and the phases of the new QMF blocks should be continuous at the point ^ ^ for ^) - this tab ^ + 1) - of this new QMF block, which is equivalent to continuity at points μ · Μ ^ (μεΝ) in the time domain.
[0080] Furthermore, in the HBE scheme, in the phase manipulation step, the manipulation is performed repeatedly for sets of QMF coefficients, while in the step of generating QMF coefficients new sets of QMF coefficients are generated. In this case, the phases are modified with respect to the block based on the following criteria.
[0081] Assuming that the original phases are φ, / k) for given coefficients QMF X (u, k), for u = 0, ..., 2L / M-1 and k = 0, ..., M- 1. Each original QMF block is sequentially converted to a new QMF block as shown in (b) in Figure 9, where the new QMF blocks are shown with different fill patterns.
[0082] Still, ψ ^<sup>1-1</sup>) ^) means the phase information of the nth new QMF block for n = 1, ..., L / M, u = 0, ..., L / M-1 and k = 0, 1, ..., M- 1. These new phases depending on whether the new block is spatially displaced or not are marked as follows.
[0083] Suppose the first new QMF X block<sup>(1)</sup>(u, k) (u = 0, ..., L / M-1) is not displaced. Thus, the new phase information ψυ<sup>(1)</sup>(^ is the same as $<sub>at</sub>(K). This means that ψυ ^ ΚΚ ^ φ, χ ^ for u = 0, ..., L / M-1 and k = 0, 1, ..., M-1.
[0084] For the second new QMF X block<sup>(2)</sup>(u, k) (u = 0, ..., L / M-1), it is displaced with the measure of the amount of shift s of the time window (e.g. 2 time windows, as shown in Fig. 9). In this case, the instantaneous frequencies at the beginning of the block should match the frequencies in the s-th time window in the first new QMF X block<sup>(1)</sup>(u, k). Thus, the instantaneous frequencies for the first time window X<sup>(2)</sup>(u, k) should be the same as those for the second time window in the original QMF block. This means that ψ0<sup>(2)</sup>(Κ) = ψ0<sup>(1)</sup>(Κ) +5 Δφι ^).
[0085] Furthermore, because the phases for the first time window have changed, the remaining phases are adapted to keep the original instantaneous frequencies. This means that Ψυ ^^ ψιμί ^ Μ + Δφυ + ι ^) for u = 1, ..., L / M-1, where Δφ ^ Κ ^ φ ^ -φ ^ Κ) represents the original momentary frequencies for the original block QMF.
[0086] The same phase modification rules apply to subsequent synthesis blocks. This means that for the mth new QMF block (m = 3, ..., L / M), its phases ψ, / ™) ^) are determined as shown below.
<img file="PL2581905T3_D0001.tif" />
^<sup>m</sup>\ k) = i / Ą<sup>m</sup> (fc) + s 4> u<sup>m)</sup>W = + Al / Wu-iW for u = 1, ..., L / M-1.
[0087] After enabling the amplitude information from the original block, the above new phases result in new L / M blocks.
[0088] In this case, in the HBE scheme in the present embodiment, in the phase manipulation step, different manipulations are performed depending on the QMF subband indicator. More specifically, the above phase modification method can be designed differently for odd and even QMF subbands, respectively.
[0089] It is based on the fact that for a tone signal, its instantaneous frequency in the field
QMF is related in different ways to the phase difference Δφ (η, k) = $ (n, kM (n-1, k).
[0090] More specifically, it has been found that the instantaneous frequency ω (η, k) can be determined using equation 6 below.
[Calculation 6] princ arg (A <p (n, k,))
--H kk even π
princ arg (A <p (n, / c) - π)
--H kk odd
Tl (equation 6) [0092] In equation 6, princ arg (a) is the main angle a, given by equation 7 below.
[0093] [Calculations 7] (equation 7) [0094] In the equation, mod (a, b) means modulation a on b.
[0095] As a result, e.g. in the above phase modification method, the phase difference can be obtained as in Equation 8 below.
[0096] [Calculations 8] princ arg (<p<sub>at</sub>(fc) - <p<sub>at</sub>_i (/ c)) k even princ arg (<p<sub>at</sub>(/ c) - <p<sub>at</sub>_i (/ c) - τι) k odd (equation 8) [0097] In addition, in the HBE scheme in the present embodiment, in the step of generating QMF coefficients, new sets of QMF coefficients are added with an overlap to generate QMF coefficients corresponding to the time-extended audio signal . To be more precise, to reduce the number of calculations, QMF synthesis operations do not apply directly to every single new QMF block. Instead, it applies to the results of adding these new QMF blocks.
[0098] It should be noted that, like the STFT-based stretching method, the new QMF coefficients are optionally subjected to synthesizing windowing prior to overlap addition. Like the analytical windowing process, synthesizing windowing can be accomplished as shown below.
X ^<sup>n + 1)</sup>(u, k) = X ^<sup>n + 1)</sup>(U, k) -H<sub>0</sub>(in),
ΔψΜ = where w = mod (u, L / M) [0099] Then, due to the additivity of the QMF transform, all new L / M blocks can be added with an overlap, with a shift measure equal to s time windows, before QMF synthesis. The results of adding with the tab, Y (u, k) can be obtained from the following equation.
[0100] [Calculations 9] (equation 9) [0101] Tu, n = 0, ..., L / M-1, u = 1, ..., L / M, and k = 0, ..., M-1.
[0102] The final audio signal can be generated using QMF to Y (u, k) synthesis, which corresponds to the original signal with a modified time scale.
[0103] When comparing the QMF-based stretching method in the HBE scheme in the present embodiment with the known STFT-based stretching method, it is worth noting that the inherent time resolution of the QMF transformation significantly reduces the amount of calculations, which can only be obtained with the help of the STFT transformation string in the known stretching method based on STFT. [0104] The following analysis of the number of calculations shows the result of an estimated comparison of the number of calculations, only taking into account the number of calculations resulting from the transformation.
[0105] Suppose the amount of LFT STFT calculations is log<sub>2</sub>(L \ L, and the number of QMF transformation analysis calculation calculations is about twice that of FFT transformation, with the number of transformation calculations in the known HF spectrum generator being approximately:
[0106] [Calculation 10] <sup>L</sup>/<sub>ra</sub> 2 L log<sub>2</sub>(L) - (7-1) + (2L) log<sub>2</sub>(2L) «2 (7 - 1) + l) L log ^ L) (equation (10) [0107] For comparison, the number of transformation calculations in the HF spectrum generator are approximated as shown in equation 11 below.
[0108] [Calculations 11]
TT <sup>2</sup>^(<sup>2L</sup>/ T) '<sup>l</sup>°32(<sup>2L</sup>/ t) ~ <sup>4</sup>Vt <sup>L</sup><sup>l</sup>° 32 (7) t = 2 t = 2 (equation 11) [0109] For example, assuming L = 1024 and Ra = 128, the above comparison of the number of calculations can be included in Table 1.
[0110] [Table 1]
<td>Number of Inserts harmonics (T)</td><td>Number of transformation calculations related to stretching in time</td><td>Number of transformation calculations associated with the known stretching in time</td><td>Ratio amounts calculations</td>
<td> 3</td><td> 33335</td><td> 350208</td><td> 9,52%</td>
<td> 4</td><td> 42551</td><td> 514048</td><td> 8,28%</td>
<td> 5</td><td> 49660</td><td> 677888</td><td> 7,33%</td>
Tab. 1 Comparison of the number of calculations between the known HBE and the proposed HBE using time stretching based on QMF (Form) [0111] The form of the HBE scheme (harmonic frequency band extension method) and the decoder (audio decoder or decoder device) will be described in detail below. audio) using this scheme.
[0112] It should be noted that due to the adoption of the QMF based time stretching method, the HBE technology used in the QMF based time stretching method requires much less calculations. On the other hand, however, adopting a QMF based time stretching method also brings two possible problems that threaten to degrade sound quality.
[0113] First of all, there is a problem of degradation of sound quality for high order inserts. Let's assume that the HF spectrum consists of (T-1) inserts with appropriate stretching factors as 2, 3, ..., T. Because QMF-based stretching in time refers to the block, the reduced number of overlap-add operations in the high insert order causes deterioration in the phenomenon of stretching.
[0114] Fig. 10 is a graph showing a sinusoidal tone signal. The upper part (a) shows the stretched 2nd order insert for a pure sinusoidal tone signal, the stretched output signal being substantially pure, containing only a few components of other frequencies occurring at low amplitudes. Whereas the lower part (b) shows the extended result of the 4th order insert for the same sinusoidal tone signal.
[0115] Compared to (a), it can be seen that although the center frequency is correctly shifted in (b), the resulting output signal also contains several components with other frequencies whose amplitudes cannot be ignored. This can cause unwanted noise in the stretched output.
[0116] Secondly, a degradation problem is possible for transient signals. There are 3 potential sources of involvement for this quality degradation problem.
[0117] The first source of participation is that the transient component may be lost during resampling. Assuming a transient signal with a Dirac pulse located in the even sample, for a 4th order insert with a decimal factor of 2, such a Dirac pulse disappears in the sampled signal. As a result, the output HF spectrum has incomplete transients.
[0118] The second source of contribution are mis-arranged transients among the various inserts. Because the inserts have different resampling coefficients, a Dirac pulse located in a particular location may have a number of components located in different time windows in the QMF domain.
[0119] Fig. 11 is a diagram showing the phenomenon of poor energy distribution and distribution. For an input with a Dirac pulse (e.g. in Fig. 11, shown as a third sample, marked in gray), after re-sampling with different coefficients, its position changes to different positions.
As a result, the extended output has a noticeably suppressed transient effect.
[0120] The third source of participation is that the energies of the transient components are unevenly distributed between the various inserts. As shown in Figure 11, in the case of the 2nd order insert, the associated transient component is broken down into the 5th and 6th samples; for 3rd order inserts, for 4th ~ 6th sample; and for 4th order inserts, for the 5th ~ 8th sample. As a result, the extended output has a weaker transient effect at a higher frequency. For some relevant signals, the extended output even exhibits some earlier and later echo artifacts.
[0121] To overcome the above quality degradation problem, improved HBE technology is desirable. However, too complex a solution also increases the number of calculations. This embodiment uses a QMF-based pitch offset method to avoid a possible quality degradation problem and keep the benefit of a small amount of calculations.
[0122] As described in detail hereinafter, in the HBE scheme (harmonic frequency band expansion method) in the present embodiment, the HF spectrum generator in HBE technology in this embodiment is designed with both time stretching and pitch shifting processes in the QMF domain. Furthermore, a decoder (audio decoder or audio decoding device) using HBE in the present form is also further described.
[0123] Fig. 12 is a block diagram of a method of extending a frequency band in the present embodiment.
[0124] This method of expanding the frequency band is a method of expanding the frequency band to produce a full-range signal from a low-frequency band signal, which method includes: a first transformation step of converting the low-frequency signal to the domain of a set of quadrature mirror filters (QMF) to generate the first low frequency QMF spectrum; the step of generating the low order harmonic insert of generating the low order harmonic insert by stretching the signal in the low frequency band in the QMF domain in time; the high frequency generation step of (i) generating signals shifted in terms of pitch by applying different shift factors to low order harmonic inserts, and (ii) generating high frequency QMF spectrum from the signals; the spectrum modification step of modifying the high frequency QMF spectrum to meet high frequency energy and tonality conditions; the step of generating the full-width frequency band by generating a full-band signal by combining the modified high-frequency QMF spectrum with the first low-frequency QMF spectrum.
[0125] It should be noted that the first transformation step is performed by the 1508 TF conversion module described below, the low order harmonic insert generation step is performed by the QMF 1503 transforming unit, the time stretching unit 1504, the QMF 601 transforming unit, and the phase vocoder 603 described later. Furthermore, the high frequency generation step is performed by the pitch shift unit 1506, bandpass units 604 and 605, frequency extension units 606 and 607, and delay compensation units 608 to 610, which are described below. In addition, the spectrum modification step is performed by the HF post-processing unit 1507 described below, and the step of generating the full frequency band is performed by the adding module 1512.
[0126] In addition, the low order harmonic insert generation step includes: a second transformation step of converting the low frequency band signal to a second low frequency QMF spectrum; a band processing step of bandwidth processing the second low frequency QMF spectrum; and a stretching step of stretching the band-processed second low frequency QMF spectrum along a time dimension.
[0127] It should be noted that the second transformation step is performed by the QMF 601 conversion unit and the QMF 1503 conversion unit, the band processing step is performed by the band unit 602 discussed below, and the stretching step is performed by phase vocoder 603 and stretching unit 1504 in time. [0128] Furthermore, the second low frequency QMF spectrum has a better frequency resolution than the first low frequency QMF spectrum.
[0129] In addition, the high frequency generation step includes: an insert generation step of bandwidth processing a low order harmonic insert to generate band-processed inserts; a high order generation step of mapping each of the bandwidth processed high frequency inserts to generate high order harmonic inserts; and the summation step of summing the high order harmonics with the low order harmonics. [0130] It should be noted that the insert generation step is performed by band processing units 604 and 605, the high order generation step is performed by the frequency expansion units 606 and 607, and the adding step is performed by the adding unit 611 discussed below.
[0131] Fig. 13 is a schematic diagram of an HF spectrum generator in the HBE scheme in the present embodiment. The HF spectrum generator includes the QMF 601 conversion unit, 602, 604, ..., and 605 bandwidth processing units, phase vocoder 603, 606 units, ..., and 607 frequency extension units, 608, 609, ..., and 610 units delay compensation and addition unit 611. [0132] A given input signal in the LF band is first converted (601) to the QMF domain, its band-processed (602) QMF spectrum extends in time (603) to double the length. The QMF extended spectrum is band-processed (604 ~ 605) to produce band-limited spectra (T-2). The resulting band-limited spectra are converted (606 ~ 607) into spectra in the higher frequency band. These HF spectra are aligned in terms of delays (608 ~ 610) to compensate for potential different delays from the spectrum transformation process and added (611) to generate the final HF spectrum. It should be noted that each of the above reference numerals 601 to 611 in parentheses indicates a component of the HF spectrum generator.
[0133] It should be noted that compared to the QMF transformation (108 in Fig. 1), the QMF transformation in the HBE scheme in the present embodiment (QMF transforming unit 601) has a better frequency resolution, and the decreasing time resolution will be compensated by the subsequent stretching operation .
[0134] By comparing the HBE scheme in the present embodiment with the known scheme (Figure 2), it can be seen that the main differences are: 1) like the first embodiment, the process of stretching over time is in the QMF domain and not in the FFT domain; 2) higher order inserts are generated based on 2nd order insert; 3) the pitch shift process is also implemented in the QMF domain, not in the time domain.
[0135] Fig. 14 is a decoder diagram using the HF spectrum generator in the HBE diagram in the present embodiment. The decoder (audio decoding device) includes the demultiplexing unit 1501, the decoding unit 1502, the QMF 1503 transforming unit, the 1504 time stretching unit, the delay equalization unit 1505, the pitch 1506 shifting unit, the HF post-processing unit 1507, the TF transform unit 1508 delay equalizer, 1510 TF inverse transformation unit, and 1511 adding unit. It should be noted that in the present embodiment, the demultiplexing unit 1501 corresponds to a separation unit that separates the encoded low-frequency signal from the encoded information (bit stream). In addition, the TF inverse transform unit 1510 corresponds to the reverse transform unit that converts the full-range signal from the signal domain of the quadrature mirror filter (QMF) signal into the time domain signal.
[0136] With the decoder, the bit stream is first demultiplexed (1501) and then the LF portion (1502) is decoded. To approximate the original HF portion, the decoded LF portion (low frequency signal) is converted (1503) in the QMF domain to generate the LF QMF spectrum. The resulting QMF LF spectrum extends (1504) in time direction to generate a low order HF insert. The low order HF insert is shifted (1506) to generate high order inserts. The resulting high order inserts are combined with the delayed (1505) low order HF insert to generate the HF spectrum, and the HF spectrum is further purified (1507) by post-processing, taking some of the decoded HF parameters as the leading. At the same time, the decoded portion of LF is also converted (1508) to the domain of QMF. Finally, the purified HF spectrum is combined with the delayed (1509) LF spectrum to produce (1512) full-band QMF spectrum. The resulting full-band QMF spectrum is transformed (1510) back to the time domain to output a decoded wide-band audio signal. It should be noted that each of the reference numbers 1501 to 1512 indicates a component of the decoder.
Tone shifting method [0137] The QMF-based pitch shifting algorithm (frequency extension method in the QMF domain) for the 1506 pitch shifting unit in the HBE scheme in the present embodiment was designed by breaking the LF QMF subbands into multiple subbands, transposing these subbands to HF subbands, and by combining the resulting HF subbands to generate an HF spectrum. More specifically, the high order generation step includes: a separation step consisting of splitting each of the QMF subbands in each of the band-processed inserts into multiple subbands; a step of mapping the sub-subbands to the high frequency QMF subbands; and a step of combining the sub-band mapping results.
[0138] It should be noted that the separation step corresponds to step 1 (901 - 903) described below, the mapping step corresponds to steps 2 and 3 (904 ~ 909) described below, and the merging step corresponds to step 4 (910) described in continue.
[0139] Fig. 15 is a schematic of such a QMF-based height shift algorithm. When the data is band-processed the 2nd order insert spectrum, the HF spectrum of the t-order (t> 2) insert can be reproduced by: 1) distribution (step 1: 901 ~ 903) of the given LF spectrum, i.e. each QMF subband within the LF spectrum is spread over multiple QMF subbands; 2) scaling (step 2: 904 ~ 906) of the center frequencies of these subbands with a factor of t / 2; 3) mapping (step 3: 907 ~ 909) of these subbands to the HF subbands; 4) summing all mapped subbands to create HF subbands (step 4: 910).
[0140] Regarding step 1, several methods are available for splitting the QMF subband into multiple subbands to achieve better frequency resolution. For example, Mth band filters are used, which are used in the MPEG spatial codec. In this preferred embodiment of the invention, subband distribution is accomplished using an additional set of exponentially modulated filter assembly, as defined by Equation 12 below.
[0141] [Calculation 12] (equation 12) <sup>J</sup>q (n) = exp | 7- (q + O 5) (nn<sub>about</sub>) j '' [0142] Tu, q = -Q, -Q + 1, ..., 0, 1, ..., Q-1 in = 0, 1, ..., N (where n0 is a constant integer, N is the order of the filter assembly).
[0143] By using the above set of filters, a given subband signal, say, k-th subband signal x (n, k), is divided into 2Q of sub-band signals according to equation 13 below.
[0144] [Calculation 13] (equation 13) [0145] Tu, q = -Q, -Q + 1, ..., 0, 1, ..., Q-1. In the equation, 'conv (.)' Means the convolution function.
[0146] By means of such additional complex transformation, the frequency spectrum of one subband is further divided into 2Q frequency sub spectra. From the frequency resolution point of view, if the QMF transform has an M-band, its associated subband frequency resolution is Π / M, and its frequency subband resolution is reduced to n / (2Q ^ M). In addition, the general system shown in equation 14 is unchanging in time, i.e. is free of aliasing distortions despite downsampling and upsampling sampling.
[0147] [Calculation 14] <3-1
9q (p) q = -Q (equation 14) [0148] It should be noted that the above additional set of filters is stacked unevenly (factor q + 0.5) which means that there are no sub-bands centered around the DC values. In contrast, for an even number Q, the center frequencies of the subbands are symmetrical around zero.
[0149] Fig. 16 is a graph showing the distribution of sub-band spectra. More specifically, Fig. 16 shows this spectrum distribution of the filter assembly for the case Q = 6. The purpose of odd stacking is to facilitate subsequent sub-banding.
[0150] Regarding step 2, center frequency scaling can be simplified by taking into account the oversampling parameters of the complex QMF.
[0151] It should be noted that in the complex domain of QMF, as the bandwidths of adjacent subbands overlap, the frequency component in the overlap zone will appear in both subbands (see International Patent Application WO 2006048814).
As a result, frequency scaling can be simplified to half the number of calculations by calculating only the frequencies for these subbands within the bandwidth, i.e. the positive frequency part for the even subband or the negative part part for the odd subband.
[0153] More specifically, the kLF-subband is divided into 2Q sub-subbands. In other words, x (n, kLF) divides as shown in equation 15.
[0154] [Calculations 15] (equation (15) [0155] Then, to create the t-th insert, the center frequencies of these subbands are scaled using equation 16 below.
[0156] [Calculations 16] f £ tale = ( <sup>fc</sup>LF + O, 5 + · g) · J <sup>(equation 16)</sup> [0157] Here, q = -Q, -Q + 1, ..., -1 when kLF is odd or q = 0, 1, ..., Q-1 when kLF is even. [0158] Regarding step 3, mapping the sub-subbands to the HF subband also requires consideration of the parameters of the complex QMF transformation. In this embodiment, such a mapping process is performed in two steps; the first is directly mapping all sub-subbands to the bandwidth in the HF subband; the second mapping result based on the above is mapping of all subbands to the barrier HF subband. More specifically, the mapping step includes: a dividing step of splitting the sub-subbands of each of the QMF subbands into the damming portion of the band and the bandwidth portion of the band; a frequency calculation step of calculating the transposed center frequencies of the sub-subbands into the bandwidth portion of the band with a factor depending on the order of the insert; a first mapping step in which the subbands are mapped on the bandwidth portion to the high frequency QMF subbands according to the center frequencies; and a second mapping step in which the subbands on the banned subband are mapped to the high frequency QMF subbands according to the subband subbands of the bandwidth.
[0159] To understand the above point, it is beneficial to re-examine what relationship exists for a positive frequency pair and a negative frequency pair for the same signal component and associated subband indicators.
[0160] As said before, in the complex domain of QMF, the sinusoidal spectrum has both a positive and a negative frequency. More specifically, the sinusoidal spectrum has one of these frequencies in the bandwidth portion of one QMF subband, and the other frequency in the damming portion of the adjacent band subband. Given that the QMF transform is an odd stacked transform, such a pair of signal components can be represented in Fig. 17.
[0161] Fig. 17 is a graph showing the relationship between a bandwidth component and a bandwidth component for a sinusoidal signal in the complex domain of QMF.
[0162] Here, the gray area indicates a subband bandwidth. For any sinusoidal signal (solid line) on the bandwidth subband, its interference portion (aliasing) (dashed line) is located in the band of the adjacent subband (paired two frequency components are connected by a line with double arrows).
[0163] A sinusoidal signal with a frequency f0 as shown in equation 17.
[0164] [Calculations 17]
COG) <<sup>f</sup>»<( <sup>ι</sup>-^')·<sup>π</sup> (<sup>equation 17</sup>) [0165] The pass band component of the sinusoidal signal with the frequency f0 described earlier is located on the k-th subband, if the following equation 18 is met.
[Calculation 18] ^ - <f<sub>0</sub> < <sup>(fc</sup> ++^ <sup>π</sup> (equation 18) [0167] Furthermore, its blocking component of the band is located at k<sup>~</sup>-this subband if Equation 19 below is met.
[0168] [Calculations 19]
f.
k-1 k + 1 if if k-π (k + 0.5) π
- <fo <--- M ~<sup>J0</sup> M (fc + 0.5) "π j. <sub>j</sub> (k + 1) π
Λ4 - / i) ji.
(equation 19) [0169] If the subband is divided into 2Q subbands, the above relation is formed with a higher frequency resolution as shown in Fig. 20 below.
[0170] [Calculation 20] (k - 1) "for - Q /<sub>7</sub><q <—l if k even; or for Λ <q <—1 if k odd k<sub>q</sub> = l * <sup>Δ Δ</sup> (k + 1), for - Q <q < <sup>—</sup> even ddy; or for 0 <q <^ 2 when k odd (equation 20) [0171] Therefore, in the present embodiment, in order to map sub-bands on the bandwidth to the HF subband, it is necessary to link them with the results of mapping these sub-bands on the bandwidth . The rationale for doing so is to ensure that the frequency pairs for LF components remain in pair when shifted up to HF components.
[0172] To this end, the subband should be mapped directly to the HF subband directly. Given the center frequencies of the frequency-scaled subbands and the frequency resolution of the QMF transform, the mapping function can be described by m (k, q) as shown in Equation 21 below.
[0173] [Calculations 21] <sup>m</sup>& LF. <?<sup>)</sup> = [f ^ ^ inches <sup>(equation 21)</sup> [0174] Here, q = -Q, -Q + 1, ..., -1 if kLF is odd or q = 0, 1, ..., Q-1 if kLF is even. Here, the coefficient shown in equation 22 below means a rounding operation to obtain the nearest integer x values in the minus infinity direction.
[0175] [Calculations 22] (equation 22) [0176] In addition, due to upscaling (t / 2> 1), it is possible for one HF subband to have multiple sub-mapping sources. This means that it is possible that m (k, q1) = m (k, q2) or m (k1, q1) = m (k2, q2). Therefore, the HF subband may be a combination of multiple subbands from the LF subbands, as shown in Equation 23.
[0177] [Calculations 23]
Xpass (n, k<sub>HF</sub>) = Σ yq<sup>LF (</sup>.n) all rntk ^ p-ą ^ k ^ p (equation 23) [0178] Here, q = -Q, -Q + 1, ..., -1 if kLF is odd or q = 0, 1 ,. ..., Q-1 if kLF is even. [0179] Secondly, according to the above relationship between frequency pairs and subband indicators, a mapping function for these sub-subbands on the bandwidth band can be established as follows.
[0180] Considering the LF subband sub-band KLF, the functions mapping the sub-subband on its bandwidth are already set by the first step as: m (kLF, -Q), m (kLF, -Q + 1), ..., m (kLF, -1) for odd kLF and m (kLF, 0), m (kLF, 1), ..., m (kLF, Q-1) for even kLF, and then the part of the dam band associated with the bandwidth map to equation 24 below.
[0181] [Calculation 24] (
m (k<sub>LF</sub>, q) - 1 condition am (k<sub>LF</sub>, q) + 1 otherwise (equation 24) [0182] Here, 'condition a' refers to the situation when kLF is even and equation 25 is even or when kLF is odd and equation 26 below is even.
[0183] [Calculation 25] | G / + o, 5) t [0184] [Calculation 26] (equation 25) | <sub>t +</sub> (q + 0.5) i.e. (equation 26) [0185] In addition, as described previously, the following equation 27 denotes the rounding operation to obtain the nearest integer x values in the minus infinity direction.
[Calculations 27] (equation (27) [0186] The resulting HF subband is a combination of all associated LF subbands, as shown in Equation 28 below.
[0187] [Calculation 28] x<sub>st</sub>about<sub>P</sub>(N, k<sub>HF</sub>j = J yq<sup>LFq</sup>(n) all πϊ (/ <^ ρ ^ · ί /) = / <^ ρ (equation 28) [0188] Here, q = -Q, -Q + 1, ..., -1 if kLF is even or q = 0, 1, ..., Q-1 if kLF is odd. [0189] Finally, all mapping results on the bandwidth and the bandwidth band are combined to form the HF subband, as shown in Equation 29 below.
[0190] [Calculations 29] (equation 29) [0191] It should be noted that the above pitch shift method in the QMF domain is beneficial both in terms of deterioration in frequency quality and possible transient handling problems.
[0192] First, all inserts now have the same stretch ratio, the smallest, which significantly reduces high frequency noise (from incorrect signal components generated during time stretching). Secondly, all sources of participation for exacerbation in transient states are avoided. This means that there is no process of re-sampling the time domain; and the same stretch ratios are used for all patches, which inherently eliminates the possibility of misalignment.
[0193] It should further be noted that the present embodiment has some disadvantages in relation to frequency resolution. It should be noted that as a result of adopting sub-band filtering, the frequency resolution increases from Π / M to n / (2Q ^ M), but is still more coarse than the high resolution of the time domain sampling (Π / L). However, taking into account the lower sensitivity of the human ear to the high frequency component of a given signal, it turned out in practice that the pitch-shifted result obtained by this form does not differ perceptually from the result obtained by the sampling method. [0194] In addition to the above, compared to the HBE scheme in the first embodiment, the HBE scheme in the present embodiment also provides the advantage of further reducing the number of calculations, since only one low-order insert requires a time-stretch operation.
[0195] Again, such a reduction in the number of calculations can be roughly analyzed only considering the amount of calculations resulting from the transformation.
[0196] As assumed in the above analysis amount of calculations, the number of transformation calculations associated with the HF spectrum generator in the present embodiment is approximated as shown below.
[0197] [Calculation 30] (equation 30) [0198] Table 1 can therefore be updated as follows.
[0199] [Table 2]
<td>Number of Inserts harmonics (T)</td><td>Number of transformations calculations in HBE in this form</td><td>Number of transformations calculations in HBE in the first example</td><td>Quantity ratios calculations</td>
<td> 3</td><td> 20480</td><td> 33335</td><td> 61,4%</td>
<td> 4</td><td> 20480</td><td> 42551</td><td> 48,1%</td>
<td> 5</td><td> 20480</td><td> 49660</td><td> 41,2%</td>
Tab. 2 Comparison of the number of calculations between HBE in this form and the HBE scheme in the first example.
[0200] The present invention is a new HBE technology for audio coding at a low bit rate. Using this technology, the broadband signal can be reproduced from the low frequency signal by generating its high frequency (HF) part by stretching it over time and frequencyly extending the low frequency (LF) part in the QMF domain. Compared to the known HBE technology, the present invention provides comparable sound quality and significantly fewer calculations. Such technology can be used in applications such as mobile phones, teleconferences, etc., where the audio codec works at a low transmission speed with a small amount of calculations.
[0201] It should be noted that each of the function blocks in block diagrams (Figures 6, 7, 13, 14 etc.) is typically implemented as LSI, which is an integrated circuit. Function blocks can be implemented as separate integrated circuits or as a single integrated circuit containing part or all of them.
[0202] Although we refer to LSI in the present context, there are cases in which the terms IC, system LSI, super LSI, ultra-LSI are used due to differences in the degree of integration. [0203] Furthermore, the means for integrating the circuit are not limited to LSI, and thus it is also possible to implement with a dedicated integrated circuit or general purpose processor. It is also acceptable to use FPGA (Field Programmable Gate Array), which enables programming after LSI production, and a configurable processor in which connections and settings of the system cells within LSI can be configured.
[0204] In addition, if LSI replacing integrated circuit technology appears as a result of advances in semiconductor technology or another derivative technology, this technology can naturally be used to perform function block integration.
[0205] Furthermore, among the individual function blocks, the data storage module for coding or decoding can be made as a separate structure without being integrated into a single integrated circuit.
[Industrial applicability] [0206] The present invention relates to new harmonic bandwidth extension (HBE) technology for low speed audio coding. As part of this technology, a broadband signal can be reproduced from a low frequency signal by generating its high frequency (HF) portion by time stretching and frequency low frequency (LF) extension in the QMF domain. Compared to the known HBE technology, the present invention provides comparable sound quality and significantly fewer calculations. Such technology can be used in applications such as mobile phones, teleconferences, etc., where the audio codec works at a low transmission speed with a small amount of calculations.
[List of links] [0207]
501-503, 602, 604, 605 bandpass unit
504-506 sampling unit
507-509, 601, 1404, 1505 QMF transformation unit
510-512, 603 phase vocoder
513-515, 608-610, 1407, 1505, 1509 delay compensation unit
516, 611, 1410, 1511, 1512 adding unit
606, 607 frequency expansion unit
1401, 1501 demultiplexing unit
1402, 1502 decoding unit
1403 resampling unit in time 1405, 1504 stretching unit in time
1406, 1508 TF transformation unit
1409, 1510 reverse transformation unit TF 1506 pitch shift unit
Panasonic Intellectual Property Corporation of America Agent:
PL-PAT-2012-572
EP 2 581 905 B1
Contents2
43 members in 19 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 2010132205 | Japan | A | |
| 2010132205 | Japan | A | |
| 11792129 | European Patent Office (EPO) | A | |
| 2011003168 | Japan | W | |
| 2011003168 | Japan | W | |
| EP20110792129 | – | – | – |
| JP20100132205 | – | – | – |
| WO2011JP03168 | – | – | – |
Members43
| Document | Office | Kind | |
|---|---|---|---|
| CA2770287A1 | Canada | A1 | |
| WO2011155170A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201207840A | Taiwan Province of China | A | |
| MX2012001696A | Mexico | A | |
| AU2011263191A1 | Australia | A1 | |
| SG178320A1 | Singapore | A1 | |
| CN102473417A | China | A | |
| US2012136670A1 | United States of America | A1 | |
| AR082764A1 | Argentina | A1 | |
| EP2581905A1 | European Patent Office (EPO) | A1 | |
| KR20130042460A | Republic of Korea | A | |
| JP2013084018A | Japan | A | |
| JP5243620B2 | Japan | B2 | |
| ZA201200919B | South Africa | B | |
| JPWO2011155170A1 | Japan | A1 | |
| RU2012104234A | Russian Federation | A | |
| EP2581905A4 | European Patent Office (EPO) | A4 | |
| CN102473417B | China | B | |
| JP5750464B2 | Japan | B2 | |
| US9093080B2 | United States of America | B2 | |
| US2015248894A1 | United States of America | A1 | |
| EP2581905B1 | European Patent Office (EPO) | B1 | |
| EP3001419A1 | European Patent Office (EPO) | A1 | |
| ES2565959T3 | Spain | T3 | |
| RU2582061C2 | Russian Federation | C2 | |
| AU2011263191B2 | Australia | B2 | |
| PL2581905T3This record | Poland | T3 | |
| TWI545557B | Taiwan Province of China | B | |
| HUE028738T2 | Hungary | T2 | |
| BR112012002839A2 | Brazil | A2 | |
| KR101773631B1 | Republic of Korea | B1 | |
| BR112012002839A8 | Brazil | A8 | |
| US9799342B2 | United States of America | B2 | |
| CA2770287C | Canada | C | |
| US2017358307A1 | United States of America | A1 | |
| EP3001419B1 | European Patent Office (EPO) | B1 | |
| US10566001B2 | United States of America | B2 | |
| US2020135217A1 | United States of America | A1 | |
| MY176904A | Malaysia | A | |
| BR112012002839B1 | Brazil | B1 | |
| US11341977B2 | United States of America | B2 | |
| US2022246159A1 | United States of America | A1 | |
| US11749289B2 | United States of America | B2 |
Numbers
- Publication, DOCDB
- 2581905
- Publication, EPODOC
- PL2581905T
- Application
- 792129
- Application, DOCDB
- 11792129
- Application, EPODOC
- PL20110792129T
Titles2
- English
- BANDWIDTH EXTENSION METHOD, BANDWIDTH EXTENSION APPARATUS, PROGRAM, INTEGRATED CIRCUIT, AND AUDIO DECODING APPARATUS
- Polish
- Sposób rozszerzania pasma częstotliwości, urządzenie do rozszerzania pasma częstotliwości, program, układ scalony oraz urządzenie dekodujące audio
Classification
- CPC, 6
- G10L19/02
- G10L21/038
- G10L21/04
- G10L19/0204
- G10L19/0208
- G10L21/02
- IPC, 4
- G10L21 038
- G10L19 02
- G10L21 0388
- G10L21 04