Apparatus, method and computer program for decoding an encoded audio signal
Abstract
a core decoder 600 for decoding the encoded core signal to obtain a decoded core signal; an analyzer (602) that analyzes the decoded core signal before or after a frequency reproduction operation to provide an analysis result (603); and a frequency regenerator (604) for reproducing spectral parts not included in the decoded core signal by using the analysis result (603), the parameter data (605) and the spectral part of the decoded core signal. Apparatus for decoding an encoded audio signal comprising an encoded core signal and parameter data.

Term
7.8 yearsto projected expiry
Projected expiry 15 July 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
17 claims: 9 independent, 8 dependent
- 1인코딩된 코어 신호 및 파라미터 데이터를 포함하는 인코딩된 오디오 신호를 디코딩하기 위한 장치에 있어서, 디코딩된 코어 신호를 얻기 위해 상기 인코딩된 코어 신호를 디코딩하는 코어 디코더(600);분석 결과(603)를 제공하기 위해 주파수 재생 작업 전 또는 후 상기 디코딩된 코어 신호를 분석하는 분석기(602);및 상기 분석 결과(603), 파라미터 데이터(605) 및 상기 디코딩된 코어 신호의 스펙트럼 부분을 이용하여 상기 디코딩된 코어 신호에 포함되지 않는 스펙트럼 부분들을 재생하는 주파수 재생기(604);를 포함하는, 인코딩된 코어 신호 및 파라미터 데이터를 포함하는 인코딩된 오디오 신호를 디코딩하기 위한 장치.
- 2제1항에 있어서, 상기 분석기(614)는 주파수 재생 전 또는 후에 상기 디코딩된 코어 신호의 하나 이상의 지역적 스펙트럼 최소값들을 위치시키기 위해 주파수 재생 전 또는 후 상기 디코딩된 코어 신호를 분석하도록 구성되며, 상기 분석기(614)는 상기 로컬 스펙트럼 최소값들을 식별하는 상기 분석 결과(603)를 제공하도록 구성되며, 상기 주파수 재생기(604, 616)는 상기 스펙트럼 부분을 재생하도록 구성되며, 상기 재생된 스펙트럼 부분의 또는 상기 디코딩된 신호의 상기 스펙트럼 부분의 주파수 타일 경계들이 하나 이상의 스펙트럼 최소값들에서 설정되는 것을 특징으로 하는 장치.
- 3제1항 또는 제2항에 있어서, 상기 주파수 재생기(604)는 예비 재생 신호(703)를 발생시키도록 구성되고, 상기 분석기(602)는 아티팩트-생성 신호 부분들을 검출하기 위해 상기 예비 재생 신호를 분석(704)하도록 구성되며, 상기 주파수 재생기(604)는 상기 예비 재생 신호를 조작하는 또는 상기 재생된 신호의 아티팩트-생성 신호를 감소 또는 제거하기 위해 상기 예비 재생에 대한 파라미터들과 다른 파라미터들로 추가 재생을 수행하는 조작기(722)를 더 포함하는 것을 특징으로 하는 장치.
- 4제1항 내지 제3항 중 어느 한 항에 있어서, 상기 주파수 재생기(604)는 재생된 스펙트럼 부분들을 얻기 위해 상기 디코딩된 코어 신호의 스펙트럼 부분을 이용하여 상기 디코딩된 코어 신호에 포함되지 않는 스펙트럼 부분들을 갖는 예비 재생 신호(703)를 재생하도록 구성되며, 상기 주파수 재생기(604)는 상기 디코딩된 코어 신호 및 재생된 스펙트럼 부분 사이의 주파수 경계 근처에 또는 상기 디코딩된 코어 신호의 동일 또는 상이한 스펙트럼 부분들을 이용하여 발생된 두개의 재생된 스펙트럼 부분들 사이의 주파수 경계 근처에 아티팩트 생성 신호 부분들을 검출(704)하도록 구성되고, 상기 주파수 재생기(604)는 상기 예비 재생 신호를 조작하는 또는 상기 예비 재생 신호를 발생시키기 위해 이용되는 상기 제어 데이터와 다른 상기 조작된 제어 데이터를 이용하여 재생된 신호를 새로 발생시키기 위해 제어 데이터를 조작하는 조작기(722)를 더 포함하는 것을 특징으로 하는 장치.
- 5제4항에 있어서, 상기 주파수 재생기(604)는 미가공(raw) 스펙트럼 부분들을 얻기 위해 상기 디코딩된 신호의 하나 이상의 스펙트럼 부분들을 이용하여 상기 스펙트럼 부분들을 유도하도록 구성되는 타일 발생기(820)를 포함하며, 상기 조작기(824)는 조작된 스펙트럼 부분들을 얻기 위해 상기 주파수 타일 발생기(820) 또는 상기 미가공(raw) 스펙트럼 부분들을 조작하도록 구성되며, 상기 주파수 재생기(604)는 상기 파라미터 데이터(605)를 이용하여 상기 조작된 스펙트럼 부분들의 엔벨로프 조정을 수행하도록 구성되는 스펙트럼 엔벨로프 조정기(826)를 더 포함하는 것을 특징으로 하는 장치.
- 6제1항 내지 제5항 중 어느 한 항에 있어서, 상기 분석기(602)는 주파수 검출 범위에 위치된 음조 신호 부분들을 검출하도록 구성되며, 상기 주파수 검출 범위는 1 바크(Bark)인 미리 결정된 검출 대역폭 또는 복원 주파수 범위 또는 소스 주파수 범위의 대역폭의 20%보다 작은 미리 결정된 검출 대역폭만큼 복원 범위 내의 인접 주파수 타일들 사이에서 또는 복원 범위의 주파수 경계로부터 확장하는 것을 특징으로 하는 장치.
- 7제6항에 있어서, 상기 조작기(824)는 상기 미리 결정된 검출 대역폭에서 재생된 신호의 음조 부분들을 포함하는 스펙트럼 부분들을 감쇠시키거나 제거(708)하도록 구성되는 것을 특징으로 하는 장치.
- 8제7항에 있어서, 상기 조작기(722, 824)는 주파수로 상기 음조 부분(802)의 시작에 위치된 주파수 시작 스펙트럼 부분 및 주파수로 상기 음조 부분(802)의 끝 주파수에 위치된 끝 스펙트럼 부분을 결정하고, 보간된 신호 부분을 얻기 위해 상기 시작 주파수 및 상기 끝 주파수 사이를 보간(804)하고, 상기 보간된 신호 부분(806)에 의해 시작 주파수 및 끝 주파수 사이의 음조 부분을 교체하도록 구성되는, 장치.
- 9제7항에 있어서, 상기 조작기(822)는 상기 재생된 스펙트럼 부분들(810)의 비-음조 신호 부분 또는 상기 디코딩된 코어 신호의 비-음조 신호 부분에 의해 결정되는 에너지를 갖는 스펙트럼 라인들(808)을 랜덤으로 또는 비-랜덤으로 발생시키도록 구성되는 것을 특징으로 하는 장치.
- 10제4항 내지 제9항 중 어느 한 항에 있어서, 상기 분석기는 특정 주파수에서 상기 아티팩트-생성 신호 부분들을 검출하도록 구성되고, 상기 조작기(722, 824)는 타일 발생기를 제어하도록 구성되어 상기 타일 발생기가 상기 재생된 스펙트럼 부분의 주파수 경계 또는 상기 디코딩된 코어 신호의 스펙트럼 부분의 주파수 경계를 변경하도록 구성되고 상기 아티팩트-생성 신호 부분이 더 적게 아티팩트-생성하거나 아티팩트-생성을 하지 않도록 하는 것을 특징으로 하는 장치.
- 11제1항 내지 제10항 중 어느 한 항에 있어서, 상기 분석기(602)는 재생된 신호의 최대 주파수 경계에서 또는 상기 디코딩된 코어 신호의 동일 또는 상이한 스펙트럼 부분들을 이용하여 재생된 두개의 재생된 스펙트럼 부분들 사이의 주파수 경계에서 또는 상기 디코딩된 코어 신호의 주파수 경계에서 재생된 신호의 또는 상기 디코딩된 코어 신호의 스펙트럼 부분의 피크 부분의 중간-분할(midway-splitting)을 검출하도록 구성되며, 상기 주파수 재생기는 상기 분할이 감소되거나 제거되도록 상기 디코딩된 신호의 동일 또는 상이한 스펙트럼 부분들을 이용하여 발생된 두개의 재생된 스펙트럼 부분들 사이의 주파수 경계 또는 상기 재생된 신호 및 상기 디코딩된 코어 신호 사이의 주파수 경계를 변경하거나 상기 최대 주파수를 변경하도록 구성되는 것을 특징으로 하는 장치.
- 12상기 제1항 내지 제11항 중 어느 한 항에 있어서, 상기 주파수 재생기(604)는 타일 발생기(820)를 포함하며, 상기 타일 발생기(820)는 상기 디코딩된 코어 신호의 동일 또는 상이한 스펙트럼 부분들을 이용하여 제1스펙트럼 부분에 대한 제1주파수 타일 및 제2스펙트럼 부분에 대한 제2주파수 타일을 발생시키도록 구성되며, 상기 제2주파수 타일의 상기 하부 주파수 경계는 상기 제1주파수 타일의 상부 주파수 경계와 일치하고, 상기 분석기(602)는 상기 피크 스펙트럼 부분이 제2주파수 타일의 하부 주파수 경계 또는 제1주파수 타일의 상부 주파수 경계 또는 제1주파수 타일의 하부 주파수 경계 및 상기 디코딩된 코어 신호의 미리 결정된 갭 필링 시작 주파수(309)에 의해 클리핑(clipped)되는지 여부를 검출하도록 구성되며, 상기 조작기(824)는 상기 타일 발생기(820)가 상기 클리핑이 감소되거나 제거되도록 변경된 수정된 시작 또는 정지 주파수 경계들을 갖는 수정된 주파수 타일들을 발생시키도록 상기 타일 발생기(820)를 제어하도록 구성되는 것을 특징으로 하는 장치.
- 13제1항 내지 제12항 중 어느 한 항에 있어서, 상기 코어 디코더는 제로(zero) 표현과 다른 스펙트럼 값들에 의해 표현되는 제1스펙트럼 부분들의 제1세트를 포함하는 주파수 영역 디코딩된 스펙트럼 부분들을 얻도록 구성되고 제2스펙트럼 부분들의 제2세트는 스펙트럼 값들에 대해 제로(zero) 표현에 의해 표현되며, 상기 파라미터 정보는 상기 제2스펙트럼 부분들의 제2세트에 대해 제공되며, 상기 주파수 재생기(604)는 제1스펙트럼 부분들의 제1세트에 포함되지 않는 복원 대역 내의 스펙트럼 부분들을 재생하기 위해 제1스펙트럼 부분들의 제1세트로부터 디코딩된 스펙트럼 부분들을 이용하도록 구성되며, 상기 장치는 상기 재생된 스펙트럼 부분들 및 상기 디코딩된 코어 신호의 스펙트럼 부분들을 시간 표현으로 변환하는 주파수-시간 변환기(828)를 더 포함하는 것을 특징으로 하는 장치.
- 14제1항 내지 제13항 중 어느 한 항에 있어서, 상기 코어 디코더(600)는 변형 이산 코사인 변환(MDCT) 스펙트럼 값들을 출력하도록 구성되고, 상기 주파수-시간 변환기(828)는 이후 얻어지는 MDCT 프레임들에 오버랩-애드(overlap-add) 처리를 적용하는 역 MDCT 변환(512, 514, 516)을 수행하기 위한 프로세서를 포함하는 것을 특징으로 하는 장치.
- 15제1항 내지 제14항 중 어느 한 항에 있어서, 상기 주파수 재생기(604)는 예비 재생 신호를 발생(702)시키도록 구성되며, 상기 주파수 재생기(604)는 상기 예비 재생 신호의 음조 구성요소들을 검출(704)하도록 구성되며, 상기 주파수 재생기는 재생된 신호를 발생시키기 위해 검출(704)의 결과에 기반하여 상기 복원 범위의 인접 주파수 타일들 사이의 또는 복원 범위 및 소스 범위 사이의 전이 주파수들을 조정하도록 구성되며, 상기 재생기는 상기 전이 주파수들 주변의 검출 범위에 위치되는 음조 구성요소들(708)을 제거하도록 더 구성되며, 상기 주파수 재생기는 상기 전이 주파수들 주변의 크로스-오버 범위에서 제거된 음조 구성요소들을 갖는 신호를 크로스-오버 필터링하기 위한 크로스-오버 필터(710)를 더 포함하며, 상기 주파수 재생기는 상기 파라미터 데이터(605)를 이용하여 상기 크로스-필터의 결과를 스펙트럼 엔벨로프 성형하기 위한 스펙트럼 엔벨로프 성형기(712)를 더 포함하는 것을 특징으로 하는 장치.
- 16인코딩된 코어 신호 및 파라미터 데이터를 포함하는 인코딩된 오디오 신호를 디코딩하기 위한 방법에 있어서, 디코딩된 코어 신호를 얻기 위해 상기 인코딩된 코어 신호를 디코딩(600)하는 단계;분석 결과(603)를 제공하기 위해 주파수 재생 작업을 수행하기 전 또는 후에 상기 디코딩된 코어 신호를 분석(602)하는 단계;및 상기 분석 결과(603), 상기 파라미터 데이터(605), 및 상기 디코딩된 코어 신호의 스펙트럼 부분을 이용하여 상기 디코딩된 코어 신호에 포함되지 않는 스펙트럼 부분들을 재생(604)하는 단계;를 포함하는, 인코딩된 코어 신호 및 파라미터 데이터를 포함하는 인코딩된 오디오 신호를 디코딩하기 위한 방법.
- 17컴퓨터 또는 프로세서 상에서 수행될 때, 제16항의 방법을 수행하기 위한 컴퓨터 프로그램.
Independent claims17
213 paragraphs, as filed
Apparatus, method and computer program for decoding an encoded audio signal
The present invention relates to audio coding/decoding, and more particularly to audio coding using Intelligent Gap Filling (IGF).
Audio coding is the domain of signal compression that uses psychoacoustic knowledge to address the exploitation of redundancy and irrelevancy in audio signals. Today's audio codecs generally require about 60 kbps/channel for perceptually transparent coding of almost any kind of audio signal. New codecs aim to reduce the coding bitrate by exploiting spectral similarities in the signal using techniques such as bandwidth extension (BWE). The bandwidth extension strategy uses a low bitrate parameter setting to represent the high frequency components of the audio signal. The high-frequency spectrum is filled with spectral content from the low-frequency regions and the spectral shape, tiling and temporal duration adjusted to maintain the timbre and color of the original signal. Such bandwidth extension methods enable audio codecs to maintain excellent quality even at a low bitrate of about 24 kbps/channel.
The audio coding system of the invention efficiently codes arbitrary audio signals over a wide range of bitrates. On the other hand, for high bitrates, the system of the present invention focuses on transparency, since low bitrate perceptual annoyance is minimized. Thus, the main share of the available bitrate is used to waveform-code the very perceptually relevant structure of the signal in the encoder, and the resulting spectral gaps are filled in the decoder with the signal content approximately approximating the original spectrum. A very limited bit budget is expended to control the so-called spectral intelligent gap filling, which is parameter driven by dedicated side information transmitted from the encoder to the decoder.
The storage or transmission of audio signals is often subject to strict bitrate constraints. In the past, coders have been forced to dramatically reduce the transmitted audio bandwidth when only very low bitrates are available.
Modern audio codecs today can code wideband signals by using bandwidth extension methods ([1]). These algorithms rely on a parametric representation of the high frequency content, which is generated from the waveform coded low frequency of the decoded signal by the application of a parametric drive post-processing and transposition into the high frequency spectral region ("patching"). . In bandwidth extension strategies, the reconstruction of a high-frequency spectral region above a given so-called cross-over frequency is often based on spectral patching. In general, a high frequency region is decomposed into a number of adjacent patches, each of these patches originating from band-pass (BP) regions of the low frequency spectrum below a given crossover frequency. Modern systems implement patching within a filterbank representation, e.g., a Quadrature Mirror Filter (QMF), by copying adjacent subband coefficients from a source to a target region.
Another technique found in today's audio codecs that increases compression efficiency and thereby enables extended audio bandwidth at low bitrates is parameter driven synthetic replacement of suitable parts of the audio spectrum. For example, signal portions such as noise of the original audio signal can be replaced without substantial loss of subjective quality by artificial noise generated within the decoder and scaled by side information parameters. An example is the perceptual noise substitution included in MPEG-4 Advanced Audio Coding (AAC) ([5]).
Another offering that also enables extended audio bandwidth is the noise filling technology embedded within MPED-D Unified Speech and Audio Coding (USAC) ([7]). Spectral gaps (zeros) inferred by the quantizer's dead-zone due to too coarse quantization are then filled with artificial noise in the decoder and by parameter-driven post-processing. is scaled
Another modern system is Accurate Spectral Replacement (ASR) ([2-4]). In addition to the waveform codec, accurate spectral replacement uses a dedicated signal synthesis step that stores the perceptually significant sinusoidal portions of the signal at the decoder. In addition, the system described in [5] relies on sinusoidal modeling in the high-frequency region of the waveform coder to enable extended audio bandwidth with adequate perceptual quality at low bitrates. All these methods involve transformation of the data into the second domain, except for the modified discrete cosine transform (MDCT), and also include fairly complex analysis/synthesis steps for the preservation of high frequency sinusoidal components.
13A shows a schematic diagram of an audio encoder for a bandwidth extension technique, eg, a GAD used in high-efficiency advanced audio coding. The audio signal on line 1300 is input into a filter system comprising a low pass 1302 and a high pass 1304. The signal output by the high pass filter 1304 is a parameter extractor/coder ( 1306. The parameter extractor/coder 1306 is configured to calculate and code parameters such as, for example, a spectral envelope parameter, a noise addition parameter, a missing harmonics parameter, or an inverse filtering parameter. The extracted parameters are input into a bit stream multiplexer 1308. The low pass output signal is input into a processor, which generally includes the functionality of a down sampler 1310 and a core coder 1312. Lowpass 1302 limits the bandwidth to be encoded to a bandwidth that is significantly less than that occurring in the original input audio signal on line 1300 . This provides significant coding gain due to the fact that the entire function occurring in the core coder must operate on a signal with reduced bandwidth. For example, when the bandwidth of the audio signal on line 1300 is 20 kHz and the low-pass filter 1302 preferably has a bandwidth of 4 kHz, in order to satisfy the sampling theorem, the signal after the downsampler It is theoretically sufficient to have a sampling frequency of 8 kHz, which is a substantial reduction to the required sampling rate for the audio signal 1300, which should be at least 40 kHz.
13b shows a schematic diagram of a corresponding bandwidth extension decoder; The decoder includes a bitstream multiplexer 1320 . The bitstream demultiplexer 1320 includes an input signal for a coder decoder 1322 and an input signal for a parameter decoder 1324 . The core decoder output signal, in the example above, has a sampling rate of 8 kHz, and thus a bandwidth of 4 kHz, whereas for full bandwidth reconstruction, the output signal of the high-frequency reconstructor 1330 requires a sampling rate of at least 40 kHz. must be present at 20 kHz. To enable this, a decoder processor having the functionality of an upsampler 1325 and a filter bank 1326 is required. High frequency reconstructor 1330 then receives the low frequency signal output frequency analyzed by filterbank 1326 and reconstructs the frequency range defined by high pass filter 1304 of FIG. 13A using the parametric representation of the high frequency band. do. The high-frequency reconstructor 1330 has several functions, such as reproducing the upper frequency range using the source range in the lower frequency range, adjusting the spectral envelope, adding noise, and introducing lossy harmonics into the upper frequency range, and if When applied and calculated within the encoder, it has the operation of inverse filtering to account for that the high frequency range is generally not as tonal as the low frequency range. In high-efficiency advanced audio coding, the lossy harmonics are resynthesized on the decoder plane and located precisely in the middle of the reconstruction band. Thus, all lossy harmonic lines determined within a particular reconstruction band are not located at frequency values located in the original signal. Instead, such lossy harmonic lines are located at frequencies within the center of a particular band. Thus, when the lossy harmonic lines in the original signal are located very close to the reconstruction band boundary of the original signal, the error in frequency introduced by placing these lossy harmonics in the reconstructed signal at the center of the band is the difference between the individual parameters generated and transmitted. It is close to 50% of the reconstruction band.
Moreover, although typical audio core coders operate within the spectral domain, the core decoder nevertheless generates a time domain signal, which is then converted back into a spectral diagram by the filter bank 1326 functionality. This introduces additional processing delay and may introduce artifacts due to the tanden processiong of first converting from spectral domain to frequency domain and back generally to different frequency domains, which of course A significant amount of computational complexity and thereby bandwidth expansion requires power, which is a problem when applied to mobile devices such as cell phones, tablets or laptop computers.
Current audio codecs implement low bitrate audio coding using bandwidth extension as a component of the coding strategy. Bandwidth extension techniques are limited to replacing only high frequency content. Moreover, it does not allow perceptually significant content above a given crossover frequency to be waveform coded. Therefore, current audio codecs lose high frequency detail or timbre when bandwidth extension is implemented, because the exact alignment of the tonal harmonics of a signal is not taken into account in most systems.
Another disadvantage of the bandwidth extension systems of the current art is the need for a transformation into a new domain (eg, transformation from a transformed discrete cosine transform to an orthogonal symmetric filter domain) for the implementation of bandwidth extension of an audio signal. This leads to the complexity of synchronization, additional computational complexity and increased memory requirements.
The storage or transmission of audio signals is sometimes subject to strict bitrate constraints. In the past, coders have been forced to drastically reduce the transmitted audio bandwidth when only very low bitrates are available. Modern audio coders can now code wideband signals by using bandwidth extension methods ([1-2]). These algorithms rely on the parametric representation of the high frequency content (HF), and the waveform coded low frequency part (LF) of the decoded signal by the application of permutation ("patching") into the high frequency spectral domain and parametric drive post-processing. arises from
In bandwidth extension strategies, the reconstruction of a high-frequency spectral region above a given so-called crossover frequency is often based on spectral patching. Other schemes that serve to fill a spectral gap, for example intelligent gap filling (intelligent gap filling (IGF)), use adjacent so-called spectral tiles to regenerate portions of the audio signal high frequency spectrum. . In general, a high frequency region consists of a number of adjacent patches, each of which originates from band-pass regions of the low frequency spectrum below a given crossover frequency. Modern systems efficiently implement patching within the filterbank representation by copying a set of adjacent subband coefficients from the source to the target region. Still, for some signal content, adjacent patches within the high-frequency band and the set of reproduced (reconstructed) signals from the low-frequency band may exhibit beating, dissonance and auditory roughness. ) can cause
So, in [19], dissonant guard-band (guard-band) filtering is presented in the context of a filterbank-based BEW system. Approximately 1 at the cross-over frequency between LF and BWE (bandwidth extension)-reproduced high frequency (HF) to replace spectral content with zeros or noise and avoid the possibility of dissonance. It is proposed to effectively apply a notch filter of the Bark bandwidth.
However, the solutions proposed in [19] have some drawbacks. First, rigorous replacement of spectral content with either zeros or noise can deteriorate the perceptual quality of the signal. Furthermore, the proposed treatment (processing) is not signal adaptive and thus may compromise the perceptual quality in some cases. For example, if the signal contains transients (transients), this can cause pre- and post-echoes.
Second, dissonance may occur at transitions (transitions) between successive HF patches. The solution proposed in [19] only functions to deal with the dissonance that occurs at the cross-over frequency between the LF and BWE-regeneration HF frequencies.
Finally, as opposed to filter bank based systems as proposed in [19], BWE systems (bandwidth extension systems) may be realized in transform based embodiments such as, for example, a modified discrete cosine transform (MDCT). . Transforms such as MDCT are the so-called wobbling [20] or It is prone to cause ringing artifacts.
In particular, US Pat. No. 8,412,365 discloses, in filterbank based transformation or folding, the use of so-called guard-bands (guard-bands) that are made and inserted into one or several subband channels set to zero. disclose The number of filterbank channels is used according to the guard-bands, and the bandwidth of the guard-band is 0,5 Bark. These dissonance guard-bands are partially reconstructed using random white noise signals, ie, subbands are input with white noise instead of zero. Guard bands are inserted regardless of the current signal to be processed.
<p>It is an object of the present invention to provide an improved concept for decoding an encoded audio signal.</p>
<p>The object of the present invention is achieved by an apparatus for decoding an encoded audio signal of claim 1 , a method for decoding an encoded audio signal of claim 16 or a computer program of claim 17 .</p><p>As such, in contrast to a fixed decoder-setup, where patching or frequency tiling is performed in a fixed manner, i.e. a particular source range is taken from the core signal and certain fixed frequency boundaries are defined between the source range and the recovery range. A signal-dependent patching or tiling is performed where it is applied to establish a frequency boundary between two adjacent frequency patches and tiles within the frequency or reconstruction range of The signal can be analyzed to find a local minima, and the core range is then selected such that the frequency boundaries of the core range coincide with a local minima in the core signal spectrum.</p><p>Alternatively or additionally, the signal analysis may be performed on the preliminary reproduction signal or the preliminary frequency-patching or tiling signal, wherein, after the preliminary frequency reproduction procedure, the boundary between the core range and the recovery range is determined as they beat upon recovery. ) are analyzed to detect any artifact-generating signal parts, such as the tonal parts that are problematic in that they are fairly close to each other to generate artifacts. Alternatively or additionally, the boundaries may be examined in such a way that halfway-clipping of the tonal portion is detected and this clipping of the tonal portion creates artifacts upon restoration as-is. To avoid these procedures, the frequency boundary of two separate frequency tiles or patches of the reconstruction range and/or of the source range and/or of the reconstruction range is referred to a signal manipulator to perform reconstruction again to the newly established boundaries. It can be modified (modified) by</p><p>Additionally, or alternatively, frequency regeneration is a regeneration based on the analysis result in that the frequency boundaries are left as they are, between two individual frequency tiles or patches within the reconstruction range or between the reconstruction range and the source range. Removal or at least attenuation of problematic tonal parts close to the frequency boundaries of Such tonal portions may be adjacent tones, which may be halfway-clipped tonal portions or induce a beating artifact.</p><p>In particular, when a non-energy conservation transform such as MDCT is used, a single tone (single tone) is not directly mapped to a single spectral line. Instead, a single tone will be mapped to a group of spectral lines with specific amplitudes depending on the phase of the tone. When the patching operation clips these tonal parts, it induces artifacts after restoration even if a perfect restoration is applied in the MDCT restoration. This is due to the fact that the MDCT reconstructor requires a complete tonal pattern for the tones to accurately and finally reconstruct these tones. Due to the fact that clipping has occurred before, this is no longer possible, so a time-varying wobbling artifact will be created. Based on the analysis according to the present invention, the frequency regenerator attenuates the complete tonal portion producing artifacts, or, as previously discussed, by changing the corresponding boundary frequencies, or by applying both methods or in those tonal patterns. We will avoid this situation by reconstructing the clipped part based on some known knowledge.</p><p>Additionally or alternatively, cross-over filtering is performed to spectrally cross-over filter a first frequency tile and a second frequency tile or a first frequency tile having frequencies extending from a gap filling frequency to a first tile stop frequency. and spectrally cross-over filtering the decoded core signal.</p><p>This cross-over filtering is useful to reduce so-called filter ringing.</p><p>The advanced approach is primarily intended to be applied within bandwidth extension (BWE) based on transforms such as MDCT. Nevertheless, the techniques of the present invention are generally similar, for example within a Quadrature Mirror Filter bank (QMF) based system, in particular, for example as a real-valued QMF representation. , can be applied when the system is critically (critically, critically) sampled.</p><p>A progressive approach is that in spectral regions close to transition points (such as cross-over frequency or patch boundaries) acoustic roughness, beats and dissonances may occur if the signal content is very tonal. based on the observation that Thus, a proposed solution to the above drawbacks has been found in the state of the art, which consists in the subsequent attenuation or elimination of these components and the signal adaptive detection of tonal components in the transition regions. Attenuation or removal of these components may alternatively be achieved by zero or noise insertion, or preferably by foot-to-foot spectral interpolation of such components. Alternatively, the spectral location of the transitions may be a signal to be adaptively selected such that transition artifacts are minimized.</p><p>In addition, this technique may be used to reduce or even avoid filter ringing. Especially for transient-like signals, ringing is an audible and objectionable artifact. Filtering artifacts are caused by the so-called brick-wall nature of the filter in the transition band (steep transition from passband to stopband at cutoff frequency). Such filters can be efficiently implemented by setting one coefficient or group of coefficients to zero in the frequency domain of the time-frequency transform. So, in the case of bandwidth extension, we propose to apply a cross-over filter at each transition frequency between patches or between the first patch and the core band to reduce the ringing effect. A cross-over filter can be implemented by spectral weighting in the transform domain using suitable gain functions.</p><p>According to a further aspect of the present invention, an apparatus for decoding an encoded audio signal comprises a core decoder, for generating one or more spectral tiles having frequencies not included in the decoded core signal using a spectral portion of the decoded core signal. A cross-over filter for spectrally cross-over filtering the tile generator and the decoded core signal and for spectrally cross-over filtering the tile and additional frequency tiles or extending from the gap filling frequency to the first tile stop frequency a first frequency tile having frequencies, and a further frequency tile having a lower border frequency that is frequency-adjacent to an upper border frequency of the frequency tile.</p><p>Preferably, this procedure is intended to be applied within a bandwidth extension based transform such as MDCT. However, the present invention is generally applicable, in particular in bandwidth extension scenarios that depend on quadrature mirror filterbanks (QMF), especially when there is a real-valued QMF representation according to a frequency-to-time transform or according to a time-to-frequency transform. As an example, it is particularly applicable when the system is critically sampled.</p><p>The embodiment is particularly useful for transient-like signals, since for such transient-like signals, ringing is an audible and objectionable artifact. Filtering artifacts are caused by the so-called brick-wall characteristic of the filter in the transition band, ie a steep transition from the passband to the stopband at the cutoff frequency. Such filters can be efficiently implemented by setting one coefficient or group of coefficients to zero in the frequency domain of the time-frequency transform. Thus, the present invention relies on a cross-over filter at each transition frequency between patches/tiles or between the core band and the first patch/tile to reduce these ringing artifacts. A cross-over filter can be implemented by spectral weighting in the transform domain using suitable gain functions.</p><p>Preferably, the cross-over filter is signal-adaptive and comprises two filters, a fade-out filter applied to the low spectral region, and a fade-in filter applied to the high spectral region. Filters may be symmetrical or asymmetrical depending on the particular embodiment.</p><p>In a further embodiment, a frequency tile or frequency patch is not the only subject of cross-over filtering, but the tile generator preferably has, before performing cross-over filtering, remaining in transition ranges around the transition frequencies. Perform patch adaptation including removal or attenuation of tonal parts and setting of frequency boundaries at spectral minima.</p>
<p>An audio coding system efficiently codes arbitrary audio signals over a wide range of bitrates. On the other hand, for high bitrates, the system of the present invention focuses on transparency, since low bitrate perceptual annoyance is minimized. Thus, the main share of the available bitrate is used to waveform-code the very perceptually relevant structure of the signal in the encoder, and the resulting spectral gaps are filled in the decoder with the signal content approximately approximating the original spectrum. A very limited bit budget is expended to control the so-called spectral intelligent gap filling, which is parameter driven by dedicated side information transmitted from the encoder to the decoder.</p>
Preferred embodiments are discussed hereinafter with reference to the accompanying drawings. 1a shows an apparatus for encoding an audio signal; Fig. 1b shows an apparatus for decoding an encoded audio signal matching the encoder of Fig. 1a; Figure 2a shows one preferred implementation of a decoder. Figure 2b shows one preferred implementation of the encoder. Fig. 3a shows a schematic representation of the spectrum as generated by the spectral domain decoder of Fig. 1b; 3B shows a table showing the relationship between the scale factors for the scale factor bands and the energies for the reconstruction bands and the noise filling information for the noise filling band. 4a shows the functionality of a spectral domain encoder for applying a selection of a spectral part into a first and a second set of spectral parts; 4B shows an implementation of the functionality of FIG. 4A. Figure 5a shows the functionality of a modified discrete cosine transform encoder. Figure 5b shows the functionality of a decoder with a modified discrete cosine transform technique. 5C shows an implementation of a frequency regenerator. Fig. 6a shows an apparatus for decoding an encoded audio signal according to an embodiment. 6b is a further embodiment of an apparatus for decoding an encoded audio signal; Fig. 7a shows a preferred embodiment of the frequency regenerator of Figs. 6a or 6b; 7b shows a further embodiment of cooperation between the frequency regenerator and the analyzer. 8 shows a further embodiment of a frequency regenerator. 8b shows a further embodiment of the present invention. Fig. 9a shows a decoder with frequency regeneration technique using energy values for the reproduction frequency range. Fig. 9b shows a more detailed implementation of the frequency reproduction of Fig. 9a. Fig. 9c schematically illustrates the function of Fig. 9b. Fig. 9d shows another implementation of the decoder of Fig. 9a. Fig. 10a shows a block diagram of an encoder matching the decoder of Fig. 9a; Fig. 10B shows a block diagram for illustrating another function of the parameter calculator of Fig. 10A. Fig. 10c shows a block diagram illustrating another function of the parameter calculator of Fig. 10a. Fig. 10D shows a block diagram illustrating another function of the parameter calculator of Fig. 10A. 11A shows the spectrum of filter ringing around transients. 11B shows a spectrogram of a transient after bandwidth extension is applied. 11C shows the spectrum of the transient after applying bandwidth extension with filtering reduction. 12a shows a block diagram of an apparatus for decoding an encoded audio signal; 12B shows a (stylized) magnitude spectrum of a tonal signal, copy-up without patch/tile adaptation, further removal of artifact-generated tonal portions and copy-up with changed frequency boundaries. 12C shows an example cross-face function. 13A shows a prior art encoder with bandwidth extension. 13b shows a prior art decoder with bandwidth extension. 14a shows a further apparatus for decoding an encoded audio signal using a cross-over filter; 14B shows a more detailed description of an example cross-over filter.
6a shows an apparatus for decoding an encoded audio signal comprising an encoded core signal and parameter data. The apparatus includes a core decoder 600 for decoding the encoded core signal to obtain a decoded core signal and an analyzer 602 for analyzing the decoded core signal before or after performing a frequency reproduction operation. The analyzer 602 is configured to provide an analysis result 603 . The frequency regenerator 604 uses the analysis result 603 and the envelope data 605 for missing spectral portions, the spectral portion of the decoded core signal, and the spectral portion not included in the decoded core signal. are configured to play. As such, in contrast to the earlier embodiments, frequency reproduction is not performed signal-independently at the decoder-side, but is performed signal-dependently. This has the advantage that when there is no problem, the frequency reproduction is performed as is, but when there are problematic signal parts, it is detected by the analysis result 603 and the frequency reproducer 604 is then, for example, restored An adapted manner of frequency reproduction may be performed, which may be a change in the frequency boundary between two individual tiles/patches in the band or a change in the initial frequency boundary between the restored band and the core region. In contrast to the embodiment of guard-bands, this has the advantage that it is always performed without any signal-dependence, as in guard-band implementation, except when specific procedures are required.
Preferably, the core decoder 600 is implemented with entropy (eg, Huffman or arithmetic decoder) decoding and dequantization steps 612 as shown in FIG. 6B . The core decoder 600 then outputs the core signal spectrum and the spectrum is implemented as a spectrum analyzer rather than any analyzer that may analyze a time domain signal, as shown in FIG. 6A , analyzer 602 in FIG. 6A . ), is analyzed by a spectrum analyzer 614 . In the embodiment of FIG. 6b , the spectrum analyzer is configured to analyze the spectrum signal and local (local) minima are determined in the source band and/or in the target band, ie in frequency patches or frequency tiles. do. Frequency regenerator 604 then performs frequency reproduction where the patch boundaries are located at minimum values in the source band and/or target band, as shown at 616 .
7A is discussed below to describe a preferred embodiment of the frequency regenerator 604 of FIG. 6A. A preliminary signal regenerator 702 receives, as an input, source data from a source band, and additionally, preliminary patch information such as preliminary boundary frequencies. Thereafter, a preliminary regenerated signal 703 is generated, which is detected by a detector 704 that detects tonal components in the preliminary reconstructed signal 703 . do. Alternatively or additionally, the source data 705 may be analyzed by a detector corresponding to the analyzer 602 of FIG. 6A . Thereafter, the preliminary signal expression step will not be necessary. When there is a well-defined mapping from the source data to the reconstructed data, as will be discussed later in view of Fig. 12b, at the frequency boundary or core range between two separately generated frequency tiles. Whether there are tonal portions that are close to (closer to) an upper border of , then minimum values or tonal portions may be detected in consideration of only the source data.
If problematic tonal components are found near the frequency boundaries, a transition frequency adjuster 706 is used between the individual frequency parts generated by the same source data and one in the recovered band or between the recovered band and the core. Adjustment of a transition frequency such as a gap filling start frequency or a cross-over frequency or a transition frequency between bands is performed. The output signal of block 706 is forwarded to a remover 708 of tonal components at the boundaries. The canceller is configured by block 706 to remove residual tonal components that are still there after the transition frequency adjustment. The result of the remover 708 is then forwarded to the cross-over filter 710 to deal with the filter ringing problem, and the result of the cross-over filter 710 is then the spectral envelope of the reconstructed band. input to a spectral envelope shaping block 712 that performs spectral envelope shaping.
As discussed in the context of FIG. 7A , the detection of the tonal components of block 704 may be performed on both the preliminary reconstruction signal 703 or the source data 705 . This embodiment is shown in FIG. 7B , where a preliminary regeneration signal is generated as shown in block 718 . The signal corresponding to signal 703 of FIG. 7A is then forwarded to a detector 720 that detects artifact-generating components. Although detector 720 is configured to be a detector that detects tonal components at frequency boundaries as shown by 704 in FIG. 7A, the detector may be implemented to detect other artifact-generating components. . Such spectral components may be components other than tonal components and detection of whether artifacts were generated is performed by trying different reproductions and comparing different reproduction results to find out which provided the artifact-generating components. can be
Detector 720 now controls a manipulator 722 that manipulates the signal, ie the preliminary regeneration signal. This manipulation can be dealt with by actually processing the preliminary playback signal by line 723 or by performing a new playback now, with the modified transition frequencies, for example as shown by line 724 . have.
One embodiment of the operating procedure is that the transition frequency is adjusted as shown in FIG. 7A . As a further embodiment is shown in FIG. 8A , this may be performed in conjunction with or instead of block 706 in FIG. 7A . A detector 802 is provided to detect the start and end frequencies of the tonal portion in question. An interpolator 804 is then configured to interpolate and, preferably, complex interpolating between the beginning and the end of the tonal portion within the spectral range. Then, as shown in FIG. 8A by block 806, the tonal portion is replaced by the interpolation result.
An alternative embodiment is illustrated in FIG. 8A by blocks 808 and 810 . Instead of performing interpolation, random generation of spectral lines 808 is performed between the beginning and end of the tonal portion. Then, energy adjustment of the randomly generated spectral lines is performed as shown at 810 , and the randomly generated spectral lines are set so that the energy is similar to adjacent non-tonal spectral parts. The tonal portion is then envelope-tuned and replaced by randomly generated spectral lines. The spectral lines are generated randomly (randomly) or pseudo-randomly to provide an artifact-free replacement signal where possible.
A further embodiment is shown in FIG. 8B . A frequency tile generator located within frequency regenerator 604 of FIG. 6A is shown at block 820 . The frequency tile generator uses predetermined frequency boundaries. The analyzer then analyzes the signal generated by the frequency tile generator, which frequency tile generator 820 preferably performs a plurality of tiling operations to generate a plurality of frequency tiles. is composed Thereafter, the manipulator 824 of FIG. 8B manipulates the result of the frequency tile generator according to the analysis result output by the analyzer 822 . Manipulation may be a change in frequency boundaries or attenuation of individual parts. Thereafter, a spectral envelope adjuster 826 performs spectral envelope adjustment using the parameter information 605 as already discussed in the context of FIG. 6A .
The spectrally tuned signal output by block 826 is input to a frequency-to-time converter which additionally receives the first spectral portions, ie, a spectral representation of the output signal of the core decoder 600 . The output of the frequency-to-time converter 828 may then be used to be transmitted or stored to a loudspeaker of audio rendering.
The present invention can also be applied to known frequency regeneration procedures as shown in Figs. 13a, 13b or preferably within the context of intelligent gap filling, which is described hereinafter with respect to Figs. 1a to 5b and 9a to 10d.
1a shows an apparatus for encoding an audio signal 99 . An audio signal 99 is input into a time spectrum converter 100 for converting an audio signal having a sampling rate into a spectral representation 101 output by a time spectrum converter. The spectrum 101 is input into a spectrum analyzer 102 for analyzing the spectral representation 101 . The spectrum analyzer 102 is configured to determine a first set of first spectral portions 103 to be encoded at a first spectral resolution and another second set of second spectral portions 105 to be encoded at a second spectral resolution. do. The second spectral resolution is less than the first spectral resolution. A second set of second spectral portions 105 is input into a parameter calculator or parameter coder 104 for calculating spectral envelope information having a second spectral resolution. Furthermore, a spectral domain audio coder 105 is provided for generating a first encoded representation of first spectral portions having a first spectral resolution. Furthermore, the parameter calculator/parameter coder 106 is configured to generate a second encoded representation of the second set of second spectral portions. The first encoded representation 107 and the second encoded representation 109 are input into a bit stream multiplexer or bit stream former 108 and block 108 is finally encoded for transmission or storage on a storage device. Outputs an audio signal.
Generally, a first spectral portion such as 306 in FIG. 3A will be surrounded by two second spectral portions such as 307a, 307b. This is not the case in high-efficiency advanced audio coding where the core coder frequency range is band-limited.
Fig. 1B shows decoder matching with the encoder of Fig. 1A; A first encoded representation 107 is input into a spectral domain audio decoder 112 for generating a first decoded representation of a first set of first spectral portions, the decoded representation having a first spectral resolution. Furthermore, the second encoded representation 109 is input into the parameter decoder 114 to generate a second decoded representation of a second set of second spectral parts having a second spectral resolution lower than the first spectral resolution.
The decoder further comprises a frequency generator 116 for generating a reconstructed second spectral portion having a first spectral resolution using the first spectral portion. Frequency regenerator 116 performs a tile filling operation, ie, tiles or first spectral portions use a first set and these first spectral portions copy the first set into a reconstruction range or reconstruction band having a second spectral portion. and perform spectral envelope shaping or other operations as represented by the decoded second representation, typically output by the parametric decoder 114 , ie by using information about the second set of second spectral parts. The decoded first set of first spectral parts and the reconstructed second set of spectral parts as shown in the output of the frequency reproducer 116 on line 117 time the first decoded representation and the reconstructed second spectral part. It is input into a spectral-time converter 118 that is configured to transform into a representation 119 , the temporal representation having a certain high sampling rate.
2B shows an implementation of the encoder of FIG. 1A. The audio input signal 99 is input into the temporal spectrum converter 100 and the corresponding analysis filterbank 220 of FIG. 1A . A temporal noise shaping operation is then performed within the temporal noise shaping block 222 . Thus, the input into the spectrum analyzer 102 of FIG. 1A corresponding to the block tone mask 226 of FIG. 2B may be full spectral values when no temporal noise shaping/temporal tile shaping operation is applied or as shown in FIG. 2B . When the temporal noise shaping operation block 222 is applied as described above, it may be spectral residual values. For two-channel signals or multi-channel signals, joint channel coding 228 may additionally be performed, and thus the spectral domain encoder 106 of FIG. 1A may include a joint channel coding block 228 . In addition, an entropy coder 232 is provided for performing lossless data compression, which is also part of the encoder 106, which is the spectral diagram of FIG. 1A.
The spectrum analyzer/tonal mask 226 compares the output of the temporal noise shaping block 222 to the core band and tonal components corresponding to the first set of first spectral portions 103 and the second spectral portions ( 105) and the corresponding residual components. Block 224, denoted as intelligent gap filling parameter extraction encoding, corresponds to parameter core 104 in FIG. 1A and bitstream multiplexer 230 corresponds to bitstream multiplexer 108 in FIG. 1A.
Preferably, the analysis filterbank 222 is implemented as a modified discrete cosine transform filterbank and the modified discrete cosine transform filterbank transforms the signal 99 into the time-frequency domain with a modified discrete cosine transform acting as a frequency analysis tool. used as a guide
The spectrum analyzer 226 preferably applies a tonality mask. The tonal mask estimation step is used to separate the tonal components from the noise-like components in the signal. This allows the core coder 228 to encode all tonal components into psychoacoustic modules. The tonal mask estimation step can be implemented in a variety of different methods, preferably sine and noise-modeling for speech/audio coding ([8, 9]) or HILN (Harmonic and Individual Line) described in [10]. plus Noise) model-based audio coders. Preferably, an implementation that is easy to implement without the need to dnb the birth-death trajectory is used, but other tonal or noise detectors may also be used.
The intelligent gap filling module calculates the similarity that exists between the source region and the target region. The target area will be represented by the spectrum from the source area. The measurement of similarity between source and target regions is performed using a cross-correlation approach. the target area<i>nTar</i> It is divided into non-overlapping frequency tiles. For all tiles within the target area, from a fixed start frequency<i>nSrc</i> Source tiles are created. These source tiles overlap by a factor between 0 and 1, where 0 means 0% overlap and 1 means 100% overlap. Each of these source tiles is correlated with a target tile at various lags to find a source tile that best matches the target tile.
The number of tiles that best match is <i>tileNum</i><i>[</i><i>idx</i><i>_tar]</i> The lags stored within and most correlated with the target are <i>xcorr</i><i>_lag[</i><i>idx</i><i>_tar][</i><i>idx</i><i>_src]</i> is stored within and the sign of the correlation is <i>xcorr</i><i>_sign[</i><i>idx</i><i>_tar][</i><i>idx</i><i>_src]</i> stored within When the correlation is highly negative, the source tile needs to be multiplied by -1 before the tile filling process at the decoder. The intelligent gap filling module also handles overwriting of tonal components in the spectrum, since the tonal components are preserved using a tonal mask. The band method energy parameter is used to store energy in the target region, which makes it possible to accurately reconstruct the spectrum.
This method is a classical spectrum in that the harmonic grid of the multi-tone signal is preserved by the core coder and the gaps between the sinusoids are preserved by "shaped noise" that best matches from the source region. It has certain advantages over band replication ([1]). Another advantage of these systems compared to Accurate Spectral Replacement (ASR) [2-4] is the absence of a signal synthesis step that generates a significant part of the signal at the decoder. Instead, this task is generated by the core coder, allowing preservation of important components of the spectrum. Another advantage of the proposed system is the continuous scalability provided by the features. The use of only tileNum[idx_tar] and xcorr_lag=0 is called total granularity matching for all tiles and can be used for lower bitrates while using the variable xcorr_lag and makes it possible to better match the target and source spectra for all tiles make it
In addition, a tile selection stabilization technique that removes frequency domain artifacts such as trilling and music noise is proposed.
In the case of stereo channel pairs an additional joint stereo procedure is applied. This is necessary because for a specific destination range the signal can be a highly correlated panned sound source. In the case where the source regions selected for this particular region are poorly correlated, the spatial image may be deteriorated due to the non-correlated source regions. The encoder analyzes each destination region energy band, typically performing cross-correlation of spectral values and, if a certain threshold is exceeded, sets a joint flag for this energy band. At the decoder, the left and right channel energy bands are processed separately, unless these joint stereo flags are set. In case the joint stereo flag is set, both energies and patching are performed within the joint stereo domain. Joint stereo information for intelligent gap filling regions is signaled similarly to joint stereo information for core coding, in which case the prediction contains a flag indicating whether the direction of prediction is to be residual from downmix or vice versa.
The energies may be calculated from the transmitted energies in the left/right domain.
<i>midNrg</i>[<i>k</i>] = <i>leftNrg</i>[<i>k</i>] + <i>rightNrg</i>[<i>k</i>];
<i>sideNrg</i>[<i>k</i>] = <i>leftNrg</i>[<i>k</i>] - <i>rightNrg</i>[<i>k</i>];
where k is the frequency exponent in the transform domain.
Another solution is to calculate and transmit the energies directly in the joint stereo domain for the bands where joint stereo is active, so no additional conversion is required at the decoder side.
Source tiles are always generated according to a mid/side matrix:
<i>midTile</i>[<i>k</i>] = 0.5·(<i>leftTile</i>[<i>k</i>] + <i>rightTile</i>[<i>k</i>])
<i>sideTile</i>[<i>k</i>] = 0.5·(<i>leftTile</i>[<i>k</i>] - <i>rightTile</i>[<i>k</i>])
Energy adjustment:
<i>midTile</i>[<i>k</i>] = <i>midTile</i>[<i>k</i>] * <i>midNrg</i>[<i>k</i>];
<i>sideTile</i>[<i>k</i>] = <i>sideTile</i>[<i>k</i>] * <i>sideNrg</i>[<i>k</i>];
Joint Stereo LR Conversion
If no additional prediction parameters are coded:
<i>leftTile</i>[<i>k</i>] = <i>midTile</i>[<i>k</i>] + <i>sideTile</i>[<i>k</i>]
<i>rightTile</i>[<i>k</i>] = <i>midTile</i>[<i>k</i>] - <i>sideTile</i>[<i>k</i>]
If an additional prediction parameter is coded and the signaled direction is from the middle to the side:
<i>sideTile</i>[<i>k</i>] = <i>sideTile</i>[<i>k</i>] - <i>predictionCoeff</i>·<i>midTile</i>[<i>k</i>]
<i>leftTile</i>[<i>k</i>] = <i>midTile</i>[<i>k</i>] + <i>sideTile</i>[<i>k</i>]
<i>leftTile</i>[<i>k]</i> = <i>midTile</i>[<i>k</i>] - <i>sideTile</i>[<i>k</i>]
If the signaled direction is halfway from the side:
<i>midTile</i>[<i>k</i>] = <i>midTile</i>[<i>k</i>] - <i>predictionCoeff</i>·<i>sideTile</i>[<i>k</i>]
<i>leftTile</i>[<i>k</i>] = <i>midTile</i>[<i>k</i>] - <i>sideTile</i>[<i>k</i>]
<i>leftTile</i>[<i>k</i>] = <i>midTile</i>[<i>k</i>] + <i>sideTile</i>[<i>k</i>]
This process ensures that from the tiles used to generate highly correlated destination regions and panned destination regions, the resulting left and right channels still represent a correlated and panned sound source, even if the source regions are uncorrelated. guarantees and preserves the stereo image for those areas.
In other words, in the bitstream, joint stereo flags indicating whether left/right or middle/side should be used as examples for general joint stereo coding are transmitted. At the decoder, first, the core signal is decoded as indicated by the joint stereo flags for the core bands. Second, the core signal is stored in both right/left and middle/side, for intelligent gap filling tile filling, source tile representation to fit the target tile representation as indicated by joint stereo information for intelligent gap filling bands this is chosen
Temporal noise shaping is a standard technique and is part of advanced audio coding ([11-13]). Temporal noise shaping can be considered as an extension of the basic strategy of the perceptual coder, inserting an optional extrapolation step between the filterbank and the quantization step. The main task of the temporal noise shaping module is to hide the quantization noise produced within the temporal masking region of transient-like signals, thus leading to a more efficient coding strategy. First, temporal noise shaping computes a set of prediction coefficients using "forward prediction" in the transform domain, eg, a transformed discrete cosine transform. These coefficients are then used to flatten the temporal envelope of the signal. Since quantization affects the temporal noise shaping filtered spectrum, the quantization noise is temporally flat. By applying the inverse temporal noise shaping on the decoder side, the quantization noise is shaped according to the temporal envelope of the temporal noise shaping filter and thus the quantization noise is masked by the transients.
Intelligent gap filling is based on a transforming discrete cosine transform representation. For efficient coding, preferably long blocks of about 20 ms should be used. If the signal in such a long block contains transients, audible pre- and post-echoes occur in the intelligent gap filling spectral bands due to tile filling. Figure 7c shows a typical pre-echo effect before transient initiation due to intelligent gap filling. On the left side, the spectrogram of the original signal is shown and on the right side the spectrogram of the bandwidth extended signal without temporal noise shaping filtering.
The pre-echo effect is reduced using temporal noise shaping within the context of intelligent gap filling. Here, the temporal noise shaping is used as the temporal tile shaping because the spectral reproduction in the decoder is performed on the temporal noise shaping residual signal. In general, the necessary temporal tile shaping prediction coefficients are calculated and applied using the full spectrum on the encoder plane. The temporal noise shaping/temporal tile shaping start and stop frequencies are the intelligent gap filling start frequency (f) of the intelligent gap filling tool.<sub>IGFstart</sub>) is not affected by Compared with the legacy temporal noise shaping, the temporal tile shaping stop frequency is multiplied by the stop frequency of the intelligent gap filling tool, which is higher than the intelligent gap filling start frequency. On the decoder side the temporal noise shaping/temporal tile shaping coefficients are again applied over the full spectrum, ie the core spectrum plus the reproduced spectrum plus the tonal components from the tonal map (see Fig. 7e). The application of temporal tile shaping is necessary to shape the temporal envelope of the reproduced spectrum in order to match the envelope of the original signal again. Thus, the pre-echoes shown are reduced. Moreover, as in temporal noise shaping, we still shape the quantization noise in the signal below the intelligent gap filling start frequency.
In legacy decoders, spectral patching for an audio signal corrupts the temporal envelope of the audio signal by introducing an error in the spectral correlation at patch boundaries and thereby introducing dispersion. Thus, another benefit of performing intelligent gap filling tile filling on residual signal is that tile boundaries are uniformly correlated after application of the shaping filter, resulting in a more reliable temporal reproduction of the signal.
In the encoder of the present invention, the spectrum with the specified temporal noise shaping/temporal tile shaping, tonal mask processing and intelligent gap filling parameter estimation, there is no signal above the intelligent gap filling start frequency except for tonal components. This sparse spectrum is now coded by the core coder using the principles of arithmetic coding and predictive coding. Together with the signaling bits, these coded components form the bitstream of the audio.
Figure 2a shows a corresponding decoder implementation. The bitstream of FIG. 2a corresponding to the encoded audio signal is input into a demultiplexer/decoder which may be coupled to blocks 112 and 114 with respect to FIG. 1b. The bitstream demultiplexer separates the input audio signal into an input audio signal of a first encoded representation 107 of FIG. 1B and a second encoded representation 109 of FIG. 1B . A first encoded representation having a first set of first spectral portions is input into a joint channel decoding block 204 corresponding to the spectral domain decoder of FIG. 1B . The second encoded representation is input into a parameter decoder 114 not shown in FIG. 2A and then into a frequency regenerator 116 of 1B and a corresponding intelligent gap filling block 202 . A first set of first spectral parts necessary for frequency reproduction is input via line 203 into intelligent gap filling block 202 . In addition, a specific core decoding is applied in the tonal mask block 206 after the joint channel decoding 204 such that the output of the tonal mask corresponds to the output of the spectral domain decoder 112 . Then, combining by combiner 208, i.e., frame building, where the output of combiner 208 now has a full range spectrum, but still exists within the temporal noise shaping/temporal tile shaping filtered domain is performed. Then, in block 210, an inverse temporal noise shaping/temporal tile shaping operation is performed using the temporal noise shaping/temporal tile shaping information provided via line 109, i.e., the temporal tile shaping side information is preferably may be included in the first encoded representation generated by the spectral domain encoder 106 , which may be, for example, simple advanced audio coding or unified speech audio coding, or may be included in a second encoded representation. At the output of block 210, a full spectrum is provided, a full range of frequencies defined by the sampling rate of the original input signal up to the maximum frequency. Then, a spectral/time conversion is performed in the synthesis filterbank 212 to finally obtain an audio output signal.
Figure 3a shows a schematic representation of the spectrum. The spectrum is subdivided into scale factor bands SCB, in which there are seven scale factor bands SCB1 to SCB7 in the example shown in FIG. 3a. The scale factor bands may be an advanced audio coding scale factor defined in the advanced audio coding standard and with increasing bandwidth up to the upper frequencies as schematically shown in FIG. 3 . It is desirable to start the intelligent gap filling operation at the intelligent gap filling start frequency shown at 309 , rather than from the very beginning of the spectrum, ie, at the lower frequencies. Thus, the core frequency band extends from the spectrum lowest frequency to the intelligent gap filling frequency. Separating the high resolution spectral components 304, 305, 306, 307, the first set of first spectral parts, above the intelligent gap filling start frequency, from the lower resolution components represented by the second set of second spectral parts. For this purpose, spectral analysis is applied. Figure 3a preferably shows the spectrum input into the spectral domain encoder 106 or joint channel coder 228, i.e. the core encoder operates within a full range, but encodes a significant amount of zero spectral values, i.e. these Zero spectral values are either quantized to zero or set to zero after quantization before quantization. In any case, the core encoder operates within full scope, i.e. the spectrum can be as shown, i.e. the core decoder will not necessarily perceive any intelligent gap filling or encoding of a second set of spectral parts with low spectral resolution. No need.
Preferably, the high resolution is defined by line-wise coding of spectral lines, such as modified discrete cosine transform lines, and the second or lower resolution is defined, for example, by calculating one single spectral value per scale factor band. However, the scale factor includes some frequency lines. Thus, the second lower resolution, in terms of its spectral resolution, is generally much lower than the first or higher resolution defined by the linewise coding applied by a core encoder, such as an advanced audio coding or integrated speech audio coding core encoder. .
Regarding the scale factor or energy calculation, the situation is shown in FIG. 3B . Due to the fact that the encoder is a core encoder, and due to the fact that it may but need not be a component of the first set of spectral parts within each band, the core encoder is not only below the intelligent gap filling start frequency 309 , but also at the sampling frequency. half of, i.e. f<sub>s</sub><sub>/2</sub>A maximum frequency similar to or equal to (f<sub>IGFstop</sub>) to calculate the scale factor for each band within the core range above the intelligent gap filling start frequency. Accordingly, the encoded tonal portions 302 , 304 , 305 , 306 , 307 of FIG. 3A correspond to the high resolution spectral data together with the scale factors SCB1 to SCB7 in this embodiment. The low resolution spectral data is computed starting from the intelligent gap filling start frequency and transmitted together with the scale factors SCB1 to SCB7, energy information values E<sub>1</sub>, E<sub>2,</sub> E<sub>3</sub>, E<sub>4</sub>) corresponds to
In particular, when the core encoder is under the low bitrate state, an additional noise-filling operation may be applied in the core band, ie, a frequency lower than the intelligent gap filling start frequency, ie, in the scale factor bands SCB1 to SCB7. On the decoder side, these values quantized to zero are resynthesized and the resynthesized spectral values are NF shown at 308 in Fig. 3b.<sub>2</sub>Adjusted within their amplitudes using noise-filling energies such as The noise-filling energy, which can be given in absolute terms or relative terms with respect to a scale factor, in particular as in unified speech audio coding, corresponds to the energy of a set of zero-quantized spectral values. These noise-filling spectral lines also contain source range and energy information (E<sub>1</sub>, E<sub>2,</sub> E<sub>3</sub>, E<sub>4</sub>) a third set of third spectral parts reproduced by simple noise-filling synthesis without any intelligent gap filling operation that relies on frequency reproduction using frequency tiles from other frequencies to reproduce frequency tiles from other frequencies can be considered as
Preferably, the bands in which the energy information is calculated coincide with the scale factor bands. In other embodiments, energy information value grouping is applied and thus only a single energy information value is transmitted, for example for scale factor bands 4 and 5, but also in this embodiment the boundaries of the grouped reconstruction bands are coincides with the boundaries of the scale factor bands. If different band separation is applied, specific re-computations or synchronization calculations may be applied, as may be understood depending on the specific implementation.
Preferably, the spectral domain encoder 106 of Fig. 1A is a psychoacoustic driven encoder as shown in Fig. 4A. In general, the audio signal to be encoded after being converted to a spectral range (401 in Fig. 4a) is passed to a scale factor calculator 400, as indicated for example in the MPEG2/4 advanced audio coding standard or the MPEG1/2 layer 3 standard. do. The scale factor calculator is further controlled by a psychoacoustic model that receives the audio signal to be encoded or receives the complex spectral signal of the audio signal as in MPEG1/2 Layer 3 or MPEG Advanced Audio Coding Standard. The psychoacoustic model computes, for each scale factor band, a scale factor representing the psychoacoustic limit. Additionally, the time factors are then adjusted such that certain bitrate conditions are met, either by well-known inner or outer iterative loops or by some other suitable encoding procedure. The spectral values to be quantized on the one hand and the calculated scale factors on the other hand are then input into a quantizer processor 404 . In simple audio encoder operation, the spectral values to be quantized are weighted by scale factors, and the weighted scaled spectral values are then input into a fixed quantizer, which generally has a compression function for the upper amplitude ranges. It is then passed into an entropy encoder with specific and highly efficient coding, such as a set of zero-quantization indices for generally adjacent frequency values at the output of the quantizer processor, or conventionally referred to as a "run of zero values." There are quantization indices.
However, in the audio encoder of FIG. 1A , the quantizer processor generally receives information about the second spectral parts from the spectrum analyzer. Thus, the quantizer processor 404 determines that, at the output of the quantizer processor 404 , the second spectral portions as defined by the spectrum analyzer 102 are zero or in particular there will be a "run" of zero values in the spectrum. , to have a representation recognized by the encoder or decoder as a zero representation that can be coded very efficiently.
4B shows an implementation of a quantizer processor. The transformed discrete cosine transform spectral values may be input into a set set to zero block 410 . Then, the second spectral portions are already set to zero before weighting by the spectral factors in block 412 . In an additional implementation, block 410 is not provided, but a set-to-zero cooperation is performed at block 418 after weight block 412 . In another implementation, a set to zero operation in the set to zero block 422 may also be performed after quantization in the quantization block 420 . In this implementation, blocks 410 and 413 may not be present. Generally, at least one of blocks 410 , 418 , 422 is provided depending on the particular implementation.
Then, at the output of block 422, a quantized spectrum corresponding to that shown in FIG. 3A is obtained. The quantized spectrum is then input into an entropy coder, such as 232 in FIG. 2B , which may be a Huffman coder or an arithmetic coder, for example as defined in the Unified Speech Audio Coding Standard.
Alternatively, the set-to-zero blocks 410 , 418 , 422 provided in parallel with each other are controlled by the spectrum analyzer 424 . The spectrum analyzer preferably comprises any implementation of a well-known tonal detector or any other kind of detector operable to separate the spectrum into components to be encoded in high resolution and components to be encoded in low resolution. Such other algorithms implemented in the spectrum analyzer may be a voice activity detector, a noise detector, a voice detector or any other detector that determines according to spectral information or related metadata for resolution requirements for different spectral parts. have.
Fig. 5a shows a preferred implementation of the temporal spectral transformer 100 of Fig. 1a, for example as implemented in advanced audio coding or integrated speech audio coding. The time spectrum converter 100 includes a windower 502 controlled by a transient detector 504 . When the transient detector 504 detects a transient, a switchover from long windows to short windows is signaled to the windower. Windower 502 then computes windowed frames for overlapping blocks, each windowed frame having two values of N, typically equal to 2048 values. Then, a transform in a block transformer 506 is performed, which generally provides additional decimation, thus obtaining a spectral frame with N values equal to the transformed discrete cosine transform spectral values. A combined decimation/transformation is performed to Thus, for long window operation, the frame at the input of block 506 contains two N values equal to 2048 values and the spectral frame then has 1024 values. However, then the transition to short blocks is executed when 8 short blocks are executed, each short block having a time domain 1/8 windowed compared to the long window and each spectral block compared to the long block. It has 1/8 spectral values. Thus, when decimation is combined with the windower's 50% overlap operation, the spectrum is a significantly sampled version of the time domain audio signal 99 .
Reference is then made to FIG. 5B , which further shows a specific implementation of the operation of the frequency regenerator 118 and the spectral-time converter 118 of 1B , or blocks 208 , 212 of FIG. 2A . In Fig. 5b, a specific reconstruction band is considered, such as the scale factor band 6 in Fig. 3a. The first spectral portion within this reconstruction band, ie the first spectral portion 306 of FIG. 3A is input into a frame builder/manipulator block 510 . In addition, the reconstructed second spectral portion for the scale factor band 6 is also input into the frame builder/manipulator 510 . In addition, E in FIG. 3b for scale factor band 6<sub>3</sub>Energy information is also input into block 510 . The reconstructed second spectral portion within the reconstruction band has already been generated by frequency tile filling using the source region and the reconstruction band then corresponds to the target range. Now, then an energy adjustment of the frame is performed to obtain a finally complete reconstructed frame with the value of N as obtained, for example, at the output of the combiner 208 of FIG. 2A . Then at block 512, for example, an inverse block transform/interpolation is performed to obtain 248 time domain values for the 124 spectral values at the input of block 512 . Then, a composite windowing operation controlled again by the long window/short window indication transmitted as additional information in the encoded audio signal is executed. Then, at block 516, an overlap/add operation with a previous time frame is performed. Preferably, the transform discrete cosine transform applies a 50% overlap, so for each new time frame of 2N values, N time domain values are finally output. 50% overlap is highly desirable due to the fact that it provides significant sampling and continuous crossover from one frame to the next due to the overlap/add operation in block 516.
As shown at 301 of Fig. 3a, the noise-filling operation is additionally applied below the intelligent gap filling start frequency as well as above the intelligent gap filling start frequency, such as for the considered reconstruction band consistent with the scale factor band of Fig. 3a. can Then, the noise-filling spectral values may also be input into the frame builder/adjuster 510 and adjustment of the noise-filling spectral values may also be applied within this block or the noise-filling spectral values may be input into the frame builder/adjuster 510 . ) can already be adjusted using noise-filling energy before being input into .
Preferably, an intelligent gap filling operation, ie a frequency tile filling operation using spectral values from different parts, can be applied within the complete spectrum. Therefore, the spectral tile filling operation can be applied not only in the high band above the intelligent gap filling start frequency, but also in the low band. In addition, noise-filling without frequency tile filling can also be applied not only below the intelligent gap filling start frequency but also above the intelligent gap filling start frequency. However, high-quality and high-efficiency audio encoding can be achieved when the noise-filling operation is limited to a frequency range below the intelligent gap-filling start frequency, and when the frequency tile filling operation is in the frequency range above the intelligent gap-filling start frequency, as shown in Fig. 3a. It has been found that it can be obtained when limited to
Preferably, the target tail (TT, with frequencies greater than the intelligent gap filling start frequency) is directed to the scale factor band boundaries of the full ratio coder. The source tiles ST from which information is obtained are not bound by scale factor band boundaries, ie for frequencies lower than the intelligent gap filling start frequency. The size of the source tile must correspond to the size of the associated target tile. This is illustrated using the following example. TT[0] has the size of 10 variant discrete cosine transform bins. This corresponds exactly to the length of the two following scale factor bands (such as 4+6). Then all possible source tiles correlated with TT[0] have a length of 10 bins. The second target tile TT[1] adjacent to TT[0] has a length of 15 bins (scale factor band with a length of 7+8). Then, the source tile for this has a length of 15 bins rather than 10 bins for TT[0].
If a case arises that a target tile for a source tile having the length of the target tile cannot be found (eg when the length of the target tile is greater than the available source range), then no correlation is calculated and the source range is the target tile. It is copied several times into this target tile until (TT) is completely filled (copying alternates so that the frequency line for the lowest frequency of the second radiation immediately follows the frequency line for the highest frequency of the first radiation).
Reference is then made to FIG. 5C , which shows another preferred embodiment of the frequency regenerator 116 of FIG. 1B or the intelligent gap filling block 202 of FIG. 2A . Block 522 is a frequency tile generator that receives a target band identification as well as additionally a source band identification. Preferably, it has been determined that the scale factor band 3 of FIG. 3a on the encoder plane is very suitable for reconstructing the scale factor band 7 . Thus, the source band identification may be 2 and the target band identification may be 7. Based on this information, frequency tile generator 522 applies radiation up to a harmonic tile filling operation or any other filling operation to generate raw second portions of spectral components 523 . The raw second portions of the spectral components have a frequency resolution equal to the frequency resolution included in the first set of first spectral portions.
A first spectral portion of the reconstruction band, such as 307 in FIG. 3A , is then input into frame builder 524 and a raw second portion 523 is also input into frame builder 524 . The reconstructed frame is then adjusted by the adjuster 526 using the gain factor for the reconstruction band calculated by the gain factor calculator 528 . Importantly, however, the first spectral portion within the frame is not affected by the adjuster 526 , but only the raw second portion for the reconstructed frame is affected by the adjuster 526 . To this end, the gain factor calculator 528 analyzes the source band or raw second part 523 and additionally the energy of the adjusted frame output by the adjuster 526 when the scale factor band 7 is taken into account. energy (E<sub>4</sub>), the first spectral portion in the reconstruction band is analyzed to finally find the correct gain factor 527 .
In this context, it is very important to evaluate the high-frequency reconstruction accuracy of the present invention compared to the high-efficiency advanced audio coding. This is explained in relation to the scale factor band 7 of FIG. 3A . It is presumed that a conventional encoder such as that shown in Fig. 13a can detect the spectral portion 307 that is to be encoded with high resolution as "lossy harmonics". The energy of this spectral component can then be transmitted to the decoder along with the spectral envelope information for the reconstruction band, such as the scale factor band (7). The decoder can then regenerate the lossy harmonics. However, the spectral value from which the lossy harmonic 307 may be reconstructed by the conventional decoder of FIG. 13A may be in the middle of the band 7 at the frequency represented by the reconstruction frequency 390 . Thus, the present invention avoids the frequency error 391 that may be introduced by the conventional decoder of FIG. 13D.
In one implementation, the spectrum analyzer is also implemented to calculate similarities between the first spectral parts and the second spectral parts and based on the calculated similarities, the second spectrum part as soon as possible for the second spectral part within the reconstruction range. A first spectral portion matching the portion is determined. Then, in this variable source/destination range implementation, a parameter coder will additionally be introduced into the second encoded representation and the matching information indicates the matching source range for each destination range. On the decoder side, this information can then be used by the frequency tile generator 522 of FIG. 5C representing the generation of the raw second portion 523 based on the source band identification and the target band identification.
Moreover, as shown in Figure 3a, the spectrum analyzer is configured to analyze the spectral representation in a small amount, less than half the sampling frequency, preferably at least one quarter of the sampling frequency, or generally up to a high maximum analysis frequency.
As shown, the encoder operates without downsampling and the decoder operates without upsampling. In other words, the spectral domain audio coder is configured to generate a spectral representation having a Nyquist frequency defined by the sampling rate of the original input audio signal.
Furthermore, as shown in FIG. 3a , the spectrum analyzer is configured to analyze the spectral representation starting with a gap filling start frequency and ending with a maximum frequency expressed in terms of a maximum frequency contained within the spectrum representation, from the minimum frequency to the gap filling start frequency. the spectral portion extending to belongs to the first set of spectral portions and another spectral portion having frequency values above the gap filling frequency, such as 304, 305, 306, 307, is additionally within the first set of first spectral portions Included.
As described, the spectral domain audio decoder 112 is configured such that the maximum frequency represented by the spectral value in the first decoded representation is equal to the maximum frequency included in the temporal frequency with the sampling rate, and the second of the first spectral portions. The spectral value for the maximum frequency in one set is zero or is different from zero. In any case, the scale factor is generated or transmitted irrespective of whether all spectral values in the scale factor band for the maximum frequency in the first set of spectral components are set to zero or not as discussed in the context of Figures 3a and 3b. There is a scale factor for the band.
Accordingly, the present invention relates to other parametric techniques for increasing the compression efficiency, such as noise replacement and noise filling (these techniques are exclusively for the efficient representation of local signal content such as noise), the present invention relates to the tonal component It is desirable in that it allows accurate frequency reproduction of the So far, no state-of-the-art technology addresses the efficient parametric representation of arbitrary signal content by spectral gap filling without the limitation of fixed a-priori subdivision within the low band and high band.
Embodiments of the system of the present invention enhance conventional approaches and thereby provide high extrusion efficiency even at low bitrates, no or very little perceptual annoyance, and full audio bandwidth.
*A typical system consists of:
Full-band core coding
Intelligent gap filling (tile filling or noise filling)
Sparse tonal parts in the core selected by the tonal mask
Joint stereo pair coding for full band, including tail filling
Temporal noise shaping on tiles
Spectral whitening within the range of intelligent gap filling
A first step towards a more efficient system is to eliminate the need to transform the spectral data into a second transform domain different from the core coder. It is also useful to implement a bandwidth extension in the transformed discrete cosine transform domain because most audio codecs, such as advanced audio coding for example, use the transformed discrete cosine transform as the default transform. A second requirement of a bandwidth extension system may be the need to preserve tonal grids in which high frequency tonal components are preserved and the quality of the coded audio is thus superior to existing systems. In order to address all of the requirements for the above-mentioned bandwidth extension strategy, a new system called intelligent gap filling is proposed. Fig. 2b shows a diagram of the proposed system on the encoder side and Fig. 2a shows the system on the decoder side.
9a shows an apparatus for decoding an encoded audio signal comprising an encoded representation of a first set of first spectral parts and an encoded representation of parameter data representing spectral energies for a second set of second spectral parts; do. A first set of spectral portions is indicated at 901a in FIG. 9A , and an encoded representation of the parameter data is indicated at 901B in FIG. 9A .
The audio decoder 900 decodes the encoded representation 901a of the first set of first spectral parts to obtain a first set of decoded first spectral parts 904 , and separate for respective reconstruction bands. provided for decoding an encoded representation of the parameter data to obtain decoded parameter data 902 for a second set of second spectral portions representing energies, the second spectral portions being located within the reconstruction bands. In addition, a frequency reproducer 906 is provided for reconstructing the spectral values of the reconstruction band comprising the second spectral portion. Frequency regenerator 906 uses one first spectral portion of the first set of first spectral portions and one respective energy information for a reconstruction band, the reconstruction band comprising a first spectral portion and a second spectral portion .
The frequency regenerator 906 includes a calculator 912 for determining survival energy information comprising accumulated energy of the first spectral portion having frequencies within the reconstruction band. In addition, the frequency regenerator 906 includes a calculator 918 for determining tile energy information of further spectral portions of the reconstruction band and determining frequency values different from the first spectral portion, these frequency values being the frequencies within the reconstruction band. , and other spectral parts are generated by frequency reproduction using a first spectral part different from the first spectral part in the reconstruction band.
The frequency regenerator 906 further includes a calculator 914 for the lost energy in the reconstruction band, the calculator 914 operating using the individual energy for the reconstruction band and the survival energy generated by the block 912 . In addition, the frequency regenerator 906 includes a spectral envelope adjuster 916 for adjusting further spectral portions within the reconstruction band based on the lost energy information and the energy information generated by block 918 .
Reference is made to FIG. 9C , which shows a specific reconstruction band 920 . The reconstruction band includes a first spectral portion within the reconstruction band, such as first spectral portion 306 in FIG. 3A schematically shown at 921 . Furthermore, the remainder of the spectral values in the reconstruction band are generated using the source region, for example from the scale factor bands 1, 2, 3 below the intelligent gap filling start frequency 309 in FIG. 3A . The frequency generator 906 is configured to generate raw spectral values for the second spectral portions 922 and 923 . Then, finally frequency bands 922 , 923 to obtain the now reconstructed and adjusted second spectral parts within the reconstruction band 920 having the same spectral resolution, ie the same line distance as the first spectral part 921 . A gain factor g is calculated as shown in FIG. 9c to adjust the raw spectral values in .It is important to understand that the first spectral portion within the reconstruction band shown at 921 of FIG. 9C is decoded by the audio decoder 900 and is not affected by the envelope adjustment performed block 916 of FIG. 9B . Instead, the first spectral portion in the reconstruction band indicated at 921 remains as it is, since this spectral portion is output by the full bandwidth or full ratio audio decoder 900 via line 904 .
After that, a specific example with real numbers is described. The remaining viable energy as calculated by block 912 is, for example, 5 energy units and this energy is preferably the energy of the 4 spectral lines indicated in the first spectral portion 921 .
Moreover, the energy value E3 for the reconstruction band corresponding to the scale factor band 6 of FIG. 3b or 3a is equal to 10 units. Importantly, the energy value is not only the energy of the spectral parts 922, 923,
It contains the full energy of the reconstruction band 920 as computed on the encoder-plane, ie prior to performing the spectral analysis using, for example, a tonal mask. Thus, ten energy units cover the first and second spectral portions within the reconstruction band. Then, the energy of the source range data for blocks 922 and 923 or raw target range data for blocks 922 and 923 is assumed to be equal to 8 energy units. Thus, the energy loss of 5 units is calculated.
Based on the lost energy subdivided by the tile energy tEk, a gain factor of 0.79 is calculated. The raw spectral lines for the second spectral portions 922 and 923 are then multiplied by the calculated gain factor. Accordingly, only the spectral values for the second spectral parts 922 , 923 are adjusted and the spectral lines for the first spectral parts 921 are not affected by this envelope adjustment. After multiplying the raw spectral values for the second spectral parts 922 , 923 , the spectral line in the second spectral parts 922 , 923 in the reconstruction band 920 is composed of the first spectral parts in the reconstruction band 920 . A complete reconstruction band consisting of
Preferably, the source range for generating the raw spectral data in bands 922 , 923 is below the intelligent gap filling start frequency 309 with respect to frequency and the reconstruction band 920 is the intelligent gap filling start frequency 309 . ) is above.
Moreover, it is desirable that the reconstruction band boundaries coincide with the scale factor band boundaries. Thus, the reconstruction band, in one embodiment, has a corresponding size of the scale factor bands of the core audio decoder or is sized such that when an energy pair is applied, the energy value for the reconstruction band provides an integer of two or more of the scale factor bands. get angry Thus, when energy accumulation is assumed to be performed for scale factor band 4 , scale factor band 5 and scale factor band 6 , the low frequency boundary of reconstruction band 920 is that of scale factor band 4 . The same as the low boundary and the high frequency boundary of the reconstruction band 920 is equal to the high boundary of the scale factor band 6 .
Thereafter, Fig. 9D is explained to show another function of the decoder of Fig. 9A. The audio decoder 900 receives dequantized spectral values corresponding to first spectral parts of a first set of spectral parts, and additionally scale factors for scale factor bands as shown in FIG. 3B are inversely scaled. Block 940 is provided. The inverse scaling block 940 includes all first set of first spectral portions below the intelligent gap filling start frequency 309 of FIG. 3A , and additionally first spectral portions above the intelligent gap filling start frequency 309 , ie the first spectral portions 304, 305, 306, 307 of FIG. 3A are all located within the reconstruction band as shown at 941 of FIG. 9D. In addition, the first spectral portions in the source band used for frequency tile filling in the reconstruction band are provided to an envelope adjuster/calculator 942 and this block is additionally provided to the encoded audio signal as shown at 943 in Fig. 9d. Receives energy information for a reconfiguration band provided as parameter additional information for The envelope adjuster/calculator 942 then provides the functions of FIGS. 9B and 9C and finally outputs adjusted spectral values for the second spectral portions within the reconstruction band. These adjusted spectral values 922 , 923 of the second spectral portions within the reconstruction band and the first spectral portions within the reconstruction band indicated by line 941 of FIG. 9D jointly represent the complete spectral representation within the reconstruction band.
Reference is then made to Figures 10a to 10b to describe preferred embodiments of an audio encoder for encoding an audio signal to provide or generate an encoded audio signal. The encoder comprises a time/spectrum converter 1002 which supplies a spectrum analyzer 1004 , which is connected on the one hand to a parameter calculator 1006 and on the other hand to an audio encoder 1008 . Audio encoder 1008 provides an encoded representation of a first set of first spectral portions and does not include a second set of second spectral portions. On the other hand, the parameter calculator 1006 provides energy information for the reconstruction band comprising the first and second spectral portions. Furthermore, the audio encoder 1008 is configured to generate a first encoded representation of a first set of first spectral portions having a first spectral resolution, wherein the audio encoder 1008 is configured to generate a spectrum generated by block 1002 . Provides scale factors for all bands of representation. Additionally, as shown in FIG. 3B , the encoder provides, with respect to frequency, energy information for the reconstruction bands located at least above the intelligent gap filling start frequency 309 as shown in FIG. 3A . Thus, preferably for reconstruction bands matching the scale factor bands or matching groups of scale factor bands, there are two values, namely the corresponding scale factor from the audio encoder 1008 and additionally a parameter calculator. The energy information output by 1006 is given.
The audio encoder preferably has different frequency bandwidths, ie scale factor bands with different numbers of spectral values. Accordingly, the parameter calculator comprises a normalizer 1012 for normalizing the energies for the bandwidths that differ significantly with respect to the bandwidth of the particular reconstruction band. With this in mind, the normalizer 1012 receives as inputs the energy within the band and the number of spectral values within the band and the normalizer 1012 then outputs the normalized energy per band for the reconstruction/scale factor.
In addition, the parameter calculator 1006a of FIG. 10A includes an energy value calculator that receives control information from the core or audio encoder 1008 as shown by line 1007 of FIG. 10A . Such control information may include information on long/short blocks used by the audio encoder and/or grouping information. Thus, the information for long/short blocks and the grouping information for short windows relate to "time" grouping, and the grouping information additionally refers to spectral grouping, i.e. the grouping of the two scale factor bands into a single reconstruction band. can do. Accordingly, the energy value calculator 1006 outputs a single energy value for each grouped band comprising the first and second spectral portions when only the spectral portions are grouped.
10D shows another embodiment for implementing spectral grouping. To this end, block 1016 is configured to calculate energy values for two adjacent bands. Then, in block 1018, the energy values for the adjacent bands are compared, and when the energy values do not differ significantly or slightly as defined by, for example, a threshold, as indicated in block 1020, A single (normalized) value for both bands is generated. As shown by line 1024 , block 1018 can be bypassed. In addition, the generation of a single value for two or more bands implemented by block 1020 may be controlled by encoder bitrate control 1024 . Thus, when the bitrate is reduced, the encoded bitrate control 1024 generates a single normalized value for two or more bands even though the comparison in block 1018 is not allowed to group energy information values. control block 1020 to do so.
In case the audio encoder implements a grouping of two or more short windows, this grouping is also applied for energy information. When the core encoder performs grouping of two or more short blocks, for these two or more blocks, only a single set of scale factors is calculated and transmitted. On the decoder-side, the audio decoder then applies the same set of scale factors for both grouped windows.
Regarding the energy information calculation, the spectral values in the reconstruction band are accumulated over two or more short windows. In other words, this means that the spectral values within a specific reconstruction band for a short block and the following short block are accumulated together and only a single energy information value is transmitted for this reconstruction band comprising two short blocks. Then on the decoder-plane, the envelope adjustments described with respect to FIGS. 9A-9D are not performed individually for each short block, but together for a grouped set of short windows.
Corresponding normalization is then performed for energy value information calculation on the decoder-plane, no matter what grouping in frequency or grouping in time has been performed, on the one hand the energy information value and the quantity of spectral lines in the reconstruction band or set of grouped reconstruction bands Easy to allow to be known.
In addition, information about spectral energies, information about individual energies or individual energy information, information about survival energy or surviving energy information, information about tile energy or tile energy information, or information about lost energy Alternatively, the lost energy information may include an energy value as well as an (eg, absolute) amplitude value, a level value, or some other value from which a final energy value may be derived. As such, the information about the energy may include, for example, a value of absolute amplitude and/or an amplitude and/or a level and/or the energy value itself.
12a shows a further embodiment of an apparatus for decoding; The bitstream is received by the core decoder 1200 , which may be, for example, an AAC decoder. The result consists, for example, of performing bandwidth extension patching or tiling 1202 corresponding to the frequency regenerator 604 . Thereafter, the procedure and post-processing of patch/tile adaptation is performed, and when the patch adaptation is performed, the frequency reproducer 1202 is controlled to, for example, perform additional frequency reproduction with adjusted frequency boundaries. Moreover, when patch processing is performed, such as by removal and attenuation of tonal lines, the result is, for example, block 1206 of performing parametric-driven bandwidth envelope shaping as discussed in the context of blocks 712 or 826. is forwarded (sent, forwarded). The result is then forwarded to the synthesis transformation block 1208 to perform transformation into a final output region, which is a PCM output region, for example, as shown in FIG. 12A .
The main features of embodiments of the present invention are as follows:
A preferred embodiment is the case where the tonal spectral regions are reduced by poor selection of cross-over frequency and/or patch margins, or where the tonal components are located too close in the vicinity of the patch boundaries, see above It is based on the MDCT representing the warbling artifacts.
Figure 12b shows how the newly proposed technique reduces the artifacts found in the latest bandwidth extension methods. In Fig. 12 panel (2), a stylized magnitude spectrum of the output of a modern bandwidth extension scheme is shown. In this example, the signal is perceptually deteriorated by the beatings caused by two adjacent tones, and also by division of the tone. Both problematic spectral regions are each marked with a circle.
To overcome these problems, the new technique first detects the spectral positions of the tonal components included in the signal. Then, according to one aspect of the present invention, adjusting the transition frequencies between the LF (low frequency) and all patches by individual shifts (within the given limits) such that splitting or beating of the tonal components is minimized. thing is tried For this purpose, the transition frequency should preferably match the regional spectral minimum. This step is shown in Fig. 12b panel (2) and panel (3), where the transition frequency<img file="KR20160024924A_D0001.tif" /> is shifted to high frequencies, <img file="KR20160024924A_D0002.tif" />to derive
According to another aspect of the present invention, if problematic spectral content remains in the transition region, at least one misplaced tonal component is removed to reduce wobbling or beating artifacts at transition frequencies. . This is processed through spectral extrapolation or interpolation/filtering, as shown in Fig. 2 panel (3). The tonal component is thus removed from foot-point to foot-point, ie from its left local minimum to its right local minimum. The resulting spectrum after application of the present technology is shown in Figure 12b panel (4).
In other words, FIG. 12b shows the original signal, in the upper left corner, ie in panel 1 . In the upper right corner, ie in panel 2 , the comparative bandwidth extension signal is shown with the regions in question marked by ellipses 1220 and 1221 . In the lower left corner, ie in panel 3, two preferred patch or frequency tile processing features are shown. The splitting of tonal parts is a frequency boundary<i>f'</i><i><sub>x2</sub></i> is processed so that there is no more clipping of the corresponding tonal part. In addition, a gain function 1030 for removing tonal portions 1031 and 1032 is applied, or, alternatively, an interpolation shown by 1033 is indicated. Finally, the lower right corner of Fig. 12b, i.e. panel 4, shows on one side an enhanced signal resulting from the tile/patch frequency combination or at least attenuation of problematic tonal parts.
Panel (1) of FIG. 12b shows that, as previously discussed, the original spectrum is the cross-over or gap filling start frequency. <i>fx1</i>up to the core frequency range.
As such, the frequency <i>f</i><i><sub>x1</sub></i> is the Nyquist frequency (<i>f</i><i><sub>Nyquist</sub></i>) shows the boundary frequency 1250 between the source range 1252 and the recovery range 1254 extending between the maximum frequency and boundary frequency 1250 less than or equal to On the encoder-side, the signal is<i>f</i><i><sub>x1</sub></i> is assumed to be bandwidth-limited in , or when the technique for intelligent gap filling is applied, <i>f</i><i><sub>x1</sub></i> It is assumed that this corresponds to the gap filling start frequency 309 of FIG. 3A . Based on the above technique,<i>f</i><i><sub>x1</sub></i> The above reconstruction range will be empty (in the case of FIGS. 13A, 13B ) and will include certain first spectral parts to be encoded in high resolution as discussed in the context of FIG. 3A .
Fig. 12b, panel 2 shows a preliminary regenerated signal, eg generated by block 702 of Fig. 7a having the two parts in question. One problematic part is shown at 1220 . The frequency distance between the tonal portion in the core region shown at 1220a and the tonal portion at the beginning of the frequency tile shown at 1220b will be too small to create a beating artifact. A further problem is a halfway-clipped or split tonal portion 1226 at the upper boundary of the first frequency tile generated by the frequency tiling operation or the first patching operation shown in 1225. am. When this tonal portion 1226 is compared with the other tonal portions of FIG. 12B , it will become clear that its width is smaller than that of a normal tonal portion, which means that this tonal portion is placed in the wrong position in the source range 1252 . This means that it is divided by setting a frequency boundary between the first frequency tile 1225 and the second frequency tile 1227 . To address these issues, boundary frequencies<i>f</i><i><sub>x2</sub></i> is modified to be slightly larger as shown in panel (3) of Fig. 12b, and no clipping of this tonal portion occurs.
On the other hand, <i>f'</i><i><sub>x2</sub></i> This modified procedure does not effectively address the beating problem, so it is addressed by any other procedures or removal of tonal components by interpolation or filtering as discussed in the context of block 708 of FIG. 7A . do. As such, FIG. 12B shows the removal of tonal portions at the boundaries shown at 708 and the sequential application of the transition frequency adjustment 706 .
Another option is the transition boundary <i>f</i><i><sub>x1</sub></i> It will be set so that the tonal portion 1220a is no longer in the core range so that α is slightly lower. Then, the tonal portion 1220a is a transition frequency (transition frequency) to a lower value (lower value)<i>f</i><i><sub>x1</sub></i> was removed or reduced by setting
This procedure worked to address the issue of having a problematic tonal component 1032 . <i>f'</i><i><sub>x2</sub></i> By setting to much higher, the spectral portion where the tonal portion 1032 is located could be reproduced within the first patching operation 1225, so that two adjacent or nearby tonal portions would not have occurred.
Basically, the beating problem depends on the distance and amplitude at the frequency of adjacent tonal parts. Detectors 704, 720 or, more generally mentioned, analyzer 602 are preferably used for locating any tonal component.<i>f</i><i><sub>x1</sub></i><i>, </i><i>f</i><i><sub>x2</sub></i><i><sub>, </sub></i><i>f'</i><i><sub>x2</sub></i> The analysis of subspectral parts located at frequencies below the same transition frequency is constructed in such a way that it is analyzed. In addition, the spectral range above the transition frequency is also analyzed to detect tonal components. When the detection results in two tonal components, one to the left of the transition frequency with respect to frequency and one to the right of the transition frequency (with respect to the rising frequency), then the boundaries shown at 708 in FIG. 7A . Removal of the tonal components in is activated. The detection of the tonal component is carried out in a specific detection range extending, from a transition frequency, which is at least 20% in both directions with respect to the bandwidth of the corresponding band, preferably upward with respect to the right of the transition frequency relative to the bandwidth. and extending only 10% downward to the left of the transition frequency, that is, in the bandwidth of the source range on the one hand and in the reconstruction range on the other side, or the transition frequency is between the two frequency tiles 1225 and 1227. When it is a transition frequency, it is in the amount of the corresponding 10% of the corresponding frequency tile. In a further embodiment, the predetermined detection bandwidth is one Bark. It must be possible to remove the tonal parts within the range of 1 bark around the patch boundary, so the full detection range is 2 barks, i.e. 1 bark of the low band and 1 bark of the high band, where 1 bark of the low band is the high band. It is immediately adjacent to 1 bar of the band.
According to another aspect of the present invention, in order to reduce the filtering artifact, a cross-over filter in the frequency domain is applied to two successive spectral regions, i.e. between the first patch and the core band or between two patches. applies. Preferably, the cross-over filter is signal adaptive.
A crossover filter is a fade-out filter applied to the lower spectral region. <img file="KR20160024924A_D0003.tif" />, and a fade-in filter applied to the high spectral region. <img file="KR20160024924A_D0004.tif" />It is composed of two filters.
each filter <i>N</i>has a length of
Furthermore, the slope of both filters is <img file="KR20160024924A_D0005.tif" />Together with, to determine the notch characteristics of the cross-over filter <img file="KR20160024924A_D0006.tif" />It is characterized by a signal adaptation value called
if <img file="KR20160024924A_D0007.tif" />, the sum of both filters is equal to 1, that is, there is no notch filter characteristic in the resulting filter.
if <img file="KR20160024924A_D0008.tif" />, then both filters are completely zero.
The basic design of cross-over filters is constrained by the following equations:
<img file="KR20160024924A_D0009.tif" />
<img file="KR20160024924A_D0010.tif" />
<img file="KR20160024924A_D0011.tif" />is the frequency index. 12C shows an example of such a cross-over filter.
In this example, the following equation is <img file="KR20160024924A_D0012.tif" />is used to create:
<img file="KR20160024924A_D0013.tif" />
The following equations are for filters <img file="KR20160024924A_D0014.tif" /> and <img file="KR20160024924A_D0015.tif" />Hereafter, how it is applied is explained.
<img file="KR20160024924A_D0016.tif" /><img file="KR20160024924A_D0017.tif" />
<img file="KR20160024924A_D0018.tif" />
<img file="KR20160024924A_D0019.tif" />represents an assembled spectrum, <img file="KR20160024924A_D0020.tif" />is the transition frequency, <img file="KR20160024924A_D0021.tif" />is the low-frequency content and <img file="KR20160024924A_D0022.tif" />is high-frequency content.
Next, evidence of the benefits of this technique is presented. In the following examples the original signal is a transient-like signal, in particular a low-pass filtered version thereof, with a cut-off frequency of 22 kHz. First, this transient is a band limited to 6 kHz in the transform domain. Then, the bandwidth of the low-pass filtered original signal is extended to 24 kHz. Bandwidth extension is achieved within the transform by replicating the LF (low frequency) band three times to completely fill the available frequency range above 6 kHz.
11a shows the spectrum of this signal, which can be considered as a general spectrum of filtering artifacts spectrally surrounding the transient due to the brick-wall nature of the transform (speech peaks 1100) . Applying the inventive approach, the filtering is reduced by approximately 20 dB at each transition frequency (reduced speech peaks).
The same effect is seen in Figures 11b, 11c, in different cities. Figure 11b shows the spectrogram of the mentioned transient-like signal with filtering artifacts temporally preceding and trailing the transient after application over the described bandwidth extension technique without any filtering reduction. Each of the horizontal lines represents filtering at the transition frequency between successive patches. 6 shows the same signal after applying the inventive approach within the bandwidth extension. Through application of the ringing reduction, the filter ringing is reduced by approximately 20 dB compared to the signal indicated in the previous figure.
14A, 14B are discussed below to further illustrate the cross-over filter invention aspect already discussed in the context of having an analyzer feature. However, the cross-over filter 710 may be implemented independently of the invention discussed in the context of FIGS. 6A-7B .
14A shows an apparatus for decoding an encoded audio signal including information on parameter data and an encoded core signal. The apparatus includes a core decoder 1400 that decodes the encoded core signal to obtain a decoded core signal. The coded core signal may have limited bandwidth in the context of FIG. 13A , and the FIG. 13B embodiment or core decoder may be a full rate coder or full frequency range (full rate coder) in the context of FIGS. 1-5C or 9A-10D frequency range).
In addition, a tile generator 1404 for reproducing one or more spectral tiles having frequencies not included in the decoded core signal is generated using the spectral portion of the decoded core signal. The tiles may be second spectral portions reconstructed within the reconstructed band, for example as shown in the context of FIG. 3A , or may include first spectral portions to be reconstructed with high resolution, but alternatively, Spectral tiles may also include completely empty frequency bands when the encoder performs strict band limiting as shown in FIG. 13dptj.
In addition, the cross-over filter 1406 is configured to cross-over filter the first frequency tile and the decoded core signal with frequencies extending from the gap filling frequency 309 to the first tile stop frequency or the second frequency tile and It is provided for spectrally cross-over filtering the first frequency tile 1225 , and the second frequency tile has a lower boundary frequency that is frequency-adjacent to the upper boundary frequency of the first frequency tile 1225 .
In a further embodiment, the cross-over filter 1406 output signal is envelope-adjusted to finally obtain an envelope-adjusted regenerated signal, parametric spectral envelope information contained in the encoded audio signal as parametric side information. is input to an envelope adjuster 1408 that applies Components 1404 , 1406 , 1408 may be implemented as a frequency regenerator, for example as shown in FIG. 13B , 1B or 6A .
14B shows a further embodiment of a cross-over filter 1406 . The cross-over filter 1406 includes a fade-out subfilter for receiving the first input signal IN1, and a second fade-in subfilter 1422 for receiving the second input IN2, and includes both filters 1420 and ( The results or outputs of 1422 are provided to a combiner 1424 which is, for example, an adder (adder). An adder or combiner 1424 outputs spectral values for frequency bins. 12C shows an example cross-fade function including a fade-out subfilter characteristic 1420a and a fade-in subfilter characteristic 1422a. Both filters have a certain frequency overlap (overlap) in the example of FIG. 12c equal to 21, ie N=21. As such, for example, other frequency values of the source region 1252 are unaffected. Only the highest 21 frequency bins of the source range 1252 are affected by the fade-out function 1420a.
On the other hand, only the lowest 21 frequency lines of the first frequency tile 1225 are affected by the fade-in function 1422a.
Additionally, although it is clear from the cross-fade function that the frequency lines between 9 and 13 are affected, the fade-in function does not actually affect the frequency lines between 1 and 9 and the fade-out function 1420a is 13 and 21 have no effect on frequency functions. It only needs overlap between frequency lines 9 and 13,<i>f</i><i><sub>x1</sub></i> This means that the same cross-over frequency is located in the frequency sample or frequency bin 11 . As such, only the frequency values between the source range and the first frequency tile or the overlap of two frequency bins will be required to implement the cross-over or cross-fade function.
Depending on the particular embodiment, high or low overlap may be applied and, in addition, other fading functions away from the cosine function may be used. Moreover, as shown in Fig. 12c, it is desirable to apply a specific notch in the cross-over range. Stated differently, the energy in the boundary range will be reduced due to the fact that both filter functions do not eventually integrate, as is the case with the notch-free cross-fade function. The energy loss on the boundaries of the frequency tile, ie the first frequency tile, will be attenuated at the upper side boundary and the lower side boundary, and the energies are more concentrated in the middle of the bands. However, due to the fact that the spectral envelope adjustment takes place after being processed by the cross-over filter, the overall frequency is not touched, and as discussed in the context of FIG. defined by In other words, the calculator 918 of FIG. 9B then calculates the "already generated raw target range", which is the output of the cross-over filter. Furthermore, the energy loss due to the removal of the tonal part by the interpolation will also be compensated for due to the fact that this removal results in low tile energy and the gain factor for the fully restored band will be higher. However, on the other hand, the cross-over frequency results in more energy concentration in the middle of the frequency tile, which in turn effectively reduces artifacts, partially caused by transients as discussed in the context of FIGS. 11A-11C . make it
14B shows different input combinations. For filtering at the boundary between the source frequency range and the frequency tile, input 1 is the upper spectral portion of the core range, and input 2 is the lower spectrum of the first frequency tile or, when only a single frequency tile is present, a single frequency tile. part In addition, the input may be a first frequency tile and the transition frequency may be the upper frequency boundary of the first tile and the input to subfilter 1422 will be the lower portion of the second frequency tile. When an additional third frequency tile is present, the additional transition frequency will be the frequency boundary between the second and third frequency tiles, and when the FIG. 12 characteristic is used, the input to the fade-out subfilter 1421 is will be the upper spectral range of the second frequency tile as determined by the filter parameter, and the input to the fade-in subfilter 1422 is the lower portion of the third frequency tile and the lowest 21 spectrum, as in the example of FIG. 12C . lines will be
As shown in Fig. 12c, it is desirable to make the parameter N the same for the fade-in sub-filter and the fade-out sub-filter. However, this is not essential.<i>N</i>The value for can vary and the result will then be that the filter "notch" will be asymmetric between the lower and upper ranges. Additionally, the fade-in/fade-out functions do not necessarily have to be of the same nature as in FIG. 12C . Instead, asymmetric properties may be used.
Furthermore, it is desirable to make the cross-over filter characteristic signal-adaptive. So, based on signal analysis, the filter characteristics are adaptive. Due to the fact that a cross-over filter is particularly useful for transient signals, it is detected whether transient signals occur. When transient signals occur, a filter characteristic such as that shown in FIG. 12C may be used. However, when a non-transient signal is detected, it is desirable to change the filter characteristic to reduce the influence of the cross-over filter. For example, this can be done by setting N to zero, or<i>X</i><i><sub>bias</sub></i> can be obtained by setting to zero, whereby the sum of both filters is equal to 1, i.e., there is no notch filter characteristic in the resulting filter. Alternatively, the cross-over filter 1406 can simply be bypassed in the case of non-transient signals. However, preferably, the parameters<i>N</i>, <i>X</i><i><sub>bias</sub></i> It is desirable to change the filter characteristics relatively slowly by changing Furthermore, it is desirable that the low-pass filter only allow such relatively small filter characteristic changes, even if the signal changes more rapidly, when detected by a particular transient/tonality detector. The detector is shown at 1405 in FIG. 14A . It may receive the output signal of the tile generator 1404 or the input signal to the tile generator or may be coupled to the core decoder 1400 to obtain transient/non-transient information, such as, for example, short block indications from AAC decoding. have. Naturally, any other crossover filter other than that shown in FIG. 12C may also be used.
Then, based on transient detection, or based on tonal detection, or based on any other signal characteristic detection, the cross-over filter 1406 characteristic is changed as discussed.
Although some aspects have been described in the context of an apparatus for encoding or decoding, it is to be understood that these aspects also represent a description of a corresponding method, wherein a block or apparatus corresponds to a method step or characteristic of a method step. Similarly, aspects described in the context of a method step also represent features of a corresponding block item or a corresponding apparatus. Some or all method steps may be executed by (or using) hardware devices such as, for example, microprocessors, programmable computers or electronic circuits. In some embodiments, one or more weighted critical method steps may be performed by such an apparatus.
Some embodiments in accordance with the present invention include a data carrier having electronically readable control signals capable of cooperating with a programmable computer system, such as by performing any of the methods described herein.
In general, embodiments of the present invention may be implemented as a computer program product having a program code, the program code operable to execute any one of the methods when the computer program product runs on a computer. The program code may be stored on, for example, a machine readable carrier.
Other embodiments include a computer program for executing any of the methods described herein, stored on a machine-readable carrier.
In other words, an embodiment of the method of the present invention is thus a computer program having a program code for executing any of the methods described herein when the computer program runs on a computer.
Another embodiment of the method of the present invention is thus a data carrier (or data storage medium, or computer readable medium) recorded therein, comprising a computer program for carrying out any of the methods described herein. A data carrier, digital storage medium, or recorded medium is generally tangible and/or non-transitory.
Another embodiment of the method of the present invention is thus a data stream or sequence of signals representing a computer program for executing any of the methods described herein. A data stream or sequence of signals may be configured to be transmitted, for example, over a data communication connection, for example the Internet.
Another embodiment comprises processing means, for example a computer, or a programmable logic device, configured or adapted to carry out any of the methods described herein.
Another embodiment comprises a computer installed therein with a computer program for performing any of the methods described herein.
Another embodiment according to the present invention comprises an apparatus or system configured to deliver (eg, electronically or optically) a computer to a receiver for performing any of the methods described herein. The receiver may be, for example, a computer mobile device, a memory device, or the like. The device or system may include, for example, a file server for delivering computer programs to receivers.
In some embodiments, a programmable logic device (eg, a field programmable gate array) may be used to perform some or all of the methods described herein. In some embodiments, the field programmable gate array may cooperate with a microprocessor to perform any of the methods described herein. In general, the methods are preferably executed by some hardware device.
The embodiments described above are merely illustrative of the principles of the present invention. It will be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Accordingly, it is intended that the present invention be limited only by the scope of the patent claims and not the specific details expressed by the description of the embodiments set forth herein.
List of citations
[One] Dietz, L. Liljeryd, K. Kjrling and O. Kunz, Spectral Band Replication, a novel approach in audio coding, in 112th AES Convention, Munich, May 2002.
[2] Ferreira, D. Sinha, Accurate Spectral Replacement, Audio Engineering Society Convention, Barcelona, Spain 2005.
[3] D. Sinha, A. Ferreira1 and E. Harinarayanan, A Novel Integrated Audio Bandwidth Extension Toolkit (ABET), Audio Engineering Society Convention, Paris, France 2006.
[4] R. Annadana, E. Harinarayanan, A. Ferreira and D. Sinha, New Results in Low Bit Rate Speech Coding and Bandwidth Extension, Audio Engineering Society Convention, San Francisco, USA 2006.
[5] T. ernicki, M. Bartkowiak, Audio bandwidth extension by frequency scaling of sinusoidal partials, Audio Engineering Society Convention, San Francisco, USA 2008.
[6] J. Herre, D. Schulz, Extending the MPEG-4 AAC Codec by Perceptual Noise Substitution, 104th AES Convention, Amsterdam, 1998, Preprint 4720.
[7] M. Neuendorf, M. Multrus, N. Rettelbach, et al., MPEG Unified Speech and Audio Coding-The ISO/MPEG Standard for High-Efficiency Audio Coding of all Content Types, 132nd AES Convention, Budapest, Hungary, April, 2012 .
[8] McAulay, Robert J., Quatieri, Thomas F. Speech Analysis/Synthesis Based on a Sinusoidal Representation. IEEE Transactions on Acoustics, Speech, And Signal Processing, Vol 34(4), August 1986.
[9] Smith, JO, Serra, X. PARSHL: An analysis/synthesis program for non-harmonic sounds based on a sinusoidal representation, Proceedings of the International Computer Music Conference, 1987.
[10] Purnhagen, H.; Meine, Nikolaus, "HILN-the MPEG-4 parametric audio coding tools," Circuits and Systems, 2000. Proceedings. ISCAS 2000 Geneva. The 2000 IEEE International Symposium on , vol.3, no., pp.201,204 vol.3, 2000
[11] International Standard ISO/IEC 13818-3, Generic Coding of Moving Pictures and Associated Audio: Audio, Geneva, 1998.
[12] M. Bosi, K. Brandenburg, S. Quackenbush, L. Fielder, K. Akagiri, H. Fuchs, M. Dietz, J. Herre, G. Davidson, Oikawa: "MPEG-2 Advanced Audio Coding", 101st AES Convention , Los Angeles 1996
[13] J. Herre, Temporal Noise Shaping, Quantization and Coding methods in Perceptual Audio Coding: A Tutorial introduction, 17th AES International Conference on High Quality Audio Coding, August 1999
[14] J. Herre, Temporal Noise Shaping, Quantization and Coding methods in Perceptual Audio Coding: A Tutorial introduction, 17th AES International Conference on High Quality Audio Coding, August 1999
[15] International Standard ISO/IEC 23001-3:2010, Unified speech and audio coding Audio, Geneva, 2010.
[16] International Standard ISO/IEC 14496-3:2005, Information technology Coding of audio-visual objects Part 3: Audio, Geneva, 2005.
[17] P. Ekstrand, Bandwidth Extension of Audio Signals by Spectral Band Replication, in Proceedings of 1st IEEE Benelux Workshop on MPCA, Leuven, November 2002
[18] F. Nagel, S. Disch, S. Wilde, A continuous modulated single sideband bandwidth extension, ICASSP International Conference on Acoustics, Speech and Signal Processing, Dallas, Texas (USA), April 2010
[19] Liljeryd, Lars; Ekstrand, Per; Henn, Fredrik; Kjorling, Kristofer: Spectral translation/folding in the subband domain, United States Patent 8,412,365, April 2, 2013.
[20] Daudet, L.; Sandler, M.; "MDCT analysis of sinusoids: exact results and applications to coding artifacts reduction," Speech and Audio Processing, IEEE Transactions on , vol.12, no.3, pp. 302- 312, May 2004.
58 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58
298 members in 22 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 131773467 | European Patent Office (EPO) | – | |
| 131773509 | European Patent Office (EPO) | – | |
| 131773533 | European Patent Office (EPO) | – | |
| 131773483 | European Patent Office (EPO) | – | |
| 13177346 | European Patent Office (EPO) | A | |
| 13177350 | European Patent Office (EPO) | A | |
| 13177353 | European Patent Office (EPO) | A | |
| 13177348 | European Patent Office (EPO) | A | |
| 131893828 | European Patent Office (EPO) | – | |
| 13189382 | European Patent Office (EPO) | A | |
| 2014065118 | European Patent Office (EPO) | W |
Members298
| Document | Office | Kind | |
|---|---|---|---|
| EP2830054A1 | European Patent Office (EPO) | A1 | |
| EP2830056A1 | European Patent Office (EPO) | A1 | |
| EP2830059A1 | European Patent Office (EPO) | A1 | |
| EP2830061A1 | European Patent Office (EPO) | A1 | |
| EP2830063A1 | European Patent Office (EPO) | A1 | |
| EP2830064A1 | European Patent Office (EPO) | A1 | |
| EP2830065A1 | European Patent Office (EPO) | A1 | |
| CA2886505A1 | Canada | A1 | |
| CA2918524A1 | Canada | A1 | |
| CA2918701A1 | Canada | A1 | |
| CA2918804A1 | Canada | A1 | |
| CA2918807A1 | Canada | A1 | |
| CA2918810A1 | Canada | A1 | |
| CA2918835A1 | Canada | A1 | |
| CA2973841A1 | Canada | A1 | |
| WO2015010947A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010948A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010949A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010950A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010952A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010953A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010954A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201513098A | Taiwan Province of China | A | |
| AU2014295302A1 | Australia | A1 | |
| TW201514974A | Taiwan Province of China | A | |
| TW201517019A | Taiwan Province of China | A | |
| TW201517023A | Taiwan Province of China | A | |
| TW201517024A | Taiwan Province of China | A | |
| SG11201502691QA | Singapore | A | |
| KR20150060752A | Republic of Korea | A | |
| TW201523589A | Taiwan Province of China | A | |
| TW201523590A | Taiwan Province of China | A | |
| EP2883227A1 | European Patent Office (EPO) | A1 | |
| MX2015004022A | Mexico | A | |
| CN104769671A | China | A | |
| US2015287417A1 | United States of America | A1 | |
| JP2015535620A | Japan | A | |
| AR096985A1 | Argentina | A1 | |
| AR096988A1 | Argentina | A1 | |
| AR096989A1 | Argentina | A1 | |
| AR096990A1 | Argentina | A1 | |
| AR096991A1 | Argentina | A1 | |
| AR096992A1 | Argentina | A1 | |
| AR096993A1 | Argentina | A1 | |
| SG11201600401RA | Singapore | A | |
| SG11201600422SA | Singapore | A | |
| SG11201600464WA | Singapore | A | |
| SG11201600494UA | Singapore | A | |
| SG11201600496XA | Singapore | A | |
| SG11201600506VA | Singapore | A | |
| KR20160024924AThis record | Republic of Korea | A | |
| AU2014295295A1 | Australia | A1 | |
| AU2014295296A1 | Australia | A1 | |
| AU2014295297A1 | Australia | A1 | |
| AU2014295298A1 | Australia | A1 | |
| AU2014295300A1 | Australia | A1 | |
| AU2014295301A1 | Australia | A1 | |
| KR20160030193A | Republic of Korea | A | |
| CN105453175A | China | A | |
| CN105453176A | China | A | |
| KR20160034975A | Republic of Korea | A | |
| KR20160041940A | Republic of Korea | A | |
| CN105518776A | China | A | |
| CN105518777A | China | A | |
| KR20160042890A | Republic of Korea | A | |
| MX2016000940A | Mexico | A | |
| KR20160046804A | Republic of Korea | A | |
| CN105556603A | China | A | |
| MX2016000857A | Mexico | A | |
| MX2016000924A | Mexico | A | |
| CN105580075A | China | A | |
| EP3017448A1 | European Patent Office (EPO) | A1 | |
| US2016133265A1 | United States of America | A1 | |
| US2016140973A1 | United States of America | A1 | |
| US2016140979A1 | United States of America | A1 | |
| US2016140980A1 | United States of America | A1 | |
| US2016140981A1 | United States of America | A1 | |
| HK1211378A | Hong Kong, China | A | |
| HK1211378A1 | Hong Kong, China | A1 | |
| EP3025328A1 | European Patent Office (EPO) | A1 | |
| EP3025337A1 | European Patent Office (EPO) | A1 | |
| EP3025340A1 | European Patent Office (EPO) | A1 | |
| EP3025343A1 | European Patent Office (EPO) | A1 | |
| EP3025344A1 | European Patent Office (EPO) | A1 | |
| MX2016000854A | Mexico | A | |
| AU2014295302B2 | Australia | B2 | |
| MX2016000935A | Mexico | A | |
| MX2016000943A | Mexico | A | |
| TWI541797B | Taiwan Province of China | B | |
| MX340575B | Mexico | B | |
| US2016210974A1 | United States of America | A1 | |
| TWI545558B | Taiwan Province of China | B | |
| TWI545560B | Taiwan Province of China | B | |
| TWI545561B | Taiwan Province of China | B | |
| EP2883227B1 | European Patent Office (EPO) | B1 | |
| JP2016525713A | Japan | A | |
| JP2016527556A | Japan | A | |
| JP2016527557A | Japan | A | |
| TWI549121B | Taiwan Province of China | B | |
| JP2016529545A | Japan | A |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Ip right lapsedLapsedST27 STATUS EVENT CODE: N-4-6-H10-H13-OTH-PC1903 (AS PROVIDED BY THE NATIONAL OFFICE); TERMINATION CATEGORY : DEFAULT_OF_REGISTRATION_FEEH13 | H13 | |
| Written decision to grantGRNT | GRNT | |
| Decision to grant (after re-examination)X701 | X701 | |
| AmendmentAMND | AMND | |
| Decision to refuse applicationE601 | E601 | |
| AmendmentAMND | AMND | |
| Notification of reason for refusalE902 | E902 | |
| AmendmentAMND | AMND | |
| Request for examinationA201 | A201 | |
| AmendmentAMND | AMND |
Numbers
- Publication
- 10-2016-0024924
- Application
- 1020167001383
Titles4
- Korean
- 인코딩된 오디오 신호를 디코딩하기 위한 장치, 방법 및 컴퓨터 프로그램
- English
- APPARATUS, METHOD AND COMPUTER PROGRAM FOR DECODING AN ENCODED AUDIO SIGNAL
- Unlabeled
- 인코딩된 오디오 신호를 디코딩하기 위한 장치, 방법 및 컴퓨터 프로그램{APPARATUS, METHOD AND COMPUTER PROGRAM FOR DECODING AN ENCODED AUDIO SIGNAL}
- Unlabeled
- Apparatus, method and computer program for decoding an encoded audio signal
Classification
- CPC, 19
- G10L21/0388
- G10L19/02
- G10L19/03
- G10L19/008
- G10L19/025
- G10L19/028
- G10L21/038
- G10L19/0204
- G10L19/0212
- G10L19/022
- G10L19/032
- G10L19/06
- G10L19/18
- H03M7/30
- G10L19/0208
- H04S1/007
- G10L25/18
- G10L25/21
- G10L25/06
- IPC, 4
- G10L19 02
- G10L19 025
- G10L19 03
- G10L21 0388