Arbitrary shaping of temporal noise envelope without side-information
Abstract
The first feature is that the temporal envelope of noise can be freely shaped in the spectral region without side information. In encoding, a filtered quantification error index was applied to the quantized frequency domain representation of the discrete time domain signal as a feedback signal prior to quantization and inversely converted from frequency domain to time domain in decoding. When the filtering coefficient to be filtered affects the shaping of the quantization is in the time domain of the quantized frequency domain representation of the discrete time domain signal. This is accomplished for each of one or more frequency bins or groups of frequency bins. Another feature is frequency domain noise feedback quantization in digital audio.

Term
0.9 yearsto projected expiry
Projected expiry 10 August 2027, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
25 claims: 6 independent, 19 dependent
- 1離散時間領域信号の周波数領域表現の量子化を採用する、離散時間領域信号をエンコーディングするためのディジタルオーディオエンコーディング方法であって、 量子化誤差の指標を抽出するステップと、 フィルターされた量子化誤差の指標を生成するために前記量子化誤差の指標にフィルターを掛けるステップと、 前記フィルターされた量子化誤差の指標を、量子化の前に前記離散時間領域信号の周波数領域表現にフィードバック信号として適用するステップと、 を具備し、 前記フルターを掛けるステップでのフィルター係数により、周波数領域から時間領域に逆変換したときに、離散時間領域信号の量子化された周波数領域表現の時間領域における量子化ノイズを整形する効果がもたらされることを特徴とする方法。
- 2前記フィルターを掛けるステップでは、離散時間領域信号の周波数領域表現のスペクトルの全セグメントに亘ってフィルターされた量子化誤差の指標が変化することができるように、1以上の周波数ビン又は周波数ビンのグループの各々にフィルターされた量子化誤差の指標を生成するために前記量子化誤差の指標にフィルターを掛けることを特徴とする請求項1に記載の方法。
- 3前記フィルター係数は動的に制御可能であることを特徴とする請求項1又は請求項2に記載の方法。
- 4前記フィルター係数は、前記離散時間領域信号の指標に応答して動的に制御可能であることを特徴とする請求項3に記載の方法。
- 5前記離散時間領域信号の前記指標は、時間的信号のエンベロープを計算し、逆変換し、その結果の逆DFTを計算することを含む処理により取得することを特徴とする請求項4に記載の方法。
- 6前記フィルター係数は、前記離散時間領域信号の周波数領域表現の指標に応答して動的に制御可能であることを特徴とする請求項3に記載の方法。
- 7前記離散時間領域信号の前記周波数領域表現の指標は、線形予測コーディング(LPC)計算を含む処理により取得することを特徴とする請求項6に記載の方法。
- 8前記フィルター係数は、時間的マスキングモデルに応答することを特徴とする請求項1乃至請求項7のいずれか1項に記載の方法。
- 9前記時間的マスキングモデルは、量子化した時間的整形をおこなおうとすることを特徴とする請求項8に記載の方法。
- 10前記時間的マスキングモデルは、変換ブロック内で前記離散時間領域信号の相対的に音量の小さいセグメントから音量の大きいセグメントに前記量子化ノイズを移動させようとすることを特徴とする請求項8に記載の方法。
- 11エンコードされたビットストリームを生成するために、前記離散時間領域信号の前記量子化した周波数領域表現をエンコードするステップをさらに具備することを特徴とする請求項1乃至請求項10のいずれか1項に記載のオーディオエンコーディング方法。
- 12請求項11に記載のエンコーディング方法により生成されたビットストリームをデコードするようにしたディジタルオーディオデコーダ。
- 13離散時間領域信号の周波数領域表現を量子化し、量子化誤差を抽出し、フィルターされた量子化誤差の指標を生成するために量子化誤差の指標にフィルターを掛け、前記離散時間領域信号の量子化する前の前記周波数領域表現に、フィードバック信号としてフィルターされた量子化誤差の指標を適用し、離散時間領域信号の量子化された周波数領域表現をビットストリームにエンコードする、エンコーダによって生成されたエントロピーエンコードされたビットストリームをデコーディングするためのディジタルオーディオデコーディング方法であって、 離散時間領域信号の量子化された前記周波数領域表現又はその近似を生成するために、前記ビットストリームをデコードするステップと、 前記量子化された前記周波数領域表現又はその近似を逆量子化するステップと、 オーディオ信号を生成するために前記周波数領域表現又はその近似を時間領域に逆変換するステップと、 を具備し、前記エンコーダ中のフィルター係数は、前記オーディオ信号の量子化ノイズの整形に影響を与えることを特徴とする方法。
- 14ディジタルオーディオエンコーダにおける、周波数領域ノイズフィードバック量子化の方法であって、 量子化装置の入力信号を生成するために時間領域オーディオ信号から抽出した周波数領域信号をノイズフィードバック信号と結合させるステップと、 量子化装置の出力信号を生成するために、前記量子化装置の入力信号を量子化するステップと、 量子化誤差信号を生成するために、前記量子化装置の前記入力信号を前記量子化装置の前記出力信号と結合させるステップと、 ノイズフィードバック信号を生成するために、前記量子化誤差信号にフィルターを掛けるステップと、 を具備することを特徴とする方法。
- 15前記ノイズフィードバックのフィルター係数を動的に制御するステップを具備することを特徴とする請求項14に記載の方法。
- 16前記動的に制御するステップでは、前記周波数領域信号を抽出する時間領域オーディオ信号の指標に応答して前記ノイズフィードバックのフィルター係数を制御することを特徴とする請求項15に記載の方法。
- 17前記動的に制御するステップでは、時間的マスキングモデルに応答して前記ノイズフィードバックのフィルター係数を制御することを特徴とする請求項16に記載の方法。
- 18請求項1乃至請求項17のいずれか1項に記載の方法を実施するようにした装置。
- 19請求項18に記載の装置をコンピュータに制御させるための、コンピュータ読み取り可能な媒体に記憶させたコンピュータプログラム。
- 20請求項1乃至請求項17に記載の方法をコンピュータに実行させるための、コンピュータ読み取り可能な媒体に記憶させたコンピュータプログラム。
- 21ディジタルオーディオエンコーダで用いるための周波数領域ノイズフィードバック量子化装置であって、 量子化装置の入力信号を生成するために、時間領域オーディオ信号から抽出した周波数領域信号をノイズフィードバック信号と結合させる第1の合成器と、 量子化装置の出力信号を生成するために、前記量子化装置の入力信号を量子化する量子化装置と、 量子化誤差信号を生成するために、前記量子化装置の入力信号と前記量子化装置の出力信号とを結合する第2の合成器と、 ノイズフィードバック信号を生成するために、前記量子化誤差信号にフィルターを掛けるノイズフィードバックフィルターと、 を具備することを特徴とする装置。
- 22ノイズフィードバックフィルター係数を動的に制御する、フィルター係数制御装置をさらに具備することを特徴とする請求項21に記載の量子化装置。
- 23前記フィルター係数制御装置は、前記周波数領域信号を抽出する時間領域の1以上の指標に応答して前記ノイズフィードバックフィルター係数を制御することを特徴とする請求項22に記載の量子化装置。
- 24前記フィルター係数制御装置は、時間的マスキングモデルに応答して前記ノイズフィードバックフィルター係数を制御することを特徴とする請求項23に記載の量子化装置。
- 25前記ノイズフィードバックフィルターの次数は10から20の範囲であることを特徴とする請求項21乃至請求項24のいずれか1項に記載の量子化装置。
Independent claims25
60 paragraphs, as filed
The present invention relates to digital audio coding. In particular, the features of the present invention are a digital audio encoding method, a digital audio decoder suitable for decoding a bit stream generated by the digital audio encoding method, a digital audio decoding method, and a frequency region noise feedback quantization method in a digital audio encoder. The present invention relates to a computer program stored in a computer-readable medium for controlling such a device or method by a computer, and a frequency region noise feedback quantization device in a digital audio encoder.
The time-frequency trade-offs of spectral domain audio coding systems have led to several techniques that provide high audio coding performance while minimizing audible coding errors. Such techniques include block switching and temporal noise shaping (TNS) (see reference 1 below), both of which are MPEG-2 / 4AAC (AAC) (see reference 2 below). It is adopted. Temporal noise shaping (TNS) provides a way to ensure that the temporal envelope of noise is adjusted to minimize audible artifacts, while using a relatively long conversion block length.
Figure 1 shows a simplified block diagram of the prior art spectral domain coding system (encoder and decoder) using TNS. In the encoder part, the "time / frequency conversion" device or function 2 has a sampling frequency f.<sub>s</sub>Converts the audio signal in the time domain represented by the discrete-time sequence x [n] sampled from the sound source in the spectrum domain (or "frequency" domain). For AAC, a 2048 sample modified discrete cosine transform (MDCT) (see reference 3 below) is used. Prior to quantization by the quantization device or quantization function (Q) 4, the encoder spectrum the filter or filter function 6 (A (z)), which can represent the transfer function as A (z) in the Z region. Applies to region signals. This encoder sends the filter coefficient to the decoder as side information. The decoder part of this coder decodes the bitstream and applies a filter or filter function 8 (1 / A (z)) to the spectrum, which the transfer function can represent as 1 / A (z) in the Z region. To do. A "frequency-time conversion" device or function (which performs the reverse conversion of time-frequency conversion 2) converts a signal in the spectral domain into a signal y (n) in the discrete-time domain. For simplicity, Figure 1 ignores perceptible quantization noise and other known AAC and TNS details.
The output of the entire spectral region of the quantization device 4 using TNS can be expressed by the Z-transform region as shown in Eq. (1). This analysis and the other analyzes described below are based on a simple additive model of quantization.<maths num="1"><img file="JP2010500631A_D0001.tif" /></maths>
Here, E (z) is the quantization error, and A (z) is the transfer function of the TNS filter.
Equation (1) becomes like Equation (2) when simplified,<maths num="2"><img file="JP2010500631A_D0002.tif" /></maths>
Equation (2) shows that the convolution process (multiplication process by 1 / A (z)) in the Z region is applied to the noise added during the quantization process of the audio spectrum. Since the convolution in the spectral domain is equivalent to the multiplication in the time domain, convolution of the noise by 1 / A (z) shows that the temporal shape of the noise was multiplied by the temporal response of the inverse TNS filter. .. Therefore, with proper selection of filter A (z), the quantization noise can be adjusted by minimizing the audible artifacts generated by the reduced time resolution. TNS has been shown to significantly improve AAC performance, making TNS an important tool in AAC.
However, TNS has some limitations. That is, the encoder must transmit the filter coefficients to the decoder, and the decoder must convolve the decoded spectrum with an inverse filter. These requirements lead to the following restrictions:
1. Increased bitrate consumption to transmit filter coefficients 2. Inverse filters need to be applied to spectral means such as AC-3 where TNS cannot have backward compatibility with existing systems. 3. Inverse filter must be applied to the spectrum, increasing the complexity of the decoder According to the features of the present invention, a new technology based on noise feedback quantization (NFQ) imposes the time envelope of quantization noise in a spectral domain coding system by the TNS coding tool used for MPEG-2 / 4 · AAC. Correct while overcoming the restrictions. According to the features of the present invention, NFQ is adopted instead of TNS in the AAC system. According to the features of the present invention, NFQ can also be employed in coding systems in other spectral regions, such as AC-3 systems.
According to the features of the present invention, a digital audio encoding method for encoding a discrete-time domain signal is provided, which employs the quantization of the frequency domain representation of a discrete-time domain signal. In this method, an index of quantization error (a numerical value of the degree) is extracted, the index of this quantization error is filtered to generate a filtered index of quantization error, and this filter is performed. The quantization error index is applied as a feedback signal to the frequency region representation of the discrete time region signal prior to quantization. Here, the filter coefficient for multiplying the fluter has the effect of shaping the quantized noise in the time domain of the quantized frequency domain representation of the discrete-time domain signal when inversely converted from the frequency domain to the time domain. ..
This digital audio encoding method involves one or more frequency bins or frequency bins so that the filtered quantization error index can vary over the entire segment of the frequency domain representation of the discrete-time domain signal. Filter the quantization error index to generate a filtered quantization error index for each of the groups. The filter coefficient can be dynamically controllable. Such controllability can be made to respond to an index of a discrete-time domain signal or a frequency domain representation of a discrete-time domain signal. The filter coefficients can also be made to respond to a temporal masking model (not shown). The quantized frequency domain representation of a discrete-time domain signal subsequently produces an encoded and encoded bitstream.
According to another feature of the present invention, there is provided a digital audio decoder applied to decode a bitstream generated by the encoding method described above.
According to yet another feature of the present invention, the frequency region representation of a discrete time region signal is quantized, the quantization error is extracted, and the quantization error index is filtered to generate a filtered quantization error index. Multiplies and applies the filtered quantization error index as a feedback signal to the pre-quantized frequency region representation of the discrete time region signal, and encodes the quantized frequency region representation of the discrete time region signal into a bitstream. Provides a digital audio decoding method for decoding the entropy-encoded bitstream generated by the encoder. This decoding method decodes the bitstream to generate a quantized frequency domain representation of a discrete time domain signal or an approximation thereof, and dequantizes the quantized frequency domain representation or its approximation to a frequency domain representation or The approximation is inversely converted into the time domain to generate the audio signal, which has the effect that the filter coefficients in the encoder shape the quantization noise of the audio signal. Yet another feature of the present invention provides a method of frequency domain noise feedback quantization in a digital audio encoder. In this method, the frequency region signal extracted from the time region audio signal is combined with the noise feedback signal to generate the input signal of the quantization device, and the input signal of the quantization device is quantized to generate the output signal of the quantization device. Then, the input signal of the quantization device and the output signal of the quantization device are combined to generate a quantization error signal, and the quantization error signal is filtered to generate a noise feedback signal.
According to the features of the present invention, there is provided an apparatus suitable for performing a method of frequency domain noise feedback quantization in a digital audio encoder.
According to another feature of the present invention, there is provided a computer program stored in a computer-readable medium for allowing a computer to control the above-mentioned device or method.
Yet another feature of the present invention provides a frequency domain noise feedback quantizer for use in digital audio encoders. The frequency region noise feedback quantization device includes a first synthesizer and a quantization device that combine a frequency region signal extracted from a time region audio signal with a noise feedback signal to generate an input signal of the quantization device. The second, which generates a quantization error signal by combining the input signal of the quantization device and the output signal of the quantization device with the quantization device that generates the output signal of the quantization device by quantizing the input signal of the It includes a synthesizer and a noise feedback filter that filters the quantization error signal to generate a noise feedback signal.
Temporal shaping of quantization noise in a spectral audio coding system is important for efficient audio compression. The TNS coding tool in MPEG-2 / 4 / AAC performs temporal shaping of quantization noise, but it is limited due to the need to transmit the filter coefficient to the decoder. According to the features of the present invention, the spectrum quantization process or the spectrum quantization device is provided with a feedback circuit or a feedback process that enables temporal shaping of quantization noise in order to control it so that it can be freely shaped in a wide range. include. In addition, encoder / decoder coding systems no longer need to transmit filter coefficients to the decoder. The present invention has one or more advantages over the MPEG-2 / 4 / AAC TNS coding tool as shown below, and can be used in place of TNS.
1. Noise shaping performance is comparable to TNS 2. Encoder-only processing 3. Does not require transmission of side information 4. Backward compatible with existing audio coding systems 5. Reduce decoder complexity Another advantage of the present invention is that the feedback filter can be varied over the entire spectrum so that the temporal evolution of noise is in good agreement with the signal characteristics of the spectral group. In other words, a unique feedback filter can be adopted for each of one or more frequency bins or groups of frequency bins, where the frequency bins form a spectral group. TNS also has such capabilities, but the number of spectral regions that can be used is very limited due to the need to indicate the spectral groups required by the decoder and also to transmit the filter coefficients to the decoder. Will be done.
<figref num="1">FIG. 6 is a simplified schematic block diagram of a prior art spectral region coding system (encoder and decoder) using temporal noise shaping (TNS).</figref><figref num="2">It is a schematic block diagram showing a simplification of the modern audio coding system of the prior art, in which the input is transformed into a spectral region and the spectrally represented signal is quantized.</figref><figref num="3">FIG. 5 is a schematic functional block diagram of an embodiment of a simplified audio coding system that employs noise feedback quantization (NFQ) according to the features of the present invention.</figref><figref num="4">It is an example of the result of applying the embodiment of the present invention in which a noise feedback filter is designed for the audio content of a specific conversion block such that the segmented signal-to-noise ratio is substantially constant.</figref><figref num="5">A simple spectral region coder without NFQ applied to the input signal shown in FIG. 4 and a segmented signal-to-noise ratio with a typical NFQ system of order 10 applied.</figref><figref num="6">It is a simplified schematic block diagram of the MPEG-2 / 4 / AAC encoder of the prior art.</figref><figref num="7">It is a schematic functional block diagram of one embodiment of a simplified AAC audio coding system that employs noise feedback quantization according to the features of the present invention.</figref>
Spectral region noise feedback quantization (NFQ)] State-of-the-art audio coding technologies, including AAC (see Reference 2 below) and AC-3 (see Reference 3 below), are used to control quantization-generated noise in a perceptually appropriate manner. Quantization in the spectral region is performed. Generally, the input time waveform is converted into a spectral region by using a time / frequency conversion such as MDCT. The perceptual model is calculated in parallel with the time-frequency conversion, and the perceptual model is used to adjust the quantization noise generated in each of the output coefficients of the time-frequency conversion. FIG. 2 is a simplified schematic block diagram of a prior art audio coding system (encoder and decoder) that transforms the input into a spectral region and quantizes the spectral representation of the signal. The discrete time domain signal x (n) is applied to a time / frequency conversion or time / frequency conversion function (time / frequency conversion) 12 for generating a signal in the frequency domain (or spectrum domain). The signal in the spectrum region is quantized by the quantization device or the quantization function (Q) 14, and the signal in the frequency domain is quantized to generate Y (K). The decoder portion of the system includes an inverse or inverse conversion function (frequency / time conversion) 16 that provides an output signal in the time domain.
Generally, the conversions used in modern audio coding systems have a length of 512 samples or more to maintain good coding efficiency. For example, MPEG-2 / 4 / AAC adopts 2048 MDCT as a pseudo fixed signal. This conversion provides good coding efficiency, but the large length conversion results in diffusion of quantization noise (the quantization noise spreads throughout the conversion block) and audible signal degradation for unfixed signals. Technologies such as block switching, TNS, and gain control have been designed to address this issue.
According to the features of the present invention, noise feedback quantization (NFQ) is adopted in this spectral region in order to adjust the temporal envelope of the quantization noise generated when the quantization process is performed in the spectral region. FIG. 3 is a schematic functional block diagram of a simplified example of an audio coding system that employs NFQ according to the features of the present invention. The output (X (k)) of the time / frequency conversion process or the time / frequency converter (time / frequency conversion) 12 is a feedback process or feedback circuit that applies a filtered quantization error to the original converted signal. It is quantized by the NFQ quantization device provided with the above or the NFQ quantization function 18. The output (Y (k)) of such a device or the quantization device or quantization function (Q) 20 during processing can be expressed by Eq. (3).<maths num="3"><img file="JP2010500631A_D0003.tif" /></maths>
Here, E (k) is the quantization error, F (m) is the coefficient in the feedback filter, X (k) is the frequency domain output of the time / frequency conversion 12, k is the spectral bin index, and m is the filter tap index. is there.
Alternatively, equation (3) can be rewritten as equation (4) using the Z-transform format.<maths num="4"><img file="JP2010500631A_D0004.tif" /></maths>
Since the convolution in the spectral domain is equivalent to the multiplication in the time domain, E (z) is expressed as (1-z) as shown in Eq. (4).<sup>-1</sup>By convolving with F (z)), the time error signal is combined with the corresponding NFQ feedback and quantization configuration (1-z).<sup>-1</sup>The effect of multiplying by the temporal effect of F (z)) is obtained. This suggests that the temporal envelope of the quantization error can be freely changed by properly selecting F (z). Therefore, although two methods of generating a valid filter transfer function are described below, the present invention is considered to be useful to coding system designers in modifying the temporal envelope of quantization error. It should be understood that it is intended for other methods of extracting functions.
Reference 5 shows that F (z) must take the form shown in Eq. (5). Further, Reference 5 shows a technique for optimally solving F (z) by adding the limitation obtained by Eq. (5) as shown by Eq. (6).<maths num="5"><img file="JP2010500631A_D0005.tif" /></maths>
a<sub>0</sub>= 1<maths num="6"><img file="JP2010500631A_D0006.tif" /></maths>
here<maths num="7"><img file="JP2010500631A_D0007.tif" /></maths>
<img file="JP2010500631A_D0008.tif" />
One method is to calculate the envelope of the temporal signal, inverse transform it, and calculate the inverse DFT (discrete Fourier transform) of the calculation result as shown in Eq. (7). This method is a partial signal-to-noise ratio (signal-to-noise ratio calculated for a small number of samples) that is approximately equal for all samples in the transformation block due to the noise characteristics obtained as a result of quantization of the spectral region with feedback. Make sure that is guided. This is shown as follows.<maths num="8"><img file="JP2010500631A_D0009.tif" /></maths>
Here, E [n] is the envelope of the temporal signal, which has only positive values, and N is the length of the conversion block.
Alternatively, another way to find the preferred solution for F (z) is to employ the inverse operation of the TNS filter. For example, the filter coefficient can be generated from the impulse response of the LPC (Linear Predictive Coding) coefficient extracted from the autocorrelation of the input audio spectrum X (k).
The application and calculation of the noise feedback filter does not have to be static. It is preferable that F (z) changes with time and is revised periodically for each conversion block, for example, in order to bring about appropriate temporal shaping of quantization noise. Further, as described above, only one F (z) may be adopted for one or more frequency bins or a group of frequency bins. Thus, F (z) can have different coefficients for each bin (frequency) and conversion block (time).
Returning to the description of FIG. 3, in the NFQ apparatus or NFQ processing 18, in the synthesizer 22, the input of the quantization apparatus or the quantization function 20 is subtracted from the output to generate the quantization error signal E (k). On the other hand, the error signal is filtered by the filter or the filter function 24, and is subtracted from the frequency domain output X (k) of the time / frequency conversion 12 in the synthesizer 26. The output Y (k) of the NFQ apparatus or NFQ process 18 is applied to the frequency / time converter or process 28 that performs the reverse conversion of the time / frequency conversion 12. The filter or filter function 24 is dynamic, and the filter is (1) the time domain input signal x (n), or (2) the frequency of the input signal measured by the dynamic noise feedback filter calculator or the dynamic noise feedback filter function 30. It is preferably controlled by either index of domain version X (k).
The concrete reshaping of temporal noise allows the designer of a particular coding configuration to choose a wide range of free forms, but moves the noise from the quiet segment of the audio to the loud segment of the audio. It is preferable to let it.
[Spectral NFQ performance] FIG. 4 shows an example of the performance when the embodiment of the present invention is applied. Here, the noise feedback filter or noise feedback filter function is designed to be applied to a specific conversion block of audio content such that the resulting partial signal-to-noise ratio is nearly constant. The partial signal-to-noise ratio is defined in this example as the signal-to-noise ratio calculated for a small number of samples, less than the number of samples in the conversion block. Furthermore, in this example, the order of the noise feedback filter is set to 10. The upper part of FIG. 4 shows the input waveform in the time domain in the conversion block having a sharp transient signal. The middle row shows the output signal of the coder in the simple spectral region, where the quantization noise is spread over the entire conversion block before the start of the transient signal. The lower part shows the output of the spectral region audio coder to which the embodiment of the present invention adopting the NFQ of order 10 is applied. For the NFQ processing or NFQ system in this example, the feedback filter is calculated so that the fragmentary SNR is kept nearly constant throughout the conversion block. To demonstrate the ability of the present invention to modify the temporal envelope of quantization noise, the output of an NFQ process or NFQ system is significantly pre-echoed (preceding a transient signal, in a conversion block) compared to a configuration without NFQ. The part where the noise spreads) is reduced.
Figure 5 shows the fragmented SNR of a simple spectral region coder without NFQ and the fragmented SNR of a coder with a degree 10 NFQ configuration, with the encoder in Figure 4 Receives the input signal indicated by. Fragmented SNRs before the onset of transient signals are very small in processes or systems without NFQ, whereas fragmented SNRs remain nearly constant throughout the conversion block in processes or systems with NFQ. Has been done.
Keeping the fragmentary SNR constant shows the advantages of the embodiments of the present invention that employ NFQ, but may not reflect the example of adopting optimal noise distribution. In practice, the processor or device designer can choose the temporal noise distribution that he or she thinks is appropriate depending on the temporal characteristics of the signal. As shown below, a feature of the present invention is that a complex temporal masking model for extracting a preferred temporal noise envelope of quantization noise can be used, if desired.
[Application of Noise Feedback Quantization in Perceptual Audio Coding Systems] Figure 6 shows a simplified schematic block diagram of the prior art MPEG-2 / 4 AAC encoder. Input pulse code modulation (PCM) audio is transformed into a spectral region using 2048 MDCT32 points, and an estimate of the masking curve for that block is calculated using the psychoacoustic model 34. The scale factor is then selected to keep the noise-to-mask ratio (NMR) as low as possible (due to spectral quantization) (36). The resulting signal is quantized (38) and then entropy-coded (40). The formatter or formatting process (bitstream) 42 produces an encoded bitstream output. However, this technique ignores temporal masking within individual transformation blocks. According to the features of the present invention, a method of arranging the quantization noise in time is shown. The AAC encoder shown in FIG. 6 is added with the noise feedback quantizer 18 and the dynamic noise feedback calculation 30 as shown in FIG. 7, and the complementary TNS encoding decoding filter (that is, the filter 6 in FIG. 1) is added. And 8) are removed, and the need to transmit the TNS filter coefficient from the encoder to the decoder is removed, and as suggested earlier, the spectral domain quantization noise, the spectrum (by applying a scale factor to the spectral domain). It can be rearranged not only to suit the masking model, but also to the temporal masking model (by applying the NFQ). The dynamic noise feedback calculation 30 receives an input from either (1) the PCM time domain input or (2) the frequency domain output of the MDCT32. The temporal masking model is not shown in Figure 7.
According to the first technique, the time domain input signal must first be aliased in order to calculate the NFQ filter for use with the MDCT (eg for use in AAC encoders or encoder / decoder systems). A given time sequence for a given transformation block is x (n) n = 0,1, ... N-1 You can extract the aliased time sequence,<maths num="9"><img file="JP2010500631A_D0010.tif" /></maths>
as well as<maths num="10"><img file="JP2010500631A_D0011.tif" /></maths>
Subsequently, the preferred temporal envelope (E'[n]) of the quantization noise can be extracted. Ideally, the temporal masking effect should be manifested by the preferred envelope calculation, but the characteristics of the MDCT aliasing prevent the application of temporal masking to this problem. Therefore, the inverse operation of the temporal energy envelope can be used, that is, a lot of noise may be placed in the loud signal region.<maths num="11"><img file="JP2010500631A_D0012.tif" /></maths>
as well as<maths num="12"><img file="JP2010500631A_D0013.tif" /></maths>
The noise feedback filter is obtained by the above equations (6) and (7).
Alternatively, the NFQ filter can be generated in the same way that TNS (see Reference 1 below) generates the encoding filter and inversely transforms the result. First, calculate the MDCT of the current conversion block and<maths num="13"><img file="JP2010500631A_D0014.tif" /></maths>
Then the MDCT autocorrelation was calculated and<maths num="14"><img file="JP2010500631A_D0015.tif" /></maths>
The Levinson Durbin algorithm (see reference 1 below) can then be used to calculate linear prediction coefficients, such as those used for TNS (A (z)). The noise feedback filter is then calculated as follows.<maths num="15"><img file="JP2010500631A_D0016.tif" /></maths>
Here, F is the noise feedback filter transfer function, M is the order of the noise feedback filter, and L is the order of the prediction coefficient.
The NFQ filter obtained by equations (11) to (13) provides almost the same temporal noise shaping as temporal noise shaping as in equations (8a) to (10). That is, noise moves from the quiet segment of the audio to the louder segment of the audio.
[Planned application of noise feedback quantization] The application of noise feedback quantization according to the present invention includes at least one of the following.
-Application to existing audio coding systems: The features of the present invention can only be applied to encoders. Therefore, there is no need for side information and it can be applied to existing technologies such as MPEG-2 / 4 / AAC and AC-3.
Application to AAC for less complex decoding: Since noise feedback quantization is performed only in the encoder, replacing TNS with encoder-only noise feedback quantization can simplify the AAC decoder.
Applicable to AAC that can be scaled without loss: A new MPEG that can be scaled to a lossless audio coder (SLS) is MPEG- as a base layer so that the lossless representation of the original audio can be reconstructed. 4. Transmit the difference between the coded signal and the original signal using AAC. This does not exactly reverse the bits in the decoder, so it limits AAC to something like an integer MDCT that can exactly reverse convert, making tools like TNS unusable. However, since the NFQ according to the present invention can be applied only to the encoder, it can be used for the SLS profile of MPEG-4.
[References and incorporation as references] The following patents, patent applications, and publications are all incorporated herein by reference.
[1] "Enhancing the performance of perceptual audio coders by using temporal noise shaping (TNS)" by J. Herre and J. Johnston, published in 101st Convention Audio Engineering. Society, 1996 Proceedings 4384. [2] Details of MPEG-2 / 4 / AAC can be found in the following references. 1) ISO / IEC IS-14496 (Part 3, Audio), 1996, AAC ISO / IEC JTC1 / SC29, "Information technology-very low bitrate audio-visual coding", 2) ISO / IEC 13818-7, International Standard, 1997 "MPEG-2 advanced audio coding, AAC", 3) M. Bosi, K. Brandenburg, S. Quackenbush, L. 1996, Proc. Of the 101st AES-Convention, "ISO / IEC MPEG-2 Advanced Audio" by Fielder, K. Akagiri, H. Fuchs, M. Dietz, J. Herre, G. Davidson, and Y. Oikawa. Coding ", 4) Journal of the AES, by M. Bosi, K. Brandenburg, S. Quackenbush, L. Fielder, K. Akagiri, H. Fuchs, M. Dietz, J. Herre, G. Davidson, and Y. Oikawa. Vol.45, No.10, October 1997, pp. 789-814, "ISO / IEC MPEG-2 Advanced Audio Coding", 5) Proc. By Karlheinz Brandenburg. of the AES 17th International Conference on High Quality Audio Coding, Florence, Italy, 1999, "MP3 and AAC explained", and 6) J. Audio Eng. Soc, Vol.46, No.3, pp 164-177 March 1998, "Subjective Evaluation of State-of-the-Art Two-Channel Audio Codecs" by GA Soulodre et al. [3] IEEE Trans. Accoust. Speech Signal Processing, vol. ASSP-34 pp. 1153-1161, Oct. By J. Princen and A. Bradley. 1986, "Analysis / synthesis filter bank design based on time domain aliasing cancellation" [4] AC-3, also known as Dolby Digital (Dolby and Dolby Digital are registered trademarks of Dolby Laboratories Licensing Corporation), is defined in the "A / 52B document". (Digital Audio Compression Standard (AC-3, E-AC-3) Revision B and its predecessors, "A52 / A" document (ATSC standard: Digital Audio Compression Standard (AC-3), Revision A) and "A52 / A" (Digital Audio Compression Standard (AC-3)).
further, 1) Steve Vernon's August 1995 EEE Trans.Consumer Electronics, Vol.41, No. 3, "Design and Implementation of AC-3 Coders" 2) Mark Davis, Audio Engineering Society Preprint 3774, 95th AES Convention, October 1993, "The AC-3 Multichannel Coder" 3) Bosi et al., October 1992 Audio Engineering Society Preprint 3365, 93rd AES Convention, "High Quality," Low-Rate Audio Transform Coding for Transmission and Multimedia Applications " [5] "Least Squares Theory and Design of Optimal Noise Shaping Filters" AES 22nd International Conference on Virtual, Synthetic and Entertainment Audio by Werner Verhelst and Dreten De Koning in June 2002.
[Embodiment] The present invention can be implemented with hardware, software, or a combination of both (eg, a programmable logic array). Unless otherwise noted, the algorithms included as part of the present invention are not inherently relevant to a particular calculator or any other device. Specifically, various general purpose machines may be used with programs written according to the contents described herein, or more specialized devices (eg, integrated circuits) to carry out the required method. It may be convenient to configure. Thus, the present invention each has at least one processor, at least one storage system (including volatile and non-volatile memory and / or storage elements), at least one input device or input port, and at least one output. It can be implemented by one or more computer programs running on one or more programmable computer systems that include a device or output port. A program code is applied to the input data to perform the functions described here and output the output information. This output information is applied to one or more output devices in a known manner.
Each such program is any computer language required for communication with a computer system, including machine language, assembly, or higher-level, procedural, logical, or object-oriented languages. It can be realized even with. In any case, the language may be a compiled language or an interpreted language.
Each of such computer programs is provided by a general purpose programmable computer or a dedicated programmable computer for setting up and operating the computer when the storage medium or storage device is read by the computer in order to carry out the procedure described herein. It is preferred to store or download to a readable storage medium or device (eg, semiconductor memory or semiconductor medium, or magnetic or optical medium). The system of the present invention can also be considered running as a computer-readable storage medium composed of computer programs. Here, the storage medium causes the computer system to operate in a specifically predetermined manner in order to perform the functions described herein.
Many embodiments of the present invention have been described. However, it will be clear that many modifications can be made without departing from the spirit and technical scope of the invention. For example, some of the steps described here are independent and can therefore be performed in a different order than described.
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US11315580B2 | Cited by | United States of America | Applicant |
| US11315583B2 | Cited by | United States of America | Applicant |
| US11545167B2 | Cited by | United States of America | Applicant |
| US11127408B2 | Cited by | United States of America | Applicant |
| US12033646B2 | Cited by | United States of America | Applicant |
| US11217261B2 | Cited by | United States of America | Applicant |
| US11562754B2 | Cited by | United States of America | Applicant |
| US11380341B2 | Cited by | United States of America | Applicant |
| JP2021502597A | Cited by | Japan | Search report |
| US11386909B2 | Cited by | United States of America | Applicant |
| US11862182B2 | Cited by | United States of America | Applicant |
| US11380339B2 | Cited by | United States of America | Applicant |
| US10984809B2 | Cited by | United States of America | Applicant |
| JP2020525853A | Cited by | Japan | Search report |
| US11462226B2 | Cited by | United States of America | Applicant |
| JP2001237708A | Cites | Japan | Search report |
| JP2002542648A | Cites | Japan | Search report |
| WO2005004113A1 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| JP2005516442A | Cites | Japan | Search report |
| JP2006047561A | Cites | Japan | Search report |
| WO2006107833A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JP2008535024A | Cites | Japan | Search report |
| US5487086A | Cites | United States of America | Search report |
| JPH03201715A | Cites | Japan | Search report |
| JPH03201716A | Cites | Japan | Examiner |
16 members in 8 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 60838094 | United States of America | – | |
| 83809406 | United States of America | P | |
| 83809406 | United States of America | P | |
| 2007017811 | United States of America | W | |
| 2007017811 | United States of America | W | |
| 2006838094 | – | – | – |
| 2007017811 | – | – | – |
| US20060838094P | – | – | – |
| WO2007US17811 | – | – | – |
Members16
| Document | Office | Kind | |
|---|---|---|---|
| WO2008021247A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW200818123A | Taiwan Province of China | A | |
| WO2008021247A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008021247A9 | World Intellectual Property Organization (WIPO) | A9 | |
| EP2054882A2 | European Patent Office (EPO) | A2 | |
| CN101501761A | China | A | |
| JP2010500631AThis record | Japan | A | |
| US2010094637A1 | United States of America | A1 | |
| EP2054882B1 | European Patent Office (EPO) | B1 | |
| AT496365T | Austria | T | |
| ATE496365T1 | Austria | T1 | |
| DE602007012116D1 | Germany | D1 | |
| CN101501761B | China | B | |
| JP5096468B2 | Japan | B2 | |
| US8706507B2 | United States of America | B2 | |
| TWI456567B | Taiwan Province of China | B |
25 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of appointment of power of attorneyJAPANESE INTERMEDIATE CODE: A7423RD03 | RD03 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 2010500631
- Publication, DOCDB
- 2010500631
- Publication, EPODOC
- JP2010500631
- Application
- 2009524635
- Application, DOCDB
- 2009524635
- Application, EPODOC
- JP20090524635
Titles2
- Japanese
- サイド情報なしの時間的ノイズエンベロープの自由な整形
- English
- Free shaping of the temporal noise envelope without side information
Classification
- CPC, 4
- G10L19/03
- G10L19/008
- G10L19/032
- H04B1/665
- IPC, 1
- G10L19 02
Designated states4
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo