Untitled record
17 claims: 3 independent, 14 dependent
- 1複素予測ステレオ符号化によりステレオ信号を提供するデコーダシステムであって、 ダウンミックス信 号と 残差信 号の 第1の周波数領域表示に基づいて、前記ステレオ信号を生成するように構成されたアップミックス段階であって、各第1の周波数領域表示は多次元空間の第1の副空間で表された対応する信号のスペクトルコンテンツを表す第1のスペクトル成分を有するアップミックス段階を有し、前記アップミックス段階は、 前記ダウンミックス信号の第1の周波数領域表示に基づき、前記ダウンミックス信号の第2の周波数領域表示を計算するモジュールであって、前記第2の周波数領域表示は、前記第1の副空間に含まれない多次元空間の一部を含む、前記多次元空間の第2の副空間で表された信号のスペクトルコンテンツを表す第2のスペクトル成分を有し、前記ダウンミックス信号の第1のスペクトル成分に有限インパルス応 答フ ィルタを適用することにより前記ダウンミックス信号の第2のスペクトル成分を決定するように構成されたモジュールと、 ビットストリーム信号にエンコードされた前記ダウンミックス信号の第1と第2の周波数領域表示と、前記残差信号の第1の周波数領域表示と、複素予測係 数と に基づいてサイド信 号を 計算する重み付け加算器と、 前記ダウンミックス信号の第1の周波数領域表示と前記サイド信号とに基づいて、前記ステレオ信号を計算する和・差段階とを有し、 前記ダウンミックス信号と残差信号が前記和・差段階に直接供給されるパススルーモードで動作可能であるアップミックス段階とを有する、デコーダシステム。
- 2前記 有限インパルス応答 フィルタのインパルス応答は、前記ダウンミックス信号の第1の周波数領域表示を決定ために適用されるウィンドウ関数に応じて決まる、請求項1に記載のデコーダシステム。
- 3前記ダウンミックス信号と残差信号が時間フレームにセグメント化され、 前記アップミックス段階は、さらに、各時間フレームにおいて、そのフレームに関連する2ビットのデータフィールドを受け取り、前記データフィールドの値に応じて、アクティブモードまたはパススルーモードで動作する、請求項1に記載のデコーダシステム。
- 4前記ダウンミックス信号と残差信号が時間フレームにセグメント化され、 前記アップミックス段階は、さらに、各時間フレームにおいて、MPEGビットストリームで、そのフレームに関連するms_mask_presentフィールドを受け取り、前記ms_mask_presentフィールドの値に応じて、アクティブモードまたはパススルーモードで動作する、請求項1に記載のデコーダシステム。
- 5ビットストリーム信号に基づいて、前記ダウンミックス信号と残差信号の前記第1の周波数領域表示を提供する、前記アップミックス段階の上流に配置された逆量子化段階を有する、請求項1ないし4いずれか一項に記載のデコーダシステム。
- 6前記第1のスペクトル成分は前記第1の副空間で表された実数値を有し、 前記第2のスペクトル成分は前記第2の副空間で表された虚数値を有し、 任意的に、前記第1のスペクトル成分は離散余弦変 換ま たは修正離散余弦変 換の うちの1つにより得られ、 任意的に、前記第2のスペクトル成分は離散正弦変 換ま たは修正離散正弦変 換の うちの1つにより得られる、請求項1ないし5いずれか一項に記載のデコーダシステム。
- 7前記アップミックス段階の上流に配置された少なくとも1つの時間的ノイズシェーピン グモ ジュールと、 前記アップミックス段階の下流に配置された少なくとも1つのさらなる 時間的ノイズシェーピング モジュールと、 (a)前記アップミックス段階の上流の前記 時間的ノイズシェーピング モジュール、または(b)前記アップミックス段階の下流にある前記さらなる 時間的ノイズシェーピング モジュールのうちいずれかを選択的にアクティブ化するセレクタ装置とを有する、請求項1ないし6いずれか一項に記載のデコーダシステム。
- 8前記ダウンミックス信号は、それぞれが複素予測係数の値に関連する連続した時間フレームにパーティションされ、 前記ダウンミックス信号の第2の周波数領域表示を計算するモジュールは、前記複素予測係数の虚部の絶対値が時間フレームの所定許容値より小さいことに応じて、その時間フレームに対する出力を生成しないように自モジュールを停止するように構成されている、請求項6に記載のデコーダシステム。
- 9ダウンミックス信号の時間フレームは、それぞれが前記複素予測係数の値を伴う周波数帯域にさらにパーティションされ、 前記ダウンミックス信号の第2の周波数領域表示を計算するモジュールは、前記複素予測係数の虚部の絶対値が時間フレームの所定許容値より小さいことに応じて、その周波数帯域に対する出力を生成しないように自モジュールを停止するように構成されている、請求項8に記載のデコーダシステム。
- 10前記ステレオ信号は時間領域で表され、 前記デコーダシステムは、さらに、 逆量子化段階と前記アップミックス段階との間に配置され、 (a)パススルー段階、または (b)和・差段階として機能し、 それにより直接符号化ステレオ入力信号と同時符号化ステレオ入力信号との間の切り替えを可能とするスイッチングアセンブリと、 前記ステレオ信号の時間領域表現を計算するように構成された逆変換段階と、 前記逆変換段階の上流に配置され、これを (a)前記アップミックス段階の下流の点であって、複素予測により得られるステレオ信号が前記逆変換段階に供給される点、または (b)前記スイッチングアセンブリの下流、かつ前記アップミックス段階の上流の点であって、直接ステレオ符号化により得られるステレオ信号が前記逆変換段階に供給される点のうちいずれかに選択的に接続するように構成されたセレクタ装置とを有する、請求項1ないし9いずれか一項に記載のデコーダシステム。
- 11複素予測ステレオ符号化によりステレオ信号を提供する復号方法であって、 ダウンミックス信号と残差信号の第1の周波数領域表示を受け取るステップであって、前記第1の周波数領域表示の各々は多次元空間の第1の副空間で表された対応する信号のスペクトルコンテンツを表す第1のスペクトル成分を有する、ステップと、 制御信号を受け取るステップと、 前記制御信号の値に応じて、 (a)アップミックス段階を用いて、前記ダウンミックス信号と残差信号をアップミックスし、前記ステレオ信号を求めるステップであって、 前記ダウンミックス信号の第1の周波数領域表示に基づき、前記ダウンミックス信号の第2の周波数領域表示を計算するサブステップであって、前記第2の周波数領域表示は、前記第1の副空間に含まれない多次元空間の一部を含む、前記多次元空間の第2の副空間で表された信号のスペクトルコンテンツを表す第2のスペクトル成分を有し、前記ダウンミックス信号の第1のスペクトル成分に有限インパルス応 答フ ィルタを適用することにより前記ダウンミックス信号の第2のスペクトル成分を決定することを含むサブステップと、 ビットストリーム信号にエンコードされた前記ダウンミックス信号の第1と第2の周波数領域表示と、前記残差信号の第1の周波数領域表示と、複素予測係数とに基づいてサイド信号を計算するサブステップと、 前 記ダウンミックス信号 の第1の周波数領域表示 とサイド信号 とに和・差変換を 適用することにより、前記ステレオ信号を計算するサブステップとにより実行される、ステップ、または (b)アップミックスを停止 して、前記ダウンミックス信号と残差信号とを直接前記和・差段階に供給して前記ステレオ信号を求める ステップを有する復号方法。
- 12前記第1のスペクトル成分は前記第1の副空間で表された実数値を有し、 前記第2のスペクトル成分は前記第2の副空間で表された虚数値を有し、 任意的に、前記第1のスペクトル成分は離散余弦変 換ま たは修正離散余弦変 換の うちの1つにより得られ、 任意的に、前記第2のスペクトル成分は離散 正弦 変 換ま たは修正離散正弦変 換の うちの1つにより得られる、請求項11に記載の復号方法。
- 13前記ダウンミックス信号は、それぞれが複素予測係数の値に関連する連続した時間フレームにパーティションされ、 前記ダウンミックス信号の第2の周波数領域表示を計算するステップは、前記複素予測係数の虚部の絶対値が時間フレームの所定許容値より小さいことに応じて、その時間フレームに対する出力を生成しないように停止される、請求項12に記載の復号方法。
- 14ダウンミックス信号の時間フレームは、それぞれが前記複素予測係数の値を伴う周波数帯域にさらにパーティションされ、 前記ダウンミックス信号の第2の周波数領域表示を計算するステップは、前記複素予測係数の虚部の絶対値が時間フレームの所定許容値より小さいことに応じて、その周波数帯域に対する出力を生成しないように停止される、請求項13に記載の復号方法。
- 15前記ステレオ信号は時間領域で表され、 前記復号方法は、さらに、 前記ビットストリーム信号が直接ステレオ符号化または同時ステレオ符号化により符号化されていることに応じて、前記アップミックスするステップを省略するステップと、 前記ビットストリーム信号を逆変換して、前記ステレオ信号を求めるステップとを有する、請求項11ないし14いずれか一項に記載の復号方法。
- 16前記ビットストリーム信号が直接ステレオ符号化または同時ステレオ符号化により符号化されていることに応じて、前記ダウンミックス信号の時間領域表現を変換するステップとサイド信号を計算するステップを省略するステップと、 前記ビットストリーム信号にエンコードされた各チャネルの周波数領域表現を逆変換して、前記ステレオ信号を求めるステップとをさらに有する、請求項15に記載の復号方法。
- 17汎用コンピュータにより実行されたとき、前記汎用コンピュータに、請求項11ないし16いずれか一項に記載の方法を実行させるコンピュータプログラム。
Independent claims17
62 paragraphs, as filed
0001The inventions disclosed herein generally relate to stereo audio coding, and more specifically to stereo coding techniques using complex prediction in the frequency domain.
0002Simultaneous coding of the left (L) and right (R) channels of a stereo signal makes coding more efficient than coding L and R independently. A common approach to simultaneous stereo coding is mid / side (M / S) coding. Here, the mid (M) signal is composed by adding the L signal and the R signal, for example, the M signal is<maths num="1"><img id="000002" he="16" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>Obtained by Also, the side (S) signal is configured by pulling two channels LR, for example the S signal<maths num="2"><img id="000003" he="16" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>Obtained by In the case of M / S coding, M signal and S signal are encoded instead of L signal and R signal.
0003MPEG (Moving Picture Experts Group) AAC (Advanced Audio Coding) standard (see standard document ISO / IEC 13818-7) selects L / R stereo coding and M / S stereo coding with variable time and frequency. it can. In this way, the stereo encoder can apply L / R coding to one frequency band of the stereo signal, and M / S coding is used to encode the other frequency band of the stereo signal (variable frequency). In addition, the encoder can switch between L / R coding and M / S coding in time (variable time). In MPEG AAC, stereo encoding is done in the frequency domain, more specifically in the MDCT (Modified Discrete Cosine Transform) domain. This makes it possible to adaptively select either L / R coding or M / S coding, which is variable in frequency and time.
0004Parametric stereo coding is a method of efficiently coding a stereo audio signal as a monaural signal and a small amount of side information as a stereo parameter. This is part of the MPEG-4 audio standard (see standard document ISO / IEC 14496-3). Monaural signals can be encoded using any audio encoder. Stereo parameters are built into the accessories of the mono bit stream, making them fully forward and backward compatible. In the decoder, the monaural signal is first decoded and then the stereo signal is reconstructed using stereo parameters. The uncorrelated version of the decoded mono signal has zero cross-correlation with the mono signal. This uncorrelated signal is filtered by a decorrelator, for example, with a suitable all-pass filter that includes a delay line. Generated by filter). Basically, the uncorrelated signal has the same spectral and temporal energy distribution as the mono signal. The monaural signal is input to the upmix process along with the uncorrelated signal. This process is controlled by stereo parameters to reconstruct the stereo signal. For more detailed information, refer to Non-Patent Document 1.
0005MPEG surround (MPS; see ISO / IEC 23003-1 and Non-Patent Document 2) combines the principle of parametric stereo coding with the principle of residual coding, and is the residual that uncorrelated signals are transmitted. Replaced with to improve the perceptible sound quality. Residual coding is done by downmixing the multichannel signal and optionally extracting spatial cues. In the downmix process, the residual signal representing the error signal is calculated, encoded and transmitted. The residual signal replaces the uncorrelated signal in the decoder. In the hybrid approach, the residual signal replaces the uncorrelated signal in a constant frequency band, preferably in a relatively low band.
0006The current MPEG Unified Speech and Audio Coding (USAC) system shows two examples in Figure 1, where the decoder has a complex value quadrature mirror filter (QMF) bank located downstream of the core decoder. The QMF representation obtained as the output of this filter bank is complex and therefore double oversampled and can be configured as a downmix signal (ie mid signal) M and a residual signal D. An upmix matrix with complex value components can be used for this. The L and R signals (in the QMF region)<maths num="3"><img id="000004" he="17" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>Obtained as. Here, g is a real value gain factor and α is a complex value prediction coefficient. α is preferably selected so that the energy of the residual signal D is minimized. The gain factor can be determined by normalization, that is, the power of the sum signal is equal to the sum of the powers of the left and right signals. The real and imaginary parts of the L and R signals are redundant with each other, and in principle one can be calculated based on the other. However, it has the benefit of being able to use a spectral band replication (SBR) decoder without producing audible aliasing artifacts later. The use of oversampled signal representations is also chosen for the same reason to prevent other time or frequency adaptive signal processing (not shown) related artifacts such as mono-stereo upmix. Inverse QMF filtering is the final processing step in the decoder. The band-limited QMF representation of the signal allows the band-limited residual method and the "residual fill" method to be used. These techniques can be incorporated into this type of decoder.
0007The above coding configuration works well for low bitrates, generally below 80 kb / s, but is not optimal for high bitrates in terms of computational complexity. More specifically, at high bitrates, SBR tools are generally not used (because they do not improve coding efficiency), and then decoders without SBR steps have complex value upmix matrices. It uses a QMF filter bank, which is computationally expensive and delays (at a frame length of 1024 samples, the QMF analysis / synthesis filter bank delays 961 samples). This clearly demonstrates the need for a more efficient coding configuration.
<p num="0008"><nplcit num="1"><text>H. Purnhagen, "Low Complexity Parametric Stereo Coding in MPEG-4", Proc. Of the 7th Int. Conference on Digital Audio Effects (DAFx'04), Naples, Italy, October 5-8, 2004, pages 163-168</text></nplcit><nplcit num="2"><text>J. Herre et al., "MPEG Surround-The ISO / MPEG Standard for Efficient and Compatible Multi-Channel Audio Coding", Audio Engineering Convention Paper 7084, 122 <nd> Convention, May 5-8, 2007</text></nplcit></p>
<p num="0009"> An object of the present invention is to provide a method and an apparatus for performing stereo coding with high calculation efficiency even in a high bit rate range.</p><p num="0010"> The present invention achieves this object by providing a coder and a decoder, coding and decoding methods, and computer program products for encoding and decoding, respectively, as defined in the independent claims. Dependent terms define embodiments of the present invention.</p><p num="0011"> In the first aspect, the present invention provides the following system. That is, a decoder system that provides a stereo signal by complex predictive stereo coding: It is an upmix stage configured to generate the stereo signal based on the first frequency domain display of the downmix signal (M) and the residual signal (D), and each first frequency domain display is The upmix step comprises an upmix step having a first spectral component representing the spectral content of the corresponding signal represented by the first subspace of the multidimensional space. A module that calculates the second frequency domain display of the downmix signal based on the first frequency domain display of the downmix signal, and the second frequency domain display is included in the first subspace. A module having a second spectral component representing the spectral content of a signal represented by a second subspace of said multidimensional space, including a portion of the multidimensional space that is not. Side signals (side signals based on the first and second frequency domain indications of the downmix signal encoded into the bitstream signal, the first frequency domain indication of the residual signal, and the complex prediction factor (α). An upmix stage with a weighted adder to calculate S), and It has a sum / difference step of calculating the stereo signal based on the first frequency domain display of the downmix signal and the side signal. The upmix stage can further operate in a pass-through mode in which the downmix signal and the residual signal are directly supplied to the sum / difference stage.</p><p num="0012"> In the second aspect, the present invention provides the following system. That is, An encoder system that encodes a stereo signal with a bitstream signal by complex predictive stereo coding. An estimator that estimates complex prediction coefficients and (a) A coding step capable of transforming the stereo signal into a frequency domain representation of the downmix signal and the residual signal having a relationship determined by the value of the complex prediction factor. It has the coding stage and a multiplexer that receives output from the estimator and encodes it into the bitstream signal.</p><p num="0013"> A third and fourth aspect of the present invention provides a method of encoding a stereo signal into a bitstream and a method of decoding the bitstream into at least one stereo signal . The technical features of each method are similar to those of the encoder system and the decoder system, respectively. In the fifth and sixth aspects, the invention provides a computer program product that includes instructions to perform each method on a computer.</p><p num="0014"> The present invention benefits from the advantages of unified stereo coding in MPEG USA C systems. These advantages are preserved even at high bitrates without significantly increasing the computational complexity associated with the QMF-based approach, where SBR is not commonly utilized, and the reason this is possible is critical sampling. The MDCT transform is the basis of the MPEG USAC transform, but at least according to the present invention, complex prediction by the present invention if the coded audio bands of the downmix and residual channels are the same and the upmix process does not include uncorrelation. It can also be used for stereo coding. This means that no additional QMF conversion is needed anymore. A typical embodiment of complex predictive stereo coding in the QMF domain significantly increases the number of operations per unit time as compared to conventional L / R or M / S stereo. As such, the coding device according to the invention appears to be competitive at such bit rates because it provides high sound quality with a modest computational load.</p><p num="0015"> As one of ordinary skill in the art will notice, due to the fact that the upmix stage can also operate in pass-through mode, the decoder adapts, at the discretion of the encoder, by conventional direct or simultaneous coding, and complex predictive coding. Can be decrypted. Therefore, if the decoder cannot raise the sound quality level more aggressively than the conventional direct L / R stereo coding or simultaneous M / S stereo coding, it can be guaranteed to maintain at least the same level. Therefore, the decoder according to this aspect of the present invention can be regarded as a superset with respect to the background technology from the functional point of view.</p><p num="0016"> The advantage over QMF-based predictive coded stereo is that the signal can be completely reconstructed (except for quantization errors that can be arbitrarily reduced).</p><p num="0017"> As described above, the present invention provides a coding device that performs conversion-based stereo coding by complex prediction. Preferably, the apparatus according to the present invention is not limited to complex predictive stereo coding, and can operate in direct L / R stereo coding by background technology or simultaneous M / S passing coding, and can be used in specific applications. You can select the most suitable coding method during a particular time.</p><p num="0018"> The oversampled representation of the signal (eg, the complex representation) contains both the first and second spectral components and is used as the basis for the complex prediction according to the invention, thus computing such an oversampled representation. The module is composed of an encoder system and a decoder system according to the present invention. Spectral components refer to the first and second subspaces of a multidimensional space. It is a set of time-dependent functions of a given time length (eg, the length of a given time frame) sampled at a finite sampling frequency. It is well known that functions in this multidimensional space can be approximated by a finite weighted sum of basis functions.</p><p num="0019"> As will be apparent to those skilled in the art, encoders configured to work with decoders provide an oversampled display on which predictive coding is based to allow faithful reproduction of the encoded signal. It has the equivalent modules it provides. Such equivalent modules are modules that are the same or similar, or modules that have the same or similar transfer characteristics. In particular, the encoder and decoder modules may be similar or dissimilar units that execute computer programs that perform equivalent mathematical operations, respectively.</p><p num="0020"> In some embodiments of a decoder system or encoder system, the first spectral component has a real value represented by the first subspace and the second spectral component has an imaginary value represented by the second subspace. Have. Both the first and second spectral components constitute a complex spectral representation of the signal. The first subspace is the linear span of the first set of basis functions, the second subspace is the linear span of the second set of basis functions, some of which are with the first set of basis functions. Is linearly independent.</p><p num="0021"> In one embodiment, the module that calculates the complex representation is a real-imaginary transformation, that is, a module that calculates the imaginary number of the spectrum of a discrete-time signal based on the real spectrum representation of the signal. This transformation is based on strict or approximate mathematical relationships, such as harmonic analysis and equations from heuristic relationships.</p><p num="0022"> In certain embodiments of the decoder system or encoder system, the first spectral component is the time-frequency domain transform of the discrete time domain signal, preferably the Fourier transform, such as the Discrete Cosine Transform (DCT), the Modified Discrete Cosine Transform (MDCT). ), Discrete Cosine Transform (DST), Modified Discrete Cosine Transform (MDCT), Fast Fourier Transform (FFT), Prime-factor-based Fourier Transform, etc. In the first four cases, the second spectral component is determined by DST, MDST, DCT, and MDCT, respectively. As is well known, a linear span of a cosine that is periodic in a unit period constitutes a subspace that is not completely included in the linear span of a periodic sine in the same period. Preferably, the first spectral component is determined by MDCT and the second spectral component is determined by MDST.</p><p num="0023"> In one embodiment, the decoder system comprises at least one temporal noise shaping module (TNS module, i.e. TNS filter), which is located upstream of the upmix stage. Generally speaking, the use of TNS improves the perceived sound quality of signals with transient components, which also applies to embodiments of the decoder system of the present invention with TNS. In traditional L / R and M / S stereo coding, the TNS filter is applied just before the inverse transformation as the last processing step in the frequency domain. However, in the case of complex predictive stereo coding, it is often advantageous to apply the TNS filter to the downmix and residual signals, i.e. before the upmix matrix. In other words, TNS is applied to a linear combination of left and right channels, which has several advantages. First, it turns out that in some situations TNS is only advantageous for, for example, downmix signals. Second, TNS filtering can be omitted for residual signals, which means economical use of available bandwidth. The TNS filter coefficient only needs to be transmitted for the downmix signal. Second, the calculation of the oversampled representation of the downmix signal (eg, MDST data is derived from the MDCT data to construct a complex frequency domain representation) is required for complex predictive coding, but down. The time domain display of the mixed signal needs to be calculable. This means that the downmix signal is preferably available as a time sequence of the uniformly determined MDCT spectrum. If the TNS filter is applied in the decoder after an upmix matrix that converts the downmix / residual display to a left / right display, only a sequence of TNS residual MDCT spectra of the downmix signal is obtained. This makes efficient calculation of the corresponding MDST spectrum very difficult. This is especially true when the left / right channels use TNS filters with different characteristics.</p><p num="0024"> It should be emphasized that obtaining a time sequence of the MDCT spectrum is not an absolute criterion for obtaining a fitted MDST display to serve as the basis for complex predictive coding. In addition to experimental evidence, this fact generally indicates that the TNS, for example, several kilohertz, so that the TNS-filtered residual signal roughly corresponds to the low-frequency unfiltered residual signal. This can be explained by the fact that it applies only to higher frequencies. As described above, the present invention can be implemented as a decoder for complex predictive stereo coding in which the TNS filter is arranged outside the upstream of the upmix stage, as described below.</p><p num="0025"> In one embodiment, the decoder system comprises at least one subtle TNS module located downstream of the upmix stage. Depending on the selector device, the TNS module upstream of the upmix stage or the TNS module downstream of the upmix stage. Under certain circumstances, the calculation of the complex frequency domain representation does not require that the time domain representation of the downmix signal be computable. Moreover, as mentioned above, the decoder can selectively operate in direct or simultaneous coding mode without applying complex predictive coding, using the TNS module in the conventional place, i.e. the last processing in the frequency domain. It is better to use it as one of the steps.</p><p num="0026"> In one embodiment, the decoder system is configured to save processing resources and possibly energy by deactivating the module that computes the second frequency domain representation of the downmix signal. The downmix signal is partitioned into consecutive time blocks, each time block associated with a value of the complex prediction factor. This value is determined by the encoder working with the decoder to determine for each time block. Further, in this embodiment, the module that calculates the second frequency domain representation of the downmix signal has the absolute value of the imaginary part of the complex prediction coefficient zero or more than a predetermined tolerance for a given time block. If small, it is configured to deactivate itself. Deactivating the module means not calculating the second frequency domain representation of the downmix signal for this time block. Without deactivation, the second frequency domain display (eg, a set of MDST coefficients) is multiplied by zero, or a number on the order approximately the same as the decoder's machine epsilon (rounded) or other suitable threshold. ..</p><p num="0027"> A further development of the above embodiment saves processing resources at the sub-level of the time block in which the downmix signal is partitioned. For example, such a sub-level within a time block is the frequency band, and the encoder determines the value of the complex prediction factor for each frequency band within the time block. Similarly, the method of generating the second frequency domain representation is configured to suppress operations on frequency bands within a time block where the complex prediction factor is zero or the magnitude is less than the tolerance.</p><p num="0028"> In one embodiment, the first spectral component is a conversion factor arranged in a time block of conversion factors, each block being generated by applying a conversion of the time domain signal to a time segment. In addition, the module that calculates the second frequency domain display of the downmix signal -The first intermediate component is obtained from the first spectral component, and -The coupling of the first spectral component is formed by at least a part of the impulse response to obtain the second intermediate component. -It is configured to obtain the second spectral component from the second intermediate component. This procedure calculates the second frequency domain display directly from the first frequency domain display, as described in detail in US Pat. No. 6,980,933B2, especially in columns 8 to 28, especially in Equation 41. Can be done. As one of ordinary skill in the art will notice, the calculation is not performed in some time domain, as opposed to, for example, an inverse transformation in which different transformations follow.</p><p num="0029"> For the examples of complex predictive stereo coding according to the present invention, the computational complexity increases only slightly compared to conventional L / R or M / S stereo (caused by complex predictive stereo coding in the QMF region). It is speculated that it is significantly less than the increase). This type of embodiment, which involves a rigorous calculation of the second spectral component, results in a delay that is only a few percent longer than that caused by the QMF-based embodiment (assuming the time block length is 1024 samples, QMF. Compared to the delay of 961 samples in the analytical / synthetic filter bank).</p><p num="0030"> Preferably, at least in some of the above embodiments, the impulse response is adapted to the conversion required for the first frequency domain display, or more precisely, the conversion required by its frequency response characteristics.</p><p num="0031"> In some embodiments, the first frequency domain representation of the downmix signal applies to one or more analysis window functions (or cutoff functions such as rectangular windows, sinusoidal windows, Kaiser-Vessel windows, etc.). The resulting transformation is obtained, one of which is to achieve temporal segmentation without producing dangerous noise volume or causing undesired changes in the spectrum. In some cases, such window functions partially overlap. Next, preferably, the frequency response characteristics of the conversion depend on the characteristics of one or more of the analysis window functions described above.</p><p num="0032"> The computational load can be reduced by using an approximate second frequency domain display with further reference to embodiments characterized by calculation of a second frequency domain display in the frequency domain. Such an approximation can be achieved by not requiring completeness in the information on which the calculation is based. For example, according to the teachings of US Pat. No. 6,980,933B2, the first frequency domain data from three time blocks, namely the block simultaneous with the output block, the preceding block, and the succeeding block, is downmixed in one block. Required for exact calculation of the second frequency domain representation of the signal. Suitable approximations by omitting or replacing data from subsequent blocks and / or preceding blocks with zeros (caused by module behavior, i.e. not contributing to delay) for complex predictive coding according to the present invention. Is obtained so that the calculation of the second frequency domain representation is based on only one or two time blocks. Note that omitting the input data means, for example, rescaling the second frequency domain display in the sense that it no longer represents the same power, but as described above, it is equivalent on both the encoder and decoder sides. It can be used as the basis for complex predictive coding as long as it is calculated in the above way. Indeed, this type of rescaling is compensated for by the corresponding changes in the prediction coefficient values.</p><p num="0033"> Yet another approximation method for calculating the spectral components that form part of the second frequency domain display of the downmix signal involves combining at least two components from the first frequency domain display. The latter components are adjacent in terms of time and / or frequency. As an alternative, the latter component can be combined by finite impulse response (FIR) filtering in a relatively small number of steps. For example, in a system with a time block size of 1024, such FIR filters include 2, 3, 4 and so on taps. A description of this type of approximate calculation method can be found, for example, in US Patent Application Publication No. 2005/0197831A1. When a window function that gives a relatively small weight near each time block boundary is used, for example, a non-rectangular function, the second spectral component of the time block is based only on the combination of the first spectral component of the same time block. However, the same amount of information cannot be obtained for the outermost component. The approximation error that may occur due to such practices can be suppressed to some extent or concealed by the shape of the window function.</p><p num="0034"> In one embodiment of a decoder designed to output a time domain stereo signal, it is possible to switch between direct or simultaneous stereo coding and complex predictive coding. This can be achieved by providing the following: That is, -A switch that can selectively operate as a pass-through stage (without changing the signal) or as a sum / difference conversion; -Inverse conversion stage that performs frequency / time conversion; and -A selector device that inputs a signal encoded directly (or simultaneously) or a signal encoded by complex prediction to the inverse conversion stage. As those skilled in the art will notice, such flexibility on the side of the decoder gives the encoder the freedom to choose between conventional direct or simultaneous encoding and complex predictive encoding. Therefore, this embodiment can guarantee that at least the same level is maintained when the sound quality level of the conventional direct L / R stereo coding or simultaneous M / S stereo coding cannot be exceeded. Therefore, the decoder according to this embodiment can be regarded as a superset with respect to the related technology.</p><p num="0035"> Another group of embodiments of the decoder system performs the calculation of the second spectral component of the second frequency domain representation over the time domain. More precisely, the inverse transformation of the transformation for which the first spectral component was obtained (or can be obtained) is applied, and then a different transformation with the second spectral component as the output is performed. Specifically, MDST is performed after reverse MDCT. In such an embodiment, in order to reduce the number of transformations and inverse transformations, the output of the inverse MDCT is sent to the MDST and the output terminal of the decoding system (which may be preceded by another processing step).</p><p num="0036"> For the examples of complex predictive stereo coding according to the present invention, the computational complexity increases only slightly compared to conventional L / R or M / S stereo (caused by complex predictive stereo coding in the QMF region). It is speculated that it is significantly less than the increase).</p><p num="0037"> As a further development of the embodiment mentioned in the paragraph above, the upmix step may have a further inverse conversion step of processing the side signal. Then, the time domain display of the side signal generated by the further inverse conversion step and the time domain display of the downmix signal generated by the inverse conversion described above are supplied to the sum / difference stage. Again, from the point of view of computational complexity, the latter signal is conveniently supplied to both the sum / difference steps described above and different conversion steps.</p><p num="0038"> In one embodiment, a decoder designed to output a time domain stereo signal can switch between direct L / R stereo coding or simultaneous M / S stereo coding and complex predictive stereo coding. is there. This can be achieved by providing the following: That is, A switch that can operate as a pass-through stage or as a sum / difference stage; -A further inverse conversion step to calculate the time domain display of the side signal; The inverse conversion step is upstream of the upmix (preferably when the switch is activated and acts as a pass filter, such as when decoding a stereo signal generated by complex predictive coding). The switch is activated and acts as a sum / difference step, either to an additional sum / difference step connected to a point downstream of the switch, or (preferably as if decoding a directly encoded stereo signal). Sometimes) a selector device that connects to a combination of the downmix signal from the switch and the side signal from the weight adder. As those skilled in the art will notice, this gives the encoder the freedom to choose between conventional direct or simultaneous coding and complex predictive coding, i.e. guarantees a sound quality level that is at least equal to direct or simultaneous stereo coding. it can.</p><p num="0039"> In one embodiment, the encoder system according to the second aspect of the present invention has an estimator that estimates a complex prediction factor for the purpose of reducing or minimizing the signal power or average signal power of the residual signal. Minimization takes place over a period of time, preferably over a time segment or time block or time frame to be encoded. The square of the amplitude can be used as a measure of the instantaneous signal power, and the integral of the square of the amplitude over an hour interval can be used as a measure of the average signal power in that time interval. Preferably, the complex prediction factor can be determined for each time block and for each frequency band. That is, its value is set to reduce the average power (ie, total energy) of the residual signal in that time block and frequency band. Specifically, modules that estimate parametric stereo-encoded parameters such as IID, ICC and IPD or similar parameters provide output that allows the calculation of complex prediction coefficients by mathematical relationships known to those of skill in the art.</p><p num="0040"> In one embodiment, the coding step of the encoder system further functions as a pass-through step to allow direct stereo coding. In situations where direct stereo coding is expected to provide higher sound quality, selecting this allows the encoder system to ensure that the encoded stereo signal has at least the same sound quality as direct coding. Similarly, in situations where significant computational improvement but the heavy computational load caused by complex-predictive coding is not desirable, encoder systems have readily available options that save computational resources. The determination between simultaneous, direct real predictive coding and complex predictive coding in the coder is generally based on the principle of rate / distortion optimization.</p><p num="0041"> In one embodiment, the encoder system calculates a second frequency domain representation based directly on the first spectral component (ie, without applying the inverse transformation to the time domain and without using the time domain data of the signal). Has a module. With respect to the corresponding embodiments of the decoder system described above, this module has a similar configuration, i.e., similar but different order of processing operations, so that the encoder outputs data suitable for input on the decoder side. It is composed of. For purposes of illustrating this embodiment, the stereo signal to be encoded has or is converted to this configuration with mid and side channels, and the coding step is configured to receive a first frequency domain representation. Suppose that. The coding stage has a module that computes the second frequency domain representation of the midchannel. (The first and second frequency domain representations referenced here are as defined above; Specifically, the first frequency domain display may be an MDCT display, and the second frequency domain display may be an MDST display. The coding step further consists of a side signal and two frequency domain representations of the mid signal, and a weighted addition that calculates the residual signal as a linear combination weighted by the real and imaginary parts of the complex prediction coefficients. Have a vessel. The mid signal, or preferably its first frequency domain display, is used directly as the downmix signal. In this embodiment, the estimator further determines the value of the complex prediction factor for the purpose of minimizing the power of the residual signal or the average signal power. The final operation (optimization) is the original operation (optimization) by feedback control, and if necessary, by feedback control in which the estimator receives the residual signal obtained by the current prediction coefficient value to be adjusted, or feedforwardly. This is done by calculations performed directly on the left / right channels of the stereo signal or on the mid / side channels. Complex prediction coefficients are calculated directly (especially non-repetitively or non-feedback) based on the first and second frequency domain representations of the mid signal and the first frequency domain representation of the side signals. The feedforward method is preferred. As a reminder, after determining the complex prediction coefficients, either direct simultaneous real prediction coding is performed, taking into account the quality obtained with each option (preferably, for example, perceptual quality considering the signal vs. mask effect). Determine whether to perform complex predictive coding. Therefore, the above statement should not be construed as the absence of a feedback mechanism in the encoder.</p><p num="0042"> In one embodiment, the encoder system has a module that computes a second frequency domain representation of a mid (ie downmix) signal over time domain. The details of the embodiment relating to this embodiment are the same, at least as far as the calculation of the second frequency domain representation is concerned, and can be performed in the same manner as the corresponding decoder embodiment. In this embodiment, the coding steps include: That is: Sum / difference stage that converts a stereo signal into a mid channel and a side channel; A conversion step that provides a side channel frequency domain display and a mid channel complex value (ie, oversampled) frequency domain display; and -A weight adder that calculates the residual signal using the complex prediction coefficient as the weight. Here, the estimator receives the residual signal and, in some cases in the form of feedback control, determines the complex prediction factor that reduces or minimizes the power or average power of the residual signal. However, preferably, the estimator receives the stereo signal to be encoded and determines the prediction factor based on it. It is advantageous from the economical point of view of calculation to use the critically sampled frequency domain display of the side channel. This is because in this embodiment the side channels are not multiplied by complex numbers. Preferably, the conversion step may include an MDCT step and an MDST step configured in parallel. Both have a mid-channel time domain display as input. In this way, the oversampled frequency domain display of the mid channel and the critically sampled frequency domain display of the side channel are generated.</p><p num="0043"> As a reminder, the methods and devices disclosed in this section can be applied to the coding of signals with more than 2 channels with appropriate modifications within the capabilities of those skilled in the art, including routine experiments. Such changes to multi-channel operability can be made, for example, in line with sections 4 and 5 of the paper by J. Herre et al. Cited above.</p><p num="0044"> In yet another embodiment, the features of the two or more embodiments described above can be combined, unless apparently complementary. The fact that the two features are listed in different claims does not mean that they cannot be combined. Similarly, in yet another embodiment, features that are not necessary or essential for the desired purpose may be omitted. As an example, the decoding system according to the invention is performed without the dequantization step if the coded signal to be processed is not quantized or is already in a format suitable for processing in the upmix step. May be good.</p>
0045The present invention will be further described by way of embodiments described in the next section with reference to the accompanying drawings.<figref num="1A">It is a block diagram which shows the QMF base decoder by the background technology.</figref><figref num="1B">It is a block diagram which shows the QMF base decoder by the background technology.</figref><figref num="2">It is a block diagram which shows the MDCT-based stereo decoder system which has a complex prediction by one Embodiment of this invention. The complex representation of the channels of the decoded signal is calculated in the frequency domain.</figref><figref num="3">It is a block diagram which shows the MDCT-based stereo decoder system which has a complex prediction by one Embodiment of this invention. The complex representation of the channels of the decoded signal is calculated in the time domain.</figref><figref num="4">It is a figure which shows another embodiment of the decoder system of FIG. The position of the active TNS stage is selectable.</figref><figref num="5">FIG. 3 is a block diagram showing an MDCT-based stereo encoder system with complex prediction according to an embodiment of another aspect of the invention.</figref><figref num="6">It is a block diagram which shows the MDCT-based stereo encoder system which has a complex prediction by one Embodiment of this invention. The complex representation of the channels of the encoded signal is calculated based on its time domain representation.</figref><figref num="7">It is a figure which shows another embodiment of the encoder system shown in FIG. The system can also operate directly in L / R coding mode.</figref><figref num="8">It is a block diagram which shows the MDCT-based stereo encoder system which has a complex prediction by one Embodiment of this invention. The complex representation of the channels of the encoded signal is calculated based on its first frequency domain representation. The system can also operate directly in L / R coding mode.</figref><figref num="9">It is a figure which shows another embodiment of the encoder system shown in FIG. 7. The system further includes a TNS stage located downstream of the coding stage.</figref><figref num="10">2 and 8 are diagrams showing another embodiment of the portion indicated by the label A.</figref><figref num="11">It is a figure which shows another embodiment of the encoder system shown in FIG. The system further includes frequency domain correction devices located upstream and downstream of the coding stage, respectively.</figref><figref num="12">A graph showing listening test results at 96 kb / s from 6 subjects, showing different complexity vs. sound quality trade-off options for calculating or approximating MDST spectra. Here, the data points indicated by the label "+" indicate hidden criteria. "X" indicates a 3.5kHz band limiting anchor. "*" Indicates conventional stereo (M / S or L / R) by USAC. indicates MDCT region unified stereo coding by complex prediction with the imaginary part of the prediction coefficient disabled (ie, by real-valued prediction that does not require MDST). indicates MDCT region unified stereo coding by complex prediction that calculates the approximate value of MDST using the current MDCT frame. "" indicates MDCT region unified stereo coding by complex prediction that calculates the approximate value of MDST using the current and previous MDCT frames. indicates MDCT region unified stereo coding by complex prediction that calculates MDST using the current, previous, and next MDCT frames.</figref><figref num="13">It is a figure which shows the data of FIG. 12 as a difference score about MDCT region unified stereo coding by complex prediction which calculates the approximate value of MDST using the current MDCT frame.</figref><figref num="14A">It is a block diagram which shows one Embodiment of the decoder system by Embodiment of this invention.</figref><figref num="14B">It is a block diagram which shows the other embodiment of the decoder system by embodiment of this invention.</figref><figref num="14C">It is a block diagram which shows still another Embodiment of the decoder system by Embodiment of this invention.</figref><figref num="15">It is a flowchart which shows the decoding method by one Embodiment of this invention.</figref><figref num="16">It is a flowchart which shows the coding method by one Embodiment of this invention.</figref>
0046I. Decoder system Figure 2 is in the form of a schematic block diagram with at least one complex prediction coefficient value α = α.<sub>R</sub>+ iα<sub>I</sub>Indicates a decoding system that decodes a bitstream with. The MDCT representation of a stereo signal has a downmix M channel and a residual D channel. Real part of prediction coefficient and α<sub>R</sub>And imaginary part α<sub>I</sub>Is quantized and / or coded jointly. However, preferably, the real and imaginary parts are quantized independently and uniformly, generally with a step size of 0.1 (dimensionless number). According to the MPEG standard, the resolution of the frequency band used for the complex predictor coefficient does not have to be the same as the resolution of the scale factor band (sfb, i.e. a group of MDCT lines using the same MDCT quantization step size and quantization range). In particular, the frequency band resolution of the prediction factor is the Bark scale. It is psychoacoustically valid, such as scale). The demultiplexer 201 is configured to extract these MDCT representations and prediction factors (part of the illustrated control information) from the supplied bitstream. In fact, the TNS information in which control information above the complex prediction coefficient is encoded in the bitstream, such as the instruction to decode it in predictive mode or non-predictive mode and TNS information, is the TNS of the decoder system. Contains the values of the TNS parameters used by the synthetic) filter. When using the same set of TNS parameters for multiple TNS filters, such as both channels, it is more economical to receive this information in the form of bits indicating the identity of the set of parameters than to receive the two sets of parameters separately. Is. For example, it may also include information on whether to apply TNS before or after the upmix phase, based on two optional psychoacoustic assessments. In addition, the control information indicates the individually limited bandwidth of the downmix signal and the residual signal. For each channel, frequency bands above the bandwidth limit are not decoded and are set to zero. In some cases, the energy content of the highest frequency band is so small that it is already zero when quantized. In normal practice (see the MPEG standard max_sfb parameter), the same bandwidth limitation should be used for both the downmix signal and the residual signal. However, the residual signal has much more energy content localized to the low frequency band than the downmix signal. Therefore, by imposing a dedicated bandwidth upper limit on the residual signal, it is possible to reduce the bit rate without significantly impairing the sound quality. For example, this is tuned by two independent max_sfb parameters encoded in the bitstream, one for the downmix signal and one for the residual signal.
0047In this embodiment, the MDCT representation of a stereo signal is a fixed number of data points (eg, 1024 points), one of a plurality of constant data points (eg, 128 points or 1024 points), or a variable number. Segmented into consecutive time frames (ie, time blocks) containing points. As known to those of skill in the art, MDCTs are critically sampled. The output of the decoding system, shown in the right part of the figure, is a stereo signal in the time domain with a left L channel and a right R channel. The dequantization module 202 demultiplexes the bitstream input to the decoding system into, if necessary, two bitstreams corresponding to the downmix channel and the residual channel obtained after demultiplexing the original bitstream. It is configured to handle. The dequantized channel signal is a transformation matrix<maths num="4"><img id="000005" he="17" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>In pass-through mode corresponding to, or transformation matrix<maths num="5"><img id="000006" he="18" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>Provided to switching assembly 203 that can operate in sum and difference modes corresponding to. The decoder system includes a second switching assembly 205, as further described in the next paragraph. Both switching assemblies 203, 205 can operate frequency-selectively, like most other switches and switching assemblies in this embodiment and the embodiments described below. This allows decoding in a wide variety of decoding modes, such as frequency-dependent L / R or M / S decoding, as is known as a related technique. Therefore, the decoder according to the present invention can be regarded as a superset with respect to the related art.
0048Assuming that the switching assembly 203 is in pass-through mode, in this embodiment the dequantized channel signal is passed through each TNS filter 204. The TNS filter 204 is not essential to the operation of the decoding system and can be replaced by a pass-through element. After this, the signal is fed to a second switching assembly 205 that has the same function as the upstream switching assembly 203. When the input signal is input as described above and set to pass-through mode, the outputs of the second switching assembly 205 are the downmix channel signal and the residual channel signal. The downmix signal, represented by a temporally continuous MDCT spectrum, is fed to a real-imaginary conversion 206 configured to compute the MDST spectrum of the downmix signal. In this embodiment, one MDST frame is based on three MDCT frames, one front frame, one current (ie simultaneous) frame, and one back frame. It is symbolic that the input side of the real / imaginary conversion 206 has a delay component (Z)<sup>-1</sup>, Z) Shown.
0049The MDST display of the downmix signal obtained from the real / imaginary conversion 206 is the imaginary part α of the prediction coefficient.<sub>I</sub>Weighted by the real part α of the prediction coefficient<sub>R</sub>And the residual signal is added to the MDCT display of the downmix signal weighted by the MDCT display. The two additions and multiplications are performed by the adders and multipliers that (functionally) constitute the weighting adders 210, 211. They are fed with the value of the complex prediction factor α that was encoded in the bitstream originally received by the decoder system. One complex prediction factor is determined for each time frame. The complex prediction factor may be determined more frequently or one for each frequency band in the frame. The frequency band is an acoustically psychologically motivated partition. The complex prediction coefficients need not be determined very often, as will be described later with respect to the coding system of the present invention. The real-imaginary conversion 206 is synchronized with the weight adder so that the current MDST frame of the downmix channel signal is combined with the simultaneous MDCT frame of each of the downmix channel signal and the residual channel signal. The sum of these three signals is the side signal S = Re {αM} + D. In this equation, M includes both the MDCT and MDST indications of the downmix signal, i.e. M = M<sub>MDCT</sub>-iM<sub>MDST</sub>Is. D = DMDCT is a real number. In this way, a stereo signal having a downmix channel and a side channel is obtained, and the sum-difference conversion 207 is performed from this stereo signal.<maths num="6"><img id="000007" he="17" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>Restores the left and right channels. These signals are represented in the MDCT region. In the final step of the decoding system, an inverse MDCT209 is applied to each channel to obtain a time domain display of the left and right stereo signals.
0050The possible implementation of the real-imaginary conversion 206 is described in detail in the applicant's US Pat. No. 6,980,933B2, as described above. The transformation can be represented as a finite impulse response filter by Equation 41 described in the above document. For example, for even points<maths num="7"><img id="000008" he="37" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths><maths num="8"><img id="000009" he="14" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>Is. Another straightforward approach can be found in US Patent Application Publication No. 2005/0197831A1.
0051It is possible to further reduce the amount of input data on which the calculation is based. For the sake of explanation, the connection between the real / imaginary conversion 206 and its upstream, which is the part indicated by A in the figure, may be replaced with a simplified modification. Two of them, A'and A', are shown in FIG. Modification A'provides an approximation of the imaginary representation of the signal. Here, the MDST calculation considers only the current frame and the previous frame. See the formula above in this paragraph for X for p = 0, ..., N-1<sub>III</sub>(p) This is done by setting = 0 (index III indicates a later time frame). The MDST calculation does not cause a time delay because variant A'does not require the MDCT spectrum of the later frame as input. Obviously, this approximation suggests that the accuracy of the resulting MDST signal is somewhat reduced, but the energy of this signal is also reduced. Due to the nature of predictive coding, the latter is α<sub>I</sub>Can be completely compensated by increasing.
0052A modified example A is shown in FIG. It uses only the MDCT data of the current time frame as input. The MDST display obtained by variant A is less accurate than that obtained by variant A . On the other hand, the modified example A operates with zero delay like the modified example A , and the calculation complexity is low. As mentioned above, as long as the encoder system and the decoder system use the same approximation, the waveform coding characteristics are not affected.
0053It should be noted that the imaginary part of the complex prediction coefficient of the MDST spectrum is non-zero, i.e. α, regardless of which variant A, A or A , or a further development of these, is used.<sub>I</sub>Only the part where 0 needs to be calculated. In a practical situation, this is the absolute value of the imaginary part of the coefficient | α<sub>I</sub>It can be understood that | means that it is larger than a predetermined threshold. This predetermined threshold relates to the unit round-off of the hardware used. If the imaginary part of the coefficients of all frequency bands in a time frame is zero, then it is not necessary to calculate MDST data for that frame. Therefore, after all, the real / imaginary conversion 206 does not generate the MDST output, so that | α<sub>I</sub>It is configured to respond when the value of | is very small. This saves computational resources. However, in the embodiment in which one frame of MDST data is generated using more than the current frame, sufficient input data for the real-imaginary conversion 206 occurs when the next time frame related to the non-zero prediction coefficient occurs. As there is, the unit upstream of the conversion 206 must continue to operate without the need for the MDST spectrum, and in particular the second switching assembly 205 must continue to transfer the MDCT spectrum. This is, of course, the next time block.
0054Returning to FIG. 2, the functioning of the decoding system was described assuming that both switching assemblies 203 and 205 were set to pass-through mode, respectively. As described herein, the decoder system can also decode signals that are not predictively coded. For this use, the second switching assembly 205 is in sum-and-difference mode. Mode) is set, as shown in the figure, the selector device 208 is set to the lower position and the signal is input directly from the source point between the TNS filter 204 and the second switching assembly 205 to the inverse conversion 209. It has become like. For correct decoding, the signal properly has L / R format at the source point. Therefore, when decoding an unpredicted coded stereo signal, a second switching assembly is used to always provide the correct mid (ie downmix) signal for the real-imaginary conversion (eg, rather than more concisely with the left signal). It is preferable to set 205 to the sum / difference mode. As described above, the predictive coding is replaced by conventional direct coding or simultaneous coding of multiple frames, for example based on data rate vs. sound quality determination. The result of such a determination is sent from the encoder to the decoder in various ways, for example by the value of a dedicated indicator bit in each frame, or by the presence or absence of a predictor coefficient value. If these facts are proved, the role of the first switching assembly 203 can be easily realized. In fact, in unpredictable coding mode, the decoder system can process both direct (L / R) stereo coded signals and simultaneous (M / S) coded signals. By operating the first switching assembly 203 in either pass-through mode or sum / difference mode, it is possible to ensure that the source point is always provided with the directly encoded signal. Obviously, the switching assembly 203 converts an M / S format input signal into an L / R format output signal (supplied to an optional TNS filter 204) when functioning in the sum / difference stage.
0055The decoder system receives a signal indicating whether the decoder system decodes a time frame in predictive coding mode or non-predictive coding mode. The unpredicted mode is signaled by the value of the dedicated indicator bit in each frame, or by the presence or absence (or value of zero) of the prediction factor. Prediction modes can be signaled as well. A particularly advantageous embodiment allows fallback without overhead, but utilizes a reserved fourth value in the 2-bit field ms_mask_present (MPEG-2 AAC, see ISO / IEC 13818-7 document). It is sent every time frame and is specified as follows, as follows:<tables num="1"><img id="000010" he="56" wi="159" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>By redefining the value 11 to mean "complex predictive coding", the decoder can operate and relate in all legacy modes, especially in M / S and L / R coding modes, without compromising the bit rate. It can receive a signal indicating the complex-predicted coding mode of the frame.
0056FIG. 4 shows a decoder system with a general configuration, similar to that shown in FIG. 2, but with at least two different configurations. First, the system of FIG. 4 includes switches 404, 411 that allow the application of processing steps, including frequency domain correction, upstream and / or downstream of the upmix stage. It is, on the other hand, downstream of the dequantization module 401 and the first switching assembly 402, and upstream of the second switching assembly 405 located just upstream of the upmix stages 406, 407, 408, 409. It is realized by a set of frequency domain modifiers 403 (drawn as a TNS synthesis filter in this figure) provided with a first switch 404. The decoder system, on the other hand, has a second set of frequency domain modifiers 410, located downstream of upmix stages 406, 407, 408, 409 and upstream of inverse conversion stage 412, provided with a second switch 411. Including. Advantageously, as shown in the figure, each frequency domain modifier is connected upstream to the input side of the frequency domain modifier and downstream in parallel with the pass-through line connected to the associated switch. With this configuration, signal data is always supplied to the frequency domain modifier, enabling processing in the frequency domain based on more time frames as well as the current time frame. The determination of whether to apply the first set of frequency domain modifiers 403 or the second set of frequency domain modifiers 410 is made by the encoder (sent in a bitstream) or predictive coding is applied. It may be based on the frequency or other criteria suitable for the practical situation. As an example, when the frequency domain modifier is a TNS filter, the first set 403 is advantageous for use on certain types of signals, while the second set 410 is advantageous for use on other types of signals. If the result of this selection is encoded in a bitstream, the decoder system activates each set of TNS filters accordingly.
0057To facilitate the understanding of the decoder system shown in Figure 4, the direct (L / R) coded signal decoding is α = 0 (pseudo L / R and L / R are the same). The side channel and the residual channel are not different), the first switching assembly 402 is in pass mode, the second switching assembly is in sum / difference mode, and the first in the upmix stage. 2 Performed when the signal is in M / S format between the switching assemble 405 and the sum / difference step 409. At this time, since the upmix stage is a step that effectively passes, the first set of frequency domain modifiers or the second set of frequency domain modifiers is activated (using each switch 404, 411). It doesn't matter.
0058FIG. 3 shows a decoder system according to an embodiment of the invention that represents a different approach to the supply of MDST data required for upmixing in relation to the decoder systems of FIGS. 2 and 4. Similar to the decoder system described above, the system of FIG. 3 has an inverse quantization module 301, a first switching assembly 302 capable of operating in pass-through mode or sum-difference mode, and a TNS (synthesis) filter 303. All of these are arranged in series from the input end of the decoder system. Modules downstream of this point are selectively utilized by the two second switches 305, 310. It is preferable that these second switches operate simultaneously so that both are in the upper or lower position, as illustrated. At the output end of the decoder system, there is a sum / difference stage 312, and just upstream of it are two inverse MDCT modules 306, 311 that convert the MDCT area display of each channel to the time domain display.
0059In complex predictive decoding, the decoder system is fed a downmix / residual stereo signal and a bitstream encoded with complex predictive coefficients, the first switching assembly 302 is set to pass-through mode, and the second switches 305, 310 are up. Set to position. Downstream of the TNS filter, the two channels of the stereo signal (inverse-quantized, TNS-filtered MDCT) are treated differently. The downmix channel, on the one hand, is supplied to the multiplier and adder 308. The multiplier and adder 308 are the real part α of the prediction coefficient.<sub>R</sub>The MDCT display of the downmix channel weighted by is added to the MDCT display of the residual channel. On the other hand, it is supplied to one of multiple MDCT transform modules, 306. The time domain display of the downmix channel M is the output from the inverse MDCT conversion module 306 and is supplied to both the final sum / difference stage 312 and the MDST conversion module 307. The double use of the time domain display of the downmix channel in this way is advantageous from the viewpoint of computational complexity. The MDST display of the downmix channel thus obtained is supplied to yet another multiplier and adder 309. The multiplier and the adder 309 are the imaginary part α of the prediction coefficient.<sub>I</sub>After weighting with, this signal is added to the linear combination output from adder 308. Therefore, the output of the adder 309 is the side channel signal S = Re {αM} + D. Similarly, the multipliers and adders 308 and 309 are coupled to the decoder system shown in FIG. 2 and input MDCT and MDST displays of the downmix signal, MDCT display of the residual signal, and complex prediction coefficient values. Form a weighted multiplier signal adder. In the present embodiment, downstream of this point, only the path through the inverse MDCT transform module 311 remains before the side channel signal is supplied to the final sum / difference step 312.
0060The synchrony required in the decoder system can be achieved by making the conversion length and window shape applied in both inverse MDCT conversion modules 306, 311 the same. It has already been put to practical use in frequency-selective M / S and L / R coding. Combining an embodiment of the inverse MDCT module 306 with an embodiment of the MDST module 307 results in a delay of one frame. Therefore, five optional delay blocks 313 (or software instructions that exert this effect in the case of computer implementation) are provided, and the part of the system on the right side of the dashed line is, if necessary, on the left side. On the other hand, it can be delayed by one frame. Obviously, there are delay blocks at all intersections between the dashed line and the connection line, with the exception of the connection between the inverse MDCT module 306 and the MDST transform module 307, where there is a compensatory delay. ..
0061The calculation of MDST data in one time frame requires data from one frame in the time domain display. However, for inverse MDCT transforms, one frame (current frame), two consecutive frames (preferably the previous frame and the current frame), or three consecutive frames (preferably the previous frame, the current frame, And the later frame). Due to the well-known Time Domain Alias Cancellation (TDAC) associated with MDCT, the 3-frame option provides complete overlap of input frames, and is most (and in some cases completely) accurate, at least in frames that contain time domain aliases. Is. Obviously, the 3-frame inverse MDCT operates with a 1-frame delay. By allowing the use of an approximate time domain display as input to the MDST transformation, this delay can be avoided, thereby avoiding the need to compensate for delays between different parts of the decoder system. With the two-frame option, overlap / add enabling TDAC is done in the first half of the frame, and the alias exists only in the second half. With the 1-frame option, there is no TDAC, so aliases occur throughout the frame. However, the MDST display thus realized and used as a daytime signal in complex predictive coding can provide sufficient quality.
0062The decoding system shown in Figure 3 can also operate in two unpredictable decoding modes. To directly decode the L / R encoded stereo signal, the second switches 305, 310 are set to the lower position and the first switching assembly 302 is set to pass-through mode. Thus, this signal is in L / R format upstream of the sum / difference step 304. The sum / difference step 304 converts this into the M / S format. Inverse MDCT conversion and final sum / difference calculation are performed on this M / S format. To decode the stereo signal provided in the simultaneous M / S encoding format, the first switching assembly 302 is set to sum / difference mode and the signal is generated between the first switching assembly 302 and the sum / difference step 304. Make it L / R format. The L / R format is more suitable than the M / S format from the viewpoint of TNS filtering. The processing downstream of the sum / difference step 304 is the same as in the case of direct L / R decoding.
0063FIG. 14 (14A-14C) is a three block diagram showing a decoder according to an embodiment of the present invention. Unlike the other block diagrams attached to this application, the connecting line in FIG. 14 shows a multi-channel signal. Specifically, such connecting lines are configured to transmit stereo signals with left / right, mid / side, downmix / residual, pseudo left / pseudo right channels and other combinations.
0064FIG. 14A shows a decoder system that decodes the frequency domain representation of an input signal (shown as an MDCT representation for the purposes of this figure). The decoder system is configured to provide a time domain representation of the stereo signal as its output. This display is generated based on the input signal. The decoder system is provided with an upmix step 1410 to allow decoding of input signals encoded by complex predictive stereo coding. However, input signals encoded in other formats and, in some cases, switching between multiple encoding formats over time, are directly left / right encoded into a sequence of time frames encoded, for example, by complex predictive encoding. It is also possible to process an input signal followed by a time portion encoded by. The function of the decoder system that processes different coding formats is realized by providing a connection line (pass-through) in parallel with the upmix stage 1410. The switch 1411 provides a decoder module further downstream that either outputs from the upmix stage 1410 (lower switch position in the figure) or the unprocessed signal obtained by the connection line (upper switch position in the figure). You can choose to supply to. In this embodiment, the inverted M DCT module 1412 is located downstream of the switch. The MDCT module 1412 converts the MDCT display of the signal into a time domain display. As an example, the signal supplied to the upmix stage 1410 may be a downmix / residual format stereo signal. Next, the upmix stage 1410 is applied to obtain the side signal and perform the sum / difference calculation so that the left / right stereo signal is output (in the MDCT region).
0065FIG. 14B shows a decoder system similar to that shown in FIG. 14A. The system is configured to receive a bitstream as an input signal. The bitstream is initially processed by the coupled demultiplexer and dequantization module 1420. This coupled demultiplexer and dequantization module 1420 is multi-channel for further processing as a first output signal, as determined by the position of switch 1422, which performs the same function as switch 1411 shown in FIG. 14A. Provides an MDCT display of stereo signals. More precisely, the switch 1422 processes the first output from the demultiplexer and dequantization by the upmix stage 1421 and the inverse MDCT module 1423 (lower position) or by the inverse MDCT module 1423 only. (Upper position) Determine. The combined demultiplexer and dequantization module 1420 also output control information. In this case, the control information associated with the stereo signal is data that indicates whether the upper or lower position of switch 1422 is suitable for decoding the signal, or more abstractly, to which coding format the stereo signal is decoded. including. The control information also includes parameters that adjust the characteristics of the upmix stage, such as, for example, the value of the complex prediction coefficient α used in the complex prediction coding, as described above.
0066FIG. 14C has first and second frequency domain correction devices 1431 and 1435 located upstream and downstream of upmix stage 1433, respectively, in addition to entities similar to those shown in FIG. 14B. For the purposes of this drawing, each frequency domain correction device is illustrated by a TNS filter. However, the term frequency domain correction device can also be understood as a process other than TNS filtering that can be applied before and after the upmix phase. Examples of frequency domain correction include prediction, noise addition, bandwidth expansion, and non-linear processing. In some cases, for psychoacoustic considerations and similar reasons, including the characteristics of the signal being processed and / or the setting of such frequency domain correction devices, the frequency domain correction is not downstream of upmix stage 1433, but upstream thereof. It is more advantageous to apply in. In other cases, from the same consideration, the position downstream of the frequency domain correction is preferably upstream. Switches 1432, 1436 selectively activate the frequency domain correction devices 1431 and 1435 so that the decoder system can select the desired configuration, depending on the control information. As an example, FIG. 14C shows that the stereo signal from the coupled demultiplexer and dequantization module 1430 is first processed by the first frequency domain correction device 1431 without passing through the second frequency domain adjustment device 1435. , Then the configuration is shown which is fed to the upmix stage 1433 and finally transferred directly to the reverse MDCT module 1437. As described in the summary section of the invention, this configuration is preferred over the option of performing TNS after upmix in complex predictive coding.
0067II. Encoder system The encoder system according to the present invention will be described with reference to FIG. FIG. 5 is a block diagram showing an encoder system that encodes a left / right (L / R) stereo signal as an output bitstream by complex predictive encoding. The encoder system receives a time domain or frequency domain indication of the signal and supplies it to both the downmix stage and the prediction factor estimator. The real and imaginary parts of the prediction coefficients are supplied to the downmix stage to control the downmixing of the left and right channels and the conversion to residual channels. The downmix and residual channels are then fed to the final multiplexer MUX. If the signal is not supplied to the encoder as a frequency domain display, it is converted to such display at the downmix stage or multiplexer.
0068One of the principles of predictive coding is to convert the left / right signal to mid / side format, ie<maths num="9"><img id="000011" he="18" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>Then use the correlation that remains between these channels, ie<maths num="10"><img id="000012" he="13" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>And set. Here, α is the complex prediction coefficient to be determined, and D is the residual signal. Α can be selected to minimize the energy D = S-Re {αM} of the residual signal. Energy minimization is provided by instantaneous power, short-term energy, or long-term energy (power average). This is optimized in the sense of mean square in the case of discrete signals.
0069Real part of prediction coefficient and α<sub>R</sub>And imaginary part α<sub>I</sub>Is quantized and / or coded jointly. However, preferably, the real and imaginary parts are quantized independently and uniformly, generally with a step size of 0.1 (dimensionless number). According to the MPEG standard, the resolution of the frequency band used for the complex predictor coefficient does not have to be the same as the resolution of the scale factor band (sfb, i.e. a group of MDCT lines using the same MDCT quantization step size and quantization range). In particular, the frequency band resolution of the prediction factor is psychoacoustically valid, such as the Bark scale. As a caveat, when the conversion length changes, the resolution of the frequency band changes.
0070As mentioned above, the encoder system according to the present invention has a degree of freedom in applying predictive stereo coding. The latter case suggests a fallback to L / R or M / S coding. Such a decision can be made on a time frame or finer basis, or on a frequency band basis within a time frame. As mentioned above, the negative result of the decision is sent to the decoding entity in various ways, for example by the value of the dedicated indicator bit in each frame, or by the presence or absence (or zero value) of the predictor coefficient value. Positive decisions are sent as well. A particularly advantageous embodiment allows fallback without overhead, but utilizes a reserved fourth value in the 2-bit field ms_mask_present (MPEG-2 AAC, see ISO / IEC 131818-7 documentation). It is sent every time frame and is specified as follows, as follows:<tables num="2"><img id="000013" he="56" wi="159" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>By redefining the value 11 to mean "complex predictive coding", the encoder can operate in all legacy modes, especially in M / S and L / R coding modes, without compromising the bit rate, which is advantageous. If so, a signal indicating the signal complex predictive coding of the frame can be received.
0071Substantial decisions may be based on the data rate vs. sound quality principle. Data obtained using the psychoacoustic model included in the encoder may be used as a measure of sound quality (as is often the case with available MDCT-based audio encoders). Specifically, some encoder embodiments make rate distortion optimization selection of the prediction coefficient. Therefore, in such an embodiment, if the increase in the prediction gain does not save enough bits for coding the residual signal and the use of the bits required for coding the prediction coefficient cannot be justified, then the prediction coefficient is imaginary. The part, and in some cases the real part, is also set to zero.
0072An encoder embodiment encodes TNS-related information into a bitstream. Such information includes the values of the TNS parameters used by the TNS (synthesis) filter on the decoder side. When using the same set of TNS parameters on both channels, it is more economical to include signaling bits to indicate that the parameters are the same, rather than sending the two sets of parameters separately. For example, it may also include information on whether to apply TNS before or after the upmix phase, based on two optional psychoacoustic assessments.
0073Yet another optional feature, which is potentially beneficial in terms of complexity and bitrate, is that the encoder provides individually limited bandwidth for coding the residual signal. Configured to use. Frequency bands above this limit are not sent to the decoder and are set to zero. In some cases, the energy content of the highest frequency band is so small that it is already zero when quantized. Normal practice (see the MPEG standard max_sfb parameter) requires the use of the same bandwidth limits for both the downmix signal and the residual signal. Here, the inventor has empirically found that the residual signal has significantly more energy content localized to the low frequency band than the downmix signal. Therefore, by imposing a dedicated bandwidth upper limit on the residual signal, it is possible to reduce the bit rate without significantly impairing the sound quality. For example, this is achieved by transmitting with two independent max_sfb parameters, one for the downmix signal and one for the residual signal.
0074It should be pointed out that with reference to the decoder system shown in Figure 5, optimizations such as prediction factors, quantization and its coding, fallback to M / S or L / R mode, TNS filtering, and bandwidth limits are optimal. Although the problem of determination has been described, the same is equally applicable to the embodiments disclosed in the embodiments described with reference to subsequent drawings.
0075FIG. 6 shows another encoder system according to the invention configured to perform complex predictive stereo encoding. The system receives a time domain representation of a stereo signal, including the left and right channels, as input, divided into consecutive, and possibly overlapping, time frames. The sum / difference step 601 converts this signal into a mid channel and a side channel. Mid-channels are supplied to both MDCT module 602 and MDST module 603, and side channels are supplied only to MDCT module 604. Prediction Factor Waterdrop 605 estimates the value of the complex prediction factor for each time frame and, in some cases, for individual frequency bands within the frame, as described above. The coefficient value α is supplied as a weight to the weight adders 606 and 607. The weighting adders 606 and 607 form a residual signal D as a linear combination of the MDCT and MDST displays of the mid signal and the MDCT display of the side signals. Preferably, the complex prediction factor is supplied to the weighted adders 606, 607 represented by the same quantization scheme used when it is encoded into the bitstream. This clearly provides a more faithful reconstruction, as both encoders and decoders use the same prediction factor values. The residual signal, the mid signal (which is more appropriate to call the downmix signal when it appears in combination with the residual signal), and the prediction coefficients are supplied to the coupled quantization and multiplexer step 608. The coupled quantization and multiplexer step 608 encodes these signals and, in some cases, yet additional information as an output bitstream.
0076FIG. 7 is a modification of the encoder shown in FIG. As can be seen from the similar symbols in the figure, the encoder configuration shown in FIG. 7 is similar, but with the addition of the ability to operate directly in L / R coded fallback mode. The encoder system is activated between the complex predictive coding mode and the fallback mode by a switch 710 located just upstream of the coupled quantization and multiplexer step 709. When the switch 710 is in the upper position, the encoder operates in fallback mode. The mid-side signal is supplied to the sum / difference stage 705 from a point just downstream of the MDCT modules 702,704. The sum / difference stage 705 converts the signal into a left / right signal and then sends it to the switch 710. Switch 710 connects the signal to the coupled quantization and multiplexer step 709.
0077FIG. 8 is a diagram showing an encoder system according to the present invention. Unlike the encoder systems shown in FIGS. 6 and 7, this embodiment obtains the MDST data required for complex predictive coding directly from the MDCT data, that is, by real / imaginary conversion in the frequency domain. The real-imaginary transformation applies one of the approaches described for the decoder systems in Figures 2 and 4. It is important that the decoder calculation method matches the encoder calculation method so that faithful decoding can be performed. The same real / imaginary conversion method is used on the encoder side and the decoder side. Regarding the embodiment of the decoder, the part A having the real / imaginary conversion 804 surrounded by the broken line can be replaced by a modification similar to this or by reducing the input time frame used. Similarly, any of the approximation approaches described above can be used to simplify coding.
0078At a high level, the encoder system of FIG. 8 has a configuration different from the configuration that would be obtained by simply replacing the MDST module of FIG. 7 with a (properly connected) real-imaginary module. This architecture is clean, robust and computationally economical, and can provide the ability to switch between predictive coding and direct L / R coding. The input stereo signal is input to the MDCT conversion module 801 and the MDCT conversion module 801 outputs the frequency domain display of each channel. This is sent to both the final switch 808, which activates the encoder system between predictive coding mode and direct coding mode, and the sum / difference step 802. In direct L / R coding, or simultaneous M / S coding performed in a time frame with the prediction factor α set to zero, this embodiment only MDCT transforms, quantizes, and multiplexes the input signal. The latter two steps are performed by a coupled quantization and multiplexer step 807 located at the output end of the system to provide a bitstream. In predictive coding, each channel is further processed between the sum / difference step 802 and the switch 808. The real-imaginary conversion 804 obtains MDST data from the MDCT display of the mid signal and sends it to both the prediction coefficient estimator 803 and the weight adder 806. Similar to the encoder systems shown in FIGS. 6 and 7, another weighting adder 805 is used to combine the side signal with the weighted MDCT and MDST display of the mid signal to form a residual channel signal. The residual channel signal is encoded by the coupled quantization and multiplexer step 807 along with the mid (ie downmix) channel signal and prediction factor.
0079It is illustrated here that each embodiment of the encoder system can be combined with one or more TNS (analytical) filters with reference to FIG. As mentioned above, it is often advantageous to apply TNS filtering to downmixed signals. Thus, as shown in FIG. 9, adaptation of the encoder system of FIG. 7 to include TNS is accomplished by adding a TNS filter 911 just upstream of the coupled quantization and multiplexer step 909.
0080Instead of the right / residual TNS filter 911b, two TNS filters (not shown) configured to handle the right or residual channels may be provided just upstream of the portion of switch 910. In this way, each of the two TNS filters is always supplied with a signal for each channel, allowing TNS filtering based on more time frames than the current frame. As mentioned above, the TNS filter is an example of a frequency domain correction device, and is particularly a device based on processing more frames than the current time frame. This benefits from such an arrangement as much as or better than a TNS filter.
0081As another alternative to the embodiment shown in FIG. 9, a TNS filter for selective activation can be configured with one or more points for each channel. This is similar to the configuration of the decoder system shown in FIG. 4, where different sets of TNS filters can be connected by switches. This allows you to select the most suitable stage for TNS filtering for each time frame. In particular, it is advantageous to switch between different TNS locations with respect to switching between complex predictive stereo coding mode and other coding modes.
0082FIG. 11 shows a modified example based on the encoder system of FIG. 8 in which the second frequency domain display of the downmix signal is obtained by the real / imaginary conversion 1105. Similar to the decoder system shown in Figure 4, this encoder system also includes a frequency domain modifier module that can be selectively activated, one of which 1102 is located upstream of the downmix stage and one of 1109. It is provided downstream of it. The frequency domain modules 1102, 1109, illustrated by the TNS filter in this figure, can be connected to each signal path using four switches 1103a, 1103b, 1109a and 1109b.
0083III. Non-device embodiment Embodiments of the third and fourth aspects of the present invention are shown in FIGS. 15 and 16. Figure 15 shows how to decode a bitstream into a stereo signal and has the following steps: 1. Enter the bitstream. 2. Dequantize the bitstream, thereby obtaining the first frequency domain representation of the downmix channel and residual channel of the stereo signal. 3. Calculate the second frequency domain display of the downmix channel. 4. Calculate the side channel signal based on the three frequency domain indications of the channel. 5. Calculate the stereo signal, preferably in left / right format, based on the side and downmix channels. 6. Output the stereo signal obtained in this way. Steps 3 to 5 may be considered as an upmixing process. Each of steps 1 through 6 is similar to the corresponding function of any of the decoder systems disclosed in the previous section of this document, and implementation details can be read from that section.
0084Figure 16 shows how to encode a stereo signal into a bitstream signal and has the following steps: 1. Input a stereo signal. 2. Convert the stereo signal to the first frequency domain display. 3. Determine the complex prediction factor. 4. Downmix the frequency domain display. 5. Encode the downmix and residual channels as a bitstream with complex prediction coefficients. 6. Output a bitstream. Steps 1 through 5 are similar to the corresponding functions of any of the encoder systems disclosed in the previous section of this document, and implementation details can be read from that section.
0085Both methods can be expressed as computer-readable instructions in the form of software programs and can be executed on a computer. The scope of protection of the present invention extends to such software and computer program products for distributing such software.
0086IV. Experimental evaluation The embodiments disclosed herein were evaluated experimentally. The most important parts of the experimental material obtained in this process are summarized below.
0087The embodiments used in the experiment have the following characteristics: (i) Each MDST spectrum (in a time frame) was calculated from the current, previous, and next MDCT spectra by two-dimensional finite impulse response filtering. (ii) An acoustic psychoacoustic model from the USAC stereo encoder was used. (iii) PS parameters Instead of ICC, CLD and IPD, the real and imaginary parts of the complex prediction factor α transmitted the transmission. The real part and the imaginary part are processed separately, [-3.0, It is limited to the range of 3.0] and is quantized using a step size of 0.1. Time derivative coding and finally Huffman coding using the USAC scale factor codebook. The prediction factor was updated every 1 scale factor band, and the frequency resolution became similar to that of MPEG surround (see, for example, ISO / IEC 23003-1). This quantization and coding scheme resulted in an average bit rate of about 2 kb / s for stereoside information in a typical configuration with a target bit rate of 96 kb / s. (iv) Since the 2-bit ms_mask_present bitstream element has only three possible values, the bitstream format was modified without breaking the current USAC bitstream. By using a fourth value that indicates complex prediction, we allowed a basic mid / side coding fallback mode without wasting bits (see the previous subsection of this disclosure for this). I want).
0088It was played back with headphones, and a listening test was performed by the MUSHRA method using 8 test items with a sampling rate of 48 kHz. Each test enrolled 3, 5 or 6 subjects.
0089We evaluated the impact of different MDST approximations and showed the practical complexity vs. sound quality trade-off between these options. The results are shown in FIGS. 12 and 13. The former shows the absolute score obtained, and the latter shows the difference score for 96s USAC cplf, i.e. for MDCT region unified stereo coding by complex prediction using the current MDCT frame to calculate the approximation of MDST. It can be seen that the sound quality gain achieved by MDCT-based unified stereo coding increases when a computationally more complex approach is used to compute the MDST spectrum. Considering the average of the whole test, the single frame-based system 96s USAC cplf greatly increases the coding efficiency compared to the conventional stereo coding. Similarly, in the case of 96s USAC cp3f, i.e. MDCT region unified stereo coding with complex prediction using the current, previous, and next MDCT frames to calculate the MDST, even better results are obtained.
0090V. Embodiment Further, the present invention can be carried out as follows.
0091A decoder system that decodes a bitstream signal into a stereo signal by complex predictive stereo coding: An inverse quantization step (202,401) that provides a first frequency domain representation of the downmix signal (M) and residual signal (D) based on the bit stream, where each frequency domain representation is the first in multidimensional space. It has a first spectral component representing the spectral content of the corresponding signal represented by the subspace of, the first spectral component is the conversion coefficient arranged in the time frame of the conversion coefficient, and each block is in the time domain. The inverse quantization stage generated by applying the transformation of the signal to the time domain; and It is configured to generate the stereo signal based on the downmix signal and the residual signal located downstream of the dequantization step: A module (206; 408) that calculates a second frequency domain display of the downmix signal based on the first frequency domain display of the downmix signal, wherein the second frequency domain display is the first. The module has a second spectral component representing the spectral content of the signal represented in the second subspace of the multidimensional space, which includes a part of the multidimensional space that is not included in the subspace of. , The first intermediate component is obtained from the first spectral component; at least a part of the impulse response constitutes the coupling of the first spectral component to obtain the second intermediate component; and the second intermediate component is obtained. A module configured to obtain the second spectral component from The side signal is displayed based on the first and second frequency domain display of the downmix signal encoded by the bitstream signal, the first frequency domain display of the residual signal, and the complex prediction coefficient (α). Weighted adders to calculate (210,211; 406,407); and It has an upmix stage (206,207,210,211,406,407,408,409) having a sum / difference stage (207; 409) for calculating the stereo signal based on the first frequency domain representation of the downmix signal and the side signal.
0092Further, the present invention can be carried out as follows. That is, it is a decoder system that decodes a bitstream signal into a stereo signal by complex predictive stereo coding: An inverse quantization step (301) that provides a first frequency domain representation of the downmix signal (M) and the residual signal (D) based on the bitstream signal, each of which is the first frequency domain representation. An inverse quantization step with a first spectral component representing the spectral content of the corresponding signal represented by the first subspace of the multidimensional space; and It is configured to generate the stereo signal based on the downmix signal and the residual signal located downstream of the dequantization step: A module (306,307) that calculates the second frequency domain display of the downmix signal based on the first frequency domain display of the downmix signal, and the second frequency domain display is in the first subspace. Of the downmix signal of the first subspace of the multidimensional space having the spectral content of the signal represented by the second subspace of the multidimensional space including a portion of the multidimensional space that is not included. Inverse conversion step (306) to calculate the time domain display of the downmix signal based on the first frequency domain display; and calculate the second frequency domain display of the downmix signal based on the time domain display of the signal. A module with a conversion stage (307); The side signal is displayed based on the first and second frequency domain display of the downmix signal encoded by the bitstream signal, the first frequency domain display of the residual signal, and the complex prediction coefficient (α). Weighted adders to calculate (308,309); and It has an upmix stage (306,307,308,309,312) having a sum / difference stage (312) for calculating the stereo signal based on the first frequency domain representation of the downmix signal and the side signal.
0093Moreover, the present invention can be carried out as follows. A module that computes the second frequency domain representation of a downmix signal in a decoder system with the features described in the independent decoder system claims: An inverse conversion step (306); and calculating the time domain representation of the downmix signal and / or saturation signal based on the first frequency domain representation of each signal in the first subspace of the multidimensional space. It has a conversion step (307) that calculates a second frequency domain representation of each signal based on the time domain representation of the signal. Preferably, the inverse transform step (306) performs an inverse modified discrete cosine transform, and the transform step performs a modified discrete cosine transform.
0094In the above decoder system, the stereo signal may be represented in the time domain, and the decoder system may further include: Switching arranged between the inverse quantization step and the upmix step, which can function as either (a) a pass-through step used for simultaneous stereo coding; or (b) a sum / difference step used for direct stereo coding. Assembly (302); Further inverse conversion steps (311) arranged in the upmix step to calculate the time domain representation of the side signal; (a) Further sum / difference stages (304) connected to points downstream of the switching assembly (302) and upstream of the upmix stage; or (b) with the downmix signal obtained from the switching assembly (302). A selector device (305,310) located upstream of the inverse conversion step (306,301) configured to be selectively connected to any of the side signals obtained from the weighted adder (308,309).
0095VI. Conclusion Further embodiments of the present invention will become apparent to those skilled in the art upon reading the above description. Although the present specification and drawings disclose embodiments and examples, the present invention is not limited to these specific examples. Numerous modifications and modifications can be made without departing from the scope of the present invention as defined in the appended claims.
0096It should be noted that the methods and devices disclosed in this application can be applied to the coding of signals with more than 2 channels with appropriate modifications within the capabilities of those skilled in the art, including routine experiments. It should be emphasized that the signals, parameters, and matrices described in connection with the embodiments described may be frequency variable or frequency invariant and / or time variable or time invariant. The computational steps described can be performed on a frequency-by-frequency basis or for all frequencies at once, and all entities can be performed to have frequency-selective behavior. For the purposes of the application, any quantization scheme can be adapted by the psychoacoustics model. Furthermore, as a point to keep in mind, various sum / difference conversions, that is, conversion from downmix / residual format to pseudo L / R format, L / R-to-M / S conversion, and M / S-to-L / R All conversions are in the following format<maths num="11"><img id="000014" he="18" wi="123" file="JP6062467B2_D0001.tif" img-format="tif" img-content="drawing" /></maths>And only the gain factor g changes. Therefore, the coding gain can be corrected by appropriately selecting the decoding gain by adjusting the gain factor individually. Moreover, as will be apparent to those skilled in the art, an even number of series-arranged differences / difference transformations affect the pass-through stage, and in some cases the gain is not 1.
0097The systems and methods disclosed herein can be implemented as software, firmware, hardware or a combination thereof. Some or all components can be implemented as software executed by a digital signal processor or microprocessor, or as hardware or purpose-built integrated circuits. Such software can be distributed on computer readable media. Computer-readable media include computer storage media and communication media. As is well known to those of skill in the art, computer storage media are volatile and non-volatile, implemented by any method or technique for storing information such as computer-readable instructions, data structures, program modules and other data. Includes removable and non-removable media. Computer storage media include RAM, ROM, EEPROM, flash memory and other memory technologies, CD-ROMs, digital versatile discs (DVDs) and other optical disc storage media, magnetic cassettes, magnetic tapes, magnetic disk storage and other magnetic storage devices, or Other, but not limited to, any medium that can be used to store the desired information. Moreover, as is known to those skilled in the art, communication media generally embody data in modulated data signals such as computer readable instructions, data structures, program modules, other carrier waves and other transmission mechanisms. And include any information distribution medium. The following notes are added. (Appendix 1) A decoder system that provides a stereo signal by complex predictive stereo coding. The upmix stage is configured to generate the stereo signal based on the first frequency domain display of the downmix signal and the residual signal, each first frequency domain display is the first in multidimensional space. It has an upmix step having a first spectral component representing the spectral content of the corresponding signal represented by the subspace of. A module that calculates the second frequency domain display of the downmix signal based on the first frequency domain display of the downmix signal, and the second frequency domain display is included in the first subspace. A module having a second spectral component representing the spectral content of a signal represented by a second subspace of said multidimensional space, including a portion of the multidimensional space that is not. Weighting to calculate the side signal based on the first and second frequency domain indications of the downmix signal encoded in the bitstream signal, the first frequency domain indication of the residual signal, and the complex prediction factor. Adder and It has an upmix stage having a sum / difference stage for calculating the stereo signal based on the first frequency domain display of the downmix signal and the side signal. The upmix stage is further capable of operating in a pass-through mode in which the downmix signal and the residual signal are directly supplied to the sum / difference stage. (Appendix 2) The downmix signal and the residual signal are segmented into time frames, and the upmix stage receives, for each time frame, a 2-bit data field associated with that frame, depending on the value of the data field. And configured to operate in active mode or pass-through mode, The decoder system described in Appendix 1. (Appendix 3) The downmix signal and the residual signal are segmented into time frames. The upmix stage is further configured in the MPEG bitstream to receive an ms_mask_present field associated with each time frame and operate in active mode or pass-through mode, depending on the value of the ms_mask_present field. , The decoder system described in Appendix 1. (Appendix 4) Further having an inverse quantization step located upstream of the upmix step, which provides the first frequency domain representation of the downmix signal and the residual signal based on the bitstream signal. The decoding system described in any one of Appendix 1 to 3. (Appendix 5) The first spectral component has a real value represented by the first subspace. The second spectral component has an imaginary value represented by the second subspace. Optionally, the first spectral component is determined by either the discrete cosine transform DCT or the modified discrete cosine transform MDCT. Optionally, the second spectral component is determined by either the discrete sine transform DST or the modified discrete sine transform MDST. The decoder system according to any one of Appendix 1 to 4. (Appendix 6) With at least one temporal noise shaping TNS module located upstream of the upmix stage, With at least one additional TNS module located downstream of the upmix stage, It has a selector device that selectively activates either (a) the TNS module upstream of the upmix stage or (b) the additional TNS module downstream of the upmix stage. The decoder system according to any one of Appendix 1 to 5. (Appendix 7) The downmix signal is partitioned into consecutive time frames, and each time frame is related to the value of the complex prediction coefficient. The module that calculates the second frequency domain representation of the downmix signal deactivates itself according to the absolute value of the imaginary part of the complex prediction coefficient being less than a predetermined tolerance in the time frame. Configured not to produce output for time frames, The decoder system described in Appendix 5. (Appendix 8) The downmix signal time frame is further partitioned into frequency bands, and each frequency band is accompanied by the value of the complex prediction coefficient. The module that calculates the second frequency domain representation of the downmix signal deactivates itself depending on the absolute value of the imaginary part of the complex prediction coefficient being less than a predetermined tolerance in the frequency band of the time frame. It was configured not to generate an output for that frequency band. The decoder system described in Appendix 7. (Appendix 9) The first spectral component is a conversion coefficient arranged in the time frame of the conversion coefficient, and each block is generated by applying the conversion of the time domain signal to the time segment. The module that calculates the second frequency domain display of the downmix signal obtains the first intermediate component from the first spectral component, and at least a part of the impulse response constitutes the coupling of the first spectral component. The second intermediate component is obtained, and the second spectral component is obtained from the second intermediate component. The decoder system according to any one of Appendix 1 to 8. (Appendix 10) Part of the impulse response is based on the frequency response characteristics of the conversion. Optionally, the frequency response characteristics of the conversion depend on the characteristics of the analytical window function applied to the conversion of the signal into time segments. The decoder system described in Appendix 9. (Appendix 11) The module that calculates the second frequency domain display of the downmix signal is (a) Simultaneous time frame of the first spectral component, (b) Simultaneous and previous time frames of the first spectral component, and (c) Each time frame of the second spectral component is configured to be determined based on one of the simultaneous, previous, and subsequent time frames of the first spectral component. Decoder system according to Appendix 9 or 10. (Appendix 12) The module that calculates the second frequency domain representation of the downmix signal is an approximation determined by the combination of at least two temporally adjacent and / or frequency adjacent first spectral components. It is configured to compute an approximate second spectral representation with a typical second spectral component, The decoder system according to any one of Appendix 1 to 11. (Appendix 13) The stereo signal is represented in the time domain, and the decoder system further Between the dequantization step and the upmix step, which can function as either (a) pass-through step or (b) sum / difference step, thereby switching between direct and simultaneously coded stereo input signals. With the placed switching assembly, An inverse conversion step configured to calculate the time domain representation of the stereo signal, It is located upstream of the inverse conversion step and is (a) at a point downstream of the upmix stage where the stereo signal obtained by complex prediction is supplied to the inverse conversion step, or (b) a direct stereo code. A selector device configured to selectively connect to any of the points downstream of the switching assembly and upstream of the upmix stage to which the stereo signal obtained by the conversion is supplied to the inverse conversion stage. Have, The decoder system according to any one of Appendix 1 to 12. (Appendix 14) An encoder system that encodes a stereo signal using complex prediction as a signal having a downmix channel, a residual channel, and a complex prediction coefficient. An estimator that estimates complex prediction coefficients and The stereo that (a) converts the stereo signal into a frequency domain display of a downmix signal and a residual signal having a relationship determined by the value of the complex prediction coefficient, and (b) operates and encodes as a pass-through step. An encoder system with a coding step that can operate to feed the signal directly to the multiplexer. (Appendix 15) Complex predictive stereo encoding is configured to encode a stereo signal with a bitstream signal, and further. Further having a multiplexer that receives the output from the coding step and the estimator and encodes it with the bitstream signal. The encoder system described in Appendix 14. (Appendix 16) The estimator determines the complex prediction coefficient by minimizing the power of the residual signal over time or the average power of the residual signal. The encoder system described in Appendix 14 or 15. (Appendix 17) The stereo signal has a downmix channel and a side channel. The coding step is configured to receive a first frequency domain representation of the stereo signal, the first frequency domain representation of the spectral content of the corresponding signal represented by the first subspace of the multidimensional space. Has a first spectral component to represent The coding step further A module that calculates the second frequency domain display of the downmix channel based on the first frequency domain display of the downmix signal, and the second frequency domain display is included in the first subspace. A module having a second spectral component representing the spectral content of the signal represented by the second subspace of the multidimensional space, including a portion of the multidimensional space that is not. It has a first and second frequency domain display of the downmix channel, a first frequency domain display of the side channel, and a weighted adder that calculates a residual signal based on the complex prediction coefficient. , The estimator receives the downmix channel and the side channel and predicts the complex to minimize the power of the residual signal over a period of time or to minimize the average power of the residual signal. Determine the coefficient, The encoder system according to any one of Appendix 14 to 16. (Appendix 18) The coding step is A sum / difference step of converting the stereo signal into a simultaneously encoded stereo signal having a downmix channel and a side channel, and A conversion step that provides an oversampled frequency domain display of the downmix channel and a critically sampled frequency domain display of the side channel, wherein the oversampled frequency domain display preferably contains complex spectral components. Have a conversion stage and It has a weighted adder that calculates a residual signal based on the oversampled frequency domain display of the downmix channel, the critical sampled frequency domain display of the side channel, and the complex prediction coefficient. And The estimator receives the residual signal and determines the complex prediction factor in order to minimize the power of the residual signal or to minimize the average power of the residual signal. Preferably, the transform step comprises a modified discrete cosine transform MDCT step arranged in parallel with the modified discrete cosine transform MDST step, which together provides the oversampled frequency domain representation of the downmix channel. The encoder system according to any one of Appendix 14 to 16. (Appendix 19) A decoding method that provides a stereo signal by complex predictive stereo coding. The step of receiving the first frequency domain representation of the downmix signal and the residual signal, each of which is the spectral content of the corresponding signal represented by the first subspace of the multidimensional space. With a first spectral component representing The step of receiving the control signal and Depending on the value of the control signal (a) A step of upmixing the downmix signal and the residual signal using the upmix step to obtain the stereo signal. A sub-step of calculating a second frequency domain display of the downmix signal based on the first frequency domain display of the downmix signal, wherein the second frequency domain display is in the first subspace. A substep having a second spectral component representing the spectral content of the signal represented by the second subspace of the multidimensional space, including a portion of the multidimensional space that is not included. A sub that calculates the side signal based on the first and second frequency domain display of the downmix signal encoded in the bitstream signal, the first frequency domain display of the residual signal, and the complex prediction coefficient. Steps and A step having a sub-step for calculating the stereo signal by applying the first frequency domain display of the downmix signal and the side signal to the sum / difference conversion. (b) A decoding method having a step of interrupting the upmixing step. (Appendix 20) The first spectral component has a real value represented by the first subspace. The second spectral component has an imaginary value represented by the second subspace. Optionally, the first spectral component is determined by either the discrete cosine transform DCT or the modified discrete cosine transform MDCT. Optionally, the second spectral component is determined by either the discrete sine transform DST or the modified discrete sine transform MDST. The decoding method described in Appendix 19. (Appendix 21) The downmix signal is partitioned into consecutive time frames, and each time frame is related to the value of the complex prediction coefficient. The step of calculating the second frequency domain representation of the downmix signal is interrupted as the absolute value of the imaginary part of the complex prediction coefficient is less than a predetermined tolerance of the time frame, with respect to that time frame. To avoid producing output, The decoding method described in Appendix 20. (Appendix 22) The downmix signal time frame is further partitioned into frequency bands, and each frequency band is accompanied by the value of the complex prediction coefficient. The step of calculating the second frequency domain representation of the downmix signal is interrupted as the absolute value of the imaginary part of the complex prediction coefficient is less than a predetermined tolerance in the frequency band of the time frame, and its frequency. Avoid producing output for the band, The decoding method described in Appendix 21. (Appendix 23) The first spectral component is a conversion coefficient arranged in the time frame of the conversion coefficient, and each block is generated by applying the conversion of the time domain signal to the time segment. The step of calculating the second frequency domain display of the downmix signal is A sub-step for obtaining the first intermediate component from the first spectral component, and A sub-step of constructing the combination of the first spectral components by at least a part of the impulse response to obtain the second intermediate component, and It has a sub-step for obtaining the second spectral component from the second intermediate component. The decoding method described in Appendix 20. (Appendix 24) A part of the impulse response is based on the frequency response characteristics of the conversion. Optionally, the frequency response characteristics of the conversion depend on the characteristics of the analytical window function applied to the conversion of the signal into time segments. The decoding method described in Appendix 23. (Appendix 25) The steps for calculating the second frequency domain representation are (a) simultaneous time frames of the first spectral component, (b) simultaneous and previous time frames of the first spectral component, and (c) th. Obtaining each time frame of the second spectral component, using one of the simultaneous, previous, and subsequent time frames of one spectral component as input. The decoding method described in Appendix 24. (Appendix 26) The step of calculating the second frequency domain representation of the downmix signal is an approximation determined by a combination of at least two temporally adjacent and / or frequency adjacent first spectral components. Includes the step of calculating an approximate second spectral representation with a typical second spectral component, The decoding method according to any one of Appendix 19 to 25. (Appendix 27) The stereo signal is represented in the time domain, and the method further comprises Depending on whether the bitstream signal is encoded directly by stereo coding or by simultaneous stereo coding, a step of omitting the upmixing step and It has a step of inversely converting the bitstream signal to obtain the stereo signal. The decoding method according to any one of Appendix 19 to 26. (Appendix 28) A step of transmitting the time domain display of the downmix signal and a step of calculating the side signal according to the fact that the bitstream is encoded by direct stereo coding or simultaneous stereo coding. Steps to omit and It further comprises a step of inversely converting the frequency domain display of each channel encoded by the bitstream signal to obtain the stereo signal. The decoding method described in Appendix 27. (Appendix 29) An encoding method that encodes a stereo signal by a bitstream by complex predictive stereo coding. Steps to determine the complex prediction factor and The step of converting the stereo signal to display the first frequency domain of the downmix signal and the residual signal having a relationship determined by the complex prediction coefficient, and the first frequency domain display is in a multidimensional space. A step and a step having a first spectral component representing the spectral content of the corresponding signal represented by the first subspace. An encoding method comprising the step of encoding the downmix channel, the residual channel, and the complex prediction coefficient as the bitstream. (Appendix 30) The step of determining the complex prediction coefficient is performed in order to minimize the power of the residual signal or the average power of the residual signal over a period of time. The encoding method described in Appendix 29. (Appendix 31) A step of defining or recognizing a partition of the stereo signal in a time frame, For each time segment, there is an additional step to encode or select the stereo signal in this time segment with at least one of the direct stereo coding, simultaneous stereo coding, and complex predictive stereo coding options. And If direct stereo coding is selected, the stereo signal is converted to a frequency domain display of the left and right channels and encoded as the bitstream. If simultaneous stereo coding is selected, the stereo signal is converted to a frequency domain display of downmix channels and side channels and encoded as the bitstream. The encoding method described in Appendix 29 or 30. (Appendix 32) The option that provides the highest sound quality is selected according to the given psychoacoustic model. The encoding method described in Appendix 31. (Appendix 33) Further having a step of defining or recognizing the partition of the stereo signal in a time frame The stereo signal has a downmix channel and a side channel. The step of converting the stereo signal into the first frequency domain display of the downmix channel and the residual channel is A sub-step of calculating a second frequency domain display of the downmix signal based on the first frequency domain display of the downmix channel, wherein the second frequency domain display is in the first subspace. A substep having a second spectral component representing the spectral content of the signal represented by the second subspace of the multidimensional space, including a portion of the multidimensional space that is not included. A step of forming a residual signal based on the first and second frequency domain display of the downmix channel, the first frequency domain display of the side channel, and the complex prediction coefficient. The step of determining the complex prediction coefficient is performed for one hour frame at a time by minimizing the average power of the residual signal in each time frame. The encoding method described in Appendix 29 or 30. (Appendix 34) A step of converting the stereo signal into a simultaneously encoded stereo signal having a downmix channel and a side channel, A step of converting the downmix channel into an oversampled frequency domain display, preferably having complex spectral components. A step of converting the side channel into a critically sampled, preferably real-valued frequency domain display. It further comprises a step of calculating a residual signal based on the oversampled frequency domain display of the downmix channel, the critical sampled frequency domain display of the side channel, and the complex prediction factor. , The determination of the complex prediction factor is made by feedback control on the residual signal thus calculated in order to minimize power or average power. The encoding method described in any one of Appendix 29 to 33. (Appendix 35) The conversion of the downmix channel to the oversampled frequency domain display is performed by applying MDCT and MDST and concatenating their outputs. The encoding method described in Appendix 34. (Appendix 36) A computer program product having a computer-readable medium containing instructions for performing the method according to any one of Appendix 19 to 35 when executed by a general purpose computer.
32 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| WO2009141775A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP2005521921A | Cites | Japan |
315 members in 18 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 61322458 | United States of America | – | |
| 32245810 | United States of America | P |
Members315
| Document | Office | Kind | |
|---|---|---|---|
| CA2793140A1 | Canada | A1 | |
| CA2793317A1 | Canada | A1 | |
| CA2793320A1 | Canada | A1 | |
| CA2921437A1 | Canada | A1 | |
| CA2924315A1 | Canada | A1 | |
| CA2988745A1 | Canada | A1 | |
| CA2992917A1 | Canada | A1 | |
| CA3040779A1 | Canada | A1 | |
| CA3045686A1 | Canada | A1 | |
| CA3076786A1 | Canada | A1 | |
| CA3097372A1 | Canada | A1 | |
| CA3105050A1 | Canada | A1 | |
| CA3110542A1 | Canada | A1 | |
| CA3125378A1 | Canada | A1 | |
| CA3185301A1 | Canada | A1 | |
| WO2011124608A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2011124616A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2011124621A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2011237869A1 | Australia | A1 | |
| AU2011237877A1 | Australia | A1 | |
| AU2011237882A1 | Australia | A1 | |
| SG184167A1 | Singapore | A1 | |
| MX2012011528A | Mexico | A | |
| MX2012011530A | Mexico | A | |
| MX2012011532A | Mexico | A | |
| IL222294A0 | Israel | A0 | |
| IL222294D0 | Israel | D0 | |
| CN102884570A | China | A | |
| KR20130007646A | Republic of Korea | A | |
| KR20130007647A | Republic of Korea | A | |
| US2013028426A1 | United States of America | A1 | |
| US2013030817A1 | United States of America | A1 | |
| EP2556502A1 | European Patent Office (EPO) | A1 | |
| EP2556503A1 | European Patent Office (EPO) | A1 | |
| EP2556504A1 | European Patent Office (EPO) | A1 | |
| KR20130018854A | Republic of Korea | A | |
| CN102947880A | China | A | |
| CN103119647A | China | A | |
| JP2013524281A | Japan | A | |
| JP2013525829A | Japan | A | |
| JP2013525830A | Japan | A | |
| US2013266145A1 | United States of America | A1 | |
| AU2011237869B2 | Australia | B2 | |
| KR20140042927A | Republic of Korea | A | |
| KR20140042928A | Republic of Korea | A | |
| RU2012143501A | Russian Federation | A | |
| RU2012144366A | Russian Federation | A | |
| RU2012147499A | Russian Federation | A | |
| AU2011237882B2 | Australia | B2 | |
| AU2011237877B2 | Australia | B2 | |
| RU2525431C2 | Russian Federation | C2 | |
| KR101437896B1 | Republic of Korea | B1 | |
| CN102947880B | China | B | |
| KR101437899B1 | Republic of Korea | B1 | |
| JP2015099403A | Japan | A | |
| SG10201502597QA | Singapore | A | |
| CN102884570B | China | B | |
| RU2554844C2 | Russian Federation | C2 | |
| US9111530B2 | United States of America | B2 | |
| CN103119647B | China | B | |
| CN104851426A | China | A | |
| CN104851427A | China | A | |
| RU2559899C2 | Russian Federation | C2 | |
| KR20150113208A | Republic of Korea | A | |
| US9159326B2 | United States of America | B2 | |
| CN105023578A | China | A | |
| JP5813094B2 | Japan | B2 | |
| JP5814340B2 | Japan | B2 | |
| JP5814341B2 | Japan | B2 | |
| US2015380001A1 | United States of America | A1 | |
| KR101586198B1 | Republic of Korea | B1 | |
| JP2016026317A | Japan | A | |
| JP2016026318A | Japan | A | |
| CA2793140C | Canada | C | |
| BR112012025878A2 | Brazil | A2 | |
| US9378745B2 | United States of America | B2 | |
| IL221911A | Israel | A | |
| IL221962A | Israel | A | |
| IL245338A0 | Israel | A0 | |
| IL245338D0 | Israel | D0 | |
| IL245444A0 | Israel | A0 | |
| IL245444D0 | Israel | D0 | |
| CA2793320C | Canada | C | |
| US2016329057A1 | United States of America | A1 | |
| JP6062467B2This record | Japan | B2 | |
| KR101698438B1 | Republic of Korea | B1 | |
| KR101698439B1 | Republic of Korea | B1 | |
| KR101698442B1 | Republic of Korea | B1 | |
| KR20170010079A | Republic of Korea | A | |
| IL222294A | Israel | A | |
| JP2017062504A | Japan | A | |
| IL250687A0 | Israel | A0 | |
| IL250687D0 | Israel | D0 | |
| BR112012025863A2 | Brazil | A2 | |
| BR112012025868A2 | Brazil | A2 | |
| IL245444A | Israel | A | |
| US9761233B2 | United States of America | B2 | |
| JP6197011B2 | Japan | B2 | |
| JP6203799B2 | Japan | B2 | |
| IL253522A0 | Israel | A0 |
15 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6062467
- Application
- 40746
Titles2
- Japanese
- MDCTベース複素予測ステレオ符号化
- English
- MDCT-based complex predictive stereo coding
Classification
- CPC, 14
- G10L19/008
- G01L19/00
- G10L25/12
- G10L19/0212
- G10L19/18
- G06F3/162
- G10L19/06
- G10L19/167
- H04S3/008
- H04S2400/01
- G10L19/002
- G10L19/022
- G10L19/012
- G10L19/03
- IPC, 2
- G10L19 02
- G10L19 008
