Apparatus and method for generating an enhanced signal using independent noise-filling
17 claims: 5 independent, 12 dependent
- 1入力信号(600)から強化された信号を生成する装置であって、前記強化された信号は、強化スペクトル領域についてのスペクトル値を有し、前記強化スペクトル領域についての前記スペクトル値は、前記入力信号(600)に含まれておらず、前記装置は、前記入力信号(600)のソース領域を前記強化スペクトル領域におけるターゲット領域にマップするためのマッパー(602)であって、前記ターゲット領域についてのソース領域識別が存在し、前記マッパー(602)は前記ソース領域識別を使用して前記ソース領域を選択し、前記選択されたソース領域を前記ターゲット領域にマップするように構成される、マッパー(602)と、前記入力信号(600)の前記ソース領域におけるノイズ充填領域(302)について第1のノイズ値を生成するように構成され、かつ、前記強化スペクトル領域内の前記ターゲット領域内のノイズ領域についての第2のノイズ値を生成するように構成され、ここで前記第2のノイズ値は、前記第1のノイズ値から非相関され、または、前記ターゲット領域におけるノイズ領域に対する第2のノイズ値を生成するように構成され、ここで前記第2のノイズ値は、前記入力信号(600)の前記ソース領域における第1のノイズ値から非相関される、ノイズ・フィラー(604)と、を含む、装置。
- 2前記入力信号(600)は、前記入力信号(600)の前記ソース領域についてのノイズ充填パラメータを含む符号化された信号であり、前記ノイズ・フィラー(604)は、前記ノイズ充填パラメータを使用して前記第1のノイズ値を生成するように構成され、かつ、前記第1のノイズ値に関するエネルギー情報を使用して前記第2のノイズ値を生成するように構成される、請求項1に記載の装置。
- 3前記ノイズ・フィラー(604)は、前記第1のノイズ値を有する、前記入力信号(600)内の前記ノイズ充填領域(302)を識別し(900)、 ソース・タイル・バッファへ 前記入力信号(600)の少なくとも1つの領域 をコ ピーし(902)、ここで、前記領域は前記ソース領域を含み、前記ソース領域は前記ノイズ充填領域(302)を含み、前記ソース・タイル・バッファにおいて、前記ノイズ充填領域(302)の前記識別(90 0 )において識別された前記第1のノイズ値を、非相関のノイズ値で置換する(904)ように構成され、前記マッパー(602)は、前記非相関のノイズ値を有する前記ソース・タイル・バッファを前記ターゲット領域にマップするように構成される、請求項1に記載の装置。
- 4前記ノイズ・フィラー(604)は、前記非相関のノイズ値に関するエネルギー情報を測定し(1102)、また前記第1のノイズ値に関するエネルギー情報を測定し(1100)、前記非相関のノイズ値に関する前記エネルギー情報と前記第1のノイズ値に関する前記エネルギー情報とから導出された(1104)スケーリング値を使用して前記非相関のノイズ値をスケールする(906,1106)ように構成される、請求項1に記載の装置。
- 5前記ノイズ・フィラー(604)は、前記マッパー(602)の動作に続く前記第2のノイズ値を生成するように構成される、または、前記ノイズ・フィラー(604)は、前記マッパー(602)の動作に続く前記第1および前記第2のノイズ値を生成する(604)ように構成される、請求項1ないし請求項4のいずれかに記載の装置。
- 6前記マッパー(602)は、前記ソース領域を前記ターゲット領域にマップするように構成され、前記ノイズ・フィラー(604)は、ノイズ充填動作とサイド情報として前記入力信号(600)において送信されるノイズ充填パラメータとを使用して前記第1のノイズ値を生成することによって、スペクトル領域におけるノイズ充填を実行するように構成され、かつ、前記第1のノイズ値に関するエネルギー情報を使用して前記第2のノイズ値を生成するために前記ターゲット領域におけるノイズ充填を実行するように構成される、請求項1ないし請求項5のいずれかに記載の装置。
- 7サイド情報として前記入力信号(600)に含まれるスペクトル・エンベロープ情報を使用して前記強化スペクトル領域における前記第2のノイズ値を調整する(1202)ためのエンベロープ調整器を、さらに含む、請求項1ないし請求項6のいずれかに記載の装置。
- 8前記ノイズ・フィラー(604)は、前記入力信号(600)のサイド情報をノイズ充填動作のためにスペクトルの位置を識別するためだけに使用するように構成されるか、または、前記ノイズ・フィラー(604)は、前記ノイズ充填動作のためにスペクトルの位置を識別するために、前記ノイズ充填領域(302)におけるスペクトル値の有無にかかわらず前記入力信号(600)の時間またはスペクトル特性を解析するように構成される、請求項1ないし請求項7のいずれかに記載の装置。
- 9前記ノイズ・フィラー(604)は、前記ソース領域のみにおけるスペクトルの位置に対する入力を有するか、または、前記ソース領域および前記ターゲット領域におけるスペクトルの位置に対する入力を有する識別ベクトル(706)を使用してノイズ位置を識別するように構成される、請求項1ないし請求項7のいずれかに記載の装置。
- 10前記ノイズ・フィラー(604)は、前記識別ベクトル(706)が示すノイズ値に関するエネルギー情報を計算し(1100)、前記ターゲット領域のために挿入されたランダム値に関するエネルギー値を計算し(1102)、前記挿入されたランダム値をスケールするためのゲインファクタを計算し(1104)、前記ゲインファクタを前記挿入されたランダム値に適用する(1010,1106)ように構成される、請求項 9 に記載の装置。
- 11前記ノイズ・フィラー(604)は、 前記入力信号(600)の少なくとも1つの領域を 前記コピーする こと( 902)において、前記入力信号(600)のスペクトル部分全体をコピーするか、もしくは、前記マッパー(602)が使用可能なノイズ充填境界周波数を超える前記入力信号(600)のスペクトル部分全体を前記ソース・タイル・バッファにコピーし、かつ、前記ソース・タイル・バッファ全体に対して前記置 換すること (904)を実行するように構成され、または、前記ノイズ・フィラー(604)は、 前記入力信号(600)の少なくとも1つの領域を 前記コピーする こと( 902)において、前記マッパー(602)が前記ターゲット領域のために使用するソース領域に対して、1つ以上の特定のソース識別子(specific source identifiers)によって識別される前記入力信号(600)のスペクトル領域のみをコピーするように構成され、ここで、各異なる個別のマッピング動作について、個別のソース・タイル・バッファが使用される 、請 求項3に記載の装置。
- 12前記マッパー(602)は、前記ターゲット領域を生成するために、ギャップ充填動作を実行するように構成され、前記装置は、第1のスペクトル部分の第1のセットの第1の復号化された表現を生成するためのスペクトル領域音声デコーダ(112)であって、前記復号化された表現は、第1のスペクトル解像度を有する、スペクトル領域音声デコーダ(112)と、前記第1のスペクトル解像度より低い第2のスペクトル解像度を有する第2のスペクトル部分の第2のセットの第2の復号化された表現を生成するためのパラメトリック・デコーダ(114)と、第1のスペクトル部分と前記第2のスペクトル部分に対するスペクトル・エンベロープ情報とを使用して前記第1のスペクトル解像度を有する再構築された第2のスペクトル部分を再生成するための周波数再生器(116)と、前記再構築された第2のスペクトル部分における前記第1の復号化された表現を時間表現に変換するためのスペクトル時間コンバータ(118)と、を含み、前記マッパー(602)および前記ノイズ・フィラー(604)は、前記周波数再生器(116)において、少なくとも部分的に含まれる、請求項1ないし請求項9のいずれかに記載の装置。
- 13前記スペクトル領域音声デコーダ(112)は、スペクトル値の一連の復号化されたフレームを出力するように構成され、復号化されたフレームは、前記第1の復号化された表現であり、ここで、前記フレームは、スペクトル部分の前記第1のセットについてのスペクトル値と第2のスペクトル部分の前記第2のセットについての0表示とを含み、ここで、復号化するための装置は、前記第1のスペクトル部分の前記第1のセットと前記第2のスペクトル部分の前記第2のセットとについてのスペクトル値を含む再構築されたスペクトル・フレームを得るために、第2のスペクトル部分の前記第2のセットについての前記周波数再生器(116)によって生成されたスペクトル値と再構築バンドにおける第1のスペクトル部分の前記第1のセットのスペクトル値とを結合するためのコンバイナ(208)をさらに含み、前記スペクトル時間コンバータ(118)は、前記再構築されたスペクトル・フレームを前記時間表現に変換するように構成される、請求項12に記載の装置。
- 14入力信号(600)から強化された信号を生成するための方法であって、前記強化された信号は、強化スペクトル領域についてのスペクトル値を有し、前記強化スペクトル領域についての前記スペクトル値は、前記入力信号(600)に含まれておらず、前記方法は、前記入力信号(600)のソース領域を前記強化スペクトル領域におけるターゲット領域にマップするステップ(602)であって、前記ソース領域は、ノイズ充填領域(302)を含み、前記ターゲット領域についてのソース領域識別が存在し、前記マップするステップ(602)は、前記ソース領域識別を使用して前記ソース領域を選択するステップ、および前記選択されたソース領域を前記ターゲット領域にマップするステップを含む、マップするステップ(602)と、前記入力信号の前記ソース領域内の前記ノイズ充填領域(302)についての第1のノイズ値を生成し、かつ、前記ターゲット領域内のノイズ領域についての第2のノイズ値を生成するステップであって、ここで前記第2のノイズ値は、前記第1のノイズ値から非相関される、生成するステップ(604)、または、前記ターゲット領域におけるノイズ領域に対する第2のノイズ値を生成するステップであって、ここで前記第2のノイズ値は、前記入力信号(600)の前記ソース領域における第1のノイズ値から非相関される、生成するステップ(604)と、を含む、方法。
- 15音声信号を処理するシステムであって、前記システムは、前記音声信号から符号化された信号を生成するためのエンコーダと、請求項1ないし請求項13のいずれかに記載の、入力信号(600)から強化された信号を生成するための装置であって、前記符号化された信号は、前記強化された信号を生成するための前記装置への前記入力信号(600)を生成するために、所定の処理(700)に付される、前記装置と、を含むシステム。
- 16音声信号を処理する方法であって、前記方法は、前記音声信号から符号化された信号を生成するステップと、請求項14に記載の、入力信号(600)から強化された信号を生成する方法であって、前記符号化された信号は、前記強化された信号を生成するための方法への前記入力信号(600)を生成するために、所定の処理(700)に付される、前記方法と、を含む、方法。
- 17コンピュータ上で稼動する際に、請求項14または請求項16の方法を実行する、コンピュータ・プログラム。
Independent claims17
139 paragraphs, as filed
This application relates to signal processing and, in particular, to audio signal processing.
Perceptual coding of audio signals for the purpose of reducing the amount of data for efficient storage, or transmission of these signals is widely and practically used. In particular, the coding used when trying to achieve the lowest bit rate often leads to poor voice quality, mainly due to limitations on the encoder side of the transmitted voice signal band. In modern codecs, well-known methods exist for decoder-side signal restoration using the audio signal Band Width Extension (BWE) (eg Spectral Band Replication (SBR)).
In low bit rate coding, so-called noise filling is used. The prominent spectral region quantized to zero due to strict bit rate restrictions is filled with noise synthesized in the decoder.
Usually both techniques are performed simultaneously in the application of low bit rate coding. In addition, there are integrated solutions that simultaneously use speech coding, noise filling and spectral gap filling, such as Intelligent Gap Filling (IGF).
However, all these methods use a signal in which the baseband or core audio signal is reproduced using waveform decoding and noise filling in the first stage and the BWE or IGF processing is immediately reproduced in the second stage. It has a common point that it is executed. This is because the same noise value filled in the baseband by noise filling during reproduction to reproduce the lost portion of the high frequency band (in BWE) or fills the remaining spectral gap (in IGF). Leads to the fact that it is used to do. The use of highly correlated noise to reproduce multiple spectral regions in BWE or IGF can lead to perceptual damage.
Related topics at the cutting edge of the art include:
-SBR as a post-processor for waveform decoding [1-3]
. AAC PNS [4]
. MPEG-D USAC noise filling [5]
. G.719 and G.722.1C [6]
. MPEG-H 3D IGF [8]
The following documents and patent applications describe methods that may be relevant to this application: [1] M. Dietz, L. Liljeryd, K. Kjoerling and O. Kunz, "Spectral Band Replication, a novel approach in audio coding. , "in 112th AES Convention, Munich, Germany, 2002. [2] S. Meltzer, R. Boehm and F. Henn," SBR enhanced audio codecs for digital broadcasting such as "Digital Radio Mondiale" (DRM), "in 112th AES Convention, Munich, Germany, 2002. [3] T. Ziegler, A. Ehret, P. Ekstrand and M. Lutzky, "Enhancing mp3 with SBR: Features and Capabilities of the new mp3PRO Algorithm," in 112th AES Convention, Munich, Germany, 2002. [4] J. Herre, D. Schulz, Extending the MPEG-4 AAC Codec by Perceptual Noise Substitution, Audio Engineering Society 104th Convention, Preprint 4720, Amsterdam, Netherlands, 1998 [5] European Patent Application Publication No. 2304720 USAC Noise Filling [6] ITU-T Recommendations G.719 and G.221C [7] European Patent Application Publication No. 2704142 [8] Publication of European Patent Application No. 13177350
Audio signals processed by these methods will suffer from artificial products such as roughness, modulation distortion and unpleasantly perceived sound quality, especially at low bit rates. As a result, spectral pores in the low bandwidth and / or LF range are generated. The reason is primarily that, as described below, the reproduced components of the extended or gap-filled spectrum are based on one or more direct copies containing noise from the baseband. It is a fact. Temporary modulation resulting from the above unwanted correlations in the reproduced noise can be heard in a noisy way as perceptual roughness or unpleasant distortion. All existing methods like mp3 + SBR, AAC + SBR, USAC, G.719 and G.722.1C, and even MPEG-H 3D IGF are first copied or reflected from the core with spectral data. Perform full core decoding including noise filling before filling the spectral gap or high band.
<p>It is an object of the present invention to provide an improved concept for producing enhanced signals.</p><p>An object of the present invention is an apparatus for generating an enhanced signal according to claim 1, a method for generating an enhanced signal according to claim 13, a coding and decoding system according to claim 14, and a coding and decoding according to claim 15. It is achieved by the method of conversion, or the computer program of claim 16.</p>
<p>The present invention is an enhanced signal generated by bandwidth expansion or advanced gap filling, or an enhanced signal generated by another method having spectral values for the enhanced spectral region not included in the input signal. A significant improvement in the audio quality of the input signal is to generate a first noise value for the noise-filled region in the source spectral region of the input signal, and in the destination or target region, i.e., now the noise value. It is based on the fact that it is obtained by generating a second independent noise value for the noise region in the enhanced region, i.e., a second noise value independent of the first noise value.</p><p>Therefore, the conventional technical problem of having noise dependent on the baseband and the enhanced band due to spectral value mapping is eliminated. And the related challenges of having artificial products such as roughness, modulation distortion and unpleasantly perceived sound quality, especially at low bitrates, are eliminated.</p><p>That is, the noise filling of the second noise value is uncorrelated from the first noise value. That is, make sure that the noise value, which is at least partially independent of the first noise value, no longer produces artificial products or is at least reduced with respect to the prior art. Therefore, prior art processing of noise filling spectral values in the baseband with simple bandwidth expansion or advanced gap filling operations is uncorrelated with noise from the baseband and, for example, only changes the level. .. However, introducing uncorrelated noise values, preferably derived from separate noise processing, in the source band on the one hand and in the target band on the other, provides the best results. However, noise that is not completely uncorrelated, or is not completely independent, and is at least partially uncorrelated with an uncorrelated value of, for example, 0.5 or less when an uncorrelated value of 0 indicates a complete uncorrelation. Even the introduction of values improves the complete correlation problem of the prior art.</p><p>Therefore, the examples describe a combination of waveform decoding, bandwidth expansion, or gap filling and noise filling in a perceptual decoder.</p><p>A further advantage is that, in contrast to existing concepts, the generation of artificial products of signal distortion and perceptual roughness is avoided. Here, artificial products are currently typical for calculating bandwidth expansion or gap filling following waveform decoding and noise filling.</p><p>This is due to the changes that occur in the course of the processing steps described above in some embodiments. Performing bandwidth expansion or gap filling directly after waveform decoding is preferred, and subsequent calculation of noise filling on an already reconstructed signal using uncorrelated noise is preferred. Further preferred.</p><p>In a further embodiment, the waveform decoding and noise filling can be performed in the conventional order, and the noise value is replaced with appropriately scaled uncorrelated noise along the flow of the process. Can be done.</p><p>Therefore, the present invention is on a noise-filled spectrum by moving the noise-filling step to the very end of the processing chain and by using uncorrelated noise for patching or gap filling. Address the challenges caused by copying or mirror image movements.</p><p>Subsequently, preferred embodiments of the present invention are described with respect to the accompanying drawings.</p>
<figref num="1a">FIG. 1a illustrates a device that encodes an audio signal.</figref><figref num="1b">FIG. 1b illustrates a decoder for decoding a coded audio signal consistent with the encoder of FIG. 1a.</figref><figref num="2a">FIG. 2a illustrates a preferred embodiment of the decoder.</figref><figref num="2b">FIG. 2b illustrates a preferred embodiment of the encoder.</figref><figref num="3a">FIG. 3a exemplifies a schematic diagram of the spectrum produced by the spectral region decoder of FIG. 1b.</figref><figref num="3b">Figure 3b exemplifies a table showing the relationship between the scale factor for the scale factor band, the energy for the reconstructed band, and the noise filling information for the noise filling band.</figref><figref num="4a">FIG. 4a illustrates the function of a spectral region encoder to apply the spectral portion selection to the first and second sets of spectral moieties.</figref><figref num="4b">FIG. 4b illustrates an embodiment of the function of FIG. 4a.</figref><figref num="5a">FIG. 5a illustrates the function of the MDCT encoder.</figref><figref num="5b">Figure 5b illustrates the functionality of a decoder with MDCT technology.</figref><figref num="5c">FIG. 5c illustrates an embodiment of a frequency regenerator.</figref><figref num="6">FIG. 6 illustrates a block diagram of a device that produces an enhanced signal consistent with the present invention.</figref><figref num="7">FIG. 7 illustrates an independent noise-filled signal flow guided by selection information in a decoder consistent with an embodiment of the invention.</figref><figref num="8">FIG. 8 illustrates a signal flow of independent noise filling performed by swapped order of gap filling or bandwidth expansion and noise filling of the decoder.</figref><figref num="9">FIG. 9 illustrates a flow chart of procedures consistent with further embodiments of the present invention.</figref><figref num="10">FIG. 10 illustrates a flow chart of procedures consistent with further embodiments of the present invention.</figref><figref num="11">FIG. 11 illustrates a flowchart for explaining the scaling of random values.</figref><figref num="12">FIG. 12 illustrates a flow chart exemplifying the application of the present invention to a general bandwidth expansion or gap filling procedure.</figref><figref num="13a">Figure 13a illustrates an encoder with the calculation of bandwidth expansion parameters.</figref><figref num="13b">FIG. 13b exemplifies a decoder with bandwidth expansion performed as a postprocessor rather than an integrated procedure as in Figure 1a or 1b.</figref>
FIG. 6 illustrates a device that produces an enhanced signal, such as an audio signal, from an input signal that can also be an audio signal. The enhanced signal has a spectral value for the enhanced spectral region. Here, the spectral value for the enhanced spectral region is not included in the first input signal at the input 600, which is the input signal. The device comprises a mapper 602 for mapping the source spectral region of the input signal to the target region of the enhanced spectral region. Here, the source spectrum region includes a noise filling region.
Furthermore, the device is configured to generate a first noise value for the noise-filled region of the source spectral region of the input signal, and a second noise value for the noise region of the target region. It comprises a noise filler 604 configured to produce. Here, the second noise value, that is, the noise value in the target region is independent, uncorrelated, or uncorrelated from the first noise value in the noise filling region.
One embodiment relates to the following situation. In that situation, noise filling is actually performed in the baseband. That is, in that situation, the noise value in the source region is generated by noise filling. In a further variant, it is presumed that noise filling in the source region was not performed. Nevertheless, the source region has a noise region that is actually filled with noise, such as spectral values schematically encoded as spectral values by the source or core encoder. Mapping this noise to the enhanced region, such as the source region, also produces dependent noise in the source and target regions. To address this issue, the noise filler only fills the target area of the mapper with the noise. That is, noise filler (noise) filler) produces a second noise value for the noise region in the target region. Here, the second noise value is uncorrelated from the first noise value in the source region. This substitution or noise filling can occur in the source tile buffer or in the target itself. The noise region can be identified by the classifier by analyzing the source region or by analyzing the target region.
To achieve this goal, it will be described with reference to Figure 3A. FIG. 3A illustrates the scale factor band 301 of the input signal as the fill region. Then, the noise filler generates a first noise spectrum value in the noise filling band 301 in the decoding operation of the input signal.
Further, this noise filling band 301 is mapped to the target area. That is, the generated noise value is mapped to the target area so as to be consistent with the prior art. And, therefore, the target region depends on or correlates with the noise associated with the source region.
According to the invention, however, the noise filler 604 of FIG. 6 produces a second noise value for the destination or the noise region in the target region. Here, the second noise value is uncorrelated, uncorrelated, or independent of the first noise value in the noise filling band 301 of FIG. 3A.
Typically, the mapper for mapping the noise fill and the source spectral region to the destination region is within the integrated gap fill, as schematically illustrated from FIGS. 1A to 5C. It may be included in the range of the high frequency player. Alternatively, it can be implemented as a post-processor as illustrated in FIG. 13B and as a corresponding encoder in FIG. 13A.
Typically, the input signal follows an additional predetermined decoder process 700, which means that the input signal of FIG. 6 is obtained at the dequantization 700 or at the output of the block 700. As a result, the input to the core coder noise filler block 704 is the input 600 of FIG. The mapper of FIG. 6 corresponds to the gap filling or bandwidth expansion block 602, and the independent noise filling block 702 is also included within the noise filler 604 of FIG. Thus, blocks 704 and 702 are both noise fillers of FIG. Filler) is also included in block 604, and block 704 produces a so-called first noise value for the noise region in the noise filling region, and block 702 is the destination or noise region in the target region. To generate a second noise value. It is then derived from the noise-filled region in the baseband on the side of the bandwidth expansion performed by the mapper or gap filling or bandwidth expansion block 602. Further, as will be described later, the independent noise filling operation performed by the block 702 is controlled by the control vector PHI exemplified by the control line 706.
1. Step: Noise Identification In the first step, all spectral lines representing noise in the transmitted audio frame are identified. The identification process may be controlled by existing, transmitted information about the noise location used for noise filling [4] [5], or may be identified as an additional classifier. The result of the noise line identification is a vector containing 0s and 1s where the position of 1 indicates the spectral line representing the noise.
In mathematical conditions, this procedure can be expressed as:
<img file="JP6992024B2_D0001.tif" />
<img file="JP6992024B2_D0002.tif" />
<img file="JP6992024B2_D0003.tif" />
<img file="JP6992024B2_D0004.tif" />
2. Step: Independent noise In the second stage, a specific region of the transmitted spectrum is selected and copied to the source tile. Within this source tile, the identified noise is replaced with random noise. The energy of the inserted random noise is matched to the same energy of the original noise in the source tile.
In mathematical conditions, this procedure can be expressed as:
<img file="JP6992024B2_D0005.tif" />
<img file="JP6992024B2_D0006.tif" />
<img file="JP6992024B2_D0007.tif" />
<img file="JP6992024B2_D0008.tif" />
<img file="JP6992024B2_D0009.tif" />
<img file="JP6992024B2_D0010.tif" />
<img file="JP6992024B2_D0011.tif" />
<img file="JP6992024B2_D0012.tif" />
FIG. 8 illustrates an embodiment that follows any post-processing such as decoding of the spectral region exemplified in block 112 of FIG. 1B. Alternatively, in the post-processor embodiment exemplified by block 1326 of FIG. 13B, the input signal first follows gap filling or bandwidth expansion. That is, the input signal first follows the mapping operation, and then independent noise filling is then performed, i.e., within the entire spectrum.
<img file="JP6992024B2_D0013.tif" />
<img file="JP6992024B2_D0014.tif" />
<img file="JP6992024B2_D0015.tif" />
<img file="JP6992024B2_D0016.tif" />
<img file="JP6992024B2_D0017.tif" />
<img file="JP6992024B2_D0018.tif" />
The independent noise filling of the present invention can also be used in a stereo channel pair environment. Therefore, the encoder calculates the appropriate channel pair representation, L / R or M / S for each frequency band and any prediction factor. The decoder applies the above independent noise filling to a well-chosen representation of the channel prior to the next calculation of the final conversion of all frequencies towards the L / R representation.
The present invention is applicable or suitable for all voice applications where full bandwidth is not available or gap filling is used to fill the spectral holes. The present invention can discover usage patterns in, for example, digital radio, internet streaming and voice communication applications, for example in the distribution or broadcasting of voice content.
Next, examples of the present invention will be described with respect to FIG. 9-12. In step 900, the noise region is identified in the source range. This procedure, the procedure previously discussed for "noise identification", can rely entirely on the noise-filled side information received from the encoder side, or without spectral values for the enhanced spectral region, ie. , Without spectral values for this enhanced spectral region, can also be configured to be alternative or additively dependent on the signal analysis of the already generated input signal.
Then, in step 902, the source range already desired for direct noise filling, which is well known in the art, i.e., the complete source range, is copied to the source tile buffer.
Then, in step 904, the first noise value, that is, the direct noise value generated in the noise filling region of the input signal, is replaced in the source tile buffer by a random value. Then, in step 906, these random values are scaled in the source tile buffer to get a second noise value for the target area. Then, in step 908, the mapping operation is performed, i.e., their contents of the source tile buffers available after steps 904 and 906 are mapped to the destination range. In this way, an independent noise filling operation for the source and target ranges was obtained by the replacement operation 904 and following the mapping operation 908.
FIG. 10 illustrates a further embodiment of the present invention. Also, in step 900, the noise in the source range is identified. However; the function of this step 900 is different from the function of step 900 in FIG. This is because step 900 of FIG. 9 can affect the input signal spectrum for which the noise value has already been received, i.e. the noise filling operation has already been performed.
However, in FIG. 10, no noise filling operation was performed on the input signal, and the input signal does not yet have any noise value in the noise filling region at the input in step 902. At step 902, the source range is mapped to a destination or target range. Here, the noise filling value is not included in the source range.
In this way, noise identification in the source range of step 900 is by identifying the 0 spectral value of the signal with respect to the noise filling region and / or by using this noise filling side information from the input signal. Can be executed. That is, the encoder side generated noise filling information. Then, in step 904, the noise filling information and, in particular, the energy information that identifies the energy brought to the input signal on the decoder side are read.
Then, as illustrated in step 1006, noise filling in the source range is performed, and then, or in parallel, step 1008 is performed. That is, the random value is inserted at a position in the destination range identified by step 900 across the entire band or identified using the baseband or input signal information along with the mapping information. And the mapping information is information about which (plural) source range is mapped to which (plural) target range.
Finally, the inserted random values are scaled to obtain the second independent, uncorrelated or uncorrelated noise values.
Next, FIG. 11 is provided to illustrate further information regarding the scaling of noise fill values in the enhanced spectral region. That is, how a second noise value can be obtained from a random value is described.
At step 1100, energy information about noise in the source range is obtained. The energy information is then determined from random values, i.e., from the values generated by random or pseudo-random processing, as illustrated in step 1102. Furthermore, step 1104 describes how to calculate the scale factor, i.e., using energy information about noise in the source range, and using energy information about random values. Then, in step 1106, the random value, that is, the random value on which the energy was calculated in step 1102, is multiplied by the scale factor generated by step 1104. Therefore, the procedure exemplified in FIG. 11 corresponds to the calculation of the scale factor g previously shown in the examples. However, all these calculations can be performed in the logarithmic domain or in any other domain, and multiplication step 1106 can be replaced by addition or subtraction in the logarithmic range.
Further references are made in FIG. 12 to illustrate the embedding of the present invention within the scope of common intelligent gap filling or bandwidth expansion schemes. At step 1200, the spectral envelope information is recovered from the input signal. Spectral envelope information can be generated, for example, by the parameter extractor 1306 of FIG. 13A and can be provided by the parameter decoder 1324 of FIG. 13b. The second noise value and other values in the destination range are then scaled using this spectral envelope information as illustrated in 1202. Subsequent further post-processing 1204, in the case of bandwidth expansion, has an increased bandwidth, or has a reduced number, to obtain an enhanced signal in the final time domain, or to know. It can be done in order not to obtain spectral holes in a typical gap filling situation.
It is outlined that in this situation some modifications can be applied, especially for the embodiment of FIG. For embodiments, step 902 is performed on the entire spectrum of the input signal, or at least a portion of the spectrum of the input signal that exceeds the noise-filled boundary frequency. This frequency ensures that no noise filling is performed below this frequency, i.e. below this frequency.
The entire input signal spectrum, or complete potential source range, is then copied to source tile buffer 902, regardless of any particular source range / target range mapping information, and then processed in steps 904 and 906. , And step 908 then selects a particular source area from this source tile buffer that is specifically needed.
In other embodiments, however, only the portion of the input signal may be particularly needed, only the source range is in a single source tile buffer or based on the source range / target range information contained in the input signal. It is copied to several individual source tile buffers, i.e. to the source tile buffer associated with this audio input signal as side information. Only the source range specifically needed is processed by steps 902, 904, 906. Depending on the situation, that is, the second variant, the complexity or minimum memory requirement is always independent of the particular mapping situation. The situation may be reduced compared to the situation where at least the entire source range above the noise-filled boundary frequency is processed by steps 902, 904, 906.
References are then made to FIGS. 1a-5e to illustrate particular embodiments of the invention within the scope of frequency regenerator 116. The frequency regenerator 116 is arranged in front of the spectral time converter 118.
FIG. 1a illustrates a device that encodes an audio signal 99. The audio signal 99 is an input to the time spectrum converter 100 for converting an audio signal having a sampling rate into a spectral representation 101 output by the time spectrum converter. The spectrum 101 is an input to the spectrum analyzer 102 for analyzing the spectral representation 101. The spectrum analyzer 101 determines a first set of first spectral portions 103 to be encoded at the first spectral resolution and a second to be encoded at the second spectral resolution. It is configured to determine a different second set of spectral portions 105. The second spectral resolution is smaller than the first spectral resolution. The second set of second spectral portions 105 are inputs to a parameter calculator or parametric coder 104 for calculating spectral envelope information with a second spectral resolution. Furthermore, the spectral region audio coder 106 is provided to generate a first coded representation 107 of a first set of first spectral portions having a first spectral resolution. Furthermore, the parameter calculator / parametric coder 104 is configured to generate a second coded representation 109 of a second set of second spectral portions. The first coded representation 107 and the second coded representation 109 are inputs to the bitstream multiplexer or bitstream former 108, and block 108 is finally transmitted or Outputs an audio signal encoded for storage on the storage device.
In general, the first spectral portion, such as 306 in FIG. 3a, is surrounded by two second spectral portions, such as 307a, 307b. This is not the case in HE AAC. Here, the core coder frequency range is a limited band.
FIG. 1b exemplifies a decoder that works in tune with the encoder of FIG. 1a. The first coded representation 107 is a spectrum for generating a first decoded representation (a decoded representation with a first spectral resolution) of a first set of first spectral portions. This is an input to the region audio decoder 112. Furthermore, the second coded representation 109 produces a second decoded representation of the second set of second spectral portions having a second spectral resolution lower than the first spectral resolution. Is the input to the parametric decoder 114 for.
The decoder also has a frequency regenerator 116 for reproducing a reconstructed second spectral portion with a first spectral resolution using the first spectral portion. The frequency regenerator 116 performs a tile filling operation, i.e., using a first set of tiles or portions of a first spectral portion and a second to a reconstructed range or reconstructed band having a second spectral portion. Copies this first set of spectral parts of one, and generally by the output of the second representation decoded by the parametric decoder 114, i.e. information about the second set of second spectral parts. It is used to perform spectral envelope formation or other operations as shown. The first set of decoded parts of the first spectrum and the second set of reconstructed parts of the spectrum as shown by the output of frequency regenerator 116 on line 117 are the first decoded. An input to a spectral time converter 118 configured to convert the expressed and the reconstructed second spectral portion into a temporal representation 119 (a temporal representation with a particular high sampling rate).
FIG. 2b illustrates an embodiment of the encoder of FIG. 1a. The audio input signal 99 is an input to the analysis filter bank 220 corresponding to the time spectrum converter 100 of FIG. 1a. Then, the input to the spectral analyzer 102 of FIG. 1a corresponding to the sound mask 226, which is one block of FIG. 2b, is a complete spectral value when the temporal noise forming operation / temporal tile forming operation is not applied. , The residual spectral value when block 222, which is a TNS operation as illustrated in FIG. 2b, is applied. For two-channel or multi-channel signals, joint channel coding 228 can be additionally performed. As a result, the spectral region encoder 106 of FIG. 1a may have a joint channel coding block 228. In addition, an entropy coder 232 is provided to perform lossless data compression. And it is also part of the spectral region encoder 106 of FIG. 1a.
The spectrum analyzer / sound mask 226 outputs the output of the TNS block 222 with the core band and the sound components corresponding to the first set of the first spectral portion 103, and the second spectral portion of FIG. 1a. It is divided into the remaining components corresponding to the second set of 105 and. Block 224, shown as the IGF parameter extraction encoding, corresponds to the parametric coder 104 of FIG. 1a, and the bitstream multiplexer 230 corresponds to the bitstream multiplexer 108 of FIG. 1a.
Preferably, the analytical filter bank 222 is performed as an MDCT (Modified Discrete Cosine Transform Filter Bank), and the MDCT transforms the signal 99 into the time frequency domain with a modified discrete cosine transform that acts as a frequency analysis tool. Used for.
The spectrum analyzer 226 preferably applies a tone mask. This tone mask evaluation stage is used to separate sound components from components such as noise in the signal. This allows the core coder 228 to encode all sound components that have a psychoacoustics module. The tone mask evaluation stage can be performed in a number of different ways, preferably a sinusoidal track evaluation stage used in sine and noise modeling for speech / audio coding [8, 9], or , [10] is implemented in its function as well as the HILN model based on the voice coder. Preferably, embodiments are used that are easy to carry out without the need to maintain a life-and-death orbit, but any other tone or noise detector can be used as well.
<img file="JP6992024B2_D0019.tif" />
<img file="JP6992024B2_D0020.tif" />
In addition, tile selection stabilization techniques have been proposed that remove artificial products in the frequency domain, such as chirping and musical noise.
In the case of a stereo channel pair, additional joint stereo processing is applied. This is necessary. This is because, for a particular destination range, the signal can be highly correlated with the panned sound source. If the source regions selected for this particular region are not well correlated, then the spatial image is due to the uncorrelated source region, even though the energies are matched for the destination region. Can be compromised. While performing cross-correlation of spectral values in general, the encoder analyzes each destination region energy band and, if a certain threshold is exceeded, a joint for this energy band. Set the (joint) flag of. If this joint stereo flag is not set, the left and right channel energy bands are processed individually in the decoder. If the joint stereo flag is set, both energy and patching will be performed in the joint stereo region. In the case of prediction, the joint stereo information for the IGF region is the core code, including a flag indicating whether the direction of the prediction is residual from the downmix or vice versa. Information on a similar joint stereo that has been signaled for conversion.
<img file="JP6992024B2_D0021.tif" />
Another solution is to calculate and conduct energy directly in the joint stereo region for the band in which the joint stereo is operating. As a result, no additional energy change is required on the decoder side.
<img file="JP6992024B2_D0022.tif" />
<img file="JP6992024B2_D0023.tif" />
Joint stereo-> LR conversion:
<img file="JP6992024B2_D0024.tif" />
<img file="JP6992024B2_D0025.tif" />
<img file="JP6992024B2_D0026.tif" />
This process is from the tiles used to regenerate the highly correlated destination region and the panned destination region, even if the source regions are not correlated. The resulting left and right channels still ensure that they represent a correlated and panned sound source. Then, the stereo image is saved for such an area.
In other words, a bitstream sends a joint stereo flag that indicates whether L / R or M / S is used as an example for common joint stereo coding. .. In the decoder, the core signal is first decoded, as indicated by the joint stereo flag for the core band. Second, the core signal is stored in the L / R and M / S representations. For IGF tile filling, the source tile representation is chosen to fit the target tile representation, as indicated by the joint stereo information for the IGF band.
Temporal noise formation (TNS) is part of standard technology and AAC [11-13].
TNS can be thought of as an extension of the perceptual coder's basic scheme. Then, an arbitrary processing step is inserted between the filter bank and the quantization stage. The main task of the TNS module is to hide the generated quantization noise in a temporary masking area of transient phenomena such as signals, and thus it leads to a more efficient coding scheme. First, TNS calculates a set of prediction coefficients using "forward prediction" in the transform area (eg MDCT). These coefficients are then used to flatten the temporal envelope of the signal. Also, when the quantization affects the spectrum passed through the TNS filter, the quantization noise is temporarily flat. By applying inverse TNS filtering on the decoder side, the quantization noise is formed according to the temporal envelope of the TNS filter, so that the quantization noise is masked by a transient phenomenon.
IGF is based on the MDCT representation. For effective coding, preferably a long block of about 20 ms should be used. If the signal in such a long block contains a transient phenomenon, the audible pre- and post-echoes occur in the IGF spectral band due to tiling. Figure 7c shows a typical pre-echo effect prior to temporary initiation by IGF. On the left side, the spectrogram of the original signal is shown, and on the right side, the spectrogram of the bandwidth-extended signal without TNS filtering is shown.
<img file="JP6992024B2_D0027.tif" />
In legacy decoders, spectral patching on the audio signal corrupts the spectral correlation at the patch boundaries, thereby compromising the temporal envelope of the audio signal by leading to dispersion. Therefore, another advantage of performing IGF tile filling on the residual signal is that the tile boundaries are seamlessly correlated after the application of filter formation. And this results in a more faithful temporal reproduction of the signal.
In the encoder of the invention, the spectrum that has undergone TNS / TTS filtering, tone masking and IGF parameter evaluation lacks any signal above the IGF start frequency except for sound components. This sparse spectrum is currently encoded by a core coder that uses the principles of arithmetic and predictive coding. These coded components, along with the signaling bits, form a bitstream of audio.
FIG. 2a illustrates an embodiment of the corresponding decoder. The bitstream of FIG. 2a corresponding to the encoded audio signal is the input to the connected demultiplexer / decoder with respect to FIGS. 1b up to blocks 112 and 114. The bitstream demultiplexer divides the input audio signal into a first encoded representation 107 of FIG. 1b and a second encoded representation 109 of FIG. 1b. The first coded representation with the first set of first spectral portions is the input to the joint channel decoding block 204 corresponding to the spectral region decoder 112 of FIG. 1b. The second coded representation is the input to the parametric decoder 114 not exemplified in FIG. 2a, and then the input to the IGF block 202 corresponding to the frequency regenerator 116 of FIG. 1b. The first set of first spectral portions required for frequency reproduction is the input to the IGF block 202 via line 203. Furthermore, following the joint channel decoding 204, a particular core duplication applies to the sound mask block 206. As a result, the output of the sound mask 206 corresponds to the output of the spectral region decoder 112. Then the combination by combiner 208 is performed. That is, the output of combiner 208 has a full range spectrum, but is still frame building in the region passed through the TNS / TTS filter. Then, in block 210, the reverse TNS / TTS operation is performed using the TNS / TTS filter information provided via line 109. That is, the TTS side information is preferably included in the first coded representation produced by the spectral region encoder 106. The spectral region encoder 106 can be, for example, a direct AAC or USAC core encoder, or can be included in a second coded representation. Sampling rate of the original input signal at the output of block 210 A complete spectrum up to the maximum frequency, which is the frequency in the complete range defined in, is provided. The spectrum / time conversion is then performed in the synthetic filter bank 212 to finally obtain the audio output signal.
FIG. 3a exemplifies a schematic diagram of the spectrum. The spectrum is subdivided in the scale factor band SCB, where there are seven scale factor bands from SCB1 to SCB7, as illustrated in the example of Figure 3a. The scale factor band can be the AAC scale factor band defined in the AAC standard and has a bandwidth increasing to the upper frequency, as schematically shown in FIG. 3a. It is preferable to perform intelligent gap filling rather than from the very beginning of the spectrum. That is, it is preferable to start the IGF operation at a low frequency and at the IGF start frequency exemplified in 309. Therefore, the core frequency band extends from the lowest frequency to the IGF start frequency. Above the IGF start frequency, spectral analysis is low, where the high resolution spectral elements 304, 305, 306, 307 (the first set of first spectral parts) are represented by the second set of second spectral parts. Applied to separate from resolution elements. FIG. 3a illustrates a spectrum that is a schematic input to a spectral region encoder 106 or a joint channel coder 228. That is, the core encoder operates over the full range, but encodes a significant number of 0-spectral values. That is, these 0 spectral values are quantized to 0, set to 0 before quantization, or occur after quantization. In any case, the core encoder works in full range. That is, as the spectrum is illustrated, i.e., the core decoder does not necessarily have to be aware of any intellectual gap filling, but with a lower spectral resolution, a second set of second spectral portions. There is no need to encode.
Preferably, high resolution is defined by linear coding of spectral lines, such as MDCT lines. On the other hand, the second resolution or low resolution is defined, for example, by calculating only a single spectral value per scale factor band. Here, the scale factor band covers several frequency lines. Thus, the second low resolution is defined with respect to its spectral resolution by the linear coding typically applied by core encoders such as AAC or USAC core encoders. Much lower than the resolution.
<img file="JP6992024B2_D0028.tif" />
In particular, when the core encoder is under low bitrate conditions, additional in the core band, i.e. at frequencies below the IGF start frequency, i.e. in the scale factor band from SCB1 to SCB3. The noise filling operation can be applied in addition. In noise filling, there are multiple adjacent spectral lines quantized to zero. On the decoder side, these quantized to 0 spectral values are resynthesized, and the resynthesized spectral values are the NFs exemplified in 308 in Figure 3b.<sub>2</sub>It is adapted on those scales using noise filling energies such as. The noise-filling energies, which can be given in absolute conditions, or as in USAC, especially in relative conditions with respect to scale factors, correspond to the energies of a set of spectral values quantized to zero. These noise-filled spectral lines are the spectral values from the source range and the energy information E.<sub>1</sub>, E<sub>2</sub>, E<sub>3</sub>, E<sub>4</sub>A third spectrum reproduced by direct noise-filled synthesis without any IGF operation that relies on frequency reproduction using frequency tiles from other frequencies to reconstruct the frequency tiles using and It can also be considered to be the third set of parts.
Preferably, the band from which the energy information is calculated coincides with the scale factor band. In other embodiments, energy information value grouping is applied such that only a single energy information value is transmitted, eg, for scale factor bands 4 and 5, but in this embodiment. Even in, the boundaries of the grouped reconstructed bands coincide with the boundaries of the scale factor band. If different band separations are applied, a particular recalculation or synchronous calculation may be applied, and this can make sense depending on the particular embodiment.
Preferably, the spectral region encoder 106 of FIG. 1a is an acoustically psychologically driven encoder, as illustrated in FIG. 4a. In general, being a voice signal encoded after being transformed into a spectral range (401 in Figure 4a), as exemplified in the MPEG2 / 4 AAC standard or the MPEG1 / 2, Layer 3 standard, is a scale factor. Sent to computer 400. The scale factor calculator additionally receives that it is a quantized audio signal, or MPEG1 / 2 Layer 3 or MPEG. As is the case with the AAC standard, it is controlled by an acoustic psychological model that receives a complex spectral representation of the audio signal. The psychoacoustics model calculates, for each scale factor band, a scale factor that represents the psychoacoustics threshold. In addition, by the cooperation of well-known internal and external iterative loops, or by any other suitable coordinated coding procedure, the scale factor is, and as a result, certain bit rate conditions. Fully filled. Then, on the one hand, it is a quantized spectral value, and on the other hand, the calculated scale factor is the input to the quantizer processor 404. In direct voice encoder operation, being a quantized spectral value is weighted by a scale factor, and the weighted spectral value is then a fixed quantizer that has compression capabilities generally up to the amplitude range above. Is the input to. Then, at the output of the quantizer processor, a specific and efficient code for a set of zero quantization indexes for adjacent frequency values, or, as is also called in the art, a "run" of zero values. There is exactly the quantization index then sent into the entropy encoder that generally has the quantization.
In the audio encoder of Figure 1a, however, the quantizer processor generally receives information about the second spectral portion from the spectral analyzer. Thus, in the quantizer processor 404, at the output of the quantizer processor 404, the second spectral portion identified by the spectral analyzer 102 is 0, or a zero-valued "run" is particularly high in the spectrum. Make sure you have the representation recognized by the encoder or decoder as a zero representation that can be encoded very efficiently when present.
FIG. 4b illustrates an embodiment of a quantizer processor. MDCT spectral values can be inputs to a set of 0 blocks 410. Then the second spectral portion is already set to 0 before the weighting by the scale factor of block 412 is performed. In an additional embodiment, block 410 is not provided, but the set to 0 cooperation is performed in block 418 following weighted block 412. In a further embodiment, the set to 0 action can also be performed in the set to 0 block 422 following the quantization in quantizer block 420. In this embodiment, blocks 410 and 418 are absent. Generally, at least one of blocks 410, 418, 422 is provided according to a particular embodiment.
Then, at the output of block 422, the quantized spectrum is obtained corresponding to that illustrated in FIG. 3a. This quantized spectrum is then the input to an entropy coder, such as 232 in Figure 2b, which can be a Huffman coder or an arithmetic coder, for example as defined in the USAC standard.
The set to 0 blocks 410, 418, 422, which are provided alternately or in parallel, respectively, is controlled by the spectrum analyzer 424. The spectrum analyzer preferably comprises any embodiment of a well-known tone detector, or to separate the spectrum into components encoded by high resolution and components encoded by low resolution. Includes any different type of valid detector. Depending on the spectral information or associated metadata about the resolution requirements for different spectral parts, other similar algorithms performed in the spectral analyzer may be voice activity detectors, noise detectors, speech detectors, or any decisive factor. It can be another detector.
For example, FIG. 5a illustrates a preferred embodiment of the time spectrum converter 100 of FIG. 1a, as performed in AAC or USAC. The time spectrum converter 100 has a window 502 controlled by a temporary detector 504. When the temporary detector 504 detects a temporary phenomenon, the switch from long window to short window is signaled to windower. Windower 502 then calculates the windowed frame to overlap the block. Here, each windowed frame generally has 2N values, for example 2048 values. Conversions within the range of block transducer 506 are then performed, and this block transducer generally provides additional decline. As a result, the congruent decline / transformation is performed to obtain a spectral frame with N values, such as MDCT spectral values. Thus, for long window operations, the frame at the input of block 506 contains 2N values, such as 2048 values, and the spectral frame then has 1024 values. Then, however, the switch runs to the short block. Each short block then has a time domain value that is 1/8 windowed compared to the long window, and each spectral block has a spectral value that is 1/8 compared to the long block. By the way, eight short blocks are executed. Thus, the spectrum is a definitive sampled version of the time domain audio signal 99, when this decline is combined with the 50% overlapping behavior of the window.
References are then made to FIG. 1b or FIG. 5b, which illustrates a particular embodiment of the combined operating frequency regenerator 116 and spectral time converter 118 of blocks 208, 212 of FIG. 2a. In FIG. 5b, the particular reconstructed band is considered, for example, as scale factor band 6 in FIG. 3a. The first spectral portion in this reconstructed band, i.e., the first spectral portion 306 of FIG. 3a, is the input to the frame assembler / regulator block 510. Furthermore, the reconstructed second spectral portion for scale factor band 6 is also the input to the frame assembler / regulator 510. Furthermore, E in Figure 3b for scale factor band 6<sub>3</sub>Energy information such as is also an input to block 510. The reconstructed second spectral portion of the reconstructed band has already been generated by frequency tiling using the source range, and the reconstructed band then corresponds to the target range. Currently, frame energy adjustments are performed, for example, to finally obtain a fully reconstructed frame with N values, as obtained at the output of combiner 208 in FIG. 2a. Then, in block 512, the reverse block transformation / interpolation is performed, for example, to obtain a 248 time domain value for 124 spectral values at the input of block 512. The synthetic windowing operation is then performed in block 514, which is again controlled by the long window / short window display transmitted as side information in the coded audio signal. Then, in block 516, the overlap / add operation with the previous time frame is performed. Preferably, MDCT applies a 50% overlap. As a result, the N time domain value is finally output for each new time frame of 2N values. 50% overlap is highly preferred due to the fact that it provides significant sampling and continuous intersection from one frame to the next due to the overlap / add operation of block 516.
As illustrated in FIG. 3a 301, the noise filling operation is additionally such as for a well-thought-out reconstructed band that is consistent with the scale factor band 6 of FIG. 3a as well as below the IGF start frequency. It can be applied even if the IGF start frequency is exceeded. Then, the noise-filled spectrum value can be an input to the frame assembler / regulator 510, and the adjustment of the noise-filled spectrum value can be applied within this block, or the noise-filled spectrum. The value can already be adjusted using noise filling energy before being an input to the frame assembler / regulator 510.
Preferably, the frequency tiles that fill the IGF operation, i.e. the operation using spectral values from other parts, can be applied in the complete spectrum. Thus, the spectral tiling operation can be applied not only in the high band above the IGF start frequency, but also in the low band. Furthermore, noise tessellation without frequency tile filling can be applied not only below the IGF start frequency but also above the IGF start frequency. However, high quality when the noise filling operation is limited to a frequency range below the IGF start frequency and when the frequency tile filling operation is limited to a frequency range above the IGF start frequency as illustrated in FIG. 3a. Moreover, it is known that highly efficient voice coding can be obtained.
Preferably, the target tile (TT) (having a frequency greater than the IGF start frequency) is constrained to the scale factor band boundaries of the full rate coder. That is, for frequencies below the IGF start frequency, the source tile (ST) from which the information is taken is not constrained to the scale factor band boundary. The size of the ST must correspond to the size of the TT involved. This is illustrated using the following example. TT [0] has a length of 10 MDCT Bins. This corresponds exactly to the length of the two next SCBs (eg 4 + 6). Then also all possible STs that correlate with TT [0] have a container length of 10. The second target tile TT [1] adjacent to TT [0] has a length of 15 vessels l (SCB with a length of 7 + 8). Then the ST for it has a length of 15 containers rather than 10 containers for TT [0].
If there is a case where no one can find a TT for an ST with the length of the target tile (eg, when the length of the TT is greater than the available source range), then the correlation. Is not calculated, and the source range is copied multiple times towards this TT until the target tile TT is fully filled (copying is done in sequence, so the first The frequency line for the lowest frequency of the second copy immediately follows in frequency the frequency line for the highest frequency of the first copy).
References are then made to FIG. 5c, which illustrates a further preferred embodiment of frequency regenerator 116 of FIG. 1b or IGF block 202 of FIG. 2a. Block 522 is a frequency tile generator that not only receives the target band ID, but also receives the source band ID. Schematically, it was determined on the encoder side that scale factor band 3 in FIG. 3a is very well suited for reconstructing scale factor band 7. Thus, the source band ID is 2 and the target band ID is 7. Based on this information, the frequency tile generator 522 applies a copy-up or harmonic tiling operation or any other tile filling operation to generate the raw second part of the spectral component 523. The raw second part of the spectral component has the same frequency resolution as the frequency resolution contained in the first set of first spectral parts.
Then, the first spectral portion of the reconstructed band, such as 307 in FIG. 3a, is the input to the frame assembler 524, and the raw second portion 523 is also the input to the frame assembler 524. .. The reconstructed frame is then tuned by the regulator 526 using the gain factor for the reconstructed band calculated by the gain factor calculator 528. However, importantly, the first spectral portion of the frame is not affected by the regulator 526, but only the raw second portion for the reconstructed frame is affected by the regulator 526. For this purpose, in order to finally find the correct gain factor 527, the gain factor calculator 528 analyzes the source band or the raw second part 523 and, in addition, the first spectral part in the reconstructed band. To analyze. As a result, when scale factor band 7 is taken into account, the energy of the frame output tuned by regulator 526 is energy E.<sub>4</sub>Have.
In this situation, it is very important to evaluate the high frequency reconstruction accuracy of the present invention as compared to HE-AAC. This is described for scale factor band 7 in Figure 3a. It is envisioned that a prior art encoder, as illustrated in FIG. 13a, will detect the spectral portion 307, which is coded at high resolution as the "lost overtone". The energy of this spectral component is then transmitted to the decoder along with spectral envelope information for the reconstructed band, such as scale factor band 7. The decoder then reshapes the lost overtones. However, the spectral value at which the overtone 307 lost by the prior art decoder of FIG. 13b is being reconstructed is in the middle of band 7 at the frequency indicated by the reconstruction frequency 390. Thus, the present invention avoids frequency error 391 induced by the prior art decoder of FIG. 13d.
In embodiments, a spectral analyzer is also performed to calculate similarities between the first spectral portion and the second spectral portion, and based on the calculated similarities, a second in the reconstruction range. It is carried out to determine the first spectral portion that matches the second spectral portion as much as possible with respect to the spectral portion of. Then, in this variable source range / destination range embodiment, the parametric coder additionally, in a second coded representation, for each destination range. Introduce matching information that indicates the matching source range. On the decoder side, this information is then used by the frequency tile generator 522 in FIG. 5c, which illustrates the generation of the raw second portion 523 based on the source band ID and target band ID.
Further, as illustrated in FIG. 3a, the spectrum analyzer has a spectrum up to a maximum analysis frequency of less than half the sampling frequency, preferably at least 1/4 of the sampling frequency or generally higher. It is configured to analyze the expression.
As shown, the encoder operates without downsampling, and the decoder operates without unsampling. In other words, the spectral domain audio coder is configured to produce a spectral representation with a Nyquist frequency defined by the sampling rate of the original input audio signal.
Further, as illustrated in FIG. 3a, the spectral analyzer is configured to analyze a spectral representation that begins at the gap filling start frequency and ends at the maximum frequency represented by the maximum frequency contained in the spectral representation. Here, the spectral portion extending from the minimum frequency to the gap filling start frequency belongs to the first set of spectral portions and has frequency values above the gap filling frequency, such as 304, 305, 306, 307. Further spectral moieties are additionally included in the first set of first spectral moieties.
As outlined, the maximum frequency represented by the spectral values in the first decoded representation is such that the spectral values for the maximum frequencies in the first set of first spectral portions are 0 or different. The spectral region audio decoder 112 is configured to be equal to the maximum frequency contained in the time representation having the sampling rate. In any case, for this maximum frequency in the first set of spectral components, there is a scale factor for the scale factor band. It is then generated and transmitted regardless of whether all spectral values in this scale factor band are set to 0, as described in the context of FIGS. 3a and 3b.
Accordingly, the present invention is techniques of other parameters to increase compression efficiency, such as noise substitution and noise filling (these techniques are only for efficient representation of noise such as local signal content). It is advantageous with respect to. The present invention enables accurate frequency reproduction of sound components. To date, state-of-the-art technology has been the efficient parameter representation of any signal content with fixed a priori splitting unrestricted spectral gap filling in the high band (HF) and low band (LF). Do not state.
The embodiments of the system of the present invention improve the state-of-the-art method, which results in high compression efficiency even for low bitrates and no or no perceptual discomfort and full voice bandwidth. I will provide a.
A general system consists of the following.
-Complete band-Core coding-Intelligent gap filling (tile filling or noise filling) -Sparse sound part in the core selected by the sound mask-Code for complete band including tile filling Spectral whitening in TNS / IGF range on joint stereo pair tiles
The first step towards a more efficient system is to remove the need to transform the spectral data into a second transform region that is different from one of the core coder. Just as most audio codecs, such as AAC, use MDCT as the basic conversion, it is also useful to perform a BWE in the MDCT domain. The second requirement for the BWE system is the need to preserve even the components of the HF sound, and to preserve the sound grid, which is such an advantage in the quality of the encoded speech over existing systems. To handle both of the above requirements, a system called Intelligent Gap Filling (IGF) has been proposed. FIG. 2b shows a block diagram of the proposed system on the encoder side, and FIG. 2a shows the system on the decoder side.
The post-processing framework is then described with respect to FIGS. 13A and 13B to illustrate that the present invention can also be implemented in the high frequency reconstructor 1330 in this post-processing embodiment.
For example, as used in high efficiency AAC (HE-AAC), FIG. 13a illustrates a schematic diagram of an audio encoder for bandwidth expansion technology. The audio signal on line 1300 is an input to a filter system with low pass 1302 and high pass 1304. The signal output from the high pass filter 1304 is the input to the parameter extractor / coder 1306. The parameter extractor / coder 1306 is configured to calculate and encode parameters such as, for example, spectral envelope parameters, noise addition parameters, lost overtone parameters or vice versa filtering parameters. These extracted parameters are the inputs to the bitstream multiplexer 1308. The low-pass output signal is an input to a processor that generally has the functionality of a down sampler 1310 and a core coder 1312. The low pass 1302 limits the encoded bandwidth to significantly less bandwidth than that generated in the original input audio signal on line 1300. This provides an important coding gain due to the fact that all the functionality that occurs in the core coder must only act on the signal with the reduced bandwidth. For example, when the bandwidth of the audio signal on line 1300 is 20 kHz, and when the low pass filter 1302 has a bandwidth of 4 kHz exemplary, to the down sampler to satisfy the sampling theorem. It is theoretically sufficient that the following signal has a sampling frequency of 8 kHz. And it is a considerable reduction to the sampling rate required for the audio signal 1300, which must be at least 40kHz.
FIG. 13b exemplifies a schematic diagram of the corresponding bandwidth expansion decoder. The decoder has a bitstream multiplexer 1320. The bitstream demultiplexer 1320 extracts an input signal to the core decoder 1322 and an input signal to the parameter decoder 1324. In the above example, the core decoder output signal has a sampling rate of 8 kHz and therefore a bandwidth of 4 kHz, while for full bandwidth reconstruction, the high frequency reconstructor 1330. The output signal must be 20kHz, which requires a sampling rate of at least 40kHz. To make this possible, a decoder processor with the functionality of the upsampler 1325 and the filter bank 1326 is required. The high frequency reconstructor 1330 then receives the low frequency signal output frequency analyzed by the filter bank 1326 and is defined by the high pass filter 1304 in Figure 13a using the representation of the parameters in the high frequency band. Reconstruct the frequency range. The high frequency reconstructor 1330 is capable of regenerating the upper frequency range using the source range in the lower frequency range, adjusting the spectrum envelope, adding noise, and introducing the overtones lost in the higher frequency range. And to explain the cause of the fact that the higher frequency range is generally not as audible as the lower frequency range when applied and calculated in the encoder of Figure 13a. Has the opposite filter operation. In HE-AAC, the lost overtones are resynthesized on the decoder side and placed exactly in the center of the reconstructed band. Therefore, not all lost overtones determined in a particular reconstruction band are placed at the frequency value where they were located in the original signal. Instead, those lost overtones are placed at a frequency in the middle of a particular band. In this way, the overtone line where the original signal is lost is the boundary of the reconstructed band of the original signal. When placed very close to the field, the error in the frequency induced by placing this lost overtone of the reconstructed signal in the center of the band is about 50% of the individual reconstructed bands. be. Parameters have been generated and transmitted for the reconstructed band.
Moreover, even if a typical voice core coder operates in the spectral region, the core decoder will nevertheless generate a time domain signal that is then again converted into the spectral region by the function of filter bank 1326. .. This leads to further processing delays, which can lead to artificial products by tandem processing with a first conversion from the spectral domain to the frequency domain and another conversion to generally different frequency domains, and of course. This also requires a sufficient amount of computational complexity and therefore power. And it is especially problematic when bandwidth expansion technology is applied to mobile devices such as mobile phones, tablets or laptop computers.
Although some embodiments are described in the context of devices for encoding or decoding, it is clear that these embodiments also represent a description of the corresponding method, where the block or device is a block or device. Corresponds to the characteristics of the method step or method step. Similarly, the embodiments described in the context of the steps of the method also represent a description of the characteristics of the corresponding block, item or corresponding device. Some or all of the steps in the method may be performed (or used) by a hardware device such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps can be performed by such a device.
Depending on the specific implementation requirements, the embodiments of the present invention may be implemented in hardware or in software. The embodiment thereof is a digital storage medium having a control signal stored therein and having an electronically readable control signal, for example, a floppy (registered trademark) disk, a hard disk drive (HDD), a DVD, or a Blu-Ray (registered trademark). ), CD, ROM, PROM, EPROM, EEPROM, or can be performed using non-temporary storage media such as FLASH memory. It works with (or can work with) a programmable computer system, and as a result, each method is performed. Therefore, the digital storage medium may be computer readable.
Some embodiments according to the invention include data carriers having electronically readable control signals that can work with programmable computer systems. As a result, one of the methods described herein is performed.
Generally, an embodiment of the present invention is carried out as a computer program product having a program code, and when the computer program product runs on a computer, the program code executes one of the methods. Is effective against. The program code may be stored, for example, on a machine-readable carrier.
Other embodiments include computer programs stored in machine readable carriers and for performing one of the methods described herein.
In other words, therefore, when a computer program runs on a computer, an embodiment of the method of the invention is a computer having program code for performing one of the methods described herein. It is a program.
Accordingly, further embodiments of the methods of the invention are recorded on it and include a data carrier (or digital storage medium) comprising a computer program for performing one of the methods described herein. , Or a computer-readable medium).
Accordingly, a further embodiment of the method of the invention is a data stream or set of signals representing a computer program for performing one of the methods described herein. For example, a data stream or set of signals may be configured to be transferred over a data communication connection, eg, over the Internet.
Further embodiments include processing means configured or adapted to perform one of the methods described herein, such as a computer or programmable logic device.
Further embodiments include a computer installed on it and having a computer program for performing one of the methods described herein.
Further embodiments of the present invention are configured to transfer (eg, electronically or optically) a computer program to perform one of the methods described herein. Includes equipment or systems. The receiver may be, for example, a computer, a mobile device, a memory element, or the like. The device or system may include, for example, a file server for transferring computer programs to the receiver.
In some embodiments, programmable logic devices (eg, field programmable gate arrays) may be used to perform some or all of the functions of the methods described herein. good. In some embodiments, the field programmable gate array may work with a microprocessor to perform one of the methods described herein. In general, the method is preferably performed by some hardware device.
The embodiments described above are merely described for the purposes of the present invention. It is understood that the modifications and changes in arrangement and the details described herein will be apparent to those of ordinary skill in the art. Accordingly, it is intended to be limited only by the imminent claims and not by the specific details provided herein as description and description of the examples.
48 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| WO2013002623A2 | Cites | World Intellectual Property Organization (WIPO) |
| WO2014033131A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP201315598A | Cites | Japan |
| WO2013141638A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP2011527455A | Cites | Japan |
| JP2004053895A | Cites | Japan |
92 members in 18 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 14178777 | European Patent Office (EPO) | A | |
| 14178777 | European Patent Office (EPO) | A | |
| 141787770 | European Patent Office (EPO) | – | |
| 141787770 | – | – | – |
| EP20140178777 | – | – | – |
Members92
| Document | Office | Kind | |
|---|---|---|---|
| EP2980792A1 | European Patent Office (EPO) | A1 | |
| CA2947804A1 | Canada | A1 | |
| CA2956024A1 | Canada | A1 | |
| WO2016016144A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2016016146A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201608561A | Taiwan Province of China | A | |
| TW201618083A | Taiwan Province of China | A | |
| AR101345A1 | Argentina | A1 | |
| AR101346A1 | Argentina | A1 | |
| AU2015295547A1 | Australia | A1 | |
| SG11201700631UA | Singapore | A | |
| SG11201700689VA | Singapore | A | |
| KR20170024048A | Republic of Korea | A | |
| US2017069332A1 | United States of America | A1 | |
| AU2015295549A1 | Australia | A1 | |
| TWI575511B | Taiwan Province of China | B | |
| TWI575515B | Taiwan Province of China | B | |
| CN106537499A | China | A | |
| US2017133024A1 | United States of America | A1 | |
| CN106796798A | China | A | |
| EP3175449A1 | European Patent Office (EPO) | A1 | |
| KR20170063534A | Republic of Korea | A | |
| EP3186807A1 | European Patent Office (EPO) | A1 | |
| MX2017001231A | Mexico | A | |
| MX2017001236A | Mexico | A | |
| JP2017526004A | Japan | A | |
| JP2017526957A | Japan | A | |
| BR112017000852A2 | Brazil | A2 | |
| BR112017001586A2 | Brazil | A2 | |
| AU2015295547B2 | Australia | B2 | |
| EP3175449B1 | European Patent Office (EPO) | B1 | |
| RU2016146738A | Russian Federation | A | |
| RU2016146738A3 | Russian Federation | A3 | |
| RU2017105507A | Russian Federation | A | |
| RU2017105507A3 | Russian Federation | A3 | |
| RU2665913C2 | Russian Federation | C2 | |
| RU2667376C2 | Russian Federation | C2 | |
| AU2015295549B2 | Australia | B2 | |
| TR2018016634T4 | Türkiye | T4 | |
| TR201816634T4 | Türkiye | T4 | |
| PT3175449T | Portugal | T | |
| ES2693051T3 | Spain | T3 | |
| EP3186807B1 | European Patent Office (EPO) | B1 | |
| JP6457625B2 | Japan | B2 | |
| PL3175449T3 | Poland | T3 | |
| KR101958359B1 | Republic of Korea | B1 | |
| KR101958360B1 | Republic of Korea | B1 | |
| MX363352B | Mexico | B | |
| PT3186807T | Portugal | T | |
| EP3471094A1 | European Patent Office (EPO) | A1 | |
| CA2956024C | Canada | C | |
| JP2019074755A | Japan | A | |
| TR2019004282T4 | Türkiye | T4 | |
| TR201904282T4 | Türkiye | T4 | |
| MX365086B | Mexico | B | |
| JP6535730B2 | Japan | B2 | |
| PL3186807T3 | Poland | T3 | |
| CA2947804C | Canada | C | |
| ES2718728T3 | Spain | T3 | |
| US10354663B2 | United States of America | B2 | |
| US2019295561A1 | United States of America | A1 | |
| JP2019194704A | Japan | A | |
| US10529348B2 | United States of America | B2 | |
| CN106537499B | China | B | |
| US2020090668A1 | United States of America | A1 | |
| CN111261176A | China | A | |
| US10885924B2 | United States of America | B2 | |
| US2021065726A1 | United States of America | A1 | |
| CN106796798B | China | B | |
| CN113160838A | China | A | |
| JP6943836B2 | Japan | B2 | |
| JP2022003397A | Japan | A | |
| JP6992024B2This record | Japan | B2 | |
| US11264042B2 | United States of America | B2 | |
| JP2022046504A | Japan | A | |
| US2022148606A1 | United States of America | A1 | |
| BR112017000852B1 | Brazil | B1 | |
| BR112017001586B1 | Brazil | B1 | |
| US11705145B2 | United States of America | B2 | |
| JP7354193B2 | Japan | B2 | |
| US2023386487A1 | United States of America | A1 | |
| JP7391930B2 | Japan | B2 | |
| US11908484B2 | United States of America | B2 | |
| CN111261176B | China | B | |
| CN113160838B | China | B | |
| EP4439559A2 | European Patent Office (EPO) | A2 | |
| EP3471094B1 | European Patent Office (EPO) | B1 | |
| EP3471094C0 | European Patent Office (EPO) | C0 | |
| EP4439559A3 | European Patent Office (EPO) | A3 | |
| ES2992880T3 | Spain | T3 | |
| US12205604B2 | United States of America | B2 | |
| PL3471094T3 | Poland | T3 |
15 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6992024
- Publication, DOCDB
- 6992024
- Publication, EPODOC
- JP6992024B
- Application
- 103541
- Application, DOCDB
- 2019103541
- Application, EPODOC
- JP20190103541
Titles2
- Japanese
- 独立したノイズ充填を用いた強化された信号を生成するための装置および方法
- English
- Equipment and methods for generating enhanced signals with independent noise filling
Classification
- CPC, 5
- G10L19/028
- G10L19/0204
- G10L21/038
- G10L25/21
- G10L15/20
- IPC, 2
- G10L21 0388
- G10L19 02
