Concept of encoding an audio signal and decoding an audio signal using deterministic and noise like information
17 claims: 5 independent, 12 dependent
- 1オーディオ信号を符号化する符号器であって、前記オーディオ信号(102)のある無声フレームから予測係数(122;322)と残差信号とを導出するよう構成された分析部(120;320)と、前記無声フレームについて、確定的コードブック(deterministic codebook)に関連する第1励振信号(c(n))を定義する第1ゲインパラメータ(g c )情報と、ノイズ状信号に関連する第2励振信号(n(n))を定義する第2ゲインパラメータ(g n )情報とを計算するよう構成されたゲインパラメータ計算部(550;550')と、有声信号フレームに関連する情報(142)と前記第1ゲインパラメータ(g c )情報と前記第2ゲインパラメータ(g n )情報とに基づいて、出力信号(692)を形成するよう構成されたビットストリーム形成部(690)と、を含む符号器。
- 2請求項1に記載の符号器において、前記ゲインパラメータ計算部(550;550')は、第1ゲインパラメータ(g c )と第2ゲインパラメータ(g n )とを計算するよう構成され、前記ビットストリーム形成部(690)は前記第1ゲインパラメータ(g c )と前記第2ゲインパラメータ(g n )とに基づいて前記出力信号(692)を形成するよう構成されるか、又は前記ゲインパラメータ計算部(550;550')は、前記第1ゲインパラメータ(g c )を量子化して 量子化済み第1 ゲインパラメータ を取得し、かつ前記第2ゲインパラメータ(g n )を量子化して 量子化済み第2 ゲインパラメータ を取得するよう構成され、前記ビットストリーム形成部(690)は前記 量子化済み第1 ゲインパラメータ と前記 量子化済み第2 ゲインパラメータ とに基づいて前記出力信号(692)を形成するよう構成された、符号器。
- 3請求項1に記載の符号器において、前記予測係数(122;322)からスピーチ関連のスペクトル整形情報(162)を計算するよう構成されたフォルマント情報計算部(160)をさらに含み、前記ゲインパラメータ計算部(550;550')は、前記スピーチ関連のスペクトル整形情報(162)に基づいて前記第1ゲインパラメータ(g c )情報と前記第2ゲインパラメータ(g n )情報とを計算するよう構成された、符号器。
- 4請求項1に記載の符号器において、前記ゲインパラメータ計算部(550')は、第1ゲインパラメータ(g c )を適用することによって前記第1励振信号(c(n))を増幅し、第1の増幅された励振信号(550f)を得るよう構成された第1増幅部(550e)と、第2ゲインパラメータ(g n )を適用することによって前記第1励振信号(c(n))とは異なる前記第2励振信号(n(n))を増幅し、第2の増幅された励振信号(350g;550h)を得るよう構成された第2増幅部(350e;550g)と、前記第1の増幅された励振信号(550f)と前記第2の増幅された励振信号(350g;550h)とを結合して、結合済み励振信号(550k;550k')を得るよう構成された結合部(550i)と、合成フィルタを用いて前記結合済み励振信号(550k;550k')をフィルタリングして合成信号(350l')を取得し、前記合成信号(350i')と前記オーディオ信号(102)のフレームとを比較して比較結果を取得し、前記比較結果に基づいて前記第1ゲインパラメータ(g c )又は前記第2ゲインパラメータ(g n )を適応するよう構成された制御部(550n)と、を含み、前記ビットストリーム形成部(690)は、前記第1ゲインパラメータ(g c ) 情報 及び前記第2ゲインパラメータ(g n ) 情報 に基づいて前記出力信号(692)を形成するよう構成された、符号器。
- 5請求項1~4のいずれか一項に記載の符号器において、前記ゲインパラメータ計算部(550;550')は、スペクトル整形情報(162)に基づいて、前記第1励振信号(c(n))若しくはそれから導出された信号、又は前記第2励振信号(n(n))若しくはそれから導出された信号をスペクトル的に整形するよう構成された、少なくとも1つの整形器(350c;550b)をさらに含む、符号器。
- 6請求項1~5のいずれか一項に記載の符号器において、前記符号器は前記オーディオ信号(102)をフレームシーケンスの中でフレーム毎に符号化するよう構成され、前記ゲインパラメータ計算部(550;550')は、処理済みフレームの複数のサブフレームの各々について第1ゲインパラメータ(g c )及び第2ゲインパラメータ(g n )を決定するよう構成され、前記ゲインパラメータ計算部(550;550')は、前記処理済みフレームに関連した平均エネルギー値を決定するよう構成された、符号器。
- 7請求項1~6のいずれか一項に記載の符号器において、前記予測係数(122;322)から少なくとも第1のスピーチ関連のスペクトル整形情報を計算するよう構成されたフォルマント情報計算部(160)と、前記残差信号が前記オーディオ信号の無声フレームから決定されたか否かを判定するよう構成された判定部(130)と、をさらに含む符号器。
- 8請求項1~7のいずれか一項に記載の符号器において、前記ゲインパラメータ計算部(550;550')は、次式に基づいて第1ゲインパラメータ(g c )を決定するよう構成された制御部(550n)を含み、 ここで、cw(n)は革新的コードブックのフィルタ済み励振信号であり、xw(n)はCELP符号器において計算された知覚的目標励振であり、前記制御部(550n)は、前記第1ゲインパラメータの量子化値 と、前記第1励振 信号 及び前記第2励振 信号 の間の二乗平方根エネルギー比 とに基づいて、量子化済みノイズゲイン を決定するよう構成され、ここでLsfはサンプル内のサブフレームのサイズである、符号器。
- 9請求項4に記載の符号器において、前記第1ゲインパラメータ(g c )を量子化して量子化済み第1ゲインパラメータ を取得するよう構成された量子化部(170-1、170-2)を更に含み、前記制御部(550n)は、次式に基づいて前記第1ゲインパラメータ(g c )を決定するよう構成され、 ここで、g c は前記第1ゲインパラメータであり、Lsfはサンプル内のサブフレームのサイズであり、cw(n)は第1の整形済み励振信号であり、xw(n)は符号励振線形予測符号化信号であり、前記制御部(550n)又は前記量子化部(170-1、170-2)は、前記第1ゲインパラメータ(g c )を正規化して、次式に基づいて正規化済み第1ゲインパラメータを得るようさらに構成され、 ここで、g nc は前記正規化済み第1ゲインパラメータを示し、 は無声残差信号の全体フレームにわたる平均エネルギーの尺度であり、前記量子化部(170-1、170-2)は、前記正規化済み第1ゲインパラメータを量子化して前記量子化済み第1ゲインパラメータ を得るよう構成された、符号器。
- 10請求項9に記載の符号器において、前記量子化部(170-1、170-2)は、前記第2ゲインパラメータ(g n )を量子化して量子化済み第2ゲインパラメータ を得るよう構成され、前記ゲインパラメータ計算部(550;550')は、次式に基づいて誤差の値を決定することにより前記第2ゲインパラメータ(g n )を決定するよう構成され、 ここで、kは0.5と1との間の範囲内にある可変の減衰ファクタであり、Lsfは処理済みオーディオフレームのサブフレームのサイズに対応し、cw(n)は前記第1の整形済み励振信号を示し、xw(n)は符号励振線形予測符号化信号を示し、g n は前記第2ゲインパラメータを示し、 は量子化済み第1ゲインパラメータを示し、前記ゲインパラメータ計算部(550;550')は、現在のサブフレームについて前記誤差を決定するよう構成され、前記量子化部(170-1、170-2)は、前記誤差を最小化する前記量子化済み第2ゲインパラメータ を決定し、かつ次式に基づいて前記量子化済み第2ゲインパラメータ を取得するよう構成され、 ここで、Q(index n )は可能な値の有限集合からのスカラー値を示す、符号器。
- 11請求項10に記載の符号器において、前記結合部(550i)は、前記量子化済み第1ゲインパラメータと前記量子化済み第2ゲインパラメータとを結合して、次式 に基づいて結合済み励振信号(e(n))を得るよう構成された、符号器。
- 12予測係数(122)に関連する情報を含む受信されたオーディオ信号(1002)を復号化する復号器(1000)であって、合成信号(1062)の一部分のために、確定的コードブックから第1励振信号(1012)を生成するよう構成された第1信号生成部(1010)と、前記合成信号(1062)の前記一部分のために、ノイズ状信号から第2励振信号(1022)を生成するよう構成された第2信号生成部(1020)と、前記第1励振信号(1012)と前記第2励振信号(1022)とを結合して、前記合成信号(1062)の前記一部分のための結合済み励振信号(1052)を生成するよう構成された結合部(1050)と、前記結合済み励振信号(1052)と予測係数(122)とから前記合成信号(1062)の前記一部分を合成するよう構成された合成部(1060)と、を含む復号器。
- 13請求項12に記載の復号器において、前記受信されたオーディオ信号(1002)は、第1ゲインパラメータ(g c )と第2ゲインパラメータ(g n )とに関連する情報とを含み、前記復号器は、前記第1ゲインパラメータ(g c )を適用することによって前記第1励振信号(1012)又はそれから導出された信号を増幅して、第1の増幅済み励振信号(1012')を得るよう構成された第1増幅部(550e)と、前記第2ゲインパラメータを適用することによって前記第2励振信号(1022)又はそれから導出された信号を増幅して、第2の増幅済み励振信号(1022')を得るよう構成された第2増幅部(254;350e;550g)と、をさらに含む復号器。
- 14請求項12又は13に記載の復号器において、前記予測係数(122;322)から第1のスペクトル整形情報(1092a)と第2のスペクトル整形情報(1092b)とを計算するよう構成されたフォルマント情報計算部(160;1090)と、前記第1のスペクトル整形情報(1092a)を使用して、前記第1励振信号(1012)又はそれから導出された信号のスペクトルをスペクトル的に整形するよう構成された第1整形器(1070)と、前記第2のスペクトル整形情報(1092b)を使用して、前記第2励振信号(1022)又はそれから導出された信号のスペクトルをスペクトル的に整形するよう構成された第2整形器(1080)と、をさらに含む復号器。
- 15オーディオ信号(102)を符号化する方法(1400)であって、前記オーディオ信号(102)のある無声フレームから予測係数(122;322)と残差信号とを導出するステップ(1410)と、前記無声フレームについて、確定的コードブック(deterministic codebook)に関連する第1励振信号(c(n))を定義する第1ゲインパラメータ(g c )情報を計算し、かつノイズ状信号に関連する第2励振信号(n(n))を定義する第2ゲインパラメータ(g n )情報を計算するステップ(1420)と、有声信号フレームに関連する情報(142)と前記第1ゲインパラメータ(g c )情報と前記第2ゲインパラメータ(g n )情報とに基づいて、出力信号(692;1002)を形成するステップ(1430)と、を含む方法。
- 16予測係数(122;322)に関連する情報を含む受信されたオーディオ信号(692;1002)を復号化する方法(1500)であって、合成信号(1062)の一部分のために、確定的コードブックから第1励振信号(1012,1012')を生成するステップ(1510)と、前記合成信号(1062)の前記一部分のために、ノイズ状信号(n(n))から第2励振信号(1022;1022')を生成するステップ(1520)と、前記第1励振信号(1012;1012')と前記第2励振信号(1022;1022')とを結合して、前記合成信号(1062)の前記一部分のための結合済み励振信号(1052)を生成するステップ(1530)と、前記結合済み励振信号(1052)と予測係数(122;322)とから前記合成信号(1062)の前記一部分を合成するステップ(1540)と、を含む方法。
- 17コンピュータ上で作動されたとき、請求項 15又は16 に記載の方法を実行するためのプログラムコードを有するコンピュータプログラム。
Independent claims17
112 paragraphs, as filed
0001The present invention relates to a encoder that encodes an audio signal, particularly a speech-related audio signal. The present invention also relates to decoders and methods for decoding encoded audio signals. The present invention further relates to encoded audio signals and advanced speech unvoiced coding at low bit rates.
0002Speech coding at low bitrates can benefit from special handling for silent frames in order to maintain speech quality while reducing the bitrate. Silent frames can be perceptually modeled as random excitation shaped in both the frequency domain and the time domain. Since its waveform and excitation look and sound much like Gaussian white noise, its waveform coding can be mitigated and replaced by synthetically generated white noise. This coding will then consist of coding the time and frequency domain shape of the signal.
0003FIG. 16 shows a schematic block diagram of the parametric silent encoding scheme. The synthetic filter 1202 is configured to model the vocal tract and is parameterized by LPC (Linear Predictive Coding) parameters. A perceptually weighted filter can be derived by weighting the LPC coefficient from the derived LPC filter containing the filter function A (z). The perceptual filter fw (n) usually has a transfer function of the form: [Number 1]<img id="000002" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Where w is less than 1. Gain parameter g<sub>n</sub>Is calculated according to the following equation to obtain a synthesized energy that matches the original energy in the perceptual domain. [Number 2]<img id="000003" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, sw (n) and nw (n) indicate the input signal filtered by the perceptual filter and the generated noise, respectively. Gain g<sub>n</sub>Is calculated for each subframe of size Ls. For example, one audio signal may be divided into a plurality of frames having a length of 20 ms. Each frame may be subdivided into a plurality of subframes, for example, four subframes each having a length of 5 ms.
0004Code-excited linear prediction (CELP) coding schemes are widely used in speech communication and are a very efficient method of coding speech. CELP coding gives more natural speech quality than parametric coding, but requires higher rates. CELP synthesizes an audio signal by transporting it to a linear predictive filter called an LPC synthesis filter. The LPC synthesis filter may include the sum of two excitations of the form 1 / A (z). One excitation comes from a decrypted past excitation called an adaptive codebook. The other contribution comes from the innovative codebook, which stores fixed codes. However, at low bit rates, innovative codebooks are not well stored to efficiently model speech microstructure or silent noise-like excitation. Therefore, the perceptual quality deteriorates, especially the silent frame sounds crispy and unnatural.
0005Different solutions have already been proposed to mitigate coding artifacts at low bitrates. In Non-Patent Document 1 and Patent Document 1, the code of the innovative codebook is adaptively and spectrally shaped by emphasizing the spectral region corresponding to the formant of the current frame. This formant position and shape can be deducted directly from the LPC coefficient, which is already available on both the encoder side and the decoder side. Formant emphasis in code c (n) is performed by simple filtering according to the following equation. [Number 3]<img id="000004" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, * indicates the convolution operator, and fe (n) is the impulse response of the filter of the transfer function shown in the following equation. [Number 4]<img id="000005" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0006Here, w1 and w2 are two weighting constants that emphasize the formant structure of the transfer function Ffe (z) more or less. The resulting shaped code inherits the characteristics of the speech signal and the composite signal sounds clearer.
0007In CELP, it is also normal to add spectral gradients to the decoders of innovative codebooks. It is done by filtering the code with the following filters: [Number 5]<img id="000006" he="15" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0008Factor β is usually related to and depends on the voicing of the previous frame. That is, it changes. Voicing can be inferred from energy contributions from adaptive codebooks. If the previous frame is voiced, it is expected that the current frame will also be voiced and that the code should have more energy at low frequencies, i.e. show a negative slope. In contrast, the added spectral gradient will be positive for silent frames and more energy will be distributed towards higher frequencies.
0009The use of spectrum shaping for speech enhancement and noise reduction at the output of the decoder is a common practice. So-called formant emphasis as post-filtering consists of adaptive post-filtering, the coefficients of which are derived from the decoder's LPC parameters. The filter then looks similar to the post-filter (fe (n)) used to shape innovative excitations in some CELP coder as described above. However, in such cases, post-filtering is applied only at the end of the decoder process, not on the encoder side.
0010In traditional CELP (CELP = (code) book excitation linear prediction), the frequency shape is modeled by an LP (linear prediction) synthesis filter, while the time domain shape is the excitation sent for all subframes. Can be approximated by gain. However, long-term potentiation (LTP) and innovative codebooks are usually not suitable for modeling noise-like excitation of silent frames. To achieve good quality of silent speech, CELP requires a relatively high bit rate.
0011The characterization of voiced or unvoiced sounds may be related to segmenting the speech into multiple parts, and each of those parts may be associated with a different source model of the speech. The source model used in the CELP speech coding scheme is an adaptive harmonic excitation that simulates the airflow through the glottis and a resonance filter that models the vocal tract excited by the generated airflow. Depends on. Such models can provide good results for voiced phonemes, but for speech parts that are not produced by the glottis, especially if the vocal cords are not vibrating, such as the unvoiced phonemes "s" and "f". Can lead to inaccurate modeling.
0012Parametric speech coder, on the other hand, is also called a vocoder and employs a single source model for silent frames. This can achieve very low bit rates, but results in so-called synthetic quality, which is not as natural as the quality delivered by the CELP coding scheme at much higher rates.
0013Therefore, it becomes necessary to strengthen the audio signal.
<p num="0014"><patcit num="1"><text>[2] US Pat. No. 5,444,816, "Dynamic codebook for efficient speech coding based on algebraic codes"</text></patcit></p>
<p num="0015"><nplcit num="1"><text>[1] Recommendation ITU-T G.718: "Frame error robust narrow-band and wideband embedded variable bit-rate coding of speech and audio from 8-32 kbit / s"</text></nplcit><nplcit num="2"><text>[3] Jelinek, M .; Salami, R., "Wideband Speech Coding Advances in VMR-WB Standard," Audio, Speech, and Language Processing, IEEE Transactions on, vol.15, no.4, pp.1167,1179 , May 2007</text></nplcit></p>
<p num="0016">An object of the present invention is to improve voice quality at low bit rates and / or reduce bit rates for good voice quality.</p>
<p num="0017">This object is achieved by an independently claimed encoder, decoder, encoded audio signal, and method thereof.</p><p num="0018">The present inventors have made the following discoveries. That is, in the first aspect, the quality of the decoded audio signal, which is related to the silent frame of the audio signal, is the shaping information related to a speech and the gain parameter information about the amplification of the signal. It is a discovery that it can be improved or enhanced by making decisions in a way that can be derived from speech-related shaping information. In addition, some speech-related shaping information can be used to spectrally shape the decoded signal. This allows frequency domains of greater importance to speech, such as low frequencies below 4kHz, to be processed so that their error is less.</p><p num="0019">The present inventors also made the following discoveries. That is, in the second aspect, the first excitation signal is generated from the deterministic codebook for the frame or subframe (part) of the composite signal, and the frame or subframe (part) of the composite signal. It was discovered that the quality of the combined signal can be improved, that is, enhanced by generating the second excitation signal from the noise-like signal of the above, and further combining the first excitation signal and the second excitation to generate the coupled excitation signal. Is. In particular, for each portion of the audio signal, including the speech signal with background noise, the sound quality can be improved by adding a noisy signal. The gain parameter for amplifying the first excitation signal may be optionally determined in the encoder and the information associated with that parameter may be transmitted with the encoded audio signal.</p><p num="0020">Alternatively or additionally, the enhancement of the synthesized audio signal may be utilized, at least in part, to reduce the bit rate in encoding the audio signal.</p><p num="0021">The encoder according to the first aspect includes an analysis unit configured to derive a prediction coefficient and a residual signal from a frame of an audio signal. The encoder further includes a formant information calculator configured to calculate speech-related spectral shaping information from prediction coefficients. The encoder has a gain parameter calculation unit configured to calculate the gain parameter from the unvoiced residual signal and spectrum shaping information, and information related to the voiced signal frame and the gain parameter or the quantized gain parameter and the prediction coefficient. A bit stream forming unit configured to form an output signal based on the above is further included.</p><p num="0022">A further embodiment according to the first aspect is a coded audio signal, the prediction coefficient information about the voiced frame and the unvoiced frame of the audio signal, further information related to the voiced signal frame, and unvoiced. A encoded audio signal is provided that includes a gain parameter or a quantized gain parameter for the frame. This makes it possible to efficiently transmit speech-related information and decode the encoded audio signal to obtain a synthesized (restored) signal with high audio quality.</p><p num="0023">Another embodiment according to the first aspect provides a decoder that decodes a received signal that includes a prediction factor. The decoder includes a formant information calculation unit, a noise generation unit, a shaper, and a synthesis unit. The formant information calculation unit is configured to calculate speech-related spectrum shaping information from prediction coefficients. The noise generation unit is configured to generate a decoded noise-like signal. The shaper is configured to use the spectrum shaping information to shape the spectrum of the decoded noise-like signal or its amplified representation to obtain the shaped decoded noise-like signal. The synthesis part is<u style="single">Formatted decryption</u>It is configured to synthesize a composite signal from a noise-like signal and a prediction coefficient.</p><p num="0024">Another embodiment according to the first aspect relates to a method of encoding an audio signal, a method of decoding a received audio signal, and a computer program.</p><p num="0025">The embodiment according to the second aspect provides a encoder that encodes an audio signal. The encoder includes an analyzer configured to derive the prediction factor and the residual signal from the silent frame of the audio signal. The encoder calculates the first gain parameter information that defines the first excitation signal associated with the deterministic codebook for the silent frame, and defines the second excitation signal associated with the noise-like signal. 2 Further includes a gain parameter calculator configured to calculate gain parameter information. The encoder further includes a bitstream forming unit configured to form an output signal based on information related to the voiced signal frame, first gain parameter information, and second gain parameter information.</p><p num="0026">A further embodiment according to the second aspect provides a decoder that decodes a received audio signal that includes information related to the prediction factor. The decoder includes a first signal generator configured to generate a first excitation signal from a deterministic codebook for a portion of the synthesized signal. The decoder further includes a second signal generator configured to generate a second excitation signal from the noise-like signal for said portion of the composite signal. The decoder further includes a coupling portion and a compositing portion so that the coupling portion combines the first excitation signal and the second excitation signal to generate a coupled excitation signal for said portion of the composite signal. It is configured.</p><p num="0027">Other embodiments according to the second embodiment include information related to the prediction coefficient, information related to the deterministic codebook, information related to the first gain parameter and the second gain parameter, and voiced and unvoiced signals. Provides an encoded audio signal that includes information related to the frame.</p><p num="0028">Another embodiment according to the second aspect provides a method of encoding an audio signal, a method of decoding a received audio signal, and a computer program.</p><p num="0029">Hereinafter, preferred embodiments of the present invention will be described with reference to the accompanying drawings.</p>
0030<figref num="1">A schematic block diagram of a encoder that encodes an audio signal according to an embodiment of the first aspect is shown.</figref><figref num="2">FIG. 6 shows a schematic block diagram of a decoder that decodes a received input signal according to an embodiment of the first aspect.</figref><figref num="3">FIG. 3 shows a schematic block diagram of a further encoder that encodes an audio signal according to one embodiment of the first aspect.</figref><figref num="4">A schematic block diagram of a encoder including a gain parameter calculation unit different from that shown in FIG. 3 according to an embodiment of the first aspect is shown.</figref><figref num="5">A schematic block diagram of a gain parameter calculation unit configured to calculate first gain parameter information and shape a code excitation signal according to an embodiment of the second aspect is shown.</figref><figref num="6">FIG. 3 shows a schematic block diagram of a encoder that encodes an audio signal and includes a gain parameter calculator shown in FIG. 5, according to an embodiment of the second aspect.</figref><figref num="7">FIG. 6 shows a schematic block diagram of a gain parameter calculator that includes an additional shaper configured to shape a noise-like signal, unlike the example of FIG. 5, according to one embodiment of the second aspect.</figref><figref num="8">A schematic block diagram of an unvoiced coding scheme for CELP according to one embodiment of the second aspect is shown.</figref><figref num="9">A schematic block diagram of parametric unvoiced coding according to one embodiment of the first aspect is shown.</figref><figref num="10">FIG. 6 shows a schematic block diagram of a decoder that decodes an encoded audio signal according to an embodiment of the second aspect.</figref><figref num="11a">A schematic block diagram of a shaper having a structure different from that of the shaper shown in FIG. 2 according to an embodiment of the first aspect is shown.</figref><figref num="11b">A schematic block diagram of a further shaper having a structure further different from that of the shaper shown in FIG. 2 according to one embodiment of the first aspect is shown.</figref><figref num="12">A schematic flowchart of a method of encoding an audio signal according to an embodiment of the first aspect is shown.</figref><figref num="13">A schematic flowchart of a method of decoding a received audio signal including a prediction coefficient and a gain parameter according to an embodiment of the first aspect is shown.</figref><figref num="14">A schematic flowchart of a method of encoding an audio signal according to an embodiment of the second aspect is shown.</figref><figref num="15">A schematic flowchart of a method of decoding a received audio signal according to an embodiment of the second aspect is shown.</figref><figref num="16">It is a schematic block diagram of a parametric silent coding scheme.</figref>
0031The same or equivalent components or components having the same or equivalent function are shown using the same or equivalent reference numerals in the following description, even if they are described in different drawings.
0032In the following description, many details will be given to more fully illustrate the embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention can be practiced without these special details. In other examples, known structures and devices are shown in the form of block diagrams rather than details for the purpose of preventing ambiguity in embodiments of the present invention. In addition, the features of the different embodiments described below may be combined with each other unless otherwise stated that they cannot be combined.
0033In the following description, the modification of the audio signal will be described. The audio signal may be modified by amplifying and / or attenuating a portion of the audio signal. The portion of the audio signal may be, for example, a sequence of audio signals in the time domain and / or a spectrum in the frequency domain. For frequency domains, the spectrum may be modified by amplifying or attenuating spectral values located within or above the frequency or frequency domain. Modifying the spectrum of an audio signal can include a series of operations, such as amplification and / or attenuation of the first frequency or frequency domain, and subsequent amplification and / or attenuation of the second frequency or frequency domain. Modifications in the frequency domain may be expressed as, for example, multiplication, division, summing or other calculation of the spectral value and the gain and / or attenuation values. The modifications may be performed in sequence, for example, first multiplying the spectral value by the first multiplication value and then by the second multiplication value. Multiplying by the second multiplication value first and then by the first multiplication value can result in the same or substantially the same result. Also, the first and second multiplication values may be combined first and then applied to the spectral values as the combined multiplication value, which will also receive the same or comparable results of the operation. obtain. Thus, the modification steps configured to form or modify the spectrum of an audio signal as described below are not limited to the order in which they are described, but can be performed in a modified order. On the other hand, it is possible to receive the same result and / or effect.
0034FIG. 1 shows a schematic block diagram of a encoder 100 that encodes an audio signal 102. The encoder 100 includes a frame construction unit 110 configured to generate a frame sequence 112 based on the audio signal 102. Column 112 contains a plurality of frames, and each frame of the audio signal 102 contains a length (duration) in the time domain. For example, each frame may include a length of 10ms, 20ms or 30ms.
0035The encoder 100, the prediction coefficient from one frame of the audio signal includes an analysis unit 120 configured to derive the number and (LPC = linear predictive coefficients) 122 and the residual signal 124. The frame construction unit 110 or the analysis unit 120 is configured to determine the representation of the audio signal 102 in the frequency domain. Alternatively, the audio signal 102 may already be a representation in the frequency domain.
0036The prediction coefficient 122 may be, for example, a linear prediction coefficient. Alternatively, the non-linear prediction may be applied so that the predictor 120 determines the non-linear prediction coefficient. One of the advantages of linear prediction is that the amount of calculation required to determine the prediction coefficient can be reduced.
0037The encoder 100 includes a voiced / unvoiced determination unit 130 configured to determine whether the residual signal 124 is determined from an unvoiced audio frame. The determination unit 130 supplies the residual signal to the voiced frame coder 140 when the residual signal 124 is determined from the voiced signal frame, and when the residual signal 124 is determined from the unvoiced audio frame, the residual signal is supplied to the voiced frame coder 140. It is configured to supply the residual signal to the gain parameter calculation unit 150. In order to determine that the residual signal 124 is determined from a voiced or unvoiced signal frame, the determination unit 130 may use various methods such as autocorrelation of a sample of the residual signal. A method for determining whether a signal frame is voiced or unvoiced is provided, for example, in Standard G.718 of the ITU (International Telecommunication Union) -T (Telecommunications Standards Division). The large amount of energy allocated to the low frequencies can indicate the voiced portion of the signal. Alternatively, unvoiced signals can result in the presence of large amounts of energy at high frequencies.
0038The encoder 100 includes a formant information calculation unit 160 configured to calculate speech-related spectral shaping information from a prediction coefficient 122.
0039The speech-related spectral shaping information may take formant information into account, for example, by determining the frequency or frequency domain of the processed audio frame that contains more energy than the surrounding frame. The spectrum shaping information can divide the speech magnitude spectrum into frequency domains of formants or humps and non-formants or valleys. The formant region of the spectrum can be derived, for example, by using the Immittance Spectral Frequency (ISF) or Line Spectral Frequency (LSF) representation with a prediction factor of 122. In fact, the ISF or LSF represents the frequency at which the synthetic filter using the prediction factor 122 resonates.
0040The speech-related spectrum shaping information 162 and the silent residual are output to the gain parameter calculation unit 150, which calculates the gain parameter g from the silent residual signal and the spectrum shaping information 162.<sub>n</sub>Is configured to calculate. Gain parameter g<sub>n</sub>May be one or more scalar values. That is, the gain parameter may include a plurality of values related to the amplification or attenuation of the spectral values within the plurality of frequency domains of the spectrum of the signal to be amplified or attenuated. The decoder refers to the information in the received encoded audio signal so that multiple parts of the received encoded audio signal are amplified or attenuated in the process of decoding based on the gain parameters. Gain parameter g<sub>n</sub>May be configured to apply. The gain parameter calculation unit 150 determines the gain parameter g.<sub>n</sub>May be configured to be determined by one or more mathematical representations or decision rules that result in continuous values. For example, an operation performed digitally using a processor expresses the result of using a limited number of bits to produce a variable, and is a quantized gain.<img id="000007" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />May bring. Alternatively, the result may be further quantized according to the quantization scheme so that some quantized gain information is obtained. Therefore, the encoder 100 may include a quantization unit 170. The quantization unit 170 determines the gain parameter g.<sub>n</sub>May be configured to quantize to the closest digital value supported by the digital operation of the encoder 100. Alternatively, the quantization unit 170 is a gain factor g that has already been digitized and thus quantized.<sub>n</sub>It may be configured to apply a quantization function (linear or non-linear) to. The non-linear quantization function may take into account the logarithmic dependence of human hearing, which exhibits high sensitivity at low sound pressure levels and lower sensitivity at high sound pressure levels, for example.
0041The encoder 100 may further include an information derivation unit 180 configured to derive the prediction coefficient related information 182 from the prediction coefficient 122. Prediction factors, such as the linear prediction factor used to excite innovative codebooks, have low robustness to distortion or error. So, for example, the linear prediction factor<u style="single">Immittance</u>It is known to convert to spectral frequency (ISF) and / or derive a line spectral pair (LSP) and transmit the associated information along with the encoded audio signal. LSP and / or ISF information has a higher robustness to distortions in the transmission medium, such as errors and calculation errors. The information derivation unit 180 is an LSF and / or<u style="single">ISF</u>With respect to information, it may further include a quantized unit configured to provide quantized information.
0042Alternatively, the information derivation unit may be configured to transfer a prediction factor of 122. Alternatively, the encoder 100 may be implemented without the information derivation unit 180. Alternatively, the quantization unit may be a functional block of the gain parameter calculation unit 150 or the bitstream formation unit 190, whereby the bitstream formation unit 190 causes the gain parameter g.<sub>n</sub>And quantized gain based on it<img id="000008" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />May be derived. Alternatively, the gain parameter g<sub>n</sub>If is already quantized, the encoder 100 may be realized without the quantization unit 170.
0043The encoder 100 receives the voiced information 142 associated with each voiced frame of the voiced signal, i.e. the encoded audio signal, and provided by the voiced frame coder 140, and the quantized gain.<img id="000009" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />And the prediction coefficient related information 182, and includes a bitstream forming unit 190 configured to form an output signal 192 based on them.
0044The encoder 100 may be a part of a device including a voice coding device such as a fixed or mobile phone and a microphone for transmitting an audio signal such as a computer or a tablet PC. The output signal 192 or a signal derived from the output signal 192 may be transmitted, for example, via mobile communication (wireless) or via wired communication such as a network signal.
0045The advantage of this encoder 100 is that the output signal 192 has a quantized gain.<img id="000010" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />It may include information derived from the spectral shaping information converted into. This allows the decoding of the output signal 192 to achieve or acquire further information related to the speech so that the acquired and decoded signal has a high quality with respect to the perceived level of speech quality. As such, it becomes possible to decode the signal.
0046FIG. 2 shows a schematic block diagram of a decoder 200 that decodes the received input signal 202. The received input signal 202 may correspond to, for example, the output signal 192 supplied by the encoder 100, the output signal 192 being encoded by a high level layer encoder and transmitted via a medium. It may be an input signal 202 to the decoder 200 that has been received by a receiver that decodes at a higher layer.
0047The decoder 200 includes a bitstream deformer (demultiplexer, DE-MUX) that receives the input signal 202. The Bitstream Deformer 210 has a prediction factor of 122 and a quantized gain.<img id="000011" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />And voiced information 142. To obtain the prediction factor 122, the bitstream deformer may include an inverse information derivation unit that performs the opposite operation when compared to the information derivation unit 180. Alternatively, the decoder 200 may include an inverse information derivation unit (not shown) configured to perform the reverse operation of the information derivation unit 180. In other words, the prediction coefficients are decoded, i.e. restored.
0048The decoder 200 includes a formant information calculator 220 configured to calculate speech-related spectral shaping information from a prediction factor 122, as described above for the formant information calculator 160. The formant information calculation unit 220 is configured to provide speech-related spectrum shaping information 222. Alternatively, the input signal 202 may include speech-related spectrum shaping information 222, but instead of speech-related spectrum shaping information 222, prediction coefficients or related information such as quantized LSF and / or By transmitting ISF or the like, it is possible to lower the bit rate of the input signal 202.
0049The decoder 200 includes a random noise generator 240 configured to generate a noise-like signal, which may simply be referred to as a noise signal. The random noise generator 240 may be configured to reproduce, for example, the noise signal acquired when measuring and storing the noise signal. The noise signal may be measured and recorded, for example by generating thermal noise in a resistor or other electrical component and storing the recorded data in memory. The random noise generator 240 is configured to provide a noise (state) signal n (n).
0050The decoder 200 includes a shaper 250 including a shaping unit 252 and a variable amplification unit 254. The shaper 250 is configured to spectrally shape the spectrum of the noise signal n (n). The shaping processing unit 252 receives the spectrum shaping information related to the speech, and further, for example, by multiplying the spectrum value of the spectrum of the noise signal n (n) by the value of the spectrum shaping information, the spectrum of the noise signal n (n). Is configured to shape. This operation can also be performed in the time domain by convolving the noise signal n (n) with a filter given by the spectral shaping information. The shaping processing unit 252 is configured to provide the shaped noise signal 256 and its spectrum to the variable amplification unit 254, respectively. The variable amplification unit 254 has a gain parameter g.<sub>n</sub>Is received and the spectrum of the shaped noise signal 256 is amplified to obtain the amplified shaped noise signal 258. The amplification unit adds the gain parameter g to the spectral value of the formatted noise signal 256.<sub>n</sub>It may be configured to multiply the value of. As described above, in the shaper 250, the variable amplification unit 254 receives the noise signal n (n), supplies the amplified noise signal to the shaping processing unit 252, and the shaping processing unit 252 amplifies the noise. It may be configured to shape the signal. Alternatively, the shaping unit 252 uses the speech-related spectrum shaping information 222 and the gain parameter g.<sub>n</sub>And may be applied in sequence to the noise signal n (n), or both information may be combined, for example by multiplication or other computational method. The combined parameters may be applied to the noise signal n (n).
0051The noise-like signal n (n) shaped by the speech-related spectral shaping information or its amplified version ensures that the decoded audio signal 282 contains better speech-related (natural) voice quality. Can be. This makes it possible to obtain a high quality audio signal and / or reduce the bit rate on the encoder side and maintain or enhance the output signal 282 in the reduced range on the decoder side. To enable.
0052The decoder 200 receives the prediction coefficient 122 and the amplified shaped noise signal 258, and synthesizes the synthesizer 260 configured to synthesize the composite signal 262 from the amplified shaped noise signal 258 and the prediction coefficient 122. Including. The compositing unit 260 may include a filter and may be configured to adapt the filter to a prediction factor. The compositing unit may be configured to filter the amplified shaped noise signal 258 using a filter. The filter may be configured as a software or hardware structure and may include an infinite impulse response (IIR) or finite impulse response (FIR) structure.
0053The composite signal corresponds to the silent decoded frame of the output signal 282 of the decoder 200. The output signal 282 includes a sequence of frames that can be converted into a continuous audio signal.
0054The bitstream deformer 210 is configured to separate and supply the voiced information signal 142 from the input signal 202. The decoder 200 includes a voiced frame decoder 270 configured to provide voiced frames based on its voiced information (signal) 142. The voiced frame decoder (voiced frame processing unit) is configured to determine the voiced signal 272 based on the voiced information (signal) 142. The voiced signal 272 may correspond to the voiced audio frame and / or the voiced residual of the decoder 100.
0055The decoder 200 includes a coupling unit 280 configured to combine the unvoiced decoded frame 262 and the voiced frame 272 to obtain the decoded audio signal 282.
0056Alternatively, the shaper 250 may be implemented without an amplifier, in which case the shaper 250 is configured to shape the spectrum of the noise-like signal n (n) and further amplify the acquired signal. There is no. This can reduce the amount of information transmitted by the input signal 222, thus allowing for a reduced bit rate or shorter duration of the sequence of input signals 202. Alternatively or additionally, the decoder 200 may be configured to decode only unvoiced frames, spectrally shape the noise signal n (n) and synthesize synthetic signal 262 for voiced and unvoiced frames. By doing so, it may be configured to handle both voiced and unvoiced frames. In this case, the decoder 200 can be configured without the voiced frame decoder 270 and / or without the coupling 280, resulting in reduced complexity of the decoder 200.
0057The output signal 192 and / or the input signal 202 includes information related to the prediction factor 122, information about voiced and unvoiced frames such as a flag indicating whether the processed frame is voiced or unvoiced, and a coded voiced signal. Contains additional information related to voiced signal frames such as. The output signal 192 and / or the input signal 202 further includes a gain parameter or a quantized gain parameter for the silent frame, the silent frame having a prediction factor 122 and a gain parameter g.<sub>n</sub>,<img id="000012" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />It may be configured to be decoded based on each of the above.
0058FIG. 3 shows a schematic block diagram of a encoder 300 that encodes the audio signal 102. The encoder 300 is configured to determine the linear prediction coefficient 322 and the residual signal 324 by applying the filter A (z) to the frame construction unit 110 and the frame sequence 112 output by the frame construction unit 110. Includes the predicted unit 320. The encoder 300 includes a determination unit 130 and a voiced frame coder 140 for acquiring voiced signal information 142. The encoder 300 further includes a formant information calculation unit 160 and a gain parameter calculation unit 350.
0059The gain parameter calculation unit 350 uses the gain parameter g as described above.<sub>n</sub>Is configured to provide. The gain parameter calculation unit 350 includes a random noise generation unit 350a that generates a coded noise-like signal 350b. The gain parameter calculation unit 350 further includes a shaper 350c having a shaping processing unit 350d and a variable amplification unit 350e. The shaping processing unit 350d receives the speech-related shaping information 162 and the noise-like signal 350b, and shapes the spectrum of the noise-like signal 350b using the speech-related spectrum shaping information 162 as described above for the shaper 250. It is configured. The variable amplification unit 350e transmits the shaped noise-like signal 350f to the gain parameter g, which is a temporary gain parameter received from the control unit 350k.<sub>n</sub>It is configured to be amplified using (temp). The variable amplification unit 350e is further configured to provide the amplified shaped noise signal 350g as described above for the amplified noise signal 258. As described above for the shaper 250, the order in which the noise-like signals are shaped and amplified may be combined or changed differently from FIG.
0060The gain parameter calculation unit 350 includes a comparison unit 350h configured to compare the silent residual provided by the determination unit 130 with the amplified shaped noise-like signal 350g. The comparison section is configured to obtain a measure of similarity between the silent residuals and the amplified shaped noise signal 350g. For example, the comparison unit 350h may be configured to determine the cross-correlation of both signals. Alternatively or additionally, the comparison unit 350h may be configured to compare the spectral values of both signals at some or all frequency bins. The comparison unit 350h is further configured to acquire the comparison result 350i.
0061The gain parameter calculation unit 350 uses the gain parameter g based on the comparison result 350i.<sub>n</sub>Includes a control unit 350k configured to determine (temp). For example, if the comparison result 350i shows that the amplified shaped noise-like signal contains an amplitude or magnitude lower than the corresponding amplitude or magnitude of the silent residual, the control unit controls the amplified noise-like signal. Gain parameter g for some or all frequencies of 350g<sub>n</sub>It may be configured to increase one or more values of (temp). Alternatively or additionally, if the comparison result 350i indicates that the magnitude or amplitude of the amplified shaped noise signal is too high, that is, the loudness of the amplified shaped noise signal is too high, control. The part is the gain parameter g<sub>n</sub>It may be configured to reduce one or more values of (temp). The random noise generation unit 350a, the shaper 350c, the comparison unit 350h, and the control unit 350k have a gain parameter g.<sub>n</sub>It may be configured to perform closed-loop optimization to determine (temp). A measure of similarity between unvoiced residuals and an amplified shaped noise-like signal of 350 g, for example, if a measure expressed as the difference between both signals indicates that the similarity exceeds a certain threshold. The control unit 350k has a determined gain parameter g.<sub>n</sub>Is configured to provide. The quantization unit 370 has this gain parameter g.<sub>n</sub>Quantized and quantized gain parameter<img id="000013" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Is configured to obtain.
0062The random noise generation unit 350a may be configured to supply Gaussian noise. The random noise generator 350a is configured to operate (call) the random generator with n uniform distributions between the lower limit (minimum value) such as -1 and the upper limit (maximum value) such as +1. May be good. For example, the random noise generator 350 is configured to call the random generator three times. The digitally configured random noise generator may output a pseudo-random value, and it may be possible to obtain a sufficiently randomly distributed function by adding or superimposing a plurality of or a large number of pseudo-random functions. .. This procedure follows the Central Limit Theorem. The random noise generator 350a may be configured to call the random generator at least two, three or more times, as shown in the pseudo code below.
0063[Number 6]<img id="000014" he="40" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0064Alternatively, the random noise generator 350a may generate a noise-like signal from the memory, as described for the random noise generator 240. Alternatively, the random noise generator 350a provides, for example, electrical resistance or other means for generating a noise signal by executing some code or measuring a physical effect such as thermal noise. It may be included.
0065Shape processing department<u style="single">350d</u>May be configured to add a formant structure and a slope to the noise signal 350b by filtering the noise signal 350b using fe (n) as described above. The slope may be added by filtering the signal using a filter t (n) containing a transfer function based on the following equation. [Number 7]<img id="000015" he="15" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, the factor β may be estimated from the voicing of the previous subframe. [Number 8]<img id="000016" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Where AC is an abbreviation for adaptive codebook and IC is an abbreviation for innovative codebook. [Number 9]<img id="000017" he="15" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0066Gain parameter g<sub>n</sub>And quantized gain parameters<img id="000018" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Is capable of providing additional information that can reduce errors or mismatches between the coded signal and the corresponding decoded signal decoded by a decoder such as the decoder 200, respectively. It is something to do.
0067Regarding the judgment rule of the following formula [Number 10]<img id="000019" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />The parameter w1 may include a positive non-zero value up to 1.0, preferably at least 0.7 and up to 0.8, and even more preferably a value of 0.75. The parameter w2 may include a positive non-zero scalar value up to 1.0, preferably at least 0.8 and up to 0.93, and even more preferably 0.9. The parameter w2 is preferably greater than w1.
0068FIG. 4 shows a schematic block diagram of the encoder 400. The encoder 400 is configured to provide voiced signal information 142 as described above with respect to the encoders 100 and 300. Compared to the encoder 300, the encoder 400 includes a different gain parameter calculator 350'. The comparison unit 350h'is configured to compare the audio frame 112 with the composite signal 350l' to obtain a comparison result 350i'. The gain parameter calculator 350'includes a synthesizer 350 m'configured to synthesize a composite signal 350 l'based on the amplified shaped noise signal 350 g and a prediction factor 122.
0069Basically, the gain parameter calculation unit 350'consists at least a part of the decoder by synthesizing the combined signal 350l'. When compared to the encoder 300, which includes a comparison unit 350h configured to compare the unvoiced residuals with the amplified shaped noise-like signal, the encoder 400 combines the (possibly complete) audio frame with the composite signal. Includes a comparison unit 350h'configured for comparison. Higher accuracy can be achieved because the frames of the signal and those containing their parameters are compared to each other. Comparing both signals is more complex and more accurate because the audio frame 122 and the composite signal 350l'can contain a higher degree of complexity compared to the residual signal and amplified shaped noise information. It may require a larger amount of computation. In addition, a calculation amount is required for the calculation of the composition by the composition unit 350m'.
0070The gain parameter calculation unit 350'is a coded gain parameter g.<sub>n</sub>Or its quantized version<img id="000020" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Includes a memory 350n'configured to record coding information including. As a result, the control unit 350k can acquire the stored gain value when processing the subsequent audio frame. For example, the control unit has a first value (the first set of values), i.e. g for the previous audio frame.<sub>n</sub>Gain factor g based on or equal to the value of<sub>n</sub>It may be configured to determine the first example of (temp).
0071FIG. 5 shows the first gain parameter information g according to one embodiment of the second aspect.<sub>n</sub>The schematic block diagram of the gain parameter calculation unit 550 configured to calculate is shown. The gain parameter calculation unit 550 includes a signal generation unit 550a configured to generate an excitation signal c (n). The signal generator 550a includes a deterministic codebook and an index within it to generate the signal c (n). That is, the input information such as the prediction coefficient 122 brings about a deterministic excitation signal c (n). The signal generator 550a may be configured to generate the excitation signal c (n) according to an innovative codebook of CELP coding schemes. The codebook may be determined or trained according to the speech data measured in the preceding calibration step. The gain parameter calculator includes a shaper 550b configured to shape the spectrum of the code signal c (n) based on the speech-related shaping information 550c for the code signal c (n). The speech-related shaping information 550c may be obtained from the formant information calculation unit 160. The shaper 550b includes a shape processing unit 550d configured to receive shaping information 550c for shaping the code signal. The shaper 550b further includes a variable amplification unit 550e configured to amplify the shaped code signal c (n) and acquire the amplified shaped code signal 550f. Thus, the code gain parameter is configured to define the code signal c (n) associated with the deterministic codebook.
0072The gain parameter calculation unit 550 includes a noise generation unit 350a configured to provide a noise (like) signal n (n) and a noise gain parameter g.<sub>n</sub>Includes an amplification unit 550g configured to amplify the noise signal n (n) based on the above to obtain the amplified noise signal 550h. The gain parameter calculation unit includes a coupling unit 550i configured to combine the amplified shaped code signal 550f and the amplified noise signal 550h to obtain a coupled excitation signal 550k. The coupling portion 550i may be configured to, for example, spectrally add or multiply the spectral values of the amplified shaped code signal 550f and the amplified noise signal 550h. Alternatively, the coupling 550i may be configured to convolve both signals 550f and 550h.
0073As described above with respect to the shaper 350c, the shaper 550b may be configured such that the code signal c (n) is first amplified by the variable amplification unit 550e and then shaped by the shaping processing unit 550d. Alternatively, the shaping information 550c for the code signal c (n) is the code gain parameter information g.<sub>c</sub>And the coupling information may be applied to the code signal c (n).
0074The gain parameter calculation unit 550 includes a comparison unit 550l configured to compare the combined excitation signal 550k with the unvoiced residual signal acquired by the voiced / unvoiced determination unit 130. The comparison unit 550l is the comparison unit.<u style="single">350h</u>It may be configured to provide a comparison result, i.e. a measure of similarity between the combined excitation signal 550k and the unvoiced residual signal 550m. The code gain calculation unit uses the code gain parameter information g.<sub>c</sub>And noise gain parameter information g<sub>n</sub>Includes a control unit 550n configured to control. Code gain parameter g<sub>c</sub>And noise gain parameter information g<sub>n</sub>Is related to the frequency domain of the noise signal n (n) or the signal derived from it, or the spectrum of the code signal c (n) or the signal derived from it. It may contain a value or an imaginary value.
0075Alternatively, the gain parameter calculation unit 550 may be configured without the shaping processing unit 550d. Alternatively, the shaping processing unit 550d may be configured to shape the noise signal n (n) and provide the shaped noise signal to the variable amplification unit 550g.
0076Thus, both gain parameter information g<sub>c</sub>And g<sub>n</sub>By controlling, the similarity between the combined excitation signal 550k and the silent residual becomes high, and as a result, the code gain parameter information g<sub>c</sub>And noise gain parameter information g<sub>n</sub>The decoder that receives the information about will be able to reproduce the audio signal with good voice quality. The control unit 550n uses the code gain parameter information g.<sub>c</sub>And noise gain parameter information g<sub>n</sub>It is configured to provide an output signal 550o containing information about. For example, the signal 550o has both gain parameter information g<sub>n</sub>And g<sub>c</sub>May be included as a scalar value or a quantized value, or as a value derived from them, for example, an encoded value.
0077FIG. 6 shows a schematic block diagram of a encoder 600 that encodes the audio signal 102 and includes the gain parameter calculator 550 as shown in FIG. The encoder 600 can be obtained, for example, by modifying the encoder 100 or 300. The encoder 600 includes a first quantization unit 170-1 and a second quantization unit 170-2. The first quantization unit 170-1 is the gain parameter information g.<sub>c</sub>Quantized and quantized gain parameter information<img id="000021" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Is configured to get. The second quantization unit 170-2 is the noise gain parameter information g.<sub>n</sub>Quantized and quantized noise gain parameter information<img id="000022" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Is configured to get. The bitstream forming unit 690 includes voiced signal information 142, LPC-related information 122, and both quantized gain parameter information.<img id="000023" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />And are configured to generate an output signal 692 containing. Compared to the output signal 192, the output signal 692 has quantized gain parameter information.<img id="000024" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Has been expanded or upgraded by. Alternatively, the quantization unit 170-1 and / or 170-2 may be part of the gain parameter calculation unit 550. Further, one of the quantization units 170-1 and / or 170-2 has both quantized gain parameters.<img id="000025" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />May be configured to obtain.
0078Alternatively, the encoder 600 has code gain parameter information g.<sub>c</sub>And noise gain parameter information g<sub>n</sub>Quantized and quantized parameter information<img id="000026" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />May be configured to include one quantization unit configured to obtain. Both gain parameter information may be quantized sequentially, for example.
0079The formant information calculation unit 160 is configured to calculate speech-related spectrum shaping information 550c from a prediction coefficient 122.
0080FIG. 7 shows a schematic block diagram of the gain parameter calculation unit 550'corrected as compared with the gain parameter calculation unit 550. The gain parameter calculation unit 550'is replaced by the shaper shown in FIG. 3 instead of the amplification unit 550 g.<u style="single">350c</u>including. Shaper<u style="single">350c</u>Is configured to provide an amplified, well-formed noise signal of 350g. The coupling portion 550i is configured to combine the amplified shaped code signal 550f and the amplified shaped noise signal 350g to provide a coupled excitation signal 550k'. The formant information calculator 160 is configured to provide both speech-related formant information 162 and 550c. The speech-related formant information 550c and 162 may be the same. Alternatively, both information 550c and 162 may be different from each other. This allows individual modeling, or shaping, of the code-generated signals c (n) and n (n).
0081The control unit 550n sets the gain parameter information g for each subframe of the processed audio frame.<sub>c</sub>And g<sub>n</sub>And may be configured to determine. The control unit has gain parameter information g based on the following details.<sub>c</sub>And g<sub>n</sub>And may be configured to determine, i.e., calculate.
0082First, the average energy of the subframes may be calculated for the original short-term predictive residual signal that can be used during the LPC analysis, i.e. the silent residual signal. Its energy is averaged in the log domain by the following equation over the four subframes of the current frame. [Number 11]<img id="000027" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0083Where Lsf is the size of the subframe in the sample. In this case, the frame is divided into four subframes. The averaged energy may then be encoded using some pre-trained stochastic codebook, for example with some bits such as 3, 4 or 5. .. Stochastic codebooks have several different values that can be represented by the number of bits, for example, size 8 for a number of 3 bits, size 16 for a number of 4 bits, or size 32 for a number of 5 bits. It may contain several entries (sizes) according to. Quantized gain<img id="000028" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />May be determined from the selected codewords in the codebook. Two gain information g for each subframe<sub>c</sub>And g<sub>n</sub>Is calculated. Code g<sub>c</sub>The gain of may be calculated based on, for example, the following equation. [Number 12]<img id="000029" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, cw (n) is, for example, a fixed excitation selected from a fixed codebook included in the signal generation 550a and filtered by a perceptually weighted filter. The display xw (n) corresponds to the conventional perceptual target excitation calculated within the CELP encoder. Code gain information g<sub>c</sub>Next, the normalized gain g<sub>nc</sub>To obtain, it may be normalized based on the following equation. [Number 13]<img id="000030" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0084Normalized gain g<sub>nc</sub>May be quantized by, for example, the quantization unit 170-1. Quantization may be performed according to a linear or logarithmic scale. The logarithmic scale may include a scale with a size of 4, 5 or more bits. For example, the logarithmic scale contains a size of 5 bits. Quantization may be performed based on the following equation. [Number 14]<img id="000031" he="15" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, if the logarithmic scale contains 5 bits, Index<sub>nc</sub>May be limited between 0 and 31. Index<sub>nc</sub>May be quantized gain parameter information. code<img id="000032" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />The quantized gain of can then be expressed based on the following equation. [Number 15]<img id="000033" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0085The gain of the code may be calculated for the purpose of minimizing the mean squared error or mean squared error (MSE) of the following equation. [Number 16]<img id="000034" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, Lsf corresponds to the line spectral frequency determined from the prediction factor 122.
0086The noise gain parameter information may be determined with respect to the energy mismatch by minimizing the error based on the following equation. [Number 17]<img id="000035" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />
0087The variable k is an attenuation factor that can change depending on or based on the prediction factor, where the prediction factor is that the speech contains a small amount of background noise or even no background noise (clean speech). Allows the determination. Alternatively, if the audio signal or its frame contains a change between unvoiced and unvoiced frames, the signal may be determined as a noisy speech. The variable k can be set to a value of at least 0.85, a value of at least 0.95, or even a value of 1 for clean speech, where high-energy dynamics are perceptually important. The variable k can be set to a value of at least 0.6 and up to 0.9, preferably at least 0.7 and up to 0.85, and even more preferably 0.8, for noisy speech, in which case unvoiced. Noise excitation is made more modest to prevent variations in output energy between the frame and the unvoiced frame. These quantized gain candidates<img id="000036" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />An error (energy mismatch) may be calculated for each of the above. One frame divided into four subframes is four quantized gain candidates<img id="000037" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />May bring. One candidate that minimizes the error may be output by the control unit. The quantized gain of noise (noise gain parameter information) can be calculated based on the following equation. [Number 18]<img id="000038" he="20" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, Index<sub>n</sub>Is limited to between 0 and 3 by 4 candidates. The resulting combined excitation signal, such as the excitation signal 550k or 550k', can be obtained based on the following equation. [Number 19]<img id="000039" he="12" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Here, e (n) is a coupled excitation signal 550k or 550k'.
0088A encoder 600 containing a gain parameter calculator 550 or 550'or a modified encoder 600 may allow unvoiced coding based on the CELP coding scheme. The CELP coding scheme may be modified based on exemplary details such as dealing with unvoiced frames. LTP parameters are not transmitted because there is little periodicity in the silent frame and the resulting coding gain is very low. Adaptive excitation is set to zero. -Saving bits are reported to the fixed codebook. More pulses can be encoded for the same bit rate, thus improving quality. -Pulse coding is not sufficient to properly model the noise-like target excitation of silent frames at low rates, i.e. at rates of 6-12 kbps. A Gaussian codebook is added to the fixed codebook to build the final excitation.
0089FIG. 8 shows a schematic block diagram of an unvoiced coding scheme for CELP according to the second aspect. The modified control unit 810 includes the functions of both the comparison unit 550l and the control unit 550n. The control unit 810 controls the code gain parameter information g based on the analysis by synthesis, that is, by comparing the composite signal with the input signal shown as s (n), for example, an unvoiced residual.<sub>c</sub>And noise gain parameter information g<sub>n</sub>Is configured to determine. The control unit 810 generates an excitation for the signal generation unit (innovative excitation) 550a, and gain parameter information g.<sub>c</sub>And g<sub>n</sub>Includes a synthetic analysis filter 820 configured to provide. Block 810 for synthetic analysis is configured to compare the internally synthesized signal with the combined excitation signal 550k'by adapting the filter according to the parameters and information provided.
0090The control unit 810 is an analysis block configured to acquire the prediction coefficient as described above for the case where the analysis unit 320 acquires the prediction coefficient 122.<u style="single">830</u>including. The control unit further includes a composite filter 840 that filters the coupled excitation signal 550k, which is adapted by a filter factor 122. Further comparisons include the input signal s (n) and, for example, a composite signal that is a decoded (restored) audio signal.<img id="000040" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />And may be configured to compare. Further, a memory 350n is arranged, and the control unit 810 is configured to store the predicted signal and / or the predicted coefficient in the memory. The signal generator 850 is configured to provide an adaptive excitation signal based on the predictions stored in the memory 350n, thereby enhancing the adaptive excitation based on the previously coupled excitation signal. Becomes possible.
0091FIG. 9 shows a schematic block diagram of parametric unvoiced coding according to the first aspect. The amplified shaped noise signal may be the input signal of the composite filter 910 adapted by the determined filter coefficient (prediction coefficient) 122. The composite signal 912 output by the composite filter may be compared with the input signal s (n), which may be, for example, an audio signal. The composite signal 912 contains an error as compared with the input signal s (n). Noise gain parameter g by analysis block 920 which can correspond to gain parameter calculator 150 or 350<sub>n</sub>The error can be reduced or minimized by modifying. The adaptive codebook may be updated by storing the amplified preformed noise signal 350f in the memory 350n. As a result, the processing of voiced audio frames can also be enhanced based on the improved coding of unvoiced audio frames.
0092FIG. 10 shows a schematic block diagram of a decoder 1000 that decodes an encoded audio signal, for example the encoded audio signal 692. The decoder 1000 includes a signal generation unit 1010 and a noise generation unit 1020 configured to generate a noise-like signal 1022. The received signal 1002 contains LPC-related information, and the bitstream deformer 1040 is configured to provide a prediction factor 122 based on the prediction factor-related information. For example, the decoder 1040 is configured to extract a prediction factor 122. The signal generation unit 1010 is a signal generation unit.<u style="single">550a</u>As described above with respect to, it is configured to generate a code-excited excitation signal 1012. The coupling portion 1050 of the decoder 1000 is configured to combine the code-excited signal 1012 and the noise-like signal 1022 to obtain the coupled excitation signal 1052, as described above for the coupling portion 550. The decoder 1000 includes a synthesizer 1060 with a filter applied with a prediction factor of 122, which filters the combined excitation signal 1052 with the adapted filter to obtain the silent decoded frame 1062. It is configured to do so. The decoder 1000 also combines the unvoiced decoded frame with the voiced frame 272 to obtain the audio signal sequence 282.<u style="single">280</u>including. Unlike the decoder 200, the decoder 1000 includes a second signal generator configured to provide a code-excited excitation signal 1012. The noise-like excitation signal 1022 may be, for example, the noise-like signal n (n) shown in FIG.
0093The audio signal sequence 282 can have good quality and high similarity when compared to the encoded input signal.
0094Another embodiment provides a decoder that enhances the decoder 1000 by shaping and / or amplifying the code-generated (code-excited) excitation signal 1012 and / or the noise-like signal 1022. That is, the decoder 1000 may include a shaping processing unit and / or a variable amplification unit arranged between the signal generation unit 1010 and the coupling unit 1050 and between the noise generation unit 1020 and the coupling unit 1050, respectively. The input signal 1002 is the code gain parameter information g.<sub>c</sub>And / or information related to noise gain parameter information may be included and the decoder may include code gain parameter information g.<sub>c</sub>May be configured to adapt the amplification unit for amplifying the code-generated excitation signal 1012 or a formatted version thereof. Alternatively or additionally, the decoder 1000 may be configured to adapt, or control, an amplification unit for amplifying the noise-like signal 1022 or a formatted version thereof by using the noise gain parameter information.
0095Alternatively, the decoder 1000 is configured to shape the code-excited excitation signal 1012, as shown by the dotted line, and / or the shaper 1080, which is configured to shape the noise-like signal 1022. May include. The shaper 1070 and / or 1080 has a gain parameter g<sub>c</sub>And / or g<sub>n</sub>, And / or speech-related shaping information may be received. The shapers 1070 and / or 1080 may be formed in the same manner as the shapers 250, 350c and / or 550b described above.
0096The decoder 1000 may include formant information calculator 1090, which provides speech-related shaping information 1092 for shaper 1070 and / or 1080, as described above for formant information calculator 160. The formant information calculator 1090 may be configured to provide different speech-related shaping information (1092a; 1092b) to the shaping machine 1070 and / or 1080.
0097FIG. 11a shows a schematic block diagram of the shaper 250'which implements an alternative structure compared to the shaper 250. The shaper 250'has the shaping information 222 and the noise-related gain parameter g.<sub>n</sub>Includes a coupling part 257 that combines with and obtains coupled information 259. The modified shaping processing unit 252'is configured to shape the noise-like signal n (n) by using the combined information 259 to obtain an amplified shaped noise-like signal 258. Formatting information 222 and gain parameter g<sub>n</sub>Since both can be interpreted as multiplication factors, both multiplication factors may be multiplied using the coupling part 257 and then applied to the noisy signal n (n) in the coupled form.
0098FIG. 11b shows a schematic block diagram of the shaper 250'' that implements an alternative structure compared to the shaper 250. Compared to the shaper 250, the variable amplification unit 254 is placed first, and this is the gain parameter g.<sub>n</sub>Is configured to generate an amplified noise-like signal by amplifying the noise-like signal n (n) using. The shaping processing unit 252 is configured to shape the amplified signal using the shaping information 222 and acquire the amplified shaped signal 258.
0099Although FIGS. 11a and 11b describe variations thereof in relation to the shaper 250, the above description applies similarly to the shapers 350c, 550b, 1070 and / or 1080.
0100FIG. 12 shows a schematic flowchart of a method 1200 for encoding an audio signal according to the first aspect. This way<u style="single">1200</u>Derives the prediction coefficient and the residual signal from the audio signal frame.<u style="single">Step 1210</u>including.<u style="single">Method 1200 includes step 1220 of calculating speech-related spectral shaping information from prediction coefficients.</u>Method 1200 forms an output signal based on step 1230, which calculates the gain parameter from the unvoiced residual signal and spectrum shaping information, and the information, gain parameter or quantized gain parameter, and prediction coefficient associated with the voiced signal frame. Including steps 1240 and.
0101FIG. 13 shows a schematic flowchart of method 1300 for decoding a received audio signal including a prediction factor and a gain parameter according to the first aspect. The method 1300 includes step 1310 of calculating speech-related spectral shaping information from prediction coefficients. In step 1320, a decoded noise-like signal is generated. In step 1330, the spectrum of the decoded noise-like signal or its amplified representation is shaped and shaped using spectrum shaping information.<u style="single">Done</u>A decoded noise-like signal is acquired. In step 1340 of method 1300, preformed<u style="single">Decryption</u>The composite signal is synthesized from the noise-like signal and the prediction coefficient.
0102FIG. 14 shows a schematic flowchart of method 1400 for encoding an audio signal according to a second aspect. The method 1400 includes step 1410 of deriving the prediction factor and the residual signal from the silent frame of the audio signal. In step 1420 of method 1400, the first gain parameter information that defines the first excitation signal associated with the deterministic codebook and the second gain parameter information that defines the second excitation signal associated with the noisy signal are silent. Calculated for the frame.
0103In step 1430 of method 1400, the output signal is formed based on the information related to the voiced signal frame, the first gain parameter information, and the second gain parameter information.
0104FIG. 15 shows a schematic flowchart of a method 1500 of decoding a received audio signal according to a second aspect. The received audio signal contains information related to the prediction factor. Method 1500 includes step 1510 of generating a first excitation signal from a deterministic codebook for a portion of the synthesized signal. In step 1520 of method 1500, a second excitation signal is generated from the noise-like signal for that portion of the composite signal. Method<u style="single">1500</u>In step 1530, the first excitation signal and the second excitation signal are combined to generate a combined excitation signal for a portion of the composite signal. In step 1540 of method 1500, a portion of the combined signal is combined with the coupled excitation signal and the prediction factors.
0105In other words, each aspect of the invention proposes a new method of encoding unvoiced frames, in which formant structures and spectral gradients are added to shape randomly generated Gaussian noise. The spectral shaping is performed in the excitation domain before exciting the synthetic filter. As a result, the well-formed excitation will be updated in the memory of the long-term forecast to generate subsequent adaptive codebooks.
0106Subsequent frames that are not silent will also benefit from spectral shaping. Unlike formant enhancement in post-filtering, the proposed noise shaping is performed on both the encoder side and the decoder side.
0107Such excitation can be used directly in parametric coding schemes targeting very low bit rates. However, the present invention also proposes associating such excitation in combination with conventional innovative codebooks within the CELP coding scheme.
0108For both methods, the present invention proposes a new gain coding that is particularly efficient for both clean speech and speech with background noise. The present invention proposes several mechanisms that are as close as possible to the original energy, but at the same time avoid the overly jarring transitions of the unvoiced frame and also avoid the unwanted instability due to gain quantization.
0109The first aspect aims for unvoiced coding at rates of 2.8 and 4 kilobits per second (kbps). Unvoiced frames are detected first. This detection can be performed by conventional speech classification, as is performed in the variable rate multimode wideband (VMR-WB) known from Non-Patent Document 2.
0110Performing spectral shaping at this stage has two main advantages. First, spectral shaping takes into account the excitation gain calculation. Since the gain calculation is the only non-blind module in excitation generation, it is very advantageous to perform the gain calculation at the end of a series of operations after shaping. Second, it makes it possible to save enhanced excitation in LTP memory. Therefore, such enhancements will also be useful for subsequent unvoiced frames.
0111Quantized parts 170, 170-1 and 170-2 are quantized parameters.<img id="000041" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />Although explained that it is configured to obtain the quantized parameters, the quantized parameters may be provided as information related to them, that is, the entry is a quantized gain parameter.<img id="000042" he="10" wi="160" file="JP6366705B2_D0001.tif" img-format="tif" img-content="drawing" />May be provided as an index or identifier for an entry in a database containing.
0112Although several embodiments have been shown in the context of the device, these embodiments also represent a description of the corresponding method, in which one block or device corresponds to one method step or feature of the method step. Is clear. Similarly, aspects shown in the context of describing method steps also represent the corresponding block or item or feature of the corresponding device.
0113The decomposed signal of the present invention can be stored in a digital storage medium, or can be transmitted via a transmission medium such as a wireless transmission medium such as the Internet or a wired transmission medium.
0114Although it depends on a predetermined configuration requirement, the embodiment of the present invention can be configured by hardware or software. This configuration has electronically readable control signals stored therein and works (or collaborates) with a computer system programmable to perform each method of the invention. It can be executed using a digital storage medium such as a flexible disk, DVD, CD, ROM, PROM, EPROM, EEPROM, flash memory or the like.
0115Some embodiments according to the present invention include a data carrier having electronically readable control signals that can work with a computer system programmable to perform one of the methods described above.
0116In general, an embodiment of the present invention can be configured as a computer program product having a program code, the program code of which, when the computer program product operates on a computer, one of the methods of the present invention. Can be actuated to perform. The program code may be stored, for example, in a machine-readable carrier.
0117Other embodiments of the invention include a computer program stored in a machine-readable carrier for performing one of the methods described above.
0118In other words, one embodiment of the method of the invention is a computer program having program code for performing one of the methods described above when the computer program runs on a computer.
0119Another embodiment of the invention is a data carrier (or digital storage medium, or computer-readable medium) that includes a computer program recorded to perform one of the methods described above.
0120Another embodiment of the invention is a data stream or signal sequence representing a computer program for performing one of the methods described above. The data stream or signal sequence may be configured to be transmitted over a data communication connection such as the Internet.
0121Other embodiments include processing means configured or adapted to perform one of the methods described above, such as, for example, a computer or a programmable logical device.
0122Other embodiments include a computer on which a computer program for performing one of the methods described above is installed.
0123In some embodiments, programmable logic devices (such as rewritable gate arrays) may be used to perform some or all of the functions of the methods described above. In some embodiments, the rewritable gate array may work with a microprocessor to perform one of the methods described above. In general, such a method is preferably performed by any hardware device.
0124The embodiments described above merely illustrate the principles of the present invention. It will be apparent to those skilled in the art that the configurations and details described herein can be modified and modified. Therefore, the present invention is not limited by the specific details presented herein for the purposes of description and explanation of embodiments, but should be limited only by the appended claims.
78 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2011518345A | Cites | Japan |
| JP2015515644A | Cites | Japan |
81 members in 19 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 131893927 | European Patent Office (EPO) | – | |
| 13189392 | European Patent Office (EPO) | A | |
| 141787853 | European Patent Office (EPO) | – | |
| 14178785 | European Patent Office (EPO) | A | |
| 2014071769 | European Patent Office (EPO) | W |
Members81
| Document | Office | Kind | |
|---|---|---|---|
| CA2927716A1 | Canada | A1 | |
| CA2927722A1 | Canada | A1 | |
| WO2015055531A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015055532A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201523588A | Taiwan Province of China | A | |
| TW201528255A | Taiwan Province of China | A | |
| AR098072A1 | Argentina | A1 | |
| AR098073A1 | Argentina | A1 | |
| AU2014336356A1 | Australia | A1 | |
| AU2014336357A1 | Australia | A1 | |
| SG11201603000SA | Singapore | A | |
| SG11201603041YA | Singapore | A | |
| KR20160070147A | Republic of Korea | A | |
| KR20160073398A | Republic of Korea | A | |
| CN105723456A | China | A | |
| CN105745705A | China | A | |
| MX2016004922A | Mexico | A | |
| MX2016004923A | Mexico | A | |
| US2016232908A1 | United States of America | A1 | |
| US2016232909A1 | United States of America | A1 | |
| EP3058568A1 | European Patent Office (EPO) | A1 | |
| EP3058569A1 | European Patent Office (EPO) | A1 | |
| JP2016533528A | Japan | A | |
| JP2016537667A | Japan | A | |
| TWI575512B | Taiwan Province of China | B | |
| TWI576828B | Taiwan Province of China | B | |
| AU2014336356B2 | Australia | B2 | |
| AU2014336357B2 | Australia | B2 | |
| BR112016008544A2 | Brazil | A2 | |
| BR112016008662A2 | Brazil | A2 | |
| RU2016118979A | Russian Federation | A | |
| RU2016119010A | Russian Federation | A | |
| ZA201603158B | South Africa | B | |
| RU2644123C2 | Russian Federation | C2 | |
| RU2646357C2 | Russian Federation | C2 | |
| KR20180021906A | Republic of Korea | A | |
| MX355091B | Mexico | B | |
| MX355258B | Mexico | B | |
| KR101849613B1 | Republic of Korea | B1 | |
| JP6366705B2This record | Japan | B2 | |
| JP6366706B2 | Japan | B2 | |
| CA2927722C | Canada | C | |
| KR101931273B1 | Republic of Korea | B1 | |
| US10304470B2 | United States of America | B2 | |
| US2019228787A1 | United States of America | A1 | |
| US10373625B2 | United States of America | B2 | |
| US2019333529A1 | United States of America | A1 | |
| CN105723456B | China | B | |
| CN105745705B | China | B | |
| US10607619B2 | United States of America | B2 | |
| CN111370009A | China | A | |
| US2020219521A1 | United States of America | A1 | |
| CA2927716C | Canada | C | |
| MY180722A | Malaysia | A | |
| EP3058569B1 | European Patent Office (EPO) | B1 | |
| PT3058569T | Portugal | T | |
| EP3058568B1 | European Patent Office (EPO) | B1 | |
| US10909997B2 | United States of America | B2 | |
| EP3779982A1 | European Patent Office (EPO) | A1 | |
| PT3058568T | Portugal | T | |
| US2021098010A1 | United States of America | A1 | |
| EP3806094A1 | European Patent Office (EPO) | A1 | |
| PL3058569T3 | Poland | T3 | |
| ES2839086T3 | Spain | T3 | |
| PL3058568T3 | Poland | T3 | |
| ES2856199T3 | Spain | T3 | |
| MY187944A | Malaysia | A | |
| BR112016008544B1 | Brazil | B1 | |
| BR112016008662B1 | Brazil | B1 | |
| US11798570B2 | United States of America | B2 | |
| CN111370009B | China | B | |
| US11881228B2 | United States of America | B2 | |
| EP3779982B1 | European Patent Office (EPO) | B1 | |
| EP3779982C0 | European Patent Office (EPO) | C0 | |
| EP3806094B1 | European Patent Office (EPO) | B1 | |
| EP3806094C0 | European Patent Office (EPO) | C0 | |
| EP4632735A2 | European Patent Office (EPO) | A2 | |
| ES3042587T3 | Spain | T3 | |
| PL3779982T3 | Poland | T3 | |
| ES3044088T3 | Spain | T3 | |
| EP4632735A3 | European Patent Office (EPO) | A3 |
16 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6366705
- Application
- 2016524410
Titles2
- Japanese
- 確定的及びノイズ状情報を用いてオーディオ信号を符号化/復号化する概念
- English
- The concept of coding / decoding an audio signal using deterministic and noisy information
Classification
- CPC, 11
- G10L19/08
- G10L19/083
- G10L19/20
- G10L19/0017
- G10L19/008
- G10L19/12
- G10L19/06
- G10L2025/932
- G10L25/15
- G10L19/07
- G10L2019/0016
- IPC, 2
- G10L19 083
- G10L19 12
