Scalable decoding apparatus
8 claims: 1 independent, 7 dependent
- 1低周波帯域の符号化情報を復号して低周波帯域の復号信号を得る第1復号化手段と、 前記低周波帯域の復号信号と高周波帯域の符号化情報とから高周波帯域の復号信号を得る第2復号化手段と、 を具備するスケーラブル復号化装置であって、 前記第1復号化手段は、前記低周波帯域の復号信号の取得に加えて、前記低周波帯域の符号化情報を復号して前記低周波帯域のスペクトル情報を取得し、 前記第2復号化手段は、 前記低周波帯域の復号信号を変換して低周波帯域のスペクトルを得る変換手段と、 前記第1復号化手段により取得した前記低周波帯域のスペクトル情報に基づいて前記低周波帯域の復号信号のスペクトルハーモニクスの谷の深さを示す指標を決定し、予め用意された複数の振幅調整係数の中から、前記指標に応じて選ばれた振幅調整係数を用いて、前記変換手段により得られた 前記低周波帯域のスペクトルに対して振幅調整を施す調整手段と、 振幅調整された低周波帯域のスペクトルと前記高周波帯域の符号化情報とを用いて、高周波帯域のスペクトル を生 成する生成手段と、 を具備するスケーラブル復号化装置。
- 2前記生成手段は、振幅調整された低周波帯域のスペクトルにミラーリングを適用して前記高周波帯域のスペクトル を生 成する、 請求項1記載のスケーラブル復号化装置。
- 3前記生成手段は、前記高周波帯域の符号化情報の少なくとも一部が復号できない場合に、前記高周波帯域のスペクトル を生 成する、 請求項1記載のスケーラブル復号化装置。
- 4前記生成手段は、振幅調整された低周波帯域のスペクトルにピッチフィルタリング処理を適用して前記高周波帯域のスペクトル を生 成する、 請求項1記載のスケーラブル復号化装置。
- 5前記高周波帯域の符号化情報は、重要度の高い順に、スケールファクタ、振幅調整係数、ラグ、スペクトル残差、の順で構成され、 前記生成手段は、前記高周波帯域の符号化情報において前記スペクトル残差が欠落する場合に、前記スケールファクタ、前記振幅調整係数、前記ラグを用いて、前記高周波帯域のスペクトル を生 成する、 請求項1記載のスケーラブル復号化装置。
- 6前記高周波帯域の符号化情報は、重要度の高い順に、スケールファクタ、振幅調整係数、ラグ、スペクトル残差、の順で構成され、 前記生成手段は、前記高周波帯域の符号化情報において前記ラグおよび前記スペクトル残差が欠落する場合に、振幅調整された低周波帯域のスペクトルにミラーリングを適用して前記高周波帯域のスペクトル を生 成する、 請求項1記載のスケーラブル復号化装置。
- 7前記高周波帯域の符号化情報は、重要度の高い順に、スケールファクタ、振幅調整係数、ラグ、スペクトル残差、の順で構成され、 前記生成手段は、前記スケールファクタ、前記振幅調整係数、前記ラグ、前記スペクトル残差の少なくとも一つが欠落する場合に、欠落した情報に対応する過去の情報を用いて前記高周波帯域のスペクトル を生 成する、 請求項1記載のスケーラブル復号化装置。
- 8前記生成手段は、隣接フレームの信号間の相関に応じて、前記高周波帯域の符号化情報を用いて前記高周波帯域のスペクトルを生成する、 請求項1記載のスケーラブル復号化装置。
Independent claims8
182 paragraphs, as filed
The present invention is a scalable decoding device used when communicating voice signals and acoustic signals in mobile communication systems, packet communication systems using Internet protocols, and the like.<u style="single">In place</u>Related.
In order to effectively use radio resources and the like in mobile communication systems, it is required to compress audio signals at a low bit rate. On the other hand, users want to improve the quality of call voice and realize a call service with a high sense of presence. In order to realize this, it is desirable not only to improve the quality of the voice signal but also to encode a signal other than the voice such as an audio signal having a wider band with a high quality.
Furthermore, in an environment where a wide variety of networks coexist, communication between different networks, communication between terminals that use different services, communication between terminals with different processing performance, and communication between two parties only. However, there is a demand for a voice coding method that can flexibly support multipoint mutual communication.
Further, there is a demand for a voice coding method that is resistant to transmission line errors (especially packet loss in packet switching networks represented by IP networks).
One of the voice coding methods that satisfy such a requirement is a band-scalable voice coding method. The band-scalable audio coding method is a method of coding a voice signal hierarchically, and is a coding method in which the coding quality increases as the number of coding layers increases. Since the bit rate can be made variable by increasing or decreasing the number of coding layers, the transmission line capacity can be effectively used.
Further, in the band-scalable audio coding method, the decoder side only needs to be able to receive the coded data of the minimum basic layer, and can tolerate the loss of the coded information of the additional layer on the transmission line to some extent. Highly resistant to mistakes. Further, as the number of coding layers is increased, the frequency band of the voice signal to be coded expands. For example, a conventional telephone band voice coding method is used for the basic layer (core layer). In addition, an additional layer (extended layer) is configured so that wideband audio such as the 7 kHz band can be encoded.
In this way, in the band-scalable voice coding method, the telephone band voice signal is encoded in the core layer and the high-quality broadband signal is encoded in the expansion layer. Can also be used for high-quality broadband voice service terminals, and can also support multipoint communication including both terminals. Further, since the coded information is hierarchical, the error tolerance can be increased depending on the device of the transmission method, and it is also easy to control the bit rate on the coded side or the transmission path. For these reasons, the band-scalable voice coding method is attracting attention as a future communication voice coding method.
As an example of the band-scalable voice coding method as described above, the method described in Non-Patent Document 1 can be mentioned.
In the band-scalable speech coding method described in Non-Patent Document 1, the MDCT coefficient is encoded by a scale factor for each band and microstructure information. The scale factor is Huffman-coded and the microstructure is vector-quantized. The auditory importance of each band is calculated using the decoding result of the scale factor, and the bit allocation to each band is determined. The bandwidth of each band is uneven, and is preset so that the higher the band, the wider the band.
In addition, transmission information is classified into the following four groups. A: Core codec coding information B: High frequency scale factor coding information C: Low frequency scale factor coding information D: Coding information of spectral fine structure
In addition, the decoding side performs the following processing. <Case 1> If the information in A cannot be completely received, frame loss compensation processing is performed to generate decoded audio. <Case 2> When only the information of A is received, the decoding signal of the core codec is output. <Case 3> When the information of B is received in addition to the information of A, a high frequency band is generated by mirroring the decoding signal of the core codec, and a decoding signal having a wider band than the decoding signal of the core codec is generated. The decoded B information is used to generate the high-frequency spectral shape. Mirroring is performed in a voiced frame in such a way that the harmonic structure (wave control structure) does not collapse. In the silent frame, random noise is used to generate high frequencies. <Case 4> When the information of C is received in addition to the information of A and B, the same decoding process as in case 3 is performed only with the information of A and B. <Case 5> When the information of D is received in addition to the information of A, B, and C, the complete decoding process is performed in the band where all the information of A to D can be received, and the band where the information of D cannot be received is low. The fine spectrum is decoded by mirroring the decoded signal spectrum on the region side. Since the information of B and C is received even if the information of D is not received, the information of B and C is used for decoding the spectrum envelope information. Mirroring is performed in a voiced frame in such a way that the harmonic structure (wave control structure) does not collapse. In the silent frame, random noise is used to generate high frequencies.<nplcit num="1"><text>B. Kovesiet al, A scalable speech and audio coding scheme with continuous bitrate flexibility, in proc. IEEE ICASSP 2004, pp.I-273--I-276</text></nplcit>
<p> In the above-mentioned prior art (Non-Patent Document 1), a high region is generated by mirroring. At this time, since mirroring is performed so as not to break the wave-tuning structure, the wave-tuning structure is maintained. However, the low-frequency toning structure becomes a mirror image and appears in the high-frequency. In general, in a voiced signal, the tuning structure collapses toward higher frequencies, so that in the high frequencies, the tuning structure is not as remarkable as in the low frequencies. In other words, even if the valley of harmonics is deep in the low range, the valley of harmonics may be shallow in the high range, and in some cases, the tuning structure itself may not be clear. Therefore, in the above-mentioned conventional technique, an excessive wave-tuning structure tends to appear in the high-frequency component, and therefore, the quality of the decoded audio signal deteriorates.</p><p> An object of the present invention is to obtain a high-quality decoded audio (acoustic) signal with little deterioration of the high-frequency spectrum even when the audio (acoustic) signal is decoded by generating a high-frequency spectrum using the low-frequency spectrum. Scalable decoding equipment that can be obtained<u style="single">Place</u>To provide.</p>
<p> The scalable decoding device of the present invention includes a first decoding means for decoding low-frequency band coding information to obtain a low-frequency band decoding signal, and the low-frequency band decoding signal and high-frequency band coding information. A scalable decoding apparatus including a second decoding means for obtaining a decoding signal in a high frequency band from the same.<u style="single">In addition to acquiring the decoding signal in the low frequency band, the first decoding means decodes the coding information in the low frequency band to acquire the spectral information in the low frequency band.</u>The second decoding means includes a conversion means that converts the decoding signal in the low frequency band to obtain a spectrum in the low frequency band.<u style="single">Based on the spectrum information of the low frequency band acquired by the first decoding means, an index indicating the depth of the spectral harmonics valley of the decoded signal of the low frequency band is determined, and a plurality of amplitude adjustment coefficients prepared in advance. Obtained by the conversion means using the amplitude adjustment coefficient selected according to the index from among the above.</u>A high-frequency band spectrum using an adjusting means for adjusting the amplitude of the low-frequency band spectrum, an amplitude-adjusted low-frequency band spectrum, and the high-frequency band coding information.<u style="single">Raw</u>A configuration is adopted that includes a generation means to be formed.</p>
<p> According to the present invention, even when a voice (acoustic) signal is decoded by generating a high-frequency spectrum using a low-frequency spectrum, a high-quality decoded audio (acoustic) signal with little deterioration of the high-frequency spectrum can be obtained. Obtainable.</p>
Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
(Embodiment 1) FIG. 1 is a block diagram showing a configuration of a scalable decoding device 100 that forms, for example, a band scalable audio (acoustic) signal decoding device.
The scalable decoding device 100 includes a separation unit 101, a first layer decoding unit 102, and a second layer decoding unit 103.
The separation unit 101 receives the bit stream transmitted from the scalable coding device described later, separates it into a coding parameter for the first layer and a coding parameter for the second layer, and separates the first layer decoding unit 102. And output to the second layer decoding unit 103, respectively.
The first layer decoding unit 102 decodes the coding parameter for the first layer input from the separation unit 101, and outputs the first layer decoding signal. This first layer decoding signal is also output to the second layer decoding unit 103.
The second layer decoding unit 103 decodes the coding parameters for the second layer input from the separation unit 101 by using the first layer decoding signal input from the first layer decoding unit 102, and second. Output the layer decoding signal.
An example of the configuration of the scalable coding device 200 corresponding to the scalable decoding device 100 of FIG. 1 is shown in FIG.
In FIG. 2, the first layer coding unit 201 encodes the input audio signal (original signal) and outputs the obtained coding parameters to the first layer decoding unit 202 and the multiplexing unit 203. The first layer coding unit 201 realizes the bandwidth scalability of the first layer and the second layer by performing downsampling processing, low-pass filtering processing, and the like in coding.
The first layer decoding unit 202 generates a decoding signal of the first layer from the coding parameters input from the first layer coding unit 201 and outputs it to the second layer coding unit 204.
The second layer coding unit 204 encodes the input audio signal (original signal) using the first layer decoding signal input from the first layer decoding unit 202, and multiplexes the obtained coding parameters. Output to unit 203. The second layer coding unit 204 performs the upsampling processing of the first layer decoding signal and the upsampling processing of the first layer decoding signal according to the processing (downsampling processing and low-pass filtering processing) performed by the first layer coding unit 201. Performs phase adjustment processing for matching the phase of the first layer decoded signal with the phase of the input audio signal.
The multiplexing unit 203 multiplexes the coding parameter input from the first layer coding unit 201 and the coding parameter input from the second layer coding unit 204, and outputs a bit stream.
Next, the second layer decoding unit 103 shown in FIG. 1 will be described in more detail. FIG. 3 is a block diagram showing the configuration of the second layer decoding unit 103. The second layer decoding unit 103 includes a separation unit 301, a scaling coefficient decoding unit 302, a fine spectrum decoding unit 303, a frequency domain conversion unit 304, a spectrum decoding unit 305, and a time domain conversion unit 306.
The separation unit 301 separates the input coding parameter for the second layer into a coding parameter (scaling coefficient parameter) representing the scaling coefficient and a coding parameter (fine spectrum parameter) representing the spectral fine structure, and the scaling coefficient. Output to the decoding unit 302 and the fine spectrum decoding unit 303, respectively.
The scaling coefficient decoding unit 302 decodes the input scaling coefficient parameters to obtain the low-frequency scaling coefficient and the high-frequency scaling coefficient, outputs the decoding scaling coefficients to the spectrum decoding unit 305, and decodes the fine spectrum. It is also output to the conversion unit 303.
The fine spectrum decoding unit 303 calculates the auditory importance of each band using the decoding scaling coefficient input from the scaling coefficient decoding unit 302, and obtains the number of bits assigned to the fine spectrum information of each band. Then, the fine spectrum decoding unit 303 decodes the fine spectrum parameters input from the separation unit 301 to obtain the decoded fine spectrum information of each band, and outputs the decoded fine spectrum information to the spectrum decoding unit 305. The information of the first layer decoding signal may be used for the calculation of the auditory importance, and in that case, the output of the frequency domain conversion unit 304 is also input to the fine spectrum decoding unit 303.
The frequency domain conversion unit 304 converts the input first layer decoding signal into a spectrum parameter (for example, MDCT coefficient) of the frequency domain and outputs it to the spectrum decoding unit 305.
The spectrum decoding unit 305 includes the first layer decoding signal converted into the frequency domain input from the frequency domain conversion unit 304, and the decoding scaling coefficient (low range and high range) input from the scaling coefficient decoding unit 302. , The spectrum of the second layer decoding signal is decoded from the decoding fine spectrum information input from the fine spectrum decoding unit 303, and is output to the time domain conversion unit 306.
The time domain conversion unit 306 converts the spectrum of the second layer decoding signal input from the spectrum decoding unit 305 into a signal in the time domain, and outputs it as a second layer decoding signal.
FIG. 4 shows an example of the configuration of the second layer coding unit 204 corresponding to the second layer decoding unit 103 of FIG.
In FIG. 4, the input audio signal is input to the auditory masking calculation unit 401 and the frequency domain conversion unit 402A.
The auditory masking calculation unit 401 calculates the auditory masking for each subband having a predetermined bandwidth, and outputs this auditory masking to the scaling coefficient coding unit 403 and the fine spectrum coding unit 404.
The human auditory characteristic has an auditory masking characteristic that when a certain signal is heard, it is difficult to hear even if a sound having a frequency close to that signal comes into the ear. Based on this auditory masking characteristic, the above-mentioned auditory masking is used to allocate a small number of quantization bits to the spectrum of frequencies where the quantization distortion is hard to hear, and to allocate the number of quantization bits to the spectrum of frequencies where the quantization distortion is easy to hear. Efficient spectrum coding can be realized by allocating a large amount.
The frequency domain conversion unit 402A converts the input audio signal into a spectrum parameter (for example, MDCT coefficient) in the frequency domain, and outputs it to the scaling coefficient coding unit 403 and the fine spectrum coding unit 404. The frequency domain conversion unit 402B converts the input first layer decoding signal into a spectrum parameter (for example, MDCT coefficient) of the frequency domain, and outputs it to the scaling coefficient coding unit 403 and the fine spectrum coding unit 404.
The scaling coefficient coding unit 403 uses the auditory masking information input from the auditory masking calculation unit 401 to transmit the spectrum parameters input from the frequency domain conversion unit 402A and the first layer decoding spectrum input from the frequency domain conversion unit 402B. The difference spectrum with and is coded to obtain a scaling coefficient parameter, and the scaling coefficient parameter is output to the coding parameter multiplexing unit 405 and the fine spectrum coding unit 404. Here, an example in which the scaling coefficient parameter of the high frequency spectrum and the scaling coefficient parameter of the low frequency spectrum are output separately is shown.
The fine spectrum coding unit 404 decodes the scaling coefficient parameters (low and high frequencies) input from the scaling coefficient coding unit 403 to obtain the decoding scaling coefficient (low and high frequencies), and the frequency domain conversion unit 402A. The difference spectrum between the spectrum parameters input from and the first layer decoded spectrum input from the frequency domain converter 402B is normalized using the decoding scaling coefficients (low and high frequencies). The fine spectrum coding unit 404 encodes the normalized difference spectrum and outputs the coded difference spectrum (fine spectrum coding parameter) to the coding parameter multiplexing unit 405. At this time, the fine spectrum coding unit 404 calculates the auditory importance for each band of the fine spectrum using the decoding scaling coefficients (low region and high region), and allocates bits according to the auditory importance. The first layer decoding spectrum may be used for calculating the auditory importance.
The coding parameter multiplexing unit 405 includes a high-frequency spectrum scaling coefficient parameter and a low-frequency spectrum scaling coefficient parameter input from the scaling coefficient coding unit 403, and a fine spectrum coding parameter input from the fine spectrum coding unit 404. , Are multiplexed and output as the first spectrum coding parameter.
Next, the spectrum decoding unit 305 shown in FIG. 3 will be described in more detail. 5 to 9 are block diagrams showing the configuration of the spectrum decoding unit 305.
FIG. 5 shows a configuration for executing processing when the first layer decoding signal, all decoding scaling coefficients (low and high frequencies), and all fine spectrum decoding information are all normally received.
FIG. 6 shows a configuration for executing processing when a part of high-frequency fine spectrum decoding information is not received. It differs from FIG. 5 in that the output result of the adder A is input to the high frequency spectrum decoding unit 602. The spectrum of the band to be decoded using the unreceived high-frequency fine spectrum decoding information is pseudo-generated by the method described later.
FIG. 7 shows a configuration for executing processing when all the high-frequency fine spectrum decoding information is not received (including the case where a part of the low-frequency fine spectrum decoding information is not received). It differs from FIG. 6 in that the fine spectrum decoding information is not input to the high frequency spectrum decoding unit 702. The spectrum of the band to be decoded using the unreceived high-frequency fine spectrum decoding information is pseudo-generated by the method described later.
FIG. 8 shows a configuration for executing processing when not all fine spectrum decoding information is received and a part of the low frequency decoding scaling coefficient is not received. It differs from FIG. 7 in that the fine spectrum decoding information is not input and that there is no output from the low-frequency spectrum decoding unit 801 and the adder A does not exist. The spectrum of the band to be decoded using the unreceived high-frequency fine spectrum decoding information is pseudo-generated by the method described later.
FIG. 9 shows a configuration for executing processing when only the high frequency decoding scaling coefficient is received (including the case where some high frequency decoding scaling coefficients are not received). It differs from FIG. 8 in that there is no input of the low frequency decoding scaling coefficient and there is no low frequency spectrum decoding section. A method of pseudo-generating a high-frequency spectrum only from the received high-frequency decoding scaling coefficient will be described later.
The spectrum decoding unit 305 of FIG. 5 includes a low-frequency spectrum decoding unit 501, a high-frequency spectrum decoding unit 502, an adder A, and an adder B.
The low-frequency spectrum decoding unit 501 uses the low-frequency decoding scaling coefficient input from the scaling coefficient decoding unit 302 and the fine-spectrum decoding information input from the fine-spectrum decoding unit 303 to generate a low-frequency spectrum. Decode and output to adder A. Generally, the decoding spectrum is calculated by multiplying the fine spectrum decoding information by the decoding scaling coefficient.
The adder A adds and decodes the decoded low-frequency spectrum (residual) input from the low-frequency spectrum decoding unit 501 and the first-layer decoded signal (spectrum) input from the frequency domain conversion unit 304. Find the low frequency spectrum and output it to adder B.
The high-frequency spectrum decoding unit 502 uses the high-frequency decoding scaling coefficient input from the scaling coefficient decoding unit 302 and the fine-spectrum decoding information input from the fine-spectrum decoding unit 303 to generate a high-frequency spectrum. Decode and output to adder B.
The adder B is a combination of the decoded low-frequency spectrum input from the adder A and the decoded high-frequency spectrum input from the high-frequency spectrum decoding unit 502, and the entire range (all frequencies including the low and high frequencies). Band) spectrum is generated and output as a decoded spectrum.
In FIG. 6, only the operation of the high-frequency spectrum decoding unit 602 is different from that in FIG.
The high-frequency spectrum decoding unit 602 uses the high-frequency decoding scaling coefficient input from the scaling coefficient decoding unit 302 and the high-frequency fine-spectrum decoding information input from the fine-spectrum decoding unit 303 to achieve high frequency. Decode the spectrum of the region. At this time, since the high-frequency fine spectrum decoding information of a part of the band has not been received, the high-frequency spectrum of the corresponding band cannot be accurately decoded. Therefore, the high-frequency spectrum decoding unit 602 uses the decoding scaling coefficient, the low-frequency decoding spectrum input from the adder A, and the high-frequency spectrum that can be received and accurately decoded, in a pseudo-high manner. Generate a region spectrum. The specific generation method will be described later.
In FIG. 7, in FIGS. 5 and 6, the operation is performed when all the high-frequency fine spectrum decoding information is not received. In this case, the high frequency spectrum decoding unit 702 decodes the high frequency spectrum using only the high frequency decoding scaling coefficient input from the scaling coefficient decoding unit 302.
Further, the low-frequency spectrum decoding unit 701 uses the low-frequency decoding scaling coefficient input from the scaling coefficient decoding unit 302 and the low-frequency fine spectrum decoding information input from the fine spectrum decoding unit 303. To decode the low frequency spectrum. At this time, since the low-frequency fine spectrum decoding information of a part of the band has not been received, the decoding process is not performed for a part of the band, and the zero spectrum is set. In this case, the spectrum of the corresponding band output through the adders A and B is the first layer decoded signal (spectrum) itself.
In FIG. 8, the operation is performed when all the low-frequency fine spectrum decoding information is not received in FIG. 7. The low-frequency spectrum decoding unit 801 does not perform the decoding process because the low-frequency decoding scaling coefficient is input but the fine spectrum decoding information is not input at all.
In FIG. 9, the operation is performed when the low-frequency decoding scaling coefficient is not input at all in FIG. However, in the high frequency spectrum decoding unit 902, when a part of the decoding scaling coefficients (high frequencies) are not input, the spectrum of that band is output as zero.
Next, a method of generating a pseudo high-frequency spectrum will be described by taking FIG. 9 as an example. In FIG. 9, it is the high-frequency spectrum decoding unit 902 that pseudo-generates the high-frequency spectrum. FIG. 10 shows the configuration of the high-frequency spectrum decoding unit 902 in more detail.
The high-frequency spectrum decoding unit 902 of FIG. 10 includes an amplitude adjusting unit 1011, a pseudo spectrum generation unit 1012, and a scaling unit 1013.
The amplitude adjustment unit 1011 adjusts the amplitude of the first layer decoded signal spectrum input from the frequency domain conversion unit 304, and outputs the amplitude to the pseudo spectrum generation unit 1012.
The pseudo spectrum generation unit 1012 generates a pseudo high-frequency spectrum using the amplitude-adjusted first layer decoded signal spectrum input from the amplitude adjustment unit 1011 and outputs it to the scaling unit 1013.
The scaling unit 1013 scales the spectrum input from the pseudo spectrum generation unit 1012 and outputs it to the adder B.
FIG. 11 is a schematic diagram showing an example of the above series of processes for generating a pseudo high-frequency spectrum.
First, the amplitude of the decoded signal spectrum of the first layer is adjusted. The amplitude adjustment method can be, for example, a constant multiple in the logarithmic region (γ × S, γ is the amplitude adjustment coefficient (real number) in the range of 0 γ 1, S is the logarithmic spectrum), or a constant power in the linear region (γ). s<sup>γ</sup>, S is a linear spectrum). Further, as the adjustment coefficient for adjusting the amplitude, it is preferable to use a typical coefficient required to match the depth of the harmonics valley in the low range and the depth of the harmonics valley in the high range in the voiced sound. The adjustment coefficient may be a fixed constant, but it is an index indicating the depth of the harmonic valley of the low frequency spectrum (for example, directly the dispersion value of the spectral amplitude in the low frequency band, etc., indirectly in the first layer. It is more preferable to prepare a plurality of appropriate adjustment coefficients according to (such as the value of the pitch gain in the coding unit 201) and selectively use the corresponding adjustment coefficients according to the above indexes. It is also possible to selectively use the adjustment coefficient according to the characteristics of each vowel by using the low-frequency spectral shape (envelope) information and the pitch period information. Further, the optimum adjustment coefficient may be separately encoded on the encoder side as transmission information and transmitted.
Next, a pseudo high-frequency spectrum is generated using the amplitude-adjusted spectrum. FIG. 11 shows an example of mirroring in which the high-frequency spectrum is generated as a mirror image of the low-frequency spectrum as a generation method. In addition to mirroring, a method of shifting the amplitude-adjusted spectrum toward the high frequency axis to generate a high-frequency spectrum, and using a pitch lag obtained from the low-frequency spectrum in the frequency-axis direction with respect to the amplitude-adjusted spectrum. There is a method of generating a high frequency spectrum by performing pitch filtering processing. In either method, the generated high-frequency harmonics structure is prevented from collapsing, and the low-frequency spectrum harmonics structure and the generated high-frequency harmonics structure are continuously connected.
Finally, the amplitude is scaled for each band of the coding unit to generate a high frequency spectrum.
FIG. 12 shows a case where the spectrum information of the first layer (for example, the decoding LSP parameter) is input to the amplitude adjustment unit 1211 from the first layer decoding unit 102. In this case, the amplitude adjustment unit 1211 determines the adjustment coefficient used for the amplitude adjustment based on the input spectrum information of the first layer. In determining the adjustment coefficient, the pitch information (pitch period and pitch gain) of the first layer may be used in addition to the spectrum information of the first layer.
FIG. 13 shows a case where the amplitude adjustment coefficient is separately input to the amplitude adjustment unit 1311. In this case, the amplitude adjustment coefficient is quantized and encoded on the encoder side and transmitted.
(Embodiment 2) FIG. 14 is a block diagram showing the configuration of the second layer decoding unit 103 according to the second embodiment of the present invention.
The second layer decoding unit 103 of FIG. 14 includes a separation unit 1401, a spectrum decoding unit 1402A, an extended band decoding unit 1403, a spectrum decoding unit 1402B, a frequency domain conversion unit 1404, and a time domain conversion unit 1405. ..
The separation unit 1401 separates the coding parameters for the second layer into the first spectrum coding parameter, the extended band coding parameter, and the second spectrum coding parameter, and the spectrum decoding unit 1402A, the extended band decoding Output to unit 1403 and spectrum decoding unit 1402B, respectively.
The frequency domain conversion unit 1404 converts the first layer decoding signal input from the first layer decoding unit 102 into frequency domain parameters (for example, MDCT coefficient), and the spectrum decoding unit 1402A is used as the first layer decoding signal spectrum. Output to.
The spectrum decoding unit 1402A is the code of the first layer obtained by decoding the first spectrum coding parameter input from the separation unit 1401 to the decoded signal spectrum of the first layer input from the frequency domain conversion unit 1404. The quantization spectrum of the conversion error is added, and the spectrum is output to the extended band decoding unit 1403 as the first decoding spectrum. The spectrum decoding unit 1402A mainly improves the coding error of the first layer with respect to the low frequency component.
The extended band decoding unit 1403 decodes various parameters from the extended band coding parameters input from the separation unit 1401, and based on the first decoded spectrum input from the spectrum decoding unit 1402A, the various decoded ones. Decode and generate a high frequency spectrum using parameters. Then, the extended band decoding unit 1403 outputs the spectrum of the entire band as the second decoding spectrum to the spectrum decoding unit 1402B.
The spectrum decoding unit 1402B obtains a second decoded spectrum obtained by decoding the second spectrum coding parameter input from the separation unit 1401 into the second decoded spectrum input from the extended band decoding unit 1403. A spectrum obtained by quantizing the coding error is added and output to the time domain conversion unit 1405 as a third decoded spectrum.
The time domain conversion unit 1405 converts the third decoded spectrum input from the spectrum decoding unit 1402B into a time domain signal and outputs it as a second layer decoding signal.
In addition, in FIG. 14, it is also possible to adopt a configuration in which one or both of the spectrum decoding unit 1402A and the spectrum decoding unit 1402B are not present. In the case of the configuration without the spectrum decoding unit 1402A, the first layer decoding signal spectrum output from the frequency domain conversion unit 1404 is input to the extended band decoding unit 1403. Further, in the case of the configuration without the spectrum decoding unit 1402B, the second decoding spectrum output by the extended band decoding unit 1403 is input to the time domain conversion unit 1405.
FIG. 15 shows an example of the configuration of the second layer coding unit 204 corresponding to the second layer decoding unit 103 of FIG.
In FIG. 15, the audio signal (original signal) is input to the auditory masking calculation unit 1501 and the frequency domain conversion unit 1502A.
The auditory masking calculation unit 1501 calculates the auditory masking using the input audio signal and outputs it to the first spectrum coding unit 1503, the extended band coding unit 1504, and the second spectrum coding unit 1505.
The frequency domain conversion unit 1502A converts the input audio signal into a spectrum parameter (for example, MDCT coefficient) in the frequency domain, and transfers the input voice signal to the first spectrum coding unit 1503, the extended band coding unit 1504, and the second spectrum coding unit 1505. Output.
The frequency domain conversion unit 1502B converts the input first layer decoding signal into spectrum parameters such as MDCT and outputs the input to the first spectrum coding unit 1503.
The first spectrum coding unit 1503 uses the auditory masking input from the auditory masking calculation unit 1501 to input the input audio signal spectrum input from the frequency domain conversion unit 1502A and the first layer input from the frequency domain conversion unit 1502B. The difference spectrum from the decoded spectrum is encoded and output as the first spectrum coding parameter, and the first decoded spectrum obtained by decoding the first spectrum coding parameter is sent to the extended band coding unit 1504. Output.
The extended band coding unit 1504 uses the auditory masking input from the auditory masking calculation unit 1501 to input the input audio signal spectrum input from the frequency domain conversion unit 1502A and the first spectrum encoded unit 1503. The error spectrum from the decoded spectrum of is encoded and output as an extended band coding parameter, and the second decoded spectrum obtained by decoding the extended band coding parameter is output to the second spectrum coding unit 1505.
The second spectrum coding unit 1505 uses the auditory masking input from the auditory masking calculation unit 1501 to input the input audio signal spectrum input from the frequency domain conversion unit 1502A and the second spectrum encoded from the extended band coding unit 1504. The error spectrum with the decoded spectrum of is encoded and output as a second spectrum coding parameter.
Next, specific examples of the spectrum decoding units 1402A and 1402B of FIG. 14 are shown in FIGS. 16 and 17.
In FIG. 16, the separation unit 1601 separates the input coding parameter into a coding parameter (scaling coefficient parameter) representing the scaling coefficient and a coding parameter (fine spectrum parameter) representing the spectral fine structure, and the scaling coefficient. Output to the decoding unit 1602 and the fine spectrum decoding unit 1603, respectively.
The scaling coefficient decoding unit 1602 decodes the input scaling coefficient parameters to obtain the low-frequency scaling coefficient and the high-frequency scaling coefficient, outputs these decoding scaling coefficients to the spectrum decoding unit 1604, and decodes the fine spectrum. It is also output to the conversion unit 1603.
The fine spectrum decoding unit 1603 calculates the auditory importance of each band using the decoding scaling coefficient input from the scaling coefficient decoding unit 1602, and obtains the number of bits assigned to the fine spectrum information of each band. Then, the fine spectrum decoding unit 1603 decodes the fine spectrum parameters input from the separation unit 1601 to obtain the decoded fine spectrum information of each band, and outputs the decoded fine spectrum information to the spectrum decoding unit 1604. The information of the decoding spectrum A may be used for calculating the auditory importance. In that case, the decoding spectrum A is also configured to be input to the fine spectrum decoding unit 1603.
The spectrum decoding unit 1604 contains the input decoding spectrum A, the decoding scaling coefficient (low and high frequencies) input from the scaling coefficient decoding unit 1602, and the decoding fine spectrum input from the fine spectrum decoding unit 1603. The decoding spectrum B is decoded and output from the information.
Explaining the correspondence between FIGS. 16 and 14, when the configuration shown in FIG. 16 is the configuration of the spectrum decoding unit 1402A, the coding parameter of FIG. 16 becomes the first spectrum coding parameter of FIG. The decoding spectrum A of FIG. 14 corresponds to the first layer decoding signal spectrum of FIG. 14, and the decoding spectrum B of FIG. 16 corresponds to the first decoding spectrum of FIG. When the configuration shown in FIG. 16 is the configuration of the spectrum decoding unit 1402B, the coding parameter of FIG. 16 is the second spectrum coding parameter of FIG. 14, and the decoding spectrum A of FIG. 16 is the second spectrum of FIG. Corresponds to the decoding spectrum of FIG. 16, and the decoding spectrum B of FIG. 16 corresponds to the third decoding spectrum of FIG.
FIG. 18 shows an example of the configuration of the first spectrum coding unit 1503 corresponding to the spectrum decoding units 1402A and 1402B of FIG. FIG. 18 shows the configuration of the first spectrum coding unit 1503 in FIG. The first spectrum coding unit 1503 shown in FIG. 18 includes a scaling coefficient coding unit 403, a fine spectrum coding unit 404, a coding parameter multiplexing unit 405 shown in FIG. 4, and a spectrum decoding unit 1604 shown in FIG. Since their operations are the same as those described in FIGS. 4 and 16, the description thereof will be omitted here. Further, if the first layer decoding spectrum of FIG. 18 is replaced with the second decoding spectrum and the first spectrum coding parameter is replaced with the second spectrum coding parameter, the configuration shown in FIG. 18 is the second in FIG. It is composed of the spectrum coding unit 1505. However, in the configuration of the second spectrum coding unit 1505, the spectrum decoding unit 1604 is excluded.
FIG. 17 shows the configurations of the spectrum decoding units 1402A and 1402B when the scaling coefficient is not used. In this case, the spectrum decoding units 1402A and 1402B include an auditory importance and bit allocation calculation unit 1701, a fine spectrum decoding unit 1702, and a spectrum decoding unit 1703.
In FIG. 17, the auditory importance and bit allocation calculation unit 1701 obtains the auditory importance of each band from the input decoding spectrum A, and obtains the bit allocation to each band determined according to the auditory importance. The obtained auditory importance and bit allocation information is output to the fine spectrum decoding unit 1702.
The fine spectrum decoding unit 1702 decodes the input coding parameters based on the auditory importance and bit allocation information input from the auditory importance and bit allocation calculation unit 1701 to obtain the decoded fine spectrum information of each band. Obtained and output to the spectrum decoding unit 1703.
The spectrum decoding unit 1703 adds the fine spectrum decoding information input from the fine spectrum decoding unit 1702 to the input decoding spectrum A, and outputs it as the decoding spectrum B.
Explaining the correspondence between FIGS. 17 and 14, when the configuration shown in FIG. 17 is the configuration of the spectrum decoding unit 1402A, the coding parameter of FIG. 17 becomes the first spectrum coding parameter of FIG. 14, and FIG. 17 The decoding spectrum A of FIG. 14 corresponds to the first layer decoding signal spectrum of FIG. 14, and the decoding spectrum B of FIG. 17 corresponds to the first decoding spectrum of FIG. When the configuration shown in FIG. 17 is the configuration of the spectrum decoding unit 1402B, the coding parameter of FIG. 17 is the second spectrum coding parameter of FIG. 14, and the decoding spectrum A of FIG. 17 is the second spectrum of FIG. Corresponds to the decoding spectrum of FIG. 17, and the decoding spectrum B of FIG. 17 corresponds to the third decoding spectrum of FIG.
It should be noted that the first spectrum coding unit corresponding to the spectrum decoding units 1402A and 1402B of FIG. 17 can be configured in the same manner as the correspondence of FIGS. 16 and 18.
Next, the details of the extended band decoding unit 1403 shown in FIG. 14 will be described with reference to FIGS. 19 to 23.
FIG. 19 is a block diagram showing a configuration of the extended band decoding unit 1403. The extended band decoding unit 1403 shown in FIG. 19 includes a separation unit 1901, an amplitude adjustment unit 1902, a filter state setting unit 1903, a filtering unit 1904, a spectrum residual shape codebook 1905, a spectrum residual gain codebook 1906, and a multiplier 1907. , Scale factor decoding unit 1908, scaling unit 1909, and spectrum synthesis unit 1910.
The separation unit 1901 converts the coding parameters input from the separation unit 1401 in FIG. 14 into amplitude adjustment coefficient coding parameters, lag coding parameters, residual shape coding parameters, residual gain coding parameters, and scale factor coding. It is separated into parameters and output to the amplitude adjustment unit 1902, the filtering unit 1904, the spectrum residual shape codebook 1905, the spectrum residual gain codebook 1906, and the scale factor decoding unit 1908, respectively.
The amplitude adjustment unit 1902 decodes the amplitude adjustment coefficient coding parameter input from the separation unit 1901, and uses the decoded amplitude adjustment coefficient to obtain the first decoded spectrum input from the spectrum decoding unit 1402A in FIG. The amplitude of is adjusted, and the first decoded spectrum after the amplitude adjustment is output to the filter state setting unit 1903. For amplitude adjustment, for example, if the first decoding spectrum is S (n) and the amplitude adjustment coefficient is γ, then {S (n)}<sup>γ</sup>It is done by the method represented by. Here, S (n) is the spectral amplitude in the linear region, and n is the frequency.
The filter state setting unit 1903 uses the transfer function P (z) = (1-z).<sup>-T</sup>)<sup>-1</sup>The first decoded spectrum after amplitude adjustment is set in the filter state of the pitch filter as represented by. Specifically, the filter state setting unit 1903 substitutes the first decoded spectrum S1 [0 to Nn] after amplitude adjustment into the generation spectrum buffer S [0 to Nn], and substitutes the generated spectrum buffer after substitution into the filtering unit. Output to 1904. Here, z is a variable in the z-transform. z<sup>-1</sup>Is a complex variable and is called the delay operator. Further, T is the lag of the pitch filter, Nn is the number of effective spectrum points of the first decoded spectrum (corresponding to the upper limit frequency of the spectrum used as the filter state), and the generated spectrum buffer S [n] is n = 0 to Nw. An array variable defined by a range. Further, Nw is the number of spectrum points after band expansion, and the spectrum of (Nw-Nn) points is generated by this filtering process.
The filtering unit 1904 performs a filtering process on the generated spectrum buffer S [n] input from the filter state setting unit 1903 by using the lag coding parameter T input from the separation unit 1901. Specifically, the filtering unit 1904 generates S [n] by S [n] = S [nT] + gC [n], n = Nn ~ Nw. Here, g indicates the spectral residual gain, C [n] indicates the spectral residual shape vector, and gC [n] is input from the multiplier 1907. The generated S [Nn ~ Nw] is output to the scaling unit 1909.
The spectrum residual shape code book 1905 decodes the residual shape coding parameter input from the separation unit 1901 and outputs the spectral residual shape vector corresponding to the decoding result to the multiplier 1907.
The spectral residual gain codebook 1906 decodes the residual gain coding parameters input from the separator 1901 and outputs the residual gain corresponding to the decoding result to the multiplier 1907.
The multiplier 1907 filters the multiplication result gC [n] of the residual shape vector C [n] input from the spectral residual shape codebook 1905 and the residual gain g input from the spectral residual gain codebook 1906. Output to unit 1904.
The scale factor decoding unit 1908 decodes the scale factor coding parameter input from the separation unit 1901 and outputs the decoding scale factor to the scaling unit 1909.
The scaling unit 1909 multiplies the spectrum S [Nn ~ Nw] input from the filtering unit 1904 by the scale factor input from the scale factor decoding unit 1908 and outputs the spectrum to the spectrum synthesis unit 1910.
The spectrum synthesis unit 1910 scales the first decoded spectrum input from the spectrum decoding unit 1402A of FIG. 14 to the high region (S [Nn to Nn]) to the high region (S [Nn to Nw]). The spectrum input from is output to the spectrum decoding unit 1402B of FIG. 14 as the second decoded spectrum by substituting the respective spectra.
Next, FIG. 20 shows the configuration of the extended band decoding unit 1403 when the spectrum residual shape coding parameter and the spectrum residual gain coding parameter cannot be completely received. In this case, the information that can be completely received is the amplitude adjustment coefficient coding parameter, the lag coding parameter, and the scale factor coding parameter.
In FIG. 20, the configurations other than the separation unit 2001 and the filtering unit 2002 are the same as those in FIG. 19, and the description thereof will be omitted.
In FIG. 20, the separation unit 2001 separates the coding parameters input from the separation unit 1401 of FIG. 14 into the amplitude adjustment coefficient coding parameter, the lag coding parameter, and the scale factor coding parameter, and the amplitude adjustment unit 1902. , Filtering unit 2002, Scale factor decoding unit 1908, respectively.
The filtering unit 2002 performs a filtering process on the generated spectrum buffer S [n] input from the filter state setting unit 1903 by using the lag coding parameter T input from the separation unit 2001. Specifically, the filtering unit 2002 generates S [n] by S [n] = S [nT], n = Nn ~ Nw. The generated S [Nn ~ Nw] is output to the scaling unit 1909.
Next, FIG. 21 shows the configuration of the extended band decoding unit 1403 when the lag coding parameter cannot be received. In this case, the completely receivable information is the amplitude adjustment coefficient coding parameter and the scale factor coding parameter.
In FIG. 21, the filter state setting unit 1903 and the filtering unit 2002 in FIG. 20 are replaced with the pseudo spectrum generation unit 2102. In FIG. 21, the configurations other than the separation unit 2101 and the pseudo spectrum generation unit 2102 are the same as those in FIG. 19, and the description thereof will be omitted.
In FIG. 21, the separation unit 2101 separates the coding parameter input from the separation unit 1401 of FIG. 14 into an amplitude adjustment coefficient coding parameter and a scale factor coding parameter, and the amplitude adjustment unit 1902 and scale factor decoding. Output to unit 1908 respectively.
The pseudo spectrum generation unit 2102 pseudo-generates a high frequency spectrum using the first decoded spectrum after amplitude adjustment input from the amplitude adjustment unit 1902, and outputs the high frequency spectrum to the scaling unit 1909. Specific methods for generating the high-frequency spectrum include a method based on mirroring that generates the high-frequency spectrum as a mirror image of the low-frequency spectrum, a method of shifting the spectrum after amplitude adjustment in the high-frequency direction of the frequency axis, and a method from the low-frequency spectrum. There is a method of obtaining a pitch lag and using this pitch lag to perform pitch filtering processing in the frequency axis direction on the spectrum after amplitude adjustment. If the frame being decoded is determined to be an unvoiced frame, a pseudo spectrum may be generated using a randomly generated noise spectrum.
Next, FIG. 22 shows the configuration of the extended band decoding unit 1403 when the amplitude adjustment information cannot be received. In this case, the fully receivable information is the scale factor coding parameter. In FIG. 22, the configurations other than the separation unit 2201 and the pseudo spectrum generation unit 2202 are the same as those in FIG. 19, and the description thereof will be omitted.
In FIG. 22, the separation unit 2201 separates the scale factor coding parameter from the coding parameter input from the separation unit 1401 of FIG. 14, and outputs the scale factor coding parameter to the scale factor decoding unit 1908.
The pseudo-spectrum generation unit 2202 pseudo-generates a high-frequency spectrum using the first decoded spectrum and outputs it to the scaling unit 1909. Specific methods for generating the high-frequency spectrum include a method based on mirroring that generates the high-frequency spectrum as a mirror image of the low-frequency spectrum, a method of shifting the spectrum after amplitude adjustment in the high-frequency direction of the frequency axis, and a method from the low-frequency spectrum. There is a method of obtaining a pitch lag and using this pitch lag to perform pitch filtering processing in the frequency axis direction on the spectrum after amplitude adjustment. If the frame being decoded is determined to be an unvoiced frame, a pseudo spectrum may be generated using a randomly generated noise spectrum. The amplitude adjustment method is, for example, a constant multiple (γ × S, S is a logarithmic spectrum) in the logarithmic region, or a constant power (s) in the linear region.<sup>γ</sup>, S is a linear spectrum). Further, as the adjustment coefficient for adjusting the amplitude, it is preferable to use a typical coefficient required to match the depth of the harmonics valley in the low range and the depth of the harmonics valley in the high range in the voiced sound. The adjustment coefficient may be a fixed constant, but is an index indicating the depth of the harmonic valley of the low frequency spectrum (for example, the variance value of the spectral amplitude in the low frequency band directly, indirectly the first layer code. It is more preferable to prepare a plurality of appropriate adjustment coefficients according to the value of the pitch gain in the chemical unit 201) and selectively use the corresponding adjustment coefficients according to the above index. It is also possible to selectively use the adjustment coefficient according to the characteristics of each vowel by using the low-frequency spectral shape (envelope) information and the pitch period information. More specifically, since it is the same as the generation of the pseudo spectrum described in the first embodiment, the description here will be omitted.
FIG. 23 is a schematic diagram showing a series of operations for generating high frequency components in the configuration of FIG. 20. As shown in FIG. 23, first, the amplitude of the first decoded spectrum is adjusted. Next, the first decoded spectrum after the amplitude adjustment is used as the filter information of the pitch filter, and filtering processing (pitch filtering) is performed in the frequency axis direction to generate a high frequency component. Next, the generated high-frequency component is scaled for each band of the scaling coefficient to generate the final high-frequency spectrum. Then, the generated high-frequency spectrum and the first decoding spectrum are combined to generate a second decoding spectrum.
FIG. 24 shows an example of the configuration of the extended band coding unit 1504 corresponding to the extended band decoding unit 1403 of FIG.
In FIG. 24, the amplitude adjusting unit 2401 adjusts the amplitude of the first decoded spectrum input from the first spectrum coding unit 1503 using the input audio signal spectrum input from the frequency domain conversion unit 1502A, and adjusts the amplitude. The coding parameter of the adjustment coefficient is output, and the first decoded spectrum after the amplitude adjustment is output to the filter state setting unit 2402. The amplitude adjustment unit 2401 performs an amplitude adjustment process so that the ratio (dynamic range) of the maximum amplitude spectrum and the minimum amplitude spectrum of the first decoded spectrum approaches the high-frequency dynamic range of the input audio signal spectrum. As a method of adjusting the amplitude, for example, the above method can be mentioned. It is also possible to adjust the amplitude by using a conversion formula such as the formula (1). S1 is the spectrum before conversion and S1'is the spectrum after conversion.<maths num="1"><img file="JP4977472B2_D0001.tif" /></maths>
Here, sign () is a function that returns a positive / negative number, and γ is a real number in the range of 0 γ 1. When Eq. (1) is used, the amplitude adjustment unit 2401 prepares in advance the amplitude adjustment coefficient γ when the first decoded spectrum after the amplitude adjustment is closest to the dynamic range of the high frequency portion of the input audio signal spectrum. A plurality of candidates are selected, and the coding parameter of the selected amplitude adjustment coefficient γ is output to the multiplexing unit 203.
The filter state setting unit 2402 sets the first decoded spectrum after the amplitude adjustment input from the amplitude adjustment unit 2401 to the internal state of the pitch filter in the same manner as the filter state setting unit 1903 of FIG.
The lag setting unit 2403 sequentially outputs the lag T to the filtering unit 2404 while gradually changing the lag T within the predetermined search range TMIN to TMAX.
The spectral residual shape code book 2405 stores a plurality of spectral residual shape vector candidates, and sequentially selects a spectral residual shape vector from all or predetermined candidates according to the instruction from the search unit 2406. Output. Similarly, the spectral residual gain codebook 2407 stores a plurality of spectral residual gain candidates, and sequentially selects a spectral residual vector from all or predetermined candidates according to the instruction from the search unit 2406. And output.
The multiplication unit 2408 multiplies the spectral residual shape vector candidates output from the spectral residual shape codebook 2405 and the spectral residual gain candidates output from the spectral residual gain codebook 2407, and filters the multiplied result. Output to part 2404.
The filtering unit 2404 performs filtering using the internal state of the pitch filter set by the filter state setting unit 2402, the lag T output from the lag setting unit 2403, and the gain-adjusted spectral residual shape vector. Calculate the estimated value of the input audio signal spectrum. This operation is the same as the operation of the filtering unit 1904 of FIG.
The search unit 2406 has a cross-correlation between the high frequency part of the input audio signal spectrum (original spectrum) and the output signal of the filtering unit 2404 among a plurality of combinations of lag, spectral residual shape vector, and spectral residual gain. The combination at the maximum is determined by a synthetic analysis method (AbS; Analysis by Synthesis). At this time, auditory masking is used to determine the most audibly similar combinations. In addition, a search is performed in consideration of scaling by a scale factor performed in the subsequent stage. The lag coding parameter, the spectral residual shape vector coding parameter, and the spectral residual gain coding parameter determined by the search unit 2406 are output to the multiplexing unit 203 and the extended band decoding unit 2409.
In the above-mentioned AbS coding parameter determination method, the pitch coefficient, the spectral residual shape vector, and the spectral residual gain may be determined at the same time. Alternatively, in order to reduce the amount of calculation, the pitch coefficient T, the spectral residual shape vector, and the spectral residual gain may be determined in this order.
The extended band decoding unit 2409 includes the amplitude adjustment coefficient coding parameter output from the amplitude adjustment unit 2401, the lag coding parameter output from the search unit 2406, the spectrum residual shape vector coding parameter, and the spectrum residual. Decoding processing is performed on the first decoding spectrum using the gain coding parameter, an estimated spectrum of the input audio signal spectrum (that is, the spectrum before scaling) is generated, and the spectrum is output to the scale factor coding unit 2410. The decoding procedure is the same as that of the extended band decoding unit 1403 in FIG. 19 (except for the processing of the scaling unit 1909 and the spectrum synthesis unit 1910 in FIG. 19).
The scale factor coding unit 2410 combines the high frequency part of the input audio signal spectrum (original spectrum) output from the frequency domain conversion unit 1502A, the estimated spectrum output from the extended band decoding unit 2409, and the auditory masking. The scale factor (scaling coefficient) of the estimated spectrum most suitable for hearing is encoded, and the coding parameter is output to the multiplexing unit 203.
FIG. 25 is a schematic diagram showing the contents of the bit stream received by the separation unit 101 of FIG. As shown in this figure, a plurality of coding parameters are time-multiplexed in the bit stream. The left side of FIG. 25 shows the MSB (Most Significant Bit, the most important bit in the bitstream), and the right side shows the LSB (Least Significant Bit, the least important bit in the bitstream). By arranging the coding parameters in this way, when the bitstream is partially discarded on the transmission path, the quality deterioration due to the discarding can be minimized by discarding the bitstream in order from the LSB side. FIG. 20 is used when LSB to (1) is discarded, FIG. 21 is used when LSB to (2) is discarded, and FIG. 22 is used when LSB to (3) is discarded. It is possible to perform the decoding process by the method described above. When the LSB to (4) are discarded, the decoded signal of the first layer is used as the output signal.
The method of realizing a network in which the coding parameters are preferentially discarded from the LSB side is not particularly limited. For example, it is also possible to use a packet network in which priority control is performed by prioritizing each coding parameter separated by FIG. 25 and transmitting them in separate packets.
Further, in the present embodiment, FIG. 19 shows a configuration including the spectrum residual shape codebook 1905, the spectrum residual gain codebook 1906, and the multiplier 1907, but a configuration without these can also be adopted. In this case, the encoder side does not need to transmit the coding parameter of the residual shape vector and the coding parameter of the residual gain, and can perform communication at a low bit rate. Further, the decoding processing procedure in this case differs from the description using FIG. 19 only in that there is no decoding processing for the spectrum residual information (shape / gain). That is, the decoding processing procedure is the processing procedure described with reference to FIG. 20, and the bit stream has the LSB at the position (1) in FIG. 25.
(Embodiment 3) This embodiment shows another configuration of the extended band decoding unit 1403 of the second layer decoding unit 103 shown in FIG. 14 in the second embodiment. In the present embodiment, the decoding parameter of the frame is determined by using the decoding parameter decoded from the extended band coding parameter of the frame and the previous frame and the data loss information for the received bit stream of the frame. Decodes the decoding spectrum of.
FIG. 26 is a block diagram showing the configuration of the extended band decoding unit 1403 according to the third embodiment of the present invention. In the extended band decoding unit 1403 of FIG. 26, the amplitude adjustment coefficient decoding unit 2601 decodes the amplitude adjustment coefficient from the amplitude adjustment coefficient coding parameter. The lag decoding unit 2602 decodes the lag from the lag coding parameters. The decoding parameter control unit 2603 uses each decoding parameter decoded from the extended band coding parameter, received data loss information, and each decoding parameter of the previous frame output from each buffer 2604a to 2604e, and uses the decoding parameter of the frame. Determine the decoding parameters used to decode the decoding spectrum of 2. The buffers 2604a to 2604e are buffers for storing the decoding parameters of the frame, such as the amplitude adjustment coefficient, the lag, the residual shape vector, the spectral residual gain, and the scale factor, respectively. Since the other configurations in FIG. 26 are the same as the configurations of the extended band decoding unit 1403 in FIG. 19, the description thereof will be omitted.
Next, the operation of the extended band decoding unit 1403 configured in this way will be described.
First, each decoding parameter included in the extended band coding parameter that is a part of the second layer coded data of the frame, that is, scale factor, lag, amplitude adjustment coefficient, residual shape vector, and spectral residual gain. The coding parameters of are decoded by the decoding units 1908, 2602, 2601, 1905, and 1906, respectively. Then, in the decoding parameter control unit 2603, the decoding parameters used for decoding the second decoding spectrum of the frame are determined based on the received data loss information by using each decoded decoding parameter and the decoding parameter of the previous frame thereof. To do.
Here, the received data loss information means that which part of the extended band coding parameter is used by the extended band decoding unit 1403 due to the loss (including packet loss and the case where an error is detected due to a transmission error). This is information indicating whether or not it can be done.
Then, the second decoding spectrum is decoded using the decoding parameter of the frame obtained by the decoding parameter control unit 2603 and the first decoding spectrum. Since the specific operation is the same as that of the extended band decoding unit 1403 of FIG. 19 in the second embodiment, the description thereof will be omitted.
Next, the first operation mode of the decoding parameter control unit 2603 will be described below.
In the first operation mode, the decoding parameter control unit 2603 substitutes the decoding parameter of the corresponding frequency band of the previous frame as the decoding parameter of the frequency band corresponding to the coding parameter not obtained due to the loss.
In particular, SF (n, m): Scale factor of the mth frequency band of the nth frame, T (n, m): Lag in the mth frequency band of the nth frame, γ (n, m): Amplitude adjustment coefficient of the mth frequency band of the nth frame, c (n, m): Residual shape vector of the mth frequency band of the nth frame, g (n, m): Spectral residual gain in the mth frequency band of the nth frame, m = ML ~ MH, ML: Number of the lowest frequency band of the high frequency band in the second layer, MH: Number of the highest frequency band in the high frequency band in the second layer, Then, if the received data loss information indicates that any of the above coding parameters in the mth band of the frame is lost and cannot be received, the previous frame is used as the decoding parameter corresponding to the lost coding parameter. The decoding parameter of the mth band of (n-1th frame) is output.
That is, If scale factor is lost: SF (n, m) SF (n-1, m) If the lag is lost: T (n, m) T (n-1, m) If the amplitude adjustment factor is lost: γ (n, m) γ (n-1, m) If the residual shape vector is lost: c (n, m) c (n-1, m) If spectral residual gain is lost: g (n, m) g (n-1, m) Is.
Instead of the above, either (a) or (b) below may be used. (a) In the frequency band in which any one of the above five types of parameters is lost, the corresponding parameter of the previous frame is used as a plurality of types of decoding parameters associated with all five types or any combination. (b) In the frequency band in which any one of the above five parameters is lost, the residual shape vector and / or the spectral residual gain is set to 0.
On the other hand, in the frequency band where no loss occurs, the decoding parameter decoded by using the coded parameter of the received frame is output as it is.
Then, the decoding parameters of all the high frequency bands of the frame obtained by the above are obtained. SF (n, m), T (n, m), γ (n, m), c (n, m), g (n, m): m = ML ~ MH are output as decoding parameters of the frame. ..
If all the coding parameters of the second layer are lost, the frame compensation of the second layer uses the corresponding decoding parameters of the previous frame as the extended band decoding parameters of the entire high frequency frequency of the frame. ..
Further, in the above description, in the frame in which the loss occurs, the decoding is always performed by using the decoding parameter of the previous frame, but as another form, based on the correlation of the signal between the previous frame and the frame, the decoding is performed. Decoding is performed by the method described above only when the correlation is higher than the threshold value, and when the correlation is lower than the threshold value, decoding is performed by a method closed within the frame according to the second embodiment. good. In this case, as an index showing the correlation between the signal of the previous frame and the signal of the frame, for example, spectral envelope information such as LPC parameter obtained from the coding parameter of the first layer, pitch period, pitch gain parameter, etc. Correlation coefficient and spectral distance between the previous frame and the frame, etc. calculated using information on the voiced constantness of the signal, the low-frequency decoding signal of the first layer, the low-frequency decoding spectrum of the first layer itself, etc. There is.
Next, the second operation mode of the decoding parameter control unit 2603 will be described below.
In the second operation mode, the decoding parameter control unit 2603 sets the decoding parameter of the frequency band of the previous frame and the frequency band of the previous frame and the frame with respect to the frequency band in which the data loss of the frame occurs. The decoding parameters of the adjacent frequency band are used to obtain the decoding parameters of the frequency band.
Specifically, when the received data loss information indicates that the coding parameter of the m-th band of the frame is lost and cannot be received, the previous frame (the first frame) is set as the decoding parameter corresponding to the lost coding parameter. Decoding as follows using the decoding parameters of the m-th band of n-1 frame) and the decoding parameters of the previous frame and the band adjacent to the frequency band of the frame (the same band as the previous frame and the frame). Get the parameters.
That is, If scale factor is lost: SF (n, m) SF (n-1, m) * SF (n, m-1) / SF (n-1, m-1) If the lag is lost: T (n, m) T (n-1, m) * T (n, m-1) / T (n-1, m-1) If the amplitude adjustment factor is lost: γ (n, m) γ (n-1, m) * γ (n, m-1) / γ (n-1, m-1) If spectral residual gain is lost: g (n, m) g (n-1, m) * g (n, m-1) / g (n-1, m-1) If the residual shape vector is lost: c (n, m) c (n-1, m) or 0 Is.
Instead of the above, either (a) or (b) below may be used. (a) In the frequency band in which any one of the above five types of parameters is lost, the parameters obtained according to the above are used as the plurality of types of decoding parameters associated with all five types or any combination. (b) In the frequency band in which any one of the above five parameters is lost, the residual shape vector and / or the spectral residual gain is set to 0.
On the other hand, in the frequency band where no loss occurs, the decoding parameter decoded by using the coded parameter of the received frame is output as it is.
Then, the decoding parameters of all the high frequency bands of the frame obtained by the above are obtained. SF (n, m), T (n, m), γ (n, m), c (n, m), g (n, m): m = ML ~ MH are output as decoding parameters of the frame. ..
In the above description, the example in which the adjacent frequency band of the frequency band m is m-1 has been described, but the parameter of the frequency band m + 1 may be used. However, if the coding parameter is also lost in the adjacent frequency band, the decoding parameter of another frequency band such as the nearest frequency band in which the loss does not occur may be used.
Further, as in the first operation mode, based on the correlation between the signal of the previous frame and the signal of the frame, decoding may be performed by the method described above only when the correlation is higher than the threshold value. good.
Furthermore, of the above five types of decoding parameters, only some of the parameters (scale factor, or scale factor and amplitude adjustment coefficient) use the decoding parameters calculated by the processing described above, and the other decoding parameters are set to the previous frame. Decoding may be performed using the parameters of the frequency band of the above, or decoding may be performed by the method described in the second embodiment.
Furthermore, as another mode of operation, in a system in which a plurality of coded frames are collectively multiplexed and transmitted into one packet, the future coded parameters are preferentially protected (not lost) in terms of time. There is a form to control. In this embodiment, on the receiving side, when decoding a bit stream received in a plurality of frames at once, the decoding of the coded parameters of the lost frame is performed by using the coded parameters of the frames before and after the frame. It may be performed in the same manner as in the first operation mode or the second operation mode. In that case, an interpolated value that is intermediate between the decoding parameter of the previous frame and the decoding parameter of the subsequent frame is obtained and used as the decoding parameter.
It is also possible to take the following form. (1) The decoding spectrum in the spectrum decoding unit 1402B in the second layer decoding unit 103 shown in FIG. 14 is not added to the frequency band in which the extended band coding parameter has a loss. (2) The extended band decoding unit 1403 may be configured not to include a spectrum residual shape code book, a spectrum residual gain code book, and a multiplier.
Further, in the above-described first to third embodiments, a configuration example of two layers is shown, but the number of layers may be three or more.
The embodiments of the scalable decoding device and the scalable coding device according to the present invention have been described above.
The scalable decoding device and the scalable coding device according to the present invention are not limited to the above-described first to third embodiments, and can be modified in various ways.
The scalable decoding device and the scalable coding device according to the present invention can be mounted on a communication terminal device and a base station device in a mobile communication system, thereby having a communication terminal device and a communication terminal device having the same effects as described above. Base station equipment can be provided.
Further, although the case where the present invention is configured by hardware has been described here as an example, the present invention can also be realized by software.
Further, each functional block used in the description of each of the above embodiments is typically realized as an LSI which is an integrated circuit. These may be individually integrated into one chip, or may be integrated into one chip so as to include a part or all of them.
Although it is referred to as LSI here, it may be referred to as IC, system LSI, super LSI, or ultra LSI depending on the degree of integration.
Further, the method of making an integrated circuit is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after the LSI is manufactured, or a reconfigurable processor that can reconfigure the connection and settings of circuit cells inside the LSI may be used.
Furthermore, if an integrated circuit technology that replaces an LSI appears due to advances in semiconductor technology or another technology derived from it, it is naturally possible to integrate functional blocks using that technology. There is a possibility of adaptation of biotechnology.
The main features of the scalable decoding apparatus of the present invention are shown below.
First, in the present invention, when high frequency generation is performed by mirroring, mirroring is performed after adjusting the fluctuation width of the original low frequency spectrum to be mirrored, so that information regarding the adjustment of the fluctuation width may or may not be transmitted. good. As a result, the wave-tuning structure that matches the actual high-frequency spectrum can be approximated, and it is possible to avoid generating an excessive wave-tuning structure.
Secondly, in the present invention, when the lag information is not received due to a transmission line error or the like, when decoding the encoded high frequency component, mirroring is performed in the same manner as in the first feature above to decode the high frequency component. Since the processing is performed, it is possible to generate a spectrum having a wave-tuning structure in the high frequency range without using lag information. In addition, the strength of the wave tuning structure can be adjusted to a reasonable level. A pseudo spectrum may be generated by using another method instead of mirroring.
Third, in the present invention, a bit stream composed of a scale factor, an amplitude adjustment coefficient, a lag, and a spectral residual is used, and if no spectral residual information is received, the scale factor, the amplitude adjustment coefficient, and When the decoding signal is generated only from the lag information and the lag information and the spectrum residual information are not received, the decoding process is performed in the same manner as the decoding according to the second feature. For this reason, the incidence of transmission line errors and loss / discard of coded information is designed to increase in the order of scale factor, amplitude adjustment coefficient, lag, and spectral residual (that is, the scale factor is the most incorrect). When the present invention is applied to a system (which has strong protection or is preferentially transmitted on a transmission line), deterioration of the quality of the decoded voice due to a transmission line error can be minimized. Further, since the decoded voice quality gradually changes for each of the above parameter units, finer scalability than before can be realized.
Fourth, in the present invention, in the extended band decoding unit, a buffer for storing the decoding parameters decoded from the extended band coding parameters used for decoding the previous frame, and the frame and the previous frame. A decoding parameter control unit that determines the decoding parameters of the frame using each decoding parameter and data loss information for the received bit stream of the frame is provided, and the first decoding spectrum of the frame and the decoding parameter control unit are provided. A second decoding spectrum is generated using the output decoding parameters. Therefore, when a part or all of the extended band coded data obtained by encoding the high frequency spectrum by using a filter having a low frequency spectrum as an internal state is lost and cannot be used for decoding. Loss compensation can be performed by using the decoding parameter of the previous frame with high similarity instead, and a high-quality signal can be decoded even when data loss occurs.
In the fourth feature, the decoding parameter control unit determines the frequency band of the previous frame and the frequency band adjacent to the frequency band of the previous frame and the frame with respect to the frequency band in which the data loss of the frame occurs. The decoding parameter of the frequency band may be obtained by using the decoding parameter of. As a result, when using the coding parameters of the preceding frame with high similarity, it is possible to utilize the relationship of the temporal change of the frequency band adjacent to the frequency band to be compensated, and it is possible to perform more accurate compensation. it can.
This specification is based on Japanese Patent Application No. 2004-322954 filed on November 5, 2004. All this content is included here.
The scalable decoding device of the present invention<u style="single">Place</u>, It can be applied to applications such as mobile communication systems and packet communication systems using Internet protocols.
<figref num="1">The block diagram which shows the structure of the scalable decoding apparatus which concerns on Embodiment 1 of this invention.</figref><figref num="2">The block diagram which shows the structure of the scalable coding apparatus which concerns on Embodiment 1 of this invention.</figref><figref num="3">The block diagram which shows the structure of the 2nd layer decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="4">The block diagram which shows the structure of the 2nd layer coding part which concerns on Embodiment 1 of this invention.</figref><figref num="5">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="6">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="7">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="8">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="9">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="10">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="11">The schematic diagram which shows the state of the process which generates the high region component by the high region spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="12">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="13">The block diagram which shows the structure of the spectrum decoding part which concerns on Embodiment 1 of this invention.</figref><figref num="14">The block diagram which shows the structure of the 2nd layer decoding part which concerns on Embodiment 2 of this invention.</figref><figref num="15">The block diagram which shows the structure of the 2nd layer coding part which concerns on Embodiment 2 of this invention.</figref><figref num="16">A block diagram showing a configuration of a spectrum decoding unit according to a second embodiment of the present invention.</figref><figref num="17">A block diagram showing a configuration of a spectrum decoding unit according to a second embodiment of the present invention.</figref><figref num="18">The block diagram which shows the structure of the 1st spectrum coding part which concerns on Embodiment 2 of this invention.</figref><figref num="19">A block diagram showing a configuration of an extended band decoding unit according to a second embodiment of the present invention.</figref><figref num="20">A block diagram showing a configuration of an extended band decoding unit according to a second embodiment of the present invention.</figref><figref num="21">A block diagram showing a configuration of an extended band decoding unit according to a second embodiment of the present invention.</figref><figref num="22">A block diagram showing a configuration of an extended band decoding unit according to a second embodiment of the present invention.</figref><figref num="23">The schematic diagram which shows the state of the process which generates the high region component in the extended band decoding part which concerns on Embodiment 2 of this invention.</figref><figref num="24">The block diagram which shows the structure of the extended band coding part which concerns on Embodiment 2 of this invention.</figref><figref num="25">The schematic diagram which shows the content of the bit stream received by the separation part of the scalable decoding apparatus which concerns on Embodiment 2 of this invention.</figref><figref num="26">A block diagram showing a configuration of an extended band decoding unit according to a third embodiment of the present invention.</figref>
27 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2003216190A | Cites | Japan | Examiner |
| JP2003323199A | Cites | Japan | Examiner |
| JP2004102186A | Cites | Japan | Examiner |
| JP2004521394A | Cites | Japan | Examiner |
| JP2003216190A | Cites | Japan | – |
| JP2003323199A | Cites | Japan | – |
| JP2004102186A | Cites | Japan | – |
| JP2004521394A | Cites | Japan | – |
| JP2003255973A | Cites | Japan | – |
| JP2003216199A | Cites | Japan | – |
14 members in 8 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004322954 | Japan | A | |
| 2004322954 | Japan | A | |
| 2004322954 | Japan | – | |
| 2005020201 | Japan | W | |
| 2005020201 | Japan | W | |
| 2006542422 | Japan | A | |
| 20042004322954 | – | – | – |
| 2005020201 | – | – | – |
| JP20040322954 | – | – | – |
| JP20060542422 | – | – | – |
| WO2005JP20201 | – | – | – |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| WO2006049205A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1808684A1 | European Patent Office (EPO) | A1 | |
| KR20070084002A | Republic of Korea | A | |
| CN101048649A | China | A | |
| JPWO2006049205A1 | Japan | A1 | |
| US2008126082A1 | United States of America | A1 | |
| RU2007116937A | Russian Federation | A | |
| EP1808684A4 | European Patent Office (EPO) | A4 | |
| RU2404506C2 | Russian Federation | C2 | |
| BRPI0517780A2 | Brazil | A2 | |
| US7983904B2 | United States of America | B2 | |
| RU2434324C1 | Russian Federation | C1 | |
| JP4977472B2This record | Japan | B2 | |
| EP1808684B1 | European Patent Office (EPO) | B1 |
19 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4977472
- Publication, DOCDB
- 4977472
- Publication, EPODOC
- JP4977472B
- Application
- 2006542422
- Application, DOCDB
- 2006542422
- Application, EPODOC
- JP20060542422
Titles2
- Japanese
- スケーラブル復号化装置
- English
- Scalable decoding device
Classification
- CPC, 4
- G10L21/038
- H04L27/06
- G10L19/24
- H03M7/30
- IPC, 4
- G10L19 035
- G10L19 005
- G10L19 02
- G10L19 00
