Binaural multi-channel decoder in the context of non-energy-conserving upmix rules
Abstract
A multi-channel decoder, which uses energy error to introduce upmixing rule information of upmixing rules to generate a stereo signal from the downmixing signal for use in accordance with the upmixing rule information and the head-related transfer function corresponding to the upmixing channel (HRTF) is based on the filter characteristics of the filter to calculate the gain factor. The one or more gain factors are used by the filter processor to filter the downmix signal, so that an energy-corrected stereo signal with the left stereo channel and the right stereo channel can be obtained.

Term
Term ended
Expired 4 September 2026, 0.1 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
17 claims: 2 independent, 15 dependent
- 1一种多声道解码器,使用参数以从下混信号产生立体信号,该下混信号从原始多声 道信号中导出,该参数包含上混规则信息,该上混规则信息可用于以上混规则对该下混信 号进行上混,该上混规则造成能量误差,该多声道解码器包括: 增益因子计算器,根据该上混规则信息以及与上混声道对应的以头部相关传递函数 HRTF为基础的滤波器特性,计算用于降低或消除该能量误差的至少一个增益因子,其中该 增益因子计算器可操作以根据该滤波器特性的组合冲激响应的能量来计算该增益因子,该 组合冲激响应是通过加上或减去各个滤波器冲激响应而计算的;以及 滤波处理器,利用该至少一个增益因子、该滤波器特性以及该上混规则信息,对该下混 信号进行滤波,以获得能量修正的立体信号。
- 2如权利要求1所述的多声道解码器,其中该滤波处理器可操作以计算针对该下混信 号的每一个声道的两个增益调整滤波器的滤波器系数,以及利用该两个增益调整滤波器中 的每一个对该下混声道进行滤波。
- 3如权利要求1所述的多声道解码器,其中该滤波处理器不利用该增益因子而操作以 计算用于该下混声道中每一个的两个滤波器的滤波器系数,并对该下混声道进行滤波,在 对该下混声道进行滤波之后,进行增益调整。
- 4如权利要求1所述的多声道解码器,其中该增益因子计算器可操作以根据具有分子 与分母的表达式来计算该增益因子,该分子具有各个滤波器冲激滤波器响应的功率组合, 而该分母具有各个滤波器冲激响应的功率的加权求和,其中在该加权求和中使用的权重因 子与该上混规则信息有关。
- 5如权利要求1所述的多声道解码器,其中该增益因子计算器可操作以根据以下方程 来计算该增益因子: 町 _血 * I]如果 β 0^ 0,σ 1;[1, 靓 其中当η设定为1时,g n 为第一声道的增益因子,其中当η设定为2时,g 2 为第二声道 的增益因子,其中E n B 为通过使用加权参数对声道冲激响应的能量进行加权所计算的加权 求和能量,而其中AE/为该上混规则所引入的该能量误差的估计,其中,α、B与。为上 混规则相关参数,g哑是最大增益因子,而其中ε为大于或等于零 的数字。
- 6如权利要求5所述的多声道解码器,其中该增益因子计算器可操作以根据以下方程 来计算Ε/及ΑΕ/: Ε =Ρ(1-+船2 -b” 3『, Ε: =0(1-+α(1 —σ·)甌2『+Ρ0”』, 其中b n , ι为对应于第一上混声道与第η个立体声道的以HRTF为基础的滤波器,其中 %2为对应于第二上混声道与第η个立体声道 的以HRTF为基础的滤波器冲激响应,其中%3为对应于第三上混声道与第η个立体声 CN 102523552 Β 道的以HRTF为基础的滤波器冲激响应, 其中如下定义为有效的 a=(l-cj3、β =(l-c 2 ) 3 σ = a + β 以及 ρ= α β 其中ci为第一预测参数,c 2 为第二预测参数,而其中该第一预测参数与该第二预测参 数构成该上混规则信息。 7.如权利要求1所述的多声道解码器,其中该增益因子计算器可操作以计算用于左立 体声道与右立体声道的公共增益因子。 如权利要求1所述的多声道解码器,其中该滤波处理器可操作以使用针对虚拟中 央、左和右方位置的左立体声道与右立体声道的以HRTF为基础的滤波器,作为该滤波器特 性,或是使用通过组合针对虚拟左前方位置与虚拟左环绕位置的HRTF滤波器、或是通过组 合针对虚拟右前方位置与虚拟右环绕位置的HRTF滤波器而导出的滤波器特性。
- 79. 如权利要求8所述的多声道解码器,其中与原始左方及左环绕声道有关或是与原始 右方及右环绕声道有关的参数包括在解码器输入信号中,以及 其中该滤波处理器可操作以使用该参数而将头部相关传递函数滤波器进行组合。
- 810. 如权利要求1所述的多声道解码器,其中该增益因子计算器可操作以根据用于立 体声道的以HRTF为基础的滤波器的声道冲激响应的能量的加权线性组合,以及从该加权 线性组合减去估计能量误差所获得的数值之间的比率,来计算该立体声道的增益因子。
- 911. 如权利要求10所述的多声道解码器,其中该增益因子计算器可操作以使用该上混 规则信息来确定该加权因子。
- 1012. 如权利要求11所述的多声道解码器,其中该上混规则信息包含至少两个预测参 数,该预测参数可用于构建上混矩阵,使得输出声道具有与相应的三个输入声道有关的能 量误差。
- 1113. 如权利要求1所述的多声道解码器,其中该滤波处理器可操作为具有下述项目作 为滤波器特性: 第一滤波器,用于对左下混声道进行滤波,以获得第一左立体输出, 第二滤波器,用于对右下混声道进行滤波,以获得第二左立体输出, 第三滤波器,用于对左下混声道进行滤波,以获得第一右立体输出, 第四滤波器,用于对右下混声道进行滤波,以获得第二右立体输出, 加法器,用于将该第一左立体输出与该第二左立体输出进行求和,以获得左方立体声 道,并用于将该第一右立体输出与该第二右立体输出进行求和,以获得右方立体声道, 其中该滤波处理器可操作以在进行求和之前或之后,对该第一或第二滤波器或对该左 方立体输出施加用于该左方立体声道的增益因子,并且在进行求和之前或之后,对该第三 和第四滤波器或对该右方立体输出施加用于该右方立体声道的增益因子。
- 1214. 如权利要求1所述的多声道解码器,其中该上混规则信息包含上混参数,该上混参 数可用于构建上混矩阵,以从两个声道至三个声道产生上混。
- 1315. 如权利要求14所述的多声道解码器,其中该上混规则被定义如下: CN 102523552 Β 其中L为第一上混声道,R为第二上混声道,以及C为第三上混声道,L o 为第一下混声 道,R o 为第二下混声道,而叫j为上混规则信息参数。
- 1416. 如权利要求1所述的多声道解码器,其中预测损失参数被包括在多声道解码器输 入信号中,以及 其中滤波处理器可操作以利用该预测损失参数将该增益因子进行缩放。
- 1517. 如权利要求1所述的多声道解码器,其中该增益计算器可操作以逐子带地计算增 益因子,以及 其中该滤波处理器可操作以逐子带地施加该增益因子。 1 如权利要求8所述的多声道解码器,其中该滤波处理器可操作以通过将HRTF滤波 器的声道冲激响应的加权或相移版本进行求和,以组合与两个声道相关联的HRTF滤波器, 其中用于对HRTF滤波器的声道冲激响应进行加权的权重因子与该声道之间的电平差异有 关,而施加的相移则与该HRTF滤波器的声道冲激响应之间的时间延迟有关。
- 1619. 如权利要求1所述的多声道解码器,其中以HRTF为基础的滤波器或HRTF滤波器 的滤波器特性为复数子带滤波器,该复数子带滤波器是通过利用复数指数调制滤波器组对 HRTF滤波器的实数数值滤波器冲激响应进行滤波而获得的。
- 1720. 一种多声道解码的方法,使用参数以从下混信号产生立体信号,该下混信号从原始 多声道信号中导出,该参数包含上混规则信息,该上混规则信息可用于以上混规则对该下 混信号进行上混,该上混规则造成能量误差,该方法包括: 根据该上混规则信息以及与上混声道相对应的以头部相关传递函数HRTF为基础的滤 波器的滤波器特性,计算至少一个增益因子,用于降低或消除该能量误差,其中该增益因子 是根据该滤波器特性的组合冲激响应的能量来计算,该组合冲激响应是通过加上或减去各 个滤波器冲激响应而计算的;以及 利用该至少一个增益因子、该滤波器特性以及该上混规则信息,对该下混信号进行滤 波,以获得能量修正的立体信号。 CN 102523552 Β
Independent claims17
305 paragraphs in 4 sections, as filed
Non-energy-saving upmix regular context stereo multi-channel decoder
[0001] This application is a divisional application of a patent application filed on December 2, 2008 with an application number of 200680054828. 9. The title of the invention is "non-energy-saving upmix regular-texture stereo multi-channel decoder".
Technical field
[0002] The present invention relates to a method of head-related transfer function (HRTF) filtering to perform binaural decoding of multi-channel audio signals based on available downmix signals and additional control data.
Background technique
[0003] The recent development of audio coding has provided a reconstruction method that can perform multi-channel representation of an audio signal based on a stereo (or mono) signal and corresponding control data. These methods are different from earlier matrix solutions based on, for example, Dolby Prologic, because they transmit additional control data to control the reconstruction based on the transmitted mono or stereo channels. They are also called surround channels. Upmix.
[0004] Therefore, such a parametric multi-channel audio decoder, such as MPEG Surround, reconstructs N channels according to M transmission channels and additional control data, where N>M. The additional control data represents a lower data transmission rate compared to the transmission of all N channels, which makes the coding action very efficient, and at the same time ensures that it is compatible with M channel devices and N channels. Compatibility of the two devices. [Refer to J. Breebaart et al. MPEG spatial audio coding/MPEG Surround: overview and current status, Proc. 119<sup>th</sup> AES convention, New York, USA, October 2005, Preprint 6447.]
[0005] These parametric surround encoding methods generally include a parametric surround signal based on channel intensity difference (CLD) and inter-channel harmony/correlation (ICC). These parameters describe the power ratio in the upmixing process and the correlation between channel pairs. Other channel prediction coefficients (CPC) are also used in the prior art to predict the intermediate or output channels during the upmixing step.
[0006] Other developments in audio coding have also provided a way to obtain an impression of multi-channel signals throughout stereo headphones. This is generally done by using the original multi-channel signal and a head related transfer function (HRTF) filter to down-mix the multi-channel signal into stereo.
[0007] Alternatively, for reasons of computational efficiency and for reasons of audio quality, it is of course also useful to avoid generating stereo signals with a left stereo channel and a right stereo channel.
[0008] However, the problem is how to combine the original head related transfer function (HRTF) filters. In addition, in an upmixing rule method based on an energy loss effect, in other words, when the input signal of the multi-channel decoder includes a downmix signal having, for example, a first downmix channel and a second downmix channel, and additionally has a space The parameters, when used for upmixing in a non-energy conservation mode, can also cause problems. Such parameters are also known as prediction parameters or CPC parameters, for example. Compared with the channel degree difference parameter (CLD), these parameters have the property that they cannot be calculated to reflect the energy distribution between the two channels, but they can be calculated to perform the best possible waveform coincidence, and naturally cause Energy error (for example, energy loss). Therefore, when generating the prediction parameter, the energy conservation property of upmixing cannot be taken into account, and only the most likely time or time of the reconstructed signal compared with the original signal can be concerned. The waveforms in the subband domain are consistent.
[0009] When you want to simply perform head related transfer function (HRTF) filtering based on the transmitted spatial prediction parameters
CN 102523552 Β
If the channel prediction is not performed well in the linear combination of the receivers, it will receive a particularly serious artifact. In this case, even a slight linear correlation can cause unwanted coloring of the stereo output spectrum.<sub>o</sub>It has been found that when the original channel carries paired uncorrelated signals with comparable intensities, such artifacts are often formed.
Summary of the invention
[0010] The objective of the present invention is to provide an efficient and quantitative acceptable concept for multi-channel decoding to obtain a stereo signal, which can be used in, for example, headphone reproduction of multi-channel signals .
[0011] According to the first aspect of the present invention, this goal is achieved by using a multi-channel decoder that uses upmixing rules to upmix the downmix signal using upmixing rule information parameters from the original The downmix signal derived from the multichannel signal generates a stereo signal, and the upmixing rule causes an energy error. The multichannel decoder includes: a gain factor calculator, which is based on the upmixing rule information and corresponding to the upmixing channel The filter characteristic based on the head related transfer function (HRTF) calculates at least one gain factor to reduce or eliminate the energy error; and a filter processor which uses the at least one gain factor, the filter characteristic and the The upmix rule information is used to filter the downmix signal to obtain an energy-corrected stereo signal.
[0012] According to the second aspect of the present invention, this goal is achieved by using a multi-channel decoding method, which uses upmixing rules including upmixing rules to upmix the downmix signal using upmixing rule information parameters, from The downmix signal derived from the original multi-channel signal generates a stereo signal. The upmixing rule causes an energy error. The method includes: according to the upmixing rule information and a head related transfer function (HRTF) corresponding to the upmixing channel ) Based on the filter characteristics, calculate at least one gain factor to reduce or eliminate the energy error; and use the at least one gain factor, the filter characteristics and the upmix rule information to filter the downmix signal to Obtain energy-corrected stereo signals.
[0013] Another aspect of the present invention is related to a computer program having computer readable codes that implement the multi-channel decoding method when executed on a computer.
[0014] The present invention is based on the fact that the upmixing rule information that causes energy error upmixing can be used more favorably to filter the downmix signal, and to obtain a stereo signal without the need to fully express the multichannel signal, and then apply multiple channels. A head related transfer function (HRTF) filter. Instead, according to the present invention, the upmixing rule information related to the energy error effect upmixing rule can be advantageously used to avoid the stereoscopic performance of the downmixed signal, and according to the present invention, the gain factor can be calculated and then the downmixed signal can be filtered. When used, the calculation of this gain factor can reduce or completely eliminate the energy error.
[0015] Specifically, the gain factor is not only related to, for example, the upmix rule information of the prediction parameter, but more importantly, it is also related to the head related transfer function (HRTF)-based filtering corresponding to the upmix channel. The upmixing rules of the upmix channel are known. Specifically, these upmix channels never exist in the preferred embodiment of the present invention, so the stereo channels do not need to be calculated, for example, for the first presentation of three middle channels. However, although the upmix channel itself never exists in the preferred embodiment, it is still possible to derive or provide a head related transfer function (HRTF) based filter corresponding to the upmix channel. It has been found that the energy error introduced by the upmixing rule affected by this energy loss does not correspond to the upmixing rule information transmitted from the encoder to the decoder, but is related to the head related transfer function (HRTF) based The filter is related, so when the gain factor is generated, the filter based on the head related transfer function (HRTF) also affects the calculation of the gain factor.
[0016] In view of this, the present invention will explain the interdependence between upmixing rule information such as prediction parameters, so that the channel is represented by the head related transfer function (HRTF) as the basic filter's prediction parameters and specific Representation
CN 102523552 Β
(appearance) will be the result of upmixing using the upmixing rule.
[0017] Therefore, the present invention provides a solution to solve the spectrum coloring phenomenon generated when the prediction upmixing is combined with the parametric multi-channel audio stereo decoding.
[0018] A preferred embodiment of the present invention includes the following features: an audio decoder that generates a stereo audio signal from spatial parameters related to M decoded signals and the establishment of N (N>M) channels, the decoder includes gain calculation It is used to estimate two compensation gains in many sub-bands from the spatial parameter subsets related to the establishment of P intermediate channels from the P-pair stereo sub-band filter, and includes a gain adjuster to modify the spatial parameters in many sub-bands. The M pairs of stereo subband filters obtained by linear combination of the ρ pair of stereo subband filters in the subband, the modification includes each pair of the M pairs of stereo subband filters and the gain calculated by the gain calculator The two gains are multiplied.
Description of the drawings
[0019] The present invention will now be described by describing examples and with reference to the accompanying drawings, which does not limit the scope or spirit of the present invention, in which:
[0020] FIG. 1 depicts a parametric multi-channel signal stereo synthesis using a head related transfer function (HRTF) correlation filter;
[0021] FIG. 2 depicts a parametric multi-channel signal stereo synthesis using combined filtering;
[0022] FIG. 3 describes the composition of the parameter/filter combiner of the present invention;
[0023] FIG. 4 describes the structure of MPEG Surround spatial decoding;
[0024] FIG. 5 depicts the frequency spectrum of a decoded stereo signal without gain compensation of the present invention;
[0025] FIG. 6 depicts the frequency spectrum of the stereo signal decoding of the present invention;
[0026] FIG. 7 depicts traditional stereo synthesis using head related transfer function (HRTF);
[0027] FIG. 8 depicts a dynamic image compression standard (MPEG) surround encoder;
[0028] FIG. 9 depicts the cascade of a dynamic image compression standard (MPEG) surround decoder and a stereo synthesizer;
[0029] FIG. 10 depicts a conceptual three-dimensional (3D) stereo decoder for a specific configuration;
[0030] FIG. 11 depicts a spatial encoder for a specific configuration;
[0031] FIG. 12 depicts a spatial (moving image compression standard surround (MPEG Surround)) decoder;
[0032] FIG. 13 describes the use of four filters to perform two down-mix channel filtering to obtain a stereo signal without gain factor correction;
[0033] FIG. 14 depicts a five-channel setting, illustrating the spatial setting of different head related transfer functions (HRTF) 1-10;
[0034] FIG. 15 depicts the state of FIG. 14 when the channels representing L, Ls, and R, Iζ have been combined;
[0035] FIG. 16a depicts the settings of FIG. 14 or 15 when the maximum combination of head related transfer function (HRTF) has been implemented, and only the four filters of FIG. 13 are left;
[0036] FIG. 16b depicts the upmixing rule determined by the encoder of FIG. 20, which has upmixing coefficients that cause non-energy conservation upmixing;
[0037] FIG. 17 describes how to combine the head related transfer function (HRTF) to finally obtain four filters based on the head related transfer function (HRTF);
[0038] FIG. 18 depicts a preferred embodiment of the multi-channel decoder of the present invention;
[0039] FIG. 19a depicts filtering based on head related transfer function (HRTF) without gain correction;
CN 102523552 Β
Then, the first embodiment of the multi-channel decoder of the present invention with a zoom level;
[0040] FIG. 19b depicts the device of the present invention after adjustment using a filter based on a head related transfer function (HRTF), which forms a gain-adjusted filter output signal; and
[0041] FIG. 20 shows an example for an encoder that generates information for non-energy conservation upmixing rules.
Detailed ways
[0042] Before discussing the gain adjustment viewpoint of the present invention in detail, the combination of head related transfer function (HRTF) filters and the filter based on head related transfer function (HRTF) will now be discussed in connection with FIGS. 7 to 11 usage of.
[0043] In order to better describe the features and advantages of the present invention, a more detailed description will be made first. A stereo synthesis algorithm is depicted in Figure 7. A set of input channel filtering is performed by a set of head related transfer functions (HRTF). Each input signal is split into two signals (a left (L) and a right (R) component); then each of these signals is assigned to the head related transfer function (HRTF) corresponding to the desired sound source position ) Filtering. Then all the left ear signals are summed to produce a left stereo output signal, and then the right ear signals are summed to produce a right stereo output signal.
[0044] The convolution operation of the head related transfer function (HRTF) can be performed in the time domain, but due to the calculation efficiency factor, it is generally preferable to perform filtering in the frequency domain. In this case, the summation shown in Figure 7 can also be performed in the frequency domain.
[0045] In principle, the stereoscopic synthesis method as depicted in FIG. 7 can be used directly to combine with the Motion Picture Compression Standard (MPEG) surround encoder/decoder. Figure 8 shows the outline of the motion picture compression standard (MPEG) surround encoder. The multi-channel input signal is analyzed by a spatial encoder, and the spatial parameters are combined to form a monophonic or stereo downmix signal. The downmix can be coded using any traditional monophonic or stereo audio coding method. The formed downmix bit stream is combined with the spatial parameter using a multiplexer to form a complete output bit stream.
[0046] FIG. 9 shows the stereo synthesis design of the combined dynamic image compression standard (MPEG) surround decoder. The input bit stream is demultiplexed to form a spatial parameter and downmix bit stream. The latter bitstream is decoded using traditional mono or stereo decoders. The spatial decoder decodes the decoding downmix according to the transmission spatial parameters to generate a multi-channel output. Finally, the multi-channel output is processed by the stereo synthesis stage as depicted in FIG. 7 to form a stereo output signal.
[0047] However, the cascade of such a dynamic image compression standard (MPEG) surround decoder and a stereo synthesis module has at least three disadvantages:
[0048] Calculate the multi-channel signal performance in an intermediate step, then perform head related transfer function (HRTF) convolution processing and perform down-mixing in the stereo synthesis step. Although it is known that each audio signal can have a different spatial position, head related transfer function (HRTF) convolution processing should be performed on a per channel basis, but from a complexity point of view, this is a kind of Unwanted situation.
[0049] The spatial decoder operates in a filter bank (quadrature mirror phase filter (QMF)). On the other hand, the head related transfer function (HRTF) convolution processing is generally applied in the fast Fourier transform (FFT) domain. Therefore, a cascade of multi-channel quadrature mirror phase filter (QMF) synthesis filter bank, multi-channel discrete Fourier transform (DFT), and stereo inverse discrete Fourier transform (DFT) is required to form a high computational requirement system.
[0050] The coding artifacts caused by the spatial decoder to establish perceptible multi-channel reconstruction will likely be enhanced in the (stereo) stereo output.
[0051] The spatial encoder is shown in FIG. 11. The multi-channel signal is composed of Lf, Ls, C, Rf and Rs signals.
CN 102523552 Β
Represents the left front, left surround, center, right front and right surround channels, and is processed by two "one-to-two (OTT)" units, each of which produces a single-tone downmix and parameters for the two input signals. The formed downmix signal is combined with the center channel, and further processed by a "two-to-three (TTT)" encoder to generate stereo downmix and additional spatial parameters.
[0052] Generally speaking, the parameters formed by the "two-to-three (TTTT)" encoder are composed of a pair of prediction coefficients or a pair of degree difference parameters for each parameter band to describe the three The energy ratio of each input signal. The parameters of the "one-to-two (OTT)" encoder are composed of the degree of difference and coherence between the input signals for each frequency band, or the cross-correlation value.
[0053] FIG. 12 depicts a moving picture compression standard (MPEG) surround decoder. The downmix signal 10 and r0 are input into a two-to-three (TTT) module to reconstruct the center channel, the right channel and the left channel. These three channels are further processed by multiple one-to-two (OTT) modules to generate six output channels.
[0054] From the conceptual point of view shown in FIG. 10, the corresponding stereo decoder can be seen. The filter bank domain and the stereo input signal (Lo, %) are processed by a two-to-three (TTT) decoder to form three signals L, R and C. These three signals are then subjected to head related transfer function (HRTF) parameter processing. The formed six channels are summed to produce the stereo stereo output pair (L<sub>b</sub>,R<sub>b</sub>)<sub>o</sub>
[0055] The two-to-three (TTT) decoder can be described by the following matrix operations:
[0056]
[0057]
<td>'L</td><td></td><td>Cut 1</td><td></td><td>~L;</td>
<td>R</td><td>=</td><td>Plus 21</td><td>m<sub>22</sub></td><td>R<sub>o</sub>_</td>
<td>C</td><td></td><td>_Plus 31</td><td><sup>m</sup>32_</td><td></td>
Matrix item m<sub>xy</sub> It is related to the spatial parameters. The relationship between the spatial parameter and the matrix item is unique to, for example, the relationship between the 5.1 multi-channel motion picture compression standard (MPEG) surround decoder. Each of the three forming signals L, R, and C is split into two, and processed using head related transfer function (HRTF) parameters corresponding to the desired (perceived) positions of these sound source positions. For the center channel (C), the spatial parameters of the sound source position can be directly applied to form two output signals (C) and Rb(C) representing the center:
[0058]
<td>B(C)"</td><td></td><td>H(c)"</td>
<td>_Rb(G_</td><td></td><td></td>
[0059] For the left (L) channel, the head-related transfer function (HRTF) parameters from the left front and left surround channels are combined with a single head-related transfer function ( HRTF) parameter set. The resulting "composite" head related transfer function (HRTF) parameters use statistical concepts to simulate the effects of both the front and surround channels. The subsequent equation is used to generate a stereo output pair (Lb, Rb) representing the left channel:
[0060]
<td>"Lb(LY</td><td></td><td></td>
<td>R2_</td><td></td><td>_Hr(L)</td>
[0061] In a similar form, the stereo output representing the right channel can also be obtained according to the following equation:
[0062]
<td></td><td></td><td></td>
<td>_Rb(R)_</td><td></td><td>Hr(R)_</td>
[0063] Given the above-mentioned definitions for Lb(C), Rb(C), Lb(L), Rb(L), Lb(R) and Rb(R), the complete and Rb signal can be Knowing the stereo input signal, derived from a single 2X2 matrix:
<td colspan="2"></td><td>%</td><td>Square 12</td><td></td>
<td>r<sub>b</sub>.</td><td></td><td>/21</td><td>Party 22_</td><td>Λ.</td>
[0064]
CN 102523552 Β
<td>[0065]</td><td colspan="2">among them</td>
<td>[0066]</td><td>hn =</td><td>m<sub>n</sub>H<sub>L</sub> (L) +m<sub>21</sub>H<sub>L</sub> (R) +m<sub>31</sub>H<sub>L</sub> (C)</td>
<td>[0067]</td><td>h<sub>i2</sub> =</td><td>Mountain eat out (L) (R) +m<sub>32</sub>H<sub>L</sub> (C)</td>
<td>[0068]</td><td>h<sub>2</sub>i =</td><td>m<sub>n</sub>H<sub>E</sub> (L) +m<sub>21</sub>H<sub>E</sub> (R) +m<sub>31</sub>H<sub>E</sub> (C)</td>
<td>[0069]</td><td>h<sub>22</sub> =</td><td>m<sub>12</sub>H<sub>E</sub> (L) +m<sub>22</sub>H<sub>E</sub> (R) +m<sub>32</sub>H<sub>E</sub> (C)</td>
[0070] The Hx(Y) filter may be expressed as a weighted combination of parameters in the parameter form of the original head related transfer function (HRTF) filter. In order to make it possible to complete, the original head related transfer function (HRTF) filter is expressed as
[0071] Represent the (average) degree of each frequency band of the left ear impulse response;
[0072] Represent the (average) degree of each frequency band of the right ear impulse response;
[0073] Represents the (average) arrival time or phase difference between the left and right ear impulse responses.
[0074] Therefore, when the center channel input signal is known, the head related transfer function (HRTF) filter representing the left and right ears can be expressed as:
[0075]
0(C))"P,(C)e+heart)/2
Hr(C)"LPr(C)e"C)/2
[0076] Where P] (C) represents the average degree of the left ear in a known frequency band, and Φ (C) is the phase difference.
[0077] Therefore, the head related transfer function (HRTF) parameter may be simply determined by using P] and P<sub>r</sub>It is composed of the signal multiplication operation of, which corresponds to the sound source position of the center channel, and the phase difference is distributed symmetrically. This process can be performed independently for each quadrature mirror filter (QMF) group, which maps from the head related transfer function (HRTF) parameters to the quadrature mirror filter (QMF) group on one band On the other hand, the spatial parameters are also mapped to the quadrature mirror filter (QMF) group.
[0078] Similarly, when the left and right channels are known, the head related transfer function (HRTF) filter representing the left and right ears can be given by the following equation:
[0079] =Bu Ji2(1/)+P: (Ls)
[0080] H<sub>r</sub> (Q = Scrumptious"Secret)) "Need this P; (Ls)
[0081] Hj(R)=e+w")+w%(")) J w Junior 2 (Miao) Level 2 (Heart)
[0082] H<sub>r</sub> (R) = Bu; Ρ; (boat) + it P; (Rs)
[0083] Obviously, for the six original channels, the head related transfer function (HRTF) is a weighted combination representing the degree and phase difference of the parameterized head related transfer function (HRTF) filter.
<td>[0084]</td><td>The weights Wif and Wis are the same as the one-to-two (0ΤΤ) functional blocks used for the front left (Lf) and surround left (Ls)</td>
The channel intensity difference (CLD) parameter is related to:
<td>[0085]</td><td>,2 _ 10 Sarakawa. , 2 _ 1W imitation] + - | _J_ JQCLD//10</td>
<td>[0086]</td><td>The weights W "f and W" are the same as the one-to-two (OTT) functional blocks used for the front right (Rf) and surround right (Rs)</td>
The channel intensity difference (CLD) parameter is related to:
<td>[0087]</td><td>,2 _ 10®. , 2 _ 1 Yin 1 + [0. ?/10 Heart 1] + [0(s/io</td>
CN 102523552 Β
[0088] The above-mentioned solution can be well applied to a short head related transfer function (HRTF) filter, and its effective accuracy can be expressed as the average degree of each frequency band and the average phase difference of each frequency band. However, for long-echo head related transfer function (HRTF), it is not applicable.
[0089] The present invention will teach how to extend a 2X2 matrix stereo decoder to be able to handle head related transfer function (HRTF) filters of any length. In order to achieve this objective, the present invention includes the following steps:
[0090] Convert the head related transfer function (HRTF) filter response to the filter bank domain;
[0091] Obtain the overall delay difference or phase difference from the head related transfer function (HRTF) filter pair;
[0092] The response of the head related transfer function (HRTF) filter pair is morphed as a function of the vocal tract intensity difference (CLD) parameter;
[0093] Gain adjustment.
[0094] This can use six filter permutations to represent Y=L<sub>o</sub>> R<sub>o</sub>And X = six complex gains H of L, R, C<sub>Y</sub>(X) Achieved. These filters are represented from Y = L<sub>0</sub>,R<sub>0</sub>Derived from the ten filters of X = Lf, Ls, Rf, Rs, C (H, X), which describe the known head related transfer function (HRTF) in the quadrature mirror phase filter (QMF) domain Filter response. These quadrature mirror phase filters (QMF) can be achieved according to the method described below.
[0095] According to the following equation, the front and surround channel filters are formed by using a complex linear combination:
[0096] Hy(X)=gw<sub>f</sub> exp(-/Such as + gw, ^PU</>xy^H<sub>y</sub>(Xs)
[0097] The phase parameter χγ can be obtained from the main delay time difference J between the front and rear head related transfer function (HRTF) filters and the subband index η of the quadrature mirror phase filter (QMF) group through Defined by the following equation: (1) <sub>Γ</sub> ί 7l\ Tl~\-[0098] \ 2
Φχγ ~ early <sup>τ</sup>ΧΥ
[0099] This phase parameter has a dual role in filter shaping. First, it implements the delay compensation of the two filters before superimposing, thereby forming a combined response that simulates the delay time corresponding to the source position between the front and rear speakers. Second, it makes the required gain compensation factor g more stable, and exhibits a slower change in frequency compared with the simple superposition case performed by using = 0.
[0100] The gain factor g is defined using the same anharmonic and extra power rules, as in the case of the parameter head related transfer function (HRTF),
[0101] Ρ<sub>γ</sub> (Ύ)<sup>2</sup> = W; horse (for Q2 + just horse (Xsf
[0102] where
[0103] Ρ<sub>γ</sub> (Ύ)<sup>2</sup> = g\^PY (AT)<sup>2</sup> + ^PY {XsY + 2w<sub>f</sub>w<sub>s</sub>P<sub>Y</sub> (Xf)P<sub>Y</sub> (Xs)<sub>PxY</sub>)
[0104] And P XY is the real value of the normalized complex cross-correlation between the following filters
[0105] exp (-j Φ χγ) Η<sub>γ</sub> (Xf) and Η<sub>γ</sub> (Xs)
[0106] In the case of simple superposition with =0, the value of Pxγ is a function of frequency, presenting an unstable oscillation mode, which causes the need for extensive gain adjustment. In actual implementation, the value of the gain factor g needs to be limited, and the residual spectrum coloring effect of the signal cannot be avoided.
[0107] In contrast, the retardation-based phase compensation molding taught by the present invention is used to form a smooth behavior of Pxγ as a function of frequency. This value is usually close to one for the natural head related transfer function (HRTF) derived by the filter pair, because the main difference lies in the delay and amplitude, and the purpose of the phase parameter is to consider the quadrature mirror.
CN 102523552 Β
Phase filter (QMF) delay difference in the group domain.
[0108] Another advantageous alternative to the phase parameter is to use the normalized complex cross-correlation phase angle between the following two filters to be calculated
[0109] H<sub>Y</sub>(Xf) and H,Xs)
[0110] And using standard unwrapping (unwrapping) technology to expand the phase value into a function of the quadrature mirror phase filter (QMF) group subband pointer n. This choice causes Pxγ to never become a negative value, so the compensation gain g satisfies 1/V2<g<l for all subbands. In addition, the selection of this phase parameter allows the front and surround channel filter shaping to be performed when the main delay time difference jy is not available.
[0111] The signals to be considered below are sub-band samples from a modulation filter bank or windowed Fast Fourier Transform (FFT) analysis of discrete-time signals, or from discrete-time signals. It can be understood that these subbands must be converted back to the discrete time domain using the corresponding synthesis filter bank operation.
[0112] FIG. 1 describes a process of stereo synthesis of parametric multi-channel signals using filters related to a head related transfer function (HRTF). The spatial decoding 101 generates a multi-channel signal including N channels (M<N) according to M transmission channels and transmission space parameters. Then, these N channels are converted to two output channels representing stereo listening by means of a head related transfer function (HRTF) filter. The head related transfer function (HRTF) filter 102 superimposes the filtering results of each input channel, where one head related transfer function (HRTF) filter is used for the left ear, and the other head related transfer function (HRTF) The filter is used for the right ear. In general, it requires 2N filters. However, when listening through N speakers, the parameterized multi-channel signal can achieve a high-quality listener experience, and the subtle internal correlation between the N signals will cause an artifact of the stereo listening. These artifacts are dominated by the difference in the spectral content of the reference stereo signal defined by the original N channel head related transfer function (HRTF) filtering before encoding. Another disadvantage of this cascade is that the overall computational cost for stereo synthesis is an addition to the cost required for each component 101 and 102.
[0113] FIG. 2 depicts a stereo synthesis of parametric multi-channel signals performed by the combined filtering method taught by the present invention. The transmission space parameters are split from 201 into two sets, set 1 and set 2. Here, set 2 includes the relevant parameters for establishing the P intermediate channels from the M transmission channels, and set 1 includes the parameters from the P channels. The middle channel establishes the relevant parameters of the N channels. The pre-combiner 202 in the prior art uses weights to combine the 2N selected pairs of subband filters related to the head related transfer function (HRTF), and the weights are combined with parameter set 1 and the selected filter. Pair related. The result of this pre-combination is to produce 2P stereo subband filters, which represent a stereo filter pair for each of the P middle channels. The combiner 203 of the present invention uses the weights related to both the parameter combination 2 and the 2P stereo subband filters to combine the 2P stereo subband filters into a set of 2M stereo subband filters. In contrast, the prior art linear combiner can apply weights that are only related to the parameter set 2. The formed 2M filter group is composed of stereo filter pairs for each of the M transmission channels. The combined filtering unit 204 obtains a two-channel output contribution pair for each of the M transmission channels by using the corresponding filter pair filtering method. Then, the sum of all M contributions is performed to form a two-channel output in the subband domain.
[0114] FIG. 3 depicts the components of the combiner 203 of the present invention, which is used for the combination of spatial parameters and stereo filters. The linear combiner 301 uses weights derived from known spectral parameters to combine the 2P stereo subband filters into 2M stereo filters, where the spatial parameters are the establishment of the P from the M transmission channels. The relevant parameters of the middle channel. Specifically, this linear combination simulates the upmixing from the M transmission channels to the P middle channels, and then the concatenation of stereo filtering from the P sources. The gain adjuster 303 uses a method of applying a common left gain to each filter corresponding to the left ear output and a method of applying a common right gain to each filter corresponding to the right ear output to modify
CN 102523552 Β
It is coming from the 2M stereo filters output by the linear combiner 301. These gains come from the gain calculator 301 that derives gains from the spatial parameters and the 2P stereo filters. The purpose of the gain adjustment performed by the components 302 and 303 of the present invention is to compensate the unintended spectrum coloring effect caused by the linear combiner 301 when the P middle channels have spatial decoding linear correlation. The gain calculator 302 taught by the present invention includes a device for estimating the energy distribution of the P middle channels with the spectral parameter as a function.
[0115] FIG. 4 describes the structure of MPEG Surround spatial decoding in the case of a stereo transmission signal. The analysis sub-bands of the M = 2 transmission signals are provided to the two-to-three (2-3) functional block 401, which outputs P = 3 intermediate signals, combining the left, combined right, and combined center signals. This upmix is related to the subset of transmission space parameters, which corresponds to set 2 in FIG. 2. These three intermediate signals are then provided to three one-to-two (1 one 2) functional blocks 402-404, which generate a total of N = 6 signals 405: left front turtle), left surround (1J, right Front (mound), right surround (), center (c), and low frequency enhancement (lfe). This upmix is related to the subset of the transmission space parameters, which corresponds to set 1 in Figure 2. The final multi-channel The digital audio output is established by passing the six subband signals to the six synthesis filter banks.
[0116] FIG. 5 describes the problem solved by the gain compensation of the present invention. The reference head related transfer function (HRTF) filtered stereo output spectrum representing the left ear is depicted by a solid graph. The dotted line graph depicts the frequency spectrum of the corresponding decoded signal generated by the method in FIG. 2, where the combiner 203 is composed of only the linear combiner 301. As can be seen, in the frequency range of 3-4 kHz and 11-13 kHz, there is a large amount of spectrum energy loss relative to the expected reference spectrum. There is also a small amount of spectrum increase near 1 kHz and 10 kHz.
[0117] FIG. 6 describes the advantages of using the gain compensation of the present invention. The solid graph is the same as the reference frequency spectrum in FIG. 5, and the dotted graph now depicts the frequency spectrum of the decoded signal generated by the method of FIG. 2, where the combiner 203 is composed of all the components in FIG. As can be seen, compared with the two curves in FIG. 5, a significantly improved spectral agreement is obtained between the two curves.
[0118] In the following text, the mathematical description of the gain compensation of the present invention will be roughly explained. For discrete complex signals X and y, the complex inner product and square modulus (energy) are defined as
[0119] <nervous0 = work * hatred) evil gang) χ = BU=<χ, X) = correcting class research <sup>r=</sup>HI<sup>2=</sup>U^=EI^)r
<img file="CN102523552B_D0001.tif" />
[0120] where the gangster) is a complex conjugate signal of y(k).
[0121] The original multi-channel signal is composed of N channels, and each channel has a filter pair associated with the stereo head related transfer function (HRTF). However, it will be assumed here that the parameterized multi-channel signal is built using the intermediate step of predictive upmixing from the M transmission channels to the P prediction channels. This structure is used in the moving image compression standard surround (MPEG Surround) as described in FIG. 4. It is assumed that the original set of 2N head related transfer function (HRTF) correlation filters has been reduced to a filter pair representing each of the P prediction channels by using the prior art precombiner 202, where M WPWN. The P predicted channel signal pairs, ρ = 1, 2, ..., P, are intended to approximate the P channel signals χ<sub>ρ</sub>,ρ=1,2,...,P, these signals are derived from the original N channels through partial downmixing. In the moving image compression standard surround (MPEG Surround), these signals are combined left, combined right, and combined and scaled center/low frequency enhancement (lfe) channels. Assume that the head-related transfer function corresponding to the signal is 13
CN 102523552 Β
The (HRTF) filter pair is described by the sub-band filter ρ representing the left ear and the sub-band filter ρ representing the right ear. The reference stereo output signal can therefore be calculated by linear addition of filtered signals for η = 1, 2.
Ρ
[0122] Ding" ) = £(®,p*Xp)(Q
P=1 (2)
[0123] where the asterisk indicates the convolution calculation in the time direction. The subband filter can be given in the form of a finite impulse response (FIR) filter, an infinite impulse response (IIR), or derived from a parameterized fandly of the filter.
[0124] In the encoder, the downmix is formed by applying the MXP downmix matrix D to the column vector signal formed by Xp, p=1, 2, ..., P, and in the decoding The prediction in the device utilizes the application of the PXM prediction matrix C to the M transmission downmix channels z<sub>ffl</sub>.m= ..., Μ formed by the column vector signal method,
Μ
[0125] Opposite partner) = work spoon swelling," partner) heart (3)
[0126] Knowing the two matrices at the decoder and ignoring the coding effect of the downmix channel, the combined effect of the prediction can be calculated using the following formula:
P
[0127] Opposite partner)= ΣX session partner), (4)
[0128] where ap is the term of the matrix product A=CD.
[0129] A straightforward method for generating stereo output at the decoder is to simply insert the prediction signal into (2) to form p
[0130] Nine groups) = Jia) group) p=i(5)
[0131] In terms of calculation, the stereo filter is combined with the prediction upmix in advance, so (5) can be written as
Μ
[0132] With (£) = heart (6)
[0133] The combined filter is defined as
Ρ
[0134] Gonggu "No." Chou) ρ=ι(7)
[0135] This equation describes the function of the linear combiner 301, which uses the coefficient c derived from the spatial parameter<sub>p</sub>, <sub>ffl</sub>And the stereo subband filter b<sub>n</sub>, <sub>p</sub>combination. When the original P signals x<sub>p</sub>With a numerical order basically defined by M, the prediction can be designed to perform well, and the approximation period f is established. This can happen, for example, if only M of the P channels are active, or if important signal components are derived from amplitude panning. In this case, the decoded stereo signal (5) has a good agreement with the reference (2). On the other hand, in the general case, and especially when the original P signals are not correlated, there will be a large amount of prediction loss, and the output derived from (5) can be significantly different from the output derived from (2). )'S energy is different. Since this difference is different in different frequency bands, the final audio output also suffers from spectral coloring artifacts as described in FIG. 5. The present invention teaches how to avoid this problem by gain compensation for the output, which is based on the following equation
[0136] y<sub>n</sub> = g<sub>n</sub>-y<sub>n</sub>
CN 102523552 Β
[0137] In terms of calculation, according to the gain adjuster 303 watts (Qiao) = partner) to change the combined filter, gain compensation can be advantageously performed. Then, the modified combined filter becomes
Μ
<td>[0138]</td><td></td><td>(9)</td>
<td>[0139]</td><td>The optimal value of the compensation gain in (8) is</td><td></td>
<td>[0140]</td><td>-fed</td><td>(10)</td>
<td>[0141]</td><td colspan="2">The purpose of the gain calculator 302 is to estimate the gains based on the information available in the decoder.</td>
The various tools used in this project will now be described. The information available here is represented by matrix items a»q and subband filters bn,ρ related to the head related transfer function (HRTF). First, the subsequent approximation will be assumed to be used for the inner product between the signals x, y that have been filtered using the subband filters b, d related to the head related transfer function (HRTF),
[0142] <b*x, d*y> heart<b, d><x, y> (11)
[0143] This approximation is based on the fact that usually the maximum energy of the filter is concentrated on the dominant single tap, and then it is presupposed that the time step of the application time-frequency conversion (step) and the head related transfer function (HRTF) The main delay difference of the filter is large enough in comparison. Apply the combination of approximation (11) and (2) to form
[0144] |child"« £<Jiuji><6,_) ρ as in (12)
[0145] The next approximation involves assuming that the original signal is uncorrelated, that is, for p#q, <x<sub>p</sub>,x<sub>q</sub>> = 0<sub>o</sub>Then (12) is simplified to ρ
[one] ικιι<sup>2</sup>«ΣΙΚ,ΙΙΊΜ<sup>2</sup>,、
1(13)
[0147] For the decoding energy, the result corresponding to (12) is
[0148] Jie ""Heart Σ0 rise" <Yes, ) Now (14)
[0149] Insert the prediction signal (4) in (14), and apply the assumption that the original signal is irrelevant to obtain the<sup>2</sup>Eat for
Ρ=1
<img file="CN102523552B_D0002.tif" />
(15)
[0150]
[0151] Next, in order to be able to calculate the compensation gain given by the quotient (10), estimate the energy distribution II x<sub>p</sub>ll<sup>2</sup>,p =
1, 2,..., Ρ, Ρ is the number of original channels up to any factor. The present invention teaches how to calculate the prediction matrix c corresponding to the hypothesis by using the function of the energy distribution<sub>fflOdel</sub>To complete this work, these channels are not related to each other, and the goal of the encoder is to minimize the prediction error. If possible, then by solving the nonlinear equation system C<sub>fflodel </sub>=C estimates the energy distribution. For the prediction parameters that form an equation system that does not have a solution, the gain compensation factor is set to gn=lo. This inventive step will be described in detail for the most important special cases in subsequent paragraphs.
[0152] The computational load increased by (15) can be reduced in the case of P=M+1 by using the following extension method (for example, refer to PCT/EP2005/011586)
[0153] (X<sub>p</sub>,x<sub>g</sub>) = (x<sub>p</sub>,x<sub>q</sub>) + ^EV<sub>p</sub>-V<sub>q</sub> (16)
[0154] where v is a unit vector with component Vp, such that Dv=0, and AE is the predicted loss energy
CN 102523552 Β
[0155]
[0156]
[0157]
[0158] \E = Ε -total =
<img file="CN102523552B_D0003.tif" />
(17) The calculation of (15) can then be replaced by the application of (16) in (14). As a result, ""ΤΑ""-Bisun Nine (18) Next, we will discuss predicting three from two channels. Preferential treatment of the soundtrack. The cases of Μ=2 and P=3 are used in the case of the moving image compression standard surround (MPEG Surround). The signal is combined left x]=1, combined right x<sub>2</sub> = r and (zoom) combined center/low frequency enhancement (lfe) channel x<sub>3</sub> = c<sub>o</sub>The downmix matrix is
[0159]
[0160] And the real number parameters C1, c are transmitted by two<sub>2</sub>The constructed prediction matrix is
[0161] + qc<sub>2</sub> -1 q _] 2 + C?
- C] 1 - (19) (20)
[0162] Under the assumption that the original sound channel is irrelevant, the prediction matrix for minimizing the prediction error is as follows
[0163]
[0164]
[0165]
[0166] model
LC + RC + LR
LC + LR
-RC
RC
-LC
RC + LR
LC makes Jiandi=C, obtains the (unnormalized) energy distribution taught by the present invention
<td colspan="2">'L~\</td><td>Xi-σ;</td>
<td>R</td><td>=</td><td>¢/(1-σ)</td>
<td>C</td><td></td><td>_ Ρ _</td>
Where a = (l-cJ/3, β
<img file="CN102523552B_D0004.tif" />
(21) (22) = α+Β and ρ=αΒ. This applies in the variable scope defined below
[0167] α> 0, 0> 0, ο <1 (23)
[0168] The prediction error can be obtained by the same scaling method
[0169] ΔΕ = 3ρ(1-σ)(24)
[0170] Since P=3=2+1=M+1, the method described in (16)-( can also be applied. The unit vector is
[v<sub>15</sub>v<sub>2</sub>,v<sub>3</sub>]=[1,1,-1]/73, and has the following definition
[0171] ΔΕ? =p(lb)=|b"i+b"2-you3(25)
[0172] and
[0173] £=0(1 -b)|0""+a(l -σ)|0",2+τφ"3(26)
[0174] In the preferred embodiment of the gain calculator 302, the calculated compensation gain for each ear n=1, 2 can be expressed as
[0175]
CN 102523552 Β
& = Bu Xun{j JE: f Ί;]two} iD^a> 0, β>0,σ<1;
(1, otherwise (27)
[0176] Here, ε>0 is a small number whose purpose is to stabilize the equation near the edge of the variable parameter range, and g28 is the upper limit of the applied compensation gain. The gain of (27) is different for the left ear and the right ear with n=1, 2. A variant of this method is to use a common gain go = gi = g, where
[0177] ί / Ε^+Ε^ + ε & = min [red, + calendar_encounter+subscribe, if qO, #aO,cf<] otherwise (28)
[0178] The modified gain factor of the present invention can coexist with available direct multi-channel gain compensation without involving any head related transfer function (HRTF) related issues.
[0179] In the dynamic image compression standard surround (MPEG Surround), the compensation for the prediction loss has been applied in the decoder by multiplying the factor 1/ρ by the upmix matrix C, where 0<p W 1 is a part of the transmission space parameter. The gains of (27) and (28) have been replaced by the products Pgn and Pg, respectively. This compensation is applied to the stereo decoding studied in Figures 5 and 6. This is also the reason why the prior art decoding method of FIG. 5 has an enlarged frequency spectrum compared with the reference. For the subbands corresponding to those frequency regions, the gain compensation of the present invention effectively uses the smaller value derived from equation (28) to replace the transmission parameter gain factor 1/P.
[0180] Furthermore, because the case of P=1 corresponds to a successful prediction, a more conservative variant of the gain compensation taught by the present invention will cause the stereo gain compensation for P=1 to fail.
[0181] In addition, the present invention can also be used with residual signals. In the dynamic image compression standard surround (MPEG Surround), an additional prediction residual signal Z3 can be transmitted, so that the original P=3 signal can be reproduced more accurately. In this case, the gain compensation is replaced by the addition of the stereo residual signal that will now be depicted. The prediction upmix enhanced by the residual signal is formed according to the following equation
[0182] For (') = Σ spoon, back group) + xi·Z3 group) <sup>m=1</sup>(29)
[0183] where net generation]=[1, 1, -l]/3o uses yuan/ to replace the pair in to form a corresponding combined filter,
[0184] Ji's = £ Chou, m*Zm), (30)
[0185] The combined filter for m=1 and 2 is defined by , and the combined filter used for the residual addition is defined as
[0186] h"3=g(b"j+b"2-b"3)(3])
[0187] The complete structure of this decoding mode therefore uses the method of setting P=M=3 and uses the modification of the combiner
203 only performs the linear combination defined by (7) and (31), and is described with Figure 2.
CN 102523552 Β
[0188] FIG. 13 depicts a modified representation of the result of the linear combiner 301 in FIG. 3. The result of this combiner is four filters h based on the head related transfer function (HRTF)<sub>n</sub>,h<sub>12</sub>,h<sub>21</sub>With ti??. As described with FIGS. 16a and 17, it is more clear that these filters correspond to the filters indicated by 15, 16, 17, and 18 in FIG. 16a.
[0189] FIG. 16a shows the head of a listener with a left ear or a left stereo point and a right ear or a right stereo point. When Figure 16a is only related to the stereo solution, the filters 15, 16, 17, 18 are general head-related transfer functions, which can be measured separately, or obtained through the Internet, or targeted at the listener. Obtained from a textbook at a different position between the left channel speaker and the right channel speaker.
[0190] However, because the present invention teaches a multi-channel stereo decoder, the filters described using 15, 16, 17, 18 are not pure head-related transfer function (HRTF) filters, but are based on head-related transfer function (HRTF) filters. The HRTF-based filter not only reflects the nature of the head-related transfer function (HRTF), but is also related to the spatial parameters. Especially when discussing in conjunction with Figure 2, it is related to the spatial parameter set 1 and the spatial parameter Set 2 is related.
[0191] FIG. 14 shows the criteria used in FIG. 16a to represent a filter based on the head related transfer function (HRTF). In particular, describe a situation where a listener is located in a sweet spot between five speakers in a five-channel speaker setup. For example, this setup can be found in a general surround-home or movie theater entertainment system. For each channel, there are two head-related transfer functions (HRTF), which can be converted into a channel impulse response with a filter with the head-related transfer function (HRTF) as the transfer function. In particular, it is well known in the art that a filter based on the head related transfer function (HRTF) can be responsible for the sound propagation in the listener's head. Therefore, for example, the head in Figure 14 The correlation transfer function 1 (HRTF1) is responsible for the situation where the sound emitted from the speaker J reaches the right ear after passing near the listener's head. In contrast, from the left surround speaker L<sub>s</sub>The emitted sound almost directly reaches the left ear, and is only partially affected by the position of the ear on the head, the shape of the ear, and so on. Therefore, it is obvious that the head related transfer function 1 (HRTF 1) and the head related transfer function 2 (HRTF 2) are different from each other.
[0192] The same holds for the head-related transfer function 3 (HRTF 3) and the head-related transfer function 4 (HRTF 4) of the left channel, because the relationship between the binaural and the left channel L is different of. The same applies to other head-related transfer functions (HRTF), although it is obvious from Figure 14 that the head-related transfer function 5 (HRTF 5) for the center channel is the same as the head-related transfer function 6 (HRTF). 6) Almost the same or exactly the same as each other, unless the head related transfer function (HRTF) data can be used to adjust the asymmetry of each listener.
[0193] As stated above, these head related transfer functions (HRTFs) have been determined for simulated heads and can be downloaded for any specific "average head" and speaker settings.
[0194] Now, as 171 and 172 in FIG. 17 become obvious, a combination method is used to combine the left channel with the left surround channel to obtain the point indicated by L in FIG. 15 The two filters on the left are based on the head related transfer function (HRTF). The same step is also performed for the right side, as described with R in FIG. 15, which forms the head-related transfer function 13 (HRTF 13) and the head-related transfer function 14 (HRTF 14). For this purpose, also refer to items 173 and 174 in FIG. 17. However, it should be noted here that for each head related transfer function (HRTF) in the combined items 171, 172, 173, and 174, it is considered to reflect the left (L) channel and Ls sound between the original settings. The inter-channel intensity difference parameter between channels, or between the right (R) channel and the Rs channel of the original multi-channel setting. In particular, these parameters define the weighting factors for linear combination of head related transfer functions (HRTF).
[0195] As previously described, when combining head related transfer functions (HRTF), a phase factor can also be applied. The phase factor uses the time delay between the combined head related transfer functions (HRTF) or spreads the phase difference. Set
CN 102523552 Β
Righteousness. However, the phase factor is not related to the transmission parameters.
[0196] Therefore, head-related transfer functions 11, 12, 13 and 14 (HRTF 11, 12, 13, 14) are not real head-related transfer function (HRTF) filters, but are based on head-related transfer functions. (HRTF)-based filters, so these filters are only related to the head related transfer function (HRTF), and have nothing to do with the transmission signal. As an alternative, the parameters eld and cld are used to calculate these head-related transfer functions 11, 12, 13 and 14 (HRTF 11, 12, 13, 14), and the head-related transfer function 11 , 12, 13 and 14 (HRTF 11, 12, 13, 14) are also related to the transmission signal.
[0197] Now, the situation of FIG. 15 is obtained, which still has three channels instead of two transmission channels as included in the preferred downmix signal. Therefore, the six head-related transfer functions 11, 12, 5, 6, 13, 14 (HRTF 11, 12, 5, 6, 13, 14) must be combined into the four heads as described in Figure 16a. Partial related transfer functions 15, 16, 17, 18 (HRTF 15, 16, 17, 18).
[0198] For this purpose, the head related transfer functions 11, 5, 13 (HRTF 11, 5, 13) are combined using the left upmixing rule, which can be clearly known from the upmixing matrix in FIG. 16b. In particular, as shown in FIG. 16b and in the function block 175, the left upmixing rule contains parameters called i, m2i, and Bbl. This left upmixing rule is only used to multiply the left channel in the matrix equation of FIG. 16. Therefore, these three parameters are also called the left upmix rule.
[0199] As depicted in the functional block 176, the same head-related transfer functions 11, 5, and 13 (HRTF 11, 5, 13) are now combined using the right upmixing rule, in other words, in Fig. 16b In the embodiment, the parameters 2, 12 and 12 are all used to multiply the right channel R° in Fig. 16b.
[0200] Therefore, the head-related transfer function 15 (HRTF 15) and the head-related transfer function 17 (HRTF 17) are produced. Similarly, using the left parameters of the upmix to call 1, 1¾, and 1¾ to combine the head-related transfer functions 12, 6, 14 (HRTF 12, 6, 14) in Figure 15 to obtain the head-related transfer function 16 (HRTF 16 ). Use the head-related transfer functions 12, 6, 14 (HRTF 12, 6, 14) and use the upmix right parameters or right upmix rules specified by the calls 2,1¾ and 1¾ to make the corresponding combinations to obtain the graph The head related transfer function 18 (HRTF) in 16a.
[0201] It should be emphasized again that although the original head related transfer function (HRTF) in FIG. 14 has nothing to do with the transmission signal, the new filters 15, 16, and 16 based on the head related transfer function (HRTF) 17, 18 is now related to the transmission signal, because the spatial parameters contained in the multi-channel signal are used to calculate these filters 15, 16, 17, 18.
[0202] Finally, in order to obtain the stereo left channel and the stereo right channel Rb, the outputs of the filters 15 and 17 must be combined in the adder 130a. Similarly, the outputs of filters 16 and 18 must be combined in adder 130b. These adders 130a. 130b reflect the superposition of two signals in the human ear.
[0203] Next, FIG. 18 will be discussed. Figure 18 shows a preferred embodiment of the multi-channel decoder of the present invention for generating a stereo signal using a downmix signal derived from an original multi-channel signal. The downmix signal is at ζ] and z<sub>2</sub>Description, or use "L" and "R" to indicate. In addition, the downmix signal has parameters related to it. The parameters at least represent the difference in channel level between left and left surround, or the difference in channel level between right and right surround, and information related to the upmix rule. .
[0204] Naturally, when the original multi-channel signal is only a three-channel signal, no eld] or cld is transmitted.<sub>r</sub>, And as previously described, only the parameter side information will become the information of the upmixing rule, and this upmixing rule will cause energy errors in the upmixing signal. Therefore, although the waveform of the upmix signal matches the original waveform as much as possible when performing non-stereoscopic rendering, the energy of the upmix channel is different from the energy of the corresponding original channel.
[0205] In the preferred embodiment of FIG. 18, the upmixing rule information uses two upmixing parameters cP®, cpc<sub>2</sub>Reflected. However, any other upmixing rule information can also be applied and indicated by a specific number of bits. in particular,
CN 102523552 Β
A predetermined table at the decoder can be used to indicate a specific upmix scheme and upmix parameters, so it is only necessary to transmit the table index from the encoder to the decoder. Alternatively, different upmixing schemes can also be used, such as upmixing from two to more than three channels. Alternatively, more than two predicted upmixing parameters may be transmitted, which then requires corresponding and different downmixing rules consistent with the upmixing rules, as discussed in detail with respect to FIG. 20.
[0206] Regardless of this preferred embodiment for the upmixing rule, any upmixing rule that can be upmixed to produce an upmixing channel energy loss effect set can be applied, which corresponds to the original signal The set waveforms are consistent.
[0207] The multi-channel decoder of the present invention includes a gain factor calculator 180 for calculating at least one gain factor H, g<sub>r</sub>Or g, to reduce or eliminate energy errors. The gain factor calculator calculates the gain factor based on the upmix rule information and the filter characteristics based on the head related transfer function (HRTF) corresponding to the upmix channel to be obtained when the upmix rule is applied. However, as described before, in the three-dimensional rendering, this upmixing action is not performed. However, as discussed in connection with the functional blocks 175, 176, 177, and 178 of FIG. 15 and FIG. 17, head-related transfer function (HRTF)-based filters corresponding to these upmix channels are used.
[0208] As previously discussed, when 1 or r is inserted in place of n, the gain factor calculator 180 can calculate the different gain factors H and P as depicted in equation (27). Alternatively, the gain factor calculator 180 may generate a single gain factor for the two channels specified by equation (28).
[0209] What is important is that the gain factor calculator 180 of the present invention not only calculates the gain factor according to the upmixing rules, but also according to the filter characteristics based on the head related transfer function (HRTF) corresponding to the upmixing channel. This reflects the fact that the filter itself is also related to the transmission signal and is affected by energy errors. Therefore, the energy error is not only caused by the upmixing rule information such as the prediction parameters CPC] and CPC?, but is also affected by the filter itself.
[0210] Therefore, in order to obtain a well-adjusted gain correction, the gain factor of the present invention is not only related to the prediction parameter, but also related to the filter corresponding to the upmix channel.
[0211] The gain factor and downmix parameters and a filter based on the head related transfer function (HRTF) are used in the filter processor 182 to filter the downmix signal to obtain an energy-corrected stereo signal, It has a left stereo channel L<sub>b</sub>And has a right stereo channel R<sub>B</sub>o
[0212] In a preferred embodiment, the difference between the gain factor and the total energy contained in the channel impulse response corresponding to the upmix channel filter relative to this total energy, and the estimated upmix energy The error AE is related to the relationship. ΔE can preferably be calculated by combining the channel impulse responses corresponding to the upmix channel filters, and then calculating the energy of the combined channel impulse response. Because all the numbers used for Jiu and Gr in Fig. 18 are positive values, it is clearer according to the definition of AE and E, that is, both gain factors are greater than Κ. This reflects the experience described in Fig. 5, that is Most of the time, the energy of the stereo signal is less than the energy of the original multi-channel signal. It should also be noted that even when multi-channel gain compensation is applied, in other words, when the factor P is used in most signals, energy loss is still caused.
[0213] FIG. 19a depicts a preferred embodiment of the filter processor 182 in FIG. 18. In particular, FIG. 19a depicts the situation when the combined filters 15, 16, 17, and 18 of FIG. 16a are used without gain compensation in the functional block 182a, and the filter output signal is similar to that depicted in FIG. 13 plus. Then, the output of the function block 182a is input to the scaling function block 182b to use the gain factor calculated by the function block 180 to scale the output.
[0214] Alternatively, the filter processor may be constructed as shown in FIG. 19b. Here, the head related transfer functions 15 to 18 (HRTF 15-18) are calculated as described in the functional block 182c. Therefore, the calculator 182c performs head related transfer function (HRTF) combination without any gain adjustment. Next, a filter adjuster 182d is provided, which uses the gain factor calculated by the present invention. The filter adjuster forms an adjustment filter as shown in the functional block 180e, which
CN 102523552 Β
The middle function block 180e uses the adjustment filter to perform filtering, and performs subsequent summation of the corresponding filter output as shown in FIG. 13. Therefore, to obtain the gain-corrected stereo channel 5 and output, the post-scaling process as shown in Figure 19a is not required.
[0215] Generally speaking, as already described in conjunction with equations 16, 17 and 18, the estimated upmix error ΔE can be used for gain calculation. This approximation is particularly useful when the number of upmix channels is equal to the number of downmix channels+1. Therefore, in the case of two downmix channels, this approximation can work well for three upmix channels. Alternatively, when there are three downmix channels, this approximation also works well for a scheme with four upmix channels.
[0216] However, it should be noted that the gain factor calculation based on the upmix error estimate can also be performed in the following example case: where three downmix channels are used to perform five channel predictions. Alternatively, a prediction-based upmix can also be used, and the two downmix channels are upmixed into four upmix channels. Regarding the estimated upmix energy error AE, not only can this estimation error be directly calculated as shown in equation (25) for this preferred case, but also some information related to the actually occurring upmix error can be transmitted in the bitstream. However, even in other cases that are different from the special cases described in conjunction with equations (25) to (28), it can be based on the head-related transfer function (HRTF)-based filter used for the upmix channel, Use the prediction parameter to calculate the value E/. When considering equation (26), it is obvious that this equation can also be simply applied to the 2/4 prediction upmix scheme for the energy of the impulse response of the filter based on the head related transfer function (HRTF) The weighting factor of is also changed accordingly.
[0217] In view of this, it is obvious that the general structure of equation (27), that is, according to E<sup>β</sup>/(Ε<sup>β</sup>-ΔΕ<sup>β</sup>The method of calculating the gain factor based on the relationship of) can also be applied to other situations.
[0218] Next, the schematic implementation of the prediction-based encoder structure shown in FIG. 20 will be discussed, which can be used to generate downmix signals L and R and upmix rule information transmitted to the decoder, so that the decoder The gain compensation can be performed in the context of the stereo filter processor.
[0219] The downmixer 191 receives five original channels, or alternative reception is as with L<sub>s</sub>With R<sub>s</sub>The three original channels described. The downmixer 191 can work according to a predetermined downmixing rule. In this case, it is not necessary to indicate the downmixing rule described by the line segment 192. Naturally, the error minimizer 193 can change the downmixing rule to minimize the error between the reconstructed channel at the output of the upmixer 194 and the corresponding original input channel.
[0220] Therefore, the error minimizer 193 may change the downmix rule 192 or the upmixer rule 196 so that the reconstructed channel has a minimized prediction loss AE. In the error minimizer 193, the optimization problem can be solved by any known algorithm, and it is preferable to operate in a subband-wise manner, so that the reconstructed channel and the input channel are operated in a subband-wise manner. Minimize the difference between.
[0221] As stated before, the input channel can be the original channel L, *R, output, Co alternative, the input channel can be three channels L, R, C, where the input sound Lanes L and R can be derived using the corresponding one-to-two (OTT) functional blocks described in FIG. 11. Alternatively, when the original signal has only L, R, and C channels, these channels can also be regarded as "original channels".
[0222] FIG. 20 additionally describes that in addition to transmitting two prediction parameters, any upmixing rule information can also be used, as long as the decoder in the position can use the upmixing rule information for upmixing. Therefore, the upmixing rule information can also be entries in the lookup table or any information related to upmixing.
[0223] The present invention therefore provides an effective way to perform stereo decoding of multi-channel audio signals by means of head-related transfer function (HRTF) filtering based on available downmix signals and additional control data. The present invention provides a solution to the problem of spectrum coloring generated when combining prediction upmixing and stereo decoding.
CN 102523552 Β
[0224] According to the specific implementation requirements of the method of the present invention, the method of the present invention can be implemented in hardware or software. The realization can be carried out using a digital storage medium, especially a disc, a versatile digital disc (DVD) or a compact disc (CD), which has electronically readable control signals stored on it, and is compatible with programmable The computer systems cooperate to perform the method of the present invention. Generally speaking, the present invention is therefore a computer program product, which has program code stored on a machine-readable medium, and when the computer program is executed on a computer, the program code is operable to perform the method of the present invention. In other words, the method of the present invention is therefore a computer program that has program code that performs at least one method of the present invention when the computer program is executed on a computer.
[0225] Although it has been specifically illustrated and described with reference to specific embodiments before, those skilled in the art can understand that other changes in form and details can be made without departing from the spirit and viewpoint of the present invention. It is also understood that suitable and different changes can be made in different embodiments without departing from the broad concepts disclosed herein and encompassed by the appended claims.
CN 102523552 Β
Contents4
55 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| CN1758336A | Cites | China |
| US20030035553A1 | Cites | United States of America |
| WO2004028204A2 | Cites | World Intellectual Property Organization (WIPO) |
| CN1497585A | Cites | China |
| WO2006048203A1 | Cites | World Intellectual Property Organization (WIPO) |
67 members in 13 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 60803819 | United States of America | – | |
| 80381906 | United States of America | P | |
| 80381906 | United States of America | P | |
| 60803819 | – | – | – |
| US20060803819P | – | – | – |
Members67
| Document | Office | Kind | |
|---|---|---|---|
| US2007280485A1 | United States of America | A1 | |
| WO2007140809A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200803190A | Taiwan Province of China | A | |
| KR20090007471A | Republic of Korea | A | |
| EP2024967A1 | European Patent Office (EPO) | A1 | |
| CN101460997A | China | A | |
| HK1124156A | Hong Kong, China | A | |
| HK1124156A1 | Hong Kong, China | A1 | |
| JP2009539283A | Japan | A | |
| EP2216776A2 | European Patent Office (EPO) | A2 | |
| KR101004834B1 | Republic of Korea | B1 | |
| TWI338461B | Taiwan Province of China | B | |
| EP2024967B1 | European Patent Office (EPO) | B1 | |
| EP2216776A3 | European Patent Office (EPO) | A3 | |
| AT503244T | Austria | T | |
| ATE503244T1 | Austria | T1 | |
| US2011091046A1 | United States of America | A1 | |
| DE602006020936D1 | Germany | D1 | |
| HK1146975A | Hong Kong, China | A | |
| HK1146975A1 | Hong Kong, China | A1 | |
| US8027479B2 | United States of America | B2 | |
| SI2024967T1 | Slovenia | T1 | |
| JP4834153B2 | Japan | B2 | |
| CN102523552A | China | A | |
| CN102547551A | China | A | |
| CN102523552BThis record | China | B | |
| EP2216776B1 | European Patent Office (EPO) | B1 | |
| US2014343954A1 | United States of America | A1 | |
| CN102547551B | China | B | |
| ES2527918T3 | Spain | T3 | |
| US8948405B2 | United States of America | B2 | |
| MY157026A | Malaysia | A | |
| US9699585B2 | United States of America | B2 | |
| US2017272885A1 | United States of America | A1 | |
| US2018091914A1 | United States of America | A1 | |
| US2018098169A1 | United States of America | A1 | |
| US2018098170A1 | United States of America | A1 | |
| US2018109897A1 | United States of America | A1 | |
| US2018109898A1 | United States of America | A1 | |
| US2018132051A1 | United States of America | A1 | |
| US2018139558A1 | United States of America | A1 | |
| US2018139559A1 | United States of America | A1 | |
| US9992601B2 | United States of America | B2 | |
| US10015614B2 | United States of America | B2 | |
| US10021502B2 | United States of America | B2 | |
| US10085105B2 | United States of America | B2 | |
| US10091603B2 | United States of America | B2 | |
| US10097940B2 | United States of America | B2 | |
| US10097941B2 | United States of America | B2 | |
| US10123146B2 | United States of America | B2 | |
| US2019110149A1 | United States of America | A1 | |
| US2019110150A1 | United States of America | A1 | |
| US2019110151A1 | United States of America | A1 | |
| US2019116443A1 | United States of America | A1 | |
| US10412524B2 | United States of America | B2 | |
| US10412525B2 | United States of America | B2 | |
| US10412526B2 | United States of America | B2 | |
| US10469972B2 | United States of America | B2 | |
| US2020021937A1 | United States of America | A1 | |
| MY180689A | Malaysia | A | |
| US10863299B2 | United States of America | B2 | |
| US2021195357A1 | United States of America | A1 | |
| US11601773B2 | United States of America | B2 | |
| US2023209291A1 | United States of America | A1 | |
| US12052558B2 | United States of America | B2 | |
| MY207129A | Malaysia | A | |
| MY209726A | Malaysia | A |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Grant of patent or utility modelGrantedC14 | C14 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 102523552
- Publication, DOCDB
- 102523552
- Publication, EPODOC
- CN102523552B
- Application
- 2011104025254
- Application, DOCDB
- 201110402525
- Application, EPODOC
- CN201110402525
Titles2
- Chinese
- 非节能上混规则脉络立体多声道解码器
- English
- Non-energy-saving upmix regular context stereo multi-channel decoder
Classification
- CPC, 10
- G10L19/008
- H04S7/30
- H04S2400/01
- H04S2420/01
- H04S2420/03
- H03M7/30
- G11B20/10
- H04N21/439
- H04S7/307
- H04S2400/03
- IPC, 2
- H04S7 00
- G10L19 008