Apparatus and method for audio signal envelope encoding, processing and decoding by modelling a cumulative sum representation employing distribution quantization and coding
Abstract
A device for generating an audio signal envelope from one or more coded values is provided. The device includes: an input interface (1610) for receiving one or more coded values; and an envelope generator (1620) for generating an audio signal envelope according to the one or more coded values. The envelope generator (1620) is used to generate an aggregation function according to one or more coded values, where the aggregation function includes a plurality of aggregation points, wherein each of the aggregation points includes a parameter value and an aggregation value, wherein the aggregation function increases monotonically, and Each of the one or more coded values indicates at least one of a parameter value and an aggregate value of one of the aggregation points of the aggregation function. In addition, the envelope generator (1620) is used to generate the audio signal envelope so that the audio signal envelope includes a plurality of envelope points, wherein each of the envelope points includes a parameter value and an envelope value, and wherein the audio signal The envelope point of the envelope is assigned to each of the aggregation points of the aggregation function so that the parameter value of the envelope point is equal to the parameter value of the aggregation point. In addition, the envelope generator (1620) is used to generate the audio signal envelope so that the envelope value of each of the envelope points of the audio signal envelope depends on the aggregation value of at least one aggregation point of the aggregation function.

Term
7.7 yearsto projected expiry
Projected expiry 10 June 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
15 claims: 7 independent, 8 dependent
- 1一种用于从一个或多个编码值生成音频信号包络的装置,包括: 输入接口(1610),用于接收所述一个或多个编码值;以及 包络生成器(1620),用于依据所述一个或多个编码值生成所述音频信号包络; 其中所述包络生成器(1620)用于依据所述一个或多个编码值生成聚合函数,其中所 述聚合函数包括多个聚合点,其中所述聚合点中的每个包括参数值和聚合值,其中所述聚 合函数单调递增,并且其中所述一个或多个编码值中的每个指示所述聚合函数的所述聚合 点中的一个的所述参数值和所述聚合值中的至少一个, 其中所述包络生成器(1620)用于生成所述音频信号包络,以使得所述音频信号包络 包括多个包络点,其中所述包络点中的每个包括参数值和包络值,并且其中对于所述聚合 函数的所述聚合点中的每个,所述音频信号包络的所述包络点中的一个被分配给所述聚合 点,以使得所述包络点的所述参数值等于所述聚合点的所述参数值,并且 其中所述包络生成器(1620)用于生成所述音频信号包络,以使得所述音频信号包络 的所述包络点中的每个的所述包络值取决于所述聚合函数的至少一个聚合点的所述聚合 值。
- 2根据权利要求1所述的装置,其中所述包络生成器(1620)用于通过为所述一个或多 个编码值中的每个依据所述编码值确定所述聚合点中的一个,以及通过依据所述一个或多 个编码值中的每个的所述聚合点应用插值以获得所述聚合函数,以确定所述聚合函数。
- 3根据权利要求1或2所述的装置,其中所述包络生成器(1620)用于确定所述聚合函 数在所述聚合函数的多个聚合点处的一阶导数。
- 4根据前述权利要求中任一项所述的装置,其中所述包络生成器(1620)用于依据所 述编码值生成所述聚合函数,以便所述聚合函数具有连续的一阶导数。
- 5根据前述权利要求中任一项所述的装置,其中所述包络生成器(1620)用于通过确 定第一差值和第二差值的比值以确定所述音频信号包络,所述第一差值为所述聚合函数的 所述聚合点中的第一聚合点的第一聚合值(c(k+l))和所述聚合函数的所述聚合点中的第 二聚合点的第二聚合值(c(k-l) ;c(k))之间的差值,以及所述第二差值为所述聚合函数的 所述聚合点中的所述第一聚合点的第一参数值(f(k+l))和所述聚合函数的所述聚合点中 的所述第二聚合点的第二参数值(f(k-l) ;f(k))之间的差值。
- 6根据权利要求5所述的装置,其中所述包络生成器(1620)通过应用 刃=I:;];; 1 ])以确定所述音频信号包络; 其中t订t(k)指示所述聚合函数在第k个编码值处的导数, 其中c(k+l)为所述第一聚合值, 其中f(k+l)为所述第一参数值, 其中c(k-l)为所述第二聚合值, 其中f(k-l)为所述第二参数值, 其中k为指示所述一个或多个编码值中的一个的索引的整数, 其中c(k)为所述第一聚合值, 其中c(k+l)-c(k-l)为所述两个聚合值c(k+l)和c(k-l)的所述第一差值,以及 其中f(k+l)-f(k-l)为所述两个参数值f (k+1)和f(k-l)的所述第二差值。 7.根据权利要求5所述的装置,其中所述包络生成器(1620)用于通过应用 咖=ο”;鳥二常 + 鳥器卯以确定所述音频信号包络, 其中t订t(k)指示所述聚合函数在所述第k个编码值处的导数, 其中c(k+l)为所述第一聚合值, 其中f(k+l)为所述第一参数值, 其中c(k)为所述第二聚合值, 其中f (k)为所述第二参数值, 其中c(k-l)为所述聚合函数的所述聚合点中的第三聚合点的第三聚合值, 其中f(k-l)为所述聚合函数的所述聚合点中的所述第三聚合点的第三参数值, 其中k为指示所述一个或多个编码值中的一个的索引的整数, 其中c(k+l)-c(k)为所述两个聚合值c(k+l)和c(k)的所述第一差值,以及 其中f (k+1)-f (k)为所述两个参数值f (k+1)和f(k)的所述第二差值。 &根据前述权利要求中任一项所述的装置,其中所述输入接口 (1610)用于接收一个 或多个分裂值作为所述一个或多个编码值, 其中所述包络生成器(1620)用于依据所述一个或多个分裂值生成所述聚合函数,其 中所述一个或多个分裂值中的每个指示所述聚合函数的所述聚合点中的一个的所述聚合 值, 其中所述包络生成器(1620)用于生成所述重建的音频信号包络,以使得所述一个或 多个分裂点将所述重建的音频信号包络划分成两个或更多个音频信号包络部分,其中预定 义的分配规则为所述两个或更多个信号包络部分中的每个信号包络部分,依据所述信号包 络部分,定义信号包络部分值,并且 其中所述包络生成器(1620)用于生成所述重建的音频信号包络,以使得对于所述两 个或更多个信号包络部分中的每个,其信号包络部分值的绝对值大于其他信号包络部分中 的每个的所述信号包络部分值的绝对值的一半。
- 79. 一种用于确定用于对音频信号包络进行编码的一个或多个编码值的装置,包括: 聚合器(1710),用于为多个参数值中的每个确定聚合值,其中对所述多个参数值排序, 以使得当所述多个参数值中的第二参数值与所述多个参数值中的第一参数值不同时,所 述第一参数值在所述第二参数值之前或之后,其中包络值被分配给所述参数值中的每个, 其中所述参数值中的每个的所述包络值取决于所述音频信号包络,并且其中所述聚合器 (1710)用于为所述多个参数值中的每个参数值,依据所述参数值的所述包络值并依据在所 述参数值之前的多个参数值中的每个的所述包络值,确定所述聚合值;以及 编码单元(1720),用于依据所述多个参数值的聚合值中的一个或多个确定一个或多个 编码值。
- 810. 根据权利要求9所述的装置,其中所述聚合器(1710)用于为所述多个参数值中的 每个参数值,通过对所述参数值的所述包络值和在所述参数值之前的所述参数值的所述包 络值进行相加以确定所述聚合值。
- 911. 根据权利要求9或10所述的装置,其中所述参数值中的每个的所述包络值指示以 所述音频信号包络作为信号包络的音频信号包络的谱值的η次幕,其中η为大于0的偶数。
- 1012. 根据权利要求9或10所述的装置,其中所述参数值中的每个的所述包络值指示时 域中表示的并以所述音频信号包络作为信号包络的音频信号包络的幅值的η次幕,其中η 为大于0的偶数。
- 1113. 根据权利要求9-12中任一项所述的装置,其中所述编码单元(1720)用于依据所述 参数值的所述聚合值中的一个或多个并依据指示多少个值将被所述编码单元(1720)确定 作为所述一个或多个编码值的编码值数,确定所述一个或多个编码值。
- 1214. 根据权利要求13所述的装置,其中所述编码单元用于根据谁) = min:砒)罟® :确定所述一个或多个编码值, 其中c(k)指示待被所述编码单元确定的第k个编码值, 其中j指示所述多个参数值中的第j个参数值, 其中a(j)指示被分配给所述第j个参数值的所述聚合值, 其中max (a)指示作为被分配给所述参数值中的一个的所述聚合值中的一个的最大 值,其中被分配给所述参数值中的一个的所述聚合值均不大于所述最大值,并且 其中min] α。)晋® ”旨示作为所述参数值中的一个的最小值,为此 呼为最小。
- 1315. 一种用于从一个或多个编码值生成音频信号包络的方法,包括:: 接收所述一个或多个编码值;以及 依据所述一个或多个编码值生成所述音频信号包络, 其中通过依据所述一个或多个编码值生成聚合函数,进行生成所述音频信号包络,其 中所述聚合函数包括多个聚合点,其中所述聚合点中的每个包括参数值和聚合值,其中所 述聚合函数单调递增,并且其中所述一个或多个编码值中的每个指示所述聚合函数的所述 聚合点中的一个的所述参数值和所述聚合值中的至少一个, 其中生成所述音频信号包络被进行,以使得所述音频信号包络包括多个包络点,其中 所述包络点中的每个包括参数值和包络值,并且其中对于所述聚合函数的所述聚合点中的 每个,所述音频信号包络的所述包络点中的一个被分配给所述聚合点,以使得所述包络点 的所述参数值等于所述聚合点的所述参数值,并且 其中生成所述音频信号包络被进行,以使得所述音频信号包络的所述包络点中的每个 的所述包络值取决于所述聚合函数的至少一个聚合点的所述聚合值。
- 1416. 一种用于确定用于对音频信号包络进行编码的一个或多个编码值的方法,包括: 为多个参数值中的每个确定聚合值,其中对所述多个参数值排序,以使得当所述多个 参数值中的第一参数值与所述多个参数值中的第二参数值不同时,所述第一参数值在所述 二参数值之前或之后,其中包络值被分配给所述参数值中的每个,其中所述参数值中的每 个的所述包络值取决于所述音频信号包络,并且其中所述聚合器(1710)用于为所述多个 参数值中的每个参数值,依据所述参数值的所述包络值并依据在所述参数值之前的多个参 数值中的每个的所述包络值,确定所述聚合值;以及 依据所述多个参数值的聚合值中的一个或多个确定一个或多个编码值。
- 1517. 一种计算机程序,当被在计算机或信号处理器上执行时,用于实现权利要求15或 16所述的方法。
Independent claims15
492 paragraphs, as filed
Apparatus and method for encoding, processing, and decoding audio signal envelopes accumulated and represented by applying distributed quantization and coding modelingTechnical field
[0001] The present invention relates to a device and method for audio signal envelope encoding, processing, and decoding, and more particularly, to a device and method for audio signal envelope encoding, processing, and decoding using distributed quantization and encoding .
Background technique
[0002] Linear predictive coding (LPC) is a typical tool used to model the spectral envelope of the core bandwidth in a speech codec. The most common domain used to quantify LPC models is the line spectrum frequency (LSF) domain. It is based on the decomposition of LPC polynomials into two polynomials, and its roots are on the unit circle, so that they can be described only by their angle or frequency.
Summary of the invention
[0003] The object of the present invention is to provide an improved concept for the encoding and decoding of audio signal envelopes. The object of the present invention is achieved by the device according to claim 1, the device according to claim 9, the method according to claim 15, the method according to claim 16, and the computer program according to claim 17.
[0004] An apparatus for generating an audio signal envelope from one or more encoded values is provided. The device includes: an input interface for receiving one or more coded values; and an envelope generator for generating an audio signal envelope according to the one or more coded values. The envelope generator is used to generate an aggregation function based on one or more coded values, where the aggregation function includes multiple aggregation points, where each of the aggregation points includes a parameter value and an aggregation value, where the aggregation function increases monotonically, and one or more of them Each of the coded values indicates at least one of a parameter value and an aggregate value of one of the aggregation points of the aggregation function. In addition, the envelope generator is used to generate the audio signal envelope so that the audio signal envelope includes a plurality of envelope points, where each of the envelope points includes a parameter value and an envelope value, and where the envelope of the audio signal envelope An envelope point is assigned to each of the aggregation points of the aggregation function so that the parameter value of the envelope point is equal to the parameter value of the aggregation point. In addition, the envelope generator is used to generate the audio signal envelope so that the envelope value of each of the envelope points of the audio signal envelope depends on the aggregation value of at least one aggregation point of the aggregation function.
[0005] According to an embodiment, the envelope generator may, for example, be used to determine one of the aggregation points according to the code value for each of one or more coded values and to determine one of the aggregation points according to the one or more coded values. Apply interpolation to obtain the aggregation function for each aggregation point of to determine the aggregation function.
[0006] In one embodiment, the envelope generator may, for example, be used to determine the first derivative of the aggregation function at a plurality of aggregation points of the aggregation function.
[0007] According to an embodiment, the envelope generator may, for example, be used to generate an aggregate function according to the coded value, so that the aggregate function has a continuous first derivative.
[0008] In one embodiment, the envelope generator may, for example, be used to determine the audio signal envelope by applying edge period)={{Τ] [;
[0009] where t set t(k) indicates the derivative of the aggregated signal envelope at the k-th coded value, where c(k) is the aggregation function
The aggregation value of the kth aggregation point of the number, and where f (k) is the parameter value of the kth aggregation point of the aggregation function.
[0010] According to an embodiment, the input interface may be used to receive one or more split values as one or more coded values. The envelope generator may be used to generate an aggregation function according to one or more split values, where each of the one or more split values indicates an aggregation value of one of the aggregation points of the aggregation function. In addition, the envelope generator can be used to generate the reconstructed audio signal envelope such that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts, where the predefined The allocation rule is for each signal envelope part of two or more signal envelope parts, and the signal envelope part value is defined according to the signal envelope part. In addition, the envelope generator can be used to generate the reconstructed audio signal envelope so that for each of the two or more signal envelope parts, the absolute value of the signal envelope part value is greater than the other signal envelope parts Half of the absolute value of the signal envelope part value of each.
[0011] In addition, an apparatus for determining one or more encoding values for encoding an audio signal envelope is provided. The device includes: an aggregator for determining an aggregate value for each of a plurality of parameter values, wherein the plurality of parameter values are sorted so that when the second parameter value among the plurality of parameter values is compared with the second parameter value among the plurality of parameter values When the first parameter value is different, the first parameter value is before or after the second parameter value, wherein the envelope value is assigned to each of the parameter values, wherein the envelope value of each of the parameter values depends on the audio signal Envelope, and the aggregator is used to determine for each parameter value in a plurality of parameter values according to the envelope value of the parameter value and according to the envelope value of each of the plurality of parameter values preceding the parameter value Aggregate value. In addition, the device includes a coding unit for determining one or more coded values according to one or more of the aggregated values of the multiple parameter values.
[0012] According to an embodiment, the aggregator may, for example, be used for each parameter value of a plurality of parameter values to pass the envelope value of the parameter value and the envelope value of the parameter value before the parameter value Add them to determine the aggregate value.
[0013] In one embodiment, the envelope value of each of the parameter values may, for example, indicate the energy value of the audio signal envelope with the audio signal envelope as the signal envelope.
[0014] According to an embodiment, the envelope value of each of the parameter values may, for example, indicate n times the spectrum value of the audio signal envelope with the audio signal envelope as the signal envelope, where n is greater than 0 The even number.
[0015] In one embodiment, the envelope value of each of the parameter values may, for example, indicate n times the amplitude of the audio signal envelope represented in the time domain and using the audio signal envelope as the signal envelope Screen, where η is an even number greater than 0.
[0016] According to an embodiment, the coding unit may, for example, be used to determine the number of coding values as one or more coding values according to one or more of the aggregate values of the parameter values and depending on how many values are indicated by the coding unit , Determine one or more encoding values.
[0017] In an embodiment, the coding unit may, for example, be used according to c(^=min/'
V Ν) determining one or more coded values;
[0018] where c(k) indicates the k-th encoding value to be determined by the encoding unit, where j indicates the j-th parameter value among a plurality of parameter values, and where a(j) indicates that it is assigned to the j-th parameter value The aggregate value of, where max (a) indicates the maximum value of one of the aggregate values assigned to one of the parameter values, wherein none of the aggregate values assigned to one of the parameter values is greater than the maximum value, and
[0019] where min/qu)-you beer small indicator as the minimum value of one of the parameter values, for this
-k NJ
Yibuchong 1 is the smallest.
[0020] In addition, a method for generating an audio signal envelope from one or more encoded values is provided. The method includes:
[0021]-receive one or more coded values; and
[0022]-Generate an audio signal envelope based on one or more encoding values.
[0023] The generation of an audio signal envelope is performed by generating an aggregation function according to one or more encoding values, where the aggregation function includes a plurality of aggregation points, wherein each of the aggregation points includes a parameter value and an aggregation value, and the aggregation function is monotonically increasing , And each of the one or more coded values indicates at least one of a parameter value and an aggregate value of one of the aggregation points of the aggregation function. In addition, generating the audio signal envelope is performed so that the audio signal envelope includes a plurality of envelope points, wherein each of the envelope points includes a parameter value and an envelope value, and wherein the envelope point of the audio signal envelope is Assign each of the aggregation points of the aggregation function so that the parameter value of the envelope point is equal to the parameter value of the aggregation point. In addition, generating the audio signal envelope is performed so that the envelope value of each of the envelope points of the audio signal envelope depends on the aggregation value of at least one aggregation point of the aggregation function.
[0024] In addition, a method for determining one or more encoding values used to encode an audio signal envelope is provided. The method includes:
[0025] Determine an aggregate value for each of the multiple parameter values, wherein the multiple parameter values are sorted so that when the first parameter value in the multiple parameter values is different from the second parameter value in the multiple parameter values , The first parameter value is before or after the second parameter value, where the envelope value is assigned to each of the parameter values, where the envelope value of each of the parameter values depends on the audio signal envelope, and where the aggregate The device is used to determine the aggregate value for each parameter value in the plurality of parameter values according to the envelope value of the parameter value and according to the envelope value of each of the plurality of parameter values preceding the parameter value; and
[0026]-Determine one or more coded values according to one or more of the aggregated values of the multiple parameter values.
[0027] In addition, a computer program is provided, which, when executed on a computer or a signal processor, implements one of the above-mentioned methods.
[0028] An apparatus for decoding to obtain a reconstructed audio signal envelope is provided. The device includes: a signal envelope reconstructor for generating a reconstructed audio signal envelope according to one or more split points; and an output interface for outputting the reconstructed audio signal envelope. The signal envelope reconstructor is used to generate the reconstructed audio signal envelope, so that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts, where the predefined distribution rule is Each of the two or more signal envelope parts defines a signal envelope part value according to the signal envelope part. In addition, the signal envelope reconstructor is used to generate the reconstructed audio signal envelope so that for each of the two or more signal envelope parts, the absolute value of the signal envelope part value is greater than that of the other signal envelope parts. Half of the absolute value of each of the signal envelope part values.
[0029] According to one embodiment, the signal envelope reconstructor may, for example, be used to generate the reconstructed audio signal envelope such that for each of the two or more signal envelope parts, the signal envelope part The absolute value of the value is greater than 90% of the absolute value of the signal envelope part value of each of the other signal envelope parts.
[0030] In one embodiment, the signal envelope reconstructor may, for example, be used to generate a reconstructed audio signal envelope such that for each of the two or more signal envelope parts, its signal envelope The absolute value of the part value is greater than 99% of the absolute value of the signal envelope part value of each of the other signal envelope parts.
[0031] In another embodiment, the signal envelope reconstructor 110 may, for example, be used to generate a reconstructed audio signal packet
Envelope so that the value of the signal envelope part of each of the two or more signal envelope parts is equal to the signal envelope of each of the other signal envelope parts of the two or more signal envelope parts Partial value.
[0032] According to an embodiment, the signal envelope part value of each of the two or more signal envelope parts may, for example, depend on one or more energy values of the signal envelope part Or one or more power values. Alternatively, the signal envelope part value of each of the two or more signal envelope parts depends on any other value suitable for reconstructing the original or target level of the signal envelope part.
[0033] Envelope scaling can be achieved in many ways. Specifically, it can correspond to signal energy or spectral quality or the like (absolute size), or it can be a scale factor or gain factor (relative size) . Therefore, it can be encoded as an absolute value or a relative value, or it can be encoded as a previous value or a combination of previous values by difference. In some cases, the scaling can also be unrelated to other available data, or Can be inferred from other available data. The envelope should be reconstructed to its original or target level. Therefore, in general, the value of the signal envelope part depends on any value suitable for the original or target level of the audio signal envelope to be reconstructed.
[0034] In one embodiment, the device may, for example, further include: a split point for decoding one or more code points to obtain the position of each of the one or more split points according to the decoding rule decoder. The split point decoder may, for example, be used to analyze the total number of positions indicating the total number of possible split point positions, the number of split points indicating the number of one or more split points, and the number of split point states. In addition, the split point decoder may, for example, be used to generate an indication of the location of each of the one or more split points using the total number of positions, the number of split points, and the number of split point states.
[0035] According to one embodiment, the signal envelope reconstructor may, for example, be used to indicate the total energy value of the reconstructed audio signal envelope or according to the original or target level suitable for reconstructing the audio signal envelope. Any other value generates the reconstructed audio signal envelope.
[0036] In addition, an apparatus for decoding to obtain a reconstructed audio signal envelope according to another embodiment is provided. The device includes: a signal envelope reconstructor for generating a reconstructed audio signal envelope according to one or more split points; and an output interface for outputting the reconstructed audio signal envelope. The signal envelope reconstructor is used to generate the reconstructed audio signal envelope so that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts, wherein the predefined distribution rule is Each of the two or more signal envelope parts defines the value of the signal envelope part according to the signal envelope part. A predefined envelope part value is assigned to each of two or more signal envelope parts. The signal envelope reconstructor is used to generate the reconstructed audio signal envelope so that for each signal envelope part of two or more signal envelope parts, the absolute value of the signal envelope part value of the signal envelope part Is greater than 90% of the absolute value of the predefined envelope part value assigned to the signal envelope part, and makes the absolute value of the signal envelope part value of the signal envelope part smaller than that assigned to the signal envelope part 110% of the absolute value of the predefined envelope value.
[0037] In one embodiment, the signal envelope reconstructor is used to generate the reconstructed audio signal envelope such that the signal envelope part value of each of the two or more signal envelope parts is equal to the value assigned to the The pre-defined envelope part value of the signal envelope part.
[0038] In one embodiment, the predefined envelope part values of the at least two signal envelope parts are different from each other.
[0039] In another embodiment, the predefined envelope part value of each of the signal envelope parts is different from the predefined envelope part value of each of the other signal envelope parts.
[0040] In addition, an apparatus for reconstructing audio signals is provided. The device includes: a device for decoding to obtain a reconstructed audio signal envelope of an audio signal according to one of the above embodiments, and a device for audio according to the audio signal
A signal generator that generates an audio signal based on the signal envelope and other signal characteristics of the audio signal. Other signal characteristics are different from the audio signal envelope.
[0041] In addition, an apparatus for encoding an audio signal envelope is provided. The device includes: an audio signal envelope interface for receiving an audio signal envelope; and an audio signal envelope interface for receiving two or more audio signals for each of at least two split point configurations according to a predefined distribution rule At least one audio signal envelope part in the envelope part is a split point determiner that determines the value of the signal envelope part. Each of the at least two split point configurations includes one or more split points, wherein the one or more split points of each of the two or more split point configurations divide the audio signal envelope into two or more split points Envelope parts of multiple audio signals. The split point determiner is used to select one or more split points in one of the at least two split point configurations as one or more selected split points to encode the audio signal envelope, wherein the split point determiner is used to select one or more split points according to the at least two split point configurations. The signal envelope part value of each of at least one of the two or more audio signal envelope parts of each of the split point configurations selects one or more split points.
[0042] According to an embodiment, the signal envelope part value of each of the two or more signal envelope parts may, for example, depend on one or more energy values of the signal envelope part Or one or more power values. Alternatively, the signal envelope part value of each of the two or more signal envelope parts depends on any other value suitable for reconstructing the original or target level of the audio signal envelope.
[0043] As already mentioned, the scaling of the envelope can be achieved in many ways. Specifically, it can correspond to signal energy or spectral quality or the like (absolute size), or it can be a scale factor or gain factor (relative size). Therefore, it can be coded as an absolute value or a relative value, or it can be coded as a previous value or a combination of previous values by difference. In some cases, scaling can also be irrelevant to other available data, or can be inferred from other available data. The envelope should be reconstructed to its original or target level. Therefore, in general, the value of the signal envelope portion depends on any value suitable for the original or target level of the audio signal envelope to be reconstructed.
[0044] In one embodiment, the device may, for example, further include: a split point encoder for encoding the position of each of the one or more split points to obtain one or more code points. The split point encoder may, for example, be used to encode the position of each of the one or more split points by encoding the number of split point states. In addition, the split point encoder may, for example, be used to provide a total number of positions indicating the total number of possible split point positions and a number of split points indicating the number of one or more split points. The number of split point states, the number of total positions, and the number of split points together indicate the position of each of the one or more split points.
[0045] According to an embodiment, the apparatus may, for example, further include: an energy determiner for determining the total energy of the audio signal envelope and encoding the total energy of the audio signal envelope. Alternatively, the device may, for example, be further used to determine any other value suitable for reconstructing the original or target level of the audio signal envelope.
[0046] In addition, an apparatus for encoding an audio signal is provided. The device includes: a device for encoding for encoding an audio signal envelope of an audio signal according to one of the above embodiments; and a secondary signal feature encoder for encoding other signal characteristics of the audio signal , Other signal characteristics are different from the audio signal envelope.
[0047] In addition, a method for decoding to obtain a reconstructed audio signal envelope is provided. The method includes:
[0048]-generating a reconstructed audio signal envelope according to one or more split points; and
[0049] Output the reconstructed audio signal envelope.
[0050] The generation of the reconstructed audio signal envelope is performed such that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts, wherein the predefined distribution rule is two For each signal envelope part of one or more signal envelope parts, the signal envelope part value is defined according to the signal envelope part. In addition, generating heavy
The created audio signal envelope is executed so that for each of the two or more signal envelope parts, the absolute value of the signal envelope part value is greater than the signal envelope part of each of the other signal envelope parts Half of the absolute value of the value.
[0051] In addition, a method for decoding to obtain a reconstructed audio signal envelope is provided. The method includes:
[0052]-generating a reconstructed audio signal envelope according to one or more split points; and
[0053] Output the reconstructed audio signal envelope.
[0054] The generation of the reconstructed audio signal envelope is performed so that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts, wherein the predefined distribution rule is two For each signal envelope part of one or more signal envelope parts, the signal envelope part value is defined according to the signal envelope part. A predefined envelope part value is assigned to each of two or more signal envelope parts. In addition, the generation of the reconstructed audio signal envelope is performed so that for each of the two or more signal envelope parts, the absolute value of the signal envelope part value of the signal envelope part is greater than that of the signal envelope part. 90% of the absolute value of the predefined envelope part value assigned to the signal envelope part, and make the absolute value of the signal envelope part value of the signal envelope part smaller than the predefined value assigned to the signal envelope part 110% of the absolute value of the envelope part value.
[0055] In addition, a method for encoding an audio signal envelope is provided. The method includes:
[0056]-Receive audio signal envelope;
[0057] Determine the signal envelope for at least one of the two or more audio signal envelope parts for each of the at least two split point configurations according to a predefined allocation rule Partial value, where each of the at least two split point configurations includes one or more split points, wherein the one or more split points of each of the two or more split point configurations divide the audio signal envelope into Two or more audio signal envelope parts;
[0058] and
[0059]-Select one or more split points in one of the at least two split point configurations as one or more selected split points to encode the audio signal envelope, wherein each of the at least two split point configurations is A signal envelope part value of each of at least one audio signal envelope part of two or more audio signal envelope parts is performed to select one or more split points.
[0060] In addition, a computer program is provided, which when it is executed on a computer or a signal processor, is used to implement one of the above-mentioned methods.
[0061] The heuristic but slightly inaccurate description of line spectrum frequency 5 (LSF5) is so, they describe the distribution of signal energy along the frequency axis. There is a high probability that LSF5 will reside at a frequency where the signal has a lot of energy. Based on this discovery, the embodiment adopts the exploratory description academically and quantifies the actual distribution of signal energy. Since LSF only applies this idea approximately, according to the embodiment, the LSF concept is omitted, and on the contrary, the frequency distribution is quantized, so that a smooth envelope shape can be created from this distribution. Hereinafter, this inventive concept is referred to as distributed quantization.
[0062] The embodiment is based on the quantization and coding of the spectral envelope used in speech and audio coding. The embodiment may, for example, be applied to the envelope of the core bandwidth and the bandwidth extension method.
[0063] According to embodiments, standard envelope modeling techniques (eg, scale factor band [3, 4] and linear prediction model [1]) can be replaced and/or improved, for example.
[0064] The purpose of the embodiment is to obtain quantification that combines the advantages of the linear prediction method and the method based on scale factor bands while removing their disadvantages.
[0065] According to an embodiment, an idea is provided, on the one hand, it has a smooth and precise spectral envelope, on the other hand, it can be encoded with a small number of bits (optionally at a fixed bit rate) and further with Reasonable computational complexity is realized
Now.
Description of the drawings
[0066] Hereinafter, embodiments of the present invention will be described in more detail with reference to the accompanying drawings, in which:
[0067] FIG. 1 shows a device for decoding to obtain a reconstructed audio signal envelope according to an embodiment;
[0068] FIG. 2 shows a device for decoding according to another embodiment, wherein the device further includes a split point decoder;
[0069] FIG. 3 shows an apparatus for encoding an audio signal envelope according to an embodiment;
[0070] FIG. 4 shows a device for encoding an audio signal envelope according to another embodiment, wherein the device further includes a split point encoder;
[0071] FIG. 5 shows an apparatus for encoding an audio signal envelope according to another embodiment, wherein the apparatus for encoding an audio signal envelope further includes an energy determiner;
[0072] FIG. 6 shows three signal envelopes described by a constant energy block according to an embodiment;
[0073] FIG. 7 shows a cumulative representation of the spectrum of FIG. 6 according to an embodiment;
[0074] FIG. 8 shows the interpolated spectral quality envelope of the original representation and the cumulative quality domain representation;
[0075] FIG. 9 shows a decoding process for decoding the split point position according to an embodiment;
[0076] FIG. 10 shows a pseudo code for realizing the decoding of the split point position according to an embodiment;
[0077] FIG. 11 shows an encoding process for encoding a split point according to an embodiment;
[0078] FIG. 12 depicts a pseudo code for encoding the position of a split point according to an embodiment of the present invention;
[0079] FIG. 13 shows a split point decoder according to an embodiment;
[0080] FIG. 14 shows an apparatus for encoding an audio signal according to an embodiment;
[0081] FIG. 15 shows an apparatus for reconstructing audio signals according to an embodiment;
[0082] FIG. 16 illustrates an apparatus for generating an audio signal envelope from one or more encoded values according to an embodiment;
[0083] FIG. 17 illustrates an apparatus for determining one or more encoding values used to encode an audio signal envelope according to an embodiment;
[0084] FIG. 18 shows an aggregation function according to the first example; and
[0085] FIG. 19 shows an aggregation function according to a second example.
Detailed ways
[0086] FIG. 3 shows an apparatus for encoding an audio signal envelope according to an embodiment.
[0087] The device includes: an audio signal envelope interface 210 for receiving an audio signal envelope.
[0088] In addition, the device includes a split point determiner 220, and the split point determiner 220 is configured to provide two or more audio signals for each of the at least two split point configurations according to a predefined allocation rule. At least one audio signal envelope part in the envelope part determines the value of the signal envelope part.
[0089] Each of the at least two split point configurations includes one or more split points, wherein the one or more split points of each of the two or more split point configurations divide the audio signal envelope into two One or more audio signal envelope parts. The split point determiner 220 is configured to select one or more split points of one of the at least two split point configurations as one or more selected split points to encode the audio signal envelope, wherein the split point determiner 220 is used to One or more split points are selected according to the signal envelope part value of each of at least one of the two or more audio signal envelope parts in each of the at least two split point configurations .
[0090] The split point configuration includes one or more split points, and is defined by the split points. For example, the audio signal envelope may include 20 samples: 0,, 19, and a configuration with two split points may be defined by the first split point located at the position of sample 3 and the second split point located at the position of sample 8. , For example, the split point configuration can be indicated by the tuple (3; 8). If only one split point should be determined, a single split point indicates the split point configuration.
[0091] One or more suitable split points should be determined as one or more selected split points. To this end, consider at least two split point configurations, where each split point configuration includes one or more split points. Choose one or more split points for the most suitable split point configuration. Determine whether one split point configuration is more suitable than another split point configuration based on the value of the signal envelope part determined according to the predefined allocation rule.
[0092] In an embodiment where the split point configuration has N split points, every possible split point configuration with split points may be considered. However, in some embodiments, not all possible split point configurations are considered, but only two split point configurations are considered. The split point of the most suitable split point configuration is selected as one or more selected split points.
[0093] In embodiments where only a single split point should be determined, each split point configuration includes only a single split point. In an embodiment where two split points should be determined, each split point configuration includes two split points. Similarly, in an embodiment where N split points should be determined, each split point configuration includes N split points.
[0094] The split point configuration with a single split point divides the audio signal envelope into two audio signal envelope parts. The split point configuration with two split points divides the audio signal envelope into three audio signal envelope parts. The split point configuration with N split points divides the audio signal envelope into N+1 audio signal envelope parts.
[0095] There is a predefined allocation rule that allocates a signal envelope part value to each of the audio signal envelope parts. The predefined distribution rules depend on the envelope part of the audio signal.
[0096] In some embodiments, the split point is determined so that each of the audio signal envelope parts obtained by dividing the audio signal envelope by one or more split points has a substantially equal distribution by a predefined distribution rule. The value of the envelope part of the signal. Therefore, since one or more splitting points depend on the audio signal envelope and distribution rules, if the distribution rules and splitting points are known at the decoder, the audio signal envelope can be estimated at the decoder. For example, as shown in Figure 6.
[0097] In FIG. 6(a), a single split point for the signal envelope 610 should be determined. Therefore, in this example, the different possible split point configurations are defined by a single split point. In the embodiment of FIG. 6(a), the split point 631 is found as the best split point. The split point 631 divides the audio signal envelope 610 into two signal envelope parts. The rectangular block 611 represents the energy of the first signal envelope portion defined by the split point 631. The rectangular block 612 represents the energy of the second signal envelope portion defined by the split point 631. In the example of FIG. 6(a), the upper edges of blocks 611 and 612 represent an estimate of the signal envelope 610. This estimate can be formed at the decoder, for example, using the information of the split point 631 (for example, if the only split point has the value s = 12, then the split point s is located at position 12), as to where the signal envelope starts (Here, point 638) and information about where the signal envelope ends (here, point 639). The signal envelope can start and end at a fixed value, and this information can be acquired as fixed information at the receiver. Alternatively, this information can be transmitted to the receiver. On the decoder side, The decoder can reconstruct the estimate of the signal envelope, so that the signal envelope part obtained by splitting the audio signal envelope by the splitting point 631 obtains the same value assigned by a predefined rule. In FIG. 6(a), the signal envelope portion of the signal envelope defined by the upper edges of the blocks 611 and 612 obtains the same value assigned by the allocation rule, and represents a good estimate of the signal envelope 610. In addition to using the split point 631, the value 621 can also be used as the split point. In addition, in addition to the start value 638, the value 628 can also be used as the start value, and in addition to the end value 639, the end value 629 can also be used as the end value. However, not only the abscissa value is coded, but the ordinate value is also coded, which requires more coding resources, and this is not necessary.
[0098] In FIG. 6(b), three split points for the signal envelope 640 should be determined. So in this example, there are three
The split point defines different possible split point configurations. In the embodiment of FIG. 6(b), the split points 661,662,663 are found as the best split points. The split points 661, 662, 663 divide the audio signal envelope 640 into four signal envelope parts. The rectangular block 641 represents the energy of the first signal envelope portion defined by the split point. The rectangular block 642 represents the energy of the second signal envelope portion defined by the split point. The rectangular block 643 represents the energy of the third signal envelope portion defined by the split point. The rectangular block 644 represents the energy of the fourth signal envelope portion defined by the split point. In the example of Figure 6(b), blocks 641,642, The upper edge of 643,644 represents an estimate of the signal envelope 640. This estimate can be formed at the decoder, for example, using information about splitting points 661, 662, 663, information about where the signal envelope starts (here, point 668), and information about where the signal envelope part ends Information (here, point 669). The signal envelope can start and end at a fixed value, and this information can be acquired as fixed information at the receiver. Alternatively, this information can be transmitted to the receiver. On the decoder side, the decoder can reconstruct the estimate of the signal envelope, so that the signal envelope part obtained by splitting the audio signal envelope by the splitting points 661, 662, 663 obtains the same value assigned by the predefined allocation rule. In FIG. 6(b), the signal envelope portion of the signal envelope defined by the upper edges of the blocks 641, 642, 643, 644 obtains the same value assigned by the allocation rule, and represents a good estimate of the signal envelope 640. In addition to using split points 661, 662, and 663, values 651, 652, and 653 can also be used as split points. Furthermore, in addition to the start value 668, the value 658 may also be used as the start value, and in addition to the end value 669, the end value 659 may be used as the end value. However, not only the abscissa value is coded, but the ordinate value is also coded, which requires more coding resources, and this is not necessary.
[0099] In FIG. 6(c), four split points for the signal envelope 670 should be determined. Therefore, in this example, four split points define different possible split point configurations. In the embodiment of FIG. 6(c), the split points 691, 692, 693, and 694 are found to be the best split points. Split points 691, 692, 693, 694 divide the audio signal envelope 670 into five signal envelope parts. The rectangular block 671 represents the energy of the first signal envelope portion defined by the split point. The rectangular block 672 represents the energy of the second signal envelope portion defined by the split point. The rectangular block 673 represents the energy of the third signal envelope portion defined by the split point. The rectangular block 674 represents the energy of the fourth signal envelope portion defined by the split point. The rectangular block 675 represents the energy of the fifth signal envelope portion defined by the split point. In the example of FIG. 6(c), the upper edges of the blocks 671, 672, 673, 674, 675 represent the estimate of the signal envelope 670. This estimate can be formed at the decoder, for example, using information about splitting points 691, 692, 693, 694, information about where the signal envelope starts (here, point 698), and information about where part of the signal envelope is The ending message (here, point 699). The signal envelope can start and end at a fixed value, and this information is available as fixed information at the receiver. Alternatively, this information can be transmitted to the receiver. On the decoder side, the decoder can reconstruct the signal packet The signal envelope part obtained by splitting the audio signal envelope by the splitting points 691, 692, 693, 694 obtains the same value assigned by the predefined distribution rule. In Figure 6(c), the signal envelope portion of the signal envelope defined by the upper edges of the blocks 671, 672, 673, 674, 675 obtains the same value assigned by the allocation rule, and represents a good estimate of the signal envelope 670 . In addition to using split points 691, 692, 693, and 694, values 681, 682, 683, and 684 can also be used as split points. In addition, in addition to the start value 698, the value 688 may be used as the start value, and in addition to the end value 699, the end value 689 may be used as the end value. However, not only the abscissa value is coded, but the ordinate value is also coded, which requires more coding resources, and this is not necessary.
[0100] As for other specific embodiments, the following examples can be considered:
[0101] The signal envelope expressed in the spectral domain should be coded. The signal envelope may, for example, include n spectral values (e.g., n = 33).
[0102] At this time, different signal envelope parts can be considered. For example, the first signal envelope part can include the first 10 spectral values<sub>V1</sub>(i = 0,-, 9, taking i as the index of the spectrum value), and the second signal envelope part can include the last 23 spectrum values (i
=10,…,32)。
[0103] In an embodiment, the predefined allocation rule may be, for example, having a spectrum value v. , %..., v" of the spectral signal envelope part m, the signal envelope part value p(m) is the energy of the spectral signal envelope part, such as: upperbound
[0104] p{m) = V ν;'j-lowerbomul
[0105] where lowerbound is the lower limit of the signal envelope part m, and where upperbound is the upper limit of the signal envelope part m.
[0106] The signal envelope part value determiner 110 may assign a signal envelope part value to one or more audio signal envelope parts according to this formula.
[0107] At this time, the split point determiner 220 is used to determine one or more signal envelope part values according to a predefined allocation rule. In particular, the split point determiner 220 is used to determine one or more signal envelope part values according to the allocation rule, so that the signal envelope part value of each of the two or more signal envelope parts (approximately) Equal to the signal envelope part value of each of the other signal envelope parts of the two or more signal envelope parts.
[0108] For example, in certain embodiments, the split point determiner 220 may be used to determine only a single split point. In this embodiment, for example, according to the formula /^)=Σ<and, =Σ>", two signal envelope parts are defined by the split point s, such as signal envelope part 1 (m=1) and signal packet Network part 2 (m = 2);
[0109] where n indicates the number of samples of the audio signal envelope, such as the number of spectral values of the audio signal envelope. In the above example, n can be, for example, n=33.
[0110] The signal envelope part value determiner 110 may assign this signal envelope part value p(1) to the audio signal envelope part 1 and assign this signal envelope part value p(2) to the audio signal envelope part 2.
[0111] In some embodiments, the signal envelope part values p(1) and p(2) are determined. However, in some embodiments, only one of the two signal envelope part values is considered. For example, if the total energy is known, it is sufficient to determine the split point so that P(l) is approximately 50% of the total energy.
[0112] In some embodiments, s(k) may be selected from a set of possible values (for example, from a set of integer index values, such as {0; 1; 2; ···; 32}). In other embodiments, s(k) may be selected from a set of possible values (e.g., from a set of frequency values indicating a set of frequency bands).
[0113] In an embodiment where more than one splitting point should be determined, a formula representing the accumulated energy (the sample energy accumulated until the splitting point s) can be considered:
[0114]
[0115]
[0116]
[0117]
[0118]
Σ If N split points should be determined, then determine the split points s (1), s (2), tone 7V+1 where totalenergy is the total energy of the signal envelope.
, S (Ν), so that:
In one embodiment, the split point s (k) can be chosen such that
<img file="CN105431902A_D0001.tif" />
v,<sup>2</sup> -k) iotalenergy
-N + 1 is the smallest.
[0119] Therefore, according to an embodiment, the split point determiner 220 may, for example, be used to determine one or more split points s(k) so as to minimize /iron;
[0120] where totalenergy indicates total energy, and where k indicates the k-th splitting point of one or more splitting points, and where N indicates the number of one or more splitting points.
[0121] In another embodiment, if the split point determiner 220 is used to select only a single split point s, the split point determiner 220 may test all possible split points s=1,...,32.
[0122] In some embodiments, the split point determiner 220 may select the best value for the split point s, such as
Μ j = £-£ ν; the smallest split point s.
:,Two 0; ::.
[0123] According to an embodiment, the signal envelope part value of each of the two or more signal envelope parts may, for example, depend on one or more energy values of the signal envelope part Or one or more power values. Alternatively, the signal envelope part value of each of the two or more signal envelope parts may, for example, depend on any other value suitable for reconstructing the original or target level of the audio signal envelope.
[0124] According to an embodiment, the audio signal envelope may, for example, be represented in the spectral domain or the time domain.
[0125] FIG. 4 shows an apparatus for encoding an audio signal envelope according to another embodiment, wherein the apparatus further includes an apparatus for encoding one or more split points (for example, according to an encoding rule) to obtain Split point encoder 225 for one or more code points.
[0126] The split point encoder 225 may, for example, be used to encode the position of each of the one or more split points to obtain one or more code points. The split point encoder 225 may, for example, be used to encode the position of each of the one or more split points by encoding the number of split point states. In addition, the split point encoder 225 may, for example, be used to provide a total number of positions indicating the total number of possible split point positions and a number of split points indicating the number of one or more split points. The number of split point states, the number of total positions, and the number of split points together indicate the position of each of the one or more split points.
[0127] FIG. 5 shows an apparatus for encoding an audio signal envelope according to another embodiment, wherein the apparatus for encoding an audio signal envelope further includes an energy determiner 230.
[0128] According to an embodiment, the apparatus may, for example, further include an energy determiner (230) for determining the total energy of the audio signal envelope and for encoding the total energy of the audio signal envelope.
[0129] However, in another embodiment, the device may, for example, be used to determine any other value suitable for reconstructing the original or target level of the audio signal envelope. In addition to the total energy, a number of other values are suitable for reconstructing the original or target level of the audio signal envelope. For example, as already mentioned, the envelope scaling can be achieved in many ways, it can correspond to the signal energy or spectral quality or the like (absolute size), or it can be a scale factor or gain factor (relative size), Therefore, it can be encoded as an absolute value or a relative value, or it can be encoded as a previous value or a combination of previous values by difference. In some cases, the scaling can also be unrelated to other available data, or Can be inferred from other available data. The envelope should be reconstructed to its original or target level.
[0130] FIG. 14 shows an apparatus for encoding an audio signal. The device includes: a device 1410 for encoding according to one of the above-mentioned embodiments to encode the audio signal envelope of the audio signal by generating one or more split points; and other signal characteristics for the audio signal The secondary signal feature encoder 1420 for encoding. other
The signal characteristics are different from the audio signal envelope. Those skilled in the art realize that the audio signal itself can be reconstructed from the signal envelope of the audio signal and from other signal characteristics of the audio signal. For example, the signal envelope may, for example, indicate the energy of the samples of the audio signal. Other signal characteristics may, for example, indicate whether for each sample in the time domain audio signal, the sample has a positive value or a negative value.
[0131] FIG. 1 shows an apparatus for decoding to obtain a reconstructed audio signal envelope according to an embodiment.
[0132] The device includes a signal envelope reconstructor 110 for generating a reconstructed audio signal envelope according to one or more split points.
[0133] In addition, the device includes an output interface 120 for outputting the reconstructed audio signal envelope.
[0134] The signal envelope reconstructor 110 is used to generate a reconstructed audio signal envelope such that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts.
[0135] The predefined allocation rule is for each signal envelope part of two or more signal envelope parts, and the signal envelope part value is defined according to the signal envelope part.
[0136] In addition, the signal envelope reconstructor 110 is used to generate the reconstructed audio signal envelope so that for each of the two or more signal envelope parts, the absolute value of the signal envelope part value is greater than the other Half of the absolute value of the signal envelope part value of each of the signal envelope parts.
[0137] As for the absolute value a of the signal envelope part value X, it is expressed as:
[0138] If X 2 0, then a = x;
[0139] If x<0, then a=-Xo
[0140] If all the signal envelope part values are positive, this above-mentioned concept implies that the reconstructed audio signal envelope is generated so that for each of the two or more signal envelope parts, its signal The envelope part value is greater than half of the signal envelope part value of each of the other signal envelope parts.
[0141] In a specific embodiment, the signal envelope part value of each of the signal envelope parts is equal to the signal envelope of each of the other signal envelope parts of the two or more signal envelope parts Partial value.
[0142] However, in the more general embodiment of FIG. 1, the audio signal envelope is reconstructed so that the signal envelope part values of the signal envelope part do not have to be completely equal. On the contrary, a certain degree of error (a certain range) is allowed.
[0143] The concept "so that for each of two or more signal envelope parts, the absolute value of the signal envelope part value is greater than the value of the signal envelope part of each of the other signal envelope parts. "Half of the absolute value" can, for example, be understood to mean that as long as the maximum absolute value of all signal envelope part values is not twice the minimum absolute value of all signal envelope part values, the requirement is met.
[0144] For example, a set of four signal envelope part values {0. 23; 0, 28; 0, 19; 0, 30} satisfies the above requirements because 0.30<2*0.19=0.38. However, another set of four signal envelope values {0. 24; 0, 16; 0, 35; 0, 25} does not meet the requirements of the requirement, because 0.35>2*0. 16 = 0.32 .
[0145] On the decoder side, the signal envelope reconstructor 110 is used to reconstruct the reconstructed audio signal envelope, so that the audio signal envelope part obtained by dividing the reconstructed audio signal envelope by the split point has a substantially equal signal envelope Partial value. Therefore, the signal envelope part value of each of the two or more signal envelope parts is greater than the signal envelope part value of each of the other signal envelope parts in the two or more signal envelope parts Half of it.
[0146] In this embodiment, the value of the signal envelope part of the signal envelope part should be substantially equal, but not necessarily equal.
[0147] It is expected that the value of the signal envelope part of the signal envelope part should be exactly equal to indicate how the decoder should reconstruct the signal. When the signal envelope part is reconstructed so that the value of the signal envelope part is completely equal, it is strictly limited to the signal on the decoder side.
The degree of freedom for reconstruction.
[0148] The greater the deviation that can exist between the signal envelope part values, the greater the degree of freedom that the decoder has to adjust the audio signal envelope according to the specifications of the decoder side. For example, when encoding the spectral audio signal envelope, some decoders may preferably place more energy on the lower frequency band, while other decoders may preferably place more energy on the higher frequency band. on. And, by allowing a certain error, a limited number of rounding errors such as caused by quantization and/or dequantization can be allowed.
[0149] In an embodiment in which the signal envelope reconstructor 110 performs the reconstruction quite accurately, the signal envelope reconstructor 110 is used to generate the reconstructed audio signal envelope such that for two or more signal envelope parts For each of the signal envelope parts, the absolute value of the signal envelope part value is greater than 90% of the absolute value of the signal envelope part value of each of the other signal envelope parts.
[0150] According to an embodiment, the signal envelope reconstructor 110 may, for example, be used to generate a reconstructed audio signal envelope such that for each of the two or more signal envelope parts, its signal envelope The absolute value of the part value is greater than 99% of the absolute value of the signal envelope part value of each of the other signal envelope parts.
[0151] However, in another embodiment, the signal envelope reconstructor 110 may, for example, be used to generate a reconstructed audio signal envelope such that the signal of each of the two or more signal envelope parts The envelope part value is equal to the signal envelope part value of each of the other signal envelope parts of the two or more signal envelope parts.
[0152] In an embodiment, the signal envelope part value of each of the two or more signal envelope parts may, for example, depend on one or more energies of the signal envelope part Value or one or more power values.
[0153] According to an embodiment, the reconstructed audio signal envelope may, for example, be represented in the spectral domain or the time domain.
[0154] FIG. 2 shows a device for decoding according to another embodiment, wherein the device further includes a split point decoder 105 for decoding one or more code points according to a decoding rule To obtain one or more split points.
[0155] According to an embodiment, the signal envelope reconstructor 110 may, for example, be used to indicate the total energy value of the reconstructed audio signal envelope according to the total energy value or according to the original or target level suitable for reconstructing the audio signal envelope. Any other value of, generates the reconstructed audio signal envelope.
[0156] At this time, in order to illustrate the present invention in more detail, specific embodiments are provided.
[0157] According to a specific embodiment, the idea is to split the frequency band into two parts so that the two halves have the same energy. This idea is described in Figure 6(a), where the envelope is described by a constant energy block, that is, the overall shape.
[0158] This idea can then be applied recursively so that both halves can be further split into two halves with the same energy. This method is shown in Figure 6(b).
[0159] More generally, the spectrum can be divided into N blocks so that each block has 1/N energy. In Figure 6(c), N =
5 shows this.
[0160] In order to reconstruct these block-like constant spectral envelopes in the decoder, the frequency boundaries of the blocks and, for example, the total energy can be transmitted. Then the frequency boundary corresponds to the LSF representation of LPC only in the exploratory sense.
[0161] So far, the energy envelope abs(x) of the signal X has been provided<sup>2</sup>explanation of. However, in other embodiments, the amplitude envelope abs(X), some other power abs(x) of the spectrum, or the performance of any perceptual excitation (such as volume) are modeled. In addition to energy, you can refer to the term "spectrum" Quality" and assume that it describes a suitable representation of the spectrum. The only important thing is that the cumulative sum of the spectrum representation can be calculated, that is, the representation has only positive values.
[0162] However, if the sequence is not positive, it can be converted into a positive sequence by adding a large enough constant, by calculating its cumulative sum or by other suitable operations. Similarly, you can convert complex-valued sequences, for example:
[0163] 1) Two sequences, one of which is a pure real number and the other is a pure imaginary number; or
[0164] 2) Two sequences, where the first represents the amplitude and the second represents the phase. Then, the two sequences can be modeled as separate envelopes in both cases.
[0165] It is not necessary to limit the model to a spectral envelope model, and any envelope shape can be described with the current model. For example, Transient Noise Shaping (TNS) [6] is a standard tool in audio codecs, which models the instantaneous envelope of a signal. Since our method models the envelope, it can also be applied to time-domain signals as well.
[0166] Similarly, the bandwidth extension (BWE) method applies the spectral envelope to model the spectral shape of higher frequencies, and the proposed method can therefore also be applied to BWE<sub>O</sub>
[0167] FIG. 17 illustrates an apparatus for determining one or more encoding values used to encode an audio signal envelope according to an embodiment.
[0168] The apparatus includes an aggregator 1710, which is used to determine an aggregate value for each of a plurality of parameter values. The multiple parameter values are sorted so that when the first parameter value of the multiple parameter values is different from the second parameter value in the multiple parameter values, the first parameter value is before or after the second parameter value.
[0169] An envelope value may be assigned to each of the parameter values, where the envelope value of each of the parameter values depends on the audio signal envelope, and where an aggregator is used to set each parameter of a plurality of parameter values Value, the aggregate value is determined according to the envelope value of the parameter value and according to the envelope value of each of the multiple parameter values before the parameter value.
[0170] In addition, the device includes an encoding unit 1720, which is configured to determine one or more encoding values according to one or more of the aggregated values of the multiple parameter values. For example, the encoding unit 1720 may generate the aforementioned one or more split points as one or more encoded values, as described above.
[0171] FIG. 18 shows an aggregate function 1810 according to the first example.
[0172] Among other things, FIG. 18 shows 16 envelope points of the audio signal envelope. For example, reference numeral 1824 indicates the 4th envelope point of the audio signal envelope, and reference numeral 1828 indicates the 8th envelope point. Each envelope point includes parameter values and envelope values. In other words, in the xy coordinate system, the parameter value can be regarded as the X component of the envelope point, and the envelope value can be regarded as the y component of the envelope point. Therefore, as can be seen from FIG. 18, the parameter value of the fourth envelope point 1824 is 4, and the envelope value of the fourth envelope point is 3. As another example, the parameter value of the eighth envelope point 1828 is 8, and the envelope value of the fourth envelope point is 2. In other embodiments, if considering, for example, the spectral envelope, the parameter value will not indicate the index number as in FIG. 18, but may, for example, indicate the center frequency of the spectral band, so that, for example, the first parameter value may be 300 Hz, The second parameter value can be 500 Hz or the like. Or, for example, in other embodiments, the parameter value may indicate a point in time if one considers, for example, an instantaneous envelope.
[0173] The aggregation function 1810 includes a plurality of aggregation points. For example, consider the fourth aggregation point 1814 and the eighth aggregation point 1818. Each aggregation point includes parameter values and aggregation values. Similarly to the above, in the xy coordinate system, the parameter value can be regarded as the X component of the aggregation point, and the aggregation value can be regarded as the y component of the aggregation point. In FIG. 18, the parameter value of the fourth aggregation point 1814 is 4, and the aggregation value of the fourth aggregation point 1818 is 7. As another example, the parameter value of the 8th envelope point is 8, and the envelope value of the 4th envelope point is 13.
[0174] The aggregation value of each aggregation point of the aggregation function 1810 depends on the envelope value of the envelope point having the same parameter value as the considered aggregation point, and further depends on the value of the multiple parameter values preceding the parameter value. The envelope value of each. In the example in Figure 18, regarding the fourth aggregation point 1814, its aggregation value depends on the envelope value of the fourth envelope point 1824 (because this envelope point has the same parameter values as the aggregation point), and further depends on The envelope values at the envelope points 1821, 1822, and 1823 (because the parameter values of these envelope points 1821, 1822, and 1823 are before the parameter values of the envelope point 1824).
[0175] In the example of FIG. 18, the aggregate value of each aggregation point is determined by summing the envelope value of the corresponding envelope point and the envelope value of the envelope point before it. Therefore, the aggregation value of the fourth aggregation point is 1+2+1+3 = 7 (because the envelope value of the first envelope point is 1, the envelope value of the second envelope point is 2, and the third The envelope value of each envelope point is 1, and the envelope value of the fourth envelope point is 3). Correspondingly, the aggregation value of the 8th aggregation point is 1+2+1+3+1+2+1+2=13.
[0176] The aggregate function increases monotonically. This means that each aggregation point (with the preceding item) of the aggregation function has an aggregation value greater than or equal to the aggregation value of the aggregation point immediately before it. For example, regarding aggregation function 1810, for example, the aggregation value of the fourth aggregation point 1814 is greater than or equal to the aggregation value of the third aggregation point, and the aggregation value of the eighth aggregation point 1818 is greater than or equal to the aggregation value of the seventh aggregation point 1817. Value, and so on, and this applies to all aggregation points of the aggregation function.
[0177] FIG. 19 shows another example of the aggregation function, here, the aggregation function 1910. In the example of FIG. 19, the sum of the square of the envelope value of the corresponding envelope point is calculated by The square of the envelope value is summed to determine the aggregation value of each aggregation point. Therefore, for example, in order to obtain the aggregation value of the fourth aggregation point 1914, the square of the envelope value of the corresponding envelope point 1924 and the square of the envelope values of the preceding envelope points 192M922 and 1923 are summed, Get 2<sup>2</sup>+1<sup>2</sup>+2<sup>2</sup>+1<sup>2</sup>= 10ο Therefore, the aggregation value of the fourth aggregation point 1914 in FIG. 19 is 10. In FIG. 19, reference numerals 1931, 1933, 1935, and 1936 indicate the squares of the envelope values of the respective envelope points, respectively.
[0178] It can also be seen from FIGS. 18 and 19 that the aggregation function provides an effective way to determine the split point. The split point is an example of a coded value. In Figure 18, the maximum aggregation value of all split points (this can be, for example, the total energy) is 20.
[0179] For example, if only one split point should be determined, the parameter value of the aggregation point may, for example, be selected as a split point equal to or close to 10 (50% of 20). In Figure 18, this parameter value will be 6, and the single split point will be 6.
[0180] If three split points should be determined, the parameter values of the aggregation points can be selected as split points equal to or close to 5, 10, and 15 (25%, 50%, and 75% of 20), respectively. In Figure 18, these parameter values will be 3 or 4, 6 and 11. Therefore, the selected split points will be 3, 6, and 11, or 4, 6 and IL. In other embodiments, non-integer values may be allowed as split points, then, in Figure 18, the determined split point will be Yes, such as 3. 33, 6 and llo
[0181] Therefore, according to some embodiments, the aggregator may, for example, be used for each parameter value in a plurality of parameter values, through the envelope value of the parameter value and the parameter value before the parameter value. The values are added together to determine the aggregate value.
[0182] In an embodiment, the envelope value of each of the parameter values may, for example, indicate the energy value of the audio signal envelope with the audio signal envelope as the signal envelope.
[0183] According to an embodiment, the envelope value of each of the parameter values may, for example, indicate n times the spectrum value of the audio signal envelope with the audio signal envelope as the signal envelope, where n is greater than 0. The even number.
[0184] In an embodiment, the envelope value of each of the parameter values may, for example, indicate n of the amplitude of the audio signal envelope expressed in the time domain and taking the audio signal envelope as the signal envelope. The second scene, where η is an even number greater than 0.
[0185] According to an embodiment, the coding unit may, for example, be used to determine the number of coding values as one or more coding values according to one or more of the aggregate values of the parameter values and depending on how many values are indicated by the coding unit. , To determine one or more encoding values.
[0186] In an embodiment, the coding unit may, for example, be used to determine one or more coding values according to cU)=min/α(/)-jinba*;
[0187] where c(k) indicates the k-th encoding value to be determined by the encoding unit, where j indicates the j-th parameter value among a plurality of parameter values, and a(j) indicates the value assigned to the j-th parameter value The aggregate value of where max(a) indicates as being assigned
The maximum value of one of the aggregate values for one of the parameter values, where none of the aggregate values assigned to one of the parameter values is greater than the maximum value, and
[0188] Among them,
<img file="CN105431902A_D0002.tif" />
Indicates the minimum value as one of the parameter values, for which reason is the minimum value.
[0189] FIG. 16 illustrates an apparatus for generating an audio signal envelope from one or more encoded values according to an embodiment.
[0190] The device includes: an input interface 1610 for receiving one or more coded values; and an envelope generator 1620 for generating an audio signal envelope according to the one or more coded values.
[0191] The envelope generator 1620 is configured to generate an aggregation function according to one or more coded values, where the aggregation function includes a plurality of aggregation points, wherein each of the aggregation points includes a parameter value and an aggregation value, and the aggregation function increases monotonically.
[0192] Each of the one or more coded values indicates at least one of a parameter value and an aggregation value of one of the aggregation points of the aggregation function. This means that the parameter value of one of the specified aggregation points or the aggregation value of one of the specified aggregation points or the parameter value and aggregation value of one of the aggregation points of the specified aggregation function in each of the coded values. In other words, each of the one or more coded values indicates a parameter value and/or aggregation value of one of the aggregation points of the aggregation function.
[0193] In addition, the envelope generator 1620 is used to generate an audio signal envelope such that the audio signal envelope includes a plurality of envelope points, wherein each of the envelope points includes a parameter value and an envelope value, and wherein For each of the aggregation points of the aggregation function, one of the envelope points of the audio signal envelope is assigned to the aggregation point so that the parameter value of the envelope point is equal to the parameter value of the aggregation point. In addition, the envelope generator 1620 is used to generate the audio signal envelope so that the envelope value of each of the envelope points of the audio signal envelope depends on the aggregate value of at least one aggregation point of the aggregation function.
[0194] According to an embodiment, the envelope generator 1620 may, for example, be used for determining one of the aggregation points for each of one or more coded values according to the coded value and by determining one of the aggregation points according to the one or more coded values Interpolation is applied to each of the aggregation points to obtain an aggregation function to determine the aggregation function.
[0195] According to an embodiment, the input interface 1610 may be used to receive one or more split values as one or more encoded values. The envelope generator 1620 may be used to generate an aggregation function according to one or more split values, where each of the one or more split values indicates an aggregation value of one of the aggregation points of the aggregation function. In addition, the envelope generator 1620 may be used to generate a reconstructed audio signal envelope such that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts. The predefined allocation rule is for each signal envelope part of two or more signal envelope parts, and the signal envelope part value is defined according to the signal envelope part. In addition, the envelope generator 1620 can be used to generate the reconstructed audio signal envelope so that for each of the two or more signal envelope parts, the absolute value of the signal envelope part value is greater than that of the other signal envelope parts. Half of the absolute value of the signal envelope part value of each of the parts.
[0196] In an embodiment, the envelope generator 1620 may, for example, be used to determine the first derivative of the aggregation function at a plurality of aggregation points of the aggregation function.
[0197] According to an embodiment, the envelope generator 1620 may, for example, be used to generate an aggregate function according to the coded value, so that the aggregate function has a continuous first derivative.
[0198] In other embodiments, the LPC model can be obtained from the quantized spectral envelope. By taking the power spectrum abs(x)<sup>2 </sup>The inverse Fourier transform of to obtain the autocorrelation. From this autocorrelation, the LPC model can be easily calculated by traditional methods. This LPC model can then be used to create a smooth envelope.
[0199] According to some embodiments, the block may be modeled by using spline interpolation or other interpolation methods to obtain a smooth envelope. Interpolation is most conveniently done by modeling the cumulative sum of spectral quality.
[0200] FIG. 7 shows the same spectra as FIG. 6, but with their cumulative mass. Line 710 indicates the cumulative quality line of the original signal envelope. Point 721 in (a), 751, 752, 753 in (b), and 781, 782, 783, and 784 in (c) indicate where the split point should be.
[0201] In (a), the step size between points 738, 721, and 729 on the y-axis is constant. Similarly, in , the step size between points 768, 751, 752, 753 and 759 on the y-axis is constant. Similarly, in (c), the step size between points 798, 781, 782, 783, 784, and 789 on the y-axis is constant. The dotted line between points 729 and 739 indicates the total value.
[0202] In (a), the point 721 indicates the position of the split point 731 on the x-axis. In (b), points 751, 752, and 753 indicate the positions of split points 761, 762, and 763 on the x-axis, respectively. Similarly, in (c), points 781, 782, 783, and 784 indicate the positions of split points 791, 792, 793, and 794 on the x-axis, respectively. The dotted lines between points 729 and 739, points 759 and 769, and points 789 and 799 indicate the total value, respectively.
[0203] It should be noted that points 721 indicating the location of split points 731; 761, 762, 763; 791, 792, 793, and 794; 751, 752, 753; 781, 782, 783, and 784 are always in the original The cumulative quality line of the signal envelope, and the step size on the y-axis is constant.
[0204] In this domain, the cumulative spectral quality can be interpolated by any traditional interpolation algorithm.
[0205] In order to obtain a continuous representation in the original domain, the cumulative domain must have a continuous first derivative. For example, interpolation can be done using a spline function, so that for the k-th block, the end points of the spline function are kE/N and (k+1)E/N, where E is the total mass of the spectrum. In addition, you can specify the derivative of the spline function at the end point to obtain a continuous envelope in the original domain.
[0206] One possibility is to specify the derivative (t set t) for the split point k as:
[0207] (A) = _-_-_-τ jμ+1)-/(/ -1)
[0208] where c(k) is the accumulated energy at the aggregation point k, and f(k) is the frequency of the aggregation point k.
[0209] More generally, the points k1, k, and k+1 can be any type of coded value.
[0210] According to an embodiment, the envelope generator 1620 is used to determine the audio signal envelope by determining the ratio of the first difference and the second difference. The first difference is the first aggregation value (c(k+l)) of the first aggregation point in the aggregation points of the aggregation function and the second aggregation value of the second aggregation point in the aggregation points of the aggregation function (c( kl) or c(k)). The second difference is the first parameter value (f(k+l)) of the first aggregation point in the aggregation point of the aggregation function and the second parameter value of the second aggregation point in the aggregation point of the aggregation function ( The difference between f(kl) or f(k)).
[0211] In a specific embodiment, the envelope generator 1620 is used to determine the envelope of the audio signal by applying /"barrier)ng#Yu Er Xi Er%;
[0212] where t and t(k) indicate the derivative of the aggregate function at the k-th coded value, where c(k+1) is the first aggregate value, and f(k+1) is the first parameter value , Where c(kl) is the second aggregate value, where f(k-1) is the second parameter value, where k is an integer representing the index of one or more coded values, where c(k+ l)-c(kl) is the first difference between the two aggregate values c(k+l) and c(kl), and where f(k+l)-f(kl) is the two parameter values f(k +l) and the second difference of f (kT).
[0213] For example, c(k+1) is the first aggregate value assigned to the k+1th coded value. f(k+l) is assigned to the
The first parameter value of k+1 coded values. c(kl) is the second aggregate value assigned to the k-1th coded value. f (k-1) is the second parameter value assigned to the k-1th coded value.
[0214] In another embodiment, the envelope generator 1620 is used to determine the audio signal envelope by applying «W)=0.5·ί <-1)-heart)-<%-[)] ;
[0215] where t set t(k) indicates the derivative of the aggregate function at the k-th coded value, where c(k+1) is the first aggregate value, and f(k+1) is the first parameter value , Where c(k) is the second aggregation value, where f(k) is the second parameter value, where c(kl) is the third aggregation value of the third aggregation point in the aggregation function, where f (kl) is the third parameter value of the third aggregation point in the aggregation point of the aggregation function, where k is an integer representing the index of one of the one or more coded values, where c(k+1)-c( k) is the first difference between the two aggregate values c(k+1) and c(k), and where f (k+1)-f (k) are the two parameter values f (k+1) and f The second difference of (k).
[0216] For example, c(k+1) is the first aggregate value assigned to the k+1th coded value. f(k+1) is the first parameter value assigned to the k+1th coded value. c(k) is the second aggregate value assigned to the k-th coded value. f(k) is the second parameter value assigned to the k-th coded value. c(kl) is the third aggregate value assigned to the k-1th coded value. f (k-1) is the third parameter value assigned to the k-1th coded value.
[0217] By specifying that the aggregate value is assigned to the k-th coded value, this means that, for example, the k-th coded value indicates the aggregated value, and/or the k-th coded value indicates the parameter of the aggregated point to which the aggregated value belongs value.
[0218] By specifying that the parameter value is assigned to the k-th coded value, this means that, for example, the k-th coded value indicates the parameter value, and/or the k-th coded value indicates the aggregation of the aggregation point to which the parameter value belongs value.
[0219] In a specific embodiment, for example, the coded values k1, k, and k+1 are split points as described above.
[0220] For example, in an embodiment, the signal envelope reconstructor 110 of FIG. 1 may, for example, be used to generate an aggregation function according to one or more split points, where the aggregation function includes a plurality of aggregation points, and the aggregation points are Each of includes a parameter value and an aggregate value, where the aggregate function increases monotonically, and each of the one or more split points represents at least one of the parameter value and the aggregate value of one of the aggregate points of the aggregate function.
[0221] In this embodiment, the signal envelope reconstructor 110 may, for example, be used to generate an audio signal envelope such that the audio signal envelope includes a plurality of envelope points, wherein each of the envelope points includes a parameter Value and an envelope value, and where the envelope point of the audio signal envelope is assigned to each of the aggregation points of the aggregation function, so that the parameter value of the envelope point is equal to the parameter value of the aggregation point.
[0222] Furthermore, in this embodiment, the signal envelope reconstructor 110 may, for example, be used to generate an audio signal envelope such that the envelope value of each of the envelope points of the audio signal envelope depends on the aggregation The aggregate value of at least one aggregate point of the function.
[0223] In a specific embodiment, the signal envelope reconstructor 110 may, for example, be used to determine the audio signal envelope by determining the ratio of the first difference and the second difference, where the first difference is the aggregation function The first aggregation value of the first aggregation point in the aggregation points (c(k+l)) and the second aggregation value of the second aggregation point in the aggregation function of the aggregation point (c(kl); c(k)) The second difference is the first parameter value (f(k+1)) of the first aggregation point in the aggregation point of the aggregation function and the second aggregation point in the aggregation point of the aggregation function The difference between the second parameter values (f(kl); f(k)). To this end, the signal envelope reconstructor 110 may be used to implement one of the concepts described above as explained for the envelope generator 1620. [0224] The left and rightmost sides cannot use the above equations for derivatives because c(k) and f(k) are not available outside their defined range. Then, these c(k) and f(k) outside the range of k can be replaced by the value at the end point to
Make
<td>[0225]</td><td>Λ1 )-./(0)</td>
<td>[0226]</td><td>as well as</td>
<td>[0227]</td><td>////(.¥ 1)=/(A<sup>r</sup>- 1)-./(;¥- 2)</td>
<td>[0228]</td><td>Since there are four constraints (cumulative mass and derivative at the two end points), the corresponding spline function can be selected</td>
It is a fourth-order polynomial.
<td>[0229][0230]</td><td>Figure 8 shows an example of the interpolated spectral quality envelope in the (a) original and (6) cumulative quality domains. In (a), the original signal envelope is indicated by 810, and the interpolated spectrum quality envelope is indicated by 820. The split points are respectively determined by</td>
831, 832, 833 and 834 instructions. 838 indicates the beginning of the signal envelope, and 839 indicates the end of the signal envelope.
<td>[0231]</td><td>In (b), 840 indicates the accumulated original signal envelope, and 850 indicates the accumulated spectral quality envelope. Split</td>
The points are indicated by 861, 862, 863 and 864 respectively. The locations of the split points are indicated by points 851, 852, 853, and 854 on the accumulated original signal envelope 840, respectively. On the x-axis, 868 indicates the beginning of the original signal envelope, and 869 indicates the end of the original signal envelope. The line between 869 and 859 indicates the total value.
<td>[0232]</td><td>The embodiment provides a concept for encoding the frequency of the separated block. Frequency represents an ordered list of scalar fk,</td>
That is, f<sub>k</sub><f<sub>k+lo</sub>If there are K+1 blocks, there are K split points.
(N\
[0233] Further, if there are N quantization levels, there are possible quantizations. For example, for 32 quantities)
With quantization levels and 5 split points, there are 201376 possible quantizations that can be encoded with 18 bits.
[0234] It should be observed that the transient steering decorrelator (TSD) tool in MPEG USAC [5] has the similar problem of encoding K positions in the range of 0 to N-1, thereby the same or Similar enumeration techniques can be used to encode the frequency of the current problem. The advantage of this coding algorithm is that it has constant bit consumption.
[0235] Optionally, in order to further improve the accuracy or reduce the bit rate, a traditional vector quantization technique, such as a technique for LSF quantization, may be used. Using this method, a higher quantization level can be obtained, and the quantization of the average distortion can be optimized. The disadvantage is that, for example, the codebook needs to be stored. On the contrary, the TSD method uses algebraic enumeration of clusters.
[0236] In the following, an algorithm according to the embodiment is described.
[0237] First, consider a general application scenario.
[0238] In particular, the practical application of the proposed distribution quantization method for encoding the spectral envelope is described below in an SBR-like scenario.
[0239] According to some embodiments, the encoder is used to:
[0240]-Calculate the spectral amplitude or energy value of the HF band from the original audio signal; and/or
[0241]-Calculate a predefined (or arbitrary, transmitted) number of K subband indexes that split the spectral envelope into K+1 equal quality blocks; and/or
[0242]-Use the same algorithm as in TSD [5] to encode the index; and/or
[0243] Quantize and encode the total quality of the HF band (eg via Huffman), and write the total quality and index into the bitstream.
[0244] According to some embodiments, the decoder is used to:
[0245]-read the total quality and index from the bitstream, and then decode; and/or
[0246] Approximately estimate the smooth cumulative quality curve by spline interpolation; and/or
[0247] Solve the first derivative of the cumulative mass curve to reconstruct the spectral envelope.
[0248] Some embodiments include other optional additions:
[0249] For example, some embodiments provide warping capabilities: reducing the number of possible quantization levels results in a reduction in the bits required to encode the split point, and additionally reduces computational complexity. For example, before applying distribution quantization, this effect can be developed by warping the spectral envelope with the aid of psychoacoustic features or simply by summing adjacent frequency bands in the encoder. On the decoder side, after reconstructing the spectral envelope from the split point index and the total mass, the envelope must be buckled by inverse features.
[0250] Some other embodiments provide adaptive envelope transformation: as mentioned before, there is no need to adjust the energy of the spectral envelope (ie, the abs(x) of the signal X).<sup>2</sup>) Application of distribution quantization, but every other representation (positive, real value) (such as abs(x), sqrt(abs(x)), etc.) can be realized. In order to be able to develop fitting features of different shapes represented by various envelopes, it is reasonable to use adaptive transform technology. Therefore, before applying the distribution quantization, the detection of the best matching transform (fixed, predefined set) for the current envelope is performed as a preprocessing step. The transformation used must be transmitted and transmitted through the bit stream to enable correct re-transformation on the decoder side.
[0251] A further embodiment is used to support the adaptive number of blocks. In order to obtain higher flexibility of the proposed model, it is advantageous to be able to convert between different numbers of blocks for each spectral envelope. The number of currently selected blocks can be any one of the predefined set to minimize the bits that need to be explicitly transmitted or transmitted to support higher flexibility. On the one hand, this reduces the overall bit rate. As for a stable envelope shape, high adaptability is not required. On the other hand, a smaller number of blocks results in a larger block quality, thereby supporting a more accurate fit of a strong single peak with a steep slope.
[0252] Some embodiments are used to provide envelope stabilization. Since the proposed distribution quantization model has higher flexibility than methods such as scale factor band-based methods, fluctuations between time-adjacent envelopes can lead to undesirable instability. In order to counteract this effect, a signal adaptive envelope stabilization technique is applied as a post-processing step: for a stable signal part that is expected to have only a small amount of fluctuation, the envelope is stabilized by smoothing adjacent envelope values in time. For signal parts that naturally include strong time changes (eg, transient or rational/starting/offset caused by friction), no or only weak smoothing is applied.
[0253] Hereinafter, an algorithm for realizing envelope distribution quantization and encoding according to an embodiment is described.
[0254] In an SBR-like scenario, the actual implementation of the proposed distributed quantization method for encoding the spectral envelope is described. The following description of the algorithm refers to steps on the encoder and decoder sides that can be executed to process a specific envelope.
[0255] Below, the corresponding encoder is described.
[0256] For example, envelope determination and preprocessing can be performed as follows:
[0257]-Determine the spectral energy target envelope curve (eg, represented by 20 subband samples) and its corresponding total energy; [0258]-By averaging the subband values in pairs, apply spectral warping to reduce the total number of values (For example, average the first 8 subband values, and therefore reduce the total from 20 to 16);
[0259]-Apply envelope amplitude transformation to make a better match between envelope model performance and perceptual quality standards (eg, extract the fourth root of each subband value, and then = rice).
[0260] For example, distributed quantization and coding can be performed as follows:
[0261] Multiple determinations of subband indexes that split the envelope into a predefined number of equal mass blocks (eg, repeat the determination 4 times to split the envelope into 3, 4, 6 and 8 blocks);
[0262]-Complete reconstruction of the envelope of the distribution quantification ("comprehensive analysis" method, see below);
[0263] Determine and determine the number of blocks that lead to the most accurate description of the envelope (eg, by comparing the cross-correlation of the quantified envelope with the original envelope);
[0264]-By comparing the original and distributed quantized envelope and according to the change of the total energy, the volume is corrected;
[0265]-Use the same algorithm as in the TSDX tool (see ) to encode the split index;
[0266]-Transmit the number of blocks used for distributed quantization (eg, 4 blocks of a predefined number, transmitted by 2 bits);
[0267]-Quantize and encode the total energy (eg, using Huffman coding).
[0268] Now, the corresponding decoder is described.
[0269] For example, decoding and inverse quantization can be performed as follows:
[0270]-Decode the number of blocks used for distributed quantization and decode the total energy;
[0271]-Use the same algorithm as in the TSDX tool (see [5]) to decode the split index;
[0272] Approximately estimate the smooth cumulative quality curve through spline interpolation;
[0273]-Reconstruction of the spectral envelope from the cumulative domain by the first derivative (eg by taking the difference of consecutive samples).
[0274] For example, post-processing can be performed as follows:
[0275]-Apply envelope stabilization to offset subsequent fluctuations between envelopes caused by quantization errors (e.g., through the temporal smoothing of the reconstructed subband values, "to = Bu & Ren Ru, k + for inclusion The frame of the transient signal part α = 0.1, otherwise α = 0.25);
[0276] Restore the envelope transform according to the application in the encoder;
[0277]-Recover envelope warpage according to the application in the encoder.
[0278] In the following, effective encoding and decoding of split points are described. The split point encoder 225 of FIGS. 4 and 5 may, for example, be used to implement effective encoding as described below. The split point decoder 105 of FIG. 2 may, for example, be used to implement effective decoding as described below.
[0279] In the embodiment shown in FIG. 2, the apparatus for decoding further includes a split point decoder 105, which is used to decode one or more code points according to a decoding rule to obtain one or Multiple split points. The split point decoder 105 is used to analyze the total number of positions indicating the total number of possible split point positions, the number of split points indicating the number of split points, and the number of split point states. In addition, the split point decoder 105 is used to generate an indication of one or more positions of the split point using the total number of positions, the number of split points, and the number of split point states. In a specific embodiment, the split point decoder 105 may, for example, be used to generate an indication of two or more positions of the split point using the total number of positions, the number of split points, and the number of split point states.
[0280] In the embodiment shown in FIG. 4 and FIG. 5, the device further includes a split point encoder 225, which is used to encode the position of each of the one or more split points to Obtain one or more code points. The split point encoder 225 is used to encode the position of each of the one or more split points by encoding the number of split point states. In addition, the split point encoder 225 is used to provide the total number of positions indicating the total number of possible split point positions and the number of split points indicating the number of one or more split points. The number of split point states, the number of total positions, and the number of split points together indicate the position of each of the one or more split points.
[0281] FIG. 15 is an apparatus for reconstructing audio signals according to an embodiment. The device includes: a device 1510 for decoding according to one of the above-mentioned embodiments or according to the following embodiments to obtain a reconstructed audio signal envelope of an audio signal; Other signal characteristics of the audio signal are generated by the signal generator 1520 of the audio signal, and other signal characteristics are different from the audio signal envelope. As outlined above, those skilled in the art
Realize that the audio signal itself can be reconstructed from the signal envelope of the audio signal and from other signal characteristics of the audio signal. For example, the signal envelope may, for example, indicate the energy of the samples of the audio signal. Other signal characteristics may, for example, indicate whether each sample of the time domain audio signal has a positive value or a negative value.
[0282] Some specific embodiments are based on: in the decoding device of the present invention, the total number of positions indicating the total number of possible split point positions and the number of split points indicating the total number of split points can be obtained. For example, the encoder may transmit the total number of positions and/or the number of split points to the device for decoding.
[0283] Based on these assumptions, some embodiments implement the following concepts:
[0284] Let N be the (total) number of possible split point locations, and
[0285] Let P be the (total) number of split points.
[0286] Assume that both the device for encoding and the device for decoding know the values of N and P.
[0287] Knowing N and P, it can be inferred that there are only different combinations of possible split point positions.
[0288] For example, if the possible split point positions are numbered from 0 to N-1, and if P = 8, then the first possible combination of split point positions and events will be (0, 1, 2, 3, 4,5,6,7), the second possible combination will be (0, 1, 2, 3, 4, 5, 6, 8), (N, and so on, until the combination (N-& N-7 , N-6, N-5, N-4, N-3, N-2, N-1), so there are a total of different groups.
[0289] A further discovery is applied: the number of split point states can be encoded by the device for encoding, and the number of split point states is transmitted to the decoder. If each of the possible combinations is represented by a unique split point state number, and if the device used for decoding knows which split point state number represents which combination of split point positions, the device used for decoding can use N, P and the state number of the split point decode the position of the split point. For a large number of typical values of N and P, compared with other concepts, this coding technique uses fewer bits to encode the position of the split point of the event.
[0290] In other words, by encoding the discrete number P of the position Pk in the range of [0···Ν-1], the problem of encoding the position of the split point can be solved to use as few bits as possible, So that for k work h, the position will not overlap Pk work P h. Since the order of the positions has no effect, it is concluded that the number of unique combinations of positions is the binomial coefficient
<td>Ν'</td><td></td><td>( /</td><td>Υτνγη</td>
<td></td><td>The number of bits required is therefore: hits = ceil</td><td>log?</td><td rowspan="2"></td>
<td>Ρ </td><td>ο:</td><td>\ I</td>
[0291] Some embodiments apply a position-by-position decoding concept. The decoding concept of position by position. The idea is based on the following findings:
[0292] Assume that N is the (total) number of possible split point positions, and P is the number of split points (this means that N can be the total number of positions FSN, and P can be the number of split points ESON) ο Consider the first possible split Point location. Two situations can be distinguished:
[0293] If the first possible split point position is a position that does not include the split point, then with regard to the remaining N-1 possible split point positions, there are only P for the remaining N-1 possible split point positions. Different possible combinations of split points.
[0294] However, if the possible split point positions are positions that include the split point, then, regarding the remaining N-1 possible split point positions, there are only the remaining PT possible for the remaining N-1 split points The location of the split point J = [pj-[p J different possible combinations.
[0295] Based on this discovery, the embodiment is further based on the discovery that all combinations of the first possible split point positions where there is no split point should be coded through the number of split point states less than or equal to the threshold. In addition, all combinations of the first possible split point positions where the split point is not located here should be coded by the number of split point states greater than the threshold. In an embodiment, the number of states of all split points may be a positive integer or 0, and the appropriate threshold for the position of the first possible split point may be ρν-η ρ.
\ <sup>1</sup> J
[0296] In an embodiment, it is determined by testing whether the first possible split point position of the frame includes a split point, and whether the number of split point states is greater than a threshold (optionally, by testing whether the number of split point states is greater than or equal to, less than or If it is equal to or less than the threshold value, the encoding/decoding process of the embodiment can also be implemented).
[0297] After analyzing the position of the first possible split point, use the adjusted value to continue to decode the position of the second possible split point. In addition to adjusting the number of split point positions considered (minus 1), the number of split points is also reduced by 1 and the number of split point states is adjusted. In the case where the number of split point states is greater than the threshold, the part related to the position of the first possible split point is deleted from the number of split point states. The decoding process can be continued for other possible split point positions in a similar manner.
[0298] In an embodiment, the discrete number P of the position Pk in the range of [0···N-1] is coded so that the position does not overlap PkH Ph for k×h. Here, each unique combination of positions on a given range is called a state, and each possible position within this range is called a possible split point position (pspp). According to an embodiment of the apparatus for decoding, the first possible split point position within the range is considered. If the possible split point position does not have a split point, then the range can be reduced to N-1, and the number of possible states can be reduced to. Conversely, if the state is greater than]ρ;, it can be concluded that the first possible split Point location, there is a split point. The following decoding algorithm can be derived from this:
[0299] For each pspp h
[0300] If state> [ <sub>p</sub> J
[0301] Assign the split point to pspp h (NhH
[0302] Update the remaining state st ate; = state-I <sub>p</sub> |
[0303] Decrease the number of remaining positions P: = P-1
[0304] End
[0305] End
[0306] The calculation of the binomial coefficients on each iteration is expensive. Therefore, according to an embodiment, the following rules can be used to update the binomial coefficients with the values obtained from the previous iteration.
[0307]
<td></td><td>pv-1)</td><td>Ν </td><td></td><td></td>
<td></td><td></td><td>- with</td><td></td><td>—</td>
<td></td><td>I p</td><td>Ν - Ρ</td><td></td><td></td>
P -1 r
[0308] Using these formulas, the cost of each update of the binomial coefficients is only one multiplication operation and one division operation.
On the contrary, the clearly estimated cost for each iteration is P times of multiplication and division.
[0309] In this embodiment, to initialize the binomial coefficients, the total complexity of the decoder is P multiplication operations and division operations. For each iteration, there is one multiplication operation, division operation, and if statement. Code position, there is a multiplication operation, addition operation and division operation. Note that it is theoretically possible to reduce the number of divisions required for initialization to 1. However, in practice, this method will result in very large integers that are difficult to handle. In the worst case, the complexity of the decoder is N+2P division operations and N+2P multiplication operations, P addition operations (if MAC- operations are used, they can be ignored), and N if statements.
[0310] In an embodiment, the encoding algorithm adopted by the apparatus for encoding does not need to iterate over all possible split point positions, but only those that have positions assigned to them. therefore,
[0311] For each position p<sub>h</sub>,h = 1...P
[0312] Update sialc state :=+ [Blade; 1
[0313] In the worst case, the complexity of the encoder is P·(P-1) times of multiplication operations and P·(P-1) times of division operations and PT times of addition operations.
[0314] FIG. 9 illustrates a decoding process according to an embodiment of the present invention. In this embodiment, decoding is performed on a position-by-position basis.
[0315] In step 110, the value is initialized . The device for decoding stores the state number of the splitting point received as an input value, in the form of a variable S. In addition, the (total) number of split points indicated by the number of split points is stored in the form of a variable P. In addition, the total number of possible split point positions contained in the frame indicated by the total number of positions is stored in the form of a variable N.
[0316] In step 120, for all possible split point positions, initialize the value of spSepData[t] with 0. The bit array spSepData is the output data to be generated. It indicates for each possible split point location t, whether the possible split point location includes the split point (spSepData[t] = 1) or whether it does not include the split point (spSepData[t] = 0) ο In step 120, 0 Initialize the corresponding values of all possible split point positions.
[0317] In step 130, the variable k is initialized with the value N-1. In this embodiment, the numbers of the N possible split point positions are 0, 1, 2,..., N-lo set k=NT, which means that the possible split point position with the highest number is considered first.
[0318] In step 140, consider whether k20. If k<0, the decoding of the split point position has been completed and the process is terminated, otherwise, the process is continued with step 150.
[0319] In step 150, it is tested whether p>k<sub>o</sub>If ρ is greater than k, it means that all remaining possible split point positions include the split point. The process continues at step 230, where all the spSepData field values of the remaining possible split point positions l, -, k are set to 1, indicating that each of the remaining possible split point positions includes a split point. In this case, the process then terminates. However, if step 150 finds that p is not greater than k, then in step 160 the decoding process continues.
[0320] In step 160, the calculated value c=(:). c is used as the threshold.
[0321] In step 170, it is tested whether the actual value of the split point state number s is greater than or equal to c, where c is the threshold just calculated in step 160.
[0322] If s is less than c, this means that the considered possible split point location (with split point k) does not include the split point. In this case, no further action is needed, because for this possible split point location, spSepData[k] has been set to 0 in step 140. Then the process continues with step 220. In step 220, k is set to k: = k-1,
And consider the location of the next possible split point.
[0323] However, if the test in step 170 shows that s is greater than or equal to c, this means that the possible split point position k considered includes the split point. In this case, update the split point state number s in step 180 and set it to the value s: =s-Co. In addition, set spSepData[k] to 1 in step 190 to indicate the possible split point position k includes the split point. In addition, in step 200, ρ is set to p1, indicating that the remaining possible split point positions to be checked now only include p_1 possible split point positions with split points.
[0324] In step 210, it is tested whether p is equal to zero. If ρ is equal to 0, the remaining possible split point positions do not include the split point, and the decoding process is completed.
[0325] Otherwise, at least one of the remaining possible split point positions includes an event, and the process continues in step 220. In step 220, the next possible split point position (k-1) continues the decoding process.
[0326] The decoding process of the embodiment shown in FIG. 9 generates an array spSepData as an output value, which indicates for each possible split point position k, whether the possible split point position includes a split point (spSepData[k] = 1) Or whether it is not included (spSepData [k] = 0).
[0327] FIG. 10 shows a pseudo code for encoding the split point position according to an embodiment.
[0328] FIG. 11 shows an encoding process for encoding a split point according to an embodiment. In this embodiment, encoding is performed on a position-by-position basis. The purpose of the encoding process according to the embodiment shown in FIG. 11 is to generate the number of split point states.
[0329] In step 310, the value is initialized. Initialize p_s with 0. By continuously updating the variable p_ s, the number of split point states is generated. When the encoding process is completed, p_s will carry the number of split point states. Step 310 also initializes the variable k by setting k to k:=number of splitting points-1.
[0330] In step 320, the variable "pos" is set to pos:=spPos[k], where spPos is an array containing the positions of the possible split point positions including the split point.
[0331] The split point positions in the array are stored in ascending order.
[0332] In step 330, a test is performed to test whether k^pos<sub>o</sub>If this situation is true, the process is terminated. Otherwise, the process continues in step 340.
[0333] In step 340, the calculated value g=
[0334] In step 350, the variable p_s is updated and set to p_s:=p_s+c.
[0335] In step 360, set k to k:=kl<sub>o</sub>
[0336] Then, in step 370, a test is performed to test whether k20. In this case, consider the next possible split point location kl. Otherwise, the process is terminated.
[0337] FIG. 12 depicts a pseudo code for encoding the split point position according to an embodiment of the present invention.
[0338] FIG. 13 shows a split point decoder 410 according to an embodiment.
[0339] The total number of positions FSN indicating the total number of possible split point positions, the number of split points ESON indicating the (total) number of split points, and the number of split point states ESTN are provided to the split point decoder 410. The split point decoder 410 includes a splitter 440. The splitter 440 is adapted to split the frame into a first partition including a first set of possible split point positions and a second partition including a second set of possible split point positions, and wherein for each partition, it is determined individually to include The possible split point location of the split point. Thus, by repeatedly splitting the partition into smaller partitions, the location of the split point can be determined.
[0340] The "partition-based" decoding of the split point decoder 410 of this embodiment is based on the following concept:
[0341] Partition-based decoding is based on this idea: the set of all possible split point positions is split into two partitions A and B, and each partition includes a set of possible split point positions, where partition A includes one possible split Point location, and where partition B includes one possible split point location, and make N<sub>a</sub>+N<sub>b</sub>= No The set of all possible split point positions can be arbitrarily split into two partitions, preferably so that partitions A and B have almost the same total number of possible split point positions (e.g., make Shan= or ratio=N<sub>b</sub>-1). By dividing the set of all possible split point positions into two partitions, the task of determining the actual split point position is also divided into two subtasks, namely, determining the actual split point position in frame partition A and determining the actual split point position in frame partition B. The location of the split point.
[0342] In this embodiment, it is again assumed that the split point decoder 105 knows the total number of possible split point positions, the total number of split points, and the number of split point states. In order to solve the two subtasks, the split point decoder 105 should also know the number of possible split point positions of each partition, the number of split points in each partition, and the number of split point states of each partition (the split point of each partition). The number of states at this moment can be referred to as the "number of split point states").
[0343] Because the split point decoder itself divides the set of all possible split points into two partitions, it knows that partition A includes one possible split point position and partition B includes N possible split point positions. Based on the following findings, determine the actual number of split points in each of the two partitions:
[0344] Because the set of all possible split point positions has been divided into two partitions, at this time each of the actual split point positions is located in either partition A or partition B. In addition, assuming that P is the number of split points of the partition, N is the total number of possible split point positions of the partition, and f(P, N) is a function that returns the number of different combinations of split point positions, then the number of possible split points The number of different combinations of the entire set of locations split (which has been divided into partition A and partition B) is: [0345]
<td>Number of split points in partition A</td><td>Number of split points in partition B</td><td>The number of different combinations in the entire set of split point positions using this configuration</td>
<td>0</td><td>P</td><td>MO,NJ-f(^N<sub>b</sub>)</td>
<td>1</td><td>P-1</td><td></td>
<td>:air</td><td>P-2</td><td><2,N<sub>a</sub>)-f(P-2,Nb)</td>
<td></td><td></td><td></td>
<td>P</td><td>0</td><td>llP.Na)-ίϊΟ-.Nb)</td>
[0346] Based on the above considerations, according to an embodiment, all combinations using the first configuration should be coded with the number of split point states less than the first threshold. In the first configuration, partition A has 0 split points, and partition B There are P split points in it. The number of split point states can be coded as a positive integer value or 0. Because there are only f(0,Nj·f(P,Nj) combinations in the first configuration, a suitable first threshold may be f(0,Na)·f(P,Nj.
[0347] All combinations using the second configuration should be coded with the number of split point states greater than or equal to the first threshold and less than or equal to the second threshold. In the second configuration, partition A has 1 split point, and partition A has 1 split point. B has P-1 split points. Because only f(l,N<sub>a</sub>) -f(Pl,N<sub>b</sub>) Combinations, a suitable second threshold can be f(0,N<sub>a </sub>)·F(P,Nj+f(l,Na)·f(Pl,Nb). Similarly, determine the number of split point states used for combinations of other configurations.
[0348] According to an embodiment, decoding is performed by separating the set of all possible split point positions into two partitions A and B. Then, it is tested whether the number of split point states is less than the first threshold. In a preferred embodiment, the first threshold may be f(0,Na)·f(P,N<sub>b</sub>)<sub>o</sub>
[0349] If the number of split point states is less than the first threshold, it can be deduced that partition A includes 0 split points, and partition B includes all P split points. Then pair the two partitions with each determined value representing the number of split points in the corresponding partition
To decode. In addition, the first split point state number is determined for the partition A, and the second split point state number is determined for the partition B, and the first split point state number and the second split point state number are respectively used as the new split point state number. In this document, the number of split point states of a partition can be referred to as the "number of split point sub-states".
[0350] However, if the number of split point states is greater than or equal to the first threshold, the number of split point states may be updated. In a preferred embodiment, the number of split point states can be updated by subtracting a certain value from the number of split point states (preferably, the first threshold, such as f(0, Nj · f(P, Nj)) can be updated. In the following In one step, it is tested whether the number of updated split point states is less than the second threshold. In a preferred embodiment, the second threshold may be f(l, N<sub>a</sub>) -f(Pl,N<sub>b</sub>)<sub>o</sub>If the number of split point states is less than the second threshold, it can be obtained that partition A has 1 split point and partition B has P-1 split points.
[0351] Then, the two partitions are decoded with the respectively determined number of split points of each partition. The first split point sub-state number is used for the decoding of partition A, and the second split point sub-state number is used for the decoding of partition B. However, if the number of split point states is greater than or equal to the second threshold, the number of split point states can be updated. In a preferred embodiment, a certain value can be subtracted from the number of split point states (preferably, f(l, N<sub>a</sub>) -f(Pl,N<sub>b</sub>)) to update the number of split point states. Similarly, it is possible to apply this decoding process to the distribution of the remaining split points with respect to the two partitions.
[0352] In an embodiment, the number of split point sub-states used for partition A and the number of split point sub-states used for partition B can be used for the decoding of partition A and the decoding of partition B, where the two events are determined by division Number of sub-states:
[0353] Number of split point states/f (number of split points in partition B, N<sub>b</sub>)。
[0354] Preferably, the number of split point sub-states of partition A is the integer part of the above division, and the number of split point sub-states of partition B is the remainder of this division. The number of split point states applied to this division may be the original number of split point states of the frame or the updated number of split point states, such as being updated by subtracting one or more thresholds, as described above.
[0355] To illustrate the above-mentioned concept of partition-based decoding, consider the case where the set of all possible split point positions has two split points. In addition, if f(P, N) is still a function that returns the number of different combinations of split point positions of the partition, where P is the number of split points of the frame partition, and N is the total number of split points of the partition. Bei U, for each of the possible distributions of locations, yield the following number of possible combinations:
[0356]
<td>Location in partition A</td><td>Location in partition B</td><td>The number of combinations in this configuration</td>
<td>0</td><td>2</td><td>f(0,Na)· f(2, N<sub>b</sub>)</td>
<td>1</td><td>1</td><td>f(l,N<sub>a</sub>) · F(l,N<sub>b</sub>)</td>
<td>2</td><td>0</td><td>f(2,Na)· f(0, N<sub>b</sub>)</td>
[0357] Therefore, it can be concluded that if the number of split point states of the encoding of the frame is less than f(0, Nj · f(2, Nj, the positions of the split points need to be distributed to 0 and 2. Otherwise, from the number of split point states Subtract f(0,Nj .f(2,Nj, and compare the result with f(l,Nj .f(l,Nj. If the result is small, the position distribution is 1 and 1. Otherwise, only the distribution remains 2 and 0, the position distribution is 2 and 0.
[0358] Hereinafter, a pseudo code is provided according to an embodiment, and the pseudo code is used to decode the position of the split point (here: "sp"). In this pseudo code, "sp_a" is the number of split points in the (hypothetical) partition A, and "sp_b" is the number of split points in the (hypothetical) partition B. In this pseudo code, the number of split point states (e.g., updated) can be referred to as "state". The number of split point states of partitions A and B is still jointly coded in the "state" variable. According to the joint coding scheme of the embodiment, the number of split point sub-states of A (referred to here as "state_a") is divided by state/f (sp_b, NJ's
Integer part, the number of split point states of B (here called state_bO is the remainder of the division. Therefore, the length of the two partitions (the total number of split points of the partition) and the number of encoding positions can be determined by the same method (The number of split points in the partition) to decode:
[0359] Function X = decodestate (state, sp, Ν)
[0360] 1. Split the vector into two partitions of length Na and Nb·
[0361] 2. for sp_a from 0 to sp
[0362] a. sp_b = sp-sp_a
[0363] b. If state<f (sp_a, Na) *f (sp_b, Nb) then
[0364] Jump out of the for-loop·
[0365] c. state: = state-f (sp_a, Na) *f (sp_b, Nb)
[0366] 3. The number of possible states for partition B is
[0367] no_states_b = f (sp_b, Nb)
[0368] 4. State_a and state_b of states and partitions A and B are the integer part and remainder of the division state/no_states_b, respectively.
[0369] 5. If Na>l, obtain the decoding vector of partition A recursively through xa = decodestate (state_a, sp_a, Na)
[0370] Otherwise (Na == 1), the vector χ& is a scalar
[0371] xa = state_a can be set.
[0372] 6. If Nb>l, obtain the decoding vector of partition B recursively through xb = decodestate (state_b, sp_b, Nb),
[0373] Otherwise (Nb == 1), the vector xb is a scalar
[0374] xb = state_b can be set.
[0375] 7. Combine xa and xb by using X=[xa xb] to obtain the final output X.
[0376] The output of this algorithm is a vector that is (1) at each coding position (ie, split point location) and (0) at other locations (ie, at possible split point locations that do not include the split point).
[0377] Hereinafter, a pseudo code is provided according to an embodiment, and the pseudo code is used to encode the position of the split point using a similar variable name in a manner similar to the above:
[0378] Function state = encodestate (χ, Ν)
[0379] 1. Split the vector into two partitions xa and xb of length Na and Nb.
[0380] 2. Count the split points in partitions A and B as sp_a and sp_b, and set sp = sp_a+sp_b<sub>o</sub>
[0381] 3. Set state to 0
[0382] 4. For k from 0 to sp_a~l
[0383] a. state: = state+f (k, Na)(sp-k, Nb)
[0384] 5. If Na>1, encode partition A through state"=encodestate (xa, Na);
[0385] No Bay! J (Na = = 1) <sub>?</sub>Set state_a = xa<sub>o</sub>
[0386] 6. If Nb>1, encode partition B through state_b = encodestate (xb, Nb);
[0387] Otherwise (Nb == 1), set state_b=xb.
[0388] 7. Joint Coding of States
[0389] state: = state+state_a^f (sp_b, Nb) +state_b.
<td>[0390]</td><td>Here, it is assumed that, similar to the decoding algorithm, each encoding position (ie, the split point</td>
Position), and all other elements are (0) (that is, possible split point positions that do not include the split point) ο
<td>[0391] 9 soil</td><td>Standard methods can be used to easily implement the above recursive formula expressed in pseudocode in a non-recursive form.</td>
<td>/mysteriousο[0392]</td><td>According to an embodiment, the function f(P,N) may be implemented as a lookup table. When the position does not overlap (such as the current</td>
In the text), the function f(p, N) of the number of states is a binomial function that can be simply calculated online, namely
<td>[0393]</td><td></td>
<td>[0394]</td><td>According to an embodiment of the present invention, both the encoder and the decoder have a for-loop. In the for-loop, k</td>
Calculate the product of f (pk, Na) * f (k, Nb). For efficient calculations, this can be written as:
[0395] <sup>ν</sup>· <...'Lp-k\pk-\ip-k-2)...\ Φ-1)(/:-2)...1 =Ν. (Ν"- 1)(Ν" - 2)...( Ν* - ο -on+1). Gong Tayi 2)···(Μ-on+1). p-( + l.
(p -k + 1)(factory-k\p-k -1)... 1 (k-1 )(.t-2)... 1 N" _ p -k + 1 k =J\<sub>P</sub> -/< <sub>+</sub> l, -1, N) Ji: Ten·
Λ\..-p-k+\ k
[0396] In other words, each iteration through three multiplication operations and one division operation can calculate the term for subtraction/addition operations (in steps 2b and 2c of the decoder and in step 4a of the encoder).
[0397] Returning to FIG. 1, alternative embodiments implement the apparatus for decoding to obtain the reconstructed audio signal envelope of FIG. 1 in a different manner. In this embodiment, as explained before, the device includes: a signal envelope reconstructor 110 for generating a reconstructed audio signal envelope according to one or more split points; and for outputting the reconstructed audio signal envelope The output interface 120.
[0398] In addition, the signal envelope reconstructor 110 is used to generate a reconstructed audio signal envelope such that one or more split points divide the reconstructed audio signal envelope into two or more audio signal envelope parts, wherein The predefined allocation rule is for each signal envelope part of two or more signal envelope parts, and the signal envelope part value is defined according to the signal envelope part.
[0399] In an alternative embodiment, however, a predefined envelope part value is assigned to each of two or more signal envelope parts.
[0400] In this embodiment, the signal envelope reconstructor 110 is used to generate the reconstructed audio signal envelope such that for each of the two or more signal envelope parts, the signal envelope The absolute value of the signal envelope part value of the signal envelope part is greater than 90% of the absolute value of the predefined envelope part value assigned to the signal envelope part, and makes the absolute value of the signal envelope part value of the signal envelope part The value is less than 110% of the absolute value of the predefined envelope part value assigned to the signal envelope part. This allows a certain deviation from the predefined envelope part value.
[0401] In a specific embodiment, however, the signal envelope reconstructor 110 is used to generate the reconstructed audio signal envelope such that the signal envelope part value of each of the two or more signal envelope parts is Equal to the predefined envelope part value assigned to the envelope part of the signal.
[0402] For example, it is possible to receive three split points that divide the audio signal envelope into four audio signal envelope parts. The allocation rule can specify that the pre-defined envelope value of the first signal envelope is 0.15, and the pre-defined value of the second signal envelope is 0.15.
The defined envelope part value is 0.25, the predefined envelope part value of the third signal envelope part is 0.25, and the predefined envelope part value of the fourth signal envelope part is 0.35.
[0403] When three split points are received, the signal envelope reconstructor 110 reconstructs the signal envelope according to the above-mentioned concept.
[0404] In another embodiment, a split point that divides the audio signal envelope into two audio signal envelope parts may be received. The allocation rule may specify that the predefined envelope part value of the first signal envelope part is P, and the predefined envelope part value of the second signal envelope part is l-ρ. For example, if P = 0.4, then 1-p = 0.6. In addition, when three split points are received, the signal envelope reconstructor 110 reconstructs the signal envelope according to the above-mentioned concept.
[0405] This alternative embodiment of applying predefined envelope part values can apply each of the above-mentioned concepts.
[0406] In an embodiment, the predefined envelope part values of at least two signal envelope parts are different from each other.
[0407] In another embodiment, the predefined envelope part value of each of the signal envelope parts is different from the predefined envelope part value of each of the other signal envelope parts.
[0408] Although some aspects have been described in the context of a device, it is obvious that these aspects also represent a description of a corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent descriptions of items or features of corresponding blocks or corresponding devices.
[0409] The decomposition signal of the present invention may be stored on a digital storage medium, or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (eg, the Internet).
[0410] According to certain implementation requirements, the embodiments of the present invention may be implemented in hardware or software. A digital storage medium with electronically readable control signals stored thereon, such as a floppy disk, DVD, CD, ROM, PROM, EPROM, EEPROM, or flash memory, can be used to execute the implementation, the electronically readable control signals and (or can be combined with) Programmable computer systems cooperate to execute various methods.
[0411] Some embodiments according to the present invention include a non-transitory data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0412] Generally, the embodiments of the present invention can be implemented as a computer program product having a program code, the program code being operable to perform one of the methods when the computer program product is executed on a computer. The program code can be stored on a machine-readable carrier, for example.
[0413] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine-readable carrier.
[0414] In other words, an embodiment of the method of the present invention is therefore a computer program with program code for executing one of the methods described herein when the computer program is executed on a computer.
[0415] A further embodiment of the present invention is therefore a data carrier (or a digital storage medium or a computer-readable medium), which includes a computer program recorded thereon for performing one of the methods described herein.
[0416] A further embodiment of the invention is therefore a data stream or signal sequence, which represents a computer program for performing one of the methods described herein. The data stream or signal sequence may for example be configured to be transmitted via a data communication connection (for example, via the Internet).
[0417] A further embodiment includes a processing device (eg, a computer or a programmable logic device) that is configured or adapted to perform one of the methods described herein.
[0418] A further embodiment includes a computer having a computer program installed thereon for performing one of the methods described herein.
[0419] In some embodiments, a programmable logic device (for example, a field programmable gate array) may be used to perform the
Some or all of the functions of the described method. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably executed by any hardware device.
[0420] The above-mentioned embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and changes in the configuration and details described herein are obvious to other skilled in the art. Therefore, it is only limited by the scope of the appended patent claims, and is not limited by the specific details presented in the description and explanation of the embodiments herein.
[0421] References
[0422] [l]Makhoul, John. Linear prediction: A tutorial review. <sup>/Z</sup>IEEE 63.4 Proceedings (1975): 561-580.
[0423] [2] Soong, Frank, and B. Juang. Line spectrum pair (LSP) and speech data compression. Acoustics, speech and signal processing, IEEE International Conference, ICASSP' 84.. Volume 9. IEEE, 1984.
[0424] [3] Pan, Davis. A tutorial on MPEG/Audio compression. Multimedia, IEEE
2. 2(1995) :60-74.
[0425] [4] M. Neuendorf, P. Gournay, M. Multrus, J. Lecomte, B. Bessette, R. Geiger, S.
Bayer, G. Fuchs, J. Hilpert, N. Rettelbach, R. Salami, G. Schuller, R. Lefebvre, B. Grill. Unified speech and audio coding scheme for high quality at low bitrates. Acoustics, speech and signal processing, 2009. ICASSP 2009. IEEE International Conference, (pp. 1-4). IEEE. April 2009.
[0426] [5] Kuntz, A., Disch, S., Backstrom, T., &Robi 11 iard, J. The Transient Steering Decorrelator Tool in the Upcoming MPEG Unified Speech and Audio Coding Standard. Audio Engineering Society Conference 131, 2011 October.
[0427] [6] Herre, Jurgen, and James D. Johnston. Enhancing the performance of perceptual audio coders by using temporal noise shaping (TNS).'Audio Engineering Society Conference 101. 1996.
37 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN101390158A | Cites | China | A | Search report | 1-17 |
| CN1121620A | Cites | China | A | Search report | 1-17 |
| CN1377499A | Cites | China | A | Search report | 1-17 |
| CN1486486A | Cites | China | A | Search report | 1-17 |
| US5710863A | Cites | United States of America | A | Search report | 1-17 |
57 members in 18 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 131713141 | European Patent Office (EPO) | – | |
| 13171314 | European Patent Office (EPO) | A | |
| 141670703 | European Patent Office (EPO) | – | |
| 14167070 | European Patent Office (EPO) | A | |
| 2014062034 | European Patent Office (EPO) | W |
Members57
| Document | Office | Kind | |
|---|---|---|---|
| CA2914418A1 | Canada | A1 | |
| CA2914771A1 | Canada | A1 | |
| WO2014198724A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2014198726A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2014280256A1 | Australia | A1 | |
| AU2014280258A1 | Australia | A1 | |
| SG11201510162WA | Singapore | A | |
| SG11201510164RA | Singapore | A | |
| CN105340010A | China | A | |
| KR20160022338A | Republic of Korea | A | |
| KR20160028420A | Republic of Korea | A | |
| CN105431902AThis record | China | A | |
| MX2015016789A | Mexico | A | |
| EP3008725A1 | European Patent Office (EPO) | A1 | |
| EP3008726A1 | European Patent Office (EPO) | A1 | |
| MX2015016984A | Mexico | A | |
| US2016148621A1 | United States of America | A1 | |
| US2016155451A1 | United States of America | A1 | |
| JP2016524186A | Japan | A | |
| JP2016526695A | Japan | A | |
| AU2014280256B2 | Australia | B2 | |
| AU2014280258B2 | Australia | B2 | |
| AU2014280258B9 | Australia | B9 | |
| CA2914418C | Canada | C | |
| EP3008725B1 | European Patent Office (EPO) | B1 | |
| RU2015156490A | Russian Federation | A | |
| RU2015156587A | Russian Federation | A | |
| HK1223725A | Hong Kong, China | A | |
| HK1223725A1 | Hong Kong, China | A1 | |
| HK1223726A | Hong Kong, China | A | |
| HK1223726A1 | Hong Kong, China | A1 | |
| BR112015030672A2 | Brazil | A2 | |
| BR112015030686A2 | Brazil | A2 | |
| EP3008726B1 | European Patent Office (EPO) | B1 | |
| ZA201600080B | South Africa | B | |
| ES2635026T3 | Spain | T3 | |
| KR101789083B1 | Republic of Korea | B1 | |
| JP6224233B2 | Japan | B2 | |
| JP6224827B2 | Japan | B2 | |
| KR101789085B1 | Republic of Korea | B1 | |
| PT3008726T | Portugal | T | |
| ES2646021T3 | Spain | T3 | |
| MX353042B | Mexico | B | |
| MX353188B | Mexico | B | |
| PL3008726T3 | Poland | T3 | |
| US9953659B2 | United States of America | B2 | |
| RU2660633C2 | Russian Federation | C2 | |
| CA2914771C | Canada | C | |
| US2018204582A1 | United States of America | A1 | |
| RU2662921C2 | Russian Federation | C2 | |
| US10115406B2 | United States of America | B2 | |
| CN105340010B | China | B | |
| MY170179A | Malaysia | A | |
| CN105431902B | China | B | |
| US10734008B2 | United States of America | B2 | |
| BR112015030672B1 | Brazil | B1 | |
| BR112015030686B1 | Brazil | B1 |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent grantGrantedGR01 | GR01 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 105431902
- Application
- 800332950
Titles2
- Chinese
- 用于通过应用分布量化和编码建模累积和表示的音频信号包络编码、处理和解码的装置和方法
- English
- Apparatus and method for encoding, processing and decoding audio signal envelope accumulated and expressed by applying distributed quantization and encoding modeling
Classification
- CPC, 5
- G10L19/06
- G10L19/032
- G10L19/0204
- H03M7/30
- G10L19/0208
- IPC, 2
- G10L19 06
- G10L19 032