Multichannel audio coding
Abstract
A method for decoding M encoded audio channels representing N audio channels, where N is two or more, and a set of one or more spatial parameters, the method comprising: a) receiving said M encoded audio channels and said set of spatial parameters, b) obtaining N audio signals from said M encoded channels, in which each audio signal is divided into a plurality of frequency bands, where each band comprises one or more spectral components, and c) generate a multichannel output signal from the N audio signals and spatial parameters, whereby M is two or more, at least one of said N audio signals is a correlated signal obtained from a weighted combination of at least two of said M encoded audio channels, said set of spatial parameters includes a first parameter indicative of the amount of a signal without correlation to be mixed with a correlated signal, and step c) includes obtaining at least one uncorrelated signal from said correlated signal, and controlling the proportion of said at least one correlated signal with respect to said at least one uncorrelated signal in at least one channel of said multichannel output signal, in response to one or some of said spatial parameters, wherein said control it is at least in part, in accordance with said first parameter.

Term
Term ended
Projected expiry passed 28 February 2025, 1.6 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
12 claims: 6 independent, 6 dependent
- 1ES 2 324 926 T3 ES 2 324 926 T3 CLAIMS REIVINDICACIONES 1. A method for decoding M encoded audio channels representing N audio channels, where N is two or more, and a set of one or more spatial parameters, the method comprising:1. Un método para descodificar M canales de audio codificados que representan N canales de audio, donde N es dos o más, y un conjunto de uno o más parámetros espaciales, comprendiendo el método: a) receive said M coded audio channels and said set of spatial parameters, a) recibir dichos M canales de audio codificados y dicho conjunto de parámetros espaciales, b) obtain N audio signals from said M coded channels, in which each audio signal is divided into a plurality of frequency bands, where each band comprises one or more spectral components, and b) obtener N señales de audio a partir de dichos M canales codificados, en los que cada señal de audio se divide en una pluralidad de bandas de frecuencia, donde cada banda comprende uno o más componentes espectrales, y c) generar una señal de salda multicanal a partir de las N señales de audio y de los parámetros espaciales, por lo que c) generate a multichannel output signal from the N audio signals and the spatial parameters, therefore M es dos o más, al menos una de dichas N señales de audio es una señal con correlación obtenida a partir de una combinación ponderada de al menos dos de dichos M canales de audio codificados, dicho conjunto de parámetros espaciales incluye un primer parámetro indicativo de la cantidad de una señal sin correlación a mezclar con una señal con correlación, y el paso c) incluye la obtención de al menos una señal sin correlación a partir de dicha señal con correlación, y el control de la proporción de dicha al menos una señal con correlación con respecto a dicha al menos una señal sin correlación en al menos un canal de dicha señal de salida multicanal, como respuesta a uno o algunos de dichos parámetros espaciales, donde dicho control es al menos en parte, conforme con dicho primer parámetro. M is two or more, at least one of said N audio signals is a correlated signal obtained from a weighted combination of at least two of said M coded audio channels, said set of spatial parameters includes a first parameter indicative of the amount of an uncorrelated signal to mix with a correlated signal, and step c) includes obtaining at least one uncorrelated signal from said correlated signal, and controlling the proportion of said at least one correlated signal with respect to said at least one uncorrelated signal in at least one channel of said multichannel output signal, in response to one or some of said spatial parameters, wherein said control it is at least in part compliant with said first parameter.
- 8The method of any of claims 1-7, further comprising varying the magnitudes of the spectral components, in at least one of said N audio signals, in response to one or more of said spatial parameters. 8. El método de cualquiera de las reivindicaciones 1 - 7, que comprende además la variación de las magnitudes de los componentes espectrales, en al menos una de dichas N señales de audio, como respuesta a uno o algunos de dichos parámetros espaciales.
Independent claims6
424 paragraphs in 25 sections, as filed
ES 2 324 926 T3
DESCRIPTION
Multi-channel audio decoding.
Technical field
The invention is generally related to audio signal processing. The invention is particularly useful in processing low and very low bit rate audio signals. More particularly, there are aspects of the invention that are related to a decoding method for audio signals, in which a plurality of audio channels are represented by means of a composite monophonic ("mono") audio channel and auxiliary information. ("Sidechain" or side chain). Alternatively, the plurality of audio channels is represented by a plurality of audio channels and by side chain information. The invention is defined by the appended claims.
Previous technique
In the AC-3 digital audio encoding and decoding system, channels can be selectively combined or “coupled” at high frequencies, when the system is greedy for bits. The details of the AC-3 system are well known in the art (see for example: ATSC standard A52 / a: Digital Audio Compression Standard (AC-3), Revision A, Committee on Advanced Television Systems, 20 August 2001. Document A / 52A is available on the World Wide Web at http://www.atsc.org/standards.html.
The frequency above which the AC-3 system combines channels on demand is called the "coupling" frequency. Above the coupling frequency, the coupled channels are combined into a "coupling" channel or composite channel. The encoder generates "coupling coordinates" (amplitude scale factors) for each subband above the coupling frequency on each channel. The coupling coordinates indicate the ratio of the original energy of each coupled channel subband to the energy of the corresponding subband in the composite channel. Below the coupling frequency, the channels are discreetly scrambled. The phase polarity of a coupled channel subband can be reversed prior to combining the channel with one or more of the other coupled channels, in order to reduce cancellation of components of the out-of-phase signal. The composite channel, together with the side chain information including, on a subband basis, the coupling coordinates and whether the channel phase is reversed, are sent to the decoder. In practice, the coupling frequencies used in commercial embodiments of the AC3 system have been rated from about 10 kHz to about 2500 Hz. US Patents 5,583,962; 5,633,981, 5,727,119, 5,909,664 and 6,021,386 include teachings related to combining multiple audio channels into a composite channel, and auxiliary or side chain information and retrieval from them of an approach to multiple channels. originals.
Relevant prior art methods are known from WO 03/090280 A1, which describes the encoding of N input audio channels to a monophonic audio signal and a set of spatial parameters including a parameter representing a measure of the waveform similarity of the N input audio channels, which are, the difference in level between channels and a selected difference between the time difference between channels and the phase difference between channels. The monophonic audio signal is obtained from the N input audio channels, adding the N input audio channels after a phase correction according to the time difference between channels. The corresponding decoding method disclosed in this document synthesizes the replica of the N audio input channels, of the monophonic audio signal and of the spatial parameters.
Description of the invention
There are aspects of the present invention that can be considered as improvements over the "coupling" techniques of the AC-3 encoding and decoding system and also over other techniques in which multiple audio channels are combined, either in a composite monophonic signal. or on multiple audio channels, along with related auxiliary information and from which multiple audio channels are reconstructed. There are aspects of the present invention that can be considered as improvements over techniques for downmixing multiple audio channels into a mono audio signal or multiple audio channels, and for uncorrecting multiple audio channels obtained from a single audio signal. mono audio channel or from multiple audio channels.
There are aspects of the invention that can be employed in an N: 1: N spatial audio coding technique (where "N" is the number of audio channels), or in an M: 1 spatial audio coding technique. : N (where “M” is the number of encoded audio channels and “N” is the number of decoded audio channels) that improve the coupling on the channel, providing, among other things, improved phase compensation, uncorrelation mechanisms and signal-dependent variable time constants. Some aspects of the present invention may be employed in the N: x: N and M: x: N techniques of spatial audio coding, where "x" can be 1 or greater than 1. The objectives include reducing coupling cancellation artifacts in the encoding process, adjusting the relative phase between channels before downmixing, and improving the spatial dimensionality of the reproduced signal, restoring phase angles and degrees of displacement. correlation in the decoder. Some aspects of the invention, when embodied in practical embodiments, should allow the coupling
ES 2 324 926 T3 to continuous channel rather than on demand, and lower coupling frequencies than, for example, in the AC-3 system, thereby reducing the required data rate.
Description of the drawings
Figure 1 is an idealized block diagram, showing the main functions or devices of an N: 1 encoding configuration, embodying aspects of the present invention.
Figure 2 is an idealized block diagram showing the main functions or devices of a 1: N decoding configuration, embodying aspects of the present invention.
Figure 3 shows an example of a simplified conceptual organization of bins (sub-channels or intervals) and sub-bands along a (vertical) axis of frequencies and blocks and a frame along the (horizontal) axis of times. The figure is not to scale.
Figure 4 is in the form of a hybrid flow diagram and a functional block diagram showing encoding steps or devices that perform functions of an encoding configuration embodying aspects of the present invention.
Figure 5 is in the form of a hybrid flow diagram and a functional block diagram showing decoding steps or devices that perform functions of a decoding configuration embodying aspects of the present invention.
Figure 6 is an idealized block diagram showing the main functions or devices of a first N: x encoding configuration, embodying aspects of the present invention.
Figure 7 is an idealized block diagram showing the main functions or devices of an x: M decoding configuration, embodying aspects of the present invention.
Figure 8 is an idealized block diagram showing the main functions or devices of a first alternative of decoding x: M configuration, embodying aspects of the present invention.
Figure 9 is an idealized block diagram showing the main functions or devices of a second alternative x: M decoding configuration, embodying aspects of the present invention.
Best mode of carrying out the invention
Basic encoder N: 1
Referring to Figure 1, an N: 1 encoding device or function is illustrated embodying aspects of the present invention. The figure is an example of a function or structure performed by a basic encoder that embodies aspects of the invention. Other functional or structural configurations that implement aspects of the invention may be employed, including alternative and / or equivalent functional or structural configurations, described below.
Two or more channels of audio input are applied to the encoder. Although, in principle, aspects of the invention may be practiced by analog, digital, or hybrid analog / digital embodiments, the examples presented herein are digital embodiments. Thus, the input signals can be time samples that may have been obtained from analog audio signals. Time samples can be encoded as pulse code modulation (PCM) signals. Each linear PCM audio input channel is processed by a device or filter bank function that has both in-phase and quadrature output, such as a Discrete Fourier Transform (DFT) in 512-point forward windows ( as implemented with the Fast Fourier Transform (FFFT)). The filter bank can be considered as a transform from the time domain to the frequency domain.
Figure 1 shows a first PCM channel input (channel "1") applied to a filter bank function or device, "Filter bank" 2, and a second PCM channel input (channel "n") applied, respectively. , to another filter bank function or device, “Filter bank” 4. There may be “n” input channels, where “n” is a positive integer equal to two or more. Thus, there are also “n” filter banks, each receiving a unique channel from the “n” input channels. For simplicity of presentation, Figure 1 shows only two input channels, "1" and "n".
When a filter bank is implemented by means of an FFT, the input signals in the time domain are segmented into consecutive blocks and are normally processed in overlapping blocks. The discrete frequency outputs of the FFT (transformation coefficients) are called bins, each of them having a complex value with real and imaginary parts corresponding, respectively, to the in-phase and quadrature components. The contiguous bins of the transformation can be grouped into sub-bands, which approximate the critical frequency bands of the human ear, and most of the side chain information produced by the encoder,
ES 2 324 926 T3 as will be described, can be calculated and transmitted on a sub-band basis, in order to minimize processing resources and to reduce the bit rate. Multiple successive blocks can be framed in the time domain, with individual block values averaged, or combined or otherwise accumulated across each frame, to minimize the sidechain data rate. In examples described herein, each filter bank is implemented by means of an FFT, contiguous transformation bins are grouped into subbands, blocks are grouped into frames, and the sidechain data is sent in once per frame basis.
Alternatively, the sidechain data can be sent based on more than once per frame (eg, once per block). See for example figure 3 and its description, below. As is well known, there is a trade-off between the frequency at which sidechain information is sent and the required bit rate.
A suitable practical implementation of aspects of the present invention may employ fixed length frames of about 32 milliseconds, when a 48 kHz sample rate is employed, with each frame having six blocks at intervals of about 5.3 milliseconds each ( using, for example, blocks that have a duration of about 10.6 milliseconds, with 50% overlap). However, neither such times nor the use of fixed-length frames, nor their division into a fixed number of blocks, are critical to putting aspects of the invention into practice, provided that the information described in this specification, which is sent based on to frames, be sent no less frequently than about 40 milliseconds. Frames can be arbitrary in size, and their size can vary dynamically. Variable lengths of the blocks can be used as in the AC-3 system mentioned above. With that understanding reference is made herein to "frames" and "blocks".
In practice, if the mono- or multichannel composite signal (or signals), or if the mono- or multichannel composite signal (or signals) and the discrete low-frequency channels are encoded, for example by means of a perceptual encoder, As described below, it is convenient to use the same frame and block configuration as used in the perceptual encoder. Furthermore, if the encoder employs variable block lengths, such that there is, from time to time, a switch from one block length to another, it would be desirable for one or more of the sidechain information described herein to be updated when such block switching takes place. In order to minimize the increase in overhead data when updating the side chain information when such a switch takes place, the resolution of the updated side chain frequency can be reduced.
Figure 3 shows an example of a simplified conceptual organization of bins and sub-bands along a (vertical) axis of frequencies and blocks and a plot along a (horizontal) axis of times. When the bins are divided into sub-bands, which approximate critical bands, the sub-bands of the lowest frequency have the fewest number of bins (for example, one) and the number of bins per sub-band increases as it increases. the frequency.
Returning to Figure 1, a version in the frequency domain of each of the n input channels of the time domain, generated by the respective filter bank of each channel (Filter banks 2 and 4 in this example) are added together ("downmixed" to obtain a monophonic ("mono") composite audio signal using an additive combination device or function "Additive Combiner" 6.
Downmix can be applied to the entire frequency bandwidth of the audio input signals or, optionally, can be limited to frequencies above a given "coupling" frequency, insofar as artifacts from the downmixing process may become more audible at mid to low frequencies. In such cases, the channels can be transmitted discreetly, below the coupling frequency. This strategy may be desirable even when process artifacts are not a problem, because the mid / low frequency sub-bands constructed by grouping transformation bins into critical band type sub-bands (size practically proportional to the frequency) tend to have a small number of transformation bins at low frequencies (a bin at very low frequencies) and can be directly encoded with fewer or fewer bits than what is required to send a mono audio signal downmixed with information from the chain side. A coupling or transition frequency as low as 4 kHz, 2300 Hz, 1000 Hz, or even the lowest part of the frequency band of the audio signals applied to the encoder, may be acceptable for some applications, particularly those where a very low bit rate is important. Other frequencies can provide a useful balance between bit savings and listener acceptance. The choice of a particular coupling frequency is not critical to the invention. The coupling frequency can be variable and, if it is variable, it can depend, for example, directly or indirectly on the characteristics of the input signal.
Before downmixing, it is an aspect of the present invention to improve the phase angle settings of the channels, opposite each other, in order to reduce the cancellation of the phase-shifted components of the signal, when the channels are combined, and to provide an improved composite mono channel. This can be achieved by varying the “absolute angle” of some or all of the transform bins on some of the channels in a controlled time. For example, all transform bins that represent audio that is above the coupling frequency, thus defining a frequency band of interest, can be time-controlled as needed on each channel, or when using a channel as a reference, on all channels except the reference channel.
ES 2 324 926 T3
The "absolute angle" of a bin can be thought of as the angle of the magnitude-and-angle representation of each complex-value transformation bin generated by a filter bank. The controllable variation of the absolute angles of the bins in a channel is done by means of an angle rotation function or device ("Angle Rotation"). Angle Rotation 8 processes the output of Filter Bank 2, prior to its application to the down-mix sum provided by Additive Combiner 6, while Angle Rotation 10 processes the output of Filter Bank 4, prior to its application. application to Additive Combiner 6. It will be appreciated that, under certain signal conditions, no angle rotation may be required for a particular transform bin over a period of time (the period of time of a frame, in the examples described herein). Below the coupling frequency, the channel information can be discreetly encoded (not illustrated in Figure 1).
In principle, it is possible to achieve an improvement in the adjustments of the phase angles of the channels, one with respect to the other, by changing the phase of each transformation bin or sub-band, by the negative value of its absolute phase angle, in each block along the frequency band of interest. While this substantially prevents cancellation of out-of-phase components of the signal, it tends to cause artifacts that may be audible, particularly if the resulting mono composite signal is heard in isolation. Therefore, it is desirable to employ the "minimum treatment" principle by varying the absolute angles of the bins in one channel only, as necessary, to minimize offset cancellation in the downmix process and minimize spatial image collapse. of the multichannel signals reconstituted by the decoder. Techniques for determining such angle variations are described below. Such techniques include time and frequency smoothing and the manner in which the signal processing responds to the presence of a transient.
Power normalization can also be performed on a bin basis in the encoder, to further reduce any remaining offset cancellation on isolated bins, as described in more detail below. As also described in more detail below, energy normalization can also be performed on a sub-band basis (in the decoder) to ensure that the energy of the composite mono signal is equal to the sums of the energies of the contributing channels.
Each input channel has an audio analyzer function or device ("Audio Analyzer") associated with it, to generate the sidechain information for that channel and to control the amount or degree of angle rotation applied to the channel before be applied to sum 6 of the down-mix. The filter bank outputs of channels 1 and n are applied to Audio Analyzer 12 and Audio Analyzer 14, respectively. Audio Analyzer 12 generates the side chain information for channel 1 and the amount of phase angle rotation for channel 1. Audio Analyzer 14 generates the side chain information for channel n and the amount of rotation of the angle for channel n. It will be understood that such references herein to "angle" refer to phase angle.
Sidechain information for each channel, generated by an audio analyzer for each channel, can include:
an Amplitude Scale Factor (“Amplitude SF”), an Angle Control Parameter, a De-correlation Scale Factor (“De-correlation SF”), a Transient Flag, and optionally, a De-correlation Flag. Interpolation.
Such side chain information may be characterized as "spatial parameters", indicative of the spatial properties of the channels and / or indicative of signal characteristics that may be relevant to spatial processing, such as transients. In each case, the sidechain information applies to an individual sub-band (except for the Transient Flagger and Interpolation Flagger, each of which applies to all sub-bands within a channel) and it can be updated once per frame, as in the examples described below, or when a block switch occurs in a related encoder. Additional details of the various spatial parameters are set out below. The angle rotation for a particular channel of the encoder can be taken as the Reverse Polarity Angle Control Parameter which is part of the side chain information.
If a reference channel is used, that channel may not require an Audio Analyzer or, alternatively, may require an Audio Analyzer that generates only Amplitude Scale Factor side chain information. It is not necessary to send an Amplitude Scale Factor if that scale factor can be deduced with sufficient precision by means of a decoder from the Amplitude Scale Factors of the other, non-reference channels. It is possible to deduce in the decoder the approximate value of the Amplitude Scale Factor of the reference channel, if the normalization of the energy in the encoder ensures that the scale factors across the channels within any sub-band squared add substantially 1, as described below. The approximate value of the Amplitude Scale Factor of the reference channel may have errors
ES 2 324 926 T3 as a result of the relatively rough quantization of the amplitude scale factors that result in image variations in the reproduced multichannel audio. However, in a low data rate environment, such artifacts may be more acceptable than using the bits to send the Amplitude Scale Factor of the reference channel. However, in some cases it may be desirable to employ an audio analyzer for the reference channel that generates at least Amplitude Scale Factor side chain information.
Figure 1 shows in dotted line an optional input to each audio analyzer, from the PCM time domain input to the channel's audio analyzer. This input can be used by the Audio Analyzer to detect a transient in a period of time (the period of a block or frame, in the examples described herein) and to generate a transient indicator (for example, a "Flag of Transients ”of one bit) in response to a transient. Alternatively, as described below, in the comments to step 408 of FIG. 4, a transient may be detected in the frequency domain, in which case the Audio Analyzer does not need to receive a time domain input.
The mono composite audio signal and sidechain information for all channels, (or all channels except the reference channel) can be stored, transmitted or stored and transmitted to a decoding process or device ("Decoder"). . Prior to storage, transmission or storage and transmission, the various audio signals and various sidechain information may be multiplexed and packed into one or more bit strings suitable for the storage, transmission, or storage and transmission medium (s). Mono composite audio can be applied to an encoding process or device that reduces the data rate, such as, for example, a perceptual encoder or a perceptual encoder and an entropy encoder (for example, an arithmetic or Huffman encoder) (sometimes referred to as a "lossless" encoder) prior to storage, transmission, or storage and transmission. Furthermore, as mentioned above, composite mono audio and related side chain information can be obtained from multiple input channels only for audio frequencies above a certain frequency (a "coupling" frequency). In that case, the audio frequencies below the coupling frequency on each of the multiple input channels can be stored, transmitted, or stored and transmitted as discrete channels, or they can be combined or processed in some way other than described. in this memory. Such discrete or somehow combined channels can also be applied to a data-reducing encoding process or device such as, for example, a perceptual encoder or a perceptual encoder and an entropy encoder. Mono composite audio and discrete multichannel audio can all be applied to an integrated perceptual encoding or perceptual and entropy encoding process or device.
The particular way in which the sidechain information is conveyed in the encoder bitstream is not critical to the invention. If desired, the sidechain information can be carried in such a way that the bitstream is compatible with legacy decoders (ie, the bitstream is backward compatible). Many suitable techniques are known to do this. For example, many encoders generate a bit string that has unused or null bits that are ignored by the decoder. An example of such a configuration is set forth in US Patent 6,807,528 B1 by Truman et al., Entitled "Adding data to a Compressed Data Frame", dated October 19, 2004. Such bits can be replaced by side chain information. Another example is that the side chain information can be steganographically encoded in the encoder bit stream. Alternatively, the sidechain information may be stored or transmitted separately from the backward compatible bitstream, by any technique that allows such information to be transmitted or stored together with a mono / stereo bitstream compatible with legacy decoders.
Basic 1: N and 1: M decoder
Referring to Figure 2, a decoding device or function ("Decoder") is illustrated embodying aspects of the present invention. The figure is an example of a function or structure that behaves like a basic decoder that embodies an aspect of the invention. Other functional or structural configurations that implement aspects of the invention may be employed, including alternative and / or equivalent functional or structural configurations described below.
The decoder receives the mono composite audio signal and side chain information for all channels or all channels except the reference channel. If necessary, the composite audio signal and related sidechain information are demultiplexed, unpacked, and / or decoded. Decoding can use a table query. The objective is to obtain from the mono channels of composite audio, a plurality of individual audio channels that approximate the respective audio channels applied to the Encoder of figure 1, conditioned to the bit rate reduction techniques of the present invention described herein.
Naturally, you can choose not to recover all the channels applied to the encoder or to use only the composite mono signal. The channels recovered by a Decoder practicing the present invention are particularly useful in relation to the channel multiplication techniques of the cited and incorporated applications in that the recovered channels not only have useful inter-channel amplitude relationships, but also they also have useful phase relationships between channels. Another alternative for channel multiplication is to use
ES 2 324 926 T3 a matrix decoder to obtain additional channels. The inter-channel amplitude and phase conservation aspects of the present invention make the output channels of a decoder embodying aspects of the present invention particularly suitable for application to an amplitude- and phase-sensitive matrix decoder. Many such matrix decoders employ wideband control circuitry that works properly only when the signals applied to them are stereo over the entire bandwidth of the signals. Thus, if aspects of the present invention are embodied in an N: 1: N system in which N is 2, the two channels recovered by the decoder can be applied to a 2: M active matrix decoder. Such channels may have been discrete channels below a coupling frequency, as mentioned above. Many suitable active matrix decoders are known in the art, including, for example, matrix decoders known as "Pro Logic" and "Pro Logic II" decoders ("Pro Logic" is a trademark of Dolby Laboratories Licensing Corporation). Aspects of Pro Logic decoders are disclosed in US Patents 4,799,260 and 4,941,177. Aspects of Pro Logic II decoders are disclosed in the US pending application Serial Number 09 / 532,711 from Fosgate, entitled “Method for Deriving at Least Three Audio Signals from Two Input Audio Signals”. minus three audio signals from two audio input signals ”), filed on March 22, 2000, and published as WO 01/41504 on June 7, 2001, and in the pending United States application with the serial number 10 / 362,786 of Fosgate, entitled "Method for Apparatus for Audio Matrix Decoding" ("Method for an audio matrix decoding apparatus"), filed February 25, 2003 and published as US 2004/0125960 A1 on July 1, 2004. Some aspects of the operation of the Dolby Pro Logic and Pro Logic II decoders are explained, for example, in documents available on the Dolby Laboratories website (www.dolby.com): “Dolby Surround Pro Logic Decoder Principles of Operation” (“ Dolby Surrounding Pro Logic Decoder Principles of Operation ”) by Roger Dressler and Jim Hilson's“ Mixing with Dolby Pro Logic II Technology ”.
Referring again to Figure 2, the received composite mono audio channel is applied to a plurality of signal paths, from which a respective channel is obtained from each of the multiple recovered audio channels. Each channel obtaining path includes, in any order, an amplitude adjustment function or device ("Amplitude Adjustment") and an angle rotation function or device ("Angle Rotation").
The Amplitude Adjustments apply gains or losses to the composite mono signal so that, under certain signal conditions, the relative magnitudes (or energies) of the output channels obtained from it are similar to those of the channels to the signal. decoder input. Alternatively, under certain signal conditions, when "random" angle variations are imposed, as described below, a controllable amount of "random" variations can also be imposed on the amplitude of a recovered channel, in order to improve its de-correlation with respect to other of the recovered channels.
Angle Rotation applies phase rotations so that, under certain signal conditions, the relative phase angles of the output channels, obtained from the composite mono signal, are similar to those of the channels at the input of the signal. encoder. Preferably, under certain signal conditions, a controllable amount of "random" variations is also imposed on the angle of a recovered channel, in order to improve its decorrelation with respect to other recovered channels.
As discussed in more detail below, "random" angle amplitude variations can include not only pseudo-random and certainly random variations, but also deterministically generated variations that have the effect of reducing cross-correlation between channels. This is discussed in more detail below in the comments to step 505 of FIG. 5A.
Conceptually, Amplitude Adjustment and Angle Rotation for a particular channel scale the DFT coefficients of composite mono audio to obtain reconstructed transform bin values for the channel.
The Amplitude Adjustment for each channel can be controlled at least by the scale factor of the Amplitude of the recovered side chain for the particular channel or, in the case of the reference channel, from the Amplitude Scale Factor of the chain. side channel recovered for the reference channel or from an Amplitude Scale Factor deduced from the Amplitude Scale Factors of the side chain recovered from the other channels, other than the reference channel. Alternatively, to reinforce the de-correlation of the recovered channels, the Amplitude Adjustment can also be controlled by a Random Amplitude Scale Factor Parameter, obtained from the De-correlation Scale Factor of the recovered side chain, to a particular channel, and the recovered sidechain Transient Flag for the particular channel.
The Angle Rotation for each channel can be controlled at least by the Angle Control Parameter of the recovered side chain (in which case, the Angle Rotation in the decoder can substantially undo the angle rotation provided by the Angle Rotation of the encoder). To reinforce the de-correlation of the recovered channels, an Angle Rotation can also be controlled by means of the Random Angle Control Parameter obtained from the De-correlation Scale Factor of the recovered side chain for a particular channel and from the Transient Flag of the recovered side chain for the particular channel. The
ES 2 324 926 T3
Random Angle Control Parameter for a channel and, if used, the Random Amplitude Scale Factor for a channel, can be obtained from the Recovered De-correlation Scale Factor for the channel and the recovered Transient Flag for the channel, by means of a controllable décorrelation device or function ("Controllable Décorrelation Device").
Referring to the example of Figure 2, the recovered mono composite audio is applied to a first channel audio recovery path 22, which obtains the channel 1 audio, and to a second channel audio recovery path 24, which obtains channel n audio. The audio path 22 includes an Amplitude Adjustment 26, an Angle Rotation 28 and, if desired a PCM output, an inverse filter bank function or device ("Inverse Filter Bank") 30. Similarly, the audio path 24 includes an Amplitude Adjustment 32, an Angle Rotation 34 and, if a PCM output is desired, an inverse filter bank function or device ("Reverse Filter Bank") 36. As in the case of Figure 1, only two channels are illustrated for simplicity of presentation, it being understood that there may be more than two channels.
The side chain information retrieved for the first channel, channel 1, may include an Amplitude Scale Factor, an Angle Control Parameter, a De-correlation Scale Factor, a Transient Flagger and, optionally, an Amplitude Flagger. Interpolation, as stated above with respect to the description of a Basic Encoder. The Amplitude Scale Factor is applied to the Amplitude Setting 26. If the optional Interpolation Flagger is used, an optional frequency interpolator or an interpolator function (“Interpolator”) 27 can be used in order to interpolate the Angle Control Parameter at the frequency (for example, in the bins of each sub-band of a channel). Such interpolation can be, for example, a linear interpolation of the angles of the bins between the centers of each sub-band. The state of a one-bit Interpolation Flag selects whether or not to use interpolation across frequencies, as explained in more detail below. The Transient Flag and the Décorrelation Scale Factor are applied to a Controllable Décorrelation Device 38, which generates a Random Angle Control Parameter in response thereto. The state of a one-bit Transient Flag selects one of two multiple random angle decorrelation modes, as explained in more detail below. The Angle Control Parameter, which can be interpolated across frequencies using the Interpolation Flag and Interpolator, and the Random Angle Control Parameter are added by means of an additive combiner or combination function 40, with in order to provide a control signal for Angle Rotation 28. Alternatively, the Controllable Décorrelation Device 38 may also generate a Random Amplitude Scaling Factor, in response to the Transient Flag and the Décorrelation Scaling Factor, in addition to generating a Random Angle Control Parameter. The Amplitude Scale Factor can be added to such a Random Amplitude Scale Factor by means of an additive combiner or combining function (not illustrated) in order to provide the control signal for the Amplitude Adjustment 26.
Similarly, the side chain information retrieved for the second channel, channel n, may also include an Amplitude Scaling Factor, an Angle Control Parameter, a Décorrelation Scaling Factor, a Transient Flag, and optionally , an Interpolation Flag, as described above in connection with the description of a basic encoder. The Amplitude Scale Factor is applied to Amplitude Setting 32. A frequency interpolator or an interpolation function (“Interpolator”) 33 can be used in order to interpolate the Angle Control Parameter across frequencies. As with channel 1, the Interpolation Flag state selects whether or not to use interpolation across frequencies. The Transient Flagger and the Décorrelation scale factor are applied to a Controllable Décorrelation Device 42 which generates a Random Angle Control Parameter in response thereto. As with channel 1, the one-bit Transient Flag state selects one of two multiple random angle decorrelation modes, as explained in more detail below. The Angle Control Parameter and the Random Angle Control Parameter are summed by means of an additive combiner or combination function 44, in order to provide a control signal for the Angle Rotation 34. Alternatively, as described above with respect to channel 1, the Controllable Décorrelation Device 42 may also generate a Random Amplitude Scaling Factor in response to the Transient Flag and the Décorrelation Scaling Factor, in addition to generating a Control Parameter. of Random Angles. The Amplitude Scale Factor and the Random Amplitude Scale Factor can be added by means of an additive combiner or combination function (not illustrated) in order to provide the control signal for the Amplitude Adjustment 32.
Although the process or topology as just described is helpful for understanding, essentially the same results can be obtained with alternative processes or topologies that achieve the same or similar results. For example, the order of Amplitude Adjustment 26 (32) and Angle Rotation 28 (34) may be reversed and / or there may be more than one Angle Rotation, one that responds to the Angle Control Parameter and another that responds to the Random Angle Control Parameter. The Angle Rotation can also be considered as three instead of one or two functions or devices, as in the example of figure 5 described below. If a Random Amplitude Scale Factor is used, there may be more than one Amplitude Adjustment, one that responds to the Amplitude Scale Factor and one that responds to the Random Amplitude Scale Factor. Because the human ear has a greater sensitivity to amplitude than to phase, if a Random Amplitude Scaling Factor is used, it may be desirable to scale its effect relative to the effect of the Random Angle Control Parameter, of so that its effect on amplitude is less than the effect that the Random Angle Control Parameter has on phase angle. As another alternative process or topology, the Décorrelation Scale Factor can be
ES 2 324 926 T3 used to control the ratio of the random phase angle as a function of the basic phase angle (instead of adding a parameter representing the random phase angle to a parameter representing the basic phase angle) and, if is also used, the ratio of the random amplitude variation as a function of the basic amplitude variation (instead of adding a scale factor that represents a random amplitude to a scale factor that represents the basic amplitude) (that is, a fade and appear simultaneous variable in each case).
If a reference channel is used, as previously discussed with respect to the basic encoder, the Angle Rotation, Controllable Decorrelation Device and Additive Combiner for that channel can be omitted, considering that the side chain information for the channel Reference may only include the Amplitude Scale Factor (or, alternatively, if the side chain information does not contain an Amplitude Scale Factor) for the reference channel, it can be deduced from the Amplitude Scale Factors of the other channels, when the energy normalization in the encoder ensures that the scale on channels within a squared sub-band add up to 1). An Amplitude Setting is available for the reference channel and is controlled by an Amplitude Scale Factor of the reference channel. If the Amplitude Scale Factor of the reference channel is obtained from the side chain, or is deduced at the decoder, the recovered reference channel is a scaled version of the amplitude of the mono composite channel. It does not require an angle rotation because it is the reference for the rotations of the other channels.
Although adjusting the relative amplitude of the recovered channels can provide a modest degree of decorrelation, using amplitude adjustment alone is likely to result in a reproduced sound field substantially lacking in spatial or imaging quality for many conditions. signal (for example, a "collapsed" sound field). Amplitude adjustment can affect interaural level differences in the ear, which is only one of the directional psychoacoustic cues employed by the ear. Thus, in accordance with aspects of the invention, certain angle adjustment techniques may be employed, depending on signal conditions, to provide additional decorrelation. Reference may be made to Table I which provides abbreviated comments useful for understanding the multiple angle adjustment décorrelation techniques or modes of operation that may be employed in accordance with aspects of the invention. Other décorrelation techniques may be employed as described below, with respect to the examples in Figures 8 and 9, instead of, or in addition to, the techniques in Table I.
In practice, applying angle rotations and magnitude alterations can result in a circular convolution (also known as a cyclic or periodic convolution). Although it is generally desirable to avoid circular convolution, the unwanted audible artifacts that result from circular convolution are somewhat reduced by complementary angle variation in an encoder and a decoder. Furthermore, the effects of circular convolution can be tolerated in low cost implementations of aspects of the present invention, particularly those in which downmixing to mono or multiple channels takes place only in part of the audio frequency band, such as such as above 1500 Hz (in which case the audible effects of circular convolution are minimal). Alternatively, circular convolution can be avoided or minimized by any suitable technique, including for example an appropriate use of zero padding. One way to use zero padding is to transform the proposed frequency domain variation (representing angle rotations and amplitude scale) to the time domain, window (with arbitrary windows), fill with zeros, then transform back to frequency domain and multiply by the version of the frequency domain of the audio to be processed (the need for the audio will not be done in windows).
(Table goes to next page)
ES 2 324 926 T3
TABLE I
Angle Adjustment Decorrelation Techniques
<td></td><td>Technique 1</td><td>Technique 2</td><td>Technique 3</td>
<td>Signal type (typical example)</td><td>Spectrally static source</td><td>Complex continuous signals</td><td>Complex impulse signals (transients)</td>
<td>Effect on decorrelation</td><td>Performs low-frequency de-correlation and steady-state signal components</td><td>Performs decorelation of complex pulse signal components</td><td>Performs decorelation of high-frequency pulse signal components</td>
<td>Transient effect present in the frame</td><td>Operates with shortened time constant</td><td>Does not operate</td><td>Opera</td>
<td>What is done</td><td>Slowly vary (frame by frame) the angle of the bin in a channel</td><td>Add to technique 1 angle a random time-invariant angle based on bins in a channel</td><td>Adds to Technique 1 angle a random quick-change angle (block by block) based on subbands in a channel</td>
<td>Controlled or scaled by</td><td>The basic phase angle is controlled by the Angle Control Parameter</td><td>Random angle amount is scaled directly by the Decorrelation SF, the same scale for the entire sub-band, scale updated in each frame</td><td>Random angle amount is indirectly scaled by the SF of Descorrelation, the same scale for the entire subband, scale updated in each frame</td>
<td>Frequency resolution of angle variation</td><td>Sub-band (same variation or interpolated value, applied to all bins of each sub-band)</td><td>Bin (different random variation value, applied to each bin)</td><td>Sub-band (same random variation value applied to all bins of each sub-band; different random variation value applied to each sub-band of the channel)</td>
<td>Time resolution</td><td>Frame (variation values updated in each frame)</td><td>Random variation values remain the same and do not change</td><td>Block (updated random variation values in each block)</td>
For signals that are substantially spectrally static, such as, for example, a note from a piccolo, a first technique ("Technique 1") restores the angle of the received mono composite signal relative to the angle of each of the other channels. recovered, at the original angle (subject to frequency and time detail level and quantization) at the original angle of the channel with respect to the other channels at the decoder input. Phase angle differences are useful, particularly for providing decorrelation of low-frequency signal components, below about 1500 Hz, where the ear follows individual cycles of the audio signal. Preferably, Technique 1 works under all signal conditions, to provide basic angle variation.
For high-frequency signal components, above around 1500 Hz, the ear does not follow individual sound cycles, but instead responds to waveform envelopes (based on the critical band). Therefore, above about 1500 Hz a better decorrelation is provided with differences in signal envelopes, rather than differences in phase angle. Applying the angle variations of
ES 2 324 926 T3 phase only according to Technique 1, the envelopes of the signals are not altered sufficiently to effect the decorrelation of the high frequency signals. The second and third techniques ("Technique 2" and "Technique 3", respectively) add a controllable amount of random variations to the angle determined by Technique 1, under certain signal conditions, thereby causing a controllable amount of random variations of envelopes, which reinforces decorrelation.
Random phase angle changes are a desirable way to cause random changes in signal envelopes. From the interaction of a particular combination of amplitudes and phases of spectral components within a sub-band, a particular envelope results. Although changing the amplitudes of spectral components within a subband changes the envelope, large changes in amplitudes are required to obtain a significant change in the envelope, which is undesirable because the human ear is sensitive to variations in the envelope. spectral amplitude. In contrast, the change in the phase angles of the spectral components has a greater effect on the envelope than the change in the amplitudes of the spectral components, the spectral components are no longer organized in the same way, so that the reinforcements and subtractions defining the envelope take place at different times, thereby changing the envelope. Although the human ear has some sensitivity to envelope, it is relatively phase deaf, so the overall sound quality remains substantially similar. However, for some signal conditions, a bit of randomness in the amplitudes of the spectral components, coupled with random phases of the spectral components can provide improved random conditions of the signal envelopes, provided that such random amplitudes do not create unwanted audible artifacts.
Preferably, a certain controllable amount or degree of Technique 2 or Technique 3 works in conjunction with Technique 1, under certain signal conditions. The Transient Flagger selects Technique 2 (there are no transients present in the frame or block, depending on whether the Transient Flagger has been sent at the frame or block rate) or Technique 3 (there are transients present in the frame or block). block). Therefore, there are multiple modes of operation, depending on whether or not there is a transient present. Alternatively, furthermore, under certain signal conditions, a controllable amount or degree of random amplitude also operates, along with amplitude scaling, which seeks to restore the original amplitude of the channel.
Technique 2 is suitable for complex continuous signals that are rich in harmonics, such as a group of violins in an orchestra. Technique 3 is suitable for complex transient impulse signals, such as clapping or castanets, etc. (Technique 2 blurs clapping claps over time, making it unsuitable for such signals.) As explained in more detail below, in order to minimize audible artifacts, Technique 2 and Technique 3 have different time and frequency resolutions to apply random angle variations, Technique 2 is selected when a transient, while Technique 3 is selected when a transient is present.
Technique 1 slowly varies (frame by frame) the angle of a bin in a channel. The amount or degree of this basic variation is controlled by the Angle Control Parameter (there is no variation if the parameter is zero). As explained in more detail below, the same parameter or an interpolated parameter is applied to all bins in each subband, and the parameter is updated on each frame. Consequently, each sub-band of each channel can have a phase variation with respect to the other channels, providing a degree of decorrelation at low frequencies (below about 1500 Hz). However, Technique 1, by itself, is unsuitable for a transient signal such as clap. For such signal conditions, the channels played back can present an annoying unstable comb filter effect. In the case of a clap, essentially no decorrelation is provided by adjusting only the relative amplitude of the recovered channels, because all channels tend to have the same amplitude during the period of a frame.
Technique 2 works when no transients are present. Technique 2 adds to the angle variation of Technique 1 a random variation of the angle that does not change with time, based on bin by bin (each bin has a different random variation) in a channel, making the envelopes of the channels are different from each other, thus providing the complex signal decorrelations between the channels. By keeping the random values of the phase angle constant over time, it is avoided that block or frame artifacts may result in the block-by-block or frame-by-frame alteration of the phase angles of the bins. Although this technique is a very useful decorelation tool when no transients are present, it can temporarily blur a transient (resulting in what is known as “pre-noise”, the post-transient blurring is masked by the transient). The amount or degree of additional variation provided by Technique 2 is scaled using the Décorrelation Scale Factor (there is no additional variation if the scale factor is zero). Ideally, the amount of random phase angle added to the basic angle variation (from Technique 1) according to Technique 2, is controlled by the Décorrelation Scale Factor in a way that minimizes chirping artifacts of the audible signal. . Such minimization of gurgle artifacts is a result of the manner in which the Décorrelation Scale Factor is obtained and the application of appropriate time smoothing, as described below. Although a different additional random angle variance value is applied to each bin, and that variance value does not change, the same scale is applied across a subband and the scale is updated on each frame.
ES 2 324 926 T3
Technique 3 works in the presence of a transient in the frame or block, depending on the speed at which the Transient Flag is sent. It varies all the bins of each sub-band in a channel from block to block, with a unique random angle value, common to all the bins of the sub-band, making not only the envelopes, but also the amplitudes and phases of the signals of one channel change with respect to the other channels from block to block. These changes in frequency and time resolution by randomizing the angle, reduce steady-state signal similarities between channels, and provide channel decorrelation, without substantially causing "pre-noise" artifacts. The change in frequency resolution of the angle randomness, from very fine (all different bins of a channel) of Technique 2, to approximate (all bins within a sub-band equal, but each sub-band different) from Technique 3, it is particularly useful for minimizing “pre-noise” artifacts. Although the ear does not respond to pure angle changes directly at high frequencies, when two or more channels are acoustically mixed on their way from the speakers to the listener, the phase differences can cause amplitude changes (comb filter effects) that can be audible and objectionable, and these are broken down by Technique 3. The pulse characteristics of the signal minimize block rate artifacts that might otherwise occur. Thus, Technique 3 adds to the phase variation of Technique 1 a rapidly changing random variation of the angle (block by block) on a sub-band to sub-band basis in a channel. The amount or degree of additional variation is indirectly scaled, as described below, by the Décorrelation Scale Factor (there is no additional variation if the scale factor is zero). The same scale is applied across a sub-band and the scale is updated on each frame.
Although angle adjustment techniques have been characterized as three techniques, this is a semantic matter and can also be characterized as two techniques: (1) a combination of Technique 1 and a variable degree of Technique 2, which can be zero. , and (2) a combination of Technique 1 and a variable degree of Technique 3, which can be zero. For the convenience of presentation, the techniques are treated as if they were three techniques.
Aspects of multiple mode décorrelation techniques and modifications can be employed to provide décorrelation of audio signals obtained, such as upmixing, from one or more audio channels, even when such audio channels are not obtained from an encoder, in accordance with aspects of the present invention. Such arrangements, when applied to a mono audio channel, are sometimes referred to as "pseudo-stereo" devices and functions. Any suitable device or function (an "up-mixer") can be used to obtain multiple signals from a mono audio channel or from multiple audio channels. Once such multiple channels of audio have been obtained by an up-mixer, one or more of them can be decorelated with respect to one or more of the other audio signals obtained, applying the multiple decorelation techniques described in this specification. In such an application, each obtained audio channel to which the décorrelation techniques are applied can be switched from one operating mode to another, detecting transients in the obtained audio channel itself. Alternatively, the operation of the technique when there are transients (Technique 3) can be simplified so as not to provide any variation of the phase angles of the spectral components, when a transient is present.
Side chain information
As mentioned above, the side chain information may include: an Amplitude Scale Factor, an Angle Control Parameter, a Décorrelation Scale Factor, a Transient Flag, and optionally an Interpolation Flag. Such side chain information for a practical embodiment of aspects of the present invention can be summarized in Table 2 below. Typically, the sidechain information can be updated once per frame.
(Table goes to next page)
ES 2 324 926 T3
TABLE 2
Sidechain information characteristics for a channel
<td>Chain information 1 a. ti € - r ci I »</td><td>Range of values</td><td>Represents (is a measure of)</td><td>Levels of quantification</td><td>Main purpose</td>
<td>For me-</td><td>0 -. + 2n</td><td>Smoothed average of</td><td>δ bits</td><td>Provides</td>
<td>another of</td><td></td><td>times in each sub-</td><td> (64</td><td>angle rotation</td>
<td>Subband Angle Control</td><td></td><td>band of the difference between the angle of each bin in the sub-band for a channel and that of the corresponding bin of the sub-band of a reference channel</td><td>levels)</td><td>basic for each bin of the channel</td>
<td>Factor</td><td> 0 -. 1</td><td>The permanence</td><td>3 bits</td><td>Make the change</td>
<td>of</td><td>The Factor</td><td>spectral of the</td><td> (8</td><td>scale of the</td>
<td>Scale</td><td>Scale of</td><td>characteristics of the</td><td>levels)</td><td>variations</td>
<td>of</td><td>Uncover it-</td><td>signal over time</td><td></td><td>random from</td>
<td>Desire-</td><td>tion of the</td><td>in a sub-band of</td><td></td><td>angle, added to</td>
<td>rrela-</td><td>sub-band is</td><td>a channel (the Factor</td><td></td><td>the rotation of the</td>
<td>tion of</td><td>tall</td><td>of permanence</td><td></td><td>basic angle and, yes</td>
<td>Sub-</td><td>only if</td><td>Spectral) and</td><td></td><td>are used, carried out</td>
<td>band</td><td>the Spectral Permanence Factor and the Angle Consistency Factor between Channels are low</td><td>consistency in the same sub-band of a channel of the angles of bins, with respect to the corresponding bins of a reference channel (the Angle Consistency Factor between channels)</td><td></td><td>also the scaling of the Random Amplitude Scaling Factor added to the basic Amplitude Scaling Factor and optionally the degree of reverb.</td>
<td>Factor</td><td>0 to 31</td><td>Energy or Amplitude in</td><td>5 bits</td><td>Make the change</td>
<td>of</td><td>(integer) 0</td><td>sub-band of a</td><td> (32</td><td>scale of the</td>
<td>Scale</td><td>is the</td><td>channel, with respect to</td><td>levels)</td><td>breadth of</td>
<td>of</td><td>amplitude plus</td><td>energy or amplitude</td><td>Level of</td><td>bins of a sub-</td>
<td>Amplitude</td><td>high and 31 is</td><td>for the same sub-</td><td>detail</td><td>band on a channel</td>
<td>by Sub-</td><td>The amplitude</td><td>band in all</td><td>1.5 dB,</td><td></td>
<td>band</td><td>more low</td><td>channels</td><td>so the range is 31 * 1.5 = 4 6.5 dB plus end value = off</td><td></td>
<td>Signaling</td><td> 1, 0</td><td>Presence of a</td><td>1 bit (2</td><td>Determine what</td>
<td>zador of</td><td>(True-</td><td>transitory in the</td><td>levels)</td><td>technique is used</td>
<td>Transient</td><td>ro / False) (polarity is arbitrary)</td><td>frame or block</td><td></td><td>to add random angle variations, both angle variations and amplitude variations</td>
<td>Signaling</td><td> 1, 0</td><td>A spectral peak</td><td>1 bit (2</td><td>Determine if the</td>
<td>zador of</td><td>(True-</td><td>near the limit of</td><td>levels)</td><td>angle rotation</td>
<td>Interpo-</td><td>ro / False)</td><td>a sub-band or the</td><td></td><td>basic © stá</td>
<td>lation</td><td>(polarity is arbitrary)</td><td>phase angles within a channel have a linear progression</td><td></td><td>interpolated as a function of frequency</td>
ES 2 324 926 T3
In each case, a channel's sidechain information applies to a single sub-band (except the Transient Flagger and Interpolation Flagger, each of which applies to all sub-bands of a channel). and it can be updated once per frame. Although timing resolution (once per frame), frequency resolution (sub-band), value ranges, and indicated quantization levels have been found to provide useful performance and a useful compromise between low bit rate data and performance, it will be appreciated that these time and frequency resolutions, ranges of values, and levels of quantization are not critical, and that other resolutions may be used, ranges and levels for putting aspects of the invention into practice. For example, the Transient Flagger and / or Interpolation Flagger, if employed, can be updated once per block, with only a minimal increase in data overhead in the sidechain. In the case of the Transient Marker, doing so has the advantage that switching from Technique 2 to Technique 3 and vice versa is more accurate. Furthermore, as mentioned above, the side chain information can be updated when a block switch of a related encoder takes place.
It will be noted that Technique 2, described above (see also Table 1), provides a frequency resolution of the bins rather than a resolution of the subband frequency (i.e., a phase angle variation is applied different pseudo-random to each bin, rather than each sub-band) even when the same Sub-band Décorrelation Scale Factor is applied to all bins in a sub-band. It will also be noted that Technique 3, described above (see also Table 1), provides a resolution of the frequency of blocks (i.e., a different random phase angle variation is applied to each block, rather than to each frame ) even though the same Sub-band Décorrelation Scale Factor is applied to all bins in a sub-band. Such resolutions, higher than the resolution of the side chain information, are possible because the random variations of the phase angles can be generated in a decoder and do not need to be known in the encoder (this is the case even when the encoder applies also a random variation of the phase angle to the coded mono composite signal, an alternative described below). In other words, it is not necessary to send sidechain information that has a level of detail of bins or blocks, even though decorelation techniques employ such a level of detail. The decoder may employ, for example, one or more random bin phase angle lookup tables. Obtaining time and / or frequency resolutions for a decorrelation greater than the information rates of the side chain are among one of the aspects of the present invention. Thus, the decorrelation by means of random phases is performed with a fine resolution of frequencies (bin by bin) that does not change with time (Technique 2), or with an approximate resolution of frequencies (band by band) ((or a resolution frequency fine (bin by bin) when using frequency interpolation, as described in more detail below)) and fine time resolution (block rate) (Technique 3).
It will also be appreciated that as increasing degrees of random phase variations are added to the phase angle of a recovered channel, the absolute phase angle of the recovered channel differs more and more from the original absolute phase angle of that channel. One aspect of the present invention is the appreciation that the resulting absolute phase angle of the recovered channel need not coincide with that of the original channel when the signal conditions are such that the random phase variations add up according to aspects of the present invention. For example, in extreme cases, when the Decorelation Scale Factor causes the highest degree of random phase variation, the phase variation caused by Technique 2 or Technique 3 exceeds the basic phase variation caused by Technique 1. . However, this is not a concern in that a random phase variation is audible the same as random phases different from the original signal that give rise to a Decorelation Scale Factor that causes the addition of some degree of random phase variations. .
As mentioned above, random amplitude variations can be employed in addition to random phase variations. For example, the Amplitude Adjustment can also be controlled by means of a Random Amplitude Scale Factor Parameter, obtained from the Recovered Side Chain Decorrelation Scale Factor for a particular channel, and the Transient Flag of the side chain retrieved for that particular channel. Such random amplitude variations can operate in two modes in a manner analogous to the application of random phase variations. For example, in the absence of a transient, a random variation of the amplitude that does not change with time can be added on a bin by bin basis, (different from bin to bin) and, in the presence of a transient (in the plot or block), a random amplitude variation that varies on a block-by-block basis (different from block to block) and changes from sub-band to sub-band (the same variation for all bins in a sub-band; different from from sub-band to sub-band). Although the amount or degree to which random amplitude variations are added can be controlled by the Décorrelation Scale Factor, it is believed that a particular value of the scale factor should result in less amplitude variation than the corresponding random phase variation that results from the same value of the scale factor, in order to avoid audible artifacts.
When a Transient Flagger is applied to a frame, the timing resolution with which the Transient Flagger selects Technique 2 or Technique 3 can be improved by providing a supplemental transient detector in the decoder, in order to provide resolution finer than the frame rate or even the block frame rate. Such a supplementary transient detector can detect when a transient occurs in the mono or multichannel composite audio signal, received by the decoder, and such detection information is sent to each Controllable Decorelation Device (such as 38, 42 in figure 2). Then, upon receiving a Transient Flagger for its channel, the Controllable Décorrelation Device switches from Technique 2 to Technique 3, upon receiving the local transient detection indication from the decoder. Thus, a substantial improvement in temporal resolution is possible without increasing the bit rate of the side chain, albeit with precision.
ES 2 324 926 T3 reduced spatial (the encoder detects transients in each input channel, before its downmix, while the detection in the decoder is done after the downmix).
As an alternative to sending the sidechain information on a frame-by-frame basis, the sidechain information can be updated on each block, at least for highly dynamic signals. As mentioned above, updating the Transient Flagger and / or Interpolation Flagger in each block, results in only a small increase in sidechain data overhead. In order to achieve such an increase in temporal resolution for other sidechain information, without substantially increasing the sidechain data rate, a block floating point differential encoding configuration can be used. For example, the consecutive blocks of the transformation can be collected in groups of six on a frame. The total information of the side chain can be sent for each channel of the sub-band in the first block. In the subsequent five blocks, only differential values can be sent, each being the difference between the amplitude and angle of the current block, and the equivalent values of the previous block. This results in a very low data rate for static signals, such as a piccolo note. For more dynamic signals, difference values with a wider range are required, but with less precision. Thus, for each group of five differential values, an exponent can be sent first, using for example 3 bits, then the differential values are quantized with precision, for example 2 bits. This arrangement reduces the mean sidechain data rate in the worst case by a factor of about two. A further reduction can be obtained by omitting the sidechain data for a reference channel (because they can be obtained from other channels), as previously studied, and using for example arithmetic coding. Alternatively, or in addition, differential encoding as a function of frequency can be employed, for example sending angle or amplitude differences in the sub-band.
Although sidechain information is sent on a frame-by-frame basis, or more frequently, it can be useful to interpolate sidechain values across the blocks of a frame. Linear interpolation over time can be used, in the same way that linear interpolation is done on frequencies, as described below.
A suitable implementation of aspects of the present invention employs process steps or devices that implement the respective process steps and are functionally related as set forth below. Although the encoding and decoding steps listed below may be performed by sequences of computer software instructions operating in the order of the steps listed below, it will be understood that equivalent or similar results may be obtained with steps ordered in other ways, having note that certain amounts are obtained from previous ones. For example, multi-chained computer software instruction sequences can be employed such that certain sequences of steps are carried out in parallel. Alternatively, the steps described may be implemented as devices that perform the described functions, the various devices having functions and functional interrelationships such as those described hereinafter.
Coding
The encoder or encoding function can collect an amount of data from a frame before it gets the sidechain information and downmixes the audio channels of the frame to generate a single channel of monophonic (mono) audio. (in the manner of the example of Figure 1, described above), or multiple audio channels (in the manner of the example of Figure 6, described below). By doing so, the sidechain information can be sent first to a decoder, which allows the decoder to begin decoding immediately upon receipt of the mono or multi-channel audio information. The steps of a coding process ("coding steps") can be described as follows. With regard to the encoding steps, reference is made to Figure 4, which is in the nature of a hybrid diagram of a flow chart and a functional block diagram. Through step 419, Figure 4 shows encoding steps for one channel. Steps 420 and 421 apply to all multiple channels that are combined to provide a mono composite signal output or put together in a matrix to provide multiple channels, as described below in relation to the example of Figure 6.
Step 401. Detect transients
to. Perform transient detection of PCM values on an input audio channel.
b. Set a Transient Flag to True if a transient is present in any block of a frame for the channel.
Comments Regarding Step 401
The Transient Flag forms a part of the side chain information and is also used in Step 411, as described below. A transient resolution finer than the decoder's block rate can improve decoder performance. Although, as described above, a
ES 2 324 926 T3
Transients of a block rate instead of a frame rate can form a part of the sidechain information with a modest increase in bit rate, a similar result can be achieved, albeit with reduced spatial precision, without increasing the side chain bit rate, detecting the occurrence of transients in the mono composite signal received at the decoder.
There is one transient flag per channel and per frame, which, as obtained in the time domain, necessarily applies to all sub-bands within that channel. Transient detection can be performed in the same way as used in an AC-3 encoder to control the decision of when to switch between long and short duration blocks of audio, but with higher sensitivity and with a Transient Flagger. True for any frame where the Transient Flag for a block is True (an AC-3 encoder detects transients on a block basis). Although not critical, a sensitivity factor of 0.2 has been found to be a suitable value in a practical embodiment of aspects of the present invention.
As another alternative, transients can be detected in the frequency domain rather than the time domain (see Comments to Step 408). In that case, Step 401 can be skipped and an alternate frequency domain step can be employed, as described below.
Step 402. Window and DFT
Multiply overlapping blocks of PCM time samples by a time window and convert them to complex frequency values, via a DFT as implemented by an FFT.
Step 403. Convert complex values to Magnitude and Angle
Convert each complex value (a + jb) from transform bin to frequency domain and angle representation, using standard complex manipulations.
to. Magnitude = square root of (a<sup>2</sup> + b<sup>2</sup>)
b. Angle = arctg (b / a)
Comments Regarding Step 403
Some of the following Steps use or may alternatively use the energy of a bin, defined as the magnitude before the square (that is, energy = (a<sup>2</sup> + b<sup>2</sup>))
Step 404. Calculate the energy of the sub-band
to. Calculate the energy of the sub-band per block, adding energy values of bins within each sub-band (a sum in the frequencies).
b. Calculate the energy of the sub-band per frame and averaging or accumulating the energy of all blocks in a frame (an average / accumulation over time).
c. If the encoder coupling frequency is below about 1000 Hz, apply the averaged energy in the frame or the accumulated energy in the frame to a time smoother that operates on all subbands below that frequency and above the coupling frequency.
Comments Regarding Step 404c
Time smoothing may be useful to provide interframe smoothing in low frequency subbands. In order to avoid discontinuities that cause artifacts between bin values at the limits of the sub-band, it may be useful to apply progressively decreasing time smoothing, from the sub-band of lower frequencies that encompasses the coupling frequency and through above it (where smoothing can have a significant effect) through a higher frequency sub-band in which the effect of time smoothing can be measured, but it is inaudible, although almost audible. A suitable time constant for the lower frequency range sub-band (where the sub-band is a single bin if the sub-bands are critical bands) may be in the range of 50 to 100 milliseconds, for example. The progressively decreasing time smoothing can continue up through a subband spanning around 1000 Hz, where the time constant can be around 10 milliseconds, for example.
ES 2 324 926 T3
Although a first-order smoother is suitable, the smoother can be a two-stage smoother that has a variable time constant that shortens its attack and decay time in response to a transient. In other words, the steady-state time constant can be scaled according to frequency and can also be variable in response to transients. Alternatively, such smoothing can be applied in Step 412.
Step 405. Calculate the sum of the magnitudes of the bins
to. Calculate the sum per block of the bin magnitudes (Step 403) of each subband (sum across the frequencies).
b. Calculate the sum per frame of the bin magnitudes of each subband by averaging or accumulating the magnitudes from Step 405a across the blocks of a frame (an average / accumulation over time). These sums are used to calculate an Inter-Channel Angle Consistency Factor in Step 410 below.
c. If the encoder coupling frequency is below about 1000 Hz, apply the frame-averaged or frame-accumulated sub-band magnitudes to a time smoother that operates on all sub-bands below that frequency and above the coupling frequency.
Comments regarding Step 405c: See comments regarding Step 404c, except that in the case of Step 405c, time smoothing can alternatively be performed as part of Step 410.
Step 406. Calculate the Relative Phase Angle of the Bins Between Channels
Calculate the relative phase angle between channels of each transform bin of each block, subtracting from the angle of each bin in Step 403, the corresponding bin angle of a reference channel (for example, the first channel). The result, as with other angle additions or subtractions in this specification, is taken in radians of modulus (π, -π) by adding or subtracting 2π until the result is within the desired range from -π to + π.
Step 407. Calculate Sub-band Phase Angle Between Channels
For each channel, calculate a phase angle between channels averaged at frame rate and weighted in amplitude, for each sub-band as follows:
to. For each bin, construct a complex number from the magnitude from Step 403 and the relative phase angle of the bins between channels from Step 406.
b. Add the complex numbers constructed in Step 407a across each subband (one sum across the frequency).
Comment regarding Step 407b: For example, if a subband has two bins and one of the bins has a complex value of 1 + j1 and the other bin has a complex value of 2 + j2, its complex sum is 3 + j3 .
c. Average or accumulate per block the sum of complex numbers for each subband of step 407b, across the blocks of each frame (an average or accumulation over time).
d. If the encoder coupling frequency is below about 1000 Hz, apply the frame-averaged or frame-accumulated complex value of the sub-band to a time smoother that operates on all sub-bands below that. frequency and above the coupling frequency.
Comments regarding Step 407d: See comments regarding Step 404c, except that in the case of Step 407d, time smoothing can alternatively be performed as part of steps 407e or 410.
and. Calculate the magnitude of the complex result from Step 407d as in Step 403.
Comments regarding Step 407e: This magnitude is used in Step 410a below. In the simple example given in step 407b, the magnitude of 3 + j3 is the square root of (9 + 9) = 4.24.
F. Calculate the angle of the complex result as in Step 403.
ES 2 324 926 T3
Comments regarding Step 407f: In the simple example given in Step 407b, the angle of 3 + j3 is arctg (3/3) = 45 degrees = π / 4 radians. This subband angle is time-smoothed depending on the signal (see step 413) time-dependent and quantized (see Step 414) to generate the side chain information of the Sub Angle Control Parameter. -band, as described below.
Step 408. Calculate the Spectral Permanence Factor of the Bin
For each bin, calculate the Spectral Permanence Factor of the Bin in the range 0 to 1, as follows:
to. Let x<sub>m</sub> = bin magnitude of the current block calculated in Step 403.
b. Let y<sub>m</sub> = corresponding magnitude of the bin from the previous block.
c. If x<sub>m</sub> > and<sub>m</sub>, then the Dynamic Amplitude Factor of the bin = (and<sub>m</sub>/ x<sub>m</sub>)<sup>2</sup>,
d. Otherwise, yes and<sub>m</sub> > x<sub>m</sub>, then the Dynamic Amplitude Factor of the bin = (x<sub>m</sub>/Y<sub>m</sub>)<sup>2</sup>,
and. Otherwise, if ym = xm, then the Spectral Permanence Factor of the bin = 1.
Comments Regarding Step 408
"Spectral permanence" is a measure of the extent to which spectral components (ie, spectral coefficients or bin values) change over time. A Bin Spectral Permanence Factor of 1 indicates that there is no change during a given period of time.
Spectral permanence can also be considered as an indicator of whether a transient is present. A transient can cause a sudden rise and fall of the spectral amplitude (bin) in a period of time of one or more blocks, depending on its position with respect to the blocks and their limits. Consequently, a change of the Bin Spectral Permanence Factor from a high value to a low value during a small number of blocks can be considered as an indication of the presence of a transient in a block or blocks having the lower value. An additional confirmation of the presence of a transient, or an alternative to employing the Bin Spectral Permanence Factor, is to observe the phase angles of the bins within the block (for example, at the phase angle output from Step 403) . Because a transient is likely to occupy a single temporal position within a block and have the dominant energy in the block, the existence and position of a transient can be indicated by a substantially uniform delay in the phase from bin to bin in the block, that is, a substantially linear ramp of phase angles as a function of frequency. A further confirmation or alternative is to observe the amplitudes of the bins in a small number of blocks (for example, in the magnitude output of Step 403), that is, looking directly if there is a sudden rise and fall of the spectral level.
Alternatively, Step 408 may look at three consecutive blocks instead of a single block. If the encoder coupling frequency is below about 1000 Hz, Step 408 can watch more than three consecutive blocks. The number of consecutive blocks that can be taken into consideration varies with frequency, such that the number gradually increases as the frequency of the sub-band decreases. If the Bin Spectral Permanence Factor is obtained from more than one block, the detection of a transient, as just described, can be determined by independent steps that respond only to the number of blocks useful to detect transients.
As a further alternative, the energies of the bins can be used instead of the magnitudes of the bins.
As a further alternative, Step 408 may employ an "event decision" detection technique, as described below in the comments following Step 409.
Step 409. Calculate the Spectral Permanence Factor of the Sub-band.
Calculate a Spectral Permanence Factor of the Sub-band on a scale from 0 to 1, forming an amplitude-weighted average of the Spectral Permanence Factor of the Bin, within each sub-band through the blocks of a frame, as follows :
to. For each bin, calculate the product of the Spectral Permanence Factor of the Bin from Step 408 and the magnitude of the bin from Step 403.
b. Add the products within each sub-band (a sum across the frequencies).
c. Average or accumulate the sum of Step 409b over all blocks in a frame (an average / accumulation over time).
ES 2 324 926 T3
d. If the encoder coupling frequency is below about 1000 Hz, apply the frame averaged or frame cumulative sum of a sub-band to a time smoother that operates on all sub-bands below that. frequency and above the coupling frequency.
Comments regarding step 409d: See comments regarding step 404c, except that in the case of Step 409d, there is no suitable subsequent step in which smoothing over time can alternatively be performed.
and. Divide the results of Step 409c or Step 409d, as appropriate, by the sum of the magnitudes of the bins (Step 403) within a subband.
Comment regarding Step 409e: Multiplying by the magnitude from Step 409a and dividing by the sum of the magnitudes from Step 409e provide an amplitude weighting. The output of Step 408 is independent of absolute amplitude and, if not weighted in amplitude, can cause the output of Step 409 to be controlled by very small amplitudes, which is undesirable.
F. Scale the result to obtain the Spectral Permanence Factor of the Sub-band establishing the correspondence map of the range from {0.5 ... 1} to {0 ... 1}. This can be done by multiplying the result by 2, subtracting 1, and limiting results less than zero to a value of 0.
Comment regarding step 409f: Step 409f may be useful to ensure that a noise channel results in a Subband Spectral Permanence Factor of zero.
Comments regarding steps 408 and 409
The goal of Steps 408 and 409 is to measure spectral permanence, changes in spectral composition over time in a sub-band of a channel. Alternatively, aspects of an "event decision" detection may be employed to measure spectral permanence in lieu of the solution just described in relation to Steps 408 and 409. The magnitudes of the complex coefficient of the FFT of each bin are calculated and normalized (the largest magnitude is set to a value of one, for example). Then, the magnitudes of the corresponding bins (in dB) from consecutive blocks are subtracted (ignoring the signs), the differences between bins are added and, if the sum exceeds a threshold, the block limit is considered as a limit of a auditory event. Alternatively, block-to-block changes in amplitude can also be considered in conjunction with spectral magnitude changes (looking at the amount of normalization required).
If aspects of the built-in event detection applications are employed to measure spectral permanence, normalization may not be required and changes in spectral magnitude (changes in amplitude would not be measured if normalization is omitted), are preferably considered. based on a sub-band. Instead of performing Step 408, as indicated above, the decibel differences of the spectral magnitude between the corresponding bins of each sub-band can be summed, in accordance with the teachings of such applications. Afterwards, each of those sums, which represent the degree of spectral change from block to block, can be scaled so that the result is the spectral permanence factor that has a range from 0 to 1, where the value 1 indicates the permanence higher, a 0 dB change from block to block for a given bin. A value of 0, indicating the lowest dwell, can be assigned to decibel changes equal to or greater than a suitable amount, such as 12 dB, for example. These results, a Bin Spectral Permanence Factor, can be used by Step 409 in the same way that Step 409 uses the results from Step 408, as described above. When Step 409 receives a Bin Spectral Permanence Factor obtained by employing the alternative event decision detection technique just described, the Sub-band Spectral Permanence Factor from Step 409 can also be used. as an indicator of a transient. For example, if the range of values produced by Step 409 is 0 to 1, a transient can be considered to be present when the Spectral Permanence Factor of the Sub-band is a small value, such as, for example, 0, 1, which indicates a substantial lack of spectral permanence.
It will be appreciated that the Bin Spectral Permanence Factor produced by Step 408 and by the alternative to Step 408 just described, each inherently provides a variable threshold to some extent, insofar as they are based on relative changes. from block to block. Optionally, it may be useful to supplement such inherence by specifically providing a threshold variation in response to, for example, multiple transients in a frame or a large transient between smaller transients (for example, a high transient proceeding from a mid-level clap down). In the case of the last example, an event listener may initially identify each clap as an event, but a high transient (for example, the beat on a drum) may be desirable to change the threshold, so that only the beat on the drum is identified as an event.
ES 2 324 926 T3
Alternatively, a metric of randomness can be employed, rather than a measure of spectral permanence over time.
Step 410. Calculate the Consistency Factor of the Angle Between Channels
For each sub-band that has more than one bin, calculate the Angle Consistency Factor between Channels as follows:
to. Divide the magnitude of the complex sum from Step 407e by the sum of the magnitudes from Step 405. The resulting “raw” Angle Consistency Factor is a number in the range 0 to 1.
b. Calculate a correction factor: let n = the number of values in the sub-band that contribute to the two quantities from the previous step (in other words, “n” is the number of bins in the sub-band). If n is less than 2, let the Angle Consistency Factor be 1 and go to Steps 411 and 413.
c. Let r = Expected Random Variation = 1 / n. Subtract r from the result of Step 410b.
d. Normalize the result of Step 410c by dividing by (1-r). The result has a maximum value of 1. Limit the minimum value to 0, as necessary.
Comments Regarding Step 410
The Angle Consistency Between Channels is a measure of the similarity of the phase angles between channels within a sub-band in the period of a frame. If all the bin angles between channels of the sub-band are equal, the Angle Consistency Factor is 1.0; whereas, if the angles between channels are randomly distributed, the value approaches zero.
The Subband Angle Consistency Factor indicates if there is a ghost image between the channels. If the consistency is low, then it is desirable to de-correlate the channels. A high value indicates a merged image. The combination of images is independent of other characteristics of the signal.
It will be noted that the Sub-band Angle Consistency Factor, although it is an angle parameter, is determined indirectly from two quantities. If the angles between channels are equal, adding the complex values and then taking the magnitude gives the same result as taking all the magnitudes and adding them, so that the quotient is 1. If the angles between channels are scattered, adding the complex values (for example adding vectors with different angles) results in at least partial cancellation, so that the magnitude of the sum is less than the sum of the magnitudes, and the quotient is less than 1.
Here is a simple example of a sub-band that has two bins:
Suppose the two complex values of the bins are (3 + j4) and (6 + j8). (The same angle in each case: angle equal to arctg (imag / real), so that angle1 = arctg (4/3) and angle2 = arctg (8/6) = arctg (4/3)). Adding the complex values, sum = (9 + j12), a magnitude whose square root of (81 + 144) is = 15.
The sum of the magnitudes is the magnitude of (3 + j4) + magnitude of (6 + j8) = 5 + 10 = 15. The quotient is therefore 15/15 = 1 = consistency (before normalization 1 / n would also be 1 after normalization) (Normalized Consistency = (1- 0.5) / (1 - 0.5) = 1, 0 =.
If one of the previous bins had a different angle, for example the second bin has a complex value (6 - j8), which has the same magnitude 10. The complex sum is now (9 - j4), which has a magnitude equal to the square root of (81 + 16) = 9.85, so the quotient is 9.85 / 15 = 0.66 = consistency (before normalization). To normalize, subtract 1 / n = 1/2 and divide by (1 - 1 / n) (normalized consistency = (0.66 - 0.5) / (1 - 0.5) = 0.32).
Although the technique described above for determining the Subband Angle Consistency Factor has been found useful, its use is not critical. Other suitable techniques may be employed. For example, a standard deviation of the angles could be calculated using standard formulas. In either case, it is desirable to employ amplitude weighting to minimize the effect of small signals on the calculated consistency value.
Also, an alternative derivation of the Subband Angle Consistency Factor may use energy (the squares of the magnitudes) instead of the magnitude. This can be accomplished by squaring the magnitude of Step 403 before being applied to Steps 405 and 407.
ES 2 324 926 T3
Step 411. Obtaining the Decorelation Scale Factor of the Sub-band
Obtain a Decorelation Scale Factor of the frame rate for each sub-band, as follows:
to. Let x = Spectral Permanence Factor of the frame rate from Step 409f.
b. Let y = Angle Consistency Factor of the screen rate from Step 410e.
c. So, the Frame Rate Sub-Band Decorrelation Scale Factor = (1 - x) * (1 - y), a number between 0 and 1.
Comments regarding Step 411
The Subband Decorrelation Scale Factor is a function of the spectral permanence of the signal characteristics over time, in a sub-band of a channel (the Spectral Permanence Factor) and the consistency in the same sub -band of a channel of bin angles, with respect to the corresponding bins of a reference channel (the Consistency Factor of the Angle between Channels). The Subband Décorrelation Scale Factor is high only if both the Spectral Permanence Factor and the Angle Consistency Factor between Channels are low.
As explained above, the Décorrelation Scale Factor controls the degree of décorrelation of the envelope provided in the decoder. Signals that exhibit spectral permanence over time should preferably not be décorrelated by altering their envelopes, regardless of what happens on other channels, because it can result in audible artifacts, that is, signal fluctuations or gurgling.
Step 412. Obtaining the Sub-band Amplitude Scale Factors
From the values of the frame energy of the sub-band of Step 404, and from the values of the energy of the sub-band frame of all other channels (as can be obtained by a step corresponding to the Step 404 or equivalent thereof), obtain the Sub-band Amplitude Scale factors of the frame rate, as follows:
to. For each sub-band, add the energy values per frame, in all the input channels.
b. Divide each subband energy value (from Step 404) by the sum of the energy values in all input channels (from Step 412a), to create values in the range 0 to 1.
c. Convert each ratio to dB, in the range of - to 0.
d. Divide by the level of detail of the scale factor, which can be fixed at 1.5 dB, for example, change the sign to obtain a non-negative value, limit to a maximum value that can be, for example, 31 (i.e. , 5-bit precision) and round to the nearest integer to create the quantized value. These values are the Sub-band Amplitude Scale Factors in the frame rate, and are carried as part of the side chain information.
and. If the encoder coupling frequency is below about 1000 Hz, apply the frame-averaged or frame-accumulated sub-band magnitudes to a time smoother that operates on all sub-bands below that frequency and above the coupling frequency.
Comments regarding step 412e: See comments regarding step 404c, except that in the case of Step 412e, there is no suitable subsequent step in which time smoothing can alternatively be performed.
Comments for Step 412
While the level of detail (resolution) and precision of quantization reported herein have been found to be helpful, they are not critical and there are other values that can provide acceptable results.
Alternatively, amplitude can be used instead of energy to generate the Amplitude Scale Factors of the Sub-band. If you use the amplitude, you would use dB = 20 * log (ratio of amplitudes), otherwise using energy, it is converted to dB by means of dB = 10 * log (ratio of energies), where the ratio of amplitudes = square root (energy ratio).
ES 2 324 926 T3
Step 413. Sub-band Phase Angles between Channels with Time Smoothing, depending on the Signal
Apply signal-dependent temporal smoothing to the inter-channel angles at the frame rate of a sub-band, obtained in Step 407f:
to. Let v = Spectral Permanence Factor of the Sub-band from Step 409d.
b. Let w = corresponding Angle Consistency Factor from Step 410e.
c. Let x = (1 - v) * w. This is a value between 0 and 1, which is high if the Spectral Permanence Factor is low and the Angle Consistency Factor is high.
d. Let y = 1 - x. y is high if the Spectral Permanence Factor is high and if the Angle Consistency Factor is low.
and. Let z = y<sup>exp</sup>, where exp is a constant, which can be = 0.1. z is also in the range from 0 to 1, but with a trend towards 1, corresponding to a slow time constant.
F. If the Transient Flag (Step 401) for the channel is on, set z = 0, corresponding to a fast time constant in the presence of a transient.
g. Calculate lim, a maximum allowable value of z, lim = 1 - (0.1 * w). This has a travel from 0.9 if the Angle Consistency factor is high, to 1.0 if the Angle Consistency Factor is low (0).
h. Limit z by lim as necessary: if (z> lim), then z = lim.
i. Smooth the angle of the sub-band from Step 407f using the value of z and a current smoothed value of the angle held for each sub-band. If A = angle from Step 407f and RSA = value of the smoothed current angle as in the previous block, and NewRSA is the new value of the smoothed current angle, then: NewRSA = RSA * z + A * (1-z). The value of RSA is subsequently set equal to NewRSA before processing the next block. The new RSA is the time-dependent smoothed angle output from the signal from Step 413.
Comments Regarding Step 413
When a transient is detected, the sub-band angle update time constant is set to 0, allowing a rapid change of the sub-band angle. This is desirable because it allows the normal angle update mechanism to use a relatively slow range of time constants, minimizing errant images during static or near-static signals, but rapidly changing signals are dealt with with fast time constants.
Although other smoothing parameters and techniques can be used, a Step 413 implementation of a first order smoother has been found to be suitable. If implemented as a first-order / low-pass smoother filter, the variable “z” corresponds to the forward feedback coefficient (sometimes referred to as “ff0”), while “(1-z)” corresponds to a backward feedback coefficient (sometimes referred to as "fb1").
Step 414. Quantify Subband Phase Angles Between Smoothed Channels
Quantify the sub-band phase angles between time-smoothed channels, obtained in Step 413i, to obtain the Sub-band Angle Control Parameter:
to. If the value is less than 0, add 2π, so that all the angle values to be quantized are in the range from 0 to 2π.
b. Divide by the level of detail (resolution) of the angle, which can be 2π / 64 radians, and round to an integer. The maximum value can be set to 63, corresponding to the 6-bit quantization.
Comments Regarding Step 414
The quantized value is treated as a non-negative integer, so an easy way to quantify the angle is to map the angle to a non-negative floating point number ((add 2π if it is less than 0, making the range is 0 to (less than) 2n)), scale to the level of detail (resolution) and round to an integer. Similarly, de-quantize that integer (which could be done differently with a simple query to
ES 2 324 926 T3 a table), it can be achieved by scaling by the inverse of the angle detail factor, converting a non-negative integer to a non-negative floating point angle (again, the range is 0 to 2π), after which it can be normalized back to the ± π range for further use. Although such quantization of the Subband Angle Control Parameter has been found useful, such quantization is not critical and other quantifications may provide acceptable results.
Step 415. Quantify Sub-band Décorrelation Scale Factors
Quantify the Sub-band Décorrelation Scale Factors produced by Step 411, for example to 8 levels (3 bits), multiplying by 7.49 and rounding to the nearest integer. These quantized values are part of the side chain information.
Comments Regarding Step 415
Although such quantification of the Subband Décorrelation Scale Factors has been found useful, quantification using example values is not critical and other quantifications may provide acceptable results.
Step 416: De-quantize the Subband Angle Control Parameters
De-quantize the Sub-band Angle Control Parameters (see Step 414), to be used before downmixing.
Comment Regarding Step 416
The use of quantized values in the encoder helps to maintain synchronization between the encoder and the decoder.
Step 417. Distribute the Dequantized Sub-band Angle Control Parameters at the Frame rate, through the Blocks
In preparing for downmix, distribute the dequantized Subband Angle Control Parameters once per frame from Step 416, over time, to the subbands of each block within the frame.
Comment to Step 417
The same frame value can be assigned to each block in the frame. Alternatively, it may be useful to interpolate the Subband Angle Control Parameter values in the blocks of a frame. Linear interpolation over time can be used, in the way that linear interpolation is done on frequencies, as described below.
Step 418. Interpolate Block Subband Angle Control Parameters into Bins
Distribute the Block Subband Angle Control Parameters from Step 417, for each of the channels across the frequency in the bins, preferably using linear interpolation, as described below.
Comment Regarding Step 418
If linear frequency interpolation is employed, Step 418 minimizes bin-to-bin phase angle shifts across a subband boundary, thus minimizing spectral doubling artifacts. Such linear interpolation can be enabled, for example, as described below after the description of Step 422. The sub-band angles are calculated independently of one another, each representing an average across a sub-band. band. Thus, there may be a large change from one sub-band to the next. If the net value of the angle for a sub-band is applied to all the bins of the sub-band (a “rectangular” distribution of the sub-band), all the phase shift from one sub-band to the neighboring sub-band has place between two bins. If there is a strong signal component there, there may be severe spectral doubling, possibly audible. Linear interpolation between the centers of each sub-band, for example, extends the phase angle shift in all
ES 2 324 926 T3 the bins of the sub-band, minimizing the change between any pair of bins, so that, for example, the angle at the lower end of a sub-band coincides with the angle at the upper end of the sub-band that is below it, while keeping the global average the same as that of the given calculated sub-band angle. In other words, instead of rectangular sub-band distributions, the sub-band angle distribution can be trapezoidal in shape.
For example, suppose the bottom coupled subband has a bin and a subband angle of 20 degrees, the next subband has three bins and a subband angle of 40 degrees, and the third subband it has five bins and a 100 degree subband angle. Without any interpolation, suppose that the first bin (a sub-band) varies by an angle of 20 degrees, the next three bins (another sub-band) vary by an angle of 40 degrees, and the next five bins (a sub-band) additional) are varied at an angle of 100 degrees. In that example, there is a maximum change of 60 degrees, from bin 4 to bin 5. With linear interpolation, the first bin continues to vary by an angle of 20 degrees, the next three bins vary by around 30, 40, and 50 degrees; and the next five bins vary by approximately 67, 83, 100, 117, and 133 degrees. The mean variation of the angle of the subband is the same, but the maximum change from bin to bin is reduced to 17 degrees.
Optionally, changes in amplitude from subband to subband, relative to this and other steps described herein, such as Step 417, can also be handled in a similar interpolation manner. However, it may not be necessary to do so, because it tends to be a more natural continuity in amplitude from one subband to the next.
Step 419. Apply the Phase Angle Rotation to the Bins Transformation Values for the Channel
Apply phase angle rotation to each bin transform value, as follows:
to. Let x = bin angle for this bin, as calculated in Step 418.
b. Let y = - x;
c. Calculate z, a complex scale factor of the phase rotation of unit magnitude with angle y, z = cos (y) + j sin (y)
d. Multiply the value of the bin (a + jb) by z.
Comments Regarding Step 419
The phase angle rotation applied to the encoder is the inverse of the angle obtained from the Sub-band Angle Control Parameter.
Phase angle adjustments, as described herein, in an encoder or encoding process, prior to downmix (Step 420), have several advantages: (1) minimize channel cancellations that are added to a composite mono signal or in multiple channel arrays, (2) minimize dependence on energy normalization (Step 421), and (3) perform a pre-compensation of reverse rotation of the decoder phase angle, thereby reducing spectral doubling.
The phase correction factors can be applied in the encoder, by subtracting each sub-band phase correction value from the angles of each bin transform value in that sub-band. This is equivalent to multiplying each complex value of bin by a complex number with a magnitude of 1.0 and an angle equal to the negative value of the phase correction factor. Note that a complex number of magnitude 1, angle A, is equal to cos (A) + j sin (A). This last quantity is calculated once for each sub-band of each channel, where A = -phase correction for that sub-band, then multiplied by each complex value of the bin signal to materialize the phase-shifted bin value.
The phase variation is circular, resulting in a circular convolution (as mentioned above). Although circular convolution can be benign for some continuous signals, it can create spurious spectral components for certain complex continuous signals (such as a piccolo) or can cause transients to blur if different phase angles are used for different subbands. Consequently, a suitable technique or the Transient Flagger can be employed to avoid circular convolution so that, for example, when the Transient Flagger is True, the results of the angle calculation can be substituted and all the sub-bands of a channel can use the same phase correction factor such as zero, or a random value.
ES 2 324 926 T3
Step 420. Down Mix
Downmix to mono, summing the corresponding complex transform bins across the channels, to produce a composite mono channel, or downmix for multiple channels producing a matrix with the input channels, for example in the manner of the example of Figure 6, as described below.
Comments Regarding Step 420
In the encoder, once the transform bins of all the channels have been varied in phase, the channels are added, bin by bin, to create the mono composite audio signal. Alternatively, the channels can be applied to an active or passive matrix that provides a simple sum to one channel, as in the N: 1 encoding of Figure 1, or to multiple channels. The matrix coefficients can be real or complex (real and imaginary).
Step 421. Normalize
To avoid cancellation of isolated bins and an overemphasis on in-phase signals, normalize the amplitude of each bin of the composite mono channel to have substantially the same energy as the sum of the contributing energies, as follows:
to. Let x = the sum through the channels of the energies of the bins (that is, the squares of the magnitudes of the bins calculated in Step 403).
b. Let y = energy of the corresponding bin of the composite mono channel, calculated according to Step 403.
c. Let z = scale factor = square root (x / y). If x = 0, then y is 0 and z is set to 1.
d. Limit z to a maximum value of, for example, 100. If z is initially greater than 100 (implying strong cancellation from downmix), add an arbitrary value, for example, 0.01 * square root (x ) to the real and imaginary parts of the compound mono bin, which will ensure that it is large enough to be normalized by the next step.
and. Multiply the value of the complex mono compound bin by z.
Comments Regarding Step 421
Although it is generally desirable to use the same phase factors for both encoding and decoding, even the optimal choice of sub-band phase correction value can result in one or more audible spectral components within the sub-band, for be canceled during the encoding downmix process, because the phase shift in step 419 is performed on a subband rather than a bin basis. In this case, a different phase factor can be used for isolated bins in the encoder, if it is detected that the sum of the energies of such bins is much less than the sum of energies of the individual channel bins at that frequency. Generally, it is not necessary to apply such an isolated correction factor to the decoder, considering that isolated bins normally have very little effect on the overall image quality. A similar normalization can be applied if multiple channels are used instead of a mono channel.
Step 422. Assemble and Package in Bit Chain (s)
The side chain information of the Amplitude Scale Factors, Angle Control Parameters, Decorrelation Scale Factors, and Transient Flags for each channel, along with the composite common mono audio or multiple channels in matrices, are multiplexed as is desired and packed into one or more bit strings suitable for storage, transmission or storage and transmission medium (s).
Comment Regarding Step 422
Mono composite audio or multi-channel audio can be applied to a data rate reduction encoding process or device, such as, for example, a perceptual encoder or a perceptual encoder and an entropy encoder (for example, an arithmetic or Huffman encoder) (sometimes referred to as a "lossless" encoder) before being packaged. Furthermore, as mentioned above, composite mono audio (or multi-channel audio) and related side chain information can be obtained from multiple input channels only for audio frequencies above a certain frequency (frequency of "coupling"). In that case, the audio frequencies below the coupling frequency on each of the multiple input channels can be stored, transmitted, or stored and transmitted as
ES 2 324 926 T3 discrete channels, or they can be combined or processed in some way other than that described in this specification. Discrete or otherwise combined channels may also be applied to a data reduction encoding process or device, such as, for example, a perceptual encoder or a perceptual encoder and an entropy encoder. Composite mono audio (or multichannel audio) and discrete multichannel audio can all be applied to an integrated perceptual encoding or a perceptual and entropy encoding process or device before being packaged.
Optional Interpolation Flag (Not illustrated in figure 4)
Interpolation across the frequency of the basic phase angle variations, provided by the Sub-band Angle Control Parameters, can be enabled at the Encoder (Step 418) and / or the Decoder (Step 505, plus go ahead). The optional sidechain parameter of the Interpolation Flag may be used to enable interpolation at the Decoder. The Interpolation Flag or an enable flag similar to the Interpolation Flag can be used in the Encoder. Note that because the Encoder has access to the data at the bin level, it can use different interpolation values than the Decoder, which interpolates the Subband Angle Control Parameters into the side chain information.
The use of such interpolation across the frequency in the Encoder or Decoder can be enabled if, for example, one of the following two conditions is true:
Condition 1. If a strong and isolated spectral peak is located at or near the boundary of two subbands that have substantially different phase rotation angle assignments.
Reason: Without interpolation, a large phase shift at the boundary can introduce a gurgle in the isolated spectral component. By using interpolation to spread the band-to-band phase shift through the bin values within the band, the amount of change at the sub-band boundaries is reduced. Thresholds for spectral peak strength, closeness to a limit, and difference in phase rotation from sub-band to sub-band can be empirically adjusted to satisfy this condition.
Condition 2. If, depending on the presence of a transient, the channel phase angles (non-transient) between channels, or the absolute phase angles within a channel (transient), comprise a good adaptation to a linear progression.
Reason: Using interpolation to reconstruct the data tends to provide a better fit to the original data. Note that the slope of the linear progression need not be constant across all frequencies, but only within each sub-band, since the angle data will continue to be transported to the decoder on a sub-band basis; and that forms the input to Interpolator 418. The degree to which the data provide a good fit to satisfy this condition can also be determined empirically.
Other conditions, such as those determined empirically, can benefit from frequency interpolation. The existence of the two conditions just mentioned can be determined as follows:
Condition 1: If a strong and isolated spectral peak is located at or near the boundary of two subbands that have substantially different phase rotation angle assignments:
for the Interpolation Flag to be used by the decoder, the Subband Angle Control Parameters (output from Step 414), and to enable Step 418 within the Encoder, the output from Step 413 before quantization can be used to determine the angle of rotation from sub-band to sub-band.
For both the Interpolation Flag and the enable within the Encoder, the magnitude output from Step 403, the current DFT magnitudes, can be used to find isolated peaks at the subband boundaries.
Condition 2. Depending on the presence of a transient, if the phase angles (non-transient) between channels or the absolute phase angles within a channel (transient), comprise a good adaptation to a linear progression:
if the Transient Flag is not true (not transient), use the relative phase angles of bins between channels from Step 406, to adapt it to a determination of linear progression, and if the Transient Flag is true (transient), use the absolute phase angles of the channel from Step 403.
ES 2 324 926 T3
Decoding
The steps in the decoding process ("decoding steps") can be described as follows. With regard to the decoding steps, reference is made to Figure 5, which is of a hybrid flow diagram and functional block diagram nature. For simplicity, the figure shows the obtaining of the side chain information components for a channel, it being understood that the side chain information components must be obtained for each channel, unless the channel is the reference channel for such components, as explained elsewhere.
Step 501. Unpack and Decode the Sidechain Information
Unpack and Decode (including dequantization), as necessary, the components of the sidechain data (Amplitude Scale Factors, Angle Control Parameters, Decorelation Scale Factors and Transient Flag) for each frame of each channel (one channel is illustrated in Figure 5). Table queries can be used to decode Amplitude Scale Factors, Angle Control Parameter, and Decorelation Scale Factors.
Comment regarding Step 501: As explained above, if a reference channel is used, the side chain data for the reference channel may not include the Angle Control Parameters, Decorrelation Scale Factors, and Transient Signaling.
Step 502. Unpack and Decode the Mono Composite or Multichannel Audio Signal
Unpack and decode, as necessary, the information from the composite mono or multichannel audio signal, to provide the DFT coefficients for each transform bin of the mono or multichannel composite audio signal.
Comment Regarding Step 502
Step 501 and Step 502 can be considered part of a single unpacking and decoding step. Step 502 can include a passive or active matrix.
Step 503. Distribute Angle Parameter Values Across the Blocks
The values of the Sub-band Angle Control Parameter of the block are obtained from the values of the Sub-band Angle Control Parameter of the dequantized frame.
Comment Regarding Step 503
Step 503 can be implemented by distributing the same parameter value for each block in the frame.
Step 504. Distribute the Subband Décorrelation Scale Factor across the Blocks
The values of the Sub-band Décorrelation Scale Factor are obtained from the values of the Sub-band Décorrelation Scale Factor of the dequantized plot.
Comment Regarding Step 504
Step 504 can be implemented by distributing the same scale factor value to each block in the frame.
Step 505. Interpolate Linearly Across Frequency
Optionally, obtain the angles of the bins from the angles of the sub-band of the blocks, of Step 503 of the decoder, by means of linear interpolation through the frequency, as described above in relation to Step 418 of the encoder . The linear interpolation from Step 505 can be enabled when the Interpolation Flag is used and is True.
ES 2 324 926 T3
Step 506. Add Random Phase Angle Compensation (Technique 3)
According to Technique 3, described above, when the Transient Flag indicates a transient, add to the Block Subband Angle Control Parameter provided by Step 503, which may have been linearly interpolated in frequency by Step 505, a random offset value scaled by the Décorrelation Scaling Factor (scaling can be indirect, as set out in this Step):
to. Let y = Decorrelation Scale Factor of the Sub-band of the block.
b. Let z = y<sup>exp</sup>, where exp is a constant, for example = 5. z will also be in the range of 0 to 1, but with a trend towards 0, reflecting a propensity for low levels of random variation, unless the Scale Factor value of Descorrelation is high.
c. Let x = a random number between +1.0 and 1.0, independently chosen for each sub-band of each block.
d. Then, the value added to the Block Sub-band Angle Control Parameter, to add a random angle compensation value, according to Technique 3, is x * pi * z.
Comments Regarding Step 506
As will be appreciated by those of ordinary skill in the art, the "random" angles (or "random amplitudes" if the amplitudes also scale) to convert the scale by means of the Décorrelation Scale Factor, may include not only pseudo-random variations and truly random, but also deterministically generated variations which, when applied to phase angles or phase angles and amplitudes, they have the effect of reducing the cross-correlation between channels. Such "random" variations can be obtained in various ways. For example, a pseudo-random number generator with various root values can be used. Alternatively, truly random numbers can be generated using a hardware random number generator. Considering that as random angle resolution an angle resolution of only about one degree may be sufficient, tables of random numbers having two or three decimal places (for example, 0.84 or 0.844) can be used. Preferably, the random values (between -1.0 and +1.0 with reference to Step 505c, above) are statistically uniformly distributed throughout each channel.
Although the indirect non-linear scale conversion of Step 506 has been found useful, it is not critical and other scale conversions may be employed, in particular other values for the exponent may be employed, in order to obtain similar results.
When the Sub-band Décorrelation Scale Factor is 1, a full range of random angles from -π to + π is added (in which case the Sub-band Angle Control Parameter values generated by Step 503, they become irrelevant). As the Subband Decorrelation Scale factor value decreases toward zero, the random angle compensation also decreases toward zero, causing the output of Step 506 to shift toward the values of the Angle Control Parameter of the Subband. Sub-band generated by Step 503.
If desired, the encoder described above can also add a scaled random offset according to Technique 3 to the angle shift applied to the channel prior to downmix. By doing so, the cancellation of spectral doubling in the decoder can be improved. It can also be beneficial in improving encoder and decoder timing.
Step 507. Add the Random Phase Angle Compensation (Technique 2)
According to Technique 2, described above, when the Transient Flag does not indicate a transient, for each bin, add to all the Sub-band Angle Control Parameters in a frame provided by Step 503 (Step 505 works only when the Transient Flag indicates a transient), a different random offset value scaled by the Décorrelation Scaling Factor (scaling can be direct, as set in this step):
to. Let y = Decorrelation Scale Factor of the Sub-band.
b. Let x = random number between +1.0 and -1.0, independently chosen for each bin in each frame.
c. Then, the value added to the Block's Bin Angle Control Parameter, to be added to the random offset value of the angle, according to Technique 3 is x * pi * y.
ES 2 324 926 T3
Comments Regarding Step 507
See comments above regarding Step 505, regarding random angle compensation.
Although the direct scale conversion in Step 507 has been found helpful, it is not critical and other scale conversions can be used.
To minimize temporal discontinuities, the unique random value of the angle for each bin in each channel preferably does not change over time. The random values of the angle of all the bins in a sub-band are scaled by the same value of the Sub-band Décorrelation Scaling Factor, which updates at the frame rate. Thus, when the value of the Sub-band Décorrelation Scale Factor is 1, a full range of random angles from -π to + π is added (in which case the block sub-band angle values, obtained from from the values of the angle of the sub-band of the dequantized frame, they become irrelevant). As the Subband Décorrelation Scale Factor value decreases toward zero, the random angle offset also decreases toward zero. Unlike in Step 504, the scale conversion in this Step 507 can be a direct function of the value of the Sub-band Décorrelation Scale Factor. For example, a Subband Décorrelation Scale Factor value of 0.5 proportionally reduces each random angle variation by 0.5.
The random value of the scaled angle can then be added to the bin angle from Step 506 of the decoder. The Décorrelation Scale Factor value is updated once per frame. In the presence of a Transient Flagger for the frame, this step is skipped to avoid transient pre-noise artifacts.
If desired, the encoder described above can also add scaled random compensation, according to Technique 2, to the angle variation applied prior to downmix. Doing so can improve the cancellation of spectral doubling in the decoder. It can also be beneficial in improving encoder and decoder timing.
Step 508. Normalize Amplitude Scale Factors
Normalize Amplitude Scale Factors across the channels so that they add up to 1 when squared.
Comment Regarding Step 508
For example, if two channels have unquantized scale factors of -3.0 dB (= 2 * 1.5 dB Level of Detail), (0.70795) the sum of squares is 1.002. Dividing each one by the square root of 1.002 = 1.001 gives two values of 0.7072 (-3.01 dB).
Step 509. Raise Sub-band Scale Factor Levels (Optional)
Optionally, when the Transient Flag indicates that there is no transient, apply a slight additional elevation to the Sub-band Scale Factor levels, depending on the Sub-band Décorrelation Scale Factor Levels: multiply each Factor Subband Amplitude Scale Scale normalized by a small factor (eg, 1 + 0.2 * Subband Décorrelation Scale Factor). When the Transient Flag is True, skip this step.
Comment Regarding Step 509
This step can be useful because decoder decoder step 507 can result in slightly reduced levels in the final reverse filter bank processing.
Step 510. Distribute Subband Width Values Across the Bins
Step 510 can be implemented by distributing the same value of the subband amplitude scale factor for each bin in the subband.
Step 510a. Add Random Amplitude Compensation (Optional)
Optionally, apply a random variation to the normalized Sub-band Décorrelation Scale Factor, depending on the levels of the Sub-band Décorrelation Scale Factor and the Transient Flag. In the absence of a transient, add a Random Amplitude Scale Factor that does not change over time on a bin-by-bin basis (different from bin to bin) and, in the presence of a transient (in the frame or block), add a Factor of
ES 2 324 926 T3
Random Amplitude Scale that changes on a block-by-block basis (different from block to block) and changes from sub-band to sub-band (same variation for all bins in a sub-band; different from sub-band to subband). Step 510a is not illustrated in the drawings.
Comment Regarding Step 510a
Although the degree to which random amplitude variations are added can be controlled by the Décorrelation Scale Factor, it is believed that a particular value of the scale factor would result in less amplitude variation than the corresponding random phase variation resulting from the same value of the scale factor, in order to avoid audible artifacts.
Step 511. Up Mix
to. For each bin of each output channel, construct a complex upmix scale factor from the amplitude in Step 508 of the decoder and the angle of the bin in Step 507 of the decoder: (amplitude * (cos (angle) + j sin (angle)).
b. For each output channel, multiply the complex value of the bin and the complex scale factor of the upmix to generate a complex value of the output bin of the upmix of each bin in the channel.
Step 512. Perform the inverse DFT (optional)
Optionally, perform an inverse DFT transform on the bins of each output channel, to obtain multi-channel PCM output values. As is well known, relative to such an inverse DFT transform, the individual blocks of the time samples are windowed, and the contiguous blocks are overlapped and added together in order to reconstruct the final output PCM audio signal in continuous time.
Comments Regarding Step 512
A decoder in accordance with the present invention may not provide PCM outputs. In the case where the decoder process is employed only above a given coupling frequency, discrete MDCT coefficients are sent for each channel below that frequency, it may be desirable to convert the DFT coefficients obtained by upmix steps 511a and 511b of the decoder to the MDCT coefficients, so that they can be combined with the lower frequency discrete MDCT coefficients and re-quantized in order to provide, for example, a bit string compatible with an encoding system that has a large number of installed users, such as the bit string of the AC-3 SP / DIF standard for application to an external device on which a inverse transform. An inverse DFT transform can be applied to some of the output channels to provide PCM outputs.
Section 8.2.2 of Document A / 52A with a Sensitivity Factor “F” added
8.2.2 Transient Detection
Transients are detected on channels with full bandwidth, in order to decide when to switch to short audio blocks, to improve pre-echo performance. High-pass filtered versions of the signals are examined for an increase in energy from one time segment of one sub-block to the next. Sub-blocks are examined on different time scales. If a transient is detected in the second half of an audio block on a channel, that channel switches to a short block. A channel that is block-switched uses the D-45 exponent strategy (that is, the data has a less approximate frequency resolution in order to reduce the data overhead resulting from increased temporal resolution).
The transient detector is used to determine when to switch from a long transform block (length 512) to the short block (length 256). It operates on 512 samples for each audio block. This is done in two passes, where each pass processes 256 samples. Transient detection is broken down into four steps: 1) high-pass filtering, 2) segmentation of the block into submultiples, 3) detection of peak amplitudes within each sub-block segment, and 4) threshold comparison. The transient detector outputs a blksw [n] flag for each full bandwidth channel, which when set to "one" indicates the presence of a transient in the second half of the input block of length 512 for the corresponding channel.
1) High-pass filtering: The high-pass filter is implemented as a two-quadruple direct form II cascade IIR filter, with a cutoff of 8 kHz.
ES 2 324 926 T3
2) Block segmentation: The block of 256 high-pass filtered samples is segmented into a hierarchical tree of levels, in which level 1 represents the block of length 256, level 2 is two segments of length 128, and level 3 are four segments of length 64.
3) Peak detection: The sample with the largest magnitude is identified for each segment at each level of the hierarchical tree. The peaks for a single level are found as follows
P [j] [k] = max (x (n)) for n = (512 x (k-1) / 2<sup>TO</sup>j), (512 x (k-1) / 2<sup>TO</sup>j) + 1, ... (512 xk / 2<sup>TO</sup>j) - 1 and k = 1, ..., 2<sup>TO</sup>(j-1);
where: x (n) = the nth sample of the block of length 256 j = 1, 2, 3 is the hierarchical level number k = number of the segment within level j
Note that P [j] [0], (that is, k = 0) is defined to be the peak of the last segment at level j of the tree, calculated immediately before the current tree. For example, P [3] [4] in the preceding tree is P [3] [0] in the current tree.
4) Comparison of thresholds: The first stage of the threshold buyer checks if there is a significant signal level in the current block. This is done by comparing the global value of the peak P [1] [1] of the current block with a “silence threshold”. If P [1] [1] is below this threshold, then a long block is forced. The value of the silence threshold is 100/32768. The next stage of the comparator checks the relative peak levels of contiguous segments at each level of the hierarchical tree. If the peak ratio of two contiguous segments on a particular level exceeds a predefined threshold for that level, a flag is set to indicate the presence of a transient in the current 256-long block. The relationships are compared as follows:
Mag (P [j] [k) x T [j]> (F * mag (P [j] [(k-1)])) [Note the sensitivity factor “F”] where: T [j] is the predefined threshold for level j, defined as:
T [1] = 0.1
T [2] = 0.075
T [3] = 0.05
If this inequality is true for any two peaks of the segments at any level, then a transient is indicated for the first half of the input block of length 512. The second pass of this process determines the presence of transients in the second half of the block. input length 512.
Coding N: M
There are aspects of the present invention that are not limited to N: 1 encoding as described in relation to Figure 1. More generally, there are aspects of the present invention applicable to the transformation of any number of input channels (n channels input) to any number of output channels (m output channels) in the manner of Figure 6 (ie, N: M encoding).
Because in many common applications, the number of input channels n is greater than the number of output channels m, the N: M encoding configuration of Figure 6 will be referred to as "downmix" for convenience of description.
Referring to the details of figure 6, instead of adding the outputs of the Angle Rotation 8 and the Angle Rotation 10 of the additive combiner 6, as in the configuration of figure 1, those outputs can be applied to a device or 6 'matrix downmix function ("Matrix Downmix"). The Downmix Matrix 6 'can be a passive or active matrix, providing a simple summation to one channel, as in the N: 1 encoding of Figure 1, or to multiple channels. The matrix coefficients can be real or complex (real and imaginary). Other devices and functions of Figure 6 may be the same as in the configuration of Figure 1, and have the same reference numerals.
The Down Mix Matrix 6 'may provide a frequency dependent hybrid function, such that it provides, for example, m<sub>n</sub>_<sub>AND</sub> channels in the frequency range f1 to f2 and m<sub>AND</sub>_<sub>f3</sub> channels in the range
ES 2 324 926 T3 frequencies from f2 to f3. For example, below a coupling frequency of, for example, 1000 Hz, the Down Mix Matrix 6 'can provide two channels, and above the coupling frequency, the Down Mix Matrix 6' can provide one channel. . By employing two channels below the coupling frequency, better spatial fidelity can be obtained, especially if the two channels represent horizontal directions (to accommodate the horizontality of human ears).
Although Figure 6 shows the generation of the same sidechain information for each channel as in the configuration of Figure 1, it may be possible to omit certain sidechain information when more than one channel is provided by the Matrix output. 6 'Down Mix. In some cases, acceptable results may be obtained when only the amplitude scale factor side chain information is provided by the configuration of Figure 6. Additional details regarding the side chain options are discussed below, in connection with the descriptions in Figures 7, 8 and 9.
As just mentioned above, the multiple channels generated by the Downmix Matrix 6 'need not be less than the number of input channels n. When the purpose of an encoder such as that in Figure 6 is to reduce the number of bits for transmission or storage, the number of channels produced by the downmix matrix 6 'is likely to be less than the number of channels n of entrance. However, the configuration of Figure 6 can also be used as a "bottom mixer". In that case, there may be applications in which the number of channels m produced by the Downmix Matrix 6 'is greater than the number of input channels n.
The encoders that have been described with reference to the examples of Figures 2, 5 and 6 may also include their own local decoder or decoder function, in order to determine whether the audio information and the side chain information, when used decode with such a decoder, they would provide adequate results. The results of such a determination could be used to improve the parameters using, for example, a recursive process. In a block encoding and decoding system, recursion calculations could be performed, for example, on each block before the end of the next block, in order to minimize the delay in the transmission of a block of audio and data information. its associated spatial parameters.
A configuration in which the encoder also includes its own decoder or decoder function could also be advantageously employed when spatial parameters are not stored or sent only for certain blocks. If improper decoding resulted from not sending the spatial parameter side chain information, such side chain information would be sent for the particular block. In this case, the decoder can be a modification of the decoder or decoding function of Figures 2, 5 or 6, in that the decoder would have the ability to retrieve the information from the spatial parameter side chain, for frequencies above the coupling frequency, from the incoming bit string, but also from generating the side chain information of simulated spatial parameters, from the stereo information below the coupling frequency.
In a simplified alternative to such encoder examples incorporating a local decoder, instead of having a local decoder or decoder function, the encoder could simply run a check to determine if there was any content of a signal below the coupling frequency ( determined in any suitable way, e.g. the sum of the energy in frequency bins across the frequency range) and, if not, it would send or store spatial parameter side chain information, rather than not if the energy was above the threshold. Depending on the encoding scheme, low signal information below the coupling frequency could also result in more bits available to send the sidechain information.
M: N decoding
Figure 7 illustrates a more generalized form of the configuration of figure 2, where a matrix upmix function or device ("Upmix Matrix") 20 receives the 1 am channels generated by the configuration of figure 6. The Up Mix Matrix 20 may be a passive matrix. It may be, but need not be, the conjugate transposition (i.e., the complement) of the Downmix 6 'Matrix of the Figure 6 configuration. Alternatively, the Up Mix Matrix 20 may be an active matrix, a variable, or a passive matrix in combination with a variable matrix. If an active matrix decoder is employed, in its idle or inactive state, it may be the complex conjugate of the Down Mix Matrix or it can be independent of the Down Mix Matrix. The side chain information can be applied as illustrated in Figure 7, so as to control Amplitude Adjustment, Angle Rotation, and (optionally) interpolating devices or functions. In that case, the Upmix Matrix, if it is an active matrix, works independently of the sidechain information and responds only to the channels applied to it. Alternatively, some or all of the side chain information can be applied to the active matrix to aid its function. In that case, some or all of the Amplitude Adjustment, Angle Rotation and Interpolators functions or devices may be omitted. The Decoder example of Figure 7 may also employ the alternative, or apply a degree of random amplitude variation under certain signal conditions, as described above in connection with Figures 2 and 5.
ES 2 324 926 T3
When the Upmix Matrix 20 is an active matrix, the configuration of FIG. 7 may be characterized as a "hybrid matrix decoder" to operate in a "hybrid matrix encode / decode system." In this context, "hybrid" refers to the fact that the decoder can obtain some measure of control information from the input audio signal (that is, the active matrix responds to the spatial information encoded in the channels applied to ella) and an additional measure of the control information from the side chain information of the spatial parameters. Other elements of Figure 7 are like those of the configuration of Figure 2 and have the same reference numerals.
Suitable active matrix decoders for use in a hybrid matrix decoder may include active matrix decoders, such as those mentioned above and incorporated by reference, including, for example, matrix decoders known as "Pro Logic" and "Pro Logic II decoders. ”(“ Pro Logic ”is a trademark of Dolby Laboratories Licensing Corporation).
Alternative decorrelation
Figures 8 and 9 show variations of the generalized decoder of figure 7. In particular, both the configuration of figure 8 and the configuration of figure 9 show alternatives to the decoding technique of figures 2 and 7. In figure 8 , the respective functions or decorrelation devices ("Decorrelation Devices") 46 and 48 are in the time domain, each of them following the respective Inverse Filter Bank 30 and 36 in its channel. In figure 9, the respective functions or decorelation devices ("Decorelation Devices") 50 and 52 are in the frequency domain, each one preceding the respective Inverse Filter Bank 30 and 36 in its channel. In both configurations of Figure 8 and Figure 9, each of the Décorrelation Devices (46, 48, 50, 52) has a unique characteristic, such that their outputs are mutually décorrelated with respect to each other. The Décorrelation Scale Factor can be used to control, for example, the ratio of the correlated signal to the uncorrelated signal provided on each channel. Optionally, the Transient Flagger can also be used to vary the mode of operation of the Decorrelation Device, as explained below. In both configurations of Figure 8 and Figure 9, each Decorrelation Device may be a Schroeder-type reverberator, having its own unique filtering characteristic, in which the amount or degree of reverberation is controlled by the scale factor of decorrelation (implemented, for example, controlling the degree to which the Décorrelation Device output forms part of a linear combination of the Décorrelation Device input and output). Alternatively, other controllable décorrelation techniques may be employed, either alone or in combination with each other or with a Schroeder-type reverberator. Schroeder-type reverberators are well known and their origin can be investigated in two journal papers: "Colorless Artificial Reverberation" by MR Schroeder and BF Logan, IRE Transactions on Audio, vol. AU-9, pages 209-214, 1961 and in MR Schroeder's "Natural Sounding Artificial Reverberation", Journal AES., July 1962, vol. 10. no. 2, pages 219-223.
When the Décorrelation Devices 46 and 48 operate in the time domain, as in the configuration of FIG. 8, a single Décorrelation Scale Factor (ie, broadband) is required. This can be achieved in any of a number of ways. For example, a single Décorrelation Scale Factor can be generated in the encoder of Figure 1 or Figure 7. Alternatively, if the encoder of Figure 1 or Figure 7 generates Décorrelation Scale Factors based on a sub-band, the Sub-band Décorrelation Scale Factors can be summed in amplitude or power in the encoder of figure 1 or figure 7, or in the decoder of figure 8.
When the Decorelation Devices 50 and 52 operate in the frequency domain, as in the configuration of Figure 9, they can receive a decorelation scale factor for each sub-band or groups of sub-bands and, simultaneously, provide a degree provided decorrelation to such sub-bands or groups of sub-bands.
The Décorrelation Devices 46 and 48 of FIG. 8 and the Décorrelation Devices 50 and 52 of FIG. 9 may optionally receive the Transient Marker. In the time-domain Décorrelation Devices of Figure 8, the Transient Flagger can be used to vary the mode of operation of the respective Décorrelation Device. For example, the Decorrelation Device may function as a Schroeder-type reverberator in the absence of a transient flag, but upon receipt and for a short period of time thereafter, for example 1 to 10 milliseconds, it operates as a fixed delay. Each channel can have a predetermined fixed delay or the delay can be varied in response to a plurality of transients within a short period of time. In the Decorelation Devices of the frequency domain of figure 9, the transient flag can also be used to vary the operating mode of the respective Decorelation Device. However, in this case, the reception of a transient flagger can trigger, for example, a short increase in amplitude (several milliseconds) in the channel in which the flagger has occurred.
In both configurations of Figures 8 and 9, an interpolator 27 (33), controlled by the optional Transient Flagger, can provide an interpolation across the frequency of the phase angles delivered to the Rotation Angle 28 (33) of the way described above.
ES 2 324 926 T3
As mentioned above, when two or more channels are sent in addition to the sidechain information, it may be acceptable to reduce the number of sidechain parameters. For example, it may be acceptable to send only the Amplitude Scaling Factor, in which case the decoder and the decorrelation and angle devices or functions may be omitted (in that case, Figures 7, 8 and 9 are reduced to the same setting ).
Alternatively, only the Amplitude Scale Factor, the Décorrelation Scale Factor, and optionally the Transient Flag may be sent. In that case, any of the configurations of Figures 7, 8 and 9 can be used (omitting the Angle Rotation 28 and 34 in each of them).
As another alternative, only the amplitude scale factor and the angle control parameter can be sent. In that case, any of the configurations of Figures 7, 8 or 9 may be employed (omitting the Decorrelation Device 38 and 42 of Figure 7, and 46, 48, 50 and 52 of Figures 8 and 9).
As in Figures 1 and 2, the configurations of Figures 6-9 are intended to show any number of input and output channels although, for simplicity of presentation, only two channels are illustrated.
It is contemplated to cover by the present invention any and all modifications, variations or equivalents that fall within the scope of the appended claims.
Contents25
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
129 members in 17 offices
Priority claims15
| Document | Office | Kind | Date |
|---|---|---|---|
| 20040549368P | United States of America | – | |
| 54936804 | United States of America | P | |
| 54936804 | United States of America | P | |
| 20040579974P | United States of America | – | |
| 57997404 | United States of America | P | |
| 57997404 | United States of America | P | |
| 20040588256P | United States of America | – | |
| 58825604 | United States of America | P | |
| 58825604 | United States of America | P | |
| 549368P08001529 | – | – | – |
| 579974P | – | – | – |
| 588256P | – | – | – |
| US20040549368P | – | – | – |
| US20040579974P | – | – | – |
| US20040588256P | – | – | – |
Members129
| Document | Office | Kind | |
|---|---|---|---|
| AU2005219956A1 | Australia | A1 | |
| CA2556575A1 | Canada | A1 | |
| CA2808226A1 | Canada | A1 | |
| CA2917518A1 | Canada | A1 | |
| CA2992051A1 | Canada | A1 | |
| CA2992065A1 | Canada | A1 | |
| CA2992089A1 | Canada | A1 | |
| CA2992097A1 | Canada | A1 | |
| CA2992125A1 | Canada | A1 | |
| CA3026245A1 | Canada | A1 | |
| CA3026267A1 | Canada | A1 | |
| CA3026283A1 | Canada | A1 | |
| WO2005086139A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200537436A | Taiwan Province of China | A | |
| EP1721312A1 | European Patent Office (EPO) | A1 | |
| IL177094D0 | Israel | D0 | |
| KR20060132682A | Republic of Korea | A | |
| HK1092580A1 | Hong Kong, China | A1 | |
| CN1926607A | China | A | |
| US2007140499A1 | United States of America | A1 | |
| BRPI0508343A | Brazil | A | |
| JP2007526522A | Japan | A | |
| WO2007109338A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200742275A | Taiwan Province of China | A | |
| US2008031463A1 | United States of America | A1 | |
| EP1721312B1 | European Patent Office (EPO) | B1 | |
| AT390683T | Austria | T | |
| ATE390683T1 | Austria | T1 | |
| EP1914722A1 | European Patent Office (EPO) | A1 | |
| US2008102119A1 | United States of America | A1 | |
| DE602005005640D1 | Germany | D1 | |
| WO2008085222A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008085222A3 | World Intellectual Property Organization (WIPO) | A3 | |
| SG149871A1 | Singapore | A1 | |
| HK1119820A1 | Hong Kong, China | A1 | |
| EP1914722B1 | European Patent Office (EPO) | B1 | |
| DE602005005640T2 | Germany | T2 | |
| AT430360T | Austria | T | |
| ATE430360T1 | Austria | T1 | |
| AU2005219956B2 | Australia | B2 | |
| EP2065885A1 | European Patent Office (EPO) | A1 | |
| DE602005014288D1 | Germany | D1 | |
| AU2009202483A1 | Australia | A1 | |
| EP2083883A2 | European Patent Office (EPO) | A2 | |
| ES2324926T3This record | Spain | T3 | |
| CN101552007A | China | A | |
| HK1128100A1 | Hong Kong, China | A1 | |
| US2009299756A1 | United States of America | A1 | |
| EP2065885B1 | European Patent Office (EPO) | B1 | |
| AT475964T | Austria | T | |
| ATE475964T1 | Austria | T1 | |
| EP2224430A2 | European Patent Office (EPO) | A2 | |
| DE602005022641D1 | Germany | D1 | |
| EP2224430A3 | European Patent Office (EPO) | A3 | |
| IL177094A | Israel | A | |
| HK1142431A1 | Hong Kong, China | A1 | |
| CN1926607B | China | B | |
| US2011184389A1 | United States of America | A1 | |
| CN102169693A | China | A | |
| CN102176311A | China | A | |
| EP2224430B1 | European Patent Office (EPO) | B1 | |
| AT527654T | Austria | T | |
| ATE527654T1 | Austria | T1 | |
| KR101079066B1 | Republic of Korea | B1 | |
| MY145083A | Malaysia | A | |
| JP4867914B2 | Japan | B2 | |
| US8170882B2 | United States of America | B2 | |
| AU2009202483B2 | Australia | B2 | |
| AU2012208987A1 | Australia | A1 | |
| AU2012208987B2 | Australia | B2 | |
| CA3026276A1 | Canada | A1 | |
| CA3035175A1 | Canada | A1 | |
| TWI397902B | Taiwan Province of China | B | |
| CN101552007B | China | B | |
| CA2556575C | Canada | C | |
| TW201329959A | Taiwan Province of China | A | |
| TW201331932A | Taiwan Province of China | A | |
| CN102169693B | China | B | |
| CN102176311B | China | B | |
| US8983834B2 | United States of America | B2 | |
| TWI484478B | Taiwan Province of China | B | |
| US2015187362A1 | United States of America | A1 | |
| TWI498883B | Taiwan Province of China | B | |
| US9311922B2 | United States of America | B2 | |
| US2016189718A1 | United States of America | A1 | |
| US2016189723A1 | United States of America | A1 | |
| CA2808226C | Canada | C | |
| SG10201605609PA | Singapore | A | |
| US9454969B2 | United States of America | B2 | |
| US9520135B2 | United States of America | B2 | |
| US2017076731A1 | United States of America | A1 | |
| US9640188B2 | United States of America | B2 | |
| US2017148456A1 | United States of America | A1 | |
| US2017148457A1 | United States of America | A1 | |
| US2017148458A1 | United States of America | A1 | |
| US9672839B1 | United States of America | B1 | |
| US2017178650A1 | United States of America | A1 | |
| US2017178651A1 | United States of America | A1 | |
| US2017178652A1 | United States of America | A1 | |
| US2017178653A1 | United States of America | A1 |
Numbers
- Publication
- 2324926
- Publication, DOCDB
- 2324926
- Publication, EPODOC
- ES2324926T
- Application
- 8001529
- Application, DOCDB
- 08001529
- Application, EPODOC
- ES20080001529T
Titles2
- Spanish
- DESCODIFICACION DE AUDIO MULTICANAL.
- English
- DESCO DIFFICATION OF MULTICHANNEL AUDIO.
Classification
- CPC, 12
- G10L19/008
- G10L19/06
- H04S3/02
- H04S5/00
- G10L19/26
- G10L19/0204
- H04S3/00
- H04S3/008
- G10L19/02
- G10L19/005
- G10L19/018
- G10L19/025
- IPC, 3
- G10L19 00
- H04S3 02
- H04S5 00