Improving transient performance of low bit rate audio coding systems by reducing pre-noise
Abstract
A method for reducing distortion artifacts that precede a signal transient in a train of audio frequency signals subsequent to a reverse transformation, in the decoder or in a low-speed bit-transfer audio coding system based on transformation , which employs coding blocks, whose method comprises receiving metadata information that is useful in reducing the duration of the transient pre-noise, whose metadata information includes the location of transients, and alter the time duration of at least a portion of said distortion artifacts, in response to said metadata information, such that the time duration of said distortion artifacts is reduced .

Term
Term ended
Projected expiry passed 25 April 2022, 4.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
2 claims: 1 independent, 1 dependent
- 1ES 2 298 394 T3 REIVINDICACIONES 1. Un método para reducir los artefactos de distorsión que preceden a un transitorio de señal en un tren de señales de audiofrecuencia subsiguiente a una transformación inversa, en el descodificador o en un sistema de codificación de audiofrecuencia a baja velocidad de transferencia de bits basado en transformación, que emplea bloques de codificación, cuyo método comprende recibir información de metadatos que es útil en la reducción de la duración del pre-ruido del transitorio, cuya información de metadatos incluye la ubicación de transitorios, y alterar la duración de tiempo de al menos una parte de dichos artefactos de distorsión, en respuesta a dicha información de metadatos, de tal manera que se reduce la duración de tiempo de dichos artefactos de distorsión.
- 2El método de la reivindicación 1, en el que dicha información de metadatos incluye también una o más de:la longitud del bloque (o de los bloques) de codificador de audiofrecuencia, la relación entre los límites de bloque de codificador con los datos de audiofrecuencia, y una longitud deseada del pre-ruido del transitorio.
Independent claims2
145 paragraphs in 10 sections, as filed
ES 2 298 394 T3
DESCRIPTION
Improvement of transient sessions of audio-frequency signal coding systems at low bit rate by pre-noise reduction.
Technical field
The invention relates generally to low bit rate digital transform coding and decoding of information representing audio signals such as music signals or speech signals. More particularly, the invention relates to the reduction of distortion artifacts preceding a signal transient ("pre-noise").
Background in the Prior Art
Time escalation
The term "time scaling" refers to the alteration of the evolution or duration in time of an audio frequency signal at the same time that its spectral content (perceived timbre) or perceived tone (where the tone is a characteristic in association relationship with periodic audio signals). Tone scaling refers to the modification of the spectral content or perceived tone of an audio frequency signal while its evolution or duration is not affected over time. Time scaling and pitch scaling are dual methods of each other. For example, a digitized audio signal tone could be increased by 5% without affecting its duration over time by time-scaling it by 5% (that is, increasing the signal's time duration) and then extracting of sample information at a 5% faster rate of variation of the samples (for example, by resampling), thereby maintaining its original time duration. The resulting signal has the same length of time as the original signal, but with a modified tone or spectral characteristics. Resampling is not an essential step in time scaling or pitch scaling, unless you want to keep a constant output sample rate or keep the input and output sample rates the same.
In aspects of the present invention, time scaling processing of audio frequency signal streams is employed. However, as mentioned above, time scaling could also be performed using pitch scaling techniques, since they are dual to each other. Thus, although the term "time scaling" is used herein, techniques that employ pitch scaling could also be employed to obtain time scaling.
Low bit rate audio signal encoding
Among those in the field of signal processing, there is considerable interest in minimizing the amount of information required to represent a signal without a perceptible loss in signal quality. By reducing information requirements, signals impose less information capacity requirements on communication channels and storage media. With respect to digital encoding techniques, the minimum information requirements are synonymous with the minimum binary bit requirements.
Some prior techniques for encoding audio signals intended for human hearing attempt to reduce information requirements without producing any audible degradation by exploiting psychoacoustic effects. The human ear exhibits frequency analysis properties that resemble highly asymmetric tuned filters that have varying center frequencies. The human ear's ability to detect distinct tones generally increases with the difference in frequencies between the tones; however, the ear's resolving ability remains substantially constant for frequency differences smaller than the bandwidth of the aforementioned filters. Therefore, the frequency resolution capacity of the human ear varies according to the bandwidth of these filters throughout the entire audio frequency spectrum. The effective bandwidth of such an auditory filter is referred to as a critical band. A dominant signal within a critical band is more likely to mask the audibility of other signals anywhere within the critical band than other signals at frequencies outside that critical band. A dominant signal could mask other signals that occur not only at the same time as the masking signal, but also occur before and after the masking signal The duration of pre- and post-masking effects within a critical band it depends on the amplitude of the masking signal, but usually the pre-masking effects are of a much shorter duration than the post-masking effects. See, generally, the Audio Engineering Handbook K. Blair Benson editors, Mc-Graw-Hill, San Francisco 1988, pages 1.40-1.42 and 4.8-4.10
Signal recording and transmission techniques that divide the signal's useful bandwidth into frequency bands with bandwidths approaching the ear's critical bands can take better advantage of psycho-acoustic effects than wider-band techniques. . Techniques that exploit psychoacoustic masking effects can encode and reproduce a signal that is indistinguishable from the original input signal using a lower bit rate than required by modified pulse modulation (PCM) coding.
ES 2 298 394 T3
Critical band techniques comprise dividing the signal bandwidth into frequency bands, processing the signal from each frequency band, and reconstructing a replica of the original signal from the processed signal from each frequency band. Two of these techniques are subband coding and transform coding. Subband and transform encoders can reduce the requirements for information transmitted in particular frequency bands where the resulting coding imprecision (noise) is psychoacoustically masked by neighboring spectral components without degrading the subjective quality of the encoded signal.
A bank of digital band pass filters could implement sub-band coding. Transform coding could be implemented by any of several discrete time-domain-to-frequency-domain transforms that a bank of digital bandpass filters implements. The remaining description refers more particularly to transformation codes, therefore the term "sub-band" refers in this case to selected parts of the total bandwidth of the signal, whether implemented by a sub-band encoder or by a transform encoder. A subband as implemented by a transform encoder is defined by a set of one or more adjacent transform coefficients; hence the sub-band bandwidth is a multiple of the bandwidth of the transformation coefficient, The bandwidth of a transformation coefficient is directly proportional to the sampling rate of the input signal and inversely proportional to the number of coefficients generated by the transformation to represent the input signal.
Psycho-acoustic masking could be carried out more easily using transform codes if the sub-band bandwidth across the entire audible spectrum is approximately half the critical bandwidth of the human ear in the same parts of the spectrum. This is because the critical bands of the human ear have variable center frequencies that are adapted to auditory stimuli, while the sub-band and transform encoders typically have fixed sub-band center frequencies. To optimize the utilization of the psycho-acoustic masking effects, any distortion artifacts resulting from the presence of a dominant signal should be limited to the sub-band that contains the dominant signal. If the bandwidth of the sub-band is approximately half or less than half the critical band and if the selectivity of the filter is sufficiently high, effective masking of undesirable distortion products is likely even for signals whose frequency is near the edge of the sub-band pass bandwidth. If the bandwidth of the sub-band is more than half that of a critical band, there is a possibility that the dominant signal may cause the critical band of the ear to be offset from the sub-band encoder in such a way that it is not they mask some of the undesirable distortion products outside the ear's critical bandwidth. This effect is most objectionable at low frequencies, where the ear's critical band is narrower.
The probability that a dominant signal can cause the ear's critical band to de-center from an encoder subband and thus "discover" other signals in the same encoder subband is generally higher at low frequencies, where the critical band the ear is narrower. In transformation codes, the narrowest possible sub-band is a transformation coefficient, so the psycho-acoustic masking could be performed more easily if the bandwidth of the transformation coefficient does not exceed half the bandwidth of the critical band of maximum narrowness of the ear. Increasing the length of the transformation could decrease the bandwidth of the transformation coefficient. A drawback of increasing the length of the transformation is an increase in the complexity of the processing to compute the transformation and in encoding larger numbers of narrower subbands. Other drawbacks are discussed later.
Of course, psycho-acoustic masking could be achieved using wider sub-bands if the center frequency of these sub-bands can be changed to follow the components of the dominant signal in the same way that the center frequency of the band changes. ear criticism.
The ability of a transform encoder to exploit psycho-acoustic masking effects also depends on the selectivity of the filter bank implemented by the transform. The term filter "selectivity", as used herein, refers to two characteristics of subband bandpass filters. The first is the bandwidth of the regions between the filter bandpass and the attenuated bands (the width of the transition bands). The second is the attenuation level in the attenuated bands. Thus, filter selectivity refers to the steepness of the filter response curve within the transition bands (steepness of the transition band roll-off), and the level of attenuation in the attenuated bands (depth of attenuated band rejection).
Filter selectivity is directly affected by numerous factors including the three factors discussed below: block length, window weighting functions, and transformations. In a very general sense, block length affects encoder frequency and temporal resolution, and windowing and transforms affect encoding gain.
Low bit rate / block length audio coding
The input signal to be encoded is sampled and segmented into "signal sample blocks" prior to sub-band filtering. The number of samples contained in the signal sample block is the block length of the signal sample.
ES 2 298 394 T3
It is common for the number of coefficients generated by a transform filter bank (transform length) to be equal to the block length of signal samples, but it is not necessary. An overlapping block transform could be used, and is sometimes described in the art as a length N transform that transforms blocks of signal samples with 2N samples. It can also be described as a transformation of length 2N that generates only exclusive coefficients N. Since all the transformations described herein can be considered to have lengths equal to the block length of signal samples, they are generally used in the I present the two lengths as synonymous with one another.
The block length of signal samples affects the temporal and frequency resolution of a transform encoder. Transform encoders that use shorter block lengths have poorer frequency resolution, because the bandwidth of the discrete transform coefficient is wider and the filter selectivity is lower (slower rate of change of the attenuation of the transition band and a lower level of attenuated band rejection). This degradation in filter performance causes the energy of a single spectral component to spread into neighboring transformation coefficients. This undesirable dispersion of spectral energy is the result of degraded filter behavior called "side lobe leakage".
Transform encoders that use longer block lengths have poorer temporal resolution, because quantization errors cause a transform encoder / decoder system to "smear" the frequency components of a sampled signal across the entire length of the block. of signal samples. The distortion artifacts present in the signal recovered from the inverse transformation are most audible as a result of large changes in signal amplitude that occur over a time interval much shorter than the length of the signal sample block. These changes in amplitude are referred to herein as "transients." Said distortion manifests as noise in the form of a transient echo or oscillation just before (pre-transient or "pre-noise" noise) or just after (post-transient noise) the transient. Pre-noise is of particular interest because it is highly audible and, unlike post-transient noise, it is minimally masked (a transient provides only minimal temporal pre-masking). Pre-noise occurs when the high frequency components of the audio transient material temporarily stain across the length of the audio encoder block in which it occurs. The present invention is substantially concerned with minimizing pre-noise. Post-transient noise is typically substantially masked. and it is not the object of this invention.
Fixed block length transform encoders use a compromise block length that offsets temporal resolution against frequency resolution. A short block length degrades subband filter selectivity, which could result in a nominal passband filter bandwidth exceeding the ear critical bandwidth at lower frequencies or at all frequencies. Even if the nominal sub-band bandwidth is narrower than the ear critical bandwidth, degraded filter characteristics manifested as a wide transition band and / or poor attenuated band rejection could result in significant signal artifacts. outside the ear's critical bandwidth. Conversely, a large block length could improve filter selectivity, but reduce temporal resolution, which could result in audible signal distortion occurring outside of the ear's temporal psycho-acoustic masking range.
Window weighting function
Discrete transformations do not produce a perfectly precise set of frequency coefficients, because they work with only one finite-length segment of the signal, the signal sample block. Strictly speaking, discrete transformations produce a time-frequency representation of the input signal in the time domain rather than a true representation in the frequency domain, which would require infinite block lengths of signal samples. However, for convenience of description, the output of the discrete transformations is referred to herein as a representation in the frequency domain. In reality, the discrete transformation assumes that the sampled signal only has frequency components whose periods are a submultiple of the signal sample block length. This is equivalent to a hypothesis that the signal of finite length is periodic. Of course, the hypothesis in general is not true. The assumed periodicity creates discontinuities at the edges of the signal sample block that cause the transformation to create phantom spectral components.
One technique that minimizes this effect is to reduce the discontinuity before transformation by weighting the signal samples in such a way that the samples near the edges of the signal sample block are zero or very close to zero. Samples in the center of the signal sample block generally pass unchanged, that is, weighted by a factor of one. This weighting function is called an "analysis window." The shape of the window directly affects the selectivity of the filter.
As used herein, the term "analysis window" refers only to the window selection function performed prior to the application of direct transformation. The analysis window is a time domain function. If no compensation is provided for "window" effects, the recovered or "synthesized" signal is distorted according to the shape of the analysis window. A compensation method known as overlap-add is well known in the art. This method requires the encoder to transform blo
ES 2 298 394 T3 which input signal samples overlap. By carefully designing the analysis window such that two adjacent windows are added to the unit through the flap, the effects of the window are exactly offset. The shape of the window significantly affects. See generally Harris's paper entitled "On the Use of Windows for Harmonic Analysis with the Discrete Fourier Transform," IEEE Proceedings, Volume 66, January, 1978, pp. 51-83. As a general rule, "smoother" shaped windows and larger flap ranges provide better selectivity. For example, a Kaisser-Bessel window generally provides greater filter selectivity than a sinusoidal tapered rectangular window.
When used with certain types of transformations such as the discrete Laplace transform (hereinafter DFT), the overlap-add method increases the number of bits required to represent the signal, because the part of the signal contained in the overlap interval it must be transformed and transmitted twice, once for each of the two overlapping signal sample blocks. Signal analysis / synthesis for systems using such overlap-add transformation is not critically sampled. The term "critically sampled" refers to a signal analysis / synthesis that over a period of time generates the same number of frequency coefficients as the number of input signal samples it receives. Hence, for systems that are non-critically sampled , it is desirable to design the window with as small an overlap interval as possible to minimize the information requirements of the encoded signal.
Some transformations also require that the synthesized output of the inverse transformation be windowed selected. The synthesis window is used to shape each synthesized signal block. Therefore, the synthesized signal is weighted by both an analysis window and a synthesis window. This two-stage weighting is mathematically similar to weighting the original signal once through a window whose shape is equal to a sample-by-sample product of the analysis and synthesis windows. Therefore, in order to use the overlap-add method to compensate for window distortion, both windows must be designed in such a way that the product of the two sums unifies across the overlap-add interval.
Although there is no single criterion that can be used to establish a window optimization, a window is generally considered "good" if the selectivity of the filter used with the window is considered "good". Therefore, a well-designed analysis window (for transformations using only one analysis window) or a pair of analysis / synthesis windows (for transformations using one analysis window and one synthesis window) can reduce lobe leakage. side.
Block switching
A common solution that overcomes the trade-off between time resolution and frequency resolution in fixed block length transform encoders is the use of transient detection and block length switching. In this solution, the presence and location of audio signal transients are detected using various transient detection methods. When transient audio signals are detected that are likely to introduce pre-noise when encoded using a large audio encoder block length, the low bit rate encoder is switched from the most efficient long block length to a low bit rate encoder. less efficient shorter block length. Although this reduces the frequency resolution and encoding performance of the encoded audio signal, it also reduces the length of transient pre-noise introduced by the encoding process, improving the perceived quality of the audio signal after decoding to low bit rate. Techniques for block length switching are described in US Patent Numbers 5,394,473; 5,848,391; and 6,226,608. Although the present invention reduces pre-noise without the complexity and drawbacks of block switching, it could be used in conjunction with - and in addition to - block switching.
The document prepared by Vafin R. and collaborators, entitled "Modification of transients for an efficient coding of audio frequency", INTERNATIONAL CONFERENCE OF THE INSTITUTE OF ELECTRICAL AND ELECTRONIC ENGINEERS (IEEE) OF 2001 ON ACOUSTICS, SIGNALS AND VOICE PROCESSING. PROCEEDINGS. May 7-11, 2001, pages 3285-3288 describes modifying, in an audio frequency parametric code, the location of estimated transients such that transients can occur only at locations specified by a grid. The grid is defined by a restricted segmentation in which the segments are defined by multiples of whole numbers of a predefined minimum segment size.
WO 00/45378 describes a method for spectral envelope coding in which, in the vicinity of transients, temporal resolution is increased at the expense of frequency resolution. In the coding system that handles the time slots of an input signal, this is achieved by changing the length of the respective time slots.
Description of the invention
In accordance with one aspect of the present invention, a method for reducing distortion artifacts preceding a signal transient in an audio frequency signal train subsequent to inverse transformation in the decoder of a speed audio coding system bit transfer based on
ES 2 298 394 T3 transformation that uses coding blocks, comprises altering the time duration of at least a part of the distortion artifacts, in response to said metadata information, in such a way as to reduce the time duration of the artifacts distortion. The metadata information includes the location of transients.
By such treatment, referred to herein as "post-treatment," audio frequency quality improvements could be achieved whether pre-treatment is used or not. Any audio signal that has undergone low bit rate audio encoding and decoding could be analyzed to identify the location of transient signals and estimate the duration of pre-noise transient signal artifacts. Then, time-scaled post-processing could be performed on the audio signal in order to remove pre-noise from the transient signal or to reduce its duration.
There are several compensation techniques to reduce alterations in the evolution of audio-frequency signal trains over time. These time-scaled compensation techniques also have the beneficial result of keeping the number of audio samples constant.
A first time-scaling compensation technique, which is useful in relation to pre-treatment, is applied before direct transformation. Applies a time offset scaling to the audio signal stream that follows the transient, the time scaling having the opposite direction to the direction of the time scaling used to change the position of the transient and preferably having substantially the same duration as the transient. time scaling of the transient change. For convenience of description, this type of compensation will be referred to herein as "sample number compensation", because it is capable of keeping the number of samples of audio signals constant but it is not capable of fully restoring evolution. original timing of the audio signal train (temporarily out of place for transients and parts of the audio signal train that are near the transient). Preferably, the time scaling that provides sample number compensation closely follows the transient, such that it temporarily masks it.
Although the sample number compensation leaves the transient shifted from its original temporal position, the fact is that it restores the audio signal train that follows the time compensation scaling to its original relative temporal position. In this way, the audibility probability of the transient time change is reduced, but not eliminated, because the transient is still out of its original position. However, this could provide a significant reduction in audibility and has the advantage that it is done prior to low bit rate audio encoding, allowing the use of a standard decoder, unmodified. As explained below, a complete recovery of the time evolution of the audio signal train can only be accomplished by processing in the decoder or after the decoder. In addition to reducing the possibility of transient time shift, time-scaling compensation before direct transformation has the advantage of keeping the number of samples of audio signals constant, which could be important for treatment and / or for the operation of the hardware that implements the treatment.
In order to provide optimal time-scaling compensation prior to direct transformation, information regarding the location of the transient and the temporal duration of the transient time change should be employed through the compensation process.
If the transient time change is applied after blocking (but before applying the direct transformation) it is necessary to use compensation for the number of samples within the same block in which the transient time change is performed in order to keep the same the block length. Therefore, it is preferred to perform transient time change and sample number compensation before blocking.
Sample number compensation could also be employed after inverse transformation (either in decoder or after decoding) in connection with post-processing. In this case, information useful for performing the compensation could be sent to the compensation process from the decoder (which information could have originated from the encoder and / or the decoder).
A more complete recovery of the time evolution of the audio signal train could be performed together with the restoration of the original number of audio samples after the inverse transformation (either at the decoder or after decoding), by applying an offset time scaling to the audio signal stream before the transient in the opposite direction to the time scaling used to change the position of the transient and preferably of the same substantial duration as the time scaling of the transitory change. For convenience of description, this type of compensation will be referred to hereinafter as "compensation for evolution over time." This time scaling compensation has the significant advantage of restoring the entire audio signal train, including the transient, to its original relative time position. This greatly reduces the audibility probability of the timescaling processes, but is not eliminated, because the two timescalling processes alone could cause audible artifacts.
ES 2 298 394 T3
In order to provide optimal compensation for the evolution over time, various information such as the location of the transient, the location of the ends of the block, the duration of the transient time change, or the duration of the pre-noise is useful. . The duration of the pre-noise is useful to ensure that the time scaling of the time evolution compensation does not occur during the pre-noise, which would possibly thereby extend the time duration of the pre-noise. The duration of the transient time change is useful if you want to restore the RF signal stream to its original relative temporal position and keep the number of samples constant. The location of the transient is useful because the duration of the pre-noise could be determined from the original location of the transient with respect to the ends of the coding blocks. The duration of the pre-noise could be estimated by measuring a signal parameter, such as the high frequency content, or a default value could be used. If compensation is done at the decoder or after decoding, the encoder could send useful information as metadata along with the encoded audio signal. When performed after decoding, metadata could be sent to the compensation process from the decoder (the information of which could have originated from the encoder and / or the decoder).
As mentioned above, the post-treatment to reduce the duration of the pre-noise artifact could also be applied as an additional step to an audio signal encoder that performs time-scaling pre-treatment and optionally provides information. metadata. Such post-treatment would act as an additional means of quality improvement by reducing the pre-noise that may still remain after post-treatment.
Pre-processing might be preferred in encoder systems employing professional encoders where cost, complexity, and time delay are relatively immaterial compared to post-processing in relation to a decoder, which is typically a device. consumer with less complexity.
The quality improvement technique of a low bit rate audio signal coding system could be implemented using any current suitable time scaling technique. A suitable technique is described in international patent application PCT / US02 / 04317, filed February 12, 2002, titled "High Quality Timing and Tone Scaling of Audio Frequency Signals." Said application designates the United States and other entities. As stated above, since time scaling and pitch shifting are dual methods of each other, time scaling could also be implemented using any suitable pitch scaling technique, as well as any that may be available in the future. . A pitch scaling followed by an extraction of information from the audio signal samples at a suitable rate that is different from the rate of change of the input sample results in a time-scaled version of the audio signal with the same spectral content or pitch of the original audio signal, and is applicable to the present invention.
As noted in the Low Bit Rate Audio Coding Background Summary, the selection of block duration in an audio signal coding system is a compromise between frequency resolution and temporal resolution. In general, a longer block duration is preferred, as it provides higher encoder performance (generally provides a higher quality of perceived audio signals with a reduced number of data bits) compared to a shorter block duration. However, the transient signals and pre-noise signals they generate counteract the gain in quality at longer block durations by introducing audible detrimental effects. It is for this reason that block switching or fixed shorter block durations are used in practical low bit rate audio encoder applications. However, the application of the time scaling pretreatment according to the present invention to audio data that is to undergo low bit rate encoding of audio signals and / or that has undergone post treatment could reduce the duration of pre-noise transients. This enables longer audio signal coding block durations to be used, thereby providing higher coding performance and improving the perceived audio signal quality without adaptively changing the block durations. However, pre-noise reduction according to the present invention could also be employed in coding systems using block duration switching. In such systems, there could be some pre-noise even for the minimum window size. The larger the window, the longer and consequently the more audible the pre-noise. Typical transients provide approximately 5 msec. pre-masking, which is translated to 240 samples at a sample rate of 48 kHz. If a window has more than 256 samples, which is common in a block switching arrangement, the invention brings some benefit.
Audio-frequency coding of transient pre-noise artifacts
Figures 1a-1e show examples of transient pre-noise artifacts generated by a fixed block length audio encoding system. Figure 1A shows six blocks, overlapped by 50%, of selected window for audio-frequency coding and of fixed length from 1 to 6. In this figure and in all other figures herein, each window is contiguous with an audio coding block and is referred to as "windowed block", "window", or "block". In this figure and in certain other figures herein, the windows are generally presented in the form of a Kaiser-Bessel window.
ES 2 298 394 T3
Other figures show windows in the form of semicircles for greater simplicity in presentation. The shape of the window is not critical to the present invention. Although the length of the window blocks of Figure 1a and other figures is not critical to the invention, the fixed length window blocks are typically in the range of 256 to 2048 samples in length. The four audio-frequency signal examples in Figures 1b to 1e illustrate, respectively, the effects of temporal relationships between windowed blocks for audio-frequency coding and transient pre-noise artifacts.
Figure 1b illustrates the relationship between the location of a transient signal in a stream of input audio signals to be encoded and the boundaries of the 50% overlapping windowed blocks. Although a fixed block length with overlap of 50% has been shown, the invention is applicable to both fixed and variable block length encoding systems and to blocks having an overlap other than 50%, including the non-overlapped blocks as described. described below with reference to Figures 2a through 5b.
Figure 1c shows the output of an audio signal stream from the audio coding system for the case of an audio signal stream input as shown in Figure 1b. As shown in Figures 1b and 1c, the transient is located between the end of the window block 3 and the end of the window block 4. Figure 1c illustrates the location and length of the transient pre-noise introduced by the low bit rate audio coding process relative to the location of the transient and the end of the windowed block 2. Note that the pre-noise is prior to the transient and is limited to windowed blocks 4 and 5, sample blocks in which the transient is located. Thus, the pre-noise extends back to the beginning of the windowed block 4.
In a similar way to Figures 1b and 1c, Figures 1d and 1e show, respectively, the relationship between an audio input signal train containing a transient located between the end of windowed block 2 and the end of block 3 with window and the pre-noise introduced into the audio-frequency output signal train by the audio-frequency coding system. Since pre-noise is limited to windowed blocks 3 and 4, within which the transient is located, pre-noise extends back to the beginning of windowed block 3. In this case, the pre-noise has a longer duration because the transient is closer to the end of the windowed block 3 than the transient of Figures 1b and 1c to the end of the windowed block 4. The ideal transient location is to follow very close to the end of the last block such that the pre-noise extends backward only to the next end of the previous block (about half the block length in the case of this example of 50% block flap).
It should be noted that the examples in Figures 1a-1e do not explicitly take into account the effects of gradual transition at the boundaries of the coding window. In general, as audio coding windows become progressively narrower, the pre-noise artifacts scale accordingly, and their audibility is reduced. For simplicity of presentation, scaling of pre-noise artifacts in the ideal waveforms of the figures herein has not been shown.
As suggested in Figures 1a-1e and shown in more detail in Figures 2A, 2B, 3A, 3B, 4A, 4B, 5A, and 5B, an audio encoder transient pre-noise artifact could be minimized if the location of transient signals is carefully placed before audio coding.
Examples of repositioning the location of a transient in order to reduce pre-noise are shown in Figures 2a, 2b, 3a, 3b, 4a, 4b, 5a and 5b for the non-overlapping block cases (Figures 2a and 2b ), with a block flap less than 50% (Figures 3a and 3b), with a block flap of 50% (Figures 4a and 4b), and a block flap greater than 50% (Figures 5a and 5b). In each case, unless the original position of the transient is equidistant between two successive block ends (in which case there is no preference), it is preferred to change the transient to a position that closely follows the closest block end. Whether the change is to the end of the previous block or to the end of the next block, and whether it is to the nearest block end or not, the resulting pre-noise is substantially the same. However, by tentatively changing the transient to a location that closely follows the nearest block end, disruption to the time evolution of the audio signal train is minimized. However, in some cases, the switch to the most distant block end could also be inaudible. Furthermore, even in the event that a shift to the most distant block end is audible, time evolution compensation could be employed, as discussed below, to reduce or eliminate such audibility.
Figures 2a and 2b present a series of ideal blocks with non-overlapping windows. In Figure 2a, an initial transient location is, as shown by a solid line arrow, closer to the end of the last window than it is to the end of the next window. The pre-noise for the initial transient location extends backward in time to the edge of the beginning of the window, as shown. If it is desired to minimize the degree of temporal change of the transient, it should be shifted "left" (backward in time) to a location that closely follows the end of the last windowed block, as shown. Although the resulting pre-noise still extends back to the beginning of the windowed block, this length is very short compared to the resulting pre-noise from the initial transient location. In this and other figures, the distance of the changed transient from the windowed block end has been exaggerated for clarity of presentation. In Figure 2b, the initial position of the transient is closer to the end of the next window than it is to the end of the previous window. Thus, if you want to minimize the degree of change
ES 2 298 394 T3 transient time, it should be shifted "clockwise" (later in time) to a location that closely follows the end of the next windowed block, as shown. Note that the improvement in pre-noise reduction increases when the initial position of the transient goes to a later time of the windowed block.
Figures 3a and 3b present a series of ideal window blocks that overlap by less than 50%. In Figure 3a, an initial transient location is, as shown by a solid line, closer to the end of the last window than to the end of the next window. The pre-noise for the initial transient location extends backward in time to the edge of the beginning of the window, as shown. If it is desired to minimize the degree of temporal change of the transient, it should be shifted "to the left" to a location that closely follows the end of the last windowed block, as shown. The resulting pre-noise still extends back to the beginning of the windowed block, but its length is short compared to the pre-noise resulting from the initial location of the transient. In Figure 3b, the initial position of the transient is closer to the end of the next window than it is to the end of the previous window. Thus, if it is desired to minimize the degree of temporal change of the transient, it should be shifted "to the right" to a location that closely follows the end of the next windowed block, as shown. Note that the improvement in pre-noise reduction increases because the initial position of the transient is later in the interval between successive window blocks.
Figures 4a and 4b present a series of ideal window blocks that overlap by 50%. In Figure 4a, an initial transient location is, as shown by the arrow drawn with a solid line, closer to the end of the last window than to the end of the next window. The pre-noise for the initial transient location extends backward in time to the edge of the beginning of the window, as shown. If it is desired to minimize the degree of temporal change of the transient, it should be shifted "to the left" to a location that closely follows the end of the last windowed block, as shown. The resulting pre-noise still extends back to the beginning of the windowed block, but its length is short compared to the pre-noise resulting from the initial location of the transient. In Figure 4b, the initial position of the transient is closer to the end of the next window than it is to the end of the previous window. Thus, if it is desired to minimize the degree of temporal change of the transient, it should be shifted "to the right" to a location that closely follows the end of the next windowed block, as shown. Note that the improvement in pre-noise reduction increases because the initial position of the transient is later in the interval between successive windowed blocks, the same as in the case of blocks overlapped by less than 50%.
Figures 5a and 5b present a series of ideal window blocks that overlap by more than 50%. In Figure 5a, an initial transient location is, as shown by the solid line arrow, closer to the end of the last window than to the end of the next window. The pre-noise for the initial transient location extends backward in time to the edge of the beginning of the window, as shown. If it is desired to minimize the degree of temporal change of the transient, it should be shifted "to the left" to a location that closely follows the end of the last windowed block, as shown. The resulting pre-noise still extends back to the beginning of the windowed block, but this length is still somewhat shorter than the pre-noise resulting from the initial location of the transient. In Figure 5b, the initial position of the transient is closer to the end of the next window than it is to the end of the previous window. Thus, if it is desired to minimize the degree of temporal change of the transient, it should be shifted "to the right" to a location that closely follows the end of the next windowed block, as shown. Note that the improvement in pre-noise reduction increases because the initial position of the transient is later in the interval between ends of successive window blocks, the same as in the case of blocks overlapped by 50%.
Note that the improvement in pre-noise reduction is the maximum for blocks that do not overlap, and that it decreases as the block flap increases.
Description of the drawings
Figures 1a-1e are a series of ideal waveforms illustrating examples of transient pre-noise artifacts generated by a fixed block length audio signal encoder system for two cases of input signal conditions.
Figures 2a and 2b present a series of ideal non-overlapping windowed blocks illustrating the initial and transiently changed temporary locations, along with the pre-noise for those locations, for the case of an initial position that is closer. of the end of the last window than of the end of the next window, and for the case of an initial position that is closer to the end of the next window than to the end of the previous window, respectively.
Figures 3a and 3b show a series of ideal windowed blocks with a flap less than 50% illustrating initial and transient changed temporary locations, together with the pre-noise for those locations, for the case of an initial position that is closer to the end of the last window than to the end of the next window,
ES 2 298 394 T3 and for the case of an initial position that is closer to the end of the next window than to the end of the previous window, respectively.
Figures 4a and 4b show a series of ideal windowed blocks with a 50% flap illustrating initial and transient changed temporary locations, along with the pre-noise for those locations, for the case of an initial position that is more closer to the end of the last window than the end of the next window, and for the case of an initial position that is closer to the end of the next window than to the end of the previous window, respectively.
Figures 5a and 5b show a series of ideal windowed blocks with a flap greater than 50% illustrating initial and transient changed temporary locations, together with the pre-noise for those locations, for the case of an initial position that is closer to the end of the last window than the end of the next window, and for the case of an initial position that is closer to the end of the next window than to the end of the previous window, respectively.
Figure 6 is a flow chart showing steps to be taken to reduce transient pre-noise artifacts by time scaling prior to low bit rate encoding.
Figure 7 is a conceptual representation of an input data buffer used for transient detection.
Figures 8a-8e are a series of ideal waveforms illustrating an example of time-scaled audio-frequency pre-treatment in accordance with aspects of the present invention when a transient exists in an audio-frequency coding block and is located closer to from the end of the last window block than from the end of the next window block.
Figures 9a-9e are a series of ideal waveforms illustrating an example of time-scaled audio treatment when a transient exists in a windowed audio coding block and is located approximately T samples ahead of one end of block.
Figures 10a-10d are a series of ideal waveforms illustrating time scaling for the case of multiple transients.
Figures 11a-11f are a series of ideal waveforms illustrating intelligent time-scaling compensation using metadata carried in a radio frequency signal stream.
Figure 12 is a flow chart of a time scaled post-treatment in conjunction with a low bit rate audio decoder.
Figures 13a-13c are a series of ideal waveforms illustrating an example of post-processing for a single transient in order to reduce pre-noise artifacts present after decoding.
Figure 14 is a flow chart of a post-processing process to improve the perceived quality of audio that has undergone low-bit-rate encoding without time-scaling pre-treatment.
Figures 15a-15c are a series of ideal waveforms demonstrating the technique of using a default value to time-scale the audio signal before each transient to reduce pre-noise without performing noise compensation. number of samples.
Figures 16a-16c are a series of ideal waveforms demonstrating the technique of using a calculated pre-noise duration to time-scale the audio signal before each transient, in order to reduce the duration of pre-noise. noise with compensation for the number of samples and evolution over time.
Optimal mode to carry out the invention
Pre-treatment overview with time scaling
Figure 6 is a flow chart illustrating a method of time scaling audio signals prior to low bit rate audio coding in order to reduce the amount of transient pre-noise (i.e. , "Pre-treatment"). This method handles input audio signals in blocks of N samples, where N could correspond to a number greater than or equal to the number of audio samples used in the audio coding block. Treatment sizes with N greater than the audio coding block size might be desirable to provide additional audio frequency data outside of the audio coding block for use in time-scaled treatment. This additional data could be used, for example, to compensate for the number of samples for the time-scaled treatment performed in order to improve the location of a transient.
ES 2 298 394 T3
The first step 202 in the process of Figure 6 checks the availability of N samples of audio data for time-scaled treatment. These audio data samples could be, for example, a file on a PC-based hard disk or a data buffer on a hardware device. The audio data could also have been provided by a low bit rate audio coding process that calls the processor with time scaling before audio coding. If N samples of audio data are available, they are passed (step 204) and then used by the time-scaled pre-treatment process in the following steps.
The third step 206 in the pre-treatment process is to detect the location of transient audio data signals that are likely to introduce pre-noise artifacts. Many different processes are available to perform this function, and their specific implementation is not critical as long as it provides accurate detection of transient signals that are likely to introduce pre-noise artifacts. There are many audio coding processes that perform transient detection of audio signals, and this step can be bypassed if the audio coding process provides the transient information to the subsequent time-scaled processing block 210 along with input audio data.
Transient detection
A suitable method for performing transient detection of audio signals is as follows. The first stage in transient detection analysis is to filter the input data (treating the data samples as a function of time). The input data could, for example, be filtered with a 2nd order harmonic high pass filter with a 3 dB cutoff frequency of approximately 8kHz. The characteristics of the filter are not critical. This filtered data is then used in transient analysis. Filtering the input data isolates high frequency transients and makes them easier to identify. The filtered input data is then processed in sixty-four sub-blocks (in the case of a signal sample block of 4,096 samples) of approximately 1.5 msec. (or 64 samples at 44.1 kHz) as shown in Figure 7. Although the actual size of the sub-block in question is not limited to 1.5 msec. and could vary, this size provides a good compromise between real-time treatment requirements (because larger block sizes require less treatment overhead) and transient location resolution (smaller blocks provide more detailed information on the location of transients). The use of 4096 sample signal sample blocks and the use of 64 sample sub-blocks is merely an example and is not critical to the invention.
The next stage of the transient detection treatment is to perform a low-pass filtering of the values of the maximum absolute data contained in each sub-block of 64 samples. This processing is done to smooth out the absolute maximum data and provide a general indication of the mean peak values in the input buffer to which the actual peak value of the sub-buffer can be compared. The method described below is a method of doing the smoothing.
To smooth the data, each 64-sample sub-block is scanned for the signal value of maximum absolute data. The absolute maximum data signal value is then used to calculate a smoothed moving average peak value. Filtered high-frequency moving averages for each k-th sub-buffer hi_mavg (k) respectively are calculated using equations 1 and 2.
for buffer k = 1: 1: 64 hi_mavg (k) = hi_mavg (k-1) + (high frequency peak value in buffer k) - h¡_ mavg (k-1)) * AVG_WHT) (1 ) end where hi_mavg (0) is set equal to hi_mavg (64) of the previous input buffer for continuous processing. In the current implementation, the AVG: WHT parameter is set equal to 0.25. This value was decided after following an experimental analysis using a wide range of common audio material.
The transient detection process then compares the peak value in each sub-block with the set of smoothed moving average peak values to determine if a transient exists. Although there are a number of methods for comparing these two measurements, the solution outlined below was taken because it allows the comparison to be tuned through the use of a scale factor that has been configured to perform under optimal conditions as determined by an analysis of a wide range. range of audio signals.
The peak value in the k-th sub-block, for the filtered data, is multiplied by the high-frequency scaling value HI_FREQ_SCALE, and compared to the smoothed calculated moving average peak value of each k.
ES 2 298 394 T3
If a sub-block scaled peak value is greater than the moving average value, a transient is signaled as being present. These comparisons are outlined later in Equations 3 and 4.
for buffer k = 1: 1: 64 if (((high frequency peak value in buffer k) * (HI_FREQ_SCALE)> hi_mavg (k)) (2) signal high frequency transient in sub-block k = TRUE end end
Following the detection of transients, several corrective checks were made to determine if the transient signaling for a 64-sample sub-block should be removed (TRUE to FALSE reset). These checks were performed in order to reduce false transient detections. First, if the high-frequency peak values fall below a minimum peak value, then the transient is removed (to cater for low-level transients). Second, if the peak value in a sub-block triggers a transient, but is not significantly higher than the previous sub-block, which would also have triggered a transient signaling, then the transient present in the sub-block is removed. current. This reduces a deterioration of information at the location of a transient.
Referring again to Figure 6, the next step 208 in the process is to determine whether there are transients in the current N sample input data set. If there are no transients, the input data could be downloaded as output (or re-passed back to a low data rate audio encoder) without time scaling treatment performed. If transients do exist, the number of transients that exist in the current N samples of audio data and their location (or locations) are passed to the time-scaled audio treatment part 210 of the process for temporal modification of input audio data. The result of a suitable time scale treatment is set forth in connection with the description of Figures 8A-8E. Note that the process requires information from the encoder as to, for example, the location of the windowed sample blocks with respect to the audio data signal stream. If, optionally, the metadata information with time scaling is downloaded as output (as shown in Figure 6), in the event that there are no transients, it would indicate that no pre-treatment has been performed. The time-scaled metadata could include, for example, time-scaled parameters such as the location and the amount of time-scaling performed and, if the time-scaling technique has employed the gradual transition of spliced audio segments, the length of gradual transition. The metadata contained in the encoded audio-frequency bitstream could also include information about transients, including their location after and / or before and after a temporal change. The audio data is output in step 212.
Audio frequency pre-treatment
Figures 8a-8e illustrate an example of time-scaled audio pre-treatment in accordance with aspects of the present invention when there is a transient in an audio coding block that is located closer to the end of the last windowed block than to the end of the window. end of next window block. For this example, a 50% block overlap is assumed, as in Figures 1a-1e and Figures 4a and 4b. As noted above, to reduce the magnitude of transient pre-noise introduced by low bit rate audio coding, it is desired to adjust the time evolution of the input audio signal such that the transient signal is located closely following the end of the last window block. Such a change in transient location is preferred, because it minimizes disruption to the time evolution of the signal train while optimally limiting the length of the transient pre-noise. However, as discussed above, a change to location that closely follows the end of the next windowed block also optimally limits the length of the transient pre-noise but does not minimize the disruption to evolution in the signal train time. In some cases, the difference in interruption may be of little or no audible significance, particularly if compensation for evolution over time is also used. Thus, a change to either of the two closest block ends is contemplated in the present example and other examples herein. As mentioned above, the transient time that changes the time scaling need not be fulfilled within a single block, unless the processing is carried out after the encoder has divided the audio signal stream into blocks.
Figure 8a shows three consecutive windowed coding blocks overlapped by 50%. Figure 8b presents the relationship between the original input audio data stream, which contains a single transient, and
ES 2 298 394 T3 windowed audio-frequency coding blocks. The beginning of the transient is T samples after the end of the preceding block. Since the transient is closer to the end of the preceding block than to the end of the next block, it is preferred to shift the transient to the left to a location that closely follows the end of the preceding block by applying time compression which has the effect to eliminate the T samples prior to the transient. Figure 8c presents two regions in the audio train where audio frequency time scaling could be performed. The first region corresponds to the audio samples located before the transient, where reducing the duration of the audio frequency by T samples "slides" or changes the position of the transient from the left to the desired location closely following the end of the transient. preceding block by providing time compression. As seen in Figures 2A to 5B and other figures to be described later, the spacing of the transient from the block end in Figures 8d and 8e has been exaggerated for clarity of presentation. The second region shows the region where time scaling could optionally be performed after the transient to increase the duration of the audio frequency by T samples by providing time expansion, such that the total length of the audio data remains at N samples. Although the removal of T samples and the optional addition of T sample number compensation has been shown to occur within a windowing audio encoding sample block, this is not essential - the scaling compensation process time does not need to occur within a single audio coding block, unless the transient time shift occurs after the encoder has divided the audio signal stream into blocks. The optimal location for such a timescale process could be determined by the timescale move process used. Since the transient could provide useful post-masking, preferably the time scaling with compensation for the number of samples is performed very close to the transient.
Figure 8d demonstrates the resulting signal stream if time scaling processing is performed on the input audio data stream by reducing the time duration of the input audio data stream by T samples in the area located before the transient and no time scale expansion is performed with compensation for number of samples after the transient signal. As discussed above, small variations in the time course of an audio signal are not discernible for most listeners. Thus, the number of time-scaled audio data stream samples is not required to equal the number of input samples, N; it might be sufficient to just treat the audio train before the transient. Figure 8e illustrates the case when the audio data stream before the transient is reduced in duration by T samples and the audio data stream following the transient is increased by T samples, thereby keeping N audio-frequency samples inside and outside the time-scaled processing block and restoring the evolution in time of the audio-frequency signal train except for the transient and the parts of the signal train very close to the transient. The variations in signal wavelengths of Figures 8b-8e are intended to show schematically that the number of samples contained in the audio data stream varies for the described conditions. When reducing the number of audio samples, as in Figure 8d, additional samples may need to be acquired before additional audio coding can be performed. This could mean extracting more samples from a file or waiting for more audio signals to be buffered in a real-time system.
Figures 9a-9e illustrate an example of time-scaled audio frequency processing when a transient exists in a windowed audio coding block and is positioned approximately T samples ahead of a block end. To reduce the amount of transient pre-noise introduced by low bit rate audio coding while minimizing transient shift, it is preferred to temporarily adjust the input audio signal such that the transient follow the end of the next block very closely. In the case of 50% overlapping blocks, a shift to the end of the next block end (or the end of the previous block) limits the pre-noise of the transient to the first half of an audio coding block, rather than disperse the pre-noise of the transient throughout the entire block and the previous audio frequency block.
Figure 9a shows three consecutive windowed coding blocks, overlapped by 50%. Figure 9b shows the relationship between the original input audio data, which contains a single transient, and the audio blocks. The beginning of the transient is T samples before the end of the next block. As the transient is closer to the end of the next block than it is to the end of the previous block, it is preferred to shift the transient to the right to a location that closely follows the end of the next block by applying a time expansion that has the effect of adding T samples before the transient. Figure 9c shows two regions where audio frequency time scaling could be performed. The first region corresponds to the audio samples placed before the transient, where increasing the duration of the audio frequency in T samples slides the position of the transient to the desired location very close after the end of the next block. Figure 9 also presents the region in which time scaling could be performed after the transient, to reduce the duration of the audio frequency in T samples, such that the total length of the audio data stream, N samples, remains constant. Figure 9d demonstrates the result if time scaling processing is performed on the input audio data stream by increasing the time duration of the input audio data stream by T samples in the time region before of the transient but without performing a time scale expansion with compensation for the number of samples after the transient signal. As discussed above, for most of the
ES 2 298 394 T3 listeners are not discernible small variations in the time evolution of an audio signal. Therefore, the number of audio stream samples after the time scaling is not required to be equal to the input, N. It might be sufficient to treat the audio frequency before the transient.
Figure 9e illustrates the case when the audio frequency before the transient is increased in duration by T samples and the audio frequency following the transient is reduced by T samples, thereby maintaining a constant number of audio samples before and after the time scaling. . As in the other figures, the block end transient spacing of Figures 9d and 9e has been exaggerated for clarity of presentation.
Time-scaled audio frequency treatment for multiple transients
Depending on the length of the audio encoding block size and the content of the audio data being encoded, it is possible that an input audio data stream being processed contains within the N samples being processed. , more of a transient signal that could introduce pre-noise artifacts. As mentioned above, the N samples being processed could include more than one audio coding block.
Figures 10a-10d illustrate treatment solutions when two transients occur in an audio coding block. In general, two or more transients could be handled in the same way as a single transient, with the earliest transient in the audio data stream being treated as the transient of interest.
Figure 10a shows three consecutive windowed coding blocks, overlapped by 50%. Figure 10b shows the case where two transients contained in the input audio frequency fork the end of an audio frequency coding block. For this case, the earliest transient introduces the most perceptible pre-noise, because a part of the pre-noise that results from the second transient is post-masked by the first transient. To minimize pre-noise artifacts, the input audio signal could be time-scaled to shift the first transient to the right in such a way that the audio frequency before the first transient has been expanded on the T-time scale. samples, where T is the number of samples that brings the first transient to a position that closely follows the end of the next block.
In order to compensate for number of samples for the time scale expansion treatment before the first transient of Figure 10b and to optimize the post-masking of the pre-noise resulting from the second transient by shifting the transients very closely together In time, the audio signal following the first transient and before the second transient is preferably time scaled to reduce in duration by T samples. As illustrated in Figure 10b, there is sufficient audio frequency treatment data between the first and second transients to perform the time scale treatment. However, in some cases the second transient may be so close to the first transient that there is not enough audio data to perform the time scale treatment between them. The amount of audio data required between transients depends on the time scaling process used for the treatment. If insufficient audio data exists between the two transients, it may be necessary to timescale the audio data following the second transient in order to provide compensation for the number of samples. In order to perform expansion of the audio data after the second transient, it might be necessary for the time scaling process to have access to a wider segment of audio data than the number of samples contained in a block used in the audio encoding process, as mentioned above.
Figure 10c illustrates the case in which the first transient is closer to the end of the last block than to the end of the next block and all transients (in this case two) are close enough together that the pre-noise resulting from the first transient is substantially post-masked by the first transient. In this way, the audio stream before the first transient is time-scale compressed by T samples, such that the first transient is switched to a location just after the end of the previous block. Compensation for number of samples to restore the original number of samples, in the form of time scale expansion, could be done on the audio data stream that follows the second transient.
Figure 10d illustrates the case in which the first transient is closer to the end of the next block than to the end of the previous block and all the transients (in this case, two) are close enough to each other that the pre-noise resulting from the second it is substantially post-masked by the first transient. In this way, the audio stream before the first transient is time-scaled by T samples, such that the first transient is switched to a location just after the end of the next block. Compensation by number of samples, in the form of timescale compression, could optionally be performed on the audio frequency data stream that follows the second transient.
For the case of multiple transients, if it is desired to compensate for evolution in time to pre-treat in an almost perfect way, metadata information could be transported with each audio-frequency block encoded in a similar way to the case of a single transient previously described. .
ES 2 298 394 T3
Compensation for evolution over time, controlled by metadata, pre-treatment with time scaling
As mentioned above, it could be convenient to apply, subsequent to the inverse transformation by a decoder, a compensatory time scaling to the audio-frequency signal train after the transient, in such a way that the evolution in time of the audio-frequency signal train treated is substantially the same as that of the original audio signal train, thereby restoring the original time evolution of the signal train. However, in experimental studies it has been shown that the majority of listeners do not perceive small temporal modifications of the audio frequency, and therefore compensation for evolution of time may not be necessary. Also, on average, transients advance and retard equally and, therefore, over a sufficiently long period of time, the cumulative effect without compensation for time evolution could be negligible. Another idea to consider is that, depending on the type of time scaling used for the pre-treatment, the additional time evolution compensation processing could introduce audible artifacts in the audio frequency. These artifacts could arise because time-scaling processing, in many cases, is not a perfectly reversible process. In other words, reducing the audio frequency by a fixed amount using a time scaling process and then expanding in time the same audio frequency could introduce audible artifacts.
An advantage of audio-frequency processing containing time-scaling transient material is that the time-scaling artifacts could be masked by the time-masking properties of the transient signals. An audio transient provides temporary forward and backward masking. Transient audio material "masks" audible material both before and after the transient, such that the audio that directly precedes and follows is not perceptible to a listener. Pre-masking has been measured, is relatively short and lasts only a few milliseconds, while post-masking could last more than 100 milliseconds. Therefore, treatment with compensation for time evolution and time scaling could be inaudible due to temporary post-masking effects. Thus, if performed, it is advantageous to perform time evolution compensation and time scaling within temporally masked regions.
Figures 11a-11f show an example where an intelligent time evolution compensation has been performed that follows an inverse transformation in the decoder using metadata information. Metadata greatly reduces the amount of analysis required to perform time evolution compensation, because it indicates where the time escalation treatment should take place and the duration of the time escalation required. As explained above, the time evolution compensation processing is intended to return the decoded audio signal to its original time course in which the signal stream, including the transient, has its original location in the audio stream . Figure 11a shows three consecutive 50% overlapping windowed coding blocks. Figure 11b presents a stream of audio input signals before pre-treatment having a transient T samples after a block end. Figure 11c shows that the input audio signal stream is treated by removing T samples before the transient to shift the transient to an earlier location. The T samples are added after the transient in order to leave the number of samples of audio data unchanged (compensation of number of samples). Figure 11d presents the modified audio signal train in which the transient has been changed to an earlier location and the audio frequency following the transient has been changed back to its original location. Figure 11e shows the required time scaling and time evolution compensation regions in which the removal of T samples (time compression) is compensated for by adding T samples (time expansion) and the addition of T samples (time expansion). time) is compensated by removing T samples (time compression). The result, presented in Figure 11f, is a "near perfect" output signal that has the same time evolution as the input signal of Figure 11a (subject mainly to imperfections in time scaling processes).
Post-treatment with time scaling to reduce transient pre-noise
As demonstrated in a number of previous examples, even with the optimal location of a transient in an audio coding block, some pre-noise is still introduced by the low bit rate audio coding system process. . As noted above, longer audio coding blocks are preferable over shorter coding blocks, because they provide higher frequency resolution and increased coding gain. However, even if the transients are optimally located by time scaling before audio coding (pre-treatment), as the length of the audio coding block increases, so does the pre-noise. Pre-masking of transient temporal pre-noise is on the order of 5 milliseconds, which corresponds to 240 audio samples sampled at 48 kHz. This implies that, for encoders with block sizes greater than about 512 samples, the transient pre-noise begins to be audible even with optimal location (only half is masked in the case of the 50% overlapped block). (This does not take into account the transient pre-noise reduction caused by window edge effects in the encoder blocks.)
Although transient pre-noise cannot be totally eliminated from a low bit-rate coding system, it is possible to perform a time-scaled post-treatment (alone or in addition to a pre-treatment) on data. audio frequency devices that have undergone reverse transformation in a decoder
ES 2 298 394 T3 transformation based low bit rate audio frequency to reduce the amount of transient pre-noise whether pre-treatment is also applied or not. The time-scaled post-treatment could be performed either in conjunction with a low bit rate audio decoder (i.e. as part of the decoder and / or by receiving metadata from the decoder and / or encoder via the decoder) or as a stand-alone after-treatment. The use of metadata is preferred because useful information such as the location of transients with respect to audio coding blocks, as well as the audio coding block length (s) are readily available and could be passed to the post process. -treatment through metadata. However, post-treatment could be used without interaction with a low bit rate audio decoder. Both methods are described later.
Post-treatment with time scaling in conjunction with a low bit rate audio decoder (receiving metadata)
Figure 12 is a flow diagram of a process for performing time scaling post-processing in conjunction with a low bit rate audio decoder to reduce transient pre-noise artifacts. The process illustrated in Figure 12 assumes that the input data is low bit rate encoded audio data (step 802). Following decoding of the compressed data to an audio signal (step 804), the audio signal corresponding to a block (or blocks) is sent to the time scaler (step 806) along with metadata information that is useful for reducing the duration of transient pre-noises. This information could include, for example, the location of transients, the length of the audio encoder block (or blocks), the ratio of the encoder block limits to the audio data, and a desired length of the preset. transient noise. If the location of the transients relative to the audio encoder block limits is available, the location of the pre-noise artifact could be accurately estimated and reduced by post-treatment. Since transients do provide some temporary pre-masking, it may not be necessary to completely eliminate transient pre-noise. By giving the time-scaled post-treatment process a desired length of pre-noise, some control over the amount of pre-noise remaining at the audio output could be achieved by step 808. The results of a suitable treatment with timescale for step 806 are described below in connection with the description of Figures 13a-13c
Note that post-treatment could be useful whether or not a pre-treatment has been applied prior to encoding. Regardless of where the transient is located with respect to the block ends, there is some transient pre-noise. For example, it is at least half the length of the audio coding window for the 50% overlap case. Larger window sizes could still introduce audible artifacts. By performing post-processing, it is possible to reduce the length of the pre-noise even further than it has been reduced by optimally locating the transient with respect to the block ends prior to quantization by the encoder.
Figures 13a-13c illustrate an example of post-treatment for a single transient in order to reduce the pre-noise artifact present after the inverse transformation. Depending on the length of the coding block, the pre-noise, even after pre-treatment, if any, could have a longer time that could be masked by the temporal masking effects of the transient. However, as shown in Figure 13b, by using the transient location metadata information from the decoder, an audio frequency region containing the pre-noise could be identified in which pre-noise could be reduced in length by time-scaling the audio signal to reduce pre-noise per T samples. The number T could be chosen in such a way that the length of the pre-noise is minimized to take advantage of the pre-masking, or it could be chosen in order to eliminate the pre-noise completely or almost completely. If it is desired to keep the same number of samples as in the original signal, the audio signal following the transient could be time-scaled by + T samples. Alternatively, as shown in connection with the example of Figure 16A, such sample number compensation could be applied before pre-noise, which has the advantage of also providing a time evolution compensation.
It should be noted that, if the post-treatment is performed in conjunction with the time-scaled pre-treatment, the amount of additional disruption to the time evolution of the output audio signal train could be minimized. As the time-scaled pretreatment described above reduces the pre-noise length to N / 2 samples for the case of a 50% flap (where N is the length of the audio-frequency coding block), the introduction of less than N / 2 samples of additional time evolution interruption in the output audio frequency compared to the original input audio signal. In the absence of pre-treatment, the pre-noise can reach up to N samples, the length of the coding block for the case of a 50% flap.
In some low bit rate audio coding systems, the location of the signal transients may not be readily available if the encoder does not carry the location information. If that is the case, the decoder or time scaling process could, using any number of transient detection processes or the efficient method described above, perform transient detection.
ES 2 298 394 T3
For multiple transients, the same concepts apply as for pre-treatment, as described above.
Post-treatment with time escalation without pre-treatment
As mentioned above, in some cases it might be desirable to improve the perceived quality of an audio signal that has undergone low bit rate coding using compression systems that do not implement pre-noise time-scaling processing. transitory (pretreatment). Figure 14 schematizes a process to carry it out.
The first step 1402 checks the availability of N audio data samples that have undergone low bit rate audio frequency encoding and decoding. These audio data samples could belong to a file on a PC hard drive or to a data buffer on a hardware device. If N samples of audio data are available, they are passed to the time-scaled post-treatment process by step 1404.
The third step 1406 in the timescaled post-treatment process is the identification of the location of transient audio data signals that are likely to introduce pre-noise artifacts. Many different processes are available to perform this function, and their specific implementation is not important as long as it provides accurate detection of transient signals that are likely to introduce pre-noise artifacts. However, the process described above is an effective and accurate method that could be used.
The fourth step 1408 is to determine if there are transients in the current N-sample input data pool as detected by step 1406. If there are no transients, the input data could be downloaded as output by step 1414 without performing any time escalation treatment. If transients exist, the number of transients and their location (or locations) are passed to the transient pre-noise estimation process step (1410) of the process to identify the location and duration of the transient pre-noise.
The fifth and sixth stages (1410) in the treatment involve estimating the location and duration of transient pre-noise artifacts and reducing their length with time-scaled processing 1412. Since, by definition, pre-noise artifacts noise are limited to the regions that precede transients in the audio frequency data, the scan area is limited by the information provided by the transient detection process. As shown in Figure 1, the pre-noise length is limited from a minimum of N / 2 to a maximum of N samples, where N is the number of audio samples in an audio coding block overlapped by 50 %. Thus, when N are 1,024 samples and the audio frequency is sampled at 48 kHz, the transient pre-noise could range from 10.7 msec. up to 21.3 msec. prior to the onset of the transient, depending on the location of the transient in the audio signal train, which significantly exceeds any temporal masking that might be expected from transient signals. Alternatively, instead of estimating the length of the pre-noise artifacts preceding a transient, step 1410 could be applied assuming the pre-noise artifacts have a default length.
Two solutions could be implemented for transient pre-noise reduction. The first assumes that all transients contain pre-noise, and therefore the audio signals before each transient could be time-scaled (compressed in time) by a predetermined amount (default) based on an expected magnitude. transient pre-noise. If this technique is used, a time scale expansion of the audio frequency could be performed before the temporal pre-noise, to provide a compensation for the number of samples for the time-scaling process with time compression used to reduce the length of the audio. pre-noise, and to provide compensation for time evolution (the time expansion before the pre-noise that compensates for the time compression within the pre-noise leaves the transient at or near its original temporal location). However, if the exact location of the pre-noise is not known, such a sample number compensation process could inadvertently increase the duration of parts of the pre-noise component.
Figures 15a-15c demonstrate a technique that uses a default value to time-scale the audio signal before each transient in order to reduce the pre-noise duration, but compensation for number of samples is not performed. As shown in Figure 15a, an audio signal from a low bit rate audio decoder has a transient preceded by a pre-noise. Figure 15b shows a default processing length that is used as the amount of time compression to be performed by the time scaling process. Figure 15c shows the resulting stream of audio signals having reduced pre-noise. In this example, time evolution compensation has not been performed to return the transient to its original location in the audio data stream. However, in a similar way to the previous treatment examples, if a constant number of samples from input to output is desired, a time scale expansion process could be performed following the transient, similar to the example in Figure 13b or possibly before pre-noise as described below in connection with the example of Figures 16a-16c. However, when applying a default processing length, the provision of such compensation before pre-noise runs the risk of performing the time scale expansion process within the pre-noise (thus undesirably increasing the pre-noise). pre-noise length) if the actual pre-noise length exceeds the default length. In addition, in some cases, the post-treatment may not have access to the
ES 2 298 394 T3 audio train before pre-noise - audio frequency could have already been downloaded as output in order to reduce waiting time.
A second technique of pre-noise reduction with post-treatment, illustrated in Figures 16a-16c, involves performing an analysis of the pre-noise resulting from a transient to determine its length and process the audio frequency so that only the transient is treated. pre-noise segment. As noted above, transient pre-noise occurs when the high-frequency components of the audio-frequency transient material are temporarily contaminated throughout an entire block as a result of the quantization process performed in the encoder. Therefore, a simple method of detection is to high-pass filter the audio frequency before a transient and measure the high-frequency energy. The onset of transient pre-noise is identified when the high-frequency noise-like pre-noise related to and caused by the transient exceeds a predetermined threshold value. When the size and location of the transient pre-noise are known, a timescale compensated expansion could be performed prior to the pre-noise timescale downscaling to return the audio signal to its original time course and restore the time evolution of the audio signal train substantially to its original condition. The invention is not limited to employing high frequency detection. Other techniques could be used to detect or estimate the length of the pre-noise.
In Figure 16a, a stream of audio signals from a low bit rate audio decoder has a transient preceded by a pre-noise. Figure 16 shows a time compression treatment length that is used as the amount of timescaled reduction to be performed by the timescaling process based on an estimated pre-noise length measured by high-frequency audio content. frequency in the block. Figure 16b also presents the use of time expansion by T samples in order to restore the original time evolution of the signal train and also to restore the original number of samples. Figure 16c shows the resulting audio signal stream having reduced pre-noise along with the original time course and the same number of samples as the original signal stream.
The present invention and its various aspects could be implemented as software functions implemented in digital signal processors, general purpose programmed digital computers, and / or special purpose digital computers. The interfaces between the analog and digital signal trains could be realized in appropriate hardware and / or as functions in software and / or in microprogram.
Contents10
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8063809B2 | Cited by | United States of America | Applicant |
22 members in 14 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 20010290286P | United States of America | – | |
| 29028601 | United States of America | P | |
| 29028601 | United States of America | P | |
| 02769666290286P | – | – | – |
| US20010290286P | – | – | – |
Members22
| Document | Office | Kind | |
|---|---|---|---|
| CA2445480A1 | Canada | A1 | |
| WO02093560A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1386312A1 | European Patent Office (EPO) | A1 | |
| MXPA03010237A | Mexico | A | |
| KR20040034604A | Republic of Korea | A | |
| US2004133423A1 | United States of America | A1 | |
| JP2004528597A | Japan | A | |
| CN1552060A | China | A | |
| HK1070457A1 | Hong Kong, China | A1 | |
| CN1312662C | China | C | |
| US7313519B2 | United States of America | B2 | |
| AU2002307533B2 | Australia | B2 | |
| EP1386312B1 | European Patent Office (EPO) | B1 | |
| AT387000T | Austria | T | |
| ATE387000T1 | Austria | T1 | |
| DE60225130D1 | Germany | D1 | |
| ES2298394T3This record | Spain | T3 | |
| DK1386312T3 | Denmark | T3 | |
| DE60225130T2 | Germany | T2 | |
| JP4290997B2 | Japan | B2 | |
| KR100945673B1 | Republic of Korea | B1 | |
| CA2445480C | Canada | C |
Numbers
- Publication
- 2298394
- Publication, DOCDB
- 2298394
- Publication, EPODOC
- ES2298394T
- Application
- 2769666
- Application, DOCDB
- 02769666
- Application, EPODOC
- ES20020769666T
Titles2
- Spanish
- MEJORA DE SESIONES TRANSITORIAS DE SISTEMAS DE CODIFICACION DE SEÑALES DE AUDIOFRECUENCIA A BAJA VELOCIDAD DE TRANSFERENCIA DE BITS POR REDUCCION DE PRE-RUIDOS.
- English
- IMPROVING TRANSITIONAL SESSIONS OF LOW-SPEED AUDIO FREQUENCY SIGNAL CODING SYSTEMS FOR BIT TRANSFER DUE TO REDUCTION OF LOSSES.
Classification
- CPC, 5
- G10L19/02
- G10L21/04
- G10L19/0212
- G10L19/022
- G10L19/025
- IPC, 2
- G10L21 04
- G10L19 02