Device and method for determining an estimated value
Abstract
To the Determining an estimated value for a Need for information units for encoding a signal in addition to the permissible noise for a Frequency band and an energy of frequency bands in addition a Measure of the distribution the energy included in the frequency band. This is a better estimate for the Need for information units obtained so that efficient and can be coded accurately.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
13 claims: 11 independent, 2 dependent
- 1Claims Patentkrav 1 Device for determining an estimate (s) for a need for information units by encoding a signal comprising audio or video information, wherein 1 Anordning for bestemmelse av et estimat (pe) for et behov for informasjonsenheter ved koding av et signal som omfatter audio- eller videoinformasjon, der 5 the signal has a plurality of frequency bands (b), characterized in that it comprises means (102) for providing a target (nb (b)) for a permissible interference in a frequency band (b) to the signal, wherein the frequency band (b) comprises at least two spectral values in a spectral representation of the signal, and a measure (e (b)) for an energy in the signal in the frequency band, io · means (106) for calculating a target (nl (b)) for a distribution of the energy (e ( (b)) in the frequency band (b);wherein the distribution of the energy in the frequency band deviates from a completely uniform distribution, where the means (106) for calculating the target (nl (b)) is the distribution of the energy (e (b)) designed to determine, as a measure of the distribution of the energy, an estimate for a number of spectral values of a magnitude greater than or equal to a predetermined threshold, or of a magnitude less than or equal to this threshold;where the threshold value is an exact or estimated quantizer value which means that in a quantizer (1014), values less than or equal to the quantizer value in quantized to zero, and 5 signalet har flere frekvensbånd (b), karakterisert ved at den omfatter • midler (102) for å kunne tilveiebringe et mål (nb(b)) for en tillatt interferens i et frekvensbånd (b) til signalet, der frekvensbåndet (b) omfatter minst to spektralverdier i en spektralrepresentasjon av signalet, og et mål (e(b)) for en energi i signalet i frekvensbåndet, io · midler (106) for beregning av et mål (nl(b)) for en distribusjon av energien (e(b)) i frekvensbåndet (b), der distribusjonen av energien i frekvensbåndet avviker fra en fullstendig uniform distribusjon, der • midlene (106) for beregning av målet (nl(b)) er distribusjonen av energien (e(b)) er innrettet til å kunne bestemme, som et mål for distribusjonen av energien, et i5 estimat for et antall av spektralverdier med en størrelse som er større enn eller lik en forutbestemt terskelverdi, eller med en størrelse som er mindre enn eller lik denne terskelverdi, der terskelverdien er en eksakt eller estimert kvantisererverdi som medfører at i en kvantiserer (1014) vil verdier som er mindre enn eller lik kvantisererverdien i kvantisert til null, og 20 · Means (104) for calculating the estimate (s) using the target (nb (b)) for the interference, the target for the energy and the target for the distribution of the energy. 20 · midler (104) for beregning av estimatet (pe) ved å benytte målet (nb(b)) for interferensen, målet for energien og målet for distribusjonen av energien.
- 33 Apparatus according to one of the preceding claims, characterized in that the means (106) for calculating are arranged to be able to calculate a form factor according to the following equation:3 Anordning ifølge ett av de foregående krav, karakterisert ved at midlene (106) for beregning er innrettet til å kunne beregne en formfaktor i overensstemmelse med følgende likning: jM ») = £«) i, k = kOffset [b] where X (k) is a spectral value with frequency index k, where kOffset is a first spectral value with band b and where ffac (b) is the form factor. jM»)= £ «)i, k=kOffset[b] hvor X(k) er en spektral verdi med frekvensindeks k, der kOffset er en første spektralverdi med bånd b, og der ffac(b) er formfaktoren. 35 35
- 44 Device according to one of the preceding claims, characterized in that the means (106) for calculating are arranged to be able to include in the calculation a fourth root of the rate between the energy in the frequency band and a width of the frequency band or the number of spectral values in the frequency band. 4 Anordning ifølge ett av de foregående krav, karakterisert ved at midlene (106) for beregning er innrettet til å kunne ta med i beregningen en fjerde rot av raten mellom energien i frekvensbåndet og en bredde til frekvensbåndet eller antall spektralverdier i frekvensbåndet.
- 55 Device according to one of the preceding claims, characterized in that the means (106) for calculating are arranged to calculate the target for the distribution of the energy according to the following equations:5 Anordning ifølge ett av de foregående krav, karakterisert ved at midlene (106) for beregning er innrettet til å beregne målet for distribusjonen av energien i overensstemmelse med følgende likninger: nl (b) = ffacLb), ° g kOffset {b + \} - \ jM ») = Σ Vi ^ Wi. nl(b) = ffacLb) ,°g kOffset{b+\}-\ jM»)= Σ Vi^Wi. k = kOffset [b) where X (k) is a spectral value with frequency index k, where kOffset is a first spectral value in a band (b), where ffac (b) is a form factor, where nl (b) represents the measure of the distribution of the energy of the band b, where e (b) is a signal energy in the band b, and io width (b) is a width of the band. k=kOffset[b) der X(k) er en spektralverdi med frekvensindeks k, der kOffset er en første spektralverdi i et bånd (b), der ffac(b) er en formfaktor, der nl(b) representerer målet for distribusjonen av energien i båndet b, der e(b) er en signalenergi i båndet b, og der io width(b) er en bredde til båndet.
- 66 Device according to one of the preceding claims, characterized in that the means (104) for calculating the estimate are arranged to use a quotient of the energy in the frequency band and the interference in the frequency band. 6 Anordning ifølge ett av de foregående krav, karakterisert ved at midlene (104) for beregninger av estimatet er innrettet til å benytte en kvotient av energien i frekvensbåndet og interferensen i frekvensbåndet.
- 77 Apparatus according to any one of the preceding claims, characterized in that the means (104) for calculating the estimate are arranged to calculate the estimate using the following terms:7 Anordning ifølge ett av de foregående krav, karakterisert ved at midlene (104) for beregning av estimatet er innrettet til å beregne estimatet ved å benytte følgende uttrykk: pe = «/ (å) log 2 b pe = «/(å)-log 2 b ^ nb (b) where pe is the estimate, where nl (b) represents the measure of the distribution of the energy in band b, where e (b) is the energy of the signal in band b, where nb (b) is the allowable interference in band b and where s is an additive link preferably equal to 1.5. ^nb(b) der pe er estimatet, der nl(b) representerer målet for distribusjonen av energien i båndet b, der e(b) er energien til signalet i båndet b, der nb(b) er den tillatte interferens i båndet b, og der s er et additivt ledd fortrinnsvis lik 1,5. 25 25
- 88 Device according to one of the preceding claims, characterized in that the means (104) for calculating the estimate are arranged to calculate the estimate according to the following equation:8 Anordning ifølge ett av de foregående krav, karakterisert ved at midlene (104) for beregning av estimatet er innrettet til å beregne estimatet i overensstemmelse med følgende likning: pe = Σ -log 2 b k nb(b) pe = Σ -log 2 bk nb (b) 30 where nl (b) = ffac (b) (e (b) Υ> · 25 \ widtf {b) / and where kOffset (b + l) -lk = kOffset (b) where pe is the estimate that nl (b) represents the measure of the distribution of the energy in band b, where e (b) is the energy of the signal in band b, where nb (b) is the permissible interference in band b, where s is an additive link preferably equal to 1.5, where X (k) is a spectral value with frequency index k, kOffset being a first spectral value in the band b, ffac (b) being a form factor, and witdh (b) being a width of the band. 30 der nl(b) = ffac(b) ( e(b) Υ>·25 \widtf{b) / og der kOffset(b+l)-l k=kOffset(b) der pe er estimatet, der nl(b) representerer målet for distribusjonen av energien i båndet b, der e(b) er energien til signalet i båndet b, der nb(b) er den tillatte interferens i båndet b, der s er et additivt ledd fortrinnsvis lik 1,5, der X(k) er en spektralverdi med frekvensindeks k, der kOffset er en første spektralverdi i båndet b, der ffac(b) er en formfaktor, og der witdh(b) er en bredde til båndet.
- 99 Device according to one of the preceding claims, characterized in that the signal is 9 Anordning ifølge ett av de foregående krav, karakterisert ved at signalet er 5 given as a spectral representation of spectral values. 5 gitt som en spektralrepresentasjon av spektralverdier.
- 1010 A method for determining an estimate when required for information units by encoding a signal comprising audio or video information, wherein the signal has multiple frequency bands, characterized in that the method comprises the steps of:10 Fremgangsmåte for å bestemme et estimat ved behov for informasjonsenheter ved koding av et signal omfattende audio- eller videoinformasjon, der signalet har flere frekvensbånd, karakterisert ved at fremgangsmåten omfatter trinnene: io · å tilveiebringe (102) et mål (nb(b)) for en tillatt interferens i et frekvensbånd (b) til signalet, der frekvensbåndet omfatter minst to spektralverdier i en spektralrepresentasjon av signalet, og et mål (e(b)) for energien i signalet i frekvensbåndet (b), • å beregne (106) et mål (nl(b)) for en distribusjon av energien i frekvensbåndet (b), providing (102) a target (nb (b)) for a permissible interference in a frequency band (b) of the signal, the frequency band comprising at least two spectral values in a spectral representation of the signal, and a target (e (b)) for the energy of the signal in the frequency band (b);• calculating (106) a target (nl (b)) for a distribution of the energy in the frequency band (b);
- 1115 where the distribution of the energy in the frequency band deviates from a completely uniform distribution, where an estimate of a number of spectral values with a magnitude greater than or equal to a predetermined threshold value or with a magnitude less than or equal to this threshold value is determined as a measure (nl (b)) for the distribution of energy, where the threshold value is an exact or estimated quantizer value 15 der distribusjonen av energien i frekvensbåndet avviker fra en fullstendig uniform distribusjon, der et estimat for et antall av spektralverdier med en størrelse som er større enn eller er lik en forutbestemt terskelverdi eller med en størrelse som er mindre enn eller lik denne terskelverdi bestemmes som et mål (nl(b)) for distribusjonen av energien, der terskelverdien er en eksakt eller estimert kvantisererverdi
- 1220 which means that values less than or equal to the quantizer value are quantized to zero in a quantizer (1014), and • to calculate (104) the estimate (s) by using the measure (nb (b)) for the interference, the target (e) (b)) for the energy, and the target (nl (b)) for the distribution of the energy. 20 som medfører at verdier som er mindre enn eller lik kvantisererverdien, kvantiseres til null i en kvanti serer (1014), og • å beregne (104) estimatet (pe) ved å benytte målet (nb(b)) for interferensen, målet (e(b)) for energien, og målet (nl(b)) for distribusjonen av energien.
Independent claims11
88 paragraphs in 2 sections, as filed
(74) Agent
<td> (54) (56)</td><td>Designation cited publications</td>
<td> (57)</td><td>Summary</td>
Apparatus and method for making an estimate
EP 0 446 037
US 2002103637
In order to determine an estimated value related to an information unit requirement by encoding a signal, a measure (nl (b)) for the distribution of the energy in the frequency band is taken into account (102, 104, 106) as well as the allowable interference in the frequency band. and the energy of said frequency band. In this way, a better estimate of the information unit requirement is obtained, so that the signal can be coded more efficiently and accurately.
<img file="NO338917B1_D0001.tif" />
Discipline
The invention relates to encoders for encoding a signal comprising audio and / or video information, and in particular the estimation of the need for information units for encoding this signal.
Background
Prior art is presented in EP 0 446 037 which presents a hybrid coding technique for high quality audio signal coding, using a subband filter technique further enhanced to obtain a large number of sub bands. Sub-band noise masking thresholds are then determined using a new tonality measure applicable to individual frequency bands or frequencies. Based on the thresholds determined in such a way, incoming signals are encoded to achieve high quality with reduced bit rates.
US 2002 103637 discloses digital audio coding systems using high-frequency reconstruction methods. The document teaches how the overall performance of such systems is improved by adjusting over time the crossover frequency between the low band coded by a core codec and the high band coded by an HFR system.
Different ways of establishing instantaneous optimal crossover frequency choices are introduced.
In the following, a code according to the prior art will be presented. A 20 audio signal to be coded is output to input 1000. This audio signal is first fed to a scaling step 1002, where a so-called AAC gain control is performed to establish the level of the audio signal. Side information of the scaling of is given to a bitstream formatter 1004, represented by the arrow between block 1002 and block 1004. The scaled audio signal is then output to an MDCT filter bank 1006. Using the AAC encoder, the filter bank will infiltrate a modified, discrete cosine transformation with 50% overlapping windows, where the window length is determined in block 1008.
The block 1008 is generally provided to be able to window transient signals with relatively short windows, and to be able to window share signals that tend to be stationary with relatively long windows. Thus, a higher level of time resolution (at the expense of frequency resolution) of the transient signal is achieved, due to the relatively short windows, while for signals that tend to be stationary, a higher frequency resolution (at the expense of time resolution) is achieved. because of longer windows, which is often preferred since it provides a higher coding gain. At the output of filter bank 1006, successive blocks of spectral values may be emitted which may be MDCT coefficients, Fourier coefficients or subband signals, depending on the implementation of the filter bank, where each subband signal has a specific limited bandwidth specified by the respective subband channel in the source band 100, comprises a specific number of subband samples.
In the following, an example will be presented of an instance where the filter bank delivers temporal, successive blocks of MDCT spectral coefficients, which generally represent successive, short-term spectra of the audio signal to be encoded, to the input 1000. A block of MDCT spectral values is then fed. to a TNS5 processing block 1010 (TNS = temporary noise shaping), where a temporary noise shaping will be performed. The TNS technique is used to shape the temporal shape of the quantization noise in each window of the formation. This is achieved by applying a filtering process to parts of the spectral data for each channel. The coding is then performed on a window basis. More specifically, the following steps are performed when applying the TNS tool to a window of spectral data, i.e. to a block of spectral values.
First, a frequency range is selected for the TNS tool. A suitable choice would be a frequency range of 1.5 kHz with a filter, up to the highest possible scaling factor band. It should be noted that this frequency range will depend on a sampling rate, as specified in the AAC standard (ISO / IEC 14496-3: 2001 (E)).
Then, an LPC (linear predictive coding) calculation is performed, more precisely using the spectral MDTC coefficients present in the selected target frequency range. In order to achieve increased stability, coefficients corresponding to frequencies below 2.5 kHz will be omitted from that process. Common LPC procedures derived from speech processing can be used for the LPC calculation, such as the known
Levinson-Durbin algorithms. The calculation is performed for the maximum permissible way of the noise forming filter.
In the LPC calculation, the expected prediction gain PG will be obtained. In addition, the reflection coefficients or Parcor coefficients are provided.
If the prediction gain does not exceed a certain threshold, it will
The TNS tool is not used. In this case, control information will be written into the bitstream so that a decoder will know that no TNS processing has been performed.
However, if the prediction gain exceeds this threshold, the TNS processing will be used.
In the next step, the reflection coefficients are quantized. The order of the noise forming filter used is determined by removing all the reflection coefficients with an absolute value less than a threshold value measured from the 'tail' to the range of reflection coefficients. The number of residual reflection coefficients will be on the order of the noise forming filter. An appropriate threshold will be 0.1.
The remaining reflection coefficients would typically be transformed into linear prediction coefficients, and this technique is also known as an '' up-transformation '' procedure.
The calculated LPC coefficients are then used as noise forming filter coefficients for the encoder, i.e. as prediction filter coefficients. This FIR filter will be used with the filtration in the specified target frequency range. An autoregressive filter is used in the decoding, while a so-called moving averaging filter is used in the coding. Then, the TNS tool page information is output to the bitstream formatter, as indicated by the arrow between the TNS processing block 1010 and the bitstream formatter 1004 of FIG. 3.
Thereafter, several optional tools not shown in FIG. 3, such as a long-term prediction tool, an intensity / switching tool, a prediction tool, a noise replacement tool, until finally the signal arrives at an averaging / side encoder 1012 will be active when the audio signal is to be coded is a multi-channel signal, i.e. a stereo signal with a left channel and a right channel. Thus far, that is, upstream from block 1012 in FIG. 3, the left and right stereo channels have been processed, i.e. scaled, transformed by the filter bank, whether they have undergone a TNS processing or not, etc., separated from each other.
In the mediation / page encoder it will be essentially verified if the mediation / page coding makes sense, ie if it will give a coding gain whatsoever.
The mid-side coding will give a coding gain that the left and right channels tend to be equal, since the middle channel, i.e. the sum of the left and right channels, in this case is almost equal to the left channel or right channel , except for scaling by a factor of 1/2, while the side channel will only assume very small values since it is equal to the difference between a left and a right channel. As a consequence, it is obvious that when the left and right channels are about the same, the difference will be about zero, or assume only very small values that will hopefully be quantized to zero in a subsequent quantization 1014, and they can thus, transmitted in a very efficient manner, since an antopi encoder 1016 is located downstream of the quantizer 1014.
The quantizer 1014 is provided with a permissible interference-per-scaling factor band at a psychoacoustic portion 1020. The quantizer is operated in an iterative manner, i.e., an external iteration loop will first be invoked, which will then invoke an internal iteration loop. In general, a quantization of a block of values will first be performed at the input of quantizer 1014, starting with initial values for the step size of the quantizer. In particular, the inner loop will quantify the MDCT coefficients, then specifically the number of bits used in this process. The outer loop will calculate the distortion and the modified energy for the coefficients using the scaling factor and then invoke an inner loop. This process is repeated until a specific condition is met. For each iteration in the outer iteration loop, the signal will be reconstructed to then calculate the interference caused by the quantization, and then compare it with the allowed interference indicated by the psychoacoustic portion 1020. In addition, the scaling factor for the frequency bands still considered after this comparison will be interfered be increased by one or more steps for each iteration, more precisely for each iteration in the outer iteration loop.
As soon as a situation arises where the interference caused by the quantization is below the permissible interference determined by the psychoacoustic part, and if the bit requirements are met at the same time, it will precisely say that the maximum bit rate is not exceeded, the iteration, ie. the analysis-by-synthesis method is terminated and the scaling factors obtained are encoded as illustrated in block 1014, whereupon they are output in encoded form to bitstream formatter 1004 as indicated by the arrow between block 1014 and block 1004. The quantized values will then be delivered to an entropy encoder 1016, which will typically perform an entropy coding for different scaling factor bands using multiple Huffman code tables, thus transferring the quantized values to a binary io format. As is well known, entropy coding in the form of Huffman coding includes relying on coding tables that are set up on the basis of expected signal statistics, and where frequently occurring values are given a shorter code word than less frequently occurring values. The entropy-encoded values are then provided as main information to the bitstream formatter 1004, which in turn will output the encoded audio signal on the output side in accordance with a specific bitstream syntax.
Audio signal data reduction now includes known techniques that are subject to a variety of international standards (eg ISO / MPEG-1, MPEG-2, AAC, MPEG-4).
The above-mentioned methods have in common that the input signal is given a compact, data-reduced representation by means of a so-called encoder, where it benefits from perception-related effects (psychoacoustics, psycho-optics). To achieve this, a spectral analysis of the signal is usually performed, whereupon the corresponding signal components are quantized taking into account a perception model, and coded as a so-called bitstream in the most compact way possible.
In order to estimate, before the actual quantization, how many bits a particular signal part to be encoded will require, so-called perceptual entropy (PE) can be used. PE will also provide a measure of how difficult it is for the encoder to encode a particular signal or parts thereof.
The deviation of PE from the actual number of bits required will be critical to the quality of the estimate.
The perceptual entropy and / or each estimate of the need for information units for encoding a signal can further be used to estimate whether the signal is transient or stationary, since the transient signals will require more bits of encoding than stationary signals. For example, the estimation of a transient property of a signal is used to be able to make a window length decision, as indicated in block
1008 in FIG. 3.
Fig. 6 shows the perceptual entropy calculated in accordance with ISO / IEC IS 13818-7 (MPEG-2 advanced audio coding (AAC)). The equation shown in FIG. 6 is used in the calculation of the perceptual entropy, ie a band-wise perceptual entropy. In this equation, the parameter pe represents the perceptual entropy. The width (b) further represents the number of spectral coefficients in the respective band B. Furthermore, e (b) is the energy of the signal in that band. Finally, nb (b) is the corresponding masking threshold, or more generally the permissible interference may be applied to the signal, e.g. by quantization, so that a listener will hear nothing, or just a minor interference.
The bands may be derived from the band splitting in the psychoacoustic model (block 1020 in Fig. 3), or they may be the so-called scaling factor bands (scfb) used in the quantization. The psychoacoustic masking threshold is the energy value that the quantization error should not exceed.
Thus, the illustration of FIG. 6 how well a perceptual entropy determined in this way will serve as an estimate of the number of bits required by coding. The respective perceptual entropy is plotted depending on the bits used in an AAC encoder at different bit rates for each individual block. The test piece used contains a typical mix of music, speech and individual instruments.
i5 Ideally, the points would accumulate along a straight line passing through the origin.
The distribution of points and deviations from the ideal line makes it clear that the estimate is inaccurate.
This deviation will thus be the disadvantage of the concept of FIG. 6, which will be the result of, for example, a value that is too high for the per20-septal entropy, which means that it will signal to the quantizer that more bits are needed than is actually needed. This will cause the quantizer to quantize too finely, ie. that it will not fully benefit from the target of the permissible interference, which will result in reduced coding gain. On the other hand, if the value of the perceptual entropy is set too low, it will be signaled to the quantizer that fewer bits are required than what is actually required for coding the signal. This, in turn, will result in the quantizer quantizing too roughly, which would immediately result in, if no countermeasures are taken, an audible interference in the signal. A countermeasure would be the quantifier requiring one or more iteration loops, which would increase the encoder computation time.
To improve the calculation of the perceptual entropy, a constant, for example 1.5, could be introduced into the algorithm expression, as shown in FIG. 7. A better result can then already be obtained, i.e., a slight deviation up or down will occur, and it will be seen that when a constant introduced in the logarithm expression is taken into account, the case where the perceptual entropy signals a optimistic need for bits, actually be reduced. However, it can be seen clearly in FIG. 7 that too high a number of bits will be significantly signaled, which will cause the quantizer to always quantize too finely, that is, the bit demand is assumed to be greater than it actually is, which in turn results in a reduced coding gain. The constant of the logarithm expression will be a rough estimate of bits required for page information.
Thus, the introduction of a constant in the logarithm expression will provide an enhancement to the band-perceptual entropy, as illustrated in FIG. 6, since this will result in a greater probability that the bands with a very small distance between energy and masking threshold are taken into account, since a certain amount of bits will also be required for transmission of spectral coefficients quantized to zero.
Another very computationally time-intensive calculation of perceptual entropy is illustrated in FIG. 8. Fig. 8 shows cases where the perceptual entropy is calculated in a linear fashion. However, the disadvantage lies in the higher computational complexity for the linear calculation. Here spectral coefficients X (k) are used instead of the energy. Where coffset (b) indicates the first index in band b. 8 is compared with FIG. 7 there will clearly be a reduction in the upward impact in the range from 2,000 to 3,000 bits. The estimate for pe would therefore be more accurate, ie not too pessimistic, but rather approach the optimum, so that the coding gain is increased compared to the calculation method shown in FIG. 6 and 7, and / or so that the number of iterations in the quantizer can be reduced.
The computation time required for the equation shown in FIG. 8, however, will be a disadvantage in the linear calculation of the perceptual entropy.
Such a drawback with a long computation time will not necessarily play any particular role since the encoder is run on a powerful PC or a powerful workstation. But the picture will be quite different if the encoder is installed in a portable device, such as a mobile UMTS phone, which on the one hand must be small and affordable, and on the other hand must have a low power demand, and which in additions must work quickly to make it possible to encode an audio signal or a video signal sent via the UMTS25 connection.
It is an object of this invention to provide an efficient and yet accurate concept for determining an estimate of a need for information units by encoding a signal.
This object is achieved by a device according to claim 1, a method according to claim 12, or a computer program according to claim 13.
The present invention is based on the findings that a frequency-band calculation of the estimate for a need for information units must be limited for computational time reasons, but that in order to arrive at an accurate estimate, the distribution of the energy in the frequency band to be calculated in a band-wise manner must be taken into account. .
Thus, following the quantizer, the entropy coder is implicitly drawn into the determination of the estimate of the need for information units. The entropy coding means that a smaller amount of bits is required for transmission of small spectral values than for the transmission of larger spectral values. The entropy coder is particularly effective when spectral values quantized to zero can be transmitted. Since this will typically be the case, the password for transmitting a spectral line quantized to zero will be the shortest, while the password for transmitting a larger quantized spectral line will be correspondingly longer. In order to obtain a particularly efficient concept for transmitting a frequency of spectral values quantized to zero, even mileage coding can be used, which means that for the series of zeros in a spectral value quantized to zero, not even on average a single bit can be required.
It has been found that the band-perceptual entropy calculation for determining the estimation of the information unit used in the prior art will completely ignore the downstream operating mode of the entropy coder if the distribution of the io energy in the frequency band deviates from a completely uniform distribution.
According to the invention, in order to reduce the inaccuracies in the band-wise calculation, the energy distributed in the band is therefore taken into account.
Depending on the implementation, a measure of the distribution of the energy in the frequency band can be determined on the basis of the actual amplitudes or by estimating the frequency lines that are not quantified to zero by the quantizer. This target, also referred to as "nl", where nl stands for "number of active lines", is preferred for rainy-season efficiency reasons. However, the number of spectral lines quantized to zero, or a finer division, can also be taken into account, where this estimation the more accurate its more information about the downstream entropy code that is included in the calculation.
If the entropy coder is constructed on the basis of Huffman code tables, the properties of these code tables can be integrated in a particularly good way, since the code tables due to signal statistics are not calculated on the site but are determined independently of the signal in question.
However, by a particularly efficient calculation, depending on the time constraints, the target for the distribution of the energy in the frequency band is determined by the lines that survive the quantization, i.e. the number of active lines.
The present invention is advantageous in that an estimate is made for the need for information scope that is both more accurate and more efficient than the prior art.
The present invention can further be adapted to different applications, since several properties of the entropy coder can always be included in the estimate for the bit requirement, depending on the desired accuracy of the estimate, but also at the expense of an increased calculation time.
Otherwise, reference is made to the claims which disclose the aspects of the invention as a device in independent claims 1 and to claims 2 to 9, the method in claim 10 and the computer software in claim 11 which performs said method.
Brief description of the figures
Preferred embodiments of the present invention will be explained in more detail hereinafter with reference to the accompanying drawings, in which:
FIG. 1 is a block diagram showing the inventive device for determining an estimate; FIG. 2a shows a preferred embodiment of the means for calculating a target for distributing the energy in the frequency band; FIG. Figure 2b shows a preferred embodiment of the means for calculating the estimate of the need for bits; 3 is a circuit diagram showing a known audio encoder; FIG. 4 is a principle illustration of the interpretation of the importance of the energy distribution in a band for the determination of the estimate; 5 is a diagram showing the estimation calculation in accordance with the present invention; FIG. Figure 6 is a diagram showing the estimate calculation in accordance with ISO / IEC IS 13818-7 (AAC); 7 is a diagram showing the estimation calculation of a constant; FIG. 8 is a diagram showing a line-by-line estimation calculation when introducing a constant.
Detailed Description of Embodiments
Referring to FIG. 1, the inventive device for determining an estimate of a need for information units by encoding a signal will now be described.
The signal, which may be an audio and / or video signal, is output to an input 100. The signal is preferably already present as a spectral representation with spectral values. However, this is not absolutely necessary, as some calculations with a time signal can also be performed with, for example, a corresponding bandpass filtering.
The signal is then provided to means 102 to provide a measure of permissible interference in a frequency band of the signal. The permissible interference may be determined, for example, by means of a psychoacoustic model, as explained in connection with FIG. 3 (block 1020). The means 102 can furthermore also be operated to provide the measure of the energy of the signal in the frequency band. It is a prerequisite for band-wise calculation that a frequency band for which an allowed interference or signal energy is indicated contains at least two spectral lines in the spectral representation of the signal. In typical, standardized audio coders, the frequency band will preferably be a scaling factor stretch band, since the quantizer immediately needs an estimate of the bit need to be able to determine whether the quantization performed meets a bit criterion.
The means 102 are arranged to be able to supply both the allowed interference nb (b) and the signal energy e (b) of the signal in the band of means 104 for calculating the estimate of the need for bits.
According to the invention, the means 104 for calculating the estimate for the need for 5 bits are arranged to be able to include in the calculation a target nl (b) for a distribution of the energy in the frequency band, in addition to the allowable interference and the signal energy, where the distribution of the energy in the the frequency band differs from a completely uniform distribution. The target for the distribution of energy is calculated in means 106, where these means 106 require at least one band, namely the current frequency band for the audio or video signal io either as a bandpass signal or directly as a result of spectral lines, in order to be able to perform for example, a spectral analysis of the band, so as to provide the target for the distribution of the energies in the frequency band.
The audio or video signal can, of course, also be provided to the means 106 as a bit signal, the means 106 then performing a band filtering as well as an analysis of the band. Alternatively, the audio or video signal provided to the means 106 may already be present in the frequency range, for example as MDCT coefficients, or also as a bandpass signal in the filter bank comprising a smaller number of bandpass filters compared to an MDCT filter bank.
In a preferred embodiment, the means 106 for calculation are arranged to take into account the relevant magnitudes of spectral values in the frequency band when calculating the estimate.
The means for calculating the energy distribution target may further be adapted to determine, as a measure of the energy distribution, the number of spectral values of a magnitude greater than or equal to a predetermined threshold value, or of a magnitude less than or equal to this threshold. Where the threshold value is preferably an estimated quantizer value, set such that values less than or equal to the quantizer value are quantized to zero in a quantizer. In this case, the energy target will be equal to the number of active lines, that is, the number of lines that have survived or are not equal to zero after quantization.
In FIG. 2a shows a preferred embodiment of the means 106 for calculating the measurements for the distribution of the energy in the frequency band. The target for the distribution of the energy in the frequency band is indicated in FIG. 2a with nl (b). The form factor ffac (b) will already be a measure of the distribution of the energy in the frequency band. From block 106, it will be seen that the target of spectral distribution nl is determined from the form factor ffac (b) by weighting with the fourth root of the signal energy e (b) divided by the bandwidth width (b) and / or the number of lines in the scaling factor band b. In this context, it should be pointed out that the form factor is also an example of a size indicating a measure of the distribution of the energies, whereas nl (b), in contrast, is an example of a size representing an estimate of the number of lines relevant for quantization.
The form factor ffac (b) is calculated by determining the size of the spectral line and a subsequent rooting of this spectral line as well as a subsequent summation of these '' roots' for the spectral lines in the band.
FIG. 2b shows a preferred embodiment of the means 104 for calculating the estimate pe, where in FIG. 2b is also distinguished between different cases, namely when the logarithm of element 2 of the rate of energy of the permissible interference is greater than a constant cl or equal to this constant. In this case, the upper alternative is selected in block 104, that is, the target of the spectral distribution nl is multiplied by the logartime expression.
If, on the other hand, it is determined that the logarithm with the base number 2 of the signal energy rate of the allowed interference is less than the value c1, the lower alternative in block 104 of FIG. 2b, which additionally comprises an additive constant c2 as well as a multiplicative constant c3 calculated from constants c2 and c1.
In the following, the inventive concept will be illustrated with reference to FIG.
4a and fig. 4b. Fig. 4a shows a band in which four equal spectral lines are present. Thus, the energy in this band will be uniformly distributed over the band. In contrast, FIG. 4b is a situation where the energy of the band is in a spectral line, while the other three spectral lines are equal to zero. The tape of FIG. For example, 4b could have been present before quantization, or it could have been obtained after quantification if the spectral lines were set to zero in FIG. 4b is less than the first quantizer value before quantization and is thus set to zero by the quantizer, ie they do not '' survive ''.
The number of active lines in FIG. 4b is similar to one in which the parameter n1 of FIG. 4b is set to the square root of 2. In contrast, the value nl, i.e. the target spectral distribution of the energy, in FIG. 4a is calculated to 4. This means that the spectral distribution of the energy is more uniform if the target for the distribution of the spectral energy is large.
It should be noted that the band-wise calculation of the perceptual entropy according to the prior art does not take into account differences between these two cases. More specifically, the cases of FIG. 4a and 4b if the energy level is the same in both of these bands.
It is obvious that the case of FIG. 4b can be encoded with only one relevant line and with fewer bits, since the three spectral lines set to zero can be transmitted very efficiently.
In general, the simpler quantization unit of the case of FIG. 4b is based on the fact that after quantization and a lossless coding, the smaller values, and especially the value quantized to zero, will require fewer bits in transmission.
According to the invention, it will thus be taken into account that how the energy is distributed in the band. As explained, this is done by replacing the number of lines per. bands in the known equation (Fig. 6) with an estimate of the number of lines not equal to zero after quantization. This estimation is shown in FIG. 2a.
It should be further noted that the form factor shown in FIG. 2a is also needed elsewhere in the encoder. For example, in quantization block 1014 for determining step 5 the size of the quantization. If the form factor has already been calculated elsewhere, it will not be necessary to perform this calculation again for the current bit estimation, so that the inventive concept for an improved estimation of the target for required bits will succeed with a minimum of control calculation.
As already explained, the X (k) spectral coefficient to be quantified is io, while the variable kOffset (b) indicates the first index of the band b.
It can be seen from Figures 4a and 4b that the spectrum of Figs. 4b gives a value for n1 of 4, while the spectrum of FIG. 4b gives a value of 1.41. Thus, by means of the form factor, a measure for the quantization of the spectral field structure in the band can be provided.
Thus, the new formula for calculating an improved band-wise perceptual entropy is based on multiplying the measure of the spectral distribution of the energy by the logarithm expression where the signal energy e (b) is in the counter and the allowed interference in the denominator and where a constant can be introduced. in the logarithm expression as needed, as already illustrated in FIG. 7. This constant may, for example, be equal to 1.5, but it may also be equal to zero, as in the case of FIG. 2b, where, for example, this can be determined empirically.
Referring again to FIG. 5, where the perceptual entropy calculated in accordance with the invention is apparent, namely in the plot relative to required bits. A higher accuracy of estimation relative to the comparable examples of Figures 6, 7 and 8 is evident. The modified bandwise calculation according to the invention will also do at least as well as the linear calculation.
The method according to the invention can be implemented in the hardware or in the software, as the case may be. Implementation can be done in a digital storage medium, in particular a floppy disk or CD with electronically readable control signals capable of working with a programmable computer system so that progress can be made. Thus, in general, the invention also comprises a computer program product having a program encoder stored in a machine-readable carrier for carrying out the inventive method, wherein the computer program product is run in a computer. In other words, the invention can thus also be realized as a computer program with a program code for performing the method, when the computer program is run in a computer.
Contents2
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| EP0446037A2 | Cites | European Patent Office (EPO) | A | Search report | 1-11 |
| US2002103637A1 | Cites | United States of America | A | Search report | 1-11 |
41 members in 19 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 102004009949 | Germany | A | |
| 102004009949 | Germany | A | |
| 2005001651 | European Patent Office (EPO) | W | |
| 2005001651 | European Patent Office (EPO) | W | |
| 102004009949 | – | – | – |
| 200501651 | – | – | – |
| DE20041009949 | – | – | – |
| WO2005EP01651 | – | – | – |
Members41
| Document | Office | Kind | |
|---|---|---|---|
| AU2005217507A1 | Australia | A1 | |
| CA2559354A1 | Canada | A1 | |
| WO2005083680A1 | World Intellectual Property Organization (WIPO) | A1 | |
| DE102004009949A1 | Germany | A1 | |
| DE102004009949B4 | Germany | B4 | |
| EP1697931A1 | European Patent Office (EPO) | A1 | |
| NO20064432L | Norway | L | |
| KR20060121978A | Republic of Korea | A | |
| IL176978D0 | Israel | D0 | |
| HK1093813A1 | Hong Kong, China | A1 | |
| CN1938758A | China | A | |
| US2007129940A1 | United States of America | A1 | |
| BRPI0507815A | Brazil | A | |
| JP2007525715A | Japan | A | |
| US7318028B2 | United States of America | B2 | |
| RU2006134638A | Russian Federation | A | |
| AU2005217507B2 | Australia | B2 | |
| KR100852482B1 | Republic of Korea | B1 | |
| RU2337414C2 | Russian Federation | C2 | |
| EP2034473A2 | European Patent Office (EPO) | A2 | |
| CN1938758B | China | B | |
| JP4673882B2 | Japan | B2 | |
| CA2559354C | Canada | C | |
| EP1697931B1 | European Patent Office (EPO) | B1 | |
| AT532173T | Austria | T | |
| ATE532173T1 | Austria | T1 | |
| DK1697931T3 | Denmark | T3 | |
| ES2376887T3 | Spain | T3 | |
| IL176978A | Israel | A | |
| EP2034473A3 | European Patent Office (EPO) | A3 | |
| NO338917B1This record | Norway | B1 | |
| BRPI0507815B1 | Brazil | B1 | |
| EP2034473B1 | European Patent Office (EPO) | B1 | |
| PT2034473T | Portugal | T | |
| EP3544003A1 | European Patent Office (EPO) | A1 | |
| PL2034473T3 | Poland | T3 | |
| ES2739544T3 | Spain | T3 | |
| EP3544003B1 | European Patent Office (EPO) | B1 | |
| PT3544003T | Portugal | T | |
| PL3544003T3 | Poland | T3 | |
| ES2847237T3 | Spain | T3 |
Numbers
- Publication
- 338917
- Publication, DOCDB
- 338917
- Publication, EPODOC
- NO338917B
- Application
- 4432
- Application, DOCDB
- 20064432
- Application, EPODOC
- NO20060004432
Titles2
- Norwegian
- Apparat og fremgangsmåte for å komme frem til et estimat
- English
- Apparatus and method for making an estimate
Classification
- CPC, 4
- G10L19/002
- G10L19/025
- G10L19/02
- G10L25/03
- IPC, 4
- G10L19 002
- G10L19 025
- H04B1 66
- H04N7 26