Coding of spectral coefficients of a spectrum of an audio signal
Abstract
A coding efficiency of coding spectral coefficients of a spectrum of an audio signal is increased by en/decoding a currently to be en/decoded spectral coefficient by entropy en/decoding and, in doing so, performing the entropy en/decoding depending, in a context-adaptive manner, on a previously en/decoded spectral coefficient, while adjusting a relative spectral distance between the previously en/decoded spectral coefficient and the currently en/decoded spectral coefficient depending on an information concerning a shape of the spectrum. The information concerning the shape of the spectrum may have a measure of a pitch or periodicity of the audio signal, a measure of an inter-harmonic distance of the audio signal's spectrum and/or relative locations of formants and/or valleys of a spectral envelope of the spectrum, and on the basis of this knowledge, the spectral neighborhood which is exploited in order to form the context of the currently to be en/decoded spectral coefficients may be adapted to the thus determined shape of the spectrum, thereby enhancing the entropy coding efficiency.

Term
8.1 yearsto projected expiry
Projected expiry 17 October 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
19 claims: 10 independent, 9 dependent
- 1Patent claims Zastrzeżenia patentowe 1. A decoder (40) configured to decode the spectral coefficients (14) of the spectrum (12) of the audio signal (18), the spectral coefficients being at the same time point, the decoder configured to sequentially low to high frequency decode the spectral coefficients and decode the spectral coefficient (x) to be decoded on-line from the spectral coefficients by entropy decoding depending on, in a context-adaptive manner, previously decoded spectral coefficients (o) from the spectral coefficients, with adjustment of the relative spectral distance (28) between the previously decoded spectral coefficient (o) and the spectral coefficient (x) to be decoded on a current basis depending on the spectrum shape information. 1. Dekoder (40) skonfigurowany do dekodowania współczynników (14) widmowych widma (12) sygnału (18) audio, przy czym współczynniki widmowe należą do tego samego momentu czasu, dekoder jest skonfigurowany do sekwencyjnego, od częstotliwości niskich do wysokich, dekodowania współczynników widmowych i dekodowania współczynnika widmowego (x), który ma być bieżąco dekodowany ze współczynników widmowych poprzez dekodowanie entropijne w zależności od, w sposób adaptacyjny względem kontekstu, poprzednio zdekodowanych współczynników widmowych (o) ze współczynników widmowych, z dopasowaniem względnej odległości (28) widmowej między poprzednio zdekodowanym współczynnikiem widmowym (o) i współczynnikiem widmowym (x), który ma być bieżąco dekodowany w zależności od informacji dotyczącej kształtu widma.
- 5A decoder according to any one of the preceding claims, wherein the decoder (40) is configured such that the entropy decoding relationship comprises a plurality of previously decoded spectral coefficients (o), the spectral spacing of the spectral positions of which is adjusted depending on the spectrum shape information. 5. Dekoder według dowolnego z poprzednich zastrzeżeń, przy czym dekoder (40) jest skonfigurowany w taki sposób, że zależność dekodowania entropijnego obejmuje wiele poprzednio zdekodowanych współczynników widmowych (o), których rozstaw widmowy pozycji widmowych jest dopasowywany w zależności od informacji dotyczącej kształtu widma.
- 6A decoder according to any of the preceding claims, wherein the decoder (40) is configured such that the spectrum shape information is a pitch measure (60) of an audio signal and the decoder is configured to match the relative spectral distance (28) between the previously decoded spectral coefficient. (o) and the spectral coefficient (x) to be decoded on a current basis depending on the pitch measure, such that the relative spectral distance increases with increasing pitch, or the spectral shape information is a measure (60) of periodicity of the audio signal and the decoder is configured to match the relative spectral distance (28) between the previously decoded spectral coefficient (o) and the spectral coefficient (x ) to be decoded on a regular basis depending on the measure of periodicity, such that the relative spectral distance decreases with increasing periodicity, or the information regarding the spectrum shape is an interharmonic distance measure of the spectrum (12) of the audio signal and the decoder (40) is configured to match the relative spectral distance between the previously decoded spectral coefficient (o) and the spectral coefficient (x) to be decoded on a current basis measure of the interharmonic distance, so that the relative spectral distance increases with the increase of the interharmonic distance, or the spectrum shape information comprises the relative locations of the formants (70) and / or the valleys (72) of the spectral envelope of the spectrum and the decoder is configured to match the relative spectral distance between the previously decoded spectral coefficient and the spectral coefficient to be decoded on a current basis depending on the location in this way, that the relative spectral distance increases with increasing spectral distance (74) between valleys in the spectral envelope and / or between formants in the spectral envelope. 6. Dekoder według dowolnego z poprzednich zastrzeżeń, przy czym dekoder (40) jest skonfigurowany w taki sposób, że informacją dotyczącą kształtu widma jest miara (60) wysokości tonu sygnału audio i dekoder jest skonfigurowany do dopasowywania względnej odległości (28) widmowej między poprzednio zdekodowanym współczynnikiem widmowym (o) i współczynnikiem widmowym (x), który ma być bieżąco dekodowany w zależności od miary wysokości tonu, tak że względna odległość widmowa zwiększa się ze wzrostem wysokości tonu, albo informacją dotyczącą kształtu widma jest miara (60) okresowości sygnału audio i dekoder jest skonfigurowany do dopasowywania względnej odległości (28) widmowej między poprzednio zdekodowanym współczynnikiem widmowym (o) i współczynnikiem widmowym (x), który ma być bieżąco dekodowany w zależności od miary okresowości, tak że względna odległość widmowa zmniejsza się ze wzrostem okresowości, albo informacją dotyczącą kształtu widma jest miara odległości interharmonicznej widma (12) sygnału audio i dekoder (40) jest skonfigurowany do dopasowywania względnej odległości widmowej między poprzednio zdekodowanym współczynnikiem widmowym (o) i współczynnikiem widmowym (x), który ma być bieżąco dekodowany w zależności od miary odległości interharmonicznej, tak że względna odległość widmowa zwiększa się ze wzrostem odległości interharmonicznej, albo informacja dotycząca kształtu widma obejmuje względne lokalizacji formantów (70) i/lub dolin (72) obwiedni widmowej widma i dekoder jest skonfigurowany do dopasowywania względnej odległości widmowej między poprzednio zdekodowanym współczynnikiem widmowym i współczynnikiem widmowym, który ma być bieżąco dekodowany w zależności od lokalizacji w taki sposób, że względna odległość widmowa zwiększa się ze wzrostem odległości (74) widmowej między dolinami w obwiedni widmowej i/lub między formantami w obwiedni widmowej.
- 7A decoder according to any of the preceding claims, wherein the decoder is configured to, in decoding the spectral coefficient to be continuously decoded by entropy decoding, deriving an estimate (56) of a probability distribution for the spectral coefficient to be currently decoded by subjecting the previously decoded spectral coefficient of a scalar function (82) and the use of probability distribution estimation for entropy decoding. 7. Dekoder według dowolnego z poprzednich zastrzeżeń, przy czym dekoder jest skonfigurowany do, w dekodowaniu współczynnika widmowego, który ma być bieżąco dekodowany za pomocą dekodowania entropijnego, wyprowadzania estymacji (56) rozkładu prawdopodobieństwa dla współczynnika widmowego, który ma być bieżąco dekodowany, poprzez poddawanie poprzednio zdekodowanego współczynnika widmowego funkcji skalarnej (82) i użycie estymacji rozkładu prawdopodobieństwa dla dekodowania entropijnego.
- 8A decoder according to any of the preceding claims, wherein the decoder is configured to use arithmetic decoding as entropy decoding. 8. Dekoder według dowolnego z poprzednich zastrzeżeń, przy czym dekoder jest skonfigurowany do zastosowania dekodowania arytmetycznego jako dekodowania entropijnego.
- 9A decoder according to any of the preceding claims, wherein the decoder is configured to decode the spectral coefficient to be decoded on an ongoing basis with the spectral and / or temporal prediction of the spectral coefficient to be decoded on-line and the spectral and / or temporal prediction correction by means of a prediction residual obtained by entropy decoding. 9. Dekoder według dowolnego z poprzednich zastrzeżeń, przy czym dekoder jest skonfigurowany do dekodowania współczynnika widmowego, który ma być bieżąco dekodowany za pomocą predykcji widmowej i/lub czasowej współczynnika widmowego, który ma być bieżąco dekodowany i korekcji predykcji widmowej i/lub czasowej za pomocą reszty predykcji pozyskanej poprzez dekodowanie entropijne.
- 10An audio transform decoder comprising a decoder configured to decode spectral coefficients of the spectrum of the audio signal according to any of the preceding claims. 10. Transformatowy dekoder audio zawierający dekoder skonfigurowany do dekodowania współczynników widmowych widma sygnału audio według dowolnego z poprzednich zastrzeżeń.
- 16An encoder (10) configured to encode the spectral coefficients (14) of the spectrum (12) of the audio signal (18), the spectral coefficients being at the same time point, the encoder being configured to sequentially low to high frequency encode the spectral coefficients and encode the spectral coefficient (x) to be coded on an ongoing basis from the spectral coefficients by entropy coding depending on, in a context-adaptive manner, the previously encoded spectral coefficients (o) from the spectral coefficients, with adjustment of the relative spectral distance (28) between the previously encoded spectral coefficient and the currently encoded spectral coefficient depending on the spectral shape information. 16. Koder (10) skonfigurowany do kodowania współczynników (14) widmowych widma (12) sygnału (18) audio, przy czym współczynniki widmowe należą do tego samego momentu czasu, koder jest skonfigurowany do sekwencyjnego, od częstotliwości niskich do wysokich, kodowania współczynników widmowych i kodowania współczynnika widmowego (x), który ma być bieżąco kodowany ze współczynników widmowych poprzez kodowanie entropijne w zależności od, w sposób adaptacyjny względem kontekstu, poprzednio zakodowanych współczynników widmowych (o) ze współczynników widmowych, z dopasowaniem względnej odległości (28) widmowej między poprzednio zakodowanym współczynnikiem widmowym i obecnie kodowanym współczynnikiem widmowym w zależności od informacji dotyczącej kształtu widma.
- 17A method for decoding spectral coefficients (14) of the spectrum (12) of an audio signal (18), wherein the spectral coefficients belong to the same time moment, the method comprises sequentially, from low to high frequencies, decoding the spectral coefficients and decoding the spectral coefficient (x), which is to be decoded on an ongoing basis from the spectral coefficients by entropy decoding depending on, in a context-adaptive manner, previously decoded spectral coefficients (o) from the spectral coefficients, with adjustment of the relative spectral distance (28) between the previously decoded spectral coefficient (o) and the spectral coefficient (x) to be decoded on a current basis depending on the spectrum shape information. 17. Sposób dekodowania współczynników (14) widmowych widma (12) sygnału (18) audio, przy czym współczynniki widmowe należą do tego samego momentu czasu, sposób zawiera kolejno, od częstotliwości niskich do wysokich, dekodowanie współczynników widmowych i dekodowanie współczynnika widmowego (x), który ma być bieżąco dekodowany ze współczynników widmowych poprzez dekodowanie entropijne w zależności od, w sposób adaptacyjny względem kontekstu, poprzednio zdekodowanych współczynników widmowych (o) ze współczynników widmowych, z dopasowaniem względnej odległości (28) widmowej między poprzednio zdekodowanym współczynnikiem widmowym (o) i współczynnikiem widmowym (x), który ma być bieżąco dekodowany w zależności od informacji dotyczącej kształtu widma.
- 18A method of coding the spectral coefficients (14) of the spectrum (12) of an audio signal (18), wherein the spectral coefficients belong to the same time point, the method comprises sequentially from low to high frequencies, coding the spectral coefficients and coding the spectral coefficient (x) to be be coded on an ongoing basis from spectral coefficients via entropy coding depending on, in a context-adaptive manner, the previously encoded spectral coefficients (o) from the spectral coefficients, with adjustment of the relative spectral distance (28) between the previously encoded spectral coefficient and the currently encoded spectral coefficient depending on the spectral shape information. 18. Sposób kodowania współczynników (14) widmowych widma (12) sygnału (18) audio, przy czym współczynniki widmowe należą do tego samego momentu czasu, sposób zawiera kolejno od częstotliwości niskich do wysokich, kodowanie współczynników widmowych i kodowanie współczynnika widmowego (x), który ma być bieżąco kodowany ze współczynników widmowych poprzez kodowanie entropijne w zależności od, w sposób adaptacyjny względem kontekstu, poprzednio zakodowanych współczynników widmowych (o) ze współczynników widmowych, z dopasowaniem względnej odległości (28) widmowej między poprzednio zakodowanym współczynnikiem widmowym i obecnie kodowanym współczynnikiem widmowym w zależności od informacji dotyczącej kształtu widma.
Independent claims10
175 paragraphs in 2 sections, as filed
Description
The present invention relates to a method for coding the spectral coefficients of an audio signal spectrum, useful for example in various transform audio codecs.
[0002] Contextual arithmetic coding is an efficient way to code noiselessly the spectral coefficients of the transform [1] encoder. The context uses the mutual information between the spectral coefficient and the already encoded coefficients in its vicinity. Context is available on both the encoder and decoder sides and does not require any additional information to be sent. Thus, entropy contextual coding has the potential to provide greater gain compared to memory-free entropy coding. In practice, however, the context design is severely limited due to, among other things, memory requirements, computer complexity, and channel fault tolerance. These constraints limit the performance of contextual entropy coding and result in lower coding gain especially for tonal signals where the context must be too constrained to use the harmonic structure of the signal.
[0003] Moreover, in low-latency transform audio coding, windows with small tabs are used to reduce the algorithmic delay. As a direct consequence of this, leaks in the MDCT are important for tonal signals and lead to more quantization noise. Tonal signals may be handled by combining a transform with frequency domain prediction as is performed in MPEG2 / 4-AAC [2] or in time domain prediction [3].
[0004] Document US 2013/0117015 A1 states with respect to the entropy coding of the spectral values of the MDCT spectrum that the context for decoding / coding the current spectral value or the tuple of the spectral value of the current k frame is determined based on the previously encoded spectral values located in a certain spectral-time vicinity of the currently decoded / coded spectral value (tuples of the spectral value). The template is shown as an L-shaped area covering the four spectral value positions in the immediate vicinity of the spectral time of the currently decoded / coded spectral value. The L-shaped pattern includes one spectral value in the current k-frame, namely encoded immediately before the current spectral value, i.e., immediately adjacent spectrally spectral value at the immediately lower spectral position. However, the template also includes spectral positions within the previous k-1 box. According to the concept set out in document D1, the frequency scale of this previous k-1 frame is "compressed" or "enlarged" relative to the frame rate scale k. "Relative pinch / zoom" is the result of a frame-specific temporal distortion which is, according to document D1, applied to the temporal audio samples of each frame before it is subjected to MDCT transform to obtain a spectrum of spectral values for each frame. The idea of this document is to re-implement the shifting of the correct spectral-time neighborhood of this part of the spectral-time template in order to contextualize the currently decoded / coded spectral value that extends into the temporal previous frame, the displacement being the result of a difference in the temporal distortion applied on the current frame on the one hand, and on the previous frame on the other.
[0005] In EP 2 650 878 A1 it is suggested to change the arrangement of the spectral coefficients of the audio spectrum so as to obtain a smooth spectrum. This "interleaved" or "rearranged" spectrum is then entropy coded.
[0006] It would be advantageous to have an encoding concept which increases the encoding efficiency. Accordingly, it is an object of the present invention to provide a concept for coding spectral coefficients of the spectrum of an audio signal, which increases the coding efficiency. The goal was achieved by the subject of the pending patent claims.
[0007] The underlying finding of the present application is that the coding efficiency for the coding of the spectral coefficients of the spectrum of an audio signal can be increased by coding / decoding the spectral coefficient to be continuously coded by entropy coding / decoding and in doing so, performing entropy coding / decoding depending on, contextually adaptively, the previously encoded / decoded spectral coefficient, with simultaneous adjustment of the relative spectral distance between the previously encoded / decoded spectral coefficient and the currently encoded / decoded spectral coefficient depending on the spectrum shape information. The spectrum shape information may include a measure of pitch or periodicity of an audio signal, a measure of the interharmonic distance of the spectrum of audio signals, and / or the relative locations of formants and / or troughs of the spectral envelope of the spectrum, and based on this knowledge, the spectral neighborhood that is used to form the context of the currently coded / decoded spectral coefficients can be adjusted to the spectrum shape determined in this way, thus increasing the efficiency of entropy coding.
[0008] Preferred embodiments are the subject of the dependent patent claims and preferred embodiments of the present application are described below with reference to the figures, among which
Fig. 1 is a schematic diagram illustrating a spectral coefficient encoder and its mode of operation in coding the spectral coefficients of an audio signal spectrum;
Fig. 2 is a schematic diagram showing a spectral coefficient decoder fitted with a spectral coefficient encoder in Fig. 1;
Fig. 3 shows a block diagram of a potential internal structure of the spectral coefficient encoder of Fig. 1 according to an embodiment;
Fig. 4 shows a block diagram of a potential internal structure of the spectral coefficient decoder of Fig. 2 according to an embodiment;
Fig. 5 is a schematic diagram of a spectrum whose coefficients are to be coded / decoded to illustrate the fitting of the relative spectral distance as a function of a pitch measure or periodicity of an audio signal or an interharmonic distance measure;
Fig. 6 is a schematic diagram illustrating a spectrum whose spectral coefficients are to be encoded / decoded according to an embodiment where the spectrum is spectrally shaped according to a perceptually weighted Linear Prediction (LP) synthesis filter, namely its inverse, with showing the adjustment of the relative spectral distance as a function of an interformant distance measure according to an embodiment;
Fig. 7 schematically shows a portion of a spectrum to illustrate a context template surrounding a spectral coefficient to be currently encoded / decoded and adjusting the spectral spread of the context templates depending on the spectral shape information according to an embodiment;
Fig. 8 is a schematic diagram illustrating the mapping of one or more reference values of the spectral coefficients of context template 81 using a scalar function so as to obtain a probability distribution estimate for use in coding / decoding the current spectral coefficient according to an embodiment;
Fig. 9a illustrates schematically the use of indirect signaling for synchronization to fit the relative distance between an encoder and a decoder;
Fig. 9b is a schematic diagram illustrating the use of direct signaling to synchronize to adjust the relative spectral distance between an encoder and a decoder;
Fig. 10a shows a block diagram of a transform audio encoder according to an embodiment;
Fig. 10b shows a block diagram of a transform audio decoder matching the encoder of Fig. 10a;
Fig. 11a shows a block diagram of a transform audio encoder using frequency domain spectral shaping according to an embodiment;
Fig. 11b shows a block diagram of a transform audio decoder matching the encoder of Fig. 11a;
Fig. 12a shows a block diagram of an audio encoder with transform code excitation according to an embodiment;
Fig. 12b shows an audio decoder with linear prediction transform code excitation according to an embodiment;
Fig. 13 shows a block diagram of a transform audio encoder according to another embodiment;
Fig. 14 shows a block diagram of a transform audio decoder matching the embodiment of Fig. 13;
Fig. 15 is a schematic diagram illustrating a conventional context or a context template covering the vicinity of a currently coded / decoded spectral coefficient;
Figs. 16a -c show modified context template configurations or a mapped context according to embodiments of the present invention;
Fig. 17 schematically illustrates a harmonic spectrum plot so as to illustrate the advantage of using the mapped context from any one of Figs. 16a through 16c compared to the context template definition of Fig. 15 for a harmonic spectrum;
Fig. 18 is a flowchart of the relative spectral distance optimization algorithm D for context mapping according to an embodiment.
[0009] Fig. 1 shows a spectral coefficient encoder 10 according to an embodiment. The encoder is configured to encode the spectral coefficients of the spectrum of the audio signal. Fig. 1 illustrates successive spectra in the form of a spectrogram 12. More specifically, the spectral coefficients 14 are shown as rectangles arranged spectrally along the time axis t and the frequency axis f. While it is possible that the spectral-time resolution is constant, Fig. 1 shows that the spectral-temporal resolution may vary with time with one such time instant as illustrated in Fig. 1 at reference number 16. The spectrogram 12 may be the result of a spectral distribution of a transform applied to the audio signal 18 at different times, such as an overlap transform, such as for example, a critically sampled transform such as the MDCT or some other critically sampled transformation. Within this range, the spectrogram 12 may be received by the spectral coefficient encoder 10 as a spectrum 20 consisting of a sequence of transform coefficients all belonging to the same time point. The spectra 20 thus represent spectral segments of the spectrogram and are shown in Fig. 1 as single columns of spectrogram 12. Each spectrum consists of a sequence of transform coefficients 14 and has been received from the respective time frame 22 of the audio signal 18 using, for example, some window function 24. In particular, time frames 22 are sequentially ranked at said times and are associated with the time sequence of the spectra 20. They may, as shown in Fig. 1, overlap, as may the corresponding transformation windows 24. That is, as used herein, "spectrum" means spectral coefficients belonging to the same time instant and is therefore a frequency distribution. A "spectrogram" is a distribution of the time-frequency made on successive spectra, where the word "spectra" is the plural of the word "spectrum". Sometimes, however, "spectrum" is used as a synonym for a spectrogram. A "transform factor" is used as a synonym for "spectral factor" if the original signal is in the time domain and the transform is a frequency transform.
[0010] As just outlined, the spectral coefficient encoder 10 serves to encode the spectral coefficients 14 of the spectrogram 12 of the audio signal 18, and for this purpose the encoder may, for example, use a predetermined encoding / decoding order that passes, for example, through spectral coefficients 14 along a spectral time path which, for example, scans the low to high frequency spectral coefficients 14 within spectrum 20 and then continues with the spectral coefficients of the time of the next spectrum 20 as outlined in Fig. 1 at reference number 26.
[0011] As outlined in more detail below, the encoder 10 is configured to encode the spectral coefficient to be encoded as indicated by the small cross in Fig. 1 with entropy encoding depending on a context-adaptive one or more the number of previously encoded spectral coefficients, visually marked with a small circle in Fig. 1. In particular, the encoder 10 is configured to adjust the relative spectral distance between the previously encoded spectral coefficients and the currently encoded in dependence on the spectrum shape information. Regarding the relationship and information regarding the spectral shape, details are presented below together with considerations on the advantages of adjusting the relative spectral distance as a function of the information just mentioned.
[0012] In other words, the spectral coefficient encoder 10 codes the spectral coefficients 14 sequentially into the data stream 30. As will be outlined in more detail below, the spectral coefficient encoder 10 may be part of a transform encoder which, in addition to the spectral coefficients 14, encodes additional information into the data stream 30, such that the data stream 30 allows the reconstruction of the audio signal 18.
Fig. 2 shows the spectral coefficient decoder 40 matching the spectral coefficient encoder 10 in Fig. 1. The spectral coefficient decoder 40 functionality is substantially the inverse of the spectral coefficient encoder 10 in Fig. 1: the spectral coefficient decoder 40 sequentially decodes the spectral coefficients 14 of the spectrum 12 using, for example, decoding order 26. When decoding the spectral coefficient to be currently decoded, indicated by a small cross in Fig. 2, by means of entropy decoding, the spectral coefficient decoder 40 performs entropy decoding depending on, in a context-adaptive manner, one or more previously decoded spectral coefficients, also indicated by a small circle in Fig. 2. In this way, the spectral coefficient decoder 40 adjusts the relative spectral distance 28 between the previously decoded spectral coefficient and the spectral coefficient to be decoded on-going in dependence on said spectrum shape information 12. In the same way as indicated above, spectral coefficient decoder 40 may be part of a transform decoder configured to reconstruct an audio signal 18 from a data stream 30, from which spectral coefficient decoder 40 decodes spectral coefficients 14 using entropy decoding. The latter transform decoder, as part of the reconstruction, subjects the spectrum 12 to an inverted transform, such as, for example, an inverted overlap transform, which, for example, results in a reconstruction of a sequence of overlapping windowed time frames 22, which, by means of an overlapping operation, removes, for example, aliasing resulting from transformation of the spectral distribution.
[0014] As described in more detail above, the advantages of adjusting the relative spectral distance 28 depending on the spectrum shape information 12 are based on the ability to improve the probability distribution estimation used to encode / decode the entropy of the current spectral coefficient x. The better the probability distribution estimation is. the more efficient the entropy coding, i.e., the more compressed. The "probability distribution estimate" is an estimate of the actual probability distribution of the current spectral coefficient 14, ie, a function that assigns a probability to each value of the domain of values that the current spectral coefficient 14 may assume. Due to the dependence of the adjustment of the distance 28 on the shape of the spectrum 12, the estimate of the probability distribution can be made to more closely match the actual probability distribution, since the use of information about the shape of the spectrum 12 makes it possible to obtain an estimate of the probability distribution from the spectral neighborhood of the current spectral coefficient x, which enables a more accurate estimation of the distribution the probability of the current spectral coefficient x. Details in this regard are set out below together with examples of information about the spectrum 12.
Before we turn to specific examples of the above-mentioned spectrum shape information 12, Figs. 3 and 4 show potential internal structures of spectral coefficient encoder 10 and spectral coefficient decoder 40, respectively. In particular, as shown in Fig. 3, the spectral coefficient encoder 10 may consist of a probability distribution estimation derivation module 42 and an entropy coding engine 44, similarly, the spectral coefficient decoder 40 may consist of a probability distribution estimation derivation module 52 and an entropy decoding engine 54. Probability distribution estimation derivation modules 42 and 52 operate in the same way: they derive, based on the value of one or more previously decoded / encoded spectral coefficients o, the probability distribution estimate 56 for decoding / entropy coding of the current spectral coefficient x. Specifically, the coding / decoding engine 44/54 receives an estimate of the probability distribution from the derivation module 42/52 and performs entropy coding / decoding regarding the current spectral coefficient x, respectively.
[0016] The 44/54 entropy coding / decoding engine may use, for example, a variable length encoding such as Huffman encoding to encode / decode the current spectral coefficient x, and within this range the 44/54 engine may use different VLC tables (Variable Length Coding, variable-length encoding) for different probability distribution estimates 56. Alternatively, the 44/54 engine may use arithmetic encoding / decoding with respect to the current spectral coefficient x with an estimate 56 of the probability distribution controlling the partitioning of the probability intervals of the current probability interval representing the internal state of the 44/54 encoding / decoding arithmetic engines, each partial interval being assigned to another potential value in the target value range, which can be assumed by the current spectral coefficient x. As will be outlined in more detail below, the entropy coding engine and the entropy decoding engine 44 and 54 may use an exit mechanism to map the overall range of spectral coefficients 14 to a limited range of integers, i.e., the target range. such as [0 ... 2<sup>N-</sup>1]. A set of integers in the target range, i.e., {0, ..., 2<sup>N-1</sup>} defines, together with the output symbol {esc}, the symbol alphabet of the arithmetic encoding / decoding 44/54 engine, i.e., {0, ..., 2 N<sup>-1</sup>, esc}. For example, the entropy coding engine 44 subjects the intrinsic spectral coefficient x to division by 2 as often, if at all, to bring the spectral coefficient x to the target interval [0 ... 2 mentioned above.<sup>N</sup>-1] of, for each partition, the encoding of the output symbol into data stream 30 followed by arithmetic encoding of the remainder of the division - or the original spectral value in case no division is necessary - into bitstream 30.
Entropy decoding engine 43 would in turn implement the output mechanism as follows: it would decode the current spectral coefficient x from bitstream 30 as a sequence of 0, 1 or more esc output symbols followed by a non-esc symbol (non-esc), i.e., as one of the sequences {a}, {esc, a}, {esc, esc, a}, ..., where a is a symbol different from the exit symbol. Entropy decoding engine 54 would acquire, by arithmetically decoding a symbol other than the exit symbol, a value in the target interval, e.g. [0 ... 2N-1], and would obtain the factor x by calculating the value of the current spectral factor to be equal to a + 2 times the number of symbols exit.
[0017] There are various possibilities for using the probability distribution estimation 56 as well as its application to the sequence of symbols used to represent the current spectral coefficient x: the estimation of the probability distribution can, for example, be applied to any symbol carried in the data stream for the spectral coefficient x that is, the non-esc symbol as well as any esc symbol, if any. Alternatively, probability distribution estimation 56 is only used for the first or first two or first n <N of the sequence of 0 or more esc symbols followed by a non-esc symbol, using, for example, some default probability distribution estimate for any consecutive one from a sequence of symbols, such as an equal probability distribution.
[0018] Fig. 5 shows an illustrative spectrum 20 from the spectrogram 12. In particular, the modules of the spectral coefficients are plotted in Fig. 5 in arbitrary units along the y axis, while the horizontal x axis corresponds to frequency in arbitrary units. As already stated, spectrum 20 in Fig. 5 corresponds to a spectral segment above the spectrogram of the audio signal at a certain point in time, spectrogram 12 consisting of a sequence of such spectra 20. Fig. 5 it also shows the spectral position of the current spectral coefficient x.
As will be outlined in more detail below, while spectrum 20 may be an unweighted spectrum of an audio signal, according to the embodiments sketched further below, for example, spectrum 20 is already perceptually weighted using a transfer function which corresponds to the inverse of the perceptual filter function. synthesis. However, the present invention is not limited to the specific case outlined further below.
[0020] In any event, Fig. 5 shows the spectrum 20 with a certain periodicity along the frequency axis, which is manifested in the pattern of more or less equal local distances between the maxima and minima in the spectrum in the frequency direction. For illustrative purposes only, Fig. 5 shows a pitch or periodicity measure 60 of an audio signal defined by the spectral distance between local spectral peaks between which the current spectral coefficient x is located. Naturally, measure 60 may be defined and otherwise determined, such as the average pitch between local highs and / or local minima, or a frequency distance equivalent to the time delay measured as a function of the autocorrelation of signal 18 in the time domain.
[0021] According to an embodiment, measure 60 is information (or consists of information) about the shape of a spectrum. Encoder 10 and decoder 40, or more precisely, probability distribution estimation derivation module 42/52, for example, adjusts the relative spectral distance between the previous spectral coefficient o and the current spectral coefficient x as a function of this measure 60. For example, relative spectral distance 28 may vary depending on measure 60 such that distance 28 increases as measure 60 increases. For example, it may be preferable to set distance 28 equal to measure 60 or an integer multiple thereof.
[0022] As described in more detail below, there are different possibilities as to how information about the shape of the spectrum 12 is made available to a decoder. Generally, this information, such as the measure 60, can be signaled to the decoder directly and only the encoder 10 or the probability distribution estimation derivation module 42 would actually derive the spectrum shape information or the determination of the spectrum shape information would be performed at the encoder and decoder side in parallel based on the former. the decoded part of the spectrum or could be inferred from other information already stored in the bitstream.
[0023] Using another expression, measure 60 can also be interpreted as an "interharmonic distance measure" because the above-mentioned local maxima or peaks in the spectrum may form harmonics to one another.
[0024] Fig. 6 shows another example of spectrum shape information against which spectral distance 28 may be adjusted - either exclusively or together with another measure such as measure 60 as previously described. In particular, Fig. 6 shows an illustrative case where spectrum 12 represented by spectral coefficients coded / decoded by an encoder 10 and a decoder 40, a spectral segment of which is shown in Fig. 6, is weighted using the inverse of the perceptually weighted synthesis filter function. That is, the original and finally reconstructed spectrum of the audio signal is shown in Fig. 6 at reference number 62. The pre-emphasized version is shown at reference number 64 as a multipoint line. The linearly predicted spectral envelope of the pre-emphasized version 64 is shown by a dotted line 66 and its perceptually modified version, i.e., the perceptual synthesis filter transfer function is shown in Fig. 6 at reference 68 using a two-point line. The spectrum 12 may be the result of filtering an enhanced version of the original spectrum 62 of the audio signal by the inverse of the function 68 of the perceptually weighted synthesis filter. In this case, both the encoder and the decoder can access the spectral envelope 66, which in turn can have more or less emphasized formants 70 or valleys 72. According to an alternative embodiment of the present invention, the spectrum shape information is at least partially defined based on o the locations of these formants 70 and / or valleys 72 of the spectral envelope 66 of the spectrum 12. For example, the spectral distance 74 between formants 70 may be used to set the aforementioned relative spectral distance 28 between the current spectral coefficient x and the previous spectral coefficient o. For example, the distance 28 may preferably be set equal to or as an integer multiple of the distance 74, but there are also alternative solutions.
[0025] Instead of the LP-based envelope as shown in Fig. 6, the spectral envelope may also be defined differently. For example, an envelope can be defined and transmitted in the data stream using scaling factors. Other methods of transferring the envelope can also be used.
[0026] Due to the adjustment of the distance 28 as outlined above with reference to Figs. 5 and 6, the value of the "reference" spectral coefficient o provides a substantially better indication for estimating the probability distribution for the current spectral coefficient x compared to other spectral coefficients that lie , for example, spectrally closer to the current spectral coefficient x. In this respect, it should be noted that context modeling is in most cases a compromise between the complexity of entropy coding on the one hand and coding efficiency on the other hand. Accordingly, the embodiments described so far suggest adjusting the relative spectral distance 28 depending on the spectrum shape information such that, for example, the distance 28 increases with an increase in measure 60 and / or an increase in interformant distance 74. However, the number of previous coefficients o based on which the matching of the entropy encoding / decoding context is performed may be constant, i.e., it may not increase. The number of previous spectral coefficients o based on which context matching is performed may, for example, be constant regardless of a change in the spectrum shape information. This means that adjusting the relative spectral distance as outlined above leads to better or more efficient entropy encoding / decoding without significantly increasing the overhead in implementing context modeling. Merely adjusting the spectral distance itself increases the context modeling overhead.
[0027] To illustrate the point just mentioned in more detail, reference is made to Fig. 7, which shows a spectral-time portion from spectrogram 12, the spectral-time portion comprising the current spectral coefficient 14 to be encoded / decoded. Moreover, Fig. 7 shows a template of exemplary five previously encoded / decoded spectral coefficients o based on which context modeling for coding / decoding the current spectral coefficient is performed. The template is positioned at the location of the current spectral coefficient x and indicates adjacent reference spectral coefficients o. Depending on the spectral shape information mentioned above, the spectral spacing of the spectral positions of these reference spectral coefficients o is adjusted. This is illustrated in Fig. , for example, scaling the spectral spread of the spectral positions of the reference spectral coefficients as a function of the alignment 80. That is, Fig. 7 shows that the number of reference spectral coefficients involved in context modeling, i.e., the number of reference spectral coefficients of the template surrounding the current spectral coefficient xi identifying the reference spectral coefficients o, is constant irrespective of any change in the spectral shape information. Only the relative spectral distance between the reference spectral coefficients and the current spectral coefficient is adjusted according to 80, and naturally the distance between the reference spectral coefficients themselves. However, it should be noted that the number of the reference spectral coefficients o is not necessarily constant. According to an embodiment, the number of reference spectral coefficients may increase with increasing relative spectral distance. The reverse is also possible.
[0028] It should be noted that Fig. 7 shows the illustrative case where the modeling of the context for the current spectral coefficient x also includes the previously coded / decoded spectral coefficients corresponding to the earlier spectral / time frame. However, this is also to be understood only as an example and the dependence on such previously preceding previously encoded / decoded spectral coefficients may be omitted according to another embodiment. Fig. 8 illustrates how the module 42/52 for deriving a probability distribution estimate may, based on one or more reference spectral coefficients o, derive an estimate of the probability distribution for the current spectral coefficient. As shown in Fig. 8, for this purpose one or more reference spectral coefficients o may be subjected to a scalar function 82. Based on a scalar function, for example, one or more reference spectral coefficients o are mapped to an index indicating a probability distribution estimate to be used for the current spectral coefficient x among the set of available probability distribution estimates. As mentioned above, the available estimates of the probability distribution may, for example, correspond to different probability interval divisions for the symbol alphabet in the case of arithmetic coding or different variable-length coding tables when using variable-length coding.
[0029] Before continuing to describe the potential integration of the spectral coefficient encoders / decoders described above into the corresponding transform encoders / decoders, several possibilities will be discussed here as to how the embodiments described so far may vary. For example, the exit mechanism briefly outlined above with reference to Fig. 3 and Fig. 4 has been selected for illustrative purposes only, and may be omitted according to an alternative embodiment. In the embodiment described below, an exit mechanism is used. Moreover, as will become clear from the description of the more specific embodiments outlined below, instead of encoding / decoding spectral coefficients individually, they may be encoded / decoded in n-tuple units, i.e., in n spectrally units of immediately adjacent spectral coefficients. In this case, the determination of the relative spectral distance may also be performed in units such as n-tuples or in units of single spectral coefficients. With respect to the scalar function 82 of Fig. 8, it should be noted that the scalar function may be an arithmetic function or a logic operation. In addition, special measures may be taken with regard to these reference spectral coefficients o which, for example, are unavailable due to, for example, exceeding the frequency range of the spectrum or, for example, lying in a part of the spectrum sampled by the spectral coefficients at a spectral-time resolution different from the spectral-time resolution with which the spectrum is sampled at the time moment corresponding to the current spectral coefficient. The values of the unavailable reference spectral values o may be replaced with default values, for example, and then entered into the scalar function 82 together with other (available) reference spectral coefficients. Another way of performing entropy encoding / decoding using spectral distance adjustment described above is as follows: for example, the current spectral coefficient may be given to binarization. For example, the spectral coefficient x may be mapped to a sequence of containers which are then entropy-encoded using an adaptive adjustment of the relative spectral distance. During decoding, the bands will be entropy-decoded until a valid container sequence is encountered, which can then be remapped to the corresponding values of the current spectral coefficient x.
[0030] Furthermore, matching the context depending on one or more of the previous spectral coefficients o may be performed differently from that shown in Fig. 8. In particular, the scalar function 82 may be used to index one of the set of available contexts and each context. could have an estimate of the probability distribution associated with it. In this case, the estimation of the probability distribution associated with a certain context may be adapted to the actual spectral coefficient statistics each time the currently coded / decoded spectral coefficient x has been assigned to the appropriate context, namely the values of that current spectral coefficient x.
[0031] Finally, Figs. 9a and 9b show how the different possibilities of the method for deriving spectrum shape information can be synchronized between an encoder and a decoder. Fig. 9a shows the possibility whereby indirect signaling is used to synchronize the output of spectrum shape information between an encoder and a decoder. In this case, both on the coding and decoding side, information derivation is performed based on the previously coded portion or the previously decoded portion of the 30 bit stream, respectively, the output on the coding side is indicated by the reference number 83, and the output on the decoding side is indicated by using reference numeral 84. Both outputs may be provided by e.g. outputs 42 and 52 alone.
[0032] Fig. 9b illustrates the possibility where direct signaling is used to communicate spectrum shape information from an encoder to a decoder. Pinout 83 on the coding side can even analyze the original audio signal including its components which are not available on the decoding side due to coding losses. Rather, direct signaling in the data stream is used to provide information regarding the spectrum shape on the decoding side. In other words, pin 84 on the decoding side uses direct signaling on the data stream 30 to access spectrum shape information. Direct signaling may include differential encoding. As will be described in more detail below, for example, the LTP (Long Term Prediction) delay parameter already available in the data stream for other purposes may be used as spectrum shape information. Alternatively, however, the direct signaling of Fig. 9b may be a measure 60 of differential encoding with respect to, i.e., different from, an already available delay parameter LPT. There are many other possibilities to provide information on the spectrum shape on the decoding side. In addition to the alternative embodiments outlined above, it should be noted that the coded / decoded spectral coefficients may, in addition to entropy coding / decoding, include spectral and / or temporal prediction of the currently coded / decoded spectral coefficient. The rest of the prediction may then be entropy coded / decoded as described above.
[0033] After the various embodiments for the spectral coefficient encoder and decoder have been described, some embodiments are described below regarding the preferred ways of incorporating them into the transform encoder / decoder.
[0034] Fig. 10a, by way of example, shows a transform audio encoder according to an embodiment of the present invention. The audio transform encoder of Fig. 10a is designated generally by the reference numeral 100 and includes a spectral computer 102 followed by a spectral coefficient encoder 10 of Fig. 1. Spectrum computer 102 receives the audio signal 18 and based on it calculates a spectrum 12, the spectral coefficients of which are encoded by the spectral coefficient encoder 10 as described above into the data stream 30. Fig.
10b shows the structure of the corresponding decoder 104: the decoder 104 comprises a combination of a spectral coefficient decoder 40 formed as outlined above, and in the case of Figs. spectral to the time domain appropriately only implements its reciprocal. The spectral coefficient encoder 10 may be configured to losslessly encode the input spectrum 20. In comparison, the spectrum computer 102 may introduce coding losses due to quantization.
[0035] For spectral quantization noise shaping, a spectral computer 102 may be implemented as shown in Fig. 11a. In this case, spectrum 12 is spectrally shaped using scaling factors. In particular, according to Fig. 11a, the spectrum computer 102 comprises a combination of a transformer 108 and a spectral shaper 110, in which the transformer 108 subject the input audio signal 18 to a spectral distribution transform so as to obtain an unshaped spectrum 112 of the audio signal 18, with spectral shaping module 110 shaping the spectrally shaping spectrum 112 by the scaling factors 114 obtained from the scaling factor calculator 116 of the spectrum computer 112 such as, to obtain the spectrum 12 which is finally encoded by the coder 10 of the spectral coefficients. For example, spectral shaper 110 acquires one scaling factor 114 per scale factor band from scale factor calculator 116 and divides each spectral factor of the respective scaling factor band by a scaling factor associated with the respective scaling factor band so as to obtain spectrum 12. The scaling factor calculator 116 may be controlled by the perceptual model so as to determine the scaling factors based on the audio signal. Alternatively, the determination module 116 may determine the scale factors based on linear prediction analysis such that the scale factors represent a transfer function dependent on a linear prediction synthesis filter defined by the linear prediction factor information. The linear prediction coefficient information 118 is encoded into the data stream 30 together with the spectral coefficients of the spectrum 20 by the encoder 10. For completeness, Fig. 11a shows the quantizer 120 as positioned downstream of the spectral shaper 110, so as to obtain the spectrum 12 with quantized spectral coefficients that are they are then losslessly encoded by the spectral coefficient encoder 10.
Fig. 11b shows a decoder corresponding to the encoder in Fig. 10a. In this case, the spectrum to the time domain computer 106 includes a scaling factor calculator 122 that reconstructs the scaling factors 114 from the linear prediction coefficient information 118 contained in the data stream 30 such that the scaling factors represent a transfer function dependent on a defined linear prediction synthesis filter defined. by the information 118 of the linear prediction coefficients. Spectral shaper 12 shapes the spectrum 12 decoded by decoder 40 from data stream 30 according to the scaling factors 114, i.e., spectral shaper 124 scales the scaling factors within each spectral band using the scaling factor of the corresponding scaling factor band. In this way, a reconstruction of the unshaped spectrum 112 of the audio signal 18 is produced at the output of the spectral shaper 124 and as shown in Fig. 11b with dashed lines, applying an inverted transform on the spectrum 112 with an inverted transformer 126, so that reconstruction of the audio signal 18 in the time domain is optional.
[0037] Fig. 12a illustrates the embodiment of the transform audio encoder of Fig. 11a in more detail when using linear prediction based spectrum shaping. In addition to the components shown in Fig. 11a, the encoder of Fig. 12a includes a pre-enhancement filter 128 configured to initially subject the internal audio signal 18 to pre-enhancement filtering. The pre-enhancement filter 128 may, for example, be implemented as a FIR filter. The carry-over function of the pre-emphasizing filter 128 may, for example, represent a high-pass transfer function. According to an exemplary embodiment, the pre-enhancement filter 128 is implemented as an nth order highpass filter, such as, for example, a one order highpass filter having a transfer function H (z) = 1 - αz-.<sup>1</sup> with a being set to 0.68, for example. Accordingly, a pre-emphasized version 130 of the audio signal 18 appears at the output of the pre-enhancement filter 128. Moreover, Fig. 12a shows the scaling factor determination module 116 as being built up of LP (Linear Prediction) analyzer 132 and linear prediction to scale factor converter 134. LPC 132 calculates the linear prediction coefficient information 118 based on the pre-emphasized version of the audio signal 18. In this way, the linear prediction coefficients represent a spectral envelope based on the linear prediction of the audio signal 18 or, more specifically, a pre-emphasized version 130 thereof. The operating mode of the LP analyzer 132 may, for example, include windowing the input signal 130 so as to obtain an LP analysis of the sequence of the windowed portions of the signal 130, an autocorrelation determination, so as to determine the autocorrelation of each windowed portion of the delay windowing, which is optional, to apply a delay window function on the autocorrelations. The estimation of the linear prediction parameter may then be performed on the autocorrelations or at the output of the delay window, i.e., a windowed autocorrelation function. The estimation of a linear prediction parameter may, for example, include implementing the Wiener-Levinson-Durbin algorithm or other suitable algorithm on (delayed windowed) autocorrelations so as to derive linear prediction coefficients per autocorrelation, i.e., per windowed portion of signal 130. That is, the LPC 118 coefficients appear at the output of the LP 132 analyzer. The LP 132 analyzer may be configured to quantize the linear prediction coefficients for input into the data stream. The quantization of the linear prediction coefficients may be performed in a domain other than the domain of the linear prediction coefficients, such as, for example, in the spectral pairs of the lines or in the spectral frequency domain of the lines. But other algorithms than the Wiener-Levinson-Durbin algorithm may also be used.
[0038] The linear prediction to scale factor converter 134 converts the linear prediction factors into scale factors 114. The converter 134 may determine the scaling factors 140 to correspond to the inverse of the linear prediction synthesis filter 1 / A (z) defined by the linear prediction coefficient information 118. Alternatively, converter 134 determines a scaling factor to conform to a perceptually motivated modification of this linear prediction synthesis filter, such as, for example, 1 / A (yz), with γ = 0.92 10%, for example. A perceptually motivated modification of the linear prediction synthesis filter, ie, 1 / A (y · ζ) can be called a "perceptual model.
[0039] For illustrative purposes, Fig. 12a shows another feature which is however optional for the embodiment of Fig. 12a. This element is an LTP (Long Period Prediction) filter positioned upstream of the transformer 108 to predict the audio signal over a long period. Preferably, the LP analyzer 132 operates on a version unfiltered by the long-term prediction filter. In other words, the LTP filter 136 performs LTP prediction on the audio signal 18 or its pre-enhanced version 130 and outputs the LTP rest version 138 such that the transformer 108 performs the transformation on the pre-enhanced and LTP-predicted residual signal 138. For example, the LTP filter may be implemented as a FIR filter and the LTP filter 136 may be controlled by LTP parameters including, for example, LTP prediction gain and LTP delay. Both LTP parameters 140 are encoded in the data stream. The LTP gain represents, as described in more detail below, an example of measure 60 as it indicates the pitch or periodicity that, without LTP filtering, would be fully visible in spectrum 12 and appear in spectrum 12 with a gradually reduced intensity using LTP filtering. some degree of reduction depending on the LTP gain parameter that controls the LTP filtering power through the LTP 136 filter.
[0040] Fig. 12b shows, for completeness, a decoder matching the encoder of Fig. 12a. In addition to the components of Fig. 11b and the scaling factor determining module 122 being implemented as LPC converter 142 to scaling factors, the decoder of Fig. 12b includes after the inverted transformer 126 an overlap member 144 subjecting the inverted transforms output by inverted transformer 126 to an overlap operation, thus obtaining a pre-emphasized and LTP reconstruction of the filtered version 138, which is then subjected to LTP post-filtering, the LTP post-filter 146, is one whose transfer function corresponds to that of the inverse LTP filter 136. For example, the LTP post-filter 146 may be implemented as an IIR filter. In the sequence downstream of the LTP post-filter 146, in Fig. 12b after it, the decoder of Fig. 12b includes a de-emphasis filter 148 that performs de-emphasis filtering on a time-domain signal using a transfer function corresponding to the inverse of the transfer function of the pre-enhancement filter 128. The de-emphasis filter 148 may also be implemented as an IIR filter. The audio signal 18 is produced at the output of the enhancement filter 148.
[0041] In other words, the embodiments described above provide the ability to encode tonal and frequency-domain signals by fitting the context structure of an entropy encoder, such as an arithmetic context encoder, to shape the spectrum of the signal, such as the periodicity of the signal. The embodiments described above frankly extend the context beyond the notion of neighborhood and propose an adaptive context construction based on the spectrum shape of the audio signals, such as based on pitch information. Such pitch information may be transmitted to the decoder additionally or may already be available in other coding modules such as the above-mentioned LTP gain. The context is then mapped to point to already coded coefficients that are related to the current coefficient to be encoded by a multiplicity of the distance or in proportion to the fundamental frequency of the input signal.
It should be noted that the LTP pre / post filter concept applied according to Figs. 12 and 12b can be replaced by the harmonic post filter concept where the harmonic post filter in the decoder is controlled by LTP parameters including pitch (or delay pitch). tone) sent from an encoder to a decoder via the data stream. The LTP parameters can be used as a reference for differential transmission of the above-mentioned spectrum shape information to the decoder using direct signaling.
[0043] As an embodiment outlined above, the prediction for tonal signals can be omitted, thus for example avoiding introducing undesirable interframe relationships. On the other hand, the above spectral coefficient coding / decoding concept can also be combined with any prediction technique as the prediction residuals still exhibit some harmonic structures.
In other words, the embodiments described above are illustrated again with reference to the following figures, in which Fig. 13 shows a general block diagram of an encoding operation using the spectral distance matching concept outlined above. In order to facilitate consistency between the description below and the description so far, the references are partially reused.
[0045] The input signal is first provided for noise shaping / prediction in a TD (Time Domain) module 200 TD. The module 200 includes, for example, one or both of the elements 128 and 136 in Fig. 12a. This module 200 can be bypassed or it can perform short term prediction by applying LPC coding, and / or as shown in Fig. 12a, long term prediction. Any kind of prediction can be predicted. If one of the time-domain processing operations uses and transmits pitch information as briefly outlined above by the LTP delay parameter output by 136 LTP filter, such information may then be passed to the context arithmetic encoder module for context mapping based on the pitch. .
[0046] Then, the remainder and the shaped time domain signal 202 is transformed by the transformer 108 into the frequency domain by a time-frequency transform. A DFT or MDCT transformation can be used. The length of the transformation can be adaptive and for small delays areas with small overlaps with previous and next transformation windows (and small overlaps of areas (cf. 24) will be used. We will use the MDCT as an illustrative example throughout the remainder of the document.
[0047] The transformed signal 112 is then frequency-shaped by a module 204, which is thus performed by, for example, a scaling factor determination module 116 and a spectral shaper 110. This can be done with the frequency response of the LPC coefficients and with the scaling factors controlled by the psychoacoustic model. Time Noise Shaping (TNS) or frequency domain prediction using and transmitting pitch information is also possible. In such a case, the pitch information may be communicated to the context arithmetic encoder module in view of the pitch-based context mapping. The latter possibility can also be used in the above embodiments of Figs. 10a to 12b, respectively.
[0048] The output spectral coefficients are then quantized by a quantizing step 120 before noiseless encoding by an entropy encoder 10 based on the context. As described above, the latter module 10 uses, for example, the pitch estimation of the input signal as information relating to the spectrum of the audio signal. Such information may be inherited from one of shaping / prediction modules 200 or 204 that was previously either in the time domain or the frequency domain. If the information is not available, a dedicated pitch estimation on the input signal may be performed, such as with a pitch estimation module 206, which then outputs the pitch information to the 30 bit stream.
Fig. 14 shows a general block diagram of a decoding operation matching Fig. 13. It comprises the inverse processing described in Fig. 13. Pitch information - which is used in Figs. 13 and 14 as an example of spectral shape information. - is first decoded and passed to the arithmetic decoder 40. As described above, the information is then passed on to other modules that need this information.
[0050] Specifically, in addition to the pitch information decoder 208 which decodes the pitch information from the data stream 30 and is therefore responsible for the derivation process 84 of Fig. 9b, the decoder of Fig. 14 comprises, following the context-based decoder 40 and in the order they were named, a dequantizer 210, a frequency-domain inverse shaping / prediction module 212, an inverted transformer 214, and a TD inverse noise shaping / prediction module 216, all of which are connected in series with each other so as to reconstruct from the spectrum 12 whose spectral coefficients are decoded by the decoder 40 from the 30 bit stream, the time-domain audio signal 18. When mapping the elements of Fig. 14 to those shown in, for example, Fig. 12b, the inverted transformer 214 includes the inverted transformer 126 and the overlap step 144 of Fig. 12b. Additionally, Fig. 14 shows that dequantization can be applied to the decoded spectral coefficients output by the encoder 40 using, for example, a quantization step function equal to all spectral lines. Moreover, Fig. 14 shows that a module 212 such as a TNS (Temporal Noise Shaping) may be located between the noise shaper 124 and 126. Inverse noise shaping / prediction in time domain module 216 includes elements 146 and / or 148 in Fig. 12b.
[0051] In order to re-justify the advantages provided by the embodiments of the present invention, Fig. 15 shows a conventional context for entropy coding of spectral coefficients. The context includes the boundary area of the past neighborhood of the current encoding coefficients. That is, Fig. 15 shows an example of entropy coding spectral coefficients using context matching as is, for example, used in MPEG USAC. Fig. 15 it thus illustrates the spectral coefficients in a manner similar to Figs. 1 and 2, but with the grouping or division of the spectral coefficients of the spectral proximity into clusters, called n-tuples of the spectral coefficients. In order to distinguish such n-tuples from individual spectral components, but in consistency with the description above, these n-tuples are indicated by the reference number 14 '. Fig. 15 distinguishes between already encoded / decoded n-tuples on the one hand and not yet encoded / decoded n-tuples by showing the form of one using rectangular outlines and the other using circular outlines. In addition, the n-tuple 14 'to be currently encoded / decoded is shown using hatching and circular contour, while the already encoded / decoded n-tuples 14' located by means of a fixed neighborhood template positioned on the n-tuple to be coded currently are also marked with hatch, but rectangular contour hatch. Thus, according to the example of Fig. 15, the neighborhood context template has identified six n-tuples 14 'adjacent to the n-tuple to be processed on the fly, namely n-tuples at the same instant in time but in the immediately adjacent lower spectral line (s), namely at n-tuples on the same spectral line (s) but at the immediately preceding moment in time, namely c1, n-tuples on the immediately adjacent higher spectral line at the immediately preceding moment in time, namely, c2, etc. That is, the context template used in Fig. 15 identifies a reference n-tuple 14 'at a constant relative distance from the n-tuple to be processed currently, namely immediate neighbors. According to Fig. 15, the spectral coefficients are illustratively considered in blocks of n, called n tuples. Combining n consecutive values allows you to use the relationship between the coefficients. Larger sizes exponentially increase the size of the n-tuple alphabet to encode and therefore the size of the codebook. The dimension n = 2 is used for illustration throughout the remainder of the description and represents a tradeoff between coding gain and codebook size. In all embodiments, the encoding considers, for example, the character separately. In addition, the 2 most significant bits and the remaining least significant bits of each coefficient can also be handled individually. For example, context matching may only be applied to the 2 Most Significant Bit (MBS) of the unsigned spectral values. The sign and the least significant bits can be taken to be evenly spaced. Along with the 16 combinations of 2-tuple MSBs, an output symbol, ESC, is added to the alphabet to indicate that one additional LSB should be expected by the decoder. As many ESC symbols as additional LSBs are sent. A total of 17 symbols form the code alphabet. The present invention is not limited to the described symbol generation method.
By transposing the latter specific details into the description of Figs. 3 and 4, it means the following: the alphabet of the symbols of the entropy coding / decoding engine 44 and 54 may include the values {0, 1, 2, 3} plus the output symbol and the input spectral coefficient. to be encoded is divided by 4 if it exceeds 3 as often as necessary to ensure it is less than 4 with the output symbol coding per division. Thus, 0 or more output symbols followed by an actual non-esc symbol are coded for each spectral coefficient, with only the first two of these symbols, for example, being coded using the context adaptability described above. By transferring this idea to 2-tuples, i.e. pairs of spectrally adjacent coefficients, the alphabet of symbols can include 16 pairs of values for this 2-tuple, namely {(0, 0), (0, 1), (1, 0), ..., (1, 1)} and the escape symbol esc (where esc is the abbreviation of the exit symbol), i.e., 17 symbols in total. Each input n-tuple of spectral coefficients containing at least one coefficient exceeding 3 is subjected to a division by 4 applied to each coefficient of the given 2-tuple. On the decoding side, the number of output symbols times 4, if any, is added to the residual value obtained from non-esc symbol.
[0053] Fig. 16 shows a mapping configuration of the mapped context resulting from the modification of the concept of Fig. 15 according to the concept outlined above that the relative spectral distance 28 of the reference spectral coefficients is adjusted depending on the spectrum shape information, such as for example by taking into account the periodicity or pitch information of a signal. In particular, Fig. 16a to 16c show that the distance D, which corresponds to the above-mentioned relative spectral distance 28, within the context can be roughly estimated by D0 given by the following formula:
; /. Έ '2N TO - -<sup>χ</sup> -z - Lj Ig where fs is the sampling rate, N is the MDCT size and L is the delay period in the samples. In the example of Fig. 16 (a), the context indicates n-tuples distant from the current n-tuple to be encoded by the D-fold. Fig. 16 (b) links the conventional neighborhood context to a harmonic context. Finally, Fig. 16 (c) shows an example of an interframe mapped context without dependence on previous frames. That is, Fig. 16a shows that, in addition to the possibilities set out above with reference to Fig. 7, fitting the relative spectral distance depending on the spectral shape information can be applied to all of the fixed number of reference spectral coefficients belonging to the context template. Fig. 16b shows that, according to another example, only a subset of these reference spectral coefficients is displaced according to adaptivity 80, such as for example only the spectrally outermost from the low frequency side of the context template, here C3 and Cs. The remaining reference spectral coefficients, here Co to C4, may be positioned at fixed positions with respect to the currently processed spectral coefficient, namely in immediately adjacent spectral-time positions with respect to the spectral coefficient to be processed on an ongoing basis. And finally, Fig. 16c shows that the possibility that only these previously encoded spectral coefficients are used as reference context template coefficients that are located at the same time instant as the spectral coefficients to be encoded on-the-fly.
[0054] Fig. 17 shows an illustration of how the mapped context of Figs. 16a-c may be more efficient than the conventional context of Fig. 15, which cannot predict the highly harmonic tone of the X-spectrum (cf. 20).
[0055] Next, we will detail the potential context mapping mechanism and provide demonstrative embodiments for performance estimation and D distance coding.
[0056] For illustrative purposes, in the following sections we use the interframe mapped context of Fig. 16c.
First embodiment: 2-tuple encoding and mapping
[0057] First, the optimal distance is searched for in a way that minimizes the number of bits needed to code the currently quantized spectrum x [] of size N. The initial distance can be estimated using a function D0 of the delay period L found in the previously performed pitch estimation. The scope of the search may be as follows:
DO - Δ <D <DO + Δ
[0058] Alternatively, the range can be corrected by taking the fold D0 into account. It becomes the extended scope
<img file="PL3058566T3_D0001.tif" />
where M is a multiplicative coefficient belonging to a finite set of F. For example, M may take the values 0.5, 1 and 2 for the purpose of half pitch and double pitch exploration. Finally, exhaustive D searches can also be performed. In practice, this latter approach can be too complex. Fig. 18 shows an example of a search algorithm. This search algorithm may be, for example, part of the acquisition method 82 or both of the acquisition methods 82 and 84 on the coding and decoding side.
[0059] Cost is included in the cost when no context mapping is performed. If no distance leads to a better cost, no mapping is performed. A flag is sent to the decoder to signal when the mapping is performed.
[0060] If the optimal distance Dopt is found, it should be transmitted. If L has already been transmitted by another encoder module, the matching parameters mi for the above-mentioned direct signaling of Fig. 9b have to be transmitted such that
<img file="PL3058566T3_D0002.tif" />
[0061] Otherwise, the absolute value Dopt must be transmitted. Both alternatives have been discussed above with reference to Fig. 9b. For example, if we take into account MDCT of size N = 256 and fs = 12800 Hz, we can cover the pitch frequency between 30 Hz and 256 Hz by limiting D between 2 and 17. At integer resolution, D can be encoded by 4 bits, and by 5 bits for a resolution of 0.5 and 6 bits for a resolution of 0.25.
[0062] The cost function may be computed as the number of bits needed to encode x [] with D used to generate a context mapping. This cost function is usually complex to obtain because it requires the encoding of the arithmetic spectrum or at least having a good estimation of the required number of bits. Since this cost function can be complex in computing for each candidate D, we alternatively propose to obtain the cost estimate directly from the derivation of the context mapping from the value of D. When deriving the context mapping, it is easy to compute the norm difference of the adjacent mapped context. Since context is used in an arithmetic encoder to predict n-tuples for encoding, and since the context is computed in our preferred embodiment based on the L1-norm, the sum of the norm difference between adjacent mapped contexts is a good mapping performance indicator given D. First is computed the norm of each 2-tuple for x [] as follows:
Aor (i = 0; i <N / 2; i ++) {>> ^ normVect [i] = pow (abs (x [2 ^] NORM pow (abs (normVećt [2 * i + ^^^
[0063] With NORM = 1 in the preferred embodiment because we include norm-L1 in calculating the context. In this section, we describe a context mapping which works with a resolution of 2, i.e., one mapping per 2 tuple. The resolution is r = 2 and the context mapping table is N / 2. The pseudo-code for generating context mapping and calculating the cost function is shown below
Input resolution
Input: normVect [N / r]
Output contextMapping [N / r] mMp ui = (int) (rnFNr)); / k- = H; - 'w meanDiffNorm = oldNorm ^ O;
/ * Detect Harmonics spectrum 7 while (i <= N / mpreroll) f 7 'for (o = 0; o <preroll; o ++ ^ meanDiffNorm + = abs (norm Vect [i] -oldNorm);
oldNorm = normVect [i]; ..... yy lndexPermutation [k ++] = i; > y> iy 1 tu / * Detect valleys from spectrum 7> 11 '·> yy Slidelndex = k; yy>.
= 0; i? y 71 for (o = 0; o <k; o + -prerollj {fy ę for (; i <lndexPermutation [o]; i ++) {(i meanDiffNorm + = abs (normVect [i] -oldNorm); oldNorm = normVect [i];, lndexPermutation [Slidelndex ++] = i; r '} ......
/ * skip tonal component * / y1 1 yy i + = preroll; } o / y yy \ x
7 * Detect taił of spectrum * / // ż p- - yyyy for (i = Slidelndex; i <numVect; ++ i) {· y · meanDiffNorm + = abs (normVect [i] -oldNorni); oldNorm = normVect [i]; ....... y lndexPermutation [i] = i; hand / Jp
[0064] Once the optimal distance D has been calculated, the index permutation table is also inferred, which gives the harmonic positions, troughs and spectrum end.
The context mapping rules are then inferred as: for (i = 0; i <'Wr; i ++)' {/ ^
[0065] This means that for a 2-tuple index and in the spectrum (x [2 *], x [2 * + 1]), the past context will be considered with the 2-tuple indexes contextMapping [i-1], contextMapping [i2] ... ContextMapping [iI], where I is the context size in the range of 2 tuples. If one or more of the previous spectra is also included for the context, 2 tuples for those spectra included in the past context will have the indexes contextMapping [i + I], ..., contextMapping [i + 1], contextMapping [i], contextMapping [and -1], contextMapping [iI], where 2I + 1 is the context size of the previous spectrum.
[0066] The IndexPermutation table also provides additional interesting information as it collects indices of tonal components followed by indices of non-tonal components. Consequently, we can expect the corresponding amplitudes to decrease. This can be exploited by detecting the last index in 5 IndexPermutaion that corresponds to a nonzero 2 tuple. This index corresponds to (lastNz / 2-1), where lastNz is computed as:
for (lastNz = (N-2); lastNz> = O; lastNz - = 2) / 'i (if ((x [2 * lndexPermutaion [lastNz / 2]]! = 0) 11 (x [2 * indexPermutaion [ lastNz / 2] +1]! = 0)) break; "xz ;;. <· ·<sup>ζ</sup> lastNz + = 2; nii lastNz / 2 is coded on ceil (log2 (/ V / 2)) bits before the spectral components, γ
Arithmetic encoder pseudo code:
[0067]
Input: spectrum x [N] z. <. <oZ Zz<sup>;</sup>
Input: contextMapping [N / 2] ZZ <ZZ input: lastNz ... z Z: z ZZ? .Z ZZ Z Output: coded bitstream Z; From t z. s Z • ΖΖ · \ ·; ZZ7 Zt 'i ^. while ((i <N / 2) && (contextMapping (i]> AastNz / 2)) {. 'z' z cont ^ = -1; f zz zzzz ZZ break; i; y
i) a ^ ai = abs (x [2A]) ;; A b = bl = Abs (^^ t = (context [contextMapping [i-2] «6) + context [contextMapping [i-1]; while ( (a1> = 4) \\ (b1> = 4)) ..........
; A y: / * encode escape pki = proba_modeHookup [t]; y;: a ari_encode (cum__proba [pki], 16.17);
(a1) »= 1; -: ^ 7; (
A: yy (b1) >> = - 1; 7d> y 7 y '/ * encode LSBs * / / 7 A 7 x7 · 7 y 7ari_encode (cum ^ quiproba, a1 & 1,2T ^^ ari ^ ncode (cum ^ quiproba, b18A, 2); y; y7 yy} y aA .....: /. 7 / * encode MSSs * /::
pki = proba_model_lookup [t]; and γ ari_encode (cum_proba [pki], a1 + 4 * b1.17); y / * encode signs * / ya
A 7A y; yy ari_encode (cum_equiproba, x [2 * i]> 0.2); a y7 7 / χ · α 'i ^<sup>}</sup> 7/ 7;^;7
7lf (b> O} 7yA: /: 7: / :)): · ·: L · X
A ^ iy-byT / * Update context * / /: :)::
context [contextMapping ^
[0068] The cum_probe tables are different cumulative models acquired during offline training on a large training set. In this particular case, it contains 17 symbols. proba_model_lookup (\ is a lookup table mapping context t into the cumulative pki probability model. This table is also obtained through the training phase. cum_equiprob [] is a cumulative probability table for alphabet 2 symbols having equal probability.
Second embodiment: 2-tuple with 1-tuple mapping
[0069] In this second embodiment, the spectral components are still coded
2 tuple by 2 tuple but contextMapping now has a resolution of 1 tuple. This means that there are much more possibilities and flexibility in context mapping. The mapped context may then be better suited to the given signal. The optimal distance is searched in the same way as it was done in section 3, but this time with a resolution r = 1. For this, normVect [] has to be computed for each MDCT line:
for (i = 0; i <N; i ++) {;; / normVect [i] = pow (abs (x [2 * i] NORM,);
[0070] The resulting context mapping is then given by a table of N dimension. LastNz is computed as in the previous sections and the encoding can be described as follows:
JnputlastNz
Input: contextMapping [N]> <'<- /.
Input: spectrum x [N] output: coded bitstream tocal: context [N / 2] o -i for / k = O, / = 0; k </ astnz '; k + = 2) {7 i '/ *' Next coefficient to-codę * / ię. 'while (contextMapping [i]> = lastnz) / ++;
j / * Next coefficient to codeN / 2 while (contextMappm ^ / ++; / / * Get context for the lowest index7 iii i i__min = min (contextMapping [a1_i], contextMapping [b17]); t = context [(i_min / 2) -2] & lt; 6 + context [(i_min / 2) -1];
7 * Init Gurrent 2-tuple encoding 7 li ii: iii a-a1 = abs (x [a1__i]); Ii ii iiyb = b1 = abs (x [b1J]);
7whiie ((aT> -4) f ii liii i 7 * encode escape symbol * / iii; i ii pki = proba_model_lookup [t]; i 1 i · ari_encode (cum_proba [pki], 16.16); i (ai) »= 1; ...... Żi ii i> Z (b1)» = 'tiii · / * encode LSBs7 il iii ari_encode (cum_equiproba, a1 & 1,2); 1 ari_encode (cum_equiproba, b1 & 1,2); ii ;
o ..... iii-ii / * encode MSBs7 itiii pki = proba_model_lookup [t]; ii - i + ii 1 ari_encode (cum_proba [pki], ai + 4 * b1.16);
/ * encode signs7 ii \ i; ii if (a> 0) ari_encode (cum_equipm ^^ if (b> 0) arFencode (cum_equiproba, x [2 * i + 1]> 0.2);
/ * update context7 f 7 7 if (contextMapping [a 1Α]! = (contextk context [contextMapping [a1_i] / 2] = min (a + a, power (2,6)); context [contextMapping [b1_i] / 2 ] = min (b + b, power (} else {; i ii -ii context [ćontexMapping [a 1_i] / 2] = min (a + b, power (2,6));
context [contextMapping [b1J] / 2] -min (a + b, power (2.6));
[0071] Contrary to the previous section, two non-consecutive spectral coefficients can be combined in the same 2-fold. For this reason, the context mapping for two 2 tuple elements may point to two different indexes in the context table. In the preferred embodiment, we select the mapped context with the lowest index, but you may also have another principle like averaging the two mapped contexts. For the same reason, updating the context should also be done differently. If 2 elements are consecutive in the spectrum, we use the conventional method of computing context. Otherwise, the context is updated separately for 2 items including only its own module.
[0072] The decoding consists of the following steps:
• Flag decoding to check if context mapping is performed • Context mapping decoding, by decoding either Dopt or parameter matching parameter to obtain Dopt for D0 • lastNz decoding • Quantized spectrum decoding as follows:
Input: lastNz?
Input contextMapping [N] '/ you
Input: coded bitstream 7 (local: context [N / 2]. <Output: quantized spectrum x [Ń] y for (k = 0, i = 0; k <lastnz; k + = 2) {r <sub>and</sub>= b = 0; tat F / * Next coefficient to code * / i while (contextMapping [i]> = lastnz) x [i ++] = 0;
/ * Next coefficient to codę * / / ęwhile (contextMapping [i]> = iastnż) / * Get contexf for the lowest index * / ...... 't = cont ^ / * Init current 2-tuple en co ding V <'. 7 i \; a = a1 = abs (x [a1J]); : - / -: 0 ^ ó b = b1 - abs (x [b1J]); : about /
AMSBs'decoding * / yy yy7; <for (lev = O ;;):: ć: /; pki = proba_modenookup [t]; > 'i. · / R = an_decode (cum_proba [pki], 16); '.7 y y-pf.: If (r <16) {........... ...... <. A and. . break; 'f? if ALSBsdecodingA --3 = (3) + ań_decode (cum__equiproba, 2) ^ (1 ^)) b = (b) + ari_decode (cum_equiproba, 2) «(lev)); - ::; ':
: lev + = 1; : _ _ ...... ........ ? / i. '· Ć
Ab1 = i >> 2; H ^, a1 = r & 0x3; :;:: i: - e: A a + = (a1) «lev; q. p 'i ii iy i. b + = (b1) «lev; '' 'and fi? ii '7? i? / qjpdate'Confexf7iiiy. y; ? and if (contextMapping [a1_ii! ~ (contextMapping [b 1f] -i)) {context [contextMapping [a 1_i] / 2] = min (a + a, power (2,6));
context [contextMapping [b1J] / 2] = min (b + b, power (2.6));
:} eise {..... . j. -yx. cont & xt [contextMapping [a 1A] / 2] = nm ^^ 6)!<sup>1</sup> context [contextMapping [b1__i] / 2] = min (a + b, power (2.6));
y./kf: Adecode isignsHA / UF if (a> 0) <sub>r</sub>2) +1); '' 'Z if (b> 0) - bb * (- 2<sup>loam</sup>anAdeGOde (cufp_eqliipr () or A) +1); y
7 * -Store decoded data * /: x [a1_i] <sup>=</sup> and;
: <sup>χ</sup>[Α1_ί] - b; u :: - I
[0073] Thus, the above embodiments, inter alia, have disclosed, for example, pitch-based context mapping for entropy coding such as arithmetic of tonal signals.
[0074] While certain aspects have been described in the context of an apparatus, it is evident that these aspects also represent a description of a corresponding method wherein the block or device corresponds to a method step or a feature of a method step. Likewise, aspects described in the context of a method also represent a description of a corresponding block or item or feature of the corresponding device. Some or all of the steps of the method may be performed by a hardware device or by a hardware device such as, for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some or all of the most important steps of the method may be performed by such a device.
[0075] The inventive encoded audio signal can be stored on a digital storage medium or it can be transmitted via wireless transmission means or wired transmission means such as the Internet.
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. Implementation may be via a non-transient storage medium such as digital storage media such as floppy disks, DVDs, Blue-Ray Discs, CDs, ROMs, PROMs, EPROMs, EEPROMs or FLASHs containing electronically readable signals stored thereon controls that interact (or are capable of such interaction) with the programmed computer system, such that the appropriate method is performed. In this way, the digital storage medium can be computer readable.
[0077] Some embodiments may include a data carrier containing electronically readable control signals that are operable to interact with a programmable computer system, such that one of the methods described herein is performed.
[0078] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operable to perform one of the inventive methods when the computer program product runs on a computer. For example, the program code may be stored on a machine-readable medium.
[0079] Other embodiments include a computer program for performing one of the methods described herein stored on a machine-readable medium.
In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program product runs on a computer.
[0081] Another embodiment is, therefore, a data medium (or a digital storage medium or a computer readable medium) containing a computer program recorded thereon for performing one of the methods described herein. The data medium, digital storage medium, or recorded medium are typically real and / or non-transient.
[0082] Another embodiment, therefore, is a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may e.g. be configured to be transmitted over a data link, e.g.
[0083] Another embodiment comprises processing means, for example a computer or programmable logic device configured or adapted to perform one of the methods described herein.
[0084] Another embodiment includes a computer in which the computer program for performing one of the methods described herein is installed.
[0085] Another embodiment comprises an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver, for example, may be a computer, mobile device, memory device, etc. The device or system may, for example, include a file server for transmitting the computer program to the receiver.
[0086] In some embodiments, a programmable logic device (e.g., a user programmable logic table) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the user programmable logic table may interact with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0087] The above described embodiments are merely illustrative for the principles of the present invention. It should be understood that modifications and variations to the systems and details described herein are apparent to those skilled in the art. It is the intention, therefore, to be limited only by the scope of the following patent claims, and not by the specific details presented for the description and explanation of the present embodiments of the invention.
Literature
[0088]
[1] Fuchs, G .; Subbaraman, V .; Multrus, M., Efficient context adaptive entropy coding for real-time applications, Acoustics, Speech and Signal Processing (ICASSP), 2011 IEEE International Conference on, vol., Pages 493, 496, May 22-27, 2011
[2] ISO / IEC 13818, Part 7, MPEG-2 AAC
[3] Juin-Hwey Chen; Dongmei Wang, Transform predictive coding of wideband speech signals, Acoustics, Speech, and Signal Processing, 1996. ICASSP-96. Conference
Proceedings., 1996 IEEE International Conference on, vol. 1, no., Pages 275, 278 vol. 1, May 7-10, 1996
Fraunhofer Gesellschaft zur Forderung der angewandten Forschung e. V., Germany
Proxy:
EP 3 058 566 B1
Z-16890/18
Contents2
50 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50
37 members in 17 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 13189391 | European Patent Office (EPO) | A | |
| 13189391 | European Patent Office (EPO) | A | |
| 14178806 | European Patent Office (EPO) | A | |
| 14178806 | European Patent Office (EPO) | A | |
| 14792420 | European Patent Office (EPO) | A | |
| 2014072290 | European Patent Office (EPO) | W | |
| 2014072290 | European Patent Office (EPO) | W | |
| 13189391 | – | – | – |
| 14178806 | – | – | – |
| 147924203 | – | – | – |
| EP20130189391 | – | – | – |
| EP20140178806 | – | – | – |
| EP20140792420 | – | – | – |
| WO2014EP72290 | – | – | – |
Members37
| Document | Office | Kind | |
|---|---|---|---|
| CA2925734A1 | Canada | A1 | |
| WO2015055800A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201521015A | Taiwan Province of China | A | |
| AR098074A1 | Argentina | A1 | |
| AU2014336097A1 | Australia | A1 | |
| KR20160060085A | Republic of Korea | A | |
| SG11201603046RA | Singapore | A | |
| MX2016004806A | Mexico | A | |
| CN105723452A | China | A | |
| EP3058566A1 | European Patent Office (EPO) | A1 | |
| US2016307576A1 | United States of America | A1 | |
| JP2017501427A | Japan | A | |
| AU2014336097B2 | Australia | B2 | |
| TWI578308B | Taiwan Province of China | B | |
| EP3058566B1 | European Patent Office (EPO) | B1 | |
| RU2016118776A | Russian Federation | A | |
| RU2638734C2 | Russian Federation | C2 | |
| US9892735B2 | United States of America | B2 | |
| KR101831289B1 | Republic of Korea | B1 | |
| PT3058566T | Portugal | T | |
| ES2660392T3 | Spain | T3 | |
| US2018122387A1 | United States of America | A1 | |
| MX357135B | Mexico | B | |
| CA2925734C | Canada | C | |
| PL3058566T3This record | Poland | T3 | |
| JP6385433B2 | Japan | B2 | |
| US10115401B2 | United States of America | B2 | |
| JP2018205758A | Japan | A | |
| US2019043513A1 | United States of America | A1 | |
| CN105723452B | China | B | |
| CN111009249A | China | A | |
| JP6748160B2 | Japan | B2 | |
| US10847166B2 | United States of America | B2 | |
| JP2020190751A | Japan | A | |
| MY181965A | Malaysia | A | |
| CN111009249B | China | B | |
| JP7218329B2 | Japan | B2 |
Numbers
- Publication
- 3058566
- Publication, DOCDB
- 3058566
- Publication, EPODOC
- PL3058566T
- Application
- 14792420
- Application, DOCDB
- 14792420
- Application, EPODOC
- PL20140792420T
Titles2
- English
- CODING OF SPECTRAL COEFFICIENTS OF A SPECTRUM OF AN AUDIO SIGNAL
- Polish
- Kodowanie współczynników widmowych widma sygnału audio
Classification
- CPC, 5
- G10L19/0017
- G10L19/00
- G10L19/032
- H03M7/30
- G10L19/02
- IPC, 2
- G10L19 032
- G10L19 00