Temporal stereo signal processing method for forming scaled bit stream
Abstract
The method involves transforming the two temporal channels (l,r) into the frequency domain, and forming a spectral mono channel (M) by combining the spectral L and R channels. Three predictions are performed for the spectral L,R and M channels to obtain filtered channels (L'.R',M'). The filtered channels are encoded to obtain the mono-layer of the scaled bit-stream. The encoded filtered mono channel is decoded to obtain an encoded/decoded mono channel (M). The filtered L' and R' channels and the encoded/decoded mono channel are subjected to a prediction using similar prediction coefficients. Spectral stereo signals (L, R, M, S) are formed for the stereo layer of the scaled bit stream by comparing the processed mono channel (M') with the processed filtered channels L' and R'. and/or by combining them. Independent claims are included for a method of processing a temporal stereo signal, a method of decoding an audio bit stream and an apparatus for processing a temporal stereo signal.

Term
Term ended
Projected expiry passed 30 June 2018, 8.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
28 claims: 9 independent, 19 dependent
- 1A method for processing a temporal Stereosi signal continuously, the first time a (l) and a time union second (r) channel, a scaled bitstream ( 100 ) With a monolayer and a stereo layer to obtain, with the following steps:Transform ( 14 . 16 ) Of the time the first (l) and the time second (r) channel in the Frequency Ranges rich;Form ( 28 . 30 ) Of a spectral channel mono (M) by Combination of spectral first (L) and the spektra len second (R) channel;Carry out ( 18 . 20 . 32 ) Of a first, second and third prediction over the frequency of the spectral first channel (L), the spectral second channel (R) and the spectral mono channel (M), one he filtered first channel (L '), a filtered second channel (R') or to obtain a filtered mono channel (M ');encoding ( 34 ) Of the filtered mono channel (M ') to the Monolayer ( 36 ) The scaled bit stream ( 100 ) to receive;Decoding ( 34 ) The encoded filtered mono channel, an encoded / decoded mono channel (M '') to it hold;To treat ( 38 . 40 ) Of the filtered first (L ') and the second channel (R ') and the coded / decoded Mono-channel (M ''), such that the three treated Ka channels a prediction with similar Prädiktionskoeffi coefficients are subjected;and Form ( 22 . 24 . 46 a, 46 b, 44 a, 44 b, 48 a, 48 b, 50 a, 50 b, 52 ) A first and a second spectral stereo signal (L '', R '', M v , S) for the stereochemistry of the layer scaled bit stream ( 100 ) By comparing ( 46 a, 46 b, 22 . 24 . 52 ) Of the treated single channel (M '' ') with the treated filtered first (L ') and second channel (R ') and / or a combination of the treated filtered first (L ') and second (R') channel.
- 9A method for processing a temporal Stereosi signal continuously, the first time a (l) and a time union second (r) channel, a scaled bitstream ( 100 ) With a monolayer and a stereo layer to obtain, with the following steps:Form ( 142 a, 142 b, 144 ) A temporal mono channel (M) from the first (l) and the second (R) channel;encoding ( 140 ) Of the time channel mono (m) to the to obtain monolayer of the scaled bit stream;Transform ( 14 . 16 ) Of the first (l) and the second (R) channel in the frequency domain;Forming a spectral mono channel (M) by Kombina tion from the spectral first (L) and the spectral second (R) channel;Decoding ( 140 ) And transforming ( 150 ) The encoded temporal mono channel in the frequency domain to a encoded / decoded spectral mono channel (M CD ) To it hold;Carry out ( 18 ) A first prediction about the Fre sequence with the spectral first channel (L), a ge filtered first channel (L ') to obtain;Carry out ( 20 ) A second prediction about the Fre sequence with the spectral second channel (R) by a ge filtered second channel (R ') to obtain;Carry out ( 172 . 174 ) A third prediction about the frequency with the encoded / decoded spectral Mono channel (M CD ) And with the spectral channel mono (M), wherein prediction of the first ( 18 ) or second ( 20 ) Prediction can be used to provide a L / R-filtered coded / decoded mono channel (M CD''' ) And a L / R-filtered mono channel (M '' ') to receive;Compare ( 154 . 156 ) Of the L / R-coding filtered th / decoded mono channel (M CD''' ) With the L / R-ge filtered mono channel (M '' ') to an L / R-comparison Mono channel (M iv CD ) to obtain;andForm ( 22 . 24 . 46 a, 46 b, 44 a, 44 b, 48 a, 48 b, 50 a, 50 b, 52 ) A first and a second spectral Stereosi gnals (L '', R '', M v , S) for the stereo layer of ska profiled bitstream by comparing the L / R-Ver equalization-mono channel (M CD iv ) With the filtered first Channel (L '), the filtered second channel (R') and with a combination ( 48 a, 48 b, 50 a, 50 b, 52 ) from the filtered first (L ') and said filtered second (R ') channel.
- 11The method according to any one of the preceding claims, wherein the step of forming a first and a second spectral stereo signal (L '', R '', M v , S) fol lowing substeps:Subtract ( 50 a) double ( 44 a) treated filtered coded / decoded mono channel (M '' ') or double ( 44 a) comparison mono channel (M iv ) of the filtered left channel (L ');Subtract ( 46 b) twice ( 48 b) filtered co coded / decoded mono channel (M ') or the double ( 48 b) comparison mono channel (M iv ) Of the filtered right channel (R ');Compare ( 22 . 24 ) The subtraction with egg nem threshold;andUse ( 22 . 24 ) The subtraction when he stes and second spectral stereo signal (L ', R' ') if the threshold is exceeded, otherwise, Use ( 22 . 24 ) Of the filtered left channel (L ') and the filtered right channel (R ') as the first and a second spectral stereo signal (L ', R' ').
- 15A method for decoding a using a Prediction over the frequency encoded audio bitstream ( 100 ), The page information ( 126 ) Which on the audio bit stream ( 100 ) Underlying coding point, with the following steps:demultiplexing ( 102 ) Of the audio bit stream ( 100 ), A Monolayer and a stereo layer to obtain, on Because of the side information;Decoding ( 108 ) Of the monolayer using a by the side information specific Decodieralgo algorithm to obtain a decoded mono channel;requantizing ( 104 . 106 ) Stereo layer, a he stes and a second spectral stereo signal (L ', R' ', M v , S) to obtain;To treat ( 110 . 112 ) Of the first and second stereo signal (L '', R '', M v , S) and the decoded mono channel, such that the two stereo signals and the de coded mono channel a prediction with similar pre diction coefficients are subjected;Combine ( 114 . 116 . 118 . 120 . 122 ) Of the treated Mono-channel (M '' ') with the treated first and two th spectral stereo signal (L ', R' ') to a gefil failed left channel (L ') and a filtered right to obtain channel (R '), due to the Seiteninformatio NEN;Carry out ( 130 . 132 ) An inverse prediction about the frequency with the filtered left channel (L ') and the filtered right channel (R '), a spectral left (L) and a spectral right (R) channel to obtained by using the second or third in the page information existing Prädiktionskoeffi cient;andInverse transforming ( 134 . 136 ) Of spectral left (L) and the spectral right (R) channel in the Zeitbe rich to obtain a temporal stereo signal a time left (l) and a temporal computationally having ten (r) channel.
- 18A method for decoding a using a Prediction over the frequency encoded audio bitstream ( 100 ), The page having information on the the audio bit stream (BS) underlying coding towards point, comprising the steps of:demultiplexing ( 102 ) Of the audio bit stream ( 100 ), A Monolayer and a stereo layer to obtain, on Because of the side information;Decoding ( 108 ) Of the monolayer using a by the side information specific Decodieral algorithm to obtain a decoded mono channel;Transform ( 162 ) Of the decoded mono channel in the Frequency domain to a spectral decoded mono channel (M CD ) to obtain;requantizing ( 104 . 106 ) Stereo layer, a he stes and a second spectral stereo signal (L ', R' ', M v , S) to obtain;Carry out ( 178 ) A prediction over the frequency the decoded mono channel (M CD ) To an L / R-gefil screened mono channel (M CD''' ) To obtain, using of the page information ( 126 ) Existing first or second prediction coefficients at a Prediction over the frequency with the left (L) or Right (R) channel WUR determined during encoding the;Combine ( 120 ;122 . 116 . 118 ) Of the L / R-filtered encoded / decoded mono channel (M CD''' ) with him first and second spectral stereo signal (L '', R ''), a filtered left (L ') and right (R) Ka nal to obtain, due to the side information;Carry out ( 130 . 132 ) An inverse prediction about the frequency with the filtered left channel (1 ') and the filtered right channel (R '), a spectral left (L) and a spectral right (R) channel to obtained by using the second or third in the page information existing Prädiktionskoeffi cient;andInverse transforming ( 134 . 136 ) Of spectral left (L) and the spectral right (R) channel in the Zeitbe rich to obtain a temporal stereo signal a time left (l) and a temporal computationally having ten (r) channel.
- 22An apparatus for processing a temporal stereo signal that a time first (l) and a time second (r) channel, a ska profiled bitstream ( 100 ) With a mono-layer and a Stereo layer to obtain, having the following features:a means for transforming ( 14 . 16 ) The time union first (l) and the time the second (r) channel into the frequency domain;means for forming ( 28 . 30 ) A spectral Mono channel (M) by combining the spectral first (L) and the second spectral (R) channel;Means for performing ( 18 . 20 . 32 ) A he first, second and third prediction over the frequency the spectral first channel (L), the spectral second channel (R) and the spectral mono channel (M) to a filtered first channel (L '), a filtered second channel (R ') or a filtered monaural channel (M ') to obtain;means for encoding ( 34 ) The filtered Mo nokanals (M ') to the monolayer ( 36 ) The scaled bitstream ( 100 ) to obtain;means for decoding ( 34 ) The encoded ge filtered mono channel to a coded / decoded Mono channel (M '') to obtain;a means for treating ( 38 . 40 ) Of gefilte th first (L ') and second channel (R') as well as the CO coded / decoded mono channel (M ''), such that the three treated channels a prediction with similar Prediction are subjected;andmeans for forming ( 22 . 24 . 46 a, 46 b, 44 a, 44 b, 48 a, 48 b, 50 a, 50 b, 52 ) A first and a second spectral stereo signal (L '', R '', M v , S) for the stereo layer of the scaled bit stream ( 100 ) by Compare ( 46 a, 46 b, 22 . 24 . 52 ) Of the treated Mo nokanals (M '' ') with the treated filtered first (L ') and second channel (R') and / or a combination from the treated filtered first (L ') and second (R ') channel.
- 24An apparatus for processing a temporal Stereosi signal continuously, the first time a (l) and a time union second (r) channel, a scaled bitstream ( 100 ) With a monolayer and a stereo layer to obtain, having the following features:means for forming ( 142 a, 142 b, 144 ) one temporal mono channel (m) from the first (l) and the second (r) channel;means for encoding ( 140 ) The time Mo nokanals (m) to the monolayer of the scaled bit stream to obtain;a means for transforming ( 14 . 16 ) Of it first (l) and the second (R) channel in the Frequency Ranges rich;means for forming a spectral Monoka Nalles (M) by a combination of the first spectral (L) and the spectral second (R) channel;means for decoding ( 140 ) And Transformie ren ( 150 ) The encoded temporal mono channel in the Frequency range to an encoded / decoded spectral tral mono channel (M CD ) to obtain;means for performing ( 18 ) A first Prediction over the frequency using the first spectral Channel (L) to produce a filtered first channel (L ') to receive;means for performing ( 20 ) A second Prediction over the frequency with the second spectral Channel (R) to produce a filtered second channel (R ') to receive;means for performing ( 172 . 174 ) a third prediction over the frequency with the encoder th / decoded spectral mono channel (M CD ) And the spectral mono channel (M), wherein Prädiktionskoeffizien th of the first ( 18 ) Or second ( 20 ) Prediction ver be used to a L / R filtered encoded / de coded mono channel (M CD''' ) Or a L / R-filtered Mono channel (M '' ') to obtain;means for comparing ( 154 . 156 ) Of the L / R filtered encoded / decoded mono channel (M CD''' ) with the L / R-filtered mono channel (M '' ') to form a L / R comparison mono channel (M iv CD ) to obtain;andmeans for forming ( 22 . 24 . 46 a, 46 b, 44 a, 44 b, 48 a, 48 b, 50 a, 50 b, 52 ) A first and a second spectral stereo signal (L '', R '', M v , S) for the stereo layer of the scaled bit stream by Ver same of the L / R comparison mono channel (M CD iv ) with the filtered first channel (L '), the filtered second Channel (R ') and with a combination ( 48 a, 48 b, 50 a, 50 b, 52 ) Gefil from the filtered first (L ') and the failed second (R ') channel.
- 25An apparatus for decoding a using a Prediction over the frequency encoded audio bitstream ( 100 ), The page information ( 126 ) Which on the audio bit stream ( 100 ) Underlying coding point, with the following characteristics:a means for demultiplexing ( 102 ) Of Audiobit current ( 100 ), A monolayer and a stereo layer to obtain, due to the side information;means for decoding ( 108 ) Of the monolayer using the page information specific decoding algorithm to a decoded Mono channel to obtain;means for requantizing ( 104 . 106 ) of the Stereo layer to a first and a second spectral Stereo signal (L '', R '', M v , S) to obtain;means for performing ( 110 ) An inverse Prediction over the frequency with the decoded mono channel using the page information ( 126 ) Existing first prediction coefficients with a prediction of the mono channel during Codie proceedings were determined to an unfiltered coding th / decoded mono channel (M '') to obtain;means for performing ( 112 ) A Prädik tion on the frequency with the unfiltered decoded th / decoded mono channel (M ''), a L / R gefilte th mono channel (M '' ') to obtain, by using in the page information ( 126 ) Existing second or third prediction that at a Prediction over the frequency of the left (L) or computationally th (R) channel were determined during encoding;means for combining ( 114 . 116 . 118 . 120 . 122 ) Of the L / R-filtered mono channel (M '' ') it with the first and second spectral stereo signal (L '', R ''), a filtered left channel (L ') and a gefil failed to obtain right channel (R '), due to the Page information;means for performing ( 130 . 132 ) A in verses prediction over the frequency with the filtered left channel (L ') and the filtered right channel (R '), a spectral left (L) and a spektra len right (R) channel, obtained by using the second and third in the side information EXISTING which prediction;andmeans for inverse transforming ( 134 . 136 ) of spectral left (L) and the spectral Right (R) Channel in the time domain a temporal Stereosi signal to get that and a time left (l) a temporal right (r) channel has.
- 27The apparatus for decoding a using a Prediction over the frequency encoded audio bitstream ( 100 ), The page having information on the the audio bit stream (BS) underlying coding towards have, with the following features:a means for demultiplexing ( 102 ) Of Audiobit current ( 100 ), A monolayer and a stereo layer to obtain, due to the side information;means for decoding ( 108 ) Of the monolayer using the page information specific decoding algorithm to a decoded Mono channel to obtain;a means for transforming ( 162 ) Of the deco coded mono channel in the frequency domain to a spectral decoded mono channel (M CD ) to obtain;means for requantizing ( 104 . 106 ) of the Stereo layer to a first and a second spectral Stereo signal (L '', R '', M v , S) to obtain;a means for treating ( 110 . 112 ) of the first and second stereo signal (L '', R '', M v , S) and the decoded mono channel, such that the two Stereosi signals and the decoded mono channel with a prediction similar prediction are subjected;means for combining ( 114 . 116 . 118 . 120 . 122 ) Of the treated mono channel (M '' ') with the behan punched first and second spectral stereo signal (L '', R '') to a filtered left channel (L ') and to obtain a filtered right channel (R '), on Because of the side information;andmeans for inverse transforming ( 134 . 136 ) of spectral left (L) and the spectral Right (R) Channel in the time domain a temporal Stereosi signal to get that and a time left (l) a temporal right (r) channel has.
Independent claims9
106 paragraphs, as filed
The present invention relates to the encoding and Decoding of audio signals and in particular to bit rate scalable encoder or decoder, the stereo and mono signals can be processed, wherein at least in the stereo encoding a temporal noise shaping (TNS TNS = Temporal Noise Shaping) is implemented.
Scalable audio coders are coders who grew modular builds are. So is the aspiration, existing to use speech coder, the signals z. B. 8 kHz sampled, processed and data rates in for example from 4.8 to 8 kilobits per second output. These known coders such. as known to those skilled Coders G.729, G.723, FS1016, CELP or parametric Models of MPEG-4 Audio-VM, are used mainly to co decode speech signals and are generally for Co dieren of higher quality music signals unsuitable because they are usually sampled at 8 kHz signals are designed so they only an audio bandwidth can encode a maximum of 4 kHz. However, they show in general speed operation and low Computational effort.
For audio coding of music signals, for example, to achieve HIFI quality or CD quality, is therefore at a scalable coder a speech with a Audio encoder combines the signals with higher sampling rate such. as can 48kHz encoding. Of course It is also possible, by the above-mentioned speech coder to replace another encoder, for example, by a music / audio coder according to the standards MPEG1, MPEG2 or MPEG4.
Such derailleur with a speech a higher-quality audio encoder used commonly the method of differential coding in the time domain. On Input signal, for example, a sampling rate of 48 has kHz, is on by means of a downsampling filter Herun suitable for speech sampling ter-scanned. Now the down-sampled signal is coded. The encoded signal can be directly a Bitstromfor matiereinrichtung be supplied in order to be transmitted. However, it contains only signals with a bandwidth of z. B. a maximum of 4 kHz. The coded signal is further again and decoded by an up-sampling filter up-from palpated. The signal now obtained has, however, because the downsampling filter only with payload a bandwidth of eg 4 kHz. It must also be noted that the spectral content of the upsampled encoded / decoded signal in the lower band to 4 kHz not exactly the first 4 kHz band of abgetaste 48kHz th input signal corresponds as encoder generally introduce coding errors.
As already mentioned, a scalable coder both a well-known speech and an audio encoder on the signals with higher sampling rates can handle. To signal components of the input signal over to whose frequencies are above 4 kHz wear, ei ne difference of the input signal at 8 kHz and the coding th / decoded upsampled output of the Speech for each discrete time sample educated. This difference can then by a known are quantized and coded audio coder, as for Professionals is known. It should be noted that the difference signal in the audio encoder, the signals may encode at higher sampling rates, is fed, in lower frequency range apart from coding errors of is speech coder very much smaller than the original. In the spectral range of the bandwidth above the upsampled coded / decoded output signal of the speech coder is located, corresponds to the difference signal in the essentially the true input signal with z. B. 48 kHz sampled.
In the first stage, ie the level of the speech, is therefore usually a coder with low sampling frequency used, since in general a very low bit rate of the encoded signal is sought. Currently working more Encoder, also said encoder, with bit rates from we Nigen kilobit (two to 8 kilobits or even higher). The same also allow a maximum sampling frequency of 8 kHz, since in any case no more audio bandwidth in this clotting gene bit rate is possible and the coding at low is sampling frequency with respect to the computational cost effective. The maximum possible audio bandwidth is 4 kHz and is limited in practice to approximately 3.5 kHz. Target now in the further stage, ie, in step with the audio encoder a bandwidth improvement can be achieved, it must know tere stage with a higher sampling work. to Adaptation of sampling are decimation and In terpolationsfilter used for downloading or upsampling.
For some time, it is known to further reduce the dataset called the. TNS technique in high quality use audio coding (J. Herre, JD Johnston, "Enhancing the Performance of Perceptual Audio Coders by Using Temporal Noise Shaping (TNS) ", 101st AES Convention, Los Angeles 1996, Preprint 4384). The TNS technique (TNS = Temporal Noise Shaping = temporal noise shaping) allowed generally speaking, by means of a predictive coding of Spectral shaping of the temporal fine structure of the Quantization noise. The TNS technique is based on a consistent application of the dualism between time and Frequency range. From the art it is known that the car correlation function of a time signal when in the Frequency range is transformed, the spectral power tight precisely indicates this time signal. The dual case to results when the autocorrelation function of the spectrum a signal is formed and in the time domain transfor grammed is. The transformed or back into the time domain transformed autocorrelation function is also known as Qua drat the Hilbert envelope of the time signal called. The Hilbert envelope of a signal is thus directly related to the associated autocorrelation function of its spectrum. The squared Hilbert envelope curve of a signal and the spectral the same power density thus represent dual aspects The time domain and in the frequency domain. When the Hilbert Envelope of a signal for each over Teilbandpaßsignal a range of frequencies remains constant, then it is also the autocorrelation between neighboring spectral values be constant. This means in fact that the series of Spectral is stationary versus frequency, wes semi predictive coding techniques are used efficiently can to represent this signal, under Ver use a common set of Prädiktionskoeffizien th.
To illustrate these facts, reference is made to <b>Fig.</b> 8A and <b>Fig.</b> referenced 8B. <b>Fig.</b> 8A shows a short off cut from a temporally strongly transient "Kastagnet th "signal a period of about 40 ms. This signal was divided into several Teilbandpaßsignale, each part bandpass signal has a bandwidth of 500 Hz. <b>Fig.</b> 8B shows now the Hilbert envelope for these bandpass signals with Center frequencies ranging from 1500 Hz to 4000 Hz. Out Clarity were all envelopes to their maximum amplitude normalized. Obviously, the shapes are all Teilhüllkurven very strongly correlated, which is why a common predictor within this frequency range can be used to encode the signal efficiently. Similar observations can be made for speech signals are experiencing the effect of the glottal excitation pulses over the entire frequency range due to the nature of the is human speech production mechanism available.
<b>Fig.</b> Thus 8B shows that the correlation of adjacent values similarly, for example at a frequency of 2000 Hz as for example, a frequency of 3000 Hz and 1000 Hz is.
An alternative way to understand the property of spectral Prädiktierbarkeit of transient signals can from in <b>Fig.</b> 7 Table shown are obtained. At top left in the table is a continuous-time Signal u (t) is shown, which has a sinusoidal course. The is the spectrum U (f) compared with this signal, which consists of a single Dirac pulse. The optimal Coding for this signal is the encoding of Spectral data or spectral values, since all the Time signal, only both the magnitude and the phase the Fourrierkoeffizienten needs to be transferred to the to fully reconstruct the time signal. A Codie ren of spectral data corresponding to the same time a Prädik tion in the time domain. A predictive coding would here must therefore take place in the time domain. The sinusoidal So time signal has a flat temporal envelope, which a maximum non-flat envelope in the frequency range equivalent.
Now the opposite case will be considered, in which the Time signal u (t) is a maximum transient signal in the form of a Dirac-pulse in the time domain. A Dirac-pulse in Time domain corresponds to a "flat" power spectrum, during the phase spectrum in accordance with the temporal position of the Pulse rotates. Obviously, this signal the above-mentioned traditional methods, such as, for example, the Transform coding or coding of spectral data or a linear predictive coding of the time domain . This signal can be data, a problem best and effective are encoded in the time domain, since only the temporal position and the performance of the Dirac pulse must be transmitted, what of the consistent application Dualism results in a predictive coding in Frequency range a suitable method for efficiently represents coding.
It is very important, not the predictive coding of Spectral coefficients over frequency with the well-known dual concept of prediction of spectral coefficients of confuse one block to the next, the already imple is mented and also in the above mentioned article (M. Bosi, K. Brandenburg, S. Quakenbush, L. Fielder, K. Akagiri, H. Fuchs, M. Dietz, J. Herre, G. Davidson, Yoshiaki Oikawa: "ISO / IEC MPEG-2 Advanced Audio Coding", 101st AES Con Convention, Los Angeles 1996, Preprint 4382) is described. In the prediction of spectral coefficients of a block to the next, which ent a prediction about the time speaking, the spectral resolution is increased, while a Prediction of spectral values over the frequency, the time reso- lution increases. A spectral coefficient at at for example 1000 Hz can therefore by Spektralkoeffizien th at, for example, 900 Hz in the same block or frame be determined.
The considerations presented so led to an effi obtain cient coding for transient signals. Predictive coding techniques can, taking into account the Duality between time and frequency domain substantially analogously to the known prediction of one Spek tralkoeffizienten for spectral same Frequency in the next block are treated. Since the spectral power spectral density and the squared Hilbert envelope a signal are dual to each other, is a reduction a residual signal energy or a prediction gain depends of a Flachheitsmaß the squared envelope of Signal as opposed to a spectral Flachheitsmaß when received conventional prediction method. The poten tielle coding gain increases with transienteren signals at.
Possible Prädiktionsschemen itself provides both the pre diction schematic closed loop, which also return wärtsprädiktion is called, and the prediction scheme open loop, also known forward prediction we then. When spectral prediction scheme with CLOSED sener loop (backward prediction) is the envelope of the Error flat. In other words, the Fehlersignalener energy evenly distributed over time.
In a forward prediction, as described in <b>Fig.</b> represented 9 is, however, occurs a temporal shaping of the Quan tomation introduced on noise. Too prädizierender Spectral coefficient x (f) is a summation point <b>600</b> to out. The same spectral coefficient is also a pre predictor <b>610</b> supplied, the output signal with a negative Sign also the summation point <b>600</b> is supplied. The input signal in a quantizer <b>620</b> is therefore the difference between the spectral value x (f) and by Prädik tion calculated spectral value x<sub>p</sub>(F). In the forward the total error energy in the decoded prediction remain the same spectral coefficient. The temporal shape the quantization error is, however, as time formed at the output of the decoder appear, since the pre diction was applied to the spectral coefficients, whereby the quantization noise in time did under the neuter signal is applied, and thus can be masked. In this way, problems of temporal Mas marking z. B. with transient signals or voice signals avoided.
This type of predictive coding of spectral values is therefore as the TNS or temporal noise shaping technology designated. is on To illustrate this technique<b>Fig.</b> referenced 10A. Top left in<b>Fig.</b> 10A is a located Time course of a strongly transient time signal. the time extending the section of a DCT spectrum top right in <b>Fig.</b> 10A juxtaposed. The lower left display from <b>Fig.</b> 10 shows the resulting frequency response of a TNS synthesis filter, which is calculated by the LPC operation was (LPC = Linear Prediction Coding). It should be noted, that the (normalized) frequency coordinates in this diagram the time coordinates due to the time-domain and frequency area duality match. Obviously, the LPC Calculating a "source model" of the input signal, since the frequency response of the LPC synthesis filter of the calculated Envelope of strongly transient time signal is similar. In<b>Fig.</b> 10A bottom right is a representation of the spectral residual values, that is the input signal of the quantizer <b>620</b> in <b>Fig.</b> 9, shown above the frequency. A comparison between the residual spectral values according to the prediction and the spotting tralwerten shows in direct time-frequency transformation, that the spectral residual values a significantly lower have energy than the original spectral values. at the example shown, corresponds to the reduction of the energy the residual spectral values a Gesamtprädiktionsgewinn of about 12 dB.
For the meaning of the left lower representation in <b>Fig.</b> 10A Note the following. In classic application of pre diction on time domain signals is the frequency response of Synthesis filter an approximation of the Betragssspektrums Input signal. The synthesis filter (re) generates certain measured the spectral shape of the signal from a residual signal with approximately "white" spectrum. When applying the pre diction on spectral signals, as in the TNS technique the case, the frequency response of the synthesis filter a Approximation of the envelope of the input filter. The frequency transition of the synthesis filter is not the Fourier transform of the impulse response, as it is in the classical case, but the inverse Fourier transform. The TNS synthesis filter (Right) speak generates the envelope curve of the signal from a residual signal with an approximately "white" (ie flat) envelope. So is the bottom left of Figure<b>Fig.</b> 10A thus modeled by the TNS synthesis filter Envelope of the input signal. This here is a loga rithmische representation of the envelope approximation of the overlying map shown Kastagnettensignals.
Subsequently a coding noise was in the spectral Residual values introduced, such that in each coding band with a width of eg 0.5 Bark a signal / Rau rule ratio of about 13 dB resulted. The from the Introduction of quantization error resulting signals in the time domain are in <b>Fig.</b> 10B. The left representation in <b>Fig.</b> 10B shows the error signal due to the Quantization noise when used TNS technique while In the right diagram the TNS technique from comparative purposes was not used. As expected the error signal is left diagram is not evenly distributed over the block, but in the area concentrated in the a high is signal component exists which this quantization is rushing optimal collapse. In the right case, however, is the quantization noise introduced uniformly in the block, that is, distributed over time, which results in that in the front the area where actually no or almost no signal is also noise is present to hear his , while in the region in which high signal components are present, a relatively small noise is present, by the marking options of the signal does not be fully utilized.
In the following, a simple, that is not scalable Barer, described audio encoder a TNS filter has.
An implementation of a TNS filter <b>804</b> in an encoder is in <b>Fig.</b> 11A. The same is between Ana lysefilterbank <b>802</b> and a quantizer <b>806</b> arranged. The time-discrete input signal is at the in <b>Fig.</b> 11A shown encoder in an audio input <b>800</b> fed while the quantized audio signal and quantized Spek tralwerte or quantized residual spectral values ei nem output <b>808</b> are output, which a redundancy Codie rer can be connected downstream. The input signal is therefore transformed into spectral values. Based on the calculation Neten spectral is a conventional linear prediction running bill, which, for example, by forming the Autocorrelation matrix of the spectral and USAGE tion of a Levinson-Durbin recursion takes place. <b>Fig.</b> 11B shows a detailed view of the TNS filter <b>804</b>, At a filter input <b>810</b> are the spectral values x (1). , ., x (i). , ., X (n) is fed. It may happen that single Lich a certain frequency range in transient signals has, while in turn a different frequency range more stationary nature. This fact is at the TNS-Fil ter <b>804</b> by an input switch <b>812</b> and by a output switch <b>814</b> taken into account, wherein the switch First, however, for a parallel-to-serial or serial- ensure to-parallel conversion of the data to be processed. Depending on whether a certain frequency range instatio is när and a certain coding gain by the TNS technique promises is only this spectral region TNS-processed, which is done in that the input switch <b>812</b> For example, at the spectral value x (i) star tet and z. B. to the spectral value x (i + 2) is running. Of the inner portion of the filter again comprises the forward prediction structure, that is, the predictor <b>610</b> and the Sum mationspunkt <b>600</b>,
The computation for determining the filter coefficients of the TNS filter or for determining the prediction coefficients is performed as follows. Making the Autokorre lationsmatrix and using the Levinson-Durbin Rekur sion is the highest allowable order of Rauschfor flow filter, z. B. 20, carried out. If the calculated Prediction gain exceeds a certain threshold, TNS processing is activated.
The order of the noise shaping filter used for the current block is then subsequently removing all the coefficients with a sufficiently small absolute value determined by the end of the coefficient arrays. To this Way are the orders of TNS filters usual , on the order of 4-12 for a speech signal.
If for a range of spectral values x (i) Example , a sufficiently high coding gain is determined, is the same process, and it is the output of the TNS Filter not the spectral value x (i) but the spectral Residual value X<sub>R</sub>(I) output. This has a much ge ringere amplitude than the original spectral value x (i), As seen from <b>Fig.</b> 10A is visible. The to the decoder transmitted page information obtained thus additionally to the usual side information, a flag that the Using TNS indicates and if necessary, Infor mation about the target frequency range and also about the TNS filter that was used for encoding. The Filterda th can be represented as quantized filter coefficients will.
In analogy to the coder with TNS filter is now on a Decoder received that an inverse TNS filter having.
In the decoder, in which <b>Fig.</b> 12A is outlined, is for each channel a TNS coding undone. Spectral Residual values X<sub>R</sub>(I) be in the inverse quantizer <b>216</b> re-quantized and in an inverse TNS filter <b>900</b> turned fed, guiding its construction in <b>Fig.</b> 12B is illustrated. The inverse TNS filter <b>900</b> provides as an output signal again Spectral values in a synthesis filter bank <b>218</b> in the Time domain are transformed. The TNS filter<b>900</b> includes turn an input switch <b>902</b> and an output switch <b>908</b>Which first again to parallel to serial Conversion or for serial-parallel conversion to the processing ended data serve. The input switch<b>902</b> considered also a target frequency range to possibly used zuzu only residual spectral values of an inverse TNS coding lead while not TNS-coded spectral values to a exit <b>910</b> be passed through unchanged. The inverse Prediction in turn comprises a predictor <b>906</b> as a summation point <b>904</b>, However, the same are in the lower retired connected, follow the TNS filter. A spektra ler residual value passes through the input switch <b>902</b> to the Summation point <b>904</b>At which the same with the output signal predictor <b>906</b> is summed. The predictor provides a Output an estimated spectral value x<sub>p</sub>(I). Of the Spectral value x (i) is the output switch to the off output gear of the inverse TNS filter. The TNS-related Side information is thus decoded in the decoder, wherein the side information includes a flag that the Using TNS indicating and, if necessary, infor mation regarding the target frequency range. additionally ent hold the side information further the Filterkoeffizien th of the prediction, the method of encoding a block or "frames" was used.
The TNS process can be so composed as follows grasp. An input signal is in a spectral Dar position by means of a high-resolution analysis filter bank transformed. Subsequently, a linear prediction in the running frequency range, between the Frequency moderately adjacent spectral values. This linear prediction can interact as a filtering process to filter the spectral are interpreted, which is performed in the spectral domain. Thus, the original spectral values by the Prediction error, ie by the residual spectral values, replaced. These residual spectral values are as übli che spectral values quantized and encoded to the decoder transmitted by decodes the values again and inversely are quantized. Before applying the inverse filter Bank (synthesis filter bank) is provided an encoder for the taken prediction inverse prediction, ie an addition the predicted signal with residue, made in which the inverse prediction filter to the transmitted pre diction error signal, ie the re-quantized spectral residual spectral values, is applied.
By applying this technique, it is possible, the time Liche envelope of the quantization noise of the A adjust output signal. This allows better off Use of the marker of the error signals with signals that a pronounced time fine structure or a pronounced have transient nature. In the case of transient signals avoids the TNS technique the so-called. "pre-echoes" in which the Quantization noise before the "stop" a sol chen signal appears.
In a scalable audio coder is as already It was mentioned in the first stage, an encoder with low he used sampling frequency as a very generally low bit rate of the encoded signal is desired. In the second stage will then preferably an audio coder which, although at higher bit rates encoded, but a significantly higher bandwidth needs and thus Audiosig signals can encode with much higher sound quality than the Speech. Usually, a frame to be encoded audio signal, which is present in a high sampling rate, first on a low sampling rate, for example by means of a Down sampling filter lowered. The sampling rate in the signal is then reduced in the encoder of the first stage fed, whereby the output of this encoder is written directly into the bitstream of the scalable to s audio encoder exits. This encoded signal nied engined bandwidth is decoded again and then by for example by means of an upsampling filter again brought to the high sampling rate, and then in the frequency domain transformed. transfor Also in the frequency range mized, the original at the input of the encoder input audio signal. There are now facing two audio signals, but the former with the coding errors of the encoder provided the first stage. These two signals in the Frequency range can then fed to a subtraction element be to obtain a signal which only the differentiator enz represents both signals. In a switching module, which also may be implemented as frequency-selective switch as is described later, it can be determined whether it is more favorable, the difference of the two input signals, or but the original transformed into the frequency domain to process audio signal directly. The output the switching module is in any case, for example, a be known quantizer / coder supplied, which, when works according to a MPEG standard, on the one hand a Quanti tion, taking into account a psychoacoustic Mo dells performs, and then secondly a Entro pie coding preferably using the Huffman-Co consolidation effected with the quantized spectral values. The Output of the quantizer and coder will next the output of the first stage in the encoder Bitstream written.
A disadvantage of the prior art is the fact that so far no encoding and decoding concept is known, the the combination of the temporal noise shaping technique (TNS) with a scalable stereo allowed. How to Use It already described, provides a scalable Stereoco decodes the possibility of at least a mono signal and a Stereo signal separately from each other to be able to decode, whereby great flexibility is achieved. An Implementa tion of the art of the temporal noise shaping (TNS) would addition to scalability, the data reduction or Compression without loss of quality in both the mono- push ahead even when stereo signal.
The object of the present invention is to provide a to provide encoding and decoding concept that next high flexibility and a high Datenmengenreduzie tion allowed.
This object is achieved by method of processing a temporal stereo signal according to claim 1 or 9, by Ver go for decoding using a Prädik tion of a frequency encoded audio bitstream of claim 15 or 18, by means for processing a time- huge stereo signal according to claim 22 or 24, and by Means for decoding a using a Prediction over the frequency encoded audio bit stream according Claim 25 or 27 dissolved.
A scalable stereo coder with TNS technique in accordance with a first embodiment of the present invention, processing tet completely in the frequency domain. This means that a Mono-channel formed in the frequency domain and using a psychoacoustic coder is encoded. This has the Advantage that the mono channel has a temporal noise molding can be applied. In order to mono channel with the To link the two stereo channels, but has the temporal noise shaping of the mono channel undo ge are making. To equal proportions between the Stereoka ducts and to obtain the decoded mono channel, the Mo must nokanal a temporal noise shaping using the Prediction coefficients of the left or right channel un be subjected, thus a difference between the left Channel and the mono channel or a difference between the right channels and the mono channel can be formed.
It should be noted that the scalable bit stream of stereo signals the two stereo channels L and R and the mono- or center channel M own prediction about the frequency, ie a TNS processing, subjected who the can. For this purpose, there are three options:
<ul><li>1. For each channel L, R M and a separate "is completeness ended "prediction performed. This gives for each Ka nal own prediction and also an opti painting prediction gain. The price is a but be more elaborate encoder or decoder, since on the one three full predictors are necessary and to the front of a combination of two channels by adding, Subtracting or comparing a more complex treatment of the signals has to be performed, ie the prediction a channel must be reversed and this Ka nal must then by the prediction of be "filtered" other channel, ie an "incomplete permanent "prediction undergo.</li><li>2. The counterpart to that single for all three channels used Lich a set of prediction becomes. Thus for example for the left channel L a "complete" prediction are carried out, the residual spectral values L 'and left prediction results. The right (R) and the center channel M would then an "incomplete" prediction undergo in the use the L-prediction coefficients an L-filtered right and an L-filtered lin ken channel to receive. However, this solution provides mei least a lower prediction gain, but leads to a substantial simplification of the encoder or De encoder, since only a complete predictor needed is and a very simple "treatment" in the form of a simple forwarding without inverse prediction or he Neute prediction is required under point 1, since for all channels, only one set of prediction exists.</li><li>3. A compromise between point 1 and point 2 is only perform two complete predictions, z. B. with a stereo channel L or R and Mono Channel M. In the treatment of the signals L and R and M or L, and R prior to their combination must then only the M-Prädik tion be undone and the obtained therefrom Signal with the L or R prediction "gefil are tert ". The other stereo channel is also only an incomplete prediction with Prädiktionskoef coefficients of a channel subjected. Although this does a slightly reduced profit, but results in a ver justifiable expense in the encoder or decoder.</li></ul>
In the embodiment of the present invention, in which a psycho-acoustic mono encoder is used, an ajar at point 3 solution. If a Center / selected page processing, is the right channel R be at least generally similar to the left channel. Then it is sufficient only to a complete prediction perform channel and the other channel with the ermittel th prediction filtering. differ contrast L and R strong, then it is preferred that Prädikti mentation coefficient of the dominant channel for filtering, to use that prediction, the other channel.
The stereo "worst case" is that the left and right channel on the one hand signally uncorrelated and on the other hand are equally dominant, ie approximately the same amount have energy. In this case, however, no middle / side- Coding. In addition, this case also prohibits a differen ence coding so anyway ge to Simulcastverarbeitung must be attacked.
An essential point of scalability is that not the mono and the stereo signal independently be transmitted, but that the stereo signal to co dieren is only the difference of the original Ste reosignals includes the mono signal. To adjust to Kings but NEN, which signal component is already encoded in mono signal must when comparing the mono signal on stereo channels same conditions prevail, such that a aussagefä hige difference can be formed.
Frequency selective switching devices are preferably used to determine frequency band as if it günsti ger is, as to be encoded stereo signal, the difference between tween the mono signal and one stereo channel or stereo channel to use themselves. Such a situation may in occur when the mono signal greatly from one stereo channel differs. Here, it is understood in the sense of As tenkompression cheaper, not the difference signal to neh men, but the stereo channel itself.
Furthermore, it is preferable also in the sense of pos ciation as high data compression, durchzu an MS-decision lead, ie frequency band, declare whether With te-side coding or a left-right coding günsti ger is.
The encoder according to the first embodiment of the prior invention thus is a scalable Stereoco coder with a psychoacoustic Monocodierer. The for Encoder of the first embodiment of the present Invention analog decoder essentially makes the case the coding steps carried out to reverse, wherein with respect to the temporal noise shaping safe again has found that having at each link of the mono channel a stereo channel present the same conditions, ie that only signals are compared, where identical prediction coefficients are assigned.
Preferably, the encoder according to the first embodiment of the present invention is a core codec he are panded to also adjacent to the mono-stereo scalability create a separate mono scalability. This means, that the corresponding encoder, a first mono sublayer and a second single sublayer and a stereo layer can multiplex a single bitstream. Of course However, all said layers corresponding to the Kon concept of scalability again be themselves in a basically arbitrary number be divided by sub-layers. Of the Core coder is preferably one of the above described encoder lower bit rate, and therefore the same on the input side a downsampling filter and output each having an upsampling filter to the data rate of original stereo signal at the data rate of the core codec adapt. Typically, the core codec is as Sprachco coder executed, the example only in the region of , 0 to 4 kHz encoded, wherein the psychoacoustic Mo nocodierer then ver the range of the signal above 4 kHz remains. In addition, the encoder of the second monolayer also the coding error of the core codec berücksichti gene, such that a mono signal with excellent quality from the mono signal with a low bit rate and mono signal can be assembled with a high bit rate. is Here an essential point is that two when comparing Signals, make sure always that the comparison underlying signals with similar and even better with same prediction coefficients have been processed to To make a meaningful difference. The analog to Decoder, as in the first case, makes the in Co dation is leading steps to reverse.
According to a second embodiment of the present invention includes an encoder only a mono-core codec and no psychoacoustic Monocodierer. Such Co coder delivers when the core codec as with speech low bit rate is performed in its bandwidth reduced mono signal and a stereo signal with full band width. This encoder is available in the applications be of advantage where no mono signal at full bandwidth nö is kind, and can be processed, for example, when the receiver decoder only mono signals with limited can handle bandwidth.
As with all scalable coding method, however, it is low when the bitstream also the high quality Ste reosignal exists with full bandwidth, if at for example thinking of a transfer to many decoder is, some of which only mono signals with limited can decode bandwidth, while other stereo signals can handle full bandwidth.
The analogous thereto decoder comprises analogously no psychoacoustic mono decoder but merely a Core decoder and corresponding TNS-functional units to the comparison between mono and stereo signals to recon construction of the stereo signal back to the same conditions have.
Preferred embodiments of the present invention are below with reference to the accompanying drawing calculations explained in more detail. Show it:
<b>Fig.</b> 1 a scalable stereo coder with TNS a Mo noschicht;
<b>Fig.</b> 2 a decoder for signals which by means of the Co coder of <b>Fig.</b> 1 have been encoded;
<b>Fig.</b> 3 a scalable stereo coder with TNS a he Mono first sublayer and a second mono part layer;
<b>Fig.</b> 4 a decoder for decoding by means of the in <b>Fig.</b> 3 shown coder coded signals;
<b>Fig.</b> 5 a scalable stereo coder with a TNS bandwidth-limited monolayer;
<b>Fig.</b> 6 a decoder for decoding by means of the in <b>Fig.</b> 5 shown coder coded signals;
<b>Fig.</b> 7 is a table illustrating the duality between the time and the frequency domain;
<b>Fig.</b> 8A is an example of a transient signal;
<b>Fig.</b> 8B Hilbert envelope of Teilbandpaßsignalen due of in <b>Fig.</b> 6A transient time signal;
<b>Fig.</b> 9 shows a schematic diagram of the prediction in the frequency Area;
<b>Fig.</b> 10A is an example to illustrate the TNS technique;
<b>Fig.</b> 10B is a comparison of the time course ei nes introduced quantization noise with (Left) and without (right) TNS technique;
<b>Fig.</b> 11A is a simplified block diagram of unskalier th encoder having a TNS filter;
<b>Fig.</b> 11B is a detailed view of the TNS filter of <b>Fig.</b> 11A;
<b>Fig.</b> 12A is a simplified block diagram of unskalier th decoder, the inverse TNS filter on includes; and
<b>Fig.</b> 12B is a detailed representation of the inverse TNS filter by <b>Fig.</b> 12A.
<b>Fig.</b> 1 shows a scalable stereo coder TNS, a Monolayer produced at full bandwidth, according to a he first embodiment of the present invention. It should Note, however, that it is by no means conclusive, that the psychoacoustic mono coder full bandwidth coded. The bandwidth can be smaller, which by zero set of spectral values above a certain frequency can be achieved. but usually the bandwidth the psychoacoustic mono coder is larger than that of the Core Coders.
As usual, will be temporal signals with lowercase draws while spectral signals or spectral values with Capital letters are marked. The encoder in <b>Fig.</b> 1 is illustrated schematically, includes a first entrance <b>10</b> for a first (left) stereo channel L and a second input for a second (right) stereo channel r. The time input signals l, r are using a Modified Discrete Cosine Transform (MDCT) <b>14</b>. <b>16</b> transformed into the frequency domain.
It should be noted that only preferably a modifi ed discrete cosine transform is used because is the same set in the newer MPEG standards. It However, it is obvious that any other possible liabilities, such. as filter banks or other transformations, can be used to transform a Time signal to accomplish in the frequency domain.
As from <b>Fig.</b> 1 can be seen, the left and the right channel equal processed substantially, in both channels is provided a TNS-block, ie a block TNS-L <b>18</b> for the left channel and a block TNS-R <b>20</b> for the right channel. The outputs of the TNS-blocks<b>18</b>. <b>20</b> are each in a frequency-selective switching means (FSS) fed with a frequency-selective switching input direction <b>22</b> is provided for the left channel, while a frequency-selective switching means <b>24</b> for the right Channel is used. The outputs of frequenzse selective switching means are, among other signals, which will be discussed later, in one block MS provi input flow, it is decided whether a left- Right stereo processing or a mid-side-Stereover processing is cheaper.
As from <b>Fig.</b> 1 can be seen, the MS operates Bestim flow completely in the frequency domain using conventional psycho acoustic stereo coder output side to the block MS analysis <b>26</b> are connected. Such encoders are in<b>Fig.</b> 1 no longer shown. However, the same are in the known art and need not be described further will. The same result, however, roughly a Quanti tion by, such that the introduced quantization noise under the masking threshold of the signal remains, said then with minimal Bitaufwand quantized Spek tralwerte typically using the Huffman Co dation are encoded to finally a bit stream condition, which is compressed to the maximum.
The following is one of the mono signal processing went. In the in<b>Fig.</b> 1 embodiment shown is a mono signal M formed in the frequency range in which the spectral first channel L and the spectral second channel R by means of an adder <b>28</b> are summed, the sum from L and R then by a multiplier <b>30</b> is multiplied by the factor of 0.5 to a mono signal to arise. The thus obtained mono signal M is in a Block TNS-M <b>32</b> a prediction over the frequency unterzo gen, after which the output of block TNS-M <b>32</b> a M-coder / decoder (codec) <b>34</b> is supplied. The block M codec <b>34</b> preferably comprises a psychoacoustic Co coder, for example after the AAC standard (AAC = Advanced Audio Coding), the received mono signal with a maximum full bandwidth coded to the same as a monolayer <b>36</b> admit of.
However, the mono signal in the monolayer <b>36</b> coded, To compare with the stereo signals, ie a produce scalability, which must in the monolayer <b>36</b> coded mono signal in the block M-codec <b>34</b> decoded again in order 'to obtain the encoded / decoded signal M'. Because the decoded signal earlier in the block TNS-M <b>32</b> a prediction has been subjected to the frequency, namely with prediction that at this pre diction were recovered and stored in side information were, it must be treated, that is, these prediction about the frequency must again by means of a block TNS<sup>-1</sup>-M back be done consistently. The output of block TNS<sup>-1</sup>-M is Thus the encoded / decoded mono signal without prediction processing, ie unfiltered, before.
As has already been mentioned several times, this signal is now are compared with the left and right channel. To must it by means of a block TNS-L / R <b>40</b> a prediction about frequency using the prediction coefficients for the left or right Ka signal are subjected, ie using the prediction coefficients in the block <b>18</b> (TNS-L) or block <b>20</b> (TNS-R) were obtained. The L / R-filtered coded / decoded mono signal, which is now on node <b>42</b> is applied, is now to both the first (left L) as well as compared to the second (right-R) stereo channel will. For this purpose it is using multiplier<b>44</b>a, <b>44</b>b with the Factor 2 multiplied and to the negative input of Addie insurer <b>46</b>a for the left branch or to a negative input ei nes adder <b>46</b>b applied for the right branch. At the exit the adder <b>46</b>a is thus the difference between the filtered left channel and twice the coded / de coded and L-filtered mono channel. is Similarly, at the output of the adder <b>46</b>b the difference between the ge filtered right channel and twice the R-filtered encoded / decoded to mono channel.
The frequency-selective switching means <b>22</b>. <b>24</b> determine Now, if it is convenient to process the difference further or the left and right channel itself. Preferably finds this decision instead of a frequency-selective, so that for each frequency range, for example, for each psy choakustische frequency group, it can be determined which Signal for the coding is more favorable.
To also perform a mid-side coding to Kings nen, is for each channel a further adder <b>48</b>a or <b>48</b>b provided, wherein by means of the adder <b>48</b>a and a wide ren Multipliziers <b>50</b>a, the multiplication by the Fak tor 0.5 performs, the mid-signal M is formed, the the sum of the left and right channels multiplied by the Factor 0.5 corresponds. By means of the adder<b>48</b>b is dage gen the side-signal S formed, that is the difference formed from the left channel or right channel, said Result is also multiplied by the factor 0.5. The side-signal, ie the output of the multiplying this insurer <b>50</b>b, is thus unchanged block MS-determination <b>26</b> supplied. The middle signal, ie the output of multiplier <b>50</b>a, however, is by means of a mid-Addie insurer <b>52</b> with the L / R-filtered coded / decoded mono Signal compared, ie only the difference between the middle signal and the coded / decoded Mo NOSIGNAL block MS-determination <b>26</b> supplied. The output signal of the middle-adder thus gives only at the the encoding / decoding in block M-codec <b>34</b> imported Error.
In the following, the functioning of the in <b>Fig.</b> 1 outlined encoder received. A temporal stereo signal having a first time (L) and a temporal second (r) channel, by means of MDCT Filterban ken <b>14</b>. <b>16</b> transformed into the frequency domain to a spectral first channel L and a spectral second Ka nal R to obtain. From the spectral first channel and the spectral second channel is by summer <b>28</b> and the multipliers <b>30</b> a spectral mono channel M formed, the a prediction over the frequency in the block M-TNS <b>32</b> U.N is subjected. The prediction coefficients obtained in this way for M prediction are in the page of information Bitstream at the output (not shown) of the encoder of <b>Fig.</b> 1 written. At the output of block TNS-M<b>32</b> is therefore a filtered mono channel M 'before.
Analogously, both the spectral first channel L and the spectral second channel R by means of a block TNS-L <b>18</b> or TNS-R <b>20</b> a prediction over the frequency subjected to a filtered first channel L 'or a filtered second channel R 'to obtain. The above in the prediction the frequency obtained by the spectral left channel pre be diction coefficients as well as in the Prädik tion it to the frequency with the spectral right channel preserved prediction also in the Seitenin formations of the bitstream written.
As has already been explained in detail at the outset, yields a prediction over the frequency both Prädiktionskoeffi coefficients that are written in the page information and represent a rough shape of the signal, as well as spectral residual values ( "residual spectrum"), at the output a TNS predictor rest. The original signal can then using the residual spectral values, ie the Output of a TNS-block, and the Prädiktionskoeffi cient to restore.
The scalable encoding or Deco invention exploding in several places a comparison, Example , in the form of a difference between spectral Residual values performed. This comparison of the spectral However, residual values brings only a maximum coding gain, if the corresponding to the spectral residual values Prediction are the same. therefore, when For example, a TNS-filtered mid-signal is present, ie consisting of spectral mid-residual values to spectral mid-prediction correspond, and if this TNS-filtered signal with a center-TNS filtered left signal to be compared, so are for TNS-filtered left signal left Prädiktionskoeffi coefficients and spectral left residual values. It would be from Co dier profit considerations make little sense, the spectral Links residual values to ver with the spectral mid-residual values same as the underlying links Prädiktionskoeffi cient or mid-prediction differently are. According to the invention therefore have similar possible Ver ratios are created. In this case, the Dif could reference to an FSS level greater than the original spectrum be, which is not the difference signal but the Origi nalspektrum would be chosen, which greatly ver the coding gain deteriorated.
This can either be done by the TNS-filtered Mid signal an inverse prediction is subjected. Now there is an unfiltered mid signal. To this ungefil shouldered mid signal to the left-prediction coefficients refer, ie to calculate spectral mid-residual values, the unfiltered to the left prediction give mid signal, a simple prediction with be already calculated in the example left Prädiktionskoeffizien th, are carried out. This L-filtered mid signal now comprises the residual spectral values, which together with the Left-prediction unfiltered mid signal would result. Now the residual spectral values can L-filtered center-signal with the residual spectral values comparing the TNS-filtered left signal, since both Spektralrestwerte the same Prädiktionskoeffi cient relate. Alternatively, however, it is also pos Lich, the TNS-filtered left signal an inverse TNS to undergo filtering to an unfiltered signal Links to obtain, and then this signal with a prediction of the to undergo mid-prediction, such that the spectral Links residual values as well as the spectral be mid-residual values in the mid-prediction attracted are.
For the reasons mentioned above must, therefore, the output signal of the M encoder / decoder of an inverse prediction by the TNS<sup>-1</sup>-M-Block <b>38</b> be subjected to a (Unfiltered) encoded / decoded mono channel to give.
By Übrerbrückungszweig <b>39</b> it is ensured that the inverse TNS filtering in block <b>38</b> not by a impaired Simulcast / difference switching the FSS 156 is, that is, the inverse TNS filtering is functioning properly.
This unfiltered encoded / decoded mono channel is now but in the frequency-selective switching means <b>22</b> or. <b>24</b> to the left and right channels, ie the spectral compared residual spectral values of the left and right channel will. To achieve this, the coded / decoded Mono channel for comparison with the TNS-filtered left Signal in the block <b>40</b> a TNS filtering with the Left Prediction that the block <b>18</b> were calculated and are in the side information, are subjected. Alternatively, the coded / decoded mono channel M '' for Comparison with the filtered second channel R 'in the fre -selective circuit means <b>24</b> also in the block <b>40</b> a prediction with the R-Prädiktionskoeffizien th, which in the block TNS-R <b>20</b> were determined and in the Page information are to be subjected. This (be acted) L / R-filtered mono channel M '' 'is connected to node <b>42</b> at. For reasons of clarity that is at the node<b>42</b> at Signal lying as L / R-filtered mono channel M '' 'designated net, which means that the mono-channel with either the L- is filtered or the R-prediction. It will preferably always the prediction of the channel with the greater overall energy use. However, it is pos Lich, from frame to frame of the prediction coefficients of a channel on the prediction of the walls toggle ren channel, a frame is known a Processing unit from z. B. 1024 time samples is.
It is not mandatory that two to be combined signals the exactly identical prediction are based. Thus, even spectral residual values on like Prediction coefficients are based are combined, without having to accept substantial Codiergewinneinbußen. Here, a compromise can be selected. If z. B. Fully constant predictions (<b>18</b>. <b>20</b>) For L and R performed wor the are, the consequential residual spectral can values without inverse prediction and renewed incomplete combined prediction of a channel. treatment the signals prior to their combination here comprises therefore the Che fen whether the prediction are similar enough what at similar L and R channels will be true, and the un Forward altered when the prediction coefficients are similar, or performing corresponding inverse Predictions and incomplete predictions when the Pre diction coefficients are not similar. The decision threshold can be a number of factors, such as, for. example, the Codierge winn, the signal strength or the reasonable efforts in Co coder or decoder, depend.
To simplify could for the prediction over the Fre frequency of the left and right channel, only one set of Prediction coefficients are used, ie the pre diction coefficient, at a TNS filtering of the lin ken channel were calculated. Then, the prediction would coefficients of the blocks <b>18</b>. <b>20</b> equal and therefore the signal at node <b>42</b>, Ie the L / R-filtered mono channel M '' ', in indeed include only one set of spectral residual values would, as it throughout the encoder in this case, only M-Prädik tion coefficients and, for example, L-Prädiktionskoeffi will cient.
The frequency-selective switching means <b>22</b>. <b>24</b> check if it is cheaper, the filtered first channel L 'and the filtered second channel R 'or the difference of gefil failed left channel L 'and L / R-filtered mono channel or the difference of the filtered right channel and the L / R-filtered mono channel to be processed further.
It is not always low, a difference processing use. The lead frequency-selective switching means therefore a so-called simulcast differential switching through. It is then unfavorable, a differential signal further verar BEITEN when the difference signal has a higher energy than the corresponding other signal at the input of frequenzselek tive switching device <b>22</b> or. <b>24</b> having. Since in principle Lich used as mono encoder any encoder may be, it may happen that the encoder certain through the stereo coder difficult to be encoded signal portions produced. If differential coding but not gun stig is because the energy content of the difference signal is greater than the energy content of the filtered first or second is the channel, is apart from a differential coding and switched to the simulcast operation.
Since the difference in the frequency domain, ie selective spectral value, takes place, it is readily pos Lich, a frequency-selective simulcast or Differenzco perform consolidation. The difference in the spectrum thus permits a simple frequency-selective choice of Frequency ranges which are to be differentially encoded. In principle, to a switch from a differential a simulcast coding for each spectral value individually occur. However, this would require a too large amount of Be require teninformationen. Therefore, it is preferred, in for example a frequency groups comparing the Ener gies of Differenzspektralwerte and the transformed perform left and right channel. Alternatively certain frequency bands from the outset set be such. B. 8 tapes to each 500 kHz in the example. On Compromise when determining the frequency bands is therein, the amount of side information to be transmitted, ie, whether active in a frequency band differential coding or not, be weighed against the benefits that from a frequent possible differential encoding grows.
<b>Fig.</b> 2 shows an outline illustration of a decoder, a by in <b>Fig.</b> 1 encoder shown coded to decode signal. The decoder of<b>Fig.</b> 2 to summarizes a bit stream input to which a scaled bit stream applied, ie a bitstream such as a Monosi signal and a stereo signal, wherein the mono signal un can be decoded depending on the stereo signal. The Bit power input <b>100</b> fitting bitstream BS is in a Demul plexer <b>102</b> fed, of the stereo layer from the Mo noschicht separated, and the additional information page extracted from the bitstream BS. In analogy to<b>Fig.</b> 1 be is the stereo layer behind the demultiplexer <b>102</b> out a preferably AAC encoded representation of a first and a second stereo signal, said first Stereosi signal into a first stereo decoder <b>104</b> is decoded, while the second stereo signal in a second stereo decoder coder <b>106</b> is decoded.
The two stereo decoder <b>104</b> and <b>106</b> are in <b>Fig.</b> 2 as designates L / M-requantizer or as an R / S-requantizer. This should make it clear that the stereo signal either Left-right or center-side coding may be. It is known that the left-right coding and the center-Sei te-coding not only varies from one block to the next can be, but also within a block frequency selectively. The setting, in which frequency range intra half of a block an MS-encoding is performed, by MS analysis <b>26</b> (<b>Fig.</b> 1) laid down that such a forms called MS-mask. If a left-right encoding in the received and de-multiplexed stereo layer bitstream ago is, are the stereo decoder <b>104</b> in analogy to <b>Fig.</b> 1 the first spectral stereo signal L '', while the second stereo coder <b>106</b> after decoding and Requantisie tion of a second spectral signal stereo signal R '' gives. If, however, is a mid / side coding, so are the stereo decoder <b>104</b> the first stereo signal, the signal M<sup>v</sup> , while the second stereo coder <b>106</b> as a second spectral signal Stereo outputs side signal S.
The by the demultiplexer <b>102</b> Monolayer is recovered hand in a mono requantizer <b>108</b> input to the coded mono signal to decode from the monolayer. In Analogously to the description of the blocks <b>104</b> and <b>106</b> will also be the block <b>108</b> called requantizer. Further above it was found that the M-Codec <b>34</b> wherein in <b>Fig. 22000 00070 552 00004 21881 001000280000000200012000285912188900040 0002019829284</b> 1 illustrated embodiment as psychoacoustic AAC Codec is performed. This means that the mono-Requanti niser <b>108</b> similar to the two Stereodecodierern <b>104</b> and <b>106</b> is constructed.
To be able to reconstruct the stereo signal again, must in the output of mono requantizer <b>108</b> still This M-TNS filtering be repealed. This ge happens in block TNS<sup>-1</sup>-M <b>110</b>, On output of block TNS<sup>-1</sup>-M <b>110</b> Thus is the encoded / decoded (ungefilter te) mono channel M '' to. This signal may by means of a blocks <b>111</b> be transformed into the time domain, as decoded mono channel are output and a Recipients are processed, the only for themselves a mono signal is interested. In analogy to<b>Fig.</b> 1 has the encoded / decoded mono channel M '' a L / R filtering under place to enable the residual spectral values of the mono channel to the same prediction as the residual spectral are values of the left and right channel based. Just Then, differences or sums useful formed who the, that only then is a combination or one meaningful Comparison possible. This is done in the block TNS-R / L<b>112</b>, The output of block TNS-R / L is thus the L / R-filtered Mono channel M '' 'to. The notation L / R or R / L to a optional use of R-prediction or L-prediction point. The L / R-filtered Mono channel is now a summer <b>114</b> supplied to the The case of a mid / side coding for the first stereo signal M<sup>v</sup> to be added. The result is then the "true" Mid signal, the relative <b>Fig.</b> 1, the signal at the output of multiplier <b>50</b>a.
The in <b>Fig.</b> 2 decoder shown further comprises two inverse frequency-selective switching means <b>116</b>. <b>118</b>, Wherein the inverse frequency-selective switching means <b>116</b> for United processing of the left, ie the first channel L provided , while the inverse frequency-selective switching means <b>118</b> for the processing of the second or right channel R serves. The inverse frequency-selective switching means<b>116</b> and <b>118</b> is each a totalizer <b>120</b> or. <b>122</b> pre on, such that an inverse frequency selective switch device as input both a spectral stereo signal L '', R '' and the sum of the spectral Stereosi gnals L '', R '' and by a multiplier <b>124</b> ver double "true" center-signal (corresponding to the output signal of the multiplier <b>50</b>a in <b>Fig.</b> 1) receives. In the verses frequency-selective switching means <b>116</b>. <b>118</b> who to by corresponding page information <b>126</b> controlled to the present in the coding ratios, ie Differential or Simulcastcodierung in a frequency band, replicate.
The inverse frequency-selective switching means <b>116</b> and <b>118</b> give when the page information <b>126</b> kor directly be driven, a (decoded) filtered first channel L 'and a (decoded) filtered second Channel R 'from. In a block MS<sup>-1</sup><b>128</b> is the middle / side- Coding reversed that by block MS-Be humor <b>26</b> (<b>Fig.</b> 1) was introduced. This means that in the presence of a left-right-coding the Eingangssi are signals L ', R' passed through unchanged, while at Existence of a mid-side coding by simple ad dition and subtraction from the mid signal and the side signal S of the (decoded) filtered first channel L 'and the (Decoded) filtered second channel R 'are calculated. To undo the TNS filtering the filtered first channel of an inverse TNS filtering with the block TNS<sup>-1</sup>-L <b>130</b> subjected. Analogously, the right channel an inverse prediction over the frequency subjected to the by block TNS<sup>-1</sup>-R <b>132</b> in <b>Fig.</b> 2 schematically Darge provides is. At this point it should be noted that the filtered first channel L 'as well as the filtered second Channel R 'residual spectral values of the first channel L and the second channel R, which together with the first corre relevant TNS prediction spectral first Channel L and the spectral second channel R result. The TNS prediction coefficients for the first channel and for L the second channel R are as described in <b>Fig.</b> 2 by the Be teninformationenleitungen <b>126</b> is shown, from the Be extracted teninformationen and TNS<sup>-1</sup>-Blöcken <b>130</b> and <b>132</b> supplied.
Finally, the time the first channel l and since to obtain union second channel r, have the spectral Channels by means of an inverse filter bank to the time domain be transformed as indicated by the blocks MDCT<sup>-1</sup>-L <b>134</b> and MDCT<sup>-1</sup>-R <b>136</b> is illustrated in block diagram form.
As has been repeatedly stated, the encoder according to a first embodiment of the present invention, the in <b>Fig.</b> 1 is shown, a scalable TNS stereo coder with a monolayer, wherein the monolayer layer preferably as well as the stereo layer with maxi times full bandwidth is encoded as the M codec <b>34</b> as psychoacoustic AAC encoder is performed. Therefore there the mono-requantizer <b>108</b> the decoder in <b>Fig.</b> 2 a Mono channel full bandwidth. The scalability be is where in <b>Fig.</b> 1 shown encoder and the analog in <b>Fig.</b> 2 decoder shown therein, for decoding under Koen select a stereo layer and a monolayer NEN.
In the following, the in <b>Fig.</b> 3 encoder shown be registered, the service is a scalable stereo coder TNS, wherein the monolayer of a first mono-layer part and consisting of a second mono-layer part. This encoder is not only with respect to stereo / mono scalable, son countries here is also the monolayer in a first mono-part layer and scaled in a second mono-sublayer. sliding like elements in the <b>Fig.</b> 1 and 3 are in <b>Fig.</b> 3 by corresponding reference numerals. As far as the Operation of these elements are not of the associated With <b>Fig.</b> 1 differs described, is this Ele elements no longer received.
In contrast to the in <b>Fig.</b> 1 shown coder according to the comprises first embodiment of the present invention the in <b>Fig.</b> 3 shown a so-called core codec encoder <b>140</b>Which is usually a coder with low Bitra te, z. B. a CELP speech coding system. The core codec<b>140</b> providing a first mono-layer part, wherein these mono- Sublayer usually a bandwidth of only 0 to 4 kHz will have. The core codec as input signal a temporal mono channel m, which is formed by both the time left channel l also the time right channel r by means of a multiplier <b>142</b>a or. <b>142</b>b be halved, whereupon the time halved left channel and the right channel by time halved an adder <b>144</b> are added to the time Mono channel m to receive.
The temporal mono channel m is not like the time Liche left channel L and the right channel r time with the Stereo sampling before. To the bit rate of the first mono-part layer compared to the bit rate of stereo layer to redu adorn, the temporal mono channel m is by means of a Down sampling filter <b>144</b> filtered. The output of the down- sampling filter <b>144</b> by means of the core codec <b>140</b> in front existing core encoder encodes and the first mono part layer <b>146</b> into a bit stream multiplexer (not shown) output. To put that in first mono sublayer already co be-coded information at the secondary coding to take into account is that encoded the core coder Signal within the core codec <b>140</b> again decoded and filtered by means of an up-sampling filter, such that the Output of the up-sampling filter <b>148</b> same Abtastra tenverhältnisse having as the temporal first channel l and the time second channel r.
The output of the up-sampling filter <b>148</b> it will then by means of a MDCT filter bank <b>150</b> into the frequency domain transformed to an encoded / decoded spectral Mono channel M<sub>CD</sub> to obtain. This coded / decoded spectral tral mono channel is now a TNS filtering within a Blocks TNS-M <b>152</b> subjected. Here, either a full permanent new Prädiktionskoeffizientenberechnung Runaway will lead, or it may already in the Seiteninfor mation existing prediction that by the TNS-M filtering in block <b>32</b> were obtained hergenommen will. In any event, for the prediction over the Frequency with the encoded / decoded spectral mono channel M<sub>CD</sub> and the spectral mono channel M behind the multiplier <b>30</b> the same prediction coefficients are used so that the output signals of the blocks <b>32</b> and <b>152</b>, That is, the residual spectral values can be compared.
This comparison takes place by means of an adder <b>154</b> and a frequency-selective switching means <b>156</b> instead of. At the Output of the adder <b>154</b> thus is the "rest" of the Mono channel to which the up to the maximum bandwidth frequency Core codec <b>140</b> only the by the core codec <b>140</b> coding error introduced includes, and the excess of the maximum Bandwidth of core codec <b>140</b> the full mono signal comprises. The frequency-selective switching means <b>156</b> surely again whether it is more favorable, a difference coding or a employ simulcast coding or processing. On Off gang of frequency-selective switching means <b>156</b> is so with a comparative mono channel M<sub>CD</sub>'' Before, by Verglei chen of the filtered encoded / decoded spectral Mono channel M<sub>CD</sub>'And the filtered mono channel M' WUR receive de. In analogy to<b>Fig.</b> 1 is the comparison mono channel M<sub>CD</sub>'' In the M-Codec <b>36</b> fed and an inverse TNS Filtering with the M-prediction <b>38</b> subjected to obtain a coded / decoded mono channel.
If <b>Fig.</b> 1 with <b>Fig.</b> 3 is compared, so remains festzu provide that the coded / decoded mono channel M '' in <b>Fig.</b> 1 and in <b>Fig.</b> 3 above the core codec bandwidth identical are, while those signals below the core codec different bandwidth frequency is that of coding te / decoded mono channel M '' of <b>Fig.</b> 3 only still the the core codec <b>140</b> introduced coding errors comprises while the coded / decoded mono channel M '' of <b>Fig.</b> 1 the entire Mono signal includes. In certain cases it may, however, be that formed by the core codec <b>140</b> introduced coding error is already greater than the mono signal, wherein in the case, the frequency-selective switching means <b>156</b> no will choose difference processing, but a simulcast Processing.
<b>Fig.</b> 4 shows the to <b>Fig.</b> 3 analog decoder. Compared to the in <b>Fig.</b> 2 shown in the decoder comprises <b>Fig.</b> 4 Decoder shown, a stereo and two mono layer can decode some layers, in addition a core decoder <b>160</b>, A MDCT filter bank <b>162</b>, One block TNS-M <b>164</b>, a adder <b>166</b> and an inverse frequency-selective switching facility <b>168</b>, In addition, the core decoder<b>160</b> on Upsamling filter <b>170</b> downstream.
In the <b>Fig.</b> additional decoder elements 4 shown are explained below. The demultiplexer<b>102</b> separated the stereo layer and monolayer and leads in particular a separation of the first monolayer and the second sub-layer Mo noteilschicht by. The output of mono Requanti sierers <b>108</b> is now the second decoded mono sublayer, during the first one-part layer in the core decoder <b>160</b> is fed, which is identical to the core decoder in Core codec <b>140</b> works. The output of the core Deco DERS is in the up-sampling filter <b>170</b> entered to slide che Abtastfrequenzverhältnisse between he decoded Mono first sublayer and the second decoded mono part layer manufacture.
Thus, there are two optional possibilities for output a mono signal. The first single sublayer, as shown in<b>Fig.</b> 4, output from the core decoder. This signal then has a sampling frequency corresponding to the Core codec. Alternatively or simultaneously, the signal can at Output of the up-sampling filter <b>170</b> as Core-time signal be used. This mono signal corresponding to the first Monolayer, but with the difference that its sample frequency of the left and right stereo channel in front of the Encode equivalent.
The of the upsampling filter <b>170</b> filtered signal is by the MDCT filter bank <b>162</b> in the frequency range trans formed to turn the encoded / decoded spectral Mono channel M<sub>CD</sub> (please refer <b>Fig.</b> 3) to obtain. This signal is in the block <b>164</b> TNS-filtered, the TNS Filterkoeffi cient of the page information <b>126</b> be used, for example, by the TNS predictor <b>152</b> or <b>32</b> from <b>Fig.</b> 3 were determined in the encoder. At the output of the block <b>164</b> then is the filtered encoded / decoded spectral Mono channel M '<sub>CD</sub> in that in the adder <b>166</b> as well as the decoded second mono sublayer is entered. The Ad coder <b>166</b> in turn feeds the inverse frequency-selective switching means <b>168</b>That in analogy to the inverse fre -selective switching means <b>116</b> and <b>118</b> depending on the side information is controlled to the encoder in imported frequency wise selections again to undo do. At the output of the inverse frequency-selective switching facility <b>168</b> then is the filtered mono channel M 'on, by the inverse predictor TNS<sup>-1</sup>-M <b>110</b> an inverse Prediction is subjected to the frequency to the coding th / decoded mono channel M '' to obtain. The more Ver processing is for the in <b>Fig.</b> 2 processing described identical.
<b>Fig.</b> 5 shows an encoder according to a second embodiment of the present invention, which encoder a scalable stereo coder TNS is the monolayer only the output of the core codec <b>140</b> has, ie the no AAC Monocodierer <b>34</b> includes. The time Monoka channel m is a filtering in the downsampling filter <b>144</b> under and then subjected to the core codec <b>140</b> encoded to a Mono layer to give. The monolayer is then within the Core codec <b>140</b> decoded again and by upsampling filter <b>148</b> filtered and then purified by the filter bank <b>150</b> in converted to frequency domain to the coded / decoded spectral mono channel M<sub>CD</sub> to obtain.
In contrast to the in <b>Fig.</b> 3 shown Ausführungsbei However, play will be no "independent" prediction about the frequency of the coded / decoded mono spectral channel M<sub>CD</sub> or a prediction about the frequency with "M-Prädik tion coefficient "performed, but already a pre diction over the frequency by means of L or R Prädiktionsko efficient, in blocks <b>18</b> or. <b>20</b> were calculated. These L / R prediction is by a block TNS-L / R <b>172</b> sym metabolised. This means that once the TNS-L / R Prädik tion coefficient "gone" is, and that no M-Prädik tion is performed. Therefore, also found in<b>Fig.</b> 5 held TNS-M prediction <b>32</b> (<b>Fig.</b> 3) a TNS-L / R prediction instead, as indicated by block <b>174</b> is indicated. At the exit the TNS-L / R block <b>172</b> Thus is the L / R-filtered Co ied / decoded mono channel M<sub>CD</sub>'' ', While at the output the TNS-L / R block <b>174</b> the L / R-filtered mono channel is present. The signal M '' 'and the signal M' ''<sub>CD</sub> are both at L or based R-prediction and can therefore by the adder <b>154</b> are compared, such that the fre -selective switching means <b>156</b> a differential mode can choose or a simulcast operation. As in to connection with <b>Fig.</b> 3 was discussed, the core codec has egg ne maximum bandwidth that much clotting generally ger than the full stereo bandwidth. Therefore, the Off output signal of the frequency-selective switching means <b>156</b>, ie the L / R comparison mono channel M<sub>CD</sub><sup>iv</sup>, Up to the maximum Core Coderfrequenz generally the encoding / Decodie comprise approximately error of core codec, and the maxi painting core coder-frequency full mono channel. The further substantial processing substantially corresponding to the in to connection with the <b>Fig.</b> 1 and 3 described Vorgehenswei sen.
<b>Fig.</b> 6 shows the to <b>Fig.</b> 5 analog decoder. Compared to <b>Fig.</b> 4 comprises <b>Fig.</b> 6 no mono requantizer <b>108</b>, there the in <b>Fig.</b> 5 encoder shown also no M-codec <b>34</b> on pointed. The monolayer in the in<b>Fig.</b> 6 shown Deco coder the output of the core coder corresponds, is in an analog core decoder <b>160</b> again decoded and by an up-sampling filter <b>170</b> filtered same Abtastfrequenzverhältnisse of mono and the stereo signal to obtain. The output of the up-sampling filter<b>170</b> will be by means of a MDCT filter bank <b>162</b> in the frequency area transformed to the encoded / decoded spektra len mono channel M<sub>CD</sub> to obtain. In contrast to<b>Fig.</b> 5 is in <b>Fig.</b> 6, however, no prediction over the frequency by means of M-prediction coefficients performed, but a pre diction over the frequency using the R or the L-prediction coefficients in the side information <b>126</b> stored. This fact is by block TNS-R / L <b>178</b> in <b>Fig.</b> 6 shown schematically. At the exit therefore the block TNS-R / L is the L / R-coding filtered te / decoded mono channel M<sub>CD</sub>'' 'Which on the one hand in a adder <b>180</b> is fed and on the other in a Mul tiplizierer <b>182</b>To discuss the adder <b>122</b> and <b>120</b> with the first spectral stereo signal L '' or to the second to be spectral stereo signal R '' compared. The second Input of the adder <b>180</b> is the first spectral Stereo signal M<sup>v</sup> applied to the middle signal in this '' 'To make the case the L / R-filtered mono channel M, when a mid-side coding was present. The output the adder <b>180</b>Which, like the first spectral Ste reosignal M<sup>v</sup> in another inverse frequency selective switching means <b>182</b> is fed corresponds, as it in connection with <b>Fig.</b> was shown 1, the Output of the multiplier <b>50</b>a, that is the L / R-gefil failed complete mono channel. Further processing in the encoder from <b>Fig.</b> 6 is again analogous to the processing in the decoders of <b>Fig.</b> 2 and 4. FIG.
In summary, therefore, it can be said that the encoder according to the present invention, at least a monolayer and have a stereo layer, the monolayer to addition may be scaled, in the form of a first Mono sublayer low bandwidth and in the form of a second monolayer in AAC quality. Skilled artisans will however obvious that the stereo layer further can be scaled, for example, a Bandbreitenco to achieve consolidation of up to 12 kHz, which is approximately the hi-fi Quality corresponds to and beyond a bandwidth coding up to 20 kHz in the other stereo scaling To achieve layer, which is about a compact disc (CD) Quality meets.
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
Every citation, both waysCites: the store holds 4 of 5
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8311815B2 | Cited by | United States of America | Applicant |
| EP1484841A4 | Cited by | European Patent Office (EPO) | Search report |
| EP1484841A1 | Cited by | European Patent Office (EPO) | Search report |
| WO2005083678A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US7340391B2 | Cited by | United States of America | Applicant |
| CN113066472A | Cited by | China | Search report |
| US7283957B2 | Cited by | United States of America | Applicant |
| WO0223528A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| NO339114B1 | Cited by | Norway | Search report |
| EP0785631A2 | Cites | European Patent Office (EPO) | Search report |
| US5481614A | Cites | United States of America | Search report |
| DE69018989T2 | Cites | Germany | Search report |
| EP785631A2 | Cites | European Patent Office (EPO) | Search report |
| J. Herre, J.D. Johnston:"Enhancing the Performanceof Perceptual Audio Coders by Using Temporal NoiseShaping (TNS)", In: 101 st AES Convention, Los Angeles 1996, Preprint 4384 | Non-patent | – | Search report |
| Rohrecker, L.:"Ein Tonsignalcodierer Hoher Quali- tät mit einer Datenrate von 2*64kBit/s durch ein adaptives 4-subbandverfahren", In: Rundfunktechni-sche Mitteilungen, 1989, Jg.33, H.4, S. 145-148 | Non-patent | – | Search report |
| Brandenburg K., Grill B.:"First Ideas on Scalable Audio Coding", In: 9th AES-Convention, San Fran- cisco 1995, Vorabdruck 3924, S. 1-6 | Non-patent | – | Search report |
| J. Herre, J.D. Johnston:"Enhancing the Performanceof Perceptual Audio Coders by Using Temporal NoiseShaping (TNS)", In: 101 st AES Convention, Los Angeles 1996, Preprint 4384 | Non-patent | – | Search report |
| Brandenburg K., Grill B.:"First Ideas on Scalable Audio Coding", In: 9th AES-Convention, San Fran- cisco 1995, Vorabdruck 3924, S. 1-6 | Non-patent | – | Search report |
| Rohrecker, L.:"Ein Tonsignalcodierer Hoher Quali- tät mit einer Datenrate von 2*64kBit/s durch ein adaptives 4-subbandverfahren", In: Rundfunktechni-sche Mitteilungen, 1989, Jg.33, H.4, S. 145-148 | Non-patent | – | Search report |
7 priority claims, no other members on record
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 19821943 | Germany | A | |
| 19821943 | Germany | A | |
| 19821943 | Germany | – | |
| 19829284 | Germany | A | |
| 198219431 | – | – | – |
| DE1998121943 | – | – | – |
| DE1998129284 | – | – | – |
4 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Expiry of rightR071 | R071 | |
| No opposition during term of oppositionOpposition8364 | 8364 | |
| Grant after examinationD2 | D2 | |
| Request for examination as to paragraph 44 patent lawOP8 | OP8 |
Numbers
- Publication
- 19829284
- Publication, DOCDB
- 19829284
- Publication, EPODOC
- DE19829284
- Application
- 19829284
- Application, DOCDB
- 19829284
- Application, EPODOC
- DE1998129284
Titles2
- German
- Verfahren und Vorrichtung zum Verarbeiten eines zeitlichen Stereosignals und Verfahren und Vorrichtung zum Decodieren eines unter Verwendung einer Prädiktion über der Frequenz codierten Audiobitstroms
- English
- Temporal stereo signal processing method for forming scaled bit stream
Classification
- IPC, 1
- H04S1 00