Audio decoding
Abstract
An audio decoder comprises a receiver (801) for receiving input data comprising an N-channel signal corresponding to a down-mixed signal of an M-channel audio signal, M>N, having complex valued subband encoding matrices applied in frequency subbands and parametric multi-channel data. A subband filter bank (805) generates real- valued frequency subbands for the N-channel signal. A matrix processor (809) determines real- valued subband decoding matrices for compensating the application of the encoding matrices in response to the parametric multi-channel data. A compensation processor (807) generates down-mix data corresponding to the down-mixed signal by a matrix multiplication of the real-valued subband decoding matrices and data of the N-channel signal in the at least some real- valued frequency subbands. The down-mix data can be used to regenerate the down-mixed signal and the M-channel audio signal. The decoder may compensate for MPEG Matrix Surround Compatibility operations performed at the encoder using real- valued frequency subbands.
Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
1 claim: 1 independent, 0 dependent
- 1Claims Zastrzeżenia patentowe 1. Audio decoder (715) containing:1. Dekoder (715) audio zawierający: - środki (801) do odbierania danych wejściowych zawierających sygnał N-kanałowy odpowiadający sygnałowi downmixu M-kanałowego sygnału audio, M>N, mające macierze kodowania podpasma o wartościach zespolonych zastosowane w podpasmach częstotliwości i parametryczne dane wielokanałowe powiązane z sygnałem downmixu;i znamienny tym, że dodatkowo zawiera: means (801) for receiving input data including an N-channel signal corresponding to the downmix signal of the M-channel audio signal, M> N, having subband bandwidth matrices used in the frequency subbands and parametric multi-channel data associated with the downmix signal;and characterized in that it additionally contains: - means (805) for generating frequency subbands for a N-channel signal, wherein at least some of the frequency subbands are real frequency subbands;- środki (805) do generowania podpasm częstotliwości dla sygnału Nkanałowego, przy czym co najmniej niektóre z podpasm częstotliwości są podpasmami częstotliwości o wartościach rzeczywistych;- determining means (809) for determining the real-time decoding matrix for compensating the use of the coding matrix in response to parametric multi-channel data;and - środki (809) wyznaczania do wyznaczania macierzy dekodowania podpasma o wartościach rzeczywistych dla kompensacji stosowania macierzy kodowania w odpowiedzi na parametryczne dane wielokanałowe;oraz - means (807) for generating downmix data corresponding to the downmix signal by matrix multiplication of the real-time subband decoding matrix and the N-channel signal data in at least some of the real frequency subbands. - środki (807) do generowania danych downmixowanych odpowiadających sygnałowi downmixu poprzez mnożenie macierzowe macierzy dekodowania podpasma o wartościach rzeczywistych i danych sygnału N-kanałowego w co najmniej niektórych podpasmach częstotliwości o wartościach rzeczywistych. 2. Audio decoder (715) according to claim 1;The method of claim 1, wherein the determining means (809) is adapted to determine the inverse subband matrix with complex values of the coding matrix and for determining the decoding matrix in response to the inverse matrix. 2. Dekoder (715) audio według zastrz. 1, w którym środki (809) wyznaczania są przystosowane do wyznaczania odwrotnych macierzy podpasma o wartościach zespolonych macierzy kodowania i do wyznaczania macierzy dekodowania w odpowiedzi na macierze odwrotne. 3. Audio decoder (715) according to claim 1. The method of claim 2, wherein the determining means (809) is adapted to determine each matrix coefficient having the real values of the decoding matrix in response to the absolute value of the respective matrix coefficients, the inverse matrix. 3. Dekoder (715) audio według zastrz. 2, w którym środki (809) wyznaczania są przystosowane do wyznaczania każdego współczynnika macierzy o wartościach rzeczywistych macierzy dekodowania w odpowiedzi na wartość absolutną odpowiednich współczynników macierzy, macierzy odwrotnych. 4. Audio decoder (715) according to claim 1. The method of claim 3, wherein the determining means (809) is adapted to determine each real-valued matrix coefficient substantially as the absolute value of the respective matrix coefficient, the inverse array. 4. Dekoder (715) audio według zastrz. 3, w którym środki (809) wyznaczania są przystosowane do wyznaczania każdego współczynnika macierzy o wartościach rzeczywistych zasadniczo jako wartości absolutnej odpowiedniego współczynnika macierzy, macierzy odwrotnych. 5. Audio decoder (715) according to claim 1;The method of claim 1, wherein the determination means (809) is adapted to determine the decoding matrix in response to the subpacket transfer matrices being a multiplication of the respective decoding matrix and coding matrix. 5. Dekoder (715) audio według zastrz. 1, w którym środki (809) wyznaczania są przystosowane do wyznaczania macierzy dekodowania w odpowiedzi na macierze transferu podpasma będące mnożeniem odpowiednich macierzy dekodowania i macierzy kodowania. 6. Audio decoder (715) according to claim 1;The method according to claim 5, wherein the determining means (809) are adapted kd yznanznnain mnzikznz kkOdedynain in dkpdyikkni an minry mdkdUg jdkzaid mnzikznz trnasferg. 6. Dekoder (715) audio według zastrz. 5, w którym środki (809) wyznaczania są przystosowane kd yznanznnain mnzikznz kkOdedynain w dkpdyikkni an minry mdkdUg jdkzaid mnzikznz trnasferg. 1. Audio decoder (715) in the order of zz ^ z. 5, in which the matrixes transtenj kondego pokpnsmn are prick knock 1. Dekodek (715) audio wekług zzu^z. 5, w którym macierze transtenj kondego pokpnsmn są knak prnkn gknik G jkst mnzikrną kkOokownain pokpnsmn, n H jkst mnzikrną Ookownain pokpnsmn, n rrokOi wynanznnain are prnystosownak at the choice of co-mating gknik G jkst mnzikrną kkOokownain pokpnsmn, n H jkst mnzikrną Ookownain pokpnsmn, n rrokOi wynanznnain są prnystosownak ko wyboru współznyaaiOów mnzikrny g.11 £ 5l2 .8) 21 §22 _ in a way that the mines of the monks and the measses will spit Orazkrigm. g.11 £5l2 .8)21 §22 _ w tnOi sposób, żk minry mozy pn i mozy spkłainją Orytkrigm. 0. DkOokkr (155) ngkio increment nnstrn. 1, in the Otterym minrn mokdłd jkst wyanznnan w okpowikkni an 0. DkOokkr (155) ngkio wkkłgg nnstrn. 1, w Otórym minrn mokdłd jkst wynanznnan w okpowikkni an 9. Audio decoder (755) according to claim 1. The method of claim 7, wherein the means (809) determines that there is a prismatic coefficient that is co-ordinate, and that the wetlands of pn and p22 are nnsnkaizno equal to jkkaorzi. 9. Dekoder (755) audio według zastrz. 7, w którym środki (809) wyznaczanńa kokntOowo są prnystosownak ko wyborg współznyaaiOów mnzikrny prny więnnzh, żk mokgły pn i p22 są nnsnkaizno rówak jkkaorzi. 59. Audio decoder by zzc ^ z. 1 in which syydownmixu and paaymerroyzne knak wikokocowk are ngokak nk stnaknrkkm MPEG Sgrrogak. 59. Dekodek audio według zzc^z. 1, w którym syydownmixu i paaymerroyzne knak wikloOcacłowk są ngokak nk stnaknrkkm MPEG Sgrrogak. 55. Audio decoder (715) based on 1, in which the decoding of jess mnemonic Ookownain MPEG Mntrix Sgrrogak Compntibility, n srchy sounded N-On-line Jkst SyglMM MPEG Mntrix Sgrrogak Compntiblk. 55. Dekoder (715) audio wedkug zasUz. 1, w którym dekodowania jess mnzikrną Ookownain MPEG Mntrix Sgrrogak Compntibility, n pikrwsny syganł N-Onanłowy jkst syganłkm MPEG Mntrix Sgrrogak Compntiblk. 52. Sposób kkOokownain ngkio, prny znym sposób obkjmgjk: 52. Way kkOokownain ngkio, empty way obkjmgjk: - okbikrnaik (5595) knayzh wkjrziowyzh nnwikrnjązyzh syganł NOnanłowy okpowinknjązy syganłowi kowamixg M-Onanłowkgo syganłg ngkio, M> N, mnjązyzh mnzikrnk Ookownain pokpnsmn of wnrtorzinzh nkspoloayzh nnstosownak in pokpnsmnzh znęstotliworzi and pnrnmktryznak Knak wikloOcacłowk powiąnnak n syganłkm kowamixg;and characterized in that the obljmgjk: - okbikrnaik (5595) knayzh wkjrziowyzh nnwikrnjązyzh syganł NOnanłowy okpowinknjązy syganłowi kowamixg M-Onanłowkgo syganłg ngkio, M>N, mnjązyzh mnzikrnk Ookownain pokpnsmn o wnrtorzinzh nkspoloayzh nnstosownak w pokpnsmnzh znęstotliworzi i pnrnmktryznak knak wikloOcacłowk powiąnnak n syganłkm kowamixg;i znamienny tym, że kokntOowo obkjmgjk: - gkakrowcaik (5593) pokpnsm znęstotliworzi kln syganłg NOnanłowkgo, prny znym zo anjmaikj aikOtórk n pokpnsm znęstotliworzi są pokpnsmnmi znęstotliworzi o wnrtorzinzh rnkznywistyzh;- gkakrowcaik (5593) pokpnsm, frequented by NOnanłowkgo, anneutical and anonymous, and often by far-feverish people, they are often buried in a wide variety of ways;- wynanznnaik (5595) mnzikrny kkOokownain pokpnsmn o wnrtorzinzh rnkznywistyzh kln Oompkasnzji stosownain mnzikrny Ookownain w okpowikkni an pnrnmktryznak knak wikloOcacłowk;ornn - wynanznnaik (5595) mnzikrny kkOokownain pokpnsmn o wnrtorzinzh rnkznywistyzh kln Oompkasnzji stosownin mnzikrny Ookownain in okpowikkni an pnrnmktryznak knak wikloocacłowk;ornn - generating (1507) downmix data corresponding to the downmix signal by matrix multiplication of the real-time subband decoding matrix and the N-channel signal data in at least some of the real-frequency subbands. - generowanie (1507) danych downmixowanych odpowiadających sygnałowi downmixu poprzez mnożenie macierzowe macierzy dekodowania podpasma o wartościach rzeczywistych i danych sygnału N-kanałowego w co najmniej niektórych podpasmach częstotliwości o wartościach rzeczywistych. 13. A receiver (703) for receiving an N-channel signal, wherein the receiver (703) comprises: 13. Odbiornik (703) do odbioru sygnału N-kanałowego, przy czym odbiornik (703) zawiera: - means (801) for receiving input data comprising a N-channel signal corresponding to a downmix signal of the M-channel audio signal, M> N, having complexity subband matrix matrices used in the frequency subbands and parametric multi-channel data associated with the downmix signal;and characterized in that it additionally contains: - środki (801) do odbioru danych wejściowych zawierających sygnał Nkanałowy odpowiadający sygnałowi downmixu M-kanałowego sygnału audio, M>N, mających macierze kodowania podpasma o wartościach zespolonych zastosowane w podpasmach częstotliwości i parametryczne dane wielokanałowe powiązane z sygnałem downmixu;i znamienny tym, że dodatkowo zawiera: - means (805) for generating frequency subbands for a N-channel signal, wherein at least some of the frequency subbands are real frequency subbands;- środki (805) do generowania podpasm częstotliwości dla sygnału Nkanałowego, przy czym co najmniej niektóre z podpasm częstotliwości są podpasmami częstotliwości o wartościach rzeczywistych;- determining means (809) for determining the real-time decoding matrix for compensating the use of the coding matrix in response to parametric multi-channel data;- środki (809) wyznaczania do wyznaczania macierzy dekodowania podpasma o wartościach rzeczywistych dla kompensacji stosowania macierzy kodowania w odpowiedzi na parametryczne dane wielokanałowe;- means (807) for generating downmix data corresponding to the downmix signal by matrix multiplication of the real-time subband decoding matrix and the N-channel signal data in at least some of the real frequency subbands. - środki (807) do generowania danych downmixowanych odpowiadających sygnałowi downmixu poprzez mnożenie macierzowe macierzy dekodowania podpasma o wartościach rzeczywistych i danych sygnału N-kanałowego w co najmniej niektórych podpasmach częstotliwości o wartościach rzeczywistych. 14. A transmission system (700) for transmitting an audio signal, the transmission system comprising: 14. Układ (700) transmisji do transmitowania sygnału audio, przy czym układ transmisji zawiera: - transmiter (701) zawierający: - transmitter (701) containing: - means (709) for generating an N-channel downmix signal of the M-channel audio signal, M> N, - środki (709) do generowania N-kanałowego sygnału downmixu M-kanałowego sygnału audio, M>N, - means (709) for generating parametric multi-channel data associated with the downmix signal, - środki (709) do generowania parametrycznych danych wielokanałowych powiązanych z sygnałem downmixu, - means (709) for generating the first N-channel signal by using a complex band coding matrix of subband for the N-channel downmix signal on the frequency subbands, - środki (709) do generowania pierwszego sygnału N-kanałowego poprzez zastosowanie macierzy kodowania podpasma o wartościach zespolonych dla N-kanałowego sygnału downmixu w podpasmach częstotliwości, - means (709) for generating a second N-channel signal including a first N-channel signal and parametric multi-channel data, and - środki (709) do generowania drugiego sygnału N-kanałowego zawierającego pierwszy sygnał N-kanałowy i parametryczne dane wielokanałowe, oraz - means (711) for transmitting a second channel signal to the receiver (703);and - środki (711) do transmitowania drugiego sygnału Nkanałowego do odbiornika (703);oraz - a receiver (703) containing: - odbiornik (703) zawierający: - means (801) for receiving a second N-channel signal, and a transmission system characterized in that the receiver further comprises: - środki (801) do odbierania drugiego sygnału N-kanałowego, i układ transmisji znamienny tym, że odbiornik zawiera dodatkowo: - means (805) for generating frequency subbands for the first N-channel signal, wherein at least some of the frequency subbands are real frequency subbands;- środki (805) do generowania podpasm częstotliwości dla pierwszego sygnału N-kanałowego, przy czym co najmniej niektóre z podpasm częstotliwości są podpasmami częstotliwości o wartościach rzeczywistych;- determining means (809) for determining the real-time decoding matrix for compensating the use of a coding matrix in response to parametric multi-channel data;and - środki (809) wyznaczania do wyznaczania macierzy dekodowania podpasma o wartościach rzeczywistych dla kompensacji zastosowania macierzy kodowania w odpowiedzi na parametryczne dane wielokanałowe;oraz - means (807) for generating downmix data corresponding to the N-channel downmix signal by matrix multiplication of the real-value sub-band decoding matrix and the N-channel signal data in at least some of the real-frequency subbands. - środki (807) do generowania danych downmixowanych odpowiadających N-kanałowemu sygnałowi downmixu poprzez mnożenie macierzowe macierzy dekodowania podpasma o wartościach rzeczywistych i danych sygnału Nkanałowego w co najmniej niektórych podpasmach częstotliwości o wartościach rzeczywistych. 15. A method for receiving an audio signal, the method comprising: 15. Sposób odbierania sygnału audio, przy czym sposób obejmuje: - odbieranie (1501) danych wejściowych zawierających sygnał Nkanałowy odpowiadający sygnałowi downmixu M-kanałowego sygnału audio, M>N, mających macierze kodowania podpasma o wartościach zespolonych zastosowane w podpasmach częstotliwości i parametryczne dane wielokanałowe powiązane z sygnałem downmixu;i znamienny tym, że dodatkowo obejmuje: - receiving (1501) input data comprising a N-channel signal corresponding to a downmix signal of the M-channel audio signal, M> N, having complexity subband matrix matrices used in the frequency subbands and parametric multi-channel data associated with the downmix signal;and characterized in that it further comprises: - generating (1503) frequency subbands for a N-channel signal, wherein at least some of the frequency subbands are real frequency subbands;- generowanie (1503) podpasm częstotliwości dla sygnału Nkanałowego, przy czym co najmniej niektóre z podpasm częstotliwości są podpasmami częstotliwości o wartościach rzeczywistych;- determining (1505) a real-character sub-band decoding matrix to compensate for the use of a coding matrix in response to parametric multi-channel data;and - wyznaczanie (1505) macierzy dekodowania podpasma o wartościach rzeczywistych dla kompensacji stosowania macierzy kodowania w odpowiedzi na parametryczne dane wielokanałowe;oraz - generating (1507) the downmix data corresponding to the downmix signal by matrix multiplication of the real-value subband decoding matrix and the N-channel signal data in at least some of the real-frequency subbands. - generowanie (1507) downmixowanych danych odpowiadających sygnałowi downmixu poprzez mnożenie macierzowe macierzy dekodowania podpasma o wartościach rzeczywistych i danych sygnału N-kanałowego w co najmniej niektórych podpasmach częstotliwości o wartościach rzeczywistych. 16. A method for transmitting and receiving an audio signal, the method comprising: 16. Sposób transmitowania i odbierania sygnału audio, przy czym sposób obejmuje: - in the transmitter (701) implementation of the stages: - w transmiterze (701) realizację etapów: - generating an N-channel downmix signal of the M-channel audio signal, M> N, - generowania N-kanałowego sygnału downmixu M-kanałowego sygnału audio, M>N, - generating parametric multi-channel data associated with the downmix signal, - generowania parametrycznych danych wielokanałowych powiązanych z sygnałem downmixu, - generating the first N-channel signal by using a subband band encoding matrix with complex values for the N-channel downmix signal on the frequency subbands, - generowania pierwszego sygnału N-kanatowego poprzez zastosowanie macierzy kodowania podpasma o wartościach zespolonych dla Nkanałowego sygnału downmixu w podpasmach częstotliwości, - generating a second N-channel signal including a first N-channel signal and parametric multi-channel data, and - generowania drugiego sygnału N-kanałowego zawierającego pierwszy sygnał N-kanałowy i parametryczne dane wielokanałowe, oraz - transmitowania drugiego sygnału N-kanałowego do odbiornika (703);oraz - transmitting the second N-channel signal to the receiver (703);and - in the subgroup (033), the implementation of the etppas: - w obbiorniuu (033) realizccję etppów: - odbierania (1501) drugiego sygnału N-kanałowego;a sposób jest znamienny tym, że odbiornik dodatkowo realizuje etapy: - receiving (1501) a second N-channel signal;and the method is characterized in that the receiver additionally performs the steps of: - generating (1503) frequency subbands for the first N-channel signal, wherein at least some of the frequency subbands are real frequency subbands;- generowania (1503) podpasm częstotliwości dla pierwszego sygnału Nkanałowego, przy czym co najmniej niektóre z podpasm częstotliwości są podpasmami częstotliwości o wartościach rzeczywistych;- determining (1505) a real-character sub-band decoding matrix to compensate for the use of a coding matrix in response to parametric multi-channel data;- wyznaczania (1505) macierzy dekodowania podpasma o wartościach rzeczywistych dla kompensacji stosowania macierzy kodowania w odpowiedzi na parametryczne dane wielokanałowe;- generating (1507) downmixed data corresponding to the downmix signal by multiplying the matrix of the real-time decoding subbands and the N-channel data in at least some of the real-frequency subbands. - generowania (1507) downmixowanych danych odpowiadających Nkanałowemu sygnałowi downmixu poprzez mnożenie macierzowe macierzy dekodowania podpasma o wartościach rzeczywistych i danych sygnału N-kanałowego w co najmniej niektórych podpasmach częstotliwości o wartościach rzeczywistych. 17. A computer program (production) for performing the method defined in any of claims 1-8. 12, 15, 16. 17. Program komputerowy (wytwór) do wykonywania sposobu określonego w dowolnym z zastrz. 12, 15, 16. 18. Astring) 770) to a refreshing aadio containing dekooee) 7 (^) specified in claim 1. 18. Uroąąoenie )770) do odtwótoaaia aadio zawierane dekooee )7(^) okreelony w zastrzeżeniu 1. Koninklijke Philips N. V., Holandia Dolby International AB, Holandia Koninklijke Philips NV, Netherlands Dolby International AB, The Netherlands Pełnomocnik: Proxy: Input multi-channel PCM bit signal Wejściowy wielokanałowy sygnał PCM bitów MPEG Surround bitstream Strumień bitów MPEG Surround FIG. 1 Stan techniki FIG. 1 State of the art Ul ul Ν Ν Ι- * Ι-* Ul ul Μ ο Μ ο Μ Μ Ο Ο Ο Ο Ο Ο Μ υ Μ υ Μ Μ Ο Ο A core bit stream (e.g. AAC) Rdzeniowy strumień bitów (np. AAC) Downmixowany The downmix PCM mono / stereo PCM mono / stereo MPEG Surround bitstream Strumień bitów MPEG Surround AND I Core coder Koder rdzeniowy Przestrzenny strumień bitów Spatial stream of bits Output multi-channel PCM signal Wyjściowy wielokanałowy sygnał PCM NJ NJ AND-1 ul I—1 Ul FIG. 2 Stan techniki FIG. 2 State of the art Z-15202/16 Z-15202/16 Kodo. Kodo. macierzowe matrix J-► J—► Deko. Deko. macierzowe matrix FIG. 3 Stan techniki ϊ FIG. 3 State of the art ϊ Ν Ν Ι- * Ι-* Ul ul Μ ο Μ ο Μ Μ Ο Ο Ο Ο Ο Ο Μ υ Μ υ Μ Μ Ο Ο Input multi-channel PCM signal Wejściowy wielokanałowy sygnał PCM MPEG Surround bitstream Strumień bitów MPEG Surround AND- I— Ul ul FIG. 4 I Stan techniki o FIG. 4 I The state of the art o And «I_l, O I « I_l, O Ul o Ul o Η 00 Η 00 O H· OH · Rdzeniowy strumień bftów (np. AAC) The core stream of bft (e.g., AAC) Downmix stereo PCM compatible matrix Downmix stereo PCM kompatybilny macierzowo MPEG Surround bitstream Strumień bitów MPEG Surround Core coder ι Koder rdzeniowy ι Koder kompatybilności macierzowej ~r~ The array compiler ~ r ~ Parametry przestrzenne Spatial parameters Output multi-channel PCM signal Wyjściowy wielokanałowy sygnał PCM U1 U1 FIG. 5 F Stan techniki o FIG. 5 F The state of the art o ι ο I_l, O ι ο I_l, O Ul o Ul o Η 00 Η 00 O H· OH · M PCM input samples M próbek wejściowych PCM FIG. 6 FIG. 6 Jedna szczelina K próbek o wartościach zespolonych i M-K o wartościach rzeczywistych w domenie podpasmowej £ o One K-gap of complex values and MK of real-valued values in the sub-band domain of £ o N N 1 · + 1·+ Ul ul M M ABOUT O M M ABOUT O ABOUT O ABOUT O SI SI -AT -U SI SI ABOUT O EP 1 999 747 B1 Z-15202/16 EP 1 999 747 B1 Z-15202/16 7/15 7/15 EP 1 999 747 B1 Z-15202/16 EP 1 999 747 B1 Z-15202/16 8/15 γΙ \ < 8/15 γΙ\< 0© © 0 v >> v>> \ ji2- ϋ ·6 mA kO \ji2- ϋ·6 mA kO Tn % Tn% Μ Ό v> * 0 V \ -A 'έ Μ Ό v> *0 V\ -A 'έ Φ φ Φ φ ^ \ 6 ΕΡ 1 999 747 Bl Z-15202/16 u / 15, Ό ^\6 ΕΡ 1 999 747 Bl Z-15202/16 u/15 ,Ό AND I Ό Ό V * <** θ ' V* < ** θ' Ο * - * Ο*-* Hearing, Posłuch, EP 1 999 747 B1 Z-15202/16 EP 1 999 747 B1 Z-15202/16 15/15 15/15
218 paragraphs in 4 sections, as filed
[0001] The invention relates to audio decoding, and more particularly, but not exclusively, to decoding MPEG Surround signals.
[0002] In recent decades, the digital coding of different source signals has become more and more important when signal representation and digital communication have in the increasing manner replaced analog representation and communication. For example, the distribution of media content, such as video and music, is increasingly based on the coding of digital content.
[0003] In addition, in the last decade, there has been a trend towards multi-channel audio, and in particular towards spatial audio going beyond conventional stereo signals. For example, traditional stereo recordings contain only two channels, while modern advanced audio systems typically use five or six channels, as in popular 5.1 Surround sound systems. This provides a more engaging listening experience where the user can be surrounded by sound sources.
[0004] Various techniques and standards have been developed for the transmission of such multi-channel signals. For example, six discrete channels representing 5.1 Surround can be transmitted in accordance with standards such as Advanced Audio Coding (AAC) or Dolby Digital.
[0005] However, to ensure backward compatibility, it is known to downmix a larger number to a smaller number of channels, and in particular often downmixing 5.1 surround sound signals to a stereo signal enables stereo signals to be reproduced by older decoders (stereo), and 5.1 signals by surround decoders. .
[0006] One example is the backwards compatible MPEG2 encoding method. The multi-channel signal is downmixed to a stereo signal. Additional signals are encoded as multi-channel data in the secondary data portion, allowing a multi-channel MPEG2 decoder to generate a multi-channel signal representation. The MPEG1 decoder will ignore the auxiliary data and thus will only decode the stereo downmix. The main disadvantage of the coding method used in MPEG2 is that the additional data rate required for additional signals is the same order of magnitude as the bit rate required to encode the stereo signal.
The additional bit rate for stereo extension to multi-channel audio is therefore significant.
[0007] Other existing methods for backward compatible multi-channel transmission without additional multi-channel information can typically be characterized as matrix-surround methods. Examples of matrix surround encoding include methods such as Dolby Prologic II and Logic-7. A common principle of these methods is that they use matrix multiplication of multiple channels of the input signal through the appropriate matrix, thus generating an output signal with a lower number of channels. In particular, the matrix coder typically applies a phase shift of the surround channels prior to mixing them with the front and middle channels.
[0008] Another reason for channel conversions is the coding efficiency. It has been found that e.g. audio signals of surround sound can be coded as stereo channel audio signals connected to a bit stream of parameters describing the spatial properties of the audio signal. The decoder can play stereo stereo audio signals with a very satisfactory level of accuracy. In this way, you can achieve significant bit rate savings.
[0009] There are several parameters that can be used to describe the spatial properties of audio signals. One such parameter is the cross-channel cross-correlation, such as cross-correlation between the left channel and the right channel for stereo signals. Another parameter is the power ratio of channels. In so-called (parametric) (en) coders) of spatial audio, such as an MPEG surround encoder, these and other parameters are extracted from the original audio signals so as to produce an audio signal with a reduced number of channels, e.g. only one channel plus a set of parameters describing the spatial properties of the original audio signal. In so-called (parametric) spatial audio decoders, the spatial properties described by the transmitted spatial parameters are installed again.
[0010] Such spatial audio coding preferably uses a cascade or hierarchical tree-based structure including standard units at the encoder and decoder. In the encoder, these standard units may be downmixers, connecting channels to a lower number of channels, such as a 2 to 1, 3 to 1, 3 to 2 downmixer, etc., while the corresponding standard units of the decoder may be upmixers dividing the channels into a higher number of channels, such as upmixery 1 to 2, 2 to 3.
[0011] Fig. 1 shows an example of an encoder for encoding multichannel audio signals according to an approach that is currently standardized by MPEG under the name MPEG Surround. The MPEG Surround circuit encodes a multi-channel signal as a mono or stereo downmix in the company of a set of parameters. The downmix signal may be encoded by a previous audio encoder, such as an MP3 or AAC encoder. The parameters represent a spatial image of a multi-channel audio signal and can be coded and input in a backward compatible manner to an earlier audio stream.
[0012] On the decoder side, the core bit stream is first decoded which generates a mono or a stereo downmix signal. Earlier decoders, i.e. decoders that do not use MPEG surround decoding, can still decode the downmix signal. However, if an MPEG Surround decoder is available, the spatial parameters are reinstalled, resulting in a multi-channel representation that is perceptually close to the original multi-channel input signal. An example of an MPEG surround decoder is illustrated in Fig. 2.
[0013] In addition to the basic spatial coding / decoding illustrated in Fig. 1 and Fig. 2, the MPEG Surround system offers a rich set of features that allow a wide field of applications. One of the distinguishing features is called matrix compatibility (Matrix Compatibility) or compatibility with matrix (matrix) Surround (Matrix (ed) Surround Compatibility).
[0014] A general outline of MPEG Surround is shown in J. Breebart et al., "MPEG Spatial Audio Coding / MPEG Surround: Overview and Current Status", Audio Engineering Society Convention Paper, presented at 119<sup>th</sup> Convention, New York, USA, October 2005, pp. 1-17.
[0015] Examples of traditional matrix surround systems are Dolby Pro Logic I and II and Circle Surround. These systems operate as shown in Fig. 3. The multi-channel input PCM signal is converted to a so-called matrix downmix signal, typically using matrix 5 (.1) to 2. The underlying idea of matrix array systems is that the front channels and The surround (rear) speakers are mixed in the stereo downmix signal accordingly in phase and not in phase. To a certain extent, this enables the reverse on the decoder side to lead to the reconstruction of multi-channel.
[0016] In matrix matrix systems, a stereo signal may be transmitted using traditional channels intended for stereo transmission. Hence, as in the MPEG Surround system, matrix surround systems also offer some form of backward compatibility. However, due to the specific phase characteristics of the stereo downmix signal resulting from matrix surround-encoding, these signals often do not have high-quality sound when they are listened to as a stereo signal, e.g. from loudspeakers or headphones.
[0017] In the matrix surround decoder, matrix M to N (where eg M = 2, and N = 5 (.1)) is used to generate a multi-channel PCM output signal. However, in general, the matrix system N to M (where N> M) is not reversible and therefore the matrix Surround systems are generally unable to accurately reconstruct the original multi-channel output PCM signals that tend to contain very noticeable artifacts.
[0018] Unlike such traditional matrix-matrix systems, the compatibility of matrix surround in MPEG Surround is achieved by using a 2x2 matrix for composite values of the samples in the MPEG Surround encoder frequency subbands, after MPEG Surround coding. An example of such an encoder is illustrated in Fig. 4. The 2x2 matrix is generally a matrix with complex values with coefficients depending on spatial parameters. In such a system, the spatial parameters are variable both in time and in frequency, and as a result the 2x2 matrix is also variable both in time and in frequency. Accordingly, the combined matrix operation is typically used on time-frequency tiles.
[0019] The use of the matrix matrix compatibility functionality in the MPEG Surround encoder allows the resulting stereo signal to be compatible with a signal generated by conventional matrix matrix coders, such as Dolby ProLogic ™. This makes the earlier decoders can decode the surround signal. In addition, the matrix matrix compatibility operation may be inverted to a MPEG compatible decoder, thus enabling the generation of a high quality multi-channel signal.
[0020] The array matrix compatibility matrix may be described as follows:
<td>^ ΜΠΧ</td><td>= Η</td><td>'L</td><td></td><td></td><td></td><td>L</td>
<td>_ ^ ΜΠΧ _</td><td></td><td>R</td><td></td><td>_<sup>h</sup>2 \</td><td>h<sub>22</sub>_</td><td>R</td>
where L, R is a conventional MPEG stereo downmix, Lmtx, Rmtx is a downmix encoded in the matrix surround and where hxy are complex coefficients determined in response to multi-channel parameters.
[0021] The main advantage of providing matrix-compatible stereo signals using a 2x2 matrix is that the matrices may be inverse. As a result, the MPEG Surround decoder can still provide the same quality of the output audio regardless of whether a matrix compatible downmix is used in the encoder or not. An example of a compatible MPEG surround decoder is illustrated in Fig. 5.
The reverse processing on the decoder side in the ordinary MPEG surround decoder can thus be determined by means of:
<td>'L</td><td>= H "'</td><td>/ μπύ</td><td></td><td>11 ^, D</td><td>^ 12 D</td><td>^ ΜΠΧ</td>
<td>R</td><td></td><td>Αμτχ _</td><td></td><td>_ ^ 21, d</td><td>^ 22, D _</td><td>_ & MTX.</td>
[0023] In this way, since H can be reversed, the encoder operation with matrix compatibility can be reversed.
[0024] In the MPEG Surround system, processing, including matrix compatibility operations, takes place in the frequency domain. More specifically, the so-called complex exponential modulated quad mirror filter banks (QMS) to divide the frequency axis into a number of bands.
[0025] In many ways, this type of QMF banks can be compared to a discrete Fourier transform (DFT) with Overlap-Add or its efficient counterpart, Fast Fourier Transform (FFT) Fourier Transfrom). The QMF bank, as well as the DFT bank, share the following desired properties for signal manipulation:
· Frequency representation is over sampled. Thanks to this property it is possible to use manipulations, such as correction (scaling of individual bands) without aliasing distortion. Critically sampled representations, such as, for example, the well-known Modified Discrete Cosine Transform (MDCT), which is e.g. used in AAC, do not undergo this feature. Hence, variable in time and frequency modifications of MDCT coefficients before synthesis lead to aliasing, which in turn generates audible artifacts in the output signal.
· Frequency representation has complex values. In contrast to the real-valued representation, a complex-valued representation enables a simple modification of the phase of the signals.
[0026] Although there are many advantages over the critically sampled representation with real values in terms of signal manipulation, a significant disadvantage compared to such representations is the computational complexity. Much of the complexity of the MPEG Surround decoder results from the analysis of QMF and synthesis filter banks and the associated signal processing of complex values.
[0027] Accordingly, it has been proposed to implement real-world processing parts for a so-called low power decoder (LP). To this end, a complex modulated filter bank has been replaced by a real-world cosmic exchange filter bank, followed by a partial extension to the domain with complex values for low-frequency bands. Such a filter bank is shown in Fig. 6.
[0028] In normal operation, the MPEG Surround decoder uses real-valued processing for complex domain subband samples or, in the case of LP, applies them to real-valued subfield samples. However, the matrix compatibility feature in the decoder involves phase rotation to restore the original stereo downmix in the frequency domain. These phase rotations are realized by means of processing with complex values. In other words, matrix H '<sup>1 </sup>Compatible matrix decoding in a natural way has complex values in order to introduce the required phase revolutions. Accordingly, in such systems the operation compatible with the matrix surround can not be reversed in the part with the actual values of the LP representation in the frequency domain, which leads to a reduced quality of decoding.
[0029] Hence, improved audio decoding would be beneficial.
[0030] Accordingly, the invention is intended to advantageously attenuate, alleviate or eliminate one or more of the above-mentioned disadvantages, alone or in any combination.
According to a first aspect of the invention, an audio decoder is provided including: means for receiving input data including an N-channel signal corresponding to a downmix signal of the M-channel audio signal, M> N, having subband banding matrixes used in the frequency subbands and parametric data multi-channel related to the downmix signal; means for generating frequency subbands for the N-channel signal, wherein at least some of the frequency subbands are real frequency subbands; means for determining, for determining the real-time decoding matrix to compensate for the use of a coding matrix in response to parametric multi-channel data;
[0032] The invention may enable improved and / or facilitated decoding. In particular, the invention can allow a significant reduction in complexity while providing high quality audio. The invention may, for example, allow at least a partial reversal of the matrix multiplication effects of a subband with composite values on the decoder side using real frequency subbands.
[0033] As a specific example, the invention may e.g. allow partial reversal of MPEG compatible MPEG (MPEG) Compatible Encoding in the MPEG Surround decoder using real-valued frequency subbands.
[0034] The decoder may include means for generating a downmix signal in response to the downmix data and may further include means for generating the Mk11 audio signal in response to the downmix data and the parametric multi-channel data. The invention may, in such embodiments, generate a valid multichannel audio signal, at least partly based on frequency subbands with real values.
[0035] For each frequency subband a different decoding matrix may be determined.
[0036] According to an optional feature of the invention, the determining means are adapted to determine the inverse subband matrix with complex values of the coding matrix and to determine the decoding matrix in response to the inverse matrix.
[0037] This may allow particularly efficient performance and / or improved decoding quality.
According to an optional feature of the invention, the means for determining is adapted to determine each matrix coefficient having real values of the decoding matrix in response to the absolute value of the respective matrix coefficient, the inverse matrix.
[0039] This may allow particularly efficient performance and / or improved decoding quality. Each matrix coefficient with real values, a decoding matrix, can be determined in response to the absolute value of only the corresponding matrix coefficient, the inverse matrix, without taking into account any other matrix coefficient. A suitable matrix coefficient can be a matrix coefficient in the same place of the inverse matrix for the same frequency subband.
According to an optional feature of the invention, the means for determining is adapted to determine each true-value matrix coefficient substantially as the absolute value of the respective matrix coefficient, the inverse matrix.
[0041] This may allow for particularly efficient performance and / or improved decoding quality.
[0042] According to an optional feature of the invention, the means for determining is adapted to determine the decoding matrix in response to the subpacket transfer matrices being a multiplication of the respective decoding matrix and the coding matrix.
[0043] This may allow particularly efficient performance and / or improved decoding quality. Suitable decoding and coding matrices may be coding and decoding matrices for the same frequency subband. The determining means may in particular be adapted to select the values of the coefficients of the decoding matrix in such a way that the transfer matrices have desirable properties.
[0044] According to an optional feature of the invention, the means for determining is adapted to determine a decoding matrix in response to a measure of the module only of the transfer matrix.
[0045] This may allow particularly efficient performance and / or improved decoding quality. In particular, the determining means may be adapted to ignore the phase measure when determining the decoding matrix. This can reduce complexity while maintaining a small degradation of perceptual audio quality.
[0046] According to an optional feature of the invention, the transfer matrices of each subband are given by where G is the subband decoding matrix and H is the subband banding matrix and the determining means are adapted to select the matrix coefficients
<td>P-</td><td>pn</td><td>Mo '</td><td>-GH</td><td></td><td>£ l2</td><td></td><td>/<sup>p</sup>2</td>
<td></td><td>_Pit</td><td>P22.</td><td></td><td>_A21</td><td><22.</td><td>_ ^ 21</td><td>^ 22 _</td>
2.1 22 _ &> 21 & 22_ in such a way that the power measures P12 and P21 of the power meet the criterion.
[0047] This may allow particularly efficient performance and / or improved decoding quality. The decoding matrix can be selected in such a way as to lead to a power measure below a threshold value (which can be determined in response to constraints or other parameters) or can e.g. be selected as a decoding matrix leading to a minimum power measure.
[0048] According to an optional feature of the invention, the size of the module is determined in response to
<img file="PL1999747T3_D0001.tif" />
[0049] This may allow particularly efficient performance and / or improved decoding quality.
[0050] According to an optional feature of the invention, the determining means are further adapted to select matrix coefficients at constraints that the pn and p22 modules are substantially equal to unity.
[0051] This may allow particularly efficient performance and / or improved decoding quality.
[0052] According to an optional feature of the invention, the downmix signal and the parametric multi-channel data are compliant with the MPEG Surround standard.
[0053] The invention, for a signal compatible with MPEG Surround, may allow for particularly efficient, low complexity and / or higher audio quality decoding.
According to an optional feature of the invention, the coding matrix is an MPEG Matrix Surround Compatibility matrix and the first N-channel signal is an MPEG Matrix Surround Compatibility signal.
[0055] The invention may allow for a particularly efficient, low complexity and / or improved audio quality, and in particular may enable low complexity decoding for efficient compensation of the MPEG Matrix Surround Compatibility operation performed by the encoder.
According to another aspect of the invention, an audio decoding method is provided, the method comprising: receiving input data including a N9 channel signal corresponding to a downmix signal of the M-channel audio signal, M> N, having complex band subband matrixes used in the frequency subbands and parametric multi-channel data associated with the downmix signal; generating frequency subbands for the N-channel signal, wherein at least some of the frequency subbands are real frequency subbands; determining a true-to-true subband decoding matrix to compensate for the use of a coding matrix in response to parametric multi-channel data;
According to another aspect of the invention, a receiver is provided for receiving a N-channel signal, the receiver comprising: means for receiving input data comprising an N-channel signal corresponding to a downmix signal of the M-channel audio signal, M> N, having sub-band encoding matrices with values complex devices used in the frequency subbands and parametric multi-channel data associated with the downmix signal; means for generating frequency subbands for the N-channel signal, wherein at least some of the frequency subbands are real frequency subbands; determining means for determining a real-value decoding matrix for compensating the use of a coding matrix in response to parametric multi-channel data;
According to another aspect of the invention, a transmission system for transmitting an audio signal is provided, wherein the transmitter comprises: a transmitter including: means for generating an N-channel downmix signal of the M-channel audio signal, M> N, means for generating parametric multi-channel related data with a downmix signal, means for generating the first N-channel signal by using a subband bandwidth matrix for the N-channel downmix signal on the frequency subbands, means for generating a second N-channel signal including an N-channel signal and parametric multi-channel data, and means for transmission a second N-channel signal to the receiver; and a receiver comprising: means for receiving a second N-channel signal, means for generating frequency subbands for the first N-channel signal, wherein at least some of the frequency subbands are real frequency subbands; means for determining the real-time decoding matrix for the use of a coding matrix in response to parametric multi-channel data, and means for generating downmix data corresponding to the downmix signal by matrix multiplication of the real-valued subband decoding matrix and the N-channel data in at least some real frequency subbands.
[0059] The second N-channel signal may include an additional associated channel including parametric multi-channel data.
According to another aspect of the invention, there is provided a method of receiving an audio signal from a scalable audio audio stream, the method comprising: receiving input data including an N-channel signal corresponding to a downmix signal of an M-channel audio signal, M> N, having a subband matrix matrix with complex values used in the frequency subbands and parametric multi-channel data associated with the downmix signal; generating frequency subbands for the N-channel signal, wherein at least some of the frequency subbands are real frequency subbands; determining a true-to-true subband decoding matrix to compensate for the use of a coding matrix in response to parametric multi-channel data;
[0061] According to another aspect of the invention, a method is provided for transmitting and receiving an audio signal, the method comprising:
in the transmitter implementation of stages:
generating an N-channel downmix signal of the M-channel audio signal, M> N, generating parametric multi-channel data associated with the downmix signal, generating the first N-channel signal by using a real-valued subband bandwidth matrix for the N-channel downmix signal on the frequency subbands, generating a second N-channel signal including the first N-channel signal and parametric multi-channel data, and transmitting the second N-channel signal to the receiver;
and in the receiver implementing the steps of: receiving a second N-channel signal;
generating frequency subbands for the first N-channel signal, wherein at least some of the frequency subbands are real frequency subbands;
determining a real-value decoding matrix for compensating the use of a coding matrix in response to parametric multi-channel data;
generating downmix data corresponding to the N-channel downmix signal by multiplying the matrix of the real-value decoding matrix and the N-channel data in at least some of the real-frequency subbands.
[0062] These and other aspects, features and advantages of the invention will be apparent and explained with reference to the embodiment (s) described below.
[0063] Embodiments of the invention will be described, by way of example only, with reference to the drawings in which
Fig. 1 shows an example of an encoder for encoding multichannel audio signals according to the prior art;
Fig. 2 shows an example of a decoder for decoding multichannel audio signals according to the prior art;
Fig. 3 shows an example of a surround / decode system of the surround matrix according to the prior art;
Fig. 4 shows an example of an encoder for encoding multichannel audio signals according to the prior art;
Fig. 5 shows an example of a decoder for decoding multichannel audio signals according to the prior art;
Fig. 6 shows an example of a filter bank for generating frequency subbands with complex values and real values;
Fig. 7 shows a transmission system for communicating an audio signal according to some embodiments of the invention;
Fig. 8 shows a decoder according to some embodiments of the invention;
Figs. 9-14 show decoder operation characteristics according to some embodiments of the invention; and
Fig. 15 shows a decoding method according to some embodiments of the invention.
The following description focuses on the embodiments of the invention applicable in the decoder for decoding an encoded MPEG Surround signal encompassing a coding compatible with the matrix surround (Matrix Surround Compatibility). However, it should be noted that the invention is not limited to this application, but can be used in many other coding standards.
[0065] Fig. 7 shows a 700 transmission system for communicating an audio signal according to some embodiments of the invention. The transmission system 700 comprises a transmitter 701, which is coupled to the receiver 703 via a network 705, which, in particular, may be the Internet.
In a particular example, the transmitter 701 is a signal recording apparatus and the receiver 703 is a signal player device, but it should be noted that in other embodiments the transmitter and receiver may be used in other applications and for other purposes.
[0067] In a particular example in which the signal recording function is supported, the transmitter 701 includes an analog-digital converter 707 that receives an analogue multi-channel signal that is converted into a digital multi-channel PCM (Pulse Coded Modulated) signal. code) by sampling and analog-to-digital conversion.
The transmitter 701 is coupled to the encoder 709 of Fig. 1, which encodes the PCM signal according to the MPEG Surround coding algorithm, which encompasses the functionality of the Matrix Surround Compatibility compatible coding. The coder 709 may be, for example, the prior art decoder of Fig. 4. In the example, the encoder 709 specifically generates an MPEG stereo downmix signal compatible with the matrix-type stereo surround (MPEG Matrix Surround Compatible stereo down-mixed signal).
[0069] Thus, the encoder 709 generates a signal given by
<td>Ιμτχ</td><td>= Η</td><td>L</td><td></td><td></td><td>1CM</td><td>L</td>
<td>β-ΜΤΧ _</td><td></td><td>R</td><td></td><td>_AND<sub>21</sub></td><td>^ 22 _</td><td>R</td>
where L, R is a conventional MPEG Surround stereo downmix, and Lmtx, Rmtx is a downmixed coded output compatible with the matrix surround of encoder 709. In addition, the signal generated by the encoder 709 includes parametric multi-channel data generated by MPEG Surround encoding. In addition, hxy are complex coefficients determined in response to multi-channel parameters. As will be appreciated by those skilled in the art, the processing carried out by the encoder 709 is performed on subbands of complex values and using complex operations.
[0070] The encoder 709 is coupled to a network transmitter 711 that receives the encoded signal and connects to the network 705. The network transmitter 711 may transmit the encoded signal to the receiver 703 via the network 705.
[0071] The receiver 703 includes a network interface 713 that connects to the network 705 and that is adapted to receive the encoded signal from the transmitter 701.
[0072] The network interface 713 is coupled to the decoder 715. The decoder 715 receives the encoded signal and decodes it according to a decoding algorithm. In the example, the decoder 715 reproduces the original multi-channel signal. In particular, the decoder 715 first generates compensated downmix stereo corresponding to the downmix generated by MPEG surround coding before performing operations compatible with the MPEG surround. The decoded multi-channel signal is then generated from this downmix and received parametric multi-channel data.
In a particular example in which the signal reproduction function is supported, the receiver 703 further includes a signal player 717 that receives the decoded multi-channel audio signal from the decoder 715 and presents it to the user. In particular, the signal player 717 may comprise an analog-to-digital converter, amplifiers and loudspeakers required to output the decoded audio signal.
[0074] Fig. 8 shows the decoder 715 in more detail.
[0075] The decoder 715 includes a receiver 801 that receives the signal generated by the encoder 709. As previously mentioned, the signal is a stereo signal that corresponds to a downmixed signal that has been processed by multiplying complex values of samples on the frequency subbands of the complex matrix H on complex values. In addition, the received signal includes parametric multi-channel data that corresponds to the downmix signal. In particular, the received signal is an encoded MPEG Surround signal by processing compatibility with matrix surround (compatibility) processing.
[0076] The receiver 801 additionally provides core decoding of the received signal to generate a downmix PCM signal.
[0077] The receiver 801 is coupled to a parametric data processor 803 that acquires parametric multi-channel data from the received signal.
[0078] The receiver 801 is additionally coupled to a filter bank 805 that converts the received stereo signal into the frequency domain. In particular, the subband filter bank 805 generates multiple frequency subbands. At least some of these frequency subbands are real frequency subbands. The subband filter bank 805 may in particular correspond to the functionality shown in Fig. 6. In this way, the subband filter bank 805 may generate K subbands with complex values and MK subbands with real values. Real-valued subbands will typically be higher frequency subbands, such as subbands above 2 kHz. The use of real-valued subbands makes it much easier to generate subbands, as well as operations performed on samples in these subbands. In this way,
[0079] The subband filter bank 805 is coupled to an offset processor 807 that generates downmix data corresponding to the downmix signal. In particular, the compensation processor 807 is compensated for the matrix matrix compatibility operation by looking for the multiplication reversal by the code matrix H on the frequency subbands of the encoder 709. This compensation is performed by multiplying the subband data values by the G decoding subband matrix. However, in contrast to processing at encoder 709, the matrix multiplication in the real-valued sub-bands of the decoder 715 is implemented only in the real domain. In this way, not only the sample values are real-valued samples, but the matrix coefficients, the decoding matrix G, are also real-valued coefficients.
[0080] The compensation processor 807 is coupled to the matrix processor 809, which determines the decoding matrices to be applied to the subbands. For K samples with complex values, the G decoding matrix can be simply determined as the reciprocal of the coding matrix H in the same subband. However, for real-valued subbands, the matrix processor 809 determines matrix coefficients with real values that can provide efficient compensation for coding matrix operations.
[0081] In this way, the output of the compensation processor 807 corresponds to a subband representation of the downmix signal encoded according to MPEG Surround. Accordingly, the effects of a compatibility operation on the matrix surround can be substantially reduced or eliminated.
[0082] The compensation processor 807 is coupled to the synthesis filter subband 811, which generates a decoded downmix signal from the PCM MPEG time domain decoded down-mix signal from the subband representation. In a particular example, the sub-band synthesis filter bank 811 thus forms an equivalent for the sub-band filter bank 805 in converting the signal back to the time domain.
The subbands synthesis filter bank 811 is provided to the multichannel decoder 813, which is further coupled to the parametric data processor 803, the multi-channel decoder 813 receives the PCM downmix signal in the time domain and parametric multi-channel data and generates the original multi-channel signal.
[0084] In the example, the synthesis filter subband bank 811 converts to the time domain the subband signal on which the matrix operations were performed. The multi-channel decoder 813 thus receives an encoded MPEG surround signal comparable to a signal that would be received if no operations compatible with the matrix surround in the decoder were used. Therefore, the same multi-channel MPEG decoding algorithm can be used for signals compatible with the matrix surround and for signals incompatible with the matrix surround. However, in other embodiments, the multi-channel decoder 813 may act directly on the subband samples after compensation by the compensation processor 807. In that cases,
[0085] In this way, to reduce complexity, it is often advantageous to stay in the sub-band domain, providing a compensated signal to the multichannel decoder 813. As such, it is possible to avoid the complexity of the synthesis filter subband 811 and the analysis filterbank that are part of the multi-channel decoder
813.
[0086] In fact, if it is possible, it is advantageous not to move between the frequency domain and the time domain, because it is computationally expensive. Hence, in some set-top boxes, according to some embodiments of the invention, after converting a signal to the sub-domain (frequency) (which in turn was determined by decoding the core bit stream and applying filter banks on the resulting PCM signals), the inverse of the matrix surround is used in the 807 processor compensation (if applicable, i.e., if it is signaled in the bit stream), then the resulting signals in the sub-band domain are used directly to reconstruct multichannel signals (sub-band domain). Ultimately,
[0087] In this way, in the arrangement of Fig. 7, the encoder 709 can generate a signal compatible with the matrix surround that can be decoded by the earlier matrix-matrix decoders, such as Dolby Pro Logic ™ decoders. Although this requires distortion of the original downmixed signal encoded according to MPEG Surround by the Compound Matrix Surround Compensation operation, this operation can be effectively removed in the MPEG multi-channel decoder, which allows generation of a correct representation of the original multi-channel signal using parametric data.
[0088] In addition, the decoder 715 allows compensation for compatibility operations with matrix surround on real-valued frequency subbands instead of having to use frequency subbands, thereby greatly reducing the complexity of the decoder 715 when achieving high quality audio.
[0089] The following will describe examples of determining the appropriate matrix coefficients for a decoding matrix.
[0090] The encoder 709 performs a compatibility operation with the matrix surround by applying the following code matrix with complex values in each subband (it should be understood that each subband has a different code matrix):
<td>F-ΙΤΧ</td><td>= Η</td><td>'L</td><td></td><td></td><td>1 εΓ</td><td>L</td>
<td>β-ΜΤΧ _</td><td></td><td>R</td><td></td><td></td><td>h "22 J</td><td>R</td>
where L, R is a conventional stereo downmix, and Lmtx, Rmtx is a downmix coded according to the matrix surround. The matrix H of coding is given by:
_ 1 - w, + y, y - 2w, + 2w, _ 1<sup>-</sup> IN<sup>-</sup> J '<sup>IN</sup>2 J1 - 2w<sub>2</sub> + 2in2
<img file="PL1999747T3_D0002.tif" />
<img file="PL1999747T3_D0003.tif" />
φ? (ΐ - 2 in, + 2 in,<sup>2</sup>) where wi and w2 depend on spatial parameters generated by MPEG Surround encoding. In particular:
<img file="PL1999747T3_D0004.tif" />
<img file="PL1999747T3_D0005.tif" />
where wi, ti wi, t are non-normalized weights which are defined as:
<img file="PL1999747T3_D0006.tif" />
where CLD and CLD represent differences in channel levels (expressed in dB) of channel pairs respectively left-front, left-surround and right-front, right-surround. these<sub>with</sub> mtx and ci, mtxsą matrix coefficients, which are a function of coefficients c and c of prediction used to acquire the intermediate signals, left L, central C and right R from the downmixed signals of the left Ldmx and the right Rdmx in the decoder, as follows:
<td>L</td><td></td><td>c, +2 c<sub>2</sub> -1</td>
<td>R</td><td>=</td><td>c, - / c<sub>2</sub> + 2</td>
<td>C</td><td></td><td>1-c, 1- c<sub>2</sub> J</td>
ci, mtx and ci, mtx are designated as follows:
<img file="PL1999747T3_D0007.tif" />
-1 <c<sub>x</sub> <-0.5
-0.5 <c<sub>x</sub> <1, in other cases, respectively x = {0, 1}.
[0091] Alternatively, the MPEG Surround decoder supports a mode in which coefficients oi and c2 represent power ratios respectively left to left with center and right to right with center. In this case, different functions for a, mtx and C2, mtx apply.
[0092] Thus, for each time-frequency plate, complex matrix coding H is applied to the complexed sample. If the front signals would be dominant in the original input multi-channel signal, weights W1 and W2 would be close to zero. As a result, the matrix's downmix will be close to the input stereo downmix. If the surround signals were dominant in the original multi-channel input signal, weights W1 and W2 would be close to one. As a result, the downmixed matrix matrix signal will contain a strongly incompatible phase version of the original stereo downmix provided by the MPEG surround encoder.
[0093] A big advantage of providing matrix compatible stereo signals using a 2x2 matrix is that the matrices may be inverse. As a result, the MPEG Surround decoder can still provide the same output audio quality regardless of whether the encoder used a matrix compatible downmix or not.
[0094] Inverse processing on the decoder side in the MPEG surround decoder where all of the frequency subbands are subbands of complex values (e.g., using a QMF bank with complex modulation) is given by
L
R
<img file="PL1999747T3_D0008.tif" />
together with
<img file="PL1999747T3_D0009.tif" />
<img file="PL1999747T3_D0010.tif" />
<img file="PL1999747T3_D0011.tif" />
<img file="PL1999747T3_D0012.tif" />
where <sup>N</sup> - h \<sub>}</sub>h<sub>22</sub> h<sub>] 2</sub>h<sub>2}</sub>.
[0095] However, such a reverse operation requires that complex values be used and hence can not be used in the decoder 715 of Fig. 7 because (at least in part) it uses real-valued subbands. Accordingly, the matrix processor 809 generates a real-valued decoding matrix that can be used to significantly reduce the effects of the coding matrix.
[0096] The total effect of the coding and decoding matrix in each subband may be represented by the P matrix of the transfer given as
<td>P =</td><td>'Mo</td><td>Mo '</td><td>= GH =</td><td>bullfinch</td><td>& 12</td><td></td><td>/ ^ 2</td>
<td></td><td>.P21</td><td>P22.</td><td></td><td>-S21</td><td><sup>S</sup>N__</td><td>_ ^ 2l</td><td>^ 22.</td>
where H represents the encoder matrix and G represents the decoder matrix.
[0097] Ideally G = H '<sup>1</sup>, so that P = H '<sup>1</sup> · H = I, unit matrix. Due to the fact that all the hx scales of the encoder matrix H are complex, the matrix may not be inverted in the decoder for real-valued subbands.
[0098] Subbands with real values are typically at higher frequencies, like subbands above 2 kHz. At such frequencies, the phase relationships are perceptually less significant and therefore the matrix processor 809 determines the decoding matrix factors that have the corresponding characteristics of the module (power) without taking into account the phase characteristics. In particular, matrix processor 809 can determine matrix coefficients with real values which leads to low values of modules or power of pn and p2i and crosstalk expressions assuming or constraining that \ pii | ~ 1 and \ Ρ22 \ ~ 1.
[0099] In some embodiments, the matrix processor 809 may determine the inverse matrix H'i subband with complex values of the coding matrix, and then can determine the real-valued decoding matrix G from the matrix coefficients of the matrix. In particular, each coefficient in G can be determined from the coefficient of H'i, which is in the same place. For example, a real-valued coefficient can be determined from the value of the corresponding factor coefficient with H'1. In fact, in some embodiments, the matrix processor can determine coefficients in H'i, and then determine the coefficients in G as the absolute value of the corresponding matrix coefficient from inverse H'1 matrix [0100] In this way, the matrix processor 809 can determine the Sn1 Sn S22 bull.
as
Su = hu, D = | <
<img file="PL1999747T3_D0013.tif" />
| Μ73 (1-2η<sub>! +</sub>2 »·) '
<img file="PL1999747T3_D0014.tif" />
§22 = <sup>h</sup>
22, D
IW 'where <sup>N</sup> - ^ 11 ^ 22 <sup>h</sup>\ 2<sup>h</sup>22 '[0101] It can be shown that this solution fully satisfies the constraints mentioned above (| / Λ 1 | = \ P22 \ = 1 and | p12 | = \ pn \ = 0) for specific cases W1 = W2 = 0 and W1 = W2 = 1.
[0102] Fig. 9 shows the module of the main expression (10log10 \ pn \ 2) of a transfer matrix for this solution. Fig. 10 shows the phase angle pn and Fig. 11 crosstalk (10log10 \ p21 \ 2).
[0103] In particular, Fig. 9 shows the deviation in dB of the main module expression pn of the matrix with respect to the ideal value \ p11 \ = 1 in the functions W1 and W2. It can be seen that the maximum deviation from the ideal case is less than 1 dB. Fig. 10 shows the angle pn as a function of W1 and W2. As can be expected in the case of a difference from the ideal case with complex values, the phase differences reach 90 degrees. Fig. 11 shows a matrix crosstalk matrix expression module P21 measured in dB as a function of W1 and W2 weights. It should be noted that other elements of the transfer matrix can be obtained by exchanging wt and w2.
[0104] In some embodiments, the array processor 809 may determine a decoding matrix G for the subband in response to a G = H sub band matrix P = G'H. In particular, the matrix processor may select coefficients in G in such a way that a given characteristic is obtained for P.
[0105] Again, since the phase values for the real-valued subbands have a rather low perceptual weighting, only the P-characteristic is taken into account by the example decoder 715. The matrix processor 809 can get a high-quality operation by selecting the coefficients of the decoding matrix such that the power measure pn and P21 fulfilled a condition such as for example that the power metric be minimized or that the power metric be below the given criterion. The matrix processor 809 may, for example, search over a wide range of possible real-valued coefficients and select those that lead to the lowest power measure for Pn and P21. In addition, the evaluation may be subject to other constraints, such as constraints, to make P11 and P22 substantially equal to unity (e.g., between 0.9 and 1.1).
[0106] In some embodiments, the matrix processor 809 may perform a mathematical algorithm for determining appropriate values of real-valued coefficients for a decoding approach. A specific such case is described below, the algorithm being aimed at minimizing total crosstalk: | p12 |<sup>2</sup> + | p21 |<sup>2</sup> with the restriction | p11 |<sup>2</sup> = 1 and | p22 |<sup>2</sup> = 1.
[0107] This problem can be solved by means of standard mathematical tools of multidimensional analysis. In particular, it is appropriate to use methods with a Lagrange multiplier, which, for each row of vector v in G, translates into the problem of the eigenvalue of the matrix in the form vA = λνΒ with the requirement of normalization q (v) = 1 given by the square form q. Matrices A and B and square characters q depend on the entries of the complex matrix H.
[0108] The following is the solution for v = [gn # 2]. It is also trivial to solve v = [g21 g22] by replacing the ma and w2 variables in the following solution. Lagrange arrays A and B are defined as:
73
V3 3 where q \ and qi are defined as:
l-2w, + 2w,<sup>2</sup> '
- 2 nd<sub>2</sub> + 2w2 [0109] Eigenvalues are found by det (A-ZB) = 0, which leads to the roots of the square polynomial:
<img file="PL1999747T3_D0015.tif" />
where
2a
<img file="PL1999747T3_D0016.tif" />
<img file="PL1999747T3_D0017.tif" />
Now you can designate two potential solutions:
(A-A<sub>2</sub>B) v<sub>at</sub> =<sup>0</sup> [0110] The final solution is given by v = ov /, where i is 1 or 2, so that \ pn t |<sup>2</sup> = 1 and with minimum crosstalk. First, it is calculated as:
<img file="PL1999747T3_D0018.tif" />
Next, the \ p12 \ 2 crosstalk is calculated for both solutions:
<img file="PL1999747T3_D0019.tif" />
[0111] The index i which produces the minimum crosstalk gives v = c · ν. Without additional proof it is stated that regardless of the variables W1 and w2, the index i is always equal to 2.
[0112] To ensure completeness, the complete solution for G in terms of analytical equations is given below. The following variables are defined:
~ in,<sup>2</sup>
1-2vv, + 2w,<sup>2</sup>' <sup>q2</sup> l-2w<sub>2</sub> + 2w<sub>2</sub>'= q, +? 2>
<img file="PL1999747T3_D0020.tif" />
Then, the variable b is calculated as:
b = 1 - 5 p - j-Mp<sup>2</sup> + (47-14) 7 + 1,
The two root elements ra and the poems of both rows of matrix G are calculated as:
<img file="PL1999747T3_D0021.tif" />
[0113] The determination may then be unscaled Vtemp, 1 and Vtemp, 2 as:
^ temp, 1.1 <sup>1</sup>
V, temp, 2.2 <sup>v</sup>temp, 2, \ <- <2<sup>r</sup>p
The normalization constants c are calculated as:
<img file="PL1999747T3_D0022.tif" />
And finally, matrix G is given by:
G = <sup>C</sup>1 'temp) ^ 2 V temp, 2 [0114] Figures 12, 13 and 14 illustrate the operation for this solution. Fig. 12 shows the deviation in dB of the main module expression pn 1 of the matrix from the ideal value \ pn 1 | = 1 in the function mi w2. It can be noticed that due to the constraints imposed for this solution, the module is always identical to the ideal value p 1 \ = 1.
[0115] Fig. 13 shows the angle pn 1 as a function of wi and W2. It should be noted that due to the crosstalk imposed by all real solutions, also in this case the phase differences reach 90 degrees.
[0116] Fig. 14 shows the module of the expression p1 of the crosstalk matrix, measured in dB as a function of w and w wattages.
[0117] As shown in Fig., The solution of bringing the decoding matrix coefficients to absolute values of reverse coding matrix coefficients deviates only +/- 1 dB from the more complicated crosstalk minimization approach, both in terms of enhancing the main expression and crosstalk suppression.
[0118] FIG. 15 illustrates a method for audio decoding in accordance with some embodiments of the invention.
[0119] In step 1501, the decoder receives input data including a N-channel signal corresponding to a downmix signal from the M-channel audio signal, M> N having subband banding matrixes applied to the frequency subbands and parametric multi-channel data associated with the downmix signal.
[0120] After step 1501, step 1503 follows in which the subbands are generated for the N-channel signal. At least some of the frequency subbands are real frequency subbands.
[0121] Step 1503 follows step 1505, in which the real-time decode arrays for the compensation of the use of the coding matrix are determined in response to the parametric multi-channel data.
[0122] Step 1505 follows step 1507, in which the downmix data corresponding to the downmix signal is generated by matrix multiplication of the real-time decode subbands and N-channel data data on at least some of the real-frequency subbands.
[0123] It should be noted that the above description for purposes of clarity has provided embodiments of the invention with respect to various functional units and processors. However, it is evident that any appropriate functional separation between different functional units or processors may be used without departing from the invention. For example, illustrated functionalities to be implemented by separate processors or controllers may be implemented by the same processor or drivers. Hence, references to specific functional entities should be seen only as references to appropriate means to provide described functionalities, rather than as indicating a rigid logical or physical structure or organization.
[0124] The invention may be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. The invention may optionally be implemented at least in part as computer software running on one or more data processors and / or digital signal processors. The elements or components of an embodiment of the invention may be implemented physically, functionally and logically in any suitable manner. In fact, the functionalities can be implemented in a single unit, in many units or as part of other functional units. As such, the invention may be implemented in a single unit or it may be physically and functionally distributed among different entities and processors.
[0125] Although the present invention has been described in connection with some embodiments, it should not be limited to the particular form presented herein. Rather, the scope of the present invention is limited only by the appended claims. In addition, although a feature may appear to be described in connection with specific embodiments, one skilled in the art will recognize that various features of the described embodiments may be combined in accordance with the invention. In the claims, the term "contain" does not exclude the presence of other elements or steps.
[0126] In addition, although they were mentioned individually, a plurality of means, elements or method steps may be implemented e.g. by a single unit or a single processor. In addition, although individual features may be included in the various claims, they may be potentially preferably combined, and the inclusion in various claims does not imply that the combination of features is not possible and / or advantageous. Also, the inclusion of a feature in one category of claims does not imply a limitation to this category, but rather indicates that the element may accordingly have equally good application in other categories of claims. In addition, the order of the features in the patent claims does not impose any specific order in which the features must work, in particular, the order of the individual steps in the method claim does not impose the necessity of performing the steps in this order. Rather, the steps can be implemented in any suitable order. In addition, the use of the singular does not exclude the plural. Accordingly, the use of a singular or the phrases "first", "second" ("a", "an", "first", "second"), etc. does not exclude the plural. The reference markings in the claims, provided only as an explanatory example, should not be construed as limiting the scope of the claims in any way. the use of a singular or the phrases "first", "second" ("a", "an", "first", "second"), etc. does not exclude the plural. The reference markings in the claims, provided only as an explanatory example, should not be construed as limiting the scope of the claims in any way. the use of a singular or the phrases "first", "second" ("a", "an", "first", "second"), etc. does not exclude the plural. The reference markings in the claims, provided only as an explanatory example, should not be construed as limiting the scope of the claims in any way.
Koninklijke Philips NV, Netherlands Dolby International AB, The Netherlands
Proxy:
EP 1 999 747 B1 Z-15202/16
Contents4
20 members in 11 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 06111916 | European Patent Office (EPO) | A | |
| 07735236 | European Patent Office (EPO) | A | |
| 06111916 | – | – | – |
| 077352367 | – | – | – |
| EP20060111916 | – | – | – |
| EP20070735236 | – | – | – |
Members20
| Document | Office | Kind | |
|---|---|---|---|
| WO2007110823A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200746046A | Taiwan Province of China | A | |
| KR20080105135A | Republic of Korea | A | |
| EP1999747A1 | European Patent Office (EPO) | A1 | |
| CN101484936A | China | A | |
| US2009240505A1 | United States of America | A1 | |
| JP2009536360A | Japan | A | |
| RU2008142752A | Russian Federation | A | |
| HK1135791A | Hong Kong, China | A | |
| KR101015037B1 | Republic of Korea | B1 | |
| RU2420814C2 | Russian Federation | C2 | |
| BRPI0709235A2 | Brazil | A2 | |
| CN101484936B | China | B | |
| JP5154538B2 | Japan | B2 | |
| US8433583B2 | United States of America | B2 | |
| TWI413108B | Taiwan Province of China | B | |
| EP1999747B1 | European Patent Office (EPO) | B1 | |
| PL1999747T3This record | Poland | T3 | |
| BRPI0709235B1 | Brazil | B1 | |
| BRPI0709235B8 | Brazil | B8 |
Numbers
- Publication
- 1999747
- Publication, DOCDB
- 1999747
- Publication, EPODOC
- PL1999747T
- Application
- 7735236
- Application, DOCDB
- 07735236
- Application, EPODOC
- PL20070735236T
Titles2
- English
- AUDIO DECODING
- Polish
- Dekodowanie audio
Classification
- CPC, 4
- H04S3/008
- G10L19/008
- G10L19/0208
- G10L25/18