Multi parametrisation based multi-channel reconstruction
Abstract
For a multi-channel reconstruction of audio signals based on at least one base channel, an energy measure is used for compensating energy losses due to an predictive upmix. The energy measure can be applied in the encoder or the decoder. Furthermore, a decorrelated signal is added to output channels generated by an energy-loss introducing upmix procedure. The energy of the decorrelated signal is smaller than or equal to an energy error introduced by the predictive upmix. Thus, problems occurring for prediction based up-mix methods such as up-mixing signals that are coded with High Frequency Reconstruction techniques are solved, so that the correct correlation between the up-mixed channels is obtained or the up-mix is adapted to arbitrary down-mixes.
Term
Term ended
Projected expiry passed 28 October 2025, 0.9 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1Zastrzeżenia patentowe 1. Wielokanałowy syntezator do generowania przynajmniej trzech wyj ściowych kanałów audio (1100) z wykorzystaniem sygnału wej ściowego zawieraj ącego przynajmniej jeden kanał bazowy (1102), gdzie kanał bazowy uzyskany jest z pierwotnego sygnału wielokanałowego (101, 102, 103), przy czym sygnał wej ściowy zawiera ponadto przynajmniej dwa różne parametry ”up-mixu” (1108) oraz wskazanie trybu działania urządzenia realizuj ącego proces ”up-mixu” (1005), informuj ące w pierwszym stanie, że wykorzystana ma być pierwsza reguła ”up-mixu” oraz informuj ące w drugim stanie, że wykorzystana ma być druga reguła ”up-mixu”, zawieraj ący urządzenie realizuj ące proces ”up-mixu” (1104), przetwarzaj ące co najmniej jeden kanał bazowy przy wykorzystaniu przynajmniej dwóch różnych parametrów ”up-mixu” (1108) w oparciu o pierwszą albo drugą regułę „up-mixu” w odpowiedzi na wskazanie trybu działania urządzenia realizuj ącego proces ”up-mixu” (1005), tak że uzyskiwane są co najmniej trzy kanały wyj ściowe, znamienny tym, że pierwsza reguła ”up-mixu” jest regułą ”up-mixu” predykcyjnego (109), zaś druga reguła ”up-mixu” jest regułą ”up-mixu” obejmuj ącą parametry ”up-mixu” zależne od energii (1003). 2. Wielokanałowy syntezator według zastrz. 1 znamienny tym, że urządzenie realizuj ące proces ”up-mixu” (1104) umożliwia podczas przeprowadzania ”up-mixu” obliczanie, w zależności od wskazania trybu działania urządzenia realizuj ącego proces ”up-mixu” (1005), parametrów dla pierwszej albo drugiej reguły ”up-mixu” z wykorzystaniem przynajmniej dwóch różnych parametrów ”up-mixu” (1108), w zależności od wskazania trybu działania urządzenia realizuj ącego proces ”up-mixu” (1005). 3. Wielokanałowy syntezator według zastrz. 1 albo 2 znamienny tym, że wskazanie trybu działania urządzenia realizującego proces ”up-mixu” (1005) wskazuje sygnalizację trybu działania urządzenia realizującego proces ”up-mixu” zależną od częstotliwości albo zależną od podzakresu pasma albo zależną od czasu albo zależną od ramki, oraz że urządzenie realizujące proces ”up-mixu” umożliwia przeprowadzenie ”up-mixu” przynajmniej jednego kanału bazowego z wykorzystaniem różnych reguł ”up-mixu” dla różnych zakresów częstotliwości lub przedziałów czasu, zgodnie ze wskazaniem trybu działania urządzenia realizuj ącego proces ”up-mixu” (1005). 4. Wielokanałowy syntezator według zastrz. 1 znamienny tym, że druga reguła ”upmixu” określona jest w następujący sposób:gdzie L to wartość energii lewego kanału wejściowego, gdzie C to wartość energii centralnego kanału wej ściowego, gdzie R to wartość energii prawego kanału wej ściowego i gdzie α to parametr określony przez „down-mix”. 5. Wielokanałowy syntezator według jednego z zastrz. 1 do 4 znamienny tym, że druga reguła ”up-mixu” zapewnia, że poddany procesowi „down-mixu” kanał prawy nie jest dodawany do poddanego procesowi „up-mixu” kanału lewego i odwrotnie. 6. Wielokanałowy syntezator według jednego z zastrz. 1 do 5 znamienny tym, że pierwsza reguła ”up-mixu” określona jest w oparciu o dopasowywanie kształtów fal pierwotnego sygnału wielokanałowego i kształtów fal sygnałów generowanych z wykorzystaniem pierwszej reguły ”up-mixu” . 7. Wielokanałowy syntezator według jednego z zastrz. 1 do 6 znamienny tym, że pierwsza albo druga reguła ”up-mixu” określona jest w następujący sposób: Λ(βί)' Z C 1) JjiCpCz) Z( c 2, c i), gdzie funkcje f1, f2, f3 to funkcje przesyłanych dwóch różnych parametrów ”upmixu” c1, c2, przy czym funkcje te określone są w następuj ący sposób: ./2(^¾) = θ gdzie α jest parametrem ze zbioru liczb rzeczywistych. 8. Wielokanałowy syntezator według jednego z zastrz. 1 do 7 znamienny tym, że obejmuje ponadto moduł SBR (1614) do rekonstrukcji pasma przynajmniej jednego kanału bazowego, nie zawartego w przesyłanym kanale bazowym przy wykorzystaniu części przynajmniej jednego kanału bazowego zawartego w sygnale wejściowym, przy czym syntezator wielokanałowy umożliwia zastosowanie drugiej reguły ”up-mixu” do rekonstruowanego pasma przynajmniej jednego kanału bazowego oraz zastosowanie pierwszej reguły ”up-mixu” do pasma kanału bazowego zawartego w sygnale wejściowym. 9. Wielokanałowy syntezator według zastrz. 8 znamienny tym, że wskazanie trybu działania urządzenia realizującego proces ”up-mixu” (1005) ma postać sygnalizacji SBR (1606) zawartej w sygnale wej ściowym. 10. Wielokanałowy syntezator według jednego z powyższych zastrzeżeń znamienny tym, że sygnał wejściowy zawiera miarę energii (1106) informującą o błędzie energii uzależnionym od reguły ”up-mixu” wprowadzaj ącej straty energii, oraz że urządzenie realizujące proces ”up-mixu” umożliwia zastosowanie reguły ”upmixu” wprowadzającej straty energii jako pierwszej albo drugiej reguły ”up-mixu” oraz generowanie przynajmniej trzech kanałów wyjściowych tak, że błąd energii jest przynajmniej częściowo kompensowany w oparciu o miarę energii. 11. Wielokanałowy syntezator według jednego z powyższych zastrzeżeń znamienny tym, że urządzenie realizuj ące proces ”up-mixu” umożliwia uzyskanie miary energii (1106) z sygnału wejściowego oraz wykorzystanie miary energii jako wskazania trybu działania urządzenia realizuj ącego proces ”up-mixu” (1005), tak że urządzenie realizujące proces ”up-mixu” może zastosować regułę ”up-mixu” wprowadzającą straty energii w odpowiedzi na obecność miary energii (1106) w sygnale wejściowym. 12. Wielokanałowy syntezator według zastrz. 11 znamienny tym, że miara energii informuje o stosunku energii rezultatu zastosowania ”up-mixu” przeprowadzonego z wykorzystaniem reguły ”up-mixu” wprowadzającej straty energii do energii pierwotnego sygnału wielokanałowego albo wskazuje stosunek różnicy energii do energii pierwotnego sygnału wielokanałowego albo wskazuje bezwzględną wartość błędu energii. 13. Wielokanałowy syntezator według jednego z powyższych zastrzeżeń znamienny tym, że urządzenie realizuj ące proces ”up-mixu” obejmuje kalkulator (1600) tworzący w odpowiedzi na wskazanie trybu działania urządzenia realizującego proces ”upmixu” (1005) macierz ”up-mixu” w oparciu o przynajmniej dwa parametry ”up-mixu” oraz informacje dotyczące reguły ”down-mixu” zastosowanej do utworzenia przynajmniej jednego kanału bazowego z pierwotnego sygnału wielokanałowego. 14. Wielokanałowy syntezator według jednego z zastrz. 10 do 13 znamienny tym, że urządzenie realizujące proces ”up-mixu” (1104) obejmuje ponadto urządzenie dekoreluj ące (501, 502, 503, 501', 503') do generowania zdekorelowanego sygnału z przynajmniej jednego kanału bazowego lub z sygnałów wyjściowych reguły ”up-mixu” wprowadzaj ącej straty energii, oraz że urządzenie realizuj ące proces ”up-mixu” umożliwia wykorzystanie zdekorelowanego sygnału w taki sposób, że energia zdekorelowanego sygnału w kanale wyj ściowym jest mniejsza lub równa błędowi energii, który może być uzyskany w oparciu o miarę energii. 15. Wielokanałowy syntezator według zastrz. 14 znamienny tym, że gdy energia zdekorelowanego sygnału jest mniejsza niż błąd energii, urządzenie realizujące proces ”up-mixu” umożliwia wzmocnienie sygnału generowanego z wykorzystaniem reguły ”upmixu” tak, że połączona energia wzmocnionego sygnału i dodanego sygnału zdekorelowanego jest równa energii sygnału pierwotnego. 16. Wielokanałowy syntezator według zastrz. 14 albo 15 znamienny tym, że energia dodanego sygnału zdekorelowanego określona jest przez współczynnik dekorelacji, przy czym wysoka wartość współczynnika dekorelacji zbliżona do wartości 1 oznacza, że dodana powinna być mniejsza ilość sygnału zdekorelowanego, podczas gdy mniejsza wartość współczynnika dekorelacji zbliżona do wartości 0 oznacza, że dodana powinna być większa ilość sygnału zdekorelowanego, gdzie miara dekorelacji uzyskiwana jest z sygnału wejściowego. 17. Wielokanałowy syntezator według jednego z powyższych zastrzeżeń znamienny tym, że sygnał wejściowy zawiera poza dwoma różnymi parametrami ”upmixu”, informacje związane z procesem ”down-mixu” odpowiedzialnym za przynajmniej jeden kanał bazowym, przy czym urządzenie realizuj ące proces ”up-mixu” umożliwia wykorzystanie dodatkowych informacji związanych z procesem ”down-mixu” w celu utworzenia macierzy ”up-mixu” (802). 18. Koder do przetwarzania wielokanałowego sygnału audio obejmuj ący: generator parametrów (104, 1001, 1520, 1522, 1414, 1416) do generowania konkretnych reprezentacji parametrycznych wybranych z wielu różnych reprezentacji parametrycznych w oparciu o informacje dostępne dla kodera, przy czym reprezentacja parametryczna użyteczna jest podczas procesu ”up-mixu” jednego lub wielu kanałów bazowych dla rekonstrukcji wielokanałowego sygnału wyjściowego, oraz interfejs wyjściowy (1408) przeznaczony do udostępniania na wyjściu wygenerowanej reprezentacji parametrycznej oraz informacji bezpośrednio lub w sposób ukryty wskazuj ącej konkretną reprezentację parametryczną wybraną z wielu różnych reprezentacji parametrycznych, znamienny tym, że wiele różnych reprezentacji parametrycznych obejmuje pierwszą reprezentację parametryczną odpowiadającą schematowi ”up-mixu” predykcyjnego opartego na dopasowywaniu kształtu fali (104) oraz drugą reprezentacj ę parametryczną odpowiadającą nie opartej na dopasowywaniu kształtu fali regule ”upmixu” zawieraj ącej parametry ”up-mixu” uzależnione od energii (1001). 19. Koder według zastrz. 18 znamienny tym, że reguła ”up-mixu” nie oparta na dopasowywaniu kształtu fali jest regułą ”up-mixu” zachowującą energię. 20. Koder według jednego z zastrz. 18 do 19 znamienny tym, że pierwsza reprezentacja parametryczna jest reprezentacją parametryczną, której parametry uzyskane zostały w wyniku procesu optymalizacji, zaś druga reprezentacja parametryczna uzyskana została przez obliczenie (1520) energii kanałów pierwotnych oraz przez obliczenie parametrów (1522) w oparciu o połączenia energii. 21. Koder według jednego z zastrz. 18 do 20 znamienny tym, że obejmuje ponadto moduł replikacyjnego poszerzania pasma (1512, 1514) do generowania informacji dotyczących replikacyjnego poszerzania bocznego pasma dla przynajmniej jednego pasma zawartego w pierwotnym sygnale wejściowym, które nie jest zawarte w kanale bazowym udostępnianym na wyj ściu przez koder, przy czym informacje dotyczące replikacyjnego poszerzania bocznego pasma w sposób ukryty wskazuj ą konkretną reprezentacj ę parametryczną. 22. Koder według jednego z zastrz. 18 do 21 znamienny tym, że obejmuje ponadto kalkulator miary energii (1402) do obliczania miary energii (ρ) uzależnionej od różnicy energii pomiędzy wielokanałowym sygnałem wejściowym lub przynajmniej jednym kanałem bazowym uzyskanym z wielokanałowego sygnału wejściowego, a poddanym procesowi „up-mixu” sygnałem, generowanym z wykorzystaniem procesu ”up-mixu” wprowadzaj ącego straty energii;przy czym interfejs wyjściowy (1408) umożliwia udostępnianie na wyjściu przynajmniej jednego kanału bazowego po jego skalowaniu (401, 402) z wykorzystaniem współczynnika skalowania (403) uzależnionego od miary energii, lub udostępnianie na wyjściu miary energii. 23. Koder według zastrz. 22 znamienny tym, że miara energii (ρ) udostępniana na wyjściu przez interfejs wyjściowy wykorzystywana jest do wskazywania w sposób ukryty konkretnej reprezentacji parametrycznej. 24. Koder według jednego z zastrz. 18 do 23 znamienny tym, że obejmuje ponadto kontroler reprezentacji parametrycznych do sterowania generatorem parametrów albo interfejsem wyjściowym dla określenia, która reprezentacja parametryczna wybrana z wielu różnych reprezentacji parametrycznych ma być generowana lub udostępniana na wyj ściu. 25. Koder według jednego z zastrz. 18 do 24 znamienny tym, że kontroler reprezentacji parametrycznych umożliwia określenie zdarzenia w koderze lub obliczenie funkcji docelowej. 26. Koder według zastrz. 25 znamienny tym, że zdarzenie w koderze jest obliczeniem informacji dotyczących replikacyjnego poszerzania pasma, dzięki czemu kontroler może sterować interfejsem wyjściowym w taki sposób, by udostępniał on na wyjściu drugą reprezentację parametryczną, odpowiadającą pasmu nie zawartemu w kanale bazowym, oraz by udostępniał on na wyjściu pierwszą reprezentację parametryczną odpowiadającą pasmu zawartemu w kanale bazowym. 27. Koder według jednego z zastrz. 18 do 25 znamienny tym, że kontroler reprezentacji parametrycznych umożliwia wykorzystywanie w funkcji docelowej, wartości lub kombinacji wartości uzyskanych w oparciu o jakość ”up-mixu” , przepływność ”downmixu”, możliwości obliczeniowe kodera lub dekodera albo pobór energii urządzeń zasilanych bateryjnie, przy czym funkcja docelowa informuje, że dla pewnego podzakresu pasma lub ramki pierwsza parametryzacja jest lepsza niż druga parametryzacja. 28. Koder według dowolnego z zastrz. znamienny tym, że interfejs wyjściowy umożliwia udostępnianie na wyjściu różnych reprezentacji parametrycznych dla różnych zakresów częstotliwości lub przedziałów czasu. 29. Koder według dowolnego z zastrz. 18 do 28 znamienny tym, że obejmuje ponadto kalkulator miary energii do obliczania miary energii w oparciu o stosunek energii sygnału poddanego procesowi „up-mixu”, generowanego w procesie ”up-mixu” z przynajmniej jednego kanału bazowego z wykorzystaniem reguły ”up-mixu” wprowadzającej straty energii oraz energii pierwotnego sygnału wielokanałowego. 30. Koder według dowolnego z zastrz. 18 do 29 znamienny tym, że obejmuje ponadto urządzenie realizuj ące proces ”down-mixu” (1410) do obliczania przynajmniej jednego kanału bazowego, przy czym interfejs wyjściowy (1408) umożliwia udostępnianie na wyj ściu przynajmniej jednego kanału bazowego. 31. Sposób generowania przynajmniej trzech wyjściowych kanałów audio (1100) z wykorzystaniem sygnału wej ściowego zawieraj ącego przynajmniej jeden kanał bazowy (1102), gdzie kanał bazowy uzyskiwany jest z pierwotnego sygnału wielokanałowego (101, 102, 103), przy czym sygnał wyjściowy zawiera ponadto przynajmniej dwa różne parametry ”up-mixu” (1108) oraz wskazanie trybu działania urządzenia realizującego proces ”up-mixu” (1005) informujące w pierwszym stanie, że wykorzystana ma być pierwsza reguła ”up-mixu” oraz informujące w drugim stanie, że wykorzystana ma być druga reguła ”up-mixu”, obejmuj ący: przeprowadzenie procesu ”up-mixu” (1104) przynajmniej jednego kanału bazowego z wykorzystaniem przynajmniej dwóch różnych parametrów ”up-mixu” (1108) w oparciu o pierwszą lub drugą regułę ”up-mixu” zależnie od wskazania trybu działania urządzenia realizującego proces ”up-mixu” (1005), tak, że uzyskiwane są przynajmniej trzy kanały wyj ściowe, znamienny tym, że pierwsza reguła ”up-mixu” jest regułą ”up-mixu” predykcyjnego (109), zaś druga reguła ”up-mixu” jest regułą ”up-mixu” obejmuj ącą parametry ”up-mixu” zależne od energii (1003). 32. Sposób przetwarzania wejściowego wielokanałowego sygnału audio obejmujący: generowanie (104, 1001, 1520, 1522, 1414, 1416) konkretnych reprezentacji parametrycznych wybranych z wielu różnych reprezentacji parametrycznych w oparciu o informacje dostępne dla kodera, przy czym reprezentacja parametryczna użyteczna jest podczas procesu ”up-mixu” jednego lub wielu kanałów bazowych w celu rekonstrukcji wielokanałowego sygnału wyj ściowego, oraz udostępnianie na wyjściu (1408) wygenerowanej reprezentacji parametrycznej oraz informacji bezpośrednio lub w sposób ukryty wskazującej konkretną reprezentację parametryczną wybraną z wielu różnych reprezentacji parametrycznych, znamienny tym, że wiele różnych reprezentacji parametrycznych obejmuje pierwszą reprezentację parametryczną odpowiadającą schematowi ”up-mixu” predykcyjnego opartego na dopasowywaniu kształtu fali (104) oraz drugą reprezentacj ę parametryczną odpowiadaj ącą nie opartej na dopasowywaniu kształtu fali regule ”upmixu” zawierającej parametry ”up-mixu” uzależnione od energii (1001). 33. Zakodowany wielokanałowy sygnał informacyjny audio zawieraj ący konkretną reprezentacj ę parametryczną wybraną z wielu różnych reprezentacji parametrycznych, przy czym reprezentacja parametryczna użyteczna jest podczas procesu ”up-mixu” jednego lub wielu kanałów bazowych dla rekonstrukcji wielokanałowego sygnału wyj ściowego oraz zawierający informację bezpośrednio lub w sposób ukryty wskazuj ącą konkretną reprezentacj ę parametryczną wybraną z wielu różnych reprezentacji parametrycznych znamienny tym, że wiele różnych reprezentacji parametrycznych obejmuje pierwszą reprezentację parametryczną odpowiadającą schematowi ”up-mixu” predykcyjnego opartego na dopasowywaniu kształtu fali (104) oraz drugą reprezentacj ę parametryczną odpowiadaj ącą nie opartej na dopasowywaniu kształtu fali regule ”upmixu” zawierającej parametry ”up-mixu” uzależnione od energii (1001). 34. Nadaj ący się do odczytania za pomocą odpowiedniego urządzania nośnik, na którym przechowywany jest zakodowany wielokanałowy sygnał informacyjny określony w zastrz. 33. 35. Nadajnik lub rejestrator dźwięku zawieraj ący koder określony w dowolnym z zastrzeżeń od 18 do 30. 36. Odbiornik lub odtwarzacz dźwięku zawierający syntezator określony w dowolnym z zastrzeżeń od 1 do 17. 37. System przesyłowy obejmujący nadajnik określony w zastrz. 35 oraz odbiornik określony w zastrz. 36. 38. Sposób przesyłania lub rejestrowania dźwięku, obejmujący sposób określony w zastrz. 32. 39. Sposób odbierania lub odtwarzania dźwięku, obejmujący sposób określony w zastrz. 31. 40. Sposób odbierania zgodnie z zastrz. 39 i nadawania zgodnie z zastrz. 38. 41. Program komputerowy zawieraj ący komputerowy kod programu wykonuj ący, podczas uruchomienia programu na komputerze, wszystkie etapy sposobu określonego w dowolnym z zastrz. 31, 32, 38, 39 albo 40. CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK: EP 1 738 353 Z-4667/07 (stan techniki) Fig.1 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 Fig.2 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 202 9. =^1+^((1-^/^^1^ ź = 1, r, c Fig.3 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 Fig.4 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 ν = 1/7(1+2α 2 ) γ = ^1/(ρ 3 -1) 109 reguła ”up-mixu” nie zapewniająca zachowania energii dekorelacja sygnału(ów) poddanego(ych) procesowi ”down-mixu” dekorelacja sygnału(ów) przewidywanego(ych) Fig.5 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 Fig.6 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Fig. 7 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 801 urządzenie realizujące proces ”down-mixu” („down-mix” zmodyfikowany”) 8.02 informacje dotyczące "”down-mixu” zmodyfikowany kształt fali („down-mix” artystyczny”) urządzenie realizujące proces ”down-mixu” /kalkulator parametrów 107 108 104 Fig. 8 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 Fig. 9 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 104 (dla ”up-mixu” predykcyjnego) 1002 1004 109 Fig. 10 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 1102 (e. g. co najmniej jeden kanał bazowy 1104 1100 1106 miara energii URZĄDZENIE REALIZUJĄCE PROCES ”UP-MIXU” wykorzystujące macierz ”up-mixu” wprowadzającą straty energii {e. g. ρ.κ) ->| -*>r -►c przynajmniej trzy kanały wyjściowe (posiadające taką samą energię jak sygnał pierwotny) 1108 dwa różne parametry „up-mixu” (e. g. c 1v c a 0( c tl c 2 ) Fig. 11 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 przeprowadzana przez koder korekcja zgodna częściowo lub zachowania wynalazkiem energii dekoder Fig. 12 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 Fig. 13 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 401,402 Z-4667/07 miara energii Fig. 14a Fig. 14b CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 1502 1520 Z-4667/07 Fig. 15a 1005 σ Fig. 15b CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 1108 Z-4667/07 1102 A 'pełne pasmo z kanałem(ami) bazowym(i) informacje dot. energii 1612 1602 przesyłane parametry „up-mixu” (c 11 ,c 22 - oparte na kształcie fali) (c 1 , c 2 oparte na wartości energii) 1108 1106 miara energii do korekcji dolnej części pasma Fig. 16a D: sześć zmiennych (wcześniej określonych i dostępnych dla dekodera) C: - przesyłane dwa parametry (np. c11, c22) - cztery parametry (np. c12, c21, c31, c32) obliczone przez przedstawiony na fig. 16a kalkulator z wykorzystaniem czterech równań uzyskanych z powyższego równania macierzy (oparte na kształcie fali) Fig. 16b CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 Fig. 17 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK EP 1 738 353 Z-4667/07 Fig. 18 CODING TECHNOLOGIES AB, Szwecja KONINKLIJKE PHILIPS ELECTRONICS N.V., Holandia PEŁNOMOCNIK
257 paragraphs in 8 sections, as filed
TECHNICAL FIELD
The present invention relates to multi-channel audio signal reconstruction based on available stereo signal and additional control data.
TECHNICAL STATE
Recent progress in the field of coding audio signals has created the possibility of multi-channel reconstruction of audio signal form based on stereo (or mono) signal and the corresponding control data. These methods differ significantly from older matrix solutions such as Dolby Prologic, because the additional control data used to control the reconstruction - which is also referred to as the "up-mix" of the surround channels based on the transmitted mono or stereo channels, is sent here.
Parametric multi-channel audio decoders therefore reconstruct N channels based on M transmitted channels, where N> M, and additional control data. Additional control data requires a significantly lower data transmission speed than when sending additional N - M channels, which enables very efficient coding while maintaining compatibility with both Channel and N - channel devices.
These parametric surround sound coding methods usually include parameterization of the surround signal based on IID (Inter Chanel Intensity Diferrence) and ICC (Inter Chanel Coherence). These parameters describe the intensity and correlation coefficients between the channel pairs in the up-mix process. Other parameters also used in the current state of the art are predictive parameters used for prediction of intermediate or output channels during the "up-mix" process.
One of the most interesting applications of the prediction process known in the current state of the art is its use in a system that allows you to play sound in 5.1 format from two transmitted channels. In this configuration, stereo decoder is available on the decoder side, which is the result of a down-mix of multi-channel audio in 5.1 format. In this context, it is particularly interesting to be able to get the center channel sound as accurate as possible, because the center channel is usually connected in the down-mix process to both the left and right channels. This is done by estimating two predictive coefficients that determine how much of each of the two transmitted channels is used to form the central channel. These parameters are estimated for different frequency ranges, similar to the IID and ICC parameters described above.
However, since the predictive parameters do not describe the ratio of the intensity of the two signals but are based on the fitting of waveforms by the method of least squares, this method is characterized by sensitivity to any changes in the waveform of the stereo signal after calculating the predictive parameters.
The recent development of audio coding has led to the introduction of High Frequency Reconstruction methods, which are an extremely useful tool in audio codecs operating at low bit rates. One example of such methods is SBR (Spectral Band Replication) [WO 98/57436], which is used in standardized MPEG codecs such as MPEG-4 High Efficiency AAC. A common feature of these methods is that they reconstruct the decoder-side high frequencies from a narrowband signal encoded using the basic codec and a small amount of control information. As in the case of parametric reconstruction of multi-channel signals based on one or two channels, the amount of control data required for reconstruction of signal components (in the case of SBR these are high frequencies) is significantly smaller than the amount of data that would be necessary to encode the entire signal using a codec using waveform coding.
It should be understood, however, that the reconstructed broadband signal is perceptually identical to the original broadband signal, but the actual waveforms of the signals vary significantly. In addition, for codecs using the waveform coding used to encode low-bit stereo signals, pre-processing of the stereo signal is commonly used, which means that bandwidth limitation is performed in the mid / side representation of the stereo signal.
If multi-channel audio representation is desired based on a stereo signal encoded using MPEG-4 High Efficiency AAC (MPEG-4 High Efficiency AAC) or any other codec using high frequency reconstruction techniques, these and other characteristics of the codec used must be taken into account for coding a stereo signal in a down-mix process.
The article "Compatibility matrixing of multichannel bit-rate-reduced audio signals" (Ten Kate WR Th, Journal of the Audio Engineering Society, NY, US, volume 44, No. 12, December 1996, pages 1104-1119) discloses variable matrixing: for each time frame an optimal matrix is determined containing the minimum number of bits required.
What's more, it is common that for recordings available as a multi-channel audio signal, a separate stereo mix is also available, which is not a down-mix version of the multi-channel signal created automatically. He is commonly called the "artistic down-mix". Such a "down-mix" cannot be regarded as a linear combination of a multi-channel signal.
SUMMARY OF THE INVENTION
The object of the present invention is to provide the concept of a multi-channel down-mix encoder or an up-mix decoder that provides better quality of the reconstructed multi-channel output signal.
This object is met by providing a multi-channel synthesizer according to claim 1, an encoder processing a multi-channel input signal according to claim 18, a method for creating at least three output channels according to claim 31, a processing method according to claim 32 and an encoded multi-channel signal according to claim 33.
The present invention is based on the discovery that different parametric representations corresponding to different signal frequencies or time intervals are useful for encoding or decoding adapted to different situations. These situations may arise from encoder-related events such as performing SBR information calculation operations or calculating the energy level used to compensate for energy losses, as well as any other events. Other situations related to different parametric representations include up-mix quality, downmix bit rate, encoder or decoder computational capabilities, or, for example, the power consumption of battery-powered devices, which means that for some portion of the band or frame, the first parameterization is better than the second parameterization. Of course, the target function can also be a combination of different separate goals / events as described above.
Preferably, one parametric representation includes parameters corresponding to the predictive "up-mix" based on the modification of the waveform of the multi-channel down-mix signal. This is also the case when the down-mixed signal is coded using a codec that performs stereo pre-processing, high frequency reconstruction and uses other coding schemes that significantly change the waveform. The invention furthermore solves the problem that occurs when using predictive up-mix techniques for an artistic down-mix, i.e. when the signal obtained in the down-mix process is not automatically obtained from a multi-channel signal.
The present invention preferably has the following properties:
- estimation of prediction parameters is based on the modified waveform, not on the basis of the waveform of the down-mix signal;
- prediction based methods are used only in those frequency ranges for which it is beneficial;
- energy loss correction and imprecise correlation between channels introduced by the "up-mix" process based on prediction.
BRIEF DESCRIPTION OF THE DRAWINGS
The present invention will be described below based on non-limiting embodiments with reference to the accompanying drawings, in which:
Fig. 1 shows a prediction based reconstruction of three channels from two channels;
Fig. 2 shows predictive up-mix with energy compensation;
Fig. 3 shows energy compensation in the predictive up-mix process;
Fig. 4 shows a prediction parameter estimator in a down-mix signal energy encoder;
Fig. 5 shows a predictive up-mix with correlation reconstruction;
Fig. 6 shows a mixing module for mixing a de-correlated signal with an up-mix signal in an up-mix process with correlation reconstruction;
Fig. 7 shows an alternative solution of a mixing module designed to mix a signal correlated with an up-mix signal in an up-mix process with correlation reconstruction;
Fig. 8 shows estimation of a prediction parameter in an encoder;
Fig. 9 shows estimation of a prediction parameter in an encoder;
Fig. 10 shows a multi-parameter system according to the invention;
Fig. 11 shows an apparatus performing an up-mix process;
Fig. 12 is an energy graph illustrating the results of an up-mix process with introduced energy losses and favorable compensation;
Fig. 13 shows a table of energy compensation methods;
Fig. 14a is a schematic of a preferred embodiment of a multi-channel encoder;
Fig. 14b is a flowchart of the method implemented by the device of Fig. 14a;
Fig. 15a shows a multi-channel encoder that differs from the device shown in Fig. 14a by the function of spectral band replication to generate different parameterizations;
Fig. 15b is a table illustrating the generation and transmission of frequency-dependent parameterization data; and
Fig. 16a is a decoder illustrating the calculation of the up-mix process matrix coefficients;
Fig. 16b shows a detailed description of the calculation of the predictive up-mix parameters;
Fig. 17 shows a transmitter and receiver of a transmission system; and
Fig. 18 shows a sound recorder having an encoder and a sound player having a decoder.
DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS OF THE INVENTION
The invention described below merely illustrates the principles of the present invention. It will be appreciated that those skilled in the art will recognize the possibility of introducing modifications and variations of the details of the invention described herein. The invention is therefore limited only by the scope of the appended claims, and not by the specific details provided herein only to describe and explain particular embodiments of the invention.
It should be noted that the calculation of subsequent parameters, applications, the "upmix" process, the "down-mix" process or any other activities can be carried out based on selected frequency bands, ie for certain parts of the band included in the filter bank.
In order to highlight the advantages of the present invention, a more detailed description of the predictive up-mix known in the prior art will first be provided. Let's assume a three-channel up-mix process based on two channels subjected to the down-mix process, as shown in Fig. 1, where reference 101 is the primary left channel, reference 102 is the primary central channel, reference 103 is the primary right channel, reference 104 is the down-mix device and the extraction of parameters forming part of the encoder, references 105 and 106 are predictive parameters, reference 107 means the left channel subjected to the downmix process, reference number 108 means the right channel subjected to the downmix process, reference number 109 means the predictive up-mix module, and references 110, 11 and 112 denote respectively reconstructed left, middle and central channels.
Let's assume the following definitions, where X is a 3 x L matrix containing as lines three signal segments l (k), r (k), c (k), k = 0, L-1.
Similarly, let two signals l0 (k), r0 (k) subjected to the down-mix process form X0 lines. The down-mix process is described by the equation
X0 = DX (1) where the down-mix matrix is described as
<img file="PL1738353T3_D0001.tif" />
The preferred choice of down-mix matrix is
».-G i:] which means that the down-mix left channel l0 (k) contains only l (k) and ac (k) and r<sub>0</sub>(k) contains only r (k) and ac (k). Such a down-mix matrix is advantageous because it assigns the same central channel content to the left and right down-mix channels, and because it does not assign any part of the right primary signal to the down-mix left signal and vice versa.
"Up-mix" is described by the equation
X = CX<sub>0</sub> (4) where C is a 3 x 2 up-mix matrix.
The up-mix predictive known in the current state of the art is based on the solution of an over definite system
CX0 = X (5) due to C least squares method. This leads to normal equations
CX0X0 * = XX0 * (6)
Multiplying (6) from the left by D gives us DCX0X * 0 = X0X * 0, which in the general case, where X0X * 0 = DXX * D * is impersonal
DC = I2 (7)
Where In is a unitary matrix n. This equation reduces the C parameter space to two dimensions.
With the above data, the up-mix matrix
<img file="PL1738353T3_D0002.tif" />
it can be completely determined by the decoder if the "down-mix" D matrix is known and two elements of the C matrix are transmitted, for example c11 and c22.
The residual (prediction error) signals are described by the equation
<img file="PL1738353T3_D0003.tif" />
Multiplying from the left by D leads to
DXr = (D-DCD) X = 0 (9) due to (7). It follows that there exists a 1 x L line vector signal xr, such that (10)
Xr = vx
Where v is a 3 x 1 unit vector connecting the nucleus (empty space) D. For example, for "down-mix" (3)
-and
-a (11)
<img file="PL1738353T3_D0004.tif" />
Generally, when v = [vi, v<sub>r</sub>, Vc]<sup>T</sup> and X ['/ (k). r (k). CFW]<sup>T</sup> this means that for the weight vector the residue is common to all three channels, / (A-) = / (A) + v<sub>;</sub>x<sub>r</sub>(*) r (fc) = r (*) + v, x, (*) c (fr) = ć (A :) + Vc<sub>r</sub>(fc) (12)
Due to the principle of orthogonality, the residual signal xr (k) is orthogonal to all three predicted signals
Solved problems and improvements introduced by preferred embodiments of the present invention
When using the prediction-based up-mix described above in the prior art, the following problems are obvious:
- This method is based on matching waveforms, determining the error by the method of least squares, which is not possible in systems in which the waveform of the signal subjected to the "down-mix" process is not preserved.
- This method does not provide the correct correlation structure between the reconstructed channels (which will be described below).
- This method does not allow reconstruction of the right amount of energy in the reconstructed channels.
Energy Compensation
As noted above, one of the problems associated with prediction-based multi-channel signal reconstruction is that the prediction error corresponds to the energy losses of the three reconstructed channels. Below, the theoretical basis for such an energy loss is described and a solution to the problem provided by the preferred embodiments of the invention is provided. First, the theoretical problem analysis is described below, followed by the preferred embodiment of the present invention according to the theory presented.
Let E, E and E<sub>r</sub> will be the sum of the energies of the original signals in X, respectively
Λ predicted signals in X and prediction errors in X signals<sub>r</sub>. It follows from the principle of orthogonality that
<img file="PL1738353T3_D0005.tif" />
Total prediction gain can be described as E
<img file="PL1738353T3_D0006.tif" />
but it is more convenient to consider the parameter
<img file="PL1738353T3_D0007.tif" />
Consequently, ρ<sup>2</sup> e [0, 1] is a measure of the total relative energy of the predictive up-mix.
Given the ρ data, it is possible to correct each of the channels by using compensating gain so that || źj - H for z = l, r, c. In particular, the target energy is described by the formula (12),
<img file="PL1738353T3_D0008.tif" />
so the equation must be solved
<img file="PL1738353T3_D0009.tif" />
Because v is a unit vector
<img file="PL1738353T3_D0010.tif" />
by definition ρ (14) and equation 13 we get
<img file="PL1738353T3_D0011.tif" />
Taking into account all equations, we get the gain ι / z (, il-p<sup>2</sup> & ll + v? <sup>μ</sup> (19) g, = r><sup>2</sup> | Κ ·
It is clear that with this method, apart from transmitting ρ, the energy distribution in the decoded channels must be calculated by the decoder. In addition, only energies are correctly reconstructed, while the extragonal correlation structure is bypassed.
It is possible to obtain a gain value that ensures the conservation of total energy, but does not ensure that the energies of individual channels are correct. Common gain for all channels gz = g ensuring the conservation of total energy is obtained by determining the equation g<sup>2</sup>E = E. We receive
<img file="PL1738353T3_D0012.tif" />
Due to linearity, this gain can be used in the encoder to process "down-mixed" signals, so there is no need to send an additional parameter.
Fig. 2 shows a preferred embodiment of the present invention which allows reconstruction of three channels while maintaining the correct energy of the output channels. The l0 and r0 signals subjected to the down-mix process are introduced into the module implementing the up-mix process 201 together with the c1 and c2 predictive parameters. The module implementing the "up-mix" process reproduces the "up-mix" C matrix based on the "downmix" D matrix and the obtained predictive parameters. The three output channels from module 201 are input to module 202 together with the correction parameter ρ. The gain of the three channels is corrected as a function of the transmitted parameter ρ, after which the corrected energy channels are output.
Fig. 3 shows the implementation of the correction module 202 in more detail. Three up-mix channels are fed into the correction module 304, as well as to modules 301, 302 and 303, respectively. The energy estimating modules 301 - 303 estimate the energy of three channels upmix process, after which the measured energy is introduced into the correction module 304. The control signal received from the encoder 304 (corresponding to the prediction gain) is also introduced into the module 304. The equation (19) described above is implemented in the correction module.
In an alternative implementation of the present invention, the energy correction may be performed in the encoder. Fig. 4 shows the implementation of the encoder in which the gain of the down-mixed signals l0 107 and r0 108 is corrected in modules 401 and 402 according to the gain value calculated in module 403. The gain value is obtained using the equation described above ( twenty). As described above, this is an advantage of this embodiment of the present invention because it is not necessary to calculate the energy of the three reconstructed channels based on the predictive up-mix. However, it is only ensured here that the total energy of the reconstructed channels is correct. The correct energy for each channel is not provided here.
Below the down-mix device in Fig. 4 is an example of a preferred down-mix matrix corresponding to equation (3). In the device implementing the "down-mix" process, however, any "down-mix" matrix can be used, which is described by the equation (2).
As it will be presented later in this description, in this case a device performing the "down-mix" process having three channels as input, and as two channels as output, at least two additional "upmix" parameters c1, c2 are required. If the D-down matrix D is mixed or not fully available to the decoder, it is necessary to send additional information from the encoder to the decoder in addition to parameters 105 and 106.
Correlation structure
One of the problems associated with the up-mix process known in the art is that it does not allow reconstruction of the correct correlation between the reconstructed channels. The reason is that, as described above, the central channel is predicted as a linear combination of the down-mix left channel and the down-mix right channel, while the left and right channels are reconstructed by subtracting the provided central channel from down-mix left and right channels. It is clear that the prediction error results in the remainder of the original central channel in the intended left and right channel. As a result, the correlations between the three channels are not the same in the case of reconstructed channels as in the original three channels.
It is apparent from the preferred embodiment of the invention that the predicted three y-channels should be combined with de-correlated signals, taking into account the measured prediction error.
The theoretical foundations for obtaining the correct correlation structure will be described below. The specific residue structure can be used to reconstruct the full 3 x 3 XX * structure by replacing the residue in the decoder with a de-correlated signal xd.
First, note that normal equations (6) lead to XrX0<sup>*</sup> = 0, so
<img file="PL1738353T3_D0013.tif" />
Hence because
<img file="PL1738353T3_D0014.tif" />
Where equations (10) and 17 have been substituted for the last equation.
Let x<sub>d</sub> will be a signal de-correlated from all decoded ones
Λ Λ Λ signals so that
Xx '= 0.
An enhanced signal
<img file="PL1738353T3_D0015.tif" />
has a correlation matrix
YY '= XX * + vv || jc, ||<sup>2</sup> (24)
For complete reconstruction of the primary correlation matrix (22) it is sufficient
<img file="PL1738353T3_D0016.tif" />
If xd is obtained by de-correlating the down-mix signal, say
- Vo <sup>+ years</sup>a) 'and then using the γ gain, the equation should be met
<img file="PL1738353T3_D0017.tif" />
This gain can be calculated by the encoder. However, if the better specified parameter ρ can be used<sup>2</sup> e [0, 1] from equation (14), estimation E and must be carried out in the decoder.
In view of the above, a more preferred alternative is to generate xd using three decorrelators
<img file="PL1738353T3_D0018.tif" />
because then || x<sub>d</sub>||<sup>2</sup> = γ<sup>2</sup> E, due to which equation (25) is satisfied by the selection
<img file="PL1738353T3_D0019.tif" />
Fig. 5 shows one embodiment of the present invention enabling predictive up-mix of three channels from two downmixed channels while maintaining the correct correlation structure between the channels. The modules 109, 110, 111 and 112 in Fig. 5 are the same as in Fig. 1, and will no longer be described. Three "down-mix" signals from module 109 are introduced into decorrelators 501, 502 and 503. They generate mutually de-correlated signals. The de-correlated signals are added together and introduced into the mixing modules 504, 505 and 506, where they are mixed with signals from module 109. Mixing signals subjected to the predictive up-mix process with their de-correlated form is an essential feature of the present invention. Fig. 6 shows one embodiment of mixing modules 504, 505 and 506. In this embodiment, the level of the de-correlated signal is corrected in module 601 based on the control signal γ. The de-correlated signal is then added together in module 602 with the signal subjected to the predictive up-mix process.
In a third preferred embodiment of the invention, the upmixed channels are processed in decorrelators 501, 502 and 503. The de-correlated signal can also be generated by the decorrelator 501 ', which receives as input signal a channel subjected to the predictive down-mix process or even all channels subjected to the "down-mix" process. In addition, in the case of more than one channel subjected to the "down-mix", as shown in Fig. 5, the decorrelation signal may also be generated by separate decorrelators for the left base channel 10 and the right base channel r0, and a combination of output signals from these separate decorrelators. This function is basically the same as the function shown in Fig. 5, but differs from the function shown in Fig. 5 in that the base channels are used before the up-mix process.
In addition, it has been noted in the description of Fig. 5 that mixing modules 504, 505 and
506 they not only receive the coefficient γ, which is the same for all three channels, because this coefficient depends only on the energy measure ρ, but also receive the coefficients vl, vc and vr characteristic for individual channels, which are determined and described using equations (10 ) and (11). However, this parameter need not be sent from the encoder to the decoder if the decoder knows what down-mix matrix is being used in the encoder. Instead of such a solution, the parameters in the matrix v, as shown in equations (10) and (11) are preferably pre-programmed in mixing modules 504, 505 and 506, so that the weight coefficients characteristic for individual channels do not have to be sent (if required, they can of course be sent).
It can be seen in Fig. 6 that the correction device 601 corrects the energy of the de-correlated signal using the product of γ and the channel-specific coefficient vz, where z is l, r or c. In this context, it has been noticed that equation (26a) ensures that energy xd is equal to the sum of energy subjected to the up-mix process prediction of left, right and center channels. The device 601 can therefore be implemented simply as a calculator using the GI scaling factor. However, if the de-correlated signal is generated in an alternative way, the mixing modules 504, 505 and 506 must perform the correction of the absolute energy of the de-correlated signal added by the summing device 602, so that the energy of the signal added by the summing device 602 is equal to the energy of the residual signal, for example energy that is lost in the predictive up-mix process that does not ensure energy conservation.
As for the down-mix coefficient vz characteristic for a given channel, vz, the comments above with reference to Fig. 6 also apply to the embodiment of Fig. 7.
It should also be noted that the embodiments of Figures 6 and 7 are based on the finding that at least part of the energy lost in the predictive up-mix process is added using a de-correlation signal. To obtain the correct signal energy and the correct proportion of "dry" signal component (uncorrelated) and "wet" signal component (decorrelated), make sure that the "dry" signal input to the 504 mixing module is not pre-scaled. For example, if the base channels have been pre-corrected in the decoder (as shown in Fig. 4), the initial correction shown in Fig. 4 must be compensated by multiplying the channel by the (relative) energy measure ρ before entering the signal into the mixing module 504, 505 or 506. In addition, the same process must be carried out when such energy correction was performed in the decoder before the input of signals subjected to the "down-mix" to the device performing the up-mix predictive 109, as shown in Fig. 5.
If only a portion of the residual energy is to be covered by a de-correlated signal, the pre-correction must only be partially removed by pre-scaling the signal fed into the mixing module 504, 505 or 506 using a factor depending on ρ, which is, however, closer to unity than the factor itself ρ. Of course, this precompensation factor for pre-scaling depends on the signal generated by the encoder κ input to 605 shown in Fig. 7. If it is necessary to carry out this partial pre-scaling process, the weight factor used in G2 is not necessary. In this case, the branch from input 604 to summing device 602 is the same as in Fig. 6.
Control of the degree of decorrelation
In a preferred embodiment of the invention, it is disclosed that the amount of de-correlation introduced into the predicted signals subjected to the "up-mix" process can be controlled by the encoder while maintaining the correct output energy at all times. The reason for this fact is that in the typical example of a "conversation" containing dry speech in the central channel and ambient sounds in the left and right channel, the replacement of the prediction error in the central channel with a de-correlated signal may be undesirable.
In a preferred embodiment of the present invention, an alternative to the one illustrated in Fig. 5 may be used. The following will describe how the present invention allows separation of the issues of total energy conservation and proper correlation reconstruction, as well as control using the K parameter of the degree of de-correlation being introduced.
Let us assume that the signal subjected to the down-mix process has been subjected to the process of gain compensation in order to save energy (20), thanks to which χ, decoded signal is obtained. On this basis, a de-correlated signal is generated to the same total energy || d ||<sup>2</sup> = E / ρ<sup>2</sup>, for example using three decorrelators as described in the previous chapter. The total "up-mix" is then determined by the following equation
<img file="PL1738353T3_D0020.tif" />
where κ e [ρ, 1] is the transmitted parameter. κ = 1 corresponds to the conservation of total energy without adding a de-correlated signal, while. κ = ρ corresponds to the full reconstruction of the 3 x 3 correlation structure
<img file="PL1738353T3_D0021.tif" />
so that the total energy is saved for each κ e [ρ, 1], as it is visible after converting the traces (sum of diagonal values) into the matrix in equation (30). The correct energies of individual channels are obtained, however, only if κ = ρ.
Fig. 7 shows the embodiment of the mixing modules 504, 505 and 506 visible in Fig. 5 according to the theoretical principles described above. In this solution of mixing modules, the control parameter γ is introduced to modules 702 and 701. The gain factor used in module 702 corresponds to κ according to the above equation (29), and the gain factor used in module 701 corresponds to the above equation (29)
The above-described embodiment of the present invention allows the system to use a detection mechanism in the encoder that allows estimation of the degree of decorrelation added in the predictive up-mix process. The implementation described with reference to fig. 7 uses the addition of the indicated amount of de-correlated signal and energy correction, thanks to which the total energy of three channels is correct, while it is still possible to replace any part of the prediction error with a de-correlated signal.
This means that, for example, in the case of three signals containing ambient sounds, which is the case, for example, in classical music where there is a lot of ambient sounds, the encoder can detect the lack of a "dry" center channel and allow the decoder to replace the entire prediction error with a de-correlated signal, thanks to which the sounds the surroundings are reconstructed in three channels in a way which cannot be carried out using only prediction-based methods known in the art. In addition, in the case of a signal containing a "dry" central channel, for example speech in the central channel and ambient sounds in the left and right channels, the encoder detects that replacing the prediction error with a de-correlated signal is not appropriate from a psychoacoustic point of view and instead allows the decoder to correct ę levels of three reconstructed channels, so that the energy of the three channels is correct. Of course, these extreme cases described above correspond to the two possible results of the invention. The invention is not limited to the extreme examples described above.
Adaptation of predictive coefficients to modified waveform signals
As described above, predictive parameters are estimated by minimizing the mean square error for the original three X channels and the "down-mix" D matrix. In many cases, it cannot be assumed that the down-mix signal can be described as product of the "down-mix" matrix D and matrix X describing the original multi-channel signal.
One obvious example of this situation is the use of the so-called "artistic down-mix", where the two-channel "down-mix" cannot be described as a linear combination of a multi-channel signal. Another example is a situation in which a "down-mix" signal is coded using a perceptual audio codec that uses pre-processing of the stereo signal or other tools that increase the codec's performance. It is well known that, in the current state of the art, many perceptual audio codecs are based on mid / side stereo coding, where the side signal is limited by bit rate, resulting in an output signal with a narrower stereo image than it is space for the signal used for encoding.
Fig. 8 illustrates a preferred embodiment of the present invention in which the extraction of parameters from the multi-channel signal by the encoder also includes access to the modified down-mix signal. The modified "down-mix" is implemented here by the 801 device. If only two C matrix parameters are transmitted, the decoder must know the D matrix to ensure that the up-mix process can be performed and to obtain the lowest mean square error value for all up-mix channels. In the present embodiment, however, it has been disclosed that it is possible to replace the down-mixed 10 and r0 signals in a down-mix encoder with the 1'0 and r'0 signals obtained using the D matrix, which is not necessarily the same as assumed by the decoder. The use of an alternative "down-mix" to estimate the parameters in the encoder only guarantees the correct reconstruction of the central channel in the decoder. Transmission of additional information from the encoder to the decoder allows for a more correct "up-mix" of three channels. In the extreme case, all six elements of the C matrix can be transmitted. In the present embodiment, however, it has been disclosed that only part of the C matrix can be transmitted if it is accompanied by information about the D matrix used by the 802 module.
As noted earlier, perceptual audio codecs for low bit rate stereo coding use mid / side coding. In addition, pre-processing of the stereo signal is commonly used to limit the side-signal energy due to bitrate constraints. This is done on the basis of a psychoacoustic model that assumes that when limiting the width of a stereo signal, the presence of coding artifacts is better than audible distortions associated with quantization and limiting bandwidth.
So, in the case of pre-processing of the stereo signal, the downmix equation (3) can be expressed as
<img file="PL1738353T3_D0022.tif" />
where γ is the side attenuation. As noted earlier, in order to be able to reconstruct the three channels, the decoder must know the D matrix. Thus, in the present embodiment, an attenuation factor should be transmitted to the decoder.
Fig. 9 illustrates another embodiment of the present invention in which the down-mix signal 10 and r0 from device 104 is fed into a pre-processing module 901 that limits the side signal (10-r0) in the representation of the signal subjected to the " down-mix ”in the mid / side representation based on the γ coefficient. This parameter is sent to the decoder.
Parameterization of HFR codec signals ( high frequency reconstruction )
If the prediction-based up-mix is used in high frequency reconstruction methods such as SBR [WO 98/57436], the predictor parameters estimated by the encoder do not match the broadband signal reconstructed by the decoder. The present embodiment of the invention discloses the use of three channels from two channels for reconstruction, not based on the wave shape of the alternative "up-mix" structure. The proposed "up-mix" process was developed to reconstruct the correct energy of all channels subjected to the "upmix" process in the case of uncorrelated noise signals.
Let's assume that the down-mix matrix used is D<sub>and</sub> described by equation (3).
Let's also assume that we now define the "up-mix" matrix C. The "up-mix" is therefore described by the equation
<img file="PL1738353T3_D0023.tif" />
In the case of attempts to reconstruct the correct energy subjected to the "upmix" process of l (k), r (k) and c (k) signals, where energies are L, R and C, the "up-mix" matrix is chosen so that the diagonal elements and XX * are the same according to the equation
<img file="PL1738353T3_D0024.tif" />
The corresponding down-mix matrix equation is
<img file="PL1738353T3_D0025.tif" />
Choosing a diagonal element equal to a diagonal element
XX * leads to three equations describing the relationship between the elements in matrix C and L, R and C.
<img file="PL1738353T3_D0026.tif" />
Based on the above, an up-mix matrix can be created. It is preferable to specify an up-mix matrix that does not add the down-mix right channel to the up-mix left channel and vice versa. Therefore, the preferred up-mix matrix may be
C = 0 7
We get the following matrix C:
c =
<td>\ ~ Rp</td><td>Λ 0</td>
<td>\ L + a<sup>2</sup>C</td><td></td>
<td rowspan="2"> 0</td><td>rc</td>
<td>And VR +<sup>2</sup>C</td>
<td>And c</td><td> 1 <sup>c</sup></td>
<td>Y + R + 4a<sup>2</sup>C</td><td>\ L + R + 4a<sup>2</sup>CJ</td>
(39) (40)
<img file="PL1738353T3_D0027.tif" />
<img file="PL1738353T3_D0028.tif" />
It can be shown that C matrix elements can be reconstructed by the decoder based on two transmitted parameters ic, = -.
<sup>1</sup> Λ
Fig. 10 shows a preferred embodiment of the present invention. Elements 101 - 112 shown here are the same as in Fig. 1, and will no longer be described. The three original signals 101 - 103 are fed into the estimator 1001. This device estimates two parameters, e.g.
L <sup>and</sup> on the basis of which it is possible for the decoder to create a C matrix. These parameters together with the parameters from the device 104 are entered into the selection module 1002. In one preferred embodiment of the invention, the output of the selection module 1002 has parameters from device 104 available if these parameters correspond to a frequency range encoded using a codec based signal waveform analysis, and those from device 1001 if these parameters correspond to a range frequency reconstructed by HFR (high frequency reconstruction). At the output of the dialer 1002, information 1005 is also available to show which parameterization is used for each signal frequency range.
In the decoder, module 1004 receives the transmitted parameters and, depending on the indications contained in parameter 1005, directs them to the device implementing "up-mix" based on prediction 109 or the device implementing "up-mix" based on energy value 1003. In the module implementing "up -mix "based on the energy value of 1003, the" up-mix "C matrix has been implemented described by equation (40).
The up-mix matrix C described by equation (40) has the same weights (δ) to obtain the estimated (decoder) signal c (k) from two downmixed signals l0 (k) and r0 (k). Based on the observed fact that the relative amount of c (k) signal can be different in the two down-mix signals l0 (k) and r0 (k) (ie C / L is not equal to C / R), the following overall up-mix matrix may be considered:
<img file="PL1738353T3_D0029.tif" />
In order to estimate c (k) in this embodiment of the invention, it is also required to send two control parameters c1 and c2, which are equal, for example, c1 = a<sup>2</sup>C / (L + a<sup>2</sup>X) and c2 = a<sup>2</sup>X / (R + a<sup>2</sup>C). The possible implementation of the function f and up-mix matrix is then described by the equations:
<td> /(0,^)=71-¾<sup>2</sup></td><td> (42)</td>
<td> /(0,,¾) = 0</td><td> (43)</td>
<td> /(0,,¾)</td><td> (44)</td>
Signaling of different parameterization for the SBR range according to the present invention is not limited to SBR only. The parameterization described above can be used in any other frequency range in which the prediction based upmix prediction error is considered too large. Module 1002 can therefore output the parameters from device 1001 or 104, depending on many criteria, such as coding of transmitted signals, prediction error, etc.
A preferred method of improved prediction-based multi-channel reconstruction includes the extraction of various multi-channel parameterizations for different frequency ranges in the encoder and the application of these parameters in the decoder to the respective frequency ranges for the reconstruction of multi-channel sound.
Another preferred embodiment of the present invention includes a method of improved multichannel prediction-based reconstruction, comprising encoder extraction of information about the down-mix process used, followed by sending this information to the decoder, and carried out in the up-mix decoder based on the obtained predictive parameters and information about the "down-mix" that aims to reconstruct multi-channel sound.
Another preferred embodiment of the present invention includes a method of improved multichannel prediction based reconstruction in which the energy of the down-mix signal is corrected in the encoder according to the prediction error obtained for the obtained predictive up-mix parameters.
Another preferred embodiment of the present invention relates to a method of improved multichannel prediction based reconstruction in which compensation of energy losses associated with a prediction error is performed in the decoder, the compensation being a reinforcement of "up-mix" channels.
A further embodiment of the present invention relates to a method of improved multichannel prediction based reconstruction in which, in a decoder, energy losses associated with a prediction error are replaced by a de-correlated signal.
Another preferred embodiment of the present invention relates to a method of improved prediction-based multi-channel reconstruction in which, in the decoder, some of the energy losses associated with the prediction error are replaced by a de-correlated signal, and some of the energy losses are leveled by amplifying the channels under the "up-mix" process. This part of the energy lost is preferably determined by the encoder.
A further preferred embodiment of the present invention is a device intended for improved prediction-based multi-channel reconstruction and comprising elements correcting the signal subjected to the "down-mix" process in accordance with the prediction error obtained for the obtained predictive up-mix parameters.
A further preferred embodiment of the present invention is a device for improved prediction-based multi-channel reconstruction and comprising compensating elements for energy losses due to prediction error, wherein the compensation consists of amplifying the channels subjected to the "up-mix" process.
A further preferred embodiment of the present invention is a device intended for improved prediction-based multi-channel reconstruction and incorporating elements replacing energy losses associated with a prediction with a de-correlated signal.
Another preferred embodiment of the present invention is a device intended for improved multichannel prediction-based reconstruction and comprising elements replacing some of the energy losses associated with the prediction error by a correlated signal and eliminating some of the energy losses by strengthening the channels undergoing the "up-mix" process.
Another preferred embodiment of the present invention is an encoder for improved prediction-based multi-channel reconstruction performing correction of the down-mix signal in accordance with the prediction error obtained for the obtained predictive up-mix parameters.
Another preferred embodiment of the present invention is a decoder designed for improved prediction-based multi-channel reconstruction that compensates for energy losses due to prediction error by amplifying the channels subjected to the "up-mix" process.
Another preferred embodiment of the present invention is a decoder for improved prediction-based multi-channel reconstruction replacing the energy losses associated with the prediction error by a de-correlated signal.
Another preferred embodiment of the present invention is a decoder designed for improved prediction-based multi-channel reconstruction replacing some of the energy losses associated with the prediction error by a de-correlated signal and leveling some of the energy losses by enhancing the upmixed channels.
Fig. 11 shows a multi-channel synthesizer for generating at least three output channels 1100 using an input signal including at least one base channel 1102, at least one base channel being obtained from the original multi-channel signal. The multi-channel synthesizer shown in Fig. 11 includes a mixing device 1104 that can be implemented as shown in any of Figs. 2 to 10. Mixer 1104 generally allows at least one base channel to "up-mix" using the "up-mix" rule, resulting in at least three output channels. The 1104 mixing device allows you to generate at least three output channels based on energy measurement 1106 and at least two different up-mix parameters 1108 using the up-mix rule introducing energy loss, so that at least three output channels have energy that is greater than signal energy resulting from the application of the up-mix rule itself introducing energy losses. In this way, regardless of the energy error depending on the up-mix rule introducing energy losses, the invention provides energy compensation, wherein energy compensation can be performed by amplifying and / or adding a de-correlated signal. At least two different 1108 up-mix parameters and 1106 energy measurement are included in the input signal.
The energy measure is preferably any value associated with energy losses associated with the up-mix rule used. It can be the total value of the energy error introduced in the up-mix process or the energy of the signal subjected to the upmix process (which usually has less energy than the original signal), or it can be a relative value, such as the ratio of the energy of the original signal to the signal energy up-mix process or the ratio of the energy error to the energy of the original signal or even the ratio of the energy error to the energy of the up-mix signal. The relative energy measure can be used as a correction factor, but is nonetheless a measure of energy, because it depends on the energy error introduced into the up-mix process of the signal generated based on the up-mix rule introducing energy loss, or by saying in other words - the "upmix" rule that does not guarantee energy conservation.
An example of an up-mix rule introducing energy losses (an up-mix rule that does not ensure energy conservation) is an "up-mix" using forward predictive factors. In the case of imprecise prediction of a frame or a fragment of the frame band, the signal subjected to the "up-mix" process of the output is burdened with a prediction error that corresponds to energy losses. Of course, the prediction error is different for different frames, because in the case of almost perfect prediction (small prediction error) only a small compensation (by amplifying or adding a de-correlated signal) needs to be carried out, while in the case of a larger prediction error (imperfect prediction) compensation. The measure of energy according to the invention varies, therefore, between a value meaning no compensation or little compensation and a value meaning big compensation.
If the energy measure is considered as a degree of statistical channel compatibility (ICC), and this approach is natural if the compensation is carried out by adding a de-correlated signal, amplified to an extent dependent on the energy measure, preferably the relative energy measure (ρ) used usually varies between values 0.8 and 1.0, where 1.0 means that the upmixed signals are de-correlated to the required degree or that there is no need to add a de-correlated signal or that the signal energy resulting from the predictive up-mix is equal to the energy of the original signal or that the prediction error is zero.
However, the present invention is also useful for other up-mix rules introducing energy losses, i.e. rules that are not based on signal waveform matching but are based on other techniques such as code book techniques, spectrum matching or any other "up-mix" rules that do not guarantee energy conservation.
Energy compensation can generally be carried out before or after applying the up-mix rule introducing energy losses. In an alternative solution, energy loss compensation can even be included in the up-mix rule, for example by replacing the original matrix coefficients using an energy measure, thanks to which a new up-mix rule is created, which is used by the implementing module up-mix process. The new "up-mix" rule is based on the "upmix" rule introducing energy losses and on the measure of energy. In other words, this embodiment of the invention relates to a situation in which energy compensation is "mixed up" in the "extended" up-mix rule, so that energy compensation and / or the addition of a de-correlated signal is performed by using one or more up-mix matrices on the input vector (one or more base channels) to obtain (after one or more operations on the arrays) the output vector ( of a reconstructed multi-channel signal comprising at least three channels).
The up-mix device preferably receives two base channels l0, r0 and provides at the output three reconstructed channels l, r and c.
Next, to illustrate exemplary energy at various locations on the encoder-decoder path, we refer to Fig. 12. Block 1200 represents the energy of a multi-channel audio signal, such as the signal shown in Fig. 1 comprising at least the left channel, right channel and center channel. In the embodiment shown in Fig. 12, it was assumed that shown in Fig. 1 input channels 101, 102, 103 are completely uncorrelated, while the module implementing the "down-mix" process allows energy conservation. In this case, the energy of one or more base channels marked with block 1202 is the same as the energy 1200 of the original multi-channel signal. If the original multi-channel signals are correlated with each other, the energy of the base channel 1202 may be less than the energy of the original multi-channel signal, for example in the case where the left and right channels (partially) cancel each other out.
For the purposes of the further description, however, it has been assumed that the energy of 1202 base channels is the same as the energy of 1200 of the original multi-channel signal.
Block 1204 represents the energy of signals subjected to the "up-mix" process, if "signals subjected to the" up-mix process "(e.g. signals 110, 111, 112 presented in Fig. 1) are generated using the" up-mix "that does not allow energy or predictive up-mix discussed with reference to Fig. 1. Because, as will be discussed below with reference to Figs. 14a and Fig. 14b, such a predictive up-mix introduces an Er energy error, the energy 1204 of the up-mix result is less than the energy of the base 1202 channels.
The 1104 up-mix device can provide output channels that have more energy than 1204. The 1104 up-mix device advantageously performs all compensation, so that the up-mix result shown in Figure 11 "Has energy marked with block 1206.
The result of using the "up-mix" whose energy is marked with block 1204 is preferably not simply amplified, as shown in Fig. 2, individual channels are not strengthened, as shown in Fig. 3, nor is it amplified by the encoder as shown in Fig. 4. Instead, the remaining Er energy that corresponds to the error associated with the predictive "up-mix" is "filled" with a de-correlated signal. In another preferred embodiment of the invention, this Er energy error is only partially compensated by a de-correlated signal, while the remainder of the energy error is compensated by scaling the result of the up-mix. The complete leveling of the energy error by the de-correlated signal is shown in Fig. 5 and Fig. 6, while the "partial" solution is shown in Fig. 7.
Fig. 13 shows many energy compensation methods, i.e. methods whose common feature is based on a measure of energy dependent on an energy error, the energy of the output channels being greater than the pure result of the predictive up-mix, i.e. the result of application (uncorrected ) up-mix rules introducing energy losses.
Number 1 in the table shown in Fig. 13 relates to the energy compensation carried out in the decoder, which is carried out after the "up-mix" process. This option is shown in Fig. 2 and further developed with reference to Fig. 3, where the upward scaling coefficients gz characteristic for individual channels are presented, which depend not only on the energy measure ρ, but also depend on the characteristic for each channel "downmix" coefficients v<sub>from</sub>where z is 1, r or c.
Number 2 in the table shown in Fig. 13 relates to the energy compensation carried out in the encoder, which is carried out after the down-mix process which is shown in Fig. 4. This embodiment of the invention is advantageous because the energy measure does not have to be transmitted from the encoder to the decoder.
The number 3 in the table shown in Fig. 13 relates to the energy compensation carried out in the decoder, which is carried out before carrying out the "up-mix" process. Looking at Fig. 2, the energy correction 202 performed in Fig. 2 after the "up-mix" process is performed in this solution before the "upmix" block 201 shown in Fig. 2. Compared to the solution shown in Fig. 2 this embodiment of the invention is easier to implement because it is not required to use the correction coefficients characteristic of individual channels shown in Fig. 3, although this may entail some reduction in quality.
Number 4 in the table in Fig. 13 relates to another embodiment of the invention in which the correction is performed in the encoder before performing the down-mix process. Looking at Fig. 1, channels 101, 102, 103 are amplified using the corresponding compensation factor, so that the output energy of the down-mix device is increased after the down-mix process is performed, which is seen in Fig. 12 and reference number 1208. In this way, the fourth embodiment of Fig. 13 has the same effect on the base channels available at the encoder output as the second embodiment of the present invention.
The number 5 in the table of Fig. 13 relates to the embodiment of Fig. 5 in which the de-correlated signal is obtained from channels generated using the "up-mix" rule of Fig. 5 not providing energy conservation 109.
The number 6 in the table in Fig. 13 relates to an embodiment of the invention in which only a portion of the residual energy is replaced by a de-correlated signal. This embodiment of the invention is illustrated in Figure 7.
The embodiment of the invention indicated in the table in Fig. 13 with the number 8 is similar to the embodiments of the invention denoted with the numbers 5 and 6, but the de-correlated signal is obtained here from the base channels before carrying out the "down-mix" process, as indicated by the block 501 'on fig. 5.
Next, the preferred embodiment of the encoder is described in detail. Fig. 14a shows an encoder for processing multi-channel input signal 1400 comprising at least two channels, and preferably comprising at least three channels 1, c, r.
The encoder includes an energy measurement calculator 1402, enabling the calculation of a measurement error depending on the difference between the energy of a 1400 multi-channel input signal or at least one base channel 1404, and subjected to the "up-mix" signal by the 1406 signal generated in the "up-mix" process, which does not ensure preservation energy 1407.
The encoder further includes an output interface 1408, providing at least one base channel after scaling (401, 402) using a scaling factor 403 dependent on the energy measure, or providing the energy measure itself.
In a preferred embodiment of the invention, the encoder includes a 1410 down-mix device designed to generate at least one base channel 1404 from the original 1400 multi-channel signal. The 1414 difference calculator and the parameter optimization module are also used to generate the up-mix parameters. 1416. These elements allow you to find the best-matched up-mix parameters 1412. In a preferred embodiment of the invention, at least two parameters from this set of best-suited up-mix parameters are made available via the output interface as output parameters. The difference calculator preferably allows the calculation of the minimum mean square error between the original multi-channel 1400 signal and the up-mix signal generated by the device performing the up-mix process for parameters input to parameter line 1412. This parameter optimization process can be carried out using several different optimization procedures that are designed to obtain the best-matched 1406 up-mix result using certain up-mix matrices contained in the device performing the up-mix 1407 process.
The operation of the encoder shown in Fig. 14a is illustrated in Fig. 14b. After completing the 1440 down-mix step of the 1410 down-mix device, the base channel or multiple base channels can be output as illustrated in block 1442. Next, the 1444 up-mix parameter optimization step is performed. , which depending on the adopted optimization strategy can be an iterative or non-iterative process. However, iterative processes are beneficial. Generally, the process of optimizing the "upmix" parameters should be implemented in such a way that the difference between the "up-mix" result and the original signal is as small as possible. Depending on the implementation, this difference may be a difference associated with individual channels or a complex difference. Generally, the effect of performing the upmix parameter optimization step 1444 is to minimize any cost function that can be obtained for individual channels or connected channels, so that a larger difference (error) is acceptable for one channel, if, for example, a better matching of the other two is obtained channels.
Then, after finding a set of the best-matched parameters, eg after finding the best-matched up-mix matrix, at least two upmix parameters selected from the set of parameters generated in step 1444 are sent to the output interface, which is designated as step 1446.
In addition, after the optimization step of the up-mix parameter, a measure of energy can be calculated which is available at the output, which is designated as step 1448. The measure of energy is generally dependent on the energy error 1210. In a preferred embodiment of the invention, the measure of energy is in the form of a coefficient ρ, which depends on the ratio of the energy of the up-mix result 1406 to the energy of the original 1400 signal, as shown in Fig. 2. In an alternative solution, the calculated and available energy measure may be in the form of the absolute value of the energy error 1210 or it may be in the form of the absolute energy of the 1406 up-mix result, which of course depends on the energy error. In this context, it should be noted that the measure of energy made available via the 1408 output interface is preferably quantized, and preferably entropy encoded using any of the well-known entropy encoders, such as an arithmetic encoder, a Hoffman encoder or a series length encoder, which is particularly useful in situations where there are many successive, same measures of energy. Alternatively or additionally, the energy measures for subsequent time intervals or frames may be differential encoded, with the differential coding preferably being carried out prior to entropy coding.
Fig. 15a shows an alternative embodiment of a down-mix device that in a preferred embodiment of the present invention is connected to the encoder shown in Fig. 14a. In the embodiment of the device shown in Fig. 15a, the SBR implementation is used, although this embodiment can also be used in cases where replication bandwidth expansion is not performed but the entire baseband bandwidth is transmitted. Shown in fig. 15a, the encoder includes a down-mix device 1500 intended for processing the original 1500 signal to obtain at least one base channel 1504. In an embodiment that does not include the use of SBR, at least one base channel 1504 is introduced into the basic encoder 1506, which in the case of a single base channel can be in the form of an AAC encoder for processing mono signals, and for example in the case of two stereo base channels it can be any encoder stereo.
At the output of basic encoder 1506, a data stream containing the encoded base channel or containing the many encoded base channels (1508) is available.
In the case where the embodiment shown in Fig. 15a uses SBR, at least one base channel 1504 is filtered prior to entering the base encoder using a low-pass filter 1510. The functionality of blocks 1510 and 1506 can, of course, be implemented in one coding device that performs filtering with using a low-pass filter and basic coding using a single coding algorithm.
The encoded base channels available at output 16508 contain only the lower baseband channel 1504 in coded form. Upper band information is calculated by the SBR 1512 spectral envelope calculator, which is connected to the SBR information encoder 1514 to generate encoded SBR information provided at output 1516.
The original signal 1502 is introduced into the energy calculator 1520, which generates channel energies (for a certain period of time of the original channels 1, c, r, with the channel energies labeled L, C, R, and the output being block 1520). The energies of the L, C, R channels are introduced into the block of the parameter calculator 1522. The parameter calculator 1522 provides two "up-mix" parameters c1, c2 at the output, which can be, for example, the parameters c1, c2 shown in Fig. 15a. Of course, the parameter calculator 1522 can also generate other (e.g., linear) energy combinations for sending to the decoder, including the energies of all input channels. Of course, the various up-mix parameters sent result in the need to use different ways of calculating the remaining elements of the upmix matrix. As noted in relation to equation (40) or equations (41 - 44), the up-mix matrix intended for the one shown in Fig. 15 the energy-oriented embodiment has at least four non-zero elements, with the elements in the third row being the same. The parameter calculator 1522 can therefore use any combination of energy L, C, R, from which it is possible to obtain four elements of the "up-mix" matrix, for example by using the equation of the "up-mix" matrix (40).
The embodiment shown in Fig. 15a illustrates an encoder that allows energy-saving "up-mix" or, generally speaking, energy-based up-mix of the entire signal band to be carried out. This means that for the encoder shown in Fig. 15a, a parametric representation given at the output by the parameter calculator 1522 is created for the entire signal. This means that for each fragment of the band of the coded base channel, a corresponding set of parameters is generated and made available at the output. If, for example, an encoded broadband base signal is considered whose band is divided into ten sub-ranges of the band, the parameter calculator may output ten parameters c1 and c2 corresponding to each of the sub-ranges of the encoded base signal. However, if the encoded base signal is in the form of a narrowband signal in an SBR environment containing, for example, only the lower five subbands, the 1522 parameter calculator provides a set of parameters for each of the lower five subbands, and additionally for each of the upper five subbands, although the signal provided at output 1508 does not contain the corresponding subbands. This is because such a subrange will be reconstructed by the decoder, which will be described below with reference to Fig. 16a.
In the preferred embodiment which has been described with reference to Fig. 10, however, the energy calculator 1520 and parameter calculator 1522 only process the upper band of the primary signal, while the parameters corresponding to the lower part of the band of the original signal are calculated using the parameter calculator of Fig. 10. that corresponds to the module in Fig. 10 performing the predictive "up-mix" 109.
Fig. 15b is a schematic representation of the parametric representation given at the output of the selector module 1002 of Fig. 10. The parametric representation of the present invention thus includes (together with or without the encoded base channel / base channels and, optionally, even without energy measure) a set of predictive parameters corresponding to the lower band range, e.g. corresponding to the subband range from 1 to and, and subband parameters corresponding to the upper band range, for example the subband range from i + 1 to N. Alternatively, the predictive parameters and the energy-based parameters can be mixed together, for example, the sub-range described using the energy-based parameters can be placed between the sub-ranges of the band described using the predictive parameters.
The frame described using only predictive parameters may further follow the frame described using only parameters based on energy values. Thus, generally speaking, the present invention described with reference to Fig. 10 relates to various parameterizations that may vary depending on the frequency as seen in Fig. 15b, or which may differ depending on the time when the frame described using only predictive parameters is followed by the frame described using only parameters based on energy values. The distribution of parameterization of the subranges of the band can of course vary for different frames, so that, for example, the subrange i has in the first frame the first set of parameters (e.g. predictive), as shown in Fig. 15b, and a second set of parameters (e.g., based on energy values) in a different frame.
The present invention is further useful when parameterizations other than the predictive parameterization in Fig. 14a or energy-based parameterization in Fig. 15a are used. Other examples of parameterization other than predictive or energy-based parameterization may also be used, provided that any endpoint or target event such as up-mix quality, down-mix rate, encoder or decoder computing performance, or, for example, consumption battery-powered devices, etc. indicate that for certain band or frame ranges one parameterization is better than other parameterization. The target function can of course be a combination of various individual goals / events described above. An example of an event may be the reconstruction of the upper band using the SBR technique, etc.
It should also be noted that the choice of calculation and transmission of parameters depending on frequency or time can be directly signaled, as illustrated in block 1005 shown in Fig. 10. In an alternative solution, the signaling can also be carried out in a hidden manner, as discussed with reference to Fig. 16a. In this case, the decoder uses predetermined rules, for example, the decoder can automatically assume that the transmitted parameters are energy-based parameters for the band sub-ranges in the upper range shown in Fig. 15b, for example the band sub-ranges that have been reconstructed using replicative bandwidth extension or high frequency reconstruction techniques.
It should also be noted that the encoder disclosed in the invention calculating one, two or even more different parameterization and the encoder selecting which parameterization to be sent, based on a decision made using all information available to the encoder (the information may be in the form used at the moment target function or may be signaling information used for other reasons, such as SBR processing and signaling) can be performed with or without transmitting the energy measure. Even if the beneficial energy correction is not carried out at all, for example, when the result of the non-energy-preservation up-mix is not subjected to energy correction or if the corresponding pre-compensation in the encoder is not carried out, disclosed in of the invention, switching between different parameterizations is useful for obtaining higher quality multi-channel output signal and / or lower bit rate.
The switching disclosed in the invention between the different parameterizations depending on the information available to the encoder can in particular be used with or without the use of a de-correlated signal, completely or partially replacing the energy error introduced as a result of the predictive upmix, which was described with reference to fig. 5-7 In this context, the addition of a de-correlated signal as described with reference to Fig. 5, is carried out only for those subranges of the band / frames for which the predictive up-mix parameters are sent, while for the subranges of the band / frames for which parameters based on energy values have been sent, other methods of de-correlation are used. Such a method can be, for example, attenuation of the "wet" signal and generation of a de-correlated signal and amplification of the de-correlated signal, thanks to which a specific degree of de-correlation is obtained, required, for example, by the transmitted measure of the degree of statistical compatibility of ICC channels, and then the "dry" signal is scaled properly correlated signals.
Next, Fig. 16a will be discussed, which illustrates the implementation in the decoder of the up-mix process module 201 according to the invention and the corresponding energy correction in block 202. As described with reference to Fig. 11, the transmitted up-mix parameters "1108 are obtained from the received input signal. These transmitted up-mix parameters are preferably entered into the 1600 calculator to calculate the remaining up-mix parameters when the 1602 upmix matrix including energy compensation is used to carry out the up-mix process and earlier or later energy correction. The process of calculating the remaining up-mix parameters is further discussed with reference to Fig. 16b.
The calculation of the up-mix parameters is based on the equation shown in Fig. 16b, which was also repeated as equation (7). In the case of an embodiment of the invention comprising three input signals / two output signals, the "downmix" matrix D comprises six variables. In addition, the up-mix matrix C also contains six variables. However, on the right side of equation (7) there are only four values. Thus, in the case of the unknown "down-mix" and the unknown "up-mix", there are twelve unknown variables from the D and C matrices, and only four equations to find these twelve variables. However, "down-mix" is known, so that the number of unknown variables is limited to the coefficients of the "up-mix" C matrix containing six variables, but there are still four equations to find these six variables. Therefore, to determine at least two variables of the up-mix matrix, preferably c11 and c22, the optimization method described with reference to step 1444 in Fig. 14b and shown in Fig. 14a is used. Then, because there are four unknowns, for example c12, c21, c31 and c32, and because there are four equations, for example, one equation for each element of the unit matrix I to the right of the equation in Fig. 16b, other unknown up-mix matrix variables can be easily calculated. These calculations are performed by the 1600 calculator to obtain the remaining "up-mix" parameters.
The up-mix matrix in the 1602 device is created based on two transmitted up-mix parameters, which is indicated by the dashed line 1604, and based on the other four up-mix parameters calculated in block 1600. Such an up matrix "mix" is then applied to the base channels introduced via line 1102. Depending on the implementation, the energy measure used to correct the lower part of the band is sent via line 1106, thanks to which it is possible to generate and share the corrected "up-mix" at the output. In the case where the predictive "up-mix" is carried out only for the lower part of the band, which is for example signaled in a hidden way via line 1606, and when on line 1108 there are up-mix parameters based on energy values, this fact it is signaled for the corresponding sub-band to the 1600 calculator and to the 1602 up-mix matrix computing device. When using energy-based parameters, it is preferable to calculate up-mix matrix elements according to equations (40) or (41). For this purpose, the transmitted parameters are used, as shown under equation (40) or the corresponding parameters, as shown under equation (41). In this embodiment of the invention, the transmitted up-mix parameters c1, c2 cannot be used directly as up-mix coefficients, but up-mix coefficients of the up-mix matrix, as shown in equation (40) or (41), must be calculated using the transmitted parameters c1 and c2.
In the case of the upper part of the band, the "up-mix" matrix determined using energy-based up-mix parameters is used in the "up-mix" process of the upper part of the multi-channel output signal band. The lower part of the band and the upper part of the band are then connected together using a device connecting the lower / upper part of the band 1608 to provide the output of the reconstructed output channels 1, r, which is the full bandwidth. As shown in fig. 16a, an upper portion of the base channel band is generated using a decoder to decode transmitted lower portions of the base channel band, the decoder being a mono decoder in the case of a monophonic base channel and a stereo decoder in the case of two stereo base channels. The decoded lower part of the base channel band or base channels is fed into the SBR 1614 which additionally receives the envelope information calculated by the device 1512 shown in Fig. 15a. Based on the lower part of the band and information about the envelope bound to the upper part of the band, the upper part of the base channel band is generated, which aims at obtaining on the 1102 line the base channels with full bandwidth, which are sent to the device calculating the "up-mix" matrix 1602.
The methods, devices or computer programs of the invention may be implemented or included in many devices. Fig. 17 shows a transmission system comprising a transmitter comprising a coder according to the invention and a receiver comprising a decoder according to the invention. The transmission channel may be in the form of a wireless or wired channel. As shown in Fig. 18, the encoder may further be included in the sound recorder or a decoder may be included in the sound player. Sound recordings from the sound recorder can be sent to the sound player via the Internet or using a data carrier sent by regular mail or via courier. Other data carriers, such as memory cards, CDs or DVDs may be used for transmission.
Depending on the individual requirements of the inventive methods, these methods can be implemented in hardware or in software. The implementation can be carried out using a digital data carrier, in particular a hard disk or a CD, containing controllable electronic signals that can cooperate with a programmable computer system enabling the implementation of the methods of the invention. In other words, the methods of the invention are in this case in the form of a computer program comprising program code implementing the methods of the invention when running the program on the computer.
CODING TECHNOLOGIES AB, Sweden KONINKLIJKE PHILIPS ELECTRONICS NV, The Netherlands
PROXY:
EP 1 738 353
Z - 4667/07
Contents8
45 members in 14 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 0402652 | Sweden | A | |
| 0402652 | Sweden | A | |
| 05797620 | European Patent Office (EPO) | A | |
| 2005011587 | European Patent Office (EPO) | W | |
| 2005011587 | European Patent Office (EPO) | W | |
| EP20050797620 | – | – | – |
| SE20040002652 | – | – | – |
| WO2005EP11587 | – | – | – |
Members45
| Document | Office | Kind | |
|---|---|---|---|
| SE0402652D0 | Sweden | D0 | |
| WO2006048203A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2006048204A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2006140412A1 | United States of America | A1 | |
| US2006165237A1 | United States of America | A1 | |
| TW200627380A | Taiwan Province of China | A | |
| TW200629961A | Taiwan Province of China | A | |
| EP1730726A1 | European Patent Office (EPO) | A1 | |
| EP1738353A1 | European Patent Office (EPO) | A1 | |
| KR20070038043A | Republic of Korea | A | |
| KR20070049627A | Republic of Korea | A | |
| CN1969317A | China | A | |
| HK1097082A1 | Hong Kong, China | A1 | |
| CN1998046A | China | A | |
| HK1097336A1 | Hong Kong, China | A1 | |
| EP1738353B1 | European Patent Office (EPO) | B1 | |
| AT371925T | Austria | T | |
| ATE371925T1 | Austria | T1 | |
| EP1730726B1 | European Patent Office (EPO) | B1 | |
| DE602005002256D1 | Germany | D1 | |
| AT375590T | Austria | T | |
| ATE375590T1 | Austria | T1 | |
| DE602005002833D1 | Germany | D1 | |
| PL1738353T3This record | Poland | T3 | |
| ES2292147T3 | Spain | T3 | |
| DE602005002833T2 | Germany | T2 | |
| PL1730726T3 | Poland | T3 | |
| ES2294738T3 | Spain | T3 | |
| JP2008517337A | Japan | A | |
| JP2008517338A | Japan | A | |
| DE602005002256T2 | Germany | T2 | |
| RU2006146947A | Russian Federation | A | |
| RU2006146948A | Russian Federation | A | |
| KR100885192B1 | Republic of Korea | B1 | |
| KR100905067B1 | Republic of Korea | B1 | |
| RU2369917C2 | Russian Federation | C2 | |
| RU2369918C2 | Russian Federation | C2 | |
| US7668722B2 | United States of America | B2 | |
| TWI328405B | Taiwan Province of China | B | |
| JP4527781B2 | Japan | B2 | |
| JP4527782B2 | Japan | B2 | |
| CN1969317B | China | B | |
| TWI338281B | Taiwan Province of China | B | |
| CN1998046B | China | B | |
| US8515083B2 | United States of America | B2 |
Numbers
- Publication, DOCDB
- 1738353
- Publication, EPODOC
- PL1738353T
- Application
- 797620
- Application, DOCDB
- 05797620
- Application, EPODOC
- PL20050797620T
Titles2
- English
- MULTI PARAMETRISATION BASED MULTI-CHANNEL RECONSTRUCTION
- Polish
- Wielokanałowa rekonstrukcja oparta na wielu parametryzacjach
Classification
- CPC, 3
- G10L19/008
- G10L19/04
- H04S2420/03
- IPC, 4
- G10L19 00
- G10L19 008
- G10L19 04
- G11B