Individual channel temporal envelope shaping for binaural cue coding schemes and the like
Abstract
At an audio encoder, cue codes are generated for one or more audio channels, wherein an envelope cue code is generated by characterizing a temporal envelope in an audio channel. At an audio decoder, E transmitted audio channel(s) are decoded to generate C playback audio channels, where C>E≧1. Received cue codes include an envelope cue code corresponding to a characterized temporal envelope of an audio channel corresponding to the transmitted channel(s). One or more transmitted channel(s) are upmixed to generate one or more upmixed channels. One or more playback channels are synthesized by applying the cue codes to the one or more upmixed channels, wherein the envelope cue code is applied to an upmixed channel or a synthesized signal to adjust a temporal envelope of the synthesized signal based on the characterized temporal envelope such that the adjusted temporal envelope substantially matches the characterized temporal envelope.

Term
Term ended
Projected expiry passed 7 September 2025, 1 year ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
40 claims: 6 independent, 34 dependent
- 1Zastrzeżenia patentowe 1. Sposób kodowania kanałów audio obejmujący:generowanie dwóch łub większej liczby kodów parametrycznych dla jednego lub większej liczby kanałów- audio, przy czym przynajmniej jeden kod parametryczny ma postać kodu parametrycznego obwiedni wygenerowanego przez opisanie obwiedni czasowej w jednym lub większej liczbie kanałów audio, dwa lub w-iększa liczba kodów- parametrycznych zawiera jeden lub większą liczbę parametrów dotyczących stopnia statystycznej zgodności kanałów- (ICC), parametrów dotyczących międzykanalowej różnicy natężeń (ICLD) i parametrów dotyczących międzykanalowej różnicy czasu (ICTD), gdzie pierwsza rozdzielczość w czasie powiązana z kodem parametrycznym obwiedni jest wyższa niż druga rozdzielczość w czasie powiązana z innym kodem(mi) parametrycznym^) i gdzie obwiednia czasowa jest opisana dla odpowiadającego jej kanału audio w dziedzinie czasu lub oddz.clnic dla różnych podzakresów pasma odpowiadającego jej kanału audio w dziedzinie podzakresu pasma;oraz przesyłanie dwóch luh większej liczby kodów- parametrycznych.
- 2Sposób według zastrzeżenia 1, znamienny tym, żc obejmuje ponadto przesyłanie E przesyłanych kanałów- audio od pcw jadających jednemu lub większej liczbie kanałów audio, gdzie E .
- 3Sposóh w-cdlug zastrzeżenia 2, znamienny tym, że jeden lub większa liczba kanałów- audio obejmuje C wejściowych kanałów audio, gdzie C E, oraz C kanałów wejściowych poddawanych jest procesowi downmixu w celu wygenerowania E przesyłanych kanałów.
- 4Sposób według zastrzeżenia 1, znamienny tym, że dwa lub większa liczba kodów parametrycznych jest przesyłana w celu umożliwienia przeprowadzenia przez dekoder procesu kształtowania obwiedni podczas dekodowania E przesyłanych kanałów w oparciu o dwa lub większą liczbę kodów parametrycznych, przy czym E przesyłanych kanałów odpowiada jednemu lub większej liczbie kanałów audio, gdzie
- 5Sposóh według zastrzeżenia 4, znamienny tym, żc kształtowanie obwiedni obejmuje modyfikację obwiedni czasowej syntetyzowanego sygnału generowanego przez dekoder w laki sposób, aby była ona zgodna z opisaną obwiednią czasową.
- 6Sposób według zastrzeżenia 1, znamienny tym, że obwiednia czasowa jest opisana jedynie dla określonych częstotliwości odpowiadającego jej kanału audio.
- 7Sposób według zastrzeżenia 1, znamienny tym, że obwiednia czasowa jest opisana jedynie dla częstotliwości odpowiadającego jej kanału audio większych niż określona częstotliwość odcięcia.
- 8Sposób według zastrzeżenia 1 znamienny lym, żc dziedzina podzakresu pasma odpowiada kwadraturow-emu filtrowi zwierciadlanemu (QMF).
- 9Sposób według zastrzeżenia 1 znamienny lym, że obejmuje ponadto określenie czy należy włączyć czy wyłączyć proces opisywania.
- 10Sposób według zastrzeżenia 9 znamienny tym, że obejmuje ponadto generow-anie i przesyłanie znacznika wlącz/wylącz w oparciu o wynik przeprowadzonego określenia, co ma na celu przekazanie do dekodera informacji, czy należy czy nie należy przeprowadzić kształtowanie obwiedni podczas dekodowania E przesyłanych kanałów odpowiadających jednemu lub większej liczbie kanałów audio, gdzie E i II. Sposóh według zastrzeżenia 9, znamienny lym, że określenie oparte jesi na analizie kanału audio w celu wykrycia transjentów w kanale audio, lak że proces opisywania jest włączany w przypadku wykrycia obecności iransjcntu.
- 1112. Sposób według zastrzeżenia 1, znamienny lym, że etap generowania kodu parametrycznego obwiedni obejmuje pierwiastkowanie (1006) lub utworzenie powiększenia, a następnie filtrowanie z wykorzystaniem fillra dolnoprzepuslow-ego (1008) próbek sygnału kanału audio lub sygnałów podzakresów pasma kanału audio w celu opisania obwiedni czasowej.
- 1213. Sposób według zastrzeżenia I albo 12, znamienny tym, że etap generowania obejmuje ponadto etap parametryzacji, kw-anlowania i kodowania uzyskanej obwiedni czasowej.
- 1314. Urządzenie do kodowania kanałów audio zawierające:elementy umożliwiające generowanie dwóch lub większej liczby kodów 7 parametrycznych dla jednego lub większej liczby kanałów audio, przy czym przynajmniej jeden kod parametryczny ma postać kodu parametrycznego obwiedni wygenerowanego przez opisanie obwiedni czasowej w jednym lub większej liczbie kanałów audio, gdzie dwa lub większa liczba kodów· parametrycznych zawiera jeden lub większą liczbc parametrów' dotyczących stopnia statystycznej zgodności kanałów (ICC), parametrów dotyczących iniędzykanałow'cj różnicy natężeń (ICLD) i parametrów dotyczących międzykanalowej różnicy czasu (ICTD), gdzie pierwsza rozdzielczość w czasie powiązana z kodem parametrycznym obwiedni jcsl wyższa niż druga rozdzielczość w czasie powiązana z innym kodem(mi) parametrycznymi i), i gdzie obwiednia czasowa jest opisana dla odpowiadającego jej kanału audio w dziedzinie czasu lub oddzielnie dla różnych podzakresów pasma odpowiadającego jej kanału audio w dziedzinie podzakresu pasma;oraz elementy umożliwiające przesyłanie dwóch lub większej liczby kodów parametrycznych.
- 1415. Urządzenie według zastrzeżenia 14, znamienne tym, że:urządzenie umożliwia kodow-anic C wejściowych kanałów audio w celu wygenerowania E przesyłanych kanałów audio, zaś elementy umożliwiające generowanie obejmują analizator obwiedni przystosowany do opisywania wejściowej obwiedni czasowej przynajmniej jednego 7. C kanałów wejściowych, elementy umożliwiające generowanie obejmują ponadto estymator kodów, przystosowany do generowania kodów parametrycznych dotyczących dwóch lub większej liczhy 7. C kanałów wejściowych, przy czyim urządzenie zawiera ponadto urządzenie przeprowadzające proces downmixu, przystosowane do przeprowadzania procesu downmixu C kanałów wejściowych w- celu wygenerowania E przesyłanych kanałów audio, gdzie OESł;a elementy umożliwiające przesyłanie są przystosowane do przesyłania informacji dotyczących dwóch lub większej liczny kodów' parametrycznych, umożliwiających dekoderowi przeprowadzenie procesów syntezy i kształtowania obwiedni podczas dekodowania E przesyłanych kanałów.
- 1516. Urządzenie według zastrzeżenia 15, znamienne tym, żc urządzenie ma postać sys.eniu wybranego z grupy obejmującej cyfrowy rejestrator wideo, cyfrowy rejestrator audio, komputer, nadajnik satelitarny, nadajnik kablowy, nadajnik wykorzystywany w r transmisji naziemnej, domowe centrum multimedialne i system kina domowego;i system obejmuje analizator obwiedni, estymator kodów i urządzenie przeprowadzające modyfikacje obwiedni.
- 1617. Program komputerowy obejmujący kod programu zapewniający po wykonaniu kodu na komputerze implementację sposobu określonego w zastrzeżeniu 1.
- 1718. Zakodowany strumień danych audio obejmujący:dwa lub większą liczbę kodów parametrycznych dla jednego lub większej liczby kanałów audio, przy czym przynajmniej jeden kod parametryczny ma postać kodu parametrycznego obwiedni wygenerowanego przez opisanie obwiedni czasowej w jednym lub większej liczbie kanałów audio, gdzie dwa lub większa liczba kodów* parametrycznych obejmują ponadto jeden lub większą liczbę parametrów dotyczących stopnia statystycznej zgodności kanałów (ICC), parametrów dotyczących międzykanalowej różnicy natężeń (ICLD) i parametrów dotyczących międzykanalowej różmicy czasu (ICTD), gdzie pierwsza rozdzielczość w czasie powiązana z kodem parametrycznym obwiedni jest wyższa niż druga rozdzielczość w czasie powiązana z innym kodem(mi) parametrycznymi i), i gdzie obwiednia czasowa jest opisana dla odpowiadającego jej kanału audio w dziedzinie czasu lub oddzielnie dla różnych podzakresów' pasma odpowiadającego jej kanału audio w* dziedzinie podzakresu pasma;a w zakodowanym strumieniu danych audio zakodowane są dwa lub większa liczba kodów* parametrycznych i E przesyłanych kanałów audio odpowiadających jednemu lub większej liczbie kanałów* audio, gdzie E^.
- 1819. Zakodowany strumień danych audio według zastrzeżenia 18 znamienny tym. że ponadto zawiera E przesyłanych kanałów· audio, gdzie E przesyłanych kanałów audio odpowiada jednemu lub większej liczbie kanałów audio.
- 1920. Sposób dekodowania E przesyłanych kanałów audio w celu wygenerowania C odtwarzanych kanałów audio, gdz.c C E \, obejmujący:odebranie kodów parainetycznych odpowiadających E przesyłanym kanałom, gdzie kody parametryczne obejmują kod parametryczny obwiedni odpowiadający opisanej obwiedni czasowej kanału audio odpowiadającemu E przesyłanym kanałom, gdzie dwa lub większa liczba kodów parametrycznych obejmuje ponadto jeden lub większą liczbę parametrów' dotyczących stopnia statystycznej zgodności kanałów (ICC), parametrów dotyczących międzykanalowej różnicy natężeń (ICLD) i parametrów dotyczących międzykanalowej różnicy czasu (ICTD), zaś pierwsza rozdzielczość w czasie powiązana z kodein parametrycznym obwiedni jest wyższa niż druga rozdzielczość w czasie powiązana z innym kodcm(mi) parametrycznym(i);przeprowadzenie procesu upmixu E przesyłanych kanałów audio w celu wygenerowania jednego lub większej liczby kanałów' poddanych procesowi upmixu;oraz przeprowadzenie procesu syntezy C odtwarzanych kanałów' przez zastosowanie kodów' parametrycznych względem jednego lub większej liczby kanałów poddanych procesowi upmixu, gdzie kod parametryczny obwiedni stosowany jest względem kanału poddanego procesowi upmixu lub kanału syntetyzowanego w celu dokonania modyfikacji obwiedni czasowej syntetyzowanego kanału w oparciu o opisaną obwiednię czasową, przy czym dokonywane Jest to przez skalowanie próbek sygnału w dziedzinie czasu lub w dziedzinie podzakresów pasma z wykorzystaniem współczynnika skalowania w ;taki sposóh, że zmodyfikowana obwiednia czasowa jest zgodna z opisaną obwiednią czasową.
- 2021. Sposób według zastrzeżenia 20, znamienny tym, że kod parametryczny obwiedni odpowiada opisanej obwiedni czasowej w pierwotnym kanale wejściowym wykorzystanym do wygenerowania E przesyłanych kanałów.
- 2122. Sposób według zastrzeżenia 21, znamienny tym, że synteza obejmuje syntezę późnej rewerberacji ICC.
- 2223. Sposób według zasirzeżenia 21, znamienny tym, że obwiednia czasowa syntetyzowanego sygnału poddawana jest modyfikacji przed przeprowadzeniem syntezy ICLD.
- 2324. Sposób według zastrzeżenia 20, znamienny tym, że:obw jednia czasowa syntetyzowanego sygnału jest opisana;oraz obwiednia czasowa syntetyzowanego sygnału poddawana jest modyfikacji w oparciu o opisaną obwiednię czasową odpowiadającą kodowi parametrycznemu obwiedni i opisaną obwiednię czasową syntetyzowanego sygnału.
- 2425. Sposóh według zastrzeżenia 24, znamienny tym. że:funkcja skalowania generowana jest w oparciu o opisaną obwiednię czasową odpowiadającą kodowi parametrycznemu obwiedni i opisaną obwiednię czasową syntetyzowanego sygnału;oraz funkcja skalowania wykorzystywana jest do przetwarzania syntetyzowanego sygnału.
- 2526. Sposób według zastrzeżenia 20 znamienny tym, że obejmuje ponadto modyfikowanie przesyłanego kanału w oparciu o opisaną obwiednię czasowy w celu wygenerowania spłaszczonego kanaki, przy czyni spłaszczony kanał poddawany jest procesom upmixu i syntezy w celu wygenerowania odpowiadającego mu odtwarzanego kanaki.
- 2627. Sposób według zastrzeżenia 20 znamienny tym, że obejmuje ponadto modyfikowanie kanału poddanego procesowi upmixu w oparciu o opisaną obwiednie czasową w celu wygenerowania spłaszczonego kanału, przy czym spłaszczony kanał poddawany jest procesowi syntezy w· celu wygenerowania odpowiadającego mu odtwarzanego kanału.
- 2728. Sposób według zastrzeżenia 20, znamienny tym, że obwiednia czasowa syntetyzowanego sygnału poddawana jest modyfikacji jedynie dla określonych częstotliwości.
- 2829. Sposób według zastrzeżenia 28, znamienny tym. że obwiednia czasowa syntetyzowanego sygnału poddawana jest modyfikacji jedynie dla częstotliwości większych niż określona częstotliwość odcięcia.
- 2930. Sposób według zastrzeżenia 20, znamienny tym. że obwiednie czasowe poddawane są modyfikacji oddzielnie dla poszczególnych podzakresów pasma syntetyzowanego sygnału.
- 3031. Sposób według zastrzeżenia 20, znamienny tym, że dziedzina pod zakresu pasma odpowiada kwadraturowenni filtrowi zwierciadlanemu (QMF).
- 3132. Sposób według zastrzeżenia 20, znamienny tym, że obwiednia czasowa syntetyzowanego sygnału poddawana jest modyfikacji w dziedzinie czasu.
- 3233. Sposób według zastrzeżenia 20 znamienny tym, że obejmuje ponadto określenie czy należy włączyć czy wyłączyć proces modyfikacji obwiedni czasowej syntetyzowanego sygnału
- 3334. Sposób według zastrzeżenia 33, znamienny tym. że określenie oparte jest o znacznik włącz/wylącz generowany przez koder audio, który wygenerował E przesyłanych kanałów.
- 3435 Sposób według zastrzeżenia 33, znamienny tym, że określenie oparte jest o analizę E przesyłanych kanałów w celu wykrycia transjentów. dzięki czemu proces modyfikacji jest włączany w przypadku wykrycia cbccności transjentu.
- 3536. Sposób według zastrzeżenia 20 znamienny tym, że obejmuje ponadto:opisywanie obw iedni czasowej przesyłanego kanału;oraz określenie, czy w procesie modyfikacji obwiedni czasowej syntetyzowanego sygnału należy wykorzystywać (1) opisaną obwiednię czasową odpowiadającą kodowi parametrycznemu obwiedni, czy (2) opisaną obwiednię czasową przesyłanego kanału.
- 3637. Sposób według zastrzeżenia 20, znamienny tym, że moc w określonym okienku syntetyzowanego sygnału po przeprowadzeniu modyfikacji obwiedni czasowej jest równa mocy w odpowiadającym mu okienku syntetyzowanego sygnału przed przeprowadzeniem modyfikacji.
- 3738. Sposób według zastrzeżenia 37, znamienny tym, że określone okienko odpowiada okienku syntezy powiązanym z jednym lub większą liczbą kodów parametrycznych nic związanych z obwiednią.
- 3839. Urządzenie przeznaczone do dekodowania E przesyłanych kanałów audio w celu wygenerowania C odtwarzanych kinałów audio, gdzie C E^ . zawierające:elementy umożliwiające odebranie kodów parametrycznych odpowiadających E przesyłanym kanałom, gdzie kody parametryczne obejmują kod parametryczny obwiedni odpowiadający opisanej obwiedni czasowej kanału audio odpowiadającemu E przesyłanym kanałom, przy czym dwa lub większa liczba kodów- parametrycznych obejmuje ponadto jeden lub większą liczbę parametrów dotyczących stopnia statystycznej zgodności kanałów (ICC), parametrów- dotyczących międzykanalowcj różnicy natężeń (ICLD) i parametrów dotyczących międzykanalowcj różnicy czasu (ICTD), a pierwsza rozdzielczość w czasie powiązana z kodem parametrycznym obwiedni jest wyższa niż druga rozdzielczość w czasie powiązana z innym kodem(mi) paramelrycznym(i);elementy przeprowadzające proces upmixu E przesyłanych kanałów audio w celu wygenerowania jednego lub większej liczby kanałów poddanych procesowi upmixu;oraz elementy przeprowadzające proces syntezy jednego hib większej liczby C odtwarzanych kanałów przez zastosowanie kodów parametrycznych do jednego lub większej liczby kanałów poddanych procesowi upmixu, gdzie kod parametryczny obwiedni stosow any jest do kanalii poddanego procesowi upmixu lub kanału syntetyzowanego, w celu dokonania modyfikacji obwiedni czasowej syntetyzowanego kanału w- oparciu o opisaną obwiednię czasową, przy czym dokonywane est to przez skalowanie próbek sygnału w dziedzinie czasu lub w dziedzinie podzakresów pasma z wykorzystaniem współczynnika skalowania w taki sposób, że zmodyfikowana obwiednia czasowa jest zgodna z opisaną obw iednią czasową.
- 3940. Urządzenie według zastrzeżenia 39, znamienne tym, że urządzenie ma postać systemu wybranego z grupy obejmującej cyfrowy odtwarzacz wideo, cyfrowy odtwarzacz audio, komputer, odbiornik satelitarny, odbiornik kablowy, odbiornik wykorzystywany w transmisji naziemnej, domowe centrum multimedialne i system kina domowego:zaś system obejmuje odbiornik, urządzenie przeprowadzające proces upmixu, syntetyzator oraz urządzenie przeprowadzające modyfikację obwiedni.
- 4041. Wytwór stanowiący program komputerowy obejmujący kod programu, przy czym po wykonaniu kodu z wykorzystaniem odpowiedniego urządzenia, urządzenie implementuje sposóh dekodowania określony w zastrzeżeniu 20. Fraunhofer-Gesellschaft zur Fordcrung der Angcwandten Forschung e.V„ Niemcy Agcrc Systems, Inc., USA Pełnomocnik Sebastian Walkiewicz Rzecznik Pitentowy ΕΡ I 803 117 BI Z - 5968/09 FIG. 1 (STAN TECHNIKI) 100 MONOFONICZNY SYGNAŁ ŹRÓDŁOWY PRZESTRZENNE DLA ŹRÓDŁA LEWY SYGNAŁ AUDIO PRAWY SYGNAŁ AUDIO 200 ΕΡ 1 803 117 BI Z - 5968/09 FIG. 3 ΕΡ I 803 117 BI Z - 5968/09 FIG. ΕΡ 1 803 117 BI Z - 5968/09 FIG. 5 El’ 1 803 117 BI Z - 5968'09 FIG. 7Δ FIG. 7B EP 1 803 1 17 BI Z - 5968/09 co FIG. ΕΡ 1 803 117 BI Z- 5968/09 FIG. 10A INFORMACJE POMOCNICZE TP FIG. 108 1002 ΕΡ 1 803 I 17 BI Z - 5968/09 EP 1 803 117 BI Z - 5968/09 FIG. 12A 1204 FIG. 12B 1206 ΕΡ 1 803 I 17 BI Z - 5968/09 FIG. 13A INFORMACJE POMOCNICZE TP FIG. 13B 1302 ΕΡ 1 803 1 17 BI Z - 5968/09 CD IL EP I 803 117 BI Z - 5968/09 FIG. 15 EP 1 803 117 BI Z - 5968'09 FIG. 16 ΕΡ 1 803 117 131 Z - 5968/09 WSPÓŁCZYNNIKI LPC FIG. 17B ODWROTNA MODYFIKACJA OBWIEDNI (ITP) WSPÓŁCZYNNIKI WIDMOWE WSPÓŁCZYNNIKI WIDMOWE WSPÓŁCZYNNIKI LPC FIG. 17C MODYFIKACJA OBWIEDNI (TP) WSPÓŁCZYNNIKI WIDMOWE WSPÓŁCZYNNIKI WIDMOWE WSPÓŁCZYNNIKI LPC ΕΡ I 803 117 BI Z - 5968/09 FIG. 18A DO BLOKÓW ITP KANAŁ KANAŁ 2 KANAŁ FIG. 188 DO BLOKÓW ITP KANAŁ KANAŁ 2 KANAŁ
Independent claims40
181 paragraphs in 8 sections, as filed
Description
FIELD OF WYN / M.A7.KU | 0ϋ01I The subject matter of this application is related to the subjects of the following published United States patent applications:
- US 2003/0026441:
- US 2003/0035553:
- US 2003/0219130:
- US 2003/0236583:
- US 2005/0180579:
- US 2005/0058304:
- US. 2005/0157883: i
- US 2006/0085200.
| (MMI2 | The following publications are also related to the subject of this application:
- F. Baumgarte and C. Fallcr. "Binaural Cuc Coding - Part I: Psychoacoustic fundamenials and design principlcs' \ IEEE Trans, on Speech and Audio Proc., Toni 11. no. 6 November 2003;
- C. Fallcr and F. Baumgarte. ..Binaural Cue Coding - Part 11; Schcmcs and applications ''. IEEE Trans, on Speech and Audio Proc., Vol. I 1, No. 6 November 2003; and
- C. Faller, "Coding of spalial audio compalible wilh diHercnl playback fonnals". Preprint 117th Conv. And Eng Soc. October 2004.
Technical field | OO (I3] The present invention relates to the encoding of audio signals and the subsequent synthesis of sound scenes from the encoded audio data
State of the art | 0004 | When a human hears an audio signal (i.e. sounds) generated by a given source, the audio signal typically reaches the recipient's right and left ear at two different times and with two different sound intensities (e.g., expressed in decibels), 5 with these time and intensity differences are functions of the differences between the paths along which the audio signal reaches the receiver's right and left ear, respectively. The recipient's brain interprets time and intensity differences, giving them the impression that the sound being heard is generated by an audio source at a given location (for example, a specific direction and distance) in relation to the recipient. An audio stage is a phenomenon perceived by a person hearing multiple different audio signals simultaneously generated by one or more different sources of a cavity at one or more different positions relative to the listener.
| 0 (K) 5] The existence of the phenomenon of processing audio signals by the brain can be used for the synthesis of sound scenes, where the audio signals from one or more sound sources are appropriately modified to generate a left and right audio signal giving the impression of hearing various sound sources located in different positions in relation to the listener.
10006 | Figure I shows a high-level block diagram of a conventional binaural signal synthesizer 100 that converts a single audio source signal (e.g., mono signal) into left and right audio signals of the binaural signal, the binaural signal being the two signals received by the drums. listener's ears. In addition to the signal from the sound source, the synthesizer 100 receives a set of spatial 'parameters' corresponding to the desired position of the sound source in relation to the listener.
In typical implementations, the set of spatial parameters comprises a 1CLD inter-channel intensity difference value (corresponding to the left and right audio signal strength difference as perceived by the left and right and ear respectively) and an ICTD inter-channel time difference value (corresponding to the difference between arrival times). left and right audio signals as received by lcw'e and, respectively
3i) right and ear). Additionally or alternatively, some synthesis techniques involve modeling the cd-dependent direction of the sound transmission function from the sound source to the eardrums, also called the head related transfer function or head transfer characteristic (HRTF). See, e.g., J. Blauert, The Psychophysics of Human Sound Localization, Μ1Γ Press, 1983.
[0007] The mono audio signal generated by a single audio source is processed using the binaural signal synthesizer 100 shown in Fig. I such that when listening to it with headphones, the sound source is localized in space by using an appropriate set of spatial parameters. (eg JCLD. ICTD and / or HRTF) capable of generating an audible signal reaching each ear. See e.g. DR Bcgaulk 3-D Sound for Virtual Rcality and Multimedia, Academic Press, Cambridge, MA, 1994.
| 0008 | Binaural signal synthesizer 100 shown in Fig. I generates the simplest type of soundstage: involving a single sound source located relative to the listener. More complex soundstages involving two or more sound sources at different positions relative to the listener can be generated using a soundstage synthesizer, which is implemented essentially as a plurality of individual binaural signal synthesizers, each of the binaural signal synthesizers generating a binaural response signal. giving a different sound source. Since each sound source is in a different position relative to the listener, a different set of spatial parameters is used to generate a binaural audio signal for each sound source.
[0009] It is an object of the present invention to provide an improved method for coding an audio signal.
[0010 | This object is achieved by providing an audio signal coding method according to claim 1, an audio signal coding device according to claim 14, manufactured by constituting a computer program according to claim 17, an encoded audio data stream according to claim 18, an audio signal decoding method according to claim 20, an audio signal decoding device according to claim 39 and a computer program product according to claim 41. US 5,812,971 (HERRE) discloses methods for post-encoding a stereo signal that is part of multi-channel audio signals using one-time circuit shaping.
[0011 | In the document "The rcfcrcncc model architccture for MPEG spalial audio coding, Proc. 118th AES conveniion, 2005, paper 6447 by J. Hcrrc et al., A method is provided which makes it possible not to use a fully discrete representation of multi-channel audio, but provides a transmission rate compatible with that of a stereo signal which is only slightly higher than the rates commonly used for mono / stereo audio. In particular, OTT elements and TTT elements are used which are based on interchannel difference in intensity parameters and parameters of the degree of statistical concordance / cross correlation which represent a time / frequency match or cross correlation between the two input channels.
| 00l2 | The technical publication "Paramctric Coding of Spatial Audio", C. Faller, Proccedings of the 7th International Conferencc on Digital. Audio Effcci, Naples, Wiochy, October 5, 2004, pages 151 to 156 discloses the technique of BCC hinaural coding. In binaural BCC coding, stereo or multi-channel audio signals are represented by one or more downmixed audio channels and side information. The side information comprises interchannel parameters corresponding to the original audio signal that are essential for the perception of spatial properties of the sound scene. The relationship between the inter-channel parameters and the characteristics of a spatial sound stage is discussed here.
BRIEF DESCRIPTION OF THE DRAWINGS Other aspects, features and advantages of the present invention will become apparent in a flea upon reading the following detailed description, appended claims, and accompanying drawings, wherein the same or similar elements are identified by reference numerals.
Fig. I is a high-level block diagram of a conventional hinaural signal synthesizer;
Fig. 2 shows a block diagram of an audio signal processing system based on universal binaural coding (BCC);
Fig. 3 is a schematic block diagram of a downmix apparatus that can be used in the Fig. 2 embodiment;
Fig. 4 shows a block diagram of a BCC synthesizer that can be used in the decoder shown in Fig. 2;
Fig. 5 is a block diagram of the BCC estimator shown in Fig. 2, according to one embodiment of the present invention;
Fig. 6 shows the generation of an ICI D Interchannel Time Difference value and an Inter-channel Time Difference value ICTD for a five-channel audio signal;
Fig. 7 shows the generation of ICC concordance data for a five-channel audio signal;
Fig. 8 shows a block diagram of an implementation of the BCC synthesizer shown in Fig. 4, which can be used in a BCC decoder to generate a stereo or multi-channel audio signal based on a single screened sum signal s (h) and spatial parameters;
Fig. 9 shows how the interchannel values of the 1CLD intensity difference and the interchannel values of the ICTD time difference vary within a subband as a function of frequency;
Fig. 10 shows a block diagram of time-domain processing incorporated into a BCC encoder, such as the BCC encoder shown in Fig. 2, according to one embodiment of the present invention;
Fig. 11 shows an exemplary application of the time-domain processing (ΓΡ) in the context of the BCC synthesizer shown in Fig. 4;
Fig. 13 shows a block diagram of frequency domain processing input into a BCC encoder, such as the BCC encoder shown in Fig. 2, according to an alternative embodiment of the present invention;
Fig. 14 shows an exemplary application of frequency domain time processing (TP) in the context of the BCC synthesizer shown in Fig. 4;
Fig. 15 shows a block diagram of frequency domain processing input into a BCC encoder, such as the BCC encoder shown in Fig. 2, according to another alternative embodiment of the present invention;
Fig. 16 shows another exemplary application of a frequency domain time processing (TP) in the context of the BCC synthesizer shown in Fig. 4;
Figs. 17 (a) to Fig. 17 (c) show block diagrams of possible implementations of the time processing analysis (TPA) shown in Figs. 15 and Fig. 16 and inverse time processing (1TP) and time processing (TP) shown in Figs. 16.
Figs. 18 (a) and Fig. 18 (c) show two exemplary modes of operation of the control unit shown in Fig. 16.
DETAILED DESCRIPTION OF THE INVENTION | (1014 | In the case of binaural coding (BCC), the encoder performs C-coding of the input audio channels to generate E transmitted audio channels, where C> E ^ In particular, in the frequency domain, two or more C input channels are provided, and for each of one h and b of more different frequency ranges in the two or more time-domain input channels, one or more parameter codes are generated. The C input channels are further downmixed to generate the E transmitted channels. In some implementations of the downmix process. at least one of the E transmitted channels is based on two or more channels selected from the C input channels, and at least one of the E transmitted channels is based on only one of the C input channels.
101115] In one embodiment, the BCC encoder includes two or more filter banks, a code estimator, and a downmixer. Two or more filterbanks enable two or more channels selected from the C input channels to be converted from the time domain to the frequency domain. The code estimator generates one or more parametric codes for each of one or more different frequency ranges on the two or more transformed input channels. The downmixer downmixes the C channels
And 5 inputs to generate E transmitted channels, where C> E ^.
10016 | In BCC binaural encoding, the E transmitted audio channels are decoded to generate the C playback audio channels. Specifically, for each of one or more different frequency ranges, one h and more of the E transmitted channels undergo an upnxu frequency-domain process to generate two or more C reproduced frequency-domain audio 'channels, where C> E ^ In each of the one or more different frequency ranges in two or more frequency-domain channels being played back, one or more parametric codes are used to generate two or more modified channels, and two luh more of the modified channels are converted from the frequency domain to the time domain. In some implementations of the upmix process, at least one of the C playback channels is based on at least one of the E transmitted channels and at least one parametric code, and at least one of the C playback channels is based on only one of the E transmitted channels and is not dependent on any parametric codes.
| 0017 | In one embodiment of a BCC decoder it includes an upmixer, a synthesizer, and one or more inverse filter banks. In each of the one or more deviating frequency ranges, the upmixer performs an upmix of one or more E of the transmitted channels. in the frequency domain to generate one µh of more C playback channels, where C> E ^.
The synthesizer uses one or more parametric codes in each of one or more differing frequency ranges on two or more frequency-domain reproduced channels to generate two or more modified channels. By using one or more inverse filter banks, two or more modified channels are converted from the frequency domain to the time domain.
| 0018 | Depending on the implementation used, the playback channel may be based on a single transmitted channel rather than a combination of two or more transmitted channels. For example, in the case of only one transmitted channel, each of the C played channels is based on that one transmitted channel. In such situations, the upmix process corresponds to copying the corresponding transmitted channel. Thus, in applications where only one transmitted channel is used, the upmixer may be implemented as a replicator copying the transmitted channel to each of the playback channels.
[0019] BCC encoders and / or decoders may be used in a number of systems and applications including, for example, digital video recorders / players, digital audio recorders / players, computers, satellite transmitters / receivers, cable transmitters / receivers, transmitters / receivers used in in terrestrial broadcasting, home multimedia centers and home theater systems.
Universal processing BCC | 0020 | Fig. 2 is a block diagram of an audio signal processing system 20 (1 based on universal binaural coding (BCC) including an encoder 202 and a decoder 204. The encoder 202 subtracts the downmix device 206 and the BCC estimator 208.
| 00211 The downmixer 206 converts C input audio channels x, (/ j) into E transmitted audio channels y, (n), where C> E ^ 30 In this expression, the signals represented by the variable n are time domain signals, Whereas the signals represented by the variable k are frequency domain signals. BCC estimator 208 generates BCC codes based on the C input channels<sup>;</sup> audio and transmits these BCCs as support information included in or outside the E band of the transmitted audio channels. BCC codes typically include one or more inter-channel time difference (1CTD) values, inter-channel intensity difference (1CLD), and degree of statistical concordance (ICC) values calculated between certain pairs of input channels as a function of frequency and time. It depends on the implementation used, between which specific pairs of input channels the BCC codes are estimated. It depends on the implementation used between which specific pairs of input channels the BCC codes are estimated between.
[0022] The ICC data corresponds to the coherence of the binaural signal which is related to the perceived width of the sound source. The wider the sound source, the smaller the coherence between the left and right channels of the resulting binaural signal. For example, the coherence of the binaural signal corresponding to the orchestral arrangement on the stage is usually lower than the coherence of the binaural signal corresponding to the individual solo violins. Generally speaking, an audio signal with less coherence is usually viewed as being more widely distributed in the sound stage. Thus, ICC data is usually related to the apparent width of the source and the degree to which the listener is surrounded by the source. See J. Blauert, The Psychophysics of Human Sound Localization, MIT Press, 1983.
Depending on the specific application, the E transmitted audio channels and the corresponding BCC codes may be directly transmitted to the decoder 204 or stored in a suitable memory for later access to the data by the decoder 204. Depending on the solution used, the term "forwarding" it may mean directly sending to a decoder or storing data to enable the decoder to access the data later. In any event, the decoder 204 receives the transmitted audio channels and side information, and performs an upmixing and BCC synthesis process using BCC codes to convert E transmitted audio channels into more than E number (typically C, but not necessarily) reproduced audio channels χ ' ^ η) used during audio playback. Depending on the implementation used, the upmix process can be performed in the time domain or the frequency domain.
| O (I24 | In addition to the BCC processing shown in Fig. 2, an audio signal processing system based on universal binaural coding (BCC) may include additional encoding and decoding steps to achieve sequential compression of the audio signals at the encoder and decompression of the audio signals at the decoder, respectively. Such audio codecs may be based on conventional audio compression / decompression techniques, for example Pulse Code Modulation (PCM), Differential Pulse Code Modulation (DPCM), or Adaptive Differential Pulse Code Modulation (ADPCM).
| 0025 | If the downmixer 206 generates a single sum signal (li. E-1), the BCC coding allows the representation of multi-channel audio signals at a bit rate only slightly higher than that required for the representation of a mono signal. The reason is that the estimated ICTD, ICLD and ICC data for a channel pair contain about two orders of magnitude less information than the waveform of the audio signal.
[0026] Not only the low code rate of the BCC was taken into account, but also its backward compatibility. A single transmitted sum signal corresponds to a mono downmix of the original stereo or multi-channel signal. For odhorians who do not have the ability to reproduce stereo and multi-channel sound, listening to the transmitted sum signal is a sensible way to present the audio material using equipment capable of reproducing a mono signal. BCC coding can therefore be used to extend existing services by changing the distribution of a mono signal to a distribution of a multi-channel signal. If auxiliary BCC information can be included in an existing transmission channel, the monophonic broadcast systems can be extended for example to reproduce a stereo or multi-channel signal. The same possibilities exist for downmixing a multi-channel audio signal to two sum signals which correspond to a stereo sound.
A certain time resolution is used when processing audio signals using BCC coding, and often! And also the word. The frequency resolution used is highly dependent on the ability of human hearing to distinguish frequencies. Psychoacoustics suggest that spatial perception is most likely based on the representation of the critical bands of an acoustic input signal. The frequency resolution is obtained using a reversible filter bank (based, for example, on a Fast Fourier Transform (FFT) or a Quadrilateral Mirror Filter (QMF)) with subbands and bandwidths equal to or proportional to the critical bandwidth of human hearing.
Universal downmix process | 0028 | In preferred implementations, the transmitted sum signal (or sum signals) comprises all components of the input audio signal. The purpose of this solution is to push the behavior of each of the signal components. Simply summing the input audio channels often results in amplification or attenuation of signal components. In other words, the power of the signal components in the "simple sum" is often greater or less than the sum of the powers of the respective signal components in each of the channels. The downmixing technique that can be used allows to smooth the sum signal such that the power of the signal components in the sum signal is approximately the same as the corresponding power on all input channels.
Fig. 3 is a hlock diagram of downmix apparatus 300, which may be in some BCC implementations 20 (1 used as the downmix apparatus of Fig. 2) Downmix apparatus 300 includes a filter bank (FB) 302. for each of the input channels rfii), downmix 304. an optional scaling / delay block 306 and inverse filter bank (1FB) 308 for each of the coded input channels v, (u).
Each of the filter banks 302 converts each frame (e.g., 20 milliseconds) of its corresponding digital time-domain input channel to a set of frequency-domain input coefficients x, (k). Downmix 304 converts each of the subbands of the C band of the corresponding input coefficients into the corresponding E subbands of the downmixed coefficients in the frequency domain. Below is the equation (1) corresponding to the downmix process k of the subband of the input coefficients (xi (k), x? (K) ..... X ((k)) in order to generate the k subband of the downmixed coefficients (y / (k)) , y ^ k) ..... yE (k)}:
<td> >|(^)</td><td></td><td></td><td> ··</td>
<td>λ w</td><td>= D<sub>cht</sub></td><td> •</td><td>(l)</td>
<td>Ml</td><td></td><td>x<sub>c</sub>(<sup>k</sup>\</td><td></td>
where DiE is the downmix matrix u C over E o real values.
| 0 (J31 | The optional scaling / delay block 306 includes a set of muhiplicators 31 (1 each of which multiplies its corresponding downmixed factor yfk) by the scaling factor e<sub>vol</sub>(k) to generate the corresponding scaled coefficient y, (k). The purpose of the scaling operation is equivalent to a generalized downmix smoothing using any weighting factors for each channel. As the input channels are independent of each other, the power pi / k) of the downmixed signal in each of the sub-bands is given by the equation (2) below:
<td></td><td></td><td>Pi<sub>vol</sub>in</td>
<td></td><td rowspan="2">= D<sub>and</sub></td><td>P% (k'r</td>
<td> • •</td><td>and •</td>
<td>PFl (*)</td><td></td><td><sub>L.</sub>Pi<sub>C.</sub>IN.</td>
(2) where Der is obtained by squaring each element of the dowmnix matrix!) <# C on E, and p<sub>vl</sub>(ł) is the power of the subband k of the input channel i.
| 0032 | If the subbands are independent of each other, the power values p<sub>s;</sub>(k) the downmixed signal is greater or less than calculated from equation (2). which is due to the gain or attenuation occurring when the signal components are in the same or different phase, respectively. To avoid this, the downmix operation described by equation (I) is performed in the sub-band after the scaling operation is performed with the multipliers 510. The scaling factors e, (k) (1 d can be obtained using the following equation (3):
<sub>eiW</sub> pn (3) where pi-i (k) is the power of the subband calculated using equation (2), and Py><sub>vol</sub>(A) is the power of the corresponding downmixed subband signal Ϋ / β ·) · | 0033 | In addition to or in lieu of using an optional scaling process, the scaffold / delay block 306 may optionally introduce delay signals.
| O (I34 | The inverse filterbanks 308 each convert the set of corresponding nn-scaling coefficients y, (k) in the frequency domain to a frame of the corresponding transmitted digital channel r, (»).
[0035] Although Fig. 3 illustrates the transformation of all C input channels into the frequency domain for subsequent downmixing, in alternative implementations one or more (but less than Ct) of the C input channels may be omitted in some or all of the operations. the processing shown in Fig. 3 and transmitted as the same number of unmodified audio channels. Depending on the specific implementation, these unmodified audio signals may or may not be used by the BCC estimator 208 shown in FIG. 2 when generating the transmitted BCC parameters.
[0036] In an implementation of the downmixer 300 capable of generating a single sum signal j (?) E = \. and signals if and c) of each subband of each input channel c are added and then multiplied by the coefficient e (k) according to equation (4) shown below:
cy (k) = e (k) £ x<sub>c</sub>(k). (4) the coefficient e (k) is given by the equation (5) below:
Ea.wz * - 1
<img file="PL1803117T3_D0001.tif" />
where p?<sub>e</sub>(k) lo is the short-term power estimate. (A ') in the time index k, and Pi [k) is the short-term power estimate i<sup>AND</sup>«^ \ The aligned subbands are backconverted into the time domain to obtain a y (n) signal sent to the RCC decoder.
Universal BCC synthesis | 0037 | Fig. 4 shows a block diagram of a synthesizer 400 that may be used as a decoder 204 in Fig. 2 in some BCC system implementations 200. The BCC 400 includes a filterbank 402 for each of the transmitted y, (n) channels. Upmix block 404, retarders 406. multipliers 408, correlation block 4H) and inverse filterbank 412 for each of the reproduced channels v ', </ rj | 0038] Each filterbank 402 converts each frame of the corresponding transmitted digital channel yfii) in the time domain into a set of input coefficients) -, (4) in the frequency domain. The upniix 404 upmixes each of the subbands of the E band of the corresponding transmitted channel coefficients into the corresponding C subbands of the frequency domain coefficients. Equation (4) corresponds to the upmix process k of the subband of the transmitted channel coefficients (v / (4).
yi (k), yRA ')) to generate a k sub-range of the upmixed coefficients
U (H ·, UW), where lo takes place as follows:
<td colspan="2"></td><td rowspan="2">WITH(*)</td>
<td>AND(*)</td><td>- u</td>
<td>V *)</td><td></td><td></td>
(6) where Uff is the real-valued E to C upmix matrix. Performing the upmixing process in the frequency domain enables the upmixing process to be applied separately in each of the differing subbands.
| 0039 | Each of the delayers 406 introduces a delay value r /, W based on the corresponding BCC code for the ICTD data, whereby the desired ICTD values are obtained between some pairs of playback channels. Each of the multipliers 408 uses a scaling factor a ^ k) based on the corresponding BCC code for the ICLD data. whereby the desired ICLD values are obtained between some pairs of playback channels. Correlation block 410 performs dc-correlation operations A based on the appropriate BCC codes for the ICC data such that desired ICC values are obtained between some pairs of playback channels. A more detailed description of the operation of correlation block 410 can be found in US 2003/0219130.
[0040] The synthesis of the ICLD values can be less cumbersome than the synthesis of the ICTD and ICC values since the synthesis of ICLD only relies on scaling the subband signals. Since ICLD parameters are the most commonly used directional parameters, it is usually more important that the ICLD values are close to the parameters of the original audio signal. ICLD data can therefore be estimated between all channel pairs. The scaling factors a, (k} (1 for each subband are preferably selected such that the power of the subband for each of the reproduced channels' is close to the corresponding power of the original input audio kanaki.
| 0041 | One approach may be the use of a relatively small number of signal modifications to synthesize ICTD and ICC values. Thus, the BCC data may not include ICTD and ICC values for all channel pairs. In this case, the BCC synthesizer 400 synthesizes the ICTD and ICC values between only some channel pairs.
[0042] Each of the inverse filterbanks 412 converts the set of corresponding synthesized frequency domain coefficients into a frame of the corresponding digital reproduced audio channel Z, Q).
[0043] Although Fig. 4 shows the transformation of all E transmitted channels into the frequency domain for subsequent upmixing and BCC processing, in alternative implementations one or more (but not all) of the E transmitted channels may be omitted in * some or all of the processing operations shown in Fig. 4. For example, one or more the transmitted channels may be unmodified channels that are not upmixed. Besides being one or more C channels being played back, the unmodified channels may in turn (but not necessarily) be used as reference channels against which BCC processing is performed to synthesize one or more other playback channels. . In either case, unmodified channels may be delayed to compensate for the processing time in the ιιριηϊχιι and / or BCC processing processes to generate the remaining playback channels.
[0044] It should be noted that although Fig. 4 shows C playback channels' synthesized from E transmitted channels, where C is also the number of original input channels, the BCC synthesis is not limited to that number of playback channels. Generally speaking, the number of playback channels can be any number, including more or less than C, so even situations are possible where the number of playback channels is equal to or less than the number of transmitted channels.
"Perception differences between audio channels | 0045 | By assuming a single sum signal, the BCC processing allows a stereo or multi-channel audio signal to be synthesized such that the iCTD values. ICLD and ICC approximate the corresponding parameters of the original audio signal. The function of ICTD, ICLD and ICC values in relation to the spatial soundstage image attributes is discussed below.
[0046] It is known from the knowledge in the field of spatial hearing that for one auditory event the values of ISTD and ICLD relate to the audible direction. Considering the binaural room impulse responses (BRIR) of a single sound source, there is a relationship between the width of the auditory event and the listener's environment, and the ECC data estimated for the early and late portions of RR1R impulse responses The relationship between ICC data and these properties is nothing direct, however, for general signals (not only BRIR impulse responses).
Stereo and multi-channel audio signals typically contain a complex mixture of signals from simultaneously active sound sources superimposed on the reflected signal components resulting from the recording of sound in confined spaces or introduced by an audio engineer to artificially increase the spatial feeling. Different sound sources and their reflections occupy different areas on the time-frequency plane. It is reflected in the values
1CTD. 1CLD and ICC which vary with time and frequency. In this case, the relationship between the instantaneous values of ICTD, ICLD and ICC and the directions of sound events and the sense of spatiality is not obvious. The strategy of some BCC processing implementations includes blindly synthesizing these parameters such that they approximate the corresponding parameters of the original audio signal.
| 0048 | Filterbanks with under bandwidth ranges twice the equivalent rectangular bandwidth (CRR) are used. Face-to-face monitoring shows that the audio signal quality in BCC processing will not significantly increase if higher frequency resolution is used.
Lower frequency resolution may be desirable as this means fewer ICTD, 1CLD and ICC values that have to be transmitted to the decoder and therefore also lower bit rate.
| 0049 | In terms of resolution over time, ICTD values. ICLD and ICC are usually specified at equal intervals. High performance can be obtained if the ICTD, ICLD and ICC values are specified approximately every 4 to 16 milliseconds. Note that if nothing is specified at very short intervals, there is no precedence effect. Assuming the classical lead / lag pair of an audio stimulus, if lead and lag are within the time interval in which only one parameter set is synthesized, the dominance of lead localization may be disregarded. In addition, BCC processing enables a sound quality with an average MUSHRA score of about 87 (Ij "excellent sound quality) to almost 100 for some audio signals.
| (IO5O | The often obtained perceptual small difference between the reference signal and the signal obtained by synthesis means that the parameters 35 related to a wide range of spatial image attributes of the sound stage are unconditionally taken into account in the synthesis of ICTD, 1CLD and ICC values at equal intervals Below are some arguments showing the relationship of ICTD, ICI D and ICC values with a wide range of spatial soundstage image attributes.
Estimating Spatial Parameters | 0051] The method for estimating ICTD, 1CLD and ICC parameters is shown below. The bit rate sufficient to transmit these spatial parameters (quantized and encoded) is only a few khps, so that using DCC processing it is possible to transmit stereo and multi-channel audio signals using bit rates close to those required for transmitting a single audio signal.
[0052] Fig. 5 is a block diagram of the BCC estimator 208 shown in Fig. 2 according to one embodiment of the present invention. The BCC estimator 208 includes filter banks (FBs) 502, which may be the same as the filter banks 302 shown in Fig. 3, and cstiming block 504, which generates ICTD, ICLD, and ICC spatial parameters for each of the differing frequency subbands generated by filter banks 502
Estimate ICTD, ICLD and ICC values for stereo signals | 0053 | The following methods are used for ICTD parameters. ICLD and ICC respectively in the signals of the subband Xj (k) and x> (A) of the two (e.g. stereo) audio channels:
- ICTD [samples]:
= argmax {o (<sup>7</sup>>
with short-term estimation of the normalized cross-correlation function expressed by equation (8) below:
Pi Ad, k)
Φ.<sub>2</sub>(d, k I = -— i *. (8) ylPiSk-dJp ^ kd.) where
d. maxLi /, 0), ί, Λΐ d<sub>2</sub> - max (i /, 0] and 0<sub>Λ</sub>ΐχ2 (<Α is a short-term estimate of the mean ijA - dAki ^ k - di).
ICLD [dB]:
Δ £ υ (Α) = 10 log<sub>lo</sub> (10)
- ICC:
(11)
Note that the absolute value of the cross-correlation is taken into account and <? I2 (A) is in the range [0, 1],
Estimating ICTD, ICLD and ICC values for multi-channel audio signals | 0054 | If there are more than two output channels, it is usually sufficient to define ICTD and ICLD values between the reference channel (for example, channel # 1) and the other channels, as shown in Fig. 6 for the case of C = 5 channels, where<sub>2</sub>(A) denote the ICTD and ICLD values, respectively, between the reference channel and channel c.
In contrast to the values of ICTD and ICLD, the ICC parameter usually has more degrees of freedom. The defined ICC value may have different values between all possible pairs of input channels. In the case of C channels 'there are C (C1) / 2 possible channel pairs, e.g. for 5 channels' there are 10 channel pairs as shown in Fig. 7 (a). However, such a scheme requires that C (Cl) / 2 ICC values are estimated and transmitted for each subband in each time index, leading to high computational complexity and high bit rate.
| 0056 | Alternatively, the ICTD and ICLD parameters determine for each subband the direction in which an auditory event is represented corresponding to the signal component in the subband. Thus, one single ICC parameter per each subband may be used to describe the overall coherence between all audio channels. Good results can be obtained by estimating and transmitting only the ICC parameters between the two highest energy channels in each subband at each time index. It is shown in Fig. 7 (b), where at time Al and k the channel pairs (3, 4) and (I, 2) are the strongest, respectively. Heuristic functions may be used to determine ICC values between other channel pairs
Synthesis of Spatial Parameters [0057] Fig. 8 is a block diagram of an implementation of a BCC synthesizer
400 4, which can be used in a BUC decoder to generate a stereo or multi-channel audio signal based on a single transmitted sum signal s (n) and spatial parameters. The sum signal is decomposed into subbands, where s (k] denotes one subband. In order to generate the corresponding subbands on each of the output channels, delays d are applied to the corresponding subband of the sum signal.<sub>(</sub>, scale factors a, and filters h<sub>f</sub>. The time index A is omitted to simplify the notation of delays, scale factors and filters. ICTD values are synthesized by introducing delays, 1CLD values by scaling, and ICC values by applying dcorclosing filters. The slate processing shown in Fig. 8 is performed independently in each subband.
ICTD synthesis
[0058] Delays d. Are determined based on ICTD parameters r<sub>) f</sub>(A ') according to the following equation (12):
<sub>=</sub> ^(<sup>max</sup>! iisc <sup>+</sup> min<sub>2s / sc</sub> r<sub>h</sub>(Aj), c = I η<sub>;</sub>(ΑΗίή 2ic <C.
Reference channel delay d, is computed in such a way as to minimize the maximum absolute value of delays d<sub>c</sub>. The fewer subband signals are modified, the lower the likelihood of artifacts. If the sampling rate does not provide a sufficiently large time resolution for ICTD synthesis, delays can be introduced more accurately by using appropriate all-pass filters.
1CLD Synthesis | 0 & lt; 159] To obtain the desired ICLD values ΔΛι<sub>2</sub>(A) between channel c and reference channel 1, gain factors a<sub>e</sub> should meet the equation (13) presented below:
a - = 10 <sup>20</sup> (H) <sup>and</sup><
Additionally, the output subbands are in an advantageous development<sup>j</sup>and thus non-analyzing such that the sum of the powers of all the output channels is equal to the power of the input sum signal. Since the total power of the original signal in each of the subbands is preserved in the sum signal, such normalization allows obtaining the absolute power in the sub / akrcsic band in each of the output channels approximately equal to its power in the original input audio signal that is input to the encoder. . Taking these constraints into account, the gain factors ^ are defined by the equation (14) shown below:
<= 'o,) otherwise
ICC synthesis | 0060 | In some embodiments, the purpose of ICC synthesis is to reduce the correlation between subbands after introducing delays and applying scaling, while not affecting the 1CTD and 1CLD values. This can be achieved by designing the filters shown in Figure 8 such that the values of 1CTD and 1CLD are changed as a function of frequency so that the average variation in each subband is zero (critical shichu band).
Figure 9 shows how ICTD and ICLD values change into subband as a function of frequency. The amplitude of the change in ICTD and ICLD values determines the degree of decorrelation and is controlled as a function of ICC. Note that the ICTD values change smoothly (as in Fig. 9 (a)) and the ICLD values change randomly (as in Fig. 9 (b)). It is possible to change the ICLD values as smoothly as the ICTD values, but this changes the timbre of the resulting audio signals.
] 0062 | Another method of ICC synthesis, which is particularly useful for multi-channel ICC synthesis, is described in more detail in the publication by C. Fuller. "Parainctric mulii-Channcl audio coding: Synthesis ofcohcrcncc cucs". IEEE Trans, on Speech and Audio Proc., 2003.
[0063] In order to obtain the desired ICC value, specific amounts of late reverberation as a function of time and frequency are added to each of the output channels. Optionally, a spectral modification can be applied, so that the obtained linear envelope corresponds to the spectral envelope of the original signal.
| 0064 | Other related and unrelated techniques to synthesize ICC parameters for stereo signals (or pairs of audio channels) are provided in E. Schuijcrs, W. Oomcn. B. den Brinker and .1. Brccbaart,, .Advances in parainctric coding for high-qualify audio Prcprint 114th Conv. Aud. Eng. Soc., March 2003 and .1 Engdegard,
H. Pumhagcn, J. Rodcn and L Liljeryd, "Synthctic ambience in parametric stereo coding", Preprint 1 1 71b Conv. And Eng. Soc., May 2004.
BCC processing - C to E | 0065] As described above, BCC processing may be implemented using more than one forwarding channel. A variant of BCC processing is described in which the C audio channels are not represented as one single channel (transmitted) but as E channels, which is referred to as C to E BCC processing. There are (at least) two motives for using BCC C to E processing:
- BCC processing with one transmission channel enables to obtain backward compatibility ensuring the possibility of modernization of existing monophonic systems, thanks to which it is possible to play multi-channel or stereo audio signals The modernized systems transmit the sum signal obtained as a result of the BCC downmix process using the existing monophonic infrastructure, by transmitting additional BCC auxiliary information C to E BCC processing enables Z-encoded -channel with backward compatibility to C-channel coding.
the BCC-C to E processing introduces a scale in terms of different degrees of reduction of the number of transmitted channels. The more audio channels are transferred, the higher the expected sound quality.
BCC 'C to E processing details such as how to determine ICTD parameters. ICLD and ICC are described in US 2005/0157883.
Shaping a single channel | 0066 | In some implementations, both single channel BCC processing and C to E BCC processing use ICTD parameter synthesis algorithms. ICLD and / or ICO. Typically it is sufficient to synthesize the ICTD, ICLD and / or ICC parameters at intervals of between 4 and 30 milliseconds. The existence of the perception and priority effect phenomena, however, means that there are specific times when the human auditory system requires parameters with higher resolution over time (for example, synthesizing every 1 to 10 milliseconds).
[0067 | A single bank of static filters usually does not allow sufficient frequency resolution to be achieved while at the same time providing sufficient resolution over time at the moments where the priority effect takes effect.
| 0068 | Some embodiments of the present invention are designed for a system that utilizes a relatively low resolution when synthesizing ICTD parameters. ICLD and / hib ICC, with additional processing then being used at those times when higher resolution over time is required. In some embodiments of the invention, the system also eliminates the need to employ a signal dependent window switching technique which is usually difficult to integrate into the system. In some embodiments, the temporal envelopes of one or more original audio input channels input to the encoder are estimated. This can for example be done directly by analyzing the signal structure over time or by examining the autocorrelation of the signal spectrum as a function of frequency. Both of these approaches will be discussed below in relation to the cited implementation examples. If, from the perceptual point of view, it is required and beneficial, the information contained in these envelopes is transmitted to the decoder (as parameter codes of the envelope).
[0069] In some embodiments, the decoder uses some kind of processing to insert the desired temporal envelopes into the output audio channels:
This can be achieved by means of time processing (TP), for example by means of a signal envelope by multiplying the time domain signal samples with a time-varying amplitude modification function. If the time resolution of the subbands is sufficiently high (at the expense of a more coarse frequency-quality resolution), a similar kind of processing may be applied to the spectral samples of the 'subbands'.
Alternatively, the frequency-dependent spectral representation of the signals may be interleaved / filtered, this may be done analogously to the prior art for quantizing noise shaping in a low bit rate audio encoder or for increasing the intensity of the encoded stereo signals. This is an advantageous solution if the filter bank has a high frequency resolution and therefore a relatively low resolution. For the weaving / filtering approach:
- The envelope shaping method is extended from intensity stereo encoding to multi-channel encoding C to E.
The technique 1a comprises a configuration in which the envelope shaping process is controlled with encoder generated parametric information (e.g. binary flags), but in fact it is done with the decoder sets of filter coefficients.
In another configuration, sets of filter coefficients are transmitted from the encoder, for example only when necessary from the point of view of audio perception and / or advantageous.
This deer also applies to the time domain / subband heritage approach. It is therefore possible to introduce criteria (e.g. Iransicnt detection and lonality estimation) to further control the transmission of envelope infor- mation.
There may be situations where it is advantageous to turn off Time Processing (TP) to eliminate the possibility of artifacts. In any event, it is preferable to select a strategy that involves disabling temporal processing by default (i.e. BCC binaural encoding follows a conventional BCC encoding scheme). Additional processing is only turned on if it is expected that higher resolution channels need to be improved 'in time', for example when a priority effect is expected to occur.
| (J072] As noted above, the on / off control may be performed using transient detection. Thus, when a transient is detected, the time processing is enabled.<sup>r</sup>e (TP). When transients are present, the priority effect is strongest. Detection of transients can be applied ahead of time, which enables efficient shaping not only of individual transients but also of signal components shortly before and after the exposure of the transient. Possible methods for detecting transients include:
- Observing the temporal envelope of the input signals input to the BCC binaural encoder or the transmitted BCC sum signal (s). A sudden increase in power indicates the presence of a transient.
- Linear Predictive Coding (LPC) gain test. If the LPC prediction gain exceeds a predetermined threshold, it can be assumed that the signal is iransjenI or has large jitter. LPC analysis is computed by autocorrelation of the spectrum.
| 0073] To avoid the possibility of artifacts in the audio signals, the time processing {TP) is preferably additionally not applied if the lonality of the transmitted signal (s) is<sup>j</sup>) the total is high.
| 0074 | In some embodiments of the present invention, the temporal envelopes of individual original audio channels are simulated by the BCC binaural encoder, so that the BCC binaural decoder can generate output channels having a temporal envelope similar (or similar from an audio perception point of view) to the temporal envelope of the original audio channels. . Some embodiments of the present invention are directed to exploiting the priority phenomenon. In some embodiments of the present invention, the transmission of parametric envelope codes in addition to other BCC codes such as 1CTD, ICLD, and / or ICC parameters is used as part of the auxiliary BCC information.
[0075 | In some embodiments of the present invention, the time resolution of the envelope parameters over time is higher than the time resolution of other BCC codes (e.g., ICTD, ICLD, ICC). This enables the envelope shaping process to be performed in the time provided by the synthesis window corresponding to the hlok length of the input channel for which other BCCs are calculated.
Implementation examples | 0076 | Fig. 10 is a block diagram of time-domain processing incorporated into a BCC encoder, such as the encoder 202 of Fig. 2, according to one embodiment of the present invention. As it is shown in Fig. 1 () (a), each of the time processing analysis (TPA) analyzers 102 estimates the temporal envelope of another primary input channel [mu] Mh), but in general it is possible to analyze any or more input channels.
| 0077 | Fig. 10 (b) is a block diagram of one possible time-domain (TPA) implementation 1002 in which the original signal samples are squared (1006) to describe the temporal envelope of an input signal and then filtered using low pass filter (1008). In alternative embodiments of the invention, the temporal envelope may be simulated using autocorrelation / LPC or other methods, such as the Il-Lilbert transform.
[0078] In FIG. 10 (a), block 1004 performs the parameterization, quantization, and coding of the cstimated temporal envelopes before sifting through the time processing (TP) information (i.e., parametric envelope codes) that is contained in the information in FIG. 2. auxiliary.
| 0079 | In one embodiment of the present invention, block 1004 includes a sensor (not shown in the drawings) to determine if the time processing (TP) in the decoder will improve audio quality such that hlok 1004 transmits temporal processing (TP) assistance information only at this time. when time processing (TP) allows for higher sound quality.
| 0080 | Fig. 11 shows an exemplary application of the TP time-domain processing in the context of a BCC synthesizer 40 (1 shown in Fig. 4). In this embodiment, a single transmitted sum signal s (n) is used, C 'base channels are generated by rcplication the sum signal, z, and the envelope shaping is applied separately to the different synthesized channels.
In alternative embodiments, the order of the processes of delaying, scaling, and other types of processing may be different. In alternative embodiments of the invention, the shaping of the envelope is not further limited to processing each channel independently. This is particularly true for convolutional / filtering implementations where the coherence of the frequency ranges is used to obtain information regarding the precise temporal structure of the signals.
[0081] The decoding block 1102 shown in FIG.<sup>j</sup> based on the TP auxiliary information transmitted from the BCC binaural encoder, each of the time processing blocks 1104 using the respective envelope information in shaping the output channel envelope.
[0082] Fig. Llib) is a block diagram of one possible time-domain TP 110-1 implementation of the TP 110-1 time processing in which the basic signal samples are squared (1106) to describe the temporal envelope b of the synthesized signal, followed by filtering using a low-pass filter (1108). In this embodiment, a scale factor (1110) is generated (e.g., which is then applied (1112) to the synthesized channel to generate an output signal having a temporal envelope substantially corresponding to the temporal envelope of the original input channel. 0083 | In alternative time processing analysis (TPA) implementations ) 1 102 and Time Processing (TP) I 104 shown in Fig. II, the temporal envelopes are described using a zoom operation, nothing but the squaring of the signal samples. In such implementations, the slope a> b may be used as the scale factor, and there is no need to perform the square root operation.
J0084] Although the scaling operation shown in Fig. 11 (a) corresponds to a time domain (TP) implementation in the time domain, time processing (TP) (as well as time processing analysis (TPA) and inverse processing (TP) (ITP)) it may also be implemented using frequency domain signals as in the embodiment shown in Fig. 16-17 (described below). Thus, as used herein, the term "scaling function" includes time and frequency domain operations such as the filtering operations shown in Figs. 17 (b) and Fig. 17 (c).
[0 () 85] Generally speaking, the Time Processing (TP) 1104 is preferably designed such that it does not modify the signal strength (i.e. signal energy). Depending on the specific implementation, this signal strength may be the short term average signal strength of each of the channels, based, for example, on the total signal strength per channel over the time determined by the synthesis window, or derived from other available power measurement means. Thus, scaling for 1CLD synthesis (e.g. using multipliers 408) may be performed before or after the envelope shaping process.
[0086 | Since full-band scaling of the BCC output signals can cause artifacts, envelope shaping can only be applied at certain frequencies, for example frequencies greater than the specified cut-off frequency fn> (for example 500 Hz). Note that the frequency range for time processing analysis (TPA) may differ from the frequency range for synthesis (TP).
[U087] Fig. 12 (a) and (b) show possible implementations of the Time Processing Analysis (TPA) 1002 shown in Fig. 10 and the Time Processing (TP) 1104 shown in Fig. 11, with only the envelope shaping being used herein. at frequencies greater than the cutoff frequency fo. In fig. 12 (a), in particular, the introduction of a high-pass filter 1202 is visible, which allows frequencies lower than fw to be filtered out before the subsequent temporal envelope description is performed. In fig. 12 (b), the addition of a two-band filter bank 1204 with a cut-off frequency f0 lies between the two subbands, with only the higher frequency portion subjected to time shaping The dual band inverse filter bank 1206 allows the recombination of the lower frequency portion with the time shaping the part with higher frequencies, so that the output signal | 0088 | is generated In fig. 13 shows a hlock diagram of frequency domain processing input into a BCC encoder, such as encoder 202 shown in Fig. 2, according to an alternative embodiment of the present invention. As shown in Fig. 13 (a), the processing of each of the time processing analysis (TPA) 1302 is applied independently across subbands, with each filter bank (FB) being the same as the corresponding filter bank (FB). ) 302 shown in Fig. 3, and block 1304 is an implementation of subband processing analogous to block 1004 shown in Fig. 10. In alternative implementations, the subbands used in processing time-processing analyzes (TPA) may differ from the subbands used in binaural BCC encoding. it is visible in fig. 13 (b), TPA time processing analysis 1302 may be implemented in a manner analogous to the TPA time processing analysis (TPA) 1002 shown in Fig. 10.
Fig. 14 illustrates an exemplary application of a TP frequency domain time processing in the context of the synthesizer 400 shown in Fig. 4. The decoding block 1402 is analogous to the decoding block 1102 shown in Fig. 1, and each time processing is (TP) 1404 is an implementation of a subband analogous to each of the Time Processing (TP) 1404 shown in Fig. 11 as shown in Fig. 14 (b).
[0090] Fig. 15 is a block diagram of frequency domain processing incorporated into a BCC encoder, such as the encoder 202 shown in Fig. 2, according to another alternative embodiment of the present invention. This scheme has the following configuration: Envelope information for each input channel is obtained by computing linear LPC predictive coding as a function of frequency <1502) and then parameterized (1504), quantized (1506) and encoded into the data stream (1508) from using an encoder. Fig. 17 (a) is an example implementation of the Time Processing Analysis (TPA) 1502 shown in Fig. 15 The auxiliary information to be sent to the multi-channel synthesizer (decoder) may be LPC filter coefficients calculated by autocorrelation, coefficients of emerging reflections, linear spectral pairs, etc., or parameters derived, for example, from LPC prediction gain, e.g., binary tags "are present". / there are no iransients ", in order to obtain the smallest possible volume of auxiliary information data.
Fig. 16 shows another example application of Frequency Domain Time Processing (TP) in the context of the BCC 400 of Fig. 4. The encoding processing of Fig. 15 and the decoding processing of Fig. 16 may be used. be implemented such that they form a configuration with a matching encoder / dckcder pair. The decoding block 1602 is analogous to the decoding block 1402 shown in Fig. 14, and each of the TP time processing (TP) 1604 is analogous to each of the TP time processing 1404 shown in Fig. 14. In the case of the Lego multi-channel synthesizer, the time processing auxiliary information is transmitted. TPs are decoded and used to control the flow of the individual channel envelope shaping. However, the synthesizer additionally includes an envelope describing (TPA) step 1606 to analyze the transmitted sum signals, inverse time processing (ITP) 1608 to "flatten the temporal perimeter of each of the base signals, with a modified envelope being applied to each of the output channels from using envelope modifying elements (TP) 1604. Depending on the particular implementation, inverse time processing (ITP) may be performed before or after the upmixing process. More specifically, lo is performed using a convolution / filtering approach, where the spectral envelope shaping is performed by applying filters based on linear LPC predictive coding as a function of frequency, as shown in Fig. 17 (a), (b) and (c), with these figures relating to TPA, IIP and TP processing, respectively. In FIG. 16, control block 1610 enables it to be determined whether envelope purification should be in-guided, and if so, whether it should be based on (1) transmitted TP temporal processing auxiliary information or (2) locally obtained data describing the envelopes generated by the envelope description (TPA) stage 1606.
| 0092 | Figures 18 (a) and (h) illustrate two exemplary modes of operation for the control unit 1610 shown in Figure 16. In the implementation of Figure 18 (a), a set of filter coefficients for the decoder are transmitted to the decoder, and envelope shaping for: z. convolution / filling is carried out based on the sent factors. If a situation is detected where the encoder shaping Iransjent is not favorable, the filter data is not transmitted and the filters are turned off (shown in Fig. 18 (a) as a transition to a unit set of filter coefficients "[1, 0, ...] ").
| 0093 | In the implementation of Fig. 18 (b), only "tag and transient / no transient" for each of the channels is transmitted, this flag being used to turn shaping based on filter coefficient sets, computed by the decoder on based on the transmitted signals subjected to the downnix process.
Further Alternative Embodiments of the Invention Although the present invention has been described in the context of BCC encoding schemes in which a single sum signal is used, the present invention can also be implemented in the context of BCC encoding schemes using two or more sum signals. In such a case, the temporal envelope of each of the differing 'fundamental' sum 'signals' may be eslimated prior to the BCC synthesis, and the differing BCC output channels may be generated based on the differing temporal envelopes depending on from which sum signals were used in the synthesis of the different output channels. An output channel synthesized from two or more deviating sum signals may be generated based on the effective temporal envelope taking into account (e.g., using a weighted average) the relative effect of the sum channels.
| 0095 | Although the present invention has been described in the context of BCC coding schemes using 1CTD, <CLD and ICC codes, the present invention may also be implemented in the context of other BCC coding schemes using only one or two of these code types (e.g., ICTD and ICC but without using ICLD) and / h and b one or more additional code types. In addition, different BCC synthesis and envelope shaping sequences may be used in different implementations of the invention. For example, if envelope shaping is applied to frequency domain signals as shown in Fig. 14 and Fig. 16, alternatively, envelope shaping may be implemented after ICTD synthesis (in embodiments of the invention involving the use of ICTD synthesis). but prior to ICLD synthesis. In other embodiments of the invention, envelope shaping may be applied to upmixed signals prior to other BCC synthesis processes. Although the present invention has been described in the context of binaural BCC encoders generating parametric envelope codes based on the original input channels, in alternative embodiments, the parametric envelope codes may be generated from the downmixed channels. which correspond to the original input channels. This enables the implementation of a processor (e.g., a separate envelope parameter encoder) which may (1) use the output of the BCC binaural encoder generating downmixed channels and certain BCC codes (e.g., ICLD, ICTD and / or ICC) or (2) describe a temporal envelope. (s) one or more downmixed channels to couple the envelope parametric codes to the BCC support information.
| (J097 | Although the present invention has been described in the context of BCC coding schemes in which parametric envelope codes are transmitted in<sup>r</sup> one or more audio channels (i.e. E transmitted channels) together with other BCC codes, in alternative embodiments the envelope parametric codes may hyc be transmitted alone or together with other BCCs to a location (i.e. decoder or mass storage device),<sup>r</sup> which already has transmitted channels and possibly other BCC codes as well.
| 0098 | Although the present invention has been described in the context of BCC coding schemes, the present invention may also be implemented in the context of other audio processing systems where audio signals are dc-correlated or in other types of audio signal processing requiring signal decorrelation.
[0099] Although the present invention has been described in the context of implementations where an input time-domain audio signal is input into an encoder and the encoder generates the transmitted time-domain audio signals, the transmitted time-domain audio signals are input into the decoder and the decoder generates the playback. Time domain audio signals are not intended to limit the invention. In other implementations of the invention, any or additional input, transmitted and reproduced audio signals may, for example, be represented in the frequency domain.
| 0100 | BCC encoders and / or decoders may be used in conjunction with or incorporated into a wide variety of applications and systems, for example television and electronic music distribution systems, cinemas, broadcasting, streaming and / or receiving. These can be systems designed to encode / decode transmissions over, for example, terrestrial lines, satellite or cable lines, the Internet, Intranet, or physical media (e.g. CDs, DVDs, semiconductor circuits, hard drives, memory cards, and the like). BCC encoders and / or decoders can also be used in games or gaming systems, for example in interactive software systems involving interaction with the user in order to provide entertainment (action games, rpg, strategic, adventure, simulation, rally, sports, dexterity, card and board) and / or educational software that can be distributed on multiple devices, platforms and media. BCC encoders and / or decoders can furthermore be components of sound recorders / players or CD-ROM / DVD systems. BCC encoders and / or decoders may also be components of personal computer applications that include digital decoding (e.g., player, decoder) and digital encoding capabilities (e.g., encoder, ripper, recorder, or jukebox).
The present invention may be implemented as integrated processing circuits, e.g. as a single integrated circuit (such as an ASIC or FPGA), a module of multiple integrated circuits, a single card, or a packet of multiple cards. It is obvious to those skilled in the art that various functions of integrated circuits may also be implemented as processing steps as part of the software. Such software can be used, for example, in digital sound processors, microcontrollers, or general purpose computers.
| 0l () 2 | The present invention may take the form of methods- and devices that employ these methods. The present invention may also take the form of a program code stored on a physical medium such as floppy disks, CD-ROMs. hard disks or any other machine readable media whereby when the program code is entered into a device such as a computer it becomes an apparatus employing the invention. The present invention can also be in the form of a program code, whether or not stored in mass storage, "loaded into the device and / or executed by the device, or transmitted over any link or medium, for example via cables. and electric wires, optical fibers or electromagnetic radiation, whereby when a program code is entered into a device such as a computer, it becomes an apparatus using the invention. When the invention is implemented in a general purpose processor, the program code segments together with the processor provide a unique device that operates in a manner analogous to logic circuits.
| 0103 | Furthermore, it should be appreciated that those skilled in the art may make various changes to the details, materials, and arrangements of parts described and illustrated to explain the nature of the invention, without these changes departing from the scope of the invention as defined by patent claims.
| 0104 | Although in the following claims, the steps of the method are given in "specific order with the corresponding labels, unless a specific order of implementation of some or all of the steps is specified in the claim", the steps need not necessarily be implemented in that order.
Fraunhofcr-Gescllschaft zur Fórdcrung der Angcwandten Forschung cV, Germany 15 Agere Systems, Inc., L'SA
Dr. Sebastian Wtelldewicz's representative
Paltntowy Rucznlk
ΕΡ 1 803 117 BI
Z-5968/09
Contents8
30 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30
34 members in 21 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 62048004 | United States of America | P | |
| 62048004 | United States of America | P | |
| 648204 | United States of America | A | |
| 648204 | United States of America | A | |
| 05792350 | European Patent Office (EPO) | A | |
| 2005009618 | European Patent Office (EPO) | W | |
| 2005009618 | European Patent Office (EPO) | W | |
| EP20050792350 | – | – | – |
| US20040006482 | – | – | – |
| US20040620480P | – | – | – |
| WO2005EP09618 | – | – | – |
Members34
| Document | Office | Kind | |
|---|---|---|---|
| US2006083385A1 | United States of America | A1 | |
| AU2005299068A1 | Australia | A1 | |
| CA2582485A1 | Canada | A1 | |
| WO2006045371A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200628001A | Taiwan Province of China | A | |
| NO20071493L | Norway | L | |
| KR20070061872A | Republic of Korea | A | |
| EP1803117A1 | European Patent Office (EPO) | A1 | |
| MX2007004726A | Mexico | A | |
| IL182236A0 | Israel | A0 | |
| CN101044551A | China | A | |
| HK1106861A1 | Hong Kong, China | A1 | |
| JP2008517333A | Japan | A | |
| BRPI0516405A | Brazil | A | |
| AU2005299068B2 | Australia | B2 | |
| RU2339088C1 | Russian Federation | C1 | |
| EP1803117B1 | European Patent Office (EPO) | B1 | |
| AT424606T | Austria | T | |
| ATE424606T1 | Austria | T1 | |
| DE602005013103D1 | Germany | D1 | |
| PT1803117E | Portugal | E | |
| DK1803117T3 | Denmark | T3 | |
| ES2323275T3 | Spain | T3 | |
| PL1803117T3This record | Poland | T3 | |
| KR100924576B1 | Republic of Korea | B1 | |
| TWI318079B | Taiwan Province of China | B | |
| US7720230B2 | United States of America | B2 | |
| JP4664371B2 | Japan | B2 | |
| IL182236A | Israel | A | |
| CN101044551B | China | B | |
| CA2582485C | Canada | C | |
| NO338919B1 | Norway | B1 | |
| BRPI0516405A8 | Brazil | A8 | |
| BRPI0516405B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 1803117
- Publication, EPODOC
- PL1803117T
- Application
- 792350
- Application, DOCDB
- 05792350
- Application, EPODOC
- PL20050792350T
Titles2
- English
- INDIVIDUAL CHANNEL TEMPORAL ENVELOPE SHAPING FOR BINAURAL CUE CODING SCHEMES AND THE LIKE
- Polish
- Kształtowanie obwiedni czasowej pojedynczych kanałów w schematach kodowania binauralnego i podobnych
Classification
- CPC, 3
- G10L19/008
- G10L19/02
- H03M7/30
- IPC, 2
- G10L19 02
- G10L19 00