Enhanced coding and parameter representation of multichannel downmixed object coding
Abstract
An audio object coder for generating an encoded object signal using a plurality of audio objects includes a downmix information generator for generating downmix information indicating a distribution of the plurality of audio objects into at least two downmix channels, an audio object parameter generator for generating object parameters for the audio objects, and an output interface for generating the imported audio output signal using the downmix information and the object parameters. An audio synthesizer uses the downmix information for generating output data usable for creating a plurality of output channels of the predefined audio output configuration.

Term
1 yearto projected expiry
Projected expiry 5 October 2027, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1Zastrzeżenia claim 1. An audio synthesizer (104) for generating output data using a coded audio object signal (95, 97), comprising:1. Syntezator audio (104) do generowania danych wyjściowych przy użyciu kodowanego sygnału obiektów audio (95, 97), obejmuj ący: an output data synthesizer (100) for generating output data useful for rendering multiple output channels of a predefined audio output configuration, representing multiple audio objects, an output data synthesizer used to use downmix information, indicating the distribution of multiple audio objects in at least two downmix channels, information about power, correlation information, indicating power characteristics and correlation characteristics of at least two downmix channels (93) and audio object parameters for audio objects, where the output data synthesizer (100) is used to transcode (502) parameters of audio objects to spatial parameters for a predefined configuration output audio using the additionally intended audio object setting (90) in the audio output configuration. syntezator danych wyjściowych (100) do generowania danych wyjściowych, przydatnych do renderowania wielu kanałów wyjściowych predefiniowanej konfiguracji wyjściowej audio, reprezentującej wiele obiektów audio, syntezator danych wyjściowych używany do wykorzystywania informacji downmiksu, wskazujących rozkład wielu obiektów audio w co najmniej dwóch kanałach downmiksu, informacje o mocy, informacje korelacji, wskazujące charakterystyki mocy i charakterystyki korelacji co najmniej dwóch kanałów downmiksu (93) oraz parametry obiektów audio dla obiektów audio, gdzie syntezator danych wyjściowych (100) jest używany do transkodowania (502) parametrów obiektów audio do parametrów przestrzennych dla predefiniowanej konfiguracji wyjściowej audio przy wykorzystaniu dodatkowo zamierzonego ustawienia obiektów audio (90) w konfiguracji wyjściowej audio. 2. Syntezator audio według zastrzeżenia 1, w którym syntezator danych wyjściowych (100) jest używany do konwersji wielu kanałów downmiksu do downmiksu stereo dla predefiniowanej konfiguracji wyjściowej audio przy użyciu macierzy konwersji, wyprowadzonej z zamierzonego rozmieszczenia obiektów audio. 2. The audio synthesizer of claim 1, wherein the output data synthesizer (100) is used to convert multiple downmix channels to a stereo downmix for a predefined audio output configuration using a conversion matrix derived from the intended arrangement of the audio objects. 3. Syntezator audio według zastrzeżenia 1, w którym parametry przestrzenne obejmują pierwszą grupę parametrów dla up-miksu Two-To-Three i drugą grupę parametrów energii dla up-miksu Three-To-Six, oraz gdzie syntezator danych wyjściowych (100) jest używany do obliczania parametrów predykcji dla macierzy predykcji TwoTo-Three przy, użyciu macierzy renderu j ące j_,_ . określonej przez zamierzone rozmieszczenie obiektów audio (90) , macierzy częściowego downmiksu, opisującej downmiksowanie kanałów wyjściowych do trzech kanałów, generowanych przez hipotetyczny proces up-miksowania Two-To-Three, oraz macierzy downmiksu. 3. The audio synthesizer according to claim 1, wherein the spatial parameters comprise a first parameter group for a Two-To-Three upmix and a second energy parameter group for a Three-To-Six up-mix, and wherein the output data synthesizer (100) is used to calculate the prediction parameters for the TwoTo-Three matrix for using a matrix of rendering j _, _. determined by the intended arrangement of audio objects (90), a partial downmix matrix describing the downmixing of output channels to three channels, generated by a hypothetical Two-To-Three upmixing process, and a downmix matrix. 4. Syntezator audio według zastrzeżenia 3, w którym parametry obiektów są parametrami predykcji obiektów i gdzie syntezator danych wyjściowych (100) jest używany do wstępnego obliczania macierzy energii w oparciu o parametry predykcji obiektów, informacje downmiksu i informacje o energii, odpowiadające kanałom downmiksu. 4. The audio synthesizer of claim 3, wherein the object parameters are object prediction parameters and wherein the output data synthesizer (100) is used to pre-calculate the energy matrix based on the prediction parameters of the objects, downmix information and energy information corresponding to the downmix channels. 5. Syntezator audio według zastrzeżenia 1, w którym syntezator danych wyjściowych (100) jest używany do generowania (165) dwóch kanałów stereo dla konfiguracji wyjściowej stereo poprzez obliczanie parametryzowanej macierzy renderującej stereo i macierzy konwersji, zależnej od parametryzowanej macierzy renderującej stereo. The audio synthesizer of claim 1, wherein the output data synthesizer (100) is used to generate (165) two stereo channels for a stereo output configuration by calculating a parameterized stereo rendering matrix and a conversion matrix depending on the parameterized stereo rendering matrix. 6. A method of audio synthesis for generating output data using a coded audio object signal (95, 97), comprising: 6. Sposób syntezy audio do generowania danych wyjściowych przy użyciu kodowanego sygnału obiektów audio (95, 97), obej muj ąca: generating output data useful for creating multiple output channels of a predefined audio output configuration representing multiple audio objects (90), where downmix information is used, indicating distribution of multiple audio objects in at least two downmix channels, power information, correlation information, indicative characteristics the power and correlation characteristics of at least two downmix channels (93) and the parameters where the downmix generator (96) is 150) information about the power of audio objects for audio objects, and where audio object parameters are transcoded (502) to spatial parameters for a predefined audio output configuration using additional object settings, audio (90) in the audio output configuration. .... generowanie danych wyjściowych, przydatnych do tworzenia wielu kanałów wyjściowych predefiniowanej konfiguracji wyjściowej audio, reprezentującej wiele obiektów audio (90), gdzie używane są informacje downmiksu, wskazujące rozkład wielu obiektów audio w co najmniej dwóch kanałach downmiksu, informacje o mocy, informacje korelacji, wskazujące charakterystyki mocy i charakterystyki korelacji co najmniej dwóch kanałów downmiksu (93) oraz parametry gdzie generator downmiksu (96) jest 150) informacji o mocy obiektów audio dla obiektów audio, i gdzie parametry obiektów audio są transkodowane (502) do parametrów przestrzennych dla predefiniowanej konfiguracji wyjściowej audio przy wykorzystaniu dodatkowo zamierzonego ustawienia obiektów, audio (90) w konfiguracji wyjściowej ..audio. .. . 7. An audio object coder (101) for generating a coded audio object signal using a plurality of audio objects (90), comprising: 7. Koder obiektów audio (101) do generowania kodowanego sygnału obiektów audio przy użyciu wielu obiektów audio (90), obejmujący: a downmix information generator (96) for generating downmix information (97), indicating the distribution of multiple audio objects in at least two downmix channels, information configured to generate and correlation information, indicating power characteristics and correlation characteristics of at least two downmix channels (93) ;generator informacji downmiksu (96) do generowania informacji downmiksu (97), wskazujących rozkład wielu obiektów audio w co najmniej dwóch kanałach downmiksu, informacj i skonfigurowany do generowania i informacji o korelacji, wskazujących charakterystyki mocy i charakterystyki korelacji co najmniej dwóch kanałów downmiksu (93);an object parameter generator (94) for generating object parameters (95) for audio objects;and an output interface (98) for generating a coded signal (99) of audio objects, wherein the encoded object signal includes downmix information, power information, correlation information, and object parameters. generator parametrów obiektów (94) do generowania parametrów obiektów (95) dla obiektów audio;oraz interfejs wyjściowy (98) do generowania kodowanego sygnału (99) obiektów audio, gdzie kodowany sygnał obiektów obejmuje informacje downmiksu, informacje o mocy, informacje o korelacji i parametry obiektów. 8. Koder obiektów audio według zastrzeżenia 7, obejmujący poza tym: The audio object coder according to claim 7, further comprising: downmikser (92) do downmiksowanie wielu obiektów audio do wielu kanałów downmiksu, gdzie ilość obiektów audio jest większa niż ilość kanałów downmiksu, i gdzie downmikser (92) jest powiązany z generatorem informacji downmiksu tak, że rozkład wielu obiektów audio w wielu kanałach downmiksu jest wykonywany zgodnie z informacjami downmiksu. downmixer (92) for downmixing multiple audio objects to a plurality of downmix channels, wherein the number of audio objects is greater than the number of downmix channels, and where the downmixer (92) is associated with the downmix information generator so that the distribution of multiple audio objects in the multiple downmix channels is performed according to the downmix information. 9. Koder obiektów audio według zastrzeżenia 7, gdzie generator informacji downmiksu (96) jest używany do obliczania informacji downmiksu tak, że informacje downmiksu wskazują, który obiekt audio jest całkowicie lub częściowo dołączony do jednego lub wielu kanałów downmiksu, oraz przy użyciu wielu wskazujących rozkład o mocy i inf ormac j i charakterystyki mocy gdy obiekt audio jest dołączony do więcej niż jednego kanału_ downmiksu, informacji dotyczących czę.ści obiektów audio, dołączonych do jednego kanału downmiksu z więcej niż jednego kanałów downmiksu. The audio object coder according to claim 7, wherein the downmix information generator (96) is used to calculate the downmix information such that the downmix information indicates which audio object is fully or partially attached to one or more downmix channels, and using a plurality of indicative distributions. with power and / or power characteristics when the audio object is connected to more than one downmix channel, information about a part of audio objects connected to one downmix channel from more than one downmix channel. 10. Sposób kodowania obiektów audio (101) do generowania kodowanego sygnału obiektów audio obiektów audio (90), obejmująca: generowanie informacji downmiksu (97), wielu obiektów audio (90) w co najmniej dwóch kanałach downmiksu, generowanie (150) informacji o korelacji, wskazujących i charakterystyki korelacji co najmniej dwóch kanałów downmiksu;A method for coding audio objects (101) for generating a coded audio object object signal (90), comprising: generating downmix information (97), multiple audio objects (90) in at least two downmix channels, generating (150) correlation information;, indicating and correlation characteristics of at least two downmix channels;generating object parameters (94) for audio objects;generowanie parametrów obiektów (94) dla obiektów audio;and generating a coded signal (99) of audio objects, wherein the encoded object signal includes downmix information, power information, correlation information, and object parameters. oraz generowanie kodowanego sygnału (99) obiektów audio, gdzie kodowany sygnał obiektów obejmuje informacje downmiksu, informacje o mocy, informacje o korelacji i parametry obiektów. 11. A coded audio object signal including downmix information indicative of the distribution of a plurality of audio objects in at least two downmix channels, power information and correlation information, indicating power characteristics and correlation characteristics of at least two downmix channels, and object parameters such that reproduction Audio objects are possible using object parameters and at least two downmix channels. 11. Kodowany sygnał obiektów audio, obejmujący informacje downmiksu, wskazujące rozkład wielu obiektów audio w co najmniej dwóch kanałach downmiksu, informacje o mocy i informacje o korelacji, wskazujące charakterystyki mocy i charakterystyki korelacji co najmniej dwóch kanałów downmiksu, oraz parametry obiektów takie, że odtwarzanie obiektów audio jest możliwe przy użyciu parametrów obiektów i co najmniej dwóch kanałów downmiksu. 12. Kodowany sygnał obiektów audio według zastrzeżenia 11, zapisany na nośniku odczytywanym przez komputer. 12. The encoded audio object signal according to claim 11, stored on a medium read by the computer. 13. A computer program for executing, when used on a computer, methods according to any one of claims 6 to 10. 13. Program komputerowy do wykonywania, podczas użycia w komputerze, metody według któregokolwiek z zastrzeżeń od 6 do 10. DOLBY INTERNATIONAL AB, Holandia DOLBY INTERNATIONAL AB, the Netherlands Pełnomocnik: Proxy: jtazewskj J Patentowy jtazewskj J Patentowy EP 2 068 ,-8945 EP 2 068, -8945 307 Bl 307 Bl N obiektów N objects Encoder parameters Parametry kodera Macierz renderującą Rendering matrix FIG1A Fig1 CO WHAT CD CD LL64 LL64 FIG 2 FIG 2 Encoder parameters Parametry kodera Γ I I y mjuci Γ II y mjuci F4 301 F4 301 302 .202 302.2202 304 304 mono (obiektów i mono (objects and Data grouping Grupowanie danych FIG 3 FIG 3 202 π 202 π FIG 4 FIG 4 102 102 FIG 5 FIG 5 Downmiks obiektów macierz konwertera Downmix objects of the converter matrix FIG 6 FIG 6 FIG 7 FIG 7 801 801 802 802 L L R obj5 R obj5 FIG 8 FIG 8 FIG 9 FIG 9 Audio synthesizer Syntezator audio predefined output configuration predefiniowanej konfiguracji wyjściowej FIG 10 FIG 10 ΊΊ ΊΊ '111 '111 -112 -112 113 113 114 114 115 115 116 116 FIG 11 FIG 11 FIG 12 FIG 12 CLD, cld2 CLD, cld2 ICC ICC, ICC2 ICC2 FIG 13A FIG 13A FIG 13B FIG 13B 142 / 142 / N obiektów N objects SS * = E SS*=E N> K> 2 N>K>2 K downmik channels 143 K kanałów downmiksu 143 X = DS \ X=DS \ M kanałów wyjściowych M output channels V = AS 1. V=A-S 1 . S ~ CX = CDS S~CX=CDS FIG 14 up-mikser przestrzenny FIG 14 up-spatial mixer 2 K prediction parameters of objects for reproducing sources by S «C» X 2K parametry predykcji obiektów do odtwarzania źródeł przez S«C»X FIG 15 (3 sources) 'μ G-μ) (1-κ) κ FIG 15 (3 źródła) ' μ G-μ) (1-κ) κ FIG 16 FIG 16
298 paragraphs in 3 sections, as filed
The present invention relates to decoding composite (harmonic) sound objects on the basis of a coded composite object signal based on the available downmix of multi-channel sound and additional control data.
BACKGROUND OF THE INVENTION [0002] The current development of audio technology allows the generation of a multi-channel analogue of an audio signal based on a stereo (or mono) signal and corresponding control data. These parametric methods of encoding surround channels usually involve their parameterization. A parametric multi-channel audio decoder (eg MPEG Surround decoder defined in ISO / IEC 23003-1 [1], [2]) reproduces, using additional control data, M channels based on K transmitted channels, where M> K. The control data includes parameterization of the multi-channel signal based on the IID parameters (Inter Channel Intensity Difference) and ICC (Inter Channel Coherence - statistical compliance of channels). These parameters are generally achieved at the coding stage, and they describe the power ratios and the correlation between channel pairs used in a process called up-mix. The use of such a coding scheme involves encoding at a significantly smaller data stream size per unit time than is the case when transmitting all ____ M channels. As a result, the coding becomes very efficient and at the same time ensures compatibility with both KANAL and M-channel devices.
[0003] A largely similar coding system is a similar encoder of audio objects [3], [4] where several audio objects are subjected to a downmix process at the encoder, and later to an upmix process performed according to the control data. The up-mix process can also be seen as separating objects that are mixed up in the downmix process. The signal created as a result of an up-mix can be rendered to one or more playback channels. Specifically, [3,4] shows how to synthesize audio channels based on the downmix (referred to as an overall signal), statistical data about source objects and data describing the desired output format. When several downmix signals are used, they consist of different subsets of objects and the upmix is performed independently for each downmix channel.
In the described new method, a method is introduced that allows an upmix to be made jointly for all downmix channels. Prior to this invention, object coding methods did not allow joint decoding of a downmix composed of more than one channel.
Literature:
[1] L. Villemoes, J. Herre, J. Breebaart, G. Hotho, S.
Disch, H. Pumhagen and K. Kjórling, "MPEG Surround: The Forthcoming ISO Standard for Spatial Audio Coding," at the 28th International AES Conference, The Futura of Audio Technology and Beyond, Pitea, Sweden, June 30 - July 2, 2006.
[2] J. Breebaart, J. Herre, L. Villemoes, C. Jin, K. Kjórling, J. Plogsties and J. Koppens, "Multi-Channels goes Mobile: MPEG Surround Binaural Rendering," on 29
International AES Conference, Audio for Mobile and Handheld Devices, Seoul, September 2-4, 2006.
[3] C. Faller, "Parametric Joint-Coding of Audio Sources," Convention Paper 6752 presented at 120 AES Convention, Paris, France, May 20-23, 2006.
[4] C. Faller, "Parametric Joint-Coding of Audio
Sources, "Patent Application PCT / EP2006 / 050904, 2006. [0005] WO 2006/048203 A2 discloses concepts for improving multi-channel prediction based reconstruction. In particular, multi-channel reconstruction takes into account the predicted energy loss occurring in the up-mix process. Specifically, the original left channel, the original middle and primary right channel being downmixed, are converted to the left downmix channel and the right downmix channel, and the left downmix channel contains only the original left channel and part of the original middle channel and the right downmix channel contains only the original right channel and part of the original middle channel This is determined by the downmix matrix The two base channels are transmitted together with two different up-mix parameters to the mixer performing the up-mix process from the energy mode.The primary left, right and center channels are reproduced, and then they are subjected to energy correction to obtain corrected left, right and center channels.
[0006] The object of the present invention is to provide an improved method for encoding and decoding audio objects. . ______________________. - ___. [0007] This object is achieved by using an audio synthesizer according to claim 1, the audio signal synthesis method according to claim 6, the audio object encoder according to claim 7, the audio object encoding method according to claim 10, the encoded audio object signal according to claim 11 or the computer program according to claim 13.
SUMMARY OF THE INVENTION [0008] A first aspect of the invention relates to an audio object encoder for generating a coded audio object signal using a plurality of such objects. The encoder consists of: a downmix generator designed to generate a downmixed signal, showing the distribution of multiple audio objects to at least two downmix channels; the object parameter generator used to generate the parameters of the audio objects: and the output signal interface used to generate the signals of the encoded audio objects using the information contained in the downmix and the parameters of the object.
[0009] A second aspect of the invention relates to a method of encoding audio objects for generating a coded audio object signal using a plurality of such objects. The method includes generating a downmix signal showing the distribution of multiple audio objects into at least two downmix channels; generating object parameters for audio objects, as well as generating a coded signal of audio objects using the information contained in the downmix and object parameters.
[0010] A third aspect of the invention relates to an audio synthesizer generating output signal data 5. using coded. signal of _audio objects.
This synthesizer includes: an output data synthesizer for generating output data, based on which it is possible to create multiple output channels of a predefined output audio configuration, reflecting a multitude of audio objects, an output data synthesizer, which uses the information contained in downmix, revealing the distribution of many objects audio for at least two downmix channels, and audio object parameters appropriate for coded objects.
[0011] A fourth aspect of the invention relates to a method of synthesizing audio that generates output data using a coded audio object signal. It includes: generating output data that can be used to create multiple output channels of a predefined output audio configuration, reflecting a plurality of audio objects, an output synthesizer based on information contained in a downmix, exposing the distribution of multiple audio objects to at least two downmix channels, and parameters an audio object for these objects.
[0012] A fifth aspect of the invention relates to a coded audio object signal comprising information contained in a downmix, showing the distribution of multiple audio objects into at least two downmix channels, and object parameters. The object parameters have the property that they allow reconstruction of audio objects using these parameters and at least two downmix channels. A sixth aspect of the invention relates to a computer program that implements, as it operates, a method for encoding audio objects or decoding them.
BRIEF DESCRIPTION OF THE FIGURES _______________________ The present invention will now be described by means of examples illustrating the non-limiting scope and nature of the invention with reference to the accompanying illustrations in which:
Fig. 1a illustrates the operation of spatial encoding of audio objects including coding and decoding,
Fig. Ib shows the operation of spatial encoding of audio objects using an MPEG Surround decoder, Fig. 2 shows the operation of an encoder for spatial coding of audio objects,
Fig. 3 shows the operation of an audio object parameter extractor operating in an energy mode;
Fig. 4 shows the operation of an audio object parameter extractor operating in a prediction-based mode, Fig. 5 shows the construction of a transcoder performing a conversion from SAOC spatial encoding to MPEG Surround
Fig. 6 shows different modes of operation of the downmix converter,
Fig. 7 shows the structure of a MPEG Surround decoder performing a stereo downmix process;
Fig. 8 is an example of a practical application using an SAOC encoder;
Fig. 9 is an embodiment of an encoder
Fig. 10 is an embodiment of a decoder
Fig. 11 is a table of various preferences regarding coder / synthesizer operating modes;
Ί
Fig. 12 shows a method for calculating certain spatial up-mix parameters;
Fig. 13a shows a method for calculating additional parameters of a spatial up-mix;
Fig. 13b. preset_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a_a
Fig. 14 shows a general view of an encoder / decoder system; Fig. 15 shows a method for calculating object prediction parameters; and
Fig. 16 shows a method for reproducing a stereo signal.
DESCRIPTION OF PREFERRED EMBODIMENTS OF THE INVENTION [0014] The following described embodiments are only intended to illustrate the principles of the present invention regarding improved coding and parameter reproduction in multi-channel encoding of objects subjected to the downmix process. It is assumed that modifications and changes introduced in the solutions and details described herein will be understandable to those familiar with this field. For this reason, the only limiting factor is the nature of the claims, not the specific details presented here, related to the description and explanation of the operation of the embodiments of the invention used herein.
[0015] The preferred embodiments show a coding scheme that combines the functionality of the object encoding scheme with the reproduction capabilities of a multi-channel decoder. Transmitted control data are associated with individual objects and therefore allow manipulation of their reproduction in terms of spatial position and level. Thus, the control data are directly related to the so-called scene description, providing information on the positioning of objects. The scene description can be controlled interactively by the listener at the decoder level, or by the producer at the encoder level.
The transcoding step given in the invention is used to process control data related to the object and a downmix signal to the control data. I. S. Signal .. downmix corresponding to the reproduction system, e.g. the MPEG Surround decoder.
[0016] In the coding scheme shown, the objects can be arbitrarily distributed over the available downmix channels in the encoder. The transcoder clearly uses the information contained in the multichannel downmix, producing a transcoded down-mix signal and control data related to the object. In this way, the upmix process in the decoder is not executed for each channel individually, as described in [3], but all downmix channels are subjected simultaneously to one upmix process. The new information scheme contained in the downmix must be part of the control data and is coded by the object encoder.
[0017] The decomposition of objects on the downmix channel space can be done automatically or according to the choice of pattern made on the encoder side. In the latter case, a downmix can be designed to be reproducible by the existing multichannel playback scheme (e.g.
stereo reproduction system). Then, playback occurs, omitting the transcoding and multi-channel decoding stages. This is an additional advantage over earlier coding schemes consisting of one downmix channel or multiple downmix channels containing subsets of source objects.
[0018] While older object encoding schemes only determined the decoding process using a single downmix channel, the present invention is not limited in this way because it provides a way to jointly decode downmixes containing more than one downmix channel. The quality of the objects obtained when separating is higher due to the increased number of downmix channels. Therefore, the invention effectively fills the gap between the coding scheme of objects with one mono downmix channel and a multi-channel coding scheme in which each object is transmitted via a separate channel. The proposed scheme thus allows flexible adaptation of quality during the separation of objects, in accordance with the application requirements and the properties of the transmission system (such as channel capacity).
[0019] In addition, the use of more than one downmix channel is advantageous because it allows for additional consideration of correlations between individual objects, as opposed to limiting the description to differences in intensity, as was the case with older object coding schemes. Earlier schemes are based on the assumption that all objects are independent and mutually unrelated (zero cross correlation), while in reality objects are often related, e.g. right and left stereo channel. By including such interdependence in the description (control data) as given in this invention, it becomes more complete and thereby further supports the ability to separate objects. [0020] A preferred embodiment includes at least one of the following features:
[0021] A system for transmitting and creating multiple individual audio objects using a multi-channel downmix and additional control data describing objects comprising: an encoder for spatial audio objects for encoding multiple audio objects into a multi-channel downmix, multi-channel downmix information and object parameters; or a spatial audio object decoder for decoding a multi-channel downmix, multi-channel downmix information, object parameters and an object rendering matrix into a second multichannel audio signal. suitable for audio reproduction.
[0022] Fig. 1 shows the operation of spatial audio object coding (SAOG) encompassing the SAOG 101 encoder and the SAOC 104 decoder. The spatial audio object encoder 101 encodes N objects creating their downmix consisting of K> 1 audio channels, depending on encoder parameters. Information on the applied weight of the downmix matrix D is passed through the SAOC encoder together with possible data on the power and correlation of the downmix. The weight of the D matrix often, but not always, remains constant in time and frequency and is therefore a relatively limited source of information. Finally, the SAOC encoder extracts the parameters of each object as a function of both time and frequency in the resolution determined by perception. The spatial audio object decoder 104 uses information in the downmix, downmix channels and object parameters (generated by the encoder) as an input signal and generate an output signal using M audio channels to present it to the recipient. The rendering of N objects on M audio channels is done using a rendering matrix that acts as an input signal for the SAOC decoder.
[0023] Fig. 1b shows the operation of spatial coding of audio objects when re-using the MPEG Surround decoder. The SAOC 104 decoder in the sense of the invention may be a SAOC transcoder for MPEG Surround 102 and an MPEG Surround 103 decoder based on a downmixed stereo signal. A user-controlled matrix A, in size
Μ χ N specifies the target rendering of N objects on N audio channels. This matrix depends on both time and frequency, and is the final output of a more user-friendly interface for manipulating "audio _________" objects (which also __________ can use the external description of the scene). If you set the speakers in the 5.1 system, the number of audio channels of the output signal is M = 6. The task of the SAOC decoder is to reproduce the signal that perceptually coincides with the initial audio objects. The SAOC transcoder on MPEG Surround 102 as the input signal treats the rendering matrix A, the downmix of the object, additional information contained in the downmix, including the weight of the downmix matrix D, and additional information about the object and generates a stereo downmix and additional MPEG surround information.
[0024] The SAOC decoder in the sense of the invention consists of a SAOC transcoder for MPEG 102 and an MPEG surround decoder 103 for stereo downmixes. The user-controlled matrix A, in size Μ χ N, specifies the target rendering of N objects on M audio channels. This matrix can be based on both time and frequency and is the final output signal of a more user-friendly interface for manipulating audio objects. If you set the speakers in the 5.1 system, the number of audio channels of the output signal is M = 6. The task of the SAOC decoder is to reproduce the signal that perceptually coincides with the initial audio objects. The SAOC transcoder on MPEG Surround 102 treats as an input signal the rendering matrix A, the downmix of the object, additional information about the downmix, including the weight of the downmix matrix D, and additional information about the object and generates a stereo downmix and additional MPEG surround information. When the transcoder is built according to the understanding of the invention, then the MPEG Surround 103 decoder used, having obtained this data, will be able to produce an M channel audio output signal with the desired properties.
[0025] Fig. 2 shows the operation of a spatial audio object encoder (SAOC) 101 in the sense of the invention. The extractor 201 and the extractor of audio objects 202 are fed N audio objects. Downmix 201 mixes these objects, which creates a downmix of objects consisting of K> 1 audio channels, according to the encoder parameters, as well as information about the downmix output sources. This information includes a description of the applied weight of the downmix matrix D, and possibly if the target audio object parameter extractor works in the prediction mode, parameters describing the power and correlation of the objects that have undergone the downmix process. As discussed in the next paragraph, the role of such parameters is to provide energy and correlation of subsets of processed audio channels in a situation where when the parameters of the objects are expressed only in relation to the downmix, the best example being the back / front directions in the 5.1 speakers setup. The parameter extractor for audio objects 202 extracts the parameters of the object in accordance with the encoder parameters. The encoder control determines on a variable time and frequency basis which of the two coder modes is used - the mode based on energy or prediction.
In the energy-based mode, the encoder parameters further include information about grouping N audio objects of the line combination
Prediction of the object in P stereo objects and N-2P mono objects. Each of the modes will be further described in figures 3 and 4.
[0026] Fig. 3 shows the operation of an audio object parameter extractor 202 operating in an energy-based mode. Grouping of 301 into P stereo objects and N-2P objects, mono is made on the basis of information contained in encoder parameters. For each frequency-time interval considered, the following operations are performed: two object power parameters are obtained and one normalized correlation for each of the P stereo objects using the stereo parameter extractor 302. One power parameter is obtained for each of the N-2P mono objects using Extractor Mono 303. The total set of N power parameters and P normalized correlation parameters are coded in module 304, together with data on the grouping, resulting in object parameters.
[0027] Fig. 4 shows an audio object parameter extractor 202 based on prediction. The following operations are performed for each specific frequency-time interval. For each N-channel object, a linear combination of the downmix channels of object K is determined, which combination corresponds to the object at least in the least squares sense. The weights of the K channel in this are called Object Prediction Coefficients (OPCs) and they are calculated by the OPC 401 extractor. The full set of NK OPCs is coded in 402 to determine the parameters of the object. The coding may involve a reduction in the total amount of OPCs based on linear interdependencies. As is apparent from the present invention, this total amount can be reduced to a maximum of {Κ · (NK), 0} as long as the weight of the downmix matrix D has a full order.
[0028] Fig. 5 shows a SAOC structure for an MPEG Surround 102 transcoder in the sense of the present invention. For each "frequency / time interval", the additional information of the downmix and the object's parameters are connected to the rendering matrix by means of the 502 parameter calculator in order to generate MPEG Surround parameters CLD, CPC, ICC and the up 2xix matrix down-pixel converter. The downmix converter 501 converts the downmix of the object into a stereo downmix by applying a matrix operation according to the G matrix. In the simplified transcoder mode for K = 2, the matrix is an identity matrix, and the downmix of the object is transmitted in unchanged form as a stereo downmix. This mode is shown in the drawing with selector switch 503 in position A, where the normal operating mode has a switch in position B.
[0029] Fig. 6 illustrates various operational modes of the downmix converter 501, in the sense of the present invention. If a downmix of an object in a bitstream output format is transmitted for the K-channel audio encoder, the bit stream is first decoded into the temporal domain sound signals by the audio decoder 601. Next, all signals take the form of a frequency domain due to the hybrid filtering assembly QMF MPEG Surround in T / F unit 602. The matrix time and frequency differentiation operation defined by the data of the converting matrix is performed in the hybrid QMF domain signals achieved by the maturation unit 603, which outputs the stereo signal in the hybrid QMF domain. The hybrid synthesis unit 604 converts the stereo signal of the hybrid domain QMF into a stereo signal of the QMF domain. The hybrid QMF domain is defined for ... _by. achieve better .. resolution, frequency, with lower frequency setting, by the following QMF subband filtering. When filtering is defined by Nyquist filters, the conversion of the QMF domain from hybrid to standard consists only of the sum of the sum of hybrid signal sub-groups, see [E. Schuijers, J. Breebart, and H. Purnhagen "Low complexity parametric stereo coding" Proc 116 AES Convention in Berlin, Germany 2004, Preprint 6073]. This signal creates the first possible downmix converter output format, as defined by selector switch 607 in position A. This QMF domain signal can be fed directly to the appropriate QEG interface of the MPEG Surround decoder and this is the most advantageous operating mode in terms of delay, complexity and quality. Another option is to use the synthesis of the filtering set QMF 605 in order to obtain a stereo signal of the time domain. With selector switch 607 in position B, the converter gives the output for a digital stereo audio signal, which can also be fed to the time domain interface of the next MPEG Surround decoder or rendered directly to the stereo player. The third option with selector switch 607 in position C is achieved by coding the stereo time domain signal with a 606 stereo audio encoder. The output format of the downmix converter is then a stereo audio bit stream, which interacts with the core decoder included in the MPEG decoder. This third operating mode is applicable when the SOAC to MPEG Surround transceiver is separated from the MPEG decoder by a connection that limits the bit rate or in the case when the user wishes to keep the particular render of the object for future playback.
[0030] Fig. 7 shows the structure of an MPEG surround decoder for a stereo downmix. The stereo downmix is converted to three intermediate channels using the Two-To-Three (TTT) [Two-to-Three] box. These intermediate channels then split into two through three One-To-Two (OTT) boxes to get six channels in a 5.1 channel configuration.
[0031] Fig. 8 shows a practical method of using an SAOC encoder. The 802 audio mixer gives the output for a stereo signal (L and R), which usually consists of a mixer input (here input channels 1-6) and optionally additional inputs return (return) for effects such as reverb etc. Mixer also gives the output for an independent channel (here channel 5). This can be achieved e.g. by commonly known mixer functions such as "direct outputs" or "extra sending" in order to give an output to the independent channel after each insertion process (such as dynamic processing and correction). The stereo signal (L and R) and the independent channel output (obj5) constitute the input to the SAOC encoder 801, which is nothing but a particular example of the SAOC 101 encoder in Fig. 1. However,
[0032] In the following text, a mathematical description of the present invention will be made. For complex discontinuous signals x, y, complex scalar product and quadratic norm (energy) are expressed by
W) - £ * (*): <*).
ł | xf = {x, x) = £ | x (i) | \ *
where y (k) is the signal of the conjugate number y (k). All signals are treated here as samples of sub-bands from the modulated filter set or FFT analysis of discontinuous time signals by means of time windows. It is understood that these subbands must be re-converted to a discontinuous time domain by appropriate synthesis of the filtering sets. The block of signal samples L shows the signal in the time interval, sucking is a time and frequency puzzle constructed for the needs of perception used to describe the signal properties. In this arrangement, the audio objects can be represented by rows N of length L
<td colspan="3">in the matrix,</td>
<td>S =</td><td>> (0) ¡(l) .. ί, (0) ^ 0) ·</td><td>- ¡(iI) ' - ^ CŁ-l></td>
<td></td><td><sub>AND</sub><0) 4Γ "<1) ..</td><td></td>
[0033] The matrix D of the downmix weight in the size KxN, where K> 1, determines the K channel downmix signals in the form of a matrix with K rows by multiplying the matrix
X = DS. P) [0034] A user-controlled matrix image matrix object of size Μ x N defines a target rendering
M channels of audio objects in the form of a matrix with M rows by multiplying the matrix
Y = AS (4) Regardless of the effects of the core audio coding, the task of the SAOC decoder is to generate the approximation in the perceptible sense for the target Y rendering of the original audio objects, along with the specific rendering matrix A, the downmix X, the downmix matrix D and the object parameters .
[0036] The object parameters in the energy mode in terms of the present invention provide information about the covariance of the initial objects. In the deterministic version suitable for the following derivation, as well as describing typical coder operations, this covariance is given in an unnormalized form by the SS * product of the matrix, where the asterisk is an operation of transferring the conjugate complex number of the matrix. Therefore, the energy mode parameters of the object provide the matrix E with a positive semi-defined Ν χ N, possibly up to the scale factor,
SS * «E. (5) [0037] Earlier encoding of audio objects often refers to a model of objects in which all objects are not interdependent. In this case the matrix E is a diagonal and contains only an approximation of the energy of the object S<sub>n</sub> = || s<sub>n</sub>||<sup>2</sup> for n = 1,2, ..., N. The object parameter extractor of Fig. 3 provides a significant refinement of the concept, especially in cases where the objects are provided as stereo signals for which the assumption of absence of correlation can not be made. The grouping of selected P-pairs of stereo objects is described in the {(u<sub>p /</sub>m<sub>p</sub>), p = 1,2, ..., P}. For these stereo pairs, the correlation is calculated (pp<sub>n</sub>s<sub>m</sub>), and the complex, real or absolute value of the normalized correlation (ICC)
<td>o =<sup>p</sup>~ MM</td><td>(6)</td>
<td>is obtained</td><td>thanks to the 302 stereo parameter extractor.</td>
<td>In the decoder,</td><td>ICC data can then be combined</td>
<td>with parameters</td><td>power to create a matrix with 2P inputs</td>
<td>diagonal.</td><td>For example, for all objects N = 3,</td>
<td colspan="2">of which the first two constitute a single pair (1,2),</td>
transmitted energy and correlation data are (Si, S<sub>2</sub>S<sub>3</sub>) and pi.2- In this case, the combination in the matrix E gives
<img file="PL2068307T3_D0001.tif" />
0, [0038] The object parameters in the prediction mode according to the present invention are intended to form a CN x K matrix of the prediction coefficient (OPC) available to the decoder, such that
S «CX = CDS. (7) [0039] In other words, for each object there is a linear combination of downmix channels, such that the object can be recovered approximately by * "(*) * c"<sub>iXi</sub>(k) + ... + c<sub>nl [</sub>x<sub>K</sub>(K). (8) [0040] In a preferred embodiment, the OPC extractor 401 solves normal equations
CXX '= SX *. (9) or, for a more attractive OPC case with real values, it solves
CRe {) 0r} = Re {SX ·).
(10) the vocal track s<sub>3</sub> . Matrix [0043] This means that the left following<sup>113</sup> and the right channel has a single path for ce [0041] In both cases, assuming the actual matrix values of downmix weight D and non-singular downmix covariance, there is a multiplication from the left of D such that
DC = I, (11) where I is the matrix of "K-sized identity." If D has a full row, it is consistent with elementary linear algebra in such a way that the set of solutions for (9) can be parameterized by a maximum of {K · (NK ) 0} parameters This is used in the complex coding in OPC data 402. The complete C prediction matrix can be reproduced in the decoder from the reduced set of parameters and the downmix matrix.
Consider, for example, for a stereo downmix (K = 2) the case of three objects (N = 3) including the stereo music track (Si, s<sub>2</sub>) and the central instrument or downmix is the following (12) downmix channel is as follows ** OPC for lu approximation s<sub>3</sub> ~ c<sub>3</sub>iX<sub>1</sub> + c<sub>32</sub>X2 and the equation (11) can in this case be solved to achieve
C ,, = 1 -C<sub>3</sub>, / - / ϊ,
C, j = -Cyifyfl,
C<sub>2l</sub> - -C<sub>3i</sub>fy / 2, and
0 1/75 oii / Tzl '
Thus, the amount of OPC that is sufficient is given by K (N-K) = 2- (3-2) = 2.
[0044] OPC 031,032 can be derived from normal equations
Ml .ta »*») tal
"[Ta» * »)» ta »* a) J
MPEG Surround SAOC Transcoder [0045] With reference to Figure 7, M = 6 configuration output channels 5.1 are (y<sub>and</sub> s<sub>2</sub>,. . ., y<sub>6</sub>) = (l<sub>f</sub>l<sub>s</sub>r<sub>f /</sub>r<sub>sf</sub> c, lfe).
The transcoder must send a stereo downmix (l<sub>0</sub>r<sub>0</sub>) and parameters for TTT and OTT boxes. Since we are now focusing on the stereo downmix, it is assumed below that K = 2. Since both object parameters and MPS TTT parameters exist in energy mode and in the prediction mode, all four combinations should be considered. The energy mode is an appropriate choice in the case where the downmix audio encoder is not a wave encoder in the considered frequency range. It is understood that the MPEG Surround parameters derived in the following text must be properly quantized and coded prior to their transmission. To further clarify the four combinations mentioned above, they include:
1. Parameters of objects in energy mode and transcoder in prediction mode
2. Parameters of objects in energy mode and transcoder in energy mode
3. Parameters of objects in the prediction mode (OPC) and transcoder in the prediction mode
4. Object parameters in the prediction mode (OPC) and transcoder in energy mode [0046] If the downmix audio encoder is a wave encoder in the considered frequency range, the object parameters may be in the energy or prediction mode, but it is good if the transcoder operates in the prediction mode . If the downmix audio encoder is not a wave encoder in the considered frequency range, both the object encoder and the transcoder should operate in energy mode. The fourth combination is less important, so the further description will only apply to the first three combinations. .
Object parameters, data in energy mode [0047] In the energy mode, data available to the transcoder is described by three matrices (D, E, A). The OTT MPEG Surround parameters are obtained by performing energy estimation and correlation in a virtual rendering, derived from the transmitted parameters and the A 6 x N matrix. The target covariance of the six channels is given by
YY * = AS (AS) '= A (SS *) A *, (13) [0048] By substituting (5) to (13) an approximation of yy ^ f = aea * is obtained.
(14) which is completely defined by the available data. Let f<sub>at</sub> means elements F. Then CLD and ICC parameters are read from
CLD<sub>0</sub>
<img file="PL2068307T3_D0002.tif" />
CLD<sub>t</sub> = 10! Og | "
<img file="PL2068307T3_D0003.tif" />
CLD<sub>1</sub> = 1Qiog<sub>i0</sub>
<img file="PL2068307T3_D0004.tif" />
<img file="PL2068307T3_D0005.tif" />
(15) (16) (Π) (18) (19) where φ is an absolute value φ (ζ) = | z | or an operator with a real value of φ (ζ) = Εο {ζ}.
[0049] Let us consider as illustrative the example of the three objects previously described in connection with equation (12).
Let the rendering matrix be given by
I 0 0 1 o 1 0 1 1 o 0 0 1 0 0 I [0050] The target rendering therefore includes placing the object 1 between the front right and right surround channels, the object 2 between the front left and left surround channels, and the object 3 on the right channel front, middle and lfe. Let's assume, for the sake of simplicity, that three objects. they are not correlated and all have the same energy, so that in this case the right side of the formula (14) takes the form and 1 0 0 0 0 1 1 0 0 0 0
0 2 11 1 0 0 1 10 0 0 0 10 1 1
0 10 1 1 [0052] By substituting the appropriate values for formulas (15) (19), one obtains
CLD, = 101og<sub>lo</sub> these)<sup>= 0, og</sup>'#<sup>3dB</sup>
<img file="PL2068307T3_D0006.tif" />
/ CC, =
CLD<sub>2</sub> = 101og<sub>)ABOUT</sub>
<img file="PL2068307T3_D0007.tif" />
ODB
In A.) _ <"(0 __L icc, = g (Zi)? 0) [0053] Consequently, the MPEG Surround decoder will be instructed to use some kind of decorrelation between the right front and right surround channels, but without the decorrelation between the left front and the left surround channel.
For TTT MPEG Surround parameters in the prediction mode, the first step is to create a reduced rendering matrix A<sub>3 </sub>in size 3xN for combined channels (l, r, qc), where = 1/72.
This means that a<sub>3</sub> = d<sub>36</sub>a, the matrix of the partial downmix of 6 to 3 is determined by <sup>=</sup> w, in, 0 0 0 0
0 Wj w<sub>2</sub> 0 0
0 0 0 gw<sub>3</sub> g<sub>3</sub> (20) [0055] Weights of the partial downmix in<sub>p</sub>, p = 1,2,3 are determined so that energy in<sub>p</sub>(s<sub>2p</sub>.<sub>1</sub>+ y2p) is equal to the sum of energy || γ<sub>2</sub>^ ill<sup>2</sup> + lly2pll<sup>2</sup> up to the limiting factor. All the data needed to derive a partial D36 downmix matrix is available in F. Then the C prediction matrix<sub>3</sub> Size 3> <2 is created so that
C<sub>3</sub>X »A, S, (21) [0056] It is good if such a matrix is output by including normal equations at the beginning
Cj (DED ') = AjEd *, [0057] possible
The solution of adjusting the normal equations gives the best wave for (21) a given model of covariance of objects E. Part of the C matrix postprocessing<sub>3 </sub>it is preferred, including row factors for the total. "or. individual loss prediction compensation based on channels.
[0058] To illustrate and explain the above steps, consider continuing the special example of rendering of the six channels mentioned above. In the conditions of elements of the matrix of F weights of the downmix are solutions of equations<sup>in</sup>t> (/ ip-Up-i <sup>+ +</sup> * Λ | »ι.2 /» - ι <sup>+</sup> Λ / up »Ρ ~, which in the specific example become the following in»,<sup>l</sup>(l + l + 2-l) = l + lw, (2 + 1+ 2-1) = 2 + 1 ^ (1 + 1 + 2-1) = 1 + 1 (in ,,<sub>wj</sub>, ^) = (1 / 72.73 / 5.1 / 72).
[0059] So that Substitution to (20) gives o
o [0060] By obtaining accuracy), the solution of the system of equations C<sub>3</sub> (DED ') = A<sub>3</sub>ED 'now going to the finished one
C
-0.3536 1.0607
1.4358-0.1394 0.3536 0.3536 [0061] The matrix C<sub>3</sub> contains the best scales to get the approximation of the desired object rendering for the connected channels (i, r, gc) from the objects downmix. This general type of operations on the matrix can not be applied by the MPEG surround decoder, which is associated with the limited space of the TTT matrix by using only two parameters. The purpose of the downmix converter according to the invention is to pre-process the downmix in such a way that the combined pre-processing effect and TTT MPEG Surround matrix is identical to the desired up-mix described by C<sub>3</sub>. ________________ [0062] In MPEG Surround, a TTT matrix for prediction (l, r, qc) with (l<sub>0 /</sub>r<sub>0</sub>) is parameterized by three parameters (α, β, γ) via ^ τττ al l-α / 7 + 2 ι-Λ (22) [0063] The matrix G of the downmix converter according to the present invention is obtained by selecting γ = 1 and solving system of equations
CmG = C ,. (23) [0064] As can be easily checked, this means that D<sub>TTT</sub>C<sub>TTT</sub> = 15 I, where I is the identity matrix two by two and
D
<img file="PL2068307T3_D0008.tif" />
(24) [0065] Thus, matrix multiplication from the left through D<sub>TTT</sub> both sides (23) leads to
G = D<sub>m</sub>C3. (25) [0066] In the general case, G will be reversible and (23) has a unique solution for C<sub>TTT</sub>, which is subordinated to DtttCttt = I · TTT parameters (α, β) are determined by this solution.
[0067] For the particular example considered earlier, it is easy to check that the solutions are given by [0 1.41421
1.7893 0.2401J "<sup>d</sup> ("· 4) - (0.3506, 0.4072).
It should be noted that the major part of the stereo downmix is interchanged between the left and right sides for this converter matrix, which reflects the fact that the rendering example places objects that are on the left downmix channel in the right part of the sound stage and vice versa. This behavior can not be obtained from an MPEG decoder
Surround in stereo mode. ___________________ [0069] If it is not possible to use a downmix converter, a suboptimal procedure can be developed as follows. For MPEG Surround, the TTT parameters in the energy mode that is required are the energy distribution of the connected channels (l, r, c). Therefore, important CLD parameters can be derived directly from the F elements via
CLD ^ =! 0log<sub>e</sub>
<img file="PL2068307T3_D0009.tif" />
CLĄrr = 10log<sub>10</sub>
<td>(11</td><td>AND'<sup>1</sup></td>
<td>Them</td><td>2 7</td>
= 10Ig, (26) (27) [0070]
In this case, it is only appropriate to use the diagonal G matrix of the downmix converter.
with positive elements for
It is practical to achieve the correct energy distribution of downmix channels before the TTT uptake. At the matrix, six for two downmix channels D<sub>2</sub>6 = D<sub>TTT</sub>D<sub>36</sub> and definitions from
Z = DED ', (28) w = d<sub>m</sub>ed;<sub>6</sub>, you simply choose
J<sup>IN</sup>J2 / 2<sub>Ώ</sub> (29) (30) [0071] A further observation is that such a diagonal downmix converter can be omitted from objects to the MPEG surround transcoder and used by activating arbitrary downmix gain parameters (ADG) of the MPEG Surround decoder. The amplification will be logistic data by ADG<sub>2</sub> = 10 logio (Wii / ζϋ) for i = 1, 2.
Object parameters, data in the prediction mode (OPC) [0072] In the prediction mode of objects, the available data is represented by three matrices (D, C, A), where C is a matrix Ν<sup>χ</sup>2 containing N pairs of OPC. Due to the relative nature of the prediction coefficients, it will be necessary to have access to approximation to the covariance matrix for the estimation of energy-based MPEG Surround parameters
2x2 downmix of objects,
XX ' »With. (31) [0073] It is good when this information is sent from the object encoder as part of the downmix further information, but can also be estimated in the transcoder based on measurements taken in the received downmix, or indirectly output from (D, C) through approximate considering the object model. With Z, the covariance of objects can be estimated by inserting a prediction model Y - CX, receiving
E = CZC, (32) and all OTT MPEG Surround and TTT energy mode parameters can be estimated from E, as with energy-based object parameters. However, the big advantage of using OPC arises in combination with the TTT MPEG Surround parameters in the prediction mode. In this case, the wave approximation D<sub>36</sub> Y «A<sub>3</sub>CX immediately gives you a reduced prediction matrix
C ^ AjC, (32) from which the remaining stages of achieving TTT (α, β) parameters and downmix converter are similar to the case of object parameters, data in energy mode. In fact, the stages of formulas (22) to (25) are completely identical. The resulting matrix G is fed to the ______ converter, and the TTT parameters (α, β) are sent to the MPEG surround decoder.
Self-use of the downmix converter for stereo rendering [0074] In all of the cases described above, the stereo downsampling converter 501 sends the approximation of the rendering of 5.1 audio object channels to the stereo downmix. This stereo rendering can be expressed by the 2 * NA matrix<sub>2</sub>, determined by A<sub>2</sub> = D<sub>26</sub>A. In many applications this downmix is interesting, and direct manipulation of stereo rendering A<sub>2</sub> he is attractive. Consider again an illustrative example of a stereo track with a superimposed central mono vocal track coded according to the special case of the method shown in figure 8 and discussed in the section near the pattern (12). The vocal volume control by the user can be performed by rendering
AND<sub>2</sub> =
0 V / T2 oi ν / 7Σ (33) where v is the control of the vocal to music ratio. The structure of the downmix converter matrix is based on
GDS »A, S. (34) [0075] For object parameters based on prediction, the approximation S & quot; CDS is simply inserted and the converter matrix G «A is obtained<sub>2</sub>C. Normal equations are solved for object parameters based on energy
G (DED ') = A<sub>2</sub>* ED.
(35) [0076] Figure 9 illustrates a preferred embodiment of an audio object encoder according to one aspect of the present invention. The encoder 101 of audio objects has already been generally described in connection with the previous figures. Audio source encoder_____ to generate. encoded. .. the object signal uses a plurality of audio objects 90 that have been shown in Figure 9 as entering the downmixer 92, and the object parameter generator 94. In addition, the audio encoder 101 includes a downmix information generator 96 for generating downmix information 97, indicating the distribution of multiple audio objects in at least two downmix channels designated in 93 as coming from the downmixer.
[0077] The object parameter generator is used for generating parameters 95 for audio objects, where object parameters are calculated in such a way that reconstruction of audio objects is possible using object parameters and at least two downmix channels 93. It is important, however, that this reconstruction does not take place on the encoder side, but on the decoder side. In spite of this, the object parameter generator on the encoder side calculates the object parameters for the objects 95 so that this full reconstruction can take place on the decoder side.
[0078] Furthermore, the audio object encoder 101 includes an output interface 98 for generating the encoded signal 99 of the audio objects using the downmix information 97 and the object parameters 95. Depending on the application, the downmix channels may also be used and coded to the encoded audio object signal. However, there may be situations in which the output interface 98 generates a scrambled signal 99 of audio objects that does not contain downmix channels. Such a situation may arise when any downmix channels to be used at the decoder side are already at the decoder side, so that the downmix information and the object parameters for the audio objects are transmitted independently of the downmix channels. Such a situation may be useful when channels of upholstic objects 93 may be purchased irrespective of parameters, objects and downmix information.
[0079] Without parameters of the downmix objects and information, the user can render the downmix channels as a stereo or multi-channel signal, depending on the number of channels in the downmix. Naturally, the user may also render a mono signal, simply adding at least two downmix channels. In order to increase the flexibility of rendering and the quality and usability of the listening experience, object parameters and downmix information allow the user to create flexible rendering of audio objects in any intended playback setting, such as a stereo system, multi-channel, or even wavefield synthesis system. Although wave field synthesis systems are not yet very popular, multi-channel systems such as 5.1 or 7.1 are gaining increasing popularity in the consumer market.
[0080] Figure 10 shows an audio synthesizer for generating output data. For this purpose, the audio synthesizer includes a synthesizer 100 of the output data. The output data synthesizer receives as an input signal the downmix information 97 and parameters 95 of the audio objects, and possibly the intended audio source data, such as setting the audio sources or user-specific loudness of a specific source, where the source should be rendered as shown in 101.
[0081] The output data synthesizer 100 is for generating output data useful for creating a plurality of output channels of a predefined audio output configuration representing a plurality of audio objects.
In particular, a synthesizer 100 of output data. it is useful for using downmix information 97 and parameters of 95 audio objects. As discussed in connection with Figure 11, the output data may be high diversity data and many applications that include special rendering of the output channels only include reconstruction of source signals or contain transcoding of parameters to spatial rendering parameters for spatial upmixer configuration without any specific rendering of output channels , but for example to save or transmit such spatial parameters.
[0082] A general use scenario of the present invention is illustrated in Figure 14. There exists an encoder side 140 that includes an encoder 101 of audio objects that receives, as an input signal, N audio objects. The preferred encoder output signal includes, besides the audio object downmix information and object parameters that are not shown in Figure 14, K downmix channels. The number of downmix channels according to the present invention is greater than or equal to two.
[0083] The downmix channels are transmitted to the decoder side
142, which comprises a spatial up-mixer 143. A spatial arrangement of the inventors,
143 may include an audio synthesizer according to when the audio synthesizer is used in the transcoder mode. However, when the audio synthesizer 101 according to figure 10 operates in a spatial up-mixer mode, then the spatial up-mixer 143 and the audio synthesizer are the same device in this embodiment. The spatial upmikser generates M output channels for playback by M speakers. These speakers are located in specific spatial locations and together represent a fixed audio output configuration. The output channel of the established audio output configuration can be seen as. digital or analog speaker signal, ___: sent _ .. "with.
up-mixer 144 to the loudspeaker input in a fixed position between a plurality of predetermined positions in a predetermined audio output configuration. Depending on the situation, the number of M output channels can be equal to two when stereo rendering is performed. However, when performing multichannel rendering, the number of M output channels is greater than two. Usually there will be a situation where the number of downmix channels is smaller than the number of output channels due to the requirements of the transmission link. In this case, M is larger than K and can be even much larger than K, twice or even more.
[0084] Figure 14 also includes a series of matrix records to illustrate the functionality of the encoder according to the invention and the decoder side of the invention. Generally, a block of sampling values is processed. For this reason, as shown in equation (2), the audio object is represented as a series of L sampling values. The matrix S has N rows, corresponding to the number of objects, and L columns corresponding to the number of samples. The matrix E is calculated according to equation (5) and has N columns and N rows. Matrix E contains object parameters when they are data in energy mode. For uncorrelated objects, the matrix E has, as given earlier in connection with equation (6), only the main diagonal elements, where the main diagonal element gives energy to the audio object. All non-diagonal elements represent, as stated earlier,
[0085] Depending on the specific embodiment, equation (2) is a time domain signal. Then generated. There is a single energy value for the entire band of audio objects. It is good, however, when the audio objects are processed by a time / frequency converter, which includes, for example, the type of transformation or the algorithm of the filter assembly. In the second case equation (2) is correct for each sub-band, so that matrix E is obtained for each sub-band and of course each time frame.
[0086] The matrix of downmix channels X has K rows and L columns and is calculated according to equation (3). As given in equation (4), M output channels are calculated using N objects by using the rendering A for N objects, objects can be recreated at the decoder side using downmix parameters and objects, and rendering can be applied directly to the reproduced object signals.
so-called matrix
Depending on the situation N [0087] Alternatively, the downmix can be directly converted into output channels without explicitly calculating the source signals. Generally, the rendering matrix A indicates the location of each source relative to the determined audio output configuration. If there are six objects and six output channels, you can place each object in each output channel, and the rendering matrix will reflect this pattern. If, however, all objects were to be placed between the two positions of the output speakers, then the rendering matrix A would look different and reflect this situation.
[0088] The rendering matrix, or more generally speaking, the intended location of the objects and also the intended relative loudness of the audio sources can generally be computed by the encoder and sent to the decoder as a so-called scene description. In other embodiments, however, this description of the scene can be generated by the user himself in order to generate ___specific _up.-mix_. for a user-specific audio output configuration. Sending the scene description is not necessarily required, but it can also be generated by the user to meet his requirements. For example, the user may want to place audio objects in places that differ from the places where the objects were located during their generation. There are also cases
[0089] Figure 9 shows the downmixer 92. The downmixer serves to downmix a plurality of audio objects to a plurality of downmix channels, wherein the number of audio objects is greater than the number of downmix channels, and where the downmixer is associated with a downmix information generator so that the distribution of multiple audio objects in a plurality Downmix channels are executed according to the downmix information. The downmix information generated by the downmix information generator 96 in Figure 9 can be created automatically or set manually. It is preferable to provide down-level information with a resolution smaller than the resolution of object parameters. Thus, bits of additional information can be saved without major loss of quality, because fixed downmix information for a specific audio element or only slowly changing downmix situation, not requiring choice according to frequency, proved to be sufficient. In one example, the downmix information represents a downmix matrix having K rows and N columns.
[0090] The value in the row of the downmix matrix has a specific value when the audio object corresponding to this value in the _downmix matrix ... is in the downmix channel represented by the row of the downmix matrix. When an audio object is attached to more than one downmix channel, the values of more than one row of the downmix matrix have a specific value. It is preferred, however, that squared values when adding for a single audio object add up to 1.0. However, other values are also possible. In addition, audio objects can be inserted into one or more downmix channels with different levels, and these levels can be indicated by weights in the downmix matrix that differ from one and do not add up to 1.0 for the given audio object.
[0091] When the downmix channels are connected to the encoded audio object signal generated by the output interface 98, the encoded audio object signal may be, for example, a time-multiplexed signal in a particular format. Alternatively, the encoded audio object signal may be any signal allowing the separation of object parameters 95, downmix information 97 and down-mix channels 93 on the decoder side. In addition, the output interface 98 may include encoders for object parameters, downmix information, or downmix channels. Encoders for downmix parameters and information may be coders and differential channel encoders and / or entropy coders, and downmixes may be mono or stereo encoders, such as MP3 or AAC encoders.
[0092] Regrets the particular application, the downmixer 92 serves to incorporate the background music stereo representation into at least two downmix channels and further inserts the vocal track into at least two downmix channels in a given ratio. In this embodiment, the first background music channel is placed in the first downmix channel and the second background music channel in the second downmix channel. This results in optimal stereo background music playback in the stereo renderer. However, the user can still modify the position of the vocal track between the left and right stereo speakers. Alternatively, the first and second background music channels may be attached to one downmix channel, and the vocal track to the second downmix channel. Thus, by eliminating one downmix channel, you can completely separate the vocal track from musica, which is especially useful for karaoke applications. However, the quality of stereo playback of the background music channels will deteriorate due to the parameterization of objects, which is of course a lossy compression method.
[0093] Downmixer 92 is adapted to perform sample addition on a sample in the time domain. This addition uses samples from audio objects that are downmixed to one downmix channel. When an audio object is to be inserted into a downmix channel with a specified percentage, the pre-determination of weights must take place prior to the sampling process. Alternatively, the summation may also take place in the frequency domain or the subband domain, i.e. in the subsequent domain after the time / frequency conversion. Thus, one can even make a downmix in the domain of the filter assembly when the time / frequency conversion is a filter group, or in the transformation domain, when the time / frequency conversion is of the FFT, MDCT or other type.
[0094] In one aspect of the present invention, the object parameter generator 94 generates energy parameters and additionally correlation parameters between the objects when the audio objects together represent a stereo signal, which is explained by the following equation (6). ^ Alternatively, object parameters are parameters of the prediction mode. Figure 15 shows the steps of an algorithm or device means for calculating these prediction parameters of audio objects. As discussed in connection with equations (7) to (12), some statistical information of the downmix channels in the matrix X and the audio objects in the matrix S must be calculated. In particular, block 150 represents the first step of measuring the real part SX * and the real part XX *. These real parts are not just numbers but matrices, which are determined in one embodiment by the entries in equation (1), when application of successive equation (12) is contemplated. In general, the values of step 150 may be calculated using the available data in the encoder 101 of the audio objects. Then, the prediction matrix C is calculated as shown in step 152. In particular, the system of equations is solved in a known manner so that all values of the prediction matrix C having N rows and K columns are obtained. Generally weight factors c that all values of the C prediction matrix having N rows and K columns are obtained. Generally weight factors c that all values of the C prediction matrix having N rows and K columns are obtained. Generally weight factors c<sub>n</sub>, and according to equation (8) are calculated in such a way that the weighted linear addition of all downmix channels reproduces the corresponding audio objects. The prediction matrix causes better playback of the audio objects when the number of downmix channels increases.
[0095] In the following, Figure 11 will be discussed in more detail. In particular, Figure 7 shows a number of types of output data useful for creating a plurality of output channels of a predetermined audio output configuration. Line 111 shows the situation in which the output data from the data synthesizer output
100 are reproduced audio sources. Input data required by
100 output synthesizer for rendering ___5 .... playing __sample of audio,. ohe.jmująą .._ inform them downmix, downmix channels and parameters of audio objects. However, for rendering the reproduced sources, it is not necessarily required to set up the output and intended setting of the audio sources in the spatial 'audio' output configuration. In this first mode, indicated by the mode number 1 in Figure 11, the output data synthesizer 100 would output the reproduced audio sources. In the case of prediction parameters as parameters of audio objects, the output data synthesizer 100 operates according to equation (7). When the object parameters are in the energy mode, then the output data synthesizer uses the inverse of the downmix matrix and the energy matrix for reproducing the source signals.
[0096] Alternatively, the output data synthesizer 100 functions as a transcoder, shown, for example, in block 102 in figure 1b. When the output synthesizer is of the transcoder type to generate spatial mixer parameters, the parameters of the audio objects, the output configuration and the intended source setting are required. In particular, the output configuration and the intended setting are provided by the rendering matrix A. However, the downmix channels are not required to generate spatial parameters of the mixer, which will be discussed in more detail with reference to Figure 12. Depending on the situation, the spatial parameters of the mixer generated by the data generator 100 output, can then be used by a direct spatial mixer, such as an MPEG Surround mixer, for upmixing downmix channels. This application example does not necessarily require modification of the downmix channels of the objects, but can provide a simple conversion matrix having only diagonal elements as discussed in equation (13). In mode 2, indicated by 112 in figure 11, output data transcriber 100 would send spatial parameters of the mixer and preferably a conversion matrix G according to equation (13), including gains that can be used as arbitrary downmix gain (ADG) parameters of the MPEG Surround decoder.
[0097] In mode number 3, indicated by 113 in the figure
11, the output data includes the spatial parameters of the mixer in the conversion matrix, shown in relation to equation (25). In this situation, the output data synthesizer 100 does not necessarily need to perform an actual downmix conversion to convert the downmix of the objects to a stereo downmix.
Another mode of operation, indicated by number 4 in line 114 in figure 11, shows synthesizer 100 of the output data of figure 10. In this situation, the transcoder is used as indicated by 102 in figure 1b and sends not only the spatial parameters of the mixer, but also additionally converted downmix. However, it is no longer necessary to send the conversion matrix G together with the converted downmix. Sending the converted downmix and spatial parameters of the mixer is sufficient, as shown in FIG. 1b.
[0099] Mode number 5 indicates a different use of the output data synthesizer 100 shown in Figure 10. In this situation, indicated by line 115 in Figure 11, the output data generated by the output data synthesizer does not include any spatial parameters of the mixer, but only the matrix G conversions, for example according to equation (35), or include a stereo output signal, indicated by 115. In this embodiment example, only stereo rendering is concerned, and no spatial parameters of the mixer are required. However, all available information is required to generate a stereo output signal. input, .shown in _____ in Figure 11.
[0100] Another mode of the output data synthesizer is indicated by the number 6 in line 116. Here, the output data synthesizer 100 generates a multi-channel output signal and the synthesizer 100 would be similar to the element 104 in figure 1b. To this end, the output data synthesizer 100 requires all available input data and outputs a multi-channel output signal having more than two output channels for rendering by an appropriate number of speakers positioned at the intended locations, according to a specific audio output configuration. Such a multi-channel signal is a 5.1, 7.1 or only 3.0 signal having a left, middle and right loudspeaker.
[0101] Next, reference is made to Figure 11 to illustrate one example of calculating a plurality of parameters from the parameterization concept of Figure 7, known from the MPEG Surround decoder. As shown, FIG. 7 shows the parameterization on the side of the MPEG surround decoder starting from the stereo downmix 70 having the left downmix channel l<sub>0</sub> and the right downmix channel<sub>0</sub>. Conceptually, both downmix channels are transmitted to the so-called Two-To-Three 71 box. It is controlled by a series of input parameters 72. The box 71 generates three output channels 73a, 73b, 73c. Each output channel is sent to the One-To-Two box. This means that channel 73a is sent to box 74a, channel 73b is sent to box 74b, and channel 73c is sent to box 74c. Each mailbox sends two output channels. The box 74a sends the left front channel lf and the left surround channel l<sub>s</sub>. In addition, the 7 4b mailbox sends the right front channel r<sub>f</sub> and the right surround channel r<sub>s</sub>. In addition, the box 74c sends the middle channel c and the low frequency gain channel lfe. What
... 5 ... important, all ... up-.mix .. from the downmix channels 70 to the output channels is performed using a matrix, and the structure of the tree from figure 7 does not have to be used step by step, but can be used by one or many operations on matrices. In addition, intermediate signals, designated 73a, 73b and 73c, are not explicitly calculated by the application, but are shown in Figure 7 for purposes of illustration only. In addition, the boxes 74a, 74b receive some resi-signals resi, res<sub>2</sub> that can be used to enter a certain randomness into the output signals.
[0102] As known from the MPEG surround decoder, the box 71 is controlled by CPC prediction parameters or CLD energy parameters<sub>T</sub>tt- Up to two prediction parameters 20 CPC1, CPC2 or at least two CLD energy parameters are required for upmixing from two to three channels<sup>1</sup>TTT and CLD<sup>2</sup>ttt. In addition, the measurement of ICC correlation<sub>TTT</sub> it can stay in box 71, which is however only a function and is not used in the embodiment of the invention. Figures 12 and 13 show the required steps and / or means for calculating all parameters CPC / CLDtttz CLD0, CLD1, ICC1, CLD2, ICC2 from object parameters 95 of figure 9, downmix information 97 from figure 9 and intended audio source positions, e.g. description scenes 101 according to figure 10. These parameters are intended for the predefined 5.1 audio output format.
[0103] Naturally, the specific calculation of parameters for this particular application can be adapted to the option placement of other output formats or parameterizations within the scope of this document. In addition, the sequence of steps or arrangement of means in figures 12 and 13 a, b are only exemplary and may change in the logical sense of mathematical equations. [010.4]. W., step 120, a rendering matrix is provided. A.
The rendering matrix indicates where to place the source of multiple sources in the context of the predefined output configuration. Step 121 shows the derivation of the matrix of partial downmix D<sub>36</sub> according to equation (20). This matrix reflects the downmix situation of the six output channels and is 3xN. When more output channels are to be generated than in 5.1 configuration, such as
8-channel output configuration (7.1), then the matrix specified in block 121 will be matrix D<sub>38</sub>. In step 122, the reduced rendering matrix A<sub>3</sub> is generated by multiplying matrix D<sub>36</sub>and a full rendering matrix according to step 120. In step 123 a downmix matrix D is inserted. This downmix matrix D can be recovered from the encoded audio object signal when it is fully appended to the signal. Alternatively, the downmix matrix can be parameterized, for example, for a specific example of downmix information and a G downmix matrix.
[0105] In addition, in step 124, the energy matrix of the objects is provided. This matrix is reflected by object parameters for N objects and can be obtained from imported audio objects or reconstructed using a specific reconstruction rule. This reconstruction principle may include entropy decoding etc.
[0106] In step 125, a "reduced" prediction matrix C is defined<sub>3</sub>. The values of this matrix can be calculated by solving the system of linear equations given in step 125. Particularly elements of matrix C<sub>3</sub> can be calculated by multiplying the equation on both sides by the inverse (DED *).
[0107] In step 126, the G conversion matrix is calculated.
The conversion matrix G has the size KxK and is generated according to equation (25.) ... In order to solve the equation in step 126, a special matrix D must be provided.<sub>TTT </sub>according to step 127. An example of this matrix is given in equation (24), and the definition can be derived from the corresponding equation for C<sub>TTT</sub>as defined in equation (22). Thus, equation (22) defines what to do in step 128. Step 129 defines the equations for calculating matrix C<sub>T</sub>tt- When the matrix C<sub>TTT</sub> will be determined according to the equation in block 129, the parameters α, β and y, which are CPC parameters, can be derived. It is good if γ is set to 1, so that the only CPC parameters to insert into block 71 are α and β.
[0108] The remaining parameters required for the diagram of Figure 7 are the parameters inserted into blocks 74a, 74b and 74c.
Calculation of these parameters is described in connection with Figure 13a. In step 130, a rendering matrix A is provided. The size of the rendering matrix A is N rows for the number of audio objects and M columns for the number of output channels. This rendering matrix includes information from the scene vector when it is used.
Generally, it includes a specific position information and a rendering matrix of the source positions in the output configuration. For example, when the rendering matrix A below the equation (19) is considered, it becomes clear how the determined location of the audio objects can be encoded in the rendering matrix. Naturally, other ways of indicating a specific position may be used, such as values different from 1. In addition, when values less than 1 on one side and greater than 1 after the other are used, it is possible to influence the volume of specific audio objects.
[0109] In one embodiment, the rendering matrix is generated on the decoder side without any information from the .coder side. This allows the user to place audio objects where they want without paying attention to the spatial relations of audio objects in the encoder settings. In another embodiment, the relative or absolute location of the audio sources may be coded on the encoder side and sent to the decoder as a kind of vector of the scene. Then, on the decoder side, this information about the location of the audio sources, which are best when they are independent of the intended sound rendering setting, are processed to obtain a rendering matrix that reflects the location of the audio sources adapted to the specific audio output configuration.
[0110] In step 131, an energy matrix of objects E is provided which has already been discussed in connection with step 124 of figure 12. This matrix has the size NxN and includes the parameters of the audio objects. In one embodiment, such an energy matrix of objects is provided for each subband and each block by a sample in the time domain or in the subband domain.
[0111] In step 132, the output energy matrix F. is calculated as the covariance matrix of the output channels. However, since the output channels are still unknown, the output energy matrix F is calculated using a rendering matrix and an energy matrix. These matrices are provided in steps 130 and 131 and are already available on the decoder side. Next, the following equations (15), (16), (17), (18) and (19) are used to calculate the parameters of the level difference of CLD channels<sub>0</sub>, CLD<sub>l</sub> CLD<sub>2</sub>, and inter-channel coherence parameters of ICCi and ICC<sub>2</sub>so that the parameters for the boxes 74a, 74b, 74c are available. Importantly, spatial parameters are calculated by combining specific elements of the energy matrix F.
[0112] After step 133, all parameters for the spatial upmixer, _______, are available. gak_____ shown schematically in Figure 7.
[0113] In previous application examples, object parameters were given as energy parameters. However, when the object parameters are given as prediction parameters, i.e. as the C-prediction matrix shown by position 124a in Figure 12, calculation of the reduced prediction matrix C<sub>3</sub> is simply a matrix multiplication, represented in block 125a and described in connection with equation (32). Matrix A<sub>3</sub>, used in block 125a, is the same matrix A<sub>3</sub>as mentioned in block 122 in figure 12.
[0114] When the object C prediction matrix is generated by the audio object encoder and sent to the decoder, some additional calculations are required for generating parameters for the boxes 74a, 74b, 74c. These additional steps are shown in Figure 13b. Again, a C-prediction matrix as shown by 124a in FIG. 13b is provided, which is the same as discussed in connection with block 124a in FIG. 12. Next, as discussed in connection with Eq. (31), a matrix of the downmix of object objects is calculated. With using a downmix transmitted, or is generated and sent as additional side information. When Z-matrix information is transmitted, the decoder does not necessarily have to perform energy calculations that inherently delay processing and increase the processing load on the decoder side. However, when these issues are not decisive for a given application, the bandwidth can be saved, and the covariance matrix Z can also be calculated using downmix samples that are obviously available at the decoder side. When step 134 is completed and the covariance matrix of the downmix of the objects is ready, the energy matrix is covered. be calculated according to step 135 using a prediction matrix C and a downmix covariance matrix or "downmix energy" Z. When step 135 is completed, all steps discussed in connection with FIG. 13 a, such as steps 132, 133, can be performed to generate all parameters for blocks 74a, 74b, 74c in figure 7. and the covariance matrix Z may also be calculated using downmix samples, which are of course available on the decoder side. When step 134 is completed and the covariance matrix of the downmix of the objects is ready, the energy matrix is covered. be calculated according to step 135 using a prediction matrix C and a downmix covariance matrix or "downmix energy" Z. When step 135 is completed, all steps discussed in connection with FIG. 13 a, such as steps 132, 133, can be performed to generate all parameters for blocks 74a, 74b, 74c in figure 7. and the covariance matrix Z may also be calculated using downmix samples, which are of course available on the decoder side. When step 134 is completed and the covariance matrix of the downmix of the objects is ready, the energy matrix is covered. be calculated according to step 135 using a prediction matrix C and a downmix covariance matrix or "downmix energy" Z. When step 135 is completed, all steps discussed in connection with FIG. 13 a, such as steps 132, 133, can be performed to generate all parameters for blocks 74a, 74b, 74c in figure 7.
[0115] Figure 16 shows another application in which only stereo rendering is used. The stereo rendering is the output signal provided by mode number 5 or line 115 in figure 11. Here the synthesizer 100 of the output data of figure 10 is not interested in any up-mix spatial parameters, but mainly a specific conversion matrix G for downmix conversion of objects to a stereo downmix, easily controlled and influenced.
[0116] In step 160 of Figure 16, a matrix of the partial downmix M for 2 is calculated. For the six output channels, the partial downmix matrix would be a downmix matrix of six to two channels, but other matrices are also available. The calculation of this partial downmix matrix can for example be derived from a partial downmix matrix D<sub>36</sub> from step 121 and matrix D<sub>TTT</sub> used in the stage
127 of Figure 12 [0117] In addition, the rendering matrix generated using the results of step 160, rendering A is shown in step 161. The matrix A is the same as discussed in connection with block 120 in figure 12.
stereo A<sub>2</sub> is a "large" matrix [0118] Then, in step 162, the stereo rendering matrix can be parameterized by placing parameters μ and k. When μ is set to 1 and κ is also set to 1, an equation (33) is obtained which allows change in vocal volume, in the example described in relation to equation (33) However, when other parameters such as μ and κ are used, the distribution of sources may also change.
[0119] Then, in step 163, the conversion matrix G is calculated using equation (33). In particular, the matrix (DED *) can be calculated and inverted, and the inverted matrix can be multiplied by the right side of the equation in block 163. Naturally, other ones can be used methods for solving the equation in block 163. Then the conversion matrix G is available, and the downmix of the X objects can be converted by multiplying the conversion matrix and the downmix of the objects as in block 164. Then the converted downmix X 'can be rendered in stereo using two stereo speakers. Depending on the application, specific values of μ, vi κ can be set to calculate the G conversion matrix. Alternatively, the G conversion matrix can be calculated using all three parameters as variables,
[0120] Preferred applications solve the problem of transmitting a series of individual audio objects (using a multi-channel downmix and additional control data describing the objects) and rendering objects in a given loudspeaker reproducing system). The technique of modifying data related to objects to control data compatible with the reproduction system is introduced. It proposes (configuration and control, further suitable coding methods based on the MPEG Surround coding scheme.
[0121] Depending on the requirements of a particular application of the methods of the invention, these methods and signals may be implemented in hardware or in software. The implementation may be performed using a digital data carrier, in particular a CD with recorded control signals for electronic reading, which may cooperate with a programmable computer system in such a way that the methods of the invention are performed. Generally, the present invention is therefore a computer program with a code stored on a machine readable medium that is configured to perform at least one of the methods of the invention when the program is running on a computer. In other words, the methods according to the invention are therefore a computer program having a code for carrying out the methods according to the invention when the program is running on a computer.
[0122] In other words, according to the application of the present case, the audio object encoder for generating a coded audio object signal, using multiple audio objects, comprises a downmix information generator for generating downmix information indicative of distribution of multiple audio objects in at least two downmix channels, a parameter generator of objects to generating object parameters for audio objects, and an output interface for generating a coded audio object signal, using downmix information and object parameters. [0123] Optionally, the output interface may generate an encoded audio signal by further using a plurality of downmix channels.
[0124] In addition or alternatively, the parameter generator may generate object parameters in a first time and frequency resolution, wherein the downmix information generator generates downmix information in a second time and frequency resolution and the second time and frequency resolution is less than the first time and frequency resolution.
[0125]. In addition, the downstream information.gr. can generate downmix information so that the downmix information is equal for the whole frequency band of the audio objects. [0126] In addition, the downmix information generator may generate downmix information such that the downmix information represents a downmix matrix defined as follows:
X = DS where D is a downmix matrix, and where X is a matrix and represents multiple downmix channels and has a number of rows equal to the number of downmix channels.
[0127] In addition, the part information may be a coefficient of less than 1 and greater than 0.
[0128] In addition, the downmixer may be used for attaching a stereo representation of background music to at least two downmix channels and inserting a vocal track into at least two downmix channels in a given ratio. [0129] Furthermore, the downmixer may perform the addition of samples of signals input to the downmix channel according to the indications of the downmix information.
[0130] In addition, the output interface may compress the downmix information and object parameters before generating the encoded audio object signal. [0131] Also, many audio objects may include stereo objects, represented by two audio objects having a specific non-zero correlation, and in which the downmix information generator generates grouping information indicative of two audio objects forming a stereo object.
[0132] In addition, the object parameter generator can generate object prediction parameters that are calculated such that the weighted addition of downmix channels for a source object controlled by prediction parameters or a source object approximates the source object. [0133] In addition, prediction parameters may be generated for the frequency band, where the audio objects cover a plurality of frequency bands.
[0134] In addition, the number of audio objects may be N, the number of downmix channels equal to K, and the number of prediction parameters calculated by the object parameter generator is equal to or less than Ν K.
[0135] In addition, the object parameter generator can calculate at most Κ (NK) object prediction parameters. [0136] In addition, the object parameter generator may include an up-mixer for mixing up multiple downmix channels using different sets of object prediction test methods, where the audio object encoder further includes an iteration controller to search for test object prediction parameters, giving the smallest deviation between the signal the source, reproduced by the upmixer, and the corresponding original source signal from among different sets of test object prediction parameters.
[0137] In addition, the output data synthesizer may be used to determine the conversion matrix using downmix information, where the conversion matrix is calculated so that at least parts of the downmix channels are converted when the audio object attached in the first downmix channel representing the first half of the stereo plane , to be played in the other half of the stereo plane.
[0138] Furthermore, the audio synthesizer may comprise a channel rendering system for rendering audio output channels for a predefined audio output configuration using spatial parameters and at least two. channels, .downmiksu. or converted ... downmix channels.
[0139] Furthermore, the output data synthesizer may additionally send output channels of a predefined audio output configuration using at least two downmix channels.
[0140] Furthermore, the output data synthesizer may be used to calculate the actual downmix weights for the partial downmix matrix such that the energy of the weighted sum of the two channels is equal to the energy of the channels within the limiting factor.
[0141] In addition, the downmix weights for the partial downmix matrix can be determined as follows:
WpC / j / ł-Up-t + blłp.lp <sup>+</sup>2 / ^ j<sub>i2?</sub>) ~ Λα-Μ / Μ + fipklp * P ~ where in<sub>p</sub> is the weight of the downmix, p is the variable integer index, fj.i is the element of the energy matrix representing the approximation of the covariance matrix of the output channels of the predefined output configuration.
[0142] In addition, the output data synthesizer may be used to calculate separate coefficients of the prediction matrix by solving a system of linear equations.
[0143] In addition, the output data synthesizer can be used to solve a system of linear equations based on:
C<sub>3</sub>(DED ') = AjED', where C<sub>3</sub> is a Two-To-Three prediction matrix, D is a downmix matrix derived from the downmix information, E is an energy matrix derived from the output source objects at least audio processing, and A<sub>3</sub> is a reduced downmix matrix, and where it indicates the complex conjugate operation.
[0144] In addition, the prediction parameters for the Two-ToThree up-mix can be derived from the parameterization of the prediction matrix so that the prediction matrix is defined using only two parameters, and where the data synthesizer is used to pre-process the two downmix channels so that the effect The initial and parameterized prediction matrix corresponds to the desired up-mix matrix.
[0145] In addition, the parameterization of the prediction matrix may be as follows:
C- = - '-ir 3 α + 2 β- \ α-1/7 + 2
1-a \ -β where the TTT index is a parameterized prediction matrix, 15 ao, β and γ are coefficients.
[0146] In addition, the downmix conversion matrix G may be calculated as follows:
G = DrrrCj, where C<sub>3</sub> is the two-To-Three prediction matrix, D<sub>TTT</sub> and C<sub>TTT</sub> they are equal to 1, and the identity matrix is two by two, and C is<sub>TTT</sub> is based on:
a + 2
<img file="PL2068307T3_D0010.tif" />
/ 7-1 / 7 + 2 \ -β where α, β and γ are constant coefficients.
[0147] In addition, prediction parameters for the Two-To25 Three up-mix can be specified as α and β, where γ is set to 1.
[0148] In addition, the output data synthesizer may be used to calculate energy parameters for a Three-Two-Six upmix using the energy matrix F, based on:
yy '~ f = aea', where A is the rendering matrix, E is the energy matrix derived from the audio source objects, Y is the channel of the output channel, and indicates the complex conjugation operation.
[0149] In addition, the output data synthesizer can be used to calculate energy parameters by combining energy matrix elements.
[0150] In addition, the output data synthesizer can be used to calculate energy parameters based on the following equations:
C £ Ą = 10Io6, / AJ
CLD, = 10l<sub>ABOUT</sub>grams LM
C £ Ą = 101og, / Aj / CC, =
P (/ m)
1CC<sub>2</sub> where φ is the absolute value φ (ζ) = | z | or an operator with a real value of φ (ζ) = Εο {ζ}, where CLD<sub>0</sub> is the parameter of the difference in levels of the first
<td>channel</td><td>CLD!</td><td>is</td><td>parameter</td><td>difference</td><td>levels</td><td>second</td>
<td>channel</td><td>cld<sub>2</sub></td><td>is</td><td>parameter</td><td>difference</td><td>levels</td><td>third</td>
<td>channel</td><td>ICCI</td><td>is</td><td colspan="2">the first parameter</td><td>energy</td><td>consistency</td>
inter-channel, ICC<sub>2</sub> is the second parameter of the inter-channel coherence energy, and fij are elements of the energy matrix F at positions ij in this matrix.
[0151] In addition, the first group of parameters may include energy parameters. And where the output data synthesizer is used to derive energy parameters by combining elements of the energy matrix F.
[0152] In addition, energy parameters can be derived based on:
CLD ^ = 101og<sub>M</sub>
<img file="PL2068307T3_D0011.tif" />
= and oiog 'and <11,112 λ
CLDjtj - 101og,<sub>0</sub>
<img file="PL2068307T3_D0012.tif" />
= 101og<sub>10</sub> FZ + / n)
1/23 + / 44 / where CLD °<sub>TTT</sub> is the first energy parameter from the first group, and CLD<sup>1</sup>TTT is the second energy parameter from the first parameter group.
[0153] Furthermore, the output data synthesizer may be used to calculate weighting factors for weighing downmix channels, which coefficients are used to control the arbitrary downmix gain coefficients in a spatial decoder.
[0154] Furthermore, the output data synthesizer may be used to calculate weighting factors based on:
= DED ',
W = D<sub>26</sub>ED ",
<img file="PL2068307T3_D0013.tif" />
where D is the downmix matrix, E is the energy matrix, derived from the audio source objects, W is the intermediate matrix, D<sub>2</sub>6 is a partial downmix matrix for downmixing from 6 to 2 channels of a predefined output configuration, and G is a conversion matrix comprising arbitrary amplification factors of a spatial decoder.
[0155] used
E = CZC '
Pose. including synthesizer - output data _ can be used to calculate energy matrix based on:
where E is the energy matrix, C is the matrix of prediction parameters, and Z is the covariance matrix of at least two downmix channels.
[0156] In addition, the output data synthesizer can be used to calculate a conversion matrix based on:
G = A<sub>r</sub>C, where G is the conversion matrix, A<sub>2</sub> is a partial rendering matrix, and C is the prediction parameter matrix.
[0157] In addition, the output data synthesizer can be used to calculate a conversion matrix based on:
G (DED> A<sub>2</sub>ED *, where G is an energy matrix derived from the sound source of paths, D is a downmix matrix derived from the downmix information, A<sub>2</sub> is a reduced rendering matrix, and indicates a complex join operation.
[0158] In addition, a parameterized stereo rendering matrix
AND<sub>2</sub> maybe μ, vi κ according to the position of the audio objects.
'μ \ ~ μ vj [l - κ IC VJ be as follows:
they are real, fixed parameters and volume of one or more source
DOLBY INTERNATIONAL AB, the Netherlands
Proxy:
<img file="PL2068307T3_D0014.tif" />
MamŁazewskł
Rzeczpospolita
Ζ-8945
ΕΡ 2,068,307 Β1
Contents3
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
57 members in 22 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 82964906 | United States of America | P | |
| 82964906 | United States of America | P | |
| 07818759 | European Patent Office (EPO) | A | |
| 07818759 | European Patent Office (EPO) | A | |
| 09004406 | European Patent Office (EPO) | A | |
| EP20070818759 | – | – | – |
| EP20090004406 | – | – | – |
| US20060829649P | – | – | – |
Members57
| Document | Office | Kind | |
|---|---|---|---|
| AU2007312598A1 | Australia | A1 | |
| CA2666640A1 | Canada | A1 | |
| CA2874451A1 | Canada | A1 | |
| CA2874454A1 | Canada | A1 | |
| WO2008046531A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200828269A | Taiwan Province of China | A | |
| EP2054875A1 | European Patent Office (EPO) | A1 | |
| NO20091901L | Norway | L | |
| MX2009003570A | Mexico | A | |
| KR20090057131A | Republic of Korea | A | |
| EP2068307A1 | European Patent Office (EPO) | A1 | |
| CN101529501A | China | A | |
| HK1126888A1 | Hong Kong, China | A1 | |
| JP2010507115A | Japan | A | |
| HK1133116A1 | Hong Kong, China | A1 | |
| RU2009113055A | Russian Federation | A | |
| KR20110002504A | Republic of Korea | A | |
| AU2007312598B2 | Australia | B2 | |
| US2011022402A1 | United States of America | A1 | |
| KR101012259B1 | Republic of Korea | B1 | |
| EP2054875B1 | European Patent Office (EPO) | B1 | |
| AU2011201106A1 | Australia | A1 | |
| UA94117C2 | Ukraine | C2 | |
| ATE503245T1 | Austria | T1 | |
| DE602007013415D1 | Germany | D1 | |
| TWI347590B | Taiwan Province of China | B | |
| RU2430430C2 | Russian Federation | C2 | |
| EP2372701A1 | European Patent Office (EPO) | A1 | |
| SG175632A1 | Singapore | A1 | |
| EP2068307B1 | European Patent Office (EPO) | B1 | |
| ATE536612T1 | Austria | T1 | |
| KR101103987B1 | Republic of Korea | B1 | |
| MY145497A | Malaysia | A | |
| ES2378734T3 | Spain | T3 | |
| AU2011201106B2 | Australia | B2 | |
| JP2012141633A | Japan | A | |
| RU2011102416A | Russian Federation | A | |
| PL2068307T3This record | Poland | T3 | |
| HK1162736A1 | Hong Kong, China | A1 | |
| CN102892070A | China | A | |
| BRPI0715559A2 | Brazil | A2 | |
| CN101529501B | China | B | |
| JP5270557B2 | Japan | B2 | |
| JP5297544B2 | Japan | B2 | |
| JP2013190810A | Japan | A | |
| CN103400583A | China | A | |
| EP2372701B1 | European Patent Office (EPO) | B1 | |
| PT2372701E | Portugal | E | |
| JP5592974B2 | Japan | B2 | |
| CA2666640C | Canada | C | |
| CN103400583B | China | B | |
| CN102892070B | China | B | |
| CA2874451C | Canada | C | |
| US9565509B2 | United States of America | B2 | |
| US2017084285A1 | United States of America | A1 | |
| NO340450B1 | Norway | B1 | |
| CA2874454C | Canada | C |
Numbers
- Publication, DOCDB
- 2068307
- Publication, EPODOC
- PL2068307T
- Application
- 20090004406
- Application, DOCDB
- 09004406
- Application, EPODOC
- PL20090004406T
Titles2
- English
- Enhanced coding and parameter representation of multichannel downmixed object coding
- Polish
- Udoskonalony sposób kodowania i odtwarzania parametrów w wielokanałowym kodowaniu obiektów poddanych procesowi downmiksu
Classification
- CPC, 10
- G10L19/20
- H04S7/30
- H04S2420/03
- G10L19/008
- G10L19/173
- H04S3/008
- H04S3/02
- H04S2400/03
- H04S5/00
- H04S2400/11
- IPC, 2
- G10L19 00
- H04S7 00