Enhanced coding and parameter representation of multichannel downmixed object coding
Abstract
Audio synthesizer (104) for generating output data using an encoded audio object signal (95, 97), comprising: an output data synthesizer (100) to generate the output data that can be used to reproduce a plurality of output channels of a predefined audio output configuration representing the plurality of audio objects, the output data synthesizer being operational to use downstream mix information indicating a distribution of the plurality of audio objects in at least two downstream mix channels, power information, correlation information indicating a power characteristic and a correlation characteristic of the at least two downstream mix channels (93), and audio object parameters for audio objects, wherein the output data synthesizer (100) is operative to transcode (502) the audio object parameters into spatial parameters for the predefined audio output configuration using also a planned positioning of the audio objects (90) in Audio output settings.

Term
1 yearto projected expiry
Projected expiry 5 October 2027, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
13 claims: 6 independent, 7 dependent
- 1ES 2 378 734 T3 ES 2 378 734 T3 CLAIMS REIVINDICACIONES 1. Audio synthesizer (104) for generating output data using an encoded audio object signal (95, 97), comprising:1. Sintetizador (104) de audio para generar datos de salida usando una señal (95, 97) de objeto de audio codificada, que comprende: an output data synthesizer (100) for generating the output data that can be used to reproduce a plurality of output channels of a predefined audio output configuration representing the plurality of audio objects, the output data synthesizer being operative to use downmix information indicating a distribution of the plurality of audio objects on at least two downmix channels, power information, correlation information indicating a power characteristic and a correlation characteristic of the at least two downmix channels (93), and audio object parameters for the audio objects, wherein the output data synthesizer (100) is operative to transcode (502) the audio object parameters into spatial parameters for the predefined audio output configuration further using a predicted positioning of the audio objects (90) in the audio output settings. un sintetizador (100) de datos de salida para generar los datos de salida que pueden usarse para reproducir una pluralidad de canales de salida de una configuración de salida de audio predefinida que representa la pluralidad de objetos de audio, siendo el sintetizador de datos de salida operativo para usar información de mezcla descendente que indica una distribución de la pluralidad de objetos de audio en al menos dos canales de mezcla descendente, información de potencia, información de correlación que indican una característica de potencia y una característica de correlación de los al menos dos canales (93) de mezcla descendente, y parámetros de objeto de audio para los objetos de audio, en el que el sintetizador (100) de datos de salida es operativo para transcodificar (502) los parámetros de objeto de audio en parámetros espaciales para la configuración de salida de audio predefinida usando además un posicionamiento previsto de los objetos (90) de audio en la configuración de salida de audio.
- 6Audio synthesizing method for generating output data using an encoded audio object signal (95, 97), comprising:6. Método de sintetización de audio para generar datos de salida usando una señal (95, 97) de objeto de audio codificada, que comprende: generar los datos de salida que pueden usarse para crear una pluralidad de canales de salida de una configuración de salida de audio predefinida que representa la pluralidad de objetos (90) de audio, en el que se usan información de mezcla descendente que indica una distribución de la pluralidad de objetos de audio en al menos dos canales de mezcla descendente, información de potencia, información de correlación que indican una característica de potencia y una característica de correlación de los al menos dos canales (93) de mezcla descendente, y parámetros de objeto de audio para los objetos de audio, y en el que los parámetros de objeto de audio se transcodifican (502) en parámetros espaciales para la configuración de salida de audio predefinida usando además un posicionamiento previsto de los objetos (90) de audio en la configuración de salida de audio. generate the output data that can be used to create a plurality of output channels of a predefined audio output configuration representing the plurality of audio objects (90), in which downmix information is used indicating a distribution of the plurality of audio objects in at least two downmix channels, power information, correlation information indicating a power characteristic and a correlation characteristic of the at least two downmix channels (93), and audio object parameters for the audio objects, and wherein the audio object parameters are they transcode (502) into spatial parameters for the predefined audio output configuration further using a predicted positioning of the audio objects (90) in the audio output configuration.
- 7Audio object encoder (101) for generating an encoded audio object signal using a plurality of audio objects (90), comprising:7. Codificador (101) de objetos de audio para generar una señal de objeto de audio codificada usando una pluralidad de objetos (90) de audio, que comprende: a downmix information generator (96) for generating downmix information (97) indicating a distribution of the plurality of audio objects on at least two downmix channels, wherein the downmix information generator (96) is configured to generate (150) power information and correlation information indicating a power characteristic and a correlation characteristic of the at least two channels (93) of down mix;un generador (96) de información de mezcla descendente para generar información (97) de mezcla descendente que indica una distribución de la pluralidad de objetos de audio en al menos dos canales de mezcla descendente, en el que el generador (96) de información de mezcla descendente está configurado para generar (150) una información de potencia y una información de correlación que indican una característica de potencia y una característica de correlación de los al menos dos canales (93) de mezcla descendente;an object parameter generator (94) for generating object parameters (95) for the audio objects;Y un generador (94) de parámetro de objeto para generar parámetros (95) de objeto para los objetos de audio;y ES 2 378 734 T3 an output interface (98) for generating the encoded audio object signal (99), the encoded object signal comprising the downmix information, the power information, the correlation information and the parameters of object. ES 2 378 734 T3 una interfaz (98) de salida para generar la señal (99) de objeto de audio codificada, comprendiendo la señal de objeto codificada la información de mezcla descendente, la información de potencia, la información de correlación y los parámetros de objeto.
- 10Audio object encoding method (101) for generating an encoded audio object signal using a plurality of audio objects, comprising:10. Método (101) de codificación de objetos de audio para generar una señal de objeto de audio codificada usando una pluralidad de objetos de audio, que comprende: generar información (97) de mezcla descendente que indica una distribución de la pluralidad de objetos (90) de audio en al menos dos canales de mezcla descendente, generar (150) una información de potencia y una información de correlación que indican una característica de potencia y una característica de correlación de los al menos dos canales de mezcla descendente;generating downmix information (97) indicating a distribution of the plurality of audio objects (90) in at least two downmix channels, generating (150) a power information and a correlation information indicating a power characteristic and a correlation characteristic of the at least two downmix channels;generar parámetros (94) de objeto para los objetos de audio;y generar la señal (99) de objeto de audio codificada, comprendiendo la señal de objeto de audio codificada la información de potencia, la información de correlación, la información de mezcla descendente y los parámetros de objeto. generating object parameters (94) for the audio objects;and generating the encoded audio object signal (99), the encoded audio object signal comprising power information, correlation information, downmix information, and object parameters.
- 11Señal de objeto de audio codificada que incluye una información de mezcla descendente que indica una distribución de una pluralidad de objetos de audio en al menos dos canales de mezcla descendente, una información de potencia y una información de correlación que indican una característica de potencia y una característica de correlación de los al menos dos canales de mezcla descendente, y parámetros de objeto, siendo los parámetros de objeto de manera que es posible la reconstrucción de los objetos de audio usando los parámetros de objeto y los al menos dos canales de mezcla descendente. eleven. Encoded audio object signal including a downmix information indicating a distribution of a plurality of audio objects in at least two downmix channels, a power information and a correlation information indicating a power characteristic and a correlation characteristic of the at least two downmix channels, and object parameters, the object parameters being such that reconstruction of the audio objects is possible using the object parameters and the at least two downmix channels.
Independent claims6
319 paragraphs in 21 sections, as filed
ES 2 378 734 T3
DESCRIPTION
Improved encoding and rendering of multi-channel downmix object encoding parameters
TECHNICAL FIELD
The present invention relates to the decoding of multiple objects from an encoded multi-object signal based on an available multi-channel downmix and additional control data.
BACKGROUND OF THE INVENTION
Recent development in audio facilitates the recreation of a multi-channel representation of an audio signal based on a stereo (or mono) signal and corresponding control data. These parametric envelope coding methods usually comprise a parameterization. A parametric multichannel audio decoder, (for example, the MPEG Surround decoder defined in ISO / IEC 23003-1 [1], [2]), reconstructs M channels based on K transmitted channels, where M> K, using the use of additional control data. The control data consists of a parameterization of the multichannel signal based on IID (Inter channel Intensity Difference; intensity difference between channels) and ICC (Inter Channel Coherence; coherence between channels). These parameters are normally extracted in the coding phase and describe power relationships and correlation between pairs of channels used in the upmixing process. Using such a coding scheme allows coding at a significantly lower data rate than all M-channel transmission, making the coding very efficient while at the same time ensuring compatibility with both K-channel devices and devices. with M channel devices.
A closely related encoding system is the corresponding audio object encoder [3], [4] in which various audio objects are down-mixed at the encoder and later down-mixed in a data-guided manner. of control. The upmix process can also be thought of as a separation of the objects that are mixed in the downmix. The resulting upmix signal can be played back on one or more playback channels. More precisely, [3,4] presents a method for synthesizing audio channels from a downmix (called a sum signal), statistical information about the source objects, and data describing the desired output format. In cases where multiple downmix signals are used, these downmix signals consist of different subsets of the objects, and the upmix is performed for each downmix channel individually.
In the new method we introduce a method where the upmix is done together for all the downmix channels. Object encoding methods, prior to the present invention, did not present a solution for joint decoding of a downmix with more than one channel.
References:
[1] L. Villemoes, J. Herre, J. Breebaart, G. Hotho, S. Disch, H. Pumhagen and K. Kjorling, “MPEG Surround: The Forthcoming ISO Standard for Spatial Audio Coding”, in 28th International aEs Conference , The Future of Audio Technology Surround and Beyond, Piteá, Sweden, June 30 - July 2, 2006.
[2] J. Breebaart, J. Herre, L. Villemoes, C. Jin, K. Kjorling, J. Plogsties and J. Koppens, “Multi-Channels goes Mobile: MPEG Surround Binaural Rendering”, at 29th International AES Conference, Audio for Mobile and Handheld Devices, Seoul, 2-4 September 2006.
[3] C. Faller, “Parametric Joint-Coding of Audio Sources”, Convention Paper 6752 presented at 120th AES Convention, Paris, France, May 20-23, 2006.
[4] C. Faller, "Parametric Joint-Coding of Audio Sources", patent application PCT / EP2006 / 050904, 2006.
WO 2006/048203 A2 discloses concepts for improved performance of prediction-based multichannel reconstruction. In particular, an energy loss introduced by a predictive upmix process is taken into account in a multichannel reconstruction. In particular, a left original channel, a center original channel, and a right original channel are down-mixed into a left down-mix channel and a right down-mix channel, where the left down-mix channel contains only the original channel. left and a part of the original center channel, and the right downmix channel only contains the original right channel and a part of the original center channel. This is defined in a downmix matrix. The two base channels are transmitted along with two different upmix parameters to an upmixer that meets a non-energy conservation upmix rule. The original reconstructed left, right, and center channels are generated and these channels are power corrected to obtain corrected left, right, and center channels.
It is an object of the present invention to provide an improved audio object encoding / decoding scheme.
ES 2 378 734 T3
This object is achieved by an audio synthesizer according to claim 1, an audio synthesizer method according to claim 6, an audio object encoder according to claim 7, an audio object encoding method according to claim 10, a encoded audio object signal according to claim 11 or a computer program according to claim 13.
SUMMARY OF THE INVENTION
A first aspect of the invention relates to an audio object encoder for generating an encoded audio object signal using a plurality of audio objects, comprising: a downmix information generator for generating downmix information indicating a distribution of the plurality of audio objects on at least two downmix channels; an object parameter generator for generating object parameters for the audio objects; and an output interface for generating the encoded audio object signal using the downmix information and the object parameters.
A second aspect of the invention relates to an audio object encoding method for generating an encoded audio object signal using a plurality of audio objects, comprising: generating downmix information indicating a distribution of the plurality of audio objects on at least two downmix channels; generate object parameters for audio objects; and generating the encoded audio object signal using the downmix information and the object parameters.
A third aspect of the invention relates to an audio synthesizer for generating output data using an encoded audio object signal, comprising: an output data synthesizer for generating the output data that can be used to create a plurality of output channels of a predefined audio output configuration representing the plurality of audio objects, the output data synthesizer being operational for use downmix information indicating a distribution of the plurality of audio objects on at least two downmix channels, and audio object parameters for audio objects.
A fourth aspect of the invention relates to an audio synthesizing method for generating output data using an encoded audio object signal, comprising: generate the output data that can be used to create a plurality of output channels of a predefined audio output configuration representing the plurality of audio objects, the output data synthesizer being operative to use downmix information indicating a distribution of the plurality of audio objects into at least two downmix channels, and audio object parameters for the audio objects.
A fifth aspect of the invention relates to an encoded audio object signal including downmix information indicating a distribution of a plurality of audio objects on at least two downmix channels and object parameters, the parameters being so that the reconstruction of the audio objects is possible using the object parameters and the at least two downmix channels. A sixth aspect of the invention relates to a computer program for performing, when run on a computer, the audio object encoding method or the audio object decoding method.
BRIEF DESCRIPTION OF THE DRAWINGS
The present invention will now be described by way of illustrative examples, which do not limit the scope or spirit of the invention, with reference to the accompanying drawings, in which:
Figure 1a illustrates the spatial audio object encoding operation comprising encoding and decoding;
Figure 1b illustrates the spatial audio object encoding operation reusing an MPEG Surround decoder;
Figure 2 illustrates the operation of a spatial audio object encoder;
Figure 3 illustrates an audio object parameter extractor operating in a power-based mode;
Figure 4 illustrates an audio object parameter extractor operating in a prediction-based mode;
Figure 5 illustrates the structure of a SAOC to MPEG Surround transcoder;
Figure 6 illustrates different modes of operation of a down-mix converter;
Figure 7 illustrates the structure of an MPEG Surround decoder for a stereo downmix;
Figure 8 illustrates a practical use case that includes an SAOC encoder;
Figure 9 illustrates an encoder embodiment;
Figure 10 illustrates a decoder embodiment;
Figure 11 illustrates a table to show different preferred decoder / synthesizer modes;
Figure 12 illustrates a method for calculating certain spatial up-mixing parameters;
Figure 13a illustrates a method for calculating additional spatial upmix parameters;
Figure 13b illustrates a method for calculating the use of prediction parameters;
Figure 14 illustrates an overview of an encoder / decoder system;
Figure 15 illustrates a method for calculating prediction object parameters; and Figure 16 illustrates a stereo reproduction method.
DESCRIPTION OF PREFERRED EMBODIMENTS
The embodiments described below are merely illustrative of the principles of the present invention for ENHANCED CODING AND REPRESENTATION OF MULTI-CHANNEL DOWNLOAD MIXING OBJECT ENCODING PARAMETERS. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. Therefore, it is only intended to be limited by the scope of the appended patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
Preferred embodiments provide an encoding scheme that combines the functionality of an object encoding scheme with the playback capabilities of a multi-channel decoder. The transmitted control data relate to the individual objects and thus allow manipulation of the reproduction in terms of level and spatial position. Therefore, the control data is directly related to the so-called scene description, giving information about the positioning of the objects. Scene description can be controlled either on the decoder side interactively by the listener or also on the encoder side by the producer. A transcoder phase, as taught by the invention, is used to convert the control data related to the object and the downmix signal into control data and a downmix signal that is related to the reproduction system, such as the MPEG Surround decoder.
In the presented coding scheme, the objects can be arbitrarily distributed in the downmix channels available in the encoder. The transcoder makes explicit use of the multichannel downmix information, providing a transcoded downmix signal and object-related control data. In this way, the upmix at the decoder is not performed for all channels individually as proposed in [3], but all the downmix channels are treated at the same time in a single upmix process. In the new scheme the multichannel downmix information must be part of the control data and is encoded by the object encoder.
The distribution of the objects in the downmix channels can be done automatically or it can be a design choice on the encoder side. In the latter case, the downmix can be designed to be suitable for reproduction by an existing multi-channel reproduction scheme (eg, a stereo reproduction system), which offers a reproduction and skips the multi-channel decoding and transcoding phase. This is an additional advantage over previous coding schemes, which consist of a single downmix channel, or multiple downmix channels containing subsets of the source objects.
While prior art object encoding schemes describe only the decoding process using a single downmix channel, the present invention does not suffer from this limitation as it provides a method for jointly decoding downmixes containing downmix of more than a channel. The achievable quality of object separation increases with a greater number of downmix channels. Thus, the invention successfully fills the gap between an object coding scheme with a single downmix mono channel and a multi-channel coding scheme in which each object is transmitted on a separate channel. Therefore, the proposed scheme allows flexible quality scaling for object separation according to application requirements and transmission system properties (such as channel capacity).
Furthermore, using more than one downmix channel is advantageous as it allows further consideration of a correlation between individual objects rather than restricting the description to intensity differences such as in prior art object coding schemes. The prior art schemes are based on the assumption that all objects are independent and uncorrelated with each other (zero cross correlation), although in real objects they are not unlikely to be correlated, such as the left and right channel of a stereo signal. Incorporating the correlation into the description (control data) as taught by the invention makes it more complete and thus further facilitates the ability to separate the objects.
ES 2 378 734 T3
Preferred embodiments comprise at least one of the following features:
A system for transmitting and creating a plurality of individual audio objects using a multichannel downmix and additional control data describing the objects, comprising: a spatial audio object encoder for encoding a plurality of audio objects in a downmix multichannel, information about multichannel downmix, and object parameters; or a spatial audio object decoder for decoding a multichannel downmix, information about the multichannel downmix, object parameters, and an object reproduction matrix into a second multichannel audio signal suitable for audio reproduction.
Figure 1a illustrates the spatial audio object coding (SAOC) operation, comprising an SAOC encoder 101 and a SAOC decoder 104. The spatial audio object encoder 101 encodes N objects in an object downmix consisting of K> 1 audio channels, according to encoder parameters. Information about the applied downmix weight matrix D is output by the SAOC encoder along with optional data concerning downmix power and correlation. Matrix D is often, but not necessarily always, constant over time and frequency, and therefore represents a relatively low amount of information. Finally, the SAOC encoder extracts object parameters for each object as a function of both time and frequency at a resolution defined by perceptual considerations. The spatial audio object decoder 104 takes the object downmix channels, the downmix information and the object parameters (generated by the encoder) as inputs and generates an output with M audio channels for presentation to the user. The reproduction of N objects on M audio channels makes use of a reproduction matrix provided as user input for the SAOC decoder.
Figure 1b illustrates the spatial audio object encoding operation reusing an MPEG Surround decoder. An SAOC decoder 104 taught by the current invention can be realized as an SAOC to MPEG Surround transcoder 102 and a stereo downmix based MPEG Surround decoder 103. A user-controlled reproduction matrix A of size M x N defines the target reproduction of the N objects at M audio channels. This matrix can be both time and frequency dependent and is the final output of a simpler user interface for manipulating audio objects (which can also make use of an externally provided scene description). In the case of a 5.1 speaker configuration the number of output audio channels is M = 6. The task of the SAOC decoder is to perceptively recreate the target reproduction of the original audio objects. The SAOC to MPEG Surround transcoder 102 takes as input the reproduction matrix A, the object downmix, the downmix sub-information including the D-down-mix weight matrix, and the object sub-information, and outputs a stereo downmix and MPEG Surround side information. When the transcoder is built in accordance with the current invention, a rear MPEG Surround decoder 103 fed with this data will produce M channel audio output with the desired properties.
A SAOC decoder taught by the current invention consists of an SAOC to MPEG Surround transcoder 102 and a stereo downmix based MPEG Surround decoder 103. A user controlled reproduction matrix A of size M x N defines the target reproduction of the N objects at M audio channels. This matrix can be both time and frequency dependent and is the final output of a simpler user interface for manipulating audio objects. In the case of a 5.1 speaker configuration the number of output audio channels is M = 6. The task of the SAOC decoder is to perceptively recreate the target reproduction of the original audio objects. The SAOC to MPEG Surround transcoder 102 takes as input the reproduction matrix A, the object downmix, the downmix sub-information including the D-down-mix weight matrix, and the object sub-information, and outputs a stereo downmix and MPEG Surround side information. When the transcoder is built in accordance with the current invention, a rear MPEG Surround decoder 103 fed with this data will produce M channel audio output with the desired properties.
Figure 2 illustrates the operation of a spatial audio object encoder (SAOC) 101 taught by the current invention. The N audio objects are fed into both a down-mixer 201 and an audio object parameter extractor 202. The downmix 201 downmixes the objects into an object downmix consisting of K> 1 channels of audio, according to the encoder parameters and also outputs downmix information. This information includes a description of the applied downmix weight matrix D and optionally whether the subsequent audio object parameter extractor operates in prediction mode, the parameters describing the power and correlation of the object downmix. As will be discussed in a later paragraph, the role of such additional parameters is to give access to the energy and correlation of subsets of reproduced audio channels in the case where the object parameters are expressed only relative to the downmix, being the main example the front / rear indications of a 5.1 speaker setup. The audio object parameter extractor 202 extracts object parameters according to the encoder parameters. The encoder control determines, based on the variation of time and frequency, which of two encoder modes is applied, the energy-based mode or the prediction-based mode. In power-based mode, the encoder parameters further contain information about a grouping of the N audio objects into P stereo objects and N-2P mono objects. Each mode will be further described by Figures 3 and 4.
ES 2 378 734 T3
Figure 3 illustrates an audio object parameter extractor 202 operating in a power-based mode. A grouping 301 into P stereo objects and N-2P mono objects is performed according to the grouping information contained in the encoder parameters. For each time-frequency interval considered, the following operations are then carried out. Two object powers and a normalized correlation are extracted for each of the P stereo objects by the stereo parameter extractor 302. A power parameter is extracted for each of the N-2P mono objects by the mono parameter extractor 303. The total set of N power parameters and P normalized correlation parameters is then encoded in 304 along with the grouping data to form the object parameters. The encoding may contain a normalization step with respect to the largest object power or with respect to the sum of extracted object powers.
FIG. 4 illustrates an audio object parameter extractor 202 operating in a prediction-based mode. For each time-frequency interval considered, the following operations are performed. For each of the N objects, a linear combination of the K object downmix channels is derived that matches the given object in a least squares sense. The K weights of this linear combination are called object prediction coefficients (OPC) and are calculated by the OPC extractor 401. The total set of NK OPCs are encoded in 402 to form the object parameters. The encoding may incorporate a reduction in the total number of OPCs based on linear interdependencies. As taught by the present invention, this total number can be reduced to max {K- (NK), 0} if the downmix weight matrix D has full rank.
Figure 5 illustrates the structure of a SAOC to MPEG Surround transcoder 102 as taught by the current invention. For each time-frequency interval, the downmix secondary information and object parameters are combined with the playback matrix by the parameter calculator 502 to form CLD, CPC, and ICC-type MPEG Surround parameters, and a matrix of 2xK size G downmix converter. The downmix converter 501 converts the object downmix to a stereo downmix by applying a matrix operation according to the G-matrices. In a simplified mode of the transcoder for K = 2, this matrix is the identity matrix and the downmix of objects are passed through without being altered as a stereo downmix. This mode is illustrated in the drawing with the selector switch 503 in position A, while the normal operating mode has the switch in position B. An additional advantage of the transcoder is its ability to be used as a standalone application in which the MPEG Surround parameters are ignored and the downmix converter output is used directly as stereo playback.
Figure 6 illustrates different modes of operation of a down-mix converter 501 as taught by the present invention. Given the downmixing of objects transmitted in the format of a bit stream output from a K channel audio encoder, this bit stream is first decoded by the audio decoder 601 into K audio signals in the audio domain. weather. These signals are then all frequency domain transformed by a hybrid MPEG Surround QMF filter bank in T / F unit 602. The time and frequency shifting matrix operation defined by the converter matrix data is performed on the resulting hybrid QMF domain signals by the matrixing unit 603 which outputs a stereo signal in the hybrid QMF domain. The hybrid synthesis unit 604 converts the signal in the stereo hybrid QMF domain to a signal in the stereo QMF domain. The hybrid QMF domain is defined in order to obtain better frequency resolution towards lower frequencies by means of a subsequent filtering of the QMF subbands. When this post-filtering is defined by Nyquist filter banks, the conversion from the hybrid to the conventional QMF domain consists of simply the sum of groups of hybrid subband signals, see [E. Schuijers, J. Breebart and H. Purnhagen "Low complexity parametric stereo coding" Proc 116th AES convention Berlin, Germany 2004, Preprint 6073]. This signal constitutes the first possible output format of the downmix converter as defined by selector switch 607 in position A. Such a signal in the QMF domain can be fed directly to the interface in the corresponding QMF domain of a decoder. MPEG Surround, and this is the most advantageous mode of operation in terms of delay, complexity and quality. The following possibility is obtained by performing a QMF filter bank synthesis 605 in order to obtain a signal in the stereo time domain. With the selector switch 607 in position B, the converter outputs a stereo digital audio signal that can also be fed to the time domain interface of an MPEG Surround back decoder, or played directly on a stereo playback device. The third possibility with the selector switch 607 in position C is obtained by encoding the stereo signal in the time domain with a stereo audio encoder 606. The output format of the downmix converter is then a stereo audio bitstream that is compatible with a core decoder contained in the MPEG decoder. This third mode of operation is suitable for the case where the SAOC to MPEG Surround transcoder is separated by the MPEG decoder by a connection that imposes restrictions on the bit rate, or in the case where the user wants to store a reproduction of a particular object for future reproduction.
Figure 7 illustrates the structure of an MPEG Surround decoder for a stereo downmix. The stereo downmix is converted to three intermediate channels using the two-to-three (TTT) box. These intermediate channels are further divided in two by the three boxes of one to two (OTT) to achieve the six channels of a 5.1 channel configuration.
Figure 8 illustrates a practical use case that includes a SAOC encoder. An 802 audio mixer outputs a stereo signal (L and R) that is typically composed by combining mixer input signals (in this case the
ES 2 378 734 T3 input channels 1-6) and optionally additional effects return inputs such as reverb etc. The mixer also outputs an individual channel (in this case channel 5) from the mixer. This can be done, for example, by means of commonly used mixer functionalities such as "direct outs" or "aux send" in order to output an individual channel after any of the insert processes (such as dynamics processing and EQ). . The stereo signal (L and R) and the individual channel output (obj5) are input to the 801 SAOC encoder, which is but a special case of the 101 SAOC encoder in Figure 1. However, it clearly illustrates a typical application where the obj5 audio object (containing, for example, speech) must undergo user-controlled level modifications on the decoder side while still being part of the stereo mix ( L and R). From the concept, it is also obvious that two or more audio objects can be connected to the “object input” panel in 801, and furthermore the stereo mix can be extended by a multi-channel mix such as a 5.1 mix.
In the text that follows, the mathematical description of the present invention will be set forth. For discrete complex signals x, y, the complex inner product and square norm (energy) is defined by i
Η<sup>2=</sup>(<sup>χ</sup>·<sup>χ</sup>) = ΣΙ<sup>χ</sup>(*) Γ where y (k) indicates the complex conjugate signal of y (k). All the signals considered in this case are subband samples from a modulated filter bank or FFT analysis with discrete time signal window function. It is understood that these subbands must be transformed back to the discrete time domain by corresponding synthesis filter bank operations. A signal block of L samples represents the signal in a time and frequency interval that is part of the perceptually motivated tiling of the time-frequency plane that is applied for the description of signal properties. In this situation, the given audio objects can be represented as N rows of length L in a matrix, '5, (0) s, (l) ... í | (£ -l)<sub>s =</sub> Μθ) MD - <sub>(2)</sub>
A (0) ^ (1) ...
The downmix weight matrix D of size K x N, where K> 1 determines the downmix signal of K channels in the form of a matrix with K rows through matrix multiplication
X = DS, (3)
The M x N size user controlled object reproduction matrix A determines the target reproduction of M channels of the audio objects in the form of a matrix with M rows through matrix multiplication
Y = AS.
(4)
Ignoring for the moment the effects of core audio coding, the task of the SAOC decoder is to generate an approximation in the perceptual sense of the target reproduction Y of the original audio objects, given the reproduction matrix A, the downmix X, the downmix matrix D and object parameters.
The object parameters in the energy mode taught by the present invention carry information about the covariance of the original objects. In a deterministic version suitable for later derivation and also descriptive of typical encoder operations, this covariance is given in non-normalized form by the matrix product SS * where the asterisk indicates the complex conjugate transpose matrix operation. Thus, the energy mode object parameters provide a semi-defined positive matrix EN x N such that, possibly up to a scale factor,
SS * »E. (5)
ES 2 378 734 T3
Prior art audio object encoding often considers an object model in which all objects are uncorrelated. In this case, the matrix E is diagonal and only contains an approximation to the object energies S<sub>n</sub> = || s<sub>n</sub>||<sup>2</sup> for n = 1,2, ..., N. The object parameter extractor according to figure 3, allows a significant refinement of this idea, particularly relevant in cases where the objects are provided as stereo signals for which the assumptions about no correlation do not hold. A grouping of P selected stereo pairs of objects is described by the sets of indices {(n<sub>p</sub>, m<sub>p</sub>), p = 1,2,., P}. For these stereo pairs the correlation (s<sub>n</sub>, s<sub>m</sub>) and the complex, real, or absolute value of the normalized correlation (ICC)
<img file="ES2378734T3_D0001.tif" />
it is extracted by stereo parameter extractor 302. In the decoder, the ICC data can then be combined with the energies to form an E matrix with 2P inputs off the diagonal. For example, for a total of N = 3 objects of which the first two consist of a single pair (1,2), the transmitted energy and the correlation data are (Si, S2, S3) and pi. 2. In this case, the combination in matrix E gives
<img file="ES2378734T3_D0002.tif" />
The object parameters in the prediction mode taught by the present invention are intended to make a matrix of object prediction coefficients (OPC) C of N x K available to the decoder so that
S «CX = CDS.
(7)
In other words, for each object there is a linear combination of the downmix channels so that the object can be roughly retrieved by
«. (*)« (*) + · .. + c<sub>nK</sub>x<sub>K</sub> (k). (8)
In a preferred embodiment, the OPC extractor 401 solves the normal equations
<img file="ES2378734T3_D0003.tif" />
or, for the more attractive real value OPC, solve
CRefXX '} = Re {SX'}. (10)
In both cases, assuming a real-valued mixdown weight matrix D, and a non-singular downmix covariance, it follows by multiplication from the left with D that
DC = I, (11) where I is the identity matrix of size K. If D has complete rank, it follows by elementary linear algebra that the set of solutions of (9) can be parameterized by parameters max {K- (NK) , 0}. This is taken advantage of in the 402 co-encoding of the OPC data. The complete prediction matrix C can be recreated in the decoder from the reduced set of parameters and the downmix matrix.
For example, consider for a stereo downmix (K = 2) the case of three objects (N = 3) that comprise a stereo music track (s1, s2) and a single instrument or voice track with center pan s3. The downmix matrix is
ES 2 378 734 T3
<img file="ES2378734T3_D0004.tif" />
<img file="ES2378734T3_D0005.tif" />
<img file="ES2378734T3_D0006.tif" />
That is, the downmix left channel is and the right channel is Los
OPC for the individual track are intended to approximate S3 -031X1 + 032X2 and equation (11) can be solved in this case for y
<img file="ES2378734T3_D0007.tif" />
<img file="ES2378734T3_D0008.tif" />
getting enough is given by K (N- K) = 2 · (3-2) = 2.
Therefore, the number of OPCs
The OPCs c<sub>31</sub>, c<sub>32</sub> can be found from the normal equations
<img file="ES2378734T3_D0009.tif" />
SAOC to MPEG Surround Transcoder
Referring to figure 7, the M = 6 output channels of configuration 5.1 are (y-ι, yi, ..., ye) = (lf, l<sub>s</sub>, rf, rs, o, lfe). The transcoder should output a stereo (/ o, ro) downmix and parameters for the TTT and OTT boxes. Since the focus is now on the stereo downmix, it will now be assumed that K = 2. Since both object parameters and MPS TTT parameters exist in both energy mode and prediction mode, all four combinations must be considered. Power mode is a suitable option, for example, in case the downmix audio encoder is not a waveform encoder in the considered frequency range. It is understood that the MPEG Surround parameters derived in the following text must be properly quantized and encoded prior to transmission. To better clarify the four combinations mentioned above, these comprise
1. Object parameters in energy mode and transcoder in prediction mode
2. Object parameters in power mode and transcoder in power mode
3. Object Parameters in Prediction Mode (OPC) and Transcoder in Prediction Mode
Four. Object Parameters in Prediction Mode (OPC) and Transcoder in Power Mode
If the audio downmix encoder is a waveform encoder in the considered frequency range, the object parameters can be in both power and prediction mode, but the transcoder should preferably operate in prediction mode. If the audio downmix encoder is not a waveform encoder in the considered frequency range, the object encoder and transcoder must both operate in power mode. The fourth combination is the one with the least relevance so the following description will address the first three combinations only.
Object Parameters Given in Power Mode
In power mode, the data available to the transcoder is described by the matrix triplet (D, E, A). The MPEG Surround OTT parameters are obtained by performing energy and correlation estimates on a virtual reproduction derived from the transmitted parameters and the A 6 x N reproduction matrix. The target six-channel covariance is given by
YY * = AS (AS) * = A (SS ') A',
Inserting (5) in (13) the approximation is obtained
YY * «¡F = AEA *, (13) (14) which is fully defined by the available data. Let's say fu are the elements of F. So, the CLD and ICC parameters are read from
ES 2 378 734 T3
CLD<sub>0</sub>
<img file="ES2378734T3_D0010.tif" />
<img file="ES2378734T3_D0011.tif" />
CLD<sub>1</sub>
<img file="ES2378734T3_D0012.tif" />
(15) (16) (17)
<img file="ES2378734T3_D0013.tif" />
<img file="ES2378734T3_D0014.tif" />
(19) where φ is either the absolute value operator φ (ζ) = \ z \ or the real value operator φ (z) = Re {z}.
As an illustrative example, consider the case of three objects previously described in relation to equation (12). Let's say the reproduction matrix is given by
1 0'
I or
A = <sup>1</sup> ° .
0 0
0 1
0 1
Objective reproduction therefore consists in placing object 1 between front-right and surround-right, object 2 between front-left and surround-left, and object 3 at front-right, center and lfe. Also suppose for the sake of simplicity that the three objects are uncorrelated and all have the same energy so that
<img file="ES2378734T3_D0015.tif" />
In this case, the right side of formula (14) becomes
1 0 0 0 0'
1 0 0 00
0 2 1 11
0 1 100
0 10 11
0 10 11
Inserting the appropriate values in the formulas (15) - (19) we obtain then
ES 2 378 734 T3
<img file="ES2378734T3_D0016.tif" />
<img file="ES2378734T3_D0017.tif" />
<img file="ES2378734T3_D0018.tif" />
<img file="ES2378734T3_D0019.tif" />
<img file="ES2378734T3_D0020.tif" />
As a consequence, the MPEG Surround decoder will be instructed to use some front right to surround right decorrelation, but not front right to surround left decorrelation.
For the MPEG Surround TTT parameters in prediction mode, the first step is to form a matrix of q = \ lj¿.
reduced reproduction A<sub>3</sub> of size 3 x N for the combined channels (l, r, qc) where. It is true that A<sub>3</sub> = D36A where the 6 to 3 partial downmix matrix is defined by
<td></td><td> ’<sup>w</sup>i</td><td><sup>W</sup>l</td><td> 0</td><td> 0</td><td> 0</td><td> 0 ‘</td><td></td>
<td></td><td> 0</td><td> 0</td><td><sup>W</sup>2</td><td><sup>W</sup>2</td><td> 0</td><td> 0</td><td> • (20)</td>
<td></td><td> 0</td><td> 0</td><td> 0</td><td> 0</td><td> <1<sup>W</sup>3</td><td>qwj</td><td></td>
The partial down-mixing weights Wp, p = 1,2,3 are adjusted so that the energy of Wp (y2p-1 + y2p) is equal to the sum of energies || y2p-i ||<sup>2</sup>+ || y2p ||<sup>2</sup> up to a limiting factor. All the data required to derive the partial downmix matrix D36 is available in F. Next, a prediction matrix C3 of size 3x2 is produced such that (21)
Such a matrix is preferably derived by first considering the normal equations
<img file="ES2378734T3_D0021.tif" />
The solution to the normal equations gives the best possible waveform match for (21) given the object covariance model E. Some post-processing of matrix C is preferable.<sub>3</sub>, including row factors for a total or individual channel based on prediction loss compensation.
To illustrate and clarify the above steps, consider a continuation of the specific six-channel playback example given above. As for the matrix elements of F, the down-mix weights are solutions to the equations
<img file="ES2378734T3_D0022.tif" />
which in the specific example becomes
ES 2 378 734 T3 ^ (1 + 1 + 21) = 1 + 1 ^ (2 + 1 + 2-1) = 2 + 1 ·, w, (1 + 1 + 2-1) = 1 + 1
So that<sub>;</sub> (^,^,^) = (1/^,7375,1/72)
Insertion in (20) gives
<td></td><td> ’ 0</td><td> 77 0 '</td>
<td>Ά - -</td><td></td><td>0 You</td>
<td></td><td> 0</td><td> 0 1</td>
Solving the system of equations C3 (DED) = A3ED then finds, (now commuting to finite precision),
<img file="ES2378734T3_D0023.tif" />
Matrix C3 contains the best weights to get an approximation to the desired object reproduction to the combined channels (l, r, qc) from the downmix of objects. This general type of matrix operation cannot be implemented by the MPEG Surround decoder, which is restricted to a limited space of TTT matrices by using only two parameters. The object of the inventive downmix converter is to pre-process the object downmix so that the combined effect of the pre-processing and the MPEG Surround TTT matrix is identical to the desired upmix described by C3.
In MPEG Surround, the TTT matrix for the prediction of (l, r, qc) from (/ 0, r0) is parameterized by three parameters (α, β, γ) by
<img file="ES2378734T3_D0024.tif" />
the
J3-1 β + 2 \ -β.
(22)
The downmix converter matrix G taught by the present invention is obtained by choosing γ = 1 and solving the system of equations
<img file="ES2378734T3_D0025.tif" />
As can be easily verified, it is true that D<sub>TTT</sub>C<sub>TTT</sub> = I, where I is the two-by-two identity matrix and
<img file="ES2378734T3_D0026.tif" />
Therefore, a matrix multiplication from the left by D<sub>TTT</sub> on both sides of (23) leads to
G = D<sub>m</sub>C<sub>3</sub>. (25)
In the generic case, G can be reversed and (23) has a unique solution for CTTT that meets DTTTCTTT = I. The TTT parameters (α, β) are determined by this solution.
ES 2 378 734 T3
For the specific example considered above, it can easily be verified that the solutions are given by
<img file="ES2378734T3_D0027.tif" />
Note that a main part of the stereo downmix is swapped between left and right for this converter matrix, reflecting the fact that the playback example puts objects that are in the left object downmix channel on the right side. from the sound scene and vice versa. Such behavior is impossible to obtain from an MPEG Surround decoder in stereo mode.
If it is impossible to apply a down-mix converter, a less than optimal procedure may be developed as follows. For the MPEG Surround TTT parameters in power mode, what is required is the power distribution of the combined channels (l, r, c). Therefore the relevant CLD parameters can be derived directly from the elements of F through
<img file="ES2378734T3_D0028.tif" />
<img file="ES2378734T3_D0029.tif" />
In this case, it is appropriate to use only a diagonal matrix G with positive inputs for the downmix converter. It is operational to get the correct power distribution of the downmix channels prior to the TTT upmix. With the six to two D channel downmix matrix<sub>26</sub> = D<sub>TTT</sub>D<sub>36</sub> and definitions from
<img file="ES2378734T3_D0030.tif" />
<img file="ES2378734T3_D0031.tif" />
it is simply chosen
<img file="ES2378734T3_D0032.tif" />
A further observation is that such a diagonal downmix converter can be bypassed from the Object to MPEG Surround transcoder and implemented by activating the arbitrary downmix gain (ADG) parameters of the MPEG Surround decoder. These gains will then occur in the logarithmic domain by means of ADG1 = 10 log10 (wi¡ / z¡i) for / = 1.2.
Object Parameters Given in Prediction Mode (OPC)
In object prediction mode, the available data is represented by the matrix triplet (D, C, A) where C is the Nx2 matrix containing the N OPC pairs. Due to the relative nature of the prediction coefficients, it will also be necessary for the estimation of energy-based MPEG Surround parameters to have access to an approximation to the 2x2 covariance matrix of the down-mixing of objects, (31)
ES 2 378 734 T3
This information is preferably transmitted from the object encoder as part of the downmix secondary information, but could also be estimated at the transcoder from measurements made on the received downmix, or indirectly derived from (D, C) by considerations of approximate object model. Given Z, the object covariance can be estimated by inserting the predictive model Y = CX, giving
E = CZC *, (32) and all MPEG Surround OTT and energy mode TTT parameters can be estimated from E as in the case of energy-based object parameters. However, the great advantage of using OPC arises in combination with MPEG Surround TTT parameters in prediction mode. In this case, the waveform approximation D36 Y = A<sub>3</sub>CX immediately gives the reduced prediction matrix
C<sub>3</sub>= A<sub>5</sub>C, (32) from which the remaining steps to achieve the parameters TTT (α, β) and the downmix converter are similar to the case of object parameters provided in power mode. In fact, the steps of formulas (22) to (25) are completely identical. The resulting matrix G is fed to the downmix converter and the TTT parameters (α, β) are transmitted to the MPEG Surround decoder.
Standalone Down Mix Converter Application for Stereo Playback
In all of the cases described above, the object to stereo downmix converter 501 outputs an approximation to a stereo downmix of the 5.1 channel reproduction of the audio objects. This stereo reproduction can be expressed by a matrix A2 2χΛ / defined by A2 = D26A. In many applications this downmix is interesting in its own right and a direct manipulation of the A2 stereo reproduction is attractive. Consider again as an illustrative example the case of a stereo track with a mono voice track with superimposed center pan encoded following a special case of the method shown in figure 8 and discussed in the section on formula (12). A user control of the voice volume can be done by playing
<img file="ES2378734T3_D0033.tif" />
v / V21 v / Vzj '(33) where v is the voice to music ratio control. The down-mix converter matrix design is based on gds »a<sub>2</sub>s.
(34)
For prediction-based object parameters, simply insert the approximation S = CDS and obtain the converter matrix G = A2C. For energy-based object parameters, the normal equations are solved
G (DED ') = A<sub>2</sub>ED '. (35)
Figure 9 illustrates a preferred embodiment of an audio object encoder in accordance with one aspect of the present invention. The audio object encoder 101 has already been generally described in connection with the preceding figures. The audio object encoder to generate the encoded object signal uses the plurality of audio objects 90 indicated in FIG. 9 when entering a down-mixer 92 and object parameter generator 94. In addition, the audio object encoder 101 includes the downmix information generator 96 for generating downmix information 97 indicating a distribution of the plurality of audio objects on at least two downmix channels indicated at 93 when they are output from the downward mixer 92.
The object parameter generator is for generating object parameters 95 for the audio objects, wherein the object parameters are calculated so that reconstruction of the audio object is possible using the object parameters and at least two channels 93 down mix. Notably, however, this reconstruction does not take place on the encoder side, but takes place on the decoder side. However, the object parameter generator on the encoder side calculates the object parameters for the objects 95 so that this total reconstruction can be performed on the decoder side.
In addition, the audio object encoder 101 includes an output interface 98 for generating the encoded audio object signal 99 using the downmix information 97 and object parameters 95. Depending on the application, downmix channels 93 can also be used and encoded into the audio object signal
ES 2 378 734 T3 coded. However, there may also be situations where the output interface 98 generates an encoded audio object signal 99 that does not include the downmix channels. This situation can arise when any downmix channel to be used on the decoder side is already on the decoder side, so that the downmix information and object parameters for the audio objects are transmitted separately from downmix channels. Such a situation is useful when the object downmix channels 93 can be acquired separately from the object parameters and the downmix information for a lesser amount of money, and the object parameters and the downmix information can be acquired. for an additional amount of money in order to provide the user on the decoder side with added value.
Without the object parameters and downmix information, a user can play the downmix channels as a stereo or multichannel signal depending on the number of channels included in the downmix. Of course, the user could also reproduce a mono signal by simply adding the at least two transmitted object downmix channels. To increase the flexibility of playing and listening quality and usability, the object parameters and the downmix information allow the user to form a flexible reproduction of the audio objects in any intended audio reproduction configuration, such as a stereo system, a multichannel system or even a wave field synthesis system. While wavefield synthesis systems are not very popular yet, multichannel systems such as 5.1 systems or 7.1 systems are becoming increasingly popular in the consumer market.
Figure 10 illustrates an audio synthesizer for generating output data. For this purpose, the audio synthesizer includes an output data synthesizer 100. The output data synthesizer receives, as input, the downmix information 97 and the audio object parameters 95, and probably the intended audio source data such as a positioning of the audio sources or a specified volume. by the user of a specific font, where the font should be when played, as indicated in 101.
The output data synthesizer 100 is for generating output data that can be used to create a plurality of output channels of a predefined audio output configuration representing a plurality of audio objects. In particular, the output data synthesizer 100 is operative for the use of the downmix information 97, and the audio object parameters 95. As discussed in connection with Figure 11 below, the output data can be data from a wide variety of different useful applications, including specific reproduction of output channels or including only a reconstruction of the source signals or that include a transcoding of parameters into spatial playback parameters for a spatial upmix setup without any specific playback of output channels, but for example to store or transmit such spatial parameters.
The general application scenario of the present invention is summarized in Fig. 14. There is an encoder side 140 that includes the audio object encoder 101 that receives, as input, N audio objects. The output of the preferred audio object encoder comprises, in addition to the downmix information and the object parameters not shown in FIG. 14, the K downmix channels. The number of downmix channels according to the present invention is greater than or equal to two.
The downmix channels are transmitted to a decoder side 142, which includes a spatial upmixer 143. The spatial uplink mixer 143 may include the inventive audio synthesizer, when the audio synthesizer is operated in a transcoder mode. However, when the audio synthesizer 101 as illustrated in FIG. 10 operates in a spatial up mixer mode, then the spatial up mixer 143 and the audio synthesizer are the same device in this embodiment. The spatial up mixer generates M output channels to be played through M speakers. These speakers are placed in predefined spatial locations and together represent the predefined audio output settings. An output channel of the predefined audio output configuration can be thought of as a digital or analog speaker signal to be sent from an output of the spatially up mixer 143 to the input of a speaker at a predefined position among the plurality of predefined positions. of the predefined audio output settings. Depending on the situation, the number of M output channels may be equal to two when performing stereo playback. However, when performing multi-channel playback, then the number of M output channels is greater than two. Typically, there will be a situation where the number of downmix channels is smaller than the number of output channels due to a requirement of a transmit link. In this case, M is greater than K and can be even much greater than K, such as doubling the size or even more.
Figure 14 further includes various matrix notations in order to illustrate the functionality of the encoder side of the invention and the decoder side of the invention. Generally, blocks of sample values are processed. Therefore, as indicated in equation (2), an audio object is represented as a line of L sample values. The matrix S has N lines that correspond to the number of objects and L columns that correspond to the number of samples. Matrix E is calculated as indicated in equation (5) and has N columns and N lines. The matrix E includes the object parameters when the object parameters are provided in the power mode. For uncorrelated objects, the matrix E has, as indicated above in connection with equation (6), only elements on the main diagonal, where an element on the main diagonal gives the energy of an audio object. All the
ES 2 378 734 T3 off-diagonal elements represent, as indicated above, a correlation of two audio objects, which is specifically useful when some objects are two channels of the stereo signal.
Depending on the specific embodiment, equation (2) is a signal in the time domain. Then, a single energy value is generated for the entire band of audio objects. Preferably, however, the audio objects are processed by a time / frequency converter that includes, for example, a type of transform or a filter bank algorithm. In the latter case, equation (2) is valid for each subband so that a matrix E is obtained for each subband and, of course, each time frame.
The downmix channel matrix X has K lines and L columns and is calculated as indicated in equation (3). As indicated in equation (4), the M output channels are calculated using the N objects by applying the so-called reproduction matrix A to the N objects. Depending on the situation, the N objects can be regenerated on the decoder side using the downmix and object parameters and the reproduction can be applied to the reconstructed object signals directly.
Alternatively, the downmix can be transformed directly to the output channels without explicit calculation of the source signals. Generally, the reproduction matrix A indicates the positioning of the individual sources with respect to the predefined audio output settings. If you had six objects and six output channels, then each object could be placed on each output channel and the reproduction matrix would reflect this scheme. However, if it is desired to place all objects between two output speaker locations, then the reproduction matrix A would appear different and reflect this different situation.
The reproduction matrix or, more generally, the predicted positioning of the objects and also a predicted relative volume of the audio sources can generally be calculated by an encoder and transmitted to the decoder as a so-called scene description. In other embodiments, however, this scene description may be self-generated to generate the user-specific upmix for the user-specific audio output setting. Therefore, a transmission of the scene description is not necessarily required, but the scene description can also be generated by the user in order to fulfill the wishes of the user. The user might want to place, for example, certain audio objects in places that are different from where these objects were when these objects were generated. There are also cases where the audio objects are designed on their own and do not have any "original" location relative to the other objects. In this situation, the relative location of the audio sources is generated by the user for the first time.
Returning to FIG. 9, a down mixer 92 is illustrated. The downmixer is for the downmix of the plurality of audio objects in the plurality of downmix channels, wherein the number of audio objects is greater than the number of downmix channels, and wherein the downmixer is coupled to the downmix information generator such that the distribution of the plurality of audio objects to the plurality of downmix channels is carried out as indicated in the downmix information. The downmix information generated by the downmix information generator 96 in FIG. 9 may be created automatically or manually adjusted. It is preferred to provide the downmix information with a resolution lower than the resolution of the object parameters. Therefore, bits of secondary information can be saved without further loss of quality, since it has been shown that fixed downmix information is sufficient for a given piece of audio or a downmix situation that only changes slowly, which does not necessarily have to be selective in frequency. In one embodiment, the downmix information represents a downmix matrix having K lines and N columns.
The value in a downmix matrix line has a certain value when the audio object corresponding to this value in the downmix matrix is on the downmix channel represented by the downmix matrix row. When an audio object is included in more than one downmix channel, the values in more than one row of the downmix matrix have a certain value. However, it is preferred that the square values when added together for a single audio object add up to 1.0. However, other values are possible as well. Additionally, audio objects can be input to one or more downmix channels with various levels, and these levels can be indicated by weights in the downmix matrix that are different from one and do not add up to 1.0 for a given audio object.
When the downmix channels are included in the encoded audio object signal generated by the output interface 98, the encoded audio object signal may for example be a time multiplexing signal in a certain format. Alternatively, the encoded audio object signal can be any signal that allows separation of the object parameters 95, the downmix information 97, and the downmix channels 93 on a decoder side. In addition, the output interface 98 may include encoders for the object parameters, the downmix information, or the downmix channels. The encoders for the object parameters and downmix information can be differential encoders and / or entropy encoders, and the encoders for the downmix channels can be mono or stereo audio encoders such as MP3 encoders or AAC encoders. . All of these encoding operations result in additional data compression in order to further decrease the required data rate for the encoded audio object signal 99.
ES 2 378 734 T3
Depending on the specific application, the downmixer 92 is operative to include the stereo representation of background music on the at least two downmix channels and further feeds the voice track on the at least two downmix channels in a predefined ratio. . In this embodiment, a first background music channel is within the first downmix channel and the second background music channel is within the second downmix channel. This results in optimal playback of stereo background music on a stereo playback device. However, the user can still modify the position of the voice track between the left stereo speaker and the right stereo speaker. Alternatively, the first and second background music channels can be included in one downmix channel and the voice track can be included in the other downmix channel. Therefore, by eliminating a downmix channel, the vocal track can be completely separated from the background music, which is particularly suitable for karaoke applications. However, the stereo reproduction quality of the background music channels will suffer due to object parameterization, which is naturally a lossy compression method.
A down-mixer 92 is adapted to perform a sample-by-sample summation in the time domain. This addition uses samples from audio objects to be down-mixed into a single down-mix channel. When an audio object is to be fed into a downmix channel by a certain percentage, pre-weighting takes place before the summation process with by samples. Alternatively, the summation can also take place in the frequency domain, or a subband domain, that is, in a domain after the time / frequency conversion. Thus, even downmixing could be done in the filterbank domain when the time / frequency conversion is a filterbank or in the transform domain when the time / frequency conversion is a type of FFT, MDCT, or whatever. another transformed.
In one aspect of the present invention, the object parameter generator 94 generates power parameters and, additionally, the correlation parameters between two objects when two audio objects together represent the stereo signal, as is clear from the equation below ( 6). Alternatively, the object parameters are prediction mode parameters. Figure 15 illustrates algorithm steps or means of a computing device to calculate these audio object prediction parameters. As discussed in connection with equations (7) to (12), some statistical information has to be computed on the downmix channels in matrix X and audio objects in matrix S. Particularly, block 150 illustrates the first stage of calculating the real part of S · X * and the real part of X · X *. These real parts are not just numbers but matrices, and these matrices are determined in one embodiment through the notations in equation (1) when the subsequent embodiment to equation (12) is considered. Generally, the values from step 150 can be calculated using data available in the audio object encoder 101. Then, the prediction matrix C is calculated as illustrated in step 152. In particular, the system of equations is solved as is known in the art so that all the values of the prediction matrix C having N lines and K columns are obtained. Generally, the weighting factors cn, i as given in equation (8) are calculated such that the weighted linear addition of all the downmix channels reconstructs a corresponding audio object as best as possible. This prediction matrix results in better reconstruction of audio objects when the number of downmix channels increases.
Figure 11 will be discussed in more detail below. In particular, Figure 7 illustrates various kinds of output data that can be used to create a plurality of output channels of a predefined audio output configuration. Line 111 illustrates a situation where the output data from the output data synthesizer 100 is reconstructed audio sources. The input data required by the output data synthesizer 100 to reproduce the reconstructed audio sources includes the downmix information, the downmix channels, and the audio object parameters. To reproduce the reconstructed sources, however, an output configuration and an intended positioning of the audio sources themselves in the spatial audio output configuration are not necessarily required. In this first mode indicated by mode number 1 in FIG. 11, the output data synthesizer 100 will output reconstructed audio sources. In the case of prediction parameters as audio object parameters, the output data synthesizer 100 operates as defined by equation (7). When the object parameters are in power mode, then the output data synthesizer uses an inverse of the downmix matrix and the power matrix to reconstruct the source signals.
Alternatively, the output data synthesizer 100 operates as a transcoder as illustrated for example at block 102 in Figure 1b. When the output synthesizer is a type of transcoder for generating spatial mixer parameters, the downmix information, the audio object parameters, the output configuration, and the intended positioning of the sources are required. In particular, the output configuration and the intended positioning are provided through the reproduction matrix A. However, the downmix channels are not required to generate the spatial mixer parameters as will be discussed in more detail in connection with the figure 12. Depending on the situation, the spatial mixer parameters generated by the output data synthesizer 100 can then be used by a direct spatial mixer such as an MPEG-surround mixer to up-mix the down-mix channels. This embodiment does not necessarily need to modify the object downmix channels, but can provide a simple conversion matrix that only has diagonal elements as discussed in equation (13). In mode 2 as indicated by 112 in FIG. 11, the output data synthesizer 100 will therefore output spatial mixer parameters and preferably the conversion matrix G as indicated in equation
ES 2 378 734 T3 (13), which includes gains that can be used as arbitrary downmix (ADG) gain parameters of the MPEG-surround decoder.
In mode number 3 as indicated by 113 of FIG. 11, the output data includes spatial mixer parameters in a conversion matrix such as the conversion matrix illustrated in connection with equation (25). In this situation, the output data synthesizer 100 does not necessarily have to perform actual downmix conversion to convert the object downmix to a stereo downmix.
A different mode of operation indicated by mode number 4 on line 114 in FIG. 11 illustrates the output data synthesizer 100 of FIG. 10. In this situation, the transcoder is operated as indicated by 102 in FIG. 1b and outputs not only spatial mixer parameters but also outputs a converted downmix. However, it is no longer necessary to output the conversion matrix G in addition to the converted downmix. Outputting the converted downmix and spatial mixer parameters is sufficient as indicated by Figure 1b.
Mode number 5 indicates another use of the output data synthesizer 100 illustrated in Figure 10. In this situation indicated by line 115 in Figure 11, the output data generated by the output data synthesizer does not include any parameters. of spatial mixer but only include a conversion matrix G as indicated by equation (35) for example or actually include the output of the stereo signals themselves as indicated in 115. In this embodiment, only a stereo reproduction is of interest and no spatial mixer parameter is required. To generate the stereo output, however, all the available input information is required as indicated in Figure 11.
Another output data synthesizer mode is indicated by mode number 6 on line 116. In this case, output data synthesizer 100 generates multi-channel output, and output data synthesizer 100 would be similar to item 104 in figure 1b. For this purpose, the output data synthesizer 100 requires all available input information and outputs a multi-channel output signal that has more than two output channels to be produced by a corresponding number of loudspeakers that are to be placed in position positions. speakers based on the predefined audio output settings. Such a multi-channel output is a 5.1 output, a 7.1 output, or just a 3.0 output that has a left speaker, a center speaker, and a right speaker.
Reference is now made to Fig. 11 to illustrate an example for calculating various parameters from the known parameterization concept of Fig. 7 of the MPEG-surround decoder. As indicated, FIG. 7 illustrates an MPEG-surround decoder side parameterization starting from stereo down-mix 70 having a left down-mix channel lü and a right down-mix channel ιό. Conceptually, both downmix channels are entered into a so-called two-to-three box 71. The two to three box is controlled by various input parameters 72. Box 71 generates three output channels 73a, 73b, 73c. Each output channel is entered in a one-to-two box. This means that channel 73a is entered in box 74a, channel 73b is entered in box 74b, and channel 73c is entered in box 74c. Each box emits two output channels. Box 74a outputs a front left channel l<sub>F</sub> and a surround left channel l<sub>s</sub>. In addition, box 74b outputs a front right rf channel and a right surround channel rs. In addition, box 74c outputs a center channel c and a low-frequency enhancement channel Ife. Notably, the entire upmix from the downmix channels 70 to the output channels is performed using a matrix operation, and the tree structure as shown in Figure 7 is not necessarily implemented stage by stage but can be implemented through a single or multiple array operations. Furthermore, the intermediate signals indicated by 73a, 73b and 73c are not explicitly calculated by a certain embodiment, but are illustrated in FIG. 7 for the sake of illustration only. In addition, cells 74a, 74b receive some residual signals res<sub>1</sub><sup>OTT</sup>, res<sub>2</sub><sup>OTT</sup> that can be used to introduce a certain randomness in the output signals.
As shown from the MPEG-surround decoder, box 71 is controlled by either CPC prediction parameters or CLDttt power parameters. For upmixing from two channels to three channels, at least two prediction parameters CPC1, CPC2 or at least two power parameters CLD are required<sup>1</sup>ttt and cld<sup>2</sup> ttt. Furthermore, the correlation measure ICCttt can be put in box 71 which is, however, only an optional feature that is not used in an embodiment of the invention. Figures 12 and 13 illustrate the necessary steps and / or means to calculate all the parameters CPC / CLDttt, CLD0, CLD1, ICC1, CLD2, ICC2 from the object parameters 95 of Figure 9, the downmix information 97 of Figure 9 and the intended positioning of the audio sources, for example scene description 101 as illustrated in Figure 10. These parameters are for the predefined audio output format of a 5.1 surround system.
Naturally, the specific calculation of parameters for this specific implementation can be adapted for other output formats or parameterizations in view of the teachings of this document. Furthermore, the sequence of steps or arrangement of means in Figures 12 and 13a, b is exemplary only and can be changed within the logical sense of mathematical equations.
At step 120, a reproduction matrix A is provided. The reproduction matrix indicates where the source of the plurality of sources is to be located in the context of the predefined output configuration. Step 121 illustrates the
ES 2 378 734 T3 derivation of the partial downmix matrix D36 as indicated in equation (20). This matrix reflects the downmix situation from six output channels to three channels and is 3xN in size. When it is intended to generate more output channels than the 5.1 configuration, such as an 8 channel (7.1) output configuration, then the matrix determined at block 121 would be a D38 matrix. In step 122, a reduced reproduction matrix A3 is generated by multiplying matrix D36 and the total reproduction matrix as defined in step 120. In step 123, the downmix matrix D is input. downstream D can be recovered from the encoded audio object signal when the matrix is fully included in this signal. Alternatively, the downmix matrix could be parameterized for example for the specific example of the downmix information and the downmix matrix G.
Furthermore, the object energy matrix is provided in step 124. This object energy matrix is reflected by the object parameters for the N objects and can be extracted from the imported or reconstructed audio objects using a certain reconstruction rule. This reconstruction rule can include entropy decoding, etc.
In step 125, the "reduced" prediction matrix C3 is defined. The values of this matrix can be calculated by solving the system of linear equations as indicated in step 125. Specifically, the elements of matrix C3 can be calculated by multiplying the equation on both sides by an inverse of (DED *).
In step 126, the conversion matrix G is calculated. The conversion matrix G has a size of KxK and is generated as defined by equation (25). To solve the equation in step 126, the specific DTTT matrix is to be provided as indicated by step 127. An example for this matrix is given by equation (24) and the definition can be derived from the corresponding equation for CTTT such as defined in equation (22). Equation (22), therefore, defines what will be done in step 128. Step 129 defines the equations to calculate the CTTT matrix. As soon as the CTTT matrix is determined according to the equation in block 129, the parameters α, β and γ can be produced, which are the CPC parameters. Preferably, γ is set to 1 so the only remaining CPC parameters entered in block 71 are α and β.
The remaining parameters needed for the schematic in Figure 7 are the parameters entered in blocks 74a, 74b, and 74c. The calculation of these parameters is discussed in connection with Figure 13a. In step 130, the reproduction matrix A is provided. The size of the reproduction matrix A is N lines for the number of audio objects and M columns for the number of output channels. This reproduction matrix includes the scene vector information, when a scene vector is used. Generally, the reproduction matrix includes the information to place an audio source at a certain position in an output configuration. When considering, for example, the reproduction matrix A under equation (19), it becomes apparent how a certain placement of audio objects can be encoded within the reproduction matrix. Naturally, other ways of indicating a certain position can be used, such as by values not equal to 1. Also, when using values that are less than 1 on the one hand and are greater than 1 on the other hand, the loudness of certain objects of audio can be influenced as well.
In one embodiment, the reproduction matrix is generated on the decoder side without any information from the encoder side. This allows a user to place the audio objects anywhere the user wishes without paying attention to a spatial relationship of the audio objects in the encoder configuration. In another embodiment, the relative or absolute location of audio sources can be encoded on the encoder side and transmitted to the decoder as a kind of a scene vector. Then, on the decoder side, this information about audio source locations that is preferably independent of an intended audio playback configuration is processed to result in a playback matrix that reflects the custom audio source locations to the specific audio output settings.
In step 131, the object energy matrix E which has already been discussed in connection with step 124 of FIG. 12 is provided. This matrix is of the size of NxN and includes the audio object parameters. In one embodiment, such an object energy matrix is provided for each subband and each block samples in the time domain or samples in the subband domain.
In step 132, the output energy matrix F is calculated. F is the covariance matrix of the output channels. Since the output channels are, however, still unknown, the output energy matrix F is calculated using the reproduction matrix and the energy matrix. These matrices are provided in steps 1 30 and 1 31 and are readily available on the decoder side. Then, the specific equations (15), (16), (17), (18) and (19) are applied to calculate the channel level difference parameters CLD0, CLD1, CLD2 and the coherence parameters between channels ICC1 and ICC2 so that the parameters for boxes 74a, 74b, 74c are available. Notably, the spatial parameters are calculated by combining the specific elements of the output energy matrix F.
After step 133, all the parameters for a space up mixer are available, such as the space up mixer as schematically illustrated in Figure 7.
ES 2 378 734 T3
In the above embodiments, the object parameters were provided as energy parameters. However, when the object parameters are provided as prediction parameters, that is, as an object prediction matrix C as indicated by element 124a in FIG. 12, the calculation of the reduced prediction matrix C3 is only one matrix multiplication as illustrated at block 125a and discussed in connection with equation (32). Matrix A3 as used in block 125a is the same matrix A3 that was mentioned in block 122 of Figure 12.
When the object prediction matrix C is generated by an audio object encoder and transmitted to the decoder, then some additional calculations are required to generate the parameters for boxes 74a, 74b, 74c. These additional steps are indicated in Figure 13b. Again, the object prediction matrix C is provided as indicated by 124a in FIG. 13b, which is the same as discussed in connection with block 124a of FIG. 12. Then, as discussed in connection with equation (31), the covariance matrix of the Z object downmix is calculated using the transmitted downmix or generated and transmitted as additional secondary information. When the information is transmitted in the Z matrix, then the decoder does not necessarily have to perform any power calculations that inherently introduce some delayed processing and increase the processing load on the decoder side. However, when these topics are not decisive for a certain application, then transmission bandwidth can be saved and the covariance matrix Z of the downmix of objects can also be calculated using the downmix samples that are naturally available in the decoder side. As soon as step 134 is completed and the covariance matrix of the descending mixture of objects is ready, the object energy matrix E can be calculated as indicated by step 135 using the prediction matrix C and the covariance matrix down-mix or “down-mix energy” Z. As soon as step 135 is completed, all the steps discussed in connection with figure 13a, such as steps 132, 133, can be performed to generate all parameters for blocks 74a, 74b, 74c of figure 7.
Figure 16 illustrates a further embodiment, in which only stereo reproduction is required. Stereo playback is the output as provided by mode number 5 or line 115 in Figure 11. In this case, the output data synthesizer 100 of FIG. 10 is not interesting in any spatial upmix parameter but is interesting mainly in a specific conversion matrix G to convert the object downmix into a useful stereo downmix and naturally easily influenced and easily controlled.
In step 160 of Figure 16, a partial downmix matrix from M to 2 is calculated. In the case of six output channels, the partial downmix matrix would be a six to two channel downmix matrix, but other downmix matrices are available as well. The calculation of this partial downmix matrix can be derived, for example, from the partial downmix matrix D36 as generated in step 121 and the matrix Dttt as used in step 127 of FIG. 12.
Furthermore, a stereo reproduction matrix A2 is generated using the result of step 160 and the "large" reproduction matrix A is illustrated in step 161. The reproduction matrix A is the same matrix that has been discussed in connection with the block 120 in figure 12.
Subsequently, in step 162, the stereo reproduction matrix can be parameterized by placement parameters µ and κ. When μ is set to 1 and κ is set to 1 as well, then equation (33) is obtained, which allows a variation of the voice volume in the example described in connection with equation (33). However, when other parameters such as μ and κ are used, then the placement of the sources can be varied as well.
Then, as indicated in step 163, the conversion matrix G is calculated using equation (33). In particular, the matrix (DED *) can be calculated, inverted, and the inverted matrix can be multiplied on the right hand side of the equation in block 163. Of course, other methods can be applied to solve the equation in block 163. Then, we have the conversion matrix G, and the downmix of objects X can be converted by multiplying the conversion matrix and the downmix of objects as indicated in block 164. Then, the converted downmix X 'can be reproduced in stereo using two stereo speakers. Depending on the implementation, certain values for μ, v, and κ can be adjusted to calculate the conversion matrix G. Alternatively, the conversion matrix G can be calculated using these three parameters as variables so that the parameters can be adjusted after step 163 as required by the user.
Preferred embodiments solve the problem of transmitting a number of individual audio objects (using multichannel downmix and additional control data describing the objects) and reproducing the objects to a given reproduction system (speaker configuration). A technique of how to modify the object-related control data into control data that is compatible with the reproduction system is introduced. It also proposes suitable encoding methods based on the MPEG Surround encoding scheme.
Depending on certain implementation requirements of the methods of the invention, the methods and signals of the invention can be implemented in hardware or in software. The implementation can be done using a digital storage medium, in particular a disk or a CD having electronically readable control signals stored therein, which can cooperate with a programmable computer system so that the
ES 2 378 734 T3 methods of the invention. Generally, the present invention is therefore a computer program product with a program code stored on a machine-readable medium, the program code being configured to perform at least one of the methods of the invention, when the program product computing runs on a computer. In other words, the methods of the invention are therefore a computer program that has a program code to perform the methods of the invention, when the computer program is run on a computer.
In other words, according to an embodiment of the present case, an audio object encoder for generating an encoded audio object signal using a plurality of audio objects, comprises a downmix information generator for generating downmix information indicating a distribution of the plurality of audio objects on at least two downmix channels; an object parameter generator for generating object parameters for the audio objects; and an output interface for generating the encoded audio object signal using the downmix information and the object parameters.
Optionally, the output interface can be operated to generate the encoded audio signal further using the plurality of downmix channels.
Additionally or alternatively, the parameter generator may be operative to generate the object parameters with a first time and frequency resolution, and wherein the downmix information generator is operative to generate the downmix information with a second time and frequency resolution, the second time and frequency resolution being smaller than the first time and frequency resolution.
Furthermore, the downmix information generator may be operative to generate the downmix information so that the downmix information is the same for the entire frequency band of the audio objects.
Furthermore, the downmix information generator may be operative to generate the downmix information such that the downmix information represents a downmix matrix defined as follows:
<img file="ES2378734T3_D0034.tif" />
where D is the downmix matrix, and where X is a matrix and represents the plurality of downmix channels and has a number of lines that is equal to the number of downmix channels.
Also, the information about a part can be a factor less than 1 and greater than 0.
Furthermore, the down-mixer may be operative to include the stereo representation of background music in the at least two down-mix channels, and to input a voice track into the at least two down-mix channels in a predefined ratio.
In addition, the downmixer may be operative to perform summation by samples of signals to be input to a downmix channel as indicated by the downmix information.
Furthermore, the output interface may be operative to perform data compression of the downmix information and object parameters before generating the encoded audio object signal.
Furthermore, the plurality of audio objects may include a stereo object represented by two audio objects that have a certain non-zero correlation, and in which the downmix information generator generates grouping information indicating the two audio objects. audio that make up the stereo object.
Furthermore, the object parameter generator may be operative to generate object prediction parameters for the audio objects, the prediction parameters being calculated such that the weighted sum of the downmix channels for a source object controlled by the parameters of the prediction or source object results in an approximation of the source object.
Furthermore, the prediction parameters can be generated per frequency band, and in which the audio objects cover a plurality of frequency bands.
Also, the number of audio objects can be equal to N, the number of downmix channels is equal to K, and the number of object prediction parameters calculated by the object parameter generator is equal to or less than NK.
Furthermore, the object parameter generator may be operative to calculate at most U (NK) object prediction parameters.
In addition, the object parameter generator may include an up-mixer to up-mix the plurality of down-mix channels using different sets of prediction parameters.
ES 2 378 734 T3 of test object; and wherein the audio object encoder further comprises an iteration controller for finding the test object prediction parameters that result in the smallest deviation between a source signal reconstructed by the up-mixer and the corresponding original source signal. between the different sets of test object prediction parameters.
Furthermore, the output data synthesizer may be operative to determine the conversion matrix using the downmix information, wherein the conversion matrix is calculated so that at least parts of the downmix channels are swapped when an object Audio included in a first downmix channel representing the first half of a stereo plane is to be played back in the second half of the stereo plane.
Furthermore, the audio synthesizer may comprise a channel player for reproducing audio output channels for the predefined audio output configuration using the spatial parameters and the at least two downmix channels or the converted downmix channels.
Furthermore, the output data synthesizer may be operative to output the output channels of the predefined audio output configuration further using the at least two downmix channels.
Furthermore, the output data synthesizer can be operative to calculate actual downmix weights for the partial downmix matrix such that an energy of a weighted sum of two channels equals the energies of the channels within a limiting factor. .
Additionally, the downmix weights for the partial downmix matrix can be determined as follows:
<img file="ES2378734T3_D0035.tif" />
where w<sub>p</sub> is a down-mix weight, p is an integer index variable, f, ¡is a matrix element of an energy matrix that represents an approximation of a covariance matrix of the output channels of the predefined output configuration.
Furthermore, the output data synthesizer may be operative to calculate separate coefficients from the prediction matrix by solving a system of linear equations.
Furthermore, the output data synthesizer can be operative to solve the system of linear equations based on:
C<sub>3</sub>(DED *) = AjED ', where C3 is the two-to-three prediction matrix, D is the downmix matrix derived from the downmix information, E is an energy matrix derived from the audio source objects, already<sub>3</sub> is the reduced downmix matrix, and where "*" indicates the complex conjugate operation.
Furthermore, the prediction parameters for the two to three upmix can be derived from a parameterization of the prediction matrix such that the prediction matrix is defined using only two parameters, and the output data synthesizer being operational for preprocessing. the at least two downmix channels so that the effect of preprocessing and the parameterized prediction matrix corresponds to a desired upmix matrix.
Also, the parameterization of the prediction matrix can be serialized as follows:
<img file="ES2378734T3_D0036.tif" />
α + 2 β- \ a-1 β + 2
1-to 1-β where the TTT index is the parameterized prediction matrix, and where α, β and γ are factors.
In addition, a downmix conversion matrix Gtal can be calculated as follows:
G - Γ> τττ ^ 3,
ES 2 378 734 T3 where C<sub>3</sub> is a two-by-three prediction matrix, where Dttt and Cttt equals 1, where I is a two-by-two identity matrix, and where C<sub>T</sub>tt is based on:
<img file="ES2378734T3_D0037.tif" />
a + 2 β- \ a- \ β + 2 \ -a \ -β where α, β and γ are constant factors.
Also, the prediction parameters for the two to three upmix can be determined as a and β, where γ is set to 1.
In addition, the output data synthesizer can be operative to calculate the energy parameters for the three to six upmix using an F energy matrix based on:
YY * «F = AEA *, where A is the reproduction matrix, E is the energy matrix derived from the audio source objects, Y is an output channel matrix, and“ * ”indicates the complex conjugate operation.
Furthermore, the output data synthesizer can be operative to calculate the energy parameters by combining elements of the energy matrix.
Furthermore, the output data synthesizer can be operative to calculate the energy parameters based on the following equations:
C £ D<sub>0</sub> = 101og<sub>I0</sub>^,
CLD<sub>X</sub> = lOlog ,,
<img file="ES2378734T3_D0038.tif" />
CLDi = 101og<sub>10</sub>
<img file="ES2378734T3_D0039.tif" />
/DC,
<img file="ES2378734T3_D0040.tif" />
<img file="ES2378734T3_D0041.tif" />
where φ is an absolute value operator φ (ζ) = | ζ | or real value rp (z) = Re {z}, where CLD<sub>0</sub> is a first channel level difference energy parameter, where CLDi is a second channel level difference energy parameter, where CLD2 is a third channel level difference energy parameter, where ICC1 is a first channel level difference parameter coherence energy between channels, and ICC2 is a second parameter of coherence energy between channels, and where fj are elements of an energy matrix F at positions ij in this matrix.
Furthermore, the first group of parameters may include energy parameters, and the output data synthesizer being operative to derive the energy parameters by combining elements of the energy matrix F.
Additionally, energy parameters can be derived based on:
ES 2 378 734 T3 cld ^
<img file="ES2378734T3_D0042.tif" />
fi \<sup>+</sup>fn + fv<sup>+</sup>fu fss + fu.
C ^ -Oo ^ -. O.og ,, ^) where CLD ° ttt is a first energy parameter of the first group and where CLD<sup>1</sup>ttt is a second energy parameter of the first group of parameters.
Furthermore, the output data synthesizer may be operative to calculate weight factors for weighting the downmix channels, the weight factors being used to control arbitrary downmix gain factors of the spatial decoder.
In addition, the output data synthesizer can be operative to calculate the weight factors based on:
Z = DED *,
W = D<sub>26</sub>ED '<sub>36</sub>,
<img file="ES2378734T3_D0043.tif" />
<img file="ES2378734T3_D0044.tif" />
where D is the downmix matrix, E is an energy matrix derived from the audio source objects, where W is an intermediate matrix, where D<sub>2</sub>e is the partial downmix matrix for 6 to 2 channel downmix of the default output configuration, and where G is the conversion matrix that includes the arbitrary downmix gain factors of the spatial decoder.
Furthermore, the output data synthesizer can be operative to calculate the energy matrix based on:
E = CZC ', where E is the energy matrix, C is the prediction parameter matrix, and Z is a covariance matrix of the at least two downmix channels.
Also, the output data synthesizer can be operative to calculate the conversion matrix based on:
G = AiC, where G is the conversion matrix, A<sub>2</sub> is the partial reproduction matrix, and C is the prediction parameter matrix.
Also, the output data synthesizer can be operative to calculate the conversion matrix based on:
G (DEDyA<sub>2</sub>ED ', where G is an energy matrix derived from the tracks' audio source, D is a downmix matrix derived from the downmix information, A<sub>2</sub> is a reduced reproduction matrix, and "*" indicates the complete conjugate operation.
ES 2 378 734 T3
Furthermore, the parameterized stereo reproduction matrix A<sub>2</sub> can be determined as follows:
TO
1-Λ-
<img file="ES2378734T3_D0045.tif" />
where μ, v, and κ are real value parameters to be adjusted based on the position and volume of one or more source audio objects.
Contents21
62 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62
57 members in 22 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 829649P | United States of America | – | |
| 82964906 | United States of America | P | |
| 82964906 | United States of America | P | |
| 829649P | – | – | – |
| US20060829649P | – | – | – |
Members57
| Document | Office | Kind | |
|---|---|---|---|
| AU2007312598A1 | Australia | A1 | |
| CA2666640A1 | Canada | A1 | |
| CA2874451A1 | Canada | A1 | |
| CA2874454A1 | Canada | A1 | |
| WO2008046531A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200828269A | Taiwan Province of China | A | |
| EP2054875A1 | European Patent Office (EPO) | A1 | |
| NO20091901L | Norway | L | |
| MX2009003570A | Mexico | A | |
| KR20090057131A | Republic of Korea | A | |
| EP2068307A1 | European Patent Office (EPO) | A1 | |
| CN101529501A | China | A | |
| HK1126888A1 | Hong Kong, China | A1 | |
| JP2010507115A | Japan | A | |
| HK1133116A1 | Hong Kong, China | A1 | |
| RU2009113055A | Russian Federation | A | |
| KR20110002504A | Republic of Korea | A | |
| AU2007312598B2 | Australia | B2 | |
| US2011022402A1 | United States of America | A1 | |
| KR101012259B1 | Republic of Korea | B1 | |
| EP2054875B1 | European Patent Office (EPO) | B1 | |
| AU2011201106A1 | Australia | A1 | |
| UA94117C2 | Ukraine | C2 | |
| ATE503245T1 | Austria | T1 | |
| DE602007013415D1 | Germany | D1 | |
| TWI347590B | Taiwan Province of China | B | |
| RU2430430C2 | Russian Federation | C2 | |
| EP2372701A1 | European Patent Office (EPO) | A1 | |
| SG175632A1 | Singapore | A1 | |
| EP2068307B1 | European Patent Office (EPO) | B1 | |
| ATE536612T1 | Austria | T1 | |
| KR101103987B1 | Republic of Korea | B1 | |
| MY145497A | Malaysia | A | |
| ES2378734T3This record | Spain | T3 | |
| AU2011201106B2 | Australia | B2 | |
| JP2012141633A | Japan | A | |
| RU2011102416A | Russian Federation | A | |
| PL2068307T3 | Poland | T3 | |
| HK1162736A1 | Hong Kong, China | A1 | |
| CN102892070A | China | A | |
| BRPI0715559A2 | Brazil | A2 | |
| CN101529501B | China | B | |
| JP5270557B2 | Japan | B2 | |
| JP5297544B2 | Japan | B2 | |
| JP2013190810A | Japan | A | |
| CN103400583A | China | A | |
| EP2372701B1 | European Patent Office (EPO) | B1 | |
| PT2372701E | Portugal | E | |
| JP5592974B2 | Japan | B2 | |
| CA2666640C | Canada | C | |
| CN103400583B | China | B | |
| CN102892070B | China | B | |
| CA2874451C | Canada | C | |
| US9565509B2 | United States of America | B2 | |
| US2017084285A1 | United States of America | A1 | |
| NO340450B1 | Norway | B1 | |
| CA2874454C | Canada | C |
Numbers
- Publication
- 2378734
- Publication, DOCDB
- 2378734
- Publication, EPODOC
- ES2378734T
- Application
- 9004406
- Application, DOCDB
- 09004406
- Application, EPODOC
- ES20090004406T
Titles2
- Spanish
- Codificación mejorada y representación de parámetros de codificación de objetos de mezcla descendente multicanal
- English
- Enhanced coding and representation of coding parameters of multichannel downlink objects
Classification
- CPC, 10
- G10L19/20
- H04S7/30
- H04S2420/03
- G10L19/008
- G10L19/173
- H04S3/008
- H04S3/02
- H04S2400/03
- H04S5/00
- H04S2400/11
- IPC, 2
- G10L19 00
- H04S7 00