Untitled record
Abstract
FIELD: information technology. SUBSTANCE: audio object encoder, designed to generate encoded object signals using a plurality of audio objects, which includes a downmixing data generator which generates downmixing parameters, having indications for the order of distribution of the plurality of audio objects on at least two downmixing channels, an audio object parameter generator which generates audio object parameters, and an output interface which generates an imported output audio signal using downmixing data and object parameters. An audio synthesiser, which uses downmixing data to generate output data, used to form a plurality of output channels for reproducing an audio signal of a given configuration. EFFECT: facilitating upmixing on all downmixing channels. 13 cl, 18 dwg
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
13 claims: 5 independent, 8 dependent
- 1Audiosintezator (104) for generating output data using an encoded audio object signal (95, 97), characterized in that it includes an output data synthesizer (100) generating the output parameters applicable for representing a plurality of output channels with a pre- predetermined configuration of audio output, displaying a plurality of audio objects, wherein the synthesizer output provides the ability to use information downmix containing instructions for allocating a plurality of audio objects in at least two downmix channels and object parameter for the audio objects, the synthesizer output (100 ) recoding (502), audio object parameters into spatial parameters for the predefined audio output configuration additionally using the given location (A) of audio objects (90) in the audio output configuration. 1. Аудиосинтезатор (104), предназначенный для генерирования выходных данных с использованием закодированного сигнала аудиообъекта (95, 97), характеризующийся тем, что включает в себя синтезатор выходных данных (100), генерирующий на выходе параметры, применимые для представления множества выходных каналов с предварительно заданной конфигурацией выходного аудиосигнала, отображающего множество аудиообъектов, при этом синтезатор выходных данных предусматривает возможность использования информации понижающего микширования, содержащей указания на распределение множества аудиообъектов, по крайней мере, по двум каналам понижающего микширования, и параметр объекта для аудиообъектов, причем синтезатор выходных данных (100) перекодирует (502) параметры аудиообъекта в пространственные параметры для предварительно заданной конфигурации выходного аудиосигнала, дополнительно используя заданное расположение (А) аудиообъектов (90) в конфигурации выходного аудиосигнала. 1. Аудиосинтезатор (104), предназначенный для генерирования выходных данных с использованием закодированного сигнала аудиообъекта (95, 97), характеризующийся тем, что включает в себя синтезатор выходных данных (100), генерирующий на выходе параметры, применимые для представления множества выходных каналов с предварительно заданной конфигурацией выходного аудиосигнала, отображающего множество аудиообъектов, при этом синтезатор выходных данных предусматривает возможность использования информации понижающего микширования, содержащей указания на распределение множества аудиообъектов, по крайней мере, по двум каналам понижающего микширования, и параметр объекта для аудиообъектов, причем синтезатор выходных данных (100) перекодирует (502) параметры аудиообъекта в пространственные параметры для предварительно заданной конфигурации выходного аудиосигнала, дополнительно используя заданное расположение (А) аудиообъектов (90) в конфигурации выходного аудиосигнала.
- 6A method of synthesizing sound comprising generating output data using an encoded audio object signal (95, 97), characterized in that it comprises generation of output data to generate a plurality of output channels with a predetermined output configuration of the audio signal, displaying a plurality of audio objects (90), using the Information downmix indicating the order of allocating a plurality of audio objects in at least two downmix channels and the parameters of the audio object for the audio objects, and the parameters audio object re-encoding (502) a spatial parameters calculated configuration with additional account data specified location (A) of audio objects (90) The audio output configuration. 6. Способ синтезирования звука, предусматривающий генерирование выходных данных с использованием закодированного сигнала аудиообъекта (95, 97), характеризующийся тем, что включает генерирование выходных данных для формирования множества выходных каналов с заданной конфигурацией выходного аудиосигнала, отображающей множество аудиообъектов (90), при этом используется информация понижающего микширования, указывающая порядок распределения множества аудиообъектов, по крайней мере, по двум каналам понижающего микширования и параметры аудиообъекта для аудиообъектов, причем параметры аудиообъекта перекодируют (502) в пространственные параметры расчетной конфигурации с дополнительным учетом данных заданного расположения (А) аудиообъектов (90) в конфигурации выходного аудиосигнала. 6. Способ синтезирования звука, предусматривающий генерирование выходных данных с использованием закодированного сигнала аудиообъекта (95, 97), характеризующийся тем, что включает генерирование выходных данных для формирования множества выходных каналов с заданной конфигурацией выходного аудиосигнала, отображающей множество аудиообъектов (90), при этом используется информация понижающего микширования, указывающая порядок распределения множества аудиообъектов, по крайней мере, по двум каналам понижающего микширования и параметры аудиообъекта для аудиообъектов, причем параметры аудиообъекта перекодируют (502) в пространственные параметры расчетной конфигурации с дополнительным учетом данных заданного расположения (А) аудиообъектов (90) в конфигурации выходного аудиосигнала.
- 7The encoder of audio objects (101) for generating a plurality of audio objects coded signals of audio objects (90), characterized in that it comprises a downmix information generator (96) for generating downmix information (97), reflecting the order of allocating a plurality of audio objects, at least between the two downmix channels;wherein the downmix information generator (96) configured to generate (150) energy characteristics (XX *) and the correlation data (SX *), reflecting power characteristics and correlation characteristics of the at least two downmix channels (93);audio object parameter generator (94);and an output interface (98) for outputting the generated signal encoded audio object (99), the encoded audio object signal comprising a downmix information, power information, information on the correlation and the object parameters. 7. Кодер аудиообъектов (101), предназначенный для генерирования закодированных сигналов аудиообъектов множества аудиообъектов (90), характеризующийся тем, что он включает в себя генератор информации понижающего микширования (96) для вырабатывания информации понижающего микширования (97), отражающей порядок распределения множества аудиообъектов, по меньшей мере, между двумя каналами понижающего микширования;причем генератор информации понижающего микширования (96) сконфигурирован с возможностью генерирования (150) энергетических характеристик (XX*) и данных корреляции (SX*), отражающих мощностные характеристики и корреляционные характеристики этих, по меньшей мере, двух каналов понижающего микширования (93);генератор параметров аудиообъекта (94);и выходной интерфейс (98), предназначенный для вывода сгенерированного закодированного сигнала аудиообъекта (99), при этом закодированный сигнал аудиообъекта содержит информацию понижающего микширования, информацию о мощности, информацию о корреляции и параметры объекта. 7. Кодер аудиообъектов (101), предназначенный для генерирования закодированных сигналов аудиообъектов множества аудиообъектов (90), характеризующийся тем, что он включает в себя генератор информации понижающего микширования (96) для вырабатывания информации понижающего микширования (97), отражающей порядок распределения множества аудиообъектов, по меньшей мере, между двумя каналами понижающего микширования;причем генератор информации понижающего микширования (96) сконфигурирован с возможностью генерирования (150) энергетических характеристик (XX*) и данных корреляции (SX*), отражающих мощностные характеристики и корреляционные характеристики этих, по меньшей мере, двух каналов понижающего микширования (93);генератор параметров аудиообъекта (94);и выходной интерфейс (98), предназначенный для вывода сгенерированного закодированного сигнала аудиообъекта (99), при этом закодированный сигнал аудиообъекта содержит информацию понижающего микширования, информацию о мощности, информацию о корреляции и параметры объекта.
- 10A method of encoding audio objects (101) to form an encoded signal a plurality of audio objects, characterized in that it includes generating downmix information (97) containing instructions for allocating a plurality of audio objects (90), the at least two downmix channels, production ( 150) power indicators (XX *) and the correlation data (SX *), reflecting power characteristics and correlation characteristics of the at least two downmix channels;formulation parameters of audio objects (94);and issue an encoded audio object signal (99), the encoded audio object signal comprises power information about the correlation of the downmix information and the object parameters. 10. Способ кодирования аудиообъектов (101) с формированием закодированного сигнала множества аудиообъектов, характеризующийся тем, что включает генерирование информации понижающего микширования (97), содержащей указания по распределению множества аудиообъектов (90), по меньшей мере, по двум каналам понижающего микширования, выработку (150) энергетических показателей (XX*) и данных корреляции (SX*), отражающих мощностные характеристики и корреляционные характеристики этих, по меньшей мере, двух каналов понижающего микширования;выработке параметров аудиообъектов (94);и выдаче закодированного сигнала аудиообъекта (99), при этом закодированный сигнал аудиообъекта содержит, информацию о мощности, информацию о корреляции, информацию понижающего микширования и параметры объекта. 10. Способ кодирования аудиообъектов (101) с формированием закодированного сигнала множества аудиообъектов, характеризующийся тем, что включает генерирование информации понижающего микширования (97), содержащей указания по распределению множества аудиообъектов (90), по меньшей мере, по двум каналам понижающего микширования, выработку (150) энергетических показателей (XX*) и данных корреляции (SX*), отражающих мощностные характеристики и корреляционные характеристики этих, по меньшей мере, двух каналов понижающего микширования;выработке параметров аудиообъектов (94);и выдаче закодированного сигнала аудиообъекта (99), при этом закодированный сигнал аудиообъекта содержит, информацию о мощности, информацию о корреляции, информацию понижающего микширования и параметры объекта.
- 11The computer-readable data carrier with stored thereon encoded audio object signal, characterized in that it comprises a downmix information defining the order of distribution of the plurality of audio objects in at least two downmix channels, power indicators (XX *) and the correlation data (SX *) reflecting power characteristics and correlation characteristics of the at least two downmix channels and object parameters, allowing in combination with at least two downmix channels to reconstruct audio objects. 11. Считываемый компьютером носитель данных с сохраненным на нем закодированным сигналом аудиообъекта, характеризующийся тем, что содержит информацию понижающего микширования, определяющую порядок распределения множества аудиообъектов, по меньшей мере, по двум каналам понижающего микширования, энергетические показатели (XX*) и данные корреляции (SX*), отражающие мощностные характеристики и корреляционные характеристики этих, по меньшей мере, двух каналов понижающего микширования, и параметры объектов, позволяющие в сочетании с, по крайней мере, этими двумя каналами низведения воссоздать аудиообъекты. 11. Считываемый компьютером носитель данных с сохраненным на нем закодированным сигналом аудиообъекта, характеризующийся тем, что содержит информацию понижающего микширования, определяющую порядок распределения множества аудиообъектов, по меньшей мере, по двум каналам понижающего микширования, энергетические показатели (XX*) и данные корреляции (SX*), отражающие мощностные характеристики и корреляционные характеристики этих, по меньшей мере, двух каналов понижающего микширования, и параметры объектов, позволяющие в сочетании с, по крайней мере, этими двумя каналами низведения воссоздать аудиообъекты.
Independent claims5
318 paragraphs, as filed
The invention relates to decoding of multiple objects by converting the encoded signal based on the multi-site available multichannel downmix and auxiliary control data.
Recent developments in audio technology makes it possible to recreate a multi-channel audio signal based on a stereo (or mono) signal and corresponding control data. These methods of parametric audio coding environment typically include parameterization. Parametric multi-channel audio decoder (eg, MPEG Surround standard ISO / TEC 23003-1, L.Villemoes, J.Herre, J.Breebaart, G.Hotho, S.Disch, H.Pumhagen, and K.Kjorling, "MPEG Surround: The Forthcoming ISO Standard for Spatial Audio Coding, "in 28th International AES Conference, The Future of Audio Technology Surround and Beyond, Pitea, Sweden, June 30-July 2, 2006; J.Breebaart, J.Herre, L.Villemoes, C. Jin ,, K.Kjorling, J.Plogsties, and J.Koppens, "Multi-Channels goes Mobile: MPEG Surround Binaural Rendering," in 29th International AES Conference, Audio for Mobile and Handheld Devices, Seoul, Sept 2-4,2006 ) reconstructs M channels based on K received channels, where M> K, with the control data. The control data represent a parametrization multichannel signal based on the signal intensity difference between the channels (IID) and inter-channel coherence coherence (ICC). Typically, these parameters are allocated to the encoding stage and describe power ratios and correlation between channel pairs used in the up-mix. The use of such coding algorithm enables coding at a data rate significantly lower than the transmission of the totality of M channels, while high encoding efficiency and simultaneously ensure compatibility with both K channel devices, and devices Channel M.
A similar coding system provides the appropriate audio object encoder [S.Faller, "Parametric Joint-Coding of Audio Sources," Convention Paper 6752 presented at the 120th AES Convention, Paris, France, May 20-23, 2006], [C.Faller, " Parametric Joint-Coding of Audio Sources, "Patent application PCT / EP2006 / 050904, 2006], where a number of audio objects mixed" down "by the encoder, and later mixed" up "using the control commands. The process of upmixing can be also seen as a separation of the objects mixed in the downmix. The resulting up-mix signal can be converted to be displayed in a single or multi-line appearance. Determine exactly the above-mentioned publications are a method of synthesis of audio channels on the basis of the downmix (called sum signal), the statistical information on the sources and the characteristics that define the desired output format. If multiple signals received downmix, these signals consist of different subsets of the objects, and the upmixing should be for each downmix channel individually. The novelty of the proposed method lies in the implementation of the upmix simultaneously for all the downmix channels. Methods of coding object submitted prior to the present invention, do not offer the option of decoding results downmix of multiple channels simultaneously.
A first aspect of the invention relates to an audio object coder that generates an encoded audio object signal using a plurality of audio objects including:
a generator of information (data) downmix generating distribution parameters plurality of audio objects in at least two downmix channels;
generator parameters of audio objects and output interface for generating the encoded audio object signal using the downmix signal and the characteristics of the object parameters.
A second aspect of the invention relates to a method of coding an audio object. for producing a coded audio object signal using a plurality of audio objects including:
generating down-mix data characterizing the distribution order of audio objects together, at least two downmix channels;
generating parameters of audio objects, and generating encoded audio object signal using the downmix information and the object parameters.
A third aspect of the invention relates to a sound synthesizer (audiosintezatoru) generating output data using an encoded audio object signal including:
synthesizer output data used to represent the number of output channels with a predetermined output configuration of the audio signal, displays the collection of audio objects, the output data synthesizer which detects characteristics downmix for allocating a plurality of audio objects in at least two downmix channels, and audio object parameters.
A fourth aspect of the invention relates to a sound synthesizing method capable of generating output data using an encoded audio object signal including:
generation of output data to generate a plurality of output channels with a predetermined output configuration of the audio signal shows the aggregate audio objects using a synthesizer output data capable of reading characteristics downmix for allocating a plurality of audio objects in at least two downmix channels, and audio object parameters.
A fifth aspect of the invention relates to the coded signal of an audio object containing characteristics downmix indicating the order of allocating a plurality of audio objects in at least two downmix channels and object parameters, allowing to reconstruct the audio objects using the object parameters and at least two channel downmix mix.
A sixth aspect of the invention relates to a computer software designed for audio object coding method or the method of decoding an audio object on the computer.
The invention is presented for illustrative purposes, not limiting thereof in either form or substance, with explanations of the accompanying drawings, wherein:
1 is a block diagram of the spatial audio object coding, including encoding and decoding;
1b is a block diagram of a spatial audio object coding using MPEG Surround decoder;
2 shows the algorithm of spatial audio object encoder;
3 is a diagram of the operation of the extractor (extractor) in the audio object parameters differentiation power mode;
4 is a diagram of the operation of the extractor (extractor) audio object parameters in the prediction mode;
5 is a diagram of an SAOC transcoder - MPEG Surround;
6 presents schematically different operation modes down mixer for down-mixing;
7 is a schematic diagram of MPEG Surround decoder for a stereo mix downlink;
8 is a diagram of a particular case of realization using the SAOC encoder;
9 is a diagram of an embodiment of the encoder;
10 is a diagram of an embodiment of a decoder;
11 is a table of the best modes of decoder / synthesizer;
Figure 12 is a block diagram of a method for calculating certain spatial upmix parameters;
13 is a block diagram of a method for calculating additional spatial upmix parameters;
at 13b is a block diagram calculation methods using prediction parameters;
Figure 14 gives a general schematic diagram of a coder / decoder;
15 is a flowchart of calculating prediction object parameters; and
Figure 16 illustrates a method for the stereo representation (rendering).
The following embodiments are no more than illustration of the principles of an improved method of coding and parametric coding of multi-channel representations of the object downmix. It is understood that for those skilled in the ability to make changes and improvements in the layout and design of the elements described is obvious. In view of this explanation and description presented embodiment, limited only by the scope of patent claims, but not by the specific details.
Preferred embodiments provide a method of coding, which combines the functionality of the encoding algorithm object with audio presentation (audiorenderinga) multi-channel decoder. Forwarded control data related to individual objects and therefore allow you to control the playback position and spatial signal level. Thus, the control information is directly related to the so-called 'scene description' giving information on the location of objects in the environment. The scene description can be controlled by the decoder or interactively by the listener or by the encoder from the sound source.
The essence of the invention consists in that the transcoder is introduced to convert (transcode) relating to the object control information and a downmix signal into control data and downmix signal intended for the playback system, e.g., MPEG Surround decoder. In the present method of encoding objects may be randomly distributed on available channels downward mixing encoder. Transcoder uses exactly the multichannel mix parameters downward, providing a transcoded downmix signal and object-related control data. This increases the mixing is not performed at the decoder for each channel individually, as suggested in [S.Faller, "Parametric Joint-Coding of Audio Sources," Convention Paper 6752 presented at the 120th AES Convention, Paris, France, May 20-23, 2006], but all downmix channels are processed simultaneously in a single upmixing process. Under the new scheme the parameters of the multi-channel down-mix to be part of the control data and encoded by the encoder object.
Distributing objects to downmix channels can be executed automatically, or it can be a constructive solution, connected to the encoder. In the latter case the system down (downward) mix can be incorporated into an existing multi-channel playback system (such as a stereo), with an emphasis on reproduction, omitting the step of transcoding and multi-channel decoding. This is another advantage over earlier coding algorithm known from prior art provides one downmix channel or multiple downmix channels containing subsets of source objects.
While object coding schemes of prior art technology describe decoding using only a single downmix channel, the present invention has no such limitation, since it offers a method for simultaneously decoding a downmix material comprising downmix signals on several channels. The quality of the separation of objects increases with the number of downmix channels. Thus the invention successfully bridges the gap between an object coding algorithm on a single multidrop downmix signal and multi-channel coding algorithm, where each object is transmitted on the dedicated channel. Thus, the proposed method allows flexible management of quality in the separation of objects, depending on the requirements, application requirements and operational properties of the transmission system (such as channel capacity).
In addition, the advantage of using more than one channel is that it allows to take into account the correlation between the different objects, in contrast to the description only takes into account the difference in the intensity of audio signals in the object coding algorithms in earlier practice. Earlier practice was based on the assumption that all objects are independent of each other and are mutually aligned (zero cross-correlation), while in reality it is unlikely that objects can not be correlated, such as the left and right channels of a stereo signal. In accordance with the concept of the present invention including a description of the correlation parameters (control data) makes it more complete and thus facilitates creating additional opportunities object separation. Preferred embodiments include at least one of the following features.
A system for transmitting and creating a plurality of individual audio objects using a multi-channel downmix and auxiliary control data describing the objects comprising:
coder spatial audio objects, encoding the plurality of audio objects for multichannel downmix, information about the multichannel downmix, and object parameters; or a decoder of spatial audio objects, data decoding multichannel downmix, information about the multichannel downmix, object parameters and audiorenderinga matrix (a matrix representation) of the object in the second channel audio signal, applicable to audio playback.
1a shows a spatial audio object coding algorithm (SAOC), comprising an SAOC encoder 101 and the SAOC decoder 104. The encoder 101 encodes the spatial audio objects N objects into an object downmix data of K> 1, audio channels according to the parameters of the encoder. Information on the application of the downmix weight matrix D Pin SAOC encoder, together with supporting data regarding the power and correlation of the downmix. The matrix D is often, but not necessarily always, constant over time and frequency, and therefore contain relatively little information. Finally, the SAOC encoder captures the parameters of each object as a function of time and frequency with depth resolution is determined on the basis of perception (perceptual coding). Spatial Decoder 104 accepts input of audio objects in a data object downmix channels, the downmix information and the object parameters (as generated by the encoder) and generates output data comprising M audio channels for presentation to the user. Audiorendering N objects into M audio channels is performed by a matrix audiorenderinga, which is a set of parameters entered by the user into the decoder SAOC.
Figure 1b shows a block diagram of the spatial audio object coding, followed using the decoder MPEG Surround. SAOC decoder 104 used in the present invention may be implemented in SAOC transcoder - MPEG Surround 102 in combination with MPEG Surround decoder 103 to the stereo downmix. Manage users audiorenderinga matrix A of dimension M × N defines a predetermined conversion ratio of N objects into M audio channels. Functions of the matrix may depend on the settings and performance of the frequency, and the final result is the most user-friendly interface for audio object management (which, moreover, can be introduced from the outside scene description). In the case of the settings for the speaker 5. 1, the number of output audio channels is M = 6. Task SAOC decoder is to perceptually recreate the original audio objects as the final result audiorenderinga. At the inlet SAOC transcoder - MPEG Surround 102 receives audiorenderinga matrix A, the object downmix information, the downmix results, including the downmix weight matrix D, and object description, and generates a stereo downmix and MPEG Surround information. If the transcoder is implemented according to the present invention, following it MPEG Surround decoder 103, receiving the data at the input, the output gives the M-channel acoustic signal with the desired characteristics.
SAOC decoder, introduced in the present invention consists of SAOC transcoder - MPEG Surround decoder 102 and MPEG Surround 103 downward to a stereo mix. Manage users audiorenderinga matrix A of dimension M × N defines a predetermined conversion ratio of N objects into M audio channels. This matrix can depend on both the settings and the frequency that is an indication of a more user-friendly management interface audio objects. When applying the settings for 5.1 speaker system is the number of output audio channels M = 6. SAOC decoder is designed to recreate the original perceptual audio objects as the final result audiorenderinga. At the inlet SAOC transcoder - MPEG Surround 102 receives audiorenderinga matrix A, the object downmix information, the downmix results, including the downmix weight matrix D, and object description, and generates a stereo downmix and MPEG Surround information. If the transcoder is implemented according to the present invention, following it MPEG Surround decoder 103, receiving the data at the input, the output gives the M-channel acoustic signal with the desired characteristics.
2 shows the algorithm of the encoder of spatial audio object (SAOC) 101 introduced by the present invention. N audio objects are introduced into a downmix unit 201, and the extractor (extractor) audio object parameters 202. The downmix unit 201 mixes the objects in the resulting data stream object downmix consisting of K> 1, the audio channels, according to encoder parameters and also outputs information on the downmix signal. This information includes a description of the applied downmix weight matrix D, and further, if the actuated sequentially audio object parameter extractor operates in prediction mode, parameters describing the power and correlation of the object downmix results.
As will be discussed in one of the following paragraphs, the role of such additional parameters is to provide access to energy and correlation indices subsets converted audio channels in cases where the object parameters are expressed only with respect to the down-mix and the main example here is the clock signals "rear / front" for 5.1 speaker systems. Extractor parameters of audio objects 202 allocates object parameters according to the encoder parameters. Tools encoder for time-frequency changes is determined by which of the two modes of the encoder used for energy or predictive basis. In the differentiation capacity encoder parameters further contains information on the grouping N audio objects into P stereo objects and N-2P monoobektov. Each mode will be described hereinafter in Figures 3 and 4.
3 is a chart of an audio object parameter extractor 202 in a mode of differentiation capacities. Grouping 301 P stereo objects and N-2P monoobektov performed according to grouping information contained in the encoder. For each given frequency-timeslot then performs the following operations. Two power exponent object and one normalized correlation are allocated extractor stereo parameters 302 for each of the P stereo objects. One indicator of energy parameter extractor 303 is allocated for each of the N-2P monoobektov. Then, the full set of N power parameters and P normalized correlation parameters are encoded in 304 together with the grouping data to form the object parameters. The encoding can include the step of normalizing given the highest power rating of the object or with the amount of power allocated object.
4 is a flowchart audio object parameter extractor 202 in the prediction mode. For each given frequency-timeslot then performs the following operations. For each of the N objects derived from a linear combination of K object downmix channels, which corresponds to a given object using the method of least squares. K weight of this linear combination are called the coefficients of the prediction object (ODP), and they are evaluated OPC extractor 401. A complete set of LFS in the number of NK 402 is encoded in the formation of object parameters. Encoding may include reducing the total number of OPC based on the linear interdependencies. The distinguishing feature of this invention is that the total number can be reduced to a maximum of {K · (NK), 0}, eating downmix weight matrix D has full rank.
5 is a diagram of an SAOC transcoder - MPEG Surround 102 according to the present invention. For each time frequency interval, information about the downmix and the object parameters are combined with the matrix audiorenderinga counter parameter 502 with the formation parameters of MPEG Surround type CLD (level difference channels), CPC (prediction coefficient of the channel) and ICC (inter-channel coherence) and G matrix converter downlink Mixing dimension 2 × K. Results downmix converter 501 converts the object downmix into a stereo downmix by matrix operation according to the matrix G. In a simplified mode of the transcoder for K-2, the matrix acts as an identity matrix and the object downmix is passed without change as a stereo downmix. In the diagram this is shown as a mode switch 503 in position A, whereas in normal operation the switch is in position B. An additional advantage of the transcoder - suitability for use as a stand-alone device where parameters are ignored MPEG Surround, and outputs the down-converter mixing used directly as stereoaudiorendering.
6 schematically presents different modes of operation of the converter 501 down-mix according to the present invention. Given the transmitted object downmix in the format of the output bitstream is a K-channel audio encoder, this Bitstream first decoded by the audio decoder 601 into K time domain audio signals. Then, these signals are transformed to the frequency domain filter bank hybrid QMF (quadrature mirror filter) MPEG Surround block T / F (Time / Frequency) 602. Working matrix varying time and frequency determined by the converter matrix data is performed on the resulting hybrid QMF domain signals matrixing unit 603 which outputs a stereo signal in the hybrid QMF domain. Hybrid synthesis unit 604 converts the stereo hybrid QMF domain in the stereo field QMF. The hybrid QMF domain is set to improve the frequency resolution in the lower frequencies by the subsequent sub-band filtering QMF. When performing a further filtration using filter banks Nyquist conversion from a standard hybrid QMF domain is a simple summation of groups of sub-band signals hybrid, see. [E.Schuijers, J.Breebart, and H.Purnhagen "Low complexity parametric stereo coding" Proc 116th AES convention Berlin. Germany in 2004, Preprint 6073]. This signal represents a first possible output format of the downmix converter that corresponds to the position and the switch 607. This QMF domain signal can be fed directly to the corresponding QMF domain interface decoder MPEG Surround, and this is the preferred mode of operation in terms of delay, complexity and quality . Another possibility is the formation of a stereo time domain using a QMF synthesis filter bank 605. At position B, the switch 607 supplies a digital stereo signal transducer which can also be introduced into the time domain interface of a subsequent MPEG Surround decoder or fed directly to a stereo playback. A third possibility when the C 607 is a switch coding domain stereo music via stereo audio encoder 606. In this case, the output format converter downmix to stereo audiobitstrim compatible with the central decoder is a component of MPEG-decoder. This third mode of operation is applicable when a transcoder SAOC - MPEG Surround decoder MPEG-blocked due compounds limiting the data rate, or when the user needs to save the image of a particular object for future playback.
7 is a schematic diagram of MPEG Surround decoder for a stereo downmix. Stereo downmix using the 'two-to-three "(TTT) is divided into three intermediate channels. Further, each of the intermediate channel with three windows, "one-to-two" (OTT) is divided into two to form a six-channel 5.1-channel configuration.
8 is a diagram of a particular case of realization using the SAOC encoder. Audio Mixer 802 outputs a stereo (left and right), which is usually created by mixing the signal at the input of the mixer (here - input channels 1-6) and arbitrary additional input from the electronic effects such as reverb, etc. Moreover, the mixer has a single individual output channel (here channel 5). This channel can be used, for example, the normal functions of the mixer, such as "direct access" or "additional shipment" to display the individual data without the involvement of any intermediate processes (such as dynamic processing and EQ). Stereo (left and right) and individual output channel (obj5) are input to the encoder 801 SAOC, which is only a special case SAOC encoder 101 in Figure 1. However, it serves as a typical example of the application when the audio object obj5 (containing, for example, speech) should be fully controlled by the user with the right adjustments at the input of the decoder, remaining, however, part of a mixed a stereo (with the left and right channels). From the concept it is also obvious that a panel "object input" ("input object") in frame 801 may be connected to two or more audio objects, and in addition to this, a stereo can be expanded to multi-channel connections such as 5.1-channel device.
The following is a brief mathematical description of the invention. For discrete complex signals x, y complex inner product and squared norm (energy) is determined by:
<img file="00000001.tif" he="18" wi="70" img-format="tif" img-content="undefined" />
where y (k) denotes the complex conjugate signal of y (k). All signals considered here are subband samples from a modulated filter bank or windowed FFT analysis (fast Fourier transform) of discrete time signals. It is understood that these subbands have to be transformed back to the discrete time domain by corresponding synthesis filter bank operations. Block signals from L samples represents the signal in a frequency-time domain, which is part of the perceptually motivated filling mosaic (tiling) the time-frequency plane are used to describe signal properties. When such a partition certain audio objects can be represented as N rows of length L in a matrix,
<img file="00000002.tif" he="25" wi="79" img-format="tif" img-content="undefined" />
Downmix weight matrix D dimension K × N,
where K> 1 determines the K-channel signal downlink mixing in the form of a matrix with K rows of the matrix multiplication
<img file="00000003.tif" he="5" wi="31" img-format="tif" img-content="undefined" />
Manage users audiorenderinga matrix of the object A dimension M × N determines the M channel audiorendering audio objects with the specified parameters in the form of a matrix with M rows of matrix multiplication
<img file="00000004.tif" he="5" wi="31" img-format="tif" img-content="undefined" />
If temporarily ignore the effects of the base stream audio coding task SAOC decoder is to generate the desired close to the perception of the result Y of the original audio objects audiorenderinga based audiorenderinga matrix A, the downmix X the results, the downmix matrix D, and object parameters.
Object parameters in energy mode according to the present invention carry information about the covariance of the original objects. The deterministic version convenient for getting consistent results and intuitive to describe the typical operations of the encoder, the covariance is presented in the form of non-normalized matrix product SS *, where the asterisk denotes the complex conjugate operation on the transposed matrix. Thus, the object parameters obtained in the power mode, providing a positive semi-definite matrix A of dimension N × N so that, perhaps up to a scaling factor,
<img file="00000005.tif" he="5" wi="33" img-format="tif" img-content="undefined" />
The prior art coding audioob projects often consider an object model, where all objects are not correlated. In this case, the matrix E is diagonal and contains only an approximation to the energy of the object Sn = || Sn || 2 with n = 1, 2, ..., N. 3, the extractor object parameters to make a significant adjustment to the idea, which is especially important in cases where the objects are represented as stereo signals for which the assumption of no correlation is not valid. Grouping P selected stereo pairs of objects described by a set of indices {(np, mp), p = 1, 2, P}. For these stereo correlation <Sn, Sm> calculated, and integrated, real or absolute value of the normalized correlation (ICC)
<img file="00000006.tif" he="13" wi="45" img-format="tif" img-content="undefined" />
isolated extractor 302. Thereafter, the stereo decoder ICC data can be combined with power conditions to form the matrix E on the 2P spaced from the diagonal elements. For example, for a total of N = 3 objects of which the first two form a single pair (1, 2), the transmitted energy and correlation data have the form:
S1, S2, S3 and p1.
In this case the union of the matrix E yields:
<img file="00000007.tif" he="21" wi="61" img-format="tif" img-content="undefined" />
Object parameters in prediction mode according to the present invention are intended to form the matrix with coefficient prediction object (ODP) dimension N × K, available for the decoder so that
<img file="00000008.tif" he="5" wi="45" img-format="tif" img-content="undefined" />
In other words for each object there is a linear combination of the downlink channel mixing so that the object can be recovered approximately according to:
<img file="00000009.tif" he="5" wi="81" img-format="tif" img-content="undefined" />
In a preferred embodiment, the object extractor prediction coefficient (OPC) 401 solves the normal equations
<img file="00000010.tif" he="5" wi="41" img-format="tif" img-content="undefined" />
or, for a more attractive valuation of real predictive coefficient object (LFS), he decides:
<img file="00000011.tif" he="7" wi="66" img-format="tif" img-content="undefined" />
In both cases, if you accept the estimated actual weight matrix downmix D and nonsingular covariance downmix, then the left multiplication with D that
<img file="00000012.tif" he="5" wi="32" img-format="tif" img-content="undefined" />
where I - identity matrix of dimension K.
If D has full rank, according to the elementary linear algebra set of solutions (9) can be parameterized max {K · (NK), 0} parameters. The principle involved in 402 when sharing data encoding ORF. Full matrix of prediction can be recovered at the decoder from the reduced set of parameters and a downmix matrix.
For example, consider the downmix obtain stereo downmix (K = 2), comprising three objects (N = 3>) - a stereo music (s1, s2) and a center panned single instrument or vocal track s3.
Downmix matrix has the form:
<img file="00000013.tif" he="14" wi="55" img-format="tif" img-content="undefined" />
That is, the left downmix channel is x1 = s1 + s3 / √2, and the right channel - x2 = s2 + s3 / √2.
Factors predicting the object (LFS) for single track seek closer to s3≈c31x1 + c32x2, and in this case, the equation (11) can be solved to give c11 = 1-c31 / √2, c12 = -s32 / √2, c21 = -c31 / √2 = 1 and c22-c32 / √2.
This implies that a sufficient number of prediction coefficients of the object (OCR) is determined by K (NK) = 2 * (3-2) = 2.
OPC c31, c32 can be found from the normal equations
<img file="00000014.tif" he="13" wi="94" img-format="tif" img-content="undefined" />
The transcoder SAOC - MPEG Surround
With regard to Figure 7, the M = 6 output channels of the 5.1 configuration are
(y1, y2, ..., y6) = (If, Is, rf, rs, c, lfe).
The transcoder should give output stereo downmix (l0, r0) and the configuration parameters for the TTT and OTT. Since attention is now focused on a stereo downmix hereinafter be assumed that K = 2. Since both the parameters of the object, and Options MPS TTT exist in the energy and prognostic mode, it is necessary to consider all four combinations.
The energy mode is effective, for example, when the downmix audio coder is not a coder in the considered wave frequency range. It is understood that the parameters of MPEG Surround, which will be mentioned below, prior to shipment should undergo proper quantization and encoding. To further clarify the above four combinations should be recalled that it is:
1) object parameters in energy mode and transcoder in prediction mode;
2) object parameters in energy mode and transcoder in energy mode;
3) object parameters in prediction mode (prediction coefficient object ODP) and transcoder in prediction mode;
4) object parameters in prediction mode (LFS) and the transcoder in energy mode.
If in the considered frequency interval of the downmix audio coder is a coder wave type, the object parameters can be fixed in a power mode and a prediction mode, and the transcoder should preferably operate in prediction mode. If in the considered frequency interval of the downmix audio coder is not a wave-type encoder, the encoder of the object and the transcoder should both operate in energy mode. The fourth combination is less relevant, so that further description will only affect the first three combinations.
Object parameters in energy mode
In power mode, the data available to the transcoder are described by three matrices (D, E, A). MPEG Surround OTT parameters are formed by evaluating the energy and correlation parameters in the virtual audiorenderinge passed parameters and matrix audiorenderinga A dimension 6 × N. The set is presented as a six-channel covariance
<img file="00000015.tif" he="7" wi="80" img-format="tif" img-content="undefined" />
Introduction (5) into (13) yields the approximation
<img file="00000016.tif" he="5" wi="77" img-format="tif" img-content="undefined" />
which is completely determined by the available data. Let fa refers to the elements of F. Then CLD and ICC parameters are determined from:
<img file="00000017.tif" he="13" wi="62" img-format="tif" img-content="undefined" />
<img file="00000018.tif" he="13" wi="62" img-format="tif" img-content="undefined" />
<img file="00000019.tif" he="13" wi="62" img-format="tif" img-content="undefined" />
<img file="00000020.tif" he="12" wi="51" img-format="tif" img-content="undefined" />
<img file="00000021.tif" he="12" wi="51" img-format="tif" img-content="undefined" />
where J> - or the absolute value of <p (z) = \ z \, or the operator of the actual value of <p (z) Pe {z}. As an illustrative example, consider the case of three objects previously described in relation to equation (12). Imagine a matrix audiorenderinga
<img file="00000022.tif" he="38" wi="28" img-format="tif" img-content="undefined" />
Thus, the problem consists in placing audiorenderinga object 1 between the right front and right panoramic position, object 2 - between the left front and left panoramic position and object 3 - front right, center and bass channel optimization (lfe). For simplicity it is also assumed that all of these three objects are uncorrelated and have equal power, so that
<img file="00000023.tif" he="19" wi="28" img-format="tif" img-content="undefined" />
In this case, the right side of formula (14) becomes
<img file="00000024.tif" he="38" wi="46" img-format="tif" img-content="undefined" />
Substituting the appropriate values into formulas (15) - (19) we obtain:
<img file="00000025.tif" he="13" wi="92" img-format="tif" img-content="undefined" />
<img file="00000026.tif" he="13" wi="92" img-format="tif" img-content="undefined" />
<img file="00000027.tif" he="13" wi="92" img-format="tif" img-content="undefined" />
<img file="00000028.tif" he="12" wi="65" img-format="tif" img-content="undefined" />
<img file="00000029.tif" he="12" wi="61" img-format="tif" img-content="undefined" />
In response MPEG Surround decoder will receive instruction on the introduction of a de-correlation between the right front and right panoramic position, but avoid de-correlation between the left front and left panoramic positioning.
For TTT MPEG Surround parameters in prediction mode, the first step would be the formation of condensed matrix audiorenderinga A3 dimension 3 × N for the combined channels (l, r, qc), where q = 1 / √2. This implies that A3 = D36A, wherein a partial downmix matrix from 6 to 3 is determined by
<img file="00000030.tif" he="19" wi="86" img-format="tif" img-content="undefined" />
Partial downmix weights wp, p = 1, 2, 3 are adjusted so that the energy of wp (y2p-y2p + 1) is the sum of the energy-y2p || 1 || 2+ || || 2 to y2p limiting factor. All the data required to derive the partial downmix matrix, D36 is available in F. Next, a prediction matrix C3 is formed of dimension 3 × 2 in such a way that
<img file="00000031.tif" he="7" wi="35" img-format="tif" img-content="undefined" />
More preferably, a matrix display, previously taking into account the normal equations C3 (DED *) = A3S.
The result of normal equations solution best meets the waveform to (21), taking into account the covariance model of the object E. It is recommended to do some post-processing matrix C3, including the row coefficients for the complete or selective compensation of projected losses channels.
To illustrate and explain the steps above, you must continue with the example audiorenderinga previously defined six channels. In considering the elements of F should be noted that the weight of the downmix are solving equations
, P = 1, 2, 3,
<IMG>
<img file="00000033.tif" he="20" wi="56" img-format="tif" img-content="undefined" />
<IMG>
So that (w1, w2, w3) = (1 / √1, √3 / 5,1 / √2).
<img file="00000034.tif" he="21" wi="58" img-format="tif" img-content="undefined" />
<IMG>
<img file="00000035.tif" he="19" wi="58" img-format="tif" img-content="undefined" />
<IMG>
The matrix C3 contains the best weights for approximation to the desired result audiorenderinga object combined channels (l, r, qc) during the downward mixing. This general type of matrix operation can not be performed by the decoder MPEG Surround, which is linked to the limited space of TTT matrices due to the use of only two parameters. Purpose converter downmix (result downmix) relating to the present invention, it is necessary to pretreat the downmix object so that the combined effect of the pretreatment and the matrix TTT MPEG Surround match the desired result upmixing (upmix), described by C3 .
<img file="00000036.tif" he="19" wi="67" img-format="tif" img-content="undefined" />
<IMG>
<img file="00000037.tif" he="5" wi="39" img-format="tif" img-content="undefined" />
<IMG>
<img file="00000038.tif" he="12" wi="50" img-format="tif" img-content="undefined" />
<IMG>
<img file="00000039.tif" he="5" wi="39" img-format="tif" img-content="undefined" />
<IMG>
In general case G is reversible, and (23) has a unique solution for CTTT, satisfying CTTTGTTT = I.
TTT parameters (α, β) are determined by the decision.
For the particular example discussed earlier can easily confirm that the solutions meet
.
<IMG>
<img file="00000041.tif" he="14" wi="149" img-format="tif" img-content="undefined" />
<img file="00000042.tif" he="14" wi="112" img-format="tif" img-content="undefined" />
<IMG>
<img file="00000043.tif" he="5" wi="38" img-format="tif" img-content="undefined" />
<img file="00000044.tif" he="5" wi="46" img-format="tif" img-content="undefined" />
<IMG>
<img file="00000045.tif" he="14" wi="75" img-format="tif" img-content="undefined" />
simply selected
<IMG>
Further observation shows that such a diagonal downmix converter can be passed on the way from the object to Transcoder MPEG Surround implemented and the introduction of the parameters of arbitrary downmix gain (ADG) decoder MPEG Surround. In this case, increments in the logarithmic domain correspond ADGi = 10log10 (wn / zn) for i = 1, 2.
<img file="00000046.tif" he="5" wi="35" img-format="tif" img-content="undefined" />
The prediction mode object available data is presented in three matrices (D, C, A), where C - N × 2 matrix comprising N pairs of prediction coefficients object tonnes. Due to the relative prediction coefficients continue to assess the energy parameters of MPEG Surround will need access to the parameters of approximation to the covariance matrix of 2 × 2 down-mixing facility,
<img file="00000047.tif" he="5" wi="39" img-format="tif" img-content="undefined" />
Preferably, if this information is made available by the object encoder as part of the information on the downlink mixing, but can also be estimated from the transcoder from measurements of the received downmix, or indirectly derived from (D, C) by approximate object model analysis. If there are Z object covariance can be estimated by administering prediction model Y = CX, resulting in
<IMG>
and all the MPEG Surround OTT parameters and energy mode TTT can be estimated from E as in the case of the power parameters of the object. However, the greatest advantage of using the prediction coefficients ODP object is shown in combination with MPEG Surround TTT parameters in prediction mode. In this case, the approximation of the waveform D36Y≈A3SKH immediately gives the reduced prediction matrix:
C3 = A3S,
by relying on the steps that remain to the formation of the TTT parameters (α, β) and downmix converter similar to the production process parameters in energy mode. In fact, steps from the formula (22) to (25) are identical.
The resulting matrix G is fed to the downmix converter results, and the TTT parameters (α, β) are forwarded to the decoder MPEG Surround.
<img file="00000048.tif" he="14" wi="77" img-format="tif" img-content="undefined" />
In all the cases described above, the inverter 501 of the object to stereo downmix output provides the data to the approximate 5.1-channel stereo down-mix as a result audiorenderinga original audio objects. This can be expressed stereoaudiorendering A2 matrix dimension 2 × N, defined as A2 = D26A. In many implementations, downmix is of independent interest, while attention is drawn to the possibility of direct control stereoaudiorenderingom A2. As an illustrative example, consider again the case with the imposition of a stereo in the center panned mono voice track, encoded by the particular case of procedures outlined in the description of figure 8 with the notes in the context of formula (12). Adjusting the dynamic range of the voice user may be through according audiorendering
<img file="00000049.tif" he="5" wi="40" img-format="tif" img-content="undefined" />
where ν - ratio control voice music. Structure transform matrix results downmix based on the expression
<img file="00000050.tif" he="5" wi="58" img-format="tif" img-content="undefined" />
For the object parameters obtained on the basis of prediction, it should only substitute approximation S≈CDS and receive matrix converter G = A2s. For the parameters of the object based on energy indicators should solve the normal equations
<IMG>
9 is a diagram of a preferred embodiment of the encoder of audio objects in accordance with one aspect of the present invention. The encoder 101 of audio objects in general have been described in the explanation of the preceding graphic diagrams. The encoder of audio objects, generating the encoded object signal uses the plurality of audio objects 90, indicated in Figure 9 as an input the down mixer 92 and an object parameter generator 94. Furthermore, the encoder of audio objects 101 includes the downmix information generator 96 for generating downmix parameters 97 fixing procedure for allocating a plurality of audio objects in at least two downmix channels, indicated in the diagram as the paths 93 coming from the down mixer 92.
Parameter generator for generating object parameters 95 of audio objects, wherein the object parameters are calculated such that the reconstruction of the audio object is possible using the object parameters and at least two downmix channels 93. It is important that the reconstruction is carried out not by the encoder, and by the decoder. However, a complete reconstruction of the part of the decoder is possible thanks to the calculation of the parameters of objects 95, carries the generator object parameters of the encoder.
Additionally, audio objects encoder 101 includes an output interface 98 for generating the encoded audio object signal 99 using the downmix information 97 and the object parameters 95. Depending on the purpose of the downmix channels 93 can also be used as the encoded audio object signal. Thus there may be situations in which the output interface 98 generates an encoded audio object signal 99 which does not include the downmix channels. This situation may arise when any downmix channels to be used by the decoder are already available to the decoder so that information on the downmix signal and the audio object parameters transmitted downmix channels separately. The benefits of such a situation can be extracted when the downmix channels facilities 93 can be bought separately from the parameters of objects and information on the downward mixing for a smaller amount of money, and the parameters of objects and information on the downmix can be purchased for additional funds in order to provide the user side decoder to obtain added value.
In the absence of the object parameters and the downmix information, the user can convert the downmix channels in the stereo or multi-channel signal depending on the number of channels involved in the downmix. Naturally, the user may also create a mono signal by simply adding at least two transmitted object downmix channels.
Object parameters and data downmix provide the user with the flexibility of acoustic transformation and improving the quality and usefulness of acoustic sound objects, allowing multipurpose audiorendering for later playback of audio content from the audio equipment of any kind - on the stereo, multi-channel systems or even systems of the wave field synthesis. If the wave field synthesis unit is still not very popular, the multi-channel 5.1 or 7.1 are increasingly being applied to the consumer market.
10 is a diagram of the audio synthesizer for generating output data. To perform its functions audiosintezator comprises output data synthesizer 100. The output data synthesizer receives an input downmix information 97 and audio object parameters 95 and, probably, intended audio source characteristics, such as spatial location of sound sources or a user-defined dynamic range of the particular source audiorenderinga result with 101.
Synthesizer output 100 for generating output data needed to generate a plurality of output channels preconfigured audio output, reconstructing a plurality of audio objects. Well synthesizer output 100 implements its functionality using the parameters downmix audio object parameters 97 and 95. According to the explanations to the 11 given below, the output data are numerous indicators for various purposes, including the rendering of a specific output channels, or simply recreate the original signals or transcoding parameters a characteristic of spatial transformation to form a spatial configuration for upmixing without audiorenderinga output channels, for example for storage or shipment of these spatial parameters.
The general scheme of the present invention displayed in Figure 14. Here, the block encoder 140 includes an encoder of audio objects 101, which receives at input N audio objects.
At the output of the technical performance of a preferred embodiment of the encoder of audio objects except the downmix information and the object parameters which are not shown in Figure 14, is formed by a number K downmix channels. In accordance with the present invention, the number of downmix channels must be greater than or equal to two.
Downmix channels are transmitted to the decoder unit 142, which includes a spatial upmixer 143. The spatial upmixer 143 may include audiosintezator being part of this invention, if audiosintezator operates in the transcoder. However, if audiosintezator 101 as shown in Figure 10, operates in the spatial upmix, in this embodiment, and the spatial upmixer 143 and audiosintezator are one and the same apparatus. Spatial upmixer generates M output channels for playback via M speakers. These speakers are placed at predetermined points of the surrounding space and collectively form an acoustic output signal of a predetermined configuration. Output channel audio output a predetermined configuration may be seen as a digital or analog electrodynamic acoustic signal broadcast from the output spatial up-mixer 143 to the input of the loudspeaker positioned in a predetermined manner among a certain number of sources configured audio output. Depending on the situation, if the runs stereoaudiorendering, M number of output channels can be equal to two. When the multi-channel audiorenderinga M number of output channels is greater than two.
In most common situation in which the number of downmix channels less than the number of output channels because of the technical requirements of data transmission paths. In such cases, the number M may be much larger than the number of K, exceeding it twice or even more times.
14 is further given to a matrix representation of the functions performed by the block encoder and decoder unit in the framework of the present invention. In most cases the sample quantities processed blocks. Therefore, as can be seen from equation (2), audio object is displayed as a series of L sample values. The matrix S has N lines corresponding to the number of objects and L columns corresponding to the number of samples. The matrix E is calculated by the equation (5), and includes N rows and N columns. The matrix E contains the parameters of the object when the object parameters are given in the power mode. For uncorrelated objects matrix E, as shown in the context of equation (6) has only the main diagonal elements, each of which displays a power audio object. All off-diagonal elements, as previously described, represents the correlation of two audio objects, which is particularly important when some objects are two channels of a stereo signal.
Depending on the features of embodiment, equation (2) represents the time domain signal. After that, the energy generated by a single figure for the entire range of audio objects. Preferably, however, if the audio objects are processed by time-frequency converter based on, for example, any conversion algorithm or filter bank, and, in the latter case, equation (2) holds for each sub-band, thereby allowing the formation of the matrix E for each subband and certainly, for each time interval.
Matrix X downmix channels has K lines and L columns and is calculated according to equation (3). As can be seen from equation (4), the M output channels are calculated using the N objects by using the so-called matrix audiorenderinga A to N objects. Depending on the situation, the N objects can be reconstructed block decoder using the results of the downmix and the object parameters, wherein audiorendering can be applied directly to the reconstructed object signals.
On the other hand, the array downmix can be directly transformed into signals of the output channels without an accurate calculation of the source signals. A Matrix audiorenderinga mainly individually positioned sources in accordance with preconfigured audio output. Suppose there are six objects and six output channels, then each entity may be associated with each output channel, and this scheme will be reflected audiorenderinga matrix. However, if you place all the objects within the acoustic space between the two speakers audiorenderinga matrix A, reflecting the new positioning, take another look.
Audiorenderinga matrix, or in a more general sense, the projected spatial localization of objects, as expected relationship dynamic range of sound sources may generally be calculated by the encoder and transmitted to the decoder as a so-called scene description. However, in other embodiments, a description of the scene may be performed directly by the user to generate his own predetermined upmix for obtaining a predetermined configuration by itself the output acoustic signals. Thus, the transmission of the scene description is not mandatory procedure is the scene description may be implemented by the user satisfaction with the achievement of its own requests. The user may, for example, at will localize certain audio objects at places other than the position in which these objects are initially located and which was generated for them. There are also cases where the audio objects are introduced as such, without the presence of the "original" and its location in relation to other, real, object. In such situations, the sound sources initially positioned relative to each other by the user.
Returning to Figure 9, consider the down-mixer 92. It is designed to reduce the phonogram when mixing a plurality of audio objects to the number of downmix channels, wherein the number of audio objects exceeds the amount of downmix channels, wherein the down-mixer coupled to the downmix information generator so that the distribution of the set audio objects over a plurality of channels of the downmix is performed in accordance with the downmix parameters. Indicators downmix information generated by the downmix generator 96 in Figure 9, may be created automatically or manually controlled. It recommended downmix data processing with less resolution than the object parameters. With this service information bits can be stored without loss of quality, since fixed downmix parameters for specific parts of the phonogram or solitary slowly changing downmix state, does not require frequency selectivity, it is quite sufficient. The variant of the invention, in which the downmix information represents a downmix matrix having K lines and N columns.
Indicator in the row of the downmix matrix has a certain value when the audio object corresponding to this index in the downmix matrix is present in the channel downmix represented among the down-mix matrix. When the audio object is included into more than one downmix channel, a particular value have more than one row of the downmix matrix. In this preference if the quadratic value in addition to the individual audio object add up to no more than 1.0. Nevertheless, other values are possible.
Additionally, audio objects can be input into one or more downmix channels with varying levels, and these levels may be indicated in the matrix downmix weights different from unity and components generally 1.0 for a particular audio object.
When the downmix channels are included in the encoded audio object signal generated by the output interface 98, the encoded audio object signal may be, for example, a multiplexed signal is time-multiplexed in a specific format. Conversely, an encoded audio object signal can be any signal which allows the decoder unit via separate object parameters 95, the parameters and the downmix channels 97 downmix 93. In addition, the output interface 98 can include encoders object parameters, information by downmix channel or downmix. Encoders for the object parameters and the downmix data may be differential encoders and / or entropy encoders, and encoders for the downmix channels can be mono or stereoaudiokodery such as MP3 encoders or AAC (Advanced Audio Codec). All these encoding operations result in a further compression of data for subsequent reduction of the data rate required for the encoded audio object signal 99.
Depending on the particular application of the down-mixer 92 provide its function stereo representation of background music, the at least two downmix channels and the introduction of the at least two downmix channel voice sound track in a predetermined ratio. With this embodiment, the first version of the background music channel extends over a first channel and a second downmix channel background music - the second downmix channel. The result of this arrangement is optimal stereo playback of background music on the stereo. The user has the ability to position the soundtrack voice between left and right stereo speakers in stereo speakers. Alternatively, the first and second background music channels can pass from one channel downmix and the voice track may be done on another channel downmix.
Thus, except for one channel downmix can be completely separate voice from background soundtrack of musical accompaniment, which, inter alia, meets the requirements of karaoke. However, the channels of a stereo playback quality background music suffers from the parameterization of the object, which, of course, is a method of lossy compression.
Down-mixer 92 is configured to sum the time domain count for readout. For such uses summation samples of audio objects, intended for a single downmix channel before downmixing. If the audio object is introduced into a downmix channel in a certain percentage before summing of samples must be performed prior weighing. In addition, the summation may be performed in the frequency domain or in the subband, i.e. in the region following the time-frequency transformation. Thus, the downmix can be performed even in the filter bank, when the time-frequency conversion is performed in the filter bank or in the transform domain when the time-frequency conversion is an FFT (fast Fourier transform, FFT), MDCT (Modified Discrete Cosine Transform , MDCT), or any other transform.
According to one aspect of the present invention, the object parameter generator 94 generates energy parameters and further - the correlation parameters between two objects when two audio object together represent the stereo signal as can be seen from the following equation (6). On the other hand, the object parameters are prediction mode parameters.
15 is a flowchart of a calculation method or audio object prediction parameters. As already explained with respect to equations (7) to (12), subject to some statistical calculation information on the downmix channels in the matrix X and the audio objects in the matrix S. In particular, block 150 shows a first step of calculating the real part S · X * and the real part of X · X *. These real parts - it is not just numbers but the matrix and the matrix in one embodiment are defined by notations in equation (1) when considering implementation, following the equation (12). In most cases, the step 150 can be calculated using information available audio objects in the encoder 101. Then, as shown in step 152, calculated prediction matrix C. In particular, as is customary in the prior art, it is necessary to solve the system of equations so that We have obtained all of the matrix prediction dimension N rows and K columns. Mainly, the weighting factors cn, i, in equation (8) is calculated so that the weighted linear summation of all downmix channels reconstructs a corresponding audio object with the highest possible quality. This matrix gives the best prediction of the result of the reconstruction of audio objects, the greater the number of downmix channels is activated.
Further detail will be considered 11. In particular, Figure 7 displays several kinds of output data used for creating a plurality of output channels with a predetermined output configuration. Line 111 displayed a situation in which the output data synthesizer output 100 are reconstructed audio sources.
The input data required output data synthesizer 100 for reconstructing the audio sources include downmix information, the downmix channels and the audio object parameters. Thus for playback reconstructed sources not necessary to create the configuration output and pre-position themselves acoustic sources within the spatial audio output configuration. In the state depicted in Figure 11 the number 1, the output data synthesizer 100 would output reconstructed form sources of sound signals. In the case of using as parameters the audio object prediction parameters output data synthesizer 100 works as defined formulated in equation (7). When the parameters of the object are recorded in the power mode, to recreate the original synthesizer output signal using the inverse matrix of downmix matrix and energy.
Alternatively, the output data synthesizer 100 may perform the functions of the transcoder as illustrated for example in block 102 in Figure 1b. When the synthesizer output signal in a transcoder generating surround mixer parameters required data downmix audio object parameters, the output configuration and planned spatial localization of sound sources. In particular, the configuration of the output signal and the planned spatial positioning provided by the matrix A. Thus audiorenderinga for generating a surround mixer parameters is not necessary in the presence of downmix channels, so that a more detailed explanation will be given in the context of Figure 12. Depending on the situation, surround mixer parameters generated by the output data synthesizer 100 can be further used directly surround mixer type MPEG Surround for upmixing the descending channel mix. With this version of embodiment downmix adjustment object is not necessarily enough to use a simple conversion matrix containing only the diagonal elements as described in relation to equation (13). The format 2 in the line 112 to the output data synthesizer 11 100 issued surround mixer parameters and, preferably, the conversion matrix G in Equation (13) including a performance gain that can be used as parameters of an arbitrary downmix gain (ADG ) decoder MPEG-surround.
In format 3 in row 113 in Figure 11 comprise the output parameters of the mixer surround a conversion matrix such as that shown in the context of equation (25). In this context, the output data synthesizer 100 does not necessarily have to actually convert the object downmix into a stereo downmix. Number 4 in line 114 in Figure 11 uses a different format of data output operation synthesizer 100, shown in Figure 10. In this case the transcoder works as element 102 in Figure 1b, and outputs not only spatial mixer parameters of sound, but the results and further converted downmix.
This eliminates the need for the withdrawal of the conversion matrix G in addition to the converted downmix. Output the converted downmix and the spatial mixer parameters is sufficient sound, which is clear from Figure 1b.
Format 5 describes another application of the output data synthesizer 100 shown in Figure 10. Under the conditions indicated in row 115 in Figure 11, the output data generated by the output data synthesizer do not include any spatial mixer parameters of sound, but only include, for example, the conversion matrix G in Equation (35) actually contain or directly output the stereo signals as shown in line 115. In this embodiment, only stereoaudiorendering interest and any spatial mixer parameters are not required sound. However, to generate a stereo output requires all available input information is available, as shown in Figure 11.
Another mode of the synthesizer output data is displayed in a format in row 6 116. In this case, the output data synthesizer 100 generates a multi-channel output and is analogous to component 104 in Figure 1b. For this output data synthesizer 100 requires all available input information on the basis of which it forms the channel output, consisting of more than two output channels to be reproduced using a corresponding amount of acoustic speakers, are localized in space in accordance with a predetermined audio output configuration. So multichannel output can be 5.1-channel output, 7.1-channel analog audio output or 3.0-channel output in the presence of left, center and right speakers.
Further reference is made to Figure 11 to illustrate an example of calculating a number of parameters, taken from a decoder MPEG-surround, on the basis of the principle of parameterization illustrated in Figure 7. As already stated, Figure 7 illustrates the process of parameterization block decoder using MPEG-Surround, starting from the input stereo downmix 70 having a left I0 and the right downmix channel r0. Schematically, both downmix channels are entered in the so-called block "two-to-three" block 71. The "two-to-three" is controlled by several input parameters 72. Unit 71 generates three output channels 73a, 73b, 73c. Each output channel is input to a "one-to-two." This means that channel 73a is introduced into the block 74a, 73b is introduced into the channel unit 74b, and 73c is introduced into the channel unit 74c. Each block has two output channels. Block 74a outputs the left and right front lf panoramic ls channels. At the same time unit 74b outputs a right-front and right panoramic rf rs channels. However, block 74c outputs a center channel and a channel with low frequency optimization (lfe). It is important that the whole process of upmixing downmix channels 70 and output channels is performed using a matrix operation, and the tree structure shown in Figure 7, is not necessarily implemented step-by-step, and may be carried out via one or a few matrix operations. Furthermore, the intermediate signals indicated as 73a, 73b and 73c, are not calculated specifically to any particular device implemented, but shown in Figure 7 only for clarity. However, the blocks 74a, 74b receive some residual signals res1 OTT, OTT res2, which can be used for insertion into the output signals to a point of randomness.
As is known from the specification of the decoder MPEG-surround, control unit 71 is carried out with or prediction parameters CPC or energy parameters CLDrrr. For the upmix from two channels to three channels are required, at least two prediction parameters SRS1, SRS2, or at least two energy parameters and. However, in the block 71 may be introduced correlation exponent, ICCTTT, which, however, is only a minor response is not necessarily to be used in the same embodiment of the technical solution of the invention. 12 and 13 is a flowchart and / or the necessary means for calculating the entire complex object parameters 95 9 - CPC / CLDTTT, CLD0, CLD1, ICC1, CLD2, ICC2, information on the downmix 9 and 97 planned spatial positioning the sound sources, for example, the scene description 101 as displayed in Figure 10. These settings are pre-defined by the format of audio output for 5.1-channel surround sound.
<IMG>
<IMG>
Naturally, such a special calculation of parameters for this particular technical solution can be adapted to different output signal or species parameterization in accordance with the concept of the invention. Moreover, the sequence of steps of the algorithm or the arrangement means 12 and 13a, b are given only as an illustrative example and may undergo changes in boundaries logic mathematical equations.
Step 120 maintains a matrix A. The matrix audiorenderinga audiorenderinga positioned in the acoustic space of each source from multiple sources based on a predetermined configuration of the output signal.
Step 121 ensures the formation of a partial downmix matrix D36 in accordance with equation (20). This matrix allows the downward mixing with six output channels to three channels and has a dimension of 3 × N. If required to generate a greater number of output channels than the 5.1 configuration, such as creating an 8-channel format of the output signal (7.1), the matrix provided in block 121, the matrix will be D38.
Step 122 provides the formation of the reduced matrix audiorenderinga A3 by multiplying the matrix and full matrix D36 audiorenderinga as determined in step 120.
Step 123 provides an introduction downmix matrix D. This matrix downmix D can be retrieved from the encoded audio object signal when the matrix is contained in the signal. Alternatively, the downmix matrix can be parametrized, for example, for the introduction of specific data on the downmix and the downmix matrix generation G.
Step 124 provides, in addition to the energy matrix of the object. This object energy matrix is reflected in the object parameters for the N objects and can be extracted from the imported audio objects or reconstructed using a certain set of rules. Such a set of rules recovery may include an entropy decoding etc.
Step 125 provides the formation of a "shortened" matrix predictions C3. The values of this matrix can be calculated by solving a system of linear equations according to step 125. In particular, the elements of matrix C3 can be calculated by multiplying both sides of the equation by the inverse of (DED *).
Step 126 calculates a conversion matrix G. The conversion matrix G dimension K * K is formed according to the equation (25). To solve the equation in step 126 requires special matrix DTTT, generated in the step 127. An example for this matrix is given in equation (24) and the definition can be obtained by starting from the corresponding equation for CTTT, as described by equation (22). Thus, the equation (22) defines the procedure in step 128. Step 129 determines the equation for calculating the matrix CTTT. Once the block based on the equation 129 will be determined matrix CTTT, can be derived the parameters α, β and γ, which are the parameters of the CPC (channel prediction coefficient). Γ is recommended to set a value of one, then only the input parameters of the CPC unit 71 will α and β.
The remaining parameters necessary for the algorithm in Figure 7 are the parameters input into blocks 74a, 74b and 74c. The calculation of these parameters is described in the context of 13A. Step 130 provides the formation of the matrix A. The dimension of the matrix audiorenderinga audiorenderinga A is N lines for the number of audio objects and M columns for the number of output channels. This matrix audiorenderinga comprises information based on a vector of the scene when the scene vector is used. In most audiorenderinga matrix includes information about a specific location for a given configuration of the output signal. If we consider audiorenderinga matrix A, for example, in the context of the following equation (19), it becomes clear how the may be encoded localization sites in the matrix structure audiorenderinga. Naturally, there can be used, and other methods well-defined position, such as by values not equal to 1. Furthermore, using the values on the one hand, less than 1 and, on the other hand, more than one can control the volume specific audio objects.
The variant of embodiment in which the matrix is formed audiorenderinga decoder module without using any information from the encoder.
This enables the user to place the audio objects arbitrarily as desired, without regard to their mutual spatial arrangement, recorded data encoder.
It is also possible version of technical solutions in which the relative or absolute positioning of the acoustic sources may be encoded by the encoder module and transmitted to the decoder in a certain stage of the vector. Then, the decoder module information regarding localization of sound sources, preferably independent of set points audiorenderinga processed to form the resulting matrix audiorenderinga reflecting spatial arrangement of sound sources which are tailored to the specific audio output configuration.
Step 131 provides the formation of the matrix E the energy performance of the object, which has already been discussed in connection with step 124 in Figure 12. This matrix has the dimension N × N, and audio object contains parameters. One embodiment of the invention provides such a matrix power object parameters for each subband and each unit time samples or subband samples.
Step 132 calculates the matrix of the energy parameters of the output signal F.
F - covariance matrix of the output channels. Since in this case the output channels remain uncertainty matrix F of the energy parameters of the output signal is calculated using the matrix and the matrix audiorenderinga energy characteristics. These matrices are formed in steps 130 and 131 with direct access to the matrices in the decoder module. Thereafter, with the use of special equations (15), (16), (17), (18) and (19) calculates the performance level difference channels CLD0, CLD1, CLD2 and the characteristics of the inter-channel coherence ICC1 and ICC2 to obtain parameters for blocks 74a , 74b, 74c. Importantly, the spatial parameters are calculated by combining the specific elements of the output energy indicators F.
In step 133 all the parameters for spatial up-mixer such as, for example, which is shown schematically in Figure 7, are prepared.
In the previously described embodiments of the invention the parameters of the object presented as energy characteristics. However, when the object parameters are prognostic representation, ie in the form of a matrix with the predictions of the objects shown at 124a, paragraph 12, to calculate the reduced prediction matrix C3 relatively simple matrix multiplication unit 125a according to the illustrations and explanations in the context of equation (32). The matrix used in block 125a is the same matrix A3, which is referred to in block 122 in Figure 12.
When the object prediction matrix C is generated by the encoder and transmitted audio objects to the decoder, requires additional computations for the preparation parameters for blocks 74a, 74b, 74c. These auxiliary steps presented in Figure 13b. Again, the object prediction matrix C is formed as a block 124a at 13b, which is similar to the description of the block 124a in Figure 12. Then, as described in connection with equation (31), the covariance matrix of the object downmix Z is calculated using the transmitted downmix or is generated and transmitted as overhead information. After transferring data matrix Z decoder need not perform any calculations of energy parameters, leading essentially to the resumption of delayed processing of certain data and to increase the total load of the decoder unit. However, when these questions are crucial to a particular application, the transmission bandwidth can be saved, and the covariance matrix of the object downmix Z can also be calculated using the down-mix samples, which are certainly available in the decoder module. Once the action step 134 will be completed and the covariance matrix of the object downmix is ready, the matrix E is the energy parameters of the object can be calculated according to the instructions of step 135 using the prediction matrix C and the downmix covariance matrix or matrices Z "energy down-mix". Upon completion of step 135 may be performed all the steps described above relating to Figs 13a - namely 132, 133, in order to form all the required parameters for blocks 74a, 74b, 74c in Figure 7.
16 is another design solution that implements only stereoaudiorendering. Stereoaudiorendering - is an output signal in accordance with the mode number 5 or line 115 11. Here, the output of the synthesizer 100 in Figure 10 does not require any spatial parameters upmix mainly it requires a special conversion matrix G, to convert the object downmix functional and, of course, quick to set up and easy to manage a stereo downmix.
Step 160 in Figure 16 includes a matrix calculation partial downmix channels M to two. In the variant with six output channels, the partial downmix matrix would serve as a downmix matrix from six to two channels, while maintaining the possibility of using other downmix matrices. Calculation of this partial downmix matrix can be, for example, by deriving from the partial downmix matrix D36, as was done in step 121 and matrix DTTT, as was done at step 127 Figure 12.
In addition, based on the result of step 160 is generated matrix stereoaudiorenderinga A2, and in step 161 is presented a "big" matrix A. The matrix audiorenderinga audiorenderinga And - this is the same matrix, which was seen in connection with block 120 12.
Next, at step 162, the matrix may be parameterized stereoaudiorenderinga localization parameters μ and κ. When setting for μ, κ and 1 values obtained by equation (33), which makes it possible to vary the dynamic range of the voice that has already been described in the example given in the context of equation (33). However, changing other parameters such as μ and κ, can vary as the location of the sources.
Then, as shown in step 163, the conversion matrix G is calculated using equation (33).
Corrections made to the description
In particular, the matrix (DED *) can be calculated, inverted and inverted matrix can be multiplied by the right side of the block 163. Of course, it can be used other ways of solving the equation block 163. Once obtained conversion matrix G, downmix object X can be converted by multiplying the conversion matrix and the object downmix, which is reflected in block 164. Thereafter, may be configured stereoaudiorendering converted downmix X 'using two stereo speakers. Depending on the technical solutions for μ, ν and κ may be assigned specific values for calculating the conversion matrix G. Alternatively, the conversion matrix G can be calculated using all these three parameters as variables so that the parameters are set in accordance with the requirements of Users 163 after passing steps.
In preferred embodiments of the invention have been found to solve the problem of transmission of multiple independent audio objects (using a multi-channel downmix and auxiliary control data describing the objects to obektyi audiorenderinga predetermined reproducing system (loudspeaker configuration)). Administered in a manner related to the conversion object the control data in the control data compatible playback system. Further features respectively etstvuyuschie coding methods based on the algorithm of encoding MPEG Surround.
Depending on the technical requirements of a particular variant of embodiment the input methods and the resultant signal may have a form of realization in hardware or in software. This part of the invention may be performed using a digital storage medium, in particular a disk or a CD, for storing in the form of electronically readable control signals, compatible with a programmable computer system so that the input can be carried out methods. Thus, in a general sense, the present invention is a computer program product with its assigned a program code stored on a machine-readable storage device, and configured to perform at least one of the inventive methods when run this software on a computer. Stated otherwise, the inventive methods are, therefore, a computer program having a program code used to implement the inventive methods when you run this program on your computer.
In other words, the design of the present invention is an encoder of audio objects, for generating a coded signal audio object as one of the plurality of audio objects, comprising in their structure information generator downmix generating information downmix maps the order of distribution of the plurality of audio objects, at least between downmix channels;
generator parameters of audio objects; and an output interface for generating the encoded audio object signal using the downmix information and the object parameters.
Alternatively, the output interface may generate an encoded audio signal using the additional set of downmix channels.
In addition, or alternatively, parameter generator characterized in that it is capable of forming characteristics of the object with the primary time and frequency resolution, and in cases where the information generator downmix has the feature information generating downmix secondary time and frequency resolution, secondary resolution for time and frequency is lower than the primary.
Furthermore, the downmix information generator is characterized in that is able to generate the downmix information such that the downmix parameters uniformly cover the entire frequency range of audio objects.
Furthermore, the downmix information generator is characterized in that is able to generate the downmix information such that the downmix information may include a downmix matrix defined as:
X = DS,
where S - matrix representing audio objects and comprising a number of rows equal to the number of audio objects,
where D - downmix matrix, and
wherein X - matrix representing a plurality of downmix channels, and comprising a number of rows equal to the number of downmix channels.
Furthermore, information about a part of the object can have an index of less than 1 and greater than 0.
Furthermore, down-mixer is characterized in that it can form a stereo representation of background music, at least two downmix channels, and to enter voice sound recording at least two downmix channels in a predetermined ratio.
Furthermore, down-mixer is characterized in that is able to perform addition of signal samples for further introduction channel downmix according to at downmixed.
In addition, the interface data output characterized in that it is capable of performing data compression on the downmix and the object parameters before generating the encoded audio object signal.
In addition, a plurality of audio objects may include a stereo object represented by two audio objects with a non-zero correlation, and contains information about the grouping formed generator downmix information indicating the two audio object forming the stereo object.
In addition, the generator parameters of the object characterized in that it is capable of generating prediction parameters of audio objects, counting them so that the weighted addition of the downmix channels to the original object, regulated by prediction parameters, or simply to the source object results in an approximation of the source object.
Furthermore, the prediction parameters may be generated based on the bandwidth, and audio objects cover the entire frequency range.
Furthermore, the number of audio objects can be equal to N, the number of downmix channels is equal to K, and the number of object prediction parameters, the object parameters calculated generator is equal to or less than N · K.
Furthermore, the object parameter generator is characterized in that it can calculate the greatest number of objects prediction parameters K · (NK).
Furthermore, the object parameter generator may include an upmixer for increasing the number of channels received downmix using various combinations of monitored parameters prediction object;
wherein a part of the up-mixer encoder audio objects includes in its design iteration controller for finding prediction parameters of the object to be tested, resulting in the minimized deviation signal reconstructed upmixer from respective original signal among the different sets of controlled parameters prediction object .
Furthermore, the synthesizer output is characterized in that it can determine the matrix conversion using information on the downmix, the transformation matrix is calculated so that, at least in part changing the location of the downmix channels, when the audio object stored in the first channel downlink mixing representing the first half stereoploskosti, to be played in the second half stereoploskosti.
Furthermore, audiosintezator may include audiorenderer channels for performing audio output channels audiorenderinga acoustic signal to obtain a predetermined configuration by using the spatial parameters and the at least two downmix channels or the converted downmix channels.
In addition, the output data synthesizer is characterized in that is able to generate output audio channels given configuration further cycling, at least two downmix channels.
In addition, the output data synthesizer is characterized in that to calculate actual downmix weights for the partial downmix matrix such that the energy of the weighted sum of two channels is equal to the energy channels within limits.
Further, the downmix weights for the partial downmix matrix may be determined as follows:
, P = 1, 2, 3,
<IMG>
wherein wp - downmix weight, p - an integer index variable, fj, i - cell power characteristics matrix representing an approximation of the covariance matrix of output channels, a predetermined output configuration.
In addition, the output data synthesizer is characterized in that to calculate separate coefficients of the prediction matrix by solving a system of linear equations.
In addition, the output data synthesizer is characterized in that is able to solve a system of linear equations based on:
<img file="00000054.tif" he="19" wi="50" img-format="tif" img-content="undefined" />
where C3 - a matrix prediction "two-to-three", D - matrix downmix derived from information in the downstream mixing, E - the matrix of power characteristics, derived on the basis of the original audio objects, and A3 - Abbreviations matrix downmix, and where "*" denotes the complex conjugate operation.
Furthermore, the prediction parameters for upmixing "two-to-three" may be prepared parameterization prediction matrix so that the prediction matrix is defined by only two parameters, and
wherein the output data synthesizer is characterized in that is able to preprocess the at least two downmix channels so that the feedback preprocessing and the parameterized prediction matrix corresponds to a desired upmix matrix.
In addition, the parameterization of the prediction matrix might look like:
<img file="00000055.tif" he="19" wi="50" img-format="tif" img-content="undefined" />
where the index TTT - parameterized prediction matrix, a α, β, and γ - coefficients.
In addition, the conversion matrix G downmix can be calculated as follows:
G = DTTTC3
wherein C3 - prediction matrix "two-to-three", and wherein CTTT DTTT equal to I, where I - the identity matrix of "two-by-two", and wherein CTTT is based on:
<IMG>
where α, β, γ - constant coefficients.
Further, prognostic parameters for upmixing "two-to-three" can be defined as α and β, wherein γ is set as 1.
In addition, the output data synthesizer is characterized in that is able to calculate the energy parameters for upmixing "three-to-six" using the energy characteristics of the matrix F based on:
YY * ≈F = AEA *,
where A - the matrix audiorenderinga, E - energy characteristics of the matrix formed on the basis of audio objects sources, Y - matrix output channel, and "*" serves as a pointer to the complex conjugate operation.
In addition, the output data synthesizer is characterized in that is able to calculate the energy parameters by combining elements of the energy characteristics.
In addition, the output data synthesizer is characterized in that to calculate the energy parameters based on the following equations:
.
<IMG>
.
<IMG>
.
<IMG>
.
<IMG>
.
<IMG>
where φ - the absolute value of φ (z) = | z | or operator of the actual value of φ (z) = Pe {z},
where CLD0 - the first energy parameter level difference channels where CLD1 - second energy parameter level difference channels where CLD2 - the third energy parameter level difference channels where ICC1 - the first energy parameter inter-channel coherence, a ICC2 - second energy parameter inter-channel coherence, and where fi, j - energetic characteristics of the matrix elements F at positions i, j in this matrix.
Moreover, the first group of parameters may include energy parameters, and wherein the output data synthesizer characterized in that it is capable of generating the energy parameters by combining elements of the energy characteristics of F.
Further, the energy parameters may be derived based on the fact that:
.
<IMG>
.
<IMG>
where - the first energy parameter of the first group, and where - the second energy parameter of the first group of parameters.
<IMG>
<IMG>
In addition, the output data synthesizer characterized in that it is capable of calculating the weighting factors for weighting the downmix channels, the weighting coefficients for controlling the arbitrary downmix gain (ADG) of the spatial decoder.
In addition, the output data synthesizer characterized in that it is capable of calculating the weighting factors based on:
<img file="00000066.tif" he="12" wi="37" img-format="tif" img-content="undefined" />
W = D26ED * 26
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10497376B2 | Cited by | United States of America | Applicant |
| US9947325B2 | Cited by | United States of America | Applicant |
| US10699722B2 | Cited by | United States of America | Applicant |
| RU2658888C2 | Cited by | Russian Federation | Search report |
| RU2646320C1 | Cited by | Russian Federation | Search report |
| RU2698775C1 | Cited by | Russian Federation | Search report |
| US10873822B2 | Cited by | United States of America | Applicant |
| RU2676415C1 | Cited by | Russian Federation | Search report |
| US10362424B2 | Cited by | United States of America | Applicant |
| US10567899B2 | Cited by | United States of America | Applicant |
| US10674299B2 | Cited by | United States of America | Applicant |
| US10891963B2 | Cited by | United States of America | Applicant |
| WO2006048203A1 | Cites | World Intellectual Property Organization (WIPO) | – |
| RU2005103637A | Cites | Russian Federation | – |
| RU2129737C1 | Cites | Russian Federation | – |
| RU2110162C1 | Cites | Russian Federation | – |
1 priority claim, no other members on record
Priority claims1
| Document | Office | Kind | Date |
|---|---|---|---|
| 60829649 | United States of America | – |
Numbers
- Publication
- 2485605
- Application
- 201110241608
Titles2
- Russian
- ??????????????????? ????? ??????????? ? ???????????????? ????????????? ??????????? ??????????????? ??????? ????? ??????????? ????????????
- English
- IMPROVED METHOD FOR CODING AND PARAMETRIC PRESENTATION OF CODING MULTICHANNEL OBJECT AFTER DOWNMIXING