Apparatus for providing one or more adjusted parameters for a provision of an upmix signal representation on the basis of a downmix signal representation, audio signal decoder, audio signal transcoder, method and computer program using an object-related parametric information
Abstract
An audio signal encoder (600) to provide a representation of downstream mixing signal (614) and parametric information related to the object (616) based on a plurality of object signals (x1 to xN), in which The audio encoder comprises: a downstream mixing device (620) configured to provide one or more downstream mixing signals depending on downward mixing coefficients (d1 to dN) associated with the object signals (x1 to xN), whereby said one or more signals of descending mixture comprises an overlap of a plurality of object signals; a complementary information provider (630) configured to provide complementary information on the relationship between objects (OLD, IOC) that describes the level differences and correlation characteristics of the object signals (x1 to xN) and complementary information on objects individual that describes one or more individual properties of the individual object signals (x1 to xN), characterized in that the complementary information on individual objects comprises an information of the object signal tone (Ni) which describes the tones of the signals of individual objects.
Term
3.6 yearsto projected expiry
Projected expiry 28 April 2030, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
6 claims: 3 independent, 3 dependent
- 1ES 2 572 083 T3 REIVINDICACIONES 1. Un codificador de señales de audio (600) para proporcionar una representación de señal de mezcla descendente (614) e información paramétrica relacionada con el objeto (616) sobre la base de una pluralidad de señales de objeto (X1 a xn), en el que el codificador de audio comprende:un dispositivo de mezcla descendente (620) configurado para proporcionar una o más señales de mezcla descendente dependiendo de coeficientes de mezcla descendente (d-ι a ón) asociados a las señales de objeto (X1 a xn), por lo que dichas una o más señales de mezcla descendente comprenden una superposición de una pluralidad de señales de objeto;un proveedor de información complementaria (630) configurado para proporcionar una información complementaria de relación entre objetos (OLD, IOC) que describe las diferencias de nivel y las características de correlación de las señales de objeto (X1 a xn) y una información complementaria sobre objetos individuales que describe una o más propiedades individuales de las señales de objeto individuales (X1 a xn), caracterizado por que la información complementaria sobre objetos individuales comprende una información de tonalidad de señal de objeto (Ni) que describe tonalidades de las señales de objetos individuales.
- 2El codificador de señal de audio de acuerdo con la reivindicación 1, en el que el codificador de señal de audio está configurado para transmitir la información de tonalidad de señal de objeto con una resolución de frecuencia mucho más basta que otros parámetros de objetos.
- 3El codificador de señal de audio de acuerdo con la reivindicación 1 o 2, en el que el codificador de señal de audio está configurado para transmitir la información de tonalidad de señal de objeto con solo una información por objeto.
- 4Un método para proporcionar una representación de señal de mezcla descendente e información paramétrica relacionada con el objeto sobre la base de una pluralidad de señales de objeto, en el que las señales de objeto son señales de objeto de audio, comprendiendo el método:proporcionar una o más señales de mezcla descendente dependiendo de coeficientes de mezcla descendente asociados a las señales de objeto, de tal manera que dichas una o más señales de mezcla descendente comprendan una superposición de una pluralidad de señales de objeto;y proporcionar una información complementaria de relación entre objetos que describe las diferencias de nivel y características de correlación de las señales de objeto;y proporcionar información complementaria sobre objetos individuales que describe una o más propiedades individuales de las señales de objetos individuales, caracterizado por que la información complementaria sobre objetos individuales comprende una información de tonalidad de señal de objeto (Ni) que describe tonalidades de las señales de objetos individuales.
- 5Un flujo de bits de audio (700) que representa una pluralidad de señales de objetos (X1 a xn) en forma codificada, comprendiendo el flujo de bits de audio:una representación de señal de mezcla descendente (710) que representa una o más señales de mezcla descendente, en el que al menos una de las señales de mezcla descendente comprende una superposición de una pluralidad de señales de objetos;y una información complementaria de relación entre objetos (720) que describe las diferencias de nivel y características de correlación de las señales de objeto;y una información complementaria sobre objetos individuales (730) que describe una o más propiedades individuales de las señales de objetos individuales;caracterizado por que la información complementaria sobre objetos individuales comprende una información de tonalidad de señal de objeto (Ni) que describe tonalidades de las señales de objetos individuales.
- 6Un programa informático adaptado para realizar el método de acuerdo con la reivindicación 4.
Independent claims6
486 paragraphs in 32 sections, as filed
ES 2 572 083 T3
Audio signal encoder, audio bitstream, method, and computer program that uses object-related parametric information
DESCRIPTION
Technical field
Embodiments according to the invention relate to an audio signal encoder, a corresponding method and an audio bit stream.
Some more embodiments relate to corresponding computer programs.
Divisional application document EP10716830.4
Background of the Invention
In the art of audio processing, audio streaming, and audio storage, there is a growing desire to handle multi-channel content in order to improve auditory impression. The use of multichannel audio content brings considerable improvements for the user. For example, a three-dimensional auditory impression can be obtained, which leads to greater user satisfaction in entertainment applications. However, multi-channel audio content is also useful in professional environments, for example in conference call applications, as speaker intelligibility can be improved by using multi-channel audio playback.
However, it is also desirable to have a good ratio of audio quality to bitstream requirements to avoid excessive resource load caused by multichannel applications.
Lately, parametric techniques have been proposed for bitstream efficient transmission and / or storage of audio scenes containing multiple audio objects, for example, Binaural Cue Coding (Type I) (see, for example, reference [BCC]), Joint Source Coding (see, for example, reference [JSC]), and MPEG Spatial Audio Object Coding (SAOC) (see, for example, references [SAOC1], [SAOC2]).
These techniques are intended to perceptually reconstruct the output audio scene rather than by parity of a waveform.
Fig. 8 illustrates a system overview of that type of system (in this case: MPEG SAOC). The MPEG SAOC 800 system depicted in Fig. 8 comprises a SAOC 810 encoder and a SAOC 820 decoder. The SAOC 810 encoder receives a plurality of object signals x1 to xn, which may be represented, for example, as time-domain signals or as time-frequency domain signals (for example, as of a series of transform coefficients of the Fourier transform type, or in the form of QMF subband signals). The SAOC 810 encoder also typically receives downmix coefficients d1 through dN, which are associated with the object signals x1 through xn. There may be separate series of downmix coefficients for each channel of the downmix signal. The SAOC 810 encoder is typically configured to obtain one channel of the downmix signal by combining the object signals x1 to xn in accordance with the associated downmix coefficients d1 to dN. There are generally fewer downmix channels than x1 to xn object signals. To result in (at least approximately) the separation (or separate processing) of the object signals on the side of the SAOC 820 decoder, the SAOC 810 encoder supplies both one or more downmix signals (designated as downmix channels ) 812 as supplementary information 814. Supplementary information 814 describes characteristics of the object signals x1 to xn, to allow for specific object processing on the decoder side.
An approach to specify supplementary information can be found for example in US 2008/0140426 A1.
The SAOC 820 decoder is configured to receive both said one or more downmix signals 812 and the supplementary information 814. In addition, the SAOC 820 decoder is typically configured to receive user interaction information and / or control information from the user. user 822, describing an intended display configuration. For example, user interaction / user control information 822 may describe a speaker configuration and convenient spatial placement of objects that produce the object signals x1 through xn.
ES 2 572 083 T3
The SAOC 820 decoder is configured to produce, for example, a plurality of downmix decoded channel signals yi to yM. The upmix channel signals may be associated, for example, with individual speakers in a multi-speaker presentation arrangement. The SAOC decoder 820 may comprise, for example, an object separator 820a, which is configured to reconstruct, at least approximately, the object signals xi to xn on the basis of said one or more downmix signals 812 and the information complementary 814, in order to thus obtain the reconstructed object signals 820b. However, the reconstructed object signals 820b may deviate to some extent from the original object signals X1 to xn, for example, because the complementary information 814 is not sufficient enough for a perfect reconstruction due to bit rate constraints. . The SAOC decoder 820 may further comprise a mixer 820c, which may be configured to receive the reconstructed object signals 820b and the user interaction information / user control information 822, and to supply, on the basis of these, the upmix channel signals y1 to yM. The mixer 820 may be configured to use the user interaction information / user control information 822 to determine the contribution of the individual reconstructed object signals 820b to the upmix channel signals y1 to yM. The user interaction information / user control information 822 may comprise, for example, presentation parameters (which are also called presentation coefficients), which determine the contribution of the individual reconstructed object signals 822 to the channel signals upmix y1 to yM.
However, it should be noted that, in many embodiments, the separation of the objects, which is indicated by the object separator 820a of Fig. 8, and the mixing, which is indicated by the mixer 820c of Fig. 8, are carried out in one step. For this purpose, general parameters describing a direct mapping of said one or more downmix signals 812 to the upmix channel signals y1 to yM can be calculated. These parameters can be calculated on the basis of supplementary information and user interaction information / user control information 820.
Taking, now, as reference Figs. 9a, 9b and 9c, different apparatuses are described for obtaining an upmix signal representation based on a downmix signal representation and object-related supplementary information. Fig. 9a illustrates a schematic block diagram of an MPEG SAOC 900 system comprising a SAOC 920 decoder. The SAOC decoder 920 comprises, as separate functional blocks, an object decoder 922 and a mixer / presenter 926. The object decoder 922 produces a plurality of reconstructed object signals 924 that are dependent on the representation of the downmix signal (for example, in the form of one or more downmix signals represented in the time domain or the time domain). time-frequency) and supplemental object-related information (for example, in the form of object metadata). The mixer / presenter 924 receives the reconstructed object signals 924 associated with a plurality of N objects and produces, based on these, one or more upmix channel signals 928. In the SAOC 920 decoder, the extraction of the object signals 924 is done independently of the mix / presentation, which results in a separation of the object decoding functionality from the mix / presentation functionality, although it involves relatively high computational complexity.
Referring now to Fig. 9b, another MPEG SAOC 930 system is briefly described, comprising an SAOC 950 decoder. The SAOC 950 decoder produces a plurality of upmix channel signals 958 that are dependent on a signal representation of downmix (for example, in the form of one or more downmix signals) and ancillary information related to an object (for example, in the form of object metadata). The SAOC 950 decoder comprises an object decoder and mixer / presenter combination, which is configured to obtain the signals from upmix channels 958 in a co-mixing process without separation of object decoding and mix / presentation, where Parameters for such a joint upmixing process depend on both the supplementary object-related information and the display information. The whole upmixing process is also dependent on the downmix information, which is considered part of the object-related supplementary information.
To summarize the above, the provision of the upmix channel signals 928, 958 can be done in a one-step process or a two-step process.
Referring now to Fig. 9c, an MPEG SAOC 960 system is described. The SAOC 960 system comprises a SAOC to MPEG Surround 980 transcoder, rather than a SAOC decoder.
The SAOC to MPEG Surround transcoder comprises a supplementary information transcoder 982, which is configured to receive the supplementary information related to objects (for example, in the form of object metadata) or, optionally, information from said one or more mixing signals. top down and information about the presentation. The side information transcoder is also configured to provide side information over MPEG Surround (for example, in the form of a
ES 2 572 083 T3
MPEG Surround) based on certain received data. Consequently, the supplementary information transcoder 982 is configured to transform an object-related (parametric) supplementary information, which is output by an object encoder, into channel-related (parametric) supplementary information, taking into account the information about the display and optionally information about the content of said one or more downmix signals.
Optionally, the SAOC to MPEG Surround transcoder 980 may be configured to manipulate said one or more downmix signals, described, for example, by the downmix signal representation, to obtain a manipulated downmix signal representation 988 . However, the downmix signal handler 986 can be omitted, so the output downmix signal 988 representation of the SAOC to MPEG Surround 980 transcoder is identical to the input downmix signal representation of the transcoder. from SAOC to MPEG Surround. The downmix signal handler 986 can be used, for example, in case the MPEG Surround supplementary information related to channels 984 does not allow the production of an adequate auditory impression based on the downmix signal representation. transcoder input from SAOC to MPEG Surround 980, which may occur in some presentation constellations.
Consequently, the SAOC to MPEG Surround 980 transcoder gives rise to the representation of the downmix signal 988 and the MPEG Surround 984 bitstream which is why a plurality of upmix channel signals can be generated, representing the audio objects according to the presentation information input into the SAOC to MPEG Surround 980 transcoder using an MPEG Surround decoder that receives the MPEG Surround 984 bitstream and the 988 downmix signal representation .
To summarize the above, different concepts can be used to decode SAOC encoded audio signals. In some cases, an SAOC decoder is used, which produces upmix channel signals (e.g. upmix channel signals 928, 958) that are dependent on the representation of the downmix signal and related parametric side information. with the object. Examples of this concept can be seen in Figs. 9a and 9b. On the other hand, the SAOC-encoded audio information can be transcoded to obtain a downmix signal representation (for example, a 988 downmix signal representation) and ancillary information related to channels (for example, stream Surround MPEG related to channels 984), which can be used by an MPEG Surround decoder to produce the intended upmix channel signals.
In the MPEG SAOC 800 system, an overview of the system which is presented in Fig. 8, general processing is carried out in a frequency-selective manner and can be described as follows within each frequency band:
• Downmixing of N object input audio signals X1 to xn is performed as part of the SAOC encoder processing. For a mono downmix, the coefficients are indicated by d1 through dN. In addition, the SAOC encoder 810 extracts supplemental information 814 that describes the characteristics of the input audio objects. In the case of MPEG SAOC, the relationships of the object powers to each other are the most basic form of such supplementary information.
• The downmix signal (or signals) 812 and supplementary information 814 are transmitted and / or stored. For this purpose, the downmix audio signal can be compressed using well-known perceptual audio encoders such as MPEG-1 Layer II or III (also known as “.mp3”), MPEG Advanced Audio Coding (AAC), or any other audio encoder.
• On the receiver side, the SAOC 820 decoder conceptually attempts to restore the original object signal ("object separation") using the transmitted side information 814 (and, of course, the one or more downmix signals 812). These rough object signals (which are also called 820b reconstructed object signals) are then mixed into a target scene represented by M audio output channels (which can be represented, for example, by the upmix and channel signals). -ι to ywi) using a display array. For a mono output, the coefficients of the display matrix are expressed by na rN.
• Indeed, the separation of the object signals is rarely performed (or even never performed), since both the separation step (indicated by the object separator 820a) and the mixing step (indicated by the mixer 820c) are combined to obtain a single transcoding step, often resulting in a huge reduction in computational complexity.
That type of scheme has been found to be tremendously efficient, both in terms of speed of
ES 2 572 083 T3 bit transmission (only a few downmix channels need to be transmitted plus some ancillary information instead of N discrete audio signals from object or a discrete system) and computational complexity (processing complexity is mainly related to the number of output channels instead of the number of audio objects). Other benefits for the user on the receiving side include the freedom to choose a presentation setting of their choice (mono, stereo, surround, virtualized playback with headphones and so on) and the user interactivity feature: the display matrix, and consequently the output scene, can be adjusted and changed interactively by the user according to his will, personal preferences or other criteria. For example, it is possible to locate the interlocutors of a group together in a spatial area to maximize discrimination from other people who are conversing. This interactivity is obtained by producing a decoder user interface:
For each transmitted sound object, its relative level and (in the case of non-mono presentation) the spatial position of the presentation can be adjusted. This can occur in real time when the user changes the position of the associated graphical interface (GUI) sliders (for example: object level = +5 dB, object position = -30 degrees).
However, it has been found that the choice of parameters to the decoder side for the provision of the upmix signal representation (eg, the upmix channel signals yi to ywi) brings with it audible impairments in some cases.
In view of this situation, the aim of the present invention is to create a concept that results in the reduction or even elimination of audible distortion by supplying a representation of the upmix signal (for example, in the form of signals upmix channels yi to yM).
Summary of the invention
This problem is solved by an audio signal encoder according to claim 1, a method according to claim 4, an audio bitstream according to claim 5 and a computer program according to claim 6.
An embodiment according to the invention relates to an audio signal encoder for producing a downmix signal representation and object-related parametric information based on a plurality of object signals. The audio encoder comprises a downmixer configured to produce one or more downmix signals that depend on downmix coefficients associated with the object signals, whereby said one or more downmix signals comprise an overlay of a plurality of object signs. The audio encoder also comprises a supplemental information producer configured to produce supplemental object relationship information describing level differences and correlation characteristics of object signals and supplemental information on individual objects describing one or more individual properties. of the individual object signals, wherein the individual object supplementary information comprises an object signal tonality information describing hues of the individual object signals. It has been found that the provision of both supplemental object-related information and supplemental individual object information by means of an audio signal encoder allows to efficiently reduce, or even avoid, audible distortions on the audio signal decoder side. on multiple channels. While the supplemental object relationship information is used to separate the object signals on the decoder side, the supplemental information on individual objects can be used to determine whether the individual characteristics of the object signals are maintained on the decoder side. indicating that the distortions are within acceptable tolerances.
The tonality of individual objects has been found to be an important quantity from a psychoacoustic point of view, allowing for the limitation of distortions on the decoder side.
Another embodiment according to the invention relates to a corresponding method.
Another embodiment according to the invention relates to an audio bit stream representing a plurality of object signals (audio) in encoded form. The audio bitstream comprises a downmix signal representation representing one or more downmix signals, wherein at least one of the downmix signals comprises an overlay of a plurality of object (audio) signals. The audio bit stream also comprises an object relationship supplementary information describing level differences and correlation characteristics of object signals and individual object supplementary information describing one or more individual properties of the individual object signals, wherein the individual object supplementary information comprises an object signal tonality information describing hues of the individual object signals.
ES 2 572 083 T3
As discussed above, such an audio bit stream allows multi-channel audio signal reconstruction, where audible distortions that would be caused by improper setting of presentation parameters can be recognized and reduced or even eliminated.
Other embodiments according to the invention relate to a computer program for implementing the above-described method.
Brief description of the figures
Embodiments according to the invention are described below with reference to the attached figures, in which:
<td>Figure 1</td><td>illustrates a schematic block diagram of an apparatus for supplying one or more adjusted parameters for producing an upmix signal representation based on a downmix signal representation and object-related parametric information;</td>
<td>Figure 2</td><td>illustrates a schematic block diagram of an MPEG SAOC system, in accordance with one embodiment of the invention;</td>
<td>Figure 3</td><td>illustrates a schematic block diagram of an MPEG SAOC system, in accordance with another embodiment of the invention;</td>
<td>Figure 4</td><td>illustrates a schematic representation of the contribution of object signals to the downmix signal and to a mixed signal;</td>
<td>Figure 5a</td><td>illustrates a schematic block diagram of a mono downmix based SAOC to MPEG Surround transcoder in accordance with one embodiment of the invention;</td>
<td>Figure 5b</td><td>illustrates a schematic block diagram of a stereo downmix based SAOC to MPEG Surround transcoder, in accordance with one embodiment of the invention;</td>
<td>Figure 6</td><td>illustrates a block schematic representation of an audio signal encoder in accordance with one embodiment of the invention;</td>
<td>Figure 7</td><td>illustrates a schematic representation of an audio bit stream in accordance with one embodiment of the invention;</td>
<td>Figure 8</td><td>illustrates a schematic block diagram of a reference MPEG SAOC system;</td>
<td>Figure 9a</td><td>illustrates a schematic block diagram of a reference SAOC system using a separate decoder and mixer;</td>
<td>Figure 9b</td><td>illustrates a schematic block diagram of a reference SAOC system using an integrated decoder and mixer; Y</td>
<td>Figure 9c</td><td>illustrates a schematic block diagram of a reference SAOC system using an SAoC to MPEG transcoder.</td>
Detailed description of the realizations
1. Apparatus for supplying one or more adjusted parameters, according to Figure 1.
Next, an apparatus 100 for supplying one or more adjusted parameters for producing an upmix signal representation based on a downmix signal representation and object-related parametric information is described below with reference to FIG. 1. Figure 1 illustrates a schematic block diagram of that apparatus 100, which is configured to receive one or more input parameters 110. The input parameters 110 can be, for example, convenient display parameters. The apparatus 100 is also configured to supply, based on these, one or more adjusted parameters 120. The adjusted parameters may be, for example, adjusted display parameters. Apparatus 100 is further configured to receive parametric information related to objects 130. The object-related parametric information 130 may be, for example, inter-object level difference information and / or inter-object correlation information describing a plurality of objects. Apparatus 100 comprises a parameter adjuster 140, which is configured to receive one or more
ES 2 572 083 T3 input parameters 110, and to supply on the basis of these, one or more adjusted parameters 120. The parameter adjuster 140 is configured to produce said one or more adjusted parameters 120 depending on said one or more input parameters 110 and the parametric information related to objects 130, thereby reducing the distortion of an upmix signal that would be caused by the use of non-optimal parameters (for example, said one or more input parameters 110) in an apparatus for supplying an upmix signal representation based on a downmix signal representation and the parametric information related to objects 130, at least in the case of the parameters of input 110 that deviate from the optimal parameters by more than a predetermined deviation.
Consequently, the apparatus 100 receives said one or more input parameters 110 and produces, on the basis of these, said one or more adjusted parameters 120. By supplying said one or more adjusted parameters 120, apparatus 100 determines, explicitly or implicitly, whether unaltered use of one or more input parameters 110 would cause unacceptably high distortions if the one or more input parameters 110 were used to control a provision of an upmix signal based on a representation of a downmix signal and the parametric information related to the objects 130. Thus, adjusted parameters 120 are generally more suitable for adjusting said apparatus for provision of the upmix signal representation than said one or more input parameters 110, at least if said one or more input parameters 110 they are chosen in a non-advantageous way.
Accordingly, apparatus 100 generally improves the perceptual impression of an upmix signal representation, which is provided by an upmix signal representation producer depending on the one or more adjusted parameters 120. The use of the object-related parametric information to adjust the one or more input parameters, to derive the one or more adjusted parameters, has been shown to bring about good results, since the quality of the upmix signal representation is typically satisfactory if said one or more adjusted parameters 120 correspond to the parametric information related to objects 130, whereas parameters that violate the convenient relationship with the parametric information related to objects 130 generally result in audible distortions. The object-related parametric information may comprise, for example, downmix parameters, describing a contribution of object signals (from a plurality of audio objects) to said one or more downmix signals. The parametric information related to the objects may also comprise, on the other hand or in addition, parameters of differences in levels of the objects and / or correlation parameters between objects, which describe the characteristics of the object signals. It has been found that both the parameters that describe encoder-side processing of the object signals, and the parameters that describe characteristics of the audio objects themselves, can be considered useful information for use by the parameter adjuster 120. However, apparatus 100 may otherwise or additionally use parametric information related to objects 130
However, it should be noted that parameter setter 140 may use additional information to supply said one or more set parameters 120 based on said one or more input parameters 110. For example, parameter setter 140 may evaluate optionally downmix coefficients, one or more downmix signals, or any additional information to further enhance the provision of one or more adjusted parameters 120.
two. System according to figure 2
The MPEG SAOC 200 system of Figure 2 will now be described in detail.
To provide a good understanding of the MPEG SAOC 200 system, an overview of the desired system specifications and design considerations is presented. Next, a structural overview of the system is presented. Furthermore, a plurality of SAOC distortion metrics are described, and the application of these SAOC distortion metrics for limiting distortions. In addition, other extensions to System 200 are described.
2.1 System design considerations
As discussed above, parametric techniques for bit rate efficient transmission of audio storage scenes containing multiple audio objects are generally efficient, both in terms of bit rate and in terms of transmission. computational complexity. Other benefits for the user of such a system on the reception side include the freedom to choose a presentation configuration of their choice (mono, stereo, surround, virtualized headphone playback, and so on) and the user interactivity feature. - The presentation matrix, and thus the output scene, can be interactively adjusted and changed according to will, personal preferences or other criteria. For example, it is possible to place speakers of a group together in a spatial area to maximize discrimination from other remaining speakers. This interactivity is obtained by including a
ES 2 572 083 T3 set-top box user interface.
For each transmitted sound object, its relative level and (in the case of non-mono presentation) the presentation spatial position can be adjusted. This can take place in real time by changing the position of associated graphical user interface (GUI) sliders by the user (for example: object level = +5 dB, object position = -30 degrees). However, it has been found that due to the parametric approach based on downmix / mix separation, the subjective quality of the displayed audio output is dependent on the display parameter settings. Changes in the relative level of objects have been found to affect the final audio quality more than changes in the position of the spatial presentation ("re-panning"). It has also been found that extreme settings for relative parameters (eg +20 dB) can even lead to inadmissible output quality. While this is simply the result of violating some of the prescriptive assumptions underlying this scheme, it is still unacceptable for a commercial product to produce poor sound and clutter depending on user interface settings. Consequently, embodiments according to the invention, such as, for example, the system 200, address this problem of avoiding these unacceptable impairments regardless of the user interface configuration (user interface configuration to be considered as " input parameters ”).
Here are some details regarding approaches to avoid SAOC distortions. The approach to limiting SAOC distortions presented here is based on the following concepts:
• Important SAOC distortions appear due to inappropriate choices of presentation coefficients (which can be considered input parameters). This choice is typically made by the user interactively (eg, via a real-time graphical user interface (GUI) for interactive applications). Therefore, an additional processing step is introduced that modifies the presentation coefficients that are provided by the user (for example, limits them based on certain calculations) and uses these modified coefficients for the SAOC presentation engine. For example, presentation coefficients that were provided by the user can be considered as input parameters, and modified coefficients for the SAOC presentation engine can be considered modified parameters.
• To control for excessive degradation of the SAOC audio output produced, it is desirable to develop a computational measure of perceptual impairment (also referred to as a DM distortion measure). This measure of distortion has been found to meet certain criteria:
o The distortion measurement should be easily calculable from internal parameters of the SAOC decoding engine. For example, it is convenient that the extra calculation of filter banks is not necessary to obtain the distortion measure.
o The distortion measurement value must be correlated with the quality of the subjectively perceived sound (perceptual degradation), that is, it must be in line with the basic principles of psychoacoustics. For this purpose, the calculation of the distortion mean can preferably be performed in a frequency-selective manner, and is generally referred to as perceptual audio coding and processing.
It has been discovered that a number of SAOC distortion measures can be defined and calculated. However, it has been found that SAOC distortion measures must preferably consider certain basic factors in order to arrive at a correct evaluation of a presented SAOC quality and therefore (although not necessarily) have certain points in common.
• Down-mixing coefficients are considered. These determine the relative mixing fractions of each audio object within the one or more downmix signals. As background information, it should be noted that the SAOC distortion that occurs has been found to be dependent on the relationship between downmix and presentation coefficients. If the relative contribution of the object defined by the presentation coefficients is substantially different from the relative contribution of the object within the downmix, then the SAOC decoding engine (which uses the modified parameters) has to perform considerable signal tuning. downmix to make it the displayed output. This has been found to result in SAOC distortion.
• Presentation coefficients are considered. These determine the relative output power of each audio object to each of the one or more output signals displayed. As background information, it should be noted that it has been found that the SAOC distortion that occurs also depends on the mutual relationship of the powers of the objects. If an object at a certain point in time has a much higher power than other objects (and if the downmix coefficient of this object is not too small) then this object dominates the downmix and plays very favorably in the output signal presented. . In contrast, weak objects are only rendered very weakly in the downmix and therefore cannot be driven to high output levels without significant distortion.
ES 2 572 083 T3 • Considering the power of object / level (relative) of each object with respect to the other. This information is described, for example, in terms of SAOC Object Level Differences (OLD). As background information, it should be noted that the SAOC distortion that occurs has been found to be dependent on the properties of individual object signals as well. For example, amplifying an object of a tonal nature in the output presented at higher levels (while other objects may be more of the noise type) results in considerable perceived distortion.
• In addition to this, other information about the properties of the original object signals can be considered. These can be transmitted by the SAOC encoder as part of the SAOC supplementary information. For example, information about the tonality or noise of each object element can be transmitted as part of the supplementary SAOC information and used in order to limit distortions.
2.2 system overview
Based on the above considerations, an overview of the MPEG SAOC 200 system is now presented to provide a good understanding of the present invention. It should be noted that the SAOC 200 system according to Figure 2 is an extended version of the MPEG SAOC 800 system according to Figure 8, so the above explanation also applies. Furthermore, it should be noted that the MPEG SAOC 200 system can be modified according to the implementation alternatives 900, 930, 960 set forth in Figures 9a, 9b and 9c, where the object encoder corresponds to the SAOC encoder, where the interaction information with the user / user control information 822 corresponds to the presentation control information / presentation coefficient.
Additionally, the SAOC decoder of the MPEG system SAOC 100 can be replaced by the separate object decoder and mixer / presenter arrangement 920, by the integrated object decoder and mixer / presenter arrangement 930 or the SAOC to MPEG Surround transcoder 980.
Now taking Figure 2 as reference, it can be seen that the MPEG SAOC 200 system comprises an SAOC 210 encoder, which is configured to receive a plurality of object signals x1 to xn, associated with a plurality of objects numbered 1 to N. The SAOC encoder 210 is also configured to receive (or otherwise obtain) downmix coefficients d1 to dN. For example, the SAOC encoder 210 may obtain a series of downmix coefficients d1 to dN for each channel of the downmix signal 212 provided by the SAOC encoder 210. The SAOC 210 encoder may be configured, for example, to obtain a weighted combination of the object signals x1 to xn in order to obtain a downmix signal, where each of the object signals x1 to xn is weighted with its coefficient downmix associated d1 to dN. The SAOC encoder 210 is also configured to obtain object relationship information, which describes a relationship between different object signals. For example, the relationship information between objects may comprise object level difference information, for example, in the form of OLD parameters and correlation information between objects, for example, in the form of IOC parameters. Consequently the SAOC encoder 200 is then configured to supply one or more downmix signals 212, each of which comprises a weighted combination of one or more object signals, weighted according to a series of downmix parameters associated with the respective downmix signal (or one channel of the multi-channel downmix signal 212). The SAOC encoder 210 is also configured to provide supplemental information 214, where the supplemental information 214 comprises inter-object relationship information (eg, in the form of level difference parameters between objects and correlation parameters between objects). Supplementary information 214 also comprises downmix parameter information, for example, in the form of downmix gain parameters and downmix channel level difference parameters. The supplemental information 214 may further comprise optional supplemental object property information, which may represent the individual properties of the objects. Details regarding optional supplemental information on object properties are described below.
The MPEG SAOC 200 system further comprises a SAOC 220 decoder, which may comprise the functionality of the SAOC 820 decoder. Consequently, the SAOC 220 decoder receives the one or more downmix signals 212 and supplementary information 214 as well as the presentation coefficients. modified (“or adjusted”, or “real”) 222 and supplies, on the basis of these, one or more upmixed signals y-ι to yN.
The MPEG SAOC 200 system further comprises an apparatus 240 for supplying one or more modified (or "adjusted" or "real") parameters, ie the modified display coefficients 222, depending on one or more input parameters, ie the parameters inputs describing presentation control information or presentation coefficients 242. Apparatus 240 is configured to also receive at least a portion of supplemental information 214. For example, apparatus 240 is configured to receive parameters 214a that describe powers of objects (eg, powers of object signals x1 through xn).
ES 2 572 083 T3
For example, the parameters 214a may comprise the object level difference parameters (also referred to as OLD). Apparatus 240 also preferably receives parameters 214b from supplemental information 214 describing downmix coefficients. For example, parameters 214b describe the downmix coefficients d1 through dN. Optionally, apparatus 240 may further receive additional parameters 214c, which constitute supplemental information on properties of individual objects.
Apparatus 240 is generally configured to provide modified presentation coefficients 222 based on input presentation coefficients 242 (which may be received, for example, from a user interface, or may be calculated, for example, depending on the user input or provided as default information), thereby reducing the distortion of the upmix signal representation that would be caused by the use of non-optimal presentation parameters by the SAOC 220 decoder. In other words, the modified presentation coefficients 222 are a modified version of the input presentation coefficients 242, where the changes are made depending on parameters 214a, 214b, such that all audible distortions in the upmix channel signals y1 to yN (which form the upmix signal representation).
The apparatus 240 for supplying the one or more adjusted parameters 242 may comprise, for example, a display coefficient adjuster 250, which receives the input display coefficients 242 and supplies, on the basis of these, the modified display coefficients 222. For this purpose, the presentation coefficient adjuster 250 may receive a distortion measure 252, which describes the distortions that would be caused by the use of the input presentation coefficients 242. The distortion measure 252 may be provided, by the For example, by the distortion calculator 260 depending on the parameters 214a, 214b and the display coefficients 242.
However, the functionalities of the presentation coefficient adjuster 250 and the distortion calculator 260 can also be integrated into a single functional unit, such that the modified presentation coefficients 222 are provided without explicit calculation of a distortion measure 252. Rather, implicit mechanisms can be applied to reduce or limit the extent of distortion.
With regard to the functionality of the MPEG SAOC 200 system, it should be noted that the upmix signal representation, which is output in the form of upmix channel signals y1 to yN, is generated with good perceptual quality since distortions audible that would be caused by an incorrect choice of user interaction information / user control information 822 in reference system 800, they are avoided by modifying or adjusting the presentation coefficients. The modification or adjustment is performed by apparatus 240 in such a way that severe perceptual impression degradations are avoided, or in such a way that perceptual impression degradations are at least reduced as compared to a case in which the SAOC decoder 220 directly use (without modification or adjustment) the input presentation coefficients 242.
The functionality of the inventive concept is briefly summarized below. Given a distortion measure (MD), excessive distortion in the audio output can be avoided by calculating the measured distortion value corresponding to the given signals, and modifying the SAOC decoding algorithm (by limiting the 212 presentation coefficients used in reality) in such a way that the measured value of the distortion does not exceed a certain threshold. A system 200 in accordance with this concept and which has been explained in some detail above is illustrated in Figure 2.
With reference to system 200, the following observations can be made:
• The desired presentation coefficients 242 are entered by the user or other interface.
• Before being applied to the SAOC 220 decoding engine, the presentation coefficients 242 are modified by a presentation coefficient adjuster 250, which makes use of one or more calculated distortion measures 252, which are supplied by a distortion 260.
• The distortion calculator 260 evaluates the information (eg parameters 214a, 214b) from the supplementary information 214 (eg the relative power of the objects / OLDs, downmix coefficients, and - optionally - information on signal properties of object). Also, it is based on the desired presentation coefficient input 242.
In a preferred embodiment, apparatus 240 is configured to modify display coefficients based on a measure of distortion. Preferably, the presentation coefficients are frequency-selectively adjusted, using, for example, frequency-selective weighting.
ES 2 572 083 T3
The modification of the presentation coefficients can be based on this frame (for example, a current frame), or the presentation coefficients can be adjusted in time not only in the form of frame by frame, but also processed / controlled in time (eg time smoothing) where possibly different attack / decay constants can be applied, or in the case of a dynamic range compressor / limiter.
In some embodiments, the distortion measurement can be selective across all frequencies.
In some embodiments, the distortion measurement may consider one or more of the following characteristics.
• Power / energy / level of each object;
• Downstream mixing coefficients;
• Presentation coefficients; and / or • Additional complementary information on object properties, if applicable.
In some embodiments, the distortion measure can be calculated for each object and combined to arrive at a total distortion.
In some embodiments, additional supplemental information about properties of the objects 214c can be evaluated. Additional supplementary information on object properties 214c can be extracted in an enhanced SAOC encoder, for example, in the SAOC 210 encoder. Additional supplemental information on object properties can be included, for example, in an enhanced SAOC bit stream, which is described with reference to Figure 7. In addition, additional supplemental information on object properties can be used to limit distortions by an improved SAOC encoder.
In a special case, noise / tonality can be used as a property of the described object by the additional supplementary information about properties of the objects. In this case, the noise / tonality can be transmitted with a much coarser frequency resolution than other parameters of the object (eg OLD) to store in the side information. In an extreme case, the supplementary information on object properties with respect to noise / tonality can be transmitted with only one information per object (eg with broadband characteristics).
2.3 SAOC distortion metric
In the following, a plurality of different distortion measures are described, which can be obtained, for example, using the distortion calculator 260. Details regarding the application of these distortion measures for limiting the presentation coefficients are described later in section 2.4.
In other words, this section summarizes various measures of distortion. These can be used individually or they can be combined to form a more complex composite distortion metric, for example by adding weighted individual distortion metric values. It should be noted here that the terms "distortion measure" and "distortion metric" designate similar quantities and in most cases it is not necessary to distinguish between them.
Next, a plurality of distortion metrics are described that can be evaluated by the distortion calculator 260 and that can be used by the presentation coefficient adjuster 250 in order to obtain the modified presentation coefficients 222 based on the presentation coefficients. input 242.
2.3.1 Distortion measurement # 1
Next, a first distortion measure (also referred to as distortion measure No. 1) is described.
For conceptual simplicity, an N-1-1 SAOC system is considered (eg, a mono downmix signal (212) and a single upmix channel (signal)). N input audio objects are downmixed into a mono signal and presented at a mono output. As presented in Figure 8 the downmix coefficients are designated di ... dN and the display coefficients are denoted by η ... rN ... In the following formulas, the time indices have been omitted for simplicity. Similarly, the frequency indices have been omitted, noting that the equations relate to the subband signals. In some of the following equations, the lower case letters indicate coefficients or signals, and the upper case letters indicate the corresponding powers, which can be observed from the context of the equations. Also, it should be noted that signals are sometimes represented by corresponding coefficients in the time-frequency domain, rather than in the time domain.
Suppose that object No. m (auditory object index m) is an object of interest, for example the most
ES 2 572 083 T3 which is increased in its relative level and thus limits the overall sound quality. Then the ideal desired output signal (upmix channel signal) is given by yv = [<sup>x</sup>m-<sup>r</sup>m] + [Σ * -d / = 1; iFm (1)
In this case, the first term is the desired contribution of the object of interest to the output signal, while the second term indicates the contributions of all other objects ("interference").
In reality, however, due to the downmixing process, the output signal is given by
NN
Λ<sup>= ί</sup>· Σν4 = Μ <] + [Σε · Μ]
<img file="ES2572083T3_D0001.tif" />
(2) that is, the downmix signal is then scaled by a transcoding coefficient, t, corresponding to the matrix "m2" of a MPEG Surround decoder. Again, this can be divided into a first term (actual contribution of the object signal to the output signal) and a second term (actual "interference" by other object signals). In this case, the SAOC system (for example, the SAOC 220 decoder, and optionally also the apparatus 240) dynamically determines the transcoding coefficient, t, such that the power of the output signal presented is actually equal to the ideal signal strength:
N
<img file="ES2572083T3_D0002.tif" />
N
<img file="ES2572083T3_D0003.tif" />
(3)
A measure of distortion (MD) can be defined by calculating the relationship between the real power contribution of object No. m and its real power contribution:
<img file="ES2572083T3_D0004.tif" />
i-1
NN
N
In this case,<img file="ES2572083T3_D0005.tif" />'indicates the power of the signal ultimately presented, and (4)
N
<img file="ES2572083T3_D0006.tif" />
is the power of the downmix signal. Note that, in the actual implementation, the X, values can be directly replaced by the corresponding Object Level Difference values (OLDj that are transmitted as part of the supplementary information SAOC 214.
For a better interpretation of the dmi, its definition can be rephrased as follows:
<img file="ES2572083T3_D0007.tif" />
N
NN
<img file="ES2572083T3_D0008.tif" />
dm<sub>1</sub> (ni) =
<img file="ES2572083T3_D0009.tif" />
(4th)
In effect, this means that the distortion metric is the ratio of the relative power contribution of the 12
<img file="ES2572083T3_D0010.tif" />
ES 2 572 083 T3 objects in the ideally presented signal (output) versus the downmix signal (input). This is consistent with the finding that the SAOC scheme works best when it does not have to alter the relative powers of the objects by high factors.
Increasing the dmi values indicates the reduction in sound quality relative to sound object No. m. It has been found that the value of dmi remains constant if all display coefficients are scaled by a common factor, or if all downmix coefficients are scaled the same. It has also been found that increasing the presentation coefficient for object No. m (increasing its relative level) results in greater distortion. The dmi values can be interpreted as follows:
• A value of 1 indicates an ideal quality with respect to object No. m;
• dmi values above 1 indicate decreased quality;
• Values of dmi below 1 do not further increase the quality relative to object No. m.
Consequently, the general measure of sound quality (i.e. the quality of all objects) can be calculated as follows:
<sup>N</sup> Γ 1 \ v (m) maxp / W;
DM, = ^ ---------- "<sup>=1</sup> (5)
In this equation, w (m) indicates an object weighting factor No. m that is related to the significance and sensitivity of the specific object within the audio scene. For example, we could then choose w (m) depending on the power / volume of the object w (m) = (r<sub>m</sub><sup>2</sup> X<sub>m</sub>)<sup>to</sup> where a can typically be chosen as 0.25 to roughly emulate the growth of the psychoacoustic volume corresponding to this object. Additionally, w (m) could take into account tonality and masking phenomena. On the other hand, w (m) can be set to 1, which facilitates the calculation of DM-i.
2.3.2 Distortion Measure # 2
An alternative distortion measure can be constructed from equation (4) to form a perceptual measure in the style of a Noise to Mask Ratio (NMR), that is, calculate the ratio between noise / interference and the masking threshold:
(m) = p
<sup>L</sup> Noise _
Mask p - p reai___ -reai <sup>msr</sup>- Aw
?) X,
----- P ---- msr / X<sub>i</sub> ^ Σ ^ Χ- ^ Σ'ίΧ \ i-1 i-1 /
<img file="ES2572083T3_D0011.tif" />
(6)
In this equation, msr is the Mask to Signal Ratio of the total audio signal that depends on its tonality. Increasing dm values<sub>2</sub> indicates greater distortion relative to sound object No. m. Once again, the value of dm<sub>2</sub> remains constant if all display coefficients are scaled by a common factor, or if all downmix coefficients are scaled in the same way. The range of values of dm<sub>2</sub> can be interpreted as follows:
• A value of 0 indicates the ideal quality with respect to object No. m;
• Increase in dm values<sub>2</sub> above 1 indicates progressive audible impairments;
• The values of dm<sub>2</sub> below 1 indicate indistinguishable quality with respect to object No. m.
Consequently, a general measure of sound scene quality (that is, the quality for all objects) can be calculated as follows:
ES 2 572 083 T3
N max [¿/ m<sub>2</sub>(m), l]
ZW<sub>2</sub> = ------------ (7) m = l
Again, w (m) indicates an object weighting factor No. m that relates to the significance / level / volume of the specific object within the audio scene, which is typically selected in terms of w (m) = (r<sub>m</sub><sup>2</sup> Xm)<sup>to</sup> with a = 0.25.
The distortion measure in equation (6) calculates the distortion in terms of difference in powers (this corresponds to a "NMR with spectral difference" measurement). On the other hand, distortion can be calculated based on a waveform leading to the following measurement that includes an additional mixed product term:
Msr P mask<sub>tütai</sub> rlfy'X, X, -2.d, r. .
i-1 i-1 V i-1 i-1
TO (<sup>8</sup>) i-1 i-1
2.3.3 Distortion Measure # 3
A third distortion measure is presented that describes the coherence between the downmix signal and the presented signal. Higher coherence results in better subjective sound quality. Also, the mapping of the audio objects can be taken into account if the IOC data is present in the SAOC decoder.
From the SAOC parameters (for example, the parameters 214a, which may comprise the level difference parameters of the objects and the correlation parameters between objects) a model of the covariance of the objects can be determined
E = a / QLD<sup>t</sup> OLD IOC
To calculate the distortion measure, a Matrix M is integrated that contains the presentation and downward mixing coefficients (M can be interpreted as a presentation matrix corresponding to an N-1-2 SAOC system) ··· άΊ
A 'N)
<img file="ES2572083T3_D0012.tif" />
The covariance between the downmixed and presented signal C is then
C = Μ Ε · Μ * = Γ<sup>11</sup> l<sup>C</sup>21
<img file="ES2572083T3_D0013.tif" />
A measure of DM3 distortion is defined as
DM<sub>3</sub> (
= 1 - min
The DM3 values can be interpreted as follows:
• The values are in the range [0 .. 1] and indicate the coherence between the downmix signal and the
ES 2 572 083 T3 filed.
• A value of 0 indicates the ideal quality.
• An increase in DM3 values indicates a decrease in quality.
2.3.4 Distortion measurement # 4
2.3.4.1 overview
This approach proposes the use, as a distortion measure, of the averaged weighted relationship between the target presentation energy (UPMIX) and the optimal downmix energy (calculated from a given DMX downmix).
For details, reference is also made to Fig. 4, which illustrates a graphical representation of the downmix (DMX), the optimal downmix power (DMX_opt), and the target presentation power (UPMIX).
2.3.4.2 Nomenclature
<td>ch = {1,2, ..., Nch} dx = {1,2} ob = {1,2, ..., Nob} pb = {1,2, ..., Npb} rch.ob. pb = r (ch, ob, pb)</td><td>index corresponding to the upmix channels index corresponding to downmix channels index corresponding to audio objects index corresponding to parameter bands display matrix corresponding to channel ch, audio object ob, and parameter band pb</td>
<td>ddx. ob. pb = d (dx, ob, pb)</td><td>downmix matrix corresponding to downmix channel dx, audio object ob, and parameter band pb</td>
<td>wob.pb = w (ob.pb)</td><td>weighting factor representing the significance / level / volume of the audio object ob corresponding to the parameter band pb</td>
<td>NRGpb = NRG (bp)</td><td>absolute energy of the object of the audio object with the highest energy corresponding to the frequency band pb</td>
<td>OLDob.pb = OLD (ob.pb)</td><td>Object level difference, which describes the intensity differences between an audio object ob and the object with the highest energy in the corresponding frequency band pb</td>
IOCob.ab.pb = IOC (obi, obj, pb) inter-object correlation, which describes the correlation between two channels of audio objects.
2.3.4.3 Algorithm
The following briefly describes the steps of an algorithm to obtain the distortion measure # 4:
• Calculation of the relative upmix and downmix energies:
F<sup>1</sup> = OTD <sub>r</sub><sup>2</sup><sup>F</sup> ch, ob, pb <sup>OLD</sup>oh.pb 'ch, ob, pb <sup>d</sup>dx, ob, pb <sup>OL</sup>-Dob, pb ' <sup>d</sup>dx, ob
Normalization of the energies, such that —2 'ch, ob, pb r ______ — __' ch, ob, pb Nob
Σ r ch, ob, pb ob_1 <sup>N</sup>ob <sup>N</sup>ob
Vf<sup>2</sup> _ 1 yd<sup>2</sup> _ 1 ¿_ ^ ch.ob.pb <sup>1</sup> ZJ <sup>d</sup>dm, ob, pb <sup>1</sup> ob_1 and ob_1 <sub>d</sub> % 2 _ ^ dm, ob, pb dm, ob, pb N<sub>ob</sub><sup>d</sup>dm, ob, pb ob_1 •
d <sup>)</sup><sup>d</sup>ch, ob, pb d ^ p *<sup>)</sup>
Optimal Down Mix Construction <sup>d</sup>ch, ob, pb for each channel and upmix band:
<sup>-</sup>ch, ob, pb <sup>d</sup>1, ob, pb <sup>+ P</sup>ch, ob, pb <sup>d</sup>2, ob, pb
-^ <sup>0</sup> α, β
The constants of multiplication <sup>-</sup>ch, ob, pb, <sup>p</sup>ch, ob, pb are calculated by solving the overdefined system of equations || d <sup>2 (or</sup>p * <sup>)</sup> _ f, .. \\ ch.ob. pb 'ch, ob, pb linear to satisfy the following condition:
• Calculation of the distortion measurement:
ES 2 572 083 T3
<img file="ES2572083T3_D0014.tif" />
N<sub>ob</sub> N<sub>cll</sub> ob = l ch = lr
ch, ób, pb ^ ch.ob.pb
2.3.4.4 Distortion control
Distortion control is achieved by limiting one or more display coefficients that depend on the distortion measure DM4.
It will be appreciated that (i) the measure is only relevant for the case of stereo downmix, and (ii) can be reduced to DM1 in the case of No. dx = 1 and No. ch = 1.
2.3.4.5 Properties
The properties of the number 4 distortion measurement calculation concept are briefly summarized below. The concept • assumes ideal transcoding • can handle stereo downmixing and • allows generalization to a multichannel presentation.
2.3.5 Distortion measurement No. 5
An alternative calculation of the transcoding coefficient t is suggested. It can be interpreted in terms of extension of t and leads to the transcoding matrix T which is characterized by the incorporation of inter-object coherence (IOC) and at the same time extends to the current metric DM No. 1 and DM N. 2 to stereo downmix and multichannel upmix. The current implementation of the transcoding coefficient t considers the parity of the power of the actually presented output signal and the power of the ideal presented signal, i.e.
N ¿2 __ / = 1 _________ <sup>1</sup> ~ N
Σ ^<sup>χ</sup>.
i = l
The incorporation of the covariance matrix E produces a modified formulation for t, that is, the transcoding matrix T, which also takes into account the coherence between objects. The elements of E are calculated from the SAOC 214 parameters, such as e<sub>t</sub> = jOLDpLDpOC.,
The transcoding matrix represents the conversion of the downmix signal to the presented output signal such that TDx ~ Rx. This is obtained by minimizing the mean square error, to give
T = RED * (DED * and
With
H = RED * g
M
NN <sup>ν</sup>ν = ΣΣ<sup>ά</sup>α<sup>ά</sup>] η<sup>β</sup>ιη 1 = 1 m = l the measure of distortion in the style of although in this case for each mix combination
ES 2 572 083 T3 descending / presentation of the object m is given by r<sup>2</sup> v dm- [mnk) = —— '' d<sup>2</sup> h.
m, nk, n
Considering (<sup>w</sup>) separately for the left and right downmix channels leads to
two / 7 \ 6/7, ¿E, l 7/7 \ ^ 777./:^1.2 dm<sub>T</sub> (m, k] = —— dm „(m, k] = —— <sup>v /</sup> dh <sup>v</sup> 'dh <sup>OR</sup>m, V<sup>l</sup>k, ly <sup>OR</sup>m, d<sup>l</sup>k, 2
It can be assumed that the best of the two downmix / upmix paths is relevant to the quality of the output presented, so the measure corresponds to the minimum value, i.e. dm- (m, Λ) = min \ dm<sub>L</sub>, dm<sub>R</sub> ]
A general measure of all the output channels, designated by the index k, can be calculated as dm<sub>5</sub> (neither)
Y<sub>J</sub>dm<sub>5</sub>(m, k) r<sup>2</sup><sub>k</sub>X<sub>m</sub> k = \
<img file="ES2572083T3_D0015.tif" />
The general measure of all objects can be obtained according to N w (ni) max [dm<sub>5</sub> (ni), 1]
DM<sub>5</sub>
N ^ w (m) m = \ w (m) = \ r<sup>2</sup>XΊ with 2 / LJ <sub>C</sub>as before.
a T in the case of y dnu
A similar extension of t
2.3.6. distortion measurement # 6
A sixth measure of distortion is described below.
Suppose that e, (t) is the square of the Hilbert envelope of object signal No. i and P, the power of object signal No. i (both typically within a subband), then we can obtain a measure N of the tonality / noise similarity from a normalized calculation of the variance of the Hilbert envelope as
N
Pi
On the other hand, the signal power / variance with Hilbert envelope difference can also be used instead of the variance of the Hilbert envelope itself. In either case, the measure describes the strength of the envelope fluctuation over time.
This measure of tonality / noise similarity, N, can be determined with respect to both the ideally presented signal mix and the actual SAOC presented sound mix, and a distortion measure can be calculated from the difference between the two, for example:
ES 2 572 083 T3 where β is a parameter (for example β = 2).
2.3.7. Calculation of the energies of the images of the source signals corresponding to the reference scene and the scene presented by SAOC
To calculate the object energies of the source image in the reference scene and that presented by SAOC used for the distortion measurements, the transcoding matrix must be taken into account <sup>T</sup> corresponding to the scene presented by SAOC as done in "Distortion measurement 5", but also the correlation of the source signals of both the reference scene and the scene presented.
Comment: The notation of the signals in capital letters reflects in this case the matrix notation of the signals, not the energies of the signals as in the previous chapters
In the case of an arbitrary origin <sup>x</sup>m the signal parts of <sup>x</sup>m from all sources <sup>x</sup>i can be calculated as follows:
All source signals are divided <sup>x</sup>i in a signal part <sup>x</sup>i \\ m that is correlated to the object of interest <sup>x</sup>my part <sup>x</sup>i ± m that does not correlate with <sup>x</sup>m. This can be achieved by subspace projection of<sup>x</sup>m over all the signs <sup>x</sup>i, that is <sup>x</sup>i = <sup>x</sup>i \\ m <sup>+ x</sup>i ± m. The correlated part is given by
T <sub>x</sub> = <sup>x</sup>m \ <sup>x</sup>IM T <sup>i || m</sup> xm<sup>T</sup> xm <sup>x</sup>m
IOCim
----— x = gx
II || 2 m oi, mm Em
2.3.7.1 Calculation of <sup>P</sup>idedl, xm from source image <sup>Y</sup> x „in the reference scene <sup>Y</sup> .
Where Y = RX y <sup>X</sup> = <sup>X</sup>± m <sup>+ X</sup>\\ m, the image <sup>Y</sup>xm from source <sup>x</sup>m corresponding to all Y = RX channels presented can be calculated by <sup>RX</sup>\ | m where
X \\ m \ = ί <sub>Y </sub>xm can be calculated according to
Y = = ( <sup>x</sup>m <sup>r</sup>ci •<sup>x</sup>1 <sub>r</sub> ch2, xi <sub>x</sub><sup>T</sup><sup>x</sup> 1 \ \ m <sub>x</sub>T <sup>x</sup> 2 \ \ m <sub>x</sub>T < <sup>Λ</sup> N \ \ m J rr \ <sup>N</sup>ch, <sup>x</sup>1 <sup>N</sup>ch, <sup>x</sup>two λ
gx <sup>T</sup>
1, mm
T gx
2, mm <sup>g</sup>N, m ' <sub>x</sub>T <sup>x</sup>m J r
chi, x2 r
ch2, x2 λ (λ „„ T
<td>• r 'chi, xn</td><td>T gx or 1, mm</td>
<td>• r</td><td>T gx or 2, mm</td>
<td>O Fn <sub>x</sub><sup>N</sup>ch-1,<sup>x</sup>N</td><td rowspan="2">T Ϋ <sup>g</sup>N, m<sup>x</sup>mj</td>
<td>rr<sup>N</sup>ch,<sup>x</sup>n-1 <sup>N</sup>ch,<sup>x</sup>NJ</td>
Therefore, the energy
<img file="ES2572083T3_D0016.tif" />
<sup>x</sup>m <sub>Y </sub>from the source image \ in the reference scene will be:
ES 2 572 083 T3
<img file="ES2572083T3_D0017.tif" />
<img file="ES2572083T3_D0018.tif" />
<img file="ES2572083T3_D0019.tif" />
<img file="ES2572083T3_D0020.tif" />
<img file="ES2572083T3_D0021.tif" />
+ ·· ιι<sup>2</sup>
<img file="ES2572083T3_D0022.tif" />
P rea! · '- V and
2.3.7.2 Calculation of __________ of the Source Image <sup>Xm</sup> in the scene presented by SAOC:
p This can be done in the same way as for ideal, x<sub>m</sub> . where T is the transcoding matrix and
D the downmix matrix, Yx<sub>m</sub> for all channels of the presented scene is:
Y<sub>x</sub><sup>x</sup>m = T<sup>05</sup>DX ..
\\ mt Ί * 12
Using
Y<sub>x</sub><sup>x</sup>m
-yes
TO
<td></td><td>5 ^ 17 ^ 12 + a / TTA</td><td>V77Aw + AAaT</td><td>T p- γ ol, ffl m</td>
<td>+ Ai</td><td>a / ^ 21 ^ 12 + VA A2</td><td>+ A Aw</td><td>T cr and 0 2, mm</td>
<td>y \] íN<sub>c</sub>pdn + V ^ v. ,,<sup>2</sup> Ai</td><td><sup>2</sup> + VA * <sup>2</sup> A2</td><td></td><td>T \ §Ν, ηι ^ ηι and</td>
P. <sub>χ</sub> Y
Therefore, the energy <sup>ugly!</sup>'“Of the source image of the reference scene must be:
<img file="ES2572083T3_D0023.tif" />
<img file="ES2572083T3_D0024.tif" />
<img file="ES2572083T3_D0025.tif" />
<img file="ES2572083T3_D0026.tif" />
δΐ, Μ
<img file="ES2572083T3_D0027.tif" />
<img file="ES2572083T3_D0028.tif" />
<img file="ES2572083T3_D0029.tif" />
<img file="ES2572083T3_D0030.tif" />
+ 'Yes ^
<img file="ES2572083T3_D0031.tif" />
2.3.7.3. Calculation of the distortion measure
The measure of distortion in the style of can be calculated for each object <sup>m</sup> and output presentation channel k as follows:
dm ^ reai
<img file="ES2572083T3_D0032.tif" />
/^+- +
<img file="ES2572083T3_D0033.tif" />
<img file="ES2572083T3_D0034.tif" />
ES 2 572 083 T3 dm<sub>7</sub> (m) <sup>N</sup>Ch
Σ <sup>dm</sup>7 (m <sup>k</sup>) <sup>r</sup>mj<sup>x</sup>mk = 1
Nch
Σ <sup>r</sup>m, k<sup>and</sup>k, kk = 1
N
Σ w (m) max [dm<sub>7</sub> (m), 1]
DM 7 = <sup>m</sup>---- n -------- Σ <sup>w</sup> (<sup>m</sup>) m = 1 with <sup>w</sup> ( <sup>m</sup> ) = [<sup>r</sup>m <sup>X</sup>m] "as before.
2.3.8 Object signal properties
An example of object signal properties that can be used, for example, by apparatus 250 or disturbance reduction device 320 to obtain a distortion measure is described below.
In SAOC processing, various audio object signals are downmixed to obtain a downmix signal which is then used to generate the final presented output. If a tonal object signal is mixed with a second more noise-like object signal of equal signal strength, the result tends to be more noise-like. The same is true if the second object signal has a higher power. Only if the second object signal is substantially less powerful than the first does the result tend to be tonal. Similarly, the tonality / noise similarity of the displayed SAOC output signal is primarily determined by the tonality / noise similarity of the downmix signal regardless of the applied presentation coefficients. To obtain a favorable subjective output quality, also the tonality / noise similarity of the actually presented signal should be close to the tonality / noise similarity of the ideally presented signal. To use this concept in the distortion measurement, it is necessary to transmit the information about the similarity of tonality / noise of each object as part of the bit stream. The tonality / noise similarity N of the output ideally presented in the SAOC decoder can then be estimated as a function of the tonality / noise similarity of each Ni object and its object power Pi, that is
N = f (N1, P1, N2, P2, Na, P3, ...) and compare it to the tonality / noise similarity of the actually presented output signal to calculate a measure of distortion. For example, you can use the following function f ():
Σ N p,
<img file="ES2572083T3_D0035.tif" />
i that combines the values of similarity of tonality / noise of the object and the powers of the objects in a single output estimating the value of similarity of tonality / noise of the mixes of the signals. The parameter α can be chosen to optimize the precision of the calculation procedure for a measure of tonality / noise similarity (for example α = 2). A suitable distortion metric based on tonality / noise similarity is described in Section 2.3.6 as distortion measure # 6.
2.4 Distortion limitation schemes
2.4.1 Overview of distortion limitation schemes
Below is a brief overview of a plurality of distortion limitation schemes. As discussed above, the presentation coefficient adjuster 250 receives the input presentation coefficients 242 and produces, based on these, a modified presentation coefficient 222 for use by the sAOc decoder 220.
Different concepts can be distinguished for the provision of modified presentation coefficients, where the concepts can also be combined in some embodiments. According to the first concept, one or more limit values of the presentation parameters are obtained in a first step depending on one or more parameters of complementary information 214 (that is, that depends on the related parametric information 20
ES 2 572 083 T3 with object 214). Next, the actual presentation coefficients "(modified or adjusted)" 222 are obtained, which depend on the desired presentation parameter 242 and one or more limit values of the presentation parameters, so that the actual presentation parameters respect the defined limits. by the limit values of the presentation parameters. Consequently, said presentation parameters that exceed the limit values of the presentation parameters are adjusted (modified) to adapt to the limit values of the presentation parameters. This first concept is easy to implement, although it sometimes leads to slightly degraded user satisfaction, since the user's choice of the desired presentation parameters 242 is out of consideration if the desired presentation parameters 242 defined by the user exceed the values. display parameter limit.
According to the second concept, the parameter adjuster calculates a combination between the square of a desired display parameter and the square of an optimal display parameter, to obtain the actual display parameter. In this case, the parameter adjuster is configured to determine a contribution of the desired presentation parameter and the optimal presentation parameter to the combination of linear depending on a predetermined threshold parameter and a distortion metric (described above).
Furthermore, it can be distinguished whether the distortion measure (distortion metric) is calculated using the properties of relationships between objects and / or the properties of individual objects. In some embodiments, only the properties of the relationship between objects are evaluated, leaving the properties of the individual objects (that relate to only one object) out of consideration. In some additional embodiments, only the properties of the individual objects are taken into account, leaving the relationship properties between objects out of the question. However, in some embodiments, a combination of object relationship properties and properties of the individual objects are evaluated.
Based on the above considerations, and also based on the above description of the different distortion measures, a number of schemes for limiting distortion are now defined, summarized in the following subsections. These distortion limiting schemes can be applied by the display coefficient adjuster 250 to obtain the modified display coefficients depending on the input display coefficients 242.
2.4.2 Distortion limitation scheme No. 1
In subsection 2.3.1 a simple distortion measure is defined by calculating the relationship between the ideal power contribution of object No. m and its real power contribution (equation 4):
<img file="ES2572083T3_D0036.tif" />
(4)
In this equation, the only variables that are under the control of the SAOC presenter are the presentation coefficients that are used in the transcoding process. Therefore, if the distortion metric obtained does not exceed a certain threshold value, T, this imposes a condition on the corresponding display matrix coefficient:
N
<img file="ES2572083T3_D0037.tif" />
N '/ -, - Ση-χ,
N
<img file="ES2572083T3_D0038.tif" />
<img file="ES2572083T3_D0039.tif" />
N
To find a solution for all <sup>m</sup> you can establish a series of linear equations Ax = b where
ES 2 572 083 T3
<td></td><td></td><td>'0 Ί</td><td></td><td><sup>-c</sup>i</td><td>d1<sup>2</sup> X 2</td><td>ld<sup>2</sup>XN<sup>Ί</sup></td>
<td>rr<sup>2</sup> Ί '1</td><td></td><td> 0</td><td></td><td>d¡ X1</td><td><sup>—C</sup>2</td><td>L dX</td>
<td>r<sup>2</sup><sup>r</sup>2</td><td>b =</td><td></td><td>A =</td><td></td><td></td><td></td>
<td>r<sup>2 </sup>_<sup>F</sup>N _</td><td></td><td>1 or 1_______________________________________________________________________</td><td>Y</td><td>dN X 1</td><td>dN X2 1</td><td> ··· <sup>-C</sup>N eleven</td>
with c = 1 IΣ d<sup>2</sup> x - td<sup>2</sup> · X | m 'Τ' I iimm I<sup>T</sup> And i = 1)
The first N rows of A are obtained directly from equation (6.1.a). In addition, a constraint is added so that the energy of the new (limited) presentation coefficients is equal to the energy of the
P<sup>2</sup> user-specified coefficients. A solution corresponding to 'm (which is considered as limit values of the presentation parameters) is then obtained as follows:
x = (A<sup>T</sup>TO )<sup>-1</sup> TO<sup>T</sup>b
From this, a first distortion limitation scheme can be seen, as follows: instead of using the coefficients from the presentation matrix 242 as they are input to the SAOC decoder from the user interface, the presentation coefficient used effectively rm ', 222 for object No. m is modified / limited (for example, by presentation coefficient adjuster 240 frame-by-frame before being used for the SAOC decoding process:
a i ~ 2 \ r = minir r) m and m 'm J
Note that the limiting process depends on the energies of the individual objects in each specific frame. The approach is straightforward and has the following minor drawbacks:
• Does not take into account object volume or perceptual masking, and • Only captures the gain increase effects of a specific object, but does not capture effects by attenuating the object gains. This could be solved by also setting a lower limit for the dm value.
2.4.3 Limitation scheme No. 2
2.4.3.1 Overview of the limitation scheme
This section describes a limiting function that takes into account the following aspects:
• the distortion measure is constrained by a limiting threshold, • the derivation of the limited presentation matrix is based on the limiting function and its distance from the initial presentation matrix.
This limiting function (or limiting scheme) can be executed, for example, by display coefficient adjuster 250 in combination with distortion calculator 260.
The distortion measure is a function of the display matrix, whereby • an initial display matrix (described, for example, by input display coefficients 242) gives a measure of initial distortion, • the measure of optimal distortion produces an optimal presentation matrix, but the distance from this optimal presentation matrix to the initial presentation matrix may not be optimal, • the distortion measure is inversely proportional to linear distance of a display matrix from the initial display matrix,
ES 2 572 083 T3 • for a given threshold, the limited presentation matrix (described, for example, by the adjusted or modified presentation coefficients 222) is derived by interpolation (for example, linear interpolation) between the initial working point and the optimum.
Furthermore, it can be assumed that the power of the signal presented at each working point is approximately constant, so <sup>N</sup>ob
Σ r<sup>2</sup> X i = 1 <sup>N</sup>ob <sup>N</sup>ob ^ ςr x -Σ<sub>r</sub><sup>2</sup> x lim, ii / j opt, iii = 1 i = 1
Limiting scheme No. described below can be used.
in combination with different distortion measures, as shown
2.4.3.2 Limitation of distortion measurement No.
For each parameter band, the distortion measure <sup>dm</sup>i <sup>(m)</sup> corresponding to an object of interest <sup>m</sup> is defined as dm<sub>1</sub> (m) =
Nob r Σ dfX.
i = 1 ____________
Nob d<sup>2</sup> Σ r<sup>2</sup> X miii = 1
The optimal display matrix is produced <sup>dm</sup>i, opt (<sup>m</sup>) = <sup>1</sup> when it fits <sup>dm</sup>i<sup>(m)</sup> at its optimal value, that is r<sup>2</sup> opt, m <sup>N</sup>ob
Σ r<sup>2</sup> X<sub>t </sub>=<sup>d</sup>mn = i—
Σ d<sup>2</sup> Xi i = 1 <sub>r</sub><sup>1</sup>
Consequently, the optimal values of the presentation matrix opt, m can be obtained using a system <sub>r</sub> ' <sub>r</sub><sup>2</sup> of equations, where 'i is replaced by opt, i.
Within the predefined threshold <sup>T</sup> corresponding to <sup>dm</sup>i <sup>(m)</sup> the limited presentation matrix is given by
T 1/2 2 \ 2 r = --------- (r - r) + r lim, m dm (m) 'm opt, m J opt, m
2.4.3.3 Limitation of distortion measurement No. 2a
The distortion measure <sup>dm</sup>2nd <sup>(m)</sup> , which is sometimes also briefly designated as " <sup>dm</sup>2 <sup>(m)</sup> ", is defined as
Nob
Nob r<sup>2</sup> X mm
Nob (2 2 2 2 Ί - o - o
Κ Σ <sup>d</sup>iX - <sup>d</sup> Σ <sup>: X</sup> ]<sup>X</sup>- Σ rx, Σ dm., (M) = -——— N ----- go ----<sup>2nd</sup> Nob Nob msr Σ r<sup>2</sup>X Σ d<sup>2</sup>X lllli = 1 i = 1 dm Xm
Nob i = 1 i = 1 <sup>N</sup>ob i = 1 i = 1 msr in the case of the object m and each parameter band. With respect to certain parameter bands<sup>bp</sup> the relationship
ES 2 572 083 T3 mask to signal msr (pb) is a function of the power of the presented signal msr (pb) =
<td>'Nob ~</td><td></td><td>~ Nob ~</td>
<td>Σ χ, μ,</td><td> =</td><td>Σ t X,</td>
<td>_ i = 1 _</td><td>k = max (pb)</td><td>_ i = 1 _</td>
[M] k = max (pb) k = max (pb ') dm (mi 1 = 0
The optimal value of the distortion measure is zero, that is ^ Gaopty<sup>1</sup> J ". This corresponds to a prefect transcoding process that does not introduce any errors. Therefore, the optimal display matrix produces<sub>r</sub><sup>2</sup> = d <sup>1</sup> 'opt, mm
Σ r- X, i = 1 ____________
Nob
Σ di = 1 dm (mi) = T where “'24' / the matrix of the limited presentation matrix, which can be described by the modified presentation coefficients 222, becomes <sup>T 1</sup> / 2 2 \ 2 r = ---------- (r - r) + r lim, mif \\ m opt, mj opt, m <sup>dm</sup>2nd <sup>(m)</sup>
2.4.3.4 Limitation of distortion measurement No. 2b
The measure of distortion <sup>dm</sup>2b <sup>(m)</sup>, which is sometimes also briefly called <sup>dm</sup>2 <sup>(m)</sup>, it can also be used by the apparatus 240 to obtain the limited presentation matrix, which can be described by the modified presentation coefficients 222, which depend on the input presentation coefficients 242.
2.4.3.5 Limitation of distortion measurement No. 4 dm<sub>4</sub> (m) =
The measure of distortion <sup>dm</sup>4 <sup>(m)</sup> is defined as follows:
m ς d2 x <sup>1—</sup>-¾— dm Σ r- X, i = 1 in the case of the m object and each parameter band and its optimal value is optimal and limited presentation matrices result in <sup>dm</sup>4, opt (<sup>m</sup>) = <sup>0</sup> . Consequently the r<sup>2</sup> opt, m = d<sup>2</sup> m
<sup>N</sup>ob
Σ t X, i = 1 ____________
Nob
Σ dx, i = 1
T — 1 r = ------ lim, m ί / \ dm ^ (m) (r<sup>1</sup> —R<sup>1</sup> ) + r<sup>1</sup> m opt, m opt, m
Accordingly, the apparatus 240 can supply the modified presentation coefficients 222 depending on the input presentation coefficients 242 and also depending on the distortion measure 252, which
ES 2 572 083 T3 can be equal to the fourth distortion measure dm<sub>4</sub> (m).
2.4.4 Limitation scheme No. 3
With respect to formula (6.1.a), the limited presentation coefficient corresponding to object m can be calculated for the measurement of distortion No. 3 as follows. With the abbreviations
NN <sup>c</sup> = ΣΣ <sup>d</sup>Chi = 1 J = 1
N <sup>c</sup>two = Σ <sup>r</sup>i<sup>and</sup>im i = 1, i = m
NN <sup>c</sup> = Σ Σ rre i = l, i = mj = 1, j = m
N
c. = Σ of i my i = 1 and
NN <sup>C</sup>5 = Σ Σ <sup>rde</sup> i = 1, ij = 1 a quadratic equation is established
<img file="ES2572083T3_D0040.tif" />
• c, e mm
- c;) + rm 2 - ((i - t)<sup>2</sup> Cc - C4C5) + (i - t)<sup>2</sup> • cc <sub>-</sub> c <sup>c</sup>1<sup>c</sup>3 <sup>c</sup>5 a r<sup>2</sup> + b · m + c = 0 mm whose (positive) solution is „-b Ub<sup>2</sup> - 4ac r = ------------ m <sup>2nd</sup> (6.2.a)
Accordingly, the apparatus 240 may comprise limit values of the display parameters. <sup>r</sup>m, and may limit the adjusted (or modified) presentation coefficients 222 in accordance with said presentation parameter limit values.
2.4.5 Other optional enhancements
The above-described concept for limiting display coefficients 222, which is implemented, individually or in combination, by apparatus 240, can be further improved. For example, you can run a generalization for the presentation of M channels. For this purpose, the sum of the squares / power of the presentation coefficients can be used instead of a single presentation coefficient.
Also, you can run a generalization to a stereo downmix. For this purpose, a sum of the squares / powers of the downmix coefficients can be used instead of a single downmix coefficient.
In some embodiments, you can combine the distortion mix across the frequency to obtain a single one that is used to control degradation. On the other hand, it may be better (and easier) in some cases to control the distortion independently for each frequency band.
Different concepts can be applied to actually effect distortion control. For example, said one or more presentation coefficients can be limited. On the other hand, or in addition, a matrix coefficient m2 (for example from a decoding of MPEG Surround) can be limited. On the other hand, or in addition, a relative gain of the object can be limited.
3. Embodiment according to Fig. 3
Another embodiment of a SAOC decoder is described below with reference to Fig. 3. For ease of understanding, a brief explanation of the underlying considerations is first presented. The output of an “Audio Object Spatial Coding” (SAOC) system (such as that which has been standardized as ISO / IEC 23003-2) may show disturbances that depend on the properties of the audio object and the relationship between the matrix display and the downmix matrix. To explain this problem, we consider here the case where the downmix and display matrices have the same dimension without loss of generality. Corresponding considerations apply if the number of channels in the displayed and downmix scene are different.
ES 2 572 083 T3
It has been found that, in general, the risk of disturbances increases when the display matrix becomes significantly different from the downmix matrix. Different types of disturbances can be distinguished:
1. Presentation imperfections, that is, the “effective” presentation matrix differs from the desired presentation matrix that is input to the SAOC decoder (the attenuation achieved in effect or the gain of an object is different from that specified in the matrix of presentation). This is typically the effect of overlapping objects in certain parameter bands.
two. Unfavorable changes and possibly temporarily variations in the timbre of an object. This disturbance is especially serious when the "leakage" mentioned in 1. only occurs locally in the case of a single parameter band.
3. Disturbances, such as modulated object signals, musical tones, or modulated noise, caused by the processing of varying time and frequency signals in the SAOC decoder.
It has been found to be advantageous to minimize all kinds of disturbances.
A general strategy to address this problem and minimize disturbances is to employ time-frequency variant post-processing of the desired display matrix before sending it to the SAOC decoder. This strategy is shown in Fig. 3.
FIG. 3 illustrates a schematic block diagram of a SAOC 300 decoder arrangement. The SAOC 300 decoder may be referred to, in a nutshell, as an audio signal decoder. The audio signal decoder 300 comprises a SAOC 310 decoder core, which is configured to receive a downmix signal representation 312 and an SAOC 314 bit stream and to produce, based on these, a description 316 of a scene displayed, for example, as a representation of a plurality of upmix audio channels.
The audio signal decoder 300 also comprises a noise reduction device 320, which may be in the form of, for example, an apparatus to supply one or more adjusted parameters depending on one or more input parameters. The disturbance reduction device 320 is configured to receive information 322 about a desired display matrix. The information 322 may take the form of, for example, a plurality of desired display parameters, which may constitute the input parameters of the disturbance reduction device. The jitter reduction device 320 is further configured to receive the representation of the downmix signal 312 and the SAOC 314 bit stream, where the SAOC 314 bit stream may carry object-related parametric information. The jitter reduction device 320 is further configured to produce a modified display matrix 324 (eg, in the form of a plurality of adjusted display parameters) depending on information 322 about the desired display matrix.
Accordingly, the core of the SAOC 310 decoder may be configured to transmit the representation 316 of the displayed scene depending on the representation of the downmix signal 312, the SAOC bit stream 314, and the modified display matrix 324.
Below are some details regarding the functionality of the audio signal decoder. It has been found that, to assess the risk of disturbances due to the potentially limited separation capabilities of the SAOC system with respect to a desired display matrix, it is desirable to take into account both the downmix signal (described by the signal representation downmix 312) and the SAOC 314 bit stream. With this information, it is possible to try to mitigate these disturbances, for example, by modifying the presentation matrix. This is done by the disturbance reduction device 320. Advanced mitigation strategies take into account both the limitations (overlap) of the time and frequency selectivity of the SAOC system and the perceptual effects, that is, they should try to make the presented signal as similar as possible to the signal of desired output while having as few audible disturbances as possible.
A preferred approach to noise reduction, which is used in the audio signal decoder 300 shown in Fig. 3, is based on an overall distortion measure which is a weighted combination of the distortion measures that evaluate the different types. of disturbances listed above. These weights determine a suitable relationship between the different types of disturbances listed above. It should be noted that the weights for these types of shocks may depend on the application in which the SAOC system is used.
In other words, the disturbance reduction device 320 may be configured to obtain measurements of
ES 2 572 083 T3 distortion corresponding to a plurality of types of disturbances. For example, the disturbance reduction device 320 may apply some of the distortion measurements dm1 to dm6 described above. On the other hand, or in addition, the disturbance reducing device 320 may use other distortion measures that describe other types of disturbances, as described within this section. Likewise, the disturbance reducing device may be configured to obtain the modified display matrix 324 based on the desired display matrix 322 using one or more of the distortion limitation schemes that have already been described in previous paragraphs (for example , in sections 2.4.2, 2.4.3 and 2.4.4), or comparable disturbance limitation schemes.
Four. Audio signal transcoders according to Figs. 5a and 5b
4.1 Transcoder of audio signals according to Fig. 5a
It should be noted that the concepts described above can be applied to both an audio signal decoder and an audio signal transcoder. Taking as reference Figs. 2 and 3, the concept has been described in combination with audio decoders. The concept of the present invention in combination with audio signal transcoders is briefly described below.
With regard to this subject, it should be noted that the similarities between audio signal decoders and audio signal transcoders have already been described with reference to Figs. 9a, 9b and 9c, which is why the explanations presented with respect to Figs. 9a, 9b and 9c are applicable to the concept of the invention.
FIG. 5a illustrates a schematic block diagram of an audio signal transcoder 500 in combination with an MPEG Surround decoder 510. As can be seen, the audio signal transcoder 500, which may consist of a SAOC to MPEG Surround transcoder, is configured to receive an SAOC 520 bit stream and to produce, based on this, an MPEG bit stream. 522 envelope without affecting (or modifying) the representation of a 524 downmix signal. The audio signal transcoder 500 comprises a SAOC analyzer 530, which is configured to receive the SAOC 520 bit stream and to extract the desired SAOC parameters from the SAOC 530 bit stream. The audio signal transcoder 500 further comprises a scene display engine 540, which is configured to receive the SAOC parameters provided by the SAOC analyzer 530 and information from the display matrix 542, which can be considered actual display (matrix) information. , and which may be represented, for example, in the form of a plurality of adjusted (or modified) display parameters. The scene display engine 540 is configured to transmit the MPEG Surround 522 bitstream that depends on said SAOC parameters and the display matrix 542. For this purpose, the scene display engine 540 is configured to calculate the parameters of the MPEG Surround 522 bit streams, which are channel-related parameters (also referred to as parametric information). Accordingly, the scene rendering engine 540 is configured to transform (or "transcode") the parameters of the SAOC 520 bitstream, constituting object-related parametric information, into the parameters of the MPEG Surround bitstream, which which constitutes a parametric information related to the channels, which depends on the display matrix 542.
The audio signal transcoder 500 further comprises generating a display matrix 550, which is configured to receive information about a desired display matrix, for example, in the form of information 552 about a playback setting and information. 554 about positions of objects. On the other hand, the presentation matrix generation 550 may receive information about the presentation parameters (eg, inputs to the presentation matrix). The presentation matrix generator is also configured to receive the SAOC 520 bit stream (or at least a subset of parametric information related to the object represented by the SAOC 520 bit stream). The display matrix generator 550 is also configured to produce the actual (adjusted or modified) display matrix 542 based on the received information. To such an extent, the display matrix generator 550 can assume the functionality of either apparatus 100 or apparatus 240.
The MPEG Surround decoder 510 is typically configured to obtain a plurality of upmix channel signals based on the downmix signal information 524 and the MPEG Surround bit stream 522 provided by the scene display engine 540.
To summarize, the audio signal transcoder 500 is configured to produce the MPEG Surround 522 bitstream such that the MPEG Surround 522 bitstream results in the production of an upmix signal representation based on the downmix signal representation 524, where the upmix signal representation is actually provided by the MPEG Surround decoder 510. The display matrix generator 550 adjusts the display matrix 542 used by the scene display engine 540 so that the upmix signal representation generated by the MPEG Surround decoder 510 does not comprise unacceptable audible distortion.
ES 2 572 083 T3
4.2 Transcoder of audio signals ce according to Fig. 5b
Fig. 5b illustrates another arrangement of an audio signal transcoder 560 and an MPEG Surround decoder 510. It should be noted that the arrangement of Fig. 5b is very similar to the arrangement of Fig. 5a, which is why they are designated identical media and signals with identical reference numerals. The audio signal transcoder 560 differs from the audio signal transcoder 500 in that the audio signal transcoder 560 comprises a downmix transcoder 570, which is configured to receive the input downmix representation 524 and to produce a representation modified downmix 574, which is fed to MPEG Surround decoder 510. Modification of the representation of the downmix signal is done to provide more flexibility in defining the desired audio result. This is because the MPEG Surround 522 bitstream cannot represent certain mappings of the input signal from the MPEG Surround 510 decoder over the signals. channel produced as output by the MPEG Surround 510 decoder. Consequently, modifying the representation of the downmix signal using the downmix transcoder 570 can lead to increased flexibility.
Once again, the display matrix generator 550 can take over the functionality of either apparatus 100 or apparatus 240, thereby ensuring that audible distortions in the upmix signal representation provided by the MPEG Surround decoder 510 are kept low enough.
5. Audio Signal Encoder according to Fig. 6
An audio signal encoder 600 is described below with reference to FIG. 6, which illustrates a schematic block diagram of such an audio signal encoder. The audio signal encoder 600 is configured to receive a plurality of object signals 612a, 612N (also designated x1 to xn) and to produce, based on these, a downmix signal representation 614 and related parametric information object 616. The audio signal encoder 600 comprises a downmix device 620 configured to produce one or more downmix signals (constituting the representation of the downmix signal 614) depending on the downmix coefficients d1 to dN associated with the object signals, such that said one or more downmix signals comprise an overlay of a plurality of object signals. The audio signal encoder 600 further comprises a supplemental information provider 630, which is configured to produce supplemental object relationship information describing the level differences and correlation characteristics of two or more object signals 612a to 612N. Supplemental information provider 630 is also configured to provide supplemental information about individual objects that describes one or more individual properties of the individual object signals. The audio signal encoder 600 then produces the object-related parametric information 616 whereby the object-related parametric information comprises both the supplemental object relationship information and the supplemental information on individual objects.
Such object-related parametric information, describing both a relationship between object signals and the individual characteristics of individual object signals, has been found to enable the provision of a multi-channel audio signal in an audio signal decoder, as discussed above. discussed above. The supplemental object relationship information can be harnessed by the audio signal decoder to receive the object-related parametric information 616 to extract, at least approximately, individual object signals from the downmix signal representation. The supplementary information on individual objects, which is also included in the parametric information related to the object 614, can be used by the audio signal decoder to check if the upmixing process leads to too strong signal distortions, so it is necessary adjust the upmix parameters (for example, presentation parameters).
Preferably, the supplemental information provider 630 is configured to supply the supplemental information about individual objects, whereby the supplemental information about individual objects describes a tonality of the individual object signals. It has been found that tonality information can be used as a safe criterion for evaluating whether or not the upmixing process brings significant distortions.
It should also be noted that the audio signal encoder 600 can be supplemented by any of the features and functionalities described in this document with respect to audio signal encoders, and that the representation of the downmix signal 614 and object-related parametric information 616 can be provided by the audio signal encoder 600 in such a way as to comprise the characteristics described with respect to the audio signal decoder of the invention.
6. Audio bit stream according to Fig. 7
ES 2 572 083 T3
An embodiment according to the invention gives rise to an audio bit stream 700, a schematic representation of which is set forth in FIG. 7. The bit stream represents a plurality of object signals in encoded form.
Bit stream 700 comprises a downmix signal representation 710 representing one or more downmix signals, wherein at least one of the downmix signals comprises an overlay of a plurality of object signals. Audio bit stream 700 further comprises supplemental object relationship information 720 describing level differences and correlation characteristics of object signals. The audio bit stream also comprises ancillary information about individual objects 730 that describes one or more individual properties of the individual object signals (which form the basis for the representation of the downmix signal 710).
The complementary information of relation between objects and the information of individual objects can be considered, as a whole, as parametric complementary information related to the object.
According to the invention, the supplementary information on individual objects describes hues of the individual object signals.
Naturally, as the audio bitstream 700 it is generally provided by an audio signal encoder as described herein, and evaluated by an audio signal decoder, which is described herein. The audio bit stream may comprise the characteristics described with respect to the audio signal encoder and the audio signal decoder. Consequently, the audio bitstream 700 may be well suited for the production of a multi-channel audio signal using an audio signal decoder, as described herein.
7. Conclution
Embodiments according to the invention offer solutions to reduce or avoid the problem of distortions explained above, which originate from the fact that the original single object signals cannot be perfectly reconstructed from the few transmitted downmix signals. There are other simple solutions that apply to fix this problem:
• A simple strategy would be to limit the relative gain range of objects to eg +/- 12 dB. Although it is true that the configuration of large object gains can lead to audible impairments (for example: increasing the gain of an object by 20 dB while leaving the other object levels at 0 dB): this is not necessary, without embargo. For example, increasing the gain of all relative levels of the objects by the same factor produces a smooth output from the system.
• A more elaborate approach would be to consider the differences in the relative levels of objects. For the presentation of two audio objects, the difference in the relative levels of both objects provides a hook for possible impairments of the presented output. However, it is not clear how this idea is generalized to more than two presented audio objects.
In view of this situation, embodiments according to the present invention offer means to address this problem and thereby prevent an unfavorable user experience. Some embodiments in accordance with the present invention may lead to even more elaborate solutions than those described in the previous section.
Consequently, a good auditory impression can be obtained using the present invention, even in the case where a user provides inappropriate presentation parameters.
In general terms, embodiments according to the invention relate to an apparatus, a method or a computer program for encoding an audio signal or for decoding an encoded audio signal, or with an encoded audio signal (for example, in audio bitstream shape) as described above.
8. Implementation Alternatives
While certain aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a characteristic of a method step. Similarly, aspects described in the context of a method step also represent a description of a corresponding block or element or feature of a respective apparatus. Some or all of the steps in the method may be executed (or used) by a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some
In embodiments, one or more of the more important steps of the method can be performed by that apparatus.
The encoded audio signal or audio bitstream of the invention may be stored on a digital storage medium or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the internet.
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or software. The implementation can be executed using a digital storage medium, for example a floppy disk, a DVD, a Blue-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, with readable control signals electronically stored therein, which cooperate (or may cooperate) with a programmable computer system in order to execute the respective method. Therefore, the digital storage medium can be computer readable.
Some embodiments according to the invention comprise a data carrier that has electronically readable control signals, capable of cooperating with a programmable computer system for the execution of the methods described herein.
In general, embodiments of the present invention may be implemented as a computer program product with a program code, where the program code is operative to perform one of the methods by running the computer program on a computer. The program code can be stored, for example, on a machine-readable carrier.
Other embodiments comprise computer program for executing one of the methods described herein, stored on a machine-readable carrier.
In other words, an embodiment of the method of the invention therefore consists of a computer program consisting of a program code to perform one of the methods described herein by running the computer program on a computer.
Another embodiment of the method of the invention therefore consists of a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded therein, the computer program to execute one of the methods described herein.
Another embodiment of the method of the invention therefore consists of a data stream or a sequence of signals representing the computer program for executing one of the methods described herein. The data stream or signal sequence can be configured, for example, to be transferred over a data communication connection, for example via the Internet.
Another embodiment comprises a processing means, for example a computer or a programmable logic device, configured or adapted to execute one of the methods described herein.
Another embodiment comprises a computer in which the computer program has been installed to execute one of the methods described herein.
In some embodiments, a programmable logic device (eg, a field of programmable gate arrays) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field of programmable gate arrays may cooperate with a microprocessor to execute one of the methods described herein. In general, the methods are preferably executed by any hardware apparatus.
The above-described embodiments are merely illustrative of the principles of the present invention. Modifications and variations of the arrangements and details described in this document are understood to be apparent to those skilled in the art. Therefore, it is only intended to be limited to the scope of the following patent claims and not to the specific details presented by way of description and explanations of the present embodiments.
References
[BCC] C. Faller and F. Baumgarte, “Binaural Cue Coding - Part II: Schemes and applications”, IEEE Trans. on Speech and Audio Proc., vol. 11, No. 6, November 2003
[JSC] C. Faller, “Parametric Joint-Coding of Audio Sources”, 120th AES Convention, Paris, 2006, prepress 6752
[SAOC1] J. Herre, S. Disch, J. Hilpert, O. Hellmuth: “From SAC To SAOC - Recent Developments in Parametric
ES 2 572 083 T3
Coding of Spatial Audio ”, 22<sup>to</sup> AES UK Regional Conference, Cambridge, UK, April 2007
[SAOC2] J. Engdegard, B. Resch, C. Falch, O. Hellmuth, J. Hilpert, A. Holzer, L. Terentiev, J. Breebaart, J. Koppens, E. Schuijers and W. Oomen: “Spatial Object Audio Coding (SAOC) - The Upcoming MPEG Standard 5 on Parametric Object Based Audio Coding ”, 124th AES Convention, Amsterdam 2008, Prepress 7377
Contents32
92 members in 19 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 173456P | United States of America | – | |
| 17345609 | United States of America | P |
Members92
| Document | Office | Kind | |
|---|---|---|---|
| AU2008314029A1 | Australia | A1 | |
| AU2008314030A1 | Australia | A1 | |
| CA2701457A1 | Canada | A1 | |
| CA2702986A1 | Canada | A1 | |
| WO2009049895A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2009049896A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2009125313A1 | United States of America | A1 | |
| US2009125314A1 | United States of America | A1 | |
| TW200926143A | Taiwan Province of China | A | |
| TW200926147A | Taiwan Province of China | A | |
| EP2076900A1 | European Patent Office (EPO) | A1 | |
| EP2082396A1 | European Patent Office (EPO) | A1 | |
| WO2009049895A9 | World Intellectual Property Organization (WIPO) | A9 | |
| MX2010004138A | Mexico | A | |
| WO2009049896A8 | World Intellectual Property Organization (WIPO) | A8 | |
| KR20100063119A | Republic of Korea | A | |
| KR20100063120A | Republic of Korea | A | |
| MX2010004220A | Mexico | A | |
| CN101821799A | China | A | |
| CN101849257A | China | A | |
| CA2760515A1 | Canada | A1 | |
| CA2852503A1 | Canada | A1 | |
| WO2010125104A1 | World Intellectual Property Organization (WIPO) | A1 | |
| JP2011501544A | Japan | A | |
| JP2011501823A | Japan | A | |
| TW201104674A | Taiwan Province of China | A | |
| AU2008314030B2 | Australia | B2 | |
| AR076434A1 | Argentina | A1 | |
| WO2009049896A9 | World Intellectual Property Organization (WIPO) | A9 | |
| RU2010112889A | Russian Federation | A | |
| RU2010114875A | Russian Federation | A | |
| AU2010243635A1 | Australia | A1 | |
| SG175392A1 | Singapore | A1 | |
| KR20120004546A | Republic of Korea | A | |
| KR20120004547A | Republic of Korea | A | |
| AU2008314029B2 | Australia | B2 | |
| KR20120018778A | Republic of Korea | A | |
| EP2425427A1 | European Patent Office (EPO) | A1 | |
| US8155971B2 | United States of America | B2 | |
| RU2452043C2 | Russian Federation | C2 | |
| US2012143613A1 | United States of America | A1 | |
| MX2011011399A | Mexico | A | |
| CN102576532A | China | A | |
| US2012213376A1 | United States of America | A1 | |
| ZA201107895B | South Africa | B | |
| US8280744B2 | United States of America | B2 | |
| JP2012525600A | Japan | A | |
| CN101821799B | China | B | |
| RU2474887C2 | Russian Federation | C2 | |
| KR101244515B1 | Republic of Korea | B1 | |
| KR101244545B1 | Republic of Korea | B1 | |
| US8407060B2 | United States of America | B2 | |
| TWI395204B | Taiwan Province of China | B | |
| HK1173551A1 | Hong Kong, China | A1 | |
| RU2011145866A | Russian Federation | A | |
| US2013138446A1 | United States of America | A1 | |
| KR101290394B1 | Republic of Korea | B1 | |
| JP5260665B2 | Japan | B2 | |
| TWI406267B | Taiwan Province of China | B | |
| KR101303441B1 | Republic of Korea | B1 | |
| US8538766B2 | United States of America | B2 | |
| AU2010243635B2 | Australia | B2 | |
| US8731950B2 | United States of America | B2 | |
| JP5554830B2 | Japan | B2 | |
| US2014229187A1 | United States of America | A1 | |
| KR101431889B1 | Republic of Korea | B1 | |
| EP2425427B1 | European Patent Office (EPO) | B1 | |
| JP2014206747A | Japan | A | |
| ES2521715T3 | Spain | T3 | |
| TW201443885A | Taiwan Province of China | A | |
| EP2816555A1 | European Patent Office (EPO) | A1 | |
| PL2425427T3 | Poland | T3 | |
| CA2760515C | Canada | C | |
| CN102576532B | China | B | |
| HK1205340A1 | Hong Kong, China | A1 | |
| RU2573738C2 | Russian Federation | C2 | |
| BRPI0816557A2 | Brazil | A2 | |
| JP5883561B2 | Japan | B2 | |
| EP2816555B1 | European Patent Office (EPO) | B1 | |
| CN101849257B | China | B | |
| TWI529704B | Taiwan Province of China | B | |
| MY157169A | Malaysia | A | |
| CA2701457C | Canada | C | |
| ES2572083T3This record | Spain | T3 | |
| CA2702986C | Canada | C | |
| PL2816555T3 | Poland | T3 | |
| TWI560706B | Taiwan Province of China | B | |
| BRPI1007777A2 | Brazil | A2 | |
| CA2852503C | Canada | C | |
| US9786285B2 | United States of America | B2 | |
| BRPI0816556A2 | Brazil | A2 | |
| BRPI0816557B1 | Brazil | B1 |
Numbers
- Publication
- 2572083
- Application
- 14180279
Titles2
- Spanish
- Codificador de señal de audio, flujo de bits de audio, método y programa informático que utiliza información paramétrica relacionada con el objeto
- English
- Audio signal encoder, audio bit stream, method and computer program that uses parametric information related to the object
Classification
- CPC, 3
- G10L19/008
- G10L19/20
- G10L19/00
- IPC, 2
- G10L19 008
- G10L19 20