Audio encoder, audio decoder, methods and computer program using jointly encoded residual signals
Summary by NHIP
Joint Residual Audio Decoding
The audio decoder generates four channel signals from a jointly encoded representation of two residual signals and two downmix signals. It performs separate multi-channel bandwidth extensions on paired signals to reconstruct audio associated with a common horizontal plane or elevation.
Claim Score by NHIP
Abstract
An audio decoder for providing at least four audio channel signals on the basis of an encoded representation is configured to provide a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and of the second residual signal using a multi-channel decoding. The audio decoder is configured to provide a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding. The audio decoder is configured to provide a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding. An audio encoder is based on corresponding considerations.

Term
7.8 yearsleft in the term
Expires 11 July 2034.
- Priority
- Filed
- Granted
- Today
- Expires
3 claims: 2 independent, 1 dependent
- 1An audio decoder for providing at least four audio channel signals on the basis of an encoded representation, wherein the audio decoder is configured to provide a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and of the second residual signal using a multi-channel decoding;wherein the audio decoder is configured to provide a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding;and wherein the audio decoder is configured to provide a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding;wherein the audio decoder is configured to perform a first multi-channel bandwidth extension on the basis of the first audio channel signal and the third audio channel signal, and wherein the audio decoder is configured to perform a second multi-channel bandwidth extension on the basis of the second audio channel signal and the fourth audio channel signal;wherein the audio decoder is configured to perform the first multi-channel bandwidth extension in order to acquire two or more bandwidth-extended audio channel signals associated with a first common horizontal plane or a first common elevation of an audio scene on the basis of the first audio channel signal and the third audio channel signal and one or more bandwidth extension parameters, and wherein the audio decoder is configured to perform the second multi-channel bandwidth extension in order to acquire two or more bandwidth-extended audio channel signals associated with a second common horizontal plane or a second common elevation of the audio scene on the basis of the second audio channel signal and the fourth audio channel signal and one or more bandwidth extension parameters.
- 2Broadest claimClaim Score 22, narrow(NHIP)A method for providing at least four audio channel signals on the basis of an encoded representation, the method comprising:providing a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and the second residual signal using a multi-channel decoding;providing a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding;and providing a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding;wherein the method comprises performing a first multi-channel bandwidth extension on the basis of the first audio channel signal and the third audio channel signal, and wherein the method comprises performing a second multi-channel bandwidth extension on the basis of the second audio channel signal and the fourth audio channel signal;wherein the first multi-channel bandwidth extension is performed in order to acquire two or more bandwidth-extended audio channel signals associated with a first common horizontal plane or a first common elevation of an audio scene on the basis of the first audio channel signal and the third audio channel signal and one or more bandwidth extension parameters, and wherein the second multi-channel bandwidth extension is performed in order to acquire two or more bandwidth-extended audio channel signals associated with a second common horizontal plane or a second common elevation of the audio scene on the basis of the second audio channel signal and the fourth audio channel signal and one or more bandwidth extension parameters.
Independent claims2
251 paragraphs in 7 sections, as filed
CROSS-REFERENCE TO RELATED APPLICATIONS
This application is a continuation of copending application Ser. No. 15/004,661, filed Jan. 22, 2016, which is a continuation of copending International Application No. PCT/EP2014/064915, filed Jul. 11, 2014, which are incorporated herein by reference in their entirety, and additionally claims priority from European Applications Nos. EP 13177376.4, filed Jul. 22, 2013, and EP 13189305.9, filed Oct. 18, 2013, both of which are incorporated herein by reference in their entirety.
Embodiments according to the invention are related to an audio decoder for providing at least four audio channel signals on the basis of an encoded representation.
Further embodiments according to the invention are related to an audio encoder for providing an encoded representation on the basis of at least four audio channel signals.
Further embodiments according to the invention are related to a method for providing at least four audio channel signals on the basis of an encoded representation and to a method for providing an encoded representation on the basis of at least four audio channel signals.
Further embodiments according to the invention are related to a computer program for performing one of said methods.
Generally speaking, embodiments according the invention are related to a joint coding of n channels.
BACKGROUND OF THE INVENTION
In recent years, a demand for storage and transmission of audio contents has been steadily increasing. Moreover, the quality requirements for the storage and transmission of audio contents has also been increasing steadily. Accordingly, the concepts for the encoding and decoding of audio content have been enhanced. For example, the so-called “advanced audio coding” (AAC) has been developed, which is described, for example, in the International Standard ISO/IEC 13818-7:2003. Moreover, some spatial extensions have been created, like, for example, the so-called “MPEG Surround”-concept which is described, for example, in the international standard ISO/IEC 23003-1:2007. Moreover, additional improvements for the encoding and decoding of spatial information of audio signals are described in the international standard ISO/IEC 23003-2:2010, which relates to the so-called spatial audio object coding (SAOC).
Moreover, a flexible audio encoding/decoding concept, which provides the possibility to encode both general audio signals and speech signals with good coding efficiency and to handle multi-channel audio signals, is defined in the international standard ISO/IEC 23003-3:2012, which describes the so-called “unified speech and audio coding” (USAC) concept.
In MPEG USAC [1], joint stereo coding of two channels is performed using complex prediction, MPS 2-1-1 or unified stereo with band-limited or full-band residual signals.
MPEG surround [2] hierarchically combines OTT and TTT boxes for joint coding of multichannel audio with or without transmission of residual signals.
However, there is a desire to provide an even more advanced concept for an efficient encoding and decoding of three-dimensional audio scenes.
SUMMARY
An embodiment may have an audio decoder for providing at least four audio channel signals on the basis of an encoded representation, wherein the audio decoder is configured to provide a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and of the second residual signal using a multi-channel decoding which exploits similarities and/or dependencies between the residual signals; wherein the audio decoder is configured to provide a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding; and wherein the audio decoder is configured to provide a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding.
Another embodiment may have an audio encoder for providing an encoded representation on the basis of at least four audio channel signals, wherein the audio encoder is configured to jointly encode at least a first audio channel signal and a second audio channel signal using a residual-signal-assisted multi-channel encoding, to acquire a first downmix signal and a first residual signal; and wherein the audio encoder is configured to jointly encode at least a third audio channel signal and a fourth audio channel signal using a residual-signal-assisted multi-channel encoding, to acquire a second downmix signal and a second residual signal; and wherein the audio encoder is configured to jointly encode the first residual signal and the second residual signal using a multi-channel encoding which exploits similarities and/or dependencies between the residual signals, to acquire a jointly encoded representation of the residual signals.
According to another embodiment, a method for providing at least four audio channel signals on the basis of an encoded representation may have the steps of: providing a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and the second residual signal using a multi-channel decoding which exploits similarities and/or dependencies between the residual signals; providing a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding; and providing a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding.
According to another embodiment, a method for providing an encoded representation on the basis of at least four audio channel signals may have the steps of: jointly encoding at least a first audio channel signal and a second audio channel signal using a residual-signal assisted multi-channel encoding, to acquire a first downmix signal and a first residual signal; jointly encoding at least a third audio channel signal and a fourth audio channel signal using a residual-signal-assisted multi-channel encoding, to acquire a second downmix signal and a second residual signal; and jointly encoding the first residual signal and the second residual signal using a multi-channel encoding which exploits similarities and/or dependencies between the residual signals, to acquire an encoded representation of the residual signals.
Another embodiment may have a computer program for performing the method according to claim <b>37</b> when the computer program runs on a computer.
Another embodiment may have a computer program for performing the method according to claim <b>38</b> when the computer program runs on a computer.
Another embodiment may have an audio decoder for providing at least four audio channel signals on the basis of an encoded representation, wherein the audio decoder is configured to provide a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and of the second residual signal using a multi-channel decoding; wherein the audio decoder is configured to provide a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding; and wherein the audio decoder is configured to provide a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding; wherein the audio decoder is configured to perform a first multi-channel bandwidth extension on the basis of the first audio channel signal and the third audio channel signal, and wherein the audio decoder is configured to perform a second multi-channel bandwidth extension on the basis of the second audio channel signal and the fourth audio channel signal; wherein the audio decoder is configured to perform the first multi-channel bandwidth extension in order to acquire two or more bandwidth-extended audio channel signals associated with a first common horizontal plane or a first common elevation of an audio scene on the basis of the first audio channel signal and the third audio channel signal and one or more bandwidth extension parameters, and wherein the audio decoder is configured to perform the second multi-channel bandwidth extension in order to acquire two or more bandwidth-extended audio channel signals associated with a second common horizontal plane or a second common elevation of the audio scene on the basis of the second audio channel signal and the fourth audio channel signal and one or more bandwidth extension parameters.
According to another embodiment, a method for providing at least four audio channel signals on the basis of an encoded representation may have the steps of: providing a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and the second residual signal using a multi-channel decoding; providing a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding; and providing a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding; herein the method includes performing a first multi-channel bandwidth extension on the basis of the first audio channel signal and the third audio channel signal, and wherein the method includes performing a second multi-channel bandwidth extension on the basis of the second audio channel signal and the fourth audio channel signal; wherein the first multi-channel bandwidth extension is performed in order to acquire two or more bandwidth-extended audio channel signals associated with a first common horizontal plane or a first common elevation of an audio scene on the basis of the first audio channel signal and the third audio channel signal and one or more bandwidth extension parameters, and wherein the second multi-channel bandwidth extension is performed in order to acquire two or more bandwidth-extended audio channel signals associated with a second common horizontal plane or a second common elevation of the audio scene on the basis of the second audio channel signal and the fourth audio channel signal and one or more bandwidth extension parameters.
Another embodiment may have a computer program for performing the method according to claim <b>41</b> when the computer program runs on a computer.
An embodiment according to the invention creates an audio decoder for providing at least four audio channel signals on the basis of an encoded representation. The audio decoder is configured to provide a first residual signal and a second residual signal on the basis of a jointly encoded representation of the first residual signal and of the second residual signal using a multi-channel decoding. The audio decoder is also configured to provide a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding. The audio decoder is also configured to provide a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding.
This embodiment according to the invention is based on the finding that dependencies between four or even more audio channel signals can be exploited by deriving two residual signals, each of which is used to provide two or more audio channel signals using a residual-signal-assisted multi-channel decoding, from a jointly-encoded representation of the residual signals. In other words, it has been found there are typically some similarities of said residual signals, such that a bit rate for encoding said residual signals, which help to improve an audio quality when decoding the at least four audio channel signals, can be reduced by deriving the two residual signals from a jointly-encoded representation using a multi-channel decoding, which exploits similarities and/or dependencies between the residual signals.
In an advantageous embodiment, the audio decoder is configured to provide the first downmix signal and the second downmix signal on the basis of a jointly-encoded representation of the first downmix signal and the second downmix signal using a multi-channel decoding. Accordingly, a hierarchical structure of an audio decoder is created, wherein both the downmix signals and the residual signals, which are used in the residual-signal-assisted multi-channel decoding for providing the at least four audio channel signals, are derived using separate multi-channel decoding. Such a concept is particularly efficient, since the two downmix signals typically comprise similarities, which can be exploited in a multi-channel encoding/decoding, and since the two residual signals typically also comprise similarities, which can be exploited in a multi-channel encoding/decoding. Thus, a good coding efficiency can typically be obtained using this concept.
In an advantageous embodiment, the audio decoder is configured to provide the first residual signal and the second residual signal on the basis of the jointly-encoded representation of the first residual signal and of the second residual signal using a prediction-based multi-channel decoding. The usage of a prediction-based multi-channel decoding typically brings along a comparatively good reconstruction quality for the residual signals. This is, for example, advantageous if the first residual signal represents a left side of an audio scene and the second residual signal represents a right side of the audio scene, because the human hearing is typically comparatively sensitive for differences between the left and right sides of the audio scene.
In an advantageous embodiment, the audio decoder is configured to provide the first residual signal and the second residual signal on the basis of the jointly-encoded representation of the first residual signal and of the second residual signal using a residual-signal-assisted multi-channel decoding. It has been found that a particularly good quality of the first and second residual signal can be achieved if the first residual signal and the second residual signal are provided using a multi-channel decoding, which in turn receives a residual signal (and typically also a downmix signal, which combines the first residual signal and the second residual signal). Thus, there is a cascading of decoding stages, wherein two residual signals (the first residual signal, which is used for providing the first audio channel signal and the second audio channel signal, and the second residual signal, which is used for providing the third audio channel signal and the fourth audio channel signal), are provided on the basis of an input downmix signal and an input residual signal, wherein the latter may also be designated as a common residual signal) of the first residual signal and the second residual signal). Thus, the first residual signal and the second residual signal are actually “intermediate” residual signals, which are derived using a multi-channel decoding from a corresponding downmix signal and a corresponding “common” residual signal.
In an advantageous embodiment, the prediction-based multi-channel decoding is configured to evaluate a prediction parameter describing a contribution of a signal component, which is derived using a signal component of a previous frame, to the provision of the residual signals (i.e., the first residual signal and the second residual signal) of a current frame. Usage of such a prediction-based multi-channel decoding brings along a particularly good quality of the residual signals (first residual signal and second residual signal).
In an advantageous embodiment, the prediction-based multi-channel decoding is configured to obtain the first residual signal and the second residual signal on the basis of a (corresponding) downmix signal and a (corresponding) “common” residual signal, wherein the prediction-based multi-channel decoding is configured to apply the common residual signal with a first sign, to obtain the first residual signal, and to apply the common residual signal with a second sign, which is opposite to the first sign, to obtain the second residual signal. It has been found that such a prediction-based multi-channel decoding brings along a good efficiency for reconstructing the first residual signal and the second residual signal.
In an advantageous embodiment, the audio decoder is configured to provide the first residual signal and the second residual signal on the basis of the jointly-encoded representation of the first residual signal and of the second residual signal using a multi-channel decoding which is operative in the modified-discrete-cosine-transform domain (MDCT domain). It has been found that such a concept can be implemented in an efficient manner, since an audio decoding, which may be used to provide the jointly-encoded representation of the first residual signal and of the second residual signal, advantageously operates in the MDCT domain. Accordingly, intermediate transformations can be avoided by applying the multi-channel decoding for providing the first residual signal and the second residual signal in the MDCT domain.
In an advantageous embodiment, the audio decoder is configured to provide the first residual signal and the second residual signal on the basis of the jointly-encoded representation of the first residual signal and of the second residual signal using a USAC complex stereo prediction (for example, as mentioned in the above referenced USAC standard). It has been found that such a USAC complex stereo prediction brings along good results for the decoding of the first residual signal and of the second residual signal. Moreover, usage of the USAC complex stereo prediction for the decoding of the first residual signal and the second residual signal also allows for a simple implementation of the concept using decoding blocks which are already available in the unified-speech-and-audio coding (USAC). Accordingly, a unified-speech-and-audio coding decoder may be easily reconfigured to perform the decoding concept discussed here.
In an advantageous embodiment, the audio decoder is configured to provide the first audio channel signal and the second audio channel signal on the basis of the first downmix signal and the first residual signal using a parameter-based residual-signal-assisted multi-channel decoding. Similarly, the audio decoder is configured to provide the third audio channel signal and the fourth audio channel signal on the basis of the second downmix signal and the second residual signal using a parameter-based residual-signal-assisted multi-channel decoding. It has been found that such a multi-channel decoding is well-suited for the derivation of the audio channel signals on the basis of the first downmix signal, the first residual signal, the second downmix signal and the second residual signal. Moreover, it has been found that such a parameter-based residual-signal-assisted multi-channel decoding can be implemented with small effort using processing blocks which are already present in typical multi-channel audio decoders.
In an advantageous embodiment, the parameter-based residual-signal-assisted multi-channel decoding is configured to evaluate one or more parameters describing a desired correlation between two channels and/or level differences between two channels in order to provide the two or more audio channel signals on the basis of a respective downmix signal and a respective corresponding residual signal. It has been found that such a parameter-based residual-signal-assisted multi-channel decoding is well adapted for the second stage of a cascaded multi-channel decoding (wherein, advantageously, the first and second downmix signals and the first and second residual signals are provided using a prediction-based multi-channel decoding).
In an advantageous embodiment, the audio decoder is configured to provide the first audio channel signal and the second audio channel signal on the basis of the first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding which is operative in the QMF domain. Similarly, the audio decoder is advantageously configured to provide the third audio channel signal and the fourth audio channel signal on the basis of the second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding which is operative in the QMF domain. Accordingly, the second stage of the hierarchical multi-channel decoding is operative in the QMF domain, which is well adapted to typical post-processing, which is also often performed in the QMF domain, such that intermediate conversions may be avoided.
In an advantageous embodiment, the audio decoder is configured to provide the first audio channel signal and the second audio channel signal on the basis of the first downmix signal and the first residual signal using an MPEG Surround 2-1-2 decoding or a unified stereo decoding. Similarly, the audio decoder is advantageously configured to provide the third audio channel signal and the fourth audio channel signal on the basis of the second downmix signal and the second residual signal using a MPEG Surround 2-1-2 decoding or a unified stereo decoding. It has been found that such decoding concepts are particularly well-suited for the second stage of a hierarchical decoding.
In an advantageous embodiment, the first residual signal and the second residual signal are associated with different horizontal positions (or, equivalently, azimuth-positions) of an audio scene. It has been found that it is particularly advantageous to separate residual signals, which are associated with different horizontal positions (or azimuth positions), in a first stage of the hierarchical multi-channel processing because a particularly good hearing impression can be obtained if the perceptually important left/right separation is performed in a first stage of the hierarchical multi-channel decoding.
In an advantageous embodiment, the first audio channel signal and the second channel signal are associated with vertically neighboring positions of the audio scene (or, equivalently, with neighboring elevation positions of the audio scene). Also, the third audio channel signal and the fourth audio channel signal are advantageously associated with vertically neighboring positions of the audio scene (or, equivalently, with neighboring elevation positions of the audio scene). It has been found that good decoding results can be achieved if the separation between upper and lower signals is performed in a second stage of the hierarchical audio decoding (which typically comprises a somewhat smaller separation accuracy than the first stage), since the human auditory system is less sensitive with respect to a vertical position of an audio source when compared to a horizontal position of the audio source.
In an advantageous embodiment, the first audio channel signal and the second audio channel signal are associated with a first horizontal position of an audio scene (or, equivalently, azimuth position), and the third audio channel signal and the fourth audio channel signal are associated with a second horizontal position of the audio scene (or, equivalently, azimuth position), which is different from the first horizontal position (or, equivalently, azimuth position).
Advantageously, the first residual signal is associated with a left side of an audio scene, and the second residual signal is associated with a right side of the audio scene. Accordingly, the left-right separation is performed in a first stage of the hierarchical audio decoding.
In an advantageous embodiment, the first audio channel signal and the second audio channel signal are associated with the left side of the audio scene, and the third audio channel signal and the fourth audio channel signal are associated with a right side of the audio scene.
In another advantageous embodiment, the first audio channel signal is associated with a lower left side of the audio scene, the second audio channel signal is associated with an upper left side of the audio scene, the third audio channel signal is associated with a lower right side of the audio scene, and the fourth audio channel signal is associated with an upper right side of the audio scene. Such an association of the audio channel signals brings along particularly good coding results.
In an advantageous embodiment, the audio decoder is configured to provide the first downmix signal and the second downmix signal on the basis of a jointly-encoded representation of the first downmix signal and the second downmix signal using a multi-channel decoding, wherein the first downmix signal is associated with the left side of an audio scene and the second downmix signal is associated with the right side of the audio scene. It has been found that the downmix signals can also be encoded with good coding efficiency using a multi-channel coding, even if the downmix signals are associated with different sides of the audio scene.
In an advantageous embodiment, the audio decoder is configured to provide the first downmix signal and the second downmix signal on the basis of the jointly-encoded representation of the first downmix signal and of the second downmix signal using a prediction-based multi-channel decoding or even using a residual-signal-assisted prediction-based multi-channel decoding. It has been found that the usage of such multi-channel decoding concepts provides for a particularly good decoding result. Also, existing decoding functions can be reused in some audio decoders.
In an advantageous embodiment, the audio decoder is configured to perform a first multi-channel bandwidth extension on the basis of the first audio channel signal and the third audio channel signal. Also, the audio decoder may be configured to perform a second (typically separate) multi-channel bandwidth extension on the basis of the second audio channel signal and the fourth audio channel signal. It has been found that it is advantageous to perform a possible bandwidth extension on the basis of two audio channel signals which are associated with different sides of an audio scene (wherein different residual signals are typically associated with different sides of the audio scene).
In an advantageous embodiment, the audio decoder is configured to perform the first multi-channel bandwidth extension in order to obtain two or more bandwidth-extended audio channel signals associated with a first common horizontal plane (or, equivalently, with a first common elevation) of an audio scene on the basis of the first audio channel signal and the third audio channel signal and one or more bandwidth extension parameters. Moreover, the audio decoder is advantageously configured to perform the second multi-channel bandwidth extension in order to obtain two or more bandwidth-extended audio channel signals associated with a second common horizontal plane (or, equivalently, a second common elevation) of the audio scene on the basis of the second audio channel signal and the fourth audio channel signal and one or more bandwidth extension parameters. It has been found that such a decoding scheme results in good audio quality, since the multi-channel bandwidth extension can consider stereo characteristics, which are important for the hearing impression, in such an arrangement.
In an advantageous embodiment, the jointly-encoded representation of the first residual signal and of the second residual signal comprises a channel pair element comprising a downmix signal of the first and second residual signal and a common residual signal of the first and second residual signal. It has been found that the encoding of the downmix signal of the first and second residual signal and of the common residual signal of the first and second residual signal using a channel pair element is advantageous since the downmix signal of the first and second residual signal and the common residual signal of the first and second residual signal typically share a number of characteristics. Accordingly, the usage of a channel pair element typically reduces a signaling overhead and consequently allows for an efficient encoding.
In another advantageous embodiment, the audio decoder is configured to provide the first downmix signal and the second downmix signal on the basis of a jointly-encoded representation of the first downmix signal and the second downmix signal using a multi-channel decoding, wherein the jointly-encoded representation of the first downmix signal and of the second downmix signal comprises a channel pair element. the channel pair element comprising a downmix signal of the first and second downmix signal and a common residual signal of the first and second downmix signal. This embodiment is based on the same considerations as the embodiment described before.
Another embodiment according to the invention creates an audio encoder for providing an encoded representation on the basis of at least four audio channel signals. The audio encoder is configured to jointly encode at least a first audio channel signal and a second audio channel signal using a residual-signal-assisted multi-channel encoding, to obtain a first downmix signal and a first residual signal. The audio encoder is configured to jointly encode at least a third audio channel signal and a fourth audio channel signal using a residual-signal-assisted multi-channel encoding, to obtain a second downmix signal and a second residual signal. Moreover, the audio encoder is configured to jointly encode the first residual signal and the second residual signal using a multi-channel encoding, to obtain a jointly-encoded representation of the residual signals. This audio encoder is based on the same considerations as the above-described audio decoder.
Moreover, optional improvements of this audio encoder, and advantageous configurations of the audio encoder, are substantially in parallel with improvements and advantageous configurations of the audio decoder discussed above. Accordingly, reference is made to the above discussion.
Another embodiment according to the invention creates a method for providing at least four audio channel signals on the basis of an encoded representation, which substantially performs the functionality of the audio encoder described above, and which can be supplemented by any of the features and functionalities discussed above.
Another embodiment according to the invention creates a method for providing an encoded representation on the basis of at least four audio channel signals, which substantially fulfills the functionality of the audio decoder described above.
Another embodiment according to the invention creates a computer program for performing the methods mentioned above.
BRIEF DESCRIPTION OF THE DRAWINGS
Embodiments of the present invention will be detailed subsequently referring to the appended drawings, in which:
<figref idref="DRAWINGS">FIG. 1</figref> shows a block schematic diagram of an audio encoder, according to an embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 2</figref> shows a block schematic diagram of an audio decoder, according to an embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 3</figref> shows a block schematic diagram of an audio decoder, according to another embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 4</figref> shows a block schematic diagram of an audio encoder, according to an embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 5</figref> shows a block schematic diagram of an audio decoder, according to an embodiment of the present invention;
<figref idref="DRAWINGS">FIGS. 6A and 6B</figref> show a block schematic diagram of an audio decoder, according to another embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 7</figref> shows a flowchart of a method for providing an encoded representation on the basis of at least four audio channel signals, according to an embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 8</figref> shows a flowchart of a method for providing at least four audio channel signals on the basis of an encoded representation, according to an embodiment of the invention;
<figref idref="DRAWINGS">FIG. 9</figref> shows as flowchart of a method for providing an encoded representation on the basis of at least four audio channel signals, according to an embodiment of the invention; and
<figref idref="DRAWINGS">FIG. 10</figref> shows a flowchart of a method for providing at least four audio channel signals on the basis of an encoded representation, according to an embodiment of the invention;
<figref idref="DRAWINGS">FIG. 11</figref> shows a block schematic diagram of an audio encoder, according to an embodiment of the invention;
<figref idref="DRAWINGS">FIG. 12</figref> shows a block schematic diagram of an audio encoder, according to another embodiment of the invention;
<figref idref="DRAWINGS">FIG. 13</figref> shows a block schematic diagram of an audio decoder, according to an embodiment of the invention;
<figref idref="DRAWINGS">FIG. 14<i>a </i></figref>shows a syntax representation of a bitstream, which can be used with the audio encoder according to <figref idref="DRAWINGS">FIG. 13</figref>;
<figref idref="DRAWINGS">FIG. 14<i>b </i></figref>shows a table representation of different values of the parameter qceIndex;
<figref idref="DRAWINGS">FIG. 15</figref> shows a block schematic diagram of a 3D audio encoder in which the concepts according to the present invention can be used;
<figref idref="DRAWINGS">FIG. 16</figref> shows a block schematic diagram of a 3D audio decoder in which the concepts according to the present invention can be used; and
<figref idref="DRAWINGS">FIG. 17</figref> shows a block schematic diagram of a format converter.
<figref idref="DRAWINGS">FIG. 18</figref> shows a graphical representation of a topological structure of a Quad Channel Element (QCE), according to an embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 19</figref> shows a block schematic diagram of an audio decoder, according to an embodiment of the present invention;
<figref idref="DRAWINGS">FIG. 20</figref> shows a detailed block schematic diagram of a QCE Decoder, according to an embodiment of the present invention; and
<figref idref="DRAWINGS">FIG. 21</figref> shows a detailed block schematic diagram of a Quad Channel Encoder, according to an embodiment of the present invention.
DETAILED DESCRIPTION OF THE INVENTION
1. Audio Encoder According to FIG.
1
<figref idref="DRAWINGS">FIG. 1</figref> shows a block schematic diagram of an audio encoder, which is designated in its entirety with <b>100</b>. The audio encoder <b>100</b> is configured to provide an encoded representation on the basis of at least four audio channel signals. The audio encoder <b>100</b> is configured to receive a first audio channel signal <b>110</b>, a second audio channel signal <b>112</b>, a third audio channel signal <b>114</b> and a fourth audio channel signal <b>116</b>. Moreover, the audio encoder <b>100</b> is configured to provide an encoded representation of a first downmix signal <b>120</b> and of a second downmix signal <b>122</b>, as well as a jointly-encoded representation <b>130</b> of residual signals. The audio encoder <b>100</b> comprises a residual-signal-assisted multi-channel encoder <b>140</b>, which is configured to jointly-encode the first audio channel signal <b>110</b> and the second audio channel signal <b>112</b> using a residual-signal-assisted multi-channel encoding, to obtain the first downmix signal <b>120</b> and a first residual signal <b>142</b>. The audio signal encoder <b>100</b> also comprises a residual-signal-assisted multi-channel encoder <b>150</b>, which is configured to jointly-encode at least the third audio channel signal <b>114</b> and the fourth audio channel signal <b>116</b> using a residual-signal-assisted multi-channel encoding, to obtain the second downmix signal <b>122</b> and a second residual signal <b>152</b>. The audio decoder <b>100</b> also comprises a multi-channel encoder <b>160</b>, which is configured to jointly encode the first residual signal <b>142</b> and the second residual signal <b>152</b> using a multi-channel encoding, to obtain the jointly encoded representation <b>130</b> of the residual signals <b>142</b>, <b>152</b>.
Regarding the functionality of the audio encoder <b>100</b>, it should be noted that the audio encoder <b>100</b> performs a hierarchical encoding, wherein the first audio channel signal <b>110</b> and the second audio channel signal <b>112</b> are jointly-encoded using the residual-signal-assisted multi-channel encoding <b>140</b>, wherein both the first downmix signal <b>120</b> and the first residual signal <b>142</b> are provided. The first residual signal <b>142</b> may, for example, describe differences between the first audio channel signal <b>110</b> and the second audio channel signal <b>112</b>, and/or may describe some or any signal features which cannot be represented by the first downmix signal <b>120</b> and optional parameters, which may be provided by the residual-signal-assisted multi-channel encoder <b>140</b>. In other words, the first residual signal <b>142</b> may be a residual signal which allows for a refinement of a decoding result which may be obtained on the basis of the first downmix signal <b>120</b> and any possible parameters which may be provided by the residual-signal-assisted multi-channel encoder <b>140</b>. For example, the first residual signal <b>142</b> may allow at least for a partial waveform reconstruction of the first audio channel signal <b>110</b> and of the second audio channel signal <b>112</b> at the side of an audio decoder when compared to a mere reconstruction of high-level signal characteristics (like, for example, correlation characteristics, covariance characteristics, level difference characteristics, and the like). Similarly, the residual-signal-assisted multi-channel encoder <b>150</b> provides both the second downmix signal <b>122</b> and the second residual signal <b>152</b> on the basis of the third audio channel signal <b>114</b> and the fourth audio channel signal <b>116</b>, such that the second residual signal allows for a refinement of a signal reconstruction of the third audio channel signal <b>114</b> and of the fourth audio channel signal <b>116</b> at the side of an audio decoder. The second residual signal <b>152</b> may consequently serve the same functionality as the first residual signal <b>142</b>. However, if the audio channel signals <b>110</b>, <b>112</b>, <b>114</b>, <b>116</b> comprise some correlation, the first residual signal <b>142</b> and the second residual signal <b>152</b> are typically also correlated to some degree. Accordingly, the joint encoding of the first residual signal <b>142</b> and of the second residual signal <b>152</b> using the multi-channel encoder <b>160</b> typically comprises a high efficiency since a multi-channel encoding of correlated signals typically reduces the bitrate by exploiting the dependencies. Consequently, the first residual signal <b>142</b> and the second residual signal <b>152</b> can be encoded with good precision while keeping the bitrate of the jointly-encoded representation <b>130</b> of the residual signals reasonably small.
To summarize, the embodiment according to <figref idref="DRAWINGS">FIG. 1</figref> provides a hierarchical multi-channel encoding, wherein a good reproduction quality can be achieved by using the residual-signal-assisted multi-channel encoders <b>140</b>, <b>150</b>, and wherein a bitrate demand can be kept moderate by jointly-encoding a first residual signal <b>142</b> and a second residual signal <b>152</b>.
Further optional improvement of the audio encoder <b>100</b> is possible. Some of these improvements will be described taking reference to <figref idref="DRAWINGS">FIGS. 4, 11 and 12</figref>. However, it should be noted that the audio encoder <b>100</b> can also be adapted in parallel with the audio decoders described herein, wherein the functionality of the audio encoder is typically inverse to the functionality of the audio decoder.
2. Audio Decoder According to FIG.
2
<figref idref="DRAWINGS">FIG. 2</figref> shows a block schematic diagram of an audio decoder, which is designated in its entirety with <b>200</b>.
The audio decoder <b>200</b> is configured to receive an encoded representation which comprises a jointly-encoded representation <b>210</b> of a first residual signal and a second residual signal. The audio decoder <b>200</b> also receives a representation of a first downmix signal <b>212</b> and of a second downmix signal <b>214</b>. The audio decoder <b>200</b> is configured to provide a first audio channel signal <b>220</b>, a second audio channel signal <b>222</b>, a third audio channel signal <b>224</b> and a fourth audio channel signal <b>226</b>.
The audio decoder <b>200</b> comprises a multi-channel decoder <b>230</b>, which is configured to provide a first residual signal <b>232</b> and a second residual signal <b>234</b> on the basis of the jointly-encoded representation <b>210</b> of the first residual signal <b>232</b> and of the second residual signal <b>234</b>. The audio decoder <b>200</b> also comprises a (first) residual-signal-assisted multi-channel decoder <b>240</b> which is configured to provide the first audio channel signal <b>220</b> and the second audio channel signal <b>222</b> on the basis of the first downmix signal <b>212</b> and the first residual signal <b>232</b> using a multi-channel decoding. The audio decoder <b>200</b> also comprises a (second) residual-signal-assisted multi-channel decoder <b>250</b>, which is configured to provide the third audio channel signal <b>224</b> and the fourth audio channel signal <b>226</b> on the basis of the second downmix signal <b>214</b> and the second residual signal <b>234</b>.
Regarding the functionality of the audio decoder <b>200</b>, it should be noted that the audio signal decoder <b>200</b> provides the first audio channel signal <b>220</b> and the second audio channel signal <b>222</b> on the basis of a (first) common residual-signal-assisted multi-channel decoding <b>240</b>, wherein the decoding quality of the multi-channel decoding is increased by the first residual signal <b>232</b> (when compared to a non-residual-signal-assisted decoding). In other words, the first downmix signal <b>212</b> provides a “coarse” information about the first audio channel signal <b>220</b> and the second audio channel signal <b>222</b>, wherein, for example, differences between the first audio channel signal <b>220</b> and the second audio channel signal <b>222</b> may be described by (optional) parameters, which may be received by the residual-signal-assisted multi-channel decoder <b>240</b> and by the first residual signal <b>232</b>. Consequently, the first residual signal <b>232</b> may, for example, allow for a partial waveform reconstruction of the first audio channel signal <b>220</b> and of the second audio channel signal <b>222</b>.
Similarly, the (second) residual-signal-assisted multi-channel decoder <b>250</b> provides the third audio channel signal <b>224</b> in the fourth audio channel signal <b>226</b> on the basis of the second downmix signal <b>214</b>, wherein the second downmix signal <b>214</b> may, for example, “coarsely” describe the third audio channel signal <b>224</b> and the fourth audio channel signal <b>226</b>. Moreover, differences between the third audio channel signal <b>224</b> and the fourth audio channel signal <b>226</b> may, for example, be described by (optional) parameters, which may be received by the (second) residual-signal-assisted multi-channel decoder <b>250</b> and by the second residual signal <b>234</b>. Accordingly, the evaluation of the second residual signal <b>234</b> may, for example, allow for a partial waveform reconstruction of the third audio channel signal <b>224</b> and the fourth audio channel signal <b>226</b>. Accordingly, the second residual signal <b>234</b> may allow for an enhancement of the quality of reconstruction of the third audio channel signal <b>224</b> and the fourth audio channel signal <b>226</b>.
However, the first residual signal <b>232</b> and the second residual signal <b>234</b> are derived from a jointly-encoded representation <b>210</b> of the first residual signal and of the second residual signal. Such a multi-channel decoding, which is performed by the multi-channel decoder <b>230</b>, allows for a high decoding efficiency since the first audio channel signal <b>220</b>, the second audio channel signal <b>222</b>, the third audio channel signal <b>224</b> and the fourth audio channel signal <b>226</b> are typically similar or “correlated”. Accordingly, the first residual signal <b>232</b> and the second residual signal <b>234</b> are typically also similar or “correlated”, which can be exploited by deriving the first residual signal <b>232</b> and the second residual signal <b>234</b> from a jointly-encoded representation <b>210</b> using a multi-channel decoding.
Consequently, it is possible to obtain a high decoding quality with moderate bitrate by decoding the residual signals <b>232</b>, <b>234</b> on the basis of a jointly-encoded representation <b>210</b> thereof, and by using each of the residual signals for the decoding of two or more audio channel signals.
To conclude, the audio decoder <b>200</b> allows for a high coding efficiency by providing high quality audio channel signals <b>220</b>, <b>222</b>, <b>224</b>, <b>226</b>.
It should be noted that additional features and functionalities, which can be implemented optionally in the audio decoder <b>200</b>, will be described subsequently taking reference to <figref idref="DRAWINGS">FIGS. 3, 5, 6 and 13</figref>. However, it should be noted that the audio encoder <b>200</b> may comprise the above-mentioned advantages without any additional modification.
3. Audio Decoder According to FIG.
3
<figref idref="DRAWINGS">FIG. 3</figref> shows a block schematic diagram of an audio decoder according to another embodiment of the present invention. The audio decoder of <figref idref="DRAWINGS">FIG. 3</figref> designated in its entirety with <b>300</b>. The audio decoder <b>300</b> is similar to the audio decoder <b>200</b> according to <figref idref="DRAWINGS">FIG. 2</figref>, such that the above explanations also apply. However, the audio decoder <b>300</b> is supplemented with additional features and functionalities when compared to the audio decoder <b>200</b>, as will be explained in the following.
The audio decoder <b>300</b> is configured to receive a jointly-encoded representation <b>310</b> of a first residual signal and of a second residual signal. Moreover, the audio decoder <b>300</b> is configured to receive a jointly-encoded representation <b>360</b> of a first downmix signal and of a second downmix signal. Moreover, the audio decoder <b>300</b> is configured to provide a first audio channel signal <b>320</b>, a second audio channel signal <b>322</b>, a third audio channel signal <b>324</b> and a fourth audio channel signal <b>326</b>. The audio decoder <b>300</b> comprises a multi-channel decoder <b>330</b> which is configured to receive the jointly-encoded representation <b>310</b> of the first residual signal and of the second residual signal and to provide, on the basis thereof, a first residual signal <b>332</b> and a second residual signal <b>334</b>. The audio decoder <b>300</b> also comprises a (first) residual-signal-assisted multi-channel decoding <b>340</b>, which receives the first residual signal <b>332</b> and a first downmix signal <b>312</b>, and provides the first audio channel signal <b>320</b> and the second audio channel signal <b>322</b>. The audio decoder <b>300</b> also comprises a (second) residual-signal-assisted multi-channel decoding <b>350</b>, which is configured to receive the second residual signal <b>334</b> and a second downmix signal <b>314</b>, and to provide the third audio channel signal <b>324</b> and the fourth audio channel signal <b>326</b>.
The audio decoder <b>300</b> also comprises another multi-channel decoder <b>370</b>, which is configured to receive the jointly-encoded representation <b>360</b> of the first downmix signal and of the second downmix signal, and to provide, on the basis thereof, the first downmix signal <b>312</b> and the second downmix signal <b>314</b>.
In the following, some further specific details of the audio decoder <b>300</b> will be described. However, it should be noted that an actual audio decoder does not need to implement a combination of all these additional features and functionalities. Rather, the features and functionalities described in the following can be individually added to the audio decoder <b>200</b> (or any other audio decoder), to gradually improve the audio decoder <b>200</b> (or any other audio decoder).
In an advantageous embodiment, the audio decoder <b>300</b> receives a jointly-encoded representation <b>310</b> of the first residual signal and the second residual signal, wherein this jointly-encoded representation <b>310</b> may comprise a downmix signal of the first residual signal <b>332</b> and of the second residual signal <b>334</b>, and a common residual signal of the first residual signal <b>332</b> and the second residual signal <b>334</b>. In addition, the jointly-encoded representation <b>310</b> may, for example, comprise one or more prediction parameters. Accordingly, the multi-channel decoder <b>330</b> may be a prediction-based, residual-signal-assisted multi-channel decoder. For example, the multi-channel decoder <b>330</b> may be a USAC complex stereo prediction, as described, for example, in the section “Complex Stereo Prediction” of the international standard ISO/IEC 23003-3:2012. For example, the multi-channel decoder <b>330</b> may be configured to evaluate a prediction parameter describing a contribution of a signal component, which is derived using a signal component of a previous frame, to a provision of the first residual signal <b>332</b> and the second residual signal <b>334</b> for a current frame. Moreover, the multi-channel decoder <b>330</b> may be configured to apply the common residual signal (which is included in the jointly-encoded representation <b>310</b>) with a first sign, to obtain the first residual signal <b>332</b>, and to apply the common residual signal (which is included in the jointly-encoded representation <b>310</b>) with a second sign, which is opposite to the first sign, to obtain the second residual signal <b>334</b>. Thus, the common residual signal may, at least partly, describe differences between the first residual signal <b>332</b> and the second residual signal <b>334</b>. However, the multi-channel decoder <b>330</b> may evaluate the downmix signal, the common residual signal and the one or more prediction parameters, which are all included in the jointly-encoded representation <b>310</b>, to obtain the first residual signal <b>332</b> and the second residual signal <b>334</b> as described in the above-referenced international standard ISO/IEC 23003-3:2012. Moreover, it should be noted that the first residual signal <b>332</b> may be associated with a first horizontal position (or azimuth position), for example, a left horizontal position, and that the second residual signal <b>334</b> may be associated with a second horizontal position (or azimuth position), for example a right horizontal position, of an audio scene.
The jointly-encoded representation <b>360</b> of the first downmix signal and of the second downmix signal advantageously comprises a downmix signal of the first downmix signal and of the second downmix signal, a common residual signal of the first downmix signal and of the second downmix signal, and one or more prediction parameters. In other words, there is a “common” downmix signal, into which the first downmix signal <b>312</b> and the second downmix signal <b>314</b> are downmixed, and there is a “common” residual signal which may describe, at least partly, differences between the first downmix signal <b>312</b> and the second downmix signal <b>314</b>. The multi-channel decoder <b>370</b> is advantageously a prediction-based, residual-signal-assisted multi-channel decoder, for example, a USAC complex stereo prediction decoder. In other words, the multi-channel decoder <b>370</b>, which provides the first downmix signal <b>312</b> and the second downmix signal <b>314</b> may be substantially identical to the multi-channel decoder <b>330</b>, which provides the first residual signal <b>332</b> and the second residual signal <b>334</b>, such that the above explanations and references also apply. Moreover, it should be noted that the first downmix signal <b>312</b> is advantageously associated with a first horizontal position or azimuth position (for example, left horizontal position or azimuth position) of the audio scene, and that the second downmix signal <b>314</b> is advantageously associated with a second horizontal position or azimuth position (for example, right horizontal position or azimuth position) of the audio scene. Accordingly, the first downmix signal <b>312</b> and the first residual signal <b>332</b> may be associated with the same, first horizontal position or azimuth position (for example, left horizontal position), and the second downmix signal <b>314</b> and the second residual signal <b>334</b> may be associated with the same, second horizontal position or azimuth position (for example, right horizontal position). Accordingly, both the multi-channel decoder <b>370</b> and the multi-channel decoder <b>330</b> may perform a horizontal splitting (or horizontal separation or horizontal distribution).
The residual-signal-assisted multi-channel decoder <b>340</b> may advantageously be parameter-based, and may consequently receive one or more parameters <b>342</b> describing a desired correlation between two channels (for example, between the first audio channel signal <b>320</b> and the second audio channel signal <b>322</b>) and/or level differences between said two channels. For example, the residual-signal-assisted multi-channel decoding <b>340</b> may be based on an MPEG-Surround coding (as described, for example, in ISO/IEC 23003-1:2007) with a residual signal extension or a “unified stereo decoding” decoder (as described, for example in ISO/IEC 23003-3, chapter 7.11 (Decoder) & Annex B.21 (Description of the Encoder & Definition of the Term “Unified Stereo”)). Accordingly, the residual-signal-assisted multi-channel decoder <b>340</b> may provide the first audio channel signal <b>320</b> and the second audio channel signal <b>322</b>, wherein the first audio channel signal <b>320</b> and the second audio channel signal <b>322</b> are associated with vertically neighboring positions of the audio scene. For example, the first audio channel signal may be associated with a lower left position of the audio scene, and the second audio channel signal may be associated with an upper left position of the audio scene (such that the first audio channel signal <b>320</b> and the second audio channel signal <b>322</b> are, for example, associated with identical horizontal positions or azimuth positions of the audio scene, or with azimuth positions separated by no more than 30 degrees). In other words, the residual-signal-assisted multi-channel decoder <b>340</b> may perform a vertical splitting (or distribution, or separation).
The functionality of the residual-signal-assisted multi-channel decoder <b>350</b> may be identical to the functionality of the residual-signal-assisted multi-channel decoder <b>340</b>, wherein the third audio channel signal may, for example, be associated with a lower right position of the audio scene, and wherein the fourth audio channel signal may, for example, be associated with an upper right position of the audio scene. In other words, the third audio channel signal and the fourth audio channel signal may be associated with vertically neighboring positions of the audio scene, and may be associated with the same horizontal position or azimuth position of the audio scene, wherein the residual-signal-assisted multi-channel decoder <b>350</b> performs a vertical splitting (or separation, or distribution).
To summarize, the audio decoder <b>300</b> according to <figref idref="DRAWINGS">FIG. 3</figref> performs a hierarchical audio decoding, wherein a left-right splitting is performed in the first stages (multi-channel decoder <b>330</b>, multi-channel decoder <b>370</b>), and wherein an upper-lower splitting is performed in the second stage (residual-signal-assisted multi-channel decoders <b>340</b>, <b>350</b>). Moreover, the residual signals <b>332</b>, <b>334</b> are also encoded using a jointly-encoded representation <b>310</b>, as well as the downmix signals <b>312</b>, <b>314</b> (jointly-encoded representation <b>360</b>). Thus, correlations between the different channels are exploited both for the encoding (and decoding) of the downmix signals <b>312</b>, <b>314</b> and for the encoding (and decoding) of the residual signals <b>332</b>, <b>334</b>. Accordingly, a high coding efficiency is achieved, and the correlations between the signals are well exploited.
4. Audio Encoder According to FIG.
4
<figref idref="DRAWINGS">FIG. 4</figref> shows a block schematic diagram of an audio encoder, according to another embodiment of the present invention. The audio encoder according to <figref idref="DRAWINGS">FIG. 4</figref> is designated in its entirety with <b>400</b>. The audio encoder <b>400</b> is configured to receive four audio channel signals, namely a first audio channel signal <b>410</b>, a second audio channel signal <b>412</b>, a third audio channel signal <b>414</b> and a fourth audio channel signal <b>416</b>. Moreover, the audio encoder <b>400</b> is configured to provide an encoded representation on the basis of the audio channel signals <b>410</b>, <b>412</b>, <b>414</b> and <b>416</b>, wherein said encoded representation comprises a jointly encoded representation <b>420</b> of two downmix signals, as well as an encoded representation of a first set <b>422</b> of common bandwidth extension parameters and of a second set <b>424</b> of common bandwidth extension parameters. The audio encoder <b>400</b> comprises a first bandwidth extension parameter extractor <b>430</b>, which is configured to obtain the first set <b>422</b> of common bandwidth extraction parameters on the basis of the first audio channel signal <b>410</b> and the third audio channel signal <b>414</b>. The audio encoder <b>400</b> also comprises a second bandwidth extension parameter extractor <b>440</b>, which is configured to obtain the second set <b>424</b> of common bandwidth extension parameters on the basis of the second audio channel signal <b>412</b> and the fourth audio channel signal <b>416</b>.
Moreover, the audio encoder <b>400</b> comprises a (first) multi-channel encoder <b>450</b>, which is configured to jointly-encode at least the first audio channel signal <b>410</b> and the second audio channel signal <b>412</b> using a multi-channel encoding, to obtain a first downmix signal <b>452</b>. Further, the audio encoder <b>400</b> also comprises a (second) multi-channel encoder <b>460</b>, which is configured to jointly-encode at least the third audio channel signal <b>414</b> and the fourth audio channel signal <b>416</b> using a multi-channel encoding, to obtain a second downmix signal <b>462</b>. Further, the audio encoder <b>400</b> also comprises a (third) multi-channel encoder <b>470</b>, which is configured to jointly-encode the first downmix signal <b>452</b> and the second downmix signal <b>462</b> using a multi-channel encoding, to obtain the jointly-encoded representation <b>420</b> of the downmix signals.
Regarding the functionality of the audio encoder <b>400</b>, it should be noted that the audio encoder <b>400</b> performs a hierarchical multi-channel encoding, wherein the first audio channel signal <b>410</b> and the second audio channel signal <b>412</b> are combined in a first stage, and wherein the third audio channel signal <b>414</b> and the fourth audio channel signal <b>416</b> are also combined in the first stage, to thereby obtain the first downmix signal <b>452</b> and the second downmix signal <b>462</b>. The first downmix signal <b>452</b> and the second downmix signal <b>462</b> are then jointly encoded in a second stage. However, it should be noted that the first bandwidth extension parameter extractor <b>430</b> provides the first set <b>422</b> of common bandwidth extraction parameters on the basis of audio channel signals <b>410</b>, <b>414</b> which are handled by different multi-channel encoders <b>450</b>, <b>460</b> in the first stage of the hierarchical multi-channel encoding. Similarly, the second bandwidth extension parameter extractor <b>440</b> provides a second set <b>424</b> of common bandwidth extraction parameters on the basis of different audio channel signals <b>412</b>, <b>416</b>, which are handled by different multi-channel encoders <b>450</b>, <b>460</b> in the first processing stage. This specific processing order brings along the advantage that the sets <b>422</b>, <b>424</b> of bandwidth extension parameters are based on channels which are only combined in the second stage of the hierarchical encoding (i.e., in the multi-channel encoder <b>470</b>). This is advantageous, since it is desirable to combine such audio channels in the first stage of the hierarchical encoding, the relationship of which is not highly relevant with respect to a sound source position perception. Rather, it is recommendable that the relationship between the first downmix signal and the second downmix signal mainly determines a sound source location perception, because the relationship between the first downmix signal <b>452</b> and the second downmix signal <b>462</b> can be maintained better than the relationship between the individual audio channel signals <b>410</b>, <b>412</b>, <b>414</b>, <b>416</b>. Worded differently, it has been found that it is desirable that the first set <b>422</b> of common bandwidth extension parameters is based on two audio channels (audio channel signals) which contribute to different of the downmix signals <b>452</b>, <b>462</b>, and that the second set <b>424</b> of common bandwidth extension parameters is provided on the basis of audio channel signals <b>412</b>, <b>416</b>, which also contribute to different of the downmix signals <b>452</b>, <b>462</b>, which is reached by the above-described processing of the audio channel signals in the hierarchical multi-channel encoding. Consequently, the first set <b>422</b> of common bandwidth extension parameters is based on a similar channel relationship when compared to the channel relationship between the first downmix signal <b>452</b> and the second downmix signal <b>462</b>, wherein the latter typically dominates the spatial impression generated at the side of an audio decoder. Accordingly, the provision of the first set <b>422</b> of bandwidth extension parameters, and also the provision of the second set <b>424</b> of bandwidth extension parameters is well-adapted to a spatial hearing impression which is generated at the side of an audio decoder.
5. Audio Decoder According to FIG.
5
<figref idref="DRAWINGS">FIG. 5</figref> shows a block schematic diagram of an audio decoder, according to another embodiment of the present invention. The audio decoder according to <figref idref="DRAWINGS">FIG. 5</figref> is designated in its entirety with <b>500</b>.
The audio decoder <b>500</b> is configured to receive a jointly-encoded representation <b>510</b> of a first downmix signal and a second downmix signal. Moreover, the audio decoder <b>500</b> is configured to provide a first bandwidth-extended channel signal <b>520</b>, a second bandwidth extended channel signal <b>522</b>, a third bandwidth-extended channel signal <b>524</b> and a fourth bandwidth-extended channel signal <b>526</b>.
The audio decoder <b>500</b> comprises a (first) multi-channel decoder <b>530</b>, which is configured to provide a first downmix signal <b>532</b> and a second downmix signal <b>534</b> on the basis of the jointly-encoded representation <b>510</b> of the first downmix signal and the second downmix signal using a multi-channel decoding. The audio decoder <b>500</b> also comprises a (second) multi-channel decoder <b>540</b>, which is configured to provide at least a first audio channel signal <b>542</b> and a second audio channel signal <b>544</b> on the basis of the first downmix signal <b>532</b> using a multi-channel decoding. The audio decoder <b>500</b> also comprises a (third) multi-channel decoder <b>550</b>, which is configured to provide at least a third audio channel signal <b>556</b> and a fourth audio channel signal <b>558</b> on the basis of the second downmix signal <b>544</b> using a multi-channel decoding. Moreover, the audio decoder <b>500</b> comprises a (first) multi-channel bandwidth extension <b>560</b>, which is configured to perform a multi-channel bandwidth extension on the basis of the first audio channel signal <b>542</b> and the third audio channel signal <b>556</b>, to obtain a first bandwidth-extended channel signal <b>520</b> and the third bandwidth-extended channel signal <b>524</b>. Moreover, the audio decoder comprises a (second) multi-channel bandwidth extension <b>570</b>, which is configured to perform a multi-channel bandwidth extension on the basis of the second audio channel signal <b>544</b> and the fourth audio channel signal <b>558</b>, to obtain the second bandwidth-extended channel signal <b>522</b> and the fourth bandwidth-extended channel signal <b>526</b>.
Regarding the functionality of the audio decoder <b>500</b>, it should be noted that the audio decoder <b>500</b> performs a hierarchical multi-channel decoding, wherein a splitting between a first downmix signal <b>532</b> and a second downmix signal <b>534</b> is performed in a first stage of the hierarchical decoding, and wherein the first audio channel signal <b>542</b> and the second audio channel signal <b>544</b> are derived from the first downmix signal <b>532</b> in a second stage of the hierarchical decoding, and wherein the third audio channel signal <b>556</b> and the fourth audio channel signal <b>558</b> are derived from the second downmix signal <b>550</b> in the second stage of the hierarchical decoding. However, both the first multi-channel bandwidth extension <b>560</b> and the second multi-channel bandwidth extension <b>570</b> each receive one audio channel signal which is derived from the first downmix signal <b>532</b> and one audio channel signal which is derived from the second downmix signal <b>534</b>. Since a better channel separation is typically achieved by the (first) multi-channel decoding <b>530</b>, which is performed as a first stage of the hierarchical multi-channel decoding, when compared to the second stage of the hierarchical decoding, it can be seen that each multi-channel bandwidth extension <b>560</b>, <b>570</b> receives input signals which are well-separated (because they originate from the first downmix signal <b>532</b> and the second downmix signal <b>534</b>, which are well-channel-separated). Thus, the multi-channel bandwidth extension <b>560</b>, <b>570</b> can consider stereo characteristics, which are important for a hearing impression, and which are well-represented by the relationship between the first downmix signal <b>532</b> and the second downmix signal <b>534</b>, and can therefore provide a good hearing impression.
In other words, the “cross” structure of the audio decoder, wherein each of the multi-channel bandwidth extension stages <b>560</b>, <b>570</b> receives input signals from both (second stage) multi-channel decoders <b>540</b>, <b>550</b> allows for a good multi-channel bandwidth extension, which considers a stereo relationship between the channels.
However, it should be noted that the audio decoder <b>500</b> can be supplemented by any of the features and functionalities described herein with respect to the audio decoders according to <figref idref="DRAWINGS">FIGS. 2, 3, 6 and 13</figref>, wherein it is possible to introduce individual features into the audio decoder <b>500</b> to gradually improve the performance of the audio decoder.
6. Audio Decoder According to FIGS.
6
A and
6
B
<figref idref="DRAWINGS">FIGS. 6A and 6B</figref> show a block schematic diagram of an audio decoder according to another embodiment of the present invention. The audio decoder according to <figref idref="DRAWINGS">FIGS. 6A and 6B</figref> is designated in its entirety with <b>600</b>. The audio decoder <b>600</b> according to <figref idref="DRAWINGS">FIGS. 6A and 6B</figref> is similar to the audio decoder <b>500</b> according to <figref idref="DRAWINGS">FIG. 5</figref>, such that the above explanations also apply. However, the audio decoder <b>600</b> has been supplemented by some features and functionalities, which can also be introduced, individually or in combination, into the audio decoder <b>500</b> for improvement.
The audio decoder <b>600</b> is configured to receive a jointly encoded representation <b>610</b> of a first downmix signal and of a second downmix signal and to provide a first bandwidth-extended signal <b>620</b>, a second bandwidth extended signal <b>622</b>, a third bandwidth extended signal <b>624</b> and a fourth bandwidth extended signal <b>626</b>. The audio decoder <b>600</b> comprises a multi-channel decoder <b>630</b>, which is configured to receive the jointly encoded representation <b>610</b> of the first downmix signal and of the second downmix signal, and to provide, on the basis thereof, the first downmix signal <b>632</b> and the second downmix signal <b>634</b>. The audio decoder <b>600</b> further comprises a multi-channel decoder <b>640</b>, which is configured to receive the first downmix signal <b>632</b> and to provide, on the basis thereof, a first audio channel signal <b>542</b> and a second audio channel signal <b>544</b>. The audio decoder <b>600</b> also comprises a multi-channel decoder <b>650</b>, which is configured to receive the second downmix signal <b>634</b> and to provide a third audio channel signal <b>656</b> and a fourth audio channel signal <b>658</b>. The audio decoder <b>600</b> also comprises a (first) multi-channel bandwidth extension <b>660</b>, which is configured to receive the first audio channel signal <b>642</b> and the third audio channel signal <b>656</b> and to provide, on the basis thereof, the first bandwidth extended channel signal <b>620</b> and the third bandwidth extended channel signal <b>624</b>. Also, a (second) multi-channel bandwidth extension <b>670</b> receives the second audio channel signal <b>644</b> and the fourth audio channel signal <b>658</b> and provides, on the basis thereof, the second bandwidth extended channel signal <b>622</b> and the fourth bandwidth extended channel signal <b>626</b>.
The audio decoder <b>600</b> also comprises a further multi-channel decoder <b>680</b>, which is configured to receive a jointly-encoded representation <b>682</b> of a first residual signal and of a second residual signal and which provides, on the basis thereof, a first residual signal <b>684</b> for usage by the multi-channel decoder <b>640</b> and a second residual signal <b>686</b> for usage by the multi-channel decoder <b>650</b>.
The multi-channel decoder <b>630</b> is advantageously a prediction-based residual-signal-assisted multi-channel decoder. For example, the multi-channel decoder <b>630</b> may be substantially identical to the multi-channel decoder <b>370</b> described above. For example, the multi-channel decoder <b>630</b> may be a USAC complex stereo predication decoder, as mentioned above, and as described in the USAC standard referenced above. Accordingly, the jointly encoded representation <b>610</b> of the first downmix signal and of the second downmix signal may, for example, comprise a (common) downmix signal of the first downmix signal and of the second downmix signal, a (common) residual signal of the first downmix signal and of the second downmix signal, and one or more prediction parameters, which are evaluated by the multi-channel decoder <b>630</b>.
Moreover, it should be noted that the first downmix signal <b>632</b> may, for example, be associated with a first horizontal position or azimuth position (for example, a left horizontal position) of an audio scene and that the second downmix signal <b>634</b> may, for example, be associated with a second horizontal position or azimuth position (for example, a right horizontal position) of the audio scene.
Moreover, the multi-channel decoder <b>680</b> may, for example, be a prediction-based, residual-signal-associated multi-channel decoder. The multi-channel decoder <b>680</b> may be substantially identical to the multi-channel decoder <b>330</b> described above. For example, the multi-channel decoder <b>680</b> may be a USAC complex stereo prediction decoder, as mentioned above. Consequently, the jointly encoded representation <b>682</b> of the first residual signal and of the second residual signal may comprise a (common) downmix signal of the first residual signal and of the second residual signal, a (common) residual signal of the first residual signal and of the second residual signal, and one or more prediction parameters, which are evaluated by the multi-channel decoder <b>680</b>. Moreover, it should be noted that the first residual signal <b>684</b> may be associated with a first horizontal position or azimuth position (for example, a left horizontal position) of the audio scene, and that the second residual signal <b>686</b> may be associated with a second horizontal position or azimuth position (for example, a right horizontal position) of the audio scene.
The multi-channel decoder <b>640</b> may, for example, be a parameter-based multi-channel decoding like, for example, an MPEG surround multi-channel decoding, as described above and in the referenced standard. However, in the presence of the (optional) multi-channel decoder <b>680</b> and the (optional) first residual signal <b>684</b>, the multi-channel decoder <b>640</b> may be a parameter-based, residual-signal-assisted multi-channel decoder, like, for example, a unified stereo decoder. Thus, the multi-channel decoder <b>640</b> may be substantially identical to the multi-channel decoder <b>340</b> described above, and the multi-channel decoder <b>640</b> may, for example, receive the parameters <b>342</b> described above.
Similarly, the multi-channel decoder <b>650</b> may be substantially identical to the multi-channel decoder <b>640</b>. Accordingly, the multi-channel decoder <b>650</b> may, for example, be parameter based and may optionally be residual-signal assisted (in the presence of the optional multi-channel decoder <b>680</b>).
Moreover, it should be noted that the first audio channel signal <b>642</b> and the second audio channel signal <b>644</b> are advantageously associated with vertically adjacent spatial positions of the audio scene. For example, the first audio channel signal <b>642</b> is associated with a lower left position of the audio scene and the second audio channel signal <b>644</b> is associated with an upper left position of the audio scene. Accordingly, the multi-channel decoder <b>640</b> performs a vertical splitting (or separation or distribution) of the audio content described by the first downmix signal <b>632</b> (and, optionally, by the first residual signal <b>684</b>). Similarly, the third audio channel signal <b>656</b> and the fourth audio channel signal <b>658</b> are associated with vertically adjacent positions of the audio scene, and are advantageously associated with the same horizontal position or azimuth position of the audio scene. For example, the third audio channel signal <b>656</b> is advantageously associated with a lower right position of the audio scene and the fourth audio channel signal <b>658</b> is advantageously associated with an upper right position of the audio scene. Thus, the multi-channel decoder <b>650</b> performs a vertical splitting (or separation, or distribution) of the audio content described by the second downmix signal <b>634</b> (and, optionally, the second residual signal <b>686</b>).
However, the first multi-channel bandwidth extension <b>660</b> receives the first audio channel signal <b>642</b> and the third audio channel <b>656</b>, which are associated with the lower left position and a lower right position of the audio scene. Accordingly, the first multi-channel bandwidth extension <b>660</b> performs a multi-channel bandwidth extension on the basis of two audio channel signals which are associated with the same horizontal plane (for example, lower horizontal plane) or elevation of the audio scene and different sides (left/right) of the audio scene. Accordingly, the multi-channel bandwidth extension can consider stereo characteristics (for example, the human stereo perception) when performing the bandwidth extension. Similarly, the second multi-channel bandwidth extension <b>670</b> may also consider stereo characteristics, since the second multi-channel bandwidth extension operates on audio channel signals of the same horizontal plane (for example, upper horizontal plane) or elevation but at different horizontal positions (different sides) (left/right) of the audio scene.
To further conclude, the hierarchical audio decoder <b>600</b> comprises a structure wherein a left/right splitting (or separation, or distribution) is performed in a first stage (multi-channel decoding <b>630</b>, <b>680</b>), wherein a vertical splitting (separation or distribution) is performed in a second stage (multi-channel decoding <b>640</b>, <b>650</b>), and wherein the multi-channel bandwidth extension operates on a pair of left/right signals (multi-channel bandwidth extension <b>660</b>, <b>670</b>). This “crossing” of the decoding paths allows that left/right separation, which is particularly important for the hearing impression (for example, more important than the upper/lower splitting) can be performed in the first processing stage of the hierarchical audio decoder and that the multi-channel bandwidth extension can also be performed on a pair of left-right audio channel signals, which again results in a particularly good hearing impression. The upper/lower splitting is performed as an intermediate stage between the left-right separation and the multi-channel bandwidth extension, which allows to derive four audio channel signals (or bandwidth-extended channel signals) without significantly degrading the hearing impression.
7. Method According to FIG.
7
<figref idref="DRAWINGS">FIG. 7</figref> shows a flow chart of a method <b>700</b> for providing an encoded representation on the basis of at least four audio channel signals.
The method <b>700</b> comprises jointly encoding <b>710</b> at least a first audio channel signal and a second audio channel signal using a residual-signal-assisted multi-channel encoding, to obtain a first downmix signal and a first residual signal. The method also comprises jointly encoding <b>720</b> at least a third audio channel signal and a fourth audio channel signal using a residual-signal-assisted multi-channel encoding, to obtain a second downmix signal and a second residual signal. The method further comprises jointly encoding <b>730</b> the first residual signal and the second residual signal using a multi-channel encoding, to obtain an encoded representation of the residual signals. However, it should be noted that the method <b>700</b> can be supplemented by any of the features and functionalities described herein with respect to the audio encoders and audio decoders.
8. Method According to FIG.
8
<figref idref="DRAWINGS">FIG. 8</figref> shows a flow chart of a method <b>800</b> for providing at least four audio channel signals on the basis of an encoded representation.
The method <b>800</b> comprises providing <b>810</b> a first residual signal and a second residual signal on the basis of a jointly-encoded representation of the first residual signal and the second residual signal using a multi-channel decoding. The method <b>800</b> also comprises providing <b>820</b> a first audio channel signal and a second audio channel signal on the basis of a first downmix signal and the first residual signal using a residual-signal-assisted multi-channel decoding. The method also comprises providing <b>830</b> a third audio channel signal and a fourth audio channel signal on the basis of a second downmix signal and the second residual signal using a residual-signal-assisted multi-channel decoding.
Moreover, it should be noted that the method <b>800</b> can be supplemented by any of the features and functionalities described herein with respect to the audio decoders and audio encoders.
9. Method According to FIG.
9
<figref idref="DRAWINGS">FIG. 9</figref> shows a flow chart of a method <b>900</b> for providing an encoded representation on the basis of at least four audio channel signal.
The method <b>900</b> comprises obtaining <b>910</b> a first set of common bandwidth extension parameters on the basis of a first audio channel signal and a third audio channel signal. The method <b>900</b> also comprises obtaining <b>920</b> a second set of common bandwidth extension parameters on the basis of a second audio channel signal and a fourth audio channel signal. The method also comprises jointly encoding at least the first audio channel signal and the second audio channel signal using a multi-channel encoding, to obtain a first downmix signal and jointly encoding <b>940</b> at least the third audio channel signal and the fourth audio channel signal using a multi-channel encoding to obtain a second downmix signal. The method also comprises jointly encoding <b>950</b> the first downmix signal and the second downmix signal using a multi-channel encoding, to obtain an encoded representation of the downmix signals.
It should be noted that some of the steps of the method <b>900</b>, which do not comprise specific inter dependencies, can be performed in arbitrary order or in parallel. Moreover, it should be noted that the method <b>900</b> can be supplemented by any of the features and functionalities described herein with respect to the audio encoders and audio decoders.
10. Method According to FIG.
10
<figref idref="DRAWINGS">FIG. 10</figref> shows a flow chart of a method <b>1000</b> for providing at least four audio channel signals on the basis of an encoded representation.
The method <b>1000</b> comprises providing <b>1010</b> a first downmix signal and a second downmix signal on the basis of a jointly encoded representation of the first downmix signal and the second downmix signal using a multi-channel decoding, providing <b>1020</b> at least a first audio channel signal and a second audio channel signal on the basis of the first downmix signal using a multi-channel decoding, providing <b>1030</b> at least a third audio channel signal and a fourth audio channel signal on the basis of the second downmix signal using a multi-channel decoding, performing <b>1040</b> a multi-channel bandwidth extension on the basis of the first audio channel signal and the third audio channel signal, to obtain a first bandwidth-extended channel signal and a third bandwidth-extended channel signal, and performing <b>1050</b> a multi-channel bandwidth extension on the basis of the second audio channel signal and the fourth audio channel signal, to obtain a second bandwidth-extended channel signal and a fourth bandwidth-extended channel signal.
It should be noted that some of the steps of the method <b>1000</b> may be preformed in parallel or in a different order. Moreover, it should be noted that the method <b>1000</b> can be supplemented by any of the features and functionalities described herein with respect to the audio encoder and the audio decoder.
11. Embodiments According to FIGS.
11
,
12
and
13
In the following, some additional embodiments according to the present invention and the underlying considerations will be described.
<figref idref="DRAWINGS">FIG. 11</figref> shows a block schematic diagram of an audio encoder <b>1100</b> according to an embodiment of the invention. The audio encoder <b>1100</b> is configured to receive a left lower channel signal <b>1110</b>, a left upper channel signal <b>1112</b>, a right lower channel signal <b>1114</b> and a right upper channel signal <b>1116</b>.
The audio encoder <b>1100</b> comprises a first multi-channel audio encoder (or encoding) <b>1120</b>, which is an MPEG surround 2-1-2 audio encoder (or encoding) or a unified stereo audio encoder (or encoding) and which receives the left lower channel signal <b>1110</b> and the left upper channel signal <b>1112</b>. The first multi-channel audio encoder <b>1120</b> provides a left downmix signal <b>1122</b> and, optionally, a left residual signal <b>1124</b>. Moreover, the audio encoder <b>1100</b> comprises a second multi-channel encoder (or encoding) <b>1130</b>, which is an MPEG-surround 2-1-2 encoder (or encoding) or a unified stereo encoder (or encoding) which receives the right lower channel signal <b>1114</b> and the right upper channel signal <b>1116</b>. The second multi-channel audio encoder <b>1130</b> provides a right downmix signal <b>1132</b> and, optionally, a right residual signal <b>1134</b>. The audio encoder <b>1100</b> also comprises a stereo coder (or coding) <b>1140</b>, which receives the left downmix signal <b>1122</b> and the right downmix signal <b>1132</b>. Moreover, the first stereo coding <b>1140</b>, which is a complex prediction stereo coding, receives a psycho acoustic model information <b>1142</b> from a psycho acoustic model. For example, the psycho model information <b>1142</b> may describe the psycho acoustic relevance of different frequency bands or frequency subbands, psycho acoustic masking effects and the like. The stereo coding <b>1140</b> provides a channel pair element (CPE) “downmix”, which is designated with <b>1144</b> and which describes the left downmix signal <b>1122</b> and the right downmix signal <b>1132</b> in a jointly encoded form. Moreover, the audio encoder <b>1100</b> optionally comprises a second stereo coder (or coding) <b>1150</b>, which is configured to receive the optional left residual signal <b>1124</b> and the optional right residual signal <b>1134</b>, as well as the psycho acoustic model information <b>1142</b>. The second stereo coding <b>1150</b>, which is a complex prediction stereo coding, is configured to provide a channel pair element (CPE) “residual”, which represents the left residual signal <b>1124</b> and the right residual signal <b>1134</b> in a jointly encoded form.
The encoder <b>1100</b> (as well as the other audio encoders described herein) is based on the idea that horizontal and vertical signal dependencies are exploited by hierarchically combining available USAC stereo tools (i.e., encoding concepts which are available in the USAC encoding). Vertically neighbored channel pairs are combined using MPEG surround 2-1-2 or unified stereo (designated with <b>1120</b> and <b>1130</b>) with a band-limited or full-band residual signal (designated with <b>1124</b> and <b>1134</b>). The output of each vertical channel pair is a downmix signal <b>1122</b>, <b>1132</b> and, for the unified stereo, a residual signal <b>1124</b>, <b>1134</b>. In order to satisfy perceptual requirements for binaural unmasking, both downmix signals <b>1122</b>, <b>1132</b> are combined horizontally and jointly coded by use of complex prediction (encoder <b>1140</b>) in the MDCT domain, which includes the possibility of left-right and mid-side coding. The same method can be applied to the horizontally combined residual signals <b>1124</b>, <b>1134</b>. This concept is illustrated in <figref idref="DRAWINGS">FIG. 11</figref>.
The hierarchical structure explained with reference to <figref idref="DRAWINGS">FIG. 11</figref> can be achieved by enabling both stereo tools (for example, both USAC stereo tools) and resorting channels in between. Thus, no additional pre-/post processing step is necessary and the bit stream syntax for transmission of the tool's payloads remains unchanged (for example, substantially unchanged when compared to the USAC standard). This idea results in the encoder structure shown in <figref idref="DRAWINGS">FIG. 12</figref>.
<figref idref="DRAWINGS">FIG. 12</figref> shows a block schematic diagram of an audio encoder <b>1200</b>, according to an embodiment of the invention. The audio encoder <b>1200</b> is configured to receive a first channel signal <b>1210</b>, a second channel signal <b>1212</b>, a third channel signal <b>1214</b> and a fourth channel signal <b>1216</b>. The audio encoder <b>1200</b> is configured to provide a bit stream <b>1220</b> for a first channel pair element and a bit stream <b>1222</b> for a second channel pair element.
The audio encoder <b>1200</b> comprises a first multi-channel encoder <b>1230</b>, which is an MPEG-surround 2-1-2 encoder or a unified stereo encoder, and which receives the first channel signal <b>1210</b> and the second channel signal <b>1212</b>. Moreover, the first multi-channel encoder <b>1230</b> provides a first downmix signal <b>1232</b>, an MPEG surround payload <b>1236</b> and, optionally, a first residual signal <b>1234</b>. The audio encoder <b>1200</b> also comprises a second multi-channel encoder <b>1240</b> which is an MPEG surround 2-1-2 encoder or a unified stereo encoder and which receives the third channel signal <b>1214</b> and the fourth channel signal <b>1216</b>. The second multi-channel encoder <b>1240</b> provides a first downmix signal <b>1242</b>, an MPEG surround payload <b>1246</b> and, optionally, a second residual signal <b>1244</b>.
The audio encoder <b>1200</b> also comprises first stereo coding <b>1250</b>, which is a complex prediction stereo coding. The first stereo coding <b>1250</b> receives the first downmix signal <b>1232</b> and the second downmix signal <b>1242</b>. The first stereo coding <b>1250</b> provides a jointly encoded representation <b>1252</b> of the first downmix signal <b>1232</b> and the second downmix signal <b>1242</b>, wherein the jointly encoded representation <b>1252</b> may comprise a representation of a (common) downmix signal (of the first downmix signal <b>1232</b> and of the second downmix signal <b>1242</b>) and of a common residual signal (of the first downmix signal <b>1232</b> and of the second downmix signal <b>1242</b>). Moreover, the (first) complex prediction stereo coding <b>1250</b> provides a complex prediction payload <b>1254</b>, which typically comprises one or more complex prediction coefficients. Moreover, the audio encoder <b>1200</b> also comprises a second stereo coding <b>1260</b>, which is a complex prediction stereo coding. The second stereo coding <b>1260</b> receives the first residual signal <b>1234</b> and the second residual signal <b>1244</b> (or zero input values, if there is no residual signal provided by the multi-channel encoders <b>1230</b>, <b>1240</b>). The second stereo coding <b>1260</b> provides a jointly encoded representation <b>1262</b> of the first residual signal <b>1234</b> and of the second residual signal <b>1244</b>, which may, for example, comprise a (common) downmix signal (of the first residual signal <b>1234</b> and of the second residual signal <b>1244</b>) and a common residual signal (of the first residual signal <b>1234</b> and of the second residual signal <b>1244</b>). Moreover, the complex prediction stereo coding <b>1260</b> provides a complex prediction payload <b>1264</b> which typically comprises one or more prediction coefficients.
Moreover, the audio encoder <b>1200</b> comprises a psycho acoustic model <b>1270</b>, which provides an information that controls the first complex prediction stereo coding <b>1250</b> and the second complex prediction stereo coding <b>1260</b>. For example, the information provided by the psycho acoustic model <b>1270</b> may describe which frequency bands or frequency bins are of high psycho acoustic relevance and should be encoded with high accuracy. However, it should be noted that the usage of the information provided by the psycho acoustic model <b>1270</b> is optional.
Moreover, the audio encoder <b>1200</b> comprises a first encoder and multiplexer <b>1280</b> which receives the jointly encoded representation <b>1252</b> from the first complex prediction stereo coding <b>1250</b>, the complex prediction payload <b>1254</b> from the first complex prediction stereo coding <b>1250</b> and the MPEG surround payload <b>1236</b> from the first multi-channel audio encoder <b>1230</b>. Moreover, the first encoding and multiplexing <b>1280</b> may receive information from the psycho acoustic model <b>1270</b>, which describes, for example, which encoding precision should be applied to which frequency bands or frequency subbands, taking into account psycho acoustic masking effects and the like. Accordingly, the first encoding and multiplexing <b>1280</b> provides the first channel pair element bit stream <b>1220</b>.
Moreover, the audio encoder <b>1200</b> comprises a second encoding and multiplexing <b>1290</b>, which is configured to receive the jointly encoded representation <b>1262</b> provided by the second complex prediction stereo encoding <b>1260</b>, the complex prediction payload <b>1264</b> proved by the second complex prediction stereo coding <b>1260</b>, and the MPEG surround payload <b>1246</b> provided by the second multi-channel audio encoder <b>1240</b>. Moreover, the second encoding and multiplexing <b>1290</b> may receive an information from the psycho acoustic model <b>1270</b>. Accordingly, the second encoding and multiplexing <b>1290</b> provides the second channel pair element bit stream <b>1222</b>.
Regarding the functionality of the audio encoder <b>1200</b>, reference is made to the above explanations, and also to the explanations with respect to the audio encoders according to <figref idref="DRAWINGS">FIGS. 2, 3, 5 and 6</figref>.
Moreover, it should be noted that this concept can be extended to use multiple MPEG surround boxes for joint coding of horizontally, vertically or otherwise geometrically related channels and combining the downmix and residual signals to complex prediction stereo pairs, considering their geometric and perceptual properties. This leads to a generalized decoder structure.
In the following, the implementation of a quad channel element will be described. In a three-dimensional audio coding system, the hierarchical combination of four channels to form a quad channel element (QCE) is used. A QCE consists of two USAC channel pair elements (CPE) (or provides two USAC channel pair elements, or receives to USAC channel pair elements). Vertical channel pairs are combined using MPS 2-1-2 or unified stereo. The downmix channels are jointly coded in the first channel pair element CPE. If residual coding is applied, the residual signals are jointly coded in the second channel pair element CPE, else the signal in the second CPE is set to zero. Both channel pair elements CPEs use complex prediction for joint stereo coding, including the possibility of left-right and mid-side coding. To preserve the perceptual stereo properties of the high frequency part of the signal, stereo SBR (spectral bandwidth replication) is applied between the upper left/right channel pair and the lower left/right channel pair, by an additional resorting step before the application of SBR.
A possible decoder structure will be described taking reference to <figref idref="DRAWINGS">FIG. 13</figref> which shows a block schematic diagram of an audio decoder according to an embodiment of the invention. The audio decoder <b>1300</b> is configured to receive a first bit stream <b>1310</b> representing a first channel pair element and a second bit stream <b>1312</b> representing a second channel pair element. However, the first bit stream <b>1310</b> and the second bit stream <b>1312</b> may be included in a common overall bit stream.
The audio decoder <b>1300</b> is configured to provide a first bandwidth extended channel signal <b>1320</b>, which may, for example, represent a lower left position of an audio scene, a second bandwidth extended channel signal <b>1322</b>, which may, for example, represent an upper left position of the audio scene, a third bandwidth extended channel signal <b>1324</b>, which may, for example, be associated with a lower right position of the audio scene and a fourth bandwidth extended channel signal <b>1326</b>, which may, for example, be associated with an upper right position of the audio scene.
The audio decoder <b>1300</b> comprises a first bit stream decoding <b>1330</b>, which is configured to receive the bit stream <b>1310</b> for the first channel pair element and to provide, on the basis thereof, a jointly-encoded representation of two downmix signals, a complex prediction payload <b>1334</b>, an MPEG surround payload <b>1336</b> and a spectral bandwidth replication payload <b>1338</b>. The audio decoder <b>1300</b> also comprises a first complex prediction stereo decoding <b>1340</b>, which is configured to receive the jointly encoded representation <b>1332</b> and the complex prediction payload <b>1334</b> and to provide, on the basis thereof, a first downmix signal <b>1342</b> and a second downmix signal <b>1344</b>. Similarly, the audio decoder <b>1300</b> comprises a second bit stream decoding <b>1350</b> which is configured to receive the bit stream <b>1312</b> for the second channel element and to provide, on the basis thereof, a jointly encoded representation <b>1352</b> of two residual signals, a complex prediction payload <b>1354</b>, an MPEG surround payload <b>1356</b> and a spectral bandwidth replication bit load <b>1358</b>. The audio decoder also comprises a second complex prediction stereo decoding <b>1360</b>, which provides a first residual signal <b>1362</b> and a second residual signal <b>1364</b> on the basis of the jointly encoded representation <b>1352</b> and the complex prediction payload <b>1354</b>.
Moreover, the audio decoder <b>1300</b> comprises a first MPEG surround-type multichannel decoding <b>1370</b>, which is an MPEG surround 2-1-2 decoding or a unified stereo decoding. The first MPEG surround-type multi-channel decoding <b>1370</b> receives the first downmix signal <b>1342</b>, the first residual signal <b>1362</b> (optional) and the MPEG surround payload <b>1336</b> and provides, on the basis thereof, a first audio channel signal <b>1372</b> and a second audio channel signal <b>1374</b>. The audio decoder <b>1300</b> also comprises a second MPEG surround-type multi-channel decoding <b>1380</b>, which is an MPEG surround 2-1-2 multi-channel decoding or a unified stereo multi-channel decoding. The second MPEG surround-type multi-channel decoding <b>1380</b> receives the second downmix signal <b>1344</b> and the second residual signal <b>1364</b> (optional), as well as the MPEG surround payload <b>1356</b>, and provides, on the basis thereof, a third audio channel signal <b>1382</b> and fourth audio channel signal <b>1384</b>. The audio decoder <b>1300</b> also comprises a first stereo spectral bandwidth replication <b>1390</b>, which is configured to receive the first audio channel signal <b>1372</b> and the third audio channel signal <b>1382</b>, as well as the spectral bandwidth replication payload <b>1338</b>, and to provide, on the basis thereof, the first bandwidth extended channel signal <b>1320</b> and the third bandwidth extended channel signal <b>1324</b>. Moreover, the audio decoder comprises a second stereo spectral bandwidth replication <b>1394</b>, which is configured to receive the second audio channel signal <b>1374</b> and the fourth audio channel signal <b>1384</b>, as well as the spectral bandwidth replication payload <b>1358</b> and to provide, on the basis thereof, the second bandwidth extended channel signal <b>1322</b> and the fourth bandwidth extended channel signal <b>1326</b>.
Regarding the functionality of the audio decoder <b>1300</b>, reference is made to the above discussion, and also the discussion of the audio decoder according to <figref idref="DRAWINGS">FIGS. 2, 3, 5 and 6</figref>.
In the following, an example of a bit stream which can be used for the audio encoding/decoding described herein will be described taking reference to <figref idref="DRAWINGS">FIGS. 14<i>a </i>and 14<i>b</i></figref>. It should be noted that the bit stream may, for example, be an extension of the bit stream used in the unified speech-and-audio coding (USAC), which is described in the above mentioned standard (ISO/IEC 23003-3:2012). For example, the MPEG surround payloads <b>1236</b>, <b>1246</b>, <b>1336</b>, <b>1356</b> and the complex prediction payloads <b>1254</b>, <b>1264</b>, <b>1334</b>, <b>1354</b> may be transmitted as for legacy channel pair elements (i.e., for channel pair elements according to the USAC standard). For signaling the use of a quad channel element QCE, the USAC channel pair configuration may be extended by two bits, as shown in <figref idref="DRAWINGS">FIG. 14<i>a</i></figref>. In other words, two bits designated with “qceIndex” may be added to the USAC bitstream element “UsacChannelPairElementConfig( )”. The meaning of the parameter represented by the bits “qceIndex” can be defined, for example, as shown in the table of <figref idref="DRAWINGS">FIG. 14</figref><i>b. </i>
For example, two channel pair elements that form a QCE may be transmitted as consecutive elements, first the CPE containing the downmix channels and the MPS payload for the first MPS box, second the CPE containing the residual signal (or zero audio signal for MPS 2-1-2 coding) and the MPS payload for the second MPS box.
In other words, there is only a small signaling overhead when compared to the conventional USAC bit stream for transmitting a quad channel element QCE.
However, different bit stream formats can naturally also be used.
12. Encoding/Decoding Environment
In the following, an audio encoding/decoding environment will be described in which concepts according to the present invention can be applied.
A 3D audio codec system, in which the concepts according to the present invention can be used, is based on an MPEG-D USAC codec for decoding of channel and object signals. To increase the efficiency for coding a large amount of objects, MPEG SAOC technology has been adapted. Three types of renderers perform the tasks of rendering objects to channels, rendering channels to headphones or rendering channels to a different loudspeaker setup. When object signals are explicitly transmitted or parametrically encoded using SAOC, the corresponding object metadata information is compressed and multiplexed into the 3D audio bit stream.
<figref idref="DRAWINGS">FIG. 15</figref> shows a block schematic diagram of such an audio encoder, and <figref idref="DRAWINGS">FIG. 16</figref> shows a block schematic diagram of such an audio decoder. In other words, <figref idref="DRAWINGS">FIGS. 15 and 16</figref> show the different algorithmic blocks of the 3D audio system.
Taking reference now to <figref idref="DRAWINGS">FIG. 15</figref>, which shows a block schematic diagram of a 3D audio encoder <b>1500</b>, some details will be explained. The encoder <b>1500</b> comprises an optional pre-renderer/mixer <b>1510</b>, which receives one or more channel signals <b>1512</b> and one or more object signals <b>1514</b> and provides, on the basis thereof, one or more channel signals <b>1516</b> as well as one or more object signals <b>1518</b>, <b>1520</b>. The audio encoder also comprises a USAC encoder <b>1530</b> and, optionally, a SAOC encoder <b>1540</b>. The SAOC encoder <b>1540</b> is configured to provide one or more SAOC transport channels <b>1542</b> and a SAOC side information <b>1544</b> on the basis of one or more objects <b>1520</b> provided to the SAOC encoder. Moreover, the USAC encoder <b>1530</b> is configured to receive the channel signals <b>1516</b> comprising channels and pre-rendered objects from the pre-renderer/mixer, to receive one or more object signals <b>1518</b> from the pre-renderer/mixer and to receive one or more SAOC transport channels <b>1542</b> and SAOC side information <b>1544</b>, and provides, on the basis thereof, an encoded representation <b>1532</b>. Moreover, the audio encoder <b>1500</b> also comprises an object metadata encoder <b>1550</b> which is configured to receive object metadata <b>1552</b> (which may be evaluated by the pre-renderer/mixer <b>1510</b>) and to encode the object metadata to obtain encoded object metadata <b>1554</b>. The encoded metadata is also received by the USAC encoder <b>1530</b> and used to provide the encoded representation <b>1532</b>.
Some details regarding the individual components of the audio encoder <b>1500</b> will be described below.
Taking reference now to <figref idref="DRAWINGS">FIG. 16</figref>, an audio decoder <b>1600</b> will be described. The audio decoder <b>1600</b> is configured to receive an encoded representation <b>1610</b> and to provide, on the basis thereof, multi-channel loudspeaker signals <b>1612</b>, headphone signals <b>1614</b> and/or loudspeaker signals <b>1616</b> in an alternative format (for example, in a 5.1 format).
The audio decoder <b>1600</b> comprises a USAC decoder <b>1620</b>, and provides one or more channel signals <b>1622</b>, one or more pre-rendered object signals <b>1624</b>, one or more object signals <b>1626</b>, one or more SAOC transport channels <b>1628</b>, a SAOC side information <b>1630</b> and a compressed object metadata information <b>1632</b> on the basis of the encoded representation <b>1610</b>. The audio decoder <b>1600</b> also comprises an object renderer <b>1640</b> which is configured to provide one or more rendered object signals <b>1642</b> on the basis of the object signal <b>1626</b> and an object metadata information <b>1644</b>, wherein the object metadata information <b>1644</b> is provided by an object metadata decoder <b>1650</b> on the basis of the compressed object metadata information <b>1632</b>. The audio decoder <b>1600</b> also comprises, optionally, a SAOC decoder <b>1660</b>, which is configured to receive the SAOC transport channel <b>1628</b> and the SAOC side information <b>1630</b>, and to provide, on the basis thereof, one or more rendered object signals <b>1662</b>. The audio decoder <b>1600</b> also comprises a mixer <b>1670</b>, which is configured to receive the channel signals <b>1622</b>, the pre-rendered object signals <b>1624</b>, the rendered object signals <b>1642</b>, and the rendered object signals <b>1662</b>, and to provide, on the basis thereof, a plurality of mixed channel signals <b>1672</b> which may, for example, constitute the multi-channel loudspeaker signals <b>1612</b>. The audio decoder <b>1600</b> may, for example, also comprise a binaural render <b>1680</b>, which is configured to receive the mixed channel signals <b>1672</b> and to provide, on the basis thereof, the headphone signals <b>1614</b>. Moreover, the audio decoder <b>1600</b> may comprise a format conversion <b>1690</b>, which is configured to receive the mixed channel signals <b>1672</b> and a reproduction layout information <b>1692</b> and to provide, on the basis thereof, a loudspeaker signal <b>1616</b> for an alternative loudspeaker setup.
In the following, some details regarding the components of the audio encoder <b>1500</b> and of the audio decoder <b>1600</b> will be described.
Pre-Renderer/Mixer
The pre-renderer/mixer <b>1510</b> can be optionally used to convert a channel plus object input scene into a channel scene before encoding. Functionally, it may, for example, be identical to the object renderer/mixer described below. Pre-rendering of objects may, for example, ensure a deterministic signal entropy at the encoder input that is basically independent of the number of simultaneously active object signals. In the pre-rendering of objects, no object metadata transmission is required. Discreet object signals are rendered to the channel layout that the encoder is configured to use. The weights of the objects for each channel are obtained from the associated object metadata (OAM) <b>1552</b>.
USAC Core Codec
The core codec <b>1530</b>, <b>1620</b> for loudspeaker-channel signals, discreet object signals, object downmix signals and pre-rendered signals is based on MPEG-D USAC technology. It handles the coding of the multitude of signals by creating channel and object mapping information based on the geometric and semantic information of the input's channel and object assignment. This mapping information describes how input channels and objects are mapped to USAC-channel elements (CPEs, SCEs, LFEs) and the corresponding information is transmitted to the decoder. All additional payloads like SAOC data or object metadata have been passed through extension elements and have been considered in the encoders rate control.
The coding of objects is possible in different ways, depending on the rate/distortion requirements and the interactivity requirements for the renderer. The following object coding variants are possible: <ul id="ul0001" list-style="none"><li id="ul0001-0001" num="0000"><ul id="ul0002" list-style="none"><li id="ul0002-0001" num="0161">1. Pre-rendered objects: object signals are pre-rendered and mixed to the 22.2 channel signals before encoding. The subsequent coding chain sees 22.2 channel signals.</li><li id="ul0002-0002" num="0162">2. Discreet object wave forms: objects are supplied as monophonic wave forms to the encoder. The encoder uses single channel elements SCEs to transfer the objects in addition to the channel signals. The decoded objects are rendered and mixed at the receiver side. Compressed object metadata information is transmitted to the receiver/renderer along side.</li><li id="ul0002-0003" num="0163">3. Parametric object wave forms: object properties and there relation to each other are described by means of SAOC parameters. The downmix of the object signals is coded with USAC. The parametric information is transmitted along side. The number of downmix channels is chosen depending on the number of objects and the overall data rate. Compressed object metadata information is transmitted to the SAOC renderer. <br /> SAOC </li></ul></li></ul>
The SAOC encoder <b>1540</b> and the SAOC decoder <b>1660</b> for object signals are based on MPEG SAOC technology. The system is capable of recreating, modifying and rendering a number of audio objects based on a smaller number of transmitted channels and additional parametric data (object level differences OLDs, inter object correlations IOCs, downmix gains DMGs). The additional parametric data exhibits a significantly lower data rate than may be used for transmitting all objects individually, making the coding very efficient. The SAOC encoder takes as input the object/channel signals as monophonic waveforms and outputs the parametric information (which is packed into the 3D-audio bit stream <b>1532</b>, <b>1610</b>) and the SAOC transport channels (which are encoded using single channel elements and transmitted).
The SAOC decoder <b>1600</b> reconstructs the object/channel signals from the decoded SAOC transport channels <b>1628</b> and parametric information <b>1630</b>, and generates the output audio scene based on the reproduction layout, the decompressed object metadata information and optionally on the user interaction information. <br /> Object Metadata Codec
For each object, the associated metadata that specifies the geometrical position and volume of the object in 3D space is efficiently coded by quantization of the object properties in time and space. The compressed object metadata cOAM <b>1554</b>, <b>1632</b> is transmitted to the receiver as side information.
Object Renderer/Mixer
The object renderer utilizes the compressed object metadata to generate object waveforms according to the given reproduction format. Each object is rendered to certain output channels according to its metadata. The output of this block results from the sum of the partial results. If both channel based content as well as discreet/parametric objects are decoded, the channel based waveforms and the rendered object waveforms are mixed before outputting the resulting waveforms (or before feeding them to a post processor module like the binaural renderer or the loudspeaker renderer module).
Binaural Renderer
The binaural renderer module <b>1680</b> produces a binaural downmix of the multichannel audio material, such that each input channel is represented by a virtual sound source. The processing is conducted frame-wise in QMF domain. The binauralization is based on measured binaural room impulse responses.
Loudspeaker Renderer/Format Conversion
The loudspeaker renderer <b>1690</b> converts between the transmitted channel configuration and the desired reproduction format. It is thus called “format converter” in the following. The format converter performs conversions to lower numbers of output channels, i.e., it creates downmixes. The system automatically generates optimized downmix matrices for the given combination of input and output formats and applies these matrices in a dowmix process. The format converter allows for standard loudspeaker configurations as well as for random configurations with non-standard loudspeaker positions.
<figref idref="DRAWINGS">FIG. 17</figref> shows a block schematic diagram of the format converter. As can be seen, the format converter <b>1700</b> receives mixer output signals <b>1710</b>, for example, the mixed channel signals <b>1672</b> and provides loudspeaker signals <b>1712</b>, for example, the speaker signals <b>1616</b>. The format converter comprises a downmix process <b>1720</b> in the QMF domain and a downmix configurator <b>1730</b>, wherein the downmix configurator provides configuration information for the downmix process <b>1720</b> on the basis of a mixer output layout information <b>1732</b> and a reproduction layout information <b>1734</b>.
Moreover, it should be noted that the concepts described above, for example the audio encoder <b>100</b>, the audio decoder <b>200</b> or <b>300</b>, the audio encoder <b>400</b>, the audio decoder <b>500</b> or <b>600</b>, the methods <b>700</b>, <b>800</b>, <b>900</b>, or <b>1000</b>, the audio encoder <b>1100</b> or <b>1200</b> and the audio decoder <b>1300</b> can be used within the audio encoder <b>1500</b> and/or within the audio decoder <b>1600</b>. For example, the audio encoders/decoders mentioned before can be used for encoding or decoding of channel signals which are associated with different spatial positions.
13. Alternative Embodiments
In the following, some additional embodiments will be described.
Taking reference now to <figref idref="DRAWINGS">FIGS. 18 to 21</figref>, additional embodiments according o the invention will be explained.
It should be noted that a so-called “Quad Channel Element” (QCE) can be considered as a tool of an audio decoder, which can be used, for example, for decoding 3-dimensional audio content.
In other words, the Quad Channel Element (QCE) is a method for joint coding of four channels for more efficient coding of horizontally and vertically distributed channels. A QCE consists of two consecutive CPEs and is formed by hierarchically combining the Joint Stereo Tool with possibility of Complex Stereo Prediction Tool in horizontal direction and the MPEG Surround based stereo tool in vertical direction. This is achieved by enabling both stereo tools and swapping output channels between applying the tools. Stereo SBR is performed in horizontal direction to preserve the left-right relations of high frequencies.
<figref idref="DRAWINGS">FIG. 18</figref> shows a topological structure of a QCE. It should be noted that the QCE of <figref idref="DRAWINGS">FIG. 18</figref> is very similar to the QCE of <figref idref="DRAWINGS">FIG. 11</figref>, such that reference is made to the above explanations. However, it should be noted that, in the QCE of <figref idref="DRAWINGS">FIG. 18</figref>, it is not necessary to make use of the psychoacoustic model when performing complex stereo prediction (while, such use is naturally possible optionally). Moreover, it can be seen that first stereo spectral bandwidth replication (Stereo SBR) is performed on the basis of the left lower channel and the right lower channel, and that that second stereo spectral bandwidth replication (Stereo SBR) is performed on the basis of the left upper channel and the right upper channel.
In the following, some terms and definitions will be provided, which may apply in some embodiments.
A data element qceIndex indicates a QCE mode of a CPE. Regarding the meaning of the bitstream variable qceIndex, reference is made to <figref idref="DRAWINGS">FIG. 14<i>b</i></figref>. It should be noted that qceIndex describes whether two subsequent elements of type UsacChannelPairElement( ) are treated as a Quadruple Channel Element (QCE). The different QCE modes are given in <figref idref="DRAWINGS">FIG. 14<i>b</i></figref>. The qceIndex shall be the same for the two subsequent elements forming one QCE.
In the following, some help elements will be defined, which may be used in some embodiments according to the invention: <ul id="ul0003" list-style="none"><li id="ul0003-0001" num="0180">cplx_out_dmx_L[ ] first channel of first CPE after complex prediction stereo decoding</li><li id="ul0003-0002" num="0181">cplx_out_dmx_R[ ] second channel of first CPE after complex prediction stereo decoding</li><li id="ul0003-0003" num="0182">cplx_out_res_L[ ] second CPE after complex prediction stereo decoding (zero if qceIndex=1)</li><li id="ul0003-0004" num="0183">cplx_out_res_R[ ] second channel of second CPE after complex prediction stereo decoding (zero if qceIndex=1)</li><li id="ul0003-0005" num="0184">mps_out_L_1[ ] first output channel of first MPS box</li><li id="ul0003-0006" num="0185">mps_out_L_2[ ] second output channel of first MPS box</li><li id="ul0003-0007" num="0186">mps_out_R_1[ ] first output channel of second MPS box</li><li id="ul0003-0008" num="0187">mps_out_R_2[ ] second output channel of second MPS box</li><li id="ul0003-0009" num="0188">sbr_out_L_1[ ] first output channel of first Stereo SBR box</li><li id="ul0003-0010" num="0189">sbr_out_R_1[ ] second output channel of first Stereo SBR box</li><li id="ul0003-0011" num="0190">sbr_out_L_2[ ] first output channel of second Stereo SBR box</li><li id="ul0003-0012" num="0191">sbr_out_R_2[ ] second output channel of second Stereo SBR box</li></ul>
In the following, a decoding process, which is performed in an embodiment according to the invention, will be explained.
The syntax element (or bitstream element, or data element) qceIndex in UsacChannelPairElementConfig( ) indicates whether a CPE belongs to a QCE and if residual coding is used. In case that qceIndex is unequal 0, the current CPE forms a QCE together with its subsequent element which shall be a CPE having the same qceIndex. Stereo SBR is used for the QCE, thus the syntax item stereoConfigIndex shall be 3 and bsStereoSbr shall be 1.
In case of qceIndex==1 only the payloads for MPEG Surround and SBR and no relevant audio signal data is contained in the second CPE and the syntax element bsResidualCoding is set to 0.
The presence of a residual signal in the second CPE is indicated by qceIndex==2. In this case the syntax element bsResidualCoding is set to 1.
However, some different and possible simplified signaling schemes may also be used.
Decoding of Joint Stereo with possibility of Complex Stereo Prediction is performed as described in ISO/IEC 23003-3, subclause 7.7. The resulting output of the first CPE are the MPS downmix signals cplx_out_dmx_L[ ] and cplx_out_dmx_R[ ]. If residual coding is used (i.e. qceIndex==2), the output of the second CPE are the MPS residual signals cplx_out_res_L[ ], cplx_out_res_R[ ], if no residual signal has been transmitted (i.e. qceIndex==1), zero signals are inserted.
Before applying MPEG Surround decoding, the second channel of the first element (cplx_out_dmx_R[ ]) and the first channel of the second element (cplx_out_res_L[ ]) are swapped.
Decoding of MPEG Surround is performed as described in ISO/IEC 23003-3, subclause 7.11. If residual coding is used, the decoding may, however, be modified when compared to conventional MPEG surround decoding in some embodiments. Decoding of MPEG Surround without residual using SBR as defined in ISO/IEC 23003-3, subclause 7.11.2.7 (FIG. 23), is modified so that Stereo SBR is also used for bsResidualCoding==1, resulting in the decoder schematics shown in <figref idref="DRAWINGS">FIG. 19</figref>. <figref idref="DRAWINGS">FIG. 19</figref> shows a block schematic diagram of an audio coder for bsResidualCoding==0 and bsStereoSbr==1.
As can be seen in <figref idref="DRAWINGS">FIG. 19</figref>, an USAC core decoder <b>2010</b> provides a downmix signal (DMX) <b>2012</b> to an MPS (MPEG Surround) decoder <b>2020</b>, which provides a first decoded audio signal <b>2022</b> and a second decoded audio signal <b>2024</b>. A Stereo SBR decoder <b>2030</b> receives the first decoded audio signal <b>2022</b> and the second decoded audio signal <b>2024</b> and provides, on the basis thereof a left bandwidth extended audio signal <b>2032</b> and a right bandwidth extended audio signal <b>2034</b>.
Before applying Stereo SBR, the second channel of the first element (mps_out_L_2[ ]) and the first channel of the second element (mps_out_R_1[ ]) are swapped to allow right-left Stereo SBR. After application of Stereo SBR, the second output channel of the first element (sbr_out_R_1[ ]) and the first channel of the second element (sbr_out_L_2[ ]) are swapped again to restore the input channel order.
A QCE decoder structure is illustrated in <figref idref="DRAWINGS">FIG. 20</figref>, which shows a QCE decoder schematics.
It should be noted that the block schematic diagram of <figref idref="DRAWINGS">FIG. 20</figref> is very similar to the block schematic diagram of <figref idref="DRAWINGS">FIG. 13</figref>, such that reference is also made to the above explanations. Moreover, it should be noted that some signal labeling has been added in <figref idref="DRAWINGS">FIG. 20</figref>, wherein reference is made to the definitions in this section. Moreover, a final resorting of the channels is shown, which is performed after the Stereo SBR.
<figref idref="DRAWINGS">FIG. 21</figref> shows a block schematic diagram of a Quad Channel Encoder <b>2200</b>, according to an embodiment of the present invention. In other words, a Quad Channel Encoder (Quad Channel Element), which may be considered as a Core Encoder Tool, is illustrated in <figref idref="DRAWINGS">FIG. 21</figref>.
The Quad Channel Encoder <b>2200</b> comprises a first Stereo SBR <b>2210</b>, which receives a first left-channel input signal <b>2212</b> and a second left channel input signal <b>2214</b>, and which provides, on the basis thereof, a first SBR payload <b>2215</b>, a first left channel SBR output signal <b>2216</b> and a first right channel SBR output signal <b>2218</b>. Moreover, the Quad Channel Encoder <b>2200</b> comprises a second Stereo SBR, which receives a second left-channel input signal <b>2222</b> and a second right channel input signal <b>2224</b>, and which provides, on the basis thereof, a first SBR payload <b>2225</b>, a first left channel SBR output signal <b>2226</b> and a first right channel SBR output signal <b>2228</b>.
The Quad Channel Encoder <b>2200</b> comprises a first MPEG-Surround-type (MPS 2-1-2 or Unified Stereo) multi-channel encoder <b>2230</b> which receives the first left channel SBR output signal <b>2216</b> and the second left channel SBR output signal <b>2226</b>, and which provides, on the basis thereof, a first MPS payload <b>2232</b>, a left channel MPEG Surround downmix signal <b>2234</b> and, optionally, a left channel MPEG Surround residual signal <b>2236</b>. The Quad Channel Encoder <b>2200</b> also comprises a second MPEG-Surround-type (MPS 2-1-2 or Unified Stereo) multi-channel encoder <b>2240</b> which receives the first right channel SBR output signal <b>2218</b> and the second right channel SBR output signal <b>2228</b>, and which provides, on the basis thereof, a first MPS payload <b>2242</b>, a right channel MPEG Surround downmix signal <b>2244</b> and, optionally, a right channel MPEG Surround residual signal <b>2246</b>.
The Quad Channel Encoder <b>2200</b> comprises a first complex prediction stereo encoding <b>2250</b>, which receives the left channel MPEG Surround downmix signal <b>2234</b> and the right channel MPEG Surround downmix signal <b>2244</b>, and which provides, on the basis thereof, a complex prediction payload <b>2252</b> and a jointly encoded representation <b>2254</b> of the left channel MPEG Surround downmix signal <b>2234</b> and the right channel MPEG Surround downmix signal <b>2244</b>. The Quad Channel Encoder <b>2200</b> comprises a second complex prediction stereo encoding <b>2260</b>, which receives the left channel MPEG Surround residual signal <b>2236</b> and the right channel MPEG Surround residual signal <b>2246</b>, and which provides, on the basis thereof, a complex prediction payload <b>2262</b> and a jointly encoded representation <b>2264</b> of the left channel MPEG Surround downmix signal <b>2236</b> and the right channel MPEG Surround downmix signal <b>2246</b>.
The Quad Channel Encoder also comprises a first bitstream encoding <b>2270</b>, which receives the jointly encoded representation <b>2254</b>, the complex prediction payload <b>2252</b><i>m </i>the MPS payload <b>2232</b> and the SBR payload <b>2215</b> and provides, on the basis thereof, a bitstream portion representing a first channel pair element. The Quad Channel Encoder also comprises a second bitstream encoding <b>2280</b>, which receives the jointly encoded representation <b>2264</b>, the complex prediction payload <b>2262</b>, the MPS payload <b>2242</b> and the SBR payload <b>2225</b> and provides, on the basis thereof, a bitstream portion representing a first channel pair element.
14. Implementation Alternatives
Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some one or more of the most important method steps may be executed by such an apparatus.
The inventive encoded audio signal can be stored on a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.
Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and/or non-transitionary.
A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are advantageously performed by any hardware apparatus.
The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
15. Conclusions
In the following, some conclusions will be provided.
The embodiments according to the invention are based on the consideration that, to account for signal dependencies between vertically and horizontally distributed channels, four channels can be jointly coded by hierarchically combining joint stereo coding tools. For example, vertical channel pairs are combined using MPS 2-1-2 and/or unified stereo with band-limited or full-band residual coding. In order to satisfy perceptual requirements for binaural unmasking, the output downmixes are, for example, jointly coded by use of complex prediction in the MDCT domain, which includes the possibility of left-right and mid-side coding. If residual signals are present, they are horizontally combined using the same method.
Moreover, it should be noted that embodiments according to the invention overcome some or all of the disadvantages of conventional technology. Embodiments according to the invention are adapted to the 3D audio context, wherein the loudspeaker channels are distributed in several height layers, resulting in a horizontal and vertical channel pairs. It has been found the joint coding of only two channels as defined in USAC is not sufficient to consider the spatial and perceptual relations between channels. However, this problem is overcome by embodiments according to the invention.
Moreover, conventional MPEG surround is applied in an additional pre-/post processing step, such that residual signals are transmitted individually without the possibility of joint stereo coding, e.g., to explore dependencies between left and right radical residual signals. In contrast, embodiments according to the invention allow for an efficient encoding/decoding by making use of such dependencies.
To further conclude, embodiments according to the invention create an apparatus, a method or a computer program for encoding and decoding as described herein.
While this invention has been described in terms of several embodiments, there are alterations, permutations, and equivalents which fall within the scope of this invention. It should also be noted that there are many alternative ways of implementing the methods and compositions of the present invention. It is therefore intended that the following appended claims be interpreted as including all such alterations, permutations and equivalents as fall within the true spirit and scope of the present invention.
REFERENCES
<ul id="ul0004" list-style="none"><li id="ul0004-0001" num="0229">[1] ISO/IEC 23003-3: 2012—Information Technology—MPEG Audio Technologies, Part 3: Unified Speech and Audio Coding;</li><li id="ul0004-0002" num="0230">[2] ISO/IEC 23003-1: 2007—Information Technology—MPEG Audio Technologies, Part 1: MPEG Surround</li></ul>
Contents7
23 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23
Every citation, both waysCites: the store holds 97 of 98
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US11451419B2 | Cited by | United States of America | Applicant |
| US10741188B2 | Cited by | United States of America | Applicant |
| US10770080B2 | Cited by | United States of America | Applicant |
| US11488610B2 | Cited by | United States of America | Applicant |
| US11929082B2 | Cited by | United States of America | Applicant |
| US12380899B2 | Cited by | United States of America | Applicant |
| US11657826B2 | Cited by | United States of America | Applicant |
| US12273221B2 | Cited by | United States of America | Applicant |
| EP1527655A2 | Cites | European Patent Office (EPO) | Applicant |
| US2005157883A1 | Cites | United States of America | Applicant |
| US2006190247A1 | Cites | United States of America | Applicant |
| US2006233379A1 | Cites | United States of America | Applicant |
| TW200627380A | Cites | Taiwan Province of China | Applicant |
| US2007067162A1 | Cites | United States of America | Applicant |
| WO2007111568A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2007174063A1 | Cites | United States of America | Applicant |
| US2008004883A1 | Cites | United States of America | Applicant |
| WO2009078681A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2009141775A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2009164223A1 | Cites | United States of America | Search report |
| JP2009508433A | Cites | Japan | Applicant |
| US2010027819A1 | Cites | United States of America | Applicant |
| TW201007695A | Cites | Taiwan Province of China | Applicant |
| US2010211400A1 | Cites | United States of America | Applicant |
| US2010228554A1 | Cites | United States of America | Applicant |
| US2010284550A1 | Cites | United States of America | Applicant |
| US2010332239A1 | Cites | United States of America | Applicant |
| JP2010540985A | Cites | Japan | Applicant |
| US2011046964A1 | Cites | United States of America | Search report |
| JP2011066868A | Cites | Japan | Applicant |
| US2011178810A1 | Cites | United States of America | Applicant |
| US2011200198A1 | Cites | United States of America | Applicant |
| US2011224994A1 | Cites | United States of America | Search report |
| JP2011501230A | Cites | Japan | Applicant |
| US2012002818A1 | Cites | United States of America | Applicant |
| KR20120029494A | Cites | Republic of Korea | Applicant |
| US2012070007A1 | Cites | United States of America | Applicant |
| US2012130722A1 | Cites | United States of America | Search report |
| WO2012158333A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2012170385A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2012275607A1 | Cites | United States of America | Applicant |
| US2012275609A1 | Cites | United States of America | Applicant |
| US2013030819A1 | Cites | United States of America | Applicant |
| US2013108077A1 | Cites | United States of America | Applicant |
| US2013124751A1 | Cites | United States of America | Applicant |
| US2013138446A1 | Cites | United States of America | Applicant |
| JP2013508770A | Cites | Japan | Applicant |
| WO2014168439A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2015162012A1 | Cites | United States of America | Search report |
| US2016071522A1 | Cites | United States of America | Search report |
| EP2194526A1 | Cites | European Patent Office (EPO) | Applicant |
| RU2449387C2 | Cites | Russian Federation | Applicant |
| GB2485979A | Cites | United Kingdom | Applicant |
| TW309691B | Cites | Taiwan Province of China | Applicant |
| US5717764A | Cites | United States of America | Applicant |
| US5970152A | Cites | United States of America | Applicant |
| US7359854B2 | Cites | United States of America | Applicant |
| US7668722B2 | Cites | United States of America | Applicant |
| US8208641B2 | Cites | United States of America | Applicant |
| US8218775B2 | Cites | United States of America | Applicant |
| US8255228B2 | Cites | United States of America | Applicant |
| US8948404B2 | Cites | United States of America | Applicant |
| TWI303411B | Cites | Taiwan Province of China | Applicant |
| US20050157883A1 | Cites | United States of America | Applicant |
| US20060190247A1 | Cites | United States of America | Applicant |
| US20060233379A1 | Cites | United States of America | Applicant |
| US20070067162A1 | Cites | United States of America | Applicant |
| US20070174063A1 | Cites | United States of America | Applicant |
| US20080004883A1 | Cites | United States of America | Applicant |
| US20090164223A1 | Cites | United States of America | Search report |
| US20100027819A1 | Cites | United States of America | Applicant |
| US20100211400A1 | Cites | United States of America | Applicant |
| US20100228554A1 | Cites | United States of America | Applicant |
| US20100284550A1 | Cites | United States of America | Applicant |
| US20100332239A1 | Cites | United States of America | Applicant |
| US20110046964A1 | Cites | United States of America | Search report |
| US20110178810A1 | Cites | United States of America | Applicant |
| US20110200198A1 | Cites | United States of America | Applicant |
| US20110224994A1 | Cites | United States of America | Search report |
| US20120002818A1 | Cites | United States of America | Applicant |
| US20120070007A1 | Cites | United States of America | Applicant |
| US20120130722A1 | Cites | United States of America | Search report |
| US20120275607A1 | Cites | United States of America | Applicant |
| US20120275609A1 | Cites | United States of America | Applicant |
| US20130030819A1 | Cites | United States of America | Applicant |
| US20130108077A1 | Cites | United States of America | Applicant |
| US20130124751A1 | Cites | United States of America | Applicant |
| US20130138446A1 | Cites | United States of America | Applicant |
| US20150162012A1 | Cites | United States of America | Search report |
| US20160071522A1 | Cites | United States of America | Search report |
| EP1527655 | Cites | European Patent Office (EPO) | Applicant |
| JP2009508433 | Cites | Japan | Applicant |
| JP2010540985 | Cites | Japan | Applicant |
| JP2011501230 | Cites | Japan | Applicant |
| JP2011066868 | Cites | Japan | Applicant |
| JP2013508770 | Cites | Japan | Applicant |
| KR1020120029494 | Cites | Republic of Korea | Applicant |
| RU2449387 | Cites | Russian Federation | Applicant |
| TW309691 | Cites | Taiwan Province of China | Applicant |
| TWI303411 | Cites | Taiwan Province of China | Applicant |
80 members in 19 offices
Priority claims22
| Document | Office | Kind | Date |
|---|---|---|---|
| 13177376 | European Patent Office (EPO) | A | |
| 13177376 | European Patent Office (EPO) | A | |
| 13177376 | European Patent Office (EPO) | – | |
| 13189305 | European Patent Office (EPO) | A | |
| 13189305 | European Patent Office (EPO) | A | |
| 13189305 | European Patent Office (EPO) | – | |
| 2014064915 | European Patent Office (EPO) | W | |
| 2014064915 | European Patent Office (EPO) | W | |
| PCTEP2014064915 | World Intellectual Property Organization (WIPO) | – | |
| 201615004661 | United States of America | A | |
| 201615004661 | United States of America | A | |
| 201615167072 | United States of America | A | |
| 13177376 | – | – | – |
| 13189305 | – | – | – |
| 15004661 | – | – | – |
| EP20130177376 | – | – | – |
| EP20130189305 | – | – | – |
| PCTEP2014064915 | – | – | – |
| PCTEP2014064915 | – | – | – |
| US201615004661 | – | – | – |
| US201615167072 | – | – | – |
| WO2014EP64915 | – | – | – |
Members80
| Document | Office | Kind | |
|---|---|---|---|
| EP2830051A2 | European Patent Office (EPO) | A2 | |
| EP2830052A1 | European Patent Office (EPO) | A1 | |
| CA2917770A1 | Canada | A1 | |
| CA2918237A1 | Canada | A1 | |
| WO2015010926A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010934A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2830051A3 | European Patent Office (EPO) | A3 | |
| TW201514972A | Taiwan Province of China | A | |
| TW201514973A | Taiwan Province of China | A | |
| AR097011A1 | Argentina | A1 | |
| AR097012A1 | Argentina | A1 | |
| SG11201600468SA | Singapore | A | |
| AU2014295282A1 | Australia | A1 | |
| AU2014295360A1 | Australia | A1 | |
| KR20160033777A | Republic of Korea | A | |
| KR20160033778A | Republic of Korea | A | |
| MX2016000939A | Mexico | A | |
| MX2016000858A | Mexico | A | |
| CN105580073A | China | A | |
| CN105593931A | China | A | |
| EP3022734A1 | European Patent Office (EPO) | A1 | |
| EP3022735A1 | European Patent Office (EPO) | A1 | |
| TWI544479B | Taiwan Province of China | B | |
| US2016247508A1 | United States of America | A1 | |
| US2016247509A1 | United States of America | A1 | |
| TWI550598B | Taiwan Province of China | B | |
| US2016275957A1 | United States of America | A1 | |
| JP2016529544A | Japan | A | |
| JP2016530788A | Japan | A | |
| JP6117997B2 | Japan | B2 | |
| ZA201601078B | South Africa | B | |
| BR112016001137A2 | Brazil | A2 | |
| BR112016001141A2 | Brazil | A2 | |
| AU2014295282B2 | Australia | B2 | |
| EP3022734B1 | European Patent Office (EPO) | B1 | |
| RU2016105702A | Russian Federation | A | |
| RU2016105703A | Russian Federation | A | |
| ZA201601080B | South Africa | B | |
| EP3022735B1 | European Patent Office (EPO) | B1 | |
| AU2014295360B2 | Australia | B2 | |
| PT3022734T | Portugal | T | |
| PT3022735T | Portugal | T | |
| ES2649194T3 | Spain | T3 | |
| ES2650544T3 | Spain | T3 | |
| KR101823278B1 | Republic of Korea | B1 | |
| PL3022734T3 | Poland | T3 | |
| PL3022735T3 | Poland | T3 | |
| KR101823279B1 | Republic of Korea | B1 | |
| US9940938B2This record | United States of America | B2 | |
| US9953656B2 | United States of America | B2 | |
| JP6346278B2 | Japan | B2 | |
| MX357667B | Mexico | B | |
| MX357826B | Mexico | B | |
| RU2666230C2 | Russian Federation | C2 | |
| US10147431B2 | United States of America | B2 | |
| RU2677580C2 | Russian Federation | C2 | |
| US2019108842A1 | United States of America | A1 | |
| US2019378522A1 | United States of America | A1 | |
| CN105580073B | China | B | |
| CN105593931B | China | B | |
| CN111105805A | China | A | |
| CN111128205A | China | A | |
| CN111128206A | China | A | |
| US10741188B2 | United States of America | B2 | |
| US10770080B2 | United States of America | B2 | |
| CA2917770C | Canada | C | |
| MY181944A | Malaysia | A | |
| US2021056979A1 | United States of America | A1 | |
| US2021233543A1 | United States of America | A1 | |
| CA2918237C | Canada | C | |
| BR112016001141B1 | Brazil | B1 | |
| US11488610B2 | United States of America | B2 | |
| BR112016001137B1 | Brazil | B1 | |
| US11657826B2 | United States of America | B2 | |
| US2024029744A1 | United States of America | A1 | |
| CN111128206B | China | B | |
| CN111105805B | China | B | |
| US12380899B2 | United States of America | B2 | |
| US20260024535A1 | United States of America | A1 | |
| CN111128205B | China | B |
77 transactions on the USPTO file
Allowed after 1 non-final rejection.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Payment of Maintenance Fee, 4th Year, Large EntityM1551 | M1551 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Email NotificationEML_NTR | EML_NTR | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Email NotificationEML_NTR | EML_NTR | |
| Printer Rush- No mailingTCPB | TCPB | |
| Mailing Corrected Notice of AllowabilityMCNOA | MCNOA | |
| Corrected Notice of AllowabilityCNOA | CNOA | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Pubs Case Remand to TCPUBTC | PUBTC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Mail PUBS Letter Withdrawing a Notice Requiring Inventors Oath or DeclarationMM327-W | MM327-W | |
| PUBS Letter Withdrawing a Notice Requiring Inventors Oath or DeclarationM327-W | M327-W | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Email NotificationEML_NTR | EML_NTR | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Email NotificationEML_NTR | EML_NTR | |
| Application ready for PDX access by participating foreign officesCCRDY | CCRDY | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Priority document has successfully retrieved via PDX/DASPD.RECVD | PD.RECVD | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Email NotificationEML_NTR | EML_NTR | |
| Application Is Now CompleteCOMP | COMP | |
| Application Is Now CompleteCOMP | COMP | |
| Filing ReceiptFLRCPT.O | FLRCPT.O | |
| Sent to Classification ContractorPGPC | PGPC | |
| FITF set to YES - revise initial settingFTFS | FTFS | |
| Cleared by OIPE CSRL194 | L194 | |
| Patent Term Adjustment - Ready for ExaminationPTA.RFE | PTA.RFE | |
| Request from applicant for the USPTO to retrieve the Priority DocumentPDREQUST | PDREQUST | |
| PTO/SB/69-Authorize EPO Access to Search ResultsSREXR141 | SREXR141 | |
| Applicants have given acceptable permission for participating foreignAPPERMS | APPERMS | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Entity Status Set To Undiscounted (Initial Default Setting or Status Change)BIG. | BIG. | |
| Initial Exam Team nnIEXX | IEXX |
4 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| Maintenance fee paymentMAFP | MAFP | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS |
Numbers
- Publication
- 09940938
- Publication, DOCDB
- 9940938
- Publication, EPODOC
- US9940938
- Application
- 15167072
- Application, DOCDB
- 201615167072
- Application, EPODOC
- US201615167072
Titles
- English
- Audio encoder, audio decoder, methods and computer program using jointly encoded residual signals
Patent term adjustment
- Applicant delay
- −123 days
- Net adjustment
- 0 days
Classification
- CPC, 8
- G10L19/008
- G10L19/0017
- G10L21/038
- H04S3/008
- H04S7/30
- H04S2400/01
- H04S2400/03
- H04S2420/03
- IPC, 6
- G10L19 008
- G10L19 00
- G10L19 038
- H04S7 00
- H04S3 00
- G10L21 038
- USPC, 2
- 704500000
- 001001000