CA2918166C

Apparatus and method for efficient object metadata coding

Abstract

An apparatus (100) for generating one or more audio channels is provided. The apparatus (100) comprises a metadata decoder (110) for receiving one or more compressed metadata signals. Each of the one or more compressed metadata signals comprises a plurality of first metadata samples. The first metadata samples of each of the one or more compressed metadata signals indicate information associated with an audio object signal of one or more audio object signals. The metadata decoder (110) is configured to generate one or more reconstructed metadata signals, so that each of the one or more reconstructed metadata signals comprises the first metadata samples of one of the one or more compressed metadata signals and further comprises a plurality of second metadata samples. Moreover, the metadata decoder (110) is configured to generate each of the second metadata samples of each reconstructed metadata signal of the one or more reconstructed metadata signals depending on at least two of the first metadata samples of said reconstructed metadata signal. Moreover, the apparatus (100) comprises an audio channel generator (120) for generating the one or more audio channels depending on the one or more audio object signals and depending on the one or more reconstructed metadata signals. Furthermore, an apparatus for generating encoded audio information comprising one or more encoded audio signals and one or more compressed metadata signals is provided.

CA2918166C, drawing sheet 1
Sheet 1 of 16

Term

7.8 yearsleft in the term

Expires 16 July 2034.

  1. Priority and filed
  2. Granted
  3. Today
  4. Expires

17 claims: 4 independent, 13 dependent

  1. 1
    Claim? 1. An apparatus for generating one or more audio channels, wherein the apparatus comprises:a metadata decoder for receiving one or more compressed metadata signals, wherein each of the one or more compressed metadata signals comprises a plurality of first metadata samples, wherein the first metadata samples of each of the one or more compressed metadata signals indicate information associated with an audio object signal of one or more audio object signals, wherein the metadata decoder is configured to generate one or more reconstructed metadata signals, so that each reconstructed metadata signal of the one or more reconstructed metadata signals comprises the first metadata samples of a compressed metadata signal of the one or more compressed metadata signals, said reconstructed metadata signal being associated with said compressed metadata signal, and further comprises a plurality of second metadata samples, wherein the metadata decoder is configured to generate the second metadata samples of each of the one or more reconstructed metadata signals by generating a plurality of approximated metadata samples for said reconstructed metadata signal, wherein the metadata decoder is configured to generate each of the plurality of approximated metadata samples depending on at least two of the first metadata samples of said reconstructed metadata signal, and an audio channel generator for generating the one or more audio channels depending on the one or more audio object signals and depending on the one or more reconstructed metadata signals, wherein the metadata decoder is configured to receive a plurality of difference values for a compressed metadata signal of the one or more compressed metadata signals, and is configured to add each of the plurality of difference values to one of the approximated metadata samples of the reconstructed metadata signal being associated with said compressed metadata signal to obtain the second metadata samples of said reconstructed metadata signal. .
  2. 4
    An apparatus according to any one of claims 1 to 3, 35 wherein at least one of the one or more reconstructed metadata signals comprises position information on one of the one or more audio object signals, or comprises a scaled representation of the position information on said one of the one or more audio object signals, and CA 02918166 2016-01-13 wherein the audio channel generator is configured to generate at least one of the one or more audio channels depending on said one of the one or more audio object signals and depending on said position information.
  3. 5
    An apparatus according to any one of claims 1 to 4. wherein at least one of the one or more reconstructed metadata signals comprises a volume of one of the one or more audio object signals, or comprises a scaled 10 representation of the volume of said one of the one or more audio object signals, and wherein the audio channel generator is configured to generate at least one of the one or more audio channels depending on said one of the one or more audio 15 object signals and depending on said volume.
  4. 6
    An apparatus according to any one of claims 1 to 5, wherein the apparatus is configured to receive random access information, wherein, for each compressed metadata signal of the one or more compressed metadata signals, the random 20 access information indicates an accessed signal portion of said compressed metadata signal, wherein at least one other signal portion of said metadata signal is not indicated by said random access information, and wherein the metadata decoder is configured to generate one of the one or more reconstructed metadata signals depending on the first metadata samples of said accessed signal portion of 25 said compressed metadata signal, but not depending on any other first metadata samples of any other signal portion of said compressed metadata signal.
  5. 7
    An apparatus for generating encoded audio information comprising one or more encoded audio signals and one or more compressed metadata signals, wherein 30 the apparatus comprises:a metadata encoder for receiving one or more original metadata signals, wherein each of the one or more original metadata signals comprises a plurality of metadata samples, wherein the metadata samples of each of the one or more 35 original metadata signals indicate information associated with an audio object signal of one or more audio object signals, wherein the metadata encoder is configured to generate the one or more compressed metadata signals, so that CA 029181«« 2016-01-13 each compressed metadata signal of the one or more compressed metadata signals comprises a first group of two or more of the metadata samples of an original metadata signal of the one or more original metadata signals, said compressed metadata signal being associated with said original metadata signal, 5 and so that said compressed metadata signal does not comprise any metadata sample of a second group of another two or more of the metadata samples of said one of the original metadata signals, and an audio encoder for encoding the one or more audio object signals to obtain the 10 one or more encoded audio signals, wherein each of the metadata samples, that is comprised by an original metadata signal of the one or more original metadata signals and that is also comprised by the compressed metadata signal, which is associated with said original metadata 15 signal, is one of a plurality of first metadata samples, wherein each of the metadata samples, that is comprised by an original metadata signal of the one or more original metadata signals and that is not comprised by the compressed metadata signal, which is associated with said original metadata 20 signal, is one of a plurality of second metadata samples, wherein the metadata encoder is configured to generate an approximated metadata sample for each of a plurality of the second metadata samples of one of the original metadata signals by conducting a linear interpolation depending on at 25 least two of the first metadata samples of said one of the one or more original metadata signals, and wherein the metadata encoder is configured to generate a difference value for each second metadata sample of said plurality of the second metadata samples of 30 said one of the one or more original metadata signals, so that said difference value indicates a difference between said second metadata sample and the approximated metadata sample of said second metadata sample.
  6. 10
    An apparatus according to any one of claims 7 to 9, 20 wherein at least one of the one or more original metadata signals comprises position information on one of the one or more audio object signals, or comprises a scaled representation of the position information on said one of the one or more audio object signals, and 25 wherein the metadata encoder is configured to generate at least one of the one or more compressed metadata signals depending on said at least one of the one or more original metadata signals.
  7. 11
    An apparatus according to any one of claims 7 to 10. wherein at least one of the one or more original metadata signals comprises a volume of one of the one or more audio object signals, or comprises a scaled representation of the volume of said one of the one or more audio object signals, and wherein the metadata encoder is configured to generate at least one of the one or more compressed metadata signals depending on said at least one of the one or more original metadata signals. CA 029181«« 2016-01-13
  8. 12
    A system, comprising:an apparatus according to any one of claims 7 to 11 for generating encoded audio 5 information comprising one or more encoded audio signals and one or more compressed metadata signals, and an apparatus according to any one of claims 1 to 6 for receiving the one or more encoded audio signals and the one or more compressed metadata signals, and for 10 generating one or more audio channels depending on the one or more encoded audio signals and depending on the one or more compressed metadata signals.
  9. 13
    A method for generating one or more audio channels, wherein the method comprises:receiving one or more compressed metadata signals, wherein each of the one or more compressed metadata signals comprises a plurality of first metadata samples, wherein the first metadata samples of each of the one or more compressed metadata signals indicate information associated with an audio object 20 signal of one or more audio object signals, generating one or more reconstructed metadata signals, so that each reconstructed metadata signal of the one or more reconstructed metadata signals comprises the first metadata samples of a compressed metadata signal of the one 25 or more compressed metadata signals, said reconstructed metadata signal being associated with said compressed metadata signal, and further comprises a plurality of second metadata samples, wherein generating the one or more reconstructed metadata signals comprises generating the second metadata samples of each of the one or more reconstructed metadata signals by generating 30 a plurality of approximated metadata samples for said reconstructed metadata signal, wherein generating each of the plurality of approximated metadata samples is conducted depending on at least two of the first metadata samples of said reconstructed metadata signal, and 35 generating the one or more audio channels depending on the one or more audio object signals and depending on the one or more reconstructed metadata signals, CA 02918166 2016-01-13 wherein the method further comprises receiving a plurality of difference values for a compressed metadata signal of the one or more compressed metadata signals, and adding each of the plurality of difference values to one of the approximated metadata samples of the reconstructed metadata signal being associated with said 5 compressed metadata signal to obtain the second metadata samples of said reconstructed metadata signal.
  10. 14
    A method for generating encoded audio information comprising one or more encoded audio signals and one or more compressed metadata signals, wherein 10 the method comprises:receiving one or more original metadata signals, wherein each of the one or more original metadata signals comprises a plurality of metadata samples, wherein the metadata samples of each of the one or more original metadata signals indicate
  11. 15
    15 information associated with an audio object signal of one or more audio object signals, generating the one or more compressed metadata signals, so that each compressed metadata signal of the one or more compressed metadata signals 20 comprises a first group of two or more of the metadata samples of an original metadata signal of the one or more original metadata signals, said compressed metadata signal being associated with said original metadata signal, and so that said compressed metadata signal does not comprise any metadata sample of a second group of another two or more of the metadata samples of said one of the 25 original metadata signals, and encoding the one or more audio object signals to obtain the one or more encoded audio signals, 30 wherein each of the metadata samples, that is comprised by an original metadata signal of the one or more original metadata signals and that is also comprised by the compressed metadata signal, which is associated with said original metadata signal, is one of a plurality of first metadata samples, 35 wherein each of the metadata samples, that is comprised by an original metadata signal of the one or more original metadata signals and that is not comprised by the compressed metadata signal, which is associated with said original metadata signal, is one of a plurality of second metadata samples, CA 029181«« 2016-01-13 wherein the method further comprises generating an approximated metadata sample for each of a plurality of the second metadata samples of one of the original metadata signals by conducting a linear interpolation depending on at least 5 two of the first metadata samples of said one of the one or more original metadata signals, and wherein the method further comprises generating a difference value for each second metadata sample of said plurality of the second metadata samples of said 10 one of the one or more original metadata signals, so that said difference value indicates a difference between said second metadata sample and the approximated metadata sample of said second metadata sample. 15. A computer-readable medium having computer-readable code stored thereon to 15 preform the method according to claim 13 or 14 when being executed on a computer or signal processor.
  12. 16
    An apparatus for encoding audio input data to obtain audio output data, comprising:an input interface for receiving a plurality of audio channels, a plurality of audio objects and metadata related to one or more of the plurality of audio objects. a mixer for mixing the plurality of objects and the plurality of channels to obtain a 25 plurality of pre-mixed channels, each pre-mixed channel comprising audio data of a channel and audio data of at least one object, and an apparatus according to any one of claims 7 to 11, 30 wherein the audio encoder of the apparatus according to any one of claims 7 to 11 is a core encoder for core encoding core encoder input data, and wherein the metadata encoder of the apparatus according to any one of claims 7 to 11 is a metadata compressor for compressing the metadata related to the one 35 or more of the plurality of audio objects.
  13. 17
    An apparatus for decoding encoded audio data, comprising:CA 2418166 2017-06-27 an input interface for receiving the encoded audio data, the encoded audio data comprising a plurality of encoded channels or a plurality of encoded objects or compressed metadata related to the plurality of encoded objects, and an apparatus according to any one of claims 1 to 6, wherein the metadata decoder of the apparatus according to any one of claims 1 to 6 is a metadata decompressor for decompressing the compressed metadata to obtain decompressed metadata, wherein the audio channel generator of the apparatus according to any one of claims 1 to 6 comprises a core decoder for decoding the plurality of encoded channels to obtain a plurality of decoded channels and for decoding the plurality of encoded objects to obtain a plurality of decoded objects, wherein the audio channel generator further comprises an object processor for processing the plurality of decoded objects using the decompressed metadata to obtain a number of output channels comprising audio data from the plurality of decoded objects and the from the plurality of decoded channels, and wherein the audio channel generator further comprises a post processor for converting the number of output channels into an output format.