EP3704875B1

Virtual rendering of object based audio over an arbitrary set of loudspeakers

Abstract

This record has no abstract on file.

EP3704875B1, drawing sheet 1
Sheet 1 of 23

Term

12.1 yearsleft in the term

Expires 24 October 2038.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

14 claims: 10 independent, 4 dependent

  1. 1
    A method (400) of rendering audio, the method comprising:deriving (402) an optimal set of a plurality of filters, wherein each of the plurality of filters is associated with a corresponding one of a plurality of loudspeakers, wherein deriving the optimal set of a plurality of filters includes: defining (404) a binaural error for an audio object as a function of the plurality of filters, wherein the binaural error is a difference between desired binaural signals related to at least one listener position and modeled binaural signals related to the at least one listener position, wherein the desired binaural signals are defined based on the audio object and the desired perceived position of the audio object, from the perspective of at least one listener position, and the modeled binaural signals are defined by modeling a playback of a plurality of rendered signals, through the plurality of loudspeakers having a plurality of loudspeaker positions, based on the at least one listener position, defining (406) an activation penalty for the audio object using the plurality of filters, wherein the activation penalty is a distance penalty which has the property that for the same amount of overall energy of the plurality of rendered signals, where overall means combining across all loudspeakers, the penalty increases when more of that energy is concentrated in loudspeakers of the plurality of loudspeakers whose nominal position is more distant from the desired perceived position of the audio object, and minimizing (408) a cost function with respect to the plurality of filters, wherein the cost function is a combination of the binaural error and the activation penalty for the plurality of filters;rendering (410) the audio object using the derived optimal set of a plurality of filters to generate a plurality of rendered signals;and outputting (412), by the plurality of loudspeakers, the plurality of rendered signals.
  2. 4
    The method of any one of claims 1-3, wherein the cost function is a combination function that is monotonically increasing in both A and B, wherein A corresponds to the binaural error and B corresponds to the activation penalty.
  3. 6
    The method of any one of claims 1-5, wherein the audio object is one of a plurality of audio objects, wherein the plurality of audio objects is rendered using the plurality of filters, and wherein each of the plurality of audio objects has an associated desired perceived position.
  4. 7
    The method of any one of claims 1-6, wherein the plurality of loudspeakers includes a first loudspeaker and a second loudspeaker, wherein the first loudspeaker has a nominal position that is at a first distance from the desired perceived position of the audio object, and wherein the second loudspeaker has a nominal position that is at a second distance from the desired perceived position of the audio object, wherein the first distance is greater than the second distance.
  5. 8
    The method of any one of claims 1-7, wherein the plurality of loudspeakers has a plurality of nominal loudspeaker positions, wherein each of the plurality of nominal loudspeaker positions is one of a first position and a second position, wherein the first position is an actual loudspeaker position of a corresponding one of the plurality of loudspeakers, and wherein the second position is other than the actual loudspeaker position.
  6. 9
    The method of any one of claims 1-8, wherein one of the plurality of loudspeakers has a nominal loudspeaker position, wherein the nominal loudspeaker position is derived by expanding one or more physical positions of the plurality of loudspeakers.
  7. 10
    The method of any one of claims 1-9, wherein the plurality of filters is independent of the audio object.
  8. 12
    The method of any one of claims 1-11, wherein the plurality of loudspeakers has a plurality of physical positions, wherein the plurality of physical positions are determined in a setup phase.
  9. 13
    A non-transitory computer readable medium storing a computer program that, when executed by a processor, controls an apparatus to execute processing including the method of any one of claims 1-11.
  10. 14
    An apparatus (300) for rendering audio, the apparatus comprising:a plurality of loudspeakers;and at least one processor, wherein the at least one processor is configured to derive an optimal set of a plurality of filters, wherein each of the plurality of filters is associated with a corresponding one of the plurality of loudspeakers, wherein deriving the optimal set of a plurality of filters includes: defining a binaural error for an audio object as function of the plurality of filters, wherein the binaural error is a difference between desired binaural signals related to at least one listener position and modeled binaural signals related to the at least one listener position, wherein the desired binaural signals are defined based on the audio object and the desired perceived position of the audio object, from the perspective of at least one listener position, and the modeled binaural signals are defined by modeling a playback of a plurality of rendered signals, through the plurality of loudspeakers having a plurality of loudspeaker positions, based on the at least one listener position, defining an activation penalty for the audio object using the plurality of filters, wherein the activation penalty is a distance penalty which has the property that for the same amount of overall energy of the plurality of rendered signals, where overall means combining across all loudspeakers, the penalty increases when more of that energy is concentrated in loudspeakers of the plurality of loudspeakers whose nominal position is more distant from the desired perceived position of the audio object, and minimizing a cost function with respect to the plurality of filters, wherein the cost function is a combination of the binaural error and the activation penalty for the plurality of filters;wherein the at least one processor is configured to render the audio object using the derived optimal set of a plurality of filters to generate a plurality of rendered signals, and wherein the plurality of loudspeakers is configured to output the plurality of rendered signals.