US20180098174A1

System and method for capturing, encoding, distributing, and decoding immersive audio

Claim Score by NHIP

Read claim 1, the broadest

Abstract

A sound field coding system and method that provides flexible capture, distribution, and reproduction of immersive audio recordings encoded in a generic digital audio format compatible with standard two-channel or multi-channel reproduction systems. This end-to-end system and method mitigates any impractical need for standard multi-channel microphone array configurations in consumer mobile devices such as smart phones or cameras. The system and method capture and spatially encode two-channel or multi-channel immersive audio signals that are compatible with legacy playback systems from flexible multi-channel microphone array configurations.

US20180098174A1, drawing sheet 1
Sheet 1 of 25

Term

9.3 yearsto projected expiry

Projected expiry 29 January 2036, counted from filing; an application has no term until it is granted.

  1. Priority and filed
  2. Published
  3. Today
  4. Projected expiry

14 claims: 2 independent, 12 dependent

  1. 1
    Broadest claimClaim Score 38, average(NHIP)A method for processing a plurality of capture microphone signals, comprising:selecting a capture microphone configuration having a plurality of capture microphones for capturing sound from at least one audio source, the capture microphone configuration defining a microphone directivity for each of the plurality of capture microphones relative to a reference direction;selecting a virtual microphone configuration having a plurality of virtual microphones for encoding spatial information about a position of the at least one audio source relative to the reference direction, the virtual microphone configuration defining a virtual microphone directivity for each of the virtual microphones relative to the reference direction;adapting the capture microphone configuration based on detection of an adverse condition for microphone performance for at least one of the plurality of capture microphones to obtain an adapted capture microphone configuration;calculating spatial encoding coefficients based on the adapted capture microphone configuration and on the virtual microphone configuration;and converting the plurality of capture microphone signals into a Spatially Encoded Signal (SES) including virtual microphone signals;wherein each of the virtual microphone signals is obtained by combining the capture microphone signals using the spatial encoding coefficients.
  2. 11
    A method for processing a plurality of capture microphone signals, comprising:decomposing the plurality of capture microphone signals into a plurality of direct components and a plurality of diffuse components;selecting a first capture microphone configuration for the direct components having a first plurality of capture microphones for capturing sound from at least one audio source, the first capture microphone configuration defining a first microphone directivity for each of the first plurality of capture microphones relative to a first reference direction;selecting a first virtual microphone configuration for the direct components having a first plurality of virtual microphones for encoding spatial information about a position of the at least one audio source relative to the reference direction, the first virtual microphone configuration defining a first virtual microphone directivity for each of the virtual microphones relative to the first reference direction;selecting a second capture microphone configuration for the diffuse components having a second plurality of capture microphones for capturing sound from at least one audio source, the second capture microphone configuration defining a second microphone directivity for each of the second plurality of capture microphones relative to a second reference direction;selecting a second virtual microphone configuration for the diffuse components having a second plurality of virtual microphones for encoding spatial information about a position of the at least one audio source relative to the second reference direction, the second virtual microphone configuration defining a second virtual microphone directivity for each of the virtual microphones relative to the second reference direction;calculating spatial encoding coefficients for the direct components based on the first capture microphone configuration for the direct components and on the first virtual microphone configuration;calculating spatial encoding coefficients for the diffuse components based on the second capture microphone configurations for the diffuse components and on the second virtual microphone configuration;converting the plurality of direct components into a direct-component Spatially Encoded Signal (SES) including virtual microphone signals for the direct components, wherein each of the virtual microphone signals for the direct components is obtained by combining the direct components using the spatial encoding coefficients for the direct components;converting the plurality of diffuse components into a diffuse-component Spatially Encoded Signal (SES) including virtual microphone signals for the diffuse components, wherein each of the virtual microphone signals for the diffuse components is obtained by combining the diffuse components using the spatial encoding coefficients for the diffuse components;and combining the direct-component Spatially Encoded Signal and the diffuse-component Spatially Encoded Signal to form an output Spatially Encoded Signal.