EP1565036A2

Late reverberation-based synthesis of auditory scenes

Abstract

A scheme for stereo and multi-channel synthesis of inter-channel correlation (ICC) (normalized cross-correlation) cues for parametric stereo and multi-channel coding. The scheme synthesizes ICC cues such that they approximate those of the original. For that purpose, diffuse audio channels are generated and mixed with the transmitted combined (e.g., sum) signal(s). The diffuse audio channels are preferably generated using relatively long filters with exponentially decaying Gaussian impulse responses. Such impulse responses generate diffuse sound similar to late reverberation. An alternative implementation for reduced computational complexity is proposed, where inter-channel level difference (ICLD), inter-channel time difference (ICTD), and ICC synthesis are all carried out in the domain of a single short-time Fourier transform (STFT), including the filtering for diffuse sound generation.

EP1565036A2, drawing sheet 1
Sheet 1 of 95

Term

Term ended

Projected expiry passed 4 February 2025, 1.6 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

10 claims: 3 independent, 7 dependent

  1. 1
    A method for synthesizing an auditory scene, comprising:processing at least one input channel to generate two or more processed input signals;filtering the at least one input channel to generate two or more diffuse signals;and combining the two or more diffuse signals with the two or more processed input signals to generate a plurality of output channels for the auditory scene.
  2. 8
    Apparatus for synthesizing an auditory scene, comprising:means for processing at least one input channel to generate two or more processed input signals;means for filtering the at least one input channel to generate two or more diffuse signals;and means for combining the two or more diffuse signals with the two or more processed input signals to generate a plurality of output channels for the auditory scene.
  3. 9
    Apparatus for synthesizing an auditory scene, comprising:a configuration of at least one time domain to frequency domain (TD-FD) converter and a plurality of filters, the configuration adapted to generate two or more processed FD input signals and two or more diffuse FD signals from at least one TD input channel;two or more combiners adapted to combine the two or more diffuse FD signals with the two or more processed FD input signals to generate a plurality of synthesized FD signals;and two or more frequency domain to time domain (FD-TD) converters adapted to convert the synthesized FD signals into a plurality of TD output channels for the auditory scene.