US7194084B2

System and method for stereo conferencing over low-bandwidth links

Summary by NHIP

Stereo conferencing over low-bandwidth links

The system encodes two spatially-separated sound field signals into a single audio stream while transmitting a relative temporal delay parameter. A decoder uses this delay to split the signal into multiple presentation channels that simulate speaker location based on the original sampling points.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

Systems and methods are disclosed for packet voice conferencing. An encoding system accepts two sound field signals, representing the same sound field sampled at two spatially-separated points. The relative delay between the two sound field signals is detected over a given time interval. The sound field signals are combined and then encoded as a single audio signal, e.g., by a method suitable for monophonic VoIP. The encoded audio payload and the relative delay are placed in one or more packets and sent to a decoding device via the packet network. The decoding device uses the relative delay to drive a playout splitter—once the encoded audio payload has been decoded, the playout splitter creates multiple presentation channels by inserting the transmitted relative delay in the decoded signal for one (or more) of the presentation channels. The listener thus perceives a speaker's voice as originating from a location related to the speaker's physical position at the other end of the conference. An advantage of these embodiments is that a pseudo-stereo conference can be conducted with virtually the same bandwidth as a monophonic conference.

US7194084B2, drawing sheet 1
Sheet 1 of 11

Term

Term ended

Expired 11 July 2020, 6.2 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

34 claims: 4 independent, 30 dependent

  1. 1
    Broadest claimClaim Score 61, broad(NHIP)An encoder comprising:a sound field signal encoder to create a digitally-encoded signal representing both a first and a second sound field signal;a stereo parameter estimator to estimate a relative temporal delay between the first sound field signal and the second sound field signal;and a packet formatter packetizing the digitally-encoded signal and a stereo decoding parameter based on the estimated relative temporal delay, the stereo decoding parameter including at least one of an explicit delay parameter, an explicit balance parameter, and an explicit arrival angle parameter.
  2. 11
    An encoder comprising:means for encoding a digital data block to represent a combination of first and second sound field signals concurrently-captured within a first time period, the first and second sound field signals representing a single sound field captured at two spatially-separated points;means for estimating, using the first and second sound field signals as captured in an approximate timeframe of the first time period, an explicit relative temporal delay between the first and second sound field signals;and means for encapsulating, in a packet format, the encoded digital data block and a stereo decoding parameter based on the relative temporal delay.
  3. 21
    A method comprising:digitally encoding a signal block to represent first and second sound field signals as concurrently-captured during a first time period, the first and second sound field signals representing a single sound field captured at two spatially-separated points;estimating a relative temporal delay between the first and second sound field signals within an approximate timeframe of the first time period;transmitting to a remote conferencing point, in packet format, both the encoded signal block and a stereo decoding parameter based on the estimated relative temporal delay, the stereo decoding parameter including at least one of an explicit delay parameter, an explicit balance parameter, and an explicit arrival angle parameter.
  4. 28
    An apparatus comprising a computer-readable medium containing computer instructions that, when executed, cause a processor or multiple communicating processors to perform a method comprising:digitally encoding a signal block to represent first and second sound field signals as concurrently-captured during a first time period, the first and second sound field signals representing a single sound field captured at two spatially-separated points;detecting a talkspurt represented in the sound field signals;estimating a relative temporal delay between the first and second sound field signals within an approximate timeframe of the first time period responsive to the detection of the talkspurt;transmitting to a remote conferencing point, in packet format, both the encoded signal block and a stereo decoding parameter based on the estimated relative temporal delay.