EP0779732A2

Multi-point voice conferencing system over a wide area network

Abstract

An interactive network system (100) communicates speech and associated information among a plurality of participants at different sites (104, 106). An example of the associated information is lip synch image information related to the speech. The system contains a speech server (110) for managing data streams set by the participants. Each participant uses a multimedia computer (114) and a modem (122) to connect to the network. Because many modems have a low bit rate, it is important to compress the speech and associated information. The server (110) receives the data streams from at least two participants and contains means (200) for combining these data streams into a single data stream having a bit rate that can be handled by the modem of the third participant. As a result, a plurality of participants can conduct speech and image communication using the network.

EP0779732A2, drawing sheet 1
Sheet 1 of 45

Term

Term ended

Projected expiry passed 6 December 2016, 9.8 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

21 claims: 4 independent, 17 dependent

  1. 1
    A system(100) for a plurality of users (104-106) conducting voice and image communication on a wide area network (112), each user being associated with a computer (114), a network access device (122) having a maximum data communication speed for connecting said computer to said network, a microphone (118) and a loudspeaker (119,120), said microphone generating speech signals in response to audio signals, which are then converted into digital speech data, said system comprising:a speech server (110) connected to said network for managing data streams sent by user computers associated with said users;an encoder (200) for running on each one of said user computers, said encoder (200) comprising: a compressor for compressing said speech data received by said encoder into compressed data, said compressor including means (222) for generating a plurality of linear predictive coding (LPC) parameters;and a bit stream encoder (232) for encoding said compressed data into an encoded data stream having a data rate below said maximum data communication speed;said bit stream encoder serving to generate a first encoded data stream having a first data rate from said speech data of a first user computer and to generate a second encoded data stream having a second data rate from said speech data of a second user computer and said server including means for combining said first and said second encoded data streams into a combined data stream having a data rate below said maximum data communication speed while said first and said second data rates have a sum above said maximum data communication speed;and a decoder (300) running on a third user computer, said decoder (300) comprising: means for receiving said combined data stream;and means for reconstructing said audio signals received by said microphones associated said first and said second user computers using information from said combined data stream.
  2. 5
    A system according to any one of claims 1 to 4, wherein said encoder (200) further comprises a voice morph means (214) for altering said speech signal.
  3. 10
    A system according to any one of claims 1 to 9, wherein said encoder (200) further comprises means for determining a first formant frequency of said speech signals using said LPC parameters;means for determining a second formant frequency of said speech signals using said LPC parameters;and each of said user computers further comprising means for displaying a lower lip position and an upper lip position using said first and second formant frequencies.
  4. 15
    A system according to claims 14, wherein said filter is a finite impulse response low-pass filter.
  5. 18
    A system according to any one of the preceding claims and further comprising means for determining a silence state in a surrounding of one of said user computers, said silence state being used by said compressor as an input for compressing said speech data, said means for determining said silence state comprising:means for generating a first source signal which is substantially a chirp signal and for causing said microphone to play said first source signal as a first audio signal;means for generating a first digital signal based on said first audio signal received by said loudspeaker;a filter for processing said first digital signal matched to said chirp signal;means for determining a bulk delay as a time when said processed first digital signal has a maximum value;means for generating a second source signal which is substantially a white noise and for causing said microphone to play said second source signal as a second audio signal;means for generating a second digital signal based on said second audio signal received by said loudspeaker;means for determining a cross-correlation function of said second source signal and said second digital signal;means for generating an auto-correlation of said second source signal;means for determining a finite impulse response as a function of said cross-correlation and said auto-correlation function;means for determining an echo cancellation energy using said finite impulse response and said bulk delay;mans for measuring acoustic energy received by said microphone;and means for measuring background noise energy;said surroundings being classified to be in said silence state when E A - E E - E B is below a predetermined value, where E A is said acoustic energy measured by said microphone, E E is said echo cancellation energy, and E B is said background noise energy.
  6. 19
    A method for determining acoustic characteristics of a room havinga microphone and a loudspeaker, said microphone being connected to a computer through an analog to digital converter and said loudspeaker being connected to said computer through a digital to analog converter, said method comprising the steps of:generating, by said computer, a first source signal which is substantially a chirp signal, converting said first source signal to a first audio signal by said digital to analog converter and said microphone;receiving said first audio signal by said loudspeaker;converting said received first audio signal by said analog to digital converter to generate a first digital signal;processing said first digital signal by a filter matched to said chirp signal;and determining a bulk delay as a time when said processed signal has a maximum value.