System and method for stereo conferencing over low-bandwidth links
Summary by NHIP
Stereo conferencing over low-bandwidth links
The system encodes two spatially-separated sound field signals into a single audio stream while transmitting a relative temporal delay parameter. A decoder uses this delay to split the signal into multiple presentation channels that simulate speaker location based on the original sampling points.
Claim Score by NHIP
Abstract
Systems and methods are disclosed for packet voice conferencing. An encoding system accepts two sound field signals, representing the same sound field sampled at two spatially-separated points. The relative delay between the two sound field signals is detected over a given time interval. The sound field signals are combined and then encoded as a single audio signal, e.g., by a method suitable for monophonic VoIP. The encoded audio payload and the relative delay are placed in one or more packets and sent to a decoding device via the packet network. The decoding device uses the relative delay to drive a playout splitter—once the encoded audio payload has been decoded, the playout splitter creates multiple presentation channels by inserting the transmitted relative delay in the decoded signal for one (or more) of the presentation channels. The listener thus perceives a speaker's voice as originating from a location related to the speaker's physical position at the other end of the conference. An advantage of these embodiments is that a pseudo-stereo conference can be conducted with virtually the same bandwidth as a monophonic conference.

Term
Term ended
Expired 11 July 2020, 6.2 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
34 claims: 4 independent, 30 dependent
- 1Broadest claimClaim Score 61, broad(NHIP)An encoder comprising:a sound field signal encoder to create a digitally-encoded signal representing both a first and a second sound field signal;a stereo parameter estimator to estimate a relative temporal delay between the first sound field signal and the second sound field signal;and a packet formatter packetizing the digitally-encoded signal and a stereo decoding parameter based on the estimated relative temporal delay, the stereo decoding parameter including at least one of an explicit delay parameter, an explicit balance parameter, and an explicit arrival angle parameter.
- 11An encoder comprising:means for encoding a digital data block to represent a combination of first and second sound field signals concurrently-captured within a first time period, the first and second sound field signals representing a single sound field captured at two spatially-separated points;means for estimating, using the first and second sound field signals as captured in an approximate timeframe of the first time period, an explicit relative temporal delay between the first and second sound field signals;and means for encapsulating, in a packet format, the encoded digital data block and a stereo decoding parameter based on the relative temporal delay.
- 21A method comprising:digitally encoding a signal block to represent first and second sound field signals as concurrently-captured during a first time period, the first and second sound field signals representing a single sound field captured at two spatially-separated points;estimating a relative temporal delay between the first and second sound field signals within an approximate timeframe of the first time period;transmitting to a remote conferencing point, in packet format, both the encoded signal block and a stereo decoding parameter based on the estimated relative temporal delay, the stereo decoding parameter including at least one of an explicit delay parameter, an explicit balance parameter, and an explicit arrival angle parameter.
- 28An apparatus comprising a computer-readable medium containing computer instructions that, when executed, cause a processor or multiple communicating processors to perform a method comprising:digitally encoding a signal block to represent first and second sound field signals as concurrently-captured during a first time period, the first and second sound field signals representing a single sound field captured at two spatially-separated points;detecting a talkspurt represented in the sound field signals;estimating a relative temporal delay between the first and second sound field signals within an approximate timeframe of the first time period responsive to the detection of the talkspurt;transmitting to a remote conferencing point, in packet format, both the encoded signal block and a stereo decoding parameter based on the estimated relative temporal delay.
Independent claims4
79 paragraphs in 6 sections, as filed
CROSS REFERENCE TO RELATED APPLICATIONS
0001This application is a continuation of U.S. patent application Ser. No. 09/614,535, filed Jul. 11, 2000, now U.S. Pat. No. 6,973,184.
FIELD OF THE INVENTION
0002This present invention relates generally to packet voice conferencing, and more particularly to systems and methods for packet voice stereo conferencing without explicit transmission of two voice channels.
BACKGROUND OF THE INVENTION
0003Packet-switched networks route data from a source to a destination in packets. A packet is a relatively small sequence of digital symbols (e.g., several tens of binary octets up to several thousands of binary octets) that contains a payload and one or more headers. The payload is the information that the source wishes to send to the destination. The headers contain information about the nature of the payload and its delivery. For instance, headers can contain a source address, a destination address, data length and data format information, data sequencing or timing information, flow control information, and error correction information.
0004A packet's payload can consist of just about anything that can be conveyed as digital information. Some examples are e-mail, computer text, graphic, and program files, web browser commands and pages, and communication control and signaling packets. Other examples are streaming audio and video packets, including real-time bi-directional audio and/or video conferencing. In Internet Protocol (IP) networks, a two-way (or multipoint) audio conference that uses packet delivery of audio is usually referred to as Voice over IP, or VoIP.
0005VoIP packets are transmitted continuously (e.g., one packet every 10 to 60 milliseconds) between a sending conference endpoint and a receiving conference endpoint when someone at the sending conference endpoint is talking. This can create a substantial demand for bandwidth, depending on the codec (compressor/decompressor) selected for the packet voice data. In some instances, the sustained bandwidth required by a given codec may approach or exceed the data link bandwidth at one of the endpoints, making that codec unusable for that conference. And in almost all cases, because bandwidth must be shared with other network users, codecs that provide good compression (and therefore smaller packets) are widely sought after.
0006Usually at odds with the desire for better compression is the desire for good audio quality. For instance, perceived audio quality increases when the audio is sampled, e.g., at 16 kHz vs. the eight kHz typical of traditional telephone lines. Also, quality can increase when the audio is captured, transmitted, and presented in stereo, thus providing directional cues to the listener. Unfortunately, either of these audio quality improvements roughly doubles the required bandwidth for a voice conference.
SUMMARY OF THE INVENTION
0007The present disclosure introduces new encoding/decoding systems and methods for packet voice conferencing. The systems and methods allow a pseudo-stereo packet voice conference to be conducted with only a negligible increase in bandwidth as compared to a monophonic packet voice conference. In addition to providing a generally more satisfying sound quality than monophonic conferencing, these systems and methods can provide a more tangible benefit when one end of a conference has multiple participants—the ability of the listener to receive a unique directional cue for each speaker on the other end of the conference. Moreover, because only a negligible increase in bandwidth over a monophonic conference is required, the present invention allows the advantages of stereo to be enjoyed over any data link that can support a monophonic conferencing data rate.
0008In the disclosed embodiments, a multichannel sound field capture system (which may or may not be part of the embodiment) captures sound field signals at spatially-separated points within a sound field. For instance, two microphones can be placed a short distance apart on a table, spatially-separated within a common VoIP phone housing, placed on opposite sides of a laptop computer, etc. The sound field signals exhibit different delays in representing a given speaker's voice, depending on the spatial relationship between the speaker and the microphones.
0009The sound field signals are provided to an encoding system, where the relative delay is detected over a given time interval. The sound field signals are combined and then encoded as a single audio signal, e.g., by a method suitable for monophonic VoIP. The encoded audio payload and the relative delay are placed in one or more packets and sent to the decoding device via the packet network. The relative delay can be placed in the same packet as the encoded audio payload, adding perhaps a few octets to the packet's length.
0010The decoding device uses the relative delay to drive a playout splitter—once the encoded audio payload has been decoded, the playout splitter creates multiple presentation channels by inserting a relative delay in the decoded signal for one (or more) of the presentation channels. The listener thus perceives the speaker's voice as originating from a location related to the speaker's actual orientation to the microphones at the other end of the conference.
BRIEF DESCRIPTION OF THE DRAWING
0011The invention may be best understood by reading the disclosure with reference to the drawing, wherein:
0012<figref idref="DRAWINGS">FIG. 1</figref> illustrates the general configuration of a packet-switched stereo telephony system;
0013<figref idref="DRAWINGS">FIG. 2</figref> illustrates a two-dimensional section of a sound field with two microphones, showing lines of constant inter-microphone delay;
0014<figref idref="DRAWINGS">FIG. 3</figref> contains a high-level block diagram for a pseudo-stereo voice encoder according to an embodiment of the invention;
0015<figref idref="DRAWINGS">FIG. 4</figref> illustrates one packet format useful with the present invention;
0016<figref idref="DRAWINGS">FIG. 5</figref> shows left and right channel voice signals along with their alignment with sampling blocks and voice activity detection signals;
0017<figref idref="DRAWINGS">FIG. 6</figref> illustrates correlation alignments for a cross-correlation method according to an embodiment of the invention;
0018<figref idref="DRAWINGS">FIG. 7</figref> illustrates left-to-right channel cross-correlation vs. sample index distance;
0019<figref idref="DRAWINGS">FIG. 8</figref> contains a high-level block diagram for a pseudo-stereo voice decoder according to an embodiment of the invention; and
0020<figref idref="DRAWINGS">FIG. 9</figref> contains a block diagram for a decoder playout splitter according to an embodiment of the invention.
DETAILED DESCRIPTION
0021In the following description, a packet voice conferencing system exchanges real-time audio conferencing signals with at least one other packet voice conferencing system in packet format. Such a system can be located at a conferencing endpoint (i.e., where a human conferencing participant is located), in an intermediate Multipoint Conferencing Unit (MCU) that mixes or bridges signals from conferencing endpoints, or in a voice gateway that receives signals from a remote endpoint in non-packet format and converts those signals to packet format. MCUs and voice gateways can typically handle more than one simultaneous conference. Note that not every endpoint in a packet voice conference need receive and transmit packet-formatted signals, as MCUs and voice gateways can provide conversion for non-packet endpoints. Such systems are also not limited to voice signals only—other audio signals can be transmitted as part of the conference, and the system can simultaneously transmit packet video or data as well.
0022As an introduction to the embodiments, the general operation of a stereo packet voice conference will be discussed. Referring to <figref idref="DRAWINGS">FIG. 1</figref>, one-half of a two-way stereo conference between two endpoints (the half allowing A to hear B<b>1</b>, B<b>2</b>, and B<b>3</b>) is depicted. A similar reverse path (not shown) allows A's voice to be heard by B<b>1</b>, B<b>2</b>, and B<b>3</b>. The number of persons present on each end of the conference is not critical, and has been selected in <figref idref="DRAWINGS">FIG. 1</figref> for illustrative purposes only.
0023The elements shown in <figref idref="DRAWINGS">FIG. 1</figref> include: two microphones <b>20</b>L, <b>20</b>R connected to an encoder <b>24</b> via capture channels <b>22</b>L, <b>22</b>R; two acoustic speakers <b>26</b>L, <b>26</b>R connected to a decoder <b>30</b> via presentation channels <b>28</b>L, <b>28</b>R, and a packet data network <b>32</b> over which encoder <b>24</b> and decoder <b>30</b> communicate.
0024Microphones <b>20</b>L and <b>20</b>R simultaneously capture the sound field produced at two spatially-separated locations when B<b>1</b>, B<b>2</b>, or B<b>3</b> talk, translate the sound field to electromagnetic signals, and transmit those signals over left and right capture channels <b>22</b>L and <b>22</b>R. Capture channels <b>22</b>L and <b>22</b>R carry the signals to encoder <b>24</b>.
0025Encoder <b>24</b> and decoder <b>30</b> work as a pair. Usually at call setup, the endpoints exchange control packets to establish how they will communicate with each other. As part of this setup, encoder <b>24</b> and decoder <b>30</b> negotiate a codec that will be used to encode capture channel data for transmission from encoder <b>24</b> to decoder <b>30</b>. The codec may use a technique as simple as Pulse-Code Modulation, or a very complex technique, e.g., one that uses subband coding, predictive coding, and/or vector quantization to decrease bandwidth requirements. In the present invention, the encoder and decoder both have the capability to negotiate a pseudo-stereo codec—this may be a combination of one of the aforementioned monophonic codecs with an added stereo decoding parameter capability. Voice Activity Detection (VAD) may be used to further reduce bandwidth. In order to provide stereo perception of Endpoint B's environment to A, the codec must either encode each capture channel separately, encode a channel matrix that can be decoded to recreate the capture channels, or use a method according to the present invention.
0026Encoder <b>24</b> gathers capture channel samples for a selected time block (e.g., 10 ms), compresses the samples using the negotiated codec, and places them in a packet along with header information. The header information typically includes fields identifying source and destination, time-stamps, and may include other fields. A protocol such as RTP (Real-time Transport Protocol) is appropriate for transport of the packet. The packet is encapsulated with lower layer headers, such as an IP (Internet Protocol) header and a link-layer header appropriate for the encoder's link to packet data network <b>32</b>, and submitted to the packet data network. This process is then repeated for the next time block, and so on.
0027Packet data network <b>32</b> uses the destination addressing in each packet's headers to route that packet to decoder <b>30</b>. Depending on a variety of network factors, some packets may be dropped before reaching decoder <b>30</b>, and each packet can experience a somewhat random network transit delay, which in some cases can cause packets to arrive in a different order than that in which they were sent.
0028Decoder <b>30</b> receives the packets, strips the packet headers, and re-orders any out-of-order packets according to timestamp. If a packet arrives too late for its designated playout time, however, the packet will simply be dropped. Otherwise, the re-ordered packets are decompressed and amplified to create two presentation channels <b>28</b>L and <b>28</b>R. Channels <b>28</b>L and <b>28</b>R drive acoustic speakers <b>26</b>L and <b>26</b>R.
0029Ideally, the whole process described above occurs in a relatively short period, e.g., 250 ms or less from the time B<b>1</b> speaks until the time A hears B<b>1</b>'s voice. Longer delays are detrimental to two-way conversation, but can be tolerated to a point.
0030A's binaural hearing capability (i.e., A's two ears) allows A to localize each speaker's voice in a distinct location within the listening environment. If the delay (and, to some extent amplitude) differences between the sound field at microphone <b>20</b>L and at microphone <b>20</b>R can be faithfully transmitted and then reproduced by speakers <b>26</b>L and <b>26</b>R, B<b>1</b>'s voice will appear to A to originate at roughly the dashed location shown for B<b>1</b>. Likewise, B<b>2</b>'s voice and B<b>3</b>'s voice will appear to A to originate, respectively, at the dashed locations shown for B<b>2</b> and B<b>3</b>.
0031From studies of human hearing capabilities, it is known that directional cues are obtained via several different mechanisms. The pinna, or outer projecting portion of the ear, reflects sound into the ear in a manner that provides some directional cues, and serves a primary mechanism for locating the inclination angle of a sound source. The primary left-right directional cue is ITD (interaural time delay) for mid-low- to mid-frequencies (generally several hundred Hz up to about 1.5 to 2 kHz). For higher frequencies, the primary left-right directional cue is ILD (interaural level differences). For extremely low frequencies, sound localization is generally poor.
0032ITD sound localization relies on the difference in time that it takes for an off-center sound to propagate to the far ear as opposed to the nearer ear—the brain uses the phase difference between left and right arrival times to infer the location of the sound source. For a sound source located along the symmetrical plane of the head, no inter-ear phase difference exists; phase difference increases as the sound source moves left or right, the difference reaching a maximum when the sound source reaches the extreme right or left of the head. Once the ITD that causes the sound to appear at the extreme left or right is reached, further delay may be perceived as an echo or cause confusion as to the sound's location.
0033ILD is based on inter-ear differences in the perceived sound level—e.g., the brain assumes that a sound that seems louder in the left ear originated on the left side of the head. For higher frequencies (where ITD sound localization becomes difficult), humans rely on ILD to infer source location.
0034For two microphones placed in the same sound field, an ITD-like signal difference can be observed. <figref idref="DRAWINGS">FIG. 2</figref> shows a two-dimensional scaled spatial plot representing one plane of a three-dimensional sound field. Microphones <b>20</b>L and <b>20</b>R are represented spaced 13 inches apart—approximately the distance that sound travels in one millisecond.
0035Now assume that the sound field signals being captured by microphones <b>20</b>L and <b>20</b>R are digitally sampled at eight kHz, or eight samples per millisecond. In the time that it takes eight samples to be gathered, sound can travel the 13 inches between microphone <b>20</b>L and <b>20</b>R. Thus a sound originating to the right of microphone <b>20</b>R would arrive at <b>20</b>R one millisecond, or eight samples, before it arrives at <b>20</b>L. The relative delay line “−8” indicates that sounds originating along that line arrive at <b>20</b>R eight samples before they arrive at <b>20</b>L, and the relative delay line “+8” indicates the same timing but a reversed order of arrival.
0036The remainder of the relative delay lines in <figref idref="DRAWINGS">FIG. 2</figref> show loci of constant relative delay. As the distance to <b>20</b>L and <b>20</b>R becomes greater than the spacing between <b>20</b>L and <b>20</b>R, the loci begin to approximate straight lines drawn at constant arrival angles. In the eight kHz sampling rate, 13-inch microphone spacing example of <figref idref="DRAWINGS">FIG. 2</figref>, 17 different integer delays are possible. Note that changing either the sampling rate or the spacing between <b>20</b>L and <b>20</b>R can vary the number of possible integer sample delays in the pattern. Non-integer delays could also be calculated with an appropriate technique (e.g., oversampling or interpolating).
0037The encoding embodiments described below have a capability to estimate inter-microphone sound propagation delay and send a stereo decoding parameter related to this delay to a companion decoder. The stereo decoding parameter can relate directly to the estimated sound propagation delay, expressed in samples or units of time. Using a lookup table or formula based on the known microphone configuration, the delay can also be converted to an arrival angle or arrival angle identifier for transmission to the decoder. An arrival-angle-based stereo decoding parameter may be more useful when the decoder has no knowledge of the microphone configuration; if the decoder has such knowledge, it can also compute arrival angle from delay.
0038In a noiseless, reflectionless environment with a single sound source, a decoder embodiment can produce highly realistic stereo information from a monophonic received audio channel and the stereo decoding parameter. One decoder uses the stereo decoding parameter to split the monophonic channel into two channels—one channel time-shifted with respect to the other to simulate the appropriate ITD for the single sound source. This method degrades for multiple simultaneous sound sources, although it may still be possible to project all of the sound sources to the arrival angle of the strongest source.
0039Like ITD, ILD can also be estimated, parameterized, and sent along with a monophonic channel. One encoder embodiment compares the signal strength for microphones <b>20</b>L and <b>20</b>R and estimates a balance parameter. In many microphone/talker configurations, the signal strength variations between channels may be slight, and thus another embodiment can create an artificial ILD balance parameter based on estimated arrival angle. The decoder can apply the balance parameter to all received frequencies, or it can limit application to those frequencies (e.g., greater than about 1.5 to 2 kHz) where ILD becomes important for sound localization.
0040Moving now from the general functional description to the more specific embodiments, <figref idref="DRAWINGS">FIG. 3</figref> illustrates an encoder <b>24</b> for a packet voice conferencing system. Left and right audio capture channels <b>22</b>L and <b>22</b>R are passed respectively through filters <b>34</b>L and <b>34</b>R. Filters <b>34</b>L and <b>34</b>R limit the frequency range of signals on their respective capture channels to a range appropriate for the sampling rate of the system, e.g., 100 Hz to 3400 Hz for an 8 kHz sampling rate. A/D converters <b>36</b>L and <b>36</b>R convert the output of filters <b>34</b>L and <b>34</b>R, respectively, to digital voice sample streams. The voice sample streams pass respectively to sample buffers <b>38</b>L and <b>38</b>R, which store the samples while they await encoding. The voice sample streams also pass to voice activity detector <b>40</b>, where they are used to generate a VAD signal.
0041Stereo parameter estimator <b>42</b> accepts samples from buffers <b>38</b>L and <b>38</b>R. Stereo parameter estimator <b>42</b> estimates, e.g., the relative temporal delay between the two sound field signals represented by the sample streams. Estimator <b>42</b> also uses the VAD signal as an enabling signal, and does not attempt to estimate relative delay when no voice activity is present. More specifics on methods of operation of stereo parameter estimator <b>42</b> will be presented later in the disclosure.
0042Adder <b>44</b> adds one sample from sample buffer <b>38</b>L to a corresponding sample from sample buffer <b>38</b>R to produce a combined sample. The adder can optionally provide averaging, or in some embodiments can simply pass one sample stream and ignore the other (other more elaborate mixing schemes, such as partial attenuation of one channel, time-shifting of a channel, etc., are possible but not generally preferred). The main purpose of adder <b>44</b> is to supply a single sample stream to signal encoder <b>46</b>.
0043Signal encoder <b>46</b> accepts and encodes samples in blocks. Typically, encoder <b>46</b> gathers samples for a fixed time (or sample period). The samples are then encoded as a block and provided to packet formatter <b>48</b>. Encoder <b>46</b> then gathers samples for the next block of samples and repeats the encoding process. Many monophonic signal encoders are known and are generally suited to perform the function of encoder <b>46</b>.
0044Packet formatter <b>48</b> constructs voice packets <b>50</b> for transmission. One possible format for a packet <b>50</b> is shown in <figref idref="DRAWINGS">FIG. 4</figref>. An RTP header <b>52</b> identifies the source, identifies the payload with a timestamp, etc. Formatter <b>48</b> may attach lower-layer headers (such as UDP and IP headers, not shown) to packet <b>50</b> as well, or these headers may be attached by other functional units before the packet is placed on the network.
0045The remainder of packet <b>50</b> is the payload <b>54</b>. The stereo decoding parameter field <b>56</b> is placed first within the payload section of the packet. A first octet of the stereo decoding parameter field represents delay as a signed 7-bit integer, where the units are time, with a unit value of 62.5 microseconds. Positive values represent delay in the right channel, negative values delay in the left. A second (optional) octet of the stereo decoding parameter field represents balance as a signed 7-bit integer, where one unit represents a half-decibel. Positive values represent attenuation in the right channel, negative values attenuation in the left. Third and fourth (also optional) octets of the stereo decoding parameter field represent arrival angle as a signed 15-bit integer, where the units are degrees. Positive values represent arrival angles to the left of straight ahead; negative values represent arrival angles to the right of straight-ahead. Following the stereo decoding parameter field, an encoded sample block completes the payload of packet <b>50</b>.
0046Several possible methods of operation for stereo parameter estimator <b>42</b> will now be described with reference to <figref idref="DRAWINGS">FIGS. 5</figref>, <b>6</b>, and <b>7</b>.
0047<figref idref="DRAWINGS">FIG. 5</figref> shows amplitude vs. time plots for time-synchronized left and right voice capture channels. Left voice sample blocks L-<b>1</b>, L-<b>2</b>, . . . , L-<b>15</b> show blocking boundaries used by signal encoder <b>46</b> of <figref idref="DRAWINGS">FIG. 3</figref> for the left voice capture channel. Right voice sample blocks R-<b>1</b>, R-<b>2</b>, . . . , R-<b>15</b> show the same blocking boundaries for the right voice capture channel. Left VAD and right VAD signals show the output of voice activity detector <b>40</b>, where detector <b>40</b> computes a separate VAD for each channel. The VAD method employed for each channel is, e.g., to detect the average RMS signal strength within a sliding sample window, and indicate the presence of voice activity when the signal strength is larger than a noise threshold. Note that the VAD signals indicate the beginning and ending points of talkspurts in the speech pattern, with a slight delay (because of the averaging window) in transitioning between on and off.
0048The on-transition times of the separate VAD signals can be used to estimate the relative delay between the left and right channels. This requires that, first, separate VAD signals be calculated, which is not generally necessary without this delay estimation method. Second, this requires that the time resolution of the VAD signals be sufficient to estimate delay at a meaningful scale. For instance, a VAD signal that is calculated once or twice per sample block will generally not provide sufficient resolution, while one that is calculated every sample generally will.
0049Stereo parameter estimator <b>42</b> receives the left and right components of the VAD signal. When one component transitions to “on”, parameter estimator <b>42</b> begins a counter, and counts the number of samples that pass until the other component transitions to “on”. The counter is then stopped, and the counter value is the delay. A negative delay occurs when the right VAD transitions first, and a positive delay occurs when the left VAD transitions first. When both VAD components transition on the same sample, the counter value is zero.
0050This delay detection method has several characteristics that may or may not cause problems in a given application. First, since it uses the onset of a talkspurt as a trigger, it produces only one estimate per talkspurt. But unless the speaker is moving very rapidly and speaking very slowly, one estimate per talkspurt is probably sufficient. Also at issue are how suddenly the talkspurt begins and how energetic the voice is—indistinct and/or soft transitions negatively impact how well this method will work in practice. Finally, if one channel receives a signal that is significantly attenuated with respect to the other, this may delay the VAD transition on that channel with respect to the other.
0051A second delay detection method is cross-correlation. One cross-correlation method is partially depicted in <figref idref="DRAWINGS">FIG. 6</figref>. Assume, as shown in <figref idref="DRAWINGS">FIG. 5</figref>, that the VAD signals turn on during the time period corresponding to sample blocks L-<b>2</b> and R-<b>2</b>. The delay can be estimated during the approximate timeframe of this time period by cross-correlation using one of several possible methods of sample selection.
0052In a first method, a cross-correlator for a given sample block time period (e.g., the L-<b>2</b> time period as shown) cross-correlates the samples in one sample stream from that sample block with samples from the other sample stream. As shown in <figref idref="DRAWINGS">FIG. 6</figref>, samples <b>0</b> to N−1 of block L-<b>2</b> (a length-N block) are used in the correlation. A sample index shift distance k determines how block L-<b>2</b> is aligned with the right sample stream for each correlation point. Thus, when k<0, L-<b>2</b> is shifted forward, such that sample <b>0</b> of block L-<b>2</b> is correlated with sample N-k of block R-<b>1</b>, and sample N−1 of block L-<b>2</b> is correlated with sample N−1 −k of block R-<b>2</b>. Likewise, when k>0, L-<b>2</b> is shifted backward, such that sample <b>0</b> of block L-<b>2</b> is correlated with sample k of block R-<b>2</b>, and sample N−1 of block L-<b>2</b> is correlated with sample k−1 of block R-<b>3</b>. For the special case k=0, which represents zero relative delay, blocks L-<b>2</b> and R-<b>2</b> are correlated directly.
0053One expression for a cross-correlation coefficient R<sub>i,k </sub>(others exist) is given below. In this expression, i is a sample index, L(i) is the left sample with index i, R(i) is the right sample with index i, N is the number of samples being cross-correlated, and k is an index shift distance.
0054<maths id="MATH-US-00001" num="00001"><math overflow="scroll"><mtable><mtr><mtd><mrow><msub><mi>R</mi><mrow><mi>i</mi><mo>,</mo><mi>k</mi></mrow></msub><mo>=</mo><mfrac><mrow><mrow><mi>N</mi><mo></mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><mrow><mrow><mi>L</mi><mo></mo><mrow><mo>(</mo><mi>j</mi><mo>)</mo></mrow></mrow><mo></mo><mrow><mi>R</mi><mo></mo><mrow><mo>(</mo><mrow><mi>j</mi><mo>+</mo><mi>k</mi></mrow><mo>)</mo></mrow></mrow></mrow></mrow></mrow><mo>-</mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><mrow><mrow><mi>L</mi><mo></mo><mrow><mo>(</mo><mi>j</mi><mo>)</mo></mrow></mrow><mo></mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><mrow><mi>R</mi><mo></mo><mrow><mo>(</mo><mrow><mi>j</mi><mo>+</mo><mi>k</mi></mrow><mo>)</mo></mrow></mrow></mrow></mrow></mrow></mrow><mrow><msqrt><mrow><mrow><mi>N</mi><mo></mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><msup><mrow><mi>L</mi><mo></mo><mrow><mo>(</mo><mi>j</mi><mo>)</mo></mrow></mrow><mn>2</mn></msup></mrow></mrow><mo>-</mo><msup><mrow><mo>(</mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><mrow><mi>L</mi><mo></mo><mrow><mo>(</mo><mi>j</mi><mo>)</mo></mrow></mrow></mrow><mo>)</mo></mrow><mn>2</mn></msup></mrow></msqrt><mo></mo><msqrt><mrow><mrow><mi>N</mi><mo></mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><msup><mrow><mi>R</mi><mo></mo><mrow><mo>(</mo><mrow><mi>j</mi><mo>+</mo><mi>k</mi></mrow><mo>)</mo></mrow></mrow><mn>2</mn></msup></mrow></mrow><mo>-</mo><msup><mrow><mo>(</mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><mrow><mi>R</mi><mo></mo><mrow><mo>(</mo><mrow><mi>j</mi><mo>+</mo><mi>k</mi></mrow><mo>)</mo></mrow></mrow></mrow><mo>)</mo></mrow><mn>2</mn></msup></mrow></msqrt></mrow></mfrac></mrow></mtd><mtd><mrow><mo>(</mo><mn>1</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><img file="US7194084B2_D0001.tif" />
0055A separate coefficient R<sub>i,k </sub>is calculated for each index shift distance k under consideration. It is noted, however, that several of the required summations do not vary with k, and need only be calculated once for a given i and N. The remaining summations (except for the summation that cross-multiplies L(i) with R(i+k)) do vary with k, but have many common terms for different values of k—this commonality can also be exploited to reduce computational load. It is also noted that if a running estimate is to be kept, e.g., since the beginning of a talkspurt, the summations can simply be updated as new samples are received.
0056<figref idref="DRAWINGS">FIG. 7</figref> contains an exemplary plot showing how R<sub>i,k </sub>can vary from a theoretical maximum of 1 (when L(i) and R(i) are perfectly correlated for a shift distance k) to a theoretical minimum of −1 (when the perfect correlation is exactly out of phase). A R<sub>i,k </sub>Of zero indicates no correlation, which would be expected when a random white noise sequence of infinite length is correlated with a second signal. When L(i) and R(i) capture the same sound field, with a dominant sound source, a positive maximum value in R<sub>i,k </sub>should indicate the relative temporal delay in the two signals, since that is the point where the two signals best match. In <figref idref="DRAWINGS">FIG. 7</figref>, the largest cross-correlation figure is obtained for a sample index shift distance of +2—thus +2 would correspond to the estimated relative temporal delay for this example.
0057With the above method, a separate estimate of relative temporal delay can be made for each sample block that is encoded by signal encoder <b>46</b>. The delay estimate can be placed in the same packet as the encoded sample block. It can be placed in a later packet as well, as long as the decoder understands how to synchronize the two and receives the delay estimate before the encoded sample block is ready for playout.
0058It may be preferable to limit the variation of the estimated relative temporal delay during a talkspurt. For instance, once an initial delay estimate for a given talkspurt has been sent to the decoder, variation from this estimate can be held relatively (or rigidly) constant, even if further delay estimates differ. One method of doing this is to use the first several sample blocks of the talkspurt to compute a single, good estimate of delay, which is then held constant for the duration of the talkspurt. Note that even if one estimate is used, it may be preferable to send it to the decoder in multiple packets in case one packet is lost.
0059A second method for limiting variation in estimated delay is as follows. After the stereo parameter estimator transmits a first delay estimate, the stereo parameter estimator continues to calculate delay estimates, either by adding more samples to the original cross-correlation summations as those samples become available, or by calculating a separate delay for each new sample block. When separate delay estimates are calculated for each block, the transmitted delay estimate can be the output of a smoothing filter, e.g., an average of the last n delay estimates.
0060The summations used in calculating a delay estimate can also be used to calculate a stereo balance parameter. Once the shift index k generating the largest cross-correlation coefficient is known, the RMS signal strengths for the time-shifted sequences can be ratioed to form a balance figure, e.g., a balance parameter B<sub>L/R </sub>can be computed in decibels as:
0061<maths id="MATH-US-00002" num="00002"><math overflow="scroll"><mtable><mtr><mtd><mrow><msub><mi>B</mi><mrow><mi>L</mi><mo>/</mo><mi>R</mi></mrow></msub><mo>=</mo><mrow><mn>10</mn><mo></mo><mstyle><mspace width="0.6em" height="0.6ex" /></mstyle><mo></mo><mrow><mi>log</mi><mo>(</mo><mfrac><mrow><mrow><mi>N</mi><mo></mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><msup><mrow><mi>L</mi><mo></mo><mrow><mo>(</mo><mi>j</mi><mo>)</mo></mrow></mrow><mn>2</mn></msup></mrow></mrow><mo>-</mo><msup><mrow><mo>(</mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><mrow><mi>L</mi><mo></mo><mrow><mo>(</mo><mi>j</mi><mo>)</mo></mrow></mrow></mrow><mo>)</mo></mrow><mn>2</mn></msup></mrow><mrow><mrow><mi>N</mi><mo></mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><msup><mrow><mi>R</mi><mo></mo><mrow><mo>(</mo><mrow><mi>j</mi><mo>+</mo><mi>k</mi></mrow><mo>)</mo></mrow></mrow><mn>2</mn></msup></mrow></mrow><mo>-</mo><msup><mrow><mo>(</mo><mrow><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mi>i</mi></mrow><mrow><mi>i</mi><mo>+</mo><mi>N</mi><mo>-</mo><mn>1</mn></mrow></munderover><mo></mo><mrow><mi>R</mi><mo></mo><mrow><mo>(</mo><mrow><mi>j</mi><mo>+</mo><mi>k</mi></mrow><mo>)</mo></mrow></mrow></mrow><mo>)</mo></mrow><mn>2</mn></msup></mrow></mfrac><mo>)</mo></mrow></mrow></mrow></mtd><mtd><mrow><mo>(</mo><mn>2</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><img file="US7194084B2_D0002.tif" />
0062Optionally, a balance parameter can be calculated only for a higher-frequency subband, e.g., 1.5 kHz to 3.4 kHz. Both sample streams are highpass-filtered, and the resulting sample streams are used in an equation like equation (2). Alternatively, once arrival angle is known, a lookup function can simply determine an appropriate ILD that a human would observe for that arrival angle. The balance parameter can simply express the balance figure that corresponds to that ILD.
0063Turning now to a discussion of a companion decoder for the disclosed encoders, <figref idref="DRAWINGS">FIG. 8</figref> shows a decoder <b>30</b>. Voice packets <b>50</b> arrive at a packet parser <b>60</b>, which splits each packet into its component parts. The packet header of each packet is used by the packet parser itself to control jitter buffer <b>64</b>, reorder out-of-order packets, etc., e.g., in one of the ways that is well understood by those skilled in the art. The stereo decoding parameter components (e.g., relative delay, balance, and arrival angle) are passed to playout splitter <b>66</b>. In addition, the encoded sample blocks are passed to signal decoder <b>62</b>.
0064Signal decoder <b>62</b> decodes the encoded sample blocks to produce a monophonic stream of voice samples. Jitter buffer <b>64</b> stores these voice samples, and makes them available for playout after a delay that is set by packet parser <b>60</b>. Playout splitter <b>66</b> receives the delayed samples from jitter buffer <b>64</b>.
0065Playout splitter <b>66</b> forms left and right presentation channels <b>28</b>L and <b>28</b>R from the voice sample stream received from jitter buffer <b>64</b>. One implementation of playout splitter <b>66</b> is detailed in <figref idref="DRAWINGS">FIG. 9</figref>. The voice samples are input to a k-stage delay register <b>70</b>, where k is the largest allowable delay in samples. The voice samples are also input directly to input <b>10</b> of a (k+1)-input multiplexer. Each stage of delay register <b>70</b> has its output tied to a corresponding input of multiplexer <b>72</b>, i.e., stage D<b>1</b> of register <b>70</b> is tied to input I<b>1</b> of multiplexer <b>72</b>, etc.
0066The delay magnitude bits that correspond to integer units of delay address multiplexer <b>72</b>. Thus, when the delay magnitude bits are 0000, input I<b>0</b> of multiplexer <b>72</b> is output on OUT, when the delay magnitude bits are 0011, input I<b>3</b> of multiplexer <b>72</b> (a three-sample-delayed version of the input) is output on OUT, etc. Note that when the delay magnitude increases by one, a voice sample will be repeated on OUT. Similarly, when the delay magnitude decreases by one, a voice sample will be skipped on OUT.
0067Switch <b>74</b> determines whether the sample-delayed voice sample stream on OUT will be placed on the left or the right output channel. When the delay sign bit is set, the delayed voice sample stream is switched to left channel <b>74</b>L. Otherwise, the delayed voice sample stream is switched to right channel <b>74</b>R. Switch <b>74</b> sends the no-delayed version of the voice sample stream to the channel that is not currently receiving the delayed version.
0068When the decoding system is to create an ILD effect in the output, additional hardware such as exponentiator <b>76</b>, switch <b>78</b>, and multipliers <b>80</b> and <b>82</b> can be added to splitter <b>66</b>. Exponentiator <b>76</b> takes the magnitude bits of the balance parameter and exponentiates them to compute an attenuation factor. The sign of the balance parameter operates a switch <b>78</b> that applies the attenuation factor to either the left or the right channel. When the balance sign bit is set, the attenuation factor is switched to left channel <b>78</b>L. Otherwise, the attenuation factor is switched to right channel <b>78</b>R. Switch <b>78</b> sends an attenuation factor of 1.0 (i.e., no attenuation) to the channel that is not currently receiving the received attenuation factor.
0069Multipliers <b>80</b> and <b>82</b> transfer attenuation to the output channels. Multiplier <b>80</b> multiplies channel <b>74</b>L with switch output <b>78</b>L to produce left presentation channel <b>28</b>L. Multiplier <b>82</b> multiplies channel <b>74</b>R with switch output <b>78</b>R to produce right presentation channel <b>28</b>R. Note that if it is desired to attenuate only high frequencies, the multipliers can be augmented with filters to attenuate only the higher frequency components.
0070The illustrated embodiments are generally applicable to use in a voice conferencing endpoint. With a few modifications, these embodiments also apply to implementation in an MCU or voice gateway.
0071MCUs are usually used to provide mixing for multi-point conferences. The MCU could possibly: (1) receive a pseudo-stereo packet stream according to the invention; (2) send a pseudo-stereo packet stream according to the invention; or (3) both.
0072When receiving a pseudo-stereo packet stream, the MCU can decode it as described in the description accompanying <figref idref="DRAWINGS">FIGS. 8 and 9</figref>. The difference would be in that the presentation channels would possibly be mixed with other channels and then transmitted to an endpoint, most likely in a packet format.
0073When sending a pseudo-stereo packet stream, the MCU must encode such a stream. Thus, the MCU must receive a stereo stream from which it can determine delay. The stereo stream could be in packet format, but would preferably use a PCM or similar codec that would preserve the left and right channels with little distortion until they reached the MCU.
0074When the MCU both receives and transmits a pseudo-stereo stream, it need not perform delay detection on a mixed output stream. For mixed channels, the received delays can be averaged, arbitrated such that the channel with the most signal energy dominates the delay, etc.
0075A voice gateway is used when one voice conferencing endpoint is not connected to the packet network. In this instance, the voice gateway connects to the endpoint over a circuit-switched or dedicated data link (albeit a stereo data link). The voice gateway receives stereo PCM or analog stereo signals from the endpoint, and transmits the same in the opposite direction. The voice gateway performs encoding and/or decoding according to the invention for communication across the packet data network with another conferencing point.
0076Although several embodiments of the invention and implementation options have been presented, one of ordinary skill will recognize that the concepts described herein can be used to construct many alternative implementations. Such implementation details are intended to fall within the scope of the claims. For example, a playout splitter can map a pseudo-stereo voice data channel to, e.g., a 3-speaker (left, right, center) or 5.1 (left-rear, left, center, right, right-rear, subwoofer) format. Alternatively, the encoder can accept more than two channels and compute more than one delay. Although a detailed digital implementation has been described, many of the components have equivalent analog implementations, for example, the playout splitter, the stereo parameter estimator, the adder, and the voice activity detector. Alternative component arrangements are also possible, e.g., the stereo parameter estimator can retrieve samples before they pass through the sample buffers, or the voice activity detector and the stereo parameter estimator can share common functionality. The particular packet and parameter format used to transmit data between encoder and decoder are application-dependent.
0077Particular device embodiments, or subassemblies of an embodiment, can be implemented in hardware. All device embodiments can be implemented using a microprocessor executing computer instructions, or several such processors can divide the tasks necessary to device operation. Thus another claimed aspect of the invention is an apparatus comprising a computer-readable medium containing computer instructions that, when executed, cause one or more processors to execute a method according to the invention.
0078The network could take many forms, including cabled telephone networks, wide-area or local-area packet data networks, wireless networks, cabled entertainment delivery networks, or several of these networks bridged together. Different networks may be used to reach different endpoints. Although the detailed embodiments use Internet Protocol packets, this usage is merely exemplary—the particular protocols selected for a given implementation are not critical to the operation of the invention.
0079The preceding embodiments are exemplary. Although the specification may refer to “an” “one”, “another”, or “some” embodiment(s) in several locations, this does not necessarily mean that each such reference is to the same embodiment(s), or that the feature only applies to a single embodiment.
Contents6
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8744065B2 | Cited by | United States of America | Applicant |
| US8428959B2 | Cited by | United States of America | Search report |
| US8335209B2 | Cited by | United States of America | Search report |
| US2009226008A1 | Cited by | United States of America | Pre-grant |
| US8577045B2 | Cited by | United States of America | Applicant |
| US2011077755A1 | Cited by | United States of America | Pre-grant |
| US9736312B2 | Cited by | United States of America | Applicant |
| US2010284310A1 | Cited by | United States of America | Pre-grant |
| US8358599B2 | Cited by | United States of America | Search report |
| US9602295B1 | Cited by | United States of America | Applicant |
| US2011085671A1 | Cited by | United States of America | Pre-grant |
| US2011191111A1 | Cited by | United States of America | Pre-grant |
| US9001182B2 | Cited by | United States of America | Applicant |
| US8363810B2 | Cited by | United States of America | Applicant |
| US2010092002A1 | Cited by | United States of America | Pre-grant |
| US2011164735A1 | Cited by | United States of America | Pre-grant |
| US2009245232A1 | Cited by | United States of America | Pre-grant |
| US8705766B2 | Cited by | United States of America | Search report |
| US2011069643A1 | Cited by | United States of America | Pre-grant |
| US9570080B2 | Cited by | United States of America | Applicant |
| US2007253558A1 | Cited by | United States of America | Pre-grant |
| US8571189B2 | Cited by | United States of America | Applicant |
| US8547880B2 | Cited by | United States of America | Applicant |
| US8208648B2 | Cited by | United States of America | Search report |
| US2011058662A1 | Cited by | United States of America | Pre-grant |
| US8144633B2 | Cited by | United States of America | Applicant |
| US2007253557A1 | Cited by | United States of America | Pre-grant |
| US4581758A | Cites | United States of America | Applicant |
| US4815132A | Cites | United States of America | Applicant |
| US6021386A | Cites | United States of America | Applicant |
| US6408327B1 | Cites | United States of America | Applicant |
| Guentchev, et al.; "Learning-Based Three Dimensional Sound Localization Using a Compact Non-Coplanar Array of Microphones"; American Association for Artificial Intelligence; 1998; 9 pages. | Non-patent | – | Applicant |
| Weinstein, et al.; Experience with Speech Communication in Packet Networks; IEEE Journal on Selected Areas in Communications, vol. SAC-1, No. 6; Dec. 1993. | Non-patent | – | Applicant |
| Guentchev, et al.; “Learning-Based Three Dimensional Sound Localization Using a Compact Non-Coplanar Array of Microphones”; American Association for Artificial Intelligence; 1998; 9 pages. | Non-patent | – | Third party observation |
| Weinstein, et al.; Experience with Speech Communication in Packet Networks; IEEE Journal on Selected Areas in Communications, vol. SAC-1, No. 6; Dec. 1993. | Non-patent | – | Third party observation |
3 members in 1 office
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 61453500 | United States of America | A | |
| 61453500 | United States of America | A | |
| 23954205 | United States of America | A | |
| 09614535 | – | – | – |
| US20000614535 | – | – | – |
| US20050239542 | – | – | – |
Members3
| Document | Office | Kind | |
|---|---|---|---|
| US6973184B1 | United States of America | B1 | |
| US2006023871A1 | United States of America | A1 | |
| US7194084B2This record | United States of America | B2 |
34 transactions on the USPTO file
Allowed without a rejection on record.
- Non-final rejections
- 0
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Correspondence Address ChangeC.AD | C.AD | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Correspondence Address ChangeC.AD | C.AD | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Mail Examiner's AmendmentMEX.A | MEX.A | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Examiner's Amendment CommunicationEX.A | EX.A | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response to Election / Restriction FiledELC. | ELC. | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Mail Restriction RequirementMCTRS | MCTRS | |
| Restriction/Election RequirementCTRS | CTRS | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Preliminary AmendmentA.PE | A.PE | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Application Is Now CompleteCOMP | COMP | |
| Application Is Now CompleteCOMP | COMP | |
| Application Return from OIPEWROIPE | WROIPE | |
| Application Return TO OIPEROIPE | ROIPE | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Cleared by OIPE CSRL194 | L194 | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Initial Exam Team nnIEXX | IEXX |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapsed due to failure to pay maintenance feeLapsedFP | FP | |
| Lapse for failure to pay maintenance feesLapsedPATENT EXPIRED FOR FAILURE TO PAY MAINTENANCE FEES (ORIGINAL EVENT CODE: EXP.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYLAPS | LAPS | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Fee payment procedureMAINTENANCE FEE REMINDER MAILED (ORIGINAL EVENT CODE: REM.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| Fee paymentFPAY | FPAY | |
| Fee paymentFPAY | FPAY | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF |
Numbers
- Publication
- 07194084
- Publication, DOCDB
- 7194084
- Publication, EPODOC
- US7194084
- Application
- 11239542
- Application, DOCDB
- 23954205
- Application, EPODOC
- US20050239542
Titles
- English
- System and method for stereo conferencing over low-bandwidth links
Patent term adjustment
- Net adjustment
- 0 days
Classification
- CPC, 2
- H04R27/00
- G10L19/008
- IPC, 3
- H04M9 08
- H04R27 00
- H04S5 00
- USPC, 4
- 379420010
- 379202010
- 379388010
- 381017000