Audio conferencing system.
10 claims: 4 independent, 6 dependent
- 1(57)【特許請求の範囲】 【請求項1】ネットワークに接続されて、該ネットワークから各々がデジタル音声サンプルのシーケンスを含む多重音声入力ストリームを受信するコンピュータ・ワークステーションであって、 各音声入力ストリームからのデジタル音声サンプルを別々のキューに記憶する手段と、 各キューから1つずつのデジタル音声サンプルを含むセットのシーケンスを形成する手段と、 各音声入力ストリームが関連する重みパラメータを有し、デジタル音声サンプルの各セットの加重合計を生成する手段と、 加重合計のシーケンスから音声出力を生成する手段とユーザ入力に応答して、多重音声ストリームの音声出力内の相対ボリュームを制御するために、前記重みパラメータを調整する手段と、 を含む、コンピュータ・ワークステーション。
- 2【請求項2】各前記音声入力ストリームが現在無音かどうかを示すビジュアル指示を提供する手段を含む、請求項1記載のコンピュータ・ワークステーション。
- 3【請求項3】前記ビジュアル指示が各前記音声入力ストリームの瞬時音響ボリュームを示す、請求項2記載のコンピュータ・ワークステーション。
- 4【請求項4】前記ビジュアル指示が、各前記音声入力ストリームの起点のビジュアル表現の近傍に表示される、請求項2または3記載のコンピュータ・ワークステーション。
- 5【請求項5】各前記音声入力ストリームに対して、デジタル音声サンプルのシーケンスから2乗平均平方根値を生成する手段を含む、請求項3または4記載のコンピュータ・ワークステーション。
- 6【請求項6】入力音声データが各々が所定数のデジタル音声サンプルを含むブロック単位で到来し、前記ビジュアル指示が音声データの各新たなブロックに対応して更新される、請求項2乃至5のいずれかに記載のコンピュータ・ワークステーション。
- 7【請求項7】前記多重音声入力ストリームの任意のストリームからの音声出力を禁止する手段を含む、請求項1乃至6のいずれかに記載のコンピュータ・ワークステーション。
- 8【請求項8】ユーザに前記重みパラメータの値のビジュアル指示を提供する手段を含み、該手段はユーザのマウス・オペレーションに応答して、前記重みパラメータを調整する、請求項1乃至7のいずれかに記載のコンピュータ・ワークステーション。
- 9【請求項9】ネットワークに接続されて、各々がデジタル音声サンプルのシーケンスを含む多重音声入力ストリームを受信するコンピュータ・ワークステーションを動作する方法であって、 各音声入力ストリームからのデジタル音声サンプルを別々のキューに記憶するステップと、 各キューから1つずつのデジタル音声サンプルを含むセットのシーケンスを形成するステップと、 各音声入力ストリームが関連する重みパラメータを有し、デジタル音声サンプルの各セットの加重合計を生成するステップと、 加重合計のシーケンスから音声出力を生成するステップと、 ユーザ入力に応答して、多重音声ストリームの音声出力における相対ボリュームを制御するために、前記重みパラメータを調整するステップと、 を含む、動作方法。
- 10【請求項10】各前記音声入力ストリームに対して、該音声ストリームの瞬時音響ボリュームを示すビジュアル指示を提供するステップを含む、請求項9記載の動作方法。
Independent claims10
112 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Industrial application field]
The present invention relates to the processing of multiple streams of voice data received over a network by a computer workstation.
【0002】
[Conventional technology]
Traditionally, voice signals have been transmitted over standard analog telephone lines. However, with the increasing number of locations provided by local area networks (LANs) and the increasing importance of multimedia communications, the use of LANs for transmitting audio signals has received much attention. This work is described, for example, in "Using Local Area Networks for Carrying Online Voice" (page 13-21) by D Cohen and "Voice Transmission over an Ethernet Backbone" (page 39-65) by P Ravasio, R Marcogliese and R Novarese. (Both "Local Computer Networks" (P Ravasio, G Hopkins and N Edited by Naffah; Northern Netherlands, 1982)). The basic principle of such a method is that the first terminal or workstation samples the audio input signal into a digital signal at a specified rate (eg 8kHz). A large number of sampled digital signals (hereinafter simply referred to as samples) are then assembled into data packets and transmitted over the network to a second terminal, which then again louds the samples at a constant rate, such as loudspeakers and others. Supply to the playback device.
【0003】
One problem with using a LAN to transmit voice data is that the transmission time across the network fluctuates. The arrival of packets at the destination node is delayed and irregular. When packets are played irregularly, this has a significant negative effect on the comprehension of the audio signal. Therefore, the LAN voice transmission method uses a certain amount of buffering on the receiving side to absorb such irregularities. Care must be taken not to cause too much delay between the original voice signal and the voice output on the destination side, which would make natural interactive two-way communication difficult (conventional telephone calls across the Atlantic). Excessive delay was hard to hear). "Adaptive Audio Playout Algorithm for Shared Packet Networks" by B Aldred, R Bowater and S Woodman (IBM Technical Disclosure Bulletine, p 255-257, Vol 36 No) 4, April 1993). The amount of buffering is properly controlled depending on the number of packets dropped (any other suitable delay measurement is available). A large number of discarded packets increases the degree of buffering, and a small number of discarded packets decreases the degree of buffering. The size of the buffer is changed by temporarily changing the playback rate (this affects the pitch, a simple technique detects periods of silence and artificially increases or decreases them moderately).
【0004】
Another important aspect of voice communication is a conference that includes multipoint communication as opposed to two-way or point-to-point communication. When implemented over traditional analog telephone lines, audio conferencing requires each participant to transmit an audio signal to a central hub. The central hub mixes the input signals, coordinates between different levels if possible, and sends each participant the sum of the signals from all other participants (excluding the signals from that participant's node). .. U.S. Pat. No. 4,650,929 discloses a centralized video / audio conferencing system that allows individuals to adjust the relative volume of other participants.
【0005】
The use of centralized mixed nodes is often referred to as a multipoint control unit (MCU) and is incorporated into several multimedia (audio and video) workstation conferencing systems. For example, U.S. Pat. No. 4,710,917 describes a multimedia conferencing system in which each participant sends audio to and receives audio from a central mixing unit. Other multimedia conferencing systems include "Distributed Multiparty Desktop Conferencing System: MERMAID" by K Watabe, S Sakata, K Maeno, H Fukuoka and To Ohmori (CSCW '90 (Proceeding of the conference on Computer-Supported Cooperative Work, 1990, Loss). "Personal Multimedia Multipoint Communications Services for Broadband" by Angeles), p27-38) and E Addeo, A Gelman and A Dayao Described in Networks "(Vol 1, IEEE GLOBECOM, 1988).
【0006】
However, the use of centralized MCUs or total nodes has some drawbacks. First, most LAN architectures are based on peer-to-peer configurations, so there is no obvious central node. In addition, the system totally depends on the continuous availability of the central node specified to operate the conference. There is also the issue of echo suppression (the central node must be careful not to include the audio signal from that node in the total signal played back to that node).
【0007】
These problems are avoided by using a distributed audio conferencing system, where each node receives a separate audio signal from any other node participating in the conference. U.S. Pat. No. 5,27001 describes such a distributed system and touches on synchronization problems caused by the fluctuating transit times of packets across the network. U.S. Pat. No. 5,27001 overcomes this problem by maintaining a separate queue of input voice packets from each source node. These effectively absorb the jitter (delay) at the arrival time by the same method as described above in the point-to-point communication. At regular intervals, one packet is read from each queue to read a set of voice packets, which are summed up for playback.
【0008】
One problem with audio conferencing systems is determining who is speaking at any given moment, as found in the MERMAID system described above. U.S. Pat. No. 4893326 describes a multimedia conferencing system in which each workstation automatically detects whether the user is speaking. This information is then supplied to the central control node, which exchanges videos so that each participant can see the current speaker on their screen. Such systems require the operation of both video and audio functions and rely on a central video exchange node, which makes them unusable in a fully distributed system.
【0009】
A distributed multimedia conferencing system is described in the "Personal Multimedia-Multipoint Teleconference System" by H Tanigawa, T Arikawa, S Masaki and K Shimamura (IEEE INFOCOM 91, Proceedings Vol 3, p1127-1134). This system is for stereo workstations where when a window containing a video signal from a conference participant is moved from right to left across the screen, the apparent source of the corresponding audio signal moves as well. Achieve sound localization. This approach provides limited support for speaker identification. A more inclusive mechanism is described in Japanese Patent Application Laid-Open No. 02-123886, where a bar graph is used to show the output audio level associated with a nearby window containing the video of the source of the sound.
【0010】
[Problems to be Solved by the Invention]
Thus, the prior art describes various audio conferencing systems. While traditional centralized telephone voice conferencing is widespread and well understood from a technical point of view, there are many challenges that must be addressed to improve the performance of voice conferencing in a desk-top environment. Remaining.
【0011】
[Means for solving problems]
Accordingly, the present invention provides a computer workstation that connects to a network and receives from the network multiplex audio input streams, each containing a sequence of digital audio samples. This workstation stores digital audio samples from each audio input stream in separate queues, forms a sequence of sets containing one digital audio sample from each queue, and the weights associated with each audio input stream. It has parameters and includes a means to generate a polymerizer for each set of digital audio samples, a means to generate an audio output from a sequence of polymerizers, and a relative volume in the audio output of a multiplex audio stream in response to user input. It is characterized by means of adjusting the weight parameters for control.
【0012】
The present invention is an additional that the provision of audio conferencing on a distributed network, where each node receives a separate audio stream from all other participants, has been achieved at great difficulty and cost in traditional centralized conferencing systems. Recognize that the function is naturally possible. In particular, each user can adjust the relative volume of all other participants according to his or her preference. This is very promising, for example, when focusing on a particular aspect of a meeting, or when there is a language problem (eg, when one person has a strong accent that others cannot understand, or when the meeting has simultaneous interpretation. Such). In addition, during the meeting, the system changes the relative volume of different participants in response to user input. To allow this control, the input audio signals are held separately and placed in different queues according to their source (queues are physically adjacent or combined, but logically separate. (Memory), then weighted by the appropriate volume control factor. They are then combined to produce the final audio output. The present invention recognizes that distributed audio conferencing systems are particularly suitable for providing individual control of relative volumes.
【0013】
Preferably, the workstation further includes means of providing visual instructions indicating whether each said voice input stream is currently silent. This overcomes one problem recognized in audio conferencing: identifying who is speaking. The visual indication may simply change the brightness (eg, switch to low brightness) of a particular form of on / off marking by a lighting or equivalent mechanism or the visual representation (eg, a portrait of a participant) at the origin of the audio input stream. However, in a preferred embodiment, each of the audio input streams is realized by displaying an instantaneous acoustic volume in the audio stream. In other words, the display provides a complete indication of the volume of the relevant participants. Volume output is calculated based on the root mean square (rms) from a sequence of digital audio samples, or, if processor power is limited, the maximum digital audio within a given number of samples. Simple algorithms such as values are used. Generally, the input audio data arrives in block units, each containing a predetermined number of digital audio samples, and the visual instructions are updated corresponding to each new block of audio data. Volume values are usually calculated in blocks.
【0014】
It is desirable that the visual instructions be displayed near the visual representation of the origin of the audio input stream, such as a video image or a still image. The former requires a complete multimedia conferencing network and the latter is offered on low bandwidth networks that cannot support the transmission of video signals. These visual instructions make it easy to identify the source of any audio, whether static or dynamic.
【0015】
Preferably, the workstation further includes means of providing the user with visual instructions for the value of the weight parameter, which adjusts the weight parameter in response to the user's mouse operation. This is implemented as a scroll bar, etc., with one scroll bar corresponding to each audio input stream and placed near the visual indication of the output volume of that stream. Further, it is convenient for the computer workstation to include means for prohibiting audio output from any stream of the multiplex audio input stream. This effectively provides the user with a complete set of volume controls for each voice input stream.
【0016】
The present invention further provides a method of operating a computer workstation connected to a network to receive a multiplex audio input stream, each containing a sequence of digital audio samples. This method stores the digital audio samples from each audio input stream in separate queues, forms a sequence of sets containing one digital audio sample from each queue, and the weights associated with each audio input stream. Relative in the audio output of a multiplex audio stream in response to user input, including the step of having parameters and generating a polymerizer for each set of digital audio samples from those parameters, and the step of generating audio output from a sequence of polymerizers. It is characterized by the steps of adjusting the weighting parameters to control the volume.
【0017】
[Example]
Figure 1 shows computer workstations A through E being linked within a local area network (LAN) 2. These workstations participate in a multi-directional conference, whereby each workstation broadcasts its audio signal to all other workstations participating in the conference. Each workstation receives a separate audio signal from any other workstation. The network shown in Figure 1 has a Token Ring architecture, where tokens circulate through workstations. Only the workstation currently in possession of the token is allowed to transmit the message to another workstation. The physical transmission time for transmitting a message in a ring is extremely short. In other words, the message transmitted by, for example, A is received by all other terminals at about the same time. This is because the token system is used to prevent conflicts arising from two nodes trying to transmit messages at the same time.
【0018】
As described in detail below, unidirectional voice communication over a LAN typically requires a bandwidth of 64 kHz. In the conference in Figure 1, each node broadcasts its audio signal to the other four nodes, which means that it requires a total bandwidth of 1.28MHz (5x4x64kHz). This happens to fall within the capabilities of standard Token Ring to support transmission rates of 4Mbit / s or 16Mbit / s, but for larger conferences bandwidth requirements are an issue and higher bandwidth is expected to be offered. It cannot support future networks.
【0019】
The present invention will be realized in many different network architectures or configurations other than Token Ring, provided that the technical requirements for bandwidth, latency, etc. required to support voice conferencing are met.
【0020】
FIG. 2 is a simplified diagram of the computer system used in the network of FIG. The computer has a system unit 10, a display screen 12, a keyboard 14 and a mouse 16. The system unit 10 includes a microprocessor 22, a semiconductor memory (ROM / RAM) 24, and a bus 26 to which data is transferred. The computer in Figure 2 is any conventional workstation, such as an IBM PS / 2 computer.
【0021】
The computer in Figure 2 is equipped with two adapter cards. One is the Token Ring Adapter Card 30. The card, along with the accompanying software, allows data to be sent and received to and from the Token Ring network shown in Figure 1. The operation of the Token Ring card is known and will not be discussed in detail here. The second card is a voice card 28, which is connected to a microphone and loudspeaker (not shown) for voice input and output.
【0022】
The voice card is shown in detail in Figure 3. The card illustrated and used in this particular embodiment is the M-Wave card provided by IBM, but other cards that perform similar functions can also be used. The card includes an A / D converter 42 that digitizes the input audio signal from the connected microphone 40. An A / D converter is connected to CODEC44, which samples the input audio signal into a 16-bit sample at a rate of 44.1 kHz (corresponding to the standard sampling rate / size of compact discs). The digitized sample is then passed through the double buffer 48 to the digital signal processor (DSP) 46 on the card (ie the CODEC loads the sample into one of the double buffers, while the other buffer. Read the previous sample from). The DSP is controlled by one or more programs stored in the semiconductor memory 52 on the card. Data is transferred to and from the main PC bus by DSP.
【0023】
The voice signal to be played out is received from the PC bus 26 by the DSP 46 and processed in the reverse direction of the voice input from the microphone. That is, the output audio signal is passed through DSP46 and double buffer 50 to CODEC44, from which it is passed to the D / A converter 54 and finally to the loudspeaker 56 or other suitable output device.
【0024】
In the specific embodiment shown, the DSP uses standard resampling techniques to sample 16-bit, 44.1 kHz sampling rates from the CODEC at an 8 kHz sampling rate corresponding to CCITT standard G.711, and μ. Μ-law scale ) (Actually logarithmic) is programmed to convert to a new digital signal with an 8-bit sample. The total bandwidth of the signal passed to the workstation for transmission to other terminals is therefore 64kHz. The DSP also performs the reverse conversion on the input signal received from the PC. That is, the signal is converted from 8 bits, 8 kHz to 16 bits, 44.1 kHz by a known resampling technique again. This conversion between the two sampling formats is only necessary for the particular choice of hardware and has no direct relation to the present invention. For example, many other voice cards include unique support for the 8kHz format, which allows the CODEC to operate at 8kHz according to the G.711 format. (Alternatively, a 44.1 kHz sample may be retained for transmission over the network, however, if there is no specific need for CD quality for the transmitted audio signal, higher bandwidth and The demand for significant processing speed improvements makes this unrealistic. For normal voice communication, a 64 kHz bandwidth signal in G.711 format is sufficient.
【0025】
Data is transferred between the voice adapter card and the workstation in blocks of 64 bytes. That is, it corresponds to audio data for 8 ms of 8-bit data sampled at 8 kHz. The workstation then processes only the entire block of data, and each data packet sent and received by the workstation typically contains a single 64-byte block of data. The choice of 64 bytes for block size is a compromise between minimizing system granularity (deriving delay) and maintaining efficiency in both internal processing at workstations and transmission over the network. In other systems, block sizes of, for example, 32 bytes or 128 bytes may be more preferred.
【0026】
Computer workstation operations related to the transmission of voice data are known in the art and will not be described in detail here. In essence, a voice card receives an input signal in analog form from a microphone or other voice source, such as a compact disc player, and produces a block of digital voice data. These blocks are then transferred to the workstation's main memory and from there to the LAN adapter card (in some architectures, the blocks are transferred directly from the voice adapter card to the LAN adapter without going through the workstation memory. It is possible to transfer to a card). The LAN adapter card generates a data packet containing digital voice data, along with header information that identifies the source and destination nodes, which is then transmitted over the network to the desired receiver. It will be appreciated that in any two-way or multi-directional communication, this transmission process is performed on the workstation at the same time as the receive process described below.
【0027】
The processing by the computer workstation for receiving voice data packets is shown in Figure 4. Each time a new packet arrives (step 402), the LAN adapter card notifies the program running on the microprocessor in the workstation and provides information to the program that identifies the source of the data packet. The program then transfers the input 64-byte audio block to a queue in main memory (step 404). As shown in FIG. 5, the queue in main memory 500 actually consists of a separate set of subqueues containing audio blocks from different source nodes. One queue contains audio blocks from one source node and another queue contains audio blocks from another source node. In FIG. 5, there are three subqueues 501, 502, and 503, which correspond to audio data from nodes B, C, and D, respectively. The number of subcues will, of course, vary according to the number of participants in the audio conference. The program uses the source node identification information in each received packet to allocate blocks of input voice data to the correct queue. Pointer P<sub>B</sub>, P<sub>C</sub>And P<sub>D</sub> Indicates the end position of the queue and is updated as new packets are added. The packet is retrieved from the bottom of the subqueue (indicated as "output" in Figure 5) for further processing. The subqueue in Figure 5 is therefore essentially a standard first-in first-out (FIFO) queue, implemented by conventional programming techniques. With the exception of support for multiple (parallel) queues, the processing of input audio blocks described so far is similar to the previous method, with equivalent buffering techniques for individual subqueues or combined queues as a whole, as needed. Allows the use of.
【0028】
The operations performed by the DSP on the voice adapter card are shown in Figure 6. The DSP is cycled and processes a new set of audio blocks every 8 milliseconds (ms) to guarantee a continuous audio output signal. The DSP reads one audio block from each subqueue corresponding to a different node by DMA access every 8ms. That is, one block is read from the bottom of queues B, C, and D shown in FIG. 5 (step 602: ie M = 3 in this case). These blocks are treated as representing simultaneous time intervals, and in the final output they are added together to produce a single audio output corresponding to that time interval. The DSP therefore efficiently performs the digital mixing function on the multiplex audio input stream. By using a lookup table, the individual samples in a 64-byte block are then converted from G.711 format (effectively logarithmic) to linear scale (step 604). Each individual sample is then multiplied by the weight parameter (step 606). There are separate weight parameters for each received audio data stream. That is, for the three subques in FIG. 5, one weight for the audio stream from node B, one for the audio stream from node C, and one for the audio stream from node D. The parameter exists. Weight parameters are used to control the relative loudness of audio signals from different sources.
【0029】
The DSP keeps a record of the root mean square (rms) of each audio stream (step 608). Usually, these rms values are obtained by generating the sum of the squares of the values in each block of audio data (that is, every 8 ms). The rms value represents the volume of each audio input stream and is used to provide volume information to the user, as described below.
【0030】
When the digital audio samples are multiplied by the appropriate weighting parameters (step 606), they are summed (step 610; this happens effectively in parallel with the processing in step 606). In this way, a single sequence of digital audio samples representing the copolymerization of the multiple input audio stream is generated. This sequence of digital audio samples is then resampled up to 44.1 kHz (step 612; as mentioned above, this is hardware dependent and not directly related to the invention) and then fed to the loudspeaker. Is passed to the CODEC (step 614).
【0031】
The actual DSP processing used to generate the volume adjustment signal is similar in result to the processing shown in FIG. 6, but has a slightly different form. These changes are usually introduced to maximize computational efficiency or reduce the demand on the DSP. For example, if the processor power is limited, volume control is performed in the conversion from the μ-law form. After the correct lookup value is found (step 604), the actual read value is obtained by moving the table up or down a predetermined number of places, depending on whether the volume of the signal is increased or decreased from its normal value. It is determined. In this case, the weight parameter effectively corresponds to the number of steps to move up and down to adjust the lookup table (this is clearly the G.711 format separating them according to the positive and negative of the original amplitude and adjusting the volume. Consider the fact that cannot be converted to the opposite polarity). The approach is computationally simple, but provides only non-continuous volume control, not continuous volume control. Alternatively, it is possible to add the logarithm of the volume control value or weight parameter to the μ law value. This approach effectively performs the multiplication in step 606 by logarithmic addition prior to the scale conversion in step 604. Addition is cheaper to use on a computer than multiplication on most processors. The result is then converted back to linear scale for mixing with other audio streams (step 604). This approach allows for precise volume control if the lookup table is created in sufficient detail (although the output is limited to 16 bits). Usually, the logarithm of the weight parameter is obtained from the lookup table or already provided in logarithmic form by the control application. Of course, only a new numerical calculation is required when adjusting the volume control, and this calculation is relatively rare.
【0032】
Similarly, if the available processing power is insufficient to perform continuous rms volume measurements, the processing is performed on any other data block, eg, the absolute difference between continuous samples. A computationally simple algorithm such as summing values is used. Here, the sum of the squared values is executed by logarithmic addition prior to step 604 (ie, before the scale conversion). A simpler approach simply uses the maximum sample value in any audio block as the volume indicator.
【0033】
FIG. 7 shows a screen 700 provided to the workstation of a user participating in a voice conference. As mentioned above, this involves receiving three different audio data streams. However, it is clear that the present invention is not limited to three participants. The screen of FIG. 7 is divided into three areas 701, 702, and 703 by a broken line, and each area represents one participant. However, in reality, these dashed lines do not appear on the screen. For each participant, there is a box 724 containing the participant's name (simply B, C, D in this example). Also, a video image of the audio source transmitted over the network with the audio, or a static bitmap (both provided by the audio source at the start of the conference or already locally present on the workstation and the name of the participant). There is an image window 720 to contain (displayed in response to). For Participant D, the video or still image is not valid and a blank window is shown. The choice of display in the image window (blank, still or video image) depends on the hardware available on the workstation, the bandwidth of the network, and the availability of relevant information.
【0034】
At the bottom of the image window is the volume display 721 (VU meter), which shows the instantaneous volume of the audio stream (calculated in block 608 in Figure 6). The length of the solid line in this display indicates the volume of the audio stream. If there is no audio signal from that participant, the solid line has zero length (ie, is not displayed). The user can therefore determine who is speaking at the conference by looking at which VU meter is active.
【0035】
Below the volume display is a volume control bar 722 that allows the user to adjust the relative volume of its participants. This is achieved by the user pressing the "+" or "-" buttons located on the edge of the bar, respectively, to increase or decrease the volume. This has the effect of increasing or decreasing the weighting parameters used in digital mixing. The indicator in the center of the volume control bar represents the current volume setting (ie the current value of the weight parameter).
【0036】
Finally, next to the name box 724 is the mouse button 723. When this button is pressed, audio output from the participant is alternately permitted and prohibited. If audio output is disabled, the weight parameter is set to 0, and if allowed, the weight parameter is restored to its previous value (ie shown on the volume control bar). If audio from participants is currently banned, this is indicated by a superimpose cross on the silence button (all three audio outputs are currently allowed in Figure 7). Here, according to the DSP processing described above, the VU meter shows 0 when the audio output is prohibited and the silence button is on. It is more straightforward to modify the system so that the VU meter, if necessary, indicates the signal level produced when audio output is actually allowed.
【0037】
FIG. 8 represents the main software components running on the workstation of FIG. 2 to provide the user interface of FIG. The workstation is controlled by an operating system 814, such as Windows provided by Microsoft. In addition, there is suitable communication software 816 on the workstation that enables LAN communication (in some cases, the communication software is effectively included in the operating system). As is known, the operating system and communication software interact with two adapter cards, namely the Token Ring adapter card and the voice adapter card, via device driver 818. The overall processing of voice is controlled by application 810. It uses the capabilities of Application Support Layer 812 and is one such example of Visual Basic provided by Microsoft. There is Basic). The purpose of the application support layer is to facilitate the development of applications, especially with respect to the user interface, but of course applications can also work directly with the operating system.
【0038】
The application uses known programming techniques to control the contents of the window box 720. For example, the VU meter 721 is provided using the features provided by Visual Basic. Visual Basic effectively plays the role of all graphics related to the meter. That is, all the application has to do is supply the relevant numbers. Since Visual Basic is interrupt driven, this is easily achieved by the DSP copying the output volume corresponding to the audio block to the workstation and then calling the interrupt. The interrupt generates an event in the application, informing the application of the new output volume to be copied to the VU meter. In fact, interrupts are used to convey the complete set of volume reads for a set of audio blocks, that is, the availability of one volume read for each audio stream. (The DSP already performs one interrupt per audio block in relation to the output audio signal generated on its workstation for transmission to the network.) Similarly, the volume control bar 722 is also provided in Visual Basic. It is a mechanism (called a "scroll bar"). Visual Basic handles all the graphics associated with the control bar, including the position of the selector, and passes the updated volume value to the application each time the volume is updated by the user. The application then writes this update to the DSP and the volume changes accordingly. The silence button 723 is another display mechanism provided by Visual Basic, which allows for simple on / off control of each audio stream. Each time the silence button is activated, the application needs to remember the previous value of the weight parameter. This will restore this value the next time the silence button is pressed.
【0039】
It will be appreciated that many changes to the user interface described above are possible. For example, VU meters are segmented or replaced by analog level meters. A simpler approach would be an on / off indicator that simply changes color in response to the presence or absence of audio output from the participant. Volume control is achieved with a dial or with a drag-and-drop slider instead of two push-buttons. A silence button may be incorporated into the volume control bar. Such changes are possible within the programming abilities of those skilled in the art.
【0040】
In the systems described above, the user effectively limits the volume control of each audio input stream, while in other systems the user provides more advanced controls such as frequent frequency control (ie treble and basic tuning). Will be done. This is achieved relatively easily by the DSP multiplying the audio signal in the time domain by an FIR or IIR filter. Frequency control is represented to the user in a manner similar to the volume control bar of FIG. 7, and changes in frequency control produce appropriate changes in FIR or IIR filter coefficients. These advanced controls are increasingly desired as the quality of audio signals transmitted over the network improves, for example in systems that use the G.721 audio transmission standard instead of G.711.
【0041】
In summary, the following matters will be disclosed with respect to the constitution of the present invention.
【0042】
(1) A computer workstation connected to a network, each receiving a multiplex audio input stream containing a sequence of digital audio samples from the network, with digital audio samples from each audio input stream in separate queues. A means to store, a means to form a sequence of sets containing one digital audio sample from each queue, and each audio input stream has associated weighting parameters to generate a copolymer for each set of digital audio samples. A means of generating an audio output from the sequence of the copolymerizer and a means of adjusting the weighting parameter in response to user input to control the relative volume in the audio output of the multiplex audio stream. , Computer workstation. (2) The computer workstation according to (1) above, comprising means for providing visual instructions indicating whether each said voice input stream is currently silent. (3) The computer workstation according to (2) above, wherein the visual instructions indicate an instantaneous acoustic volume of each voice input stream. (4) The computer workstation according to (2) or (3) above, wherein the visual instructions are displayed in the vicinity of the visual representation of the origin of each voice input stream. (5) The computer workstation according to (3) or (4) above, comprising means for generating a root mean square value of running squares from a sequence of digital voice samples for each said voice input stream. (6) The input audio data arrives in block units, each containing a predetermined number of digital audio samples, and the visual instructions are updated corresponding to each new block of audio data, (2) to (5). The computer workstation described in any of. (7) The computer workstation according to any one of (1) to (6) above, which includes means for prohibiting audio output from an arbitrary stream of the multiplex audio input stream. (8) any of the above (1) to (7), which includes means for providing the user with visual instructions for the value of the weight parameter, the means adjusting the weight parameter in response to the user's mouse operation. The computer workstation described in Crab. (9) A method of operating a computer workstation connected to a network, each receiving a multiplex audio input stream containing a sequence of digital audio samples, with a separate queue of digital audio samples from each audio input stream. A polymerizer for each set of digital audio samples, with a step to store in, a step to form a sequence of sets containing one digital audio sample from each queue, and a weighting parameter associated with each audio input stream. A step of generating, a step of generating an audio output from the sequence of the amplifier, and a step of adjusting the weight parameter in response to a user input to control the relative volume in the audio output of the multiplex audio stream. Including, how to operate. (10) The operation method according to (9) above, comprising a step of providing a visual instruction indicating an instantaneous acoustic volume of the voice stream for each voice input stream.
【0043】
[Effect of the invention]
As described above, according to the present invention, it is possible to provide a computer workstation in which each participant can identify and control the voices of other participants in a voice conference conducted via a network. it can.
[Simple explanation of drawings]
[Figure 1]
It is a figure which shows the computer network.
[Figure 2]
It is a simplified block diagram of a computer workstation used in a voice conference.
[Fig. 3]
FIG. 2 is a simplified block diagram of an audio adapter card in a computer workstation in Figure 2.
[Fig. 4]
It is a flow chart which shows the process which is executed with respect to the input voice packet.
[Fig. 5]
It is a figure which shows the queue of the input voice packet in playout waiting.
[Fig. 6]
It is a flow diagram which shows the process executed by the digital signal processor on the voice adapter card.
[Fig. 7]
It is a figure which shows the typical screen interface provided to the user of the workstation of FIG.
[Fig. 8]
It is a diagram showing the main software components running on the workstation of Figure 2.
[Explanation of symbols]
2 Local Area Network (LAN) 10 system units 12 Display screen 14 keyboard 16 mouse 22 microprocessor 24, 52 Semiconductor memory 26 PC bus 28 voice card 30 Token Ring Adapter Card 42, 54 A / D converter 44 CODEC 46 Digital Signal Processor (DSP) 48, 50 double buffer 56 Loudspeaker 500 main memory 501, 502, 503 sub queue 700 screens 720 image window 721 Volume display VU meter 722 Volume control bar 723 mouse box Silence button 724 name box 810 application 812 Application Support Layer 814 Operating system 816 Communication software 818 device driver
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
13 members in 8 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 9325924 | United Kingdom | A | |
| 9325924 | United Kingdom | A | |
| 93259240 | United Kingdom | – | |
| 9325924 | – | – | – |
| GB19930025924 | – | – | – |
Members13
| Document | Office | Kind | |
|---|---|---|---|
| GB9325924D0 | United Kingdom | D0 | |
| EP0659006A2 | European Patent Office (EPO) | A2 | |
| GB2284968A | United Kingdom | A | |
| KR950022401A | Republic of Korea | A | |
| JPH07200424A | Japan | A | |
| CN1111775A | China | A | |
| US5539741A | United States of America | A | |
| JP2537024B2This record | Japan | B2 | |
| KR0133416B1 | Republic of Korea | B1 | |
| EP0659006A3 | European Patent Office (EPO) | A3 | |
| TW366633B | Taiwan Province of China | B | |
| CN1097231C | China | C | |
| IN190028B | India | B |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS |
Numbers
- Publication
- 2537024
- Publication, DOCDB
- 2537024
- Publication, EPODOC
- JP2537024B
- Application
- 6256409
- Application, DOCDB
- 25640994
- Application, EPODOC
- JP19940256409
Titles2
- Japanese
- 音声会議システム及びその制御方法
- English
- [Title of Invention] Audio Conferencing System and Control Method thereof
Classification
- CPC, 9
- H04L12/1813
- H04M3/42161
- H04M3/567
- H04M3/568
- H04N7/15
- H04L65/1083
- H04L65/403
- H04L65/1059
- H04L65/613
- IPC, 5
- G06F15 00
- G06F13 00
- H04L12 18
- H04M3 56
- H04N7 15
