Audio conferencing system conference covering communications of multiple points
Abstract
A computer workstation receives multiple audio input streams over a network in an audio conference. The audio input streams are kept separate by storing them in different queues. Digital samples from each of the queues are transferred to an audio adapter card 28 for output. A digital signal processor 46 on the audio adapter card multiplies each audio stream by its own weighting parameter, before summing the audio streams together for output. Thus the relative volume of each of the audio output streams can be controlled. For each block of audio data, the volume is calculated and displayed to the user, allowing the user to see the volume in each audio input stream independently. The user is also provided with volume control for each audio input stream, which effectively adjusts the weighting parameter, thereby allowing the user to alter the relative volumes of each speaker in the conference.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
11 claims: 11 independent, 0 dependent
- 1一種電腦工作站可供連接至一網路,並供接收來自網路的多重音頻輸入串流,每一音頻串流各包含一序列的數位音頻樣本,該工作站包含:可將來自於各音頻輸入串流的數位音頻樣本儲存於一個分離的排序之中之裝置;可構成各包含來自每一排序的一個數位音頻樣本的一組序列之裝置;可為每一組的數位音頻樣本產生一加權和之裝置,其每一音頻輸入串流各具有一個加權參數與之相關聯;可由加權和之序列產生一音頻輸出之裝置;其特徵為其可反應於使用者輸入,而可調整該加權參數,以控制在多重音頻串流的音頻輸出之中的相對音量之裝置。
- 2如申請專利範圍第1項之電腦工作站,更包含有裝置可為該些多重音頻輸入串流中的每一個各產生一個視覺的指示,不論該串流目前是否靜止無聲。
- 3如申請專利範圍第2項之電腦工作站,其中該視覺指示更為該些多重音頻輸入串流中的每一個,各指示其在該音頻串流之中的即時音量。
- 4如申請專利範圍第2或3項之電腦工作站,其中該視覺指示係被顯示於該音頻輸入串流之來源的一個視覺代表的近旁。
- 5如申請專利範圍第3項之電腦工作站,其中該些多重音頻輸入串流中的每一個,更包含可由其中之數位音頻樣本的序列產生一移動的均方根值。
- 6如申請專利範圍第4項之電腦工作站,其中該些多重音頻輸入串流中的每一個,更包含可由其中之數位音頻樣本的序列產生一移動的均方根值。
- 7如申請專利範圍第2、3或5項之電腦工作站,其中進來的音頻資料係以區塊的方式到達,其各包含一個預定數量的數位音頻樣本,且該視覺指示係為每一個區塊的音頻資料更新一次。
- 8如前列申請專利範圍中1、2、3或5項之電腦工作站,更包含可將來自於該些多重音頻輸入串流中之任何一個的音頻輸出加以關斷的裝置。
- 9如前列申請專利範圍中1、2、3或5項之電腦工作站,更包含可為使用者提供該些加權參數數值的一視覺指示的裝置,該裝置係反應於使用者的滑鼠操作而調整該加權參數。
- 10一種操作電腦工作站之方法,該工作站被連接至一網路上以接收多重音頻輸入串流,每一個音頻串流各包含一序列的數位音頻樣本,該方法包含下列之步驟:將來自於每一音頻輸入串流的數位音頻樣本各儲存於一個分離的排序之中;構成包含了來自於每一個排序的一個數位音頻樣本組的一個序列;為數位音頻樣本的每一個組產生一加權和,每一音頻輸入串流各具有與之相關聯的一個加權參數;由加權和之序列產生一音頻輸出;其特徵為其可反應於使用者輸入,而可調整該加權參數,以控制在多重音頻串流的音頻輸出之中的相對音量之裝置。
- 11如申請專利範圍第10項之操作電腦工作站之方法,更包含為該些多重音頻輸入串流中的每一個各產生一個視覺的指示之步驟,各指示其在該音頻串流之中的即時音量。
Independent claims11
42 paragraphs, as filed
The present invention relates to the processing of multiple streams of audio data received from a network by a computer workstation.
General voice signals are transmitted via standard analog telephone lines. However, due to the increasing number of points equipped with local area networks (LANs, local area networks) and the increasing importance of multimedia communications, there is considerable interest in using LANs to carry voice signals. This type of work is, for example, in the book "Local Computer Network" (edited by P Ravasio, G Hopkins and N Naffah; North Holland, 1982), No. 13- Page 21, D. "Using Local Area Networks for Carrying Online Voice" by Cohen, and "Using Local Area Networks for Carrying Online Voice", pages 39-65, by P Ravasio, R Marcogliese, and R Novarese. "Voice Transmission over an Ethernet Backbone" has been described in these two articles. The basic principle of this type of program is to use the first terminal or workstation to digitally sample a voice input signal at a general rate (for example, 8 kHz). A number of samples can then be combined into a data packet for transmission over the network to a second terminal, which then feeds the samples to a speaker or equivalent device for constant control again The rate is played out.
One problem with using a local area network to carry voice data is that the transmission time through the network can vary. In this way, the arrival of a message packet to a target node is often delayed and irregular. If the packet is played out in an irregular manner, it will have an extremely adverse effect on the recognition of the voice signal. Therefore, the voice passing through the LAN program will use a certain degree of buffering at the receiving end in order to absorb these irregularities. But you must be careful to avoid introducing too much delay between the original voice signal and the audio output of the destination. Such too much delay will make the interaction of two-way talks difficult (the situation is similar to that of general telephone calls across the Atlantic. The excessive delay can be quite annoying, the same situation). There is a system in the IBM Technical Disclosure Bulletin in April 1993, Vol 36 No 4, pages 255-257, B Aldred, It is described in an article "Adaptive Audio Playout Algorithm for Shared Packet Networks" by R Bowater and S Woodman et al., in which the maximum allowable time is exceeded. All arriving packets are intercepted. The amount of buffering is controlled in an adaptive manner, depending on the number of packets to be intercepted (any other appropriate measure for late arrivals can be applied). If the number of intercepted packets is quite high, the amount of buffering is increased. On the contrary, if the amount of intercepted packets is low, the degree of buffering will be reduced. The size of the buffer is changed by temporarily changing the playback rate (this affects the pitch; a less noticed technique is to detect periods of silence and manually increase or decrease them in an appropriate manner) .
Another key point of audio communication is talks involving multipoint communication. Compared with this kind of multipoint communication talks, it is two-way point-to-point communication. When implemented on a traditional analog telephone line, an audio voice conference requires a central hub for each participant to send an audio signal. The central hub mixes the incoming signals, may adjust it according to different levels, and sends the sum of the signals from all other participants to each participating party (deducting the signal from the specific node). US Patent No. 4650929 discloses a centralized video/audio conferencing system in which each participant can adjust the volume relative to other participants.
The use of a centralized mixing node, usually called a multipoint control unit (MCU, multipoint control unit), has been converted to some multimedia (audio plus video) workstation conferencing systems. For example, US Patent No. 4,710,917 describes a multimedia conference system in which each participant transmits audio signals to a central mixing unit, and the central mixing unit receives the audio signals. Other multimedia conferencing systems are in K Watabe, S Sakata, K Maeno, H Fukuoka, and T Ohmori's "Distributed Multiparty Desktop Conferencing System: MERMAID" ("Distributed Multiparty Desktop Conferencing System: MERMAID"), pages 27-38, CSCW '90 (Proceedings of the Conference on Computer-Supported Cooperative Work, 1990, Los Angeles), and E Addeo, A Gelman and A Dayao, etc. "Broadband Network Personal Multimedia "Personal Multimedia Multipoint Communications Services for Broadband Networks", pages 53-57, Vol 1, IEEE GLOBECOM, 1988.
However, the use of centralized MCUs or integrated nodes has some disadvantages. First of all, the architecture of most LANs is based on a point-to-point arrangement, so there is no obvious central node. In addition, the system completely relies on the continuous availability of the selected central node for the operation of the meeting. There may also be problems with echo suppression (the central node must be careful not to include the audio signal from one node in the sum signal broadcast to that node).
If a decentralized voice conference system is used, these problems may be avoided, where each node has a separate audio signal from another node in the conference. U.S. Patent No. 5,127,001 describes such a distributed system, and discusses the synchronization problem caused by the varying transmission time of packets passing through the network. US Patent No. 5,127,001 overcomes this problem by maintaining a separate ordering of incoming audio packets for each source node. This can be used in the same way as the simple point-to-point communication described above, and effectively absorb the chattering in the arrival time. At regular time intervals, a group of audio packets are read out, and one packet will come out of each sequence, and they will be summed up for playback.
A problem with the voice meeting system, as found in the aforementioned MERMAID system, is to determine who is speaking at any given moment. US Patent No. 4,893,326 describes a multimedia conferencing system in which each workstation automatically detects whether its user is speaking. This information is then fed to a central control node, which then switches the video signal so that each participant can see the current speaker on his screen. Such a system requires both video and audio capabilities to be able to operate, and it also relies on a central video switching node, so it cannot be applied to a completely distributed system.
It is described in "Personal Multimedia-Multipoint Teleconference System" ("Personal Multimedia-Multipoint Teleconference System") by H Tanigawa, T Arikawa, S Masaki and K Shimamura, IEEE INFOCOM 91 Proceedings, Vol 3, 1127-1134 A distributed multimedia conference system. This kind of system provides sound localization for a stereo workstation. When a window containing a video signal from a meeting participant is moved across the screen from right to left, the apparent source of the corresponding audio signal is also the same To move. This approach provides limited assistance in identifying a speaker. The Japanese Abstract JP 02-123886 describes a more complex facility in which a bar graph is used to visualize the output sound level associated with a window adjacent to a video signal containing the sound source.
Conventional technology therefore describes a variety of voice conferencing systems. Although the conventional centralized telephone audio conference has been widely used from a technical point of view and is well-known to the public, there is still a lot of work to be done to increase the performance of voice conferences in a desktop environment.
Therefore, the present invention is to provide a computer workstation that can be connected to a network and receive multiple audio input streams from the network. Each audio stream contains a sequence of digital audio samples. The workstation includes: Devices that can store the digital audio samples from each audio input stream in a separate sequence; devices that each contain a set of sequences of one digital audio sample from each sequence; can be for each set A device for generating a weighted sum of digital audio samples, each audio input stream has a weighting parameter associated with it; a device that can generate an audio output from a sequence of weighted sums; its feature is that it can respond to user input, The weighting parameter can be adjusted to control the relative volume of the audio output of multiple audio streams.
The present invention recognizes that the provision of voice conferences on a distributed network, where each node receives a separate audio stream from all other participants, naturally allows for additional functionality, which was previously only available It can be achieved only at the greater difficulty and cost of a centralized conference system. In particular, each user can adjust the relative volume of all other participants according to their personal preferences. This may be necessary, for example, if they need to pay special attention to a particular topic of the meeting, or when language problems arise (for example, someones heavy accent may make it difficult for others to understand, or perhaps the meeting involves To the problem of simultaneous translation). In addition, even during the talk, the system can react to user input to change the relative volume of different participants. In order to allow such control, the incoming audio signals are kept separate from each other, and are placed into different sorts according to their sources before being weighted with appropriate volume control coefficients (the sorts are logically separated Open storage, although in fact they may still belong to adjacent or combined). Only then can they be combined to produce the final audio output. The present invention therefore considers that a decentralized voice conferencing system is particularly suitable for the provision of individual control of relative volume.
Preferably, the workstation should include a device that can generate a visual indication for each of the multiple audio input streams, regardless of whether the stream is currently silent or not. This can overcome a recognized problem in current voice talks, namely, who is currently speaking. The visual indicator can simply be some form of on/off indicator, such as a light or equivalent facility, but in the preferred embodiment, it is made as one of the multiple audio input strings that can be for each of them. The stream indicates the display of the real-time sound volume in an audio stream. In other words, the display provides a complete indication of the volume of the relevant participant. The volume output can be calculated from a sequence of digital audio samples based on a moving root mean square value, or, if the processing power is limited, a simpler algorithm can be adopted, such as using a The largest digital audio value among a predetermined number of samples. Usually, incoming audio data arrives in the form of blocks, each containing a predetermined number of digital audio samples, and the visual indication is also updated once for each new block of audio data. In this way, the volume number is usually calculated on a per-block basis.
It is also preferable that the visual indicator is displayed near a visual representation of the source of the audio input stream, such as a moving image or a still image. In the former case, a complete multimedia conference network is required, while the latter can be provided on a much lower bandwidth network that cannot support video signal transmission. Such a visual indication, whether stationary or moving, allows easy identification of any audio source.
Preferably, the workstation should further include a device that can provide the user with a visual indication of the value of the weighting parameter, and the device adjusts the weighting parameter in response to the user's mouse operation. This can be made in the form of a scroll bar, etc., with one for each audio input stream, and it is positioned adjacent to the visual indication of the output volume of the stream. If the computer workstation further includes a device that can shut off the audio output from any of the multiple audio input streams, it is more convenient. In this way, the user has a complete set of volume controls for each audio input stream.
The present invention also provides a method for operating a computer workstation, the workstation is connected to a network to receive multiple audio input streams, each audio stream contains a sequence of digital audio samples, the method includes the following steps: The digital audio samples in each audio input stream are stored in a separate sequence; a sequence containing a group of digital audio samples from each sequence is formed; a weight is generated for each group of digital audio samples And, each audio input stream has a weighting parameter associated with it; an audio output is generated from the sequence of the weighted sum; its feature is that it can react to user input, and the weighting parameter can be adjusted to control the The relative volume of the audio output of multiple audio streams.
An embodiment of the present invention will now be described by way of example with reference to the following drawings, in which: Figure 1 is a schematic diagram of a computer network; Figure 2 is a computer workstation that can be used for voice conversations Figure 3 is a simplified block diagram of an audio interface card in the computer workstation in Figure 2; Figure 4 is a flow chart showing the processing performed on an incoming audio packet; 5 shows the sequence of incoming audio packets waiting to be played out; Figure 6 is a flow chart showing the processing of the digital signal processor on the audio interface card; Figure 7 shows what is shown in Figure 2 A typical screen interface of the user of the workstation; Fig. 8 is a simplified diagram showing the main software components running on the workstation of Fig. 2.
FIG. 1 is a schematic diagram of computer workstations AE connected together in a local area network (LAN) 2. These workstations are participating in a multilateral meeting, and each workstation is broadcasting its audio signal to every other workstation in the meeting. In this way, each workstation will receive a separate audio signal from other workstations. The network shown in Figure 1 has a symbol loop architecture, in which a symbol circulates between workstations. Only the workstation that currently has a token is allowed to send a message to another workstation. It should be understood that the time required for a message to be transmitted in a loop is extremely short. In other words, for example, a message sent by A can be received by all other terminals almost simultaneously. This is why a marking system was chosen to avoid interference between two nodes when trying to send messages at the same time.
As will be described in greater detail below, one-way audio communication on a LAN will have a bandwidth of 64 kHz. In the talk in Figure 1, each node will broadcast the remaining four nodes, which means that a total of 5x4x64 kHz (1.28 MHz) bandwidth is required. It is well within the performance range of standard sign loop networks. This type of network can support a transmission rate of 4 or 16 Mbits per second. However, if a large-scale meeting is held, it can be understood that the demand for bandwidth will soon become a problem. Obviously, the future network is expected to have a much higher bandwidth.
It is noted that the present invention can be implemented on many types of networks with different structures and configurations from the token loop, but the technical requirements necessary for supporting voice conversations such as bandwidth and latency must of course still be met.
Figure 2 is a simplified block diagram of a computer workstation that can be used for voice conferencing. This computer has a system unit 10, a display screen 12, a keyboard 14 and a mouse 16. The system unit 10 includes a microprocessor 22, a semiconductor memory (ROM/RAM) 24, and a bus 26 that can transmit data. The computer in Figure 2 can be any conventional workstation, such as an IBM PS/2 computer.
The computer in Figure 2 is equipped with two interface cards. The first piece of it is a token loop interface card 30. This card, together with the attached software, can allow messages to be transmitted to or received from the token loop network as shown in Figure 1. The operation of the mark loop card is well-known to the public, so it will not be explained in detail here. The second interface card is an audio card 28, which is connected to a microphone and a speaker (not shown) for audio input and output respectively.
The details of this audio card are shown in Figure 3. The card shown in the figure is used in this particular embodiment. It is an M-Wave card available from IBM, although other interface cards with similar functions can also be used. This card contains an A/D converter 42, which can digitize the audio signal sent by an attached microphone 40. This A/D converter is attached to a CODEC 44, which can sample the incoming audio signal at a rate of 44.1 kHz to become a 16-bit sample (which corresponds to the compact CD standard) Sampling rate/size). The digitized samples are then sent to a digital signal processor (DSP) 46 on the card via the double buffer 48 (that is, the CODEC loads a sample into half of the double buffer, and the CODEC simultaneously Read the previous sample from the other half). The DSP is controlled by one or more programs stored in the semiconductor memory 52 on the card. Data can be transmitted between the DSP and the main PC bus by the DSP.
The audio signal to be played is received by the DSP 46 from the PC bus 26, and processed into an audio input by the microphone in a conversation manner. That is, the output audio signal is sent through the DSP 46 and a double buffer 50 to the CODEC 44, and from this to a D/A converter 54, and finally to the speaker 56 or other appropriate output device.
In the specific embodiment shown here, using standard resampling techniques, the DSP is programmed to convert the samples from the CODEC from 44.1 kHz, 16 bits into a new digital signal with a sampling frequency of 8 kHz. An 8-bit sample on a μ method scale (essentially a logarithmic scale) conforms to the CCITT standard G.711. The total bandwidth of the signals sent to the workstation for transmission to other terminals is therefore 64 kHz. Once again using the technique of resampling, the DSP also performs a reverse conversion on the incoming signal received by the PC, that is, the signal is converted from 8 bits, 8 kHz to 16 bits, 44.1 kHz. Note that this conversion between the two sampling formats is only necessary due to the selection of specific hardware, and is not directly related to the present invention itself. In this way, for example, because many other audio cards themselves include direct support for the 8 kHz format, the CODEC can completely operate in the 8 kHz format according to the G.711 format (in another way, 44.1 kHz The samples are still reserved for transmission on the network, although the much higher bandwidth and greatly increased processing speed required for it make it seem impossible to do, unless the audio signal it transmits has a specific CD quality is required; as far as normal voice communication is concerned, 64 kHz bandwidth signal in G.711 format is suitable).
The data is transmitted between the audio interface card and the workstation in 64-byte blocks: that is, 8 ms audio data, in terms of 8-bit data sampled at 8 kHz. The workstation then only processes the entire block of data, and each data packet sent by the workstation or received by the workstation usually contains a single 64-byte block of data. The block size is selected as 64 bytes in order to minimize the granular nature of the system (which causes delay), while maintaining the efficiency of both the processing inside the workstation and the transmission on the network. Compromise. In other systems, a block size of 32 or 128 bytes, for example, may be more appropriate.
The operation of a computer workstation regarding the transmission of audio data is widely known in the art of learning, so it will not be described in detail here. The audio interface card necessarily receives an input signal, whether it is an analog form of a microphone, or from other audio sources, such as a CD player, and generates blocks of digital audio data. These blocks are then transferred to the main memory of the workstation, and from there to the LAN interface card (in some architectures, it may be possible to transfer the blocks from the audio interface card directly into the LAN interface card , Without going through the memory of the workstation). The LAN interface card generates a data packet containing digital audio data and header information, identifies the source and destination nodes, and the packet is then sent to the desired receiver on the network. It can be understood that, in any two-way or multi-directional communication, this transmission procedure will be performed simultaneously with the receiving procedure described below.
The receiving system of audio data packets of a computer workstation is shown in Figure 4. When a new packet arrives (step 402), the LAN interface card informs a program executed on the workstation's microprocessor and provides information to the program to identify the source of the data packet. The program then transfers the incoming 64-byte audio block into a sequence in the main memory (step 404). As shown in FIG. 5, the sorting in the main memory 500 is actually composed of a set of audio blocks from each of the different source nodes, and their separate sub-sorts. In this way, one sort includes audio blocks from one source node, another sort includes audio blocks from another source node, and so on. In Figure 5, there are three sub-sequences 501, 502, and 503, corresponding to the audio data from nodes B, C, and D; the number of sub-sequences will of course depend on the number of participants in the voice conference. The program uses the information in each received packet to identify the source node in order to place the incoming audio data in the correct order. Index P<sub>B</sub>, P<sub>C</sub>With P<sub>D</sub>Point out the position of the end of the sorting respectively, and update it every time a new packet is added. These packets are then removed for further processing starting from the bottom of the sub-sequence (as shown by "out" in Figure 5). The sub-sequence in Fig. 5 is therefore a substantive first-in-west-out order, and can be achieved by using conventional programming techniques. Note that in addition to the support for multiple (parallel) sorting, the processing of incoming audio blocks described above is exactly the same as the conventional technology so far, whether it is for individual sub-sorting or combining. Sorting as a whole allows the use of equivalent buffering techniques if necessary.
The actions performed by the DSP on the audio interface card are shown in Figure 6. DSP is executed in a cycle, processing a new set of audio blocks every 8 milliseconds to ensure a continuous audio output signal. In this way, every 8 ms, DSP uses DMA to capture, from corresponding to different nodes Read out an audio block in each sub-sequence of, that is, as shown in Figure 5, a block from the bottom of the sorts B, C and D (step 602: that is, in this case, it is M=3). These blocks are treated as representing the same time interval: in the final output, they will be added together to produce a single audio output for that time interval. The DSP therefore effectively performs a digital mixing function on multiple audio input streams. Using a search table, individual samples in the 64-byte block can be converted out of the G.711 format (which is essentially a logarithmic form) and become a linear scale (step 604). Each individual sample is then multiplied by a weighted parameter (step 606). Each received audio data stream has a separate weighting parameter; that is, for the three sub-orders in Figure 5, the audio stream from node B has an audio stream, and node C The audio stream of has an audio stream, and the audio stream of node D also has an audio stream. These weighting parameters are used to control the relative volume of audio signals from different sources.
The DSP maintains a root mean square (rms) mobility record for each audio stream (step 608). Such an rms value is usually generated for the audio data of each block by using the sum of the generated sum and the square of the value in the block (that is, every 8 milliseconds). The rms value represents the volume of the individual audio input stream, and is used to provide the user with volume information in the manner described below.
Once the digital audio samples have been multiplied by the appropriate weighting parameters, they are added together (step 608; note that this can effectively be done concurrently with the processing of step 606). In this way, a single sequence of digital audio samples can be generated, representing the weighted sum of multiple input audio streams. This sequence of digital audio samples is sampled to 44.1 kHz before being sent to the CODEC (step 612) to be supplied to the speakers (step 610, although as previously described, this is related to the hardware, and It is not directly related to the present invention).
Note that the actual DSP processing used to generate the volume adjustment signal may be somewhat different from that shown in Figure 6, although in essence the final result is similar. Such changes can usually be introduced to maximize computational efficiency or reduce the need for DSP. For example, if the effectiveness of the processor is limited, the volume control can be performed when switching away from the μ-law format. In this way, after the correct search table value has been found (step 604), the actual read value can be determined by moving up or down a predetermined number of positions in the table, depending on whether the volume of the signal is normal. It depends on the value increase or decrease. In this case, the weighting parameter is actually to adjust the number of steps in the search table up or down (obviously it allows the G.711 format to separate the original amplitude according to whether it is positive or negative, and the volume adjustment You cannot convert one to another). The above method is quite simple in calculation, but only provides segmented rather than continuous volume control. According to another method, the logarithmic value of the volume control value, or the weighting parameter, can also be added to the μ-law value. This method can effectively perform the multiplication in step 606 before the logarithmic addition is used to perform the scale conversion. For most processors, this method is computationally cheaper than multiplication. The result can then be converted back to a linear scale (step 604) for mixing with other audio streams. This approach does not allow fine control of the volume, unless the search table can have sufficient detail (although, note that its output is still limited to 16 bits). Usually the logarithm of the weighting parameter can be obtained from a search table, or in another way, it may have been provided in the form of a logarithm for the purpose of control. Of course, when the volume control is adjusted, only a new logarithmic value needs to be calculated, which is quite unusual.
Similarly, if the available processor is not effective enough to perform continuous rms volume measurement, its processing may be performed every other block of data, or in another way, some calculations are simpler. Methods, such as adding up the absolute values of the differences between successive samples, can also be used. Note that the addition of the squared values can be performed before step 604 using logarithmic addition (that is, before the scale conversion). A simpler approach is to simply use the largest sample value in any audio area as an indicator of volume.
Fig. 7 shows a typical screen interface shown to the user of the workstation of Fig. 2. As previously discussed, this involves the reception of three different streams of audio data, although obviously the present invention is not limited to only three participants. The screen in Figure 7 has been divided into three areas 701, 702, and 703 by dashed lines, each representing a participant, although in reality, these dashed lines do not appear on the screen. Associated with each participant is a box 724 containing the participant's name (in this example, it is simply represented by B, G, and D). There is also an image window 720, which can be used to accommodate a video image of an audio source, or a static bitmap (whether it is provided by the audio source at the beginning of the meeting, or perhaps it has been previously stored in the workstation, and It can be revealed in response to the participants name), transmitted through the Internet along with audio signals. In the case of Participant D, there is no dynamic video image or still image available, so a blank window appears. The choice of display in the image window (blank, still or video image) depends on the hardware available on the workstation, the bandwidth of the network, and the acquisition of relevant information.
Below the image window is a volume display 721 (a VU meter), which can indicate the current volume of the audio stream (as described in block 608 of FIG. 6). The length of the solid line in this display shows the volume of the audio stream. If there is no audio signal coming from the participant, the length of the solid line is zero (it disappears). The user can therefore determine who is speaking during the meeting by checking whose VU meter is active.
Below the volume display is a volume control bar 722 that allows the user to adjust the relative volume of the participant. This is done by the user pressing the "+" or "-" buttons at both ends of the bar to thereby increase or decrease the volume respectively. This has the effect of correspondingly increasing or decreasing the weighting parameters applied to the digital blending. The indicator in the middle of the volume control bar represents the current volume setting (that is, the current value of the weighting parameter).
Finally, next to the name box 724 is a sound reduction button 723. Press this button to switch between turning off and enabling audio output from that participant. When the audio output is turned off, the weighting parameter is set to zero, and when enabled, the weighting parameter is restored to its previous value (ie, as displayed on the volume control bar). If the audio from a participant is temporarily turned off, this will be indicated by a cross on the mute button (in Figure 7 all three audio outputs are currently enabled). Note that when the DSP is used for the aforementioned processing, when the mute button is turned on, the audio output is turned off, and the VU meter is also displayed as zero. If necessary, the system will be directly modified so that the VU meter displays the level of the signal generated when the audio output is actually enabled.
FIG. 8 shows the main software components that can be executed on the workstation of FIG. 2 to provide the user interface of FIG. 7. The workstation is controlled by an operating system 814 such as the Windows operating system of a micro company. Also present on the workstation is the appropriate communication software 816 to enable LAN communication (in some cases, the communication software is actually included in the operating system). The operating system and the communication software, the symbol loop and the audio interface card, two interface cards, interact with each other through the device driver 818 known in the art. The entire audio processing is controlled by the application 810. It utilizes the function of the application support layer 812. In an implementation example, the support layer is the Visual Basic language (Visual Basic) of Microsoft Corporation. The purpose of the application support layer is to achieve the development of the application, especially with regard to the user interface, but of course it is also possible to make the application work directly with the operating system.
The application program controls the content of the window box 720 according to a known program planning technique. For example, the VU table 721 is provided by the functions in the view basic language, which is essentially responsible for all table-related graphics: all the application needs to do is to supply appropriate values. Since the view base language is interrupt-driven, it is easy to copy the output volume from the DSP to the output of an audio block input to the workstation, and then connect the interrupt. The interruption generates an event in the application, notifying it of the new output volume that can now be copied to the VU meter. In fact, the interrupt is used to notify that a complete set of volume readings for a set of audio blocks is available; that is, one volume reading for each audio stream (note that the DSP has been implemented in each audio block An interruption is related to the outgoing audio signal generated on the workstation for transmission on the network). Similarly, the volume control bar 722 is also a feature provided in the basic language of the view (called a scroll bar). The basic language of the view is responsible for all graphics related to the control bar, including the position of the selector, and simply sends an updated volume value to the application when the user makes adjustments. The application can then write this updated value into the DSP to change the volume accordingly. The mute button 723 is another display feature provided by the basic language of the view, allowing simple on/off control of each audio stream. Note that when the mute button is activated, the application needs to remember the previous value of the weighting parameter so that it can be stored when the mute button is pressed again.
It should be understood that many variations of the aforementioned user interface are possible. For example, the VU table can be segmented or replaced with an analog level table. As another and even simpler way, it is an on/off indicator that changes color based on whether there is any audio output from the participant. The volume control function can be controlled by using a rotating dial or pulling a slider instead of two buttons. The mute button can also be integrated into the volume control bar. Such changes should be within the program planning ability of those who are accustomed to this skill.
Although in the aforementioned system, the user actually limits the volume control of each audio input stream, in other systems, more advanced controls such as frequency control can also be provided to the user (that is, Adjustment of treble and bass). This can be implemented by making the DSP use a FIR or IIR filter to multiply the audio signal in the time zone. For the user, the frequency control is represented in a similar way to the control bar in FIG. 7, and the change of the frequency will also cause an appropriate change in the coefficient of the FIR/IIR filter. As the quality of the audio signal transmitted over the network improves, these advanced controls will become more necessary, for example, in systems that use G.721 instead of G.711 audio transmission standards.
Voice conference system
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| TWI795759B | Cited by | Taiwan Province of China | Examiner |
13 members in 8 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 9325924 | United Kingdom | A | |
| 9325924 | United Kingdom | A | |
| 19930025924 | – | – | – |
| GB19930025924 | – | – | – |
Members13
| Document | Office | Kind | |
|---|---|---|---|
| GB9325924D0 | United Kingdom | D0 | |
| EP0659006A2 | European Patent Office (EPO) | A2 | |
| GB2284968A | United Kingdom | A | |
| KR950022401A | Republic of Korea | A | |
| JPH07200424A | Japan | A | |
| CN1111775A | China | A | |
| US5539741A | United States of America | A | |
| JP2537024B2 | Japan | B2 | |
| KR0133416B1 | Republic of Korea | B1 | |
| EP0659006A3 | European Patent Office (EPO) | A3 | |
| TW366633BThis record | Taiwan Province of China | B | |
| CN1097231C | China | C | |
| IN190028B | India | B |
Numbers
- Publication
- 366633
- Publication, DOCDB
- 366633
- Publication, EPODOC
- TW366633B
- Application
- 84102303
- Application, DOCDB
- 84102303
- Application, EPODOC
- TW19950102303
Titles4
- Chinese
- 語音會議系統
- English
- AUDIO CONFERENCING SYSTEM
- Unlabeled
- 語音會議系統
- Unlabeled
- Voice conference system
Classification
- CPC, 9
- H04L12/1813
- H04M3/42161
- H04M3/567
- H04M3/568
- H04N7/15
- H04L65/1083
- H04L65/403
- H04L65/1059
- H04L65/613
- IPC, 5
- G06F15 00
- G06F13 00
- H04L12 18
- H04M3 56
- H04N7 15