Audio conferencing system
Abstract
A computer workstation receives multiple audio input streams over a network within an audio conference. These audio input streams are individualized by storing the streams in different queues. Digital samples from each queue are sent to the audio adapter card 28 for output. The digital signal processor 46 on the audio adapter multiplies each audio stream by its weight variable before summing the audio streams for output. So you can control the relative volume of each of the audio streams. For each audio data block, it is calculated and displayed to the user so that the user can independently view the volume of each audio stream. The user will also be able to have volume control for each audio input stream, which effectively allows the weighting variable to be adjusted, thus allowing the user to vary the relative volume of each speaker in the conference.

Term
Term ended
Expired 16 December 2014, 11.8 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
10 claims: 4 independent, 6 dependent
- 1네트워크에 접속되고 그 각각의 디지탈 오디오 샘플들의 시퀀스(sequence of digital audio sample)를 포함하는 다수의 오디오 입력 스트림(multiple audio input streams)을 상기 네트워크로부터 수신하기 위한 컴퓨터 워크스테이션에 있어서, 각각의 오디오 입력 스트림으로부터의 디지탈 오디오 샘플들을 별도의 큐(separate queue)에 저장하기 위한 수단, 각각의 큐로부터의 하나씩의 디지탈 오디오 샘플을 각각 포함하는 세트들의 시퀀스를 형성하기 위한 수단, 상기 각각의 오디오 입력 스트림이 그와 연관된 가중 변수(weighting parameter)를 가지고 있어, 디지탈 오디오 샘플들의 각각으 세트에 대해 가중합계(weighted sum)를 생성하기 쉽다. 상기 가중 힙계들의 시퀀스로부터 오디오 출력을 발생시키기 위한 수단, 및 사용자 입력에 응답하여 상기 다수의 오디오 스트림의 오디오 출력 내에서 상대적인 볼륨을 제어하도록 상기 가중 변수들을 조정하기 위한 수단을 포함하는 것을 특징으로 하는 컴퓨터 워크스테이션.
- 2제1항에 있어서, 상기 다수의 오디오 입력 스트림 각각에 대해 해당 스트림이 현재 무음 상태에 있는(silent) 지의 여부에 대한 시각적인 표시(visual indication)를 제공하기 위한 수단을 더 포함하는 것을 특징으로 하는 컴퓨터 워크스테이션.
- 3제2항에 있어서, 상기 시각적인 표시는 상기 다수의 오디오 입력 스트림 각각에 대해 해당 오디오 스트림 내의 순간적인 소리 볼륨을 표시하는 것을 특징으로 하는 컴퓨터 워트스테이션.
- 4제2항에 또는 제3항에 있어서, 상기 시각적인 표시는 해당 오디오 입력 스트림의 근원(origin)을 시각적으로 표현한 것(visual representation)에 인접하여 디스플레이되는 것을 특징으로 하는 컴퓨터 워크스테이션.
- 5제3항에 있어서, 상기 다수의 오디오 입력 스트림 각각에 대해, 디지탈 오디오 샘플들의 시퀀스로부터 이동 실효값(running root-mean-square value)을 발생시키기 위한 수단을 더 포함하는 것을 특징으로 하는 컴퓨터 워크스테이션.
- 6제2항 또는 제3항에 있어서, 인입 오디오 데이타는 블럭으로 도착하고, 상기 각각의 블럭은 선정된 수의 디지탈 오디오 샘플을 포함하고 있으며, 상기 시각적인 표시는 오디오 데이타의 새로운 블럭 각각에 대해 갱신되는 것을 특징으로 하는 컴퓨터 워크스테이션.
- 7제1항, 제2항, 제3항 및 제5항 중 어느 한 항에 있어서, 상기 다수의 오디오 입력 스트림 중 어느 하나로부터의 오디오 출력을 디스에이블시키기 위한 수단을 더 포함하는 것을 특징으로 하는 컴퓨터 워크스테이션.
- 8제1항, 제2항, 제3항, 및 제5항 중 어느 한 항에 있어서, 상기 가중 변수들의 값에 대한 시각적인 표시를 사용자에게 제공하기 위한 수단을 더 포함하고, 상기 가중 변수들의 값에 대한 시각적 표시 제공 수단은 상기 가중 변수들을 조절하기 위해 사용자 마우스 조작에 응답하는 것을 특징으로 하는 컴퓨터 워크스테이션.
- 9다수의 오디오 입력 스트림을 수신하기 위해 네트워크에 접속되고, 각각의 오디오 스트림은 디지탈 오디오 샘플들의 시퀀스를 포함하게 되는 컴퓨터 워크스테이션을 조작하는 방법에 있어서, 상기 각각의 오디오 입력 스트림으로부터의 디지탈 오디오 샘프들을 별도의 큐에 저장하는 단계, 각각의 큐로부터의 하나씩의 디지탈 오디오 샘플을 각각 포함하는 세트들의 시퀀스를 형성하는 단계, 상기 각각의 오디오 입력 스트림이 그와 연결된 가중 변수를 가지고 있어, 디지탈 오디오 샘플들의 각각의 세트에 대한 가중 합계를 생성하는 단계, 상기 가중 합계들의 시퀀스로부터 오디오 출력을 발생시키는 단계, 및 사용자 입력에 응답하여 상기 다수의 오디오 스트림의 오디오 출력 내에서 상대적인 볼륨을 제어하도록 상기 가중 변수들을 조저하는 단계를 포함하는 것을 특징으로 하는 컴퓨터 워크스테이션 조작 방법.
- 10제9항에 있어서, 상기 다수의 오디오 입력 스트림 각각에 대해, 해당 오디오 스트림에서 순간적인 소리 볼륨에 대한 시각적인 표시를 제공하는 단계를 더 포함하는 것을 특징으로 하는 컴퓨터 워크스테이션 조작 방법.
Independent claims10
60 paragraphs, as filed
[Name of invention]
audio conferencing system
[Brief Description of Drawings]
1 is a schematic diagram of a computer network;
Figure 2 is a schematic block diagram of a computer workstation for use in audio conferencing;
Fig. 3 is a schematic block diagram of an audio adapter card of the computer workstation of Fig. 2;
Fig. 4 is a flowchart showing processing executed on incoming audio packets;
Fig. 5 shows a queue of incoming audio packets waiting to be played out;
Fig. 6 is a flowchart showing processing executed by a digital signal processor on an audio adapter card;
Fig. 7 shows a typical screen interface presented to a user of the workstation of Fig. 2;
Fig. 8 is a schematic diagram showing the main software components running on the workstation of Fig. 2;
*Explanation of symbols for main parts of the drawing*
10 : system unit of computer, 12 : display screen,
14 : keyboard, 16 : mouse,
22 : microprocessor, 24, 52 : semiconductor memory,
26 : bus, 28 : audio card,
30 : token ring adapter card, 40 : microphone,
42 : A/D converter, 44 : CODEC,
46 : Digital Signal Processor (DSP), 48 : Double Buffer,
54 : D/A converter
[Detailed Description of the Invention]
The present invention relates to the processing of multiple streams of audio data received over a network with a computer workstation.
Traditionally, voice signals were transmitted over standard analog telephone lines. However, with the increasing number of areas equipped with local area networks (LANs) and the increasing importance of multimedia communication, considerable attention has been paid to the use of LANs to carry voice signals. This work has been edited by, for example, P Ravasio, G Hopkins, and N Naffah; Using Local Area Networks for Carrying Online Voice by D Cohen, North Holland, 1982, in Local Computer Networks, page Described in pages 13-21 and Voice Transmission over an Ethernet Backbone by P Ravasio, R Marcogliese, and R Novarese, pages 39-65 has been The basic principle of this scheme is that the first terminal or workstation digitally samples the voice input signal at a normal rate (eg, 8 kHz). Some samples are then assembled into data packets for transmission over the network to a second terminal, and then the samples are fed back at a constant rate to a loudspeaker or equivalent device for playout.
One of the problems with using a LAN to carry voice data is that the transmission time over the network is variable. Thus, the arrival of packets at the destination node is delayed and irregular. If the packet is played out in an irregular format, it greatly adversely affects the recognizability of the voice signal. Therefore, the voice transmitted over the LAN scheme uses some amount of buffering at the receiving end to absorb this irregularity. There is too much delay between the original voice signal and the audio output at the destination end, which makes a natural two-way conversation difficult (in the same way as is severely hampered by excess delay in a typical transatlantic telephone call). You just have to be careful to avoid it. IBM Technical Disclosure Bulletin, pages 255-257, Vol. 36, by B Aldred, R Bowater, and S Woodman, where a system in which packets arriving later than the maximum allowed are abandoned No. 4, April 1993, Adaptive Audio Playout Algorithm for Shared Packet Networks. The amount of buffering is adaptively controlled according to the number of dropped packets (other suitable delay measures may be used). When the number of abandoned packets is large, the degree of buffering is increased, and when the number of abandoned packets is small, the degree of buffering is decreased. The size of the buffer can be modified by temporarily changing the play-out speed (changes in the play-out speed affect the pitch, which are not very good techniques to detect silent periods and artificially increase them appropriately) or can be reduced).
Another important aspect of audio communication is conferencing, including multipoint communication, as opposed to two-way or point-to-point communication. When implemented over conventional analog telephone lines, audio conferencing requires each participant to transmit an audio signal to a central hub. The central hub adjusts the incoming signals to as different levels as possible, mixes them, and sends to each participant the sum of the signals from all other participants (except the signal from its own node). U.S. Patent 4,650,929 describes a centralized video/audio conferencing system in which individuals can adjust the relative volume of different participants.
The use of centralized mixing nodes, often referred to as multi-point control units (MCUs), has been taken to mean any multimedia (audio + video) workstation conferencing system. For example, US Pat. No. 4,710,917 describes a multimedia conferencing system in which each participant transmits audio to and receives audio from a central mixing unit. Another multimedia conferencing system is CSCW '90 (Proceedings of the Conference on Computer-Supported Cooperative Work, 1990, Los Angeles), pages 27-38, K Watab, S Sakata, K Manno ( K Maeno), H Fukuoka and T Ohmori My Decentralized Multi-Party Desktop Conferencing Systems: Distributed Multiparty Desktop Conferencing System (MERMAID) and Personal Multimedia Multipoint for Broadband Networks by E Addeo, A Gelman and A Dayao of IEEE GLOBECOM, pages 53-57 of 1988 It is described in Personal Multimedia Multipoint Communications Services for Broadband Networks.
However, the use of a centralized MCU or summation node has several problems. First, since most LAN's architects are based on a peer-to-peer arrangement, there is no clear central node. Moreover, the system is entirely dependent on the continuous availability of the designated central node to conduct the conference. Also, there are problems of ehp suppression (the central node must take care that the audio signal from one node is not included in the sum signal reproduced at that node).
These problems can be solved by using a distributed audio conferencing system in which each node receives one separate audio signal from each other node during the conference. US 5127001 describes such a distributed system and discusses synchronization problems caused by variable transition times of packets passing through the network. US5127001 overcomes this problem by maintaining a separate queue of incoming audio packets from each source node. They effectively absorb jitter at the time of arrival in the same way as described above for the simple point-to-point case. At regular intervals, a set of audio packets are read, one packet from each queue, and summed together for playout.
As found in the MERMAID system described above, one of the problems with audio conferencing systems is determining who is speaking at any one point in time. US 4893326 describes a multimedia conferencing system in which each workstation automatically detects whether its user is speaking. This information is then fed to a central control node, which switches the video so that each participant sees the current speaker on their screen. These systems require both video and audio manipulation capabilities and, furthermore, rely on a central video switching node, and therefore cannot be used in fully distributed systems.
Distributed multimedia conferencing systems are described by H Tanigawa, T Arikawa, S Masaki, and K. Shimamura, IEEE INFOCOM 91, pages 1127-1134 of Volume 3 of Proceedings. K Shimamura) in Personal Multimedia-Multipoint Teleconferencing System. This system provides sound localization ( sound localization). This method only serves a limited role in identifying the speaker. A more comprehensive function is described in Japanese abstract JP 02-123886 in which a bar graph is used to display the output speech level associated with an adjacent window containing video for a sound source.
As such, the prior art has proposed various audio conferencing systems. Although conventional centralized telephone audio conferencing is widespread and technically well mastered, much work remains to be done to increase the performance of audio conferencing implementations in desktop environments.
Accordingly, the present invention provides a computer workstation connected to a network and for receiving multiple audio input streams from the network, each audio input stream comprising a sequence of digital audio samples, the computer workstation comprising: Means for storing digital audio samples from each audio input stream in a separate queue, means for forming a sequence of sets each comprising one digital audio sample from each queue, each means for generating a weighted sum for each set of digital audio samples, the audio input stream having a weighting parameter associated therewith; means for generating an audio output from the sequence of weighted sums and means for adjusting the weighting variables to control a relative volume within the audio output of the plurality of audio streams in response to user input. workstation is provided.
By having audio conferencing over a distributed network where each node receives one individual audio stream from every other participant, additional functions that are difficult and expensive to achieve in a centralized conferencing system are naturally possible. In particular, each user can adjust the relative volume of all other participants according to their own personal preferences. This can be done, for example, when users need to focus on one particular situation in a conference, or when they have language problems (for example, one person has a strong accent that some people may not understand, or the conference is simultaneous interpreting). ) is highly desirable. Also, during a conference, the system responds to user input to change the relative volume of different participants. To allow this control, the incoming audio signals are kept separate and placed in different queues depending on their source before being weighted by an appropriate volume control element (where the cues are physically adjacent or combined). However, they are logically separate storage devices). The audio signals are then combined together to produce a final audio output. Accordingly, the present invention recognizes that a distributed audio conferencing system is particularly suitable for installations that individually control relative volume.
Preferably, the workstation further comprises means for visually indicating for each of said plurality of audio input streams whether the stream is currently silent. This solves one of the recognized problems of audio conferencing, the problem of determining who is speaking. The visible indication is simply in the form of some form of on/off indicator such as a light or equivalent feature, but in a preferred embodiment for each of said multiple audio input streams an indication of the instantaneous sound volume in that audio stream. implemented by the display. In other words, the display provides all indications of the volume of the participants involved. The volume output can be calculated based on the running rms value from a sequence of digital audio samples, or, if processing power is limited, a more simple one, such as using the maximum digital audio value of a predetermined number of samples. algorithm is used. In general, incoming audio data arrives in a block state, each containing a predetermined number of digital audio samples, and the visible indication is updated each time a new block of audio data arrives. In this way the volume diagram may typically be calculated on a block-by-block basis.
The visible indication is also preferably displayed adjacent to the visible representation of the origin of the audio input stream, such as a video or still image. The former requires a complete multimedia conferencing network, while the latter may be provided over a lower bandwidth network that cannot support the transmission of video signals. Any audio source can be easily identified by these visible signs, whether stationary or moving. Preferably, the workstation further comprises means for providing a user with a visual indication of the values of the weight variable, wherein the means is responsive to a user mouse operation to adjust the weight variable. This is implemented as a scroll-bar or the like, one for each audio input stream and is placed adjacent to a visible indication of the output volume of that stream. It is also convenient if the computer workstation further comprises means for disabling audio output from any one of said plurality of audio input streams. In this way, the user can effectively have a complete set of volume control for each audio input stream. The invention also provides a method of operating a computer workstation coupled to a network for receiving a plurality of audio input streams, each audio stream comprising a sequence of digital audio samples, said Storing audio samples in a separate queue, forming a sequence of sets each comprising one digital audio sample from each queue, each audio input stream having a weighting variable associated with it, generating a weighted sum for each set of audio samples, generating an audio output from the sequence of weighted sums, and responsive to user input, the weighted sum to control a relative volume within the audio output of the plurality of audio streams. There is provided a method of operating a computer workstation comprising adjusting the parameters.
Other features and advantages, including the above-described features and advantages of the present invention, will become apparent to those skilled in the art by the following detailed description set forth with reference to the accompanying drawings in which like elements are denoted by the same reference numerals. can
1 is a schematic diagram of computer workstations AE linked together by a local area network (LAN) 2 . These workstations are participating in a multipath conference, so each workstation is broadcasting its audio signal to all other workstations in the conference. In this way, each workstation receives a separate audio signal from all other individual workstations. The network shown in Figure 1 has a Token Ring architecture, in which tokens circulate along workstations. A workstation that currently owns the token is allowed to send messages to other workstations. It should be noted that the physical transmission time of a message along the ring is extremely short. In other words, for example, a message sent by A is received at all other terminals at about the same time. This is why the token system is used to prevent interference when two nodes attempt to transmit a message at the same time.
As will be explained in detail later, one-way communication of audio over a LAN typically requires a bandwidth of 64 kHz. In the conference of Figure 1, each node broadcasts its audio signal to the other 4 nodes, implying a total bandwidth requirement of 5 X 4 X 64 kHz (1.28 MHz). This is well within the capabilities of standard Token Ring, which supports transfer rates of 4 or 16 megabits per second. For larger conferences, you'll find that bandwidth demands are immediately problematic, although other networks are expected to offer higher bandwidth.
As long as technical requirements regarding bandwidth, latency, etc. required to support audio conferencing can be satisfied, the present invention can be implemented on a number of other network configurations or architectures other than token ring.
Figure 2 is a simplified schematic block diagram of a computer system that may be used in the network of Figure 1; The computer has a system unit (10), a display screen (12), a keyboard (14) and a mouse (16). The system unit 10 includes a microprocessor 22, a semiconductor memory (ROM/RAM) 24, and a bus 26 through which data is transferred. The computer of Figure 2 is any conventional workstation, such as an IBM PS/2 computer.
The computer of Figure 2 is equipped with two adapter cards. One of these is the token ring adapter card 30 . This card, along with accompanying software, allows messages to be sent to and received from the token ring network shown in FIG. Since the operation of the token ring card is well known, it will not be described in detail. The second card is an audio card 28 connected to each of a microphone and loudspeaker (not shown) for audio input and output. The audio card is shown in more detail in FIG. 3 . The card shown and used in this particular embodiment is an M-par card available from IBM, although other cards implementing analog functions may be used. The card includes an A/D converter 42 that digitizes incoming audio signals from an attached microphone 40 . An A/D converter is attached to the CODEC 44, which samples the incoming audio signal into 16-bit samples at a rate of 44.1 kHz (corresponding to the standard sampling rate/size for compact discs). The digitized samples are then passed through the dual buffer 48 to a digital signal processor (DSP) 46 on the card (i.e., the CODEC loads the samples into one half of the double buffer, while the CODEC simultaneously loads the samples from the other half). read the previous sample). The DSP is controlled by one or more programs stored in a semiconductor memory 52 on the card. Data can be transferred to and from the main PC bus by the DSP.
The audio signals to be played out are received by the DSP 46 from the PC bus 26 and are processed in the opposite manner as audio input from the microphone. That is, the output audio signals are transmitted to the CODEC 44 through the DSP 46 and the double buffer 50, and from the CODEC 44 to the D/A converter 54, and finally the loudspeaker 56 or other passed to the appropriate output device.
In the specific embodiment shown, the DSP samples samples from the CODEC from 16 bits at 44.1 kHz using standard resampling techniques, corresponding to the CCITT standard G.711, on a μ-aw scale (basically logarithmic). is programmed to convert it to a new digital signal with an 8 kHz sampling rate, which is an 8-bit sample of Therefore, the total bandwidth of the signal delivered to the workstation for transmission to another terminal is 64 kHz. The DSP also performs an inverse transform on the incoming signal received from the PC. That is, the signal is converted from an 8-bit 8 kHz to a 16-bit 44.1 kHz using a known resampling technique again. This conversion between the two sampling formats is only necessary because of a particular hardware choice and is not directly relevant to the present invention. Thus, for example, many other audio cards also include native support for the 8kHz format. In other words, CODEC uses the 8kHz format by operating in accordance with G.711 (unless the transmitted audio signal needs to have CD quality specifically, it is unlikely due to the requirement of higher bandwidth and greatly increased processing speed). However, optionally, 44.1 kHz samples may be used for transmission over the network; for normal voice communication, a 64 kHz bandwidth signal in G.711 format is suitable).
Data is transferred between the audio adapter card and the workstation in blocks of 64 bytes, that is, 8 ms of audio data for 8-bit data sampled at 8 kHz. The workstation then processes the entire block of data, and each data packet received by or sent to the workstation typically contains a single 64-byte block of data. Choosing a block size of 64 bytes is a compromise to minimize the granularity of the system (causing delay) while maintaining efficiency with respect to both internal processing at the workstation and transmission over the network. On other systems, a block size of, for example, 32 or 128 bytes is more suitable.
Since the operation of the computer workstation related to the transmission of audio data is well known in the prior art, a detailed description thereof will be omitted. Basically, an audio card receives an input audio signal, whether in analog form from a microphone or some other audio source such as a compact disc player, and generates a digital audio data block. These blocks are then transferred to the workstation's main memory, and from the main memory to the LAN adapter card (on some architectures, the blocks are transferred directly from the Audio Adapter Guard to the LAN adapter card without having to go through the workstation memory. can be transmitted). The LAN adapter card generates data packets identifying the source and destination nodes, including digital audio data along with header information, which are then forwarded over the network to the required recipient(s). In any two-way or multi-path communication, such transmission processing may be performed at the workstation simultaneously with reception processing described later.
The processing by the computer workstation regarding the reception of audio data packets is shown in FIG. Each time a new packet arrives (step 402), the LAN adapter card informs a program running on the workstation's microprocessor and provides information to the program identifying the source of the data packet. The program then transfers the incoming 64-byte audio block to a queue in main memory (step 404). As shown in Figure 5, queue 500 in main memory actually contains a set of respective sequences comprising audio blocks taken from each of the different source nodes. Thus, one queue contains audio blocks from one source node, another queue contains audio blocks from another source node, and so on. In Figure 5, there are three sequences 501, 502, 503 for audio data from each of Nodes B, C and D. The number of sequences will of course vary depending on the number of participants in the audio conference. The program uses the information identifying the source node within each packet received to assign the block of incoming audio data to the correct queue. Pointers PB, PC and PD indicate the end positions of the queues, which are updated each time a new packet is added. Packets are removed from the beginning of the sequence (as shown by "OUT" in Fig. 5) to continue processing. Therefore, the sequence of Figure 5 is essentially a standard first-in-first-out queue and can be implemented using conventional programming techniques. Apart from the support of multiple (parallel) queues, the processing of incoming audio blocks described heretofore is almost similar to the conventional method, allowing equivalent buffering techniques to be used, if necessary, in relation to individual sequences or combined queues within entries. .
The operation performed by the DSP on the audio adapter card is shown in FIG. The DSP cycles through processing a new set of audio blocks every 8 microseconds to ensure a continuous audio output signal. Therefore, every 8 microseconds the DSP uses a DMA access to read a block of audio from each sequence corresponding to a different node - i.e., one block from the leading edge of queues B, C and D shown in Figure 5. (Step 602: ie, in this case, M=3). These blocks are treated as representing simultaneous time intervals: in the final output, they are added together to produce a single audio output for that time interval. Therefore, the DSP effectively implements the digital mixing function on multiple audio input streams. Then, using a look-up table, each sample in the 64 byte block is converted from the (basically log in) G. 711 format to a linear scale (step 604). Then, each individual sample is multiplied by a weight variable. (Step 606). There is a respective weight variable for each received audio data stream; That is, in the case of the three sequences in FIG. 5, there is one weight variable for the audio stream from node B, one for the audio stream from node C, and one for the audio stream from node D. The weight variable is used to control the relative magnitude of audio signals from different sources.
The DSP maintains a running record of the effective (rms) values for each audio stream (step 608). Typically, this rms value is generated for each block of audio data (ie every 8 microseconds) by generating the root of the sum of the squares of the values in the block. The rms value represents that individual audio input stream and is used to provide volume information to the user as described below.
Once the digital audio samples are multiplied by an appropriate weighting variable, they are summed together (step 608: this can effectively be done in parallel with the processing of step 606). In this way, a single sequence of digital audio samples is created, representing the weighted sum of multiple input audio streams. This sequence of digital audio samples is then returned to 44.1 KHz (as described above, although it is hardware dependent and not directly relevant to the present invention) before being passed to the CODED (step 612) for supply to the loudspeaker. is sampled (step 610).
The actual DSP processing used to generate the volume adjusted signals is somewhat different from that in Figure 6, although the end result is effectively similar. These differences are typically introduced to maximize computational efficiency or reduce the demand on the DSP. For example, if processing power is limited, volume control can be implemented as a conversion from the μ-law format. Thus, after the correct no-look value has been placed (step 604), the actual reading can be determined by moving a predetermined number of places up and down the table, depending on whether the volume of this signal is increased or decreased from its normal value. . The weight variable in this case is the effective number of steps up or down to adjust the look-up table (either the F.711 format separates them depending on whether the original amplitudes are positive or negative, or by adjusting the volume to change one format to another? to explicitly take into account the fact that it cannot be converted to a format). The method is computationally simple, but provides only discrete control, not continuous volume control. Alternatively, it is possible to add the logarithm of a volume control value or weight variable to a μ-law plot. This method effectively performs the multiplication of step 606 before the scale transformation of step 604 using log addition, which is less computationally expensive than multiplication for most processors. The result may then be converted back to a linear scale for mixing with other audio streams (step 604). This method allows fine volume control (although the output is limited to 16 bits), providing a sufficiently detailed look-up table. Typically, the logarithm of the weight variable may be obtained from a look-up table, or may optionally be provided in advance in log form by the control application. Of course, it is necessary to compute a new logarithmic value when the volume control is adjusted, but this is relatively uncommon.
Similarly, if the available processing power was not sufficient to perform successive rms volume measurements, the process may be performed on each and every other block data, or optionally by summing the absolute values of the differences between successive samples. Similarly, a computationally rather simple algorithm can be used. The summation of squared values may be performed by logarithmic addition before step 604 (ie, before scale transformation). A simple way to do this is to simply use the maximum sample value of any audio block as a volume indicator.
7 shows a screen 700 presented to a user of a workstation involved in an audio conference. As described above, this involves the reception of three different streams of audio data, it is clear that the present invention is not limited to only three participants. The screen of FIG. 7 is divided into three areas by dotted lines, each representing one participant, and these dotted lines do not actually appear on the screen. Associated with each participant is a box 724 containing the participant's name (in this case just B, C and D). Also, a video image of an audio source delivered over an audio network, or (audio at the start of the conference) There is a video window 720 that can be used to contain a still bitmap (either supplied by the source or already local to the workstation and displayed in response to the participant's name). In the case of Participant D, there is no video or still image available, so it is shown as an empty window. The choice of display (blank, still, or video image) of the video window depends on the available hardware of the wargstation, the bandwidth of the network, and the availability of relevant information.
Below the video window is a volume display 721 (VU meter) indicating the instantaneous volume of the audio stream (which is calculated at block 608 of Figure 6). The length of the solid line in this display indicates the volume of the audio stream. If there is no audio signal from that participant, the solid line has a length of zero (ie invisible). Thus, a user can determine who is speaking during a conference by observing whose VU meter is working.
Below the volume display is a volume control bar 722 that allows the user to adjust the relative volume of the attendees. Adjustments can be made by the user pressing the + or - buttons respectively at one end of the bar to increase or decrease the volume. This has the effect of correspondingly increasing, decreasing, or decreasing the weight variable used for digital blending. The indicator in the middle of the volume control bar indicates the current volume setting (ie the current value of the weight variable).
Finally, next to the name box 724 is a silence button 723 . Pressing this button toggles between disabling and enabling audio output from that participant. When audio output is disabled, the weight variable is set to zero, and when enabled, the weight variable returns to its previous (ie, displayed on the volume control bar) value. If audio from a participant is currently disabled, this is represented by a cross overlaid on the mute button (all three audio outputs in FIG. 7 are currently disabled). Using the DSP processing described above, it is possible to directly transform the system so that when the mute button is in the disabled audio output state, the VU meter shows the signal level that would be generated if the audio output was actually enabled.
Fig. 8 shows the main software elements running on the workstation of Fig. 2 to provide the user interface of Fig. 7;
The workstation may be controlled by an operating system 814 such as Windows available from Microsoft Corporation. In addition, suitable communication software 816 to enable LAN communication is present on the Wookstation (in some cases, communication software is provided via device driver 818, in a manner known in the art, to the two adapter cards). , token ring and audio adapter card The overall processing of audio is controlled by application 810. This application functions as an Application Support Layer (812) available from Microsoft Corporation. , one of which implements the application support layer 812 is Visual Basic. The purpose of the application support layer is to facilitate the development of applications, especially in relation to user interfaces, but it is of course possible instead for applications to cooperate directly with the operating system directly.
The application controls the contents of the window box 720 in conjunction with known programming techniques. For example, a VU meter 721 is provided using the functionality provided in Visual Basic, which effectively responds to all graphs associated with the UV meter: all the application has to do is supply the relevant numerical values. Since Visual Basic is interrupt-driver, this is easily accomplished by DSP copying the output volume for the audio block into the workstation and then requesting an interrupt. The interrupt generates an event on the application and signals it as a new output volume that can now be copied onto the VU meter. In practice, an interrupt is used to indicate the availability of a complete set of volumes for a set of audio blocks, i.e. one volume for each audio stream (DSP is also used to indicate the output audio generated by that workstation for transmission over the network). It should be noted that with respect to the signal, we perform an interrupt for each audio block). Likewise, volume control bar 722 is also a feature provided in Visual Basic (called a scroll bar). Visual Basic contains the location of the selector, is responsible for all graphics related to the control bar, and then the application simply passes this updated volume diagram to the application. The application then writes this updated volume diagram to the DSP and changes the volume accordingly. Mute button 723 is another display feature given by Visual Basic, allowing simple on/off control of each audio stream. Whenever the silence button is activated, it is necessary for the application to remember the previous value of the weight parameter, so that the application can be reverted the next time the silence button is pressed.
It should be noted that many modifications to the user interface are possible. For example, a VU meter may be segmented or replaced with an analog level meter. A much simpler way is to have an on/off indicator that changes color depending on whether or not there was any audio output from that participant. The volume control function may be implemented using a dial or may be implemented as a drag and drop slider rather than two push buttons. A mute button may also be combined with a volume control bar. Such changes are within the programming capabilities of those skilled in the art.
In the system described above, the user is effectively limited to controlling the volume of each audio input stream, but in other systems the user can have more advanced controls such as frequency control (i.e., treble and base adjustment). . This can be implemented relatively easily by DSPs that serve audio signals in the time domain by means of FIR or IIR filters. The frequency control can make appropriate changes to the FIR/IIR filter coefficients to change in FIG. 7 . These advanced controls will increase in demand as they improve the quality of audio signals carried over the network, for example in systems using F.721 rather than the G.711 audio transport standard.
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
13 members in 8 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 9325924 | United Kingdom | A | |
| 9325924 | United Kingdom | A | |
| 93259240 | United Kingdom | – | |
| 93259240 | – | – | – |
| GB19930025924 | – | – | – |
Members13
| Document | Office | Kind | |
|---|---|---|---|
| GB9325924D0 | United Kingdom | D0 | |
| EP0659006A2 | European Patent Office (EPO) | A2 | |
| GB2284968A | United Kingdom | A | |
| KR950022401A | Republic of Korea | A | |
| JPH07200424A | Japan | A | |
| CN1111775A | China | A | |
| US5539741A | United States of America | A | |
| JP2537024B2 | Japan | B2 | |
| KR0133416B1This record | Republic of Korea | B1 | |
| EP0659006A3 | European Patent Office (EPO) | A3 | |
| TW366633B | Taiwan Province of China | B | |
| CN1097231C | China | C | |
| IN190028B | India | B |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapse due to unpaid annual feeLapsedLAPS | LAPS | |
| Annual fee paymentFPAY | FPAY | |
| Written decision to grantGRNT | GRNT | |
| Decision to grant or registration of patent rightE701 | E701 | |
| Request for examinationA201 | A201 |
Numbers
- Publication
- 1001334160000
- Publication, DOCDB
- 0133416
- Publication, EPODOC
- KR0133416B
- Application
- 100034643
- Application, DOCDB
- 19940034643
- Application, EPODOC
- KR19940034643
Titles2
- Korean
- 오디오 컨퍼런싱 시스템
- English
- audio conferencing system
Classification
- CPC, 9
- H04L12/1813
- H04M3/42161
- H04M3/567
- H04M3/568
- H04N7/15
- H04L65/1083
- H04L65/403
- H04L65/1059
- H04L65/613
- IPC, 5
- G06F15 00
- G06F13 00
- H04L12 18
- H04M3 56
- H04N7 15