Multi party teleconference methods and systems
4 claims: 1 independent, 3 dependent
- 1A method of controlling multi-party téléconférence calls in a télécommunications network, said method comprising:5 storing data représentative of one or more predetermined words or phrases;establishing a multi-party téléconférence call between at least three telephony devices in said network;receiving an audio stream from each of said at least three telephony 10 devices in said network;processing said received audio data streams such that the received audio data streams are combined into mixed audio data streams using one of a plurality of available mixing modes;monitoring said received audio data streams for the utterance of at least 15 one of said one or more predetermined words or phrases represented in said stored data;and altering a mixing mode applied during the processing of said received audio data streams in response to said monitoring detecting the utterance of said one or more predetermined words or phrases represented in said stored data, 20 wherein said altering comprises altering from a first mixing mode in which volume adjustment levels are applied to the received data streams in first respective levels to a second mixing mode in which the volume adjustment levels are applied in second, different, respective levels, and wherein said altering comprises cutting at least a first participant from a 25 mixed audio stream in response to detecting an utterance of a word or phrase indicating that said first participant wishes to participate in a private téléphoné call with an identified other participant.
113 paragraphs in 2 sections, as filed
Multi-party téléconférence methods and Systems
Field of the Invention
The présent invention relates to multi-party téléconférence methods and Systems.
Background of the Invention
Multi-party audio téléconférences today are a vital business tool for allowing people to “meet” without being in the same location. A multi-party audio téléconférence typically comprises a number (greater than two) of endpoint telephony devices, either wire line or wireless, and an audio téléconférence bridge. The endpoint telephony devices dial into the multi-party audio téléconférence bridge, enter an access code and can then talk to other people in the multi-party audio téléconférence call. The way this tends to work is that the multi-party audio téléconférence bridge detects several of the loudest (or “active”) speakers, mixes the audio from these speakers and sends out the mixed audio to ail participants but subtracts the audio that each individual is sending to the audio téléconférence bridge to avoid écho or feedback.
With the wide deployment of fast, reliable Internet services, many people are now choosing to work at home. It makes sense for them in terms of work/life balance and can also boost productivity. One major downside of working at home, however, is that people can become quite detached from the social aspects of working in an office. In an office, if someone passes another person’s desk and stops for a chat it is quite natural, but calling that person on the téléphoné for the same type of chat is not. Likewise, in an office, it is normal to overhear the conversations of others and join in if desired, but such an opportunity does not arise for people working from home. It is in these kinds of environments, and others, in which improved multi-party téléconférence methods could be employed to great effect.
It would therefore be désirable to provide improved multi-party téléconférence methods and Systems.
<img file="GB2492103B_D0001.tif" />
Summary of the Invention
In accordance with a first aspect of the présent invention, there is provided a method of controlling multi-party téléconférence calls in a télécommunications network, said method comprising:
storing data représentative of one or more predetermined words or phrases;
establishing a multi-party téléconférence call between at least three telephony devices in said network;
receiving an audio stream from each of said at least three telephony devices in said network;
processing said received audio data streams such that the received audio data streams are combined into mixed audio data streams using one of a plurality of available mixing modes;
monitoring said received audio data streams for the utterance of at least one of said one or more predetermined words or phrases represented in said stored data; and altering a mixing mode applied during the processing of said received audio data streams in response to said monitoring detecting the utterance of said one or more predetermined words represented in said stored data, wherein said altering comprises altering from a first mixing mode in which volume adjustment levels are applied to the received data streams in first respective levels to a second mixing mode in which the volume adjustment levels are applied in second, different, respective levels, and wherein said altering comprises cutting at least a first participant from a mixed audio stream in response to detecting an utterance of a word or phrase indicating that said first participant wishes to participate in a private téléphoné call with an identified other participant.
Hence, a speech récognition function can be used to allow a participant to alter the mixing mode applied, for example by the utterance of a word or phrase.
In accordance with a second aspect of the invention, there is provided apparatus adapted to perform the method of the first aspect of the présent invention.
In accordance with a third aspect of the invention, there is provided 5 computer software adapted to perform the method of the first aspect of the présent invention.
Further features and advantages of the invention will become apparent from the following description of preferred embodiments of the invention, given by way of example only, which is made with reference to the accompanying drawings.
<img file="GB2492103B_D0002.tif" />
Brief Description of the Drawings
Figure 1 shows a system diagram according to preferred embodiments. Figure 2 shows a block diagram according to preferred embodiments.
Figure 3 shows a flow diagram according to preferred embodiments.
Figure 4 shows a graph according to preferred embodiments.
Figure 5 shows a graph according to preferred embodiments.
Figure 6 shows a graph according to preferred embodiments.
Detailed Description of the Invention
Figure 1 shows a system diagram according to some embodiments. Figure 1 depicts a télécommunications network 110 in which audio teleconferencing server (ATS) 100 is responsible for providing multi-party audio téléconférence services to telephony devices TD1, TD2, TD3, TD4, TD5 and TD6. ATS 100 may also be responsible for providing audio téléconférence services to other telephony devices (not shown).
Each of TD1, TD2, TD3, TD4, TD5 and TD6 is associated with a user who works for or is otherwise associated with a given group or organisation. Some of telephony devices TD1, TD2, TD3, TD4, TD5 and TD6 may be located in one office of the organisation, some may be located in one or more other offices of the organisation and some may be located at home locations where the respective users résidé. A user may be associated with more than one telephony device, for example a user may hâve one telephony device located on their desk at work and another telephony device located at their home.
<img file="GB2492103B_D0003.tif" />
O
LO
Telephony device TD1 is provided with telephony services by téléphoné switch 104 which is connected to ATS 100 via network 102. Network 102 may comprise one or more Public Switch Téléphoné Networks (PSTNs) and/or the Internet.
A user of telephony device TD1 also has access to a computing device 114 with display and user input capabilities, for example a personal computer, laptop, or personal digital assistant, etc. which is located locally to telephony device TD1.
Téléphoné switch 104 is responsible for providing switching of téléphoné calls for a number of telephony devices such as telephony device TD1 including provision of dial tone, ringing tone, etc. Téléphoné switch 104 may also include the ability to select processes that can be applied to such calls, routeing for such calls based on signalling and subscriber database information, the ability to transfer control of calls to other network éléments and management functions such as provisioning, fault détection and billing. Téléphoné switch 104 may also be referred to as a local téléphoné exchange, central office, class 5 switch or softswitch. Téléphoné switch 104 includes a database 108 for storing call related data, including call state information, for example in relation to téléphoné calls incoming to or outgoing from telephony device TD1.
Telephony devices TD2, TD3, TD4, TD5 and TD6 may similarly be provided with telephony services by one or more other téléphoné switches (not shown). In this example, telephony device TD7 is provided with telephony services by the same téléphoné switch 104 as telephony device TD1, but such services could be provided by a different téléphoné switch.
Figure 2 shows a block diagram according to embodiments of the présent invention. Figure 2 shows some components of ATS 100 of Figure 1. ATS 100 comprises a processor 200 (or processors) for carrying out data processing functionality. ATS 100 includes a data store 206 for storing data représentative of predetermined words or phrases, and a data store 208 for storing user configuration data. ATS 100 has an audio monitoring module 204 with speech récognition capabilities for monitoring audio data streams received from telephony devices TD1 to TD6. ATS 100 has a web interface 210 for providing user access to various multi-party audio téléconférence fünctionalities such as display and user configuration. ATS 100 has a mixing module 202 for processing audio data streams received from telephony devices TD1 to TD6 including combining the received audio data streams into one or more mixed audio data streams according to different audio mixing modes. Mixing module 202 includes codée functionality for decoding audio data streams received from telephony devices in different data formats and recoding mixed audio data streams for transmittal out to the telephony devices in the appropriate data formats.
Some embodiments relate to handling an incoming téléphoné call of a first call type during a multi-party audio téléconférence call which is of a second call type, different from the first call type. The second call type may for example be a call that can be interrupted. The first call type may for example be a call that can not be interrupted.
ATS 100 pro vides a multi-party audio téléconférence service and a multi-party audio téléconférence call is currently being conducted between TD2, TD3, TD4, TD5 and TD6.
The user of TD1 wishes to join the multi-party audio téléconférence so dials an appropriate number associated with ATS 100 which results in an incoming multi-party audio téléconférence call setup request being received at téléphoné switch 104. Téléphoné switch 104 establishes a first call leg for transmittal of audio data associated with the multi-party audio téléconférence call between TD1 and ATS 100. ATS 100 connects TD1 to the audio téléconférence call being conducted between TD2 to TD6. Téléphoné switch 104 stores, in state information store 108, state information indicating that the call between TD1 and TD2 to TD6 is of the second call type. In this embodiment, the second call type is an interruptible multi-party audio téléconférence call type. This indicates that it can be interrupted by an incoming call to TD1. Alternatively, it may be indicative of a more general call type, for example an interruptible call type.
During the multi-party audio téléconférence call between TD1 and TD2 to TD6, TD7 dials a téléphoné dialling number associated with TD1 which results in an incoming call setup request from TD7 being received at téléphoné switch 104. TD7 is not a participant in the multi-party audio téléconférence call currently being conducted between TD1 and TD2 to TD6. Téléphoné switch 104 inspects the stored state information for the multi-party audio téléconférence call and réalisés that the multi-party audio téléconférence is of the second call type so can be interrupted with respect to TD1 by the incoming call from TD7. Téléphoné switch 104 therefore proceeds to handle the incoming téléphoné call setup request and the multi-party audio téléconférence call in accordance with the stored state information indicating that the multiparty audio téléconférence call is of the second call type.
Téléphoné switch 104 directs the incoming téléphoné call setup request to TD1 which results in TD1 ringing, i.e. TD1 enters a ringing state in relation to the incoming téléphoné call from TD7.
Téléphoné switch 104 disables the first call leg between TD1 and ATS 100. The audio téléconférence call between TD2 to TD6 carries on as normal, but without TD1 participating.
If the user of TD1 chooses to answer the incoming call, téléphoné switch 104 establishes a second call leg for transmittal of audio data associated with the answered téléphoné call between TD1 and TD7. The users of TD1 and TD7 are thus able to conduct a téléphoné call to each other. During this téléphoné call, ATS 100 continues to provide a multi-party audio téléconférence service to TD2 to TD 6 and the audio téléconférence call between TD2 to TD6 continues as normal.
When the user of TD1 or the user of TD7 terminâtes the téléphoné call between them, for example by hanging up their respective telephony device, téléphoné switch 104 receives signalling information indicating such and tears down the second call leg between TD1 and TD7. Téléphoné switch 104 establishes a third call leg for transmittal of audio data associated with the multiparty audio téléconférence call between TD1 and ATS 100. TD1 thus re-enters the multi-party audio téléconférence call being conducted between TD2 to TD6, i.e. the audio téléconférence call then has ail of TD1 to TD7 participating again.
In some embodiments, TD1 is capable of operating either in a handset mode or a speakerphone mode. Speakerphone mode is a mode of operation where a loudspeaker on TD1 outputs audio from TD1 without the user having to pick up the handset and put it to their ear/mouth. Similarly, in speakerphone mode, an extemal (i.e. extemal to the handset) microphone on TD1 picks up audio generated by the user (possibly also audio generated by others in close proximity and other background noise), without the user having to pick up the handset and put it to their ear/mouth.
When the user of TD1 initially joins the multi-party audio téléconférence, the user dials an appropriate number associated with ATS 100 without picking up the handset of TD1. TD1 thus enters the multi-party audio téléconférence in speakerphone mode, i.e. during the multi-party audio téléconférence call, TD1 opérâtes in speakerphone mode. The appropriate number associated with ATS 100 may for example be stored in a speed-dial function on TD1 such that the user of TD1 just needs to press a single button to join the multi-party audio téléconférence initially.
When téléphoné switch 104 handles the incoming call setup request received from TD7, téléphoné switch 104 instmcts TD1 to operate in handset mode. This means that TD1 will cease to pick up audio from its extemal microphone. Further, no audio data from the disabled first call leg with ATS 100 will be output by the loudspeaker of TD1. However, when téléphoné switch 104 directs the incoming téléphoné call setup request to TD1 and it enters a ringing state, TD1 is in handset mode and will émit a ringing sound from its loudspeaker.
When the user of TD1 picks up the handset of TD1 to conduct the téléphoné call with the user of TD7, TD1 will be in handset mode and continue to be in handset mode when the call with TD7 is terminated and the user of TD1 puis the handset of TD1 back on-hook. However, when TD1 re-enters the multi-party audio téléconférence call with TD2 to TD6, TD1 should be in speakerphone mode. Therefore, when téléphoné switch 104 establishes the third call leg between TD1 and ATS 100, téléphoné switch 104 instructs TD1 to operate in speakerphone mode.
Further, to avoid the user of TD1 having to perform any further operations to re-enter the multi-party audio téléconférence, other than putting the handset of TD1 back on-hook, téléphoné switch 104 applies modified signalling to the third call leg such that the third call leg will be automatically answered by TD1.
When the user of TD1 initially dials the number associated with ATS 100, this results in an incoming multi-party audio téléconférence call setup request call being received at téléphoné switch 104. In embodiments, the incoming multi-party audio téléconférence call setup request comprises an identifier associated with the multi-party audio téléconférence call being conducted between TD2 and TD7 which allows téléphoné switch 104 to recognise which multi-party audio téléconférence call the incoming multi-party audio téléconférence call setup request call relates to, i.e. one provided by ATS 100. Téléphoné switch 104 therefore forwards the incoming multi-party audio téléconférence call setup request on to ATS 100 and establishes the first call between TD1 and ATS 100 on the basis of the received identifier.
In some embodiments, the received identifier comprises a téléphoné dialling number associated with TD1, for example detected ffom a calling line identifier (CLI) field of the incoming multi-party audio téléconférence call setup request.
In another embodiment, the received identifier comprises a téléphoné dialling number reserved for multi-party audio téléconférence calls conducted via ATS 100, with téléphoné switch 104 being preconfigured to recognise multiparty audio téléconférence call setup requests to the reserved dialling number and forward them to ATS 100 accordingly.
In some embodiments, téléphoné switch 104 disables the first call leg between TD1 and ATS 100 by tearing down the first call leg between TD1 and ATS 100.
In other embodiments, téléphoné switch 104 disables the first call leg between TD1 and ATS 100 by maintaining the network resources reserved for establishing the first call leg, but not transmitting any audio data from TD1 to ATS 100 for the duration of the incoming call from TD7. In such embodiments, after the incoming call from TD7 is terminated, téléphoné switch 104 establishes the third call leg by utilising the maintained network resources. This avoids having to establish the third call leg from scratch and can help speed up re-entry back into the multi-party audio téléconférence call.
In some embodiments, if the user of TD1 chooses not to take the incoming call from TD7, or if the user of TD1 is unable to take the incoming call from TD7 such that TD1 remains in a ringing state for a predetermined time period, téléphoné switch 104 will abort the handling of the incoming call and reenable the first call leg between TD1 and ATS 100.
Figure 3 shows a flow diagram according to some embodiments. Figure 3 depicts the flow of signalling messages and audio data between various entities of Figure 1 when a multi-party audio téléconférence call is interrupted by an incoming call from a telephony device which is not participating in the multi-party audio téléconférence call.
A multi-party audio téléconférence call is currently being conducted between TD1, TD2, and TD3. Other téléphoné devices (not shown) may also participate in the multi-party audio téléconférence call.
A call leg for transmittal of audio data associated with the multi-party audio téléconférence call has been established between TD2 and ATS 100 as shown by step 3a. This call leg may be established via a téléphoné switch (not shown) which provides telephony services to TD2.
A call leg for transmittal of audio data associated with the multi-party audio téléconférence call has also been established between TD3 and ATS 100 as shown by step 3b. This call leg may be established via a téléphoné switch (not shown) which provides telephony services to TD3.
Téléphoné switch 104 establishes a call leg for transmittal of audio data associated with the multi-party audio téléconférence call between TD1 and ATS
100 as shown by steps 3 c and 3d. Téléphoné switch 104 stores state information indicating that the multi-party audio téléconférence call between TD1 to TD3 is of the second call type and is thus a multi-party audio téléconférence which can be interrupted by an incoming call to TD1.
During the multi-party audio téléconférence call between TD1, TD2 and TD3, the user of TD7 dials a téléphoné dialling number associated with TD1 which results in an incoming call setup request from TD7 being received at téléphoné switch 104 as per step 3e. TD7 is not a participant in the multi-party audio téléconférence call currently being conducted between TD1 to TD3. Téléphoné switch 104 inspects stored state information for the multi-party audio téléconférence call and recognises that the multi-party audio téléconférence is of the second call type and so can be interrupted with respect to TD1 by the incoming call from TD7. Téléphoné switch 104 therefore proceeds to handle the incoming téléphoné call setup request and the multi-party audio téléconférence call in accordance with the stored state information indicating that the multi-party audio téléconférence call is of the second call type.
Téléphoné switch 104 directs the incoming téléphoné call setup request to TD1 in step 3g which results in TD1 ringing, i.e. TD1 enters a ringing state in relation to the incoming téléphoné call from TD7. Step 3g also involves téléphoné switch 104 instructing TD1 to operate in handset mode.
In step 3f, téléphoné switch 104 disables the first call leg between TD1 and ATS 100. The audio téléconférence call between TD2 and TD3 carries on as normal, but without TD1 participating.
In step 3h, the user of TD1 chooses to answer the incoming call from TD7, such action being notified to téléphoné switch 104 in step 3i. Téléphoné switch 104 establishes a second call leg for transmittal of audio data associated with the answered téléphoné call between TD1 and TD7 in steps 3j and 3k. The users of TD1 and TD7 are thus able to conduct a téléphoné call to each other. During this téléphoné call, ATS 100 continues to provide a multi-party audio téléconférence service to TD2 and TD3 and the audio téléconférence call between TD2 and TD3 carries on as normal.
In step 31, the user of TD7 terminâtes the téléphoné call between TD1 and TD7, such action being notifïed to téléphoné switch 104 in step 3m. Téléphoné switch 104 tears down the second call leg between TD1 and TD7.
Téléphoné 104 informs ATS 100 that TD1 is re-entering the multi-party audio téléconférence call being conducted between TD2 and TD3 in step 3n.
In step 3o, téléphoné switch 104 instructs TD1 to operate in speakerphone mode. Téléphoné switch 104 establishes a third call leg for transmittal of audio data associated with the multi-party audio téléconférence call between TD1 and ATS 100 in steps 3p and 3q. During establishment of the third call leg, téléphoné switch 104 applies modified signalling to the third call leg such that the third call leg will be automatically answered by TD1, for example in conjunction with step 3o.
TD1 thus re-enters the multi-party audio téléconférence call being conducted between TD2 and TD3, i.e. the audio téléconférence call then has ail three of TD1, TD2 and TD3 participating again.
Some embodiments relate to controlling a multi-party audio téléconférence call in a télécommunications network. Such control is carried out by ATS 100 in relation to a multi-party audio téléconférence call being conducted between at least three telephony devices in the network, for example telephony devices TD1, TD2, TD3, TD4, TD5 and TD6.
Various mixing modes are illustrated in Figures 4 to 6. During the multi-party audio téléconférence, ATS 100 receives an audio data stream from each of telephony devices TD1, TD2, TD3, TD4, TD5 and TD6. Processor 200 of ATS 100 processes the received audio data streams and mixing module 202 of ATS 100 combines the received audio data streams into mixed audio data streams. Different mixed audio streams are then transmitted to the telephony devices TD1, TD2, TD3, TD4, TD5 and TD6 participating in the multi-party audio téléconférence.
To avoid feedback or echoing, a mixed audio data stream transmitted to a given telephony device during the multi-party audio téléconférence will not contain audio data received from that telephony device. For example, a mixed audio stream transmitted to TD6 will contain a mix of the audio data streams received from TD1 to TD5, but no audio data fforn the audio data stream received from TD6.
In each of the audio mixing modes shown in Figures 4 to 6, a mixed audio data stream is generated from a plurality of received audio data streams using an audio mixing mode. An audio mixing mode will generate a mixed audio data stream by combining the received audio data streams in the plurality according to respective volume adjustment levels defined for that audio mixing mode.
In accordance with embodiments, the mixing mode is normally set to a default audio mixing mode, in which each party provides a substantially equal contribution, and the volume adjustment levels for each incoming audio data stream included in any particular mixed audio data stream are equal. On receipt of an appropriate trigger, the ATS 100 switches to a non-default audio mixing mode, in which the volume adjustment levels for the respective streams are altered. On receipt of a further appropriate trigger, or timeout, the ATS 100 switches back to the default audio mixing mode.
During at least a part of the multi-party audio téléconférence call, a first mixed audio data stream is generated using a first audio mixing mode, for example the default audio mixing mode, and is transmitted to TD6. The first mixed audio data stream is generated from a plurality of audio data streams in first respective volume adjustment levels, the plurality of data streams here being at least some of the audio data streams received from TD1 to TD5. During at least a different part of the multi-party audio téléconférence call, a second mixed audio data stream is generated using a second audio mixing mode and is transmitted to TD6. The second mixed audio data stream is generated fforn a plurality of audio data streams in second respective volume adjustment levels, the plurality of data streams here being at least some of the audio data streams received from TD1 to TD5.
Preferably, in each of the audio mixing modes shown in Figures 4 to 6, the ATS 100 détermines mean input volumes for each data stream. The mean input volumes may be determined according to various known techniques, for example the mean input volume may be determined as an RMS (Root Mean Square) value of the incoming signal, to represent the power of the incoming signal. The mean input volume of each of the received audio data streams is preferably measured over a periodic or sliding window of between 0.5 to 5 seconds, preferably between 1 and 2 seconds.
Preferably, in each of the audio mixing modes shown in Figures 4 to 6, the ATS 100 détermines volume adjustment levels for each data stream. These volume adjustments may be applied as a pre-amplification levels, applied in a pre-amplifier, before mixing of each of the audio data streams in equal proportions by an audio mixer, or alternatively may be applied as mixing levels, applied in differing proportions, in an audio mixer. The mixing levels may represent fixed, linear, amplification levels or variable, non-linear, amplification levels.
Preferably, the ATS 100 détermines a normalised total output volume for each mixed data stream, and normalises the volume adjustment levels accordingly. This is shown in each of Figures 4 to 6, where the horizontal axis dénotés time and the vertical axis dénotés normalised total output volume of the mixed media data streams. The normalised total output volume of a mixed media stream may be normalised against various parameters, for example a predetermined acceptable total output volume, or against a measurement of the input volumes of the incoming data streams. For example, it may be normalised against an average of the mean input volumes of each of the received audio data streams. In the mixing modes shown in Figures 4 to 6, the normalised total output volume is normalised against an average of the mean input volumes of each of the received audio data streams, taken over a periodic or sliding window of between 0.5 to 5 seconds, preferably between 1 and 2 seconds.
In some embodiments, the normalised total output volume is, in certain audio mixing modes, relatively low, for example in the région of 0.4 to 0.9, in this embodiment 0.6. However, due to the increase in the contribution from the audio data stream received from TD1 in the mixed audio data stream from time tl onwards, the normalised total output volume of the mixed media stream can be seen to increase to a relatively high level, e.g. in the région of 0.9 to 1.5, in this embodiment 1.0. When the normalisation of the total output volume is, for example, conducted in relation to a predetermined acceptable total output volume, or against a measurement of the input volumes of the incoming data streams, this provision of a normalised output volume in default mode which is below the normalised output volume in a non-default mode provides the advantage that, when a récipient of the mixed audio stream receives the nondefault mode mixed stream, they do not need to tum down the output volume on their téléphoné speaker since this is normally set to an acceptable level comparable to the normalised, relatively high level, of the non-default mode mixed stream.
In the mixing modes shown in Figures 4 to 6, the volume adjustment levels used for each of the received audio data streams Tl to T5 are shown schematically, to illustrate their respective sizes.
In the mixing modes shown in Figures 4 to 6, the first respective volume adjustment levels comprise substantially equal volume adjustment levels, whereby each input audio data stream in the plurality (i.e. TD1 to TD5) has a substantially equal volume adjustment level applied to generate the first mixed audio data stream transmitted to TD6.
Such an initial mixed audio data stream, transmitted to TD6, is depicted in each of Figures 4 to 6, in particular between time = 0 and time = tl. It can be seen that each of the audio data streams received from TD1 to TD5 has a substantially equal volume adjustment level (represented by equal area in Figures 4 to 6) applied in the first mixed audio data stream. The audio data streams received from TD1 to TD5 make a non-zero, substantially equally adjusted, contribution, but there is no contribution from the audio data stream received from TD6.
In response to detecting a trigger during the multi-party audio téléconférence call, a second mixed audio data stream generated using a second audio mixing mode is transmitted to TD6. The second mixed audio data stream is generated from the same plurality of data streams, but the respective volume adjustment levels are different, the plurality of data streams again being the audio data streams received from TD1 to TD5.
Similarly, mixed audio data streams comprising different mixes of received audio data streams will also be transmitted to the other telephony devices TD1 to TD5.
In the exemplary audio mixing mode shown in Figure 4, the second respective volume adjustment levels comprise a relatively high volume adjustment level applied to a first audio data stream received from TD1 and relatively low, substantially equal, volume adjustment levels applied to the other audio data streams in the plurality, i.e. the audio data streams received from TD2 to TD5. The volume adjustment level applied to TD1 may for example be at least double that applied to each of TD2 to TD5. Such a mixed audio data stream is depicted in Figure 4, in particular during the period between time = tl and time = t2. It can be seen that each of the audio data streams received from TD2 to TD5 has a substantially equal volume adjustment, that there is a relatively higher contribution fforn the audio data stream received from TD 1 and that there is no contribution from the audio data stream received from TD6. In this embodiment, ail the audio data streams received from TD1 to TD5 make a non-zero contribution, but there is no contribution from the audio data stream received from TD6.
The volume adjustment levels applied to the audio data streams received from TD2 to TD5 during the period between time = tl and time = t2 are lower than the volume adjustment levels applied to the audio data streams received from the same sources, TD2 to TD5, during the period between time = 0 and time = tl. This is depicted in Figure 4 where the contributions to the mixed audio data stream of TD2 to TD5 during the period between time = 0 and time = tl can be seen to be greater (schematically represented as relatively large areas in Figure 4) than the contributions to the mixed audio data stream of TD2 to
TD5 during the period between time = tl and time = t2 (schematically represented as relatively small areas in Figure 4).
Figure 4 shows that during the period between time = 0 and time = tl, the total volume adjustment level of the mixed media stream is relatively low, for example in the région of 0.4 to 0.9, in this embodiment 0.6. However, due to the increase in the contribution from the audio data stream received from TD1 in the mixed audio data stream during the period between time = tl and time = t2 , the total volume adjustment level of the mixed media stream can be seen to increase to a relatively high level, for example in the région of 0.9 to 1.5, in this embodiment 1.0.
In embodiments, the trigger comprises the mean input volume of the first received audio data stream rising above a first predetermined threshold. This could for example be due to the user of TD1 raising their voice to make an announcement to others in the room in which the user is located. Increasing the contribution due to the voice of the user of TD1 in the mixed audio data streams will enable the announcement to be more easily distinguished in the mixed audio data streams. Increasing the contribution due to the voice of the user of TD1 in the mixed audio data streams could continue for a predetermined time period, and/or after the détection of a predetermined period of silence (e.g. detected as a signal in which mean input volume is below a predetermined threshold) from the voice of the user of TD1, for example until time t2. This can be useful for example to ensure that the volume adjustment level for TD1 is increased for the entire announcement rather than just an initial part of the announcement. An audio mixing mode similar, or identical, to that applied prior to time = tl may thereafter be used again, as shown in Figure 4.
In the exemplary audio mixing mode shown in Figure 5, the second respective volume adjustment levels comprise relatively high, substantially equal, volume adjustment levels applied to both a first audio data stream received from TD1 and a second audio data stream received from TD2 and relatively low, substantially equal, volume adjustment levels applied to the other audio data streams in the plurality, i.e. the audio data streams received from
TD3 to TD5. The volume adjustment levels applied to TD1 and TD2 may for example be at least double that applied to each of TD3 to TD5. Such a mixed audio data stream is depicted in Figure 5, in particular during the period between time = tl and time = t2. It can be seen that each of the audio data streams received from TD3 to TD5 has a substantially equal volume adjustment, that there is a relatively higher contribution from the audio data stream received from TD1 and TD2 and that there is no contribution from the audio data stream received from TD6. In this embodiment, ail the audio data streams received from TD1 to TD5 make a non-zero contribution, but there is no contribution from the audio data stream received from TD6.
The volume adjustment levels applied to the audio data streams received from TD3 to TD5 during the period between time = tl and time = t2 are lower than the volume adjustment levels applied to the audio data streams received from the same sources, TD3 to TD5, during the period between time = 0 and time = tl. This is depicted in Figure 5 where the contributions to the mixed audio data stream of TD3 to TD5 during the period between time = 0 and time = tl can be seen to be greater (schematically represented as relatively large areas in Figure 5) than the contributions to the mixed audio data stream of TD3 to TD5 during the period between time = tl and time = t2 (schematically represented as relatively small areas in Figure 5).
Figure 5 shows that during the period between time = 0 and time = tl, the total volume adjustment level of the mixed media stream is relatively low, for example in the région of 0.4 to 0.9, in this embodiment 0.6. However, due to the increase in the contribution from the audio data streams received from TD1 and TD2 in the mixed audio data stream during the period between time = tl and time = t2 , the total volume adjustment level of the mixed media stream can be seen to increase to a relatively high level, for example in the région of 0.9 to 1.5, in this embodiment 1.0.
In these embodiments, the trigger is detected at time tl and could comprise the mean input volume of one or more of the first received audio data stream and the second received audio data stream rising above a second predetermined threshold. This could for example be due to the users of TD1 and TD2 having a relatively loud conversation with each other. Increasing the volume adjustment level of the voices of the users of TD1 and TD2 in the mixed audio data streams will enable their conversation to be more easily distinguished in the mixed audio data streams thus helping them to converse more easily above contributions to the mixed audio data streams due to background noise or chatter. Increasing the volume adjustment levels of the voice of the users of TD1 and TD2 in the mixed audio data streams could continue for a predetermined time period after the détection, for example until time t2. An audio mixing mode similar, or identical, to that applied prior to time = tl may thereafter be used again, as shown in Figure 5.
Embodiments comprise storing data représentative of one or more predetermined words and monitoring the received audio data streams for the utterance of any of the one or more predetermined words represented in the stored data. Monitoring module 204 of ATS 100 includes speech récognition capabilities, which may be embodied by any suitable speech récognition engine known in the art, such that when a user ofa telephony device utters any of the one or more predetermined words, this can be detected in the audio data stream received from that user’s telephony device. In such embodiments, the trigger comprises the monitoring detecting utterance of a given one or more predetermined words represented in the stored data in a received audio data stream.
In some embodiments, the given one or more predetermined words are uttered by a user associated with TD1 and are contained in the audio data stream received from TD1. The given one or more predetermined words comprise an identifier for a user associated with TD2. Such embodiments allow two users to conduct a conversation with each other within the multi-party audio téléconférence. The volume adjustment level of the conversation between the two users is increased in the resulting mixed audio data stream which allows their conversation to be distinguished more easily. The identifier could comprise the name of one of the two users such that the conversation can be initiated by one of the users uttering the name of the other user.
In other embodiments, the given one or more predetermined words are uttered by a first user associated with TD1 and the given one or more words comprise an identifier for a second user associated with TD2. The given one or more words further comprise an indication that the first user wishes to conduct a private téléphoné call with the second user separate to the multi-party audio téléconférence call. The second mixed audio data stream is thus generated with second, different respective volume adjustment levels for the audio data streams in the plurality apart from audio data streams received from TD1 and TD2. In the private téléphoné call, the audio data stream received from TD1 is transmitted to TD2 and the audio data stream received from the TD2 is transmitted to TD1, with substantially equal volume adjustment levels applied in each case - similar to a standard two-party téléphoné call.
The users of TD1 and TD2 are thus able to hâve a private téléphoné call separate to the multi-party audio téléconférence call and the multi-party audio téléconférence call carries on between the remaining telephony devices (TD3 to TD5). The indication that the first user wishes to conduct a téléphoné call with the second user separate to the multi-party audio téléconférence call may comprise the first user uttering one or more key words, or a phrase, such as “private call” which are predetermined as being opérable to trigger a private call, plus an identifier for the second user such as “Hey Joe”.
In the exemplary audio mixing mode shown in Figure 6, the media stream transmitted to TD6 during such a private call between TD1 and TD2 is shown between time = tl and time = t2. The trigger is detected at time = tl. It can be seen that each of the audio data streams received from TD3 to TD5 has a substantially equal volume adjustment level and contributes substantially equally to the total volume of the first mixed audio data stream, and that there is no contribution from any of the audio data streams received from TD1, TD2 or TD6. Each of the audio data streams from TD1 and TD2 are thus eut from the mixed audio data stream sent to TD6. As such, the audio data streams received from TD3 to TD5 make a non-zero contribution, but there is no contribution from the audio data stream received from TD1, TD2 or TD6, in the mixed audio data stream sent to TD6. On the other hand, the audio data stream from TD1 is sent, unmixed, to the other participant in the private conversation, TD2, and the audio data stream from TD2 is sent, unmixed, to TD2, during this period.
In the exemplary audio mixing mode shown in Figure 6, the second respective volume adjustment levels comprise relatively high, substantially equal, volume adjustment levels applied to each of the audio data streams received from TD3 to TD5 and the audio data streams received from TD1 and TD 2 are eut out in the mixed audio data stream sent to TD6. The respective volume adjustment levels applied to each of the audio data streams received fforn TD3 to TD5 received during the period time = 0 to time = tl comprise relatively low, substantially equal, volume adjustment levels. The volume adjustment levels applied to each of TD3 to TD5 during the period time = tl to time = t2 may for example be at least 1.5 times that applied during the period time = 0 to time = tl. Such a mixed audio data stream is depicted in Figure 6, in particular during the period between time = tl and time = t2. It can be seen that each of the audio data streams received from TD3 to TD5 has a substantially equal volume adjustment, that there is no contribution from the audio data stream received from TD1 and TD2 and that there is no contribution from the audio data stream received from TD6. In this audio mixing mode, ail the audio data streams received from TD3 to TD5 make a non-zero contribution, but there is no contribution from the audio data stream received from TD1, TD2 and TD6.
The volume adjustment levels applied to the audio data streams received fforn TD3 to TD5 during the period between time = tl and time = t2 are higher than the volume adjustment levels applied to the audio data streams received fforn the same sources, TD3 to TD5, during the period between time = 0 and time = tl. This is depicted in Figure 6 where the contributions to the mixed audio data stream of TD3 to TD5 between time = 0 and time = tl can be seen to be lower (schematically represented as relatively small areas in Figure 6) than the contributions to the mixed audio data stream of TD3 to TD5 between time = tl and time = t2 (schematically represented as relatively large areas in Figure 6).
Figure 6 shows that between time = 0 and time = tl, the total volume adjustment level of the mixed media stream is relatively low, for example in the région of 0.4 to 0.9, in this embodiment 0.6. However, due to the cutting out of the contribution from the audio data streams received from TD1 and TD2 in the mixed audio data stream during the period between time = tl and time = t2, and the relative increase in the contributions from the audio data streams received from TD3 to TD5 in the mixed audio data stream during the period between time = tl and time = t2, the mixed media stream can be seen to remain at a relatively low level, for example in the région of 0.4 to 0.9, in this embodiment 0.6.
The audio data streams from TD1 and TD2 may be re-introduced in the mixed audio data stream sent to TD6 in response to detecting an utterance of a word or phrase, for example “end private” during said private téléphoné call indicating that the user of TD1 wishes to end the private téléphoné call with the user of TD2. An audio mixing mode similar, or identical, to that applied prior to time = tl may thereafter be used again, as shown in Figure 6.
In some embodiments, the user of TD1 has access to computing device 114. In such embodiments, TD1 may for example be a desktop téléphoné and computing device 114 may be a desktop personal computer. ATS 100 can provide information about a multi-party audio téléconférence call being conducted between TD1 and other telephony devices, for example TD2 to TD6. ATS 100 provides a log-in web page via its web interface 210. The user of TD1 is able to log-in to the web page, for example by entering the téléphoné dialling number of TD1 or other appropriate identifier which ATS 100 will recognise as being associated with the multi-party audio téléconférence call being conducted between TD1 and TD2 to TD6.
In embodiments, the web interface 210 allows display of information associated with the multi-party audio téléconférence such as visual indicators for each audio téléconférence call participant. The web interface 210 allows a user to configure various settings such that the user can configure multi-party audio téléconférence services provided via ATS 100.
Exampie of such configurable settings include the predetermined time period, the predetermined threshold volume adjustment level and the one or more predetermined words described above in relation to multi-party audio téléconférence control embodiments.
In embodiments, the web interface 210 allows triggering for switching between mixing modes. For example, instead of a user having to raise their voice in order for the second audio mixing mode to be employed to generate a mixed audio data stream, the user may instead enter appropriate user input via the web interface to instruct ATS 100 manually. Further, a user may manually trigger a direct conversation with another user (with associated boosting of their associated audio data stream in the mixed audio data stream), for example by clicking on a visual indicator associated with that user in the web interface 210. Similarly, a user may manually trigger a private conversation with another user (with their respective audio data stream not being combined into mixed audio data streams transmitted to the remaining participants of the multi-party audio téléconférence call), for example by clicking on a further visual indicator associated with that user in the web interface 210.
The above embodiments are to be understood as illustrative examples of the invention. Further embodiments of the invention are envisaged.
Whilst in the above embodiments, the multi-party téléconférence call is between a total of six participants, it should be understood that any number of participants may be accommodated within a téléconférence, and that the audio mixing modes shown may hâve any number of participants within a particular mixing mode. For example a mixing mode in which one participant is emphasized, and a mixing mode in which one or more other participants are deemphasized (but still heard) may be provided, similar to the embodiment shown in Figure 4. Further, for example a mixing mode in which two or more participants are emphasized, and a mixing mode in which one or more other participants are de-emphasized (but still heard) may be provided, similar to the embodiment shown in Figure 5. Further, for example a mixing mode in which two or more participants having a private chat are eut out, and in which one or more other participants are relatively emphasized (compared to a default mixing mode) may be provided, similar to the embodiment shown in Figure 6.
For example, telephony device TD1 could comprise a mobile telephony device and téléphoné switch 104 could comprise a mobile switching centre in a mobile télécommunications network connected to network 102. In such embodiments, instead of the display and user input capabilities of computing device 114 being employed to interface with the web server interface of ACS 100, corresponding display and user input functionality of the mobile telephony device could be employed. The invention could be implemented on a mobile telephony device as an application installed on the mobile telephony device.
In further alternative embodiments, telephony device TD1 could comprise a smart desk phone. In such embodiments, instead of the display and user input capabilities of computing device 114 being employed to interface with the web server interface of ACS 100, corresponding display and user input functionality of the smart desk phone could be employed.
In further alternative embodiments, the invention could be implemented using a softphone, for example installed on computing device 114; such embodiments need not require a separate telephony device such as TD1.
In the above embodiments, the techniques of the invention are applied to audio téléconférence services; however it should be appreciated that the invention may also be applied in relation to other téléconférence services, such as video téléconférence services.
It is to be understood that any feature described in relation to any some embodiments may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, équivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.
Contents2
14 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO0190839A2 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| EP1662766A1 | Cites | European Patent Office (EPO) | Search report |
| EP1804474A1 | Cites | European Patent Office (EPO) | Search report |
| EP1942646A2 | Cites | European Patent Office (EPO) | Search report |
| US2006023900A1 | Cites | United States of America | Search report |
| US2007263603A1 | Cites | United States of America | Search report |
| US2008037749A1 | Cites | United States of America | Search report |
| US2008159490A1 | Cites | United States of America | Search report |
| US2009097625A1 | Cites | United States of America | Search report |
| US2011044474A1 | Cites | United States of America | Search report |
| CA2242426A1 | Cites | Canada | Search report |
| GB2284968A | Cites | United Kingdom | Search report |
| US6008838A | Cites | United States of America | Search report |
| US6792092B1 | Cites | United States of America | Search report |
| US7734692B1 | Cites | United States of America | Search report |
| WO9931865A1 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| WO2001090839A2 | Cites | World Intellectual Property Organization (WIPO) | – |
| WO1999031865A1 | Cites | World Intellectual Property Organization (WIPO) | – |
| CA002242426A | Cites | Canada | – |
| US20110044474A1 | Cites | United States of America | – |
| US20090097625A1 | Cites | United States of America | – |
| US20080159490A1 | Cites | United States of America | – |
| US20080037749A1 | Cites | United States of America | – |
| US20070263603A1 | Cites | United States of America | – |
| US20060023900A1 | Cites | United States of America | – |
4 members in 2 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 201110494 | United Kingdom | A | |
| GB20110010494 | – | – | – |
Members4
| Document | Office | Kind | |
|---|---|---|---|
| GB2492103A | United Kingdom | A | |
| WO2012175964A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2012175964A3 | World Intellectual Property Organization (WIPO) | A3 | |
| GB2492103BThis record | United Kingdom | B |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Amendments to the register in respect of changes of name or changes affecting rights (sect. 32/1977)REGISTERED BETWEEN 20250619 AND 20250625732E | 732E |
Numbers
- Publication
- 2492103
- Publication, DOCDB
- 2492103
- Publication, EPODOC
- GB2492103
- Application
- 11104940
- Application, DOCDB
- 201110494
- Application, EPODOC
- GB20110010494
Titles
- English
- Multi party teleconference methods and systems
Classification
- CPC, 8
- H04M3/56
- H04M3/38
- H04L12/1822
- H04L29/06414
- H04M2203/5009
- H04M3/20
- H04M3/568
- H04L65/403
- IPC, 4
- H04M3 56
- H04L12 18
- H04L29 06
- H04M3 20
