Managing a packet switched conference call
Abstract
Method for managing a conference call, centralized, by switching packets, between a plurality of terminals (13), said method comprising a conference call server (12): - receiving data packets from all terminals (13) participating in said conference call, whose said data packets include either voice data or background noise information as well as an identifier associated with the respective terminal (13) that provides said voice data or said background noise information; - determine, based on said received data packets, at least one terminal (13) that provides voice data at that time, if any, among said terminals (13) participating in said conference call; - mixing said received voice data and said received background noise information and said mixed data is inserted into new data packets together with at least one identifier associated with one of said terminals (13) that were determined to provide data at that time of voice, if any, in such a way that said at least one identifier can be differentiated with respect to any other information included in said data packets; and - transmitting said new data packets to terminals (13) participating in said conference call.

Term
Term ended
Projected expiry passed 4 July 2022, 4.2 years ago.
- Filed
- Priority
- Published
- Projected expiry
- Today
17 claims: 7 independent, 10 dependent
- 1ES 2 296 950 Τ3 REIVINDICACIONES 1. Método para gestionar una llamada en conferencia, centralizada, por conmutación de paquetes, entre una pluralidad de terminales (13), comprendiendo dicho método en un servidor de llamadas en conferencia (12):- recibir paquetes de datos de todos los terminales (13) que participan en dicha llamada en conferencia, cuyos dichos paquetes de datos incluyen bien datos de voz o bien información de ruido de fondo así como un identificador asociado al terminal respectivo (13) que proporciona dichos datos de voz o dicha información de ruido de fondo;- determinar, basándose en dichos paquetes de datos recibidos, por lo menos un terminal (13) que proporciona en ese momento datos de voz, en caso de que hubiera alguno, de entre dichos terminales (13) que participan en dicha llamada en conferencia;- mezclar dichos datos de voz recibidos y dicha información de ruido de fondo recibida y se insertan dichos datos mezclados en paquetes de datos nuevos junto con por lo menos un identificador asociado a uno de dichos terminales (13) que estaban determinados para proporcionar en ese momento datos de voz, en caso de que hubiera alguno, de tal manera que dicho por lo menos un identificador se puede diferenciar con respecto a cualquier otra información incluida en dichos paquetes de datos;y - transmitir dichos paquetes de datos nuevos hacia unos terminales (13) que participan en dicha llamada en conferencia.
- 2Método según la reivindicación 1, en el que dichos identificadores asociados a dichos terminales (13) son unos identificadores asociados aleatoriamente a dichos terminales (13) para dicha llamada en conferencia, comprendiendo dicho método, como etapas anteriores, la recepción, en dicho servidor de llamadas en conferencia (12), de paquetes de control provenientes de dichos terminales (13) que participan en dicha llamada en conferencia, incluyendo dichos paquetes de control una correspondencia de un identificador asociado a un terminal respectivo (13) con una identificación de dicho terminal (13), y el reenvío de dicha correspondencia en paquetes de control desde dicho servidor de llamadas en conferencia (12) hacia dichos terminales (13) que participan en dicha llamada en conferencia.
- 3Método según una de las reivindicaciones anteriores, en el que dicho servidor de llamadas en conferencia (12) transmite en dichos paquetes de datos nuevos exclusivamente identificadores asociados a unos terminales (13) de los cuales se determinó que proporcionaban datos de voz.
- 4Método según la reivindicación 1 ó 2, en el que dicho servidor de llamadas en conferencia (12) incluye, en dichos paquetes de datos nuevos, unos identificadores asociados a unos terminales (13) que proporcionan en ese momento datos de voz, así como unos identificadores asociados a unos terminales que proporcionan en ese momento información de ruido de fondo, incluyéndose, en dicho paquete de datos en una posición predeterminada entre todos los identificadores incluidos, por lo menos un identificador asociado a un terminal (13) del cual se determinó que proporcionaba datos de voz.
- 5Método según la reivindicación 1 ó 2, en el que dicho servidor de llamadas en conferencia (12) incluye, en dichos paquetes de datos nuevos, unos identificadores asociados a unos terminales (13) que proporcionan en ese momento unos datos de voz así como unos identificadores asociados a unos terminales (13) que proporcionan en ese momento información de ruido de fondo, en el que por lo menos un identificador asociado a uno de dichos terminales (13) de los cuales se determinó que proporcionaban datos de voz, en caso de que hubiera alguno, se incluye en dichos paquetes de datos en una posición predeterminada entre todos los identificadores incluidos, y en el que identificadores asociados a unos terminales (13) de los cuales se determinó que proporcionaban datos de voz se separan mediante un marcador con respecto a identificadores incluidos asociados a otros terminales (13).
- 6Método según la reivindicación 5, en el que dicho marcador se corresponde con un identificador asociado a dicho servidor de llamadas en conferencia (12).
- 7Método según una de las reivindicaciones anteriores, en el que dicha llamada en conferencia se basa en el Protocolo de Transporte de Tiempo Real, en el que dichos paquetes de datos son paquetes del Protocolo de Transporte de Tiempo Real, en el que dichos identificadores asociados a dichos terminales (13) son unos identificadores de Fuentes de Sincronización, y en el que dichos identificadores se incluyen por parte de dicho servidor de llamadas en conferencia (12) en dichos paquetes de datos nuevos en un campo proporcionado en un encabezamiento de paquete correspondiente a una lista de Fuentes Contribuyentes.
- 8Método según una de las reivindicaciones anteriores, que comprende asimismo la recepción de dichos paquetes de datos nuevos transmitidos por dicho servidor de llamadas en conferencia (12) en un terminal (13) que participa en dicha llamada en conferencia y la indicación, de una identificación (32, 33) de por lo menos un terminal (13) del que se ha determinado que proporciona datos de voz, a un usuario basándose en un identificador incluido en dichos paquetes de datos nuevos recibidos. ES 2 296 950 T3
- 9Servidor de llamadas en conferencia (12) que comprende unos medios para gestionar una llamada en conferencia, centralizada, entre una pluralidad de terminales (13), incluyendo dichos medios - unos medios (15) configurados para recibir unos paquetes de datos de todos los terminales (13) que participan 5 en dicha llamada en conferencia, incluyendo dichos paquetes de datos bien datos de voz o bien información de ruido de fondo así como un identificador asociado al terminal respectivo (13) que proporciona dichos datos de voz o dicha información de ruido de fondo;- unos medios (15) configurados para determinar, basándose en dichos paquetes de datos recibidos, por lo menos un 10 terminal (13) que proporciona en ese momento datos de voz, en caso de que hubiera alguno, de entre dichos terminales (13) que participan en dicha llamada en conferencia;- unos medios (15) configurados para mezclar dichos datos de voz recibidos y dicha información de ruido de fondo recibida e insertar dichos datos mezclados en paquetes de datos nuevos junto con por lo menos un identificador 15 asociado a uno de dichos terminales (13) de los que se determinó que proporcionaban en ese momento datos de voz, en caso de que hubiera alguno, de tal manera que dicho por lo menos un identificador se puede diferenciar con respecto a cualquier otra información incluida en dichos paquetes de datos;y - unos medios (15) configurados para transmitir dichos paquetes de datos nuevos hacia terminales (13) que parti20 cipan en dicha llamada en conferencia.
- 10Servidor de llamadas en conferencia (12) según la reivindicación 9, en el que dichos identificadores asociados a dichos terminales (13) son unos identificadores asociados aleatoriamente a dichos terminales (13) para dicha 11amada en conferencia, comprendiendo asimismo unos medios (15) configurados para recibir unos paquetes de control 25 provenientes de dichos terminales (13) que participan en dicha llamada en conferencia, incluyendo dichos paquetes de control una correspondencia de un identificador asociado a un terminal (13) respectivo con una identificación de dicho terminal (13), y unos medios (15) configurados para reenviar dicha correspondencia en paquetes de control hacia dichos terminales (13) que participan en dicha llamada en conferencia. 30
- 11Servidor de llamadas en conferencia (12) según la reivindicación 9, en el que dichos medios (15) están configurados para transmitir en dichos paquetes de datos nuevos exclusivamente identificadores asociados a unos terminales (13) de los cuales se determinó que proporcionaban datos de voz.
- 12Servidor de llamadas en conferencia (12) según la reivindicación 9 ó 10, en el que dichos medios (15) están 35 configurados para incluir, en dichos paquetes de datos nuevos, unos identificadores asociados a unos terminales (13) que proporcionan en ese momento datos de voz así como identificadores asociados a unos terminales que proporcionan en ese momento información de ruido de fondo, incluyéndose, en dicho paquete de datos en una posición predeterminada entre todos los identificadores incluidos, por lo menos un identificador asociado a un terminal (13) del cual se determinó que proporcionaba datos de voz.
- 13Servidor de llamadas en conferencia (12) según la reivindicación 9 ó 10, en el que dichos medios (15) están configurados para incluir, en dichos paquetes de datos nuevos, unos identificadores asociados a unos terminales (13) que proporcionan en ese momento datos de voz así como identificadores asociados a unos terminales (13) que proporcionan en ese momento información de ruido de fondo, en el que por lo menos un identificador asociado a uno 45 de dichos terminales (13) de los cuales se determinó que proporcionaban datos de voz, en caso de que hubiera alguno, se incluye en dichos paquetes de datos en una posición predeterminada entre todos los identificadores incluidos, y en el que identificadores asociados a unos terminales (13) de los cuales se determinó que proporcionaban datos de voz se separan mediante un marcador con respecto a unos identificadores incluidos asociados a otros terminales (13).
- 14Servidor de llamadas en conferencia (12) según la reivindicación 13, en el que dicho marcador se corresponde con un identificador asociado a dicho servidor de llamadas en conferencia (12).
- 15Servidor de llamadas en conferencia (12) según una de las reivindicaciones 9 a 14, en el que dicha llamada en 55 conferencia se basa en el Protocolo de Transporte de Tiempo Real, en el que dichos paquetes de datos son paquetes del Protocolo de Transporte de Tiempo Real, en el que dichos identificadores asociados a dichos terminales (13) son identificadores de Fuentes de Sincronización, y en el que dichos medios (15) están configurados para incluir dichos identificadores en dichos paquetes de datos nuevos en un campo proporcionado en un encabezamiento de paquete correspondiente a una lista de Fuentes Contribuyentes.
- 16Aparato (15) para un servidor de llamadas en conferencia centralizadas (12), comprendiendo dicho aparato (15) unos medios configurados para - recibir paquetes de datos de todos los terminales (13) que participan en dicha llamada en conferencia, incluyendo 65 dichos paquetes de datos bien datos de voz o bien información de ruido de fondo así como un identificador asociado al terminal respectivo (13) que proporciona dichos datos de voz o dicha información de ruido de fondo;ES 2 296 950 T3 - determinar, basándose en dichos paquetes de datos recibidos, por lo menos un terminal (13) que proporciona en ese momento datos de voz, en caso de que hubiera alguno, de entre dichos terminales (13) que participan en dicha llamada en conferencia;- mezclar dichos datos de voz recibidos y dicha información de ruido de fondo recibida e insertar dichos datos mezclados en paquetes de datos nuevos junto con por lo menos un identificador asociado a uno de dichos terminales (13) de los cuales se determinó que proporcionaban en ese momento datos de voz, en caso de que hubiera alguno, de tal manera que dicho por lo menos un identificador se puede diferenciar con respecto a cualquier otra información incluida en dichos paquetes de datos;y - proporcionar dichos paquetes de datos nuevos para su transmisión hacia unos terminales (13) que participan en dicha llamada en conferencia.
- 17Terminal (13) que comprende unos medios para participar en una llamada en conferencia, centralizada, incluyendo dichos medios - unos medios para recibir paquetes de datos transmitidos por un servidor de llamadas en conferencia (12), comprendiendo dichos paquetes de datos, datos de voz y/o información de ruido de fondo mezclados proporcionados por unos terminales (13) que participan en dicha llamada en conferencia y por lo menos un identificador asociado a un terminal (13) del cual se determinó en dicho servidor de llamadas en conferencia (12) que proporcionaba en ese momento datos de voz, en el caso de que hubiera alguno;- unos medios para reconocer, en paquetes de datos recibidos, unos identificadores asociados a unos terminales (13) de los cuales se determinó en un servidor de llamadas en conferencia (12) que proporcionaban en ese momento datos de voz;y - unos medios para indicar a un usuario una identificación de unos terminales (13) que proporcionan datos de voz basándose en identificadores reconocidos asociados a unos terminales (13) de los cuales se determinó en un servidor de llamadas en conferencia (12) que proporcionaban en ese momento datos de voz.
Independent claims17
67 paragraphs in 1 section, as filed
is 2 296 950 T3
DESCRIPTION
Management of a packet-switched conference call.
Field of the invention
The present invention relates to a method for managing a centralized, packet-switched conference call between a plurality of terminals. The invention also relates to a conference call server comprising means for managing a centralized conference call, and to a terminal 10 comprising means for participating in a centralized conference call.
Background of the invention
In a conference call, a group of terminal users is connected to each other so that when one of the participating users speaks, all of the other participating users can hear the voice of the participant who is speaking. In such a communication, normally only one of the participating users is speaking at the same time, while the other users are listening. In a centralized conference call, the terminals of the participating users are not directly connected to each other, but through a conference call server. A centralized conference call can be made, for example, by means of a Voice over Internet Protocol (VoIP) conference calling application on the Internet or in the form of an audio conference in the packet-switched domain of the Services networks. of Universal Mobile Telecommunications (UMTS).
In a VoIP session, voice data is typically transported using Real Time Transport Protocol (RTP) over Internet Protocol (IP) and User Datagram Protocol (UDP). RTP has been described in detail in RFC 1889: "RTP: A Transport Protocol for Real-Time Applications", January 1996, by H. Schulzrinne et al.
An end-to-end VoIP connection is often called a VoIP tunnel. In the typical centralized conference call setup 30, VoIP tunnels are formed between each participating terminal and the conference call server.
By way of illustration, figure 1 shows the tunneling of coded voice in a centralized RTP-based conference.
Figure 1 schematically shows a centralized conference calling system in a packet-switched domain of a UMTS network 11, with a conference call server 12 connected to this network 11 and with a plurality of mobile terminals 13. The mobile terminals 13 connect to the conference call server over the UMTS network 11 using RTP tunnels 14.
At the terminals 13, the voice data produced by the respective user of the terminals 13 is first encoded and then inserted into the RTP packet payload. There are a multitude of alternative audio encoders that can be used to perform specific speech encoding. For example, the Adaptive Multi-Rate (AMR) speech codec, which is specified as the mandatory speech codec for 45 systems 3<sup>to</sup> generation, it could be used to compress the voice data carried in the RTP payload. The encoders encode the speech samples into frames, which are then transported over the RTP / UDP / IP protocols through the UMTS network 11 to the conference call server 12.
The conference call server 12 comprises an RTP mixer 15, which receives the incoming RTP packet streams 50 from the connected terminals 13, removes the RTP packing, combines the streams into a single RTP packet stream, and then sends this stream. to each of the terminals 13.
Each RTP packet transmitted between terminals 13 and conference call server 12 is associated with a header. The structure of this header, which is specified in the aforementioned RFC 1889, is illustrated in figure 2. The header comprises a V field in which it identifies the version of the RTP used and a P field for a stuffing bit. If the padding bit is set, the packet contains one or more additional padding octets at the end which are not part of the payload. The header further comprises an X field for an extension bit. If the extension bit is set, the set header is followed by exactly one header extension. On the other hand, the header comprises a CC field for the count of 60 outgoing Contributing Sources (CSRC), which contains the number of CSRC identifiers that follow the set header, and an M field corresponding to a marker bit, defining the interpretation of the bookmark for a profile. Additionally, the header comprises a PT field to identify the payload format and a field for a Sequence Number, which is incremented by one for each RTP data packet sent. The Sequence Number can be used by the receiver to detect a packet loss and to restore the sequence of the packets. The header further comprises a field corresponding to a Timestamp, which reflects the sampling instant of the first octet in the RTP data packet.
is 2 296 950 T3
Furthermore, the RTP packet headers carry a Synchronization Source identifier (SSRC) and, as mentioned above with reference to the CC field, a list of Contributing Source identifiers (CSRC).
The SSRC identifier is used to identify the synchronization source that has transmitted the RTP packet in question. An SSRC identifier that is unique for the respective RTP session is randomly associated with each possible source, that is, with each of the terminals 13 and the conference call server 12. Each terminal 13 adds the SSRC identifier associated with it to the field SSRC identifier in the RTP header of each RTP packet you assemble. Similarly, the RTP mixer 15 of the conference call server 10 12 adds the SSRC identifier associated with the conference call server 12 to the SSRC identifier field of the RTP header of each RTP packet leaving the server 12.
The CSRC list is used to identify the different sources that contribute to an RTP packet and is therefore only relevant for RTP packets assembled on the conference call server 13. Mixer 15 RTP 15 adds to the CSRC fields of the packets Outgoing RTP the SSRC identifiers of those terminals 13 that contribute to the combined outgoing VoIP flow.
Additionally, to allow a control of VoIP connections using RTP, in the aforementioned RFC 1889 a Real Time Control Protocol (RTCP) is defined. RTCP is used for example to keep both ends of a connection informed about the quality of service that they are providing and receiving. This information is sent in packets of the sender report (SR) and receiver report (RR) RtCp type. Additionally, the RTP specification defines an RTCP Source Description (SDES) packet. SDES RTCP packets can be used by the source to provide more information about itself. The CNAME or NAME SDES packets can be used for example to provide a mapping between the random SSRC identifier and the source identity. CNAME SDES packets are intended to provide canonical endpoint identifiers, while NAME SDES packets are intended to provide a real name used to describe the respective source. RTP mixer 15 is expected to combine SR and RR type RTCP packets from all terminals 13 before forwarding them. In contrast, SDES-type RTCP packets are forwarded by RTP mixer 15 to all participants in conference 13 without modification.
In a conference call, it is sometimes difficult for participating users to immediately recognize who is speaking. This in particular is a problem in the case of many users participating in a conference call, when these participating users do not know each other very well.
The aforementioned RFC 1889 states that an illustrative application is conducting an audio conference in which a mixer indicates all the speakers whose voice was combined to produce the output packet, allowing the receiver to indicate the current speaker, even though all the Audio packets contain the same SSRC identifier, that is, the one corresponding to the mixer.
However, in any reasonable VoIP use of a voice codec, the codec will emit Silence Descriptor (SID) frames that allow comfort noise generation at the receiving end, as long as the respective participant in the conference is idle, i.e. , listening. In this way, all sources will always produce a signal that is transmitted to the conference call server 12. The conference call server 12 decodes 45 the VoIP streams received from each of the participants passing them to voice or SID frames to sum them before encoding the outgoing voice and the SID frames to be transmitted to the terminals 13. This implies that the SSRC identifiers of all the terminals 13 are included by the mixer 15 in the CSRC list of the mixed outgoing RTP packets, and therefore it is impossible for the receiving terminals 13 to differentiate between active and inactive participants. . It should be noted that the inclusion of the SSRC identifiers of all participating terminals 13 in the CSRC list also has its advantages, for example, to keep each participating user updated on the number and identity of all the other participating users. in the conference.
US patent n<sup>or</sup> 6,292,979 B1 describes a telecommunications conference system with three or more telephone sets connected to a network to participate simultaneously in a conference call. Each telephone set 55 engaged in the conference call receives a list of participants in the call and an identifier associated with that call. Each telephone set generates voice data packets that include the identifier and forwards the generated packets to the network. Each telephone set produces locally a combination of the received packets.
WO 00/72560 Al relates to a teleconference bridge. The document proposes the use of a separate mixing function for each participant in a conference. Incoming audio signals are received and transmitted by an audio conference bridge system. The system performs audio mixing for the participating stations by mixing a plurality of incoming audio signals in the conference according to mixing parameters. It can also output active speaker indicators for each participating station 65 indicating, for each mixed output signal, which incoming audio signals are being mixed. Once the incoming audio signals have been properly mixed, a separate mixed audio signal is output to each participating station. Active speaker indicators can be translated by participating stations into a visual indication of which participant's voice is being heard at any given time.
is 2 296 950 T3
Summary of the invention
One of the objectives of the invention is to improve the comfort of a user participating in a voice over IP conference call.
This objective is achieved according to the invention with a method for managing a centralized, packet-switched conference call between a plurality of terminals, which comprises, as a first step, the reception, in a conference call server, of packets data from all terminals participating in the conference call. These data packets include speech data or background noise information and an identifier associated with the respective terminal that provides the speech data or background noise information. In a second step, at least one terminal is determined to provide voice data at that time, if any, among the terminals participating in the conference call based on the received data packets. Obviously, in the event that none of the users participating in the conference call speaks for a time, none of the terminals will provide voice data during that time, and no terminal that provides voice data can be determined. In a third stage, the received voice data and background noise information are mixed and inserted into new data packets together with at least one identifier associated with one of the terminals that were determined to provide data at that time. voice, if any. The identifier is included in a data packet so that it can be distinguished from any other information included. This implies in particular that the at least one identifier can be differentiated with respect to other possibly included identifiers which are not necessarily associated with terminals that provide voice data. Finally, the new data packets are transmitted by the conference call server to terminals participating in the conference call.
The object of the invention is also achieved with a conference call server comprising means 25 for carrying out the proposed method.
Additionally, the object of the invention is achieved with a terminal which comprises means for participating in a centralized conference call, said means being suitable for making use of the information transmitted according to the invention by a conference call server. To this end, the terminal comprises means 30 for receiving data packets transmitted by a conference call server. The data packets comprise mixed voice data and / or background noise information provided by terminals participating in the conference call and at least one identifier associated with a terminal which was determined in the conference call server providing voice data at that time, if any. On the other hand, the terminal comprises means for recognizing, in received data packets, identifiers associated with terminals of which it was determined in a conference call server that they were currently providing voice data. Furthermore, the terminal comprises means for signaling to a user an identification of terminals that provide voice data based on recognized identifiers associated with terminals of which it was determined in a conference call server that they were currently providing voice data.
The invention has its origin in the idea that a conference call server can be designed that can differentiate between those participants of a conference call who are active at the moment, that is, who provide voice data, and those who in that moment they are inactive, that is, they provide only background noise information. The invention also has its origin in the idea that a terminal can be designed that can signal to a user the currently active participants of a conference call, in the event that it receives corresponding information. Thus, it is proposed that a conference call server makes a determination of the currently active participants of a conference call and that the server forward a corresponding, differentiable information to the terminals participating in the conference call.
One of the advantages of the invention is that it allows an improvement of the user interface of a terminal, since the user can be presented with information transmitted about the active participant in the conference. In this way, conference call participants can always identify the active speaker among all participants.
Preferred embodiments of the invention appear from the dependent claims.
Active terminal identifiers can be transmitted by the conference call server in a variety of ways.
In a first alternative, the conference call server transmits in each combined data packet exclusively an identifier associated with those terminals that are active at that moment. One of the advantages of this approach is that receiving terminals can indicate all active speakers to their users, even in the case of multiple simultaneous speakers. However, with this approach, receiving terminals cannot keep their users updated on all participants.
In a second alternative, the conference call server transmits, in each combined data packet, identifiers corresponding to all terminals participating in the conference, although in such a way that an identifier associated with an active terminal is always presented in a position The default in the identifier list is 2 296 950 T3 ficators, for example, as the first item in the list. Although this approach constantly provides up-to-date information on all conference participants, it does not allow more than one active terminal to be indicated simultaneously. However, in a reasonable dialogue, especially over a telephone connection, only one participant will speak at a time and this problem can be considered minor.
A third alternative is offered by a refinement of the second approach. In this third approach, the conference call server always transmits again, in each combined data packet, identifiers corresponding to all terminals participating in the conference. The identifiers associated with the currently active terminals are displayed at the beginning of the identifier list. Additionally, between the 10 identifiers associated with terminals active at that moment and the identifiers associated with terminals inactive at that moment, a marker is inserted. This third approach combines the advantages of the first and second approaches, simply by introducing an additional value that must be transmitted.
The identifier associated with a respective terminal might not by itself be adequate to identify a transmitting terminal in a receiving terminal, such as the randomly distributed SSRC identifier. In this case, from all possible transmitting terminals to the conference call server and additionally to all possible receiving terminals, a correspondence of the identifiers with a clear identification of the respective terminal is preferably transmitted first. Each receiving terminal can then map a subsequently received identifier associated with a transmitting terminal with a corresponding identification of this terminal. The identification can be in particular a SIP address or a telephone number. The receiving terminal may also have the possibility to additionally map the determined identification to another type of identification. In the case that the identification is, for example, a SIP address or a telephone number, the terminal can establish a correspondence of this address or number with a name or an image stored in a directory of the receiving terminal.
In the event that the user of a terminal is presented with all the conference call participants, the active participants can be pointed to a user in any suitable manner.
The invention can be used in particular, but not exclusively in a system in which centralized conference calls are based on the RTP defined in RFC 1889 mentioned above. In this case, the data packets transmitted from the terminals to the conference call server and from the conference call server to the terminals are RTP packets. The terminal identifiers transmitted by the conference call server in the combined RTP packets may advantageously be SSRC identifiers added to the CSRC list of the RTP header. In the third alternative presented for the transmission of identifiers by the conference call server, the dialer used can be, for example, the SSRC identifier associated with the conference call server. Since the SSRC identifier associated with the conference call server is somehow transmitted in the SSRC field of the RTP header of each combined RTP packet, the receiving terminals are aware of this value and can use it to separate, in the CSRC list, the terminals active with respect to inactive terminals. In contrast, in conventional applications, the SSRC identifier 40 associated with the conference call server is included only in the SSRC field of the outbound combined RTP packets, not in the CSRC list, since the conference call server does not contribute the same to the combined RTP stream.
Each of the three alternatives presented for the transmission of identifiers by the conference call server conforms to the current RTP specification and would not harm implementations that are not designed to make use of special SSRC / CSRC management.
An exhaustive embodiment of the method according to the invention, implemented in an RTP-based system, advantageously comprises three parts. A first part consists of a mechanism for the terminals participating in a conference call to exchange RTP source identifiers and establish correspondence of said identifiers with the respective identity of each terminal or terminal user by means of SDES RTCP packets. A second part consists of a mechanism implemented in the conference call server to set the CSRC field of RTP headers according to predefined rules. A third part consists of a mechanism implemented in the participating receiving terminals to establish correlations of the identifiers of the CSRC field of the RTP packet headers with terminal or user identities, with a view to enabling a presentation of the identity of the active speaker at that point. moment to the users of the receiving terminals.
It should be noted that the number of identifiers that can be transmitted by the conference call server to the participating terminals and / or the number of participants that can be presented by the receiving terminals 60 may be limited to a predetermined value. For example, according to RFC 1889 cited above, the CSRC list is limited to a maximum number of 15 entries.
The invention can be used in particular for the Internet or UMTS packet-switched voice conferencing. In the case of UMTS, information about active participants can be displayed, for example, on the screen of a mobile terminal.
is 2 296 950 T3
Brief description of the figures
Other objects and features of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:
Fig. 1 illustrates the operating principle of a centralized, RTP-based conference calling system;
Fig. 2 illustrates the structure of an RTP header; and Fig. 3 shows a user interface of a terminal, which is making use of one of the embodiments of the method according to the invention.
Detailed description of the invention
An embodiment of the method according to the invention will now be described with reference to Figures 1 to 3.
The embodiment supports the management of VoIP conference calls and is implemented in an RTP-based system which comprises a UMTS network 11, a conference call server 12 that includes an RTP mixer 15 connected to the network 11 and a plurality of terminals 13. Terminals 13 can be connected to conference call server 12 via UMTS 11 by means of RTP tunnels 14. Thus, the system generally corresponds to the system illustrated in Figure 1, which has already been described above.
To establish a VoIP conference call in this system, the Session Initiation Protocol (SIP) is used as the signaling protocol. SIP is used in conjunction with the Session Description Protocol (SDP) to send invitations to called parties and to reach agreements on voice codecs, and so on. The users of the terminals 13 join the conference either by initiating the session themselves by sending the INVITE SIP message to the conference call server 12 or by responding to INVITE messages received through the conference call server 12.
At the beginning of an initiated conference session, the conference software in each terminal 13 sends SDES RTCP packets towards the conference call server 12. These SDES packets carry the SSRC identifier associated with the respective terminal 13 for this session and additionally, in the SDES element field, the SIP address or the telephone number of the respective terminal 13. The conference call server 12 forwards 35 the received SDES packets to each terminal 13 participating in the conference call. Based on the information found in these SDES packets, the terminals 13 prepare to map the SSRC identifiers received during the conference session with corresponding SIP addresses or telephone numbers.
When the conference session is active, all terminals 13 participating in the conference transmit RTP packets to the conference call server 12. To this end, the terminals 13 use a speech codec, for example the AMR speech codec, in such a way that they transmit at a normal speed when there is speech at the input, that is, when the user of the terminal 13 is speaking, and with a reduced speed, when the source is silent, that is, when the user of the terminal 13 is listening to other participants. In the first case, the speech codec encodes speech data and transmits it in the payload of the RTP packet. In the latter case, the speech codec produces and transmits SID frames that carry an estimate of the background noise which is necessary for the generation of comfort noise in the receiver. In this case, this receiver is conference call server 12.
The RTP mixer 15 of the conference call server 12 decodes all incoming streams, 50 to enable a sum of the decoded speech and a combined encoding of the speech. Based on the data rate used respectively, the conference call server 12 obtains as secondary information of the decoding process, an indication as to whether the decoded signal is speech or an estimate of the background noise.
Thereafter, the RTP mixer 15 of the conference call server 12 mixes the encoded voice data and background noise estimates from all sources 13 together and assembles RTP packets with a combined and encoded data stream. Each assembled RTP packet comprises an RTP header having a structure which corresponds to the structure illustrated in Figure 2, which has already been described above. Thus, each RTP header comprises a field corresponding to an SSRC identifier and a field corresponding to a CSRC list.
The RTP mixer 15 inserts the SSRC identifier associated with the conference call server 12 for the current conference call into the SSRC identifier field of the RTP headers of the outgoing RTP packets, since the conference call server 12 is the source of these RTP packets.
On the other hand, the RTP mixer 15 includes in the CSRC list of the RTP headers the SSRC identifiers associated with those terminals 13 that contribute to the combined RTP packets. Since all terminals 13 participating in the conference call always transmit RTP packets towards the conference call server 12, either with voice data or with an estimate of the background noise, the CSRC list always comprises at least
ES 2 296 950 T3 both the SSRC identifiers corresponding to all participating terminals 13. However, the RTP mixer 15 ensures that the SSRC identifiers that are associated with the actively participating terminals 13 are included as first elements in the CSRC list .
Additionally, the RTP mixer 15 also inserts the SSRC identifier associated with the conference call server 12 into the CSRC list. More specifically, the SSRC identifier associated with the conference call server 12 is included as a marker between the SSRC identifiers associated with the active terminals 13 located at the beginning of the CSRC list and the SSRC identifiers associated with the inactive terminals 13 located at the end of the list. CSRC list.
The conference call server 12 then forwards the composite stream to each participating terminal.
13.
The terminals 13 receive the RTP packets transmitted by the conference call server 12 through the UMTS network 14 and retrieve the SSRC identifiers included in the respective CSRC list of the RTP packet headers. Next, based on the previously received mapping information, the terminals 13 determine the SIP addresses or telephone numbers corresponding to the SSRC identifiers retrieved from the CSRC list. Terminals 13 do not perform such a mapping for the SSRC identifier that is associated with the conference call server 12. This SSRC identifier is recognized by terminals 13 based on the identical SSRC identifier included in the SSRC identifier field of the RTP header. The terminals 13 further determine names that are associated, in their internal address directories, with the determined SIP addresses or telephone numbers, as they are available. The determined names are then presented to a respective user on the display of the terminals 13 in the form of a list.
Figure 3 shows an embodiment of said screen 31, which presents, together with other information and options, a list 32 with the names of users participating in a conference call in progress.
Additionally, the terminals 13 determine all those SSRC identifiers from the CSRC list that are presented before the SSRC identifier associated with the conference call server 12. The names that were determined for said SSRC identifiers belong to active participants at that time and are point in the displayed list 32 on screen 31. In the example of FIG. 3, a special speaker indicator icon 33 is used to indicate the participants who are currently speaking. In the situation presented, only one participant is speaking at that moment, and next to the corresponding name "Saimi", in list 32, an icon 35 is located indicating the speaker 33.
In this way, the user of a terminal 13 can always see an identification of all the users participating in the conference call, and differentiate the participants who are currently speaking from the inactive participants.
It should be understood that the described embodiment is only one of a variety of possible embodiments of the invention.
3 sheets
Sheet 1 Sheet 2 Sheet 3
22 members in 10 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 0202625 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 0202625 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 2002IB02625 | World Intellectual Property Organization (WIPO) | – | |
| 02741069 | – | – | – |
| WO2002IB02625 | – | – | – |
Members22
| Document | Office | Kind | |
|---|---|---|---|
| WO2004006475A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2002314458A1 | Australia | A1 | |
| AU2002314458A8 | Australia | A8 | |
| US2004076277A1 | United States of America | A1 | |
| KR20050013667A | Republic of Korea | A | |
| WO2004006475A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1547332A2 | European Patent Office (EPO) | A2 | |
| JP2005531999A | Japan | A | |
| CN1871825A | China | A | |
| WO2004006475A8 | World Intellectual Property Organization (WIPO) | A8 | |
| EP1547332B1 | European Patent Office (EPO) | B1 | |
| AT377314T | Austria | T | |
| ATE377314T1 | Austria | T1 | |
| DE60223292D1 | Germany | D1 | |
| JP4064964B2 | Japan | B2 | |
| ES2296950T3This record | Spain | T3 | |
| DE60223292T2 | Germany | T2 | |
| US7483400B2 | United States of America | B2 | |
| US2009109879A1 | United States of America | A1 | |
| KR100914949B1 | Republic of Korea | B1 | |
| CN100574287C | China | C | |
| US8169937B2 | United States of America | B2 |
Numbers
- Publication
- 2296950
- Publication, DOCDB
- 2296950
- Publication, EPODOC
- ES2296950T
- Application
- 2741069
- Application, DOCDB
- 02741069
- Application, EPODOC
- ES20020741069T
Titles2
- Spanish
- GESTION DE UNA LLAMADA EN CONFERENCIA POR CONMUTACION DE PAQUETES.
- English
- MANAGEMENT OF A CALL IN CONFERENCE FOR SWITCHING PACKAGES.
Classification
- CPC, 11
- H04L12/1822
- H04M3/56
- H04M3/569
- H04M7/006
- H04M2207/18
- H04L65/403
- H04L65/1104
- H04L65/65
- H04L12/18
- H04L2012/5603
- H04L65/1101
- IPC, 6
- H04L12 64
- H04M3 56
- H04L12 18
- H04L12 70
- H04L29 06
- H04M7 00