Active speaker identification
Abstract
FIELD: radio engineering, communication. SUBSTANCE: media server can order clients providing audio based on the input level. An identifier can be associated with the client for identifying the client providing input within the event. The ordered clients can be included in a list which can be inserted into a packet header carrying the audio content. EFFECT: high accuracy of identifying audio clients. 20 cl, 4 dwg
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
20 claims: 3 independent, 17 dependent
- 1A computer-implemented method for the identification of the audio input clients, the method comprising:receiving input from a first client specific audio input, wherein the input indicates that the first specific audio input client is an active participant in the first audio event, wherein the audio event comprises active and inactive participants;receive input from a second specific customer audio input, wherein the input indicates that the second audio in the particular client is the second active member in said audio event;a first identifier associated with the first customer-specific audio input and a second identifier of the second audio in a specific client;it is determined that the first concrete client audio input is dominant in the conversation to the second customer-specific audio input;ordering the first identifier of the second identifier in the list, based on the definition of what is an active participant in the first concrete client audio input is dominant in a conversation in relation to an active participant in the second customer-specific audio input, the first identifier is placed on top of the list, and the inactive participants do not included in the list, and insert the ordered list of the first identifier and the second identifier in the packet header. 1. Компьютерно-реализуемый способ идентификации клиентов аудиоввода, содержащий этапы, на которых принимают ввод от первого конкретного клиента аудиоввода, причем данный ввод указывает, что первый конкретный клиент аудиоввода является первым активным участником в аудиособытии, при этом аудиособытие содержит активных и неактивных участников;принимают ввод от второго конкретного клиента аудиоввода, причем данный ввод указывает, что второй конкретный клиент аудиоввода является вторым активным участником в упомянутом аудиособытии;связывают первый идентификатор с первым конкретным клиентом аудиоввода и второй идентификатор со вторым конкретным клиентом аудиоввода;определяют, что первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению ко второму конкретному клиенту аудиоввода;упорядочивают первый идентификатор над вторым идентификатором в списке, основываясь на определении того, что являющийся активным участником первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению к являющемуся активным участником второму конкретному клиенту аудиоввода, при этом первый идентификатор помещается на вершину списка, причем неактивные участники не включаются в список, и вставляют список с упорядоченными первым идентификатором и вторым идентификатором в заголовок пакета. 1. Компьютерно-реализуемый способ идентификации клиентов аудиоввода, содержащий этапы, на которых принимают ввод от первого конкретного клиента аудиоввода, причем данный ввод указывает, что первый конкретный клиент аудиоввода является первым активным участником в аудиособытии, при этом аудиособытие содержит активных и неактивных участников;принимают ввод от второго конкретного клиента аудиоввода, причем данный ввод указывает, что второй конкретный клиент аудиоввода является вторым активным участником в упомянутом аудиособытии;связывают первый идентификатор с первым конкретным клиентом аудиоввода и второй идентификатор со вторым конкретным клиентом аудиоввода;определяют, что первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению ко второму конкретному клиенту аудиоввода;упорядочивают первый идентификатор над вторым идентификатором в списке, основываясь на определении того, что являющийся активным участником первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению к являющемуся активным участником второму конкретному клиенту аудиоввода, при этом первый идентификатор помещается на вершину списка, причем неактивные участники не включаются в список, и вставляют список с упорядоченными первым идентификатором и вторым идентификатором в заголовок пакета.
- 10A computer readable medium comprising computer instructions that when executed by a computer sistemeprinimat require input from the first client specific audio input, wherein the input indicates that the first specific audio input client is an active participant in the first audio event, wherein the audio event comprises active and inactive participants;receive input from a second specific customer audio input, wherein the input indicates that the second audio in the particular client is the second active member in said audio event;associate the first identifier to the first customer-specific audio input and a second identifier of the second audio in a specific customer, with each particular client audio input associated canonical name (CNAME) and identification of the synchronization source (SSRC);determine that the first customer specific audio input is dominant in the conversation to the second customer-specific audio input;organize a list of one or more active customer specific audio input in the conference, based on the definition of what is an active participant in the first concrete client audio input is dominant in a conversation in relation to an active participant in the second customer-specific audio input, so that the first identifier is located on the second identifier, wherein the first identifier is placed on top of the list, and inactive participants are not included in the list, and to insert an ordered list in the list is the source of the real-time transport protocol (RTP) in one or more audio streams. 10. Машиночитаемый носитель, содержащий машиноисполняемые команды, которые при их исполнении предписывают компьютерной системепринимать ввод от первого конкретного клиента аудиоввода, причем данный ввод указывает, что первый конкретный клиент аудиоввода является первым активным участником в аудиособытии, при этом аудиособытие содержит активных и неактивных участников;принимать ввод от второго конкретного клиента аудиоввода, причем данный ввод указывает, что второй конкретный клиент аудиоввода является вторым активным участником в упомянутом аудиособытии;связывать первый идентификатор с первым конкретным клиентом аудиоввода и второй идентификатор со вторым конкретным клиентом аудиоввода, при этом с каждым конкретным клиентом аудиоввода связывается каноническое имя (CNAME) и идентификация источника синхронизации (SSRC);определять, что первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению ко второму конкретному клиенту аудиоввода;упорядочивать список из одного или более активных конкретных клиентов аудиоввода в конференции, основываясь на определении того, что являющийся активным участником первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению к являющемуся активным участником второму конкретному клиенту аудиоввода, так что первый идентификатор располагается над вторым идентификатором, при этом первый идентификатор помещается на вершину списка, причем неактивные участники не включаются в список, и вставлять упорядоченный список в поле списка составляющих источников транспортного протокола реального времени (RTP) в одном или более аудиопотоках. 10. Машиночитаемый носитель, содержащий машиноисполняемые команды, которые при их исполнении предписывают компьютерной системепринимать ввод от первого конкретного клиента аудиоввода, причем данный ввод указывает, что первый конкретный клиент аудиоввода является первым активным участником в аудиособытии, при этом аудиособытие содержит активных и неактивных участников;принимать ввод от второго конкретного клиента аудиоввода, причем данный ввод указывает, что второй конкретный клиент аудиоввода является вторым активным участником в упомянутом аудиособытии;связывать первый идентификатор с первым конкретным клиентом аудиоввода и второй идентификатор со вторым конкретным клиентом аудиоввода, при этом с каждым конкретным клиентом аудиоввода связывается каноническое имя (CNAME) и идентификация источника синхронизации (SSRC);определять, что первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению ко второму конкретному клиенту аудиоввода;упорядочивать список из одного или более активных конкретных клиентов аудиоввода в конференции, основываясь на определении того, что являющийся активным участником первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению к являющемуся активным участником второму конкретному клиенту аудиоввода, так что первый идентификатор располагается над вторым идентификатором, при этом первый идентификатор помещается на вершину списка, причем неактивные участники не включаются в список, и вставлять упорядоченный список в поле списка составляющих источников транспортного протокола реального времени (RTP) в одном или более аудиопотоках.
- 17A computer system comprising a media server configured to accept input from a first client specific audio input, wherein the input indicates that the first specific audio input client is an active participant in the first audio event, wherein the audio event comprises active and inactive participants;receive input from a second specific customer audio input, wherein the input indicates that the second audio in the particular client is the second active member in said audio event;associate the first identifier to the first customer-specific audio input and a second identifier of the second audio in a specific client;determine that the first customer specific audio input is dominant in the conversation to the second customer-specific audio input;organize the first identifier of the second identifier in the list, based on the definition of what is an active participant in the first concrete client audio input is dominant in a conversation in relation to an active participant in the second customer-specific audio input, the first identifier is placed on top of the list, and the inactive participants do not included in the list, and to insert the ordered list of the first identifier and the second identifier in the packet header. 17. Компьютерная система, содержащая медиасервер, выполненный с возможностью принимать ввод от первого конкретного клиента аудиоввода, причем данный ввод указывает, что первый конкретный клиент аудиоввода является первым активным участником в аудиособытии, при этом аудиособытие содержит активных и неактивных участников;принимать ввод от второго конкретного клиента аудиоввода, причем данный ввод указывает, что второй конкретный клиент аудиоввода является вторым активным участником в упомянутом аудиособытии;связывать первый идентификатор с первым конкретным клиентом аудиоввода и второй идентификатор со вторым конкретным клиентом аудиоввода;определять, что первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению ко второму конкретному клиенту аудиоввода;упорядочивать первый идентификатор над вторым идентификатором в списке, основываясь на определении того, что являющийся активным участником первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению к являющемуся активным участником второму конкретному клиенту аудиоввода, при этом первый идентификатор помещается на вершину списка, причем неактивные участники не включаются в список, и вставлять список с упорядоченными первым идентификатором и вторым идентификатором в заголовок пакета. 17. Компьютерная система, содержащая медиасервер, выполненный с возможностью принимать ввод от первого конкретного клиента аудиоввода, причем данный ввод указывает, что первый конкретный клиент аудиоввода является первым активным участником в аудиособытии, при этом аудиособытие содержит активных и неактивных участников;принимать ввод от второго конкретного клиента аудиоввода, причем данный ввод указывает, что второй конкретный клиент аудиоввода является вторым активным участником в упомянутом аудиособытии;связывать первый идентификатор с первым конкретным клиентом аудиоввода и второй идентификатор со вторым конкретным клиентом аудиоввода;определять, что первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению ко второму конкретному клиенту аудиоввода;упорядочивать первый идентификатор над вторым идентификатором в списке, основываясь на определении того, что являющийся активным участником первый конкретный клиент аудиоввода является доминирующим в разговоре по отношению к являющемуся активным участником второму конкретному клиенту аудиоввода, при этом первый идентификатор помещается на вершину списка, причем неактивные участники не включаются в список, и вставлять список с упорядоченными первым идентификатором и вторым идентификатором в заголовок пакета.
Independent claims3
48 paragraphs in 4 sections, as filed
BACKGROUND
Media conference participants may have difficulty identifying other participants. Participant may be unfamiliar to the speaker's voice or face party or audio exchange may interfere with the listener. In the latter case, the listener or the speaker or not, can be confused when several participants speak simultaneously, or if there is a rapid exchange between the numerous parties. In some cases, participants may speak to attach his / her name, "Bob is ..." or student may request the identity of the previous speaker. The complexity of this problem may increase with the number of speakers or the audio input of contributing members. Despite the fact that the listener can retrieve the identity of the speaker of the "context clues" during the conversation, in some cases, participants may not know which participants provide audio input.
Furthermore, it may be desirable to minimize bandwidth consumption or the amount of bandwidth to transfer information. For example, despite the fact that the physical connection for transporting data may have additional throughput, consuming communication link resources may reduce available for other data transmission capacity, or may affect the transmission of data of an audio conference, if a user is fortunate enough to have limited network bandwidth.
Admissibility improvements media conference may be limited if the improvement is not "backward compatible." For example, if the change is not compatible with existing protocols and versions, users may need to obtain an updated version to communicate with the party that implements a modified version and / or obtain the approval of the organizations. The above situation may hinder the acceptability of the modified technology.
SUMMARY OF THE INVENTION
It describes procedures for identifying clients in an audio or audio / video event. As an example, a media server may order clients providing audio based on the input level. The identifier may be associated with the client for identifying the client providing the audio input within the event. The ordered clients may be included in the list, which can be inserted into a packet header carrying the audio content.
This Summary is provided to introduce the option selecting concepts in a simplified form that are further described below in the description of the preferred embodiments. This summary is not intended to identify key features or essential features of the claimed subject matter as it is not intended to be used as an aid in determining the scope of the claimed subject matter.
BRIEF DESCRIPTION OF DRAWINGS
Description of the preferred embodiments described with reference to the accompanying drawings. In the drawings, the leftmost digit reference number identifies the drawing in which the reference number is found for the first time. Using the same reference numbers in different instances in the description and the figures may indicate similar or identical elements.
Figure 1 - Illustration of the environment in an exemplary implementation in which technology can be used, providing the ability to identify the active speaker participant.
2 - a diagram depicting the data packet RTP comprising ordered / reordered list of active clients in the form of a list in the component sources (CSRC).
Figure 3 - a block diagram depicting a procedure in an exemplary implementation for identifying active clients.
4 - block diagram depicting a procedure in an exemplary implementation for identifying active clients within a conference real time protocol.
DESCRIPTION OF PREFERRED EMBODIMENTS
Overview
Describes methods for identifying active audio conference participants in the media event. The list of implementations facilitate or participate Audio customers may be organized on the basis of the client's contribution to the work of the session. The ID can be associated with participating customers so that customers can identify which customers to actively contribute to the event. Organized list may be inserted into the packet header for transmitting streaming data to clients of the conference. In implementations, the identification information may be included in control packets used in conjunction with data transport. The techniques discussed here can provide information about the current speaker, consuming minimal network resources and without increasing the synchronization problems.
In further embodiments, the media server to switch / mix audio streams can be configured to insert an ordered list of active clients in the packet header data. For example, a media server may include a list of active speakers, which can be ordered based on the currently active participant speaking, so that customers are provided with information about what clients are actively spoken. The list can be provided without increasing the overhead of the network to transport media.
Exemplary environment
1 - Vector environment 100 in an exemplary implementation that is configured to use active speaker identification. For example, media server 102 may identify active audio clients in spite of the mixing and switching between clients providing audio streams in a media event. Although considered processing audio data, the media server 102 may handle other type of media data including video and so on, based on the conference and the capabilities of the client devices. For example, media server 102 may manipulate audio data / video for some clients, sending audio data to clients who do not have video capability and so on.
For example, the processor 104 of the media server may determine which client or clients are actively providing audio content while mixing / switching audio streams for clients. The processor 104 of the media server may determine which clients are actively administered audio data based on the algorithm / techniques mixing / switching involved processor for generating media streams sent. The definition can be used to sort the list of customers, contributing to coming from a Media Server 102 media streams, or customers who have contributed to the withdrawal of a media server.
For audio events includes client 106 "A", 108 "B", 110 "C", 112 "D" and 114 "E", in which the client 106 "A" and 114 "E" contribute to the audio input ( Because the client 106, and E 114 continues the conversation), inactive clients 108 "B", 110 "C", 112 "D" can be given sent from media server 102 Stream "A + E", or a combination of the two speakers at the While client 106 "A" and 114 "E", respectively, taking the opposite side sent from media server 102 stream (for example, client 106 "A" takes sent by the client "E" stream, while the client 114 "E" - posted Client 106 "A" stream). Applicable client devices include, but are not limited to, phone with voice over IP-protocol (VoIP), the computing device having audio, phone switched telephone network (PSTN), phones connected through a gateway to the digital audio session, , and so on.
In some implementations, an active speaker may not be provided a signal comprising a stream sent by the speakers to avoid feedback or an echo (e.g., Client 106 "A" may not be sent an audio stream containing audio Customer "A"). You can see a few basic scenarios of identification, for example, Customer A can "talk" Client E (as if associated with the client 106 A participant says loudly, while Member "E" (associated with the client 114 E) says on a normal voice ), members of "A" and "E" are engaged in a quick exchange in which the speaker is now changes between the two participants, or Participant "A" predominant in the conversation, while Participating "E" provides relatively less input information. An example of the latter situation may include a member who adds a little to confirm the prevailing monologue of the main speaker.
In embodiments of the media server 102 may determine the dominant client (and thus a speaker) based on the number of packets received from the client, when audio content is received, packet size, energy audio level and so on. Thus, although two or more client provide content simultaneously, one active client may be designated as the dominant client (and thus speaker) based on the above factors. For example, media server 102 may determine the current active client (and associated speaker) based on including the audio content of the current packet of data received from the active client in conjunction with mixing and / or switching between the inputs received from different clients . For example, media server 102 may designate Client A 106 as the current "active" client, if Client E is not currently offered data packets. In other cases, if the client 106 A and Client 114 E are active, but the client 106 A provides audio content with high level of energy than the client 114 E (ie, participant A speaks loudly, while the E says quietly) Client A 106 can be designated as a dominant active speaker. Customers can be provided with a list of active clients, which begins with a client 106 A. This type of determination can be made during the mixing / switching audio streams client input for one or more of ongoing conferences. For example, a media server processor 102 may differentiate between the active clients cycling algorithm mixing, while the identification module 116 may be used to insert data into appropriate data packets.
With primary reference to Figure 2, in implementations, when implementing the transport layer protocol in real time (RTP) protocol and associated real-time control (RTCP), the media server 102 may identify active clients, and thus active speakers, by examining data under streams sent from the clients, including data and signaling are transported flows. If the client 106 A media server 102 may identify that the sent client audio stream coming from the client 106 A, through the study of the field synchronization source (SSRC) within the RTP packet or from the SSRC Client (identifier for the client within the session) and the canonical name (CNAME), included in the report RTCP. It may also be examined other information. SSRC also may be obtained from the RTP packet header. For example, SSRC may be aligned with the CNAME in a RTCP report.
While RTCP signaling may be used to identify missing packets, providing data transport quality and so on, the RTCP report may be obtained from the RTCP-band signals. For example, the RTCP report may include the randomly generated client SSRC, aligned to the client CNAME. Typically, CNAME is an identifier / record which is associated with the aliases used for the client device. In some cases, CNAME - a string of numbers, etc. In embodiments of the media server 102 may be assigned a SSRC within the session. In some cases, SSRC may change for a client included in a session. For example, the client SSRC may change if the client is interrupted (for example, a long pause, and then the answer) if customers SSRC conflict (there is more than one client with a common SSRC), and so on. Thus, the input data stream can be identified by SSRC in the data stream or RTCP signaling. Media server 102 may also receive signaling of RTCP canonical name for use in identifying the client.
When forming send stream (including an audio output) for SSRC and CNAME, obtained from the active client, media server 102 may identify which clients are providing audio input to the session. For example, media server 102 may associate the SSRC, CNAME, inserted into the RTCP packet c you send audio content stream (i.e., output of the media server streams carrying audio data). Returning to the previous example, the session between the client 106 "A", 108 "B", 110 "C", 112 "D" and 114 "E", in the case of mixed-signal "A + E", the media server 102 can streamline customers' A "and" E ", in accordance with which the client is currently active, which one is active and dominating in the session, and the like. The procedure may vary based on the customer, providing the audio input. In this case, the list can start with customer ID 106 A and includes client 114 E, if Client A is currently providing input, or if Client A dominates the conversation. In situations in which there is audio communication between the Client and the Client A 106 114 E, pryadok can be changed based on the currently speaking participant, as specified on a per-packet basis.
With reference to Figure 2, in a RTP configuration, the identification module 116 of the media server may insert the ordered list of SSRC of the RTP packet header to the output stream. For example, the ordered identifiers are inserted (204) in the list is the source (CSRC) in the header sent in the data stream packet. If Client A and Client E changes the current active roles, the location of the SSRC may vary from 204 (a) "Client A, Client E ..." to the 204 (b) "Clients E, Client A ...". In the previous method clients receiving the data stream (listening to clients or participants in the session), they can be informed as to what customers provide input relative contribution and so on, while avoiding additional signaling associated with the problems of synchronization and network overhead. For example, the field CSRC may be able to include up to fifteen identifiers thirty two bits each, while remaining in compliance with the technical requirements. Customers not operating in accordance with the techniques discussed here may participate, without having the advantages discussed herein. Thus creating a backward-compatible system and methods.
Despite the fact that the SSRC may identify an active customer, the use of SSRC can be problematic because SSRC may be assigned randomly, may vary due to a conflict with another client with similar SSRC, the customer is reassigned SSRC after dropping out of the session and then return to the session , and so on.
Media server 102 may insert a CNAME active client sends the client packages RTCP (for example, so that others "listen" to customers could be made aware of the CNAME and SSRC active clients). For example, the identification module 116 media server can "distribute" active client identifiers sent to "listen to customers" in the RTCP packet media server. For example, if several active clients are contributing to a conference, the media server may insert the obtained client identifiers with determinable intervals RTCP packets, sent along with data stream media server. While RTCP packets may include the CNAME in each packet. To minimize transport overhead, CNAME may be inserted into the gaps listening clients send RTCP packets. Clients receiving data RTCP media server comprising active client identifiers, may store the data in local memory so that the CNAME may be associated with data packets when the audio content is received. For example, CNAME, delivered in accordance with the SSRC, and other related information can be stored in a lookup table, etc. For example, while the data stream included in the audio content in many cases may be sent continuously, RTCP signaling may occur only intermittently, for example, c-defined intervals (e.g., 5 or 10 second intervals). Thus, a client receiving a data packet may associate a SSRC in the CSRC with a previously received CNAME. In implementations in order to identify a particular client can use a universal resource indicator globally routable user agent (GRUU).
The implementation of an active client may be informed that the Conference of the client is active. For example, the participant (associated with the active customer) may wish to know that he / she is not "talking" another party. Returning to the session between the client 106 is "A", 108 "B", 110 "C", 112 "D" and 114 "E", when, for example, the client 106 and the asset, and the client 108 "B", 110 "C" 112 "D" 114 and "E" is not active, it can be determined by Client A signal sent RTCP. Thus, along with the fact that the media server 102 may form sent to clients "B", "C", "D" and "E" media stream by passing the flows sent through the client A, as a "listener" customer Participants of the session or client A 106 may determine that no other active customer based on packages CSRC / RTCP.
In further embodiments, human-readable information may be associated with the SSRC and CNAME. For example, the user may wish to image the speaker's party was displayed on the connected monitor when the participant said. The customer implementations can be shared human-readable information about the client. For example, the data usually can be exchanged at the beginning of an event or session.
Along with the fact that the Internet (Wide Area Network) can be used for connecting clients and other components, other networks and various links are also applicable. For example, connecting the media server 102 and the client network may include a wide area network (WAN), local area network (LAN), a wireless network, a public telephone network, an intranet, and so on. The network may be configured to include a plurality of subnets.
Following discussion describes techniques that may be implemented using the previously described systems and devices. Aspects of each of the procedures may be implemented in hardware, firmware, software, or a combination thereof. The procedures are shown as a set of blocks that specify executed by one or more transaction devices, and not necessarily limited to the sequences shown respective units to perform operations.
Exemplary Procedure
Following discussion describes techniques that may be implemented using the previously described systems and devices. Aspects of each of the procedures may be implemented in hardware, firmware or software, or a combination thereof. The procedures are shown as a set of blocks that specify executed by one or more transaction devices, and not necessarily limited to the sequences shown respective units to perform operations. It also assumes the existence of many other examples.
Figure 3 shows exemplary procedures for identifying active audio input clients in media sessions. For example, the technique can be used in conference or a media conference in which some clients are missing the functionality of video and so on.
In embodiments used as a host or central point of the media server can determine (302) the audio input client in accordance with the input provided by each active client. For example, the determination may be made as part of mixing and / or switching audio client input. Thus, client A may be designated as the main active customer until the customer does not provide a different audio input. In another example, Client A may be selected if Client A and Client E are contributing but Client A audio has a higher energy level. Audio customer may have a higher energy level, if associated with the client party said in a loud voice or say in a non-stop manner, as if the client provides the dominant audio input.
Introducing audio client may be identified as "high" the client if the client is currently active, it dominates the conversation and so on. In systems RTP / RTCP, operating in accordance with these procedures, the media server may receive (304) flows of client input and associated packages RTCP (for example, packets RTCP, sent from the client), including SSRC, delivered in accordance with the CNAME for specific client, generates a stream including the audio content. For example, a media server may receive the SSRC and CNAME for the client. CNAME identifying the client in conjunction with the SSRC. Media Server can streamline (306) SSRC customer input in line with what customers currently provide audio input to dominate the conversation and so on. For example, the media server can streamline the SSRC identifier of active clients in descending order, starting with the current active "speaker", such as the customer provides the active input. In the examples, RTP may allow the identification of 15 active speakers participants included in the CRSC using thirty-two bit identifier per active client.
The media server may associate an identifier with the audio input client. For example, a media server may receive the SSRC and CNAME from an audio RTCP packet client-input. SSRC may be used to identify the audio input client in the CRSC field, included in the output stream media server.
Customers can receive / other data associated with the client audio input. For example, the receiving client (listen to the customer or client in the media event) may have a human-readable information associated with a CNAME. For example, the client may be an image member, the participant's name and so on (which is associated with the CNAME / SSRC client).
Ordered client identifiers audio input can be inserted (308) in a list in the packet header. For example, if clients "A" and "E" provide audio input (Customer A - the currently active client), CSRC field in the RTP header may include sources from the SSRC SSRC Client "A", the novice list. Thus, the listener to the client (which may include an audio input client which receives audio input from another active client) within a content stream may be communicated to the identification information of the speaker. In another example, the order of injecting the audio clients in the list may be based, at least in part, on what the client audio input dominates in the media session. Dominance factors may include energy level of the audio input, duration of the input, duration of silence periods, packet size and so on. For example, the list may commence with Client A because Client A is currently active and the Client A stream sent shows a high level of energy in comparison with one or more other audio input clients.
The media server may send (310) the listening client (session) SSRC and CNAME in the media server send streams (e.g., in RTCP packets, sent together with transported content). SSRC for the audio input clients may also be located the CSRC field in the packet header streaming data in RTP packets. For example, if five client media event, three participants say, the client SSRC and CNAME associated with the audio input clients may be included in RTCP packets media server (sent listening clients) associated with transmitting audio content RTP packets. Thus, the media server may send a SSRC and CNAME customer identifying the active audio clients. Thus, a listening client may identify the original source of audio content in accordance with the source SSRC and CNAME in the names of RTP packet. SSRC client can be updated if the client SSRC conflict with SSRC, issued by another client, or if the customer changes the address of the transmission source for any other reason. Sources SSRC and CNAME names can be stored (312) in a local memory so that a listening client may access the information throughout the media event.
Figure 4 shows an exemplary procedure for identifying active clients in a media conference. For example, these techniques may be used during a media conference in which some clients are missing the functionality of the video or may be used in audio conferencing.
In these implementations, a media server may receive (402) an active client input (audio content) as well as identifiers from the active clients. For example, contributing to an audio conference may send a client SSRC and CNAME, which identifies this client. For example, SSRC may be included in the data stream packet header in RTP and RTCP packet along with the CNAME.
It may be generated (404) an ordered list of one or more active clients within a conference. For example, mixing / switching audio input stream media server can sort the list of active clients (SSRC identifier in the RTP / RTCP), or those customers who provide input in the conference or meeting. For example, a media server mixing audio / video server (AVMCU), which receives the SSRC identification from an active client sent flow which in turn may include some data and an associated signaling portion. Then AVMCU can determine the relative location of the SSRC of active clients or other client identifier within the session. SSRC may be identified from a RTCP report, which may correspond to a client CNAME. For example, the ranking may be based on what the client is currently active. In other implementations, factors such as the energy level, the number of data packets provided, the duration of silence periods, packet size and so on, can be taken into account. For example, an ordered list can begin with an active customer, who can dominate the session because of the amount granted to a package while the second active at the same time the client is assigned to a smaller relative status.
The ordered list may be inserted (406) in the CSRC list, included in the packet headers, the sequence of operations within the media server send data stream. For example, the output of the media server includes providing active customer audio field CSRC with an ordered list of active clients SSRC identifier. As a result, we listen to the customer, that is, the client receives a stream of audio content can be reported to which customers are active, and the comparative relationship of active clients. Furthermore, SSRC and the CNAME may be included in the sent media server RTCP packets.
SSRC may be associated (408) with the CNAME of the active audio client. For example, a media server may send a RTCP packet, which includes the client CNAME related to SSRC, included in the CRSC field in the RTCP packet header. CNAME may be obtained from the RTCP packets.
Human readable information may be associated with the CNAME and / or also with the client SSRC audio input. For example, the image or the name may be associated with a CNAME client so that the image or name to appear when the client provides the associated audio content. This information may be reported at the conference, or the client can enter the human-readable information.
In further embodiments, GRUU may be associated with the SSRC for the active client. In some situations in which a client is active, but other clients are not active, the media server can provide (410) to the active client designation so that the active client is notified that no other active yet, despite the fact that the administration of active customer flow It does not return to active client. Thus, the active customer becomes aware that a member is not "talking" another party.
While RTP and RTCP are considered, and methods of the present invention may be applied to other protocols data transport algorithms.
Conclusion
Although the present invention has been described in a particular formulation structural features and / or methodological acts, it is to be understood that the present invention defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claimed invention.
Contents4
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| RU2734895C2 | Cited by | Russian Federation | Search report |
| EP1551205A1 | Cites | European Patent Office (EPO) | Search report |
| US2004076277A1 | Cites | United States of America | Search report |
| US2005025073A1 | Cites | United States of America | Search report |
| RU2006101325A | Cites | Russian Federation | Search report |
| US2004076277A1 | Cites | United States of America | – |
| RU2006101325A | Cites | Russian Federation | – |
| US2005025073A1 | Cites | United States of America | – |
18 members in 8 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 11761963 | United States of America | – | |
| 76196307 | United States of America | A | |
| 2008065441 | United States of America | W | |
| 11761963 | – | – | – |
| US2008065441 | – | – | – |
| US20070761963 | – | – | – |
| WO2008US65441 | – | – | – |
Members18
| Document | Office | Kind | |
|---|---|---|---|
| US2008312923A1 | United States of America | A1 | |
| WO2008157005A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20100021435A | Republic of Korea | A | |
| EP2163035A1 | European Patent Office (EPO) | A1 | |
| CN101689998A | China | A | |
| JP2010529814A | Japan | A | |
| RU2009146029A | Russian Federation | A | |
| EP2163035A4 | European Patent Office (EPO) | A4 | |
| US8385233B2 | United States of America | B2 | |
| RU2483452C2This record | Russian Federation | C2 | |
| US2013138740A1 | United States of America | A1 | |
| EP2163035B1 | European Patent Office (EPO) | B1 | |
| US8717949B2 | United States of America | B2 | |
| US2014177482A1 | United States of America | A1 | |
| JP5579598B2 | Japan | B2 | |
| BRPI0812128A2 | Brazil | A2 | |
| KR101486607B1 | Republic of Korea | B1 | |
| US9160775B2 | United States of America | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Official registration of the transfer of exclusive rightPC41 | PC41 |
Numbers
- Publication
- 0002483452
- Publication, DOCDB
- 2483452
- Publication, EPODOC
- RU2483452
- Application
- 200914602907
- Application, DOCDB
- 2009146029
- Application, EPODOC
- RU20090146029
Titles2
- English
- ACTIVE SPEAKER IDENTIFICATION
- Russian
- ????????????? ????????? ?????????? ?????????
Classification
- CPC, 7
- H04L65/403
- H04L65/1069
- H04L65/4038
- H04L65/608
- H04M3/569
- H04M7/006
- H04M2203/5072