Distributed real-time media composer
Abstract
System that allows the simultaneous exchange of audio, video and / or data information between a plurality of endpoints in a communications network and supported by a central unit, the plurality of endpoints comprising transmitting endpoints and receiving endpoints, in the that the system is adapted for the bi-directional exchange of multimedia streams between the central unit and the plurality of endpoints, The system is adapted to compose multimedia streams comprising partial frames, and the central unit is, based on the capacity information received from one or more of the endpoints, adapted to instruct said one or more transmitting endpoints to generate multimedia streams. which comprise partial frames adjusted to match the capabilities of the receiving endpoints participating in a session.

Term
Term ended
Projected expiry passed 15 November 2024, 1.9 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
18 claims: 10 independent, 8 dependent
- 1ES 2 336 216 T3 ES 2 336 216 T3 CLAIMS REIVINDICACIONES 1. System that allows the simultaneous exchange of audio, video and / or data information between a plurality of end points in a communications network and supported by a central unit, the plurality of end points comprising transmitting end points and receiving end points, in the that the system is adapted for the bidirectional exchange of multimedia streams between the central unit and the plurality of end points, the system is adapted to compose multimedia streams comprising partial frames, and the central unit is, based on the capacity information received from one or more of the endpoints, adapted to instruct said one or more transmitting endpoints to generate multimedia streams comprising partial frames adjusted to correspond to the capabilities of the receiving endpoints participating in a session. 1. Sistema que permite el intercambio simultáneo de informaciones de audio, video y/o datos entre una pluralidad de puntos extremos en una red de comunicaciones y soportado por una unidad central, la pluralidad de puntos extremos comprendiendo puntos extremos transmisores y puntos extremos receptores, en el que el sistema está adaptado para el intercambio bidireccional de flujos multimedia entre la unidad central y la pluralidad de puntos extremos, el sistema está adaptado para componer flujos multimedia comprendiendo tramas parciales, y la unidad central está, basándose en la información de capacidad recibida desde uno o más de los puntos extremos, adaptada para dar instrucciones a dicho uno o más puntos extremos transmisores para generar flujos multimedia que comprenden tramas parciales ajustadas para corresponderse con las capacidades de los puntos extremos receptores que participan en una sesión.
- 2System according to claim 2. Sistema según la reivindicación 1, caracterizado por el hecho de que la información de capacidad es una de las siguientes:formatos de visualización, anchos de banda de transmisión, requisitos de procesamiento o múltiples combinaciones de los mismos. 1, characterized by the fact that the capacity information is one of the following: display formats, transmission bandwidths, processing requirements, or multiple combinations thereof.
- 3System according to any of claims 1 or 2, characterized in that the central unit is adapted to compose a combined data stream from one or more single data streams sent from one or more transmitting endpoints and to route said combined stream to the receiving endpoints. 3. Sistema según cualquiera de las reivindicaciones 1 ó 2, caracterizado por el hecho de que la unidad central está adaptada para componer un flujo de datos combinados desde uno o más flujos de datos únicos enviados desde uno o más puntos extremos transmisores y para encaminar dicho flujo combinado a los puntos extremos receptores.
- 5System according to any of claims 1-4, characterized in that the communication between said central unit and the plurality of end points uses scalable compression techniques. 5. Sistema según cualquiera de las reivindicaciones 1-4, caracterizado por el hecho de que la comunicación entre dicha unidad central y la pluralidad de puntos extremos utiliza técnicas de compresión escalables.
- 7System according to any of claims 1-6, characterized in that a display, included in said plurality of end points, is divided into a grid of cells, where each said partial frame occupies one or more cells of the grid and the display can have multiple aspect ratios. 7. Sistema según cualquiera de las reivindicaciones 1-6, caracterizado por el hecho de que una visualización, incluida en dicha pluralidad de puntos extremos, está dividida en una cuadrícula de celdas, donde cada dicha trama parcial ocupa una o más celdas de la cuadrícula y la visualización puede tener varias relaciones de aspecto. ES 2 336 216 T3 ES 2 336 216 T3
- 10Procedure for simultaneously exchanging audio, video and / or data information between a plurality of end points in a communications network supported by a central unit, the plurality of end points comprising transmitting end points and receiving end points, characterized in that gives instructions to one or more of the transmitting endpoints, based on the capacity information received from one or more endpoints, to generate multimedia streams comprising partial frames adjusted to correspond to the capabilities of the receiving endpoints participating in a session;and by the bidirectional exchange of multimedia streams between the central unit and the end points, where said multimedia streams are made up of partial frames. 10. Procedimiento para intercambiar simultáneamente informaciones de audio, video y/o datos entre una pluralidad de puntos extremos en una red de comunicaciones soportado por una unidad central, la pluralidad de puntos extremos comprendiendo puntos extremos transmisores y puntos extremos receptores, caracterizado por el hecho de que da instrucciones a uno o más de los puntos extremos transmisores, basándose en la información de capacidad recibida desde uno o más puntos extremos, para generar flujos multimedia comprendiendo tramas parciales ajustadas para corresponderse con las capacidades de los puntos extremos receptores que participan en una sesión;y por el intercambio bidireccional de flujos multimedia entre la unidad central y los puntos extremos, donde dichos flujos multimedia están compuestos de tramas parciales.
- 14Method according to any of claims 10-13, characterized by the use of compression techniques during the data exchange between said central unit and the plurality of end points. 14. Procedimiento según cualquiera de las reivindicaciones 10-13, caracterizado por la utilización de técnicas de compresión durante el intercambio de datos entre dicha unidad central y la pluralidad de puntos extremos. ES 2 336 216 T3 ES 2 336 216 T3
- 15Procedimiento según cualquiera de las reivindicaciones 10-14, caracterizado por la división de una visualización en una cuadrícula de celdas, donde cada dicha trama parcial ocupa una o más celdas de la cuadrícula y la visualización puede tener varias relaciones de aspecto. fifteen. Method according to any one of claims 10-14, characterized by dividing a display into a grid of cells, where each said partial frame occupies one or more cells of the grid and the display may have several aspect ratios.
Independent claims10
163 paragraphs in 20 sections, as filed
ES 2 336 216 T3
DESCRIPTION
Composer of distributed media in real time.
Field of the invention
The present invention relates to systems that allow the simultaneous exchange of audio, video and data information through the use of telecommunications, in particular it relates to online videoconferencing and conference systems.
Background of the invention
In particular, the invention describes a system and a method that allow the simultaneous exchange of audio, video and data information between a plurality of units, using an existing telecommunications network.
There are a number of technological systems available to organize meetings between participants located in different areas. These systems may include multipoint audiovisual conferencing or video conferencing, online conferencing, and audio conferencing.
The most realistic substitute for real meetings is high-end video conferencing systems. Conventional video conferencing systems include a number of endpoints that communicate video, audio, and / or data streams in real time over and between various networks such as WAN, LAN, and circuit-switched networks. Endpoints include one or more display (s), camera (s), microphone (s), and / or data capture device (s), and a codec. These codecs encode and decode the incoming and outgoing streams, respectively.
Multimedia conferences can be divided into three main categories: centralized, decentralized and hybrid conferences, in which each category has a plurality of variations for running a conference.
Centralized conference
Traditional audiovisual multipoint conferences have a Multipoint Control Unit (MCU) with three or more connected end points. These MCUs perform switching functions to allow audiovisual terminals to intercommunicate in a conference. The central function of an MCU is to link multiple video teleconference sites (EP-endpoints) and also receive frames of digital signals from audiovisual terminals (EP), process the received signals, and retransmit the processed signals to the appropriate audiovisual terminals ( EP) as digital signal frames. Digital signals can include audio, video, data, and control information. Video signals from two or more audiovisual terminals (EPs) can be spatially mixed to form a composite video signal for viewing by teleconference participants. The MCU acts as a selective media stream router in this scenario. A part of the MCU called the Multipoint Controller (MC) controls the conference. Each endpoint has a control channel to send and receive control signals to and from the MC. The Mc acts and sends commands to the end points.
Unique voice switching flow
In a centralized conference, the MCU will receive incoming video streams from all participants. You can relay a video stream from one endpoint to all other endpoints. The endpoint stream is typically selected by the single switched voice stream solution, based on the loudest speaker, ie, the speaker. The flow is called Current View. Whereas Previous View is the video stream coming from the participant at the extreme point that was the speaker before the current speaker. In a voice switched conference, a Current View video stream is sent to everyone except the current speaker, and the Previous View is sent to the current speaker. One problem for the MCU is ensuring that all conference endpoints can receive the Current View and the Previous View.
Single flow of switching by other means
The Current View can also be controlled by sending commands between the MCU and the endpoints. Such a mechanism is called speaking turn control. An endpoint can send a floor request command to the MCU so that its video is sent to all other participants. The Previous View will typically be a voice switching view among the other participants in the conference. The Current View can be issued by sending a speaking-time broadcast command. There are other known procedures to control the Current View, as well as, among others, the control of the speaking turn and the coordination control. They both deal with single switching flows, however providing a full description of each known workaround on this matter is outside the scope of this document. However, the principle with a current view and single stream switching is the same.
ES 2 336 216 T3
Continuous presence
In a conference, you will often see more than one participant. This is accomplished in a number of ways. The MCU can combine the incoming video streams to make one or more output video streams achieve this. Combining multiple incoming low resolution video streams from the extreme points to one high resolution stream can do it. The high resolution stream is then sent from the MCU to some or all of the endpoints in the conference. The flow is called Combined View. The characteristic of low resolution streams limits the format of the high resolution stream from the MCU. Strict limitations on incoming low resolution streams are needed to ensure that all receiving endpoints can receive the combined high resolution stream. The MCU must, as long as each receiver receives the same multimedia stream, find the "least common mode" to ensure acceptable audio and vision performance at the receiver with the worst capability. Due to various screen variations, the MCU must also compensate for different screens such as 4: 3 16: 9 view; this is not possible with a common mode. This less common mode solution does not work particularly well and provides significant limitations to receivers that have a capacity that exceeds those with the worst capacity.
Resized view
A more flexible solution is to let the MCU resize all incoming video streams and thus make a view that is receivable at all endpoints that receive it. To perform resizing, the MCU needs to decode all incoming video streams. The decoded data - raw data - is then resized and transformed. The different raw data streams are combined into a composite layout and put together into a given set layout, tailored to the receiver's requirements for bit rate and encoding standard. The combined raw data stream is then encoded and you will have a new video stream containing one or more incoming streams. The solution is called Resized View. To make a resized view, the MCU must understand and have the ability to encode and decode video streams. The more endpoints in the conference, the more capacity the MCU will need to decode all incoming streams. The heavy data manipulation performed by the MCU will add additional delay to the multimedia streams and thus reduce the quality of the multimedia conference. The higher the number of endpoints, the longer the data manipulation will take. Scalability is an issue in a solution like this. The layout can be different on all set-top boxes to avoid the end user seeing the video with lag on the screen. Depending on the number of different designs, different output streams must be coded. An MCU can differentiate between the endpoints themselves or by groups of endpoints, exemplified in two groups, one for a low bit rate providing a first view and one for high bit rates providing a second view.
Decentralized conference
In a decentralized multipoint scenario, a centralized MC is needed. Each endpoint will send its media data to the other endpoints - typically by multicast. Each endpoint will mix the audio from the other endpoints, and combine or select the video streams for local display. The MC still acts as the controller for the conference, and each endpoint will have a control connection with the MC. In a decentralized conference, each endpoint must have the functionality of the MCU by functionally displaying a Current / Previous View, a Combined View, or a Resized View. The complexity of an endpoint that supports decentralized conferences is greater than that of endpoints that support centralized conferences.
Hybrid Conference
A hybrid conference uses a combination of centralized and decentralized conferencing. Some endpoints will be in a centralized conference, and others will be in a decentralized one. A hybrid conference can have a centralized treatment of one media stream, and a decentralized distribution of another. Before the start of the multimedia conference, the centralized MCU will send commands to each endpoint participating in the conference, these commands, among others, will ask the endpoint to inform the MCU of its bit rate capabilities and its processing capacity. codecs. The information received will be used by the centralized MCU to establish a multimedia hybrid conference, in which the characteristics of each endpoint are taken into account.
The term hybrid is also used when audio is mixed at the MCU and each endpoint selects and decodes one or more incoming video streams for local viewing.
Scalable signal compression
Scalable signal compression algorithms are a primary requirement of the rapidly evolving global network, involving a variety of channels with widely differing capabilities. Many applications require that data can be decided simultaneously at a variety of speeds. Some examples include applications such as multicasting in a heterogeneous network, where the channels dictate the feasible bit rates for each user. Similarly, it is motivated by the coexistence of extreme points of complexity and different costs. A compression technique is scalable if it offers a variety of decoding speeds and / or pro
ES 2 336 216 T3 cessation using the same basic algorithm, and wherein the lower speed information streams are embedded within the higher bit rate streams in a way that minimizes redundancy.
Various algorithms have been proposed that allow scalability of video communication, including frame rate (temporally scalable encoding), visual quality (SNR), and spatial scalability. Something common in these procedures is that the video is encoded in layers and the scalability comes from the decoding of one or more layers.
Temporarily scalable coding
Video is encoded in frames, and a time-scalable video encoding algorithm enables the extraction of video of multiple frame frequencies from a single encoded stream. The video is divided into multiple sets of interleaved frames. When decoding more than one set of frames, the frame rate increases.
Spatial scalable coding
The spatial scalable compression algorithm is an algorithm in which the first layer has a current resolution, and the video resolution can be improved by decoding more layers.
Scalable SNR encoding (visual quality scalable encoding)
SNR scalable compression refers to encoding a stream in such a way that videos of different qualities can be reconstructed by decoding the encoded bitstream. Scalable compression is useful in today's heterogeneous network environments where different users have different speeds, resolutions, displays, and computational capabilities.
Patent WO9918728 A1 describes a multimedia server that includes a number of different codecs. Data streams of different standards enter the server and are routed in the appropriate codec, in which the data streams are decompressed. After decompression, the signals are mixed and routed back to the appropriate codecs. The signals are recompressed to the appropriate standard for each receiver unit before exiting the server.
Summary of the invention
An object of the present invention is to provide a system and a method that eliminates the drawbacks described above. The characteristics defined in the appended claims characterize this system and procedure.
In a traditional centralized system, the endpoints will send a full scale image to an MCU, for example an encoded CIF image (352 x 288 pixels) will be sent to the MCU. To improve the quality of the conference, it would be useful to present a composite image at each endpoint. This full image can display one participant as a main fraction of a full screen, while the rest of the participants are displayed as smaller sub-images. The participant, the participant size, and the number of participants displayed at each site may depend on the processing and display capabilities and the conference situation. If each endpoint is supposed to receive composite images, the MCU has to perform heavy data manipulation as described in continuous presence and resized view. After decoding the encoded CIF data streams to video images, the MCU will compose composite images which will be recoded and sent to the appropriate endpoint.
This solution increases the demand on the capacity of the central MCU; The solution, in cases where heavy use of encoding and decoding is necessary, will incorporate annoying delay between participants in a multimedia conference.
In particular, the present invention describes an improved method and system for the exchange of information between a number of units where a central unit, based on knowledge related to a plurality of sub-units, it will instruct the sub-units to generate multimedia data streams matched to other sub-units participating in the same session in such a way that the central unit can route data streams without using its built-in codecs or minimal use of said codecs.
Brief description of the drawings
To make the invention more easily understood, the description that follows will refer to the accompanying drawings.
Figure 1 shows an example of a centralized conference, with four endpoints participating in a video conference according to the invention.
ES 2 336 216 T3
Figure 2 shows an exemplification of the invention with four participating endpoints and EP D as the current speaker and EPD as the previous speaker.
Detailed description of the invention
As stated above, all solutions have their drawbacks. A feature of a quality video conference will be that it includes the ability to display composite windows, or mixed and correct view layouts (ie 4: 3, 16: 9) without annoying time lags. All these requirements must be met with existing equipment, that is, with MCUs that are currently available. Known solutions don't scale very well as it doesn't take into account the different capacity of different endpoints in a conference. Ideally, each endpoint should receive data streams tailored to its capacity and design. Additional data processing performed in the central MCU should be minimized. The indicated weakness applies to centralized, decentralized, as well as hybrid solutions.
To overcome the aforementioned weakness, this invention takes advantage of the use of decentralized processing capacity at the participating endpoints, and also takes advantage of the use of a command language to instruct each participating endpoint on how to take part. in the conference.
The idea is simply to use the available capacity as decentralized as possible, therefore the requests on the MCU will be reduced, and it is also important for optimization that the solution scales well. The MCU, through the use of a command language, will obtain information about each endpoint capability such as a receiver and a transmitter regarding available bit rate, decoding capability, etc., therefore, the MCU will adapt data flows to each endpoint according to your specifications, resulting in well-scaled systems.
Thus, the MCU will collect information related to the encoding capacity of each endpoint in relation to the number of multiple frames it can do and what resolution, the bit rates and the frame frequencies. Also, the MCU will have knowledge related to end point designs, etc., as stated above. With this, the MCU will be able to analyze this information and personalize a conference. Therefore, the thinking is based on the knowledge of the encoding capabilities to the endpoints, that the endpoints will use their capabilities to send optimized multimedia streams. Thus, the need for processing in the MCU will be drastically reduced compared to what is normal in a conference of similar quality.
Data switching Continuous presence
Encoders can be instructed to send a larger but still reduced stream to the MCU. The instructions will be sent in a command language from the MC in the central MCU to each encoder of the participant in the multimedia conference. Endpoints can also be instructed to send that stream and a smaller stream. The MCU can then combine streams to a current view with for example the speaker in a "large" window along with the other participants in smaller windows. This speaker can receive the previous speaker in a "large" window with the rest of the participants in "small" windows. An example of a layout is 5 + 1.
The required MCU capacity can be significantly reduced by controlling the size of each stream and the bit rate used. The MC will use the encoding and decoding capabilities exchanged in a set command to select the appropriate designs. The capabilities of the encoder will restrict the size and number of partial frames from an endpoint, and the capabilities of the decoder will restrict the number and size of partial frames that an endpoint can receive. This forms the basis on which the MCU can decide its designs.
The MCU instructs the endpoints to send one or more partial frames of the session. The size of these partial frames will depend on the number of conference participants and the chosen design; the MCU will give instructions to the endpoints at the start of the session regarding the size of the partial frames. Thus, each endpoint will send a fraction of a composite image in the requested format. The MCU can also give additional commands during the session to change the layout. The amount of data that has to be encoded at the endpoint will therefore be substantially less, at least for non-speaking participants. The MCU will receive encoded images that are already in the correct format, therefore the MCU does not have to decode the incoming video streams. The MCU will only put together the composite pictures from the incoming partial frames without any decoding or encoding. This can be achieved by manipulating high-level syntax in the video stream to produce a combined frame, or by identifying, tagging, and forwarding a selection of the video streams to all endpoints, where they can be decoded separately and joined for one composite view.
Thus, the need for processing power is drastically reduced, and by avoiding processing of video streams, the delay will be reduced correspondingly.
ES 2 336 216 T3
In a centralized conference, the MCU will instruct the endpoints in the conference to make one or more partial frames. The endpoints will encode their video streams to comply with the format of these partial frames. The partial frames are then sent from the endpoint to the mCu. The MCU will combine the partial frames into one or more combined frames. The structures of these frames combined are called layouts. The patterns contain the format of the partial frames received for a given set of combined frames, and the instructions sent to each endpoint are derived from the patterns of these combined frames. Typically, one layout is defined for the Current View with a 4: 3 scale screen, and another for a 16: 9 scale screen. The combined frame for the Previous View will typically be scaled to match the receiving endpoint following the same principle as for the current view. The combined frames are sent to each endpoint in the conference with the best matching design for that specific endpoint.
In a decentralized conference, the MCU will instruct the endpoints in the conference to make one or more partial frames. The endpoints will encode their video streams to comply with the format of these partial frames. These partial frames are distributed to all end points of the conference. Each endpoint will combine the partial screens into combined screens with a set of patterns. The MC decides and signals the designs to each endpoint separately. Different endpoints in the conference may have different layouts assigned to them. Typically, some endpoints combine partial frames in a current view, while others are combined in a previous view.
The central MCU, using a command language communicated with the control channels, will request the endpoint to provide information on its capacity in terms of bit rates, designs and decompression algorithms. The central MCU, based on the responses from the decentralized MCUs, will establish a session tailored to each endpoint specification in terms of bit rates and the other parameters described above. The invention can use scalability as described above to encode multiple video streams at various bit rates and resolutions to ensure the best use of available bandwidth.
Set of marked commands
Describes the set of commands between the central MCU and each endpoint. The command set is used to instruct the coding of partial frames at the end points and the layout of video streams, and the capability set describes the range of formats that can be received at each end point. Commands for aligning or changing capabilities can also be part of the language.
First embodiment of the invention
Centralized Conference Example
The example of a centralized conference is shown in figure 4.1 The example contains a central MCU. The MCU has a conference with 4 endpoints. They have been named Endpoint A, B, C and D. Each endpoint has a bidirectional control channel, a video stream that goes from the endpoint to the MCU, and a video stream that goes from the MCU to the endpoint. The current speaker in the conference is at End Point A, and End Point A is therefore receiving a combined frame from the previous view. The rest of the endpoints in the conference are receiving different combined frames from the current view. The extreme point D is the previous speaker.
The MCU signals through the command set described above to endpoint A to produce two partial frames. These are Partial Frame 1 and Partial Frame 5. The size, format and scale of both partial frames are specifically signaled. Partial Screen 1 is part of the layout for the current 16: 9 view selected by the MCU. Partial Screen 5 is part of the layout for the current 4: 3 view also selected by the MCU. The MCU continuously receives a video stream from endpoint A that contains the format of both Partial Frame 1 and Partial Frame 5 until a new command is signaled from the MCU to endpoint A.
Similar to Endpoint A, the MCU is signaling to endpoint B to encode partial frame 2 and partial frame 6. Endpoint C must encode partial frame 3 and partial frame 7. Endpoint D must encode partial frame 4, partial frame 6 and partial frame 9.
The MCU receives all partial frames 1 to 9. With the layout for the "Combined Frame Current View 16: 9" the MCU combines Partial Frame 1, Partial Frame 2, Partial Frame 3 and Partial Frame 4. This frame combined is sent to Endpoint C and Endpoint B. Both have signaled that they can receive a scaled 16: 9 frame. With the layout for the “Current View 4: 3 Combined Screen” the MCU combines Partial Screen 5, Partial Screen 6, Partial Screen 7 and Partial Screen 8. This combined frame is sent to Endpoint D which can only receive a 4: 3 scale frame.
The combination of Partial Screen 9, Partial Screen 3 and Partial Screen 5 make the layout for the “Combined Screen 16: 9 Previous View”.
ES 2 336 216 T3
Example of a group of commands for the implementation of the invention
This example is a reduced exchange of information between participating units to illustrate how communication should be implemented. In a real-world situation, the varying capabilities of the endpoints, such as encoding standards and bandwidth, and the capabilities of the MCU can cause several series of exchanges to align capabilities, adding new endpoints on the fly can also cause realignment of capabilities during the session.
For the sake of simplicity, this exchange assumes that the capabilities of the MCU are all-encompassing and that the capabilities of the endpoints match such that alignment is not necessary. It is also a real case that all units in the session are of the same type.
In this example, multiple endpoints get the same design. In a real case, each endpoint can have different layouts and even different aspect ratios depending on its display.
Capacity exchange
The exchange between the participating units gives information related to the processing capabilities such as standards, image size, frame frequency and bandwidth.
Encoder / Decoder Capability
DECCAP- {ProcessingRate, NumberOfStreams, TotalImageSize, Bandwidth}
ENCCAP- {ProcessingRate, NumberOfStreams, TotalImageSize, Bandwidth} ProcessingRate - The ability to process video elements. These elements can be measured in MacroBlocks (MBs), which is a group of 16x16 pixels.
NumberOfStreams (Number of streams) - The number of separate streams that can be handled.
TotalImageSize - The maximum combined size of all streams, also measured in MBs. The description of the image may also contain the aspect ratio of the image.
Bandwidth - The maximum total data rate that can be sent or received.
Commands
A small group of commands that will allow the exchange of data . CODE-SEQn- {Resolution, FrameRate, Bandwidth}.
A command for an encoder that forces the encoding of a video stream with a set of limitations.
Resolution - The size of the video image measured in MBs.
FrameRate (Frame Rate) - The number of video images that can be sent per second (F / s).
Bandwidth - The number of bits per second that can be used for this video stream (Bits / s). STOP-Seqn. A command to stop encoding a particular video stream.
LAYOUT (DESIGN) {Mode, SEQ1, SEQ2, .., SEQm}
A command for a decoder that tells you how to place a number of streams on the display.
Mode - The particular design chosen, eg. eg, 5 + 1, in which the number of flows and their position on the screen are defined.
Seq1.m - The ID of the sequences that can be placed in the defined layout. The order of the sequences gives the position. If a particular position does not have any flow, SEQ0 can be used.
Request
GET-FLOOR The delivery of the current speaker to a particular endpoint.
ES 2 336 216 T3
Data exchange
VIDEO-FRAME-SEQn
The encoded video data for a frame of a particular video sequence. For simplicity, the data units for a video sequence are defined as a frame.
The example used is the one shown in figure 2, where EP A is the current speaker and EP D is the previous speaker; the exchange of additional capacity, the commands for the login, the commands to capture the word and the exchange of data are shown in the following diagrams.
Capacity exchange
<td rowspan="7">MCU</td><td>EP; DECCAP- {12000MBS, 6,</td><td rowspan="3">To EP] 396MBS, 384kBit / s 396MBS, 384kBit / s 396MBS, 768kBit / s</td><td rowspan="5">B EP '</td><td rowspan="7">C EP ...</td>
<td>w - ENCCAP- {12000MBS, 6,</td>
<td> - DECCAP- {12000MBS, 6,</td>
<td>4 ENCCAP- {12000MBS, 6,</td><td>396MBS, 768kBit / s</td>
<td> - DECCAP- {12000MBS, 6,</td><td>396MBS, 768kBit / s</td>
<td>ENCCAP- {12000MBS, 6,</td><td>396MBs, 768kBit / s</td><td></td>
<td></td><td></td><td></td>
Commands for login
MCU
EP A
EP B
EP C
EP ...
CODE-SEQA1- {8x6MB, 30F / S, 42kBit / s}
---------------<sub>k</sub>
CODE-SEQA2- {14x12MB, 30F / s
170kBit / s}
CODE-SEQB1- (8x 6MB,
30F / s, 42kBit / s}
CODE-SEQC1- {8x6MB,
42kBit / s]
LAYOÜT-15 + 1,
LAYOUT- {5 + l,
<img file="ES2336216T3_D0001.tif" />
t
SEQC1,
SEQC1,
SEQC,
SEQ0}
SEQ0
---►
SEQ0}
ES 2 336 216 T3
Commands to capture the word
B becomes the current speaker and A the previous speaker
<img file="ES2336216T3_D0002.tif" />
Data exchange
MCU
<img file="ES2336216T3_D0003.tif" />
<img file="ES2336216T3_D0004.tif" />
<img file="ES2336216T3_D0005.tif" />
<img file="ES2336216T3_D0006.tif" />
<img file="ES2336216T3_D0007.tif" />
<img file="ES2336216T3_D0008.tif" />
<img file="ES2336216T3_D0009.tif" />
<img file="ES2336216T3_D0010.tif" />
VIDEO- FRAME- YES ¡QAl ◄-VI DEO- FRAME- YES ¡QA2
VIDEO-FRAME - YES ¡QA 1
VI DEO-FRAME- YES, QA2
VIDEO-FRAME-SI ¡QBl «- ---- VIDEO-FRAME-SI ¡QBl
-----------> - Decentralized conference
Using the same situation described above, in a decentralized conference, the MC will instruct EP A to encode and transmit PF1 to EP B and C and to send PF5 to EP D. EP B will transmit PF 2 to EP A, B and C and will send PF 6 to EP D. EP C will transmit PF 3 to EP A, B and C and send PF 7 to EP D. Finally, EP D will transmit PF 4 to EP A, B and C, send PF 8 to EP D and send PF 9 to EP A.
ES 2 336 216 T3
Advantage
Some of the advantages according to the present invention are summarized below:
• Reduction of the processing requirement in the central unit. Leads to a more scalable solution.
• Reduced delay in transmission compared to transcoding.
• Reduced processing power at extreme points due to smaller overall image size. Better video quality as no resizing is required at the headunit or endpoints.
Equivalents
Although this invention has been particularly shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes can be made in the form and details thereof without departing from the scope and spirit of the invention. , as defined in the appended claims.
In the above examples, the preferred embodiments of the present invention are exemplified by the use of 4: 3 and 16: 9 scale screens in a display, however, the solution is not limited to the use of these aspect ratios, it is they can implement other known aspect ratios, such as 14: 9 or other ratios that can be divided into a grid scheme in a display.
For example, the idea of using partial frames based on the knowledge of each endpoint in a conference is extended to be used whenever there is a need to send multimedia streams between a plurality of users. The concept will be of interest within traditional broadcasts, particularly when real-time events are covered. By imagining a scenario where a plurality of cameras are used to cover an event, if each camera is transferring information to a centralized unit according to the rules negotiated between the centralized unit and the cameras, a lot of processing power can be saved in the centralized unit. Also, it would be much easier and faster to process composite / PIP frames for end users.
Another example is that the physical requirement of the described MCU and the MC is similar, any combination of Centralized and Decentralized conference can be performed, the implementations are also expected to have traditional MCUs as part of the network to be backward compatible of today's solutions.
One of the main ideas of the invention is: the use of capacity where it is. In a traditional multipoint streaming media exchange, there is a central station or centralized unit that manages the data exchange. Traditionally, this centralized unit has not negotiated with all the peripheral units that participate in the data exchange to optimize the data exchange according to the capacity of each peripheral unit, therefore, an optimized use of the entire unit is not known or is not frequent. processing capacity available.
Abbreviations and references
Endpoint: any terminal capable of joining a conference.
Media: Audio, video and similar data.
Flow: medium continuous.
Multipoint Control Unit (MCU): The entity that controls and deals with the media for 3 or more endpoints is, in a conference, a Controller
Multipoint (MC): It deals with the control of 3 or more end points in a conference.
Centralized conference: Control channels are signaled one-way or two-way between the endpoints and the MCU. Each endpoint sends its media to the MCU. The MCU mixes and matches the media and returns the media to the extreme points.
Decentralized conference: Control channels are signaled one-way or two-way between the endpoints and the MCU. The media is transported as a multicast between the endpoints, and the endpoints mix and match the media on their own.
Hybrid Conference: The MCU has a conference that is partially centralized and partially decentralized.
Speaker: The participant at the endpoint who speaks loudest between the endpoints in a conference.
ES 2 336 216 T3
Current View: The video stream from the current speaker.
Previous View: The video stream from the previous speaker.
Combined View: A high-resolution video stream derived from low-resolution video streams.
Resized view: A video stream made from another video stream by resizing.
References cited in description
This list of references cited by the applicant is only intended to aid the reader and is not part of the European patent document. Although the utmost care has been taken to carry them out, errors or omissions cannot be excluded and the EPO declines any responsibility in this regard.
Patent documents cited in the description • WO 9918728 A1 [0019]
Contents20
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
71 members in 12 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 20035078 | Norway | A | |
| 20035078 | Norway | A | |
| 0480019220035078 | – | – | – |
| NO20030005078 | – | – | – |
Members71
| Document | Office | Kind | |
|---|---|---|---|
| US6136707A | United States of America | A | |
| WO0126145A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2001005056A1 | United States of America | A1 | |
| KR20020043604A | Republic of Korea | A | |
| WO0126145A9 | World Intellectual Property Organization (WIPO) | A9 | |
| TW504795B | Taiwan Province of China | B | |
| US6518668B2 | United States of America | B2 | |
| JP2003511858A | Japan | A | |
| US2003129828A1 | United States of America | A1 | |
| US6610151B1 | United States of America | B1 | |
| NO20034775D0 | Norway | D0 | |
| NO20035078D0 | Norway | D0 | |
| US2004087171A1 | United States of America | A1 | |
| WO2004043068A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2004043074A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003274849A1 | Australia | A1 | |
| AU2003276786A1 | Australia | A1 | |
| US2004150712A1 | United States of America | A1 | |
| US2004168110A1 | United States of America | A1 | |
| WO2005041574A1 | World Intellectual Property Organization (WIPO) | A1 | |
| NO318868B1 | Norway | B1 | |
| NO318911B1 | Norway | B1 | |
| WO2005048600A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US6903016B2 | United States of America | B2 | |
| US2005122392A1 | United States of America | A1 | |
| US2005124153A1 | United States of America | A1 | |
| US2005144233A1 | United States of America | A1 | |
| US2005148172A1 | United States of America | A1 | |
| US6924226B2 | United States of America | B2 | |
| EP1676439A1 | European Patent Office (EPO) | A1 | |
| EP1683356A1 | European Patent Office (EPO) | A1 | |
| US2006166448A1 | United States of America | A1 | |
| US7105434B2 | United States of America | B2 | |
| CN1871855A | China | A | |
| CN1883197A | China | A | |
| US7199052B2 | United States of America | B2 | |
| JP2007511954A | Japan | A | |
| JP2007513537A | Japan | A | |
| US2007117379A1 | United States of America | A1 | |
| US7282445B2 | United States of America | B2 | |
| US2008026569A1 | United States of America | A1 | |
| US7474326B2 | United States of America | B2 | |
| US7509553B2 | United States of America | B2 | |
| US7550386B2 | United States of America | B2 | |
| US7561179B2 | United States of America | B2 | |
| US2009233440A1 | United States of America | A1 | |
| US2009239372A1 | United States of America | A1 | |
| CN100568948C | China | C | |
| EP1683356B1 | European Patent Office (EPO) | B1 | |
| ATE455436T1 | Austria | T1 | |
| US2010033550A1 | United States of America | A1 | |
| DE602004025131D1 | Germany | D1 | |
| US7682496B2 | United States of America | B2 | |
| ES2336216T3This record | Spain | T3 | |
| US2011068470A1 | United States of America | A1 | |
| US2011134206A1 | United States of America | A1 | |
| EP1676439B1 | European Patent Office (EPO) | B1 | |
| ATE538596T1 | Austria | T1 | |
| US8123861B2 | United States of America | B2 | |
| US2012126409A1 | United States of America | A1 | |
| US8289369B2 | United States of America | B2 | |
| US2013038677A1 | United States of America | A1 | |
| US8560641B2 | United States of America | B2 | |
| US8586471B2 | United States of America | B2 | |
| US2014061919A1 | United States of America | A1 | |
| US8773497B2 | United States of America | B2 | |
| US2014354766A1 | United States of America | A1 | |
| US2015155239A1 | United States of America | A1 | |
| US9462228B2 | United States of America | B2 | |
| US9673090B2 | United States of America | B2 | |
| US10096547B2 | United States of America | B2 |
Numbers
- Publication, DOCDB
- 2336216
- Publication, EPODOC
- ES2336216T
- Application
- 4800192
- Application, DOCDB
- 04800192
- Application, EPODOC
- ES20040800192T
Titles2
- Spanish
- COMPOSITOR DE MEDIOS DISTRIBUIDOS A TIEMPO REAL.
- English
- COMPOSER OF MEDIA DISTRIBUTED IN REAL TIME.
Classification
- CPC, 6
- H04M3/562
- H04N7/152
- H04M3/567
- H04N7/15
- H04N7/14
- H04L65/403
- IPC, 2
- H04N7 15
- H04M3 56