Distributed real-time media composer
Abstract
A system and a method allowing simultaneous exchange of audio, video or data information between a plurality of units over a communication network, supported by a central unit, wherein the central unit is, based on knowledge regarding one or more of the units, adapted to instruct said one or more units to generate multimedia data streams adjusted to fit into certain restrictions to be presented on other units participating in a same session.

Term
No projected expiry on record.
- Priority and filed
- Granted
- Today
24 claims: 12 independent, 12 dependent
- 1Claims Patentkrav 1. A system that allows simultaneous exchange of audio, video and / or data information between a plurality of devices over a communication network and supported by a central unit, characterized in that the central unit is, based on knowledge of one or more of the units, adapted to instructing said one or more devices to generate multimedia data streams adjusted to fit specific restrictions, to be presented on other devices participating in the same session. 1. System som tillater samtidig utveksling av audio-, video- og/eller datainformasjon mellom et flertall av enheter over et kommunikasjonsnettverk og støttet av en sentral enhet, karakterisert ved at sentralenheten er, basert på kunnskap med hensyn til en eller flere av enhetene, tilpasset til å instruere nevnte ene eller flere enheter for å generere multimedia datastrømmer justert for å passe inn til bestemte restriksjoner, for å bli presentert på andre enheter som deltar i den samme sesjon.
- 4System according to any one of claims 1-3, characterized in that the central unit is adapted to compose a combined data stream from one or more simple data streams sent from one or more units and to route said combined current to said other units. 4. System i henhold til et av kravene 1-3, karakterisert ved at sentralenheten er tilpasset til å komponere en kombinert datastrøm fra en eller flere enkle datastrømmer sendt fra den ene eller flere enheter og til å rute nevnte kombinerte strøm til nevnte andre enheter.
- 5System according to characterized conference session. 5. System i henhold til karakterisert konferansesesjon.
- 6A system according to any one of claims 1-4, wherein the session is a video one of claims 1-5, wherein the communication between said central unit and the units utilizes scalable compression techniques. 6. System i henhold til karakterisert et av kravene 1-4, ved at sesjonen er en video et av kravene 1-5, ved at kommunikasjonen mel25 lom nevnte sentralenhet og enhetene bruker skalerbare kompresjonsteknikker.
- 7System according to any one of claims 1-3, characterized in that said data multimedia streams are unidirectional, including subframes and are directed from the units to the central unit. 7. System i henhold til et av kravene 1-3, karakterisert ved at nevnte datamultimediestrømmer er ensrettet, inkludert delrammer og er rettet fra enhetene til sentralenheten.
- 10System according to any one of claims 3-6, characterized in that a display is divided into grid pattern of cells, wherein each of said subframes occupies one or more cells of the grid pattern and the display may have different length / width ratios. 10. System i henhold til et av kravene 3-6, karakterisert ved at et display er delt inn i rutemønster av celler, der hver av de nevnte delrammer opptar en eller flere celler av rutemønsteret og displayet kan ha forskjellige lengde-/breddeforhold.
- 13Process for exchanging audio, video and data information simultaneously between a plurality of devices in a communications network supporting a central unit, characterized by instructing one or more of the devices, based on a knowledge of these, to generate multimedia data streams adapted to fit to specific restrictions, to be presented on other entities participating in one 13. Fremgangsmåte for utveksling av audio-, video- og datainformasjon samtidig mellom et flertall av enheter i et kommunikasjonsnettverk som støtter en sentral enhet, karakterisert ved å instruere en eller flere av enhetene, basert på en kunnskap om disse, å generere multimediedatastrømmer tilpasset til å passe inn til bestemte restriksjoner, å bli presentert på andre enheter som deltar i en 5 same session. 5 samme sesjon.
- 16A method according to any of claims 13-15, characterized in 16. Fremgangsmåte i henhold til et av kravene 13-15, karakterisert ved 20 to compose in the central unit, a combined data stream composed of one or more single data streams sent from one or more units, and route said combined power to said other units. 20 å komponere i sentralenheten, en kombinert datastrøm sammensatt fra en eller flere enkeltdatastrømmer sendt fra en eller flere enheter, og rute nevnte, kombinerte strøm til nevnte, andre enheter. 25 25
- 18A method according to any one of claims 13-17, characterized by using scalable compression techniques in the exchange of data between said central unit and the units. 18. Fremgangsmåte i henhold til et av kravene 13-17, karakterisert ved å bruke skalerbare kompresjonsteknikker ved utveksling av data mellom nevnte sentralenhet og enhetene.
- 19A method according to any one of claims 15-18 19. Fremgangsmåte i henhold til et av kravene 15-18, karakterisert ved 5 dividing a display into a grid pattern of cells where each of said subframes occupies one or more cells of the grid pattern and the display may have different width / height ratios. 5 å dele et display inn til et rutemønster av celler der hver av nevnte delrammer opptar en eller flere celler av rutemønsteret og displayet kan ha forskjellig bredde-/høydeforhold.
- 22A method according to any one of claims 13, 14, 19 or 20, 22. Fremgangsmåte i henhold til et av kravene 13, 14, 19 eller 20, 20 characterized in that said data multimedia streams are unidirectional, include subframes and are directed from the units to the central unit, 20 karakterisert ved at nevnte datamultimediestrømmer er ensrettet, inkluderer delrammer og er rettet fra enhetene til sentralenheten,
Independent claims12
179 paragraphs in 8 sections, as filed
<img file="NO318911B1_D0001.tif" />
NORWAY (12) PATENT (19) NO (51) IntCl<sup>7</sup> (11) 318911 H 04 L 12/18 (13) BI
NIPO
<td> (21)</td><td>Appln</td><td> 20035078</td><td> (86)</td><td>Commencement day and</td>
<td></td><td></td><td></td><td></td><td>Appln</td>
<td> (22)</td><td>Inng.dag</td><td> 2003.11.14</td><td> (85)</td><td>Videreføringsdag</td>
<td> (24)</td><td>Løpedag</td><td> 2003.11.14</td><td> (30)</td><td>Priority None</td>
<td> (41)</td><td>Alm.tilgj</td><td> 2005.05.18</td><td></td><td></td>
<td> (45)</td><td>communicated</td><td> 2005.05.23</td><td></td><td></td>
<td> (73)</td><td>proprietor</td><td colspan="3">Tandberg Telecom AS, PO Box 92, 1325 Lysaker, NO</td>
<td> (72)</td><td>Inventor</td><td colspan="3">Tom-Ivar Johansen, Eckersbergs gate 35.0266 OSLO, NO</td>
<td></td><td></td><td colspan="3">Geir Ame Sandbakken, Vøyensvingen 19 B, 0458 Oslo, NO</td>
<td> (74)</td><td>Fullmektig</td><td colspan="3">Oslo Patentkontor AS, PO Box 7007 Majorstua, 0306 OSLO, NO</td>
<td> (54)</td><td>Designation</td><td>Distributed composition aY real-time media</td>
<td> (56)</td><td>cited</td><td></td>
<td></td><td>publications</td><td>US 6,288,739 WO 99/18728 Al EPO 691 779 A2</td>
<td> (57)</td><td>Summary</td><td></td>
A system and method permitting simultaneous exchange of audio, video and / or data information between a plurality of devices over a communication network supported by a central unit, the central unit being based on knowledge of one or more units adapted to instruct said one or more devices to generate multimedia data streams adapted to fit specific restrictions to be presented on other devices participating in the same session.
<img file="NO318911B1_D0002.tif" />
Xomroltkanal
Strem
Subframe 5
Endpoint A
Field of the Invention
The present invention relates to systems which allow simultaneous exchange of audio, video and data information using telecommunications, in particular it relates to video conferencing and web conferencing systems.
BACKGROUND OF THE INVENTION
In particular, the invention discloses a system and method which allows simultaneous exchange of audio, video and data information between a plurality of devices using existing telecommunication networks.
There are a number of technology systems available to arrange meetings between participants located in different locations. These systems may include audio visual multipoint conferencing or video conferencing, web conferencing and audio conferencing.
The most realistic substitute for real-world meetings is high-end video conferencing systems. Conventional video conferencing systems comprise a number of endpoints that communicate real-time video, audio and / or data streams over and between different networks such as WAN, LAN and line switched networks. Endpoints include one or more monitors, cameras, microphones and / or data capture devices and a codec. Said codec codes and decodes outgoing and incoming streams respectively.
Multimedia conferencing can be divided into three main categories, centralized, decentralized, and hybrid conferences, with each category having multiple variations to run a conference.
Centralized conference
Traditional audio-visual multipoint conference has a central multipoint control unit (MCU) with three or more endpoints connected. These MCUs perform switching functions to allow the audiovisual terminals to communicate in between in a conference. The central function of an MCU is to connect a plurality of video conferencing sites (sites, EP endpoints) by receiving frames with digital signals from audiovisual terminals (EP), and processing the received signals and transmitting the processed signals to the appropriate audiovisual terminals (EP ) that hit with digital signals. The digital signals may include audio, video, data and control information. Video signals from two or more audiovisual terminals (EPs) may be spatially mixed to form a composite video signal to be viewed by teleconferencing participants. The MCU serves as a selective router of media streams in this scenario. Part of the MCU called the Multipoint Controller controls the conference.
Each endpoint has a control channel for sending and receiving control signals to and from the MC. The MC acts on and sends commands to the endpoints.
Woice switch single stream voice control
In a centralized conference, the MCU will receive incoming video streams from all participants. It can forward a video stream from one endpoint to all the other endpoints. Which endpoint stream endpoint stream is selected is typical for single stream voice control based on which participant speaks the most, ie the speaker. This stream is called Current View. While Previous View is the video stream of the participant at the endpoint that is the speaker before the current speaker. In a Voice Switched Conference, a Current View video stream will be sent to anyone other than the current speaker and Previous View will be sent to the current speaker. One problem for the MCU is to ensure that Current View and Previous View can be received by all endpoints in the conference.
Switch single stream by other means
Current View can also be controlled by sending commands between the MCU and the endpoints. Such a mechanism is called floor control. An endpoint can send a floor request command to the MCU so that its video will be sent to all other participants. The previous picture will then typically be a Voice switch picture between all the other participants in the conference. Current View (current image) can be released by sending a word command. There are also other known methods for controlling the current image, including
Floor control or chair control. They both relate to single-stream control, but it is beyond the scope of this document to provide a full description of each known solution in this field. However, the principle of a current image and control of a single stream is the same.
Simultaneous presence
In a conference, one will often want to see more than one participant. This can be achieved in several ways. The MCU can combine the incoming video streams to create one or more outgoing video streams to achieve this. By combining multiple incoming low-resolution video streams from the endpoints into a high-resolution stream, one can do so. The high-resolution current is then sent from the MCU to all or some of the conference endpoints. This stream is called a Combined View. The characteristic of low-resolution currents limits the format of the high-resolution current from the MCU. Strict constraints on the incoming low-resolution stream are necessary to ensure that the combined high-resolution stream can be received by all endpoints to receive it. The MCU must, as long as each receiver receives the same multimedia stream, find the least common mode, to ensure acceptable viewing and listening characteristics of the receiver with the poorest capacity. With the large number of monitor variations, the MCU will also compensate for different monitors, such as 4: 3 or 16: 9 display. This is not possible with a common mode. This least common mode solution does not scale very well and imposes strong constraints on the receivers having a capacity that exceeds the capacity of the one with the lowest capacity.
Rescaled View
A more flexible solution is to let the MCU rescale all incoming video streams and thus create an image that can be received by all the endpoints receiving it. To do rescaling, the MCU must decode all incoming video streams. The decoded data - raw data - is then rescaled and transformed. The various raw data streams are then combined into a composite layout and put together to provide a set layout, tailored to the recipient's needs with regard to bit rate and coding standard. The combined raw data stream is then encoded and we will then have a new video stream containing one or more of the incoming streams. This solution is called the Rescaled View. To create a rescaled image, the MCU must understand and have the capability to encode and decode video streams. The more endpoints that are in the conference, the more capacity is needed from the MCU to decode all incoming streams. The heavy data manipulation performed by the MCU will add extra delay to the multimedia streams and thus reduce the quality of the multimedia conference. The higher the number of endpoints, the heavier the data manipulation will be. Scalability is important in a solution like this. The layout may be different for all decoders to prevent the end user from seeing themselves in a delayed video on the monitor. Depending on the number of different layouts, different outgoing streams will need to be coded. An MCU can differentiate between the endpoints per se or at groups of endpoints, exemplified by two groups, one for a low bit rate giving a first view and one for a high bit rate giving a second view.
Decentralized conference
In a decentralized multi-point scenario, one will only need one centralized MC. Each endpoint will send its media data to all other endpoints - typically by multicasting (broadcasting data to multiple users). Each endpoint io will mix audio from all other endpoints and will combine or select which video streams to display locally. The MC will still be the controller for the conference and each endpoint will have a control connection with the MC. In a decentralized conference, each endpoint of ice would have to have MCU functionality that displays a current / previous image (Current / Previous), combined image (Combined View) or a rescaled image (Rescaled View). The complexity of an endpoint that supports decentralized conferences is higher than that of endpoints that support centralized conferences.
Hybrid conference
A hybrid conference uses a combination of centralized and decentralized conference. Some endpoints will be in a centralized conference and others will be in a decentralized conference. A hybrid conference may have a centralized handling of one media stream and a decentralized distribution of another. Prior to starting the multimedia conference, the centralized MCU will send commands to each endpoint participating in the conference. These commands will include
request that the endpoints inform the MCU of its bitrate capacity and codec processing capacity. The information received will then be used by the centralized MCU to set up a multimedia hybrid conference taking into account the characteristics of each endpoint.
The term hybrid will also be used where audio is mixed at the MCU and each endpoint chooses to decode one or more incoming video streams for local viewing.
Scalable signal compression
Scalable signal compression algorithms are a key requirement of the rapidly evolving global network involving a variety of channels with a wide range of different capacities. Many applications require that data be simultaneously determined by a different number of rates. Examples include applications such as multicast in heterogeneous networks, where the channels dictate the usable bit rates for each user. At the same time, it is motivated by the coexistence of endpoints of different complexity and cost. A compression technique is scalable if it offers a plurality of decoding rates and / or processing needs using the same basic algorithm and where the lower rate of information flows is embedded within the higher rate of bit streams in a manner that minimizes redundancy.
Several algorithms have been proposed that allow scalability of video communications including frame rate, temporal scalable coding, visual quality (SNR) and spatial scalability. Common to these methods is that the video is encoded in layers where the scalability is due to decoding of one or more layers.
Temporally scalable coding
Video is encoded in frames and a temporarily scalable video coding algorithm allows the extraction of video for a plurality of frames from a single encoded stream. The video is divided into several interlaced sets of frames. By decoding more than one set of frames, the frame rate can be increased.
Spatially scalable coding
Spatial scalable compression algorithm is an algorithm where the first layer has a coarse resolution and the video resolution can be improved by multi-layer decoding.
SNR scalable coding (visual quality scalable coding)
SNR scalable compression refers to coding a sequence in such a way that different quality video can be reconstructed by decoding a subset of the encoded bit stream. Scalable compression is useful in today's heterogeneous network environment where different users have different rates, resolution, display and computational capabilities.
Known technique from the patent literature
From the patent literature there are examples of distributed video conferencing systems, though none that solve the problems outlined above. In particular, there are no solutions that scale well, that is, the solutions do not have the desired flexibility, they do not offer good solutions for centralized conferences, hybrid conferences and decentralized conferences, they do not provide sufficient flexibility in the use of the resources contained in a conference, ie say that there is not a good fit between available capabilities and conference results.
Publication US 6,288,739 B1 (Hales & al.) Discloses a distributed video communication system including a plurality of nodes, each of which nodes interfaces with a network comprising a message layer and a data transmission layer. Data is sent from each of the nodes of the data transmission layer in a multicast protocol so that all other nodes have access to the transmitted information sent. Each of the nodes has configuration information stored in a configuration block that is received through the message layer. A conference can be started by any of the nodes by sending information to other conference8 participants, each node being able to decide on how to extract data from the message layer, based on the received information, however, no central function for this information dissemination exists.
From WO 99/18728, a method for interconnecting multimedia data streams having different compression formats is known, in that the data streams of different compression standards reach a server, where they are sent to appropriate codecs relative to the coding standard they have.
io The signals are mixed and switched by a controller and a multipoint switch and then routed back to the correct codecs. The signals are recompressed to the appropriate standard for each user before leaving the server. That is, the system enables multimedia communication between participants adapted to the application of different standards. The system will not give any direct control over presentation layout.
Finally it should be mentioned that from EP 0 691 779 A2 a multimedia conference system is known in which there is a mention of meto20 there for resource management within multimedia conferences, but this resource management is specifically aimed at MMS application.
Summary of the Invention
It is an object of the present invention to provide a system and method that eliminates the disadvantages described above. The features defined in the appended claims characterize this system and method.
In a traditional, centralized system, the endpoints will send a full-scale image to an MCU. As an example, a coded CIF image (352 x 288 pixels) will be sent to the MCU. To improve the quality of the conference, it would be advantageous to present a composite image at each endpoint. This composite image may display a participant as the main part of a full screen, while all the other participants will be shown as smaller thumbnails (sub-pictures). Which participant, the size of the participant, and how many participants are displayed at each site (site) may depend on the processing and display capacity and conference situation. If each endpoint is assumed to receive composite images, the MCU will have to perform heavy data manipulation as described under simultaneous presence and rescaled image. After decoding the encoded CIF data stream into video images, the MCU will compose composite images that will be decoded and sent to the appropriate endpoint.
This solution places strict demands on the capacity of the central MCU and, in cases where heavy use of coding and decoding is required, will introduce an annoying delay between participants in a multimedia conference.
Specifically, the present invention discloses an improved method and system for exchanging information between a plurality of units where a central unit, based on knowledge of a plurality of subunits, will instruct the subunits to generate multimedia data streams to other subunits participating in the same session. in such a way that the central unit is capable of routing data streams without the use of its built-in codecs or the minimal use of said codecs.
Brief description of the drawings
In order to make the invention easier to understand, one in the discussion which follows will refer to the accompanying drawings.
Figure 1 shows an example of a centralized conference with four endpoints participating in a video conference according to the invention.
Figure 2 shows an exemplification of the invention with four endpoints participating and EP A is current speaker and EP D is preceding speaker.
Detailed description of the invention
As indicated above, all solutions will have their drawbacks. A characteristic of a quality video conference would be that it includes the ability to display composite windows, or combined viewing and correct layouts (ie 4: 3, 16: 9) without annoying time delays. All these requirements will be met with existing equipment, that is, with MCUs available today. The known solutions do not scale well as they do not take into account different capacities for different endpoints in a conference. Ideally, each endpoint should receive data streams tailored for its capacity and layout. Further processing by a central MCU should be minimized. The weakness indicated applies to centralized, decentralized as well as hybrid solutions.
To overcome the aforementioned weaknesses, the invention has the late advantage of using a decentralized processing power at the participating endpoints, as well as the benefit of using a command language to instruct each participant endpoint on how to participate in the conference.
The idea is simply to use the available capacity as decentralized as possible, thus reducing the requirement for the central MCU, furthermore it is important for optimization that the solution scales well. Using the command language, the MCU will get information on each endpoint's capacity as a receiver and transmitter with regard to available bitrate, encoding capacity, etc. Thus, the MCU will adapt the data streams to each endpoint according to their specifications, and as a result, will have a system that scales well.
Thus, the MCU will collect information with respect to each endpoint, coding capacity with respect to how many multiple frames can be created and at what resolution, bit rates and frames. Furthermore, the MCU will have knowledge of the layout of the endpoints, etc. as indicated above. With this information, the MCU will be able to analyze the information and customize a conference. The idea then is based on knowledge of the coding capabilities of the endpoints, that the endpoints will utilize their capacity to send optimized multimedia streams. Thus, the need for processing at the MCU will be greatly reduced compared to what is normal for a conference of similar quality.
Voice-controlled simultaneous presence
The encoders can be instructed to send a larger but still reduced power to the MCU. The instructions will be sent in a command language from the MC in the central MCU to each multimedia conference participant's codes. The endpoints may also be instructed to transmit said stream and a smaller stream. The MCU can then combine the streams of a current image {current view) with, for example, the speaker in a large window, along with the rest of the participants in smaller windows. The speaker can receive the previous speaker in a large window with the rest of the participants in small windows. A layout example is 5 + 1.
The required MCU capacity can be significantly reduced by controlling the size of each stream and the bit rate used. The MC will use encoder and decoder capabilities that are exchanged in a command set to select the appropriate layout. The encoder capabilities will limit the size and number of subframes from an endpoint, and the decoder capability will limit how many and how large subframes an endpoint can receive. This will form the basis for how the MCU can determine its layouts.
The MCU will instruct the endpoints to send one or more partial frames for the session. The size of these subframes will depend on the number of participants in the conference and the layout chosen. The MCU will provide instructions to the endpoints at the start of the session regarding the size of the subframes. Thus, each endpoint will send a fraction of a composite image in the requested format. The MCU can also issue additional commands during the session to change the layout. The amount of data that must be encoded at endpoints will be correspondingly small, at least for non-speaking participants. The MCU will receive encoded images that are already in the correct format, so the MCU will not need to decode the incoming video streams. The MCU will only assemble the composite images from the incoming subframes without any decoding or coding. This can be achieved by manipulating high-level syntax in the video stream to produce a combined frame or by identification labeling and forwarding a selection of the video streams to all endpoints where they can be separately decoded and composed for a composite image. Thus, the need for processing power will be greatly reduced and by avoiding processing of the video streams the delay will be correspondingly reduced.
In a centralized conference, the MCU will instruct the conference endpoints to create one or more subframes. The endpoints will encode their video streams to fit the format of these subframes. The subframes are then sent from the endpoint to the MCU. The MCU will combine the subframes into one or more combined frames. The structures for these combined frames are called layouts. The layouts contain the format of the subframes received for a given set of combined frames, and the instructions sent to each endpoint are derived from the layouts for those combined frames. Typical layout is defined for the current image (Current View) with a 4: 3 scaled frame and another for a 16: 9 scaled frame. The combined frames for the previous image will typically be scaled to match the endpoint receiving according to the same principles as the current image. The combined frames are sent to each endpoint of the conference given the best-fitting layout for the specific endpoint,
In a decentralized conference, MC will instruct the endpoints of the conference to create one or more partial frames. The endpoints will encode their video streams to fit the format of these subframes. These subframes are distributed to all endpoints of the conference. Each endpoint will combine the subframes into combined frames given a set of layouts. The layouts are determined and signaled by the MC to each endpoint separately. Different endpoints in the conference may have different layouts dedicated to themselves. Typically, some endpoints will combine subframes ice to a current image, while others combine up to a previous image.
The central MCU will, using a command language, communicate across the control channels, requesting the endpoints to provide information regarding its capacity relative to bit rates, layouts and compression algorithms. Based on the response of the decentralized MCUs, the central MCU will set up a session tailored to each endpoint specification with respect to bit rates and other parameters described above. The invention may use scalability as described above to encode multiple video streams at different bit rates and resolution to ensure the best use of available bandwidth.
Signaled Command Set
Describes the command set between the central MCU and each endpoint. The command set is used to instruct the encoding of subframes at the endpoints and the layout of video streams and the capability set describes the format range that can be received at each endpoint. Align or change capabilities can also be part of the language.
A first embodiment of the invention
Example of a centralized conference
The example of a centralized conference is shown in Figure 4.1. The example contains a central MCU. The MCU has a 5 conference with 4 endpoints. These are named endpoints A, B, C and D. Each endpoint has a bi-directional control channel, a video stream that goes from endpoint to MCU and a video stream that goes from MCU to the endpoint. The present speaker in the conference is at endpoint k, and endpoint io A therefore receives a combined frame of the previous image. All other endpoints in the conference receive different combined frames of the current image. Endpoint D is the previous speaker.
The MCU signals at the command set described above to ice endpoint A to produce two subframes. These are part frame (Partial Frame 1) and part frame 5 (Partial Frame 5). The size, format and scale of both subframes are specifically signaled. Subframe 1 is part of a layout for a 16: 9 current image selected by the MCU. Subframe 5 is part of the layout for the 4: 3 current image also selected by the MCU. The MCU will continuously receive an endpoint A video stream containing the format of both subframe 1 and subframe 5 until a new command is signaled from the MCU to endpoint A.
At the same time as for endpoint A, the MCU will signal to endpoint B to encode subframe 2 and subframe 6. Endpoint C should encode subframe 3 and subframe 7. Endpoint D should encode subframe 4, subframe 8 and subframe 9.
The MCU receives all subframes 1 through 9. With the Combined Frame 16: 9 Current View layout, the MCU combines subframe 1, subframe 2, subframe 3 and subframe 4. This combined frame is sent to endpoint C and endpoint B. Both have signaled that they can receive a 16: 9 scaled frame. With the layout of combined frame 4: 3 current image, the MCU will combine subframe 5, subframe 6, subframe 7 and subframe 8. This combined frame is sent to endpoint D which can only receive a 4: 3 scaled frame.
Combination of subframe 9, subframe 2, subframe 3 and subframe 5 constitutes the layout of the combined frame 16: 9 previous image.
Example of a command set for implementing the invention
This example is a reduced exchange of information between the participant units to illustrate how communication can be implemented. In a real situation, the different endpoint capabilities such as standards and bandwidth coding and capabilities for the MCU could create multiple rounds of interchange to match capabilities. Additions of new endpoints on the fly can also re-create capabilities during the session.
For simplicity, this exchange of information will assume that MCU capabilities are all-encompassing (all encompassing) and that endpoint capabilities match so that no customization is necessary. This is also a relevant case when all units in the session are of the same type.
In this example, several endpoints will have the same layout. In a real case, each endpoint would have different layout and even different aspect ratio (width-to-height ratio) according to its display.
Kapabilitetsutveksling:
The interchange between the participating units provides information regarding processing capabilities such as standards, image size, frame rate and bandwidth.
Codes / dekoderkapabilitet
DECCAP - {ProcessingRate, NumberOfStreams, TotallmageSize, Bandwidth}
ENCCAP - {ProcessingRate, NumberOfStreams, TotallmageSize, s Bandwidth}
ProcessingRate - The ability to process video elements. These elements can be measured in MacroBlocks (MBs) which is a group of 16x16 pixels (pixels).
NumberOfStreams - The number of separate streams that can be handled.
TotalmageSize - Maximum combined size of all streams, here also measured in MBs. The image description could also contain image aspect ratio
Bandwidth - Maximum total data that can be sent or received.
commands:
A small set of commands that will enable data exchange.
CODE-SEQn- (Resolution, FrameRate, Bandwidth)
A command for an encoder that forces encoding of a video stream with a set of constraints
Resolution - The size of the video image, measured here in MBs. FrameRate - The number of video images that can be sent per frame. second <F (s).
Bandwidth - Number of bits per second that can be used for this video stream (Bits / s)
STOP SEQn
A command to stop encoding a particular video stream.
LAYOUT- (Mode, SEQ1, SEQ2, SEQm}
A command for a decoder that tells you how to place a number of streams on the display.
Mode - The specially selected layout, such as 5 + 1, where the number of streams and their position on the screen is defined. SEQ1..m - ID of the sequences to be placed in the defined layout. The order of the sequences gives the position.
If a particular position should have no power, SEQO can be used.
Request:
GET-FLOOR (Get the word)
Handing over present speeches to a particular endpoint.
Data exchange:
VIDEO FRAME-SEQn
They encoded video data for a frame of a particular video sequence. For the sake of simplicity, the data units of a video drive are defined as a frame.
The example used is as shown in Figure 2, where EP A is the current speaker and EP D is the previous speaker, furthermore capability exchange, session start commands, grabbing floor commands and data exchange are shown in the following diagram.
Kapabilitetsutveksling:
MCU EP A EP B EP C EP ...
<td colspan="2">DECCAP- (12000HBS, 6,</td><td>396MB5, 3S4kBit / s</td><td></td>
<td>rt</td><td></td><td></td><td></td>
<td></td><td>ENCCAP (12000WBS, 6,</td><td>39SHBS, 364kBit / s</td><td></td>
<td> «-</td><td></td><td></td><td></td>
<td></td><td>DECCAP- [12000MB3, 6,</td><td>396MBS, 7iØkBit / s</td><td></td>
<td></td><td></td><td></td><td></td>
<td></td><td>ENCCAP <12000HBs, 6,</td><td>396HB5, 76SkBit / s</td><td></td>
<td></td><td></td><td></td><td></td>
<td>-a</td><td>DECCAP-U2000MBS, 6,</td><td>396MBS, 76BkBit / s</td><td></td>
<td>V<sup>1</sup></td><td></td><td></td><td></td>
<td></td><td>ENCCAP- {12DQOHB3,6,</td><td>396HB3, 760kBit / s</td><td></td>
<td colspan="2"></td><td></td><td></td>
Session Startup Commands:
MCU
EP A
EP B
EP C
EP
CODE-SEQA1- (flxfiHB,
30F / j, 42kBit / il
CODE * SEQA2-) Ux12h! , 3OF / s, 17 <HtBit / s (
CODE-SEQBl- (8x6MB,
30F / S, 42kBit / s)
CODE-SEQCl- <8x "MB,
LAYOUT-15 + 1, SEJQD2, SEQA1, SEQE | 1, SEQC1, SEQi
L, SEQC1, SEQC
LAYOUT- (5 + 1, SEQA2, SEQD1, SEQE1, SEQO), SEQO]
Grabbing floor commands:
B becomes current speaker and A previous speaker.
<img file="NO318911B1_D0003.tif" />
Data exchange:
<img file="NO318911B1_D0004.tif" />
Decentralized conference
If one uses the same situation as described above, in a decentralized conference, MC will instruct EP A on S code and broadcast PFI to EP B and C and to send PF5 to EP D. EP B would broadcast PR 2 to EP A, B and C and send PF 6 to EP D. EP C would broadcast PF 3 to EP A, B and C and send PF 7 to EP D. Finally, EP D would broadcast PF to EP A, B and C, it would send PF 8 to EP D and it would send PF 9 to EP A.
Benefits
Some of the advantages of the present invention are summarized below:
Reduced processing needs in the central unit result in ice a more scalable solution.
Reduced transmission delay compared to transcoding.
Reduced processing power at endpoints due to smaller total image size.
Better video quality as no rescaling is needed at the central unit or at the endpoints.
equivalents
As this invention has been particularly shown and described with reference to preferred embodiments, it will be appreciated by those skilled in the art that various variations in embodiments and details may be made therein without departing from the idea and scope of the invention as defined in the teachings of the present invention. attached requirements.
In the examples above, the preferred embodiments of the present invention are exemplified using 4: 3 io and 16: 9 scaled frames on a display, but the solution is not limited to use with these width / height ratios, other known aspect ratios comprising for example 14: 9 or other conditions that may be divisible until a grid pattern on a display can be implemented.
is As an example, the idea of using subframes based on the knowledge of each endpoint of a conference can be expanded to be used where it is necessary to send multimedia streams between a majority of users. The concept will have an interest in traditional broadcasting, especially when covering contemporary events. One can imagine a scenario where a majority of film cameras are used to cover an event. If each camera transmits its information to a centralized unit according to rules negotiated between the centralized unit and the cameras, a great deal of processing power can be saved at the centralized unit. Furthermore, it would be much easier and faster to process composite / PIP frameworks for end users.
Another example is that as the physical requirements of the described MCUs and MCs are similar, any combination of centralized and decentralized conference can be designed, it can also be expected that embodiments will have traditional MCUs as part of the network to be backward compatible with today's solutions.
One of the main ideas of this invention is: use the capacity where it exists. In traditional multi-point multimedia power exchange, a hub or centralized device manages the data exchange. Traditionally, this centralized unit will not need to negotiate with all peripherals participating in the data exchange to optimize data exchange according to each peripheral unit's capacity, thus optimum use of all available processing capabilities has not been common or unknown.
Abbreviations and references
Endpoint: Any terminal capable of attending a conference.
Media: Audio, video and similar data.
Stream: Continuous media.
Multipoint Control Unit (MCU): A unit that can control and handle media from 3 or more endpoints in a conference.
Multipoint Controller (Multipoint Controller, MC): Handle control for 3 or more endpoints in a conference.
Centralized Conference: The control channels are signaled unidirectional or bidirectional between the endpoints and the MCU. Each endpoint sends its media to the MCU. The MC15 MCU sees and combines the media - and sends the media back to the endpoints.
Centralized Conference: The control channels are signaled unidirectional or bidirectional between the endpoints and the MCU. Media is transported as multicasting between the endpoints and the endpoints mix and combine the media themselves.
Hybrid conference: The MCU holds a conference that is partially centralized and partially decentralized.
Speaker: The participant (s) at the endpoint who speaks the highest among the endpoints in a conference.
Current image: The video stream from the current speaker.
Previous image: The video stream for the previous speaker.
Combined image: A high-resolution video stream made from low-resolution video streams.
Rescaled Image: A video stream created by other rescaled video streams.
Contents8
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
71 members in 12 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 20035078 | Norway | A | |
| NO20030005078 | – | – | – |
Members71
| Document | Office | Kind | |
|---|---|---|---|
| US6136707A | United States of America | A | |
| WO0126145A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2001005056A1 | United States of America | A1 | |
| KR20020043604A | Republic of Korea | A | |
| WO0126145A9 | World Intellectual Property Organization (WIPO) | A9 | |
| TW504795B | Taiwan Province of China | B | |
| US6518668B2 | United States of America | B2 | |
| JP2003511858A | Japan | A | |
| US2003129828A1 | United States of America | A1 | |
| US6610151B1 | United States of America | B1 | |
| NO20034775D0 | Norway | D0 | |
| NO20035078D0 | Norway | D0 | |
| US2004087171A1 | United States of America | A1 | |
| WO2004043068A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2004043074A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003274849A1 | Australia | A1 | |
| AU2003276786A1 | Australia | A1 | |
| US2004150712A1 | United States of America | A1 | |
| US2004168110A1 | United States of America | A1 | |
| WO2005041574A1 | World Intellectual Property Organization (WIPO) | A1 | |
| NO318868B1 | Norway | B1 | |
| NO318911B1This record | Norway | B1 | |
| WO2005048600A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US6903016B2 | United States of America | B2 | |
| US2005122392A1 | United States of America | A1 | |
| US2005124153A1 | United States of America | A1 | |
| US2005144233A1 | United States of America | A1 | |
| US2005148172A1 | United States of America | A1 | |
| US6924226B2 | United States of America | B2 | |
| EP1676439A1 | European Patent Office (EPO) | A1 | |
| EP1683356A1 | European Patent Office (EPO) | A1 | |
| US2006166448A1 | United States of America | A1 | |
| US7105434B2 | United States of America | B2 | |
| CN1871855A | China | A | |
| CN1883197A | China | A | |
| US7199052B2 | United States of America | B2 | |
| JP2007511954A | Japan | A | |
| JP2007513537A | Japan | A | |
| US2007117379A1 | United States of America | A1 | |
| US7282445B2 | United States of America | B2 | |
| US2008026569A1 | United States of America | A1 | |
| US7474326B2 | United States of America | B2 | |
| US7509553B2 | United States of America | B2 | |
| US7550386B2 | United States of America | B2 | |
| US7561179B2 | United States of America | B2 | |
| US2009233440A1 | United States of America | A1 | |
| US2009239372A1 | United States of America | A1 | |
| CN100568948C | China | C | |
| EP1683356B1 | European Patent Office (EPO) | B1 | |
| ATE455436T1 | Austria | T1 | |
| US2010033550A1 | United States of America | A1 | |
| DE602004025131D1 | Germany | D1 | |
| US7682496B2 | United States of America | B2 | |
| ES2336216T3 | Spain | T3 | |
| US2011068470A1 | United States of America | A1 | |
| US2011134206A1 | United States of America | A1 | |
| EP1676439B1 | European Patent Office (EPO) | B1 | |
| ATE538596T1 | Austria | T1 | |
| US8123861B2 | United States of America | B2 | |
| US2012126409A1 | United States of America | A1 | |
| US8289369B2 | United States of America | B2 | |
| US2013038677A1 | United States of America | A1 | |
| US8560641B2 | United States of America | B2 | |
| US8586471B2 | United States of America | B2 | |
| US2014061919A1 | United States of America | A1 | |
| US8773497B2 | United States of America | B2 | |
| US2014354766A1 | United States of America | A1 | |
| US2015155239A1 | United States of America | A1 | |
| US9462228B2 | United States of America | B2 | |
| US9673090B2 | United States of America | B2 | |
| US10096547B2 | United States of America | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapsed by not paying the annual feesLapsedMM1K | MM1K | |
| Change of representativeCREP | CREP |
Numbers
- Publication, DOCDB
- 318911
- Publication, EPODOC
- NO318911B
- Application
- 5078
- Application, DOCDB
- 20035078
- Application, EPODOC
- NO20030005078
Titles2
- Norwegian
- Distribuert sammensetting av sanntids-media
- English
- Distributed composition of real-time media
Classification
- CPC, 6
- H04M3/562
- H04N7/152
- H04M3/567
- H04N7/15
- H04N7/14
- H04L65/403
- IPC, 2
- H04M3 56
- H04N7 15