Arrangement and method for generating continuous presence images
Abstract
Method for creating an objective continuous presence (CP) image encoded according to a video coding standard from a plurality of encoded video signals including defined macroblock orders, each comprising encoded video signals corresponding to a respective video image of the end terminal received from end terminals participating in a multipoint video conference, characterized by the fact that the procedure comprises the following steps: - decoding said encoded video signals, resulting in final terminal video images, - spatially mixing said final terminal video images, resulting in a plurality of CP images composed of regions associated respectively with each of said video images end-terminal, - encode said CP images, - rearrange the macroblocks of the encoded CP images, thereby creating said target encoded CP image.

Term
Term ended
Projected expiry passed 11 February 2025, 1.6 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
12 claims: 2 independent, 10 dependent
- 1ES 2 345 893 T3 ES 2 345 893 T3 CLAIMS REIVINDICACIONES 1. Method for creating a target continuous presence (CP) image encoded according to a video encoding standard from a plurality of encoded video signals including defined orders of macroblocks, each comprising encoded video signals corresponding to a respective video image end terminal received from end terminals participating in a multipoint video conference, characterized by the fact that the procedure comprises the following stages:1. Procedimiento para crear una imagen de presencia continua (CP) objetivo codificada según una norma de codificación de vídeo a partir de una pluralidad de señales de vídeo codificadas incluyendo órdenes definidos de macrobloques, comprendiendo cada uno señales de vídeo codificadas correspondientes a una imagen de vídeo respectiva de terminal final recibida desde terminales finales que participan en una conferencia de vídeo multipunto, caracterizado por el hecho de que el procedimiento comprende las siguientes etapas: - decoding said encoded video signals, resulting in end-terminal video images, - descodificar dichas señales de vídeo codificadas, dando como resultado imágenes de vídeo de terminal final, - mezclar espacialmente dichas imágenes de vídeo de terminal final, dando como resultado una pluralidad de imágenes CP compuestas por regiones asociadas respectivamente con cada una de dichas imágenes de vídeo de terminal final, - spatially mixing said end-terminal video images, resulting in a plurality of CP images composed of regions respectively associated with each of said end-end video images, - codificar dichas imágenes CP, - encode said CP images, - rearranging the macroblocks of the encoded CP images, thereby creating said target encoded CP image. - reorganizar los macrobloques de las imágenes CP codificadas, creando de ese modo dicha imagen CP codificada objetivo.
- 8Arrangement in a multipoint control unit (MCU) to create a target CP image encoded according to a video encoding standard from a plurality of encoded video input signals, each corresponding to a respective received end-terminal video image from end terminals participating in a multipoint video conference, characterized by 8. Disposición en una unidad de control multipunto (MCU) para crear una imagen CP objetivo codificada según una norma de codificación de vídeo a partir de una pluralidad de señales de entrada de vídeo codificadas, correspondiéndose cada una con una imagen de vídeo respectiva de terminal final recibida desde terminales finales que participan en una conferencia de vídeo multipunto, caracterizada por - decoders to decode each of said encoded video signals, resulting in end-terminal video images, - descodificadores para descodificar cada una de dichas señales de vídeo codificadas, dando como resultado imágenes de vídeo de terminal final, - a mixing and scaling unit, configured to spatially mix said end-terminal video images, resulting in a plurality of CP images composed of regions respectively associated with each of said end-end video images, - una unidad de mezcla y escalado, configurada para mezclar espacialmente dichas imágenes de vídeo de terminal final, dando como resultado una pluralidad de imágenes CP compuestas por regiones asociadas respectivamente con cada una de dichas imágenes de vídeo de terminal final, - a plurality of encoders, configured to encode said CP images, - una pluralidad de codificadores, configurados para codificar dichas imágenes CP, - a reorganization device, configured to reorganize the macroblocks of the encoded CP images, thereby creating said target encoded CP image. - un dispositivo de reorganización, configurado para reorganizar los macrobloques de las imágenes CP codificadas, creando de ese modo dicha imagen CP codificada objetivo. ES 2 345 893 T3 ES 2 345 893 T3
Independent claims2
80 paragraphs in 7 sections, as filed
ES 2 345 893 T3
DESCRIPTION
Arrangement and procedure to generate continuous presence images.
Field of the invention
The present invention relates to videoconferencing and, in particular, to the generation of Continuous Presence (CP) images in a Multipoint Control Unit (MCU).
Background of the invention
The transmission of moving images in real time is used in various applications such as, for example, video conferencing, meetings over the network, TV broadcasting and video telephony.
However, the representation of moving images requires a large amount of information since digital video is normally described by representing each pixel of an image with 8 bits (1 octet). Such uncompressed video data results in large bit volumes and cannot be transferred over communication networks and conventional transmission lines in real time due to limited bandwidth.
Therefore, allowing real-time video transmission requires a high degree of data compression. However, data compression can compromise the quality of the images. Therefore, great efforts have been made to develop compression techniques that allow real-time transmission of high-quality video over bandwidth-limited data connections.
In video compression systems, the main objective is to represent the video information with the smallest possible capacity. Capacity is defined in bits, either as a constant value or as a unit of bits / time. In both cases, the main goal is to reduce the number of bits.
The most common video encoding procedure is described in the MPEG * and H.26 * standards. Video data goes through four main processes prior to transmission, namely prediction, transformation, quantization, and entropy encoding.
The prediction process significantly reduces the number of bits required for each image of a video sequence to be transferred. Take advantage of the similarity of parts of the sequence to other parts of the sequence. Since the prediction part is known to both the encoder and the decoder, only the difference has to be transferred. This difference typically requires much less capacity to render. The prediction is primarily based on image content from previously reconstructed images where the location of the content is defined by motion vectors. The prediction process is typically done in square block sizes (eg 16x16 pixels).
Video conferencing systems allow the simultaneous exchange of audio, video, and data information between multiple conference sites. Systems known as multipoint control units (MCUs) perform switching functions to allow multiple sites to intercommunicate in a conference. The MCU connects the sites to each other by receiving conference signal frames from the sites, processing the received signals, and retransmitting the processed signals to the appropriate sites. Conference signals include audio, video, data, and control information. In a switched conference, the video signal from one of the conference sites, typically that of the loudest speaker, is broadcast to each of the participants. In a continuous presence conference, video signals from two or more sites are spatially mixed to form a composite video signal for viewing by conference participants. The composite or continuous presence image is a combined image that can include live video streams, still images, menus, or other visual images of the conference participants.
In a typical continuous presence conference, the video display is divided into a composite layout that presents areas or regions (for example, quadrants). The sites are selected in the conference settings from the sites connected in the conference for display in the regions. Common composite distributions include four, nine, or sixteen regions. The layout is selected and then set for the duration of the conference.
Some conference arrangements provide different composite signals or a video mix so that each site can view a different mix of sites. Another arrangement uses a voice activated quadrant selection to associate sites with particular quadrants. This arrangement allows conference participants to view not only fixed video mixing sites, but also a site selected based on voice activity. However, the distribution, in terms of the number of regions or quadrants, is fixed for the conference.
Referring now to Fig. 1, there is shown a schematic diagram of an embodiment of an MCU 10 of the type disclosed in US Patent 5,600,646, the disclosure of which is expressly incorporated herein.
ES 2 345 893 T3 for reference. MCU 10 further includes H.323 functionality as disclosed in US Patent 6,404,745, the disclosure of which is also expressly incorporated herein by reference. Furthermore, the video processing in the MCU has been improved, as will be described in detail in this document. The features described in this document for MCU 10 can be represented on a Tandberg MCU.
MCU 10 includes at least one Network Interface Unit (NIU) 120, at least one Bridge Processing Unit (BPU) 122, one Video Processing Unit (VPU) ) 124, a Data Processing Unit (DPU) 126 and a Main Processing Unit (HPU) 130. In addition to an Industry Standard Architecture (ISA) main control bus 132, the MCU 10 includes a network bus 134, a BPU bus 136, and an X bus 138. The network bus 134 complies with the protocol. Multi-Vendor Integration Protocol (MVIP), while the BPU 136 bus and the X bus are derived from the MVIP specification. The HPU 130 provides a management interface for MCU operations. Each of the above MCU elements is described in detail in the aforementioned US Patents 5,600,646 and 6,404,745.
H.323 functionality is provided by adding a Gateway Processing Unit (GPU) 128 and a modified BPU referred to as a BPU-G 122A. The GPU 128 uses H.323 protocols for call signaling and the creation and control of audio, video, and data streams over an Ethernet or other LAN interface 140 for endpoints. The BPU-G 122A is a BPU 122 that is programmed to process audio, video, and data packets received from the GPU 128.
The following describes the operation of an MCU at a high level, initially for circuit-switched conferencing and later for H.323 packet-switched conferences. In circuit-switched conferencing, digital data frames from H.320 circuit-switched end terminals become available on the network bus 134 through a network interface 142 to an NIU 120. The BPUs 122 process the data frames from the network bus 134 to generate data frames that become available to other BPUs 122 on the BPU 136 bus. The BPUs 122 also extract audio information from the data frames.
The BPUs 122 combine compressed video information and mixed encoded audio information into frames that are placed on the network bus 134 for transmission to respective H.320 terminals.
In cases where AV terminals operate at different transmission rates or with different compression algorithms or are to be mixed into a composite image, multiple video inputs are sent to VPU 124, where the video inputs are decompressed, they are mixed and re-compressed into a single video stream. This single video stream is then returned through the BPU 122, which switches the video stream to the appropriate end terminals.
For a packet-based H.323 conference, the GPU 128 makes audio, video, and data packets available on the network bus 134. The data packets are processed through the DPU 126. The BPU-G 122A processes audio and video packets from network bus 134 to generate audio and video broadcast mixes that are placed on network bus 134 for transmission to respective end terminals through GPU 128. In addition, the BPU-G 122A processes audio and video packets to generate data frames that become available to the BPUs 122 on the BPU 136 bus. In this way, the MCU 14 performs a gateway function whereby the BPUs 122 Common devices and the BPU-G 122A can seamlessly exchange audio and video between H.320 and H.323 terminals.
Having described the components of the MCU 10 that enable the basic bridging functions of a conference, a high-level description of the flexibility provided by the VPU 124 is provided below with reference to the functional block diagram of FIG. 2. At MCU 10, compressed video information from up to five AV terminals that are in the same conference is routed to a particular VPU 124 via the BPU 136 bus. The VPU 124 comprises five video compression processors (VCP0 to VCP4), each presenting a video decoder / encoder pair 102-i, 106-i, and pixel scaling blocks 104-i, 108-i.
A video decoder / encoder pair 102-i, 106-i is assigned to the compressed video information stream associated with each particular conference site. Each video decoder 102-i decodes the compressed video information using the algorithm that matches the encoding algorithm of its associated site. Processing to determine the frame structure, packets, and checksums that may be part of the transmission protocol may be included as part of the video decoder 102i. It should be noted that a processor-encoded video stream can be assigned to multiple sites (eg, a continuous presence application that has more than five sites in the conference). In addition, a decoder / encoder pair 102-i, 106-i can switch between the locations of a conference.
Decoded video information (e.g. pixels) is scaled up or down, if necessary, by a 104-i pixel scaling block to accommodate the pixel resolution requirements of other conference sites that will encode the scaled pixels. For example, a desktop system may encode at a resolution of 256x240 pixels, while an H.320 terminal may require a pixel resolution of 352x288 pixels for a Common Intermediate Format (CIF) image. Other common formats include the Quarter Common Intermediate Format (QCIF).
ES 2 345 893 T3
Intermediate Format) (176x144 pixels), 4CIF (704x576), SIF (352x240), 4SIF (704x480), VGA (640x480), SVGA (800x600) and XGA (1024x768).
The VPU 124 includes a pixel bus 182 and a memory 123. The system disclosed in US Patent 5,600,646 uses a time division multiplexing bus. In particular, each decoder 102-j provides pixels on the pixel bus 182 to memory 123. Each encoder 106-j can retrieve any of the images from memory 123 on the pixel bus for recoding and / or mixing or spatial composition. Another pixel scaling block 108-j is coupled between pixel bus 182 and encoder 106-j to adjust the pixel resolution of the sampled image as necessary.
A continuous presence application will now be described with reference to Figs. 3 and 4. For the sake of simplicity, the end terminals shown are H.320 terminals. In Fig. 3, data from sites 38 arrives via a communication network at the respective NIUs 120. Five sites 38 (A, B, C, D, E) are connected in the conference. Sites A and B are shown connected to a particular NIU 120 that supports multiple codec connections (eg, a T1 interface). The other sites C, D, and E are connected to NIU 120 that support only a single codec connection (eg, an ISDN interface). Each site 38 places one or more bytes of digital data on the network bus 134 as unsynchronized H.221 frame data. The BPUs 122 then determine the byte alignment and H.221 frame structure. This aligned data becomes available to all other units on the BPU 136 bus. The BPUs 122 further extract audio information from the H.221 frames and decode the audio into 16-bit PCM data. The decoded audio data becomes available on the BPU bus 136 for mixing with audio data from other locations in the conference.
The aligned H.221 frames are received by the VPU 124 to be processed by encoder / decoder elements referred to as video compression processors (VCPs). The VPU 124 has 5 VCPs (Fig. 2) which in this example are assigned respectively to the sites A, B, C, D, E. A VCP of the VPU 124 that is assigned to the site E is functionally illustrated in the Fig. 4. The compressed video information (H.261) is extracted from the H.221 frames and decoded by the VCP as an X image. The decoder video X image is placed on the 182 pixel bus through a block of scaled. FIG. 4 shows pixel bus 182 with decoded video frames from each location A, B, C, D, E successively retrieved from memory 123 identified by their respective RAM addresses. The VCP assigned to site E receives the decoded video frames from sites A, B, C, and D which are then tiled (spatially mixed) into a single composite I-picture. The mosaic I image is then encoded as H.261 video in an H.221 frame structure and placed on the BPU bus 136 (Fig. 3) for its BPU processing described above.
As can be seen from the above description, transcoding requires considerable processing resources as the raw pixel data has to be scrambled and then encoded to form a scrambled view or a continuous presence view. To avoid self-views, that is, to prevent CP views from containing an image of the respective participants to which they are transmitted, the MCU must include at least one encoder for each image of a CP view. To allow 16 CPs, the MCU must then include at least 16 encoders.
Summary of the invention
The invention is described in independent claims 1 and 8.
Additional objects and advantages are achieved by the features described in the dependent claims.
Brief description of the drawings
In order to more easily understand the invention, the following description will refer to the accompanying drawings, in which:
Figure 1. Block diagram of an MCU configuration.
Figure 2. Schematic block diagram of an embodiment of a VPU.
Figure 3. Block diagram of an MCU configuration illustrating a data flow for a continuous presence conference.
Figure 4. Block diagram illustrating the tiling of an image in a continuous presence conference.
Figure 5. Block diagram of the arrangements of a group of blocks in a CIF image.
Figure 6. Illustrates the block group layer according to the H.263 standard.
ES 2 345 893 T3
Figure 7. Illustrates the macroblock layer according to the H.263 standard.
Figure 8. Block diagrams illustrating three different continuous presence images used in one embodiment of the present invention.
Figure 9. Schematic block diagram of one embodiment of the present invention.
Figure 10. Schematic flow diagram illustrating an embodiment of the method according to the invention.
Best mode of carrying out the invention
The present invention uses the bit structure of the ITU standard H.26 * to reduce the processing time and requirements on an MCU to generate CP views without views of their own. In order to understand the characteristics of the bit structures that are used, the structure of image blocks according to the H.263 standard will be described below.
According to the H.263 standard, each image is divided into blocks representing 8x8 pixels. The blocks are arranged in macroblocks, which for the luminance part of the pixels are 16 (8x8) blocks and for the chrominance part of the pixels are 4 (2x2) blocks. A group of blocks (GOB, Group Of Blocks) normally represents 22 macroblocks, and the number of GOBs per image is 6 for sub-QCIF, 9 for QCIF and 18 for CIF, 4CIF and 16CIF. GOB numbering is done using a vertical sweep of the GOBs, starting with the upper GOB (number 0) and ending with the lower GOB. An example of the arrangement of GOBs in an image is provided for the CIF image format of Figure 5. The data for each GOB consists of a GOB header followed by data for macroblocks. The data for the GOBs is transmitted for each GOB in increasing number of GOBs. The start of a GOB is identified by a Group of Block Start Code (GBSC). The structure of the GOB layer is shown in figure 6.
The data for each macroblock consists of a macroblock header followed by data for blocks. The structure is shown in figure 7. The COD is only present in images that are not of the “INTRA” type for each macroblock of these images. A bit that when set to "0" indicates that the macroblock is encoded. If set to "1" no additional information is transmitted for this macroblock; in that case, the decoder will treat the macroblock as an INTER macroblock with a motion vector for the entire block equal to zero and no coefficient data.
If the COD is set to "0", the data part of the macroblock includes information from the respective blocks in the macroblock, and this information is represented by motion vectors indicating the position in previous images for which the included pixels are equal. .
Conventionally, the avoidance of own views in a CP image requires a special encoding for each of the participants that involves an encoder for each outgoing data stream in the MCU, as indicated in Figure 2. The present invention uses the macroblock structure of the already encoded video data to obtain a custom mix of a receiver-dependent CP image.
In the following example embodiment of the present invention, consider a conference with five end sites capturing CIP format video images and encoding the images according to the H.263 standard. In the MCU, the data stream of the respective participants is decoded by decoders associated with respective MCU inputs according to the H.263 standard. After decoding, the raw pixel data from the respective participants will be available on an internal MCU bus, ready for mixing and transcoding.
In case of five participants, it is obvious to choose a CP 4 format of the mixed image to be returned to the respective sites. The mixed format can be selected by the MCU according to the "best print" principles described in US Patent Application 10 / 601,095.
According to this exemplary embodiment of the invention, two different CP4 images, a CP image 1 and a CP image 2, are encoded by each respective encoder as illustrated in Figure 8. Image 1 CP includes the images received from the sites 1, 2, 3 and 4, while CP image 2 includes the images received from site 5 in one quadrant leaving the remaining quadrants empty. When the CP images are encoded and the encoded data is arranged in the block system described above, the quadrant boundaries coincide with the macroblock boundaries of the GOBs. Regarding CP image 1, the first 11 macroblocks of the first GOB include data from the image of site 1, while the last 11 macroblocks of the first GOB include data of the image of site 2.
According to the present invention, the MCU rearranges the macroblocks of each CP image according to the receiver. As an example, in the CP image transmitted to location 4, the last 11 macroblocks of each of the last 9 GOBs of CP image 1 are replaced by the first 11 macroblocks of each of the first 9 GOBs of CP image 2 , respectively. This results in a new decoded CP image that includes the image received from site 5 instead of the image received from site 4. This CP image is returned to site 4, thus avoiding self-view at that site.
ES 2 345 893 T3
Corresponding substitutions or reorganizations are carried out for the other four CP images associated respectively with the other sites.
Figure 9 illustrates an example of the internal architecture of an MCU according to the present invention. This architecture is used in accordance with the invention in place of the prior art VPU illustrated in FIG. 2. The pixel bus, memory, and pixel scaling units are replaced for simplicity by a mixing and scaling unit 156. Note that, alternatively, the first data bus 176 and the second data bus 177 of FIG. 9 can be merged to form a common data bus. Also note that the actual implementation may be different and that only the units relevant to the present invention are shown.
The input data streams 171, 172, 173, 174 and 175 are decoded with different decoders 151, 152, 153, 154, 155 respectively, with one decoder for each site. The encoding is carried out according to the encoding standard that is being used, in this case H.263. The decoded data is in the form of PCM data and is made accessible to a Mixing and Scaling Unit (MSU) 156 on the first data bus 176.
The MSU 156 spatially mixes the PCM data 171, 172, 173, 174 from the first, second, third, and fourth locations, creating a first CP image. A second CP image is also created by placing the PCM 175 data from the fifth site in the first quadrant, leaving the remaining quadrants empty or filled with dummy data. The PCM data of the two spatially mixed images then becomes accessible to the next first encoder 157 and second encoder 158 on the second data bus 177.
Encoders 157, 158 extract PCM data from CP images generated by MSU 156 from second data bus 177 and encode each respective image. The result of the encoding process is a plurality of macroblocks assembled in GOBs as described above. In the case of a CIF format according to the H.263 standard, a GOB contains 22 macroblocks and each image consists of 18 GOBs. After encoding the images, the GOBs are consecutively inserted into associated buffers 159, 160, respectively. The size of the buffers 159, 160 must be large enough to accommodate the GOBs of at least one image. The outputs of the buffers 159, 160 are connected to a third data bus 178 which is also connected to the inputs of the reorganization devices 161, 162, 163, 164, 165.
However, the number of bits representing an encoded image is not constant, but can vary substantially depending on the variation of the image content and movements from one image to another. The number of bits also depends on whether the image is INTRA encoded or INTER encoded, ie prediction from neighboring macroblocks in the same image or prediction from a previous image / images (s).
When the full synchronous image encoded data is inserted into the respective buffers, the reorganization devices 161, 162, 163, 164, 165 are ready to reorganize the order of the macroblocks to create the CP images required for the outputs 181, 182 , 183, 184 and 185 associates, respectively. Reorganization devices will be able to identify and isolate macroblocks using the GOB and macroblock headers. The start of each GOB is indicated by a unique start code called a GBSC (Block Group Start Code), followed by a GN (Group Number) indicating the GOB number. In the macroblock headers, COD indicates whether the macroblock is encoded or not. If COD is set to "1", there is no additional information for that macroblock, with the motion vector for the entire block equal to zero and no coefficient data. If COD is "0", there is additional data from the macroblock. Some of the additional data may be of variable length, but the different codes are defined in such a way that the length of each code is well defined.
Since the macroblocks can be identified and temporarily stored in a buffer 159, 160, the reorganization devices 161, 162, 163, 164, 165 can read the macroblocks in any order creating any variant of a CP 4 image from the image of the five sites. As an example, consider the reorganization device 164 that creates the CP image for the fourth site, that is, the output 184 of site 4. The first buffer 159 contains the encoded data of CP image 1, while the second buffer 160 contains the encoded data of CP image 2. The reorganization device 164 extracts the GOBs 1 to 9 from the CP image 1 in the same order as set in the first buffer 159 to create the new CP image. However, when GOB 10 is created, reorganization device 164 identifies and extracts the first 11 macroblocks of GOB 10 from first buffer 159 followed by the first 11 macroblocks of GOB 1 from second buffer 160. In addition, GOB 11 is created by extracting the first 11 macroblocks of GOB 11 from the first buffer 159 followed by the first 11 macroblocks of GOB 2 from the second buffer 160. The remaining 7 GOBs are created similarly, ending with GOB 18 being created by extracting the first 11 macroblocks of GOB 18 from the first buffer 159 followed by the first 11 macroblocks of GOB 9 from the second buffer 160.
The reorganization devices 181, 182, 182, 184, 185 can be pre-programmed to extract the macroblocks from the buffers 159, 160 in a constant order, or the order of the macroblocks can be controlled by a control unit (not shown) allowing a reorganization device creates multiple CP images.
ES 2 345 893 T3
Fig. 10 is a schematic block diagram illustrating an embodiment of the method according to the invention.
The process steps illustrated in FIG. 10 are included in a method of creating a target continuous presence image (CP) encoded according to a video encoding standard from a plurality of encoded video signals including defined orders of macroblocks. Each macroblock comprises encoded video signals corresponding to a respective end terminal video image, received from end terminals participating in a multipoint video conference.
The procedure begins at step 202.
Decoding step 204 is performed first, in which the video signals are decoded to the corresponding end-terminal video images.
Then, in the mixing step 206, the final terminal video images are spatially mixed for a plurality of CP images composed of regions respectively associated with each of the final terminal video images.
Then, in encoding step 208, the CP images are encoded for a plurality of encoded CP images, respectively. This stage establishes the defined orders of macroblocks corresponding to the video coding standard and a fusion of the region boundaries and the macroblock boundaries.
After the three preparatory steps 204, 206 and 208 above, the creation step 210 is carried out. In the creation step 210, the encoded target CP image is created by rearranging said macroblock orders in a predefined or controlled manner.
Advantageously, the creation step 210 comprises replacing a first number m of macroblocks representing a region width into n subsequent GOB numbers representing a region height in a first image of the plurality of CP images encoded by m numbers of macroblocks in n subsequent GOB numbers of a second image of the plurality of encoded CP images.
The procedure ends at step 212.
Alternatively, the three preparatory steps of the process described with reference to FIG. 10, that is, decoding stage 204, mixing stage 206, and encoding stage 210, can be replaced by a single requesting stage (not illustrated). In this requesting stage, the end terminals are requested to encode the respective end terminal video images according to the video standard and a determined resolution, bit rate and scaling.
In its most basic form, the method according to the invention simply includes the creation step 210.
The embodiments described so far have been limited to the creation of CIF format CP 4 images according to the H.263 standard. However, those skilled in the art will appreciate that the basic principles of the present invention can also be applied to other CP images of other formats. As an example, the present invention can in fact be used in CP 9 imaging. Therefore, the image boundaries of each GOB (in the case of CIF, H.263) are after the seventh and fourteenth macroblock (alternatively, after the eighth and fifteenth or the seventh and fifteenth).
In either case, there must be a reorganization device (or at least a reorganization procedure) for each MCU output. The number of encoders depends on the relationship between the number of regions in the CP images and the number of sites that are going to fill in the regions of the CP images. There must be a sufficient number of encoders to generate a region for each site. In the case of CP 4 and eight sites, two encoders will suffice to place a region for each site in each of the eight total quadrants. However, increasing the number of sites to nine, an additional encoder will be required to create a third CP 4 image in which the ninth region can reside.
In a more general embodiment of the invention, no decoding and no scaling / mixing is performed in the MCU. Instead, the end terminals are requested to transmit an image encoded according to a certain standard, resolution, bit rate, and scaling. The macroblocks of the incoming data streams are then loaded directly into the buffers (preferably one for each data stream), and the reorganization devices rearrange the macroblocks according to preprogrammed or controlled procedures, creating CP images without their own view that are transmitted to the respective conference sites. As an example, consider five sites participating in a conference. So a CP 4 view requires the end terminals to place their respective images in one of the quadrants of a full image before being encoded. This is requested from the end terminals along with encoding information such as standard, resolution, bit rate, and scaling. The reorganization devices can then easily reorganize the macroblocks of the incoming data streams when they are present in the respective buffers as described above.
ES 2 345 893 T3
References cited in description
This list of references cited by the applicant is only intended to aid the reader and is not part of the European patent document. Although the utmost care has been taken to carry them out, errors or omissions cannot be excluded and the EPO declines any responsibility in this regard.
Patent documents cited in the description • US 5600646 A [0011] [0012] [0021] • US 6404745 B [0011] • US 6404745 A [0012] • US 10601095 B [0034]
Contents7
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
15 members in 9 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 20040661 | Norway | A | |
| 20040661 | Norway | A | |
| 0571094620040661 | – | – | – |
| NO20040000661 | – | – | – |
Members15
| Document | Office | Kind | |
|---|---|---|---|
| NO20040661D0 | Norway | D0 | |
| WO2005079068A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2005195275A1 | United States of America | A1 | |
| NO320115B1 | Norway | B1 | |
| EP1721462A1 | European Patent Office (EPO) | A1 | |
| CN1918912A | China | A | |
| JP2007522761A | Japan | A | |
| CN100559865C | China | C | |
| US7720157B2 | United States of America | B2 | |
| EP1721462B1 | European Patent Office (EPO) | B1 | |
| AT471631T | Austria | T | |
| ATE471631T1 | Austria | T1 | |
| DE602005021859D1 | Germany | D1 | |
| ES2345893T3This record | Spain | T3 | |
| JP4582659B2 | Japan | B2 |
Numbers
- Publication, DOCDB
- 2345893
- Publication, EPODOC
- ES2345893T
- Application
- 5710946
- Application, DOCDB
- 05710946
- Application, EPODOC
- ES20050710946T
Titles2
- Spanish
- DISPOSICION Y PROCEDIMIENTO PARA GENERAR IMAGENES DE PRESENCIA CONTINUA.
- English
- PROVISION AND PROCEDURE FOR GENERATING IMAGES OF CONTINUOUS PRESENCE.
Classification
- CPC, 2
- H04N7/152
- H04N7/157
- IPC, 1
- H04N7 15