Tiling in video encoding and decoding
Abstract
This record has no abstract on file.
Term
1.5 yearsto projected expiry
Projected expiry 11 April 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
14 claims: 14 independent, 0 dependent
- 1Método compreendendo:arranjar uma primeira imagem e uma segunda imagem para formar uma imagem única, correspondendo a primeira imagem a uma primeira vista de um vídeo em múltiplas vistas e a segunda imagem correspondendo a uma segunda vista do vídeo em múltiplas vistas;gerar infirmação indicando como a primeira imagem e a segunda imagem são combinadas na imagem única, em que a informação gerada indica que pelo menos uma da primeira imagem e da segunda imagem está invertida individualmente numa ou mais de uma direcção horizontal e uma direcção vertical, sendo a primeira imagem não invertida horizontalmente e sendo a segunda imagem invertida horizontalmente, sendo a primeira imagem e a segunda imagem organizadas lado a lado, ou - sendo a primeira imagem não invertida verticalmente e sendo a segunda imagem invertida verticalmente, sendo a primeira imagem e a segunda imagem organizadas de cima a baixo;e ΡΕ2512136 codificar a imagem única e a informação gerada para formar um fluxo de bits, em que o fluxo de bits também inclui informação de upsampling indicando se pelo menos uma da primeira imagem e da segunda imagem é para sofrer upsampling.
- 2Método como definido na Reivindicação 1, em que a informação de upsampling indica adicionalmente que nem a primeira imagem nem a segunda imagem requerem upsampling.
- 3Método como definido na Reivindicação 1, em que a informação de indicação de upsampling indica que pelo menos a primeira imagem requer upsampling;e em que a organização inclui adicionalmente realizar a subamostragem para a primeira vista para produzir a primeira imagem.
- 4Método como definido na Reivindicação 3, em que a informação de indicação de upsampling indica que a primeira imagem e a segunda imagem requerem upsampling;e em que a organização inclui adicionalmente a subamostragem da primeira vista para produzir a primeira imagem, a subamostragem da segunda vista para produzir a segunda imagem, e organizar a primeira imagem e a segunda imagem na imagem única. ΡΕ2512136
- 5Método compreendendo:determinar a partir de um fluxo de bits recebido uma imagem única e informação de upsampling, em que a imagem única inclui uma primeira imagem e uma segunda imagem organizadas como a imagem única, correspondendo a primeira imagem a uma primeira vista de um video em múltiplas vistas e correspondendo a segunda imagem a uma segunda vista do video em múltiplas vistas, em que a informação de upsampling indica se pelo menos uma da primeira imagem e da segunda imagem é para sofrer upsampling;aceder a informação indicando como a primeira imagem e a segunda imagem estão combinadas na imagem única, em que a informação acedida indica que pelo menos uma da primeira imagem e da segunda imagem está individualmente invertida numa ou mais de uma direcção horizontal e uma direcção vertical: - sendo a primeira imagem não invertida horizontalmente e sendo a segunda imagem invertida horizontalmente, sendo a primeira imagem e a segunda imagem organizadas lado a lado, ou - sendo a primeira imagem não invertida verticalmente e sendo a segunda imagem invertida verticalmente, sendo a primeira imagem e a segunda imagem organizadas de cima a baixo;ΡΕ2512136 descodificar a imagem única na primeira imagem e na segunda imagem.
- 6Método como definido na Reivindicação 5, em que a informação de upsampling indica que nem a primeira imagem nem a segunda imagem requerem upsampling.
- 7Método como definido na Reivindicação 5, em que a informação de upsampling indica que pelo menos a primeira imagem requer upsampling;e em que a descodificação inclui adicionalmente inclui uma indicação de um tipo de filtro para utilizar na upsampling.
- 89. Método como definido na Reivindicação 8, em que a informação de upsampling indica adicionalmente um conjunto de coeficientes de filtro, em que cada coeficiente de filtro define um valor para um coeficiente particular de um filtro especificado pelo tipo de filtro.
- 910. Método como definido na Reivindicação 9, em ΡΕ2512136 que a informação de upsampling indica que a primeira imagem e a segunda imagem requerem upsampling;e em que a descodificação inclui adicionalmente fazer a upsampling da primeira imagem para produzir uma primeira vista, e fazer a upsampling da segunda imagem para produzir uma segunda vista.
- 1011. Método como definido em qualquer das Reivindicações 1 a 10, em que a referida informação de upsampling está formatada numa mensagem em concordância com uma sintaxe de alto nivel, em que a mensagem inclui a informação de upsampling.
- 1112. Aparelho configurado para realizar um método de acordo com uma ou mais das Reivindicações 1 a 11.
- 1213. Sinal de video formatado para incluir informação, compreendendo o sinal de vídeo:uma secção de imagem codificada incluindo uma codificação de uma imagem única, incluindo a imagem única uma primeira imagem e uma segunda imagem organizadas na imagem única, correspondendo a primeira imagem a uma primeira vista de um vídeo em múltiplas vistas e correspondendo a segunda imagem a uma segunda vista do video em múltiplas vistas;e uma secção de sinalização incluindo uma ΡΕ2512136 codificação de informação indicando como a primeira imagem e a segunda imagem estão combinadas na imagem única, em que a informação indica que pelo menos uma das múltiplas imagens está individualmente invertida numa ou mais de uma direcção horizontal e uma direcção vertical, - sendo a primeira imagem não invertida horizontalmente e sendo a segunda imagem invertida horizontalmente, sendo a primeira imagem e a segunda imagem organizadas lado a lado, ou - sendo a primeira imagem não invertida verticalmente e sendo a segunda imagem invertida verticalmente, sendo a primeira imagem e a segunda imagem organizadas de cima a baixo;em que a secção de sinalização também indica se pelo menos uma da primeira imagem e da segunda imagem é para sofrer upsampling.
- 1314. Meio legível por processador tendo nele armazenada uma estrutura de sinal de vídeo, compreendendo a estrutura de sinal de vídeo:uma secção de imagem codificada incluindo uma codificação de uma imagem única, incluindo a imagem única uma primeira imagem e uma segunda imagem organizadas na imagem única, correspondendo a primeira imagem a uma primeira vista de um vídeo em múltiplas vistas e ΡΕ2512136 correspondendo a segunda imagem a uma segunda vista do vídeo em múltiplas vistas;e uma secção de sinalização incluindo uma codificação de informação indicando como a primeira imagem e a segunda imagem estão combinadas na imagem única, em que a informação indica que pelo menos uma das múltiplas imagens está individualmente invertida numa ou mais de uma direcção horizontal e uma direcção vertical, - sendo a primeira imagem não invertida horizontalmente e sendo a segunda imagem invertida horizontalmente, sendo a primeira imagem e a segunda imagem organizadas lado a lado, ou - sendo a primeira imagem não invertida verticalmente e sendo a segunda imagem invertida verticalmente, sendo a primeira imagem e a segunda imagem organizadas de cima a baixo;em que a secção de sinalização também indica se pelo menos uma da primeira imagem e da segunda imagem é para sofrer upsampling.
- 1415. Meio legível por máquina tendo nele armazenadas instruções executáveis por máquina que, quando ΡΕ2512136 executadas, implementam um método de acordo com uma ou mais das Reivindicações 1 a 11.
Independent claims14
711 paragraphs in 14 sections, as filed
ORGANIZATION OF MOSAIC IMAGES ON VIDEO ENCODING AND DECODING
CROSSING REFERENCES TO REQUIREMENTS
RELATED
This application claims the benefit of each of (1) United States Provisional Application Serial Number 60/923 014, filed April 12, 2007, entitled Multiview Information, and (2) ) United States Provisional Application Serial Number 60/925 400, filed April 20, 2007, entitled View Tiling in MVC Coding (Case No. PU070103).
TECHNICAL FIELD
The present principles generally relate to video encoding and / or decoding.
BACKGROUND
Video monitor manufacturers may use a tiling or tiling structure
ΡΕ2512136 Different views of a single frame. Views can then be extracted from their respective locations and represented.
DE19619598 patent document entitled
Verfahren zur Speicherung oder Obertratung von stereokopischen Videosignalen discloses a process for storage or stereoscopy.
video signal transmission Patent document EP1581003 monitoring system.
discloses a
SUMMARY
According to a general aspect, access is made to a video image which includes multiple images combined into a single image. Information is accessed by indicating how the multiple images in the accessed video image are combined. The video image is decoded to provide a decoded representation of multiple combined images. The information accessed and the decoded video image are provided as the output.
According to another general aspect, information is generated indicating how multiple images, included in a video image, are combined into a single image. The video image is encoded to provide a representation of
ΡΕ2512136 encoded multiple images combined. The generated information and the encoded video image are available as output.
According to another general aspect, a signal or signal structure includes information indicating how multiple images included in a single video image are combined into the single video image. The signal or signal structure also includes an encoded representation of the multiple combined images.
In another general aspect, access to a video image comprising multiple images combined into a single image is provided. Information that indicates how multiple images are combined in the accessed video image is accessed. The video image is decoded to provide a decoded representation of at least one of multiple images. The information accessed and the decoded representation are provided as the output.
In another general aspect, access to a video image comprising multiple images combined into a single image is provided. Information that indicates how multiple images are combined in the accessed video image is accessed. The video image is decoded to provide a decoded representation of multiple combined images. A user input is received that selects at least one of the following
ΡΕ2512136 Multiple images for presentation. A decoded output of at least one selected image is provided, and the decoded output is provided based on the information accessed, decoded representation, and user input.
Details of one or more implementations are set forth in the accompanying drawings and the description below. Even if described in a particular way, it should be clear that implementations can be configured or performed in various ways. For example, an implementation may be performed as a method, or performed as a device configured to perform a set of operations, or performed as a device that stores instructions for performing a set of operations, or performed as a signal. Other aspects and features will become apparent from the following detailed description taken in conjunction with the accompanying drawings and the claims.
BRIEF DESCRIPTION OF DRAWINGS
Figure 1 is a diagram showing an example of four tiled images in a single frame;
Figure 2 is a diagram showing an example of four inverted and tiled images in a single frame;
ΡΕ2512136
Figure 3 shows a block diagram for a video encoder to which the present principles may be applied, in accordance with an embodiment of the present principles;
Figure 4 shows a block diagram for a video decoder to which the present principles may be applied, in accordance with an embodiment of the present principles;
Figure 5 is a flowchart for a method for encoding images for a plurality of images using the MPEG-4 AVC Standard, in accordance with one embodiment of the present principles;
Figure 6 is a flowchart for a method for decoding images to a plurality of images using the MPEG-4 AVC Standard, in accordance with one embodiment of the present principles;
Figure 7 is a flowchart for a method for encoding images for a plurality of views and depths using the MPEG-4 AVC Standard, in accordance with one embodiment of the present principles;
Figure 8 is a flow chart for a method for decoding images for a plurality of views and depths using the MPEG-4 AVC Standard in
ΡΕ2512136 agreement with an embodiment of the present principles;
Figure 9 is a diagram showing an example of a depth signal, in accordance with an embodiment of the present principles;
Figure 10 is a diagram showing an example of a depth signal added as a title, in accordance with an embodiment of the present principles;
Figure 11 is a diagram showing an example of 5 mosaic views in a single frame, in accordance with one embodiment of the present principles;
Figure 12 is a block diagram for an exemplary Multiple View Video Encoding (MVC) encoder to which the present principles may be applied, in accordance with one embodiment of the present principles;
Figure 13 is a block diagram for an exemplary Multiple View Video Encoding (MVC) decoder to which the present principles may be applied, in accordance with one embodiment of the present principles;
ΡΕ2512136
Figure 14 is a flowchart for a method for processing images for a plurality of views in preparation for encoding the images using the MPEG-4 AVC Standard Multi-Video Video Encoding Extension (MVC) in accordance with one embodiment. realization of these principles;
Figure 15 is a flowchart for a method for encoding images for a plurality of views using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension, in accordance with one embodiment of the present principles;
Figure 16 is a flowchart for a method for processing images for a plurality of views in preparation for decoding the images using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension in accordance with one embodiment. realization of these principles;
Figure 17 is a flowchart for a method for decoding images for a plurality of views using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension, in accordance with one embodiment of the present principles;
Figure 18 is a flow chart for a method for processing images for a plurality of views and depths in preparation for encoding images.
ΡΕ2512136 using the MPEG-4 AVC Standard multi-view video coding (MVC) extension, in accordance with one embodiment of the present principles;
Figure 19 is a flow chart for a method for encoding images for a plurality of views and depths using the MPEG-4 AVC Standard multi-view video coding extension (MVC), in accordance with one embodiment of the present principles;
Figure 20 is a flowchart for a method for processing images for a plurality of views and depths in preparation for image decoding using the MPEG-4 AVC Standard Multi-View Video Encoding (MVC) extension in accordance with a embodiment of the present principles;
Figure 21 is a flow chart for a method for decoding images for a plurality of views and depths using the MPEG-4 AVC Standard multi-view video coding extension (MVC), in accordance with one embodiment of the present principles;
Figure 22 is a diagram showing examples of pixel-level tiling in accordance with one embodiment of the present principles; and
ΡΕ2512136
Figure 23 shows a block diagram for a video processing device to which the present principles may be applied, in accordance with an embodiment of the present principles.
DETAILED DESCRIPTION
Several implementations target methods and apparatuses for visualizing tessellation in video encoding and decoding. It will thus be understood that those skilled in the art will be able to design various arrangements which, although not explicitly described or shown herein, are embodiments of the present principles and are included within their spirit and objectives.
All examples and conditional language defined herein are intended for pedagogical purposes to assist the reader in understanding the present principles and concepts with which the inventors contributed to complement art, and are to be construed as being without limitation to such examples and conditions. specifically described.
In addition, all statements made herein describing principles, aspects, and embodiments of the present principles, as well as their specific examples, are intended to encompass both their structural and functional equivalents. Additionally, there is the
122512136 is the equivalent of such known future-developed equivalents to include both presently and equivalently ie, any developed elements which perform the same function, regardless of structure,
Thus, for example, it will be understood by those skilled in the art that the block diagrams presented herein represent conceptual views of illustrative circuits implementing the present principles. Similarly, it will be understood that any organization chart, flowchart, state transition diagram, pseudo code, and the like represent various processes that may be substantially represented in computer readable media and thus performed by a computer or processor, whether or not such a computer. or processor is explicitly displayed.
The functions of the various elements shown in the figures may be provided by dedicated hardware as well as hardware capable of running software in combination with appropriate software. When provided by a computer, the functions may be provided by a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which may be shared. In addition, explicit use of the term processor or controller should not be construed as referring solely to hardware capable of running software.
ΡΕ2512136 may implicitly include, without limitation, digital signal processing hardware (DSP), read-only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage.
conceptual switches
Other conventional and / or custom hardware may also be included. Similarly, any shown in the figures are only their functions can be carried out by program logic operation, by dedicated logic, by program control interaction and by dedicated logic, or even manually, the particular technique being selectable by the user. implementer as understood more specifically from the context.
In the claims set forth herein, any element expressed as a means of performing a specified function is intended to encompass any form of
<td rowspan="2">accomplish combination</td><td rowspan="2">That in</td><td rowspan="2">occupation elements</td><td colspan="2">including by</td><td rowspan="2">example, a) a that perform this</td>
<td>in</td><td>circuit</td>
<td>function or</td><td>B)</td><td>software</td><td>in</td><td>any</td><td>including</td>
therefore, firmware, microcode, or the like, combined with appropriate circuits to perform this software to perform the function. The present principles as defined by such claims reside in the fact that the functionalities provided by the various means referred to are combined and brought together in the form that the claims claim. It is thus seen that any means which may
Providing such features are equivalent to those set forth herein.
Reference in the specification to an embodiment (or an implementation) or embodiment (or implementation) of the present principles means that a particular functionality, structure, feature, and so forth described in connection with the embodiment is included. in at least one embodiment of the present principles. Thus, where the phrase appears in one embodiment or in the embodiment appearing in various places throughout the specification are not necessarily referring to the same embodiment.
It is to be understood that the use of the terms and / or and at least one of, for example, in cases of A and / or B and at least one of A and B, is intended to include the selection of only the first listed option. (A), or selecting only the second listed option (B), or selecting
<td>selection</td><td>of both options</td><td>(A and</td><td>B) . How</td><td>An example</td>
<td>additional</td><td>, in cases of A, B,</td><td>and / or</td><td>C and at</td><td>least one of</td>
<td>A, B, and</td><td colspan="3">C, such expression has the intention of</td><td>cover the</td>
<td>selection</td><td>only from the first</td><td>option</td><td>enumerated</td><td>(A), or</td>
<td>selection</td><td>only from the second</td><td>option</td><td>enumerated</td><td>(B), or</td>
<td>selection</td><td>only from the third</td><td>option</td><td>enumerated</td><td>(C), or</td>
<td>selection</td><td>only from the first and</td><td colspan="2">of the second options</td><td>enumerated</td>
(A and B), or selecting only the first and third listed options (A and C), or selecting only the second
122512136 and the third listed options (B and C), or the selection of all three options (A and B and C). This can be extended, as readily apparent by one of ordinary skill in the art and related arts, to as many items as enumerated.
Furthermore, it is to be understood that although one or more embodiments of the present principles are described herein with respect to the MPEG-4 AVC standard, the present principles are not limited solely to this standard and thus may be used with respect to others its standards, recommendations, and extensions, particularly its standards, recommendations, and video encoding extensions, including MPEG-4 AVC standard extensions, while maintaining the spirit of the present principles.
Additionally, it is to be understood that while one or more other embodiments of the present principles are described herein with respect to the MPEG-4 AVC standard multi-view video encoding extension, the present principles are not limited solely to this extension and / or this standard and thus may be used with respect to its other standards, recommendations, and video coding extensions related to multi-view video coding, while maintaining the spirit of the present principles. Multiple view video coding (MVC) is the compression structure for coding multiple view sequences. A Multiple View Video Encoding (MVC) sequence is a
ΡΕ2512136 Set of two or more video sequences that capture the same scene from a different point of view.
Likewise, it is to be understood that while one or more other embodiments of the present principles are described herein using in-depth information with respect to the video content, the present principles are not limited to such embodiments and thus , other embodiments that do not use in-depth information may be implemented while maintaining the spirit of the present principles.
Additionally, as used herein, high level syntax refers to the syntax present in the bit stream that resides hierarchically above the macroblock layer. For example, high level syntax, as used herein, may refer to, but is not limited to, slice header level syntax, Enhancement Supplemental Information (SEI) level, Image Parameter Set level. (PPS), Supplemental Sequence Parameter Set (SPS) level, View Parameter Set (VPS), and Network Abstraction Layer (NAL) unit header level.
In the current implementation of International Multi-View Video (MVC) encoding
Organization
Electrotechnical for Standardization / International
Commission (ISO / IEC) Moving Picture
Experts Group-4 (MPEG-4) Part 10 Advanced Video Coding
ΡΕ2512136 (AVC) Standard / International Telecommunication Union,
Telecommunication Sector (ITU-T) H.264 Recommendation (hereafter the MPEG-4 AVC Standard), the reference software predicts multiple views by encoding each view with a single encoder and taking cross-view references into account. Each view is encoded as a bit stream separated by the encoder at its original resolution and subsequently all bit streams are combined to form a single bit stream which is then decoded. Each view produces a separate decoded YUV output.
Another approach to multiple view prediction involves grouping a set of views into pseudo views. In an example of this approach, we can mosaic the images from each N view of a total of M views (sampled at the same time) into a larger frame or a supersample with possible under-sampling or other operations. Turning to Figure 1, an example of four tiled views in a single frame is generally indicated by reference numeral 100. All four views are in their normal orientation.
Turning to Figure 2, an example of four inverted and tiled views in a single frame is generally indicated by reference numeral 200. The view in the upper left is in its normal orientation. The view in the upper right is reversed horizontally. The bottom left view is
ΡΕ2512136 inverted vertically. The bottom right view is reversed both horizontally and vertically.
Thus, if there are four views, then an image from each view is organized into a super frame like a mosaic. This results in a single unencoded input sequence with a higher resolution.
Alternatively, we can sub-sample the image to produce a lower resolution. Thus we created multiple sequences in which each includes different views which are arranged together in tessellation. Each of these sequences then forms a pseudo view, where each pseudo view includes N different mosaic views. Figure 1 shows a pseudo view, and Figure 2 shows another pseudo view. These pseudo views can then be encoded using coding standards of
<td>video Pattern</td><td>such as MPEG-4 AVC.</td><td>the standard</td><td>ISO / IEC</td><td>MPEG-2 and</td><td>O</td>
<td></td><td>One yet another</td><td>approach</td><td>to the</td><td>prediction</td><td>in</td>
<td colspan="2">multiple views involves</td><td colspan="3">simply code</td><td>at</td>
different views independently using a new pattern and, after decoding, tile the views as required by the reader.
Additionally, in another approach, the views may also be tiled in a pixel by pixel manner. For example, in a superview that is made up of four views, the pixel (x, y) can be from view 0, while
ΡΕ2512136 pixel (χ + 1, y) can be from view 1, pixel (x, y + 1) can be from view 2, and i pixel (x + 1, y + 1) can be from view 3.
Many monitor manufacturers use such a structure to arrange or mosaic different views in a single frame and then extract the views from their respective locations and represent them. In such cases, there is no standard way to determine if the bitstream has such a property. Thus, if a system uses the method of tiling images of different views into a larger frame, then the method of extracting the different views is itself.
However, there is no standard way to determine if the bitstream has such a property. We propose a high level syntax in order to make it easier for the representative or reader to extract such information to aid in representation or other post processing. It is also possible that the subimages have different resolutions and some upsampling is required to eventually represent the view. 0 The user may also wish to have the upsampling method indicated in the high level syntax. Additionally, parameters can also be transmitted to change the depth of focus.
In one embodiment, we propose a new Enhancement Supplemental Information (SEI) message to signal multiple view information in a bit stream.
ΡΕ2512136 compatible with MPEG-4 AVC Standard where each image includes sub-images belonging to a different view. The embodiment is intended, for example, for the easy and convenient viewing of multi-view video bitstreams on three-dimensional (3D) monitors that may use such a structure. The concept may be extended to other video coding standards and recommendations by signaling such information using high level syntax.
Furthermore, in one embodiment, we propose a signaling method of arranging views before they are sent to the multi-view video encoder and / or decoder. Advantageously, the embodiment may lead to a simplified implementation of multi-view coding, and may benefit coding efficiency. Certain views may be put together and form a pseudo or super view and then the tiled super views are treated as a normal view by a common multi-view video encoder and / or decoder, for example, as by the current implementation of the MPEG-4 AVC Standard of multi-view video encoding. A new flag is proposed in the extension of the Multiple View Video Encoding Supplementary Sequence Parameter Set (SPS) to signal the use of the pseudo-view technique. The embodiment is intended for easy and convenient viewing of multi-view video bitstreams on 3D monitors that can
ΡΕ2512136 use such a structure.
Encoding / decoding using a single view video encoding / decoding standard / recommendation
Current Implementation of Multi-Video Encoding (MVC) based on International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group-4 (MPEG-4) Part 10 Advanced Video Coding (AVC) Standard / International Telecommunication Union, Telecommunication Sector (ITU-T) H.264 Recommendation (hereafter the MPEG-4 AVC Standard), reference software predicts multiple views by encoding each view with a single encoder and taking cross-view references into account. Each view is encoded as a bit stream separated by the encoder at its original resolution and subsequently all bit streams are combined to form a single bit stream which is then decoded. Each view produces a separate decoded YUV output.
Another approach to multiple view prediction involves tiling images from each view (sampled at the same time) into a larger frame or a super frame with a possible subsampling operation. Returning to Figure 1, an example of four mosaic views in a single frame is generally indicated by reference numeral 100.
ΡΕ2512136
Figure 2, an example of four inverted and tiled views in a single frame is generally indicated by reference numeral 200. Thus, if there are four views, then an image from each view is organized into a super frame as a mosaic. This results in a single unencoded input sequence with a higher resolution. This signal can then be encoded using existing video coding standards such as the ISO / IEC MPEG-2 Standard and the
MPEG-4 AVC.
Yet another approach to multiple view prediction simply involves coding the different views independently using a new pattern and, after decoding, tiling the views as required by the reader.
Many monitor manufacturers use such a structure to arrange or mosaic different views in a single frame and then extract the views from their respective locations and represent them. In such cases, there is no standard way to determine if the bitstream has such a property. Thus, if a system uses the method of tiling images of different views into a larger frame, then the method of extracting the different views is itself.
Turning to Figure 3, a video encoder capable of performing video encoding accordingly
ΡΕ2512136 with the MPEG-4 AVC standard is generally indicated by reference number 300.
Video encoder 300 includes a frame sort buffer 310 having an output in signal communication with a noninverting input of a combiner 385. An output of combiner 385 is connected in signal communication with a first input of a transformer and quantifier 325. An output of transformer and quantizer 325 is connected in signal communication with a first input of an entropy encoder 345 and a first input of an inverter transformer and inverter quantizer 350. An output of entropy encoder 345 is connected in signal communication with a first noninverting input of a combiner 390. An output of combiner 390 is connected in signal communication with a first input of an output buffer 335.
A first output of an encoder controller 305 is connected in signal communication with a second input of frame sort buffer 310, a second input of inverter transformer and inverter quantizer 350, an input of an image type decision module. 315, an input of a macroblock type (MB-type) decision module 320, a second input of an intra prediction module 360, a second input of an unlock filter 365, a first input of a motion compensator 370, a first
2512136 input of a motion estimator 375, and a second input of a reference image buffer 380.
A second output of encoder controller 305 is connected in signal communication with a first input of an Enhancement Supplemental Information inserter.
<td>(CES) 330,</td><td>an</td><td>Monday</td><td>input from</td><td>transformer</td><td>and</td>
<td>quantifier</td><td> 325,</td><td colspan="2">a second entry</td><td>of the encoder</td><td>in</td>
<td>entropy 345,</td><td>an</td><td>Monday</td><td>entrance of</td><td>buffer memory</td><td>in</td>
<td>output 335, and</td><td>an</td><td>input</td><td colspan="2">Inserter</td><td>in</td>
Supplemental Sequence (SPS) and Image Parameter Set (PPS) Parameters 340.
A first output of image type decision module 315 is connected in signal communication with a third input of a frame buffer 310. A second output of image type decision module 315 is connected in signal communication. signal with a second input of a macro block type decision module 320.
A Supplemental Sequence Parameter Set (SPS) and Picture Parameter Set (PPS) inserter 340 output is connected in signal communication with a third noninverting input of combiner 390. A SEI 330 inserter output is connected. in signal communication with a second noninverting combiner input
390 .
ΡΕ2512136
An inverter quantizer and inverter transformer output 350 is connected in signal communication with a first noninverting input of a combiner 319. An output of combiner 319 is connected in signal communication with a first input of the prediction module 360 and a first unlock filter input 365. An unlock filter output 365 is connected in signal communication with a first input of a reference picture buffer 380. A reference picture buffer output 380 is connected in signal communication with a second motion estimator input 375 and a first input of a motion compensator 370. A first output of motion estimator 375 is connected on a signal communication. signal with a second motion compensator input 370. A second output of the motion estimator 375 is connected in signal communication with a third entropy encoder input 345.
A motion compensator output 370 is connected in signal communication with a first input of a switch 397. An intra prediction module output 360 is connected in signal communication with a second input of switch 397. A module output Macro block type decision 320 is connected in signal communication with a third input of switch 397 to provide a control input for switch 397. The third switch input 397 determines whether or not the switch data input (compared to
(Ie, the third input) is to be provided by the motion compensator 370 or the prediction module 360. The output of the switch 397 is connected in signal communication with a second noninverting input of the combiner 319 and with an inverter input from combiner 385.
Frame sort buffer 310 and encoder controller 105 inputs are available as encoder 300 inputs to receive an input image 301. In addition, an Enhancement Supplemental Information (SEI) 330 input is available. as an input from encoder 300 to receive metadata. An output buffer output 335 is available as an output from encoder 300 to output a bit stream.
Turning to Figure 4, a video decoder capable of performing video decoding in accordance with the MPEG-4 AVC standard is generally indicated by reference numeral 400.
Video decoder 400 includes an input buffer 410 having an output turned on in the communication.
<td>signal</td><td>with</td><td>an</td><td>first</td><td>input from</td><td>decoder</td><td>in</td>
<td>entropy</td><td> 445 .</td><td>An</td><td>first</td><td>output from</td><td>decoder</td><td>in</td>
<td>entropy</td><td colspan="2">445 is</td><td>turned on</td><td>Communication</td><td>signal with</td><td>an</td>
<td>first</td><td colspan="2">input</td><td>on one</td><td colspan="2">inverter transformer</td><td>and</td>
inverter quantifier 450. One transformer output
ΡΕ2512136 inverter and quantizer inverter 450 is connected in signal communication with a second noninverting input of a combiner 425. An output of combiner 425 is connected in signal communication with a second input of an unlock filter 465 and a first input of a prediction module 460. A second output of the unlocking filter 465 is connected in signal communication with a first input of a reference picture buffer 480. An output of the reference picture buffer 480 is connected in signal communication with a second input of a motion compensator 470.
A second entropy decoder 445 output is connected in signal communication with a third motion compensator input 470 and a first unlocking filter input 465. A third entropy decoder 445 output is connected in signal communication with an input of a decoder controller 405. A first output of decoder controller 405 is connected in signal communication with a second entropy decoder input 445. A second output of decoder controller 405 is connected in signal communication with a second input of inverter transformer and inverter quantizer 450. A third output of decoder controller 405 is connected in signal communication with a third input of unlocking filter 465 A fourth output of decoder controller 405 is connected in signal communication with a second input of intra prediction module 460,
ΡΕ2512136 with a first motion compensator input 470, and with a second input of the reference image buffer 480.
A motion compensator output 470 is connected in signal communication with a first input of a switch 497. An output of the prediction module 460 is connected in signal communication with a second input of switch 497. A output of the switch 497 is connected in signal communication with a first noninverting input of combiner 425.
An input buffer 410 input is available as a decoder input 400 to receive an input bit stream. A first unlock filter output 465 is available as a decoder output 400 to output an output image.
Turning to Figure 5, an exemplary method for encoding images for a plurality of views using the MPEG-4 AVC Standard is indicated generally by reference numeral 500.
Method 500 includes a start block 502 that passes control to a function block 504. Function block 504 arranges each view at a particular instant in time as a sub-image in tessellation format, and passes control to a function block 506. Function block 506 initializes a syntax element
ΡΕ2512136 num_coded_views_minusl, and passes control to a function block 508. Function block 508 initializes org_pic_width_in_mbs_minusl and org_pic_height_in_mbs_minusl syntax elements, and passes control to function block 510. Function block 510 zeroes a variable i equals , and passes control to decision block 512. Decision block 512 determines whether or not variable i is less than the number of views. If it is, then control is passed to a function block 514. Otherwise, control is passed to a function block 524.
Function block 514 initializes a view_id [i] syntax element, and passes control to a function block 516. Function block 516 initializes a num_parts [view_id [i]] syntax element, and passes control to a function block 518. Function block 518 initializes a variable j equal to zero, and passes control to a decision block 520. Decision block 520 determines whether or not the current value of variable j is less than the current value of the num_parts [view_id [i]] syntax element. If so, then control is passed to a function block 522. Otherwise, control is passed to a function block 528.
Function block 522 initializes the following syntax elements, increments variable j, and then returns control to decision block 520:
ΡΕ2512136 depth_flag [view_id [i]] [j]; flip_dir [view_id [i]] [j];
loc_left_offset [view_id [i]] [j]; loc_top_offset [view_id [i]] [j]; frame_crop_offset [view_id [i]] [j]; frame_crop_right_offset [view_id [i]] [j];
frame_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] [j].
Function block 528 initializes an upsample_view_flag [view_id [i]] syntax element, and passes control to a decision block 530. Decision block 530 determines whether or not the current value of the upsample_view_f lag syntax element [view_id [i]] is equal to one. If so, then control is passed to a 532 function block. Otherwise control is passed to a decision block.
534 .
Function block 532 initializes an upsample_filter [view_id [i]] syntax element, and passes control to decision block 534.
Decision block 534 determines whether or not the current value of the upsample_filter [view_id [i]] syntax element is three. If so then the
<td>control is</td><td>past to</td><td>a function block</td><td> 536</td><td>Case</td>
<td>contrary,</td><td>the control is</td><td>passed to a block</td><td>in</td><td>occupation</td>
<td> 540 .</td><td></td><td></td><td></td><td></td>
Function block 536 initializes the following
ΡΕ2512136 syntax elements and passes control to a 538 function block: vert_dim [view_id [i]]; hor_dim [view_id [i]]; and quantizer [view _id [i]].
Function block 538 initializes the filter coefficients for each YUV component, and passes control to function block 540.
Function block 540 increments variable i, and returns control to decision block 512.
Function block 524 writes these syntax elements in at least one Supplemental Sequence Parameter Set (SPS), Image Parameter Set (PPS), Enhancement Supplemental Information (SEI) message, Layer unit header Network Abstraction (NAL), and slice header, and passes control to a function block 526. Function block 526 encodes each figure using the MPEG-4 AVC Standard or other single view codec, and passes control to an end block.
599.
Turning to Figure 6, an exemplary method for decoding images for a plurality of views using the MPEG-4 AVC Standard is generally indicated by reference numeral 600.
Method 600 includes a start block 602 that passes control to a function block 604.
ΡΕ2512136 Function 604 recognizes the following syntax elements from at least one Supplemental Sequence Parameter Set (SPS), Image Parameter Set (PPS), Enhancement Supplemental Information (SEI) message, unit header Network Abstraction Layer (NAL), and slice header, and passes control to a function block 606. Function block 606 recognizes a num_coded_views_minusl syntax element, and passes control to function block 608. Function block 608 recognizes org_pic_width_in_mbs_minusl and org_pic_height_in_mbs_minusl syntax elements, and passes control to a block of function 610. Function block 610 initializes a variable i equal to zero, and passes control to decision block 612. Decision block 612 determines whether or not variable i is less than the number of views. If so, then control is passed to a function block 614. Otherwise, control is passed to a function block 624.
Function block 614 recognizes a view_id [i] syntax element, and passes control to a function block 616. Function block 616 recognizes a syntax element num_parts_minusl [view_id [i]], and passes control to function block 618. Function block 618 initializes a variable j equal to zero, and passes control to decision block 620. Decision block 620 determines whether or not the current value of variable j is less than the value
ΡΕ2512136 current syntax element num_parts [view_id [i]]. If yes, control is passed to a function block 622. Otherwise control is passed to a function block.
628 .
Function block 622 recognizes the following syntax elements, increments variable j, and then returns control to decision block 620:
depth_flag [view_id [i]] [j]; flip_dir [view_id [i]] [j];
loc_left_offset [view_id [i]] [j]; loc_top_offset [view_id [i]] [j]; frame_crop_left_offset [view_id [i]] [j]; frame_crop_right_offset [view_id [i]] [j];
frame_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] [j].
Function block 628 recognizes an upsample_view_flag [view_id [i]] syntax element, and passes control to a decision block 630. Decision block 630 determines whether or not the current value of the syntax element upsample_view_flag [view_id [i]] is equal to one. If yes, control is passed to a function block 632. Otherwise control is passed to a decision block.
634 .
Function block 632 recognizes an upsample_filter [view_id [i]] syntax element, and passes control to decision block 634.
ΡΕ2512136 decision block 634 determines whether or not the current value of the upsample_filter [view_id [i]] syntax element is equal to three. If so, then control is passed to function block 636. If
<td>contrary, 640.</td><td>control is passed</td><td>for</td><td>one</td><td>function block</td>
<td></td><td>0 function block 636</td><td>do the</td><td colspan="2">recognition of</td>
<td>following</td><td>syntax elements and</td><td>goes by</td><td>O</td><td>control for a</td>
<td>block</td><td>Function 638:</td><td colspan="2">vert.</td><td>_dim [view_id [i]];</td>
hor_dim [view_id [i]]; and quantizer [view_id [i]].
Function block 638 recognizes the filter coefficients for each YUV component, and passes control to function block 640.
Function block 640 increments variable i, and returns control to decision block 612.
Function block 624 decodes each image using the MPEG-4 AVC Standard or other single-image codec, and passes control to function block 626. Function block 626 separates each view of the image using high-level syntax, and passes control to an end block 699.
Turning to Figure 7, an exemplary method for encoding images for a plurality of views and
ΡΕ2512136 depths using the MPEG-4 AVC Standard is generally indicated by reference number 700.
Method 700 includes a start block 702 that passes control to a function block 704. Function block 704 arranges each view and corresponding depth at a particular point in time as a tiled sub-image, and passes control for a function block 706. Function block 706 initializes a num_coded_views_minusl syntax element, and passes control to a function block 708. Function block 708 initializes syntax elements org_pic_width_in_mbs_minusl and org_pic_height_in_mbs_minusl, and passes control to function block 710. Function block 710 initializes a variable i equal to zero, and passes control to decision block 712. Block Decision 712 determines whether or not variable i is less than the number of views. If so, then control is passed to a function block 714. Otherwise, control is passed to a function block 724.
Function block 714 initializes a view_id [i] syntax element, and passes control to a function block 716. Function block 716 initializes a num_parts [view_id [i]] syntax element, and passes control to a function block 718. Function block 718 initializes a variable j equal to zero, and passes control to decision block 720. Decision block 720 determines whether or not the current value of variable j is less than the value
ΡΕ2512136 current syntax element num_parts [view_id [i]]. If
<td>Yes,</td><td>so the control</td><td>it is past</td><td>for one</td><td>block</td><td>of function</td>
<td> 722 .</td><td>Otherwise, the</td><td>control is</td><td>past</td><td colspan="2">for a block of</td>
<td colspan="2">function 728.</td><td></td><td></td><td></td><td></td>
<td></td><td>0 block of</td><td>722 function</td><td colspan="2">initializes the</td><td>following</td>
syntax elements, increments variable j, and returns control to decision block 720:
depth_flag [view_id [i]] [j]; flip_dir [view_id [i]] [j];
loc_left_offset [view_id [i]] [j]; loc_top_offset [view_id [i]] [j]; frame_crop_left_offset [view_id [i]] [j]; frame_crop_right_offset [view_id [i]] [j];
frame_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] [j].
Function block 728 initializes an upsample_view_flag [view_id [i]] syntax element, and passes control to a decision block 730. Decision block 730 determines whether or not the current value of the upsample_view_f lag syntax element [ view_id [i]] is equal to one. If so, then control is passed to a 732 function block. Otherwise control is passed to a decision block.
734 .
Function block 732 initializes an upsample_filter [view_id [i]] syntax element, and passes control to decision block 734.
ΡΕ2512136 decision block 734 determines whether or not the current value of the upsample_filter [view_id [i]] syntax element is equal to three. If so, then control is passed to a function block 736. Otherwise control is passed to a function block.
740 .
Function block 736 initializes the following syntax elements and passes control to function block 738:
vert_dim [view_id [i]]; hor_dim [view_id [i]]; and quantizer [view_id [i]].
Function block 738 initializes the filter coefficients for each YUV component, and passes control to function block 740.
Function block 740 increments variable i, and returns control to decision block 712.
Function block 724 writes these syntax elements into at least one Supplemental Sequence Parameter Set (SPS), Image Parameter Set (PPS), Enhancement Supplemental Information (SEI) message, Layer 2 unit header. Network Abstraction (NAL), and slice header, and passes control to a function block 726. Function block 726 encodes each image using the MPEG-4 AVC Standard or other codec.
Única2512136 single view, and passes control to an end block 799.
Turning to Figure 8, an exemplary method for decoding images for a plurality of views and depths using the MPEG-4 AVC Standard is generally indicated by reference numeral 800.
Method 800 includes a start block 802 that passes control to a function block 804. Function block 804 recognizes the following syntax elements from at least one Supplemental Sequence Parameter Set (SPS), Image Parameter Set (PPS), Enhancement Supplemental Information (SEI) message, Network Abstraction Layer (NAL) unit header, and slice header, and passes control to an 806 function block. Function block 806 recognizes a num_coded_views_minusl syntax element, and passes control to function block 808. Function block 808 recognizes the org_pic_width_in_mbs_minusl and org_pic_height_in_mbs_minusl syntax elements, and passes control to a block of function 810. Function block 810 initializes a variable i equal to zero, and passes control to decision block 812. Decision block 812 determines whether or not variable i is less than the number of views. If so, then control is passed to function block 814. Otherwise, control is passed to function block 824.
ΡΕ2512136 function block 814 recognizes a view_id [i] syntax element, and passes control to a function block 816. function block 816 recognizes a syntax element num_parts_minusl [view_id [i]], and passes control to function block 818. Function block 818 initializes a variable j equal to zero, and passes control to decision block 820. Decision block 820 determines whether or not the current value of variable j is less than the current value of the num_parts [view_id [i]] syntax element. If so, then control is passed to a function block 822. Otherwise, control is passed to a function block 828.
Function block 822 recognizes the following syntax elements, increments variable j, and then returns control to decision block 820:
depth_flag [view_id [i]] [j]; flip_dir [view_id [i]] [j];
loc_left_offset [view_id [i]] [j]; loc_top_offset [view_id [i]] [j]; frame_crop_left_offset [view_id [i]] [j]; frame_crop_right_offset [view_id [i]] [j];
frame_crop_top_offset [view_id [i]] [j]; and frame_crop_bottom_offset [view_id [i]] [j].
Function block 828 recognizes an upsample_view_flag [view_id [i]] syntax element, and passes control to a decision block 830. The decision block
ΡΕ2512136
830 determines whether or not the current value of the upsample_view_flag [view_id [i]] syntax element is equal to one. If so, then control is passed to function block 832. Otherwise, control is passed to decision block 834.
Function block 832 recognizes an upsample_filter [view_id [i]] syntax element, and passes control to decision block 834.
Decision block 834 determines whether or not the current value of the upsample_filter [view_id [i]] syntax element is equal to three. If yes, then control is passed to function block 836. Otherwise control is passed to function block
840 .
Function block 836 recognizes the following syntax elements and passes control to function block 838: vert_dim [view_id [i]]; hor_dim [view_id [i]]; and quantizer [view_id [i]].
Function block 838 recognizes filter coefficients for each YUV component, and passes control to function block 840.
Function block 840 increments variable i, and returns control to decision block 812.
ΡΕ2512136 function block 824 decodes each image using the MPEG-4 AVC Standard or other single-view codec, and passes control to a function block 826. Function block 826 separates each view and corresponding image depth using the high level, and passes control to a function block 827. Function block 827 potentially synthesizes views using the extracted view and depth signals, and passes control to an end block 899.
With respect to the depth used in Figures 7 and 8, Figure 9 shows an example of a depth signal 900, where depth is provided as a pixel value for each corresponding location of an image (not shown). Additionally, Figure 10 shows an example of two depth signals included in a tiled array 1000. The upper right portion of the tile 1000 is a depth signal having depth values corresponding to the image in the upper right corner of the tile 1000. The lower right portion of the tile 1000 is a depth sign having depth values corresponding to the image in the lower left corner of the tile layout
1000 .
Turning to Figure 11, an example of 5 tiled views in a single frame is indicated generally by reference numeral 1100. The four
ΡΕ2512136 top views are in a normal orientation. The fifth view is also in a normal orientation, but is divided into two portions along the bottom part of the tessellation 1100. A left portion of the fifth view shows the top of the fifth view, and a right portion of the fifth view shows the bottom. from the fifth view.
Encoding / decoding using a multi-view video encoding / decoding standard / recommendation
Views (MVC)
Returning
Example encoding is in Figure 12, a Multiples Video generally indicated by reference numeral 1200. Encoder 1200 includes a combiner
1205 having an output connected in signal communication to a
<td>input from</td><td>one</td><td>transformer</td><td>1210. One</td><td>output</td><td>of</td>
<td>transformer</td><td> 1210</td><td>is on</td><td>Communication</td><td>signal</td><td>The</td>
<td>an entry</td><td>of</td><td>quantifier</td><td>1215. One</td><td>output</td><td>of</td>
<td>quantifier</td><td> 1215</td><td>is on</td><td>Communication</td><td>signal</td><td>The</td>
<td>an entry</td><td>on one</td><td>encoder</td><td colspan="3">entropy 1220 and a</td>
input of an inverter quantifier 1225. An output of the inverter quantifier 1225 is connected in signal communication to an input of an inverter transformer 1230. An output of inverter transformer 1230 is connected in signal communication to a first noninverting input of a combiner 1235. A combiner output 1235 is connected in signal communication to an input of an intra-predictor 1245 and an input of a filter of
ΡΕ2512136 unlock 1250. An unlock filter output 1250 is connected in signal communication to an input of a reference image store 1255 (for view i). A reference picture store output 1255 is connected in signal communication to a first input of a motion compensator 1275 and a first input of a motion estimator 1280. A motion estimator output 1280 is connected in signal communication to a second motion compensator input
1275.
An output from a reference image store 1260 (for other views) is connected in signal communication to a first input of a disparity estimator 1270 and a first input of a disparity compensator 1265. A output of a disparity estimator 1270 is connected in signal communication to a second input of the disparity compensator 1265.
An entropy decoder output 1220 is available as an output from encoder 1200. A noninverting input from combiner 1205 is available as an input from encoder 1200, and is connected in signal communication to a second input of the disparity estimator 1270, and to a second input of motion estimator 1280. An output of a switch 1285 is connected in signal communication to a second noninverting input of combiner 1235 and an inverter input of combiner
1205. Switch 1285 includes a first input connected
No. 2512136 in signal communication at a motion compensator output 1275, a second input connected in signal communication at a disparity compensator output 1265, and a third input connected in signal communication at an output of intrapredictor 1245.
A mode decision module 1240 has an output connected to switch 1285 to control which input is selected by switch 1285.
Turning to Figure 13, an exemplary Multipoint Video Encoding (MVC) decoder is indicated generally by reference numeral 1300. Decoder 1300 includes an entropy decoder 1305 having an output connected in signal communication to an input of a quantizer inverter 1310. An inverter quantizer output is connected in signal communication to an input of an inverter transformer 1315. An output of inverter transformer 1315 is connected in signal communication to a first noninverting input of a combiner 1320. An output of combiner 1320 is connected in signal communication to an input of an unlock filter 1325 and an input of an intra predictor 1330. An unlock filter output 1325 is connected in signal communication to a reference image store input 1340 (for view i). An output of reference image store 1340 is connected in signal communication to a first input of a motion compensator 1335.
ΡΕ2512136
<td colspan="4">An output from a storage of</td><td>Image</td><td>in</td>
<td>reference</td><td>1345 (for other</td><td>views)</td><td>it is</td><td>on</td><td>at</td>
<td>Communication</td><td>signal to a</td><td>first</td><td colspan="2">input from</td><td>one</td>
<td>compensating</td><td>of disparity 1350.</td><td></td><td></td><td></td><td></td>
An entropy encoder input 1305 is available as an input to decoder 1300 to receive a residue bit stream. In addition, a mode module input 1360 is also available as an input for decoder 1300 to receive control syntax for controlling which input is selected by switch 1355. Additionally, a
<td>Monday</td><td>input</td><td>of</td><td>compensating</td><td>of movement</td><td> 1335</td><td>it is</td>
<td colspan="2">available as</td><td>an</td><td>input from</td><td>decoder</td><td> 1300,</td><td>for</td>
<td>to receive</td><td>vectors</td><td>in</td><td colspan="2">movement. Similarly,</td><td colspan="2">a second</td>
disparity compensator input 1350 is available as an input to decoder 1300 to receive disparity vectors.
An output of a switch 1355 is connected in signal communication to a second noninverting input of combiner 1320. A first input of switch 1355 is connected in signal communication to a disparity compensator output 1350. A second input of switch 1355 is is connected in signal communication to an output of motion compensator 1335. A third input of switch 1355 is connected in signal communication to an output of intrapredictor 1330. An output of mode module 1360 is connected in signal communication to the
ΡΕ2512136 switch 1355 to control which input is selected by switch 1355. An unlock filter output 1325 is available as a decoder output 1300.
Turning to Figure 14, an exemplary method for processing images for a plurality of views in preparation for encoding images using the MPEG-4 AVC Standard Multi-Video Video Coding Extension (MVC) is indicated by reference numeral 1400.
Method 1400 includes a start block 1405 that passes control to a function block 1410. Function block 1410 arranges each N views out of a total of M views at a particular time point in the form of a supervisor. tiled image, and passes control to function block 1415. Function block
1415 initializes a num_coded_views_minusl syntax element, and passes control to a function block 1420. Function block 1420 initializes a view_id [i] syntax element to all views (num_coded_views_minusl +1), and passes control to a function block. function 1425. Function block 1425 initializes the interview reference dependency information for images that are anchors, and passes control to a function block 1430. Function block 1430 initializes interview reference dependency information for non-anchor images, and passes control to function block 1435. Function block 1435 initializes a
ΡΕ2512136 pseudo_view_present_flag syntax element, and passes control to a decision block 1440. Decision block 1440 determines whether or not the current value of the pseudo_view_present_flag syntax element is true. If so,
<td>So</td><td>the control</td><td>is passed to a block of</td><td>Function 1445.</td>
<td>Case</td><td>on the contrary</td><td>control is passed to a</td><td>end block</td>
<td> 1499.</td><td>0 block</td><td>of function 1445 initializes</td><td>the following</td>
syntax elements, and passes control to a function block 1450: tiling_mode; org_pic_width_in_mbs_minusl; and org_pic_height_in_mbs_minusl. Function block 1450 calls a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the end block.
1499.
Turning to Figure 15, an exemplary method for encoding images for a plurality of views using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension is indicated generally by reference numeral 1500.
Method 1500 includes a start block 1502 that has an input parameter pseudo_view_id and passes control to a function block 1504. The function block
1504 initializes a num_sub_views_minusl syntax element, and passes control to function block 1506. Function block 1506 initializes a variable i equal to zero, and passes control to a decision block
ΡΕ2512136
1508. Decision block 1508 determines whether or not variable i is less than the number of sub_views. If so, then control is passed to a function block 1510. Otherwise, control is passed to a function block 1520.
Function block 1510 initializes a sub_view_id [i] syntax element, and passes control to function block 1512. Function block 1512 initializes a num_parts_minusl [sub_view_id [i]] syntax element, and passes control to a function block 1514. Function block 1514 initializes a variable j equal to zero, and passes control to decision block 1516. Decision block 1516 determines whether or not variable j is smaller than the num_parts_minusl [sub_view_id [i]] syntax element. If so, control is passed to a function block 1518. Otherwise, control is passed to a decision block 1522.
Function block 1518 initializes the following syntax elements, increments variable j, and returns control to decision block 1516:
loc_left_offset [sub_view_id [i]] [j];
loc_top_offset [sub_view_id [i]] [j];
frame_crop_left_offset [sub_view_id [i]] [j]; frame_crop_right_offset [sub_view_id [i]] [j]; frame_crop_top_offset [sub_view_id [i]] [j]; frame_crop_bottom_offset [sub_view_id [i]] [j].
ΡΕ2512136 function block 1520 encodes the current image to the current view using multi-view video (MVC) encoding, and passes control to an end block 1599.
Decision block 1522 determines whether or not a tiling_mode syntax element is zero. If so, then control is passed to a function block 1524. Otherwise, control is passed to a function block 1538.
Function block 1524 initializes a flip_dir [sub_view_id [i]] syntax element and an upsample_view_flag [sub_view_id [i]] syntax element, and passes control to decision block 1526. Decision block 1526 determines whether, or no, the current value of the upsample_view_flag [sub_view_id [i]] syntax element is equal to one. If so, then control is passed to function block 1528. Otherwise, control is passed to decision block 1530.
Function block 1528 initializes an upsample_filter [sub_view_id [i]] syntax element, and passes control to decision block 1530. Decision block 1530 determines whether or not a value of the upsample_filter syntax element [sub_view_id [ i]] is equal to three. If yes, control is passed to a function block 1532. Otherwise control is passed to a function block.
1536 .
ΡΕ2512136 function block 1532 initializes the following syntax elements, and passes control to a function block 1534: vert_dim [sub_view_id [i]]; hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 1534 initializes the filter coefficients for each YUV component, and passes control to function block 1536.
Function block 1536 increments variable i, and returns control to decision block 1508.
Function block 1538 initializes a pixel_dist_x [sub_view_id [i]] syntax element and the flip_dist_y [sub_view_id [i]] syntax element, and passes control to a function block 1540. Function block 1540 initializes variable j equal to zero, and passes control to decision block 1542. Decision block 1542 determines whether or not the current value of variable j is less than the current value of the num_parts [sub_view_id [i]] syntax element . If so, then control is passed to function block 1544. Otherwise, control is passed to function block 1536.
Function block 1544 initializes a num_pixel_tiling_filter_coeffs_minusl [sub_view_id [i]] syntax element, and passes control to a function block 1546. Function block 1546 initializes the coefficients for all pixel tiling filters, and passes O
ΡΕ2512136 Control for function block 1536.
Turning to Figure 16, an exemplary method for processing images for a plurality of views in preparation for decoding the images using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension is indicated generally by reference numeral 1600.
Method 1600 includes a start block 1605 that passes control to a function block 1615. Function block 1615 recognizes a num_coded_views_minusl syntax element, and passes control to a function block 1620. The function block 1620 recognizes a view_id [i] syntax element for all views (num_coded_views_minusl +1), and passes control to a function block 1625. Function block 1625 recognizes inter-view reference dependency information for images that are anchors, and passes control to function block 1630. Function block 1630 recognizes inter-reference reference dependency information. views non-anchor images, and passes control to a function block 1635. Function block 1635 recognizes a pseudo_view_present_flag syntax element, and passes control to a decision block 1640. Decision block 1640 determines whether or not the current value of the pseudo_view_present_flag syntax element is true. If so, then control is passed to a 1645 function block.
Otherwise, control is passed to an end block 1699.
Function block 1645 recognizes the following syntax elements, and passes control to function block 1650: tiling_mode; org_pic_width_in_mbs_minusl; and org_pic_height_in_mbs_minusl. Function block 1650 calls a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the end block.
1699.
Turning to Figure 17, an exemplary method for decoding images for a plurality of views using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension is indicated generally by reference number 1700.
Method 1700 includes a start block 1702 that begins with the input parameter pseudo_view_id and passes control to a function block 1704. The function block
1704 recognizes a num_sub_views_minusl syntax element, and passes control to a function block 1706. Function block 1706 initializes a variable i equal to zero, and passes control to a decision block 1708. Decision block 1708 determines whether or not variable i is less than the number of sub_views. If so, then control is passed to a function block 1710. Otherwise, control is passed to a function block 1720.
ΡΕ2512136 function block 1710 recognizes a sub_view_id [i] syntax element, and passes control to a function block 1712. Function block 1712 recognizes a num_parts_minusl [sub_view_id [i]] syntax element, and passes control to function block 1714. Function block 1714 initializes a variable j equal to zero, and passes control to decision block 1716. Decision block 1716 determines whether or not variable j is smaller than the num_parts_minusl [sub_view_id [i]] syntax element. If so, then control is passed to a function block 1718. Otherwise control is passed to a decision block.
1722 .
Function block 1718 initializes the following syntax elements, increments variable j, and returns control to decision block 1716:
loc_lef_offset [sub_view_id [i]] [j];
loc_top_offset [sub_view_id [i]] [j];
frame_crop_left_offset [sub_view_id [i]] [j];
frame_crop_right_offset [sub_view_id [i]] [j];
frame_crop_top_offset [sub_view_id [i]] [j]; and frame_crop_bottom_offset [sub_view_id [i]] [j].
Function block 1720 decodes the current image to the current view using multi-view video coding (MVC), and passes control to a function block 1721. Function block 1721 separates each
Vista2512136 image view using high level syntax, and passes control to an end block 1799.
Separation of each view from the decoded image is performed using the high level syntax indicated in the bitstream. This high level syntax can indicate the exact location and possible orientation of the views (and possible corresponding depth) present in the image.
decision block 1722 determines whether or not a tiling_mode syntax element is zero. If so, then control is passed to a function block 1724. Otherwise, control is passed to a function block 1738.
Function block 1724 recognizes a flip_dir [sub_view_id [i]] syntax element and an upsample_view_flag [sub_view_id [i]] syntax element, and passes control to a decision block 1726. Decision block 1726 determines whether or not the current value of the upsample_view_flag [sub_view_id [i]] syntax element is equal to one. If so, then control is passed to a function block 1728. Otherwise, control is passed to a decision block 1730.
Function block 1728 recognizes an upsample_filter [sub_view_id [i]] syntax element, and passes control to decision block 1730.
ΡΕ2512136 Decision 1730 determines whether or not a value of the upsample_filter syntax element [sub_view_id [i]] equals three. If yes, control is passed to a function block 1732. Otherwise, control is passed to a function block 1736.
Function block 1732 recognizes the following syntax elements, and passes control to function block 1734: vert_dim [sub_view_id [i]]; hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 1734 recognizes the filter coefficients for each YUV component, and passes control to function block 1736.
Function block 1736 increments variable i, and returns control to decision block 1708.
Function block 1738 recognizes a pixel_dist_x [sub_view_id [i]] syntax element and the flip_dist_y [sub_view_id [i]] syntax element, and passes control to a 1740 function block. Function block 1740 initializes variable j equals zero, and passes control to decision block 1742. Decision block 1742 determines whether or not the current value of variable j is less than the current value of the num_parts [sub_view_id [] syntax element i]]. If so, then control is passed to a function block 1744. Otherwise, control is passed to function block 1736.
ΡΕ2512136 function block 1744 recognizes a num_pixel_tiling_filter_coeffs_minusl [sub_view_id [i]] syntax element, and passes control to a function block 1746. Function block 1776 recognizes the coefficients for all tiling filters pixel, and passes control to function block 1736.
Turning to Figure 18, an exemplary method for processing images for a plurality of views and depths in preparation for encoding images using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension is indicated generally by the reference numeral. 1800
Method 1800 includes a start block 1805 that passes control to a function block 1810. Function block 1810 arranges each N views and depth maps from a total of M views and depth maps at a particular point in time. time in the form of a tessellated super picture, and passes control to a function block 1815. Function block 1815 initializes a num_coded_views_minusl syntax element, and passes control to a function block 1820. Function block 1820 initializes a view_id [i] syntax element for all depths (num_coded_views_minusl + 1) corresponding to view_id [i], and passes control to a function block 1825. Function block 1825 initializes the information of inter-view reference dependency for
ΡΕ2512136 depth images that are anchor, and passes control to function block 1830. Function block 1830 initializes inter-view reference dependency information for non-anchor depth images, and passes control to block 1835. Function block 1835 initializes a pseudo_view_present_flag syntax element, and passes control to decision block 1840. Decision block 1840 determines whether or not the current value of the pseudo_view_present_flag syntax element is true. If so, then control is passed to function block 1845. Otherwise, control is passed to end block 1899.
Function block 1845 initializes the following syntax elements, and passes control to function block 1850: tiling_mode; org_pic_width_in_mbs_minusl; and org_pic_height_in_mbs_minusl. Function block 1850 calls a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the end block.
1899.
Turning to Figure 19, an exemplary method for encoding images for a plurality of views and depths using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension is indicated generally by reference numeral 1900.
Method 1900 includes a start block 1902 that passes control to a function block 1904. The function block
192512136 Function 1904 initializes a num_sub_views_minusl syntax element, and passes control to a function block 1906. Function block 1906 initializes a variable i equal to zero, and passes control to a decision block 1908. Decision block 1908 determines whether or not variable i is less than the number of sub_views. If so, then control is passed to a function block 1910. Otherwise, control is passed to a function block 1920.
Function block 1910 initializes a sub_view_id [i] syntax element, and passes control to a function block 1912. Function block 1912 initializes a num_parts_minusl [sub_view_id [i]] syntax element, and passes control to a function block 1914. Function block 1914 initializes a variable j equal to zero, and passes control to decision block 1916. Decision block 1916 determines whether or not variable j is smaller than the num_parts_minusl [sub_view_id [i] syntax element. If so, then control is passed to a function block 1918. Otherwise, control is passed to a decision block 1922.
Function block 1918 initializes the following syntax elements, increments variable j, and returns control to decision block 1916:
loc_left_offset [sub_view_id [i]] [j];
loc_top_offset [sub_view_id [i]] [j];
ΡΕ2512136 frame_crop_left_offset [sub_view_id [i]] [j];
frame_crop_right_offset [sub_view_id [i]] [j];
frame_crop_top_offset [sub_view_id [i]] [j]; and frame_crop_bottom_offset [sub_view_id [i]] [j].
function block 1920 encodes the current depth of the current view using multi-view video coding (MVC), and passes control to an end block 1999. The depth signal can be encoded similarly to the manner in which its corresponding signal video is encoded . For example, the depth signal for a view may be included in a tiled arrangement that includes only other depth signals, or only video signals, or both depth and video signals. The pseudo-view is then treated as a single view for MVC, and there will also presumably be other tessellations that are treated as other views for MVC.
<td></td><td> 0</td><td>block of</td><td>decision 1922</td><td>determines if,</td><td>or not a</td>
<td>element</td><td>in</td><td>syntax</td><td>tiling_mode is</td><td>equal to zero.</td><td>If so, the</td>
<td>control</td><td>is</td><td>past</td><td colspan="2">for a function block</td><td>1924. Case</td>
<td>contrary</td><td>r</td><td colspan="2">control is passed</td><td>for a block</td><td>of function</td>
1938 .
Function block 1924 initializes a syntax element flip_dir [sub_view_id [i]] and a syntax element upsample_view_flag [sub_view_id [i]], and passes control to
ΡΕ2512136 a decision block 1926. Decision block 1926 determines whether or not the current value of the upsample_view_flag [sub_view_id [i]] syntax element is equal to one. If so, then control is passed to function block 1928. Otherwise, control is passed to decision block 1930.
Function block 1928 initializes an upsample_filter [sub_view_id [i]] syntax element, and passes control to decision block 1930. Decision block 1930 determines whether or not a value of the upsample_filter syntax element [sub_view_id [ i]] is equal to three. If yes, control is passed to a 1932 function block. Otherwise control is passed to a function block.
1936.
Function block 1932 initializes the following syntax elements, and passes control to a function block 1934: vert_dim [sub_view_id [i]];
hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 1934 initializes the filter coefficients for each YUV component, and passes control to function block 1936.
Function block 1936 increments variable i, and returns control to decision block 1908.
Function block 1938 initializes a pixel_dist_x [sub_view_id [i]] syntax element and the flip_dist_y [sub_view_id [i]] element, and passes the syntax control.
122512136 for a function block 1940. Function block 1940 initializes a variable j equal to zero, and passes control to decision block 1942. Decision block 1942
<td>determines</td><td>if</td><td>or not,</td><td> 0</td><td>value</td><td colspan="2">current variable</td><td>j is smaller</td>
<td>than</td><td>O</td><td>value</td><td colspan="2">current</td><td>of</td><td>element of</td><td>syntax</td>
<td>num_parts</td><td>) sub_</td><td>view_id</td><td>[i]</td><td>]. If</td><td>Yes,</td><td>the control</td><td>it is past</td>
<td colspan="2">for a block</td><td colspan="2">of function</td><td> 1944 .</td><td>Case</td><td>on the contrary</td><td>control is</td>
<td colspan="2">passed to the</td><td>block</td><td>in</td><td>occupation</td><td> 1936</td><td> •</td><td></td>
Function block 1944 initializes a num_pixel_tiling_filter_coeffs_minusl [sub_view_id [i]] syntax element, and passes control to a function block 1946. Function block 1946 initializes the coefficients for all pixel tiling filters, and passes the control for function block 1936.
Turning to Figure 20, an exemplary method for processing images for a plurality of views and depths in preparation for decoding the images using the MPEG-4 AVC Standard Multi-Video Video Encoding (MVC) extension is indicated generally by the reference numeral. 2000
Method 2000 includes a start block 2005 that passes control to a function block 2015. Function block 2015 recognizes a num_coded_views_minusl syntax element, and passes control to a function block 2020. The function block 2020 recognizes a view_id [i] syntax element for all
ΡΕ2512136 (num_coded_views_minusl + 1) depths corresponding to view_id [i], and passes control to function block 2025. Function block 2025 recognizes interview reference dependency information for depth images that are anchors, and passes control to a function block 2030. Function block 2030 recognizes inter-view reference dependency information for non-anchor depth images, and passes control to function block 2035. Function block 2035 recognizes an input element. syntax pseudo_view_present_flag, and passes control to a decision block 2040. Decision block 2040 determines whether or not the current value of the pseudo_view_present_flag syntax element is true. If so, then control is passed to a function block 2045. Otherwise, control is passed to an end block 2099.
Function block 2045 recognizes the following syntax elements, and passes control to function block 2050: tiling_mode; org_pic_width_in_mbs_minusl; and org_pic_height_in_mbs_minusl. Function block 2050 calls a pseudo_view_info (view_id) syntax element for each coded view, and passes control to the end block.
2099.
Turning to Figure 21, an exemplary method for decoding images for a plurality of views and depths using video encoding extension
The MPEG-4 AVC Standard Multi-View (MVC) ΡΕ2512136 is generally indicated by reference number 2100.
Method 2100 includes a start block 2102 that begins with the input parameter pseudo_view_id, and passes control to a function block 2104. The function block
2104 recognizes a num_sub_views_minusl syntax element, and passes control to a function block 2106. Function block 2106 initializes a variable i equal to zero, and passes control to a decision block 2108. Decision block 2108 determines whether or not variable i is less than the number of sub_views. If so, then control is passed to a function block 2110. Otherwise, control is passed to a function block 2120.
Function block 2110 recognizes a sub_view_id [i] syntax element, and passes control to a function block 2112. Function block 2112 recognizes a syntax element num_parts_minusl [sub_view_id [i]], and passes control to function block 2114. Function block 2114 initializes a variable j equal to zero, and passes control to decision block 2116. Decision block 2116 determines whether or not variable j is smaller than the num_parts_minusl [sub_view_id [i]] syntax element. If yes, then control is passed to a function block 2118. Otherwise control is passed to a decision block.
2122 .
ΡΕ2512136 function block 2118 initializes the following syntax elements, increments variable j, and returns control to decision block 2116:
loc_left_offset [sub_view_id [i]] [j];
loc_top_offset [sub_view_id [i]] [j];
frame_crop_left_offset [sub_view_id [i]] [j];
frame_crop_right_offset [sub_view_id [i]] [j];
frame_crop_top_offset [sub_view_id [i]] [j]; and frame_crop_bottom_offset [sub_view_id [i]] [j].
function block 2120 decodes the current image using multi-view video coding (MVC), and passes control to function block 2121. function block 2121 separates each view from the image using high-level syntax, and passes the control for an end block 2199. Separation of each view using the high level syntax is as described above.
Decision block 2122 determines whether or not a tiling_mode syntax element is zero. If so, then control is passed to a function block 2124. Otherwise, control is passed to a function block 2138.
Function block 2124 recognizes a flip_dir [sub_view_id [i]] syntax element and an upsample_view_flag [sub_view_id [i]] syntax element, and passes control to a decision block 2126. The decision block
ΡΕ2512136 Decision 2126 determines whether or not the current value of the upsample_view_flag [sub_view_id [i]] syntax element is equal to one. If so, then control is passed to function block 2128. Otherwise control is passed to decision block 2130.
Function block 2128 recognizes an upsample_filter [sub_view_id [i]] syntax element, and passes control to decision block 2130. Decision block 2130 determines whether or not a value of the upsample_filter syntax element [sub_view_id [i]] is equal to three. If yes, control is passed to a function block 2132. Otherwise, control is passed to a function block 2136.
Function block 2132 recognizes the following syntax elements, and passes control to function block 2134: vert_dim [sub_view_id [i]]; hor_dim [sub_view_id [i]]; and quantizer [sub_view_id [i]]. Function block 2134 recognizes filter coefficients for each YUV component, and passes control to function block 2136.
Function block 2136 increments variable i, and returns control to decision block 2108.
element element
Function block 2138 recognizes one of the syntax pixel_dist_x [sub_view_id [i]] and one of the syntax flip_dist_y [sub_view_id [i]], and passes control to a function block 2140. The function block
ΡΕ2512136
2140 initializes variable j to zero, and passes control to decision block 2142. Decision block 2142 determines whether or not the current value of variable j is less than the current value of the num_parts syntax element [sub_view_id [i]]. If so, then control is passed to a function block 2144. Otherwise, control is passed to function block 2136.
Function block 2144 recognizes a num_pixel_tiling_filter_coeffs_minusl [sub_view_id [i]] syntax element, and passes control to function block 2146. Function block 2146 recognizes coefficients for all tiling filters pixel, and passes control to function block 2136.
Turning to Figure 22, examples of pixel-level tiling are generally indicated by reference numeral 2200. Figure 22 is further described below.
ΡΕ2512136
MOSAIC ARRAY ARRANGEMENT USING MPEG-4 AVC OR MVC
A multi-view video encoding application is free-view TV (or FTV). This application requires the user to be able to freely switch between two or more views. In order to achieve this, virtual views between two views need to be interpolated or synthesized. There are several methods for interpolating views. One method uses depth for interpolation / synthesis of views.
Each view can have an associated depth signal. Thus, depth can be considered to be another form of video signal. Figure 9 shows an example of a depth signal 900. In order to enable applications such as FTV, the depth signal is transmitted in conjunction with the video signal. In the proposed structure for tiling, the depth signal can also be added as one of the tiling arrangements. Figure 10 shows an example of depth signals added as tessellations. Depth signals / tessellations are shown on the right side of Figure 10.
Once the depth is coded as a tessellation of the entire frame, the high level syntax should indicate which tiling is the
Sinal2512136 depth signal so that the representative can use the depth signal appropriately.
Where the input sequence (as shown in Figure 1) is encoded using an MPEG-4 AVC Standard encoder (or an encoder corresponding to a different video coding standard and / or recommendation), the syntax of The proposed high-level can be present in, for example, Supplemental Sequence Parameter Set (SPS), Image Parameter Set (PPS), slice header, and / or an Enhancement Supplemental Information (SEI) message. One embodiment of the proposed method is presented in TABLE 1 where the syntax is present in a Supplemental Enhancement Information (SEI) message.
In the event that the pseudo view input sequences (such as shown in Figure 1) are encoded using an MPEG-4 AVC Standard Multi-View Video Encoding Extension (MVC) encoder (or an encoder corresponding to a multi-view video coding standard (regarding a different video coding standard and / or recommendation), the proposed high-level syntax may be present in the SPS, PPS, slice header, SEI message, or in a specified profile. An embodiment of the proposed method is presented in TABLE 1. TABLE 1 shows syntax elements present in the Supplementary Sequence Parameter Set (SPS) structure,
ΡΕ2512136 including proposed syntax elements in accordance with one embodiment of the present principles.
TABLE 1
<td>seq parameter set mvc extension {) {</td><td>ç</td><td>Desorftnr</td>
<td>in a video minus 1</td><td></td><td>us (v)</td>
<td>for <i = 0; i <= num views minus „1; i ++}</td><td></td><td></td>
<td>view id</td><td></td><td>eu (v)</td>
<td>for (i = 0; i <sup><=</sup> num vtews minus 1, «+ * ·) {</td><td></td><td></td>
<td>in an anchor refs! 0 [ij</td><td></td><td>eu (v)</td>
<td>for (í = 0; j <fuim anchor refs iG [»}; j ++}</td><td></td><td></td>
<td>anchor ref iOfi] | jl</td><td></td><td>eu (v)</td>
<td>in an anchor refs l1 [ij</td><td></td><td>eu (v)</td>
<td>for (j = 0; | <in an anchor refs H [ij; j +<sup>+</sup> )</td><td></td><td></td>
<td>artch or refj 1 [ij [jj</td><td></td><td>eu (v)</td>
<td> }</td><td></td><td></td>
<td>fior (i = 0; i num views minus 1; í ++) {</td><td></td><td></td>
<td>num non anchor refs IO [} j</td><td></td><td>eu (v)</td>
<td>for <j = 0; j <num non anchor refsJO [ij; j - * - * -)</td><td></td><td></td>
<td>no n an c hor ref l [i] [j]</td><td></td><td>eu (v)</td>
<td>num non snchor meal i1 (ij</td><td></td><td>eu (v)</td>
<td>for {j = 0; j <num rton anchor refsJ1 [ij; j ++)</td><td></td><td></td>
<td>non anchor ref H [i OJ)</td><td></td><td>eu (v)</td>
<td></td><td></td><td></td>
<td>psetKJo view presení ÍJag</td><td></td><td>u (l)</td>
<td>if {pseudo view presence flag} {</td><td></td><td></td>
<td>tiltng mode</td><td></td><td></td>
<td>org pic pic in mbs minusl</td><td></td><td></td>
<td>arg pic heig htinm bs my naked s 1</td><td></td><td></td>
<td>for (i - 0; i <num views minus' !; i ++)</td><td></td><td></td>
<td>pseudo viewjnfo (i);</td><td></td><td></td>
<td></td><td></td><td></td>
<td></td><td></td><td></td>
TABLE 2 shows syntax elements for the pseudo_view_info syntax element of TABLE 1, in accordance with one embodiment of the present principles.
ΡΕ2512136
TABLE 2
<td>pseudo víewjnfo {pseudo yiew id) {</td><td>Ç</td><td>Dscrifor</td>
<td>in a sub video minus 1 (pseudo vtew id}</td><td> 5</td><td>eu (v)</td>
<td>if (in a sub video minus 1 S = 0} {</td><td></td><td></td>
<td>for (i = 0; i <num sub views minus 1 [pseudo viewJdJ; i ++) {</td><td></td><td></td>
<td>sub view id [i}</td><td> 5</td><td>eu (v)</td>
<td>in a part minus1 | sub vtewjdf i]}</td><td> 5</td><td>eu (v)</td>
<td>for (j = 0; j <= numj> arts minus1 [stib view id [í}}; j ++} {</td><td></td><td></td>
<td>! oc left offeetjsub view td [i]} (j]</td><td> 5</td><td>eu (v)</td>
<td>loc top offset [sub view id [i}] (j]</td><td> 5</td><td>eu (v)</td>
<td>frame crapjeft offeet [sijb view id [í]} J j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop right offeet [sub videojj [i]] fj]</td><td> 5</td><td>eu (v)</td>
<td>frame crop top offeet [sub view id [i J] {j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop bottom offset [sub view id [i]] I j 1</td><td> 5</td><td>eu (v)</td>
<td> }</td><td></td><td></td>
<td>if (tiling rroxfe == 0) {</td><td></td><td></td>
<td>flip dirub view KJ [i} [j]</td><td> 5</td><td><sub>U</sub>(2)</td>
<td>upsampfe viewjag [sub vtew idl i]]</td><td> 5</td><td>u {I)</td>
<td>if (upsampi v »ew fl3g {sub vsew id [i}})</td><td></td><td></td>
<td>upsample filter (sub view jd [i] J</td><td> 5</td><td>u (2)</td>
<td>if {upsampie fiíef {sub view id {»]] == 3} {</td><td></td><td></td>
<td>vert dím (sub viaw id [ill</td><td> 5</td><td>eu (v)</td>
<td>hor dimlsub view id [i} j</td><td> 5</td><td>eu (v)</td>
<td>quanijzerfsub vtewjdíi3]</td><td> 5</td><td>eu (v)</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td><td></td>
<td>for (y = 0; y <vert dim {sub viewjdfi}] - 1; y ++) {</td><td></td><td></td>
<td>for (x = 0; x <hor dim [sub view id [i]] - 1; x ++)</td><td></td><td></td>
<td>fitter coeffsísub view id [i]] (yuv3 {y] {x]</td><td> 5</td><td>if (v)</td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td>í</td><td></td><td></td>
<td>} // if (tilting mode == O)</td><td></td><td></td>
<td>else if ftfing mode 1} {</td><td></td><td></td>
<td>set dist x | sub view id [i] 3</td><td></td><td></td>
<td>pixel! dist y | sub view id [i] J</td><td></td><td></td>
<td>for (j = 0; j <= num j33rts {sub vtew id [»]]; j ++) {</td><td></td><td></td>
<td>num pixet tiling fiiter coeffs jTiIF »us1 (sub view id [i] JJjj</td><td></td><td></td>
<td>for {coeffjdx = 0; coeffjdx <= numjíxel tiling fiiteF coeffe minus1 (sub viewjdf i Hfíl; j<sup>++</sup>}</td><td></td><td></td>
<td>pixet ti [ing fitter coeffs [sub view id [i] j03</td><td></td><td></td>
<td>} it for (j = 0; j <= num parts [subj / iewjd [i]]; j ++)</td><td></td><td></td>
<td>} else if (tiling made == 1)</td><td></td><td></td>
<td>} // for (i = 0; i <num sub views minus 1; i ++)</td><td></td><td></td>
<td>} ti tf (num sub views ffflnuB 1! = 0)</td><td></td><td></td>
<td>Ϊ</td><td></td><td></td>
ΡΕ2512136
Semantic of the syntax elements presented in TABLE 1 and TABLE 2 pseudo_view_present_flag equal to truth indicates that some view is a multiple view sub-view.
tiling_mode of 0 indicates that subviews are tiled at the image level. A value of 1 indicates that the tiling is at pixel level.
The new SEI message could use a value for the SEI payload type that had not been used in the MPEG-4 AVC Standard or an extension of the MPEG-4 AVC Standard.
The new SEI message includes several syntax elements with the following semantics.
num_coded_views_minusl plus 1 indicates the number of encoded views supported by the bitstream. The num_coded_views_minusl value is in the range 0-1023 inclusive.
org_pic_width_in_mbs_minusl plus 1 specifies the width of an image in each view in macroblock units.
The variable for image width in macroblock units is derived as follows:
ΡΕ2512136
PicWidthlnMbs = org_pic_width_in_mbs_minusl + 1
The image width variable for the luma component is derived as follows:
PicWidthlnSamplesL = PicWidthlnMbs * 16
The image width variable for the chroma component is derived as follows:
PicWidthlnSamplesC = PicWidthlnMbs * MbWidthC org_pic_height_in_mbs_minusl plus 1 specifies the height of an image in each view in macroblock units.
The variable for image height in macroblock units is derived as follows:
PicHeightlnMbs = org_pic_height_in_mbs_minusl + 1
The image height variable for the luma component is derived as follows:
PicHeightlnSamplesL = PicHeightlnMbs * 16
The image height variable for the chroma component is derived as follows:
ΡΕ2512136
PicHeightlnSamplesC = PicHeightlnMbs * MbHeightC num_sub_views_minus1 plus 1 indicates the number of encoded subviews included in the current view. The num_coded_views_minusl value is in the range 0-1023 inclusive.
sub_view_id [i] specifies the sub_view_id of the subview with the decoding order indicated by i.
num_parts [sub_view_id [i]] specifies the number of parts into which the sub_view_id [i] image is divided.
loc_left_offset [sub_view_id [i]] [j] and loc_top_offset [sub_view_id [i]] [j] specify the locations in the left and top pixel offsets, respectively, where the current part j is located in the final reconstructed view image with sub_view_id equals sub_view_id [i].
view_id [i] encoding order specifies the view_id of the view indicated by i.
with frame_crop_left_offset [view_id [i]] [j], frame_crop_right_offset [view_id [i]] [j], frame_crop_top_offset [view_id [i]] [j], and frame_crop_bottom_offset [view_id [i]] [j] images in the encoded video stream that are part of num_part je of view_id i, in terms of a
ΡΕ2512136 rectangular region specified in frame coordinates for output.
The CropUnitX and CropUnitY variables are derived as follows:
- If chroma_format_idc is 0, CropUnitX and CropUnitY are derived as follows:
CropUnitX = 1
CropUnitY = 2 - frame_mbs_only_flag
- Otherwise (chroma_format_idc equals 1, 2, or 3), CropUnitX and CropUnitY are derived as follows:
CropUnitX = SubWidthC
CropUnitY = SubHeightC * (2-frame_mbs_only_flag)
The frame crop rectangle includes luma samples with horizontal frame coordinates from the following:
CropUnitX * frame_crop_left_offset for PicWidthlnSamplesL - (CropUnitX * frame_crop_right_offset + 1) and vertical frame coordinates from CropUnitY * frame_crop_top_offset for (16 *
FrameHeightlnMbs) - (CropUnitY * frame_crop_bottom_offset + 1), inclusive. The value of frame_crop_left_offset should
ΡΕ2512136 be in the range 0 to (PicWidthlnSamplesL / CropUnitX) (frame_crop_right_offset + 1) inclusive; and the value of frame_crop_top_offset should be in the range 0 to (16 * FrameHeightInMbs / CropUnitY) - (frame_crop_bottom_offset + 1), inclusive.
When chroma_format_idc is not equal to 0, the corresponding specified samples from the two chroma series are samples that have the frame coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the frame coordinates of the specified luma samples. .
For decoded fields, the specified samples of the decoded fields are the samples that fall within the rectangle specified in the frame coordinates.
num_parts [view_id [i]] specifies the number of parts into which the view_id [i] image is divided.
depth_f lag [view_id [i]] specifies whether or not the current part is a depth signal. If depth_flag is zero then the current part is not a depth sign. If depth_flag is equal to 1, then the current part is a depth sign associated with the view identified by view_id [i].
flip_dir [sub_view_id [i]] [j] specifies the reversal direction for the current part. If flip_dir equals 0
ΡΕ2512136 indicates no inversion, if flip_dir equals 1 indicates a horizontal inversion, if flip_dir equals 2 indicates a vertical inversion, and if flip_dir equals 3 indicates an inversion in the horizontal and vertical directions.
flip_dir [view_id [i]] [j] specifies the reversal direction for the current part. If flip_dir equals 0 indicates no reversal, if flip_dir equals 1 indicates a horizontal inversion, if flip_dir equals 2 indicates a vertical reversal, and if flip_dir equals 3 indicates a reversal in directions. horizontal and vertical.
loc_left_offset [view_id [i]] [j], loc_top_offset [view_id [i]] [j] specifies the pixel offset location where the current part j is located in the final reconstructed view image with view_id equal to view_id [i] .
upsample_view_flag [view_id [i]] indicates whether the image belonging to the view specified by view_id [i] needs to be upsampled. If upsample_view_flag [view_id [i]] equals 0 specifies that the image with view_id equal to view_id [i] will not be upsampled. If upsample_view_flag [view_id [i]] equals 1 specifies that the image with view_id equals upsampling.
view_id [i] will suffer from
ΡΕ2512136 upsample_filter [view_id [i]] indicates the type of filter that is to be used for upsampling. If upsample_filter [view_id [i]] equals 0 then the filter should be used
6-tap
Stroke if upsample_filter [view_id [i]] equals 1 indicates filter should be used
4-tap
SVC, if upsample_filter [view_id [i]] equals 2 indicates that a bilinear filter should be used, if upsample_filter [view_id [i]] equals 3 indicates that proper filter coefficients will be transmitted. When upsample_fiter [view_id [i]] is not present then it is initialized to 0. In this embodiment we use our own 2D filter. It can easily be extended to a 1D filter, and to some other nonlinear filters.
vert_dim [view_id [i]] vertical 2D filter itself.
specifies dimension hor_dim [view_id [i]] specifies horizontal 2D filter itself.
dimension quantizer [view_id [i]] specifies the quantification factor for each filter coefficient.
filter_coeffs [view_id [i]] [yuv] [y] [x] specifies the quantified filter coefficients. The yuv signals the component to which the filter coefficients apply. If yuv equals 0 specifies component Y, if yuv is
ΡΕ2512136 equals 1 specifies component U, and if yuv equals 2 specifies component V.
pixel_dist_x [sub_view_id [i]] and pixel_dist_y [sub_view_id [i]] specify, respectively, the distance in horizontal and vertical direction in the final reconstructed pseudo view between neighboring pixels in view with sub_view_id [i].
num_pixel_tiling_filter_coeffs_minusl [sub_view_id [i]] [j] plus one indicates the number of coefficients when tiling mode is initialized to 1.
pixel_tiling_filter_coeffs [sub_view_id [i]] [j] signals the filter coefficients that are required to represent a filter that can be used to filter the tiled figure.
Examples of pixel-level tiling
Turning to Figure 22, two examples are respectively shown showing the composition of a pseudo view by the pixel mosaic from four views by reference numerals 2210 and 2220, respectively. The four views are collectively indicated by reference numeral 2250. The syntax values for the first example in Figure 22 are provided in TABLE 3 below.
ΡΕ2512136
TABLE 3
<td>pseudo view info {pseudo view id) {</td><td>Wer</td>
<td>numrsub views minus 1 fpseudo view id]</td><td> 3</td>
<td>sub view id [OJ</td><td> 0</td>
<td>num parts msnus1 f 0]</td><td> 0</td>
<td>Yo! Efi offeet [03 [0J</td><td> 0</td>
<td>loc top offset (0j [01</td><td> 0</td>
<td>pixet dist x [O] [O3</td><td> 0</td>
<td>pixel dist y {O] [O]</td><td> 0</td>
<td>sub view id [1)</td><td> 0</td>
<td>num parts minus1 [1]</td><td> 0</td>
<td>loc left offset [1] [0]</td><td> 1</td>
<td>loc top offset (1] [0)</td><td> 0</td>
<td>pixel dist x [1] [O]</td><td> 0</td>
<td>pixel dist y [1] [O]</td><td> 0</td>
<td>sub view id [2]</td><td> 0</td>
<td>num parts minus1 [2]</td><td> 0</td>
<td>loc left offset [2] [OJ</td><td> 0</td>
<td>loc top offset [2J [0]</td><td> 1</td>
<td>pixel dist x [2] [0]</td><td> 0</td>
<td>pixel dist y [2] [0]</td><td> 0</td>
<td>sub view id [3]</td><td> 0</td>
<td>num parts minus1 [3]</td><td> 0</td>
<td>loc left offset [3] [0]</td><td> 1</td>
<td>loc top offset [3J [0]</td><td> 1</td>
<td>pixel dist x [3] [0]</td><td> 0</td>
<td>pixel dist y [3] [0]</td><td> 0</td>
The syntax values for the second example in Figure 22 are all the same except the following two syntax elements: loc_left_offset [3] [0] equal to 5 and loc_top_offset [3] [0] equal to 3.
Offsets indicate that pixels corresponding to a view should start at a certain offset location. This is shown in Figure 22 (2220). This can be done, for example, when two views produce images in which common objects appear.
ΡΕ2512136 shifted from one view to the other. For example, if the first and second cameras (representing the first and second views) capture images of an object, the object may appear to be shifted five pixels to the right in the second view compared to the first view. This means that the pixel (i-5, j) in the first view corresponds to the pixel (i, j) in the second view. If the pixels in both views are simply pixel-by-pixel, then there may not be much correlation between neighboring pixels in the tile, and the gains in spatial coding may be small. On the other hand, by shifting the tessellation so that pixel (i-5, j) of view one is positioned close to pixel (i, j) of view two, spatial correlation can be increased and coding gain can also be increased. This is because, for example, the corresponding pixels for the object in the first and second views are being tiled next to each other.
Thus, the presence of loc_left_offset and loc_top_offset can benefit coding efficiency. Deviation information may be obtained by external means. For example, camera position information or global disparity vectors between views may be used to determine such offset information.
As a result of skewing, some pixels in the pseudo view are not pixel values assigned from either view. Continuing the above example,
ΡΕ2512136 when the pixel (i-5, j) of view one is arranged next to pixel (i, j) of view two, for values of i = 0 ... 4 there is no pixel (i-5, j) of view one for tiling so that these pixels are empty in the tiling. For pseudo view (tessellation) pixels that are not values assigned from either view, at least one implementation uses an interpolation procedure similar to the sub-pixel interpolation procedure in stroke compensation. That is, the empty pixels in the tessellation can be interpolated from neighboring pixels. Such interpolation may result in greater spatial correlation in tiling and greater coding gain for tiling.
In video encoding, we can choose a different encoding type for each image, such as images I, P, and B. For multi-view video encoding, we additionally define images that are anchors and images that are not anchors. In one embodiment, we propose that the decision to group can be made based on the image type. This cluster information is flagged in high level syntax.
Turning to Figure 11, an example of 5 tiled views in a single frame is indicated generally by reference numeral 1100. In particular, the ballroom sequence is presented with 5 tiled views in a single frame.
ΡΕ2512136
Additionally, it can be observed that the fifth view is divided into two parts so that it can be arranged in a rectangular frame. Here each view has the QVGA dimension so the total frame size is 640x600. Since 600 is not a multiple of 16 it should be extended to 608.
For this example, the possible SEI message might be as shown in TABLE 4.
TABLE 4
<td>multíview dispfay info (payíoadSize) {</td><td>Value</td>
<td>num coded vtews minus1</td><td> 5</td>
<td>org ptc width tn mbs minus 1</td><td> 40</td>
<td>org pic heiqht in mbs my nudes 1</td><td> 30</td>
<td></td><td></td>
<td>vièwjdf 0 J</td><td> 0</td>
<td>num parts (view id {0)]</td><td> 1</td>
<td></td><td></td>
<td>depth flagfview id [0}] [0 J</td><td> 0</td>
<td>flip dir {view id [O]} (0)</td><td> 0</td>
<td>icc ieft offset [view id [0 JJ10]</td><td> 0</td>
<td>loc top offeet (view td [0]] f 0]</td><td> 0</td>
<td>frame crop left offset [view id [0]] [0]</td><td> 0</td>
<td>frame crop right offset [view id [0]] [0]</td><td> 320</td>
<td>frame crop top offset (view id [0]] [0]</td><td> 0</td>
<td>frame_crop_bottom_offset (view_id [0]] [0 |</td><td> 240</td>
<td></td><td></td>
<td>upsample view flag [view id [0]]</td><td> 1</td>
<td>if (upsample view flag [viewjd [0]]) {</td><td></td>
<td>vert dim [view id [O]]</td><td> 6</td>
<td>hor dim [view id [O] J</td><td> 6</td>
<td>quantizer [view id [O] J</td><td> 32</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td>
<td>for (y = 0; y <vert dim [view id [i]] -1; y ++) {</td><td></td>
<td>for (x = 0; x <hor dim [view id [i]] - 1; x ++)</td><td></td>
<td>filter coeffs [view id [ij] [yuv] [y] [x]</td><td>XX</td>
ΡΕ2512136
<td></td><td></td>
<td></td><td></td>
<td>view id [1 J</td><td> 1</td>
<td>num parts [view id [1]]</td><td> 1</td>
<td></td><td></td>
<td>depth flag [view id [0]] [0]</td><td> 0</td>
<td>flip dir [view id [1]] [0]</td><td> 0</td>
<td>loc left offset [view id [1]] [0 J</td><td> 0</td>
<td>loc top offset [view id [1]] [0]</td><td> 0</td>
<td>frame crop left offset (view id [1]] [0]</td><td> 320</td>
<td>frame crop right offset [view id [1]] [0]</td><td> 640</td>
<td>frame crop top offset (view id [1]] [0]</td><td> 0</td>
<td>frame crop bottom offset [view id [1]] [0]</td><td> 320</td>
<td></td><td></td>
<td>upsample view flag [view id [1]]</td><td> 1</td>
<td>if (upsample view flag [view id [1]]) {</td><td></td>
<td>veri di iTi [vi w id [1J]</td><td>& u</td>
<td>hor dim [view id [1]]</td><td> 6</td>
<td>quantizer [view id [1]]</td><td> 32</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td>
<td>for (y = 0; y <vert dim [view id [i]] -1; y ++) {</td><td></td>
<td>for (x = 0; x <hor dim [view id [i]] - 1; x ++)</td><td></td>
<td>filter coeffs [view id [i]] [yuv] [y] [x]</td><td>XX</td>
<td></td><td></td>
<td></td><td></td>
<td></td><td></td>
<td>...... (similarly for view 2,3)</td><td></td>
<td></td><td></td>
<td>view id [4]</td><td> 4</td>
<td>num parts [view id [4]]</td><td> 2</td>
<td></td><td></td>
<td>depth flag [viewjd [0]] [0]</td><td> 0</td>
<td>flip dir [view id [4]] [0}</td><td> 0</td>
<td>loc left offset [view id [4]] [0]</td><td> 0</td>
<td>loc top offset [view id [4] J [0]</td><td> 0</td>
<td>frame crop left offset [view id [4]] [0]</td><td> 0</td>
<td>frame crop right offset [view id [4]] [0]</td><td> 320</td>
<td>frame crop top offset [view id [4]} [0]</td><td> 480</td>
<td>frame crop bottom offset [view id [4 J] [0]</td><td> 600</td>
<td></td><td></td>
<td>flip dir [view id [4]] [1]</td><td> 0</td>
<td>loc lefi offset [view id [4]] [1]</td><td> 0</td>
<td>loc top offset [view id [4]] [1]</td><td> 120</td>
<td>frame crop left offset [view id [4]] [1]</td><td> 320</td>
<td>frame crop right offset [view id [4]] [1)</td><td> 640</td>
<td>frame crop top offset [view id [4]] [1]</td><td> 480</td>
<td>frame crop bottom offset [view id [4]) [1]</td><td> 600</td>
ΡΕ2512136
<td></td><td></td>
<td></td><td></td>
<td>upsampte view ftag [view id [4]]</td><td> 1</td>
<td>if (upsampte view ffag {vtew id [4]]} {</td><td></td>
<td>vert dím [view id [4]]</td><td> 6</td>
<td>hor dim [view Jdí4J]</td><td> 6</td>
<td>quantszerfvíew idI4] 3</td><td> 32</td>
<td>for {yuv = 0; yuv <3; yuv ++) {</td><td></td>
<td>for (y = 0; y <vert dimfvtew id (i}] -1; y ++) {</td><td></td>
<td>for (x = 0; x <hor difn [view id [ij] - 1; x ++)</td><td></td>
<td>ft} have coeffs [v »ew td (i]} íyuv] [yjx]</td><td>XX</td>
<td></td><td></td>
TABLE 5 presents the general structure of the syntax for transmitting information in multiple views to the example presented in TABLE 4.
TABLE 5
<td>multjview dispar! ayjnfo (payloadSize) {</td><td>Ç</td><td>Descriptor</td>
<td>num codet views mintjs1</td><td> 5</td><td>eu (v)</td>
<td>org pic widfo tn mbs minus1</td><td> 5</td><td>eu (v)</td>
<td>org pic fceíght Jn mbs minus1</td><td> 5</td><td>eu (v)</td>
<td>for {ί = 0; i <= num çoded view5 minus1; i ++) {</td><td></td><td></td>
<td>viewJd [i]</td><td> 5</td><td>eu (v)</td>
<td>num parts [view id [i JJ</td><td> 5</td><td>eu (v}</td>
<td>for {j = 0; j <= num partspj; j ++) {</td><td></td><td></td>
<td>depth flag [vtew td [i JJ [j]</td><td></td><td></td>
<td>fltp dír (víew ídf <JJJ j J</td><td> 5</td><td>u {2)</td>
<td>! oc teft offset [view id [i JJ {jJ</td><td> 5</td><td>eu (v)</td>
<td>loc top offsetlviewjd [»}} {j J</td><td> 5</td><td>eu (v)</td>
<td>frame erop left offset [view id [i J] (j J</td><td> 5</td><td>uefv)</td>
<td>frame crop right offset [view id [í J1 [j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop top offset (view id [i JJ [j]</td><td> 5</td><td>eu (v)</td>
<td>frame crop bottom offset (view id [i J] [j]</td><td> 5</td><td>eu (v)</td>
<td> }</td><td></td><td></td>
<td>upsample view flag [view id [i JJ</td><td> 5</td><td>u (1)</td>
<td>if (upsample view flagfview id [i] 1)</td><td></td><td></td>
<td>upsample filter [view id (i JJ</td><td> 5</td><td>u (2)</td>
<td>if (upsample fiter [view id [i] J == 3) {</td><td></td><td></td>
<td>vert dim [view id [i] J</td><td> 5</td><td>eu (v)</td>
<td>hor dim [view id [i]]</td><td> 5</td><td>eu (v)</td>
<td>quantizer [view id [i] J</td><td> 5</td><td>eu (v)</td>
<td>for (yuv = 0; yuv <3; yuv ++) {</td><td></td><td></td>
ΡΕ2512136
<td>for {y - 0, y <vert dim [view td [S}] -1; y + *} {</td><td></td><td></td>
<td>for (x = 0; x <hor dtm | viewJdfi]] -1 x ++)</td><td></td><td></td>
<td>ftiter coeffs | view idf i} J [yuvgy] [x]</td><td> 5</td><td>if {v)</td>
<td> }</td><td></td><td></td>
<td> )</td><td></td><td></td>
<td>í</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td></td><td></td><td></td>
Referring to Figure 23, a video processing device 2300 is shown. The video processing device 2300 may be, for example, a television converter subscriber interface or other device that receives encoded video and provides, for example, video. decoded for viewing by a user or storage pair. Thus, device 2300 may make its output available to a television, computer monitor, or a computer or other processing device.
Device 2300 includes a decoder 2310 that receives a data signal 2320. Data signal 2320 may include, for example, an AVC or MVC compatible stream. The decoder 2310 decodes all or part of the received signal 2320 and outputs a decoded video signal 2330 and tiling information 2340. The decoded video 2330 and tiling information 2340 is provided to a selector 2350. Device 2300 also includes a user interface 2360 which receives input from user 2370. User interface 2360 provides an image selection signal 2380,
ΡΕ2512136 based on user input 2370 for selector 2350. Image selection signal 2380 and user input 2370 indicate which of multiple images a user wants to be displayed. The 2350 selector provides the selected image (s) as an output 2390. The 2350 selector uses the 2380 image selection information to select which of the images in the decoded video 2330 to output as the 2390 output. The selector 2350 uses the tiling information 2340 to locate the selected image or images in the decoded video 2330.
In various implementations, selector 2350 includes user interface 2360, and in other implementations no user interface 2360 is required because selector 2350 directly receives user input.
2370 without performing a separate interface function. Selector 2350 may be implemented in software or as an integrated circuit, for example. The selector 2350 may also incorporate the decoder 2310.
More generally, the decoders of various implementations described in this application may provide a decoded output that includes a complete tiling arrangement. Additionally or alternatively, the decoders may provide a decoded output that includes only one or more selected images (images or depth signals, for example) from the mosaic arrangement.
ΡΕ2512136
As noted above, high level syntax may be used to perform signaling in accordance with one or more embodiments of the present principles. High level syntax can be used, for example, but is not limited to, signaling any of the following: the number of encoded views present in the larger frame; the original width and height of all views; for each coded view, the view identifier corresponding to the view; for each coded view, the number of parts into which the frame of a view is divided; For each part of the view, the direction of inversion (which may be, for example, without inversion, horizontal inversion only, vertical inversion only or horizontal and vertical inversion); for each part of the view, the left position in pixels or number of macro blocks to which the current part belongs in the final frame for the view; for each part of the view, the position at the top of the pixel part or number of macro blocks to which the current part belongs in the final frame for the view; for each part of the view, the position on the left in the current large decoded / encoded frame of the cropped pixel window or number of macro blocks; for each part of the view, the position on the right, in the current large decoded / encoded frame, of the cropped pixel window or number of macro blocks; for each part of the view, the top position in the current large decoded / coded frame of the cropped pixel window or number of macro blocks; and, for each part of the view, the position below, in the current large low frame
ΡΕ2512136 decoded / encrypted, cropped pixel window or number of macro blocks; for each coded view whether or not the view needs to be upsampled before exiting (where, if upsampling is required, a high level syntax can be used to indicate the method for upsampling (including but not limited to, 6-tap AVC filter, 4-tap SVC filter, bilinear filter or 1D, 2D nonlinear or linear filter)).
It is to be noted that the terms encoder and decoder are connoted with general structures and are not limited to any particular functions or characteristics. For example, a decoder may receive a modulated carrier that carries an encoded bit stream, and demodulates the encoded bit stream, as well as decodes the bit stream.
Various methods have been described. Many of these methods are detailed to provide wide dissemination. It is noted, however, that variations are contemplated which may vary one or many of the specific characteristics described for such methods. Additionally, many of the features that are enumerated are known in the art and are not, accordingly, described in great detail.
Additionally, reference has been made to using high level syntax to send certain
ΡΕ2512136 Information in various implementations. It is to be understood, however, that other implementations use low level syntax, or indeed other mechanisms together (such as, for example, sending information as part of encoded data) to provide the same information (or variations of this information). ).
Various implementations provide appropriate tessellation and signaling to allow multiple views (images, more generally) to be tiled into a single image, encoded as a single image, and sent as a single image. Signaling information may allow a post processor to separate views / images from each other. Similarly, multiple images that are tiled may be viewed, but at least one of the images may be depth information. These implementations may provide one or more advantages. For example, users may wish to view multiple tiled views, and these various implementations provide an efficient way to encode and transmit or store such views by tiling them prior to encoding and transmitting / storing them with arrangement. in mosaic.
Implementations that tile multiple views in the context of AVC and / or MVC also provide additional advantages. Stroke is ostensibly only used for a single view, so they are not
Additional views are expected. However, such stroke-based implementations can provide multiple views in a stroke environment because tiled views can be arranged such that, for example, a decoder knows that tiled images belong to different views (for example, the picture in the upper left corner in pseudo view is view 1, the image in the upper right corner is view 2, etc.).
Additionally, MVC already includes multiple views, so multiple views are not expected to be included in a single pseudo view. In addition, MVC has a limit on the number of views that can be supported, and such MVC-based implementations effectively increase the number of views that can be supported by allowing (as in AVC-based implementations) additional views to be tiled. For example, each pseudo view may correspond to one of the supported views of MVC, and the decoder may know that each supported view actually includes four views in a pre-arranged tiling order. Thus, in such an implementation, the number of possible views is four times the number of supported views.
The implementations described herein may be implemented in, for example, a method or process, an apparatus, or a software program. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the
Implementation of features discussed may also be implemented in other ways (e.g., an apparatus or program). An apparatus may be implemented by, for example, appropriate hardware, software, and appropriate firmware. The methods may be implemented with, for example, an apparatus such as, for example, a processor, which relates to general processing devices, including, for example, a computer, a microprocessor, an integrated circuit, or a device. programmable logic. Processing devices also include communication devices, such as, for example, computers, mobile phones, handheld / personal digital assistants (PDAs), and other devices that facilitate information communication between end users.
Implementations of the various processes and features described herein may be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding and decoding. Examples of equipment include video encoders, video decoders, video codecs, network servers, a television converter subscriber interface, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be clear, the equipment can be mobile and even installed in a mobile vehicle.
ΡΕ2512136
Additionally, the methods may be implemented by instructions to be performed by a processor, and such instructions may be stored in a computer readable medium such as, for example, an integrated circuit, or software carrier, or other storage device such as, for example. For example, a hard disk, a floppy disk, random access memory (RAM), or read-only memory (ROM). The instructions may constitute an application program that takes shape tangibly in a processor readable medium. Of course, a processor may include a processor readable medium having, for example, instructions for carrying out a process. Such an application program may be loaded into, and executed by, a machine comprising a suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processor units (CPU), random access memory (RAM), and input / output (I / O) interfaces. The computer platform may also include an operating system and micro-instruction code. The various processes and functions described herein may be either part of the micro-instruction code or part of the application program, or any combination thereof, which may be executed by a CPU. Additionally, several other peripheral units may be connected to the computer platform such as an additional data storage unit and a printer unit.
ΡΕ2512136
As should be apparent to one skilled in the art, implementations may also produce a signal formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for carrying out a method, or data produced by one of the described implementations. Such a signal may be formatted, for example, as an electromagnetic wave (e.g. using a radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream, producing syntax, and modulating a carrier with the encoded data stream and syntax. The information that the signal carries may be, for example, analog or digital information. 0 A signal can be transmitted over a variety of different wired or wireless connections, as is known.
It is further to be understood that because some of the system constituent components and methods depicted in the accompanying drawings are preferably implemented in software, the actual connections between system components or process function blocks may differ depending on the manner in which the present principles are programmed. Given these teachings, one of ordinary skill in the pertinent art will be able to consider these and similar implementations or configurations of the present principles.
A number of implementations have been described. At the
However, it will be appreciated that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes may be substituted for those disclosed and the resulting implementations will perform at least substantially the same or same functions, at least substantially the same form (s), to achieve at least substantially the same (s). same result (s) as in the disclosed implementations.
Contents14
164 members in 21 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 92301407 | United States of America | P | |
| 92301407 | United States of America | P | |
| 92540007 | United States of America | P | |
| 92540007 | United States of America | P | |
| 923014P | – | – | – |
| 925400P | – | – | – |
| US20070923014P | – | – | – |
| US20070925400P | – | – | – |
Members164
| Document | Office | Kind | |
|---|---|---|---|
| AU2008239653A1 | Australia | A1 | |
| WO2008127676A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008127676A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008127676A9 | World Intellectual Property Organization (WIPO) | A9 | |
| MX2009010973A | Mexico | A | |
| EP2137975A2 | European Patent Office (EPO) | A2 | |
| KR20100016212A | Republic of Korea | A | |
| CN101658037A | China | A | |
| US2010046635A1 | United States of America | A1 | |
| JP2010524398A | Japan | A | |
| RU2009141712A | Russian Federation | A | |
| ZA201006649B | South Africa | B | |
| AU2008239653B2 | Australia | B2 | |
| EP2512135A1 | European Patent Office (EPO) | A1 | |
| EP2512136A1 | European Patent Office (EPO) | A1 | |
| ZA201201942B | South Africa | B | |
| BRPI0809510A2 | Brazil | A2 | |
| AU2012278382A1 | Australia | A1 | |
| CN101658037B | China | B | |
| JP5324563B2 | Japan | B2 | |
| BRPI0823512A2 | Brazil | A2 | |
| JP2013258716A | Japan | A | |
| RU2521618C2 | Russian Federation | C2 | |
| US8780998B2 | United States of America | B2 | |
| KR20140098825A | Republic of Korea | A | |
| US2014301479A1 | United States of America | A1 | |
| KR101467601B1 | Republic of Korea | B1 | |
| AU2012278382B2 | Australia | B2 | |
| JP5674873B2 | Japan | B2 | |
| EP2512135B1 | European Patent Office (EPO) | B1 | |
| KR20150046385A | Republic of Korea | A | |
| AU2008239653C1 | Australia | C1 | |
| JP2015092715A | Japan | A | |
| AU2015202314A1 | Australia | A1 | |
| EP2887671A1 | European Patent Office (EPO) | A1 | |
| US2015281736A1 | United States of America | A1 | |
| RU2014116612A | Russian Federation | A | |
| US9185384B2 | United States of America | B2 | |
| US2015341665A1 | United States of America | A1 | |
| US9219923B2 | United States of America | B2 | |
| US9232235B2 | United States of America | B2 | |
| US2016080757A1 | United States of America | A1 | |
| EP2512136B1 | European Patent Office (EPO) | B1 | |
| KR101646089B1 | Republic of Korea | B1 | |
| PT2512136TThis record | Portugal | T | |
| DK2512136T3 | Denmark | T3 | |
| US9445116B2 | United States of America | B2 | |
| AU2015202314B2 | Australia | B2 | |
| ES2586406T3 | Spain | T3 | |
| KR20160121604A | Republic of Korea | A | |
| PL2512136T3 | Poland | T3 | |
| US2016360218A1 | United States of America | A1 | |
| HUE029776T2 | Hungary | T2 | |
| US9706217B2 | United States of America | B2 | |
| JP2017135756A | Japan | A | |
| US2017257638A1 | United States of America | A1 | |
| KR20170106987A | Republic of Korea | A | |
| KR101766479B1 | Republic of Korea | B1 | |
| US9838705B2 | United States of America | B2 | |
| US2018048904A1 | United States of America | A1 | |
| RU2651227C2 | Russian Federation | C2 | |
| US9973771B2 | United States of America | B2 | |
| EP2887671B1 | European Patent Office (EPO) | B1 | |
| US9986254B1 | United States of America | B1 | |
| US2018152719A1 | United States of America | A1 | |
| ES2675164T3 | Spain | T3 | |
| PT2887671T | Portugal | T | |
| DK2887671T3 | Denmark | T3 | |
| TR201809177T4 | Türkiye | T4 | |
| LT2887671T | Lithuania | T | |
| US2018213246A1 | United States of America | A1 | |
| KR101885790B1 | Republic of Korea | B1 | |
| KR20180089560A | Republic of Korea | A | |
| PL2887671T3 | Poland | T3 | |
| HUE038192T2 | Hungary | T2 | |
| SI2887671T1 | Slovenia | T1 | |
| EP3399756A1 | European Patent Office (EPO) | A1 | |
| US10129557B2 | United States of America | B2 | |
| US2019037230A1 | United States of America | A1 | |
| RU2684184C1 | Russian Federation | C1 | |
| KR101965781B1 | Republic of Korea | B1 | |
| KR20190038680A | Republic of Korea | A | |
| US10298948B2 | United States of America | B2 | |
| US2019253727A1 | United States of America | A1 | |
| HK1255617A1 | Hong Kong, China | A1 | |
| US10432958B2 | United States of America | B2 | |
| BRPI0809510B1 | Brazil | B1 | |
| BR122018004903B1 | Brazil | B1 | |
| BR122018004904B1 | Brazil | B1 | |
| BR122018004906B1 | Brazil | B1 | |
| KR102044130B1 | Republic of Korea | B1 | |
| KR20190127999A | Republic of Korea | A | |
| JP2019201435A | Japan | A | |
| JP2019201436A | Japan | A | |
| US2019379897A1 | United States of America | A1 | |
| RU2709671C1 | Russian Federation | C1 | |
| RU2721941C1 | Russian Federation | C1 | |
| KR102123772B1 | Republic of Korea | B1 | |
| KR20200069389A | Republic of Korea | A | |
| US10764596B2 | United States of America | B2 |
Numbers
- Publication
- 2512136
- Publication, DOCDB
- 2512136
- Publication, EPODOC
- PT2512136T
- Application
- 121590608
- Application, DOCDB
- 12159060
- Application, EPODOC
- PT20120159060T
Titles2
- English
- TILING IN VIDEO ENCODING AND DECODING
- Portuguese
- ORGANIZAÇÃO DE IMAGENS EM MOSAICO EM CODIFICAÇÃO E DESCODIFICAÇÃO DE VÍDEO
Classification
- CPC, 11
- H04N19/597
- H04N19/46
- H04N19/70
- H04N19/61
- H04N19/172
- H04N19/182
- H04N2213/003
- H04N13/161
- H04N13/194
- H04N19/174
- H04N13/111
- IPC, 5
- H04N19 70
- H04N13 00
- H04N19 46
- H04N19 597
- H04N19 61