Frame packing for video coding
Summary by NHIP
Video Frame Packing
The method accesses a video picture containing multiple combined pictures and retrieves associated spatial interleaving and sampling information. This data specifies filtering directions for upsampling and defines relationships such as stereo views, 2D plus depth maps, or layer depth video formats.
Claim Score by NHIP
Abstract
Implementations are provided that relate, for example, to view tiling in video encoding and decoding. A particular implementation accesses a video picture that includes multiple pictures combined into a single picture, and accesses additional information indicating how the multiple pictures in the accessed video picture are combined. The accessed information includes spatial interleaving information and sampling information. Another implementation encodes a video picture that includes multiple pictures combined into a single picture, and generates information indicating how the multiple pictures in the accessed video picture are combined. The generated information includes spatial interleaving information and sampling information. A bitstream is formed that includes the encoded video picture and the generated information. Another implementation provides a data structure for transmitting the generated information.

Term
4.2 yearsleft in the term
Expires 20 December 2030, including 328 days of term adjustment.
- Priority
- Filed
- Granted
- Today
- Expires
4 claims: 1 independent, 3 dependent
- 1Broadest claimClaim Score 35, narrow(NHIP)A method, executed by a processor, comprising:accessing a video picture that includes multiple pictures combined into a single picture, the video picture being part of a received video stream;accessing information that is part of the received video stream, the accessed information indicating how the multiple pictures in the accessed video picture are combined, wherein the accessed information includes spatial interleaving information and sampling information, wherein the spatial interleaving information indicates spatial interleaving applied to the multiple pictures in forming the single picture and wherein the sampling information indicates one or more parameters related to an upsampling filter for restoring each of the multiple pictures to another resolution, the one or more parameters related to the upsampling filter including an indication of filtering direction, and wherein the spatial interleaving information further includes relationship information, the relationship information indicating a type of relationship that exists between the multiple pictures, the relationship information indicating that the multiple pictures are stereo views of an image, the multiple pictures are not related, the multiple pictures are a 2D image and its related depth map (2D+Z), the multiple pictures are multiple sets of a 2D+Z (MVD), the multiple pictures represent images in a layer depth video format (LDV), or the multiple pictures represent images in two sets of LDV (DES), and the multiple pictures are formatted according to a depth related format;and decoding the video picture to provide a decoded representation of at least one of the multiple pictures.
359 paragraphs in 6 sections, as filed
CROSS-REFERENCE TO RELATED APPLICATIONS
0001This application claims the benefit, under 35 U.S.C. §365 of International Application PCT/US2010/000194, filed Jan. 26, 2010, which was published in accordance with PCT Article 21(2) on Oct. 21, 2010 in English and which claims the benefit of U.S. provisional patent application Nos. 61/205,938 and 61/269,955 filed Jan. 26, 2009 and Jul. 1, 2009, respectively.
0002This application is also related to International Application No. PCT/US2008/004747, titled “Tiling In Video Encoding And Decoding” and having an International Filing Date of Apr. 11, 2008. This application is also identified as U.S. patent application Ser. No. 12/450,829.
TECHNICAL FIELD
0003Implementations are described that relate generally to the fields of video encoding and/or decoding.
BACKGROUND
0004With the emergence of 3D displays in the market, including stereoscopic and auto stereoscopic displays, there is a strong demand for more 3D content to be available. It is typically a challenging task to code the 3D content usually involving multiple views and possibly corresponding depth maps as well. Each frame of 3D content may require the system to handle a huge amount of data. In typical 3D video applications, multiview video signals need to be transmitted or stored efficiently due to limitations in transmission bandwidth, storage limitations, and processing limitation, for example. Multiview Video Coding (MVC) extends H.264/Advanced Video Coding (AVC) using high level syntax to facilitate the coding of multiple views. This syntax aids in the subsequent handling of the 3D images by image processors.
0005H.264/AVC, though designed ostensibly for 2D video, can also be used to transmit stereo contents by exploiting a frame-packing technique. The technique of frame-packing is presented simply as follows: on the encoder side, two views or pictures are generally downsampled for packing into one single video frame, which is then supplied to a H.264/AVC encoder for output as a bitstream; on the decoder side, the bitstream is decoded and the recovered frame is then unpacked. Unpacking permits the extraction of the two original views from the recovered frame and generally involves an upsampling operation to restore the original size to each view so that the views can be rendered for display. This approach is able to be used for two or more views, such as with multi-view images or with depth information and the like.
0006Frame packing may rely on the existence of ancillary information associated with the frame and its views. Supplemental enhancement information (SEI) messages may be used to convey some frame-packing information. As an example, in a draft amendment of AVC, it has been proposed that an SEI message be used to inform a decoder of various spatial interleaving characteristics of a packed picture, including that the constituent pictures are formed by checkerboard spatial interleaving. By employing the SEI message, it is possible to encode the checkerboard interleaved picture of stereo video images using AVC directly. <figref idref="DRAWINGS">FIG. 26</figref> shows a known example of checkerboard interleaving. To date, however, the SEI message contents and the contents of other high level syntaxes have been limited in conveying information relevant to pictures or views that have been subjected to frame packing.
SUMMARY
0007According to a general aspect, a video picture is encoded that includes multiple pictures combined into a single picture. Information is generated indicating how the multiple pictures in the accessed video picture are combined. The generated information includes spatial interleaving information and sampling information. The spatial interleaving information indicates spatial interleaving applied to the multiple pictures in forming the single picture. The sampling information indicates one or more parameters related to an upsampling filter for restoring each of the multiple pictures to a desired resolution. The one or more parameters related to the upsampling filter include an indication of filtering direction. A bitstream is formed that includes the encoded video picture and the generated information. The generated information provides information for use in processing the encoded video picture.
0008According to another general aspect, a video signal or video structure includes an encoded picture section and a signaling section. The encoded picture section includes an encoding of a video picture, the video picture including multiple pictures combined into a single picture. The signaling section includes an encoding of generated information indicating how the multiple pictures in the accessed video picture are combined. The generated information includes spatial interleaving information and sampling information. The spatial interleaving information indicates spatial interleaving applied to the multiple pictures in forming the single picture. The sampling information indicates one or more parameters related to an upsampling filter for restoring each of the multiple pictures to a desired resolution. The one or more parameters related to the upsampling filter includes an indication of filtering direction. The generated information provides information for use in decoding the encoded video picture.
0009According to another general aspect, a video picture is accessed that includes multiple pictures combined into a single picture, the video picture being part of a received video stream. Information is accessed that is part of the received video stream, the accessed information indicating how the multiple pictures in the accessed video picture are combined. The accessed information includes spatial interleaving information and sampling information. The spatial interleaving information indicates spatial interleaving applied to the multiple pictures in forming the single picture. The sampling information indicates one or more parameters related to an upsampling filter for restoring each of the multiple pictures to a desired resolution. The one or more parameters related to the upsampling filter includes an indication of filtering direction. The video picture is decoded to provide a decoded representation of at least one of the multiple pictures.
0010The details of one or more implementations are set forth in the accompanying drawings and the description below. Even if described in one particular manner, it should be clear that implementations may be configured or embodied in various manners. For example, an implementation may be performed as a method, or embodied as an apparatus configured to perform a set of operations, or embodied as an apparatus storing instructions for performing a set of operations, or embodied in a signal. Other aspects and features will become apparent from the following detailed description considered in conjunction with the accompanying drawings and the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
0011<figref idref="DRAWINGS">FIG. 1</figref> is a diagram showing an example of four views tiled on a single frame.
0012<figref idref="DRAWINGS">FIG. 2</figref> is a diagram showing an example of four views flipped and tiled on a single frame.
0013<figref idref="DRAWINGS">FIG. 3</figref> shows a block diagram for an implementation of a video encoder to which the present principles may be applied.
0014<figref idref="DRAWINGS">FIG. 4</figref> shows a block diagram for an implementation of a video decoder to which the present principles may be applied.
0015<figref idref="DRAWINGS">FIG. 5</figref> is a flow diagram for an implementation of a method for encoding pictures for a plurality of views using the MPEG-4 AVC Standard.
0016<figref idref="DRAWINGS">FIG. 6</figref> is a flow diagram for an implementation of a method for decoding pictures for a plurality of views using the MPEG-4 AVC Standard.
0017<figref idref="DRAWINGS">FIG. 7</figref> is a flow diagram for an implementation of a method for encoding pictures for a plurality of views and depths using the MPEG-4 AVC Standard.
0018<figref idref="DRAWINGS">FIG. 8</figref> is a flow diagram for an implementation of a method for decoding pictures for a plurality of views and depths using the MPEG-4 AVC Standard.
0019<figref idref="DRAWINGS">FIG. 9</figref> is a diagram showing an example of a depth signal.
0020<figref idref="DRAWINGS">FIG. 10</figref> is a diagram showing an example of a depth signal added as a tile.
0021<figref idref="DRAWINGS">FIG. 11</figref> is a diagram showing an example of 5 views tiled on a single frame.
0022<figref idref="DRAWINGS">FIG. 12</figref> is a block diagram for an exemplary Multi-view Video Coding (MVC) encoder to which the present principles may be applied.
0023<figref idref="DRAWINGS">FIG. 13</figref> is a block diagram for an exemplary Multi-view Video Coding (MVC) decoder to which the present principles may be applied.
0024<figref idref="DRAWINGS">FIG. 14</figref> is a flow diagram for an implementation of a method for processing pictures for a plurality of views in preparation for encoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0025<figref idref="DRAWINGS">FIG. 15</figref> is a flow diagram for an implementation of a method for encoding pictures for a plurality of views using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0026<figref idref="DRAWINGS">FIG. 16</figref> is a flow diagram for an implementation of a method for processing pictures for a plurality of views in preparation for decoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0027<figref idref="DRAWINGS">FIG. 17</figref> is a flow diagram for an implementation of a method for decoding pictures for a plurality of views using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0028<figref idref="DRAWINGS">FIG. 18</figref> is a flow diagram for an implementation of a method for processing pictures for a plurality of views and depths in preparation for encoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0029<figref idref="DRAWINGS">FIG. 19</figref> is a flow diagram for an implementation of a method for encoding pictures for a plurality of views and depths using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0030<figref idref="DRAWINGS">FIG. 20</figref> is a flow diagram for an implementation of a method for processing pictures for a plurality of views and depths in preparation for decoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0031<figref idref="DRAWINGS">FIG. 21</figref> is a flow diagram for an implementation of a method for decoding pictures for a plurality of views and depths using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard.
0032<figref idref="DRAWINGS">FIG. 22</figref> is a diagram showing tiling examples at the pixel level.
0033<figref idref="DRAWINGS">FIG. 23</figref> shows a block diagram for an implementation of a video processing device to which the present principles may be applied.
0034<figref idref="DRAWINGS">FIG. 24</figref> shows a simplified diagram of an exemplary 3D video system.
0035<figref idref="DRAWINGS">FIG. 25</figref> shows exemplary left and right depth maps for an image from different reference views.
0036<figref idref="DRAWINGS">FIG. 26</figref> shows an exemplary block diagram for spatially interleaving two constituent pictures into a single picture or frame using checkerboard interleaving.
0037<figref idref="DRAWINGS">FIG. 27</figref> shows an exemplary picture of side-by-side spatial interleaving of two constituent pictures.
0038<figref idref="DRAWINGS">FIG. 28</figref> shows an exemplary picture of top-bottom spatial interleaving of two constituent pictures.
0039<figref idref="DRAWINGS">FIG. 29</figref> shows an exemplary picture of row-by-row spatial interleaving of two constituent pictures.
0040<figref idref="DRAWINGS">FIG. 30</figref> shows an exemplary picture of column-by-column spatial interleaving of two constituent pictures.
0041<figref idref="DRAWINGS">FIG. 31</figref> shows an exemplary picture of side-by-side spatial interleaving of two constituent pictures in which the right hand picture is flipped horizontally.
0042<figref idref="DRAWINGS">FIG. 32</figref> shows an exemplary picture of top-bottom spatial interleaving of two constituent pictures in which the bottom picture is flipped vertically.
0043<figref idref="DRAWINGS">FIG. 33</figref> shows an exemplary interleaved picture or frame in which the constituent pictures represent a layer depth video (LDV) format.
0044<figref idref="DRAWINGS">FIG. 34</figref> shows an exemplary interleaved picture or frame in which the constituent pictures represent a 2D plus depth format.
0045<figref idref="DRAWINGS">FIGS. 35-38</figref> show the exemplary flowcharts of different embodiments for handling the encoding and decoding of video images using SEI messages for frame packing information.
0046<figref idref="DRAWINGS">FIG. 39</figref> shows an exemplary video transmission system to which the present principles may be applied.
0047<figref idref="DRAWINGS">FIG. 40</figref> shows an exemplary video receiving system to which the present principles may be applied.
0048<figref idref="DRAWINGS">FIG. 41</figref> shows an exemplary video processing device to which the present principles may be applied.
0049The exemplary embodiments set out herein illustrate various embodiments, and such exemplary embodiments are not to be construed as limiting the scope of this disclosure in any manner.
DETAILED DESCRIPTION
0050Various implementations are directed to methods and apparatus for view tiling in video encoding and decoding. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the present principles and are included within its spirit and scope.
0051All examples and conditional language recited herein are intended for pedagogical purposes to aid the reader in understanding the present principles and the concepts contributed by the inventor(s) to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions.
0052Moreover, all statements herein reciting principles, aspects, and embodiments of the present principles, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents as well as equivalents developed in the future, that is, any elements developed that perform the same function, regardless of structure.
0053Thus, for example, it will be appreciated by those skilled in the art that the block diagrams presented herein represent conceptual views of illustrative system components and/or circuitry. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudocode, and the like represent various processes which may be substantially represented in computer readable media and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
0054The functions of the various elements shown in the figures may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (“DSP”) hardware, read-only memory (“ROM”) for storing software, random access memory (“RAM”), and non-volatile storage.
0055Other hardware, conventional and/or custom, may also be included in the realization of various implementations. For example, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
0056In the claims hereof, any element expressed as a means for performing a specified function is intended to encompass any way of performing that function including, for example, a) a combination of circuit elements that performs that function or b) software in any form, including, therefore, firmware, microcode or the like, combined with appropriate circuitry for executing that software to perform the function. The present principles as defined by such claims reside in the fact that the functionalities provided by the various recited means are combined and brought together in the manner which the claims call for. It is thus regarded that any means that can provide those functionalities are equivalent to those shown herein.
0057Reference in the specification to “one embodiment” (or “one implementation”) or “an embodiment” (or “an implementation”) of the present principles means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment of the present principles. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” appearing in various places throughout the specification are not necessarily all referring to the same embodiment.
0058It is to be appreciated that while one or more embodiments of the present principles are described herein with respect to the MPEG-4 AVC standard, the principles described in this application are not limited to solely this standard and, thus, may be utilized with respect to other standards, recommendations, and extensions thereof, particularly video coding standards, recommendations, and extensions thereof, including extensions of the MPEG-4 AVC standard, while maintaining the spirit of the principles of this application.
0059Further, it is to be appreciated that while one or more other embodiments are described herein with respect to the multi-view video coding extension of the MPEG-4 AVC standard, the present principles are not limited to solely this extension and/or this standard and, thus, may be utilized with respect to other video coding standards, recommendations, and extensions thereof relating to multi-view video coding, while maintaining the spirit of the principles of this application. Multi-view video coding (MVC) is the compression framework for the encoding of multi-view sequences. A Multi-view Video Coding (MVC) sequence is a set of two or more video sequences that capture the same scene from a different view point.
0060Also, it is to be appreciated that while one or more other embodiments are described herein that use depth information with respect to video content, the principles of this application are not limited to such embodiments and, thus, other embodiments may be implemented that do not use depth information, while maintaining the spirit of the present principles.
0061Additionally, as used herein, “high level syntax” refers to syntax present in the bitstream that resides hierarchically above the macroblock layer. For example, high level syntax, as used herein, may refer to, but is not limited to, syntax at the slice header level, Supplemental Enhancement Information (SEI) level, Picture Parameter Set (PPS) level, Sequence Parameter Set (SPS) level, View Parameter Set (VPS), and Network Abstraction Layer (NAL) unit header level.
0062In the current implementation of multi-video coding (MVC) based on the International Organization for Standardization/International Electrotechnical Commission (ISO/IEC) Moving Picture Experts Group-4 (MPEG-4) Part 10 Advanced Video Coding (AVC) standard/International Telecommunication Union, Telecommunication Sector (ITU-T) H.264 Recommendation (hereinafter the “MPEG-4 AVC Standard”), the reference software achieves multi-view prediction by encoding each view with a single encoder and taking into consideration the cross-view references. Each view is coded as a separate bitstream by the encoder in its original resolution and later all the bitstreams are combined to form a single bitstream which is then decoded. Each view produces a separate YUV decoded output.
0063An exemplary video system supporting the production and use of 3D images is presented schematically in <figref idref="DRAWINGS">FIG. 24</figref>. The content production side of the system shows image capture by various means including, but not limited to, stereo cameras, a depth camera, multiple cameras operating simultaneously, and conversion of 2D images to 3D images. An example of depth map information (for example, Z information) captured for left and right views of the same scene is shown in <figref idref="DRAWINGS">FIG. 25</figref>. Each of these approaches not only captures the video image content, but some also generate certain depth information associated with the captured video images. Once processed and coded, all this information is available to be distributed, transmitted, and ultimately rendered. Meta-data is also generated with the video content for use in the subsequent rendering of the 3D video. Rendering can take place using 2D display systems or 3D displays. The 3D displays can vary from stereo displays to multi-view 3D displays as shown in the figure.
0064Another approach for multi-view prediction involves grouping a set of views into pseudo views. In one example of this approach, we can tile the pictures from every N views out of the total M views (sampled at the same time) on a larger frame or a super frame with possible downsampling or other operations. Turning to <figref idref="DRAWINGS">FIG. 1</figref>, an example of four views tiled on a single frame is indicated generally by the reference numeral <b>100</b>. All four views are in their normal orientation.
0065Turning to <figref idref="DRAWINGS">FIG. 2</figref>, an example of four views flipped and tiled on a single frame is indicated generally by the reference numeral <b>200</b>. The top-left view is in its normal orientation. The top-right view is flipped horizontally. The bottom-left view is flipped vertically. The bottom-right view is flipped both horizontally and vertically. Thus, if there are four views, then a picture from each view is arranged in a super-frame like a tile. This results in a single un-coded input sequence with a large resolution.
0066Alternatively, we can downsample the image to produce a smaller resolution. Thus, we create multiple sequences which each include different views that are tiled together. Each such sequence then forms a pseudo view, where each pseudo view includes N different tiled views. <figref idref="DRAWINGS">FIG. 1</figref> shows one pseudo-view, and <figref idref="DRAWINGS">FIG. 2</figref> shows another pseudo-view. These pseudo views can then be encoded using existing video coding standards such as the ISO/IEC MPEG-2 Standard and the MPEG-4 AVC Standard.
0067Yet another approach for multi-view coding simply involves encoding the different views independently using a new standard and, after decoding, tiling the views as required by the player.
0068Further, in another approach, the views can also be tiled in a pixel wise way. For example, in a super view that is composed of four views, pixel (x, y) may be from view 0, while pixel (x+1, y) may be from view 1, pixel (x, y+1) may be from view 2, and pixel (x+1, y+1) may be from view 3.
0069Many displays manufacturers use such a framework of arranging or tiling different views on a single frame and then extracting the views from their respective locations and rendering them. In such cases, there is no standard way to determine if the bitstream has such a property. Thus, if a system uses the method of tiling pictures of different views in a large frame, then the method of extracting the different views is proprietary.
0070However, there is no standard way to determine if the bitstream has such a property. We propose high level syntax in order to facilitate the renderer or player to extract such information in order to assist in display or other post-processing. It is also possible the sub-pictures have different resolutions and some upsampling may be needed to eventually render the view. The user may want to have the method of upsample indicated in the high level syntax as well. Additionally, parameters to change the depth focus can also be transmitted.
0071In an embodiment, we propose a new Supplemental Enhancement Information (SEI) message for signaling multi-view information in a MPEG-4 AVC Standard compatible bitstream where each picture includes sub-pictures which belong to a different view. The embodiment is intended, for example, for the easy and convenient display of multi-view video streams on three-dimensional (3D) monitors which may use such a framework. The concept can be extended to other video coding standards and recommendations signaling such information using high level syntax.
0072Moreover, in an embodiment, we propose a signaling method of how to arrange views before they are sent to the multi-view video encoder and/or decoder. Advantageously, the embodiment may lead to a simplified implementation of the multi-view coding, and may benefit the coding efficiency. Certain views can be put together and form a pseudo view or super view and then the tiled super view is treated as a normal view by a common multi-view video encoder and/or decoder, for example, as per the current MPEG-4 AVC Standard based implementation of Multi-view Video Coding. A new flag shown in Table 1 is proposed in the Sequence Parameter Set (SPS) extension of multi-view video coding to signal the use of the technique of pseudo views. The embodiment is intended for the easy and convenient display of multi-view video streams on 3D monitors which may use such a framework.
0073Another approach for multi-view coding involves tiling the pictures from each view (sampled at the same time) on a larger frame or a super frame with a possible downsampling operation. Turning to <figref idref="DRAWINGS">FIG. 1</figref>, an example of four views tiled on a single frame is indicated generally by the reference numeral <b>100</b>. Turning to <figref idref="DRAWINGS">FIG. 2</figref>, an example of four views flipped and tiled on a single frame is indicated generally by the reference numeral <b>200</b>. Thus, if there are four views, then a picture from each view is arranged in a super-frame like a tile. This results in a single un-coded input sequence with a large resolution. This signal can then be encoded using existing video coding standards such as the ISO/IEC MPEG-2 Standard and the MPEG-4 AVC Standard.
0074Turning to <figref idref="DRAWINGS">FIG. 3</figref>, a video encoder capable of performing video encoding in accordance with the MPEG-4 AVC standard is indicated generally by the reference numeral <b>300</b>.
0075The video encoder <b>300</b> includes a frame ordering buffer <b>310</b> having an output in signal communication with a non-inverting input of a combiner <b>385</b>. An output of the combiner <b>385</b> is connected in signal communication with a first input of a transformer and quantizer <b>325</b>. An output of the transformer and quantizer <b>325</b> is connected in signal communication with a first input of an entropy coder <b>345</b> and a first input of an inverse transformer and inverse quantizer <b>350</b>. An output of the entropy coder <b>345</b> is connected in signal communication with a first non-inverting input of a combiner <b>390</b>. An output of the combiner <b>390</b> is connected in signal communication with a first input of an output buffer <b>335</b>.
0076A first output of an encoder controller <b>305</b> is connected in signal communication with a second input of the frame ordering buffer <b>310</b>, a second input of the inverse transformer and inverse quantizer <b>350</b>, an input of a picture-type decision module <b>315</b>, an input of a macroblock-type (MB-type) decision module <b>320</b>, a second input of an intra prediction module <b>360</b>, a second input of a deblocking filter <b>365</b>, a first input of a motion compensator <b>370</b>, a first input of a motion estimator <b>375</b>, and a second input of a reference picture buffer <b>380</b>.
0077A second output of the encoder controller <b>305</b> is connected in signal communication with a first input of a Supplemental Enhancement Information (SEI) inserter <b>330</b>, a second input of the transformer and quantizer <b>325</b>, a second input of the entropy coder <b>345</b>, a second input of the output buffer <b>335</b>, and an input of the Sequence Parameter Set (SPS) and Picture Parameter Set (PPS) inserter <b>340</b>.
0078A first output of the picture-type decision module <b>315</b> is connected in signal communication with a third input of a frame ordering buffer <b>310</b>. A second output of the picture-type decision module <b>315</b> is connected in signal communication with a second input of a macroblock-type decision module <b>320</b>.
0079An output of the Sequence Parameter Set (SPS) and Picture Parameter Set (PPS) inserter <b>340</b> is connected in signal communication with a third non-inverting input of the combiner <b>390</b>. An output of the SEI Inserter <b>330</b> is connected in signal communication with a second non-inverting input of the combiner <b>390</b>.
0080An output of the inverse quantizer and inverse transformer <b>350</b> is connected in signal communication with a first non-inverting input of a combiner <b>319</b>. An output of the combiner <b>319</b> is connected in signal communication with a first input of the intra prediction module <b>360</b> and a first input of the deblocking filter <b>365</b>. An output of the deblocking filter <b>365</b> is connected in signal communication with a first input of a reference picture buffer <b>380</b>. An output of the reference picture buffer <b>380</b> is connected in signal communication with a second input of the motion estimator <b>375</b> and with a first input of a motion compensator <b>370</b>. A first output of the motion estimator <b>375</b> is connected in signal communication with a second input of the motion compensator <b>370</b>. A second output of the motion estimator <b>375</b> is connected in signal communication with a third input of the entropy coder <b>345</b>.
0081An output of the motion compensator <b>370</b> is connected in signal communication with a first input of a switch <b>397</b>. An output of the intra prediction module <b>360</b> is connected in signal communication with a second input of the switch <b>397</b>. An output of the macroblock-type decision module <b>320</b> is connected in signal communication with a third input of the switch <b>397</b> in order to provide a control input to the switch <b>397</b>. The third input of the switch <b>397</b> determines whether or not the “data” input of the switch (as compared to the control input, that is, the third input) is to be provided by the motion compensator <b>370</b> or the intra prediction module <b>360</b>. The output of the switch <b>397</b> is connected in signal communication with a second non-inverting input of the combiner <b>319</b> and with an inverting input of the combiner <b>385</b>.
0082Inputs of the frame ordering buffer <b>310</b> and the encoder controller <b>105</b> are available as input of the encoder <b>300</b>, for receiving an input picture <b>301</b>. Moreover, an input of the Supplemental Enhancement Information (SEI) inserter <b>330</b> is available as an input of the encoder <b>300</b>, for receiving metadata. An output of the output buffer <b>335</b> is available as an output of the encoder <b>300</b>, for outputting a bitstream.
0083Turning to <figref idref="DRAWINGS">FIG. 4</figref>, a video decoder capable of performing video decoding in accordance with the MPEG-4 AVC standard is indicated generally by the reference numeral <b>400</b>.
0084The video decoder <b>400</b> includes an input buffer <b>410</b> having an output connected in signal communication with a first input of the entropy decoder <b>445</b>. A first output of the entropy decoder <b>445</b> is connected in signal communication with a first input of an inverse transformer and inverse quantizer <b>450</b>. An output of the inverse transformer and inverse quantizer <b>450</b> is connected in signal communication with a second non-inverting input of a combiner <b>425</b>. An output of the combiner <b>425</b> is connected in signal communication with a second input of a deblocking filter <b>465</b> and a first input of an intra prediction module <b>460</b>. A second output of the deblocking filter <b>465</b> is connected in signal communication with a first input of a reference picture buffer <b>480</b>. An output of the reference picture buffer <b>480</b> is connected in signal communication with a second input of a motion compensator <b>470</b>.
0085A second output of the entropy decoder <b>445</b> is connected in signal communication with a third input of the motion compensator <b>470</b> and a first input of the deblocking filter <b>465</b>. A third output of the entropy decoder <b>445</b> is connected in signal communication with an input of a decoder controller <b>405</b>. A first output of the decoder controller <b>405</b> is connected in signal communication with a second input of the entropy decoder <b>445</b>. A second output of the decoder controller <b>405</b> is connected in signal communication with a second input of the inverse transformer and inverse quantizer <b>450</b>. A third output of the decoder controller <b>405</b> is connected in signal communication with a third input of the deblocking filter <b>465</b>. A fourth output of the decoder controller <b>405</b> is connected in signal communication with a second input of the intra prediction module <b>460</b>, with a first input of the motion compensator <b>470</b>, and with a second input of the reference picture buffer <b>480</b>.
0086An output of the motion compensator <b>470</b> is connected in signal communication with a first input of a switch <b>497</b>. An output of the intra prediction module <b>460</b> is connected in signal communication with a second input of the switch <b>497</b>. An output of the switch <b>497</b> is connected in signal communication with a first non-inverting input of the combiner <b>425</b>.
0087An input of the input buffer <b>410</b> is available as an input of the decoder <b>400</b>, for receiving an input bitstream. A first output of the deblocking filter <b>465</b> is available as an output of the decoder <b>400</b>, for outputting an output picture.
0088Turning to <figref idref="DRAWINGS">FIG. 5</figref>, an exemplary method for encoding pictures for a plurality of views using the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>500</b>.
0089The method <b>500</b> includes a start block <b>502</b> that passes control to a function block <b>504</b>. The function block <b>504</b> arranges each view at a particular time instance as a sub-picture in tile format, and passes control to a function block <b>506</b>. The function block <b>506</b> sets a syntax element num_coded_views_minus 1, and passes control to a function block <b>508</b>. The function block <b>508</b> sets syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1, and passes control to a function block <b>510</b>. The function block <b>510</b> sets a variable i equal to zero, and passes control to a decision block <b>512</b>. The decision block <b>512</b> determines whether or not the variable i is less than the number of views. If so, then control is passed to a function block <b>514</b>. Otherwise, control is passed to a function block <b>524</b>.
0090The function block <b>514</b> sets a syntax element view_id[i], and passes control to a function block <b>516</b>. The function block <b>516</b> sets a syntax element num_parts[view_id[i]], and passes control to a function block <b>518</b>. The function block <b>518</b> sets a variable j equal to zero, and passes control to a decision block <b>520</b>. The decision block <b>520</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, then control is passed to a function block <b>522</b>. Otherwise, control is passed to a function block <b>528</b>.
0091The function block <b>522</b> sets the following syntax elements, increments the variable j, and then returns control to the decision block <b>520</b>: depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]][j]; frame_crop_top_offset[view_id[i]][j]; and frame_crop_bottom_offset[view_id[i]][j].
0092The function block <b>528</b> sets a syntax element upsample_view_flag[view_id[i]], and passes control to a decision block <b>530</b>. The decision block <b>530</b> determines whether or not the current value of the syntax element upsample_view_flag[view_id[i]] is equal to one. If so, then control is passed to a function block <b>532</b>. Otherwise, control is passed to a decision block <b>534</b>.
0093The function block <b>532</b> sets a syntax element upsample_filter[view_id[i]], and passes control to the decision block <b>534</b>.
0094The decision block <b>534</b> determines whether or not the current value of the syntax element upsample_filter[view_id[i]] is equal to three. If so, then control is passed to a function block <b>536</b>. Otherwise, control is passed to a function block <b>540</b>.
0095The function block <b>536</b> sets the following syntax elements and passes control to a function block <b>538</b>: vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]].
0096The function block <b>538</b> sets the filter coefficients for each YUV component, and passes control to the function block <b>540</b>.
0097The function block <b>540</b> increments the variable i, and returns control to the decision block <b>512</b>.
0098The function block <b>524</b> writes these syntax elements to at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Supplemental Enhancement Information (SEI) message, Network Abstraction Layer (NAL) unit header, and slice header, and passes control to a function block <b>526</b>. The function block <b>526</b> encodes each picture using the MPEG-4 AVC Standard or other single view codec, and passes control to an end block <b>599</b>.
0099Turning to <figref idref="DRAWINGS">FIG. 6</figref>, an exemplary method for decoding pictures for a plurality of views using the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>600</b>.
0100The method <b>600</b> includes a start block <b>602</b> that passes control to a function block <b>604</b>. The function block <b>604</b> parses the following syntax elements from at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Supplemental Enhancement Information (SEI) message, Network Abstraction Layer (NAL) unit header, and slice header, and passes control to a function block <b>606</b>. The function block <b>606</b> parses a syntax element num_coded_views_minus1, and passes control to a function block <b>608</b>. The function block <b>608</b> parses syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1, and passes control to a function block <b>610</b>. The function block <b>610</b> sets a variable i equal to zero, and passes control to a decision block <b>612</b>. The decision block <b>612</b> determines whether or not the variable i is less than the number of views. If so, then control is passed to a function block <b>614</b>. Otherwise, control is passed to a function block <b>624</b>.
0101The function block <b>614</b> parses a syntax element view_id[i], and passes control to a function block <b>616</b>. The function block <b>616</b> parses a syntax element num_parts_minus1 [view_id[i]], and passes control to a function block <b>618</b>. The function block <b>618</b> sets a variable j equal to zero, and passes control to a decision block <b>620</b>. The decision block <b>620</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, then control is passed to a function block <b>622</b>. Otherwise, control is passed to a function block <b>628</b>.
0102The function block <b>622</b> parses the following syntax elements, increments the variable j, and then returns control to the decision block <b>620</b>: depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]] [j]; frame_crop_top_offset[view_id[i]][j]; and frame_crop_bottom_offset[view_id[i]][j].
0103The function block <b>628</b> parses a syntax element upsample_view_flag[view_id[i]], and passes control to a decision block <b>630</b>. The decision block <b>630</b> determines whether or not the current value of the syntax element upsample_view_flag[view_id[i]] is equal to one. If so, then control is passed to a function block <b>632</b>. Otherwise, control is passed to a decision block <b>634</b>.
0104The function block <b>632</b> parses a syntax element upsample_filter[view_id[i]], and passes control to the decision block <b>634</b>.
0105The decision block <b>634</b> determines whether or not the current value of the syntax element upsample_filter[view_id[i]] is equal to three. If so, then control is passed to a function block <b>636</b>. Otherwise, control is passed to a function block <b>640</b>.
0106The function block <b>636</b> parses the following syntax elements and passes control to a function block <b>638</b>: vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]].
0107The function block <b>638</b> parses the filter coefficients for each YUV component, and passes control to the function block <b>640</b>.
0108The function block <b>640</b> increments the variable i, and returns control to the decision block <b>612</b>.
0109The function block <b>624</b> decodes each picture using the MPEG-4 AVC Standard or other single view codec, and passes control to a function block <b>626</b>. The function block <b>626</b> separates each view from the picture using the high level syntax, and passes control to an end block <b>699</b>.
0110Turning to <figref idref="DRAWINGS">FIG. 7</figref>, an exemplary method for encoding pictures for a plurality of views and depths using the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>700</b>.
0111The method <b>700</b> includes a start block <b>702</b> that passes control to a function block <b>704</b>. The function block <b>704</b> arranges each view and corresponding depth at a particular time instance as a sub-picture in tile format, and passes control to a function block <b>706</b>. The function block <b>706</b> sets a syntax element num_coded_views_minus1, and passes control to a function block <b>708</b>. The function block <b>708</b> sets syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1, and passes control to a function block <b>710</b>. The function block <b>710</b> sets a variable i equal to zero, and passes control to a decision block <b>712</b>. The decision block <b>712</b> determines whether or not the variable i is less than the number of views. If so, then control is passed to a function block <b>714</b>. Otherwise, control is passed to a function block <b>724</b>.
0112The function block <b>714</b> sets a syntax element view_id[i], and passes control to a function block <b>716</b>. The function block <b>716</b> sets a syntax element num_parts[view_id[i]], and passes control to a function block <b>718</b>. The function block <b>718</b> sets a variable j equal to zero, and passes control to a decision block <b>720</b>. The decision block <b>720</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, then control is passed to a function block <b>722</b>. Otherwise, control is passed to a function block <b>728</b>.
0113The function block <b>722</b> sets the following syntax elements, increments the variable j, and then returns control to the decision block <b>720</b>: depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]] [j]; frame_crop_top_offset[view_id[i]][j]; and frame_crop_bottom_offset[view_id[i]] [j].
0114The function block <b>728</b> sets a syntax element upsample_view_flag[view_id[i]], and passes control to a decision block <b>730</b>. The decision block <b>730</b> determines whether or not the current value of the syntax element upsample_view_flag[view_id[i]] is equal to one. If so, then control is passed to a function block <b>732</b>. Otherwise, control is passed to a decision block <b>734</b>.
0115The function block <b>732</b> sets a syntax element upsample_filter[view_id[i]], and passes control to the decision block <b>734</b>.
0116The decision block <b>734</b> determines whether or not the current value of the syntax element upsample_filter[view_id[i]] is equal to three. If so, then control is passed to a function block <b>736</b>. Otherwise, control is passed to a function block <b>740</b>.
0117The function block <b>736</b> sets the following syntax elements and passes control to a function block <b>738</b>: vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]].
0118The function block <b>738</b> sets the filter coefficients for each YUV component, and passes control to the function block <b>740</b>.
0119The function block <b>740</b> increments the variable i, and returns control to the decision block <b>712</b>.
0120The function block <b>724</b> writes these syntax elements to at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Supplemental Enhancement Information (SEI) message, Network Abstraction Layer (NAL) unit header, and slice header, and passes control to a function block <b>726</b>. The function block <b>726</b> encodes each picture using the MPEG-4 AVC Standard or other single view codec, and passes control to an end block <b>799</b>.
0121Turning to <figref idref="DRAWINGS">FIG. 8</figref>, an exemplary method for decoding pictures for a plurality of views and depths using the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>800</b>.
0122The method <b>800</b> includes a start block <b>802</b> that passes control to a function block <b>804</b>. The function block <b>804</b> parses the following syntax elements from at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Supplemental Enhancement Information (SEI) message, Network Abstraction Layer (NAL) unit header, and slice header, and passes control to a function block <b>806</b>. The function block <b>806</b> parses a syntax element num_coded_views_minus1, and passes control to a function block <b>808</b>. The function block <b>808</b> parses syntax elements org_pic_width_in_mbs_minus1 and org_pic_height_in_mbs_minus1, and passes control to a function block <b>810</b>. The function block <b>810</b> sets a variable i equal to zero, and passes control to a decision block <b>812</b>. The decision block <b>812</b> determines whether or not the variable i is less than the number of views. If so, then control is passed to a function block <b>814</b>. Otherwise, control is passed to a function block <b>824</b>.
0123The function block <b>814</b> parses a syntax element view_id[i], and passes control to a function block <b>816</b>. The function block <b>816</b> parses a syntax element num_parts_minus1[view_id[i]], and passes control to a function block <b>818</b>. The function block <b>818</b> sets a variable j equal to zero, and passes control to a decision block <b>820</b>. The decision block <b>820</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[view_id[i]]. If so, then control is passed to a function block <b>822</b>. Otherwise, control is passed to a function block <b>828</b>.
0124The function block <b>822</b> parses the following syntax elements, increments the variable j, and then returns control to the decision block <b>820</b>: depth_flag[view_id[i]][j]; flip_dir[view_id[i]][j]; loc_left_offset[view_id[i]][j]; loc_top_offset[view_id[i]][j]; frame_crop_left_offset[view_id[i]][j]; frame_crop_right_offset[view_id[i]][j]; frame_crop_top_offset[view_id[i]][j]; and frame_crop_bottom_offset[view_id[i]][j].
0125The function block <b>828</b> parses a syntax element upsample_view_flag[view_id[i]], and passes control to a decision block <b>830</b>. The decision block <b>830</b> determines whether or not the current value of the syntax element upsample_view_flag[view_id[i]] is equal to one. If so, then control is passed to a function block <b>832</b>. Otherwise, control is passed to a decision block <b>834</b>.
0126The function block <b>832</b> parses a syntax element upsample_filter[view_id[i]], and passes control to the decision block <b>834</b>.
0127The decision block <b>834</b> determines whether or not the current value of the syntax element upsample_filter[view_id[i]] is equal to three. If so, then control is passed to a function block <b>836</b>. Otherwise, control is passed to a function block <b>840</b>.
0128The function block <b>836</b> parses the following syntax elements and passes control to a function block <b>838</b>: vert_dim[view_id[i]]; hor_dim[view_id[i]]; and quantizer[view_id[i]].
0129The function block <b>838</b> parses the filter coefficients for each YUV component, and passes control to the function block <b>840</b>.
0130The function block <b>840</b> increments the variable i, and returns control to the decision block <b>812</b>.
0131The function block <b>824</b> decodes each picture using the MPEG-4 AVC Standard or other single view codec, and passes control to a function block <b>826</b>. The function block <b>826</b> separates each view and corresponding depth from the picture using the high level syntax, and passes control to a function block <b>827</b>. The function block <b>827</b> potentially performs view synthesis using the extracted view and depth signals, and passes control to an end block <b>899</b>.
0132With respect to the depth used in <figref idref="DRAWINGS">FIGS. 7 and 8</figref>, <figref idref="DRAWINGS">FIG. 9</figref> shows an example of a depth signal <b>900</b>, where depth is provided as a pixel value for each corresponding location of an image (not shown). Further, <figref idref="DRAWINGS">FIG. 10</figref> shows an example of two depth signals included in a tile <b>1000</b>. The top-right portion of tile <b>1000</b> is a depth signal having depth values corresponding to the image on the top-left of tile <b>1000</b>. The bottom-right portion of tile <b>1000</b> is a depth signal having depth values corresponding to the image on the bottom-left of tile <b>1000</b>.
0133Turning to <figref idref="DRAWINGS">FIG. 11</figref>, an example of 5 views tiled on a single frame is indicated generally by the reference numeral <b>1100</b>. The top four views are in a normal orientation. The fifth view is also in a normal orientation, but is split into two portions along the bottom of tile <b>1100</b>. A left-portion of the fifth view shows the “top” of the fifth view, and a right-portion of the fifth view shows the “bottom” of the fifth view.
0134Turning to <figref idref="DRAWINGS">FIG. 12</figref>, an exemplary Multi-view Video Coding (MVC) encoder is indicated generally by the reference numeral <b>1200</b>. The encoder <b>1200</b> includes a combiner <b>1205</b> having an output connected in signal communication with an input of a transformer <b>1210</b>. An output of the transformer <b>1210</b> is connected in signal communication with an input of quantizer <b>1215</b>. An output of the quantizer <b>1215</b> is connected in signal communication with an input of an entropy coder <b>1220</b> and an input of an inverse quantizer <b>1225</b>. An output of the inverse quantizer <b>1225</b> is connected in signal communication with an input of an inverse transformer <b>1230</b>. An output of the inverse transformer <b>1230</b> is connected in signal communication with a first non-inverting input of a combiner <b>1235</b>. An output of the combiner <b>1235</b> is connected in signal communication with an input of an intra predictor <b>1245</b> and an input of a deblocking filter <b>1250</b>. An output of the deblocking filter <b>1250</b> is connected in signal communication with an input of a reference picture store <b>1255</b> (for view i). An output of the reference picture store <b>1255</b> is connected in signal communication with a first input of a motion compensator <b>1275</b> and a first input of a motion estimator <b>1280</b>. An output of the motion estimator <b>1280</b> is connected in signal communication with a second input of the motion compensator <b>1275</b>
0135An output of a reference picture store <b>1260</b> (for other views) is connected in signal communication with a first input of a disparity estimator <b>1270</b> and a first input of a disparity compensator <b>1265</b>. An output of the disparity estimator <b>1270</b> is connected in signal communication with a second input of the disparity compensator <b>1265</b>.
0136An output of the entropy decoder <b>1220</b> is available as an output of the encoder <b>1200</b>. A non-inverting input of the combiner <b>1205</b> is available as an input of the encoder <b>1200</b>, and is connected in signal communication with a second input of the disparity estimator <b>1270</b>, and a second input of the motion estimator <b>1280</b>. An output of a switch <b>1285</b> is connected in signal communication with a second non-inverting input of the combiner <b>1235</b> and with an inverting input of the combiner <b>1205</b>. The switch <b>1285</b> includes a first input connected in signal communication with an output of the motion compensator <b>1275</b>, a second input connected in signal communication with an output of the disparity compensator <b>1265</b>, and a third input connected in signal communication with an output of the intra predictor <b>1245</b>.
0137A mode decision module <b>1240</b> has an output connected to the switch <b>1285</b> for controlling which input is selected by the switch <b>1285</b>.
0138Turning to <figref idref="DRAWINGS">FIG. 13</figref>, an exemplary Multi-view Video Coding (MVC) decoder is indicated generally by the reference numeral <b>1300</b>. The decoder <b>1300</b> includes an entropy decoder <b>1305</b> having an output connected in signal communication with an input of an inverse quantizer <b>1310</b>. An output of the inverse quantizer is connected in signal communication with an input of an inverse transformer <b>1315</b>. An output of the inverse transformer <b>1315</b> is connected in signal communication with a first non-inverting input of a combiner <b>1320</b>. An output of the combiner <b>1320</b> is connected in signal communication with an input of a deblocking filter <b>1325</b> and an input of an intra predictor <b>1330</b>. An output of the deblocking filter <b>1325</b> is connected in signal communication with an input of a reference picture store <b>1340</b> (for view i). An output of the reference picture store <b>1340</b> is connected in signal communication with a first input of a motion compensator <b>1335</b>.
0139An output of a reference picture store <b>1345</b> (for other views) is connected in signal communication with a first input of a disparity compensator <b>1350</b>.
0140An input of the entropy coder <b>1305</b> is available as an input to the decoder <b>1300</b>, for receiving a residue bitstream. Moreover, an input of a mode module <b>1360</b> is also available as an input to the decoder <b>1300</b>, for receiving control syntax to control which input is selected by the switch <b>1355</b>. Further, a second input of the motion compensator <b>1335</b> is available as an input of the decoder <b>1300</b>, for receiving motion vectors. Also, a second input of the disparity compensator <b>1350</b> is available as an input to the decoder <b>1300</b>, for receiving disparity vectors.
0141An output of a switch <b>1355</b> is connected in signal communication with a second non-inverting input of the combiner <b>1320</b>. A first input of the switch <b>1355</b> is connected in signal communication with an output of the disparity compensator <b>1350</b>. A second input of the switch <b>1355</b> is connected in signal communication with an output of the motion compensator <b>1335</b>. A third input of the switch <b>1355</b> is connected in signal communication with an output of the intra predictor <b>1330</b>. An output of the mode module <b>1360</b> is connected in signal communication with the switch <b>1355</b> for controlling which input is selected by the switch <b>1355</b>. An output of the deblocking filter <b>1325</b> is available as an output of the decoder <b>1300</b>.
0142Turning to <figref idref="DRAWINGS">FIG. 14</figref>, an exemplary method for processing pictures for a plurality of views in preparation for encoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>1400</b>.
0143The method <b>1400</b> includes a start block <b>1405</b> that passes control to a function block <b>1410</b>. The function block <b>1410</b> arranges every N views, among a total of M views, at a particular time instance as a super-picture in tile format, and passes control to a function block <b>1415</b>. The function block <b>1415</b> sets a syntax element num_coded_views_minus1, and passes control to a function block <b>1420</b>. The function block <b>1420</b> sets a syntax element view_id[i] for all (num_coded_views_minus1+1) views, and passes control to a function block <b>1425</b>. The function block <b>1425</b> sets the inter-view reference dependency information for anchor pictures, and passes control to a function block <b>1430</b>. The function block <b>1430</b> sets the inter-view reference dependency information for non-anchor pictures, and passes control to a function block <b>1435</b>. The function block <b>1435</b> sets a syntax element pseudo_view_present_flag, and passes control to a decision block <b>1440</b>. The decision block <b>1440</b> determines whether or not the current value of the syntax element pseudo_view_present_flag is equal to true. If so, then control is passed to a function block <b>1445</b>. Otherwise, control is passed to an end block <b>1499</b>.
0144The function block <b>1445</b> sets the following syntax elements, and passes control to a function block <b>1450</b>: tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus 1. The function block <b>1450</b> calls a syntax element pseudo_view_info(view_id) for each coded view, and passes control to the end block <b>1499</b>.
0145Turning to <figref idref="DRAWINGS">FIG. 15</figref>, an exemplary method for encoding pictures for a plurality of views using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>1500</b>.
0146The method <b>1500</b> includes a start block <b>1502</b> that has an input parameter pseudo_view_id and passes control to a function block <b>1504</b>. The function block <b>1504</b> sets a syntax element num_sub_views_minus1, and passes control to a function block <b>1506</b>. The function block <b>1506</b> sets a variable i equal to zero, and passes control to a decision block <b>1508</b>. The decision block <b>1508</b> determines whether or not the variable i is less than the number of sub_views. If so, then control is passed to a function block <b>1510</b>. Otherwise, control is passed to a function block <b>1520</b>.
0147The function block <b>1510</b> sets a syntax element sub_view_id[i], and passes control to a function block <b>1512</b>. The function block <b>1512</b> sets a syntax element num_parts_minus1 [sub_view_id[i]], and passes control to a function block <b>1514</b>. The function block <b>1514</b> sets a variable j equal to zero, and passes control to a decision block <b>1516</b>. The decision block <b>1516</b> determines whether or not the variable j is less than the syntax element num_parts_minus1 [sub_view_id[i]]. If so, then control is passed to a function block <b>1518</b>. Otherwise, control is passed to a decision block <b>1522</b>.
0148The function block <b>1518</b> sets the following syntax elements, increments the variable j, and returns control to the decision block <b>1516</b>: loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]][j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]][j]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i]][j].
0149The function block <b>1520</b> encodes the current picture for the current view using multi-view video coding (MVC), and passes control to an end block <b>1599</b>.
0150The decision block <b>1522</b> determines whether or not a syntax element tiling_mode is equal to zero. If so, then control is passed to a function block <b>1524</b>. Otherwise, control is passed to a function block <b>1538</b>.
0151The function block <b>1524</b> sets a syntax element flip_dir[sub_view_id[i]] and a syntax element upsample_view_flag[sub_view_id[i]], and passes control to a decision block <b>1526</b>. The decision block <b>1526</b> determines whether or not the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to one. If so, then control is passed to a function block <b>1528</b>. Otherwise, control is passed to a decision block <b>1530</b>.
0152The function block <b>1528</b> sets a syntax element upsample_filter[sub_view_id[i]], and passes control to the decision block <b>1530</b>. The decision block <b>1530</b> determines whether or not a value of the syntax element upsample_filter[sub_view_id[i]] is equal to three. If so, the control is passed to a function block <b>1532</b>. Otherwise, control is passed to a function block <b>1536</b>.
0153The function block <b>1532</b> sets the following syntax elements, and passes control to a function block <b>1534</b>: vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]]. The function block <b>1534</b> sets the filter coefficients for each YUV component, and passes control to the function block <b>1536</b>.
0154The function block <b>1536</b> increments the variable i, and returns control to the decision block <b>1508</b>.
0155The function block <b>1538</b> sets a syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]], and passes control to a function block <b>1540</b>. The function block <b>1540</b> sets the variable j equal to zero, and passes control to a decision block <b>1542</b>. The decision block <b>1542</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, then control is passed to a function block <b>1544</b>. Otherwise, control is passed to the function block <b>1536</b>.
0156The function block <b>1544</b> sets a syntax element num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i]], and passes control to a function block <b>1546</b>. The function block <b>1546</b> sets the coefficients for all the pixel tiling filters, and passes control to the function block <b>1536</b>.
0157Turning to <figref idref="DRAWINGS">FIG. 16</figref>, an exemplary method for processing pictures for a plurality of views in preparation for decoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>1600</b>.
0158The method <b>1600</b> includes a start block <b>1605</b> that passes control to a function block <b>1615</b>. The function block <b>1615</b> parses a syntax element num_coded_views_minus1, and passes control to a function block <b>1620</b>. The function block <b>1620</b> parses a syntax element view_id[i] for all (num_coded_views_minus1+1) views, and passes control to a function block <b>1625</b>. The function block <b>1625</b> parses the inter-view reference dependency information for anchor pictures, and passes control to a function block <b>1630</b>. The function block <b>1630</b> parses the inter-view reference dependency information for non-anchor pictures, and passes control to a function block <b>1635</b>. The function block <b>1635</b> parses a syntax element pseudo_view_present_flag, and passes control to a decision block <b>1640</b>. The decision block <b>1640</b> determines whether or not the current value of the syntax element pseudo_view_present_flag is equal to true. If so, then control is passed to a function block <b>1645</b>. Otherwise, control is passed to an end block <b>1699</b>.
0159The function block <b>1645</b> parses the following syntax elements, and passes control to a function block <b>1650</b>: tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1. The function block <b>1650</b> calls a syntax element pseudo_view_info(view_id) for each coded view, and passes control to the end block <b>1699</b>.
0160Turning to <figref idref="DRAWINGS">FIG. 17</figref>, an exemplary method for decoding pictures for a plurality of views using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>1700</b>.
0161The method <b>1700</b> includes a start block <b>1702</b> that starts with input parameter pseudo_view_id and passes control to a function block <b>1704</b>. The function block <b>1704</b> parses a syntax element num_sub_views_minus1, and passes control to a function block <b>1706</b>. The function block <b>1706</b> sets a variable i equal to zero, and passes control to a decision block <b>1708</b>. The decision block <b>1708</b> determines whether or not the variable i is less than the number of sub_views. If so, then control is passed to a function block <b>1710</b>. Otherwise, control is passed to a function block <b>1720</b>.
0162The function block <b>1710</b> parses a syntax element sub_view_id[i], and passes control to a function block <b>1712</b>. The function block <b>1712</b> parses a syntax element num_parts_minus1[sub_view_id[i]], and passes control to a function block <b>1714</b>. The function block <b>1714</b> sets a variable j equal to zero, and passes control to a decision block <b>1716</b>. The decision block <b>1716</b> determines whether or not the variable j is less than the syntax element num_parts_minus1 [sub_view_id[i]]. If so, then control is passed to a function block <b>1718</b>. Otherwise, control is passed to a decision block <b>1722</b>.
0163The function block <b>1718</b> sets the following syntax elements, increments the variable j, and returns control to the decision block <b>1716</b>: loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]][j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]][j]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i]][j].
0164The function block <b>1720</b> decodes the current picture for the current view using multi-view video coding (MVC), and passes control to a function block <b>1721</b>. The function block <b>1721</b> separates each view from the picture using the high level syntax, and passes control to an end block <b>1799</b>.
0165The separation of each view from the decoded picture is done using the high level syntax indicated in the bitstream. This high level syntax may indicate the exact location and possible orientation of the views (and possible corresponding depth) present in the picture.
0166The decision block <b>1722</b> determines whether or not a syntax element tiling_mode is equal to zero. If so, then control is passed to a function block <b>1724</b>. Otherwise, control is passed to a function block <b>1738</b>.
0167The function block <b>1724</b> parses a syntax element flip_dir[sub_view_id[i]] and a syntax element upsample_view_flag[sub_view_id[i]], and passes control to a decision block <b>1726</b>. The decision block <b>1726</b> determines whether or not the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to one. If so, then control is passed to a function block <b>1728</b>. Otherwise, control is passed to a decision block <b>1730</b>.
0168The function block <b>1728</b> parses a syntax element upsample_filter[sub_view_id[i]], and passes control to the decision block <b>1730</b>. The decision block <b>1730</b> determines whether or not a value of the syntax element upsample_filter[sub_view_id[i]] is equal to three. If so, the control is passed to a function block <b>1732</b>. Otherwise, control is passed to a function block <b>1736</b>.
0169The function block <b>1732</b> parses the following syntax elements, and passes control to a function block <b>1734</b>: vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]]. The function block <b>1734</b> parses the filter coefficients for each YUV component, and passes control to the function block <b>1736</b>.
0170The function block <b>1736</b> increments the variable i, and returns control to the decision block <b>1708</b>.
0171The function block <b>1738</b> parses a syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]], and passes control to a function block <b>1740</b>. The function block <b>1740</b> sets the variable j equal to zero, and passes control to a decision block <b>1742</b>. The decision block <b>1742</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, then control is passed to a function block <b>1744</b>. Otherwise, control is passed to the function block <b>1736</b>.
0172The function block <b>1744</b> parses a syntax element num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i]], and passes control to a function block <b>1746</b>. The function block <b>1776</b> parses the coefficients for all the pixel tiling filters, and passes control to the function block <b>1736</b>.
0173Turning to <figref idref="DRAWINGS">FIG. 18</figref>, an exemplary method for processing pictures for a plurality of views and depths in preparation for encoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>1800</b>.
0174The method <b>1800</b> includes a start block <b>1805</b> that passes control to a function block <b>1810</b>. The function block <b>1810</b> arranges every N views and depth maps, among a total of M views and depth maps, at a particular time instance as a super-picture in tile format, and passes control to a function block <b>1815</b>. The function block <b>1815</b> sets a syntax element num_coded_views_minus1, and passes control to a function block <b>1820</b>. The function block <b>1820</b> sets a syntax element view_id[i] for all (num_coded_views_minus1+1) depths corresponding to view_id[i], and passes control to a function block <b>1825</b>. The function block <b>1825</b> sets the inter-view reference dependency information for anchor depth pictures, and passes control to a function block <b>1830</b>. The function block <b>1830</b> sets the inter-view reference dependency information for non-anchor depth pictures, and passes control to a function block <b>1835</b>. The function block <b>1835</b> sets a syntax element pseudo_view_present_flag, and passes control to a decision block <b>1840</b>. The decision block <b>1840</b> determines whether or not the current value of the syntax element pseudo_view_present_flag is equal to true. If so, then control is passed to a function block <b>1845</b>. Otherwise, control is passed to an end block <b>1899</b>.
0175The function block <b>1845</b> sets the following syntax elements, and passes control to a function block <b>1850</b>: tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1. The function block <b>1850</b> calls a syntax element pseudo_view_info(view_id) for each coded view, and passes control to the end block <b>1899</b>.
0176Turning to <figref idref="DRAWINGS">FIG. 19</figref>, an exemplary method for encoding pictures for a plurality of views and depths using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>1900</b>.
0177The method <b>1900</b> includes a start block <b>1902</b> that passes control to a function block <b>1904</b>. The function block <b>1904</b> sets a syntax element num_sub_views_minus1, and passes control to a function block <b>1906</b>. The function block <b>1906</b> sets a variable i equal to zero, and passes control to a decision block <b>1908</b>. The decision block <b>1908</b> determines whether or not the variable i is less than the number of sub_views. If so, then control is passed to a function block <b>1910</b>. Otherwise, control is passed to a function block <b>1920</b>.
0178The function block <b>1910</b> sets a syntax element sub_view_id[i], and passes control to a function block <b>1912</b>. The function block <b>1912</b> sets a syntax element num_parts_minus1[sub_view_id[i]], and passes control to a function block <b>1914</b>. The function block <b>1914</b> sets a variable j equal to zero, and passes control to a decision block <b>1916</b>. The decision block <b>1916</b> determines whether or not the variable j is less than the syntax element num_parts_minus1 [sub_view_id[i]]. If so, then control is passed to a function block <b>1918</b>. Otherwise, control is passed to a decision block <b>1922</b>.
0179The function block <b>1918</b> sets the following syntax elements, increments the variable j, and returns control to the decision block <b>1916</b>: loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]][j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]] [j]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i][j].
0180The function block <b>1920</b> encodes the current depth for the current view using multi-view video coding (MVC), and passes control to an end block <b>1999</b>. The depth signal may be encoded similar to the way its corresponding video signal is encoded. For example, the depth signal for a view may be included on a tile that includes only other depth signals, or only video signals, or both depth and video signals. The tile (pseudo-view) is then treated as a single view for MVC, and there are also presumably other tiles that are treated as other views for MVC.
0181The decision block <b>1922</b> determines whether or not a syntax element tiling_mode is equal to zero. If so, then control is passed to a function block <b>1924</b>. Otherwise, control is passed to a function block <b>1938</b>.
0182The function block <b>1924</b> sets a syntax element flip_dir[sub_view_id[i]] and a syntax element upsample_view_flag[sub_view_id[i]], and passes control to a decision block <b>1926</b>. The decision block <b>1926</b> determines whether or not the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to one. If so, then control is passed to a function block <b>1928</b>. Otherwise, control is passed to a decision block <b>1930</b>.
0183The function block <b>1928</b> sets a syntax element upsample_filter[sub_view_id[i]], and passes control to the decision block <b>1930</b>. The decision block <b>1930</b> determines whether or not a value of the syntax element upsample_filter[sub_view_id[i]] is equal to three. If so, the control is passed to a function block <b>1932</b>. Otherwise, control is passed to a function block <b>1936</b>.
0184The function block <b>1932</b> sets the following syntax elements, and passes control to a function block <b>1934</b>: vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]]. The function block <b>1934</b> sets the filter coefficients for each YUV component, and passes control to the function block <b>1936</b>.
0185The function block <b>1936</b> increments the variable i, and returns control to the decision block <b>1908</b>.
0186The function block <b>1938</b> sets a syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]], and passes control to a function block <b>1940</b>. The function block <b>1940</b> sets the variable j equal to zero, and passes control to a decision block <b>1942</b>. The decision block <b>1942</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, then control is passed to a function block <b>1944</b>. Otherwise, control is passed to the function block <b>1936</b>.
0187The function block <b>1944</b> sets a syntax element num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i]], and passes control to a function block <b>1946</b>. The function block <b>1946</b> sets the coefficients for all the pixel tiling filters, and passes control to the function block <b>1936</b>.
0188Turning to <figref idref="DRAWINGS">FIG. 20</figref>, an exemplary method for processing pictures for a plurality of views and depths in preparation for decoding the pictures using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>2000</b>.
0189The method <b>2000</b> includes a start block <b>2005</b> that passes control to a function block <b>2015</b>. The function block <b>2015</b> parses a syntax element num_coded_views_minus1, and passes control to a function block <b>2020</b>. The function block <b>2020</b> parses a syntax element view_id[i] for all (num_coded_views_minus1+1) depths corresponding to view_id[i], and passes control to a function block <b>2025</b>. The function block <b>2025</b> parses the inter-view reference dependency information for anchor depth pictures, and passes control to a function block <b>2030</b>. The function block <b>2030</b> parses the inter-view reference dependency information for non-anchor depth pictures, and passes control to a function block <b>2035</b>. The function block <b>2035</b> parses a syntax element pseudo_view_present_flag, and passes control to a decision block <b>2040</b>. The decision block <b>2040</b> determines whether or not the current value of the syntax element pseudo_view_present_flag is equal to true. If so, then control is passed to a function block <b>2045</b>. Otherwise, control is passed to an end block <b>2099</b>.
0190The function block <b>2045</b> parses the following syntax elements, and passes control to a function block <b>2050</b>: tiling_mode; org_pic_width_in_mbs_minus1; and org_pic_height_in_mbs_minus1. The function block <b>2050</b> calls a syntax element pseudo_view_info(view_id) for each coded view, and passes control to the end block <b>2099</b>.
0191Turning to <figref idref="DRAWINGS">FIG. 21</figref>, an exemplary method for decoding pictures for a plurality of views and depths using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard is indicated generally by the reference numeral <b>2100</b>.
0192The method <b>2100</b> includes a start block <b>2102</b> that starts with input parameter pseudo_view_id, and passes control to a function block <b>2104</b>. The function block <b>2104</b> parses a syntax element num_sub_views_minus1, and passes control to a function block <b>2106</b>. The function block <b>2106</b> sets a variable i equal to zero, and passes control to a decision block <b>2108</b>. The decision block <b>2108</b> determines whether or not the variable i is less than the number of sub_views. If so, then control is passed to a function block <b>2110</b>. Otherwise, control is passed to a function block <b>2120</b>.
0193The function block <b>2110</b> parses a syntax element sub_view_id[i], and passes control to a function block <b>2112</b>. The function block <b>2112</b> parses a syntax element num_parts_minus1[sub_view_id[i]], and passes control to a function block <b>2114</b>. The function block <b>2114</b> sets a variable j equal to zero, and passes control to a decision block <b>2116</b>. The decision block <b>2116</b> determines whether or not the variable j is less than the syntax element num_parts_minus1[sub_view_id[i]]. If so, then control is passed to a function block <b>2118</b>. Otherwise, control is passed to a decision block <b>2122</b>.
0194The function block <b>2118</b> sets the following syntax elements, increments the variable j, and returns control to the decision block <b>2116</b>: loc_left_offset[sub_view_id[i]][j]; loc_top_offset[sub_view_id[i]] [j]; frame_crop_left_offset[sub_view_id[i]][j]; frame_crop_right_offset[sub_view_id[i]][1]; frame_crop_top_offset[sub_view_id[i]][j]; and frame_crop_bottom_offset[sub_view_id[i][j].
0195The function block <b>2120</b> decodes the current picture using multi-view video coding (MVC), and passes control to a function block <b>2121</b>. The function block <b>2121</b> separates each view from the picture using the high level syntax, and passes control to an end block <b>2199</b>. The separation of each view using high level syntax is as previously described.
0196The decision block <b>2122</b> determines whether or not a syntax element tiling_mode is equal to zero. If so, then control is passed to a function block <b>2124</b>. Otherwise, control is passed to a function block <b>2138</b>.
0197The function block <b>2124</b> parses a syntax element flip_dir[sub_view_id[i]] and a syntax element upsample_view_flag[sub_view_id[i]], and passes control to a decision block <b>2126</b>. The decision block <b>2126</b> determines whether or not the current value of the syntax element upsample_view_flag[sub_view_id[i]] is equal to one. If so, then control is passed to a function block <b>2128</b>. Otherwise, control is passed to a decision block <b>2130</b>.
0198The function block <b>2128</b> parses a syntax element upsample_filter[sub_view_id[i]], and passes control to the decision block <b>2130</b>. The decision block <b>2130</b> determines whether or not a value of the syntax element upsample_filter[sub_view_id[i]] is equal to three. If so, the control is passed to a function block <b>2132</b>. Otherwise, control is passed to a function block <b>2136</b>.
0199The function block <b>2132</b> parses the following syntax elements, and passes control to a function block <b>2134</b>: vert_dim[sub_view_id[i]]; hor_dim[sub_view_id[i]]; and quantizer[sub_view_id[i]]. The function block <b>2134</b> parses the filter coefficients for each YUV component, and passes control to the function block <b>2136</b>.
0200The function block <b>2136</b> increments the variable i, and returns control to the decision block <b>2108</b>.
0201The function block <b>2138</b> parses a syntax element pixel_dist_x[sub_view_id[i]] and the syntax element flip_dist_y[sub_view_id[i]], and passes control to a function block <b>2140</b>. The function block <b>2140</b> sets the variable j equal to zero, and passes control to a decision block <b>2142</b>. The decision block <b>2142</b> determines whether or not the current value of the variable j is less than the current value of the syntax element num_parts[sub_view_id[i]]. If so, then control is passed to a function block <b>2144</b>. Otherwise, control is passed to the function block <b>2136</b>.
0202The function block <b>2144</b> parses a syntax element num_pixel_tiling_filter_coeffs_minus1 [sub_view_id[i]], and passes control to a function block <b>2146</b>. The function block <b>2146</b> parses the coefficients for all the pixel tiling filters, and passes control to the function block <b>2136</b>.
0203Turning to <figref idref="DRAWINGS">FIG. 22</figref>, tiling examples at the pixel level are indicated generally by the reference numeral <b>2200</b>. <figref idref="DRAWINGS">FIG. 22</figref> is described further below.
0204An application of multi-view video coding is Free-viewpoint TV (or FTV). This application requires that the user can freely move between two or more views. In order to accomplish this, the “virtual” views in between two views need to be interpolated or synthesized. There are several methods to perform view interpolation. One of the methods uses depth for view interpolation/synthesis.
0205Each view can have an associated depth signal. Thus, the depth can be considered to be another form of video signal. <figref idref="DRAWINGS">FIG. 9</figref> shows an example of a depth signal <b>900</b>. In order to enable applications such as FTV, the depth signal is transmitted along with the video signal. In the proposed framework of tiling, the depth signal can also be added as one of the tiles. <figref idref="DRAWINGS">FIG. 10</figref> shows an example of depth signals added as tiles. The depth signals/tiles are shown on the right side of <figref idref="DRAWINGS">FIG. 10</figref>.
0206Once the depth is encoded as a tile of the whole frame, the high level syntax should indicate which tile is the depth signal so that the renderer can use the depth signal appropriately.
0207In the case when the input sequence (such as that shown in <figref idref="DRAWINGS">FIG. 1</figref>) is encoded using a MPEG-4 AVC Standard encoder (or an encoder corresponding to a different video coding standard and/or recommendation), the proposed high level syntax may be present in, for example, the Sequence Parameter Set (SPS), the Picture Parameter Set (PPS), a slice header, and/or a Supplemental Enhancement Information (SEI) message. An embodiment of the proposed method is shown in TABLE 1 where the syntax is present in a Supplemental Enhancement Information (SEI) message.
0208In the case when the input sequences of the pseudo views (such as that shown in <figref idref="DRAWINGS">FIG. 1</figref>) is encoded using the multi-view video coding (MVC) extension of the MPEG-4 AVC Standard encoder (or an encoder corresponding to multi-view video coding standard with respect to a different video coding standard and/or recommendation), the proposed high level syntax may be present in the SPS, the PPS, slice header, an SEI message, or a specified profile. An embodiment of the proposed method is shown in TABLE 1. TABLE 1 shows syntax elements present in the Sequence Parameter Set (SPS) structure, including syntax elements proposed realized in accordance with the present principles.
0209<tables id="TABLE-US-00001" num="00001"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="1" colwidth="168pt" align="left" /><colspec colname="2" colwidth="14pt" align="center" /><colspec colname="3" colwidth="35pt" align="left" /><thead><row><entry namest="1" nameend="3" rowsep="1">TABLE 1</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row><row><entry>seq_parameter_set_mvc_extension( ) {</entry><entry>C</entry><entry>Descriptor</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> num_views_minus_1</entry><entry /><entry>ue(v)</entry></row><row><entry> for(i = 0; i <= num_views_minus_1; i++)</entry></row><row><entry> view_id[i]</entry><entry /><entry>ue(v)</entry></row><row><entry> for(i = 0; i <= num_views_minus_1; i++) {</entry></row><row><entry> num_anchor_refs_l0[i]</entry><entry /><entry>ue(v)</entry></row><row><entry> for( j = 0; j < num_anchor_refs_l0[i]; j++ )</entry></row><row><entry> anchor_ref_l0[i][j]</entry><entry /><entry>ue(v)</entry></row><row><entry> num_anchor_refs_l1[i]</entry><entry /><entry>ue(v)</entry></row><row><entry> for( j = 0; j < num_anchor_refs_l1[i]; j++ )</entry></row><row><entry> anchor_ref_l1[i][j]</entry><entry /><entry>ue(v)</entry></row><row><entry> }</entry></row><row><entry> for(i = 0; i <= num_views_minus_1; i++) {</entry></row><row><entry> num_non_anchor_refs_l0[i]</entry><entry /><entry>ue(v)</entry></row><row><entry> for( j = 0; j < num_non_anchor_refs_l0[i]; j++ )</entry></row><row><entry> non_anchor_ref_l0[i][j]</entry><entry /><entry>ue(v)</entry></row><row><entry> num_non_anchor_refs_l1[i]</entry><entry /><entry>ue(v)</entry></row><row><entry> for( j = 0; j < num_non_anchor_refs_l1[i]; j++ )</entry></row><row><entry> non_anchor_ref_l1[i][j]</entry><entry /><entry>ue(v)</entry></row><row><entry> }</entry></row><row><entry> pseudo_view_present_flag</entry><entry /><entry>u(1)</entry></row><row><entry> if (pseudo_view_present_flag) {</entry></row><row><entry> tiling_mode</entry></row><row><entry> org_pic_width_in_mbs_minus1</entry></row><row><entry> org_pic_height_in_mbs_minus1</entry></row><row><entry> for( i = 0; i < num_views_minus_1; i++)</entry></row><row><entry> pseudo_view_info(i);</entry></row><row><entry> }</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0210TABLE 2 shows syntax elements for the pseudo_view_info syntax element of TABLE 1.
0211<tables id="TABLE-US-00002" num="00002"><table frame="none" colsep="0" rowsep="0" pgwide="1"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="1" colwidth="224pt" align="left" /><colspec colname="2" colwidth="14pt" align="center" /><colspec colname="3" colwidth="35pt" align="left" /><thead><row><entry namest="1" nameend="3" rowsep="1">TABLE 2</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row><row><entry>pseudo_view_info (pseudo_view_id) {</entry><entry>C</entry><entry>Descriptor</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> num_sub_views_minus_1[pseudo_view_id]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> if (num_sub_views_minus_1 != 0) {</entry></row><row><entry> for ( i = 0; i < num_sub_views_minus_1[pseudo_view_id]; i++) {</entry></row><row><entry> sub_view_id[i]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> num_parts_minus1[sub_view_id[ i ]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> for( j = 0; j <= num_parts_minus1[sub_view_id[ i ]]; j++ ) {</entry></row><row><entry> loc_left_offset[sub_view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> loc_top_offset[sub_view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_left_offset[sub_view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_right_offset[sub_view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_top_offset[sub_view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_bottom_offset[sub_view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> }</entry></row><row><entry> if (tiling_mode == 0) {</entry></row><row><entry> flip_dir[sub_view_id[ i ][ j ]</entry><entry>5</entry><entry>u(2)</entry></row><row><entry> upsample_view_flag[sub_view_id[ i ]]</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> if(upsample_view_flag[sub_view_id[ i ]])</entry></row><row><entry> upsample_filter[sub_view_id[ i ]]</entry><entry>5</entry><entry>u(2)</entry></row><row><entry> if(upsample_fiter[sub_view_id[i]] == 3) {</entry></row><row><entry> vert_dim[sub_view_id[i]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> hor_dim[sub_view_id[i]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> quantizer[sub_view_id[i]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> for (yuv= 0; yuv< 3; yuv++) {</entry></row><row><entry> for (y = 0; y < vert_dim[sub_view_id[i]] − 1; y ++) {</entry></row><row><entry> for (x = 0; x < hor_dim[sub_view_id[i]] − 1; x ++)</entry></row><row><entry> filter_coeffs[sub_view_id[i]] [yuv][y][x]</entry><entry>5</entry><entry>se(v)</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry> } // if(tiling_mode == 0)</entry></row><row><entry> else if (tiling_mode == 1) {</entry></row><row><entry> pixel_dist_x[sub_view_id[ i ] ]</entry></row><row><entry> pixel_dist_y[sub_view_id[ i ] ]</entry></row><row><entry> for( j = 0; j <= num_parts[sub_view_id[ i ]]; j++ ) {</entry></row><row><entry> num_pixel_tiling_filter_coeffs_minus1[sub_view_id[ i ] ][j]</entry></row><row><entry> for (coeff_idx = 0; coeff_idx <=</entry></row><row><entry>num_pixel_tiling_filter_coeffs_minus1[sub_view_id[ i ] ][j]; j++)</entry></row><row><entry> pixel_tiling_filter_coeffs[sub_view_id[i]][j]</entry></row><row><entry> } // for ( j = 0; j <= num_parts[sub_view_id[ i ]]; j++ )</entry></row><row><entry> } // else if (tiling_mode == 1)</entry></row><row><entry> } // for ( i = 0; i < num_sub_views_minus_1; i++)</entry></row><row><entry> } // if (num_sub_views_minus_1 != 0)</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables><br /> Semantics of the syntax elements presented in TABLE 1 and TABLE 2 are described below:
0212pseudo_view_present_flag equal to true indicates that some view is a super view of multiple sub-views.
0213tiling_mode equal to 0 indicates that the sub-views are tiled at the picture level. A value of 1 indicates that the tiling is done at the pixel level.
0214The new SEI message could use a value for the SEI payload type that has not been used in the MPEG-4 AVC Standard or an extension of the MPEG-4 AVC Standard. The new SEI message includes several syntax elements with the following semantics.
0215num_coded_views_minus1 plus 1 indicates the number of coded views supported by the bitstream. The value of num_coded_views_minus1 is in the scope of 0 to 1023, inclusive.
0216org_pic_width_in_mbs_minus1 plus 1 specifies the width of a picture in each view in units of macroblocks.
0217The variable for the picture width in units of macroblocks is derived as follows: <br />PicWidthInMbs=org_pic_width_in_mbs_minus1+1
0218The variable for picture width for the luma component is derived as follows: <br />PicWidthInSamples<i>L</i>=PicWidthInMbs*16
0219The variable for picture width for the chroma components is derived as follows: <br />PicWidthInSamples<i>L</i>=PicWidthInMbs*MbWidth<i>C </i>
0220org_pic_height_in_mbs_minus1 plus 1 specifies the height of a picture in each view in units of macroblocks.
0221The variable for the picture height in units of macroblocks is derived as follows: <br />PicHeightInMbs=org_pic_height_in_mbs_minus1+1
0222The variable for picture height for the luma component is derived as follows: <br />PicHeightInSamples<i>L</i>=PicHeightInMbs*<b>16</b>
0223The variable for picture height for the chroma components is derived as follows: <br />PicHeightInSamples<i>C</i>=PicHeightInMbs*MbHeight<i>C </i>
0224num_sub_views_minus1 plus 1 indicates the number of coded sub-views included in the current view. The value of num_coded_views_minus1 is in the scope of 0 to 1023, inclusive.
0225sub_view_id[i] specifies the sub_view_id of the sub-view with decoding order indicated by i.
0226num_parts[sub_view_id[i]] specifies the number of parts that the picture of sub_view_id[i] is split up into.
0227loc_left_offset[sub_view_id[i]][j] and loc_top_offset[sub_view_id[i]][j] specify the locations in left and top pixels offsets, respectively, where the current part j is located in the final reconstructed picture of the view with sub_view_id equal to sub_view_id[i].
0228view_id[i] specifies the view_id of the view with coding order indicate by i.
0229frame_crop_left_offset[view_id[i]][j], frame_crop_right_offset[view_id[i]][j], frame_crop_top_offset[view_id[i]][j], and frame_crop_bottom_offset[view_id[i]][j] specify the samples of the pictures in the coded video sequence that are part of num_part j and view_id i, in terms of a rectangular region specified in frame coordinates for output.
0230The variables CropUnitX and CropUnitY are derived as follows: <ul id="ul0001" list-style="none"><li id="ul0001-0001" num="0000"><ul id="ul0002" list-style="none"><li id="ul0002-0001" num="0231">If chroma_format_idc is equal to 0, CropUnitX and CropUnitY are derived as follows: <br />CropUnit<i>X=</i>1<br />CropUnit<i>Y=</i>2−frame_mbs_only_flag</li><li id="ul0002-0002" num="0232">Otherwise (chroma_format_idc is equal to 1, 2, or 3), CropUnitX and CropUnitY are derived as follows: <br />CropUnit<i>X</i>=SubWidth<i>C </i><br />CropUnit<i>Y</i>=SubHeight<i>C</i>*(2−frame_mbs_only_flag)</li></ul></li></ul>
0233The frame cropping rectangle includes luma samples with horizontal frame coordinates from the following:
0234CropUnitX*frame_crop_left_offset to PicWidthInSamplesL−(CropUnitX*frame_crop_right_offset+1) and vertical frame coordinates from CropUnitY*frame_crop_top_offset to (16*FrameHeightInMbs)−(CropUnitY*frame_crop_bottom_offset+1), inclusive. The value of frame_crop_left_offset shall be in the range of 0 to (PicWidthlnSamplesL/CropUnitX)−(frame_crop_right_offset+1), inclusive; and the value of frame_crop_top_offset shall be in the range of 0 to (16*FrameHeightInMbs/CropUnitY)−(frame_crop_bottom_offset+1), inclusive.
0235When chroma_format_idc is not equal to 0, the corresponding specified samples of the two chroma arrays are the samples having frame coordinates (x/SubWidthC, y/SubHeightC), where (x, y) are the frame coordinates of the specified luma samples.
0236For decoded fields, the specified samples of the decoded field are the samples that fall within the rectangle specified in frame coordinates.
0237num_parts[view_id[i]] specifies the number of parts that the picture of view_id[i] is split up into.
0238depth_flag[view_id[i]] specifies whether or not the current part is a depth signal. If depth_flag is equal to 0, then the current part is not a depth signal. If depth_flag is equal to 1, then the current part is a depth signal associated with the view identified by view_id[i].
0239flip_dir[sub_view_id[i]][j] specifies the flipping direction for the current part. flip_dir equal to 0 indicates no flipping, flip_dir equal to 1 indicates flipping in a horizontal direction, flip_dir equal to 2 indicates flipping in a vertical direction, and flip_dir equal to 3 indicates flipping in horizontal and vertical directions.
0240flip_dir[view_id[i]][j] specifies the flipping direction for the current part. flip_dir equal to 0 indicates no flipping, flip_dir equal to 1 indicates flipping in a horizontal direction, flip_dir equal to 2 indicates flipping in vertical direction, and flip_dir equal to 3 indicates flipping in horizontal and vertical directions.
0241loc_left_offset[view_id[i]][j], loc_top_offset[view_id[i]][j] specifies the location in pixels offsets, where the current part j is located in the final reconstructed picture of the view with view_id equals to view_id[i]
0242upsample_view_flag[view_id[i]] indicates whether the picture belonging to the view specified by view_id[i] needs to be upsampled. upsample_view_flag[view_id[i]] equal to 0 specifies that the picture with view_id equal to view_id[i] will not be upsampled. upsample_view_flag[view_id[i]] equal to 1 specifies that the picture with view_id equal to view_id[i] will be upsampled.
0243upsample_filter[view_id[i]] indicates the type of filter that is to be used for upsampling. upsample_filter[view_id[i]] equals to 0 indicates that the 6-tap AVC filter should be used, upsample_filter[view_id[i]] equals to 1 indicates that the 4-tap SVC filter should be used, upsample_filter[view_id[i]] 2 indicates that the bilinear filter should be used, upsample_filter[view_id[i]] equals to 3 indicates that custom filter coefficients are transmitted. When upsample_filter[view_id[i]] is not present it is set to 0. In this embodiment, we use 2D customized filter. It can be easily extended to 1D filter, and some other nonlinear filter.
0244vert_dim[view_id[i]] specifies the vertical dimension of the custom 2D filter.
0245hor_dim[view_id[i]] specifies the horizontal dimension of the custom 2D filter.
0246quantizer[view_id[i]] specifies the quantization factor for each filter coefficient.
0247filter_coeffs[view_id[i]] [yuv][y][x] specifies the quantized filter coefficients. yuv signals the component for which the filter coefficients apply. yuv equal to 0 specifies the Y component, yuv equal to 1 specifies the U component, and yuv equal to 2 specifies the V component.
0248pixel_dist_x[sub_view_id[i]] and pixel_dist_y[sub_view_id[i]] respectively specify the distance in the horizontal direction and the vertical direction in the final reconstructed pseudo view between neighboring pixels in the view with sub_view_id equal to sub_view_id[i].
0249num_pixel_tiling_filter_coeffs_minus1[sub_view_id[i][j] plus one indicates the number of the filter coefficients when the tiling mode is set equal to 1.
0250pixel_tiling_filter_coeffs[sub_view_id[i][j] signals the filter coefficients that are required to represent a filter that may be used to filter the tiled picture.
0251Turning to <figref idref="DRAWINGS">FIG. 22</figref>, two examples showing the composing of a pseudo view by tiling pixels from four views are respectively indicated by the reference numerals <b>2210</b> and <b>2220</b>, respectively. The four views are collectively indicated by the reference numeral <b>2250</b>. The syntax values for the first example in <figref idref="DRAWINGS">FIG. 22</figref> are provided in TABLE 3 below.
0252<tables id="TABLE-US-00003" num="00003"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="offset" colwidth="14pt" align="left" /><colspec colname="1" colwidth="140pt" align="left" /><colspec colname="2" colwidth="63pt" align="center" /><thead><row><entry /><entry namest="offset" nameend="2" rowsep="1">TABLE 3</entry></row><row><entry /><entry namest="offset" nameend="2" align="center" rowsep="1" /></row><row><entry /><entry>pseudo_view_info (pseudo_view_id) {</entry><entry>Value</entry></row><row><entry /><entry namest="offset" nameend="2" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /><entry>num_sub_views_minus_1[pseudo_view_id]</entry><entry>3</entry></row><row><entry /><entry>sub_view_id[0]</entry><entry>0</entry></row><row><entry /><entry>num_parts_minus1[0]</entry><entry>0</entry></row><row><entry /><entry>loc_left_offset[0][0]</entry><entry>0</entry></row><row><entry /><entry>loc_top_offset[0][0]</entry><entry>0</entry></row><row><entry /><entry>pixel_dist_x[0][0]</entry><entry>0</entry></row><row><entry /><entry>pixel_dist_y[0][0]</entry><entry>0</entry></row><row><entry /><entry>sub_view_id[1]</entry><entry>0</entry></row><row><entry /><entry>num_parts_minus1[1]</entry><entry>0</entry></row><row><entry /><entry>loc_left_offset[1][0]</entry><entry>1</entry></row><row><entry /><entry>loc_top_offset[1][0]</entry><entry>0</entry></row><row><entry /><entry>pixel_dist_x[1][0]</entry><entry>0</entry></row><row><entry /><entry>pixel_dist_y[1][0]</entry><entry>0</entry></row><row><entry /><entry>sub_view_id[2]</entry><entry>0</entry></row><row><entry /><entry>num_parts_minus1[2]</entry><entry>0</entry></row><row><entry /><entry>loc_left_offset[2][0]</entry><entry>0</entry></row><row><entry /><entry>loc_top_offset[2][0]</entry><entry>1</entry></row><row><entry /><entry>pixel_dist_x[2][0]</entry><entry>0</entry></row><row><entry /><entry>pixel_dist_y[2][0]</entry><entry>0</entry></row><row><entry /><entry>sub_view_id[3]</entry><entry>0</entry></row><row><entry /><entry>num_parts_minus1[3]</entry><entry>0</entry></row><row><entry /><entry>loc_left_offset[3][0]</entry><entry>1</entry></row><row><entry /><entry>loc_top_offset[3][0]</entry><entry>1</entry></row><row><entry /><entry>pixel_dist_x[3][0]</entry><entry>0</entry></row><row><entry /><entry>pixel_dist_y[3][0]</entry><entry>0</entry></row><row><entry /><entry namest="offset" nameend="2" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0253The syntax values for the second example in <figref idref="DRAWINGS">FIG. 22</figref> are all the same except the following two syntax elements: loc_left_offset[3][0] equal to 5 and loc_top_offset[3][0] equal to 3.
0254The offset indicates that the pixels corresponding to a view should begin at a certain offset location. This is shown in <figref idref="DRAWINGS">FIG. 22</figref> (<b>2220</b>). This may be done, for example, when two views produce images in which common objects appear shifted from one view to the other. For example, if first and second cameras (representing first and second views) take pictures of an object, the object may appear to be shifted five pixels to the right in the second view as compared to the first view. This means that pixel(i−5, j) in the first view corresponds to pixel(i, j) in the second view. If the pixels of the two views are simply tiled pixel-by-pixel, then there may not be much correlation between neighboring pixels in the tile, and spatial coding gains may be small. Conversely, by shifting the tiling so that pixel(i−5, j) from view one is placed next to pixel(i, j) from view two, spatial correlation may be increased and spatial coding gain may also be increased. This follows because, for example, the corresponding pixels for the object in the first and second views are being tiled next to each other.
0255Thus, the presence of loc_left_offset and loc_top_offset may benefit the coding efficiency. The offset information may be obtained by external means. For example, the position information of the cameras or the global disparity vectors between the views may be used to determine such offset information.
0256As a result of offsetting, some pixels in the pseudo view are not assigned pixel values from any view. Continuing the example above, when tiling pixel(i−5, j) from view one alongside pixel(i, j) from view two, for values of i=0 . . . 4 there is no pixel(i−5, j) from view one to tile, so those pixels are empty in the tile. For those pixels in the pseudo-view (tile) that are not assigned pixel values from any view, at least one implementation uses an interpolation procedure similar to the sub-pixel interpolation procedure in motion compensation in AVC. That is, the empty tile pixels may be interpolated from neighboring pixels. Such interpolation may result in greater spatial correlation in the tile and greater coding gain for the tile.
0257In video coding, we can choose a different coding type for each picture, such as I, P, and B pictures. For multi-view video coding, in addition, we define anchor and non-anchor pictures. In an embodiment, we propose that the decision of grouping can be made based on picture type. This information of grouping is signaled in high level syntax.
0258Turning to <figref idref="DRAWINGS">FIG. 11</figref>, an example of 5 views tiled on a single frame is indicated generally by the reference numeral <b>1100</b>. In particular, the ballroom sequence is shown with 5 views tiled on a single frame. Additionally, it can be seen that the fifth view is split into two parts so that it can be arranged on a rectangular frame. Here, each view is of QVGA size so the total frame dimension is 640×600. Since 600 is not a multiple of 16 it should be extended to 608.
0259For this example, the possible SEI message could be as shown in TABLE 4.
0260<tables id="TABLE-US-00004" num="00004"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="1" colwidth="189pt" align="left" /><colspec colname="2" colwidth="28pt" align="center" /><thead><row><entry namest="1" nameend="2" rowsep="1">TABLE 4</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row><row><entry>multiview_display_info( payloadSize ) {</entry><entry>Value</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="1" colwidth="189pt" align="left" /><colspec colname="2" colwidth="28pt" align="char" char="." /><tbody valign="top"><row><entry> num_coded_views_minus1</entry><entry>5</entry></row><row><entry> org_pic_width_in_mbs_minus1</entry><entry>40</entry></row><row><entry> org_pic_height_in_mbs_minus1</entry><entry>30</entry></row><row><entry> view_id[ 0 ]</entry><entry>0</entry></row><row><entry> num_parts[view_id[ 0 ]]</entry><entry>1</entry></row><row><entry> depth_flag[view_id[ 0 ]][ 0 ]</entry><entry>0</entry></row><row><entry> flip_dir[view_id[ 0 ]][ 0 ]</entry><entry>0</entry></row><row><entry> loc_left_offset[view_id[ 0 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> loc_top_offset[view_id[ 0 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> frame_crop_left_offset[view_id[ 0 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> frame_crop_right_offset[view_id[ 0 ]] [ 0 ]</entry><entry>320</entry></row><row><entry> frame_crop_top_offset[view_id[ 0 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> frame_crop_bottom_offset[view_id[ 0 ]] [ 0 ]</entry><entry>240</entry></row><row><entry> upsample_view_flag[view_id[ 0 ]]</entry><entry>1</entry></row><row><entry> if(upsample_view_flag[view_id[ 0 ]]) {</entry></row><row><entry> vert_dim[view_id[0]]</entry><entry>6</entry></row><row><entry> hor_dim[view_id[0]]</entry><entry>6</entry></row><row><entry> quantizer[view_id[0]]</entry><entry>32</entry></row><row><entry> for (yuv= 0; yuv< 3; yuv++) {</entry></row><row><entry> for (y = 0; y < vert_dim[view_id[i]] − 1; y ++) {</entry></row><row><entry> for (x = 0; x < hor_dim[view_id[i]] − 1; x ++)</entry></row><row><entry> filter_coeffs[view_id[i]] [yuv][y][x]</entry><entry>XX</entry></row><row><entry> view_id[ 1 ]</entry><entry>1</entry></row><row><entry> num_parts[view_id[ 1 ]]</entry><entry>1</entry></row><row><entry> depth_flag[view_id[ 0 ]][ 0 ]</entry><entry>0</entry></row><row><entry> flip_dir[view_id[ 1 ]][ 0 ]</entry><entry>0</entry></row><row><entry> loc_left_offset[view_id[ 1 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> loc_top_offset[view_id[ 1 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> frame_crop_left_offset[view_id[ 1 ]] [ 0 ]</entry><entry>320</entry></row><row><entry> frame_crop_right_offset[view_id[ 1 ]] [ 0 ]</entry><entry>640</entry></row><row><entry> frame_crop_top_offset[view_id[ 1 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> frame_crop_bottom_offset[view_id[ 1 ]] [ 0 ]</entry><entry>320</entry></row><row><entry> upsample_view_flag[view_id[ 1 ]]</entry><entry>1</entry></row><row><entry> if(upsample_view_flag[view_id[ 1 ]]) {</entry></row><row><entry> vert_dim[view_id[1]]</entry><entry>6</entry></row><row><entry> hor_dim[view_id[1]]</entry><entry>6</entry></row><row><entry> quantizer[view_id[1]]</entry><entry>32</entry></row><row><entry> for (yuv= 0; yuv< 3; yuv++) {</entry></row><row><entry> for (y = 0; y < vert_dim[view_id[i]] − 1; y ++) {</entry></row><row><entry> for (x = 0; x < hor_dim[view_id[i]] − 1; x ++)</entry></row><row><entry> filter_coeffs[view_id[i]] [yuv][y][x]</entry><entry>XX</entry></row><row><entry>......(similarly for view 2,3)</entry></row><row><entry> view_id[ 4 ]</entry><entry>4</entry></row><row><entry> num_parts[view_id[ 4 ]]</entry><entry>2</entry></row><row><entry> depth_flag[view_id[ 0 ]][ 0 ]</entry><entry>0</entry></row><row><entry> flip_dir[view_id[ 4 ]][ 0 ]</entry><entry>0</entry></row><row><entry> loc_left_offset[view_id[ 4 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> loc_top_offset[view_id[ 4 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> frame_crop_left_offset[view_id[ 4 ]] [ 0 ]</entry><entry>0</entry></row><row><entry> frame_crop_right_offset[view_id[ 4 ]] [ 0 ]</entry><entry>320</entry></row><row><entry> frame_crop_top_offset[view_id[ 4 ]] [ 0 ]</entry><entry>480</entry></row><row><entry> frame_crop_bottom_offset[view_id[ 4 ]] [ 0 ]</entry><entry>600</entry></row><row><entry> flip_dir[view_id[ 4 ]][ 1 ]</entry><entry>0</entry></row><row><entry> loc_left_offset[view_id[ 4 ]] [ 1 ]</entry><entry>0</entry></row><row><entry> loc_top_offset[view_id[ 4 ]] [ 1 ]</entry><entry>120</entry></row><row><entry> frame_crop_left_offset[view_id[ 4 ]] [ 1 ]</entry><entry>320</entry></row><row><entry> frame_crop_right_offset[view_id[ 4 ]] [ 1 ]</entry><entry>640</entry></row><row><entry> frame_crop_top_offset[view_id[ 4 ]] [ 1 ]</entry><entry>480</entry></row><row><entry> frame_crop_bottom_offset[view_id[ 4 ]] [ 1 ]</entry><entry>600</entry></row><row><entry> upsample_view_flag[view_id[ 4 ]]</entry><entry>1</entry></row><row><entry> if(upsample_view_flag[view_id[ 4 ]]) {</entry></row><row><entry> vert_dim[view_id[4]]</entry><entry>6</entry></row><row><entry> hor_dim[view_id[4]]</entry><entry>6</entry></row><row><entry> quantizer[view_id[4]]</entry><entry>32</entry></row><row><entry> for (yuv= 0; yuv< 3; yuv++) {</entry></row><row><entry> for (y = 0; y < vert_dim[view_id[i]] − 1; y ++) {</entry></row><row><entry> for (x = 0; x < hor_dim[view_id[i]] − 1; x ++)</entry></row><row><entry> filter_coeffs[view_id[i]] [yuv][y][x]</entry><entry>XX</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0261TABLE 5 shows the general syntax structure for transmitting multi-view information for the example shown in TABLE 4.
0262<tables id="TABLE-US-00005" num="00005"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="1" colwidth="168pt" align="left" /><colspec colname="2" colwidth="14pt" align="center" /><colspec colname="3" colwidth="35pt" align="left" /><thead><row><entry namest="1" nameend="3" rowsep="1">TABLE 5</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row><row><entry>multiview_display_info( payloadSize ) {</entry><entry>C</entry><entry>Descriptor</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> num_coded_views_minus1</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> org_pic_width_in_mbs_minus1</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> org_pic_height_in_mbs_minus1</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> for( i = 0; i <= num_coded_views_minus1; i++ ) {</entry></row><row><entry> view_id[ i ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> num_parts[view_id[ i ]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> for( j = 0; j <= num_parts[i]; j++ ) {</entry></row><row><entry> depth_flag[view_id[ i ]][ j ]</entry></row><row><entry> flip_dir[view_id[ i ]][ j ]</entry><entry>5</entry><entry>u(2)</entry></row><row><entry> loc_left_offset[view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> loc_top_offset[view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_left_offset[view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_right_offset[view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_top_offset[view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> frame_crop_bottom_offset[view_id[ i ]] [ j ]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> }</entry></row><row><entry> upsample_view_flag[view_id[ i ]]</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> if(upsample_view_flag[view_id[ i ]])</entry></row><row><entry> upsample_filter[view_id[ i ]]</entry><entry>5</entry><entry>u(2)</entry></row><row><entry> if(upsample_fiter[view_id[i]] == 3) {</entry></row><row><entry> vert_dim[view_id[i]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> hor_dim[view_id[i]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> quantizer[view_id[i]]</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> for (yuv= 0; yuv< 3; yuv++) {</entry></row><row><entry> for (y = 0; y < vert_dim[view_id[i]] − 1;</entry></row><row><entry> y ++) {</entry></row><row><entry> for (x = 0; x < hor_dim[view_id[i]] − 1;</entry></row><row><entry> x ++)</entry></row><row><entry> filter_coeffs[view_id[i]] [yuv][y][x]</entry><entry>5</entry><entry>se(v)</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0263Referring to <figref idref="DRAWINGS">FIG. 23</figref>, a video processing device <b>2300</b> is shown. The video processing device <b>2300</b> may be, for example, a set top box or other device that receives encoded video and provides, for example, decoded video for display to a user or for storage. Thus, the device <b>2300</b> may provide its output to a television, computer monitor, or a computer or other processing device.
0264The device <b>2300</b> includes a decoder <b>2310</b> that receive a data signal <b>2320</b>. The data signal <b>2320</b> may include, for example, an AVC or an MVC compatible stream. The decoder <b>2310</b> decodes all or part of the received signal <b>2320</b> and provides as output a decoded video signal <b>2330</b> and tiling information <b>2340</b>. The decoded video <b>2330</b> and the tiling information <b>2340</b> are provided to a selector <b>2350</b>. The device <b>2300</b> also includes a user interface <b>2360</b> that receives a user input <b>2370</b>. The user interface <b>2360</b> provides a picture selection signal <b>2380</b>, based on the user input <b>2370</b>, to the selector <b>2350</b>. The picture selection signal <b>2380</b> and the user input <b>2370</b> indicate which of multiple pictures a user desires to have displayed. The selector <b>2350</b> provides the selected picture(s) as an output <b>2390</b>. The selector <b>2350</b> uses the picture selection information <b>2380</b> to select which of the pictures in the decoded video <b>2330</b> to provide as the output <b>2390</b>. The selector <b>2350</b> uses the tiling information <b>2340</b> to locate the selected picture(s) in the decoded video <b>2330</b>.
0265In various implementations, the selector <b>2350</b> includes the user interface <b>2360</b>, and in other implementations no user interface <b>2360</b> is needed because the selector <b>2350</b> receives the user input <b>2370</b> directly without a separate interface function being performed. The selector <b>2350</b> may be implemented in software or as an integrated circuit, for example. The selector <b>2350</b> may also incorporate the decoder <b>2310</b>.
0266More generally, the decoders of various implementations described in this application may provide a decoded output that includes an entire tile. Additionally or alternatively, the decoders may provide a decoded output that includes only one or more selected pictures (images or depth signals, for example) from the tile.
0267As noted above, high level syntax may be used to perform signaling in accordance with one or more embodiments of the present principles. The high level syntax may be used, for example, but is not limited to, signaling any of the following: the number of coded views present in the larger frame; the original width and height of all the views; for each coded view, the view identifier corresponding to the view; for each coded view, the number of parts the frame of a view is split into; for each part of the view, the flipping direction (which can be, for example, no flipping, horizontal flipping only, vertical flipping only or horizontal and vertical flipping); for each part of the view, the left position in pixels or number of macroblocks where the current part belongs in the final frame for the view; for each part of the view, the top position of the part in pixels or number of macroblocks where the current part belongs in the final frame for the view; for each part of the view, the left position, in the current large decoded/encoded frame, of the cropping window in pixels or number of macroblocks; for each part of the view, the right position, in the current large decoded/encoded frame, of the cropping window in pixels or number of macroblocks; for each part of the view, the top position, in the current large decoded/encoded frame, of the cropping window in pixels or number of macroblocks; and, for each part of the view, the bottom position, in the current large decoded/encoded frame, of the cropping window in pixels or number of macroblocks; for each coded view whether the view needs to be upsampled before output (where if the upsampling needs to be performed, a high level syntax may be used to indicate the method for upsampling (including, but not limited to, AVC 6-tap filter, SVC 4-tap filter, bilinear filter or a custom 1D, 2D linear or non-linear filter).
0268It is to be noted that the terms “encoder” and “decoder” connote general structures and are not limited to any particular functions or features. For example, a decoder may receive a modulated carrier that carries an encoded bitstream, and demodulate the encoded bitstream, as well as decode the bitstream.
0269Various methods have been described. Many of these methods are detailed to provide ample disclosure. It is noted, however, that variations are contemplated that may vary one or many of the specific features described for these methods. Further, many of the features that are recited are known in the art and are, accordingly, not described in great detail.
0270Further, reference has been made to the use of high level syntax for sending certain information in several implementations. It is to be understood, however, that other implementations use lower level syntax, or indeed other mechanisms altogether (such as, for example, sending information as part of encoded data) to provide the same information (or variations of that information).
0271Various implementations provide tiling and appropriate signaling to allow multiple views (pictures, more generally) to be tiled into a single picture, encoded as a single picture, and sent as a single picture. The signaling information may allow a post-processor to pull the views/pictures apart. Also, the multiple pictures that are tiled could be views, but at least one of the pictures could be depth information. These implementations may provide one or more advantages. For example, users may want to display multiple views in a tiled manner, and these various implementations provide an efficient way to encode and transmit or store such views by tiling them prior to encoding and transmitting/storing them in a tiled manner.
0272Implementations that tile multiple views in the context of AVC and/or MVC also provide additional advantages. AVC is ostensibly only used for a single view, so no additional view is expected. However, such AVC-based implementations can provide multiple views in an AVC environment because the tiled views can be arranged so that, for example, a decoder knows that that the tiled pictures belong to different views (for example, top left picture in the pseudo-view is view 1, top right picture is view 2, etc).
0273Additionally, MVC already includes multiple views, so multiple views are not expected to be included in a single pseudo-view. Further, MVC has a limit on the number of views that can be supported, and such MVC-based implementations effectively increase the number of views that can be supported by allowing (as in the AVC-based implementations) additional views to be tiled. For example, each pseudo-view may correspond to one of the supported views of MVC, and the decoder may know that each “supported view” actually includes four views in a pre-arranged tiled order. Thus, in such an implementation, the number of possible views is four times the number of “supported views”.
0274In the description that follows, the higher level syntax such as the SEI message is expanded in various implementations to include information about which of the plurality of spatial interleaving modes is attributable to the pictures that are tiled as the single picture. Spatial interleaving can occur in a plurality of modes such as side-by-side, top-bottom, vertical interlacing, horizontal interlacing, and checkerboard, for example. Additionally, the syntax is expanded to include relationship information about the content in the tiled pictures. Relationship information can include a designation of the left and right pictures in a stereo image, or an identification as the depth picture when 2D plus depth is used for the pictures, or even an indication that the pictures form a layered depth video (LDV) signal.
0275As already noted herein above, the syntax is contemplated as being implemented in any high level syntax besides the SEI message, such as syntax at the slice header level, Picture Parameter Set (PPS) level, Sequence Parameter Set (SPS) level, View Parameter Set (VPS) level, and Network Abstraction Layer (NAL) unit header level. Additionally, it is contemplated that a low level syntax can be used to signal the information. It is even contemplated that the information can be signalled out of band in various manners.
0276The SEI message syntax for checkerboard spatial interleaving as described in the aforementioned draft amendment of AVC is defined in Table 6 as follows:
0277<tables id="TABLE-US-00006" num="00006"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="1" colwidth="168pt" align="left" /><colspec colname="2" colwidth="14pt" align="center" /><colspec colname="3" colwidth="35pt" align="left" /><thead><row><entry namest="1" nameend="3" rowsep="1">TABLE 6</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row><row><entry>spatially_interleaved_pictures( payloadSize ) {</entry><entry>C</entry><entry>Descriptor</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> spatially_interleaved_pictures_id</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> spatially_interleaved_pictures_cancel_flag</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> if( !spatially_interleaved_pictures_cancel_flag ) {</entry></row><row><entry> basic_spatial_interleaving_type_id</entry><entry>5</entry><entry>u(8)</entry></row><row><entry> spatially_interleaved_pictures_repetition_period</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> }</entry></row><row><entry> additional_extension_flag</entry><entry>5</entry><entry>u(1)</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0278This SEI message informs the decoder that the output decoded pictures are formed by spatial interleaving of multiple distinct pictures using an indicated spatial interleaving scheme. The information in this SEI message can be used by the decoder to appropriately de-interleave and process the picture data for display or other purposes.
0279The semantics for defined values of the syntax of the spatially interleaved pictures SEI message in Table 6 are as follows:
0280spatially_interleaved_pictures_id contains an identifying number that may be used to identify the usage of the spatially interleaved pictures SEI message.
0281spatially_interleaved_pictures_cancel_flag equal to 1 indicates that the spatially interleaved pictures SEI message cancels the persistence of any previous spatially interleaved pictures SEI message in output order. spatially_interleaved_pictures_cancel_flag equal to 0 indicates that spatially interleaved pictures information follows.
0282basic_spatial_interleaving_type_id indicates the type of spatial interleaving of the multiple pictures included in the single tiled picture <ul id="ul0003" list-style="none"><li id="ul0003-0001" num="0000"><ul id="ul0004" list-style="none"><li id="ul0004-0001" num="0283">basic_spatial_interleaving_type_id equal to 0 indicates that each component plane of the decoded pictures contains a “checkerboard” based interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 26</figref>.</li><li id="ul0004-0002" num="0284">basic_spatial_interleaving_type_id equal to 1 indicates that each component plane of the decoded pictures contains a “checkerboard” based interleaving of a corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 26</figref>, and additionally indicates that the two constituent pictures form the left and right views of a stereo-view scene as illustrated in <figref idref="DRAWINGS">FIG. 26</figref>.</li></ul></li></ul>
0285spatially_interleaved_pictures_repetition_period specifies the persistence of the spatially interleaved pictures SEI message and may specify a picture order count interval within which another spatially interleaved pictures SEI message with the same value of spatially_interleaved_pictures_id or the end of the coded video sequence shall be present in the bitstream. <ul id="ul0005" list-style="none"><li id="ul0005-0001" num="0000"><ul id="ul0006" list-style="none"><li id="ul0006-0001" num="0286">spatially_interleaved_pictures_repetition_period equal to 0 specifies that the spatially interleaved pictures SEI message applies to the current decoded picture only.</li><li id="ul0006-0002" num="0287">spatially_interleaved_pictures_repetition_period equal to 1 specifies that the spatially interleaved pictures SEI message persists in output order until any of the following conditions are true: <ul id="ul0007" list-style="none"><li id="ul0007-0001" num="0288">A new coded video sequence begins.</li><li id="ul0007-0002" num="0289">A picture in an access unit containing a spatially interleaved pictures SEI message with the same value of spatially_interleaved_pictures_id is output having PicOrderCnt( )greater than PicOrderCnt(CurrPic).</li></ul></li><li id="ul0006-0003" num="0290">spatially_interleaved_pictures_repetition_period equal to 0 or equal to 1 indicates that another spatially interleaved pictures SEI message with the same value of spatially_interleaved_pictures_id may or may not be present.</li><li id="ul0006-0004" num="0291">spatially_interleaved_pictures_repetition_period greater than 1 specifies that the spatially interleaved pictures SEI message persists until any of the following conditions are true: <ul id="ul0008" list-style="none"><li id="ul0008-0001" num="0292">A new coded video sequence begins.</li><li id="ul0008-0002" num="0293">A picture in an access unit containing a spatially interleaved pictures SEI message with the same value of spatially_interleaved_pictures_id is output having PicOrderCnt( )greater than PicOrderCnt(CurrPic) and less than or equal to PicOrderCnt(CurrPic)+spatially_interleaved_pictures_repetition_period.</li></ul></li></ul></li></ul>
0294spatially_interleaved_pictures_repetition_period greater than 1 indicates that another spatially interleaved pictures SEI message with the same value of spatially_interleaved_pictures_id shall be present for a picture in an access unit that is output having PicOrderCnt( )greater than PicOrderCnt(CurrPic) and less than or equal to PicOrderCnt(CurrPic)+spatially_interleaved_pictures_repetition_period; unless the bitstream ends or a new coded video sequence begins without output of such a picture.
0295additional_extension_flag equal to 0 indicates that no additional data follows within the spatially interleaved pictures SEI message.
0296Without changing the syntax shown in Table 6, an implementation of this application provides relationship information and spatial interleaving information within the exemplary SEI message. The range of possible values for basic_spatial_interleaving_type_id is modified and expanded in this implementation to indicate a plurality of spatial interleaving methods rather than just the one checkerboard method. Moreover, the parameter basic_spatial_interleaving_type_id is exploited to indicate that a particular type of spatial interleaving is present in the picture and that the constituent interleaved pictures are related to each other. In this implementation, for basic_spatial_interleaving_type_id, the semantics are as follows:
0297Value 2 or 3 means that the single picture contains a “side-by-side” interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 27</figref>. Value 3 means additionally that the two constituent pictures form the left and right views of a stereo-view scene. For side-by-side interleaving, one picture is placed to one side of the other picture so that the composite picture includes two images side-by-side.
0298Value 4 or 5 means that the single picture contains a “top-bottom” interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 28</figref>. Value 5 means additionally that the two constituent pictures form the left and right views of a stereo-view scene. For top-bottom interleaving, one picture is placed above the other picture so that the composite picture appears to have one image over the other.
0299Value 6 or 7 means that the single picture contains a “row-by-row” interleaving or simply row interlacing of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 29</figref>. Value 7 means additionally that the two constituent pictures form the left and right views of a stereo-view scene. For row-by-row interleaving, the consecutive rows of the single picture alternate from one constituent picture to the other. Basically, the single picture is an alternation of horizontal slices of the constituent images.
0300Value 8 or 9 means that the single picture contains a “column-by-column” interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 30</figref>. Value 9 means additionally that the two constituent pictures form the left and right views of a stereo-view scene. For column-by-column interleaving, the consecutive columns of the single picture alternate from one constituent picture to the other. Basically, the single picture is an alternation of vertical slices of the constituent images.
0301In another embodiment of a syntax for use in an exemplary SEI message to convey information about an associated picture, several additional syntax elements have been included in Table 7 to indicate additional information. Such additional syntax elements are included, for example, to indicate orientation of one or more constituent pictures (for example, flip), to indicate separately whether a left-right stereo pair relationship exists for the images, to indicate whether upsampling is needed for either constituent picture, and to indicate a possible degree and direction of upsampling.
0302<tables id="TABLE-US-00007" num="00007"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="1" colwidth="168pt" align="left" /><colspec colname="2" colwidth="14pt" align="center" /><colspec colname="3" colwidth="35pt" align="left" /><thead><row><entry namest="1" nameend="3" rowsep="1">TABLE 7</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row><row><entry>spatially_interleaved_pictures( payloadSize ) {</entry><entry>C</entry><entry>Descriptor</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> spatially_interleaved_pictures_id</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> spatially_interleaved_pictures_cancel_flag</entry><entry>5</entry><entry>U(1)</entry></row><row><entry> if( !spatially_interleaved_pictures_cancel_flag ) {</entry></row><row><entry> basic_spatial_interleaving_type_id</entry><entry>5</entry><entry>U(8)</entry></row><row><entry> stereo_pair_flag</entry><entry>5</entry><entry>U(1)</entry></row><row><entry> upsample_conversion_horizontal_flag</entry><entry>5</entry><entry>U(1)</entry></row><row><entry> upsample_conversion_vertical_flag</entry><entry>5</entry><entry>U(1)</entry></row><row><entry> if (basic_spatial_interleaving_type_id == 1 or</entry></row><row><entry> basic_spatial_interleaving_type_id == 2) {</entry></row><row><entry> flip_flag</entry><entry>5</entry><entry>U(1)</entry></row><row><entry> }</entry></row><row><entry> spatially_interleaved_pictures_repetition_period</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> }</entry></row><row><entry> additional_extension_flag</entry><entry>5</entry><entry>U(1)</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables><br /> Semantics are defined as follows: <br /> basic_spatial_interleaving_type_id indicates the type of spatial interleaving of the pictures. <ul id="ul0009" list-style="none"><li id="ul0009-0001" num="0000"><ul id="ul0010" list-style="none"><li id="ul0010-0001" num="0303">basic_spatial_interleaving_type_id equal to 0 indicates that each component plane of the decoded pictures contains a “checkerboard” based interleaving of corresponding planes of two pictures as in the earlier proposal (see <figref idref="DRAWINGS">FIG. 26</figref>).</li><li id="ul0010-0002" num="0304">basic_spatial_interleaving_type_id equal to 1 indicates that each component plane of the decoded pictures contains a “side-by-side” based interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 27</figref>.</li><li id="ul0010-0003" num="0305">basic_spatial_interleaving_type_id equal to 2 indicates that each component plane of the decoded pictures contains a “top-bottom” based interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 28</figref>.</li><li id="ul0010-0004" num="0306">basic_spatial_interleaving_type_id equal to 3 indicates that each component plane of the decoded pictures contains a “row-by-row” based interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 29</figref>.</li><li id="ul0010-0005" num="0307">basic_spatial_interleaving_type_id equal to 4 indicates that each component plane of the decoded pictures contains a “column-by-column” based interleaving of corresponding planes of two pictures as illustrated in <figref idref="DRAWINGS">FIG. 30</figref>. <br /> stereo_pair_flag indicates whether the two constituent pictures have a relationship as forming the left and right views of a stereo-view scene. Value of 0 indicates that the constituent pictures do not form the left and right views. Value of 1 indicates they are related as forming the left and right views of an image. <br /> upsample_conversion_horizontal_flag indicates whether the two constituent pictures need upsampling in the horizontal direction after they are extracted from the single picture during decoding. Value of 0 indicates that no upsampling is needed. That corresponds to a sampling factor of zero. Value of 1 indicates that upsampling by a sampling factor of two is required. <br /> upsample_conversion_vertical_flag indicates whether the two constituent pictures need upsampling in the vertical direction after they are extracted from the single picture during decoding. Value of 0 indicates that no upsampling is needed. That corresponds to a sampling factor of zero. Value of 1 indicates that upsampling by a sampling factor of two is required. </li></ul></li></ul>
0308As described in more detail below, it is contemplated that many facets of the upsampling operation can be conveyed in the SEI message so that the upsampling operation is handled properly during picture decoding. For example, an additional range of factors for upsampling can be indicated; the type of upsampling filter can also be indicated; the downsampling filter can also be indicated so that the decoder can determine an appropriate, or even optimal, filter for upsampling. It is also contemplated that filter coefficient information including the number and value of the upsampling filter coefficients can be further indicated in the SEI message so that the receiver performs the preferred upsampling operation.
0309The sampling factor indicates the ratio between the original size and the sampled size of a video picture. For example, when the sampling factor is 2, the original picture size is twice a large as the sampled picture size. The picture size is usually a measure of resolution in pixels. So a horizontally downsampled picture requires a corresponding horizontal upsampling by the same factor to restore the resolution of the original video picture. If the original pictures have a width of 1024 pixels, for example, it may be downsampled horizontally by a sampling factor of 2 to become a downsampled picture with a width of 512 pixels. The picture size is usually a measure of resolution in pixels. A similar analysis can be shown for vertical downsampling or upsampling. The sampling factor can be applied to the downsampling or upsampling that relies on a combined horizontal and vertical approach.
0310flip_flag indicates whether the second constituent picture is flipped or not. Value of 0 indicates that no flip is present in that picture. Value of 1 indicates that flipping is performed. Flipping should be understood by persons skilled in the art to involve a rotation of 180° about a central axis in the plane of the picture. The direction of flipping is determined, in this embodiment, by the interleaving type. For example, when side-by-side interleaving (basic_spatial_interleaving_type_id equals 1) is present in the constituent pictures, it is preferable to flip the right hand picture in horizontal direction (that is, about a vertical axis), if indicated by an appropriate value of flip_flag (see <figref idref="DRAWINGS">FIG. 31</figref>). When top-bottom spatial interleaving (basic_spatial_interleaving_type_id equals 2) is present in the constituent pictures, it is preferable to flip the bottom picture in vertical direction (that is, about a central horizontal axis), if indicated by an appropriate value of flip_flag (se <figref idref="DRAWINGS">FIG. 32</figref>). While flipping the second constituent picture has been described herein, it is contemplated that other exemplary embodiments may involve flipping the first picture, such as the top picture in top-bottom interleaving or the left picture in side-by-side interleaving. Of course, in order to handle these additional degrees of freedom for the flip_flag, it may be necessary to increase the range of values or introduce another semantic related thereto.
0311It is also contemplated that additional embodiments may allow flipping a picture in both the horizontal and vertical direction. For example, when four views are tiled together in a single picture with one view (that is, picture) per quadrant as shown in <figref idref="DRAWINGS">FIG. 2</figref>, it is possible to have the top-left quadrant include a first view that is not flipped, while the top-right quadrant includes a second view that is flipped only horizontally, whereas the bottom-left quadrant includes a third view that is flipped vertically only, and whereas the bottom-right quadrant includes a fourth view that is flipped both horizontally and vertically. By tiling or interleaving in this manner, it can be seen that the borders at the interface between the views have a large likelihood of having common scene content on both sides of the borders from the neighboring views in the picture. This type of flipping may provide additional efficiency in compression. Indication of this type of flipping is contemplated within the scope of this disclosure.
0312It is further contemplated that the syntax in Table 7 can be adopted for use in processing 2D plus depth interleaving. Commercial displays may be developed, for example, that accept such a format as input. An exemplary syntax for such an application is set forth in Table 8. Many of the semantics have already been defined above. Only the newly introduced semantics are described below.
0313<tables id="TABLE-US-00008" num="00008"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="1" colwidth="175pt" align="left" /><colspec colname="2" colwidth="42pt" align="left" /><thead><row><entry namest="1" nameend="2" rowsep="1">TABLE 8</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row><row><entry>spatially_interleaved_pictures( payloadSize ) {</entry><entry>Descriptor</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> spatially_interleaved_pictures_id</entry><entry>ue(v)</entry></row><row><entry> spatially_interleaved_pictures_cancel_flag</entry><entry>u(1)</entry></row><row><entry> if( !spatially_interleaved_pictures_cancel_flag ) {</entry></row><row><entry> basic_spatial_interleaving_type_id</entry><entry>u(8)</entry></row><row><entry> semantics_id</entry><entry>u(1)</entry></row><row><entry> upsample_conversion_horizontal_flag</entry><entry>u(1)</entry></row><row><entry> upsample_conversion_vertical_flag</entry><entry>u(1)</entry></row><row><entry> spatially_interleaved_pictures_repetition_period</entry><entry>ue(v)</entry></row><row><entry> if (basic_spatial_interleaving_type_id == 1 or</entry></row><row><entry> basic_spatial_interleaving_type_id == 2) {</entry></row><row><entry> flip_flag</entry><entry>u(1)</entry></row><row><entry> }</entry></row><row><entry> if (semantics_id == 1) {</entry></row><row><entry> camera_parameter_set( )</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry> additional_extension_flag</entry><entry>u(1)</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0314semantics_id replaces stereo_pair_flag in the prior exemplary embodiment and is used to indicate what relationship is intended for the two interleaved pictures, in other words, what the two pictures physically mean. Value of 0 indicates that the two constituent pictures form the left and right views of an image. Value of 1 indicates that the second picture represents the corresponding depth map of the first constituent picture. <figref idref="DRAWINGS">FIG. 34</figref> shows an exemplary single picture in which the constituent video and depth images (for example, 2D+Z) are interleaved side by side. It should be appreciated that other interleaving methods can be indicated by a particular value of the basic_spatial_interleaving_type_id. Value of 2 indicates that the relationship between the two interleaved pictures is not specified. Values greater than 2 can be used for indicating additional relationships.
0315camera_parameter_set indicates the parameters of the scene related to the camera. The parameters may appear in this same SEI message or the parameters can be conveyed in another SEI message, or they may appear within other syntaxes such as SPS, PPS, slice headers, and the like. The camera parameters generally include at least focal length, baseline distance, camera position(s), Znear (the minimum distance between the scene and cameras), and Zfar (the maximum distance between the scene and cameras). The camera parameter set may also include a full parameter set for each camera, including the 3×3 intrinsic matrix, a 3×3 rotation matrix, and a 3D translation vector. The camera parameters are used, for example, in rendering and possibly also in coding.
0316When using an SEI message or other syntax comparable to that shown in Table 8, it may be possible to avoid a need for synchronization at the system level between the video and depth. This results because the video and the associated depth are already tiled into one single frame.
0317Layer depth video (LDV) format for pictures is shown in <figref idref="DRAWINGS">FIG. 33</figref>. In this figure, four pictures are interleaved—side-by-side and top-bottom—to form the composite picture. The upper left quadrant picture represents the center view layer, whereas the upper right quadrant picture represents the center depth layer. Similarly, the lower left quadrant picture represents the occlusion view layer, whereas the bottom right quadrant picture represents the occlusion depth layer. The format shown in <figref idref="DRAWINGS">FIG. 33</figref> may be used as an input format for a commercial auto-stereoscopic display.
0318The presence of LDV indicated using the syntax similar to that shown in the previous embodiments. Primarily, the semantics of semantics_id are extended to introduce an LDV option as follows:
0319<tables id="TABLE-US-00009" num="00009"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="1" colwidth="168pt" align="left" /><colspec colname="2" colwidth="14pt" align="center" /><colspec colname="3" colwidth="35pt" align="left" /><thead><row><entry namest="1" nameend="3" rowsep="1">TABLE 9</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row><row><entry>spatially_interleaved_pictures( payloadSize ) {</entry><entry>C</entry><entry>Descriptor</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> spatially_interleaved_pictures_id</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> spatially_interleaved_pictures_cancel_flag</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> if( !spatially_interleaved_pictures_cancel_flag ) {</entry></row><row><entry> basic_spatial_interleaving_type_id</entry><entry>5</entry><entry>u(8)</entry></row><row><entry> semantics_id</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> upsample_conversion_horizontal_flag</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> upsample_conversion_vertical_flag</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> spatially_interleaved_pictures_repetition_period</entry><entry>5</entry><entry>ue(v)</entry></row><row><entry> if (basic_spatial_interleaving_type_id == 1 or</entry></row><row><entry> basic_spatial_interleaving_type_id == 2) {</entry></row><row><entry> flip_flag</entry><entry>5</entry><entry>u(1)</entry></row><row><entry> }</entry></row><row><entry> if (semantics_id == 1 || semantics_id == 3) {</entry></row><row><entry> camera_parameter_set( )</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry> additional_extension_flag</entry><entry>5</entry><entry>u(1)</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables><br /> semantics_id indicates a relationship between the pictures, that is, what the interleaved pictures physically mean with respect to each other. <br /> Values from 0 to 2 mean that two pictures are related as defined above. Value of 0 indicates that the two pictures form the left and right views. This value can also indicate that the views are stereo views of a scene. Value of 1 indicates that the second picture stands for the corresponding depth map of the first picture, that is, a set of 2D+Z pictures. Value of 2 indicates that the relationship between the two interleaved pictures is not specified. <br /> Value of 3 indicates that four component pictures are interleaved and they correspond to the four component pictures of the LDV representation as shown, for example, in <figref idref="DRAWINGS">FIG. 33</figref>. Values greater than 3 can be employed for additional relationship indications.
0320Additional relationship indications contemplated herein include, but are not limited to: an indication that the multiple views are multiple sets of 2D+Z pictures, also known as MultiView plus Depth (MVD); an indication that the multiple views represent images in two sets of LDV pictures, which is known as Depth Enhanced Stereo (DES).
0321When the semantics_id is equal to 3 and the basic_spatial_interleaving_type_id is equal to 1 (side-by-side) or 2 (top-bottom), the four component pictures are interleaved as shown in <figref idref="DRAWINGS">FIG. 33</figref>.
0322When semantics_id is equal to 3 and basic_spatial_interleaving_type_id equal to 0, the four pictures are interleaved as illustrated in <figref idref="DRAWINGS">FIG. 22</figref>, either as Example 1 (element <b>2210</b>) or Example 2 (element <b>2220</b>) if the offsets of views 1, 2, and 3 relative to view 0 are additionally signaled by a particular indication such as via a new semantic in the syntax. It is contemplated that when semantics_id is equal to 3, interleaving is performed for LDV pictures as shown in Example 1 (element <b>2210</b>) of <figref idref="DRAWINGS">FIG. 22</figref>. It is further contemplated that when semantics_id is equal to 4, interleaving is performed for LDV pictures as shown in Example 2 (element <b>2220</b>) of <figref idref="DRAWINGS">FIG. 22</figref>. The interleaving for LDV with and without offset has already been described above with respect to <figref idref="DRAWINGS">FIG. 22</figref>.
0323As noted above, 3D/stereo video contents can be encoded using 2D video coding standards. A frame-packing method allows spatially or even temporally downsampled views to be packed into one frame for encoding. A decoded frame is unpacked into the constituent multiple views, which are then generally upsampled to the original resolution.
0324Upsampling and, more particularly, the upsampling filter play an important role in the quality of the reconstruction of the constituent views or pictures. Selection of the upsample filter depends typically on the downsample filter used in the encoder. The quality of each upsampled decoded picture can be improved if information about either the downsampling or upsampling filter is conveyed to the decoder either in the bitstream or by other non-standardized means. Although the characteristics of upsample filter in the decoder are not necessarily bound to the characteristics of the downsample filter in the encoder such as solely by an inverse or reciprocal relationship, a matched upsample filter can be expected to provide optimal signal reconstruction. In at least one implementation from experimental practice, it has been determined that an SEI message can be used to signal the upsample parameters to the decoder. In other implementations, it is contemplated that upsample and downsample filter parameters are indicated within an SEI message to the video decoder. Typical H.264/MPEG-4 AVC video encoders and decoders are depicted in <figref idref="DRAWINGS">FIGS. 3 and 4</figref>, respectively, as described above. These devices are suitable for practice of the aspects of various implementations disclosed herein. Additionally, other devices not tailored to or even operative with, H.264/MPEG-4 AVC are suitable for practice of various implementations described in this application.
0325As described above, a frame packing upsample SEI message is defined below within the framework of H.264/MPEG-4 AVC frame packing for signaling various upsample filters. This syntax is not intended to be limited to SEI message alone, since it is contemplated that the message information can also be provided by other means, such as, for example, being used in other high level syntaxes, such as, for example, SPS, PPS, VPS, slice header, and NAL header.
0326In an exemplary embodiment employing upsampling information, the upsample filter is selected as a 2-dimensional symmetrical FIR filter with an odd number of taps defined by k non-null coefficients represented in the following form: <br /><i>c</i><sub>k</sub>,0,<i>c</i><sub>k-j</sub>,0, . . . , <i>c</i><sub>1</sub>,1,<i>c</i><sub>1</sub>,0, . . . , 0,<i>c</i><sub>k-j</sub>,0,<i>c</i><sub>k </sub><br /> The filter coefficients c<sub>i </sub>are normalized such that the sum of all the coefficients c<sub>i </sub>is 0.5 for i in the range {1, . . . , k}.
0327The exemplary horizontal upsample process using the filter parameters shown above is described through a series of steps as presented in more detail below. The original image is depicted as an m×n matrix X. The matrix A is defined as an n×(2k+n) matrix having the following attributes:
0328<maths id="MATH-US-00001" num="00001"><math overflow="scroll"><mrow><mrow><mi>A</mi><mo>=</mo><mrow><mo>[</mo><mtable><mtr><mtd><mi>D</mi></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><msub><mn>0</mn><mrow><mrow><mo>(</mo><mrow><mi>n</mi><mo>-</mo><mi>k</mi></mrow><mo>)</mo></mrow><mo>,</mo><mi>k</mi></mrow></msub></mtd></mtr><mtr><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mrow><mo></mo><msub><mi>I</mi><mrow><mi>n</mi><mo>,</mo><mi>n</mi></mrow></msub><mo></mo></mrow></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd></mtr><mtr><mtd><msub><mn>0</mn><mrow><mrow><mo>(</mo><mrow><mi>n</mi><mo>-</mo><mi>k</mi></mrow><mo>)</mo></mrow><mo>,</mo><mi>k</mi></mrow></msub></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mi>D</mi></mtd></mtr></mtable><mo>]</mo></mrow></mrow><mo>,</mo><mrow><mrow><mi>where</mi><mo></mo><mstyle><mspace width="0.8em" height="0.8ex" /></mstyle><mo></mo><mi>D</mi></mrow><mo>=</mo><mrow><mi>adiag</mi><mo></mo><mrow><mo>(</mo><mover><mrow><mn>1</mn><mo>,</mo><mi>…</mi><mo>,</mo><mn>1</mn></mrow><mover><mi>︷</mi><mi>k</mi></mover></mover><mo>)</mo></mrow></mrow></mrow></mrow></math></maths><img file="US9036714B2_D0001.tif" /><br /> is a k×k anti-diagonal identity matrix. The original image X is converted into a new image matrix X′ using the tensor product operation as follows: <br /><i>X</i>′=(<i>XA</i>)<img file="US9036714B2_D0002.tif" />[1 0].<br /> The matrix H is defined as an n×(4k+n−1) matrix including the k non-null tap coefficients as shown above (that is, c<sub>k</sub>, 0, c<sub>k-1</sub>, 0, . . . c<sub>1</sub>, 1, c<sub>1</sub>, 0, . . . 0, c<sub>k-1</sub>, 0, c<sub>k</sub>) shifted along each successive row and padded with leading and/or trailing zeros as follows:
0329<maths id="MATH-US-00002" num="00002"><math overflow="scroll"><mrow><mi>H</mi><mo>=</mo><mrow><mrow><mo>[</mo><mtable><mtr><mtd><msub><mi>c</mi><mi>k</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><mi>⋯</mi></mtd><mtd><msub><mi>c</mi><mn>1</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><msub><mi>c</mi><mn>1</mn></msub></mtd><mtd><mi>⋯</mi></mtd><mtd><msub><mi>c</mi><mi>k</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mi>⋯</mi></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><msub><mi>c</mi><mi>k</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><mi>⋯</mi></mtd><mtd><msub><mi>c</mi><mn>1</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><msub><mi>c</mi><mn>1</mn></msub></mtd><mtd><mi>⋯</mi></mtd><mtd><msub><mi>c</mi><mi>k</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><mi>⋯</mi></mtd></mtr><mtr><mtd><mi>⋯</mi></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd></mtr><mtr><mtd><mi>⋯</mi></mtd><mtd><mn>0</mn></mtd><mtd><msub><mi>c</mi><mi>k</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><mi>⋯</mi></mtd><mtd><msub><mi>c</mi><mn>1</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><msub><mi>c</mi><mn>1</mn></msub></mtd><mtd><mi>⋯</mi></mtd><mtd><msub><mi>c</mi><mi>k</mi></msub></mtd><mtd><mn>0</mn></mtd></mtr></mtable><mo>]</mo></mrow><mo>.</mo></mrow></mrow></math></maths><img file="US9036714B2_D0003.tif" /><br /> The output image matrix Y after the horizontal upsampling is then represented as: <br /><i>Y</i><sup>T</sup><i>=HX′</i><sup>T</sup>.
0330In a similar manner, an exemplary vertical upsampling process is described as the follows. A new matrix A′ is defined similar to matrix as:
0331<maths id="MATH-US-00003" num="00003"><math overflow="scroll"><mrow><msup><mi>A</mi><mi>′</mi></msup><mo>=</mo><mrow><mrow><mo>[</mo><mtable><mtr><mtd><mi>D</mi></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><msub><mn>0</mn><mrow><mrow><mo>(</mo><mrow><mi>m</mi><mo>-</mo><mi>k</mi></mrow><mo>)</mo></mrow><mo>,</mo><mi>k</mi></mrow></msub></mtd></mtr><mtr><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mrow><mo></mo><msub><mi>I</mi><mrow><mi>m</mi><mo>,</mo><mi>m</mi></mrow></msub><mo></mo></mrow></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd></mtr><mtr><mtd><msub><mn>0</mn><mrow><mrow><mo>(</mo><mrow><mi>m</mi><mo>-</mo><mi>k</mi></mrow><mo>)</mo></mrow><mo>,</mo><mi>k</mi></mrow></msub></mtd><mtd><mstyle><mspace width="0.3em" height="0.3ex" /></mstyle></mtd><mtd><mi>D</mi></mtd></mtr></mtable><mo>]</mo></mrow><mo>.</mo></mrow></mrow></math></maths><img file="US9036714B2_D0004.tif" /><br /> The image Y is then converted to a new image matrix Y′ via the tensor product operation as shown below: <br /><i>Y′=</i>(<i>YA′</i>)<img file="US9036714B2_D0005.tif" />[1 0]<br /> The output image matrix after vertical upsampling is then represented as: <br /><i>F</i><sup>T</sup><i>=HY′</i><sup>T</sup>.<br /> The matrix F is the final image matrix after horizontal and vertical upsampling conversion.
0332Upsampling in the example above has been depicted as both a horizontal and a vertical operation. It is contemplated that the order of the operations can be reversed so that vertical upsampling is performed prior to horizontal upsampling. Moreover, it is contemplated that only one type of upsampling may be performed in certain cases.
0333In order to employ the SEI message for conveying upsampling information to the decoder, it is contemplated that certain semantics must be included within the syntaxes presented above. The syntax shown below in Table 10 includes information useful for upsampling at the video decoder. The syntax below has been abridged to show only the parameters necessary for conveying the upsampling information. It will be appreciated by persons skilled in the art that the syntax shown below can be combined with any one or more of the prior syntaxes to convey a substantial amount of frame packing information relating to relationship indication, orientation indication, upsample indication, and spatial interleaving indication, for example.
0334<tables id="TABLE-US-00010" num="00010"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="1" colwidth="189pt" align="left" /><colspec colname="2" colwidth="28pt" align="left" /><thead><row><entry namest="1" nameend="2" rowsep="1">TABLE 10</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row><row><entry /><entry>De-</entry></row><row><entry>Frame_packing_filter( payloadSize ) {</entry><entry>scriptor</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry> for (c = 0; c<3; c++)</entry><entry /></row><row><entry> {</entry></row><row><entry> number_of_horizontal_filter_parameters[c]</entry><entry>ue(v)</entry></row><row><entry> for (i=0; i< number_of_horizontal_filter_parameters[c];</entry></row><row><entry> i++)</entry></row><row><entry> {</entry></row><row><entry> h[c][i]</entry><entry>s(16)</entry></row><row><entry> }</entry></row><row><entry> number_of_vertical_filter_parameters[c]</entry><entry>ue(v)</entry></row><row><entry> for (1=0; i< number_of_horizontal_filter_parameters[c];</entry></row><row><entry> i++)</entry></row><row><entry> {</entry></row><row><entry> v[c][i]</entry><entry>s(16)</entry></row><row><entry> }</entry></row><row><entry> }</entry></row><row><entry>}</entry></row><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row></tbody></tgroup></table></tables><br /> Semantics of the syntax elements are defined below as follows: <br /> number of horizontal_filter_parameters[c] indicate order of the (4k−1)-tap horizontal filter for color component c, where c is one of three different color components. <br /> h[c][i] specifies the i<sup>th </sup>coefficient of the horizontal filter for color component c in 16-bit precision which is in the range of −2<sup>15 </sup>to 2<sup>15</sup>−1. The norm of the filter is 2<sup>16</sup>. <br /> number_of vertical_filter_parameters indicate order of the (4k−1)-tap vertical filter for color component c. <br /> v[c][i] specifies the i<sup>th </sup>coefficient of the vertical filter for color component c in 16-bit precision which is in the range of −2<sup>15 </sup>to 2<sup>15</sup>−1. The norm of the filter is 2<sup>16</sup>.
0335For this syntax, it is assumed that a known filter type is present in the decoder. It is contemplated, however, that various different types of upsample filters may be used in addition to, or in lieu of, the above-described 2-dimensional symmetrical FIR filter. Such other upsample filters include, but are not limited to, any interpolation filter designs such as a bilinear filter, a cubic filter, a spline filter, a Wiener filter, or a least squares filter. When other filter types are available for use, it is contemplated that the syntax above should be expanded to include one or more semantics to indicate information about the filter type for upsampling.
0336It will be appreciated by persons skilled in the art that the upsample filter parameters may be derived at the encoder for inclusion in the SEI message based on various inputs such as downsampling filter information or the like. In addition, it will also be appreciated that such information may simply be provided to the encoder for inclusion in the SEI message.
0337It is further contemplated that the syntax shown in Table 10 could be modified to transmit the downsample filter parameters instead of the upsample filter parameters. In this embodiment, indication of the downsample filter parameters to the decoder affords the decoding application the ability to determine the optimal upsample filter parameters based on the downsample filter parameters. Similar syntax and semantics can be used to signal the downsample filter. Semantics may be included to indicate that the parameters are representative of either a downsample filter or an upsample filter.
0338If an optimal upsample filter is to be determined by the decoder, the decoder may perform such a determination in any manner known in the art. Additionally, a decoder may determine an upsample filter that is not optimal, or that is not necessarily optimal. The determination at the decoder may be based on processing limitations or display parameter considerations or the like. As a further example, a decoder could initially select multiple different types of upsampling filters and then ultimately select the one filter that produces a result determined by some decoder criteria to be optimal for the decoder.
0339A flowchart showing an exemplary encoding technique related to the indication of upsample filter information is shown in <figref idref="DRAWINGS">FIG. 35</figref> and described briefly below. Initially, the encoder configuration is determined and the high level syntax is created. Pictures (for example, views) are downsampled by a downsample filter for packing into frames. The downsampled pictures are spatially interleaved into the frames. Upsample filter parameters are determined by the encoder or are supplied directly to the encoder. In one embodiment, the upsample filter parameters such as type, size, and coefficient values are derived from the downsample filter. For each color component (for example, YUV, RGB), the horizontal and vertical upsample filter parameters are determined in number and value. This is shown in <figref idref="DRAWINGS">FIG. 35</figref> with three loops: a first loop for the components, a second loop (nested in the first loop) for the horizontal filter parameters for a component, and a third loop (nested in the first loop) for the vertical filter parameters for a component. The upsample filter parameters are then written into an SEI message. The SEI message is then sent to the decoder either separately (out-of-band) from the video image bitstream or together (in-band) with the video image bitstream. The encoder combines the message with the sequence of packed frames, when needed.
0340In contrast to the encoding method depicted in <figref idref="DRAWINGS">FIG. 35</figref>, the alternative embodiment shown in <figref idref="DRAWINGS">FIG. 37</figref> shows that the downsample filter parameters are included in the SEI message instead of the upsample filter parameters. In this embodiment, upsample filter parameters are not derived by the encoder and are not conveyed to the decoder with the packed frames.
0341A flowchart showing an exemplary decoding technique related to the indication of upsample filter information is shown in <figref idref="DRAWINGS">FIG. 36</figref> and described briefly below. Initially, the SEI message and other related messages are received by the decoder and the syntaxes are read by the decoder. The SEI message is parsed to determine the information included therein including sampling information, interleaving information, and the like. In this example, it is determined that upsampling filter information for each color component is included in the SEI message. Vertical and horizontal upsampling filter information is extracted to obtain the number and value of each horizontal and vertical upsampling filter coefficient. This extraction is shown in <figref idref="DRAWINGS">FIG. 36</figref> with three loops: a first loop for the components, a second loop (nested in the first loop) for the horizontal filter parameters for a component, and a third loop (nested in the first loop) for the vertical filter parameters for a component. The SEI message is then stored and the sequence of packed frames is decoded to obtain the pictures packed therein. The pictures are then upsampled using the recovered upsampling filter parameters to restore the pictures to their original full resolution.
0342In contrast to the decoding method depicted in <figref idref="DRAWINGS">FIG. 36</figref>, the alternative embodiment shown in <figref idref="DRAWINGS">FIG. 38</figref> shows that the downsample filter parameters are included in the SEI message instead of the upsample filter parameters. In this embodiment, upsample filter parameters are not derived by the encoder and are not conveyed to the decoder with the packed frames. Hence, the decoder employs the received downsample filter parameters to derive upsample filter parameters to be used for restoring the full original resolution to the pictures extracted from interleaving in the packed video frames.
0343<figref idref="DRAWINGS">FIG. 39</figref> shows an exemplary video transmission system <b>2500</b>, to which the present principles may be applied, in accordance with an implementation of the present principles.
0344The video transmission system <b>2500</b> may be, for example, a head-end or transmission system for transmitting a signal using any of a variety of media, such as, for example, satellite, cable, telephone-line, or terrestrial broadcast. The transmission may be provided over the Internet or some other network.
0345The video transmission system <b>2500</b> is capable of generating and delivering compressed video with depth. This is achieved by generating an encoded signal(s) including depth information or information capable of being used to synthesize the depth information at a receiver end that may, for example, have a decoder.
0346The video transmission system <b>2500</b> includes an encoder <b>2510</b> and a transmitter <b>2520</b> capable of transmitting the encoded signal. The encoder <b>2510</b> receives video information and generates an encoded signal(s) with depth. The encoder <b>2510</b> may include sub-modules, including for example an assembly unit for receiving and assembling various pieces of information into a structured format for storage or transmission. The various pieces of information may include, for example, coded or uncoded video, coded or uncoded depth information, and coded or uncoded elements such as, for example, motion vectors, coding mode indicators, and syntax elements.
0347The transmitter <b>2520</b> may be, for example, adapted to transmit a program signal having one or more bitstreams representing encoded pictures and/or information related thereto. Typical transmitters perform functions such as, for example, one or more of providing error-correction coding, interleaving the data in the signal, randomizing the energy in the signal, and/or modulating the signal onto one or more carriers. The transmitter may include, or interface with, an antenna (not shown). Accordingly, implementations of the transmitter <b>2520</b> may include, or be limited to, a modulator.
0348<figref idref="DRAWINGS">FIG. 40</figref> shows an exemplary video receiving system <b>2600</b> to which the present principles may be applied, in accordance with an embodiment of the present principles. The video receiving system <b>2600</b> may be configured to receive signals over a variety of media, such as, for example, satellite, cable, telephone-line, or terrestrial broadcast. The signals may be received over the Internet or some other network.
0349The video receiving system <b>2600</b> may be, for example, a cell-phone, a computer, a set-top box, a television, or other device that receives encoded video and provides, for example, decoded video for display to a user or for storage. Thus, the video receiving system <b>2600</b> may provide its output to, for example, a screen of a television, a computer monitor, a computer (for storage, processing, or display), or some other storage, processing, or display device.
0350The video receiving system <b>2600</b> is capable of receiving and processing video content including video information. The video receiving system <b>2600</b> includes a receiver <b>2610</b> capable of receiving an encoded signal, such as for example the signals described in the implementations of this application, and a decoder <b>2620</b> capable of decoding the received signal.
0351The receiver <b>2610</b> may be, for example, adapted to receive a program signal having a plurality of bitstreams representing encoded pictures. Typical receivers perform functions such as, for example, one or more of receiving a modulated and encoded data signal, demodulating the data signal from one or more carriers, de-randomizing the energy in the signal, de-interleaving the data in the signal, and/or error-correction decoding the signal. The receiver <b>2610</b> may include, or interface with, an antenna (not shown). Implementations of the receiver <b>2610</b> may include, or be limited to, a demodulator. The decoder <b>2620</b> outputs video signals including video information and depth information.
0352<figref idref="DRAWINGS">FIG. 41</figref> shows an exemplary video processing device <b>2700</b> to which the present principles may be applied, in accordance with an embodiment of the present principles. The video processing device <b>2700</b> may be, for example, a set top box or other device that receives encoded video and provides, for example, decoded video for display to a user or for storage. Thus, the video processing device <b>2700</b> may provide its output to a television, computer monitor, or a computer or other processing device.
0353The video processing device <b>2700</b> includes a front-end (FE) device <b>2705</b> and a decoder <b>2710</b>. The front-end device <b>2705</b> may be, for example, a receiver adapted to receive a program signal having a plurality of bitstreams representing encoded pictures, and to select one or more bitstreams for decoding from the plurality of bitstreams. Typical receivers perform functions such as, for example, one or more of receiving a modulated and encoded data signal, demodulating the data signal, decoding one or more encodings (for example, channel coding and/or source coding) of the data signal, and/or error-correcting the data signal. The front-end device <b>2705</b> may receive the program signal from, for example, an antenna (not shown). The front-end device <b>2705</b> provides a received data signal to the decoder <b>2710</b>.
0354The decoder <b>2710</b> receives a data signal <b>2720</b>. The data signal <b>2720</b> may include, for example, one or more Advanced Video Coding (AVC), Scalable Video Coding (SVC), or Multi-view Video Coding (MVC) compatible streams.
0355The decoder <b>2710</b> decodes all or part of the received signal <b>2720</b> and provides as output a decoded video signal <b>2730</b>. The decoded video <b>2730</b> is provided to a selector <b>2750</b>. The device <b>2700</b> also includes a user interface <b>2760</b> that receives a user input <b>2770</b>. The user interface <b>2760</b> provides a picture selection signal <b>2780</b>, based on the user input <b>2770</b>, to the selector <b>2750</b>. The picture selection signal <b>2780</b> and the user input <b>2770</b> indicate which of multiple pictures, sequences, scalable versions, views, or other selections of the available decoded data a user desires to have displayed. The selector <b>2750</b> provides the selected picture(s) as an output <b>2790</b>. The selector <b>2750</b> uses the picture selection information <b>2780</b> to select which of the pictures in the decoded video <b>2730</b> to provide as the output <b>2790</b>.
0356In various implementations, the selector <b>2750</b> includes the user interface <b>2760</b>, and in other implementations no user interface <b>2760</b> is needed because the selector <b>2750</b> receives the user input <b>2770</b> directly without a separate interface function being performed. The selector <b>2750</b> may be implemented in software or as an integrated circuit, for example. In one implementation, the selector <b>2750</b> is incorporated with the decoder <b>2710</b>, and in another implementation, the decoder <b>2710</b>, the selector <b>2750</b>, and the user interface <b>2760</b> are all integrated.
0357In one application, front-end <b>2705</b> receives a broadcast of various television shows and selects one for processing. The selection of one show is based on user input of a desired channel to watch. Although the user input to front-end device <b>2705</b> is not shown in <figref idref="DRAWINGS">FIG. 41</figref>, front-end device <b>2705</b> receives the user input <b>2770</b>. The front-end <b>2705</b> receives the broadcast and processes the desired show by demodulating the relevant part of the broadcast spectrum, and decoding any outer encoding of the demodulated show. The front-end <b>2705</b> provides the decoded show to the decoder <b>2710</b>. The decoder <b>2710</b> is an integrated unit that includes devices <b>2760</b> and <b>2750</b>. The decoder <b>2710</b> thus receives the user input, which is a user-supplied indication of a desired view to watch in the show. The decoder <b>2710</b> decodes the selected view, as well as any required reference pictures from other views, and provides the decoded view <b>2790</b> for display on a television (not shown).
0358Continuing the above application, the user may desire to switch the view that is displayed and may then provide a new input to the decoder <b>2710</b>. After receiving a “view change” from the user, the decoder <b>2710</b> decodes both the old view and the new view, as well as any views that are in between the old view and the new view. That is, the decoder <b>2710</b> decodes any views that are taken from cameras that are physically located in between the camera taking the old view and the camera taking the new view. The front-end device <b>2705</b> also receives the information identifying the old view, the new view, and the views in between. Such information may be provided, for example, by a controller (not shown in <figref idref="DRAWINGS">FIG. 41</figref>) having information about the locations of the views, or the decoder <b>2710</b>. Other implementations may use a front-end device that has a controller integrated with the front-end device.
0359The decoder <b>2710</b> provides all of these decoded views as output <b>2790</b>. A post-processor (not shown in <figref idref="DRAWINGS">FIG. 41</figref>) interpolates between the views to provide a smooth transition from the old view to the new view, and displays this transition to the user. After transitioning to the new view, the post-processor informs (through one or more communication links not shown) the decoder <b>2710</b> and the front-end device <b>2705</b> that only the new view is needed. Thereafter, the decoder <b>2710</b> only provides as output <b>2790</b> the new view.
0360The system <b>2700</b> may be used to receive multiple views of a sequence of images, and to present a single view for display, and to switch between the various views in a smooth manner. The smooth manner may involve interpolating between views to move to another view. Additionally, the system <b>2700</b> may allow a user to rotate an object or scene, or otherwise to see a three-dimensional representation of an object or a scene. The rotation of the object, for example, may correspond to moving from view to view, and interpolating between the views to obtain a smooth transition between the views or simply to obtain a three-dimensional representation. That is, the user may “select” an interpolated view as the “view” that is to be displayed.
0361As presented herein, the implementations and features described in this application may be used in the context of coding video, coding depth, and/or coding other types of data. Additionally, these implementations and features may be used in the context of, or adapted for use in the context of, the H.264/MPEG-4 AVC (AVC) Standard, the AVC standard with the MVC extension, the AVC standard with the SVC extension, a 3DV standard, and/or with another standard (existing or future), or in a context that does not involve a standard. Accordingly, it is to be understood that the specific implementations described in this application that operate in accordance with AVC are not intended to be restricted to AVC, and may be adapted for use outside of AVC.
0362Also as noted above, implementations may signal or indicate information using a variety of techniques including, but not limited to, SEI messages, slice headers, other high level syntax, non-high-level syntax, out-of-band information, datastream data, and implicit signaling. Although implementations described herein may be described in a particular context, such descriptions should in no way be taken as limiting the features and concepts to such implementations or contexts.
0363Various implementations involve decoding. “Decoding”, as used in this application, may encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. Such processes may include processes typically performed by a decoder such as, for example, entropy decoding, inverse transformation, inverse quantization, differential decoding. Such processes may also, or alternatively, include processes performed by a decoder of various implementations described in this application, such as, for example, extracting a picture from a tiled (packed) picture, determining an upsample filter to use and then upsampling a picture, and flipping a picture back to its intended orientation.
0364Various implementations described in this application combine multiple pictures into a single picture. The process of combining pictures may involve, for example, selecting pictures based on for example a relationship between pictures, spatial interleaving, sampling, and changing the orientation of a picture. Accordingly, information describing the combining process may include, for example, relationship information, spatial interleaving information, sampling information, and orientation information.
0365Spatial interleaving is referred to in describing various implementations. “Interleaving”, as used in this application, is also referred to as “tiling” and “packing”, for example. Interleaving includes a variety of types of interleaving, including, for example, side-by-side juxtaposition of pictures, top-to-bottom juxtaposition of pictures, side-by-side and top-to-bottom combined (for example, to combine 4 pictures), row-by-row alternating, column-by-column alternating, and various pixel-level (also referred to as pixel-wise or pixel-based) schemes.
0366Also, as used herein, the words “picture” and “image” are used interchangeably and refer, for example, to all or part of a still image or all or part of a picture from a video sequence. As is known, a picture may be a frame or a field. Additionally, as used herein, a picture may also be a subset of a frame such as, for example, a top half of a frame, a single macroblock, alternating columns, alternating rows, or periodic pixels. As another example, a depth picture may be, for example, a complete depth map or a partial depth map that only includes depth information for, for example, a single macroblock of a corresponding video frame.
0367Reference in the specification to “one embodiment” or “an embodiment” or “one implementation” or “an implementation” of the present principles, as well as other variations thereof, mean that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment of the present principles. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout the specification are not necessarily all referring to the same embodiment.
0368It is to be appreciated that the use of any of the following “/”, “and/or”, and “at least one of”, for example, in the cases of “A/B”, “A and/or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and/or C” and “at least one of A, B, and C” and “at least one of A, B, or C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as readily apparent by one of ordinary skill in this and related arts, for as many items listed.
0369Additionally, this application or its claims may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
0370Similarly, “accessing” is intended to be a broad term. Accessing a piece of information may include any operation that, for example, uses, stores, sends, transmits, receives, retrieves, modifies, or provides the information.
0371It should be understood that the elements shown in the figures may be implemented in various forms of hardware, software, or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more appropriately programmed general-purpose devices, which may include a processor, memory and input/output interfaces. Moreover, the implementations described herein may be implemented as, for example, a method or process, an apparatus, or a software program. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented as mentioned above. The methods may be implemented in, for example, an apparatus such as, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processing devices also include communication devices, such as, for example, computers, cell phones, portable/personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.
0372Implementations of the various processes and features described herein may be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding and decoding. Examples of equipment include video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, cell phones, PDAs, other communication devices, personal recording devices (for example, PVRs, computers running recording software, VHS recording devices), camcorders, streaming of data over the Internet or other communication links, and video-on-demand. As should be clear, the equipment may be mobile and even installed in a mobile vehicle.
0373Additionally, the methods may be implemented by instructions being performed by a processor, and such instructions may be stored on a processor-readable medium such as, for example, an integrated circuit, a software carrier or other storage device such as, for example, a hard disk, a compact diskette, a random access memory (“RAM”), or a read-only memory (“ROM”). The instructions may form an application program tangibly embodied on a processor-readable medium. As should be clear, a processor may include a processor-readable medium having, for example, instructions for carrying out a process. Such application programs may be uploaded to, and executed by, a machine comprising any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware such as one or more central processing units (“CPU”), a random access memory (“RAM”), and input/output (“I/O”) interfaces. The computer platform may also include an operating system and microinstruction code. The various processes and functions described herein may be either part of the microinstruction code or part of the application program, or any combination thereof, which may be executed by a CPU. In addition, various other peripheral units may be connected to the computer platform such as an additional data storage unit and a printing unit.
0374As should be evident to one of skill in the art, implementations may also produce a signal formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream, producing syntax, and modulating a carrier with the encoded data stream and the syntax. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known.
0375It is to be further understood that, because some of the constituent system components and methods depicted in the accompanying drawings are preferably implemented in software, the actual connections between the system components or the process function blocks may differ depending upon the manner in which the present principles are programmed. Given the teachings herein, one of ordinary skill in the pertinent art will be able to contemplate these and similar implementations or configurations of the present principles.
0376A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, supplemented, modified, or removed to produce other implementations. Additionally, one of ordinary skill will understand that other structures and processes may be substituted for those disclosed and the resulting implementations will perform at least substantially the same function(s), in at least substantially the same way(s), to achieve at least substantially the same result(s) as the implementations disclosed. In particular, although illustrative embodiments have been described herein with reference to the accompanying drawings, it is to be understood that the present principles is not limited to those precise embodiments, and that various changes and modifications may be effected therein by one of ordinary skill in the pertinent art without departing from the scope or spirit of the present principles. Accordingly, these and other implementations are contemplated by this application and are within the scope of the following claims.
Contents6
51 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2019068947A1 | Cited by | United States of America | Search report |
| US10142611B2 | Cited by | United States of America | Search report |
| US10419021B2 | Cited by | United States of America | Applicant |
| US2017078567A1 | Cited by | United States of America | Pre-grant |
| US2019068947A1 | Cited by | United States of America | Search report |
| US11044454B2 | Cited by | United States of America | Search report |
| US10200603B2 | Cited by | United States of America | Search report |
| WO0225420A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| KR100535147B1 | Cites | Republic of Korea | Applicant |
| EP1501318A1 | Cites | European Patent Office (EPO) | Applicant |
| EP1581003A1 | Cites | European Patent Office (EPO) | Applicant |
| EP1667448A1 | Cites | European Patent Office (EPO) | Applicant |
| EP1729521A2 | Cites | European Patent Office (EPO) | Applicant |
| DE19619598A1 | Cites | Germany | Applicant |
| US2004028288A1 | Cites | United States of America | Applicant |
| JP2004048293A | Cites | Japan | Applicant |
| KR20050055163A | Cites | Republic of Korea | Applicant |
| US2005117637A1 | Cites | United States of America | Applicant |
| US2005134731A1 | Cites | United States of America | Search report |
| US2005243920A1 | Cites | United States of America | Applicant |
| WO2006001653A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2006041261A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| RU2006101400A | Cites | Russian Federation | Applicant |
| WO2006137006A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2006176318A1 | Cites | United States of America | Search report |
| US2006222254A1 | Cites | United States of America | Search report |
| US2006262856A1 | Cites | United States of America | Applicant |
| US2007030356A1 | Cites | United States of America | Applicant |
| US2007041633A1 | Cites | United States of America | Applicant |
| WO2007046957A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2007047736A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2007081926A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| RU2007103160A | Cites | Russian Federation | Applicant |
| US2007121722A1 | Cites | United States of America | Search report |
| WO2007126508A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2007153838A1 | Cites | United States of America | Applicant |
| JP2007159111A | Cites | Japan | Applicant |
| US2007177813A1 | Cites | United States of America | Applicant |
| US2007205367A1 | Cites | United States of America | Search report |
| US2007211796A1 | Cites | United States of America | Applicant |
| US2007229653A1 | Cites | United States of America | Applicant |
| US2007269136A1 | Cites | United States of America | Applicant |
| WO2008024345A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| JP2008034892A | Cites | Japan | Applicant |
| WO2008127676A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2008140190A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2008150111A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| WO2008156318A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2008199091A1 | Cites | United States of America | Applicant |
| US2008284763A1 | Cites | United States of America | Applicant |
| US2008303895A1 | Cites | United States of America | Applicant |
| US2009002481A1 | Cites | United States of America | Applicant |
| KR20090102116A | Cites | Republic of Korea | Applicant |
| WO2009040701A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2009092311A1 | Cites | United States of America | Applicant |
| JP2009182953A | Cites | Japan | Applicant |
| US2009219282A1 | Cites | United States of America | Applicant |
| WO2010011557A2 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US2010026712A1 | Cites | United States of America | Applicant |
| EP2096870A2 | Cites | European Patent Office (EPO) | Applicant |
| EP2197217A1 | Cites | European Patent Office (EPO) | Applicant |
| US5193000A | Cites | United States of America | Applicant |
| US5915091A | Cites | United States of America | Applicant |
| US6055012A | Cites | United States of America | Search report |
| US6157396A | Cites | United States of America | Applicant |
| US6173087B1 | Cites | United States of America | Applicant |
| US6223183B1 | Cites | United States of America | Applicant |
| US6390980B1 | Cites | United States of America | Search report |
| US7254264B2 | Cites | United States of America | Applicant |
| US7254265B2 | Cites | United States of America | Applicant |
| US7321374B2 | Cites | United States of America | Applicant |
| US7391811B2 | Cites | United States of America | Applicant |
| US7489342B2 | Cites | United States of America | Applicant |
| US7552227B2 | Cites | United States of America | Applicant |
| US8139142B2 | Cites | United States of America | Applicant |
| US8259162B2 | Cites | United States of America | Applicant |
| US8885721B2 | Cites | United States of America | Applicant |
| WO9802844A1 | Cites | World Intellectual Property Organization (WIPO) | Applicant |
| US20040028288A1 | Cites | United States of America | Applicant |
| US20050117637A1 | Cites | United States of America | Applicant |
| US20050134731A1 | Cites | United States of America | Search report |
| US20050243920A1 | Cites | United States of America | Applicant |
| US20060176318A1 | Cites | United States of America | Search report |
| US20060222254A1 | Cites | United States of America | Search report |
| US20060262856A1 | Cites | United States of America | Applicant |
| US20070030356A1 | Cites | United States of America | Applicant |
| US20070041633A1 | Cites | United States of America | Applicant |
| US20070121722A1 | Cites | United States of America | Search report |
| US20070153838A1 | Cites | United States of America | Applicant |
| US20070177813A1 | Cites | United States of America | Applicant |
| US20070205367A1 | Cites | United States of America | Search report |
| US20070211796A1 | Cites | United States of America | Applicant |
| US20070229653A1 | Cites | United States of America | Applicant |
| US20070269136A1 | Cites | United States of America | Applicant |
| US20080199091A1 | Cites | United States of America | Applicant |
| US20080284763A1 | Cites | United States of America | Applicant |
| US20080303895A1 | Cites | United States of America | Applicant |
| US20090002481A1 | Cites | United States of America | Applicant |
| US20090092311A1 | Cites | United States of America | Applicant |
| US20090219282A1 | Cites | United States of America | Applicant |
32 members in 10 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 20593809 | United States of America | P | |
| 26995509 | United States of America | P | |
| 2010000194 | United States of America | W |
Members32
| Document | Office | Kind | |
|---|---|---|---|
| WO2010085361A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2010085361A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AU2010206977A1 | Australia | A1 | |
| KR20110113752A | Republic of Korea | A | |
| US2011286530A1 | United States of America | A1 | |
| EP2389764A2 | European Patent Office (EPO) | A2 | |
| CN102365869A | China | A | |
| JP2012516114A | Japan | A | |
| ZA201104938B | South Africa | B | |
| RU2011135541A | Russian Federation | A | |
| RU2543954C2 | Russian Federation | C2 | |
| CN102365869B | China | B | |
| US9036714B2This record | United States of America | B2 | |
| CN104702960A | China | A | |
| RU2015102754A | Russian Federation | A | |
| CN104717521A | China | A | |
| CN104768031A | China | A | |
| CN104768032A | China | A | |
| US2015222928A1 | United States of America | A1 | |
| AU2010206977B2 | Australia | B2 | |
| JP5877065B2 | Japan | B2 | |
| JP2016105631A | Japan | A | |
| US9420310B2 | United States of America | B2 | |
| KR101676059B1 | Republic of Korea | B1 | |
| JP6176871B2 | Japan | B2 | |
| CN104702960B | China | B | |
| CN104768031B | China | B | |
| CN104717521B | China | B | |
| CN104768032B | China | B | |
| RU2015102754A3 | Russian Federation | A3 | |
| BRPI1007163A2 | Brazil | A2 | |
| RU2689191C2 | Russian Federation | C2 |
71 transactions on the USPTO file
Allowed after 1 non-final rejection and 1 RCE.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 1
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Email NotificationEML_NTR | EML_NTR | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Payment of Maintenance Fee, 4th Year, Large EntityM1551 | M1551 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Email NotificationEML_NTR | EML_NTR | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Preliminary AmendmentA.PE | A.PE | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Workflow - Drawings FinishedDRWF | DRWF | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Miscellaneous Incoming LetterLET. | LET. | |
| Incoming Letter Pertaining to the DrawingsLTDR | LTDR | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail PUB other miscellaneous communication to applicantMM327-D | MM327-D | |
| PUB Other miscellaneous communication to applicantM327-D | M327-D | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Response to 312 Amendment (PTO-271)MN271 | MN271 | |
| Response to Amendment under Rule 312N271 | N271 | |
| Amendment after Notice of Allowance (Rule 312)AllowedA.NA | A.NA | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Reference capture on IDSRCAP | RCAP | |
| Request for Foreign Priority (Priority Papers May Be Included)RQPR | RQPR | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Email NotificationEML_NTR | EML_NTR | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Sent to Classification ContractorPGPC | PGPC | |
| Filing ReceiptFLRCPT.O | FLRCPT.O | |
| Notice of DO/EO Acceptance MailedM903 | M903 | |
| Cleared by OIPE CSRL194 | L194 | |
| Request for Foreign Priority (Priority Papers May Be Included)RQPR | RQPR | |
| Preliminary AmendmentA.PE | A.PE | |
| 371 Completion Date371COMP | 371COMP | |
| Initial Exam Team nnIEXX | IEXX |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| Maintenance fee paymentMAFP | MAFP | |
| AssignmentAS | AS | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS |
Numbers
- Publication
- 9036714
- Application
- 13138267
Titles
- English
- Frame packing for video coding
Patent term adjustment
- A delay
- +493 daysthe office missed an examination deadline
- B delay
- +88 dayspendency past three years
- Applicant delay
- −253 days
- Net adjustment
- 328 days
Classification
- CPC, 14
- H04N21/2365
- H04N19/597
- H04N19/59
- H04N13/194
- H04N21/2383
- H04N21/434
- H04N21/4347
- H04N19/70
- H04N19/46
- H04N19/61
- H04N19/44
- H04N19/132
- H04N21/00
- H04N19/513
- IPC, 10
- H04N7 12
- H04N19 44
- H04N19 46
- H04N19 59
- H04N19 597
- H04N19 61
- H04N19 70
- H04N21 2365
- H04N21 2383
- H04N21 434