System and method for implementing efficient decoded buffer management in multi-view video coding.
Abstract
A system and method for encoding a first picture sequence and a second picture sequence into coded pictures, with the first picture sequence and the second picture sequence being different, and with at least one coded picture of the second picture sequence being predicted from at least one picture in the first picture sequence. According to various embodiments of the present invention, signal element is encoded into a coded picture of the second picture sequence. The signal element indicates whether a picture in the first picture sequence is used for prediction of the coded picture of the second picture sequence.

Term
1.1 yearsleft in the term
Expires 15 October 2027.
- Priority
- Filed
- Granted
- Today
- Expires
12 claims: 6 independent, 6 dependent
- 1CLAIMS:REIVINDICACIONES : 1. Un método para codificar vistas de una escena, dicho método estando comprende los pasos de: one. A method to encode views of a scene, said method being comprises the steps of: construct an initial reference image list based at least in part on intraview reference images and interview reference images, provide a first signaling element to reorder the interview reference images with respect to the image list of initial reference, the first signaling element being derived based at least in part on a view identifier value. construir una lista de imágenes de referencia inicial con base al menos en parte en imágenes de referencia intravista e imágenes de referencia inter-vista, proporcionar un primer elemento de señalización para reordenar las imágenes de referencia inter-vista con respecto a la lista de imágenes de referencia inicial, el primer elemento de señalización siendo derivado con base al menos en parte en un valor de identificador de vistas.
- 3A method according to claim 3. Un método de conformidad con la reivindicación 1, caracterizado porque comprende:1, characterized in that it comprises: proporcionar un segundo elemento de señalización que indica cuál del reordenamiento de imágenes de referencia intravista y del reordenamiento de imágenes de referencia intervista se va a realizar, y providing a second signaling element indicating which of the rearranged reference image rearrangement and rearranged reference image rearrangement is to be performed, and - en caso de que dicho segundo elemento de señalización o - in the event that said second signaling element or INSTITUTO ME > D!;íá l'Pf '( indique el reordenamiento de imágenes de referencia intervista, proporcionar dicho primer elemento de señalización. INSTITUTO ME> D !;íá l'Pf '(indicate the rearrangement of intervened reference images, provide said first signaling element.
- 4Un método de decodificación de un flujo de bits de vídeo codificado, una representación codificada de una pluralidad de vistas de una escena, el método estando caracterizado porque comprende:Four. A method of decoding an encoded video bitstream, an encoded representation of a plurality of views of a scene, the method being characterized in that it comprises: construct an initial reference image list based, at least in part, on intra-view reference images and inter-view reference images;construir una lista de imágenes de referencia inicial con base, al menos en parte, en las imágenes de referencia intra-vista e imágenes de referencia inter-vista;valor de identificador de vista. view identifier value.
- 7An apparatus characterized in that it comprises:7. Un aparato, caracterizado porque comprende: - means for constructing an initial reference image list based at least in part on intra-view reference images and Inter-view reference images, and means for providing a signaling element to reorder the intervened reference images with respect to the initial reference image list, the signaling element being derived based at least in part on a view identifier value. - medios para construir una lista de imágenes de referencia inicial con base al menos en parte en imágenes de referencia intra-vista e imágenes de referencia Ínter -vista, y medios para proporcionar un elemento de señalización para reordenar las imágenes de referencia intervista con respecto a la lista de imágenes de referencia inicial, el elemento de señalización siendo derivado con base al menos en parte en un valor de identificador de vista.
- 9An apparatus according to claim-ie-ation .......... 7, characterized in that it comprises:9. Un aparato según la reivind-ie-ación..........7, caracterizado porque comprende: - means for providing a second signaling element indicating which of the intra-view reference image rearrangement and the inter-view reference image rearrangement is to be performed, and for, in case said second signaling element indicates the rearrangement of the inter-view reference images, providing said first signaling element. - medios para proporcionar un segundo elemento de señalización que indica cuál del reordenamiento de imágenes de referencia intra-vista y el reordenamiento de imágenes de referencia inter-vista se va a realizar, y para, en caso de que dicho segundo elemento de señalización indique el reordenamiento de las imágenes de referencia inter-vista, proporcionar dicho primer elemento de señalización.
- 10An apparatus characterized in that it comprises:10. Un aparato, caracterizado porque comprende: - means for constructing an initial reference image list based at least in part on intra-view reference images and inter-view reference images, and - medios para construir una lista de imágenes de referencia inicial con base al menos en parte en imágenes de referencia intra-vista e imágenes de referencia inter-vista, y - means for reordering intervened reference images with respect to the initial reference image list based at least in part on a signaling element recovered from the encoded bitstream and a view identifier value. - medios para reordenar imágenes de referencia intervista con respecto a la lista de imágenes de referencia inicial con base al menos en parte en un elemento de señalización recuperado a partir de la corriente de bits codificada y un valor de identificador de vista.
Independent claims6
501 paragraphs in 24 sections, as filed
(54) Title: SYSTEM AND METHOD FOR IMPLEMENTING AN EFFICIENT ADMINISTRATION OF DECODED INTERMEDIATE MEMORY IN MULTIPLE VIEW VIDEO CODING.
(54) Title: SYSTEM AND METHOD FOR IMPLEMENTING EFFICIENT DECODED BUFFER MANAGEMENT IN MULTI-VIEW VIDEO CODING.
(57) Summary
A system and method for encoding a first image sequence and a second image sequence into encoded images, the first image sequence and the second image sequence being different, and at least one encoded image being from the predicted second image sequence. from at least one image in the first image sequence. In accordance with various embodiments of the present invention, the signal element is encoded in an encoded image of the second image sequence. The signal element indicates whether an image in the first image sequence is used for predicting the encoded image of the second image sequence.
(57) Abstract
A system and method for encoding a first picture sequence and a second picture sequence into coded pictures, with the first picture sequence and the second picture sequence being different, and with at least one coded picture of the second picture sequence being predicted from at least one picture in the first picture sequence. According to various embodiments of the present invention, signal element is encoded into a coded picture of the second picture sequence. The signal element indicates whether a picture in the first picture sequence is used for prediction of the coded picture of the second picture sequence.
_I KNOW_
SfCMTARJA M tCOMJMÍÁ
Institute
Mexican Property
Industrial
<img file="MX337935B_D0001.tif" />
PATENT TITLE NO. 337935
Owner (s): NOKIA TECHNOLOGIES ΟΥ T
Address: Karaportti 3, FI-02610, Espoo, FINLAND
Name: SYSTEM AND METHOD FOR IMPLEMENTING AN EFFICIENT ADMINISTRATION OF DECODED INTERMEDIATE MEMORY IN MULTIPLE VIEW VIDEO CODING.
Classification: lnt.CI.8: H04N13 / 00; H04N19 / 103; H04N21 / 4545
Inventor (s): YING CHEN; YE-KUi WANG; MISKA HANNUKSELA
REQUEST
Number: International filing date:
MX / a / 2014/009567 October 15, 2007
Divisional Patent Number: 322704
PRIORITY
Country: Date: <sub>Λ</sub> Á Number:
US 1 ..... October 16, 2006 ..... τ '? -' - 60 / 852,223
............ ......
Validity: Twenty years <sup>1</sup>
Expiration Date: October 15, 2027
The reference patent is granted based on articles * 1 * 8 ° fraction 6 useful fraction, and 59 of the jidustrial Property Law.
In accordance with article 23 of the Industrial Property Law, this patent has a non-extendable term of twenty years, counted from the date of filing of the International application and will be subject to the payment of the fee to keep the rights alive.
Whoever signs this title does so based on the provisions of articles 6, sections III and 7, bis 2 of the Industrial Property Law (Official Gazette of the Federation (DOF) 06/27/1991, amended on 02 / 08/1994, 10/25/1996, 12/26/1997, 05/17/1999, 01/26/2004, 06/16/2005, 01/25/2006, 06/05/2009 / 06/01/2010 , 06/18/2010, 06/28/2010, 01/27/2012 and 04/09/2012); Articles 1, 3, fraction V, subsection a), 4, and 12, sections I and III of the Regulations of the Mexican Institute of Industrial Property (DOF 14/12/1999, amended on 07/01/2002, 07/15/19) 2004, 07/28/2004 and 09/07/2007); Articles 1, 3, 4, 5 'section V Clause a), 16 sections I and III and 30 of the Organic Statute of the Mexican Institute of Industrial Property (DOF 12/27/1999, amended on 10/10/2002, 07/29/2004, 08/04/2004 and 09/13/2007); 1, 3 and 5 subsection a) of the Agreement that delegates powers to the Deputy Directors General, Coordinator, Divisional Directors, Holders of the Regional Offices, Divisional Deputy Directors, Departmental Coordinators and other subordinates of the Mexican Institute of Industrial Property. (DOF 12/15/1999, amended on 02/04/2000, 07/29/2004, 08/04/2004 and 09/13/2007).
<img file="MX337935B_D0002.tif" />
Sand; No. 550, Floor 1,
Coi. Santa María Tepeoan village. Xochimilco, CP 16020,
Mexico City
Tel. (55) 53 34 07 0G wwsv.impi gob.wx
Issue Date: March 29, 2016
THE DIVISIONAL DIRECTOR OF PATENTS
<img file="MX337935B_D0003.tif" />
NAHANNY CANAL REYES
<img file="MX337935B_D0004.tif" />
MX / 2016/23596
<img file="MX337935B_D0005.tif" />
SYSTEM AND METHOD FOR IMPLEMENTING AN ADMINI
IMPI lhSTITIJTO MEXICANO aTRACf ^ W '<sup>1</sup>
<img file="MX337935B_D0006.tif" />
OF INTERMEDIATE MEMORY DECODED IN VIDEO CODING OF
MULTIPLE VIEWS
FIELD OF THE INVENTION
The present invention relates generally to video encoding. More specifically, the present invention relates to encoded image buffer management in multi-view video encoding.
BACKGROUND OF THE INVENTION
This section is intended to provide a background or context for the invention that is mentioned in the claims. The present description may include concepts that might be desired, but have not necessarily been previously conceived or accomplished. Therefore, unless otherwise indicated herein, what is described in this section is not prior art to the description or claims of the present application and is not admitted as prior art by its mere inclusion in this section.
In multi-view video encoding, video sequences produced from different cameras, each corresponding to different views of a scene, are encoded in a single bitstream. After decoding, to show a certain view, the decoded images that
Mp
X.
IrtA.
i ¡_ \?.; o: INCA:
<img file="MX337935B_D0007.tif" />
belong that view are reconstructed visualized,
It is also possible that more than one view will be reconstructed and displayed.
Multi-view video encoding processes a wide variety of applications, including point-of-view video / television, three-dimensional (3D) TV, and probing applications. Currently, the Joint Video Team (JVT) of the International Organization for Standardization (ISO) / Moving Image Expert Group (MPEG) of the International Engineering Consortium (IEC) and the Expert Group on
Video Coding of the Telecommunication Union
International (ITU) -T is working to develop a multi-view video encoding (MVC) standard, which is becoming an extension of the ITU-T H standard. 264, also known as ISO / IEC MPEG-4 Part 10. These draft standards are referred to herein as MVC and AVC, respectively. The latest draft of the MVC standard is described in JVT-T208, Joint Multiview Video Model (JMVM) 1.0, 20a. JVT Meeting, Klagenfurt, Austria, July 2006, can be found at ftp3.itu.ch/av-arch/ivtsite/2006 07 Klagenfurt / JVT-T208.zip, and is incorporated herein by reference in its entirety.
In JMVM 1.0, for each group of images (GOP, for
<img file="MX337935B_D0008.tif" />
INSTITUTO rÁi <OE LA PKO;
INDUSTiü.
<img file="MX337935B_D0009.tif" />
its acronym in English), the images of any ..... vlata are contiguous in order of decoding. This is illustrated in Figure 1, where the horizontal direction denotes time (each instant of time being represented by Tm) and the vertical direction denotes view (each view being represented by Sn). The images of each view are grouped into GOPs, for example, images TI to T8 in figure 1 for each view forms a GOP. This decoding order arrangement is referred to as encoding the views first. It should be noted that for images in a view and in a GOP, even though their decoding order is continuous with no other images to be inserted between either image, their decoding order may change internally.
It is also possible to have a different decoding order than discussed for encoding views first. For example, images can be arranged such that images from any temporary location are contiguous in the decoding order. This arrangement is shown in Figure 2. This decoding order arrangement is called time first encoding. It should also be noted that the decoding order of the access units may not be identical to the temporal order.
A typical prediction structure (including inter-image prediction in each view and inter-view prediction) for multi-view video encoding is
I
<img file="MX337935B_D0010.tif" />
INSTITuYC ·., ΑΑΑα, Α. . ....... ii ocla> fV ygj
INDUSTRi .-. L shown in Figure 2, where predictions are indicated by arrows, and the object pointed to uses the object pointed from for prediction reference. Prediction between images within a view is also called temporal prediction, intra-view prediction, or simply interprediction.
An image of
Instant update of
Decoding (IDR) is an intracode image that causes the decoding process to mark all reference images as unused for reference immediately after decoding the IDR image. After decoding an IDR image, all of the following images encoded in the decoding order can be decoded without inter-prediction of any decoded image before the IDR image.
In AVC and MVC, the encoding parameters that are kept unchanged through an encoded video stream are included in a set of streaming parameters. In addition to the parameters that are essential to the decoding process, the stream parameter set may optionally contain Video Utilization Information (VUI), which includes parameters that are important for temporary storage, synchronization of the output of images, representation, and reservation of resources. There are two specified structures to carry parameter sets
INSTITUTE '1'
DE LA F sequence - the NAL unit of the set of parameters Se sequences containing all the data for sequence images, and the extension of the set of sequence parameters for MVC. An image parameter set contains parameters that are unlikely to change across multiple encoded images. Frequently the change of image level data is repeated in each slice header, and the image parameter sets carry the remaining image level parameters. The H.264 / AVC syntax allows many instances of sequence and image parameter sets, and each instance is identified with a unique identifier. Each slice header includes the identifier of the image parameter set that is active for decoding the image containing the slice, and each image parameter set contains the identifier of the active sequence parameter set. Consequently, the transmission of the image and sequence parameter sets does not have to be precisely synchronized with the transmission of the slices. Rather, it is sufficient that the active image and sequence parameter sets are received at any time before they are referenced, allowing transmission of parameter sets using a more reliable transmission mechanism compared to the protocols used for the data. of slices. For example, parameter sets can be included as a MIME parameter in the session description
<img file="MX337935B_D0011.tif" />
for Real-Time Protocol) RTP sessions, of course-<sup>-</sup>'' H.264 / AVC. It is recommended to use a reliable out-of-band transmission mechanism whenever possible in the application in use. If parameter sets are transmitted within the band, they can be repeated to improve error robustness.
As discussed herein, an anchor image is an encoded image in which all slices only refer to slices with the same time index, that is, only slices in other views and not slices in previous images in the current view. An anchor image is signaled by setting an anchor_pic_flag to 1. After decoding the anchor image, all subsequent encoded images in the display order are capable of decoding without inter-prediction of any decoded image prior to the anchor image. If an image in one view is an anchor image, then all images with the same time index in other views are also anchor images. Consequently, the decoding of any view can be started from a temporal index that corresponds to anchor images.
Image output timing, such as timeout logging, is not included in the integral part of AVC or MVC bit streams. However, an image order count (POC) value is derived for each image and is non-decreasing with the increase in the image position in the output order relative to the previous IDR image. or an image containing liria memory management control operation which '^ maT'üar' all images as unused for reference. Therefore, the POC indicates the order of output of the images. It is also used in the decoding process to implicitly scale motion vectors in doubly predictive slice direct modes, for weights implicitly derived in weighted prediction, and for initialization of a reference image list of slice B. Additionally, the POC is also used in the verification of the conformity of the starting order.
POC values can be encoded with one of the three modes indicated in the set of active sequence parameters. In the first mode, the selected number of the least significant bits of the POC value is included in each slice header. In the second mode, the relative increments of POC as a function of the position of the image in the decoding order in the encoded video sequence are encoded in the sequence parameter set. In addition, deviations from the POC value derived from the sequence parameter set can be indicated in the slice headings. In the third mode, the POC value is derived from the decoding order assuming that the decoding and the output order are identical. Also, only a non-reference image can appear consecutively when using the third mode.
<img file="MX337935B_D0012.tif" />
nal_ref_idc is a 2-bit syntax element in the NAL unit header. The value of nal_ref_idc indicates the relevance of the NAL unit for the reconstruction of sample values. The non-zero values of nal_ref_idc should be used for slice data partitioned NAL units and encoded slice of the Reference Images, as well as for NAL units of the parameter set. The nal_ref_idc value must equal 0 for non-reference image slice data slices and partitions and for NAL units that do not affect the reconstruction of sample values, such as NAL units of supplemental enhancement information. In the high-level design of H.264 / AVC, external specifications (i.e. any system or specification that uses or references H.264 / AVC) were allowed to specify an interpretation of non-zero values of nal_ref_idc. For example, the RTP payload format for H.264 / AVC, Request for Comments (RFC) 3984 (which can be found at www.ietf.org/rfc/rfc3984.txt and is incorporated here for reference) specified strong recommendations about using nal_ref_idc. In other words, some systems have established practices for setting and interpreting the non-zero values of nal_ref_idc. For example, an RTP mixer could set nal_ref_idc according to the NAL unit type, for example, nal_ref_idc is set to 3 for NAL IDR units. As MVC is a compatible extension towards
<img file="MX337935B_D0013.tif" />
Behind the H.264 / AVC standard, it is desirable that the H.264 / AVC knowledge system elements are also capable of manipulating MVC streams. Therefore it is undesirable that the particular non-zero value semantics of nal_ref_idc is specified differently in the MVC specification compared to any non-zero value of nal_ref_idc.
The decoded images used to predict subsequent encoded images and for future output are temporarily stored in a decoded image buffer (DPB). To efficiently use the buffer, the DPB management process, including the process of storing decoded images within the DPB, the process of reference image marking, the output and removal of decoded images from the DPB, must be specified.
The process for marking reference images in AVC is generally as follows. The maximum number of reference images used for inter-prediction, designated M, is indicated in the Active Sequence Parameter Set. When an image is decoded, it is marked as used for reference. If decoding the reference image causes more than M images to be marked as used for reference, then at least one image must be marked as unused for reference. The
<img file="MX337935B_D0014.tif" />
DPB removal process would then remove images marked as unused for DPB reference if they do not need to be removed either.
There are two types of operations for marking reference images: adaptive memory control and sliding window. The mode of operation for marking reference images is selected based on the image. Adaptive image control requires the presence of memory management control (MMCO) operation commands in the bitstream. Memory management control operations allow explicit signaling of which images are marked as unused for reference, assigning long-term indices to short-term reference images, storing current images as long-term images, the switching from a short-term image to a long-term image, and assigning the maximum allowed long-term index (MaxLongTermFrameldx) for long-term images. If the sliding window mode of operation is in use and there are M images marked as used for reference, then the short-term reference image that was the first decoded image among the short-term reference images that were marked as used for reference It is marked as unused for reference. In other words, the sliding window mode of operation results in a first-in / first-out buffering operation between
<td rowspan="2">images</td><td rowspan="2">reference Each picture</td><td rowspan="2">of of</td><td colspan="2">short term.</td><td rowspan="2">c dñ ..... uffS</td>
<td>short term</td><td>is associated</td>
<td>variable</td><td>PicNum that</td><td>I know</td><td>derives from</td><td>element of</td><td>syntax</td>
frame_num. Each long-term image is associated with a variable LongTermPicNum that is derived from the long_term_frame_idx syntax element, which is signaled by the MMCO command. PicNum is derived from the FrameNumWrap syntax element, depending on whether the box or field is encoded or decoded. For frames where PicNum equals FrameNumWrap, FrameNumWrap is derived from FrameNum, and FrameNum is derived directly from frame_num. For example, in AVC frame encoding, FrameNum is assigned the same value as frame_num, and FrameNumWrap is defined as follows:
if (FrameNum> frame_num)
FrameNumWrap = FrameNum - MaxFrameNum else
FrameNumWrap = FrameNum
LongTermPicNum is derived from the long-term box index (LongTermFrameldx) assigned for the image. For frames, LongTermPicNum is equal to LongTermFrameldx. frame_num is a syntax element in each slice header. The value of frame_num for a frame or a pair of complementary fields increases essentially by one, in arithmetic module, in relation to the frame num of the previous reference table or the pair of complementary fields of
<img file="MX337935B_D0015.tif" />
reference. In IDR images, the value of frame_nüm is
For images that contain an · eon-t-roi — d © ---- memory management operation that marks all images as unused for reference, the value of frame_num is considered to be zero after image decoding .
MMCO commands use PicNum and LongTermPicNum to indicate the target image for the command as follows. To mark a short-term image as unused for reference, the PicNum difference between the current image p and the target image r is noted in the MMCO command. To mark a long-term image as unused for reference, the LongTermPicNum of the image to be removed is pointed at the MMCO command. To store the current image p as a long-term image, a long_term_frame_idx is signaled with the MMCO command. This index is assigned to the recently stored long-term image as the value of LongTermPicNum. To change an image r from being a short-term image to a long-term image, a PicNum difference is pointed out between the current image p and the image r in the MMCO command, the long_term_f rame_idx is pointed out in the MMCO command, and the index is assigned to this long-term image.
When multiple reference images could be used, each reference image should be identified. In AVC, the identification of a reference image used for an encoded block is as follows. First, all llvl '
INSiYi'UB '
CE L ?. / ·> .: OAC reference images stored in the © PB 'paTa' future image prediction reference — i '2L χ' .. -> A
f.- j. r> V) \ - ...
as '' used for short-term reference (short-term images) or used as long-term reference (long-term images). When a coded slice is decoded, a list of reference images is built. If the coded slice is a doubly predicted slice, then a second list of reference images is also constructed. A reference image used for a coded block is then identified by the index of the reference image used in the reference image list. The index is encoded in the bitstream when more than one reference image is used.
The process of building the list of reference images is as follows. For simplicity, it is assumed that only a list of reference images is needed. First, an initial reference image list is constructed that includes all short-term and long-term images. Reordering the Reference Image List (RPLR) is then performed when the slice header contains RPLR commands. The PRLR process may reorder the reference images in a different order than the order in the initial list. Finally, the final list is constructed by keeping only a number of images at the beginning of the possibly rearranged list, the number being indicated by another syntax element in the header of
<img file="MX337935B_D0016.tif" />
slice or the image parameter set referred to by the slice.
During the initialization process, all short-term and long-term images are considered as candidates for reference image lists for the current image. Regardless of whether the current image is a imagen or P image, the long-term images are placed after the short-term images in RefPicListO (and RefPicListl available for B slices). For P images, the initial reference image list for RefPicListO contains all the short term reference images arranged in descending order of PicNum. For B images, those reference images obtained from all short-term images are ordered by means of a rule related to the current POC number and the reference image POC number for RefPicListO, the reference images with more POC Small (compared to current POC) are considered first and inserted into the RefPicListO in descending order of POC.
Then the images with the largest POC are appended in ascending order of POC. For RefPicListl (if available), reference images with a larger POC (compared to the current POC) are considered first and inserted into the RefPicListl in ascending order of POC. The images with the smallest POC are then appended in descending order of the POC. After considering all the short-term reference images, the long-term reference images are attached in
<img file="MX337935B_D0017.tif" />
ascending order of LongTermPicNum, both pa for B.
<img file="MX337935B_D0018.tif" />
The reordering process is invoked by<sup>-</sup> RPLR commands, which. include four types. The first type is a command to specify a short-term image with a smaller PicNum (compared to a temporarily predicted PicNum) that will be moved. The second type is a command to specify a short-term image with a larger PicNum to be moved. The third type is a command to specify a long-term image with a certain LongTermPicNum to be moved and the end of the RPLR cycle. If the current image is doubly predicted, then there are two cycles, one for a forward reference list and the other for a backward reference list.
The predicted PicNum named picNumLXPred is initialized as the PicNum of the current encoded image. This is set to the PicNum of the image that just moved after each reordering process for a short-term image. The difference between the PicNum of the current image that is reordered and the picNumLXPred will be pointed out in the RPLR command. The image indicated to be reordered is moved to the top of the reference image list. After the reordering process is complete, an entire list of reference images will be truncated based on the size of the active reference image list which is num_ref_idx_IX_active_minusl + l (X equals 0 or 1 corresponds
<img file="MX337935B_D0019.tif" />
INSTIT;
or;
<img file="MX337935B_D0020.tif" />
RefPicListO and RefPicListl respectively).
The hypothetical reference decoder (HRD), specified in Annex C of the H.264 / AVC standard, is used to verify the compliance of the bitstream and the decoder. The HRD contains an encoded image buffer (CPB), an instantaneous decoding process, a decoded image buffer (DPB), and an output image clip block . The CPB and instant decoding process are specified similarly to any other video encoding standard, and the output image trim block simply trims those samples from the decoded image that are outside the scope of the signaled output images. DPB was introduced in H.264 / AVC in order to control the memory resources required to decode conformance bit streams.
There are two reasons for temporarily storing decoded images, for inter-prediction references, and for reordering decoded images in the order of output. As the H.264 / AVC standard provides great flexibility for both reference image marking and output rearrangement, separate buffers for reference image temporary storage and output image temporary storage could be a waste of resources. memory. Therefore, the DPB
<img file="MX337935B_D0021.tif" />
includes a unified decoded image temporary storage process for reference images and output rearrangement. A decoded image is removed from the DPB when it is no longer used as a reference or needed for output. The maximum DPB size that can be used for bit streams is specified in the Level definitions (Annex A) of the H.264 / AVC standard.
There are two types of conformance for decoders: output timing conformance and output order conformance. For output sync compliance, a decoder must output images at identical times compared to HRD. For the conformity of the output order, only the correct order of the output image is taken into account. The output order DPB is assumed to contain a maximum number of frame buffers allowed. A frame is removed from the DPB when it is no longer used as a reference or needed for output. When the DPB is full, the earliest frame in the output order is output until at least one frame buffer is unoccupied.
Temporal scalability is done by hierarchical structure of B-image GOP using only AVC tools. A typical time scalability GOP usually includes a key image which is encoded as an I or P frame, and other images that are encoded as B images. Those B images are hierarchically encoded with
<img file="MX337935B_D0022.tif" />
base on the POC. Coding a GOP just “need the key images from the previous GOP in addition to that ^ -jjaág ^ xias ^^ n ^^^ GOP. The relative number of POCs (POC minus the previous anchor image POC) is called POCIdlnGOP in the implementation. Each POCIdlnGOP can have a form of POCIdInGOP = 2<sup>x</sup>y (where y is an odd number). Images with the same value of x belong to the same time level, which is indicated as Lx (where L = log2 (GOP_length)). Only images with the highest temporal level L are not stored as reference images. Typically, images at a temporal level can only use images at lower temporal levels as references to support temporal scalability, that is, higher temporal level images can be deposited without affecting decoding of lower temporal level images. Similarly, the same hierarchical structure can be applied in the views dimension for view scalability.
In the current JMVM, frame_num is encoded separately and flagged for each view, i.e. the value of frame_num is incremented relative to the previous reference frame or pair of complementary reference fields within the same view as the image current. Also, images in all views share the same DPB buffer. In order to globally manipulate reference image list building and reference image management, FrameNum and POC generation are redefined from the
<img file="MX337935B_D0023.tif" />
INS7 niWtóTXiAl Ή. ' Following way:
FrameNum = frame_num * (l + num_views_minus_l) + view_id PicOrderCnt () = PicOrderCnt () * (l + num_views_minus_l) + view_id;
JMVM basically follows the same reference image marking as that used for AVC. The only difference is that in JMVM the FramNum is redefined and that FrameNumWrap is redefined as follows:
if (FrameNum> frame_num * (l + num_views_minus_l) + view_id)
FrameNumWrap = FrameNum-MaxFrameNum * (l + num_views_minus_l) + view_id else
FrameNumWrap = FrameNum
In the current JMVM standard, Interview reference images are implicitly specified in the SPS (Sequence Parameter Set) extension, where the active number of Interview reference lists and the identification are specified of views of those images. This information is shared by all the images that refer to the same SPS. The process of building the reference image list first performs the initialization of the reference image list, reordering and truncating in the same way as in AVC, but taking into account all reference images
<img file="MX337935B_D0024.tif" />
stored in the DPB. Images with view identifications specified in the SPS and within the same time axis (i.e. having the same capture / exit time) are then appended to the list of references in the order in which they are listed in the SPS.
Unfortunately, previous JSVM designs give rise to various problems. First, sometimes it is desirable that the change of views decoded (by means of a decoder), transmitted (by means of a sender) or sent (by means of a gateway or MANE) at a different time index than that corresponding to anchor images. For example, a base view can be compressed for the highest encoding efficiency (time prediction is used a lot) and anchor images are infrequently encoded. Consequently, anchor images for other views also occur infrequently, because they are synchronized across all views. The current JMVM syntax does not include signaling an image from which decoding of a certain view can be started (unless all views in the time index contain an anchor image).
Second, the reference views allowed for inter-view prediction are specified for each view (and separately for anchor and non-anchor images). However, depending on the similarity between an image that is encoded and a potential image on the same time axis and in a view
<img file="MX337935B_D0025.tif" />
or not σ 'Τ7ϊ Τ' J5 .. · *
-τ .MEXICANO ': V fROVlfPAU of potential reference, inter-view prediction can be done in the encoder. The current JMVM standard uses nal_ref_idc to indicate whether an image is used for intra-view or inter-view prediction, but cannot separately indicate whether an image is used for intra-view prediction and / or inter-view prediction. Also, according to JMVM 1.0, for AVC compliant view, nai_ref_idc must be set not equal to 0 even if the image is not used for temporal prediction when used only for inter-view prediction reference. Consequently, if only that view is decoded and output, an additional DPB size is needed for the storage of such images when those images can be output as soon as they are decoded.
Third, it can be seen that the reference image marking process specified in JMVM 1.0 is basically identical to that of the AVC process, except for the redefinition of FramNum, FrameNumWrap and consequently PicNum. Therefore, several special problems arise. For example, this process cannot efficiently handle the management of decoded images that are required to be temporarily stored for interview prediction, particularly when those images are not used for temporal prediction reference. The reason is that the DPB management process specified in the AVC standard was aimed at single view encoding. In single-view encoding such as in the AVC standard, images
<img file="MX337935B_D0026.tif" />
Decodes that need to be temporarily stored for the future prediction or future output reference can be removed from the buffer when they are no longer needed for future output and temporary forecast reference. To allow removal of a reference image as soon as it is no longer needed for temporal prediction reference and future output, the reference image marking process is specified such that it can be known immediately after a reference image is already not needed for temporal prediction reference. However, when it comes to images for inter-view prediction reference, there is a lack of a way to tell immediately after an image is no longer needed for inter-view prediction reference. Consequently, interview prediction reference images can be temporarily stored unnecessarily in the DPB, which reduces the efficiency of using the buffer.
In another example, given the way to recalculate the PicNum, if the sliding window operation mode is in use and the number of short-term and long-term images equals the maximum, the short-term reference image has the smallest FrameNumWrap is marked as unused for reference. However, due to the fact that this image is not necessarily the earliest encoded image because the FrameNum order in the current JMVM does not follow the decoding order, the window reference image
<img file="MX337935B_D0027.tif" />
Slider does not operate optimally in the current JMVM. Additionally, due to the fact that the PicNum is derived from the redefined and scaled FrameNumWrap, the difference between the PicNum values of two encoded images would be scaled on average. For example, it is helpful to assume that there are two images in the same view and that they have a frame_num equal to 3 and 5, respectively. When there is only one view, i.e. the bit stream is an AVC stream, then the difference of the two PicNum values would be 2. When encoding the image having a frame_num equal to 5, if an MMCO command is needed to mark the image that has PicNum equal to 3 as not used for reference, then the difference of the two values minus 1 is equal to 1, which will be indicated in the MMCO. This value needs 3 bits. However, if there are 256 views, then the difference of the two PicNum values minus 1 would become 511. In this case, 19 bits are required for signaling the value. Consequently, MMCO commands are encoded much less efficiently. Typically, the highest number of bits equals 2 * log2 (number of views) for an MMC command from the current JMVM compared to H.264 / AVC plain view encoding.
A fourth set of problems has to do with the process of building the reference image list specified in JMVM 1.0. The reference image list initialization process considers reference images from all views prior to the
<img file="MX337935B_D0028.tif" />
reordering. However, due to the fact that 'láh liuáui.nac .— „---- of the other views used for inter-view prediction are appended to the list after truncating the list, the reference images of the other views are not appear in the reference image list after reordering and truncating in any way. Therefore, consideration of those images is not required in the initialization process. Also, illegal reference images may appear (images that have a different view_id from the current image and are not temporarily aligned with the current image) and repeated interviewed reference images may appear in the finally constructed reference image list.
The reference image list initialization process operates as listed in the following steps: (1)
All reference images are included in the initial list regardless of their view_id and whether they are temporarily aligned with the current image. In other words, the initial reference image list may contain illegal reference images (images that have a different view_id than the current image and are not temporarily aligned with the current image). However, in first view encoding, the beginning of the initial list contains reference images of the same view as the current image. (2) Both intra-view reference images and inter-view images can be reordered. After reordering, the beginning of the list may still contain images
J.
illegal referrals. (3) The list is tr'incadá7<sup>:</sup>-'- pefo — Truncated list may still contain images-of-r-ef-er & ncL ^ ™.
illegal. (4) Interview reference images are appended to the list in the order in which they appear in the SPS MVC extension.
Additionally, the reordering process of the reference image list specified in JMVM 1.0 does not allow the reordering of interview frames, which are always placed at the end of the list in the order in which they appear in the MVC extension of SPS. This results in less flexibility in constructing the reference image list, resulting in reduced compression efficiency, when the default order of inter-view reference frames is not optimal or certain inter-view reference frames have more likely to be used for prediction than certain intra-view reference tables. Additionally, similar to MMCO commands, due to the fact that PicNum is derived from redefined and scaled FrameNumWrap, VLC codewords are no longer required to encode RPLR commands that include signaling a difference between PicNum values minus 1 compared to the simple-view encoding of the H.264 / AVC standard.
SUMMARY OF THE INVENTION
The present invention provides an improved system and method for implementing efficient decoded image buffer management in encoding: ·. video views. In a modality /. __ new flag to indicate if the decoding of a view can be started from a certain image. In a more particular embodiment, this flag is signaled in the NAL unit header. In another fashion, a new flag is used to indicate whether an image is used for inter-view prediction reference, while the nal_ref_idc syntax element only indicates whether an image is used for temporal prediction reference. This flag can also be signaled in the NAL unit header. In a third fashion, a set of reference image marking methods are used to efficiently manage the decoded images. These methods can include both a sliding window and adaptive memory control mechanisms. In a fourth embodiment, a set of new reference image list construction methods are used and include both initialization and reordering of reference image lists.
These and other advantages and characteristics of the invention, together with the organization and manner of operation thereof, will become apparent from the following detailed description when considered in conjunction with the accompanying drawings, where like elements have numbers similar to throughout the various drawings described below.
<img file="MX337935B_D0029.tif" />
BRIEF DESCRIPTION OF THE DRAWINGS
<img file="MX337935B_D0030.tif" />
Fig. 1 is an image arrangement in a first view encoding arrangement;
Fig. 2 is an image arrangement of a first time encoding arrangement;
Figure 3 is an illustration of a MVC interviews and time prediction structure;
Figure 4 is a general diagram of a system in which the present invention can be implemented;
Figure 5 is a perspective view of a mobile device that can be used in the implementation of the present invention; and
Figure 6 is a schematic representation of the circuits of the mobile device of Figure 5.
DETAILED DESCRIPTION OF THE INVENTION
Figure 4 shows a generic multimedia communication system for use with the present invention. As shown in Figure 4, a data source 100 provides a source signal in an analog, uncompressed digital, or compressed digital format, or any combination of these formats. An encoder (110) encodes the source signal to an encoded media bit stream. Encoder (110) may be capable of encoding more than one type of medium, such as audio and video, or more than one encoder (110) may be required to encode different types of signal media. Encoder (110) may also obtain - one .------ input ^.
synthetically produced, such as graphics and text, or may be capable of producing encoded bit streams from synthetic media. In the following, only the processing of an encoded media bit stream of a media type is considered to simplify the description. It should be noted, however, that real-time broadcast services typically comprise multiple streams (typically at least one stream of audio, video, and text captioning). It should also be noted that the system may include many encoders, but only one encoder (110) is considered below to simplify the description without a lack of generality.
The encoded media bit stream is transferred to a storage (120). Storage 120 may comprise any type of bulk memory for storing the encoded media bit stream. The format of the encoded media bitstream in storage 120 may be an elementary standalone bitstream format, or one or more encoded media bitstreams may be encapsulated in a container file. Some systems operate live, that is, they skip storage and transfer the encoded media bit stream from encoder (110) directly to emitter (130). The encoded media bit stream is then transferred to sender 130, also called a server, based on need. The format used in the chi ϊί'ίϋ'ϊ transmission may be an elementary ^^ áutóniórnójc ^ bitstream format, a packet stream format, ο ^ τγγτγγ- · —encoded media bitstreams may be encapsulated in a container file. Encoder 110, storage 120, and emitter 130 may reside on the same physical device or may be contained on separate devices. Encoder 110 and emitter 130 can operate with live real-time content, in which case the encoded media bit stream is typically not permanently stored, but is temporarily stored for short periods of time in the encoder content (110) and / or emitter (130) to smooth variations in processing delay, transfer delay, and encoded media bit rate.
The emitter (130) sends the encoded media bit stream using a communication protocol stack. The stack may include but is not limited to a Real-time Transport Protocol (RTP), User Datagram Protocol (UDP), and Internet Protocol (IP). acronym in English). When the communication protocol stack is packet oriented, the sender 130 encapsulates the bit stream of packet encoded media. For example, when RTP is used, sender 130 encapsulates the bitstream of encoded media in RTP packets according to an RTP payload format. Typically, each media type has a dedicated RTP payload format. Duty<sup>l</sup>ar -'- '' 'i% €> ta «á5e _; _><sup>il</sup>'again that a system can contain more ”~' tte“ -a »—emitter (130), but for simplicity, the following description only considers an emitter (130).
The emitter (130) may or may not be connected to a gateway (140) through a communication network. The gateway 140 can perform different types of functions, such as moving a packet stream according to one communication protocol stack to another communication protocol stack, gathering and branching the data streams, and manipulating data streams according to the downlink and / or receiving capabilities, such as controlling the bit rate of the sent stream according to the prevailing downlink network conditions. Examples of gateways (140) include Multipoint Conference Control Units (MCUs), gateways between circuit-switched and packet-switched video telephony, Push-To-Talk-Cellular (PoC) servers ), IP encapsulators, in Digital Video Broadcasting Systems for Mobile Devices (DVB-H), or decoders that send locally broadcast transmissions to home wireless networks. When RTP is used, the gateway (140) is called the RTP mixer and acts as the endpoint of an RTP connection.
The system includes one or more receivers (150), typically capable of receiving, demodulating, and decapsulating the
ΙΜΡΐβΐ - '* ““ · ”· ^ · V
INSTITUTE v? XIO <sup>μ</sup>Ο V'7
C, ί.Λ UIGFIELAD V-:
¡.'¡ □ • JSr.UU signal transmitted to an encoded media bit stream. The encoded media bit stream is typically further processed by a decoder 160, the output of which is one or more uncompressed media streams. It should be noted that the bit stream to be decoded can be received from a remote device located within virtually any type of network. Additionally, the bit stream can be received from local hardware or software. Finally, a media synthesizer (170) can reproduce the uncompressed media streams, for example with a speaker or a screen. The receiver 150, the decoder 160, and the media synthesizer 170 can reside on the same physical device or can be included on separate devices.
Scalability in terms of bit rate, decoding complexity, and image size is a desirable property for heterogeneous and error prone environments. This property is desirable in order to counter limitations such as restrictions on bit rate, display resolution, network throughput, and computing power on a receiving device.
<td>Shall</td><td>be understood</td><td>than,</td><td>in spite of</td><td>that he</td><td>text and</td>
<td colspan="2">examples contained in</td><td>the</td><td>Present</td><td>they can</td><td>describe</td>
<td>specifically</td><td>a process</td><td>of</td><td>coding</td><td>, a</td><td>person with</td>
Experience in the art would easily understand that the same concepts and principles also apply to the corresponding decoding process and vice versa. It should be appreciated that
<img file="MX337935B_D0031.tif" />
the bit stream to be decoded can be received from a remote device located within virtually any type of network. Additionally, the bit stream can be received from local hardware or software.
The communication devices of the present invention can communicate using various transmission technologies including, but not limited to, Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM). English), Universal Mobile Telecommunications System (UMTS), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Transmission Control Protocol / Internet Protocol (TCP / IP), Short Message System (SMS) , Multimedia Messaging Service (MMS), email, Instant Messaging Service (IMS), Bluetooth, IEEE 802.11, etc. A communication device can communicate using various means including, but not limited to radio, infrared, laser, wired connection, and the like.
Figures 5 and 6 show a representative mobile device (12) on which the present invention can be implemented. However, it should be understood that the present invention is not intended to be limited to a particular type of
Gave
I urge b-ru>
ÜE ΙΛ FKOFIH'MJ
INOUSTKIAl.
<img file="MX337935B_D0032.tif" />
<img file="MX337935B_D0033.tif" />
I mobile device (12) or another electronic device -----...
Some of the features illustrated in Figures 5 and 6 could be incorporated into any or all of the devices that can be used in the system shown in the figure.
The mobile device (12) of Figures 5 and 6 includes a housing (30), a display (32) in the form of a liquid crystal display, a keyboard (34), a microphone (36), a headset (38) , a battery (40), an infrared port (42), an antenna (44), a smart card (46) in the form of a UICC in accordance with an embodiment of the invention, a card reader (48), a circuit radio interface (52), a codee circuit (54), a controller (56) and a memory (58).
The individual circuits and elements are all of a type well known in the art, for example in the range of Nokia mobile devices.
The present invention provides an improved system and method for implementing efficient decoded image buffer management in multi-view video encoding. To address the issue around the fact that the JMVM syntax does not include pointing to an image from which decoding of a certain view can be started (unless all views in the time index contain an anchor image), a new flag is signaled indicating if a view can be accessed from a certain image, that is, if the decoding of a view can be started from
<img file="MX337935B_D0034.tif" />
certain image. In one embodiment of the invention, this Joandeta is signaled in the NAL unit header. The following is an example of the syntax and semantics of the flag according to a particular modality. However, it is also possible to change the semantics of the anchor_pic_flag syntax element in a similar way instead of adding a new syntax element.
<td>nal unit header svc mvc extension () {</td><td>C</td><td>Descriptor</td>
<td>svc mvc flag</td><td>Everybody</td><td>u (l)</td>
<td>if (! svc mvc flag) {</td><td></td><td></td>
<td>priority id</td><td>Everybody</td><td>u (6)</td>
<td>discardable flag</td><td>Everybody</td><td>u (l)</td>
<td>temporary level</td><td>Everybody</td><td>u (3)</td>
<td>dependency id</td><td>Everybody</td><td>u (3)</td>
<td>quality level</td><td>Everybody</td><td>u (2)</td>
<td>layer base flag</td><td>Everybody</td><td>u (l)</td>
<td>use base prediction flag</td><td>Everybody</td><td>u (1)</td>
<td>fragmented flag</td><td>Everybody</td><td>u (l)</td>
<td>last fragment flag</td><td>Everybody</td><td>u (l)</td>
<td>fragment order</td><td>Everybody</td><td>u (2)</td>
<td>reserved zero two bits</td><td>Everybody</td><td>u (2)</td>
<td>} else {</td><td></td><td></td>
<td>view refresh flag</td><td>Everybody</td><td>u (l)</td>
<td>view subset id</td><td>Everybody</td><td>u (2)</td>
<td>view level</td><td>Everybody</td><td>u (3)</td>
<td>anchor foot flag</td><td>Everybody</td><td>u (1)</td>
<td>view id</td><td>Everybody</td><td>u (10)</td>
<td>reserved zero five bits</td><td>Everybody</td><td>u (6)</td>
<td> }</td><td></td><td></td>
<td>nalUnitHeaderBytes + = 3</td><td></td><td></td>
<td> }</td><td></td><td></td>
For a certain image in one view, all images in the same time site from other views that use inter-view prediction are called direct dependency view images, and all images in the same time site from other views that need to be decoded the current image are called dependency images.
The semantics of view_refresh_flag can be specified in four ways in one mode. A first way to specify the semantics of view_refresh_flag involves view_refresh_flag indicating that the current image and all subsequent images in the output order in the same view can be successfully decoded when all direct dependency view images of current and subsequent images in the same view is also (possibly partially) decoded without decoding any preceding image in the same view or in other views. This implies that (1) none of the dependency views falls on any preceding image in the decoding order in any view, or (2) if any of the dependency view images falls on any preceding image in the decoding order on any view, then only the restricted intra-coded areas of the images of direct dependency views of the current and subsequent images in the same view are used for interview prediction. A restricted intracoded area does not use data from neighboring intercoded areas for intra prediction.
A second way of specifying the semantics of view_refresh_flag involves view_refresh_flag indicating that the current image and all subsequent images in decoding order in the same view can be successfully decoded when all de '' view images:,? ί · <
IFU \ J '(-) ^ ílttágeiréé ^ -
1NS7 current and direct dependence on subsequent images in the same view also s cfi- ~ -deo <> di-'É-i-ea da s · completely or, in one mode, partially without decoding any preceding image.
A third way to specify the semantics of view_refresh_flag involves view_refresh_flag indicating that the current image and all subsequent images in the output order in the same view can be successfully decoded when all dependency view images of current and subsequent images in the The same view is also fully or partially decoded. This definition is analogous to an intra-image that initiates an open GOP in single-view encoding. In terms of specification text, this option can be written as follows: A view_refresh_flag equal to 1 indicates that the current image and any subsequent images in the decoding order in the same view as the current image and following the current image in the order of output does not refer to an image that precedes the current image in the decoding order in the inter-prediction process. A view_refresh_flag equal to 0 indicates that the current image or a subsequent image in the decoding order in the same view as the current image and following the current image in the output order can refer to an image that precedes the current image in the decoding order in the interprediction process.
go z ··
<img file="MX337935B_D0035.tif" />
to specify the semaicit that view_r efr is h_f 1 ag “~ í 'ηοΓΓΤρτδ — qtre subsequent images in the order
A fourth way view_refresh_flag involves the current image and all decoding images in the same view can be successfully decoded when all dependency view images of current and subsequent images in the same view are also fully or, in one mode, partially decoded. This definition is analogous to an intra image that initiates a closed GOP in single view encoding.
The view_refresh_flag can be used in a system such as the one illustrated in figure 4. In this situation, the receiver (150) has received, or the decoder (160) has decoded, only certain subset M of all the N available views, the subset excluding view A. Due to a user action, for example, receiver 150 or decoder 160 would like to receive or decode view A respectively from now on. The decoder can start decoding view A of the first image, with view_refresh_flag equal to 1 in view A. If view A was not received, then the receiver (150) can indicate to the Gateway (140) or to the sender (130) including encoded images of view A in the transmitted bit stream. The gateway (140) or the sender (130) can wait until the next image has a view refresh flag equal to 1 in view A before sending any image from view A with
--- <7¡ in order to avoid sending unnecessary images, give the · ”view A that the decoder (160) could not decode κ… - * ·· ............<sup>..</sup>
To deal with the second issue discussed above, a new flag is signaled to indicate whether a view is used for interview prediction reference, and the nal_ref_idc syntax element only indicates whether an image is used for temporal prediction reference. In a particular embodiment, this flag is signaled in the NAL unit header. The following is an example of the flag syntax and semantics.
<td>nal unit header svc mvc extension () {</td><td>C</td><td>Descriptor</td>
<td>svc mvc flag</td><td>Everybody</td><td>u (1)</td>
<td>if (! svc mvc flag) {</td><td></td><td></td>
<td>priority id</td><td>Everybody</td><td>u (6)</td>
<td>discardable flag</td><td>Everybody</td><td>utl)</td>
<td>temporary level</td><td>Everybody</td><td>u (3)</td>
<td>dependency id</td><td>Everybody</td><td>u {3)</td>
<td>quality level</td><td>Everybody</td><td>u (2)</td>
<td>layer base flag</td><td>Everybody</td><td>u (l)</td>
<td>use base prediction flag</td><td>Everybody</td><td>u (l)</td>
<td>fragmented flag</td><td>Everybody</td><td>u (l)</td>
<td>last fragment flag</td><td>Everybody</td><td>u (l)</td>
<td>fragment order</td><td>Everybody</td><td>u (2)</td>
<td>reserved zero two bits</td><td>Everybody</td><td>u (2)</td>
<td>} else {</td><td></td><td></td>
<td>Inter view refresh flag</td><td>Everybody</td><td>u (l)</td>
<td>view subset id</td><td>Everybody</td><td>u (2)</td>
<td>view level</td><td>Everybody</td><td>u (3)</td>
<td>anchor foot flag</td><td>Everybody</td><td>u (l)</td>
<td>view id</td><td>Everybody</td><td>u (10)</td>
<td>reserved zero five bits</td><td>Everybody</td><td>u (5)</td>
<td> }</td><td></td><td></td>
<td>nalUnitHeaderBytes + = 3</td><td></td><td></td>
<td> }</td><td></td><td></td>
An inter_view_refresh_flag equal to 0 indicates that the current image is not used as an interview reference image. An inter_view_refresh_flag equal to 1 indicates that the
<img file="MX337935B_D0036.tif" />
li'.S'n current image is used as an inter-view reference image. It is inferred that the value of inter_view_refresh_flag is equal when profile_idc indicates an MVC profile and view_id is 0. When decoding an image, all images that have an inter_view_refresh_flag equal to 1 and with the same time axis as the current image are called images Interviews of the current image.
The inter_view_refresh_flag can be used in a gateway (140), also referred to as a media knowledge network element (MANE). When an image is not used as an inter-view reference and intra-view reference (inter_view_refresh_flag equals 0 and nal_ref_idc equals 0), a MANE may choose not to send it without consequences on decoding the remaining bitstream. When an image is not used as an inter-view reference but is used as an intra-view reference, a MANE should deposit the image only if it also deposits the transmission of the dependent views. When an image is not used as an inter-view reference, but is used as an intra-view reference, a MANE should deposit the image only if the view where the image resides is not required or decoded.
Regarding the issue of the reference image marking process specified in JMVM 1.0 it is not capable of efficiently handling the management of decoded images that must be temporarily stored for i
Interview prediction is reused Tá: ......: - (Tita
<img file="MX337935B_D0037.tif" />
inter_view_refresh_flag.
Inter_view_refresh_flag images equal to 1 can be marked using any of three methods.
A first method of marking images with an inter_view_refresh_flag equal to 1 includes storing interviewed reference images temporarily as long-term images. In the encoding process, each image used for interview prediction is indicated in the bitstream to be marked as used for long-term reference. One way to indicate the markup as used for long-term reference is the inter_view_refresh_flag. The decoder responds to the prompt by marking the image as used for long-term reference and multi-view long-term time reference. Any memory management control operation directed at an image marked as used for long-term reference and multi-view long-term time reference is temporarily stored. When all images on the time axis are encoded or decoded, all images marked as used for long-term reference and long-term multi-view time reference are no longer marked as used for long-term reference and long-term time reference multi-view, and markup of reference images is performed again in their decoding order using either the ίί slider window operation or administering control operations (any I applied "'3" "a— --Marion _ particular). For example, if an image is used for interprediction (that is, the value of nal_ref_idc is greater than 0), it is re-marked as used for short-term reference. If the image is not used for interprediction (that is, nal_ref_idc equals 0), it is marked as unused for reference. Usually, there are only two cases for the image on a certain time axis: all images are reference images for inter-prediction, or no image is a reference image for inter-prediction. This last operation can be performed after the last VCL NAL unit on the time axis has been decoded, or before the next access unit or the next image on the subsequent time axis is decoded. In the decoding process, the operation at this stage may be implicitly triggered by the change in the time axis, or it may be explicitly signaled, for example, as an MMCO command. With this method, interviewed reference images have the same influence as long-term reference images for weighted prediction and in temporal direct mode.
A second method of marking images with an inter_view_refresh_flag equal to 1 includes marking the interviewed reference images as used for interview reference. With this method the marking of the reference image for interprediction (marking as used for
<img file="MX337935B_D0038.tif" />
short-term reference and used for long-term reference) remains unchanged compared to the AVC standard. For processes related to temporal direct mode and weighted prediction, images marked as used for inter-view reference, that is, those inter-view reference images that share the same time axis as the current image, are treated identically to those Long-term reference images. When all images on the time axis are encoded or decoded, all images marked as used for inter-view reference are no longer marked as used for inter-view reference.
It is appreciated that removal of the marking used for inter-view reference after all images on the temporal axis are processed is only one embodiment of the invention. Marking as used for interview reference could also be removed at other times in the decoding process. For example, marking as used for inter-view reference of a particular image may be removed as soon as the current image or any subsequent image ceases to depend directly or indirectly on the image according to the view dependency signaling included in the extension SPS MVC.
The operation of no longer marking the appropriate images as used for inter-view reference can be performed after the last VCL NAL unit on the time axis is decoded or before the next ia'ó ;, '! Unit is decoded. .
I heard , access or the next image on the subsequent time axis. In the decoding process, this can be triggered implicitly by the change in the time axis or it can be explicitly signaled, for example, as an MMCO command.
With this particular method, interview reference images have the same influence as long-term reference images for weighted prediction and in temporal direct mode. In other words, this method has the same effect as the first method discussed above for weighted prediction and in direct temporal mode.
In the present method, an improved sliding window mechanism can be applied to remove marking used for inter-view reference from images used only for inter-view prediction, i.e. images that have nal_red_idc equal to 0 and marked as used for inter-reference -views. This improved sliding window mechanism uses a variable, for example called num_inter_view_ref_frames, preferably signaled in the SPS extension for MVC, such that when the number of images marked as used for interview reference and having nal_ref_idc equal to 0 is equal to num_inter_view_ref_frames, then the one that was decoded first becomes unmarked as used for interview reference. Consequently, if the image also does not need to be taken (to be taken now or not intentionally taken) the
<img file="MX337935B_D0039.tif" />
The decoder can invoke a process to remove v ^ gÁ ^ i from the DPB such that a newly decoded image can be stored in the DPB.
A third method of marking images with an inter_view_reference_flag equal to 1 includes marking images after decoding all images on the same time axis / time index. Rather than marking an image immediately after decoding, this method is based on the idea that images are marked after decoding all images on the same time axis (i.e. the same time index). Sliding window or adaptive reference image marking as indicated in each of the encoded images is performed in the order in which the images were decoded. For processes related to temporal direct mode and weighted prediction, marked images of the same temporal axis as the current image are treated identically to long-term reference images. Interview reference images from the same time axis as the current image are included in the construction of the initial reference image lists and can be reordered based on their view_id or assigned long-term reference indexes first and can then revert to correlate based on the long-term benchmark.
As previously discussed, given the way to recalculate the PicNum, if the window operation mode; ; '·' ''>. · ----- - • '¿μ *' Λ slider is in use and the number of long-term images equals the maximum, the image of —r-eljare n c of short The term that the smallest FrameNumWrap has is marked as unused for reference. However, due to the fact that this image is not necessarily the earliest encoded image because the FrameNum order in the current JMVM does not follow the decoding order, the sliding window reference image markup does not operate optimally in the Current JMVM. To solve this problem, and compared to the JMVM standard, the variable FrameNum and FrameNumWrap are not redefined / scaled, that is, their definition remains unchanged compared to the AVC standard. It is designed that short-term images can be automatically managed by the first-in, first-out mechanism of the sliding window. Only a slight modification of the sliding window mechanism is required compared to JMVM 1.0. The modifications are as follows, with the new text represented in italics:
G.8.2.5.3 Marking process for decoding sliding window reference images
This process is adaptive_ref_pic_marking_mode_flag equals 0. Reference images that have current slice are considered in the calculation of numShortTerm and numLongTer.
num ref frames.
<td>invoked</td><td colspan="2">when</td>
<td>equal to 0.</td><td>Alone</td><td>the</td>
<td>same view</td><td>go what</td><td>the</td>
<td colspan="2">process including</td><td>the</td>
<td>and the value</td><td>applied</td><td>of</td>
7Γ
<img file="MX337935B_D0040.tif" />
In the above method, the total number of reference frames for the entire MVC bit stream, indicating the size of the image storage buffer used for intra-view or inter-view reference of an entire bit stream of MVC, must be equal to the sum of the values of num_ref_frames applied for all the views contained in the MVC bitstream plus the maximum number of interview reference frames to decode the MVC bitstream. Alternatively, the sliding window can be made globally for all images in all views.
For time first encoding, the sliding window process is defined as follows, with new text for JMVM 1.0 rendered in italics:
G.8.2.5.3 Marking process for decoding sliding window reference images
When numShortTerm + numLongTerm equals Max (num_ref_frames, 1), the condition must be met that numShortTerm is greater than 0, and the short-term reference box, the pair of complementary reference fields or non-pair reference field that is select by the following rule is marked as unused for reference. When it is a box or a pair of complementary fields, both fields are also marked as not used for reference.
* i <3 selection rule is: of all those Images with the smallest value of FrameNumWrap, the first one is selected in the decoding order. The decoding order of those images can be indicated by the value of view_id, or the dependency information of signaled views in the SPS of the MVC extension.
For time first encoding, the sliding window process is defined as follows, with new text for JMVM 1.0 rendered in italics:
G.8.2.5.3 Marking process for decoding sliding window reference images
When numShortTerm + numLongTerm equals Max (num_ref_frames, 1), the condition must be met that numShortTerm is greater than 0, and the short-term reference box, the pair of complementary reference fields or non-pair reference field that is select by the following rule is marked as unused for reference. When it is a box or a pair of complementary fields, both fields are also marked as not used for reference.
* The selection rule is: of all those images in the most previously decoded view, the one with the smallest FrameNumWrap value is selected. The decoding order of views can be indicated by the view id value, or the flagged view dependency information
O to
TO /
l »! / 11. i? '
INSTÍTÍ.
GAVE
<img file="MX337935B_D0041.tif" />
in the SPS of the MVC extension.
As discussed above, due to the fact that PicNum is derived from the redefined and scaled FrameNumWrap, the difference between the PicNum values of the two encoded images would be scaled on average. For example, it is useful to assume that there are two images in the same view and that they have frame_num equal to 3 and 5, respectively. When there is only one view, i.e. the bit stream is an AVC stream, then the difference of the two PicNum values would be 2. When encoding the image that has frame_num equal to 5, if an MMCO command is needed to mark the image that has PicNum equal to 3 as not used for reference, then the difference of the two values minus 1 equals 1, which it will be signaled at the MMCO. This value needs 3 bits. However, if there are 256 views, then the difference of the two PicNum values minus 1 would be 511. In this case, 19 bits are required to signal the value. Consequently, MMCO commands are encoded much less efficiently. Typically, the greatest number of bits equals 2 * log2 (number of views) for a current JMVM MMCO command compared to H.264 / AVC plain view encoding.
To solve this problem and in contrast to the JMVM standard, the FrameNum and FrameNumWrap variables are not redefined / scaled, which is the same as in the AVC standard. In most cases, it is not required from the point
<img file="MX337935B_D0042.tif" />
size view of DPB an image contains an MMCO command to remove an image that does not belong to the same view nor does it belong to the same time axis as the current image. Even some of the images are no longer necessary for reference and therefore may be marked as unused for reference. In this case, marking can be done using the sliding window process or postponed to the next encoded image with the same view_id. Therefore, MMCO commands are restricted only to mark images, as not used as a reference for images that belong to the same view or the same time axis, although the DBP may contain images from different views or different time axes.
The modification of JMVM 1.0 for marking intra-view reference images is as follows, with the changes shown in italics:
G.8.2.5.4.1 Process for marking a short-term reference image as not used as a reference
This process is' invoked when adaptive_pic_marking_mode_flag equals 1. Only reference images that have the same view_id as the current slice are considered in this process.
The syntax and semantics for marking reference images can be as follows:
<td>slice header () {</td><td>C</td><td>Descriptor</td>
<td> . . ·</td><td></td><td></td>
<td>if (nal ref idc! = 0)</td><td></td><td></td>
<td>dec ref foot marking ()</td><td> 2</td><td></td>
<td>if (inter view reference flag)</td><td></td><td></td>
<td>dec view ref foot marking mvc ()</td><td> 2</td><td></td>
<td> }</td><td></td><td></td>
<td></td><td></td><td></td>
<td>dec view ref foot marking mvc () {</td><td>C</td><td>Descriptor</td>
<td>adaptive view ref foot marking mode flag</td><td> 2</td><td>u (l)</td>
<td>if (adaptive view ref foot marking mode flag)</td><td></td><td></td>
<td>do{</td><td></td><td></td>
<td>view memory management control operation</td><td> 2</td><td>ue (v)</td>
<td>if (view memory management control operation == 1 II view memory management control operation == 2)</td><td></td><td></td>
<td>abs difference of view id minusl</td><td> 2</td><td>ue (v)</td>
<td>} while (view memory management control operation! = 0)</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
The values for the memory management control operation (view_memory_management_control_ operation) are as follows
<td>view memory management with trol operation</td><td>Memory Management Control Operation</td>
<td> 0</td><td>End cycle view memory_ management control operation</td>
<td> 1</td><td>Remove the marking of used for reference between views or mark an image as not used for reference, abs difference of view id minusl is present and corresponds to a difference to subtract from current view id</td>
<td> 2</td><td>Remove marking used for reference between views or mark image as unused for reference, abs difference of view id minusl is present and corresponds to a difference to add to current view id</td>
The adaptive_view_ref_pic_marking_mode_flag specifies whether the sliding window mechanism is in use (when it is
<img file="MX337935B_D0043.tif" />
<img file="MX337935B_D0044.tif" />
LíJ ΰ .1 and
INSTITUTE M rvLC of u '«WSWMl equal to 0) or the process of marking of images of adaptive RTTTT—— (when it is equal to 1).
The modified decoding process for marking inter-view reference images is as follows:
8.2.5.5.2 Image marking
This process is invoked when view_memory_management_control_operation equals 1.
Specify viewIDX as follows, if (view_memory_management_control_operation == l) viewIDX = CurrViewId - (difference_of_view_id_minusl + 1) else if (view_memory_management_control_operation == 2) viewIDX = CurrViewId +__difference_of_view_id_)
To allow scalability, that is, the ability to choose which views are transmitted, sent, or decoded, memory management control operations can be restricted as follows. If currTemporalLevel equals temporal_level of the current image and dependentViews is a set of views that depend on the current view, an MMCO command can only target an image that has a temporal_level equal to or greater than currTemporalLevel and is within dependentViews. To allow this, MMCO commands are appended with an indication of view_id or new MMCO commands are specified with an indication of view_id.
In order to solve these problems related to the process of building reference image lists
ΛΑ 'lV «
INSfí-IJTl t> £ ú?
ii ·. '; · described above, the FrameNum and FrameNumWrap variables are not redefined / scaled. This is the same action that occurs in the AVC standard and contrasts with the JMVM standard, where the variables are redefined / rescaled. The JMVM 1.0 modification is as shown below, with changes shown in italics:
In 8.2.4.3.1 the reordering process of reference image lists for short-term reference images, 8-38 shall be changed as:
for (cldx = num ref idx IX active minusl + 1; cldx> refIdxLX; cldx—)
RefPicListX [cldx] = RefPicListX [cldx - 1]
RefPicListX [refIdxLX ++] = short-term reference image with PicNum equal to picNumLX and view_id equal to CurrViewID nldx = refldxLX for (cldx = refldxLX; cldx <= num_ref_idx_lX_active_minusl + 1;
cldx ++) (8-38) // if (PicNumF (RefPicListX [cldx])! = pícNumLX) if (PicNumF (RefPicListX [cldx])! = picNumLX II ViewID (RefPicListX [cldx]! = CurrViewID)
RefPicListX [nldx ++] = RefPicListX [cldx]
Where CurrViewID is the view_id of the current decoding image.
Regarding the issues associated with the reference image list initialization process discussed above, these issues can be solved by appreciating that aaa-anag -
X v ΰ ¡TUT
IWS 'only tables, fields, or pairs of fields belonging to view as the current slice can be considered in the initialization process. In terms of JMVM 1.0, this language can be added to the beginning of each of subclasses 8.2.4.2.1 Initialization process for the reference image list for P and SP slices in tables through 8.2.4.2.5 Process of initialization for reference image lists in the fields.
to the same
With respect to other issues related to the reference image list building process, various methods can be used to efficiently rearrange both interviewed images and images used for intra-prediction. A first method of such methods involves listing interview reference images in front of intraview reference images, as well as specifying separate RPLR processes for interview images and images for intraview prediction. Images used for intra-view prediction are also called intra-view images. In this method, the reference image list initialization process for intra-view images as specified above is performed, followed by the RPLR reordering process and the list truncation process for intra-view images. Next, the interviewed images are appended to the list after the intraviewed images. Finally, each interviewed image can be further selected and placed in a specific entry in the list of reference images using the following <sup>INST</sup>ITUT ° .MEXICAN OF INDUSTRIAL PROPERTY
<img file="MX337935B_D0045.tif" />
syntax, semantics and decoding process, modified "of JMVM 1.0. The method is applicable to both refPicListO and refPicListl, if present.
<td>ref foot list reordering () {</td><td>C</td><td>Descriptor</td>
<td>if (slice type! = 1 && slice type! = YES) {</td><td></td><td></td>
<td> ...</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td>if (svc mvc flag)</td><td></td><td></td>
<td> {</td><td></td><td></td>
<td>view ref pie list reordering flag 10</td><td> 2</td><td>u (l)</td>
<td>if (view ref foot list reordering flag 10)</td><td></td><td></td>
<td>do {</td><td></td><td></td>
<td>view reordering gone</td><td> 2</td><td>ue (v)</td>
<td>if (view reordering idc == 0 □ view reordering idc == l)</td><td></td><td></td>
<td>abs diff view idx minusl</td><td> 2</td><td>ue (v)</td>
<td>ref idx</td><td> 2</td><td>ue (v)</td>
<td>} while (view reordering idc! = 2)</td><td></td><td></td>
<td>view ref foot list reordering flag 11</td><td> 2</td><td>u (l)</td>
<td>if (view ref pic list reordering flag 11)</td><td></td><td></td>
<td>do {</td><td></td><td></td>
<td>view reordering ide</td><td> 2</td><td>ue (v)</td>
<td>if (view reordering idc == 0 □ view reordering idc == l)</td><td></td><td></td>
<td>abs diff view idx minusl</td><td> 2</td><td>ue (v)</td>
<td>ref idx</td><td> 2</td><td>ue (v)</td>
<td>} while (view reordering idc! = 2)</td><td></td><td></td>
<td> }</td><td></td><td></td>
Regarding syntax, a view_ref_pic_list_reordering_flag_lX (X is 0 or 1) equal to 1 specifies that the view_reordering_idc syntax element is present for refPicListX. A view_ref_pic_list_reordering_flag_lX equal to 0 specifies that the view_reordering__idc syntax element is not present for refPicListX. id_ref indicates the entry that the interviewed image will place in the list of reference images.
abs_diff_view_idx_minusl plus 1 specifies the
7
..... X <sup>¡!</sup> wsr, Ϊ ex 7 '7 <í .___ r-δ <sup>i</sup>'<sup>t tA</sup>, ™ P¡'IEI? /. L>. .
industrial absolute difference between the view index of the image to be placed at the entry of the reference image list indicated by ref idx and the prediction value of the view index, abs_diff_view_idx_minusl is in the range 0 to num_multiview_refs_for_listX [view_id ]-one. The num_multiview_refs_for_listX [] refers to anchor_reference_view_for_list_X [curr_view_id] [] for an anchor image and non_anchor_reference_view_for_list_X [curr_view__id] [] for a curr_view_id equal to the view that contains the slice. An index view of an inter-views image indicates the order of view_id of the inter-views image that appears in the MVC SPS extension. For an image with a view index equal to view_index, view_id is equal to num_multiview_refs_for__listX [view_index].
abs_diff_view_idx_minusl plus 1 specifies the absolute difference between the view index of the image being moved to the current index in the list and the prediction value of list indexes. abs_diff_view_idx_minusl is in the range 0 to num_multiview_refs_for_listX [view_id] -1. The num_multiview_refs_forlistX [] refers to anchor_reference_view_for_list_X [curr_view_id] [] for an anchor image and non_anchor_reference_view_for_list_X [curr_view_id] [] for an
one vi P i ϊ - .o Mexican Institute ÁtÁC, £ CtA PKomPAD V *. f & Wj image not anchor, where curr_view_id equals vieWMfd '<sup>TO</sup>give '<lS? j view containing the current slice. An index cte '~ vi »st * & ^ d £ _ an inter-views image indicates the order of view_id of the inter-views image that appears in the MVC SPS extension. For an image with a view index equal to view_index, view_id is equal to num_multiview_refs_forlistX [view_index].
The decoding process is as follows:
The definition of NumRefIdxLXActive is made after truncation for intra-view images:
NumRefIdxLXActive = num_ref_idx_lX_active_minusl + 1 + num multiview refs for listX [view id]
G.8.2.4.3.3 Reordering process of reference image lists for interviewed images
The inputs to this process are the list of RefPicListX reference images (X being 0 or 1). The outputs of this process are a list of possibly modified reference images RefPicListX (where X 0 or 1).
The picViewIdxLX variable is derived as follows.
If view_reordering_idc equals 0 picViewIdxLX = picViewIdxLXPred (abs_diff_view_idx_minusl + l)
Otherwise (view_reordering_idc equals 1), picViewIdxLX = picViewIdxLXPred + (abs_diff_view_idx_minusl + l) picViewIdxLXPred is the prediction value for the
I
INST
Gave
<img file="MX337935B_D0046.tif" />
picViewIdxLX variable. When the process specified in this subclause is first invoked for a slice (that is, for the first occurrence of view_reordering_idc equal to 0 or 1 in the ref_pic_list_reordering () syntax), picViewIdxLOPred and picViewIdxLlPred are initially set equal to 0. After of each picViewIdxLX mapping, the value of picViewIdxLX is mapped to picViewidxLXPred.
The following procedure is carried out to place the inter-view image with the view index equal to picViewIdxLX in the index position, ref_Idx changes the position of any other remaining image for later in the list, as follows.
for (cldx = NumRefIdxLXActive; cldx> ref_Idx; cldx--)
RefPicListX [cldx] = RefPicListX [cldx - 1]
RefPicListX [ref_Idx] = cross-reference reference image with view id equal to reference_view_for_list_X [picViewIdxLX] nldx = ref_Idx + l;
for (cldx = refldxLX; cldx <= NumRefIdxLXActive; cldx ++) if (ViewID (RefPicListX [cldx])! = TargetViewID || Time (RefP icListX [cldx])! = TargetTime)
RefPicListX [nldx ++] = RefPicListX [cldx] preView_id = PicViewIDLX
TargetViewID and TargetTime indicate the view_id or time axis value of the target reference image to be reordered, and Time (foot) returns the time axis value of
<img file="MX337935B_D0047.tif" />
I!
the foot image.
INSTITUTE V ''>
Say: The mu i ?, industrial
<img file="MX337935B_D0048.tif" />
In accordance with a second method to efficiently reorient both interview images and images used for intra-prediction, the initialization process of reference image lists for intra-view images is performed as specified above, and the inter- views are appended to the end of the list in the order in which they appear in the MVC SPS extension. Next, an RPLR reordering process is applied for both intra-view and inter-view images, followed by a list truncation process. The syntax, semantics and process, as a modified JMVM-based decoding sample, are as follows:
Reordering syntax of the reference image list.
<td>ref foot list reordering () {</td><td>C '</td><td>Dfeisgariptor and</td>
<td>if (slice type! = I && slice type! =) {</td><td></td><td></td>
<td>ref pie list reordering flag 10</td><td> 2</td><td>u (l)</td>
<td>if (ref foot list reordering flag 10)</td><td></td><td></td>
<td>do {</td><td></td><td></td>
<td>reordering of pic nums gone</td><td> 2</td><td>ue (v)</td>
<td>if (reordering of pie nums idc == 0 II reordering of pie nums ide —1)</td><td></td><td></td>
<td>abs diff foot num minusl</td><td> 2</td><td>ue (v)</td>
<td>else if (reordering of pie nums ide == 2)</td><td></td><td></td>
<td>long had tweet num</td><td> 2</td><td>ue (v)</td>
<td>if (reordering of pie nums idc == 4 II reordering of pie nums ide == 5)</td><td></td><td></td>
<td>abs diff view idx minusl</td><td> 2</td><td>ue (v)</td>
<td>} while (reordering of pie nums idc! = 3)</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td>if (slice type == B II slice type == EB) {</td><td></td><td></td>
<td>ref pic list reordering flag 11</td><td> 2</td><td>u (1)</td>
<td>if (ref pie list reordering flag 11)</td><td></td><td></td>
<td>do {</td><td></td><td></td>
<td>reordering of foot nums gone</td><td> 2</td><td>ue (v)</td>
<td>if (reordering of pie nums idc == 0 II reordering of pie nums idc == l)</td><td></td><td></td>
<td>abs diff pic num minusl</td><td> 2</td><td>ue (v)</td>
<td>else if (reordering of pie nums idc == 2)</td><td></td><td></td>
<td>long term pic num</td><td> 2</td><td>ue (v)</td>
<td>if (reordering of pie nums idc == 4 II reordering of pie nums idc == 5)</td><td></td><td></td>
<td>abs diff view idx minusl</td><td> 2</td><td>ue (v)</td>
<td>} while (reordering of pie ide! = 3)</td><td></td><td></td>
<td> }</td><td></td><td></td>
<td> }</td><td></td><td></td>
G. 7.4.3.1 Reordering semantics of reference image lists
<img file="MX337935B_D0049.tif" />
INSTiT U Lt:.
Table
Reordering_of_pic_nums_idc operations to reorder reference images.
<img file="MX337935B_D0050.tif" />
lists' of
<td>reordering of pie nums ide</td><td>Rearrangement specified</td>
<td> 0</td><td>abs diff foot minusl num is present and corresponds to a difference to subtract from an image number prediction value</td>
<td> 1</td><td>abs diff foot minusl num is present and corresponds to a difference to add to a picture number prediction value</td>
<td> 2</td><td>long term foot num is present and specifies the number of long-term images for a reference image</td>
<td> 3</td><td>End of cycle to reorder initial reference image list</td>
<td> 4</td><td>abs diff view idx minusl is present and corresponds to a difference to subtract from a view index prediction value</td>
<td> 5</td><td>abs diff view idx minusl is present and corresponds to a difference to add to a view index prediction value</td>
The reordering_of_pic_nums_idc, along with abs_diff_pic_num_minusl or long_term_pic_num, specifies which reference images are re-rendered. The reordering_of_pic_nums_idc, along with abs_diff_view_idx_minusl, specifies which interviewed reference images are re-rendered. The reordering_of_pic_nums_idc values are specified in the table above. The value of the first reordering_of_pic_nums_idc that follows immediately after ref_pic_list_reordering_flag_10 or ref pic_list reordering_flag 11 is not equal to 3.
The abs_diff_view_idx_minusl plus 1 specifies the
<img file="MX337935B_D0051.tif" />
iNsrn
r.
<sub>z</sub> ,, L'DJUSÍP.IAL absolute difference between the · index of views of the image for
<img file="MX337935B_D0052.tif" />
place in the current index in the list of ^ images ”give reference and the prediction value of the index of views. abs_diff_view_idxminusl is in the range 0 to num_multiview_ref s_for__listX [view_id] -1.
num multiview reís for listX [] refers to anchor_reference_view_for_list_X [curr_view_id] [] for an anchor image non_anchor_reference_view_for_list_X [curr_view_id] [] for a non-anchor image, where curr_view_id equals the view that contains the current slice. An index view of an interview image indicates the order of the view_id of the interview image that appears in the MVC SPS extension. For an image with a view index equal to view_index, the view_id is equal to num_muitiview_refs_for_listX [view_index].
The reordering process can be described as follows:
G. 8.2.4.3.3 Reordering process of reference image lists for interviewed images.
The input to this process is an index refldxLX (where X 0 or
1) ·
The output of this process is an incremented index refldxLX.
The picViewIdxLX variable is derived as follows.
If reordering_of_pic_nums_idc equals 4 picViewIdxLX = picViewIdxLX Pred- (abs_diff_view_idx_minusl · +1)
Otherwise (reordering_of_pic_nums_idc equals 5),
<img file="MX337935B_D0053.tif" />
picViewIdxLX = picViewIdxLXPred + (abs_diff_view_idx_minusl + 1) picViewIdxLXPred is the prediction value for the picViewIdxLX variable. When the process specified in this subclause is invoked the first time for a slice (that is, for the first occurrence of reordering_of_pic_nums_idc equal to 4 or 5 in the syntax ref_pic_list_reordering ()), picViewIdxLOPred and picViewIdxLlPred are initially set equal to 0. After each picViewIdxLX assignment, the value of picViewIdxLX is assigned to picViewIdxLXPred.
The following procedure is performed to place the inter-view image with a view index equal to picViewIdxLX within the position of the index refldxLX, change the position of any other remaining image for later in the list, and increase the value of refldxLX .
for (cIdx = num_ref_idx_lX_active_minusl + 1; cldx> refldxLX; cldx--)
RefPicListX [cldx] = RefPicListX [cldx-l]
RefPicListX [refIdxLX ++] = inter-view reference image with view id equal to reference_view_for_list_X [picViewIdxLX] nldx = refldxLX for (cldx = refldxLX; cldx <= num_ref_idx_lX_active_minusl + 1;
cldx ++) if (ViewID (RefPicListX [cldx])! = TargetViewIEdTime (RefPicListX [cldx])! = TargetTime)
RefPicListX [nldx ++] = RefPicListX [cldx]
Where TargetViewID and TargetTime indicate the view_id or time axis value of the target reference image going to ·: · / ΛΛν nr iNSTil ÍJ'¡o ·, <DE i .. · · reorder, and Time (foot) returns the time axis value of the footer image.
In accordance with a third method to efficiently rearrange both interview images and images used for intra-prediction, the initial reference image list contains images marked as used as a short-term reference or used as a long-term reference and having the same view_id as current image. Also, the initial reference image list contains images that can be used for interview prediction. The images used for inter-view prediction are concluded from the extension of the set of sequence parameters can also conclude from The images for inter-view prediction are assigned certain long-term reference indices for the decoding process of this image. The long-term reference indices assigned for inter-view reference images can, for example, be the first N reference indices, and the indices for long-term intra-view images can be modified to be equal to their previous value + N for the decoding process of this image, where N represents the number of inter-viewed reference images. Alternatively, the indicated long-term benchmarks may be in the range of MaxLongTermFrameldx + 1 to MaxLongTermFrameldx + N, inclusive. As an alternative, the set extension for MVC and inter view reference flag
<img file="MX337935B_D0054.tif" />
Sequence parameters for MVC can contain ~ .xui, & syntax element named here as start_lt_index_for_rplr, and assigned long-term indices map the range start_lt_index_for_rplr, inclusive, to start_lt_index_for_rplr + N, exclusive. The long-term indices available for interview reference images can be assigned in the order of view_id, in the order of cameras, or in the order in which view dependencies are listed in the sequence parameter set extension for MVC. RPLR commands (syntax and semantics) remain unchanged compared to the H.264 / AVC standard.
For direct relationship temporal processing, for example for motion vector scaling, if both reference images are interprediction images (intra-view prediction) (that is, the reference images are not marked as used for inter-view reference ), then the AVC decoding process is followed. If one of the two reference images is an inter-prediction image and the other is an inter-view prediction image, the inter-view prediction image is treated as a long-term reference image. Otherwise (if both reference images are inter-view images), the camera view indicator or view_id values are used in place of the POC values for motion vector scaling.
For the derivation of the prediction weights for implicit weighted prediction, the following is performed
ΙΝ5ΠΤΗΪΟ
OF THE ;
ll '(> ··<sup>}</sup> process. If both reference images are 'interprediction' (inter-view prediction) images (i.e. 'no e'S't'arp' marked as used for inter-view reference), the AVC decoding process is followed. If one of the two reference images is an inter-prediction image and the other is an inter-view prediction image, then the inter-view prediction image is treated as a long-term reference image. Otherwise (that is, both images are inter-view prediction images), the values of view_id or camera order indicators are used in place of the POC values for the derivation of the weighted prediction parameters.
The present invention is described in the general context of method steps, which can be implemented in an embodiment by a program product that includes computer-executable instructions, such as program code, implemented in a computer-readable medium, and executed by computers in network environments. Examples of computer-readable media may include various types of storage media including, but not limited to, electronic device memory units, random access memory (RAM), read-only memory (ROM, compact discs (CDs), digital versatile discs (DVDs), and other internal or external storage devices. Typically, program modules include routines, programs,
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The computer executable instructions, associated data structures, and program modules represent examples of program code for executing steps of the methods described herein. The particular sequence of such executable instructions or associated data structures represent examples of corresponding acts to implement the functions described in such steps.
The software implementations of the present invention could be accomplished with standard programming techniques with rule-based logic and other logic to accomplish the various stages of database search, correlation stages, comparison stages, and decision stages. It will also be appreciated that the words component and module, as used herein and in the claims, are intended to encompass implementations using one or more lines of software code, and / or hardware implementations, and / or equipment for receiving input from manual data.
The foregoing description of embodiments of the present invention have been presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the present invention to the precise manner described, and modifications and variations are possible in light of the above teachings or may be acquired from the practice of the present invention. The modalities were
<img file="MX337935B_D0055.tif" />
l 1W P í i
INSTITUTO MR c / N-,
OE Ι.Λ the present invention chosen and described by the skilled person several invention to fit the use in order to explain the principles a and their practical application to allow the use of the present modalities and with various modifications in the art particular contemplated.
'W i - ···. ar
INSTITUTE M · '·, t rif ia .....' · INK · '<
Contents24
60 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60
43 members in 14 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 60852223 | United States of America | – | |
| 85222306 | United States of America | P | |
| 85222306 | United States of America | P | |
| 2007054200 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 2007054200 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 60852223 | – | – | – |
| IB0754200 | – | – | – |
| US20060852223P | – | – | – |
| WO2007IB54200 | – | – | – |
Members43
| Document | Office | Kind | |
|---|---|---|---|
| AU2007311476A1 | Australia | A1 | |
| CA2666452A1 | Canada | A1 | |
| CA2858458A1 | Canada | A1 | |
| CA3006093A1 | Canada | A1 | |
| WO2008047303A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2008117985A1 | United States of America | A1 | |
| US2008137742A1 | United States of America | A1 | |
| WO2008047303A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW200829034A | Taiwan Province of China | A | |
| TW200829035A | Taiwan Province of China | A | |
| MX2009003967A | Mexico | A | |
| KR20090079932A | Republic of Korea | A | |
| EP2087741A2 | European Patent Office (EPO) | A2 | |
| CN101548550A | China | A | |
| HK1133761A1 | Hong Kong, China | A1 | |
| EP2087741A4 | European Patent Office (EPO) | A4 | |
| KR20110123291A | Republic of Korea | A | |
| KR101120648B1 | Republic of Korea | B1 | |
| US8165216B2 | United States of America | B2 | |
| AU2007311476B2 | Australia | B2 | |
| AU2007311476C1 | Australia | C1 | |
| US8396121B2 | United States of America | B2 | |
| TWI396451B | Taiwan Province of China | B | |
| EP2642756A2 | European Patent Office (EPO) | A2 | |
| EP2642756A3 | European Patent Office (EPO) | A3 | |
| BRPI0718206A2 | Brazil | A2 | |
| EP2087741B1 | European Patent Office (EPO) | B1 | |
| CN101548550B | China | B | |
| ES2492923T3 | Spain | T3 | |
| CN104093031A | China | A | |
| CA2666452C | Canada | C | |
| TWI488507B | Taiwan Province of China | B | |
| MX337935BThis record | Mexico | B | |
| CN104093031B | China | B | |
| EP2642756B1 | European Patent Office (EPO) | B1 | |
| EP3379834A2 | European Patent Office (EPO) | A2 | |
| EP3379834A3 | European Patent Office (EPO) | A3 | |
| ZA200903322B | South Africa | B | |
| BRPI0718206A8 | Brazil | A8 | |
| ES2702704T3 | Spain | T3 | |
| CA2858458C | Canada | C | |
| PL2642756T3 | Poland | T3 | |
| BRPI0718206B1 | Brazil | B1 |
Numbers
- Publication
- 337935
- Publication, DOCDB
- 337935
- Publication, EPODOC
- MX337935
- Application
- 2014009567
- Application, DOCDB
- 2014009567
- Application, EPODOC
- MX20140009567
Titles
- Spanish
- SISTEMA Y METODO PARA IMPLEMENTAR UNA ADMINISTRACION EFICIENTE DE MEMORIA INTERMEDIA DECODIFICADA EN CODIFICACION DE VIDEO DE VISTAS MULTIPLES.
Classification
- CPC, 9
- H04N19/573
- H04N19/105
- H04N19/70
- H04N19/597
- H04N19/46
- H04N19/423
- H04N19/172
- H04N19/103
- H04N13/00
- IPC, 3
- H04N13 00
- H04N19 103
- H04N21 4545