Image coding method and device for buffer management of decoder, and image decoding method and device.
Abstract
The present invention relates to an image coding method and device for buffer management of a decoder, and an image decoding method and device. The image coding method according to the present invention: determines the maximum size of a buffer required for decoding each image frame in a decoder and the number of image frames requiring realignment, and latency information on an image frame which has the greatest difference between a decoding order and a display order among image frames forming an image sequence, based on the decoding order of each image frame of a coded image sequence, a decoding order of a reference frame that each image frame refers to, the display order of each image frame and a display order of the reference frame; and adds a first syntax representing the maximum size of a buffer, a second syntax representing the number of image frames requiring realignment, and a third syntax representing latency information to a set of essential sequence parameters that is a set of information related to the coding of an image sequence.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
6 claims: 6 independent, 0 dependent
- 1CLAIMS REIVINDICACIONES Habiéndose descrito la invención como antecede, se reclama como propiedad lo contenido en las siguientes reivindicaciones:Having described the invention as above, the content of the following claims is claimed as property: 1. Un aparato de decodificación de una imagen, el cual está caracterizado porque comprende: one. An image decoding apparatus, which is characterized in that it comprises: an image data extractor and encoding information to obtain a first syntax indicating a maximum size of a buffer memory needed to decode an image included in an image sequence, a second syntax indicating the maximum number of images which can precede any first image in the sequence of images in decoding order and follow any first image in exit order, and a third syntax used to obtain latency information indicating the maximum number of images that can precede any second image in the sequence of images in the output order and follow any second image in decoding order, from a stream bit;un extractor de datos de imagen e información de codificación para la obtención de una primera sintaxis que indica un tamaño máximo de una memoria de almacenamiento intermedio necesaria para decodificar una imagen incluida en una secuencia de imágenes, una segunda sintaxis que indica el número máximo de imágenes que pueden preceder a cualquier primera imagen en la secuencia de imágenes en orden de decodificación y sigan la cualquier primera imagen en orden de salida, y una tercera sintaxis usada para obtener información de latencia que indica el número máximo de imágenes que pueden preceder a cualquier segunda imagen en la secuencia de imágenes en el orden de salida y seguir la cualquier segunda imagen en orden de decodificación, a partir de un flujo de bits;a decoder to decode the bit stream;and a buffer memory for storing the decoded image frames, where the buffer memory determines a maximum size of a storage memory un decodificador para decodificar el flujo de bits;y una memoria de almacenamiento intermedio para almacenar los cuadros de imagen decodificados, en donde la memoria de almacenamiento intermedio determina un tamaño máximo de una memoria de almacenamiento IMPI IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL intermedio que almacena imagen decodificada con base en la primera sintaxis y determina si se da salida a la imagen decodificada almacenada en la memoria de almacenamiento intermedio con base en una información de latencia obtenida MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY intermediate that stores decoded image based on the first syntax and determines if the decoded image stored in the buffer memory is output based on obtained latency information
- 25 by adding the second syntax and the third syntax, and where the first syntax, the second syntax, and the third syntax are included in a parameter set, where:5 al agregar la segunda sintaxis y la tercera sintaxis, y en donde la primera sintaxis, la segunda sintaxis y la tercera sintaxis están incluidas en un conjunto de parámetros, en donde: la imagen es dividida en una pluralidad de unidades the image is divided into a plurality of units
- 310 maximum coding unit, according to information about a maximum size of a coding unit, the maximum coding unit is hierarchically divided into one or more depth coding units according to divided information, 10 de codificación máxima, de conformidad con información acerca de un tamaño máximo de una unidad de codificación, la unidad de codificación máxima se divide jerárquicamente en una o más unidades de codificación de profundidades de conformidad con información dividida,
- 415 una unidad de codificación de una profundidad actual es una de unidades de datos rectangulares divididas de una unidad de codificación de una profundidad superior, cuando la información dividida indica una división para la profundidad actual, la unidad de codificación de la fifteen a coding unit of a current depth is one of divided rectangular data units of a coding unit of a higher depth, when the divided information indicates a division for the current depth, the coding unit of the
- 520 profundidad actual se divide en unidades de codificación de una profundidad inferior, independientemente de unidades de codificación vecinas, cuando la información dividida indica una no división de la profundidad inferior, se obtienen una o más unidades de twenty current depth is divided into coding units of a lower depth, independently of neighboring coding units, when the divided information indicates a non-division of the lower depth, one or more units of
- 625 lower depth coding unit prediction. 25 predicción de la unidad de codificación de la profundidad inferior.
Independent claims6
562 paragraphs in 107 sections, as filed
(54) Title: IMAGE CODING METHOD AND DEVICE FOR HANDLING OF DECODER BUFFER, AND IMAGE DECODING METHOD AND DEVICE.
(54) Title: IMAGE CODING METHOD AND DEVICE FOR BUFFER MANAGEMENT OF DECODER, AND IMAGE DECODING METHOD AND DEVICE.
(57) Summary
The present invention relates to an image encoding method and device for buffer handling of a decoder, and an image decoding method and device. The image coding method according to the present invention. Determines the maximum size of a buffer required to decode each image frame in a decoder and the number of image frames that require realignment, and latency information about an image frame that has the largest difference between a decoding order and a order of presentation between image frames that form an image sequence, based on the decoding order of each image frame.
(57) Abstract
The present invention relates to an image coding method and device for buffer management of a decoder, and an image decoding method and device. The image coding method according to the present invention: determine the maximum size of a buffer required for decoding each image frame in a decoder and the number of image trames requiring realignment, and latency Information on an image frame which has the greatest difference between a decoding order and a display order among image trames forming an image sequence, based on the decoding order of each image frame of a coded image sequence, a decoding order of a reference frame that each image frame refers to, the display order of each image frame and a display order of the reference frame; and adds a first syntax representing the maximum size of a buffer, a second syntax representing the number of image trames requiring realignment, and a third syntax representing latency Information to a set of essential sequence parameters that is a set of Information related to the coding of an image sequence.
Institute
Mexican Property
Industrial
<img file="MX337918B_D0001.tif" />
PATENT TITLE NO. 337918
Owner (s): SAMSUNG ELECTRONICS CO .. LTD.
Address: 129, Samsung-ro, Yeongtong-gu, Suwon-s¡, Gyeonggi-do, 443-742, REPUBLICA
FROM KOREA
Name: IMAGE CODING METHOD AND DEVICE FOR HANDLING OF DECODER BUFFER, AND IMAGE DECODING METHOD AND DEVICE.
Classification: lnt.CI.8: H04N19 / 10; H04N19 / 61 Inventor (s): YOUNG-O PARK; CHAN-YUL KIM: KWANG-PYO CHOI. JEONG-HOON PARK
REQUEST
Number: International filing date:
MX / a / 2015/006962 November 23, 2012
Divisional Patent Number: 330592 $ • í 'i
I; PRIORITY 1 · 'i
Date:
Country:
US
KR November 2011 April 2, 2012
Number:,
61/563,678 ] 10-2012-0034095
Validity: Twenty years
Expiration Date: November 23, 2032
The reference patent is granted based on articles 1 * 2nd fraction V 6th fraction III. and 59 of the Industrial Property Law
In accordance with article 23 di counted from the date of pi f rights. s »
X lAa il W
Industrial Property Law the present patent will be in force for twenty years, mp | fiTogables, request for interrogation of the;
subject to the rate to keep the ace people based on the provisions of articles 6 fractions lll and 7 ° bis 2 of
05/06 / 2009,06 / 01/2010, 06/18/2010, 06/28/2010, 01/27/2012 and 04/09/2012); articles 1, 3, fraction V reformed the i Who subscribes the present title
PñSfdétfeñflrtdusíffÉiFfDfSf¡o Oficial 26/01/2004, 16/06/2005, 25/01/2006 subsection a), 4th and 12th sections I and lll of the Regulation of the Mexican Institute of Industrial Property (DOF 14/12/1999
Law of 05/07/1999,
07/01/2002, 07/15/2004, 07/28/2004 and 09/07/2007); Articles 1, 3, 4, 5, fraction V, subsection a), 16 sections I and lll and 30 of the Organic Statute of the Mexican Institute of Industrial Property (DOF 12/27/1999, amended 10/10/02, 07/29/2004, 08/04/2004 and 09/13/2007); 1st. 3rd and 5th Subsection a) of the Agreement that delegates powers to the Deputy Directors General, Coordinator, Divisional Directors, Heads of Regional Offices, Divisional Deputy Directors, Departmental Coordinators, and. other subordinates of the Mexican Institute of Industrial Property. (DOF 12/15/1999, amended on 02/04/2000, 07/29/2004, 08/04/2004 / 09/13/2007);
<img file="MX337918B_D0002.tif" />
Arenai No. 550. Floor 1.
: oi. Puebio Sat María Tepepan, Xcchimhco. CP 16020,
Mexico City eí. (55) 53 34 07 00 mvsv impijzob.mx
Issue Date: March 28, 2016
DIVISIONAL DIRECTOR OF PATENTS
N AH ANN AND CANAL REYES
<img file="MX337918B_D0003.tif" />
MX / 2016/23302
<img file="MX337918B_D0004.tif" />
30t5 / 616 <3
IMAGE CODING METHOD AND DEVICE FOR HANDLING OF DECODER BUFFER, AND IMAGE DECODING METHOD AND DEVICE
FIELD OF THE INVENTION
IΜ ΡI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0005.tif" />
The present invention relates to methods and apparatus for encoding and decoding an image, and more particularly, methods and apparatus for efficiently encoding and decoding information to control and manage a decoded illustration buffer (DPB) that stores a decoded illustration.
BACKGROUND OF THE INVENTION
In a video codec, such as ITU-T H.261,
ISO / IEC MPEG-1 visual, ITU-T H.262 (ISO / IEC MPEG-2 visual), ITU-T H.264, ISO / IEC MPEG-4 visual, or ITU-T H.264 (ISO / IEC MPEG-4 AVC), a macroblock is predictively encoded through inter-prediction or intra-prediction, and a bit stream is generated from image data encoded in accordance with a predetermined format defined by each video codec and sent .
BRIEF DESCRIPTION OF THE INVENTION
Technical problem
The present invention provides a method and apparatus for encoding an image, where information to control and handle a decoder buffer is encoded.
<img file="MX337918B_D0006.tif" />
I.IEXiCAN'O INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0007.tif" />
efficiently, and a method and apparatus for decoding an image, wherein a buffer is handled efficiently by using information to control and handle the buffer.
Technical Solution
In accordance with one aspect of the present invention, buffer size information, which is necessary to decode illustrations included in video stream, is mandatorily included in bitstream and transmitted, and a decoder may decode illustration by assigning a buffer size. necessary based on the information.
Also, in accordance with one aspect of the present invention, information used to determine when to send artwork stored in the buffer is mandatorily included in the bitstream and transmitted.
Advantageous Effects
In accordance with one or more embodiments of the present invention, system resources of a decoder can be prevented from being wasted because the buffer size information required to decode illustrations included in an image sequence is compulsorily aggregated to and transmitted with a stream of bits, and the decoder uses the buffer size information to perform decoding by allocating a buffer size as required. Also, in accordance with one or more modalities of
ΓΓΚ.Τ time determining information of the present invention,
MEXICAN INSTITUTE OF THE PROFltDAQ
INDUSTRIAL
<img file="MX337918B_D0008.tif" />
The output of a buffered artwork is mandatorily added to and transmitted with a bit stream, and a decoder can preset whether to send a precoded picture frame by using the information to determine an output time for a buffered artwork to thereby preventing an output latency of a decoded picture frame.
BRIEF DESCRIPTION OF
FIGURES
Figure 1 is a block diagram of a video encoding apparatus in accordance with an embodiment of the present invention;
Figure 2 is a block diagram of a video decoding apparatus in accordance with an embodiment of the present invention;
Figure 3 illustrates a concept of encoding units in accordance with an embodiment of the present invention;
Figure 4 is a block diagram of an image encoder based on encoding units, in accordance with an embodiment of the present invention;
Figures 5 is a block diagram of an image decoder based on encoding units, in accordance with an embodiment of the present invention;
Figure 6 is a diagram illustrating units of
<img file="MX337918B_D0009.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0010.tif" />
coding corresponding to depths, and divisions, in accordance with an embodiment of the present invention;
Figure 7 is a diagram illustrating a relationship between a coding unit and transformation units, in accordance with an embodiment of the present invention;
Figure 8 is a diagram illustrating coding information corresponding to depths, in accordance with an embodiment of the present invention;
Figure 9 is a diagram illustrating coding units corresponding to depths, in accordance with an embodiment of the present invention;
Figures 10, 11, and 12 are diagrams illustrating a relationship between encoding units, prediction units, transformation units, in accordance with an embodiment of the present invention;
Figure 13 is a diagram illustrating a relationship between a coding unit, a prediction unit, and a transformation unit, in accordance with coding mode information in Table 1;
Figure 14 is a diagram of an image encoding procedure and an image decoding procedure, which are hierarchically classified, in accordance with an embodiment of the present invention;
Figure 15 is a diagram of a structure of
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0011.tif" />
a network abstraction layer unit (NAL), in accordance with one embodiment of the present invention;
Figures 16A and 16B are reference diagrams for describing maximum size information of a decoded illustration buffer required in accordance with a decoding command during an image sequence encoding procedure;
Figure 17 is a diagram illustrating a procedure for sending a decoded illustration from
DPB, in accordance with a damping procedure in a video code field related to the present invention;
Figure 18 is a diagram for describing a procedure for sending a coded illustration from a DPB using a MaxLatencyFrames syntax, in accordance with an embodiment of the present invention;
Figures 19A through 19D are diagrams for describing a MaxLatencyFrames syntax and a num_reorder_frames syntax, in accordance with embodiments of the present invention;
Figure 2 0 is a flow chart illustrating an image encoding method in accordance with an embodiment of the present invention; and Figure 21 is a flow chart illustrating a
IMPI <Ύ! ΤυΤΟ MEXICANO DE IA PROPERTY INDU5T1IAL
<img file="MX337918B_D0012.tif" />
image decoding method in accordance with an embodiment of the present invention.
DETAILED DESCRIPTION OF THE INVENTION
In accordance with an aspect of the present invention, a method for encoding an image is provided, the method comprises: determining reference frames respectively of image frames that form an image sequence by performing motion prediction and compensation, and encoding the frames image when using certain reference frames; determining a maximum size of a buffer required to decode the image frames by a decoder and the number of image frame required to be reordered, based on an order of encoding of the image frames, an order of encoding of the image frames reference indicated by the image boxes, a display order of the image boxes, and a display order of the reference boxes; determine latency information of an image frame that has a larger difference between an encoding order and a display order, from among the image frames that make up the image sequence, based on the number of image frames required to be rearranged; adding a first syntax indicating the maximum buffer size, a second syntax indicating the number of image frames required to be sorted, and
IMPI
MEXICAN INSTITUTE OR THE INDUSTRIAL PROPERTY
<img file="MX337918B_D0013.tif" />
a third syntax indicating latency information to a mandatory sequence parameter group (SPS) which is a group of information related to encoding the image sequence.
In accordance with another aspect of the present invention, an apparatus for encoding an image is provided, the apparatus comprising: an encoder for determining respectively reference frames of image frames that form an image sequence by performing motion prediction and compensation, and encode the image frames by using the given reference frames; and an output unit for determining a maximum size of a buffer required to encode the image frames by a decoder and the number of image frames required to reorder, based on an order of encoding the image frames, an order encoding of the reference frames indicated by the image frames, an order of presentation in the image frames, and a display order of the reference frames, determine latency information of an image frame that has a larger difference between an encoding order and a display order, from among the image frames that make up the image sequence, based on the number of image frames required to be reordered, and generate a bit stream by adding a first syntax that
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0014.tif" />
indicates the maximum buffer size, a second syntax indicating the number of image frames required to be reordered, and a third syntax indicating latency information to a mandatory sequence parameter group which is a group of information that refers to encoding the image sequence.
In accordance with another aspect of the present invention, a method for decoding an image is provided, the method comprising: obtain a first syntax indicating a maximum size of a buffer required to decode each of the image frames that make up an image sequence, a second syntax indicating the number of image frames displayed after a post-decoded image frame and required to be reordered, and a third syntax indicating latency information of an image frame having a larger difference between a decoding order and a display order among the image frames that make up the image sequence, from a bit stream; set a maximum size of a buffer required to decode the image sequence by a decoder, using the first syntax; obtaining encoded data, where the image frames are encoded, from the bit stream, and obtaining decoded image frames by decoding the obtained encoded data; store the decoded image boxes in
IMPI
<img file="MX337918B_D0015.tif" />
<sup>, NST</sup>Í¡]<sup>jT</sup>O MÍXICANO DI THE PROPERTY
INDUSTRIAL____ a decoder buffer; and determine if an Image box stored in the decoder buffer is sent, using the second syntax and the third syntax, where the first syntax, the second syntax, and the third syntax are included in a mandatory sequence parameter group which is a group of information related to encoding the image sequence.
In accordance with another aspect of the present invention, an apparatus for decoding an image is provided, the apparatus comprising: image data and encoding information extractor to obtain a first syntax indicating a maximum size of a buffer required to decode each of the image frames forming an image sequence, a second syntax indicating the number of image frames presented after a post-decoded image box and required to be rearranged, a third syntax indicating latency information of an image frame having a larger difference between a decoding order and a display order of between the image frames that make up the image sequence, and encoded data where the picture frames, from a bit stream; a decoder for obtaining decoded image frames by decoding the encoded data obtained; and a buffer to store the decoded image frames, where the buffer sets ίο
ΙΜΡΪ
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL the maximum buffer size required to decode the image sequence when using the first syntax, and determines whether to send a stored image frame using the second and third syntax, and the first syntax, the second syntax, and the third syntax are included in a mandatory sequence parameter group which is a group of information related to encoding the image sequence.
Mode for the Invention
Hereinafter, illustrative embodiments of the present invention will be described in detail with reference to accompanying figures. Although the present invention is described, an image may be a still image or a moving image, and may be denoted as a video. Also, while describing the present invention, an image box may be denoted as an illustration.
FIG. 1 is a block diagram of a video encoding apparatus 100 in accordance with an embodiment of the present invention.
The video encoding apparatus 100 includes a maximum encoding unit splitter 110, an encoding unit determiner 120, and an output unit 130.
The maximum encoding unit divider 110 can divide a current illustration of an image based on a maximum encoding unit for the illustration.
<img file="MX337918B_D0016.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL MONEDAD
<img file="MX337918B_D0017.tif" />
current. If the current artwork is larger than the maximum encoding unit, image data from the current artwork can be divided into at least one maximum encoding unit. The maximum encoding unit in accordance with an embodiment of the present invention may be a data unit having a size of 32x32, 64x64, 128x128,
256x256, etc., where a shape of the data unit is a square having a width and length in squares of 2 that are greater than 8. Image data may be sent to the encoding unit determiner 120 in accordance with the at least one maximum encoding unit.
A coding unit in accordance with an embodiment of the present invention can be characterized by a maximum size and depth. Depth denotes a number of times the coding unit is spatially divided from the maximum coding unit, and as depth intensifies, coding units corresponding to depths can be divided from the maximum coding unit to a minimum coding unit . A maximum coding unit depth can be determined as a higher depth, and the minimum coding unit can be determined as a lower coding unit. Since a size of a coding unit corresponding to each depth decreases as the depth of the coding unit intensifies
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0018.tif" />
For maximum coding, a coding unit corresponding to a greater depth may include a plurality of coding units corresponding to a lower depth.
As described above, the image data in the current illustration is divided into the maximum encoding units according to a maximum encoding unit size, and each of the maximum encoding units may include encoding units that are divided in accordance with depths. Since the maximum encoding unit according to an embodiment of the present invention is divided according to depths, the image data of a spatial domain included in the maximum encoding unit can be hierarchically classified according to the depths.
A maximum depth and maximum size of one encoding unit, which limit the total number of times a height and width of the maximum encoding unit are hierarchically divided, can be predetermined.
The encoding unit determiner 120 encodes at least a divided region obtained by dividing a region of the maximum encoding unit in accordance with depths, and determines a depth to send finally encoded image data to the at least one region
<img file="MX337918B_D0019.tif" />
INSTITUTO MSXICANO DE LA PRCfISDAD
INDUSTRIAL
<img file="MX337918B_D0020.tif" />
divided. In other words, the encoding unit determiner 120 determines an encoded depth by encoding the image data in the encoding units corresponding to depths in units of the maximum encoding units in the current illustration, and by selecting a depth having the minimal coding error. The determined coded depth and
<td>the data</td><td>of image</td><td>in each of</td><td>the</td><td>units</td><td>of</td>
<td>coding</td><td>maximum is</td><td>send to the unit</td><td colspan="2">exit 130.</td><td></td>
<td colspan="2">The data from</td><td>image in each</td><td>of the</td><td>units</td><td>of</td>
<td>coding</td><td>maximum is</td><td>encode based</td><td>in the</td><td>units</td><td>of</td>
encoding corresponding to depths, in accordance with at least a depth equal to or below the maximum depth, and results of encoding the image data based on the encoding units corresponding to depths are compared. A depth having the minimum coding error can be selected after comparing coding errors of the coding units corresponding to depths. At least one encoded depth can be selected for each of the maximum encoding units.
The size of the maximum encoding unit is divided as one encoding unit is hierarchically divided according to depths, and the number of encoding units increases. Also, even if
<img file="MX337918B_D0021.tif" />
<img file="MX337918B_D0022.tif" />
MEXICAN INSTITUTE THE PROPERTY
INDUSTRIAL
<img file="MX337918B_D0023.tif" />
encoding units included in a maximum encoding unit correspond to the same depth, if each of the encoding units will be divided to a lower depth, it is determined by measuring an encoding error of the image data of each of the units coding. Thus, since even data included in a maximum encoding unit has a different encoding error corresponding to a depth, according to the location of the data, an encoded depth can be set differently according to the location of the data. data. Accordingly, at least one encoded depth can be set for a maximum encoding unit, and the image data of the maximum encoding unit can be divided according to encoding units of the at least one encoded depth.
Accordingly, the encoding unit determiner 120 in accordance with an embodiment of the present invention can determine encoding units having a tree structure included in a current maximum encoding unit. Coding units having a tree structure in accordance with an embodiment of the present invention include coding units corresponding to a given depth
ΙΜΡΪ ^
<img file="MX337918B_D0024.tif" />
to be the encoded depth, out of all the encoding units corresponding to depths included in the current maximum encoding unit.
Coding units corresponding to a coded depth can be hierarchically determined according to depths in the same region of the maximum coding unit, and can be determined independently in different regions of the maximum coding unit. Similarly, a coded depth in one current region can be determined independently of a coded depth in another region.
A maximum depth in accordance with an embodiment of the present invention is an index related to the number of times of division from a maximum code unit to a minimum code unit. A first maximum depth in accordance with an embodiment of the present invention may denote the total number of times of division from the maximum encoding unit to the minimum encoding unit. A second maximum depth in accordance with an embodiment of the present invention may denote the total number of depth levels from the maximum coding unit to the minimum coding unit. For example, when a maximum coding unit depth is 0, a coding unit depth obtained by dividing the coding unit
MEXICAN INSTITUTE DF. iJk PROPERTY
INDUSTRIAL
<img file="MX337918B_D0025.tif" />
maximum once can be set to, and a depth of one encoding unit obtained by dividing the maximum encoding unit twice can be set to 2. If a encoding unit obtained by dividing the maximum encoding unit four times is the encoding unit minimum, then there are depth levels of depths 0, 1, 2, 3 and 4. That way, the first maximum depth can be set to 4, and the second maximum depth can be set to 5.
Prediction encoding and transformation can be performed in the maximum encoding unit.
Similarly, prediction coding is performed and transformed into units of maximum coding units, based on coding units corresponding to depths and in accordance with depths equal to or less than the maximum depth.
Since the number of encoding units corresponding to depths increases any time the maximum encoding unit is divided according to depths, encoding including prediction encoding and transformation must be performed on all encoding units corresponding to custom generated depths. that intensifies a depth. For convenience of explanation, prediction coding and transformation will now be described based on a
<img file="MX337918B_D0026.tif" />
<<sup>Γ</sup>· Τ, Ί UTO MEXICANO CE <TO INDUSTRIAL PROPERTY
<img file="MX337918B_D0027.tif" />
coding unit of a current depth, included in at least one maximum coding unit.
Video encoding apparatus 100 can variously select a size or shape of a data unit for encoding image data. In order to encode the image data, operations are performed, such as prediction encoding, transformation, and entropy encoding. At this time, the same data unit can be used for all operations or different data units can be used for each operation.
For example, the video encoding apparatus 100 may select not only an encoding unit to encode the image data, but also a different data unit from the encoding unit to perform predictive encoding on image data in the encoding unit. coding.
In order to prediction encode the maximum encoding unit, prediction encoding can be performed based on an encoding unit corresponding to an encoded depth, i.e. based on an encoding unit that is no longer divided into encoding units corresponding to a shallower depth. Hereinafter, the encoding unit that is no longer divided and becomes a base unit for prediction encoding will now be indicated as a
Γ ivi ρτ
Α
INSTITUTE .W5XÍCANO D £ LA PRC?! TDAD
INDUSTRIAL
<img file="MX337918B_D0028.tif" />
'prediction unit'. Divisions obtained by dividing the prediction unit may include a data unit obtained by dividing at least one of a height and a width of the prediction unit.
For example, when a 2Nx2N encoding unit (where N is a positive integer) is no longer divided, this encoding unit becomes a prediction unit of 2Nx2N, and a size of a division can be 2Nx2N, 2NxN, Nx2N, or NxN. Examples of a division type include symmetric divisions that are obtained by symmetrically dividing a height or width of the prediction unit, divisions obtained by asymmetrically dividing the height or width of the prediction unit, such as l: non: l, divisions that they are obtained by dividing the prediction unit geometrically, and divisions that have arbitrary shapes.
A prediction unit prediction mode can be at least one of an intra mode, an inter mode, and a jump mode. For example, intra-mode or inter-mode can be performed in a division of 2Nx2N, 2NxN, Nx2N, or NxN. Also, jump mode can only be performed in a 2Nx2N split. Coding can be done independently at. one prediction unit in each coding unit, and a prediction mode having a minimum coding error can be selected.
INSTÚ'l.'TO MEXICAN PROPERTY
INDUSTRIAL
<img file="MX337918B_D0029.tif" />
Also, the video encoding apparatus 100 can perform transformation of the image data into an encoding unit based not only on the encoding unit to encode the image data, but also based on a data unit that is different. of the coding junction.
In order to perform transformation in the encoding unit, transformation can be performed based on a data unit that is less than or equal to the size of the encoding unit. For example, a data unit for transformation can include a data unit for intra mode and a data unit for inter mode.
Hereinafter, the data unit that is a transformation base may also be indicated as a transformation unit. Similarly to encoding units having a tree structure in accordance with an embodiment of the present invention, a transformation unit into an encoding unit can be recursively divided into smaller size transformation units. That way, residual data in the encoding unit can be divided according to transformation units that have a tree structure according to transformation depths.
IMPI
MEXICAN INSTITUTE DF. INDUSTRIAL PROPERTY
<img file="MX337918B_D0030.tif" />
A transformation unit in accordance with an embodiment of the present invention can also be assigned with a transformation depth denoting a number of times the height and width of a coding unit is divided to have the transformation unit. For example, a transformation depth can be 0 when a size of a transformation unit for a current 2Nx2N encoding unit is 2Nx2N, a transformation depth can be 1 when a size of a transformation unit for the current 2Nx2N encoding unit is NxN, and a transformation depth can be 2 when a transformation unit size for the current 2Nx2N encoding unit is N / 2xN / 2. That is, transformation units that have a tree structure can also be established according to transformation depths.
Coding information for each coded depth requires not only information about the coded depth, but also information related to prediction and transformation coding. Accordingly, the coding unit determiner 120 can determine not only a coded depth that has minimal coding error, but also determine a type of division in a prediction unit,
<img file="MX337918B_D0031.tif" />
mexican institute OF PROPERTY a prediction mode for each prediction unit, size of a transformation unit for 't'raiibf uiuiacién.
Coding units having a tree structure included in a maximum coding unit and a method for determining a division, in accordance with embodiments of the present invention, will be described in detail below with reference to Figures 3 to 12.
Coding unit determiner 120 can measure coding unit coding errors corresponding to depths using multiplier-based Speed Distortion Optimization
Lagrangians.
The output unit 13 0 outputs the image data of the maximum encoding unit, which is encoded based on the at least one encoded depth determined by the encoding unit determiner 120, and information on the encoding mode of each one of the depths, in a bit stream.
The encoded image data may be a result of residual encoding data from an image.
Information on the encoding mode of each of the depths may include information on the encoded depth, on the type of division in the prediction unit, the prediction mode, and the size of the transformation unit.
IMPI Mexican Institute of INDUSTRIAL PROPERTY
Coded depth information
<img file="MX337918B_D0032.tif" />
it can be defined using information divided according to depths, indicating whether encoding is to be performed in encoding units of a lower depth rather than a current depth. If a current depth of a current encoding unit is the encoded depth, then the current encoding unit is encoded using encoding units corresponding to the current depth, and divided information about the current depth can be defined in such a way that the unit Coding current depth may no longer be divided into coding units of a lower depth. Conversely, if the current depth of the current encoding unit is not the encoded depth, then units of a lesser depth must be encoded, and the divided information about the current depth can thus be defined so that the current encoding unit of
<td>the depth</td><td>current can be divided</td><td>in</td><td>units</td><td>of</td>
<td>coding</td><td>a lower depth.</td><td></td><td></td><td></td>
<td>If the</td><td>current depth no</td><td>is the</td><td colspan="2">depth</td>
<td>encoded, it</td><td>performs encoding in</td><td>the</td><td>units</td><td>of</td>
<td>coding</td><td>the bottom depth.</td><td colspan="2">Since it exists</td><td>to the</td>
minus one coding unit from the bottom depth
<img file="MX337918B_D0033.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0034.tif" />
in a current depth coding unit, coding is done repeatedly in each lower depth coding unit, and coding units having the same depth in that way can be recursively coded.
Since encoding units having a tree structure in a maximum encoding unit are to be determined and information on at least one encoding mode is determined for each encoding unit of an encoded depth, information on at least one encoding mode can be determined for a maximum encoding unit. Also, image data from the maximum encoding unit may have a different encoded depth according to the location thereof since the image data is hierarchically divided according to depths. In this way, information about an encoded depth and an encoding mode can be set for the image data.
Accordingly, the output unit 130 in accordance with an embodiment of the present invention can assign encoding information about a corresponding encoded depth and encoding mode to at least one of encoding units, prediction units, and a minimum unit. included in the maximum encoding unit.
ΙΜ ΡΙ®>
OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0035.tif" />
The minimum unit in accordance with an embodiment of the present invention is a rectangular data unit obtained by dividing a minimum coding unit of a lesser depth by 4, and can be a maximum rectangular data unit that can be included in all units of encoding, prediction units, and transformation units included in the maximum encoding unit.
For example, encoding information sent through output unit 130 may be classified into encoding information from each of the depth encoding units, and encoding information from each of the prediction units. The encoding information of each of the depth encoding units may include prediction mode information and division size information. The encoding information of each of the prediction units may include information on an estimated direction of an intermode, on a reference image index of the intermode, on a motion vector, on a chroma component of the intramode, and on an interpolation method of an intra mode.
Information about a maximum size of encoding units defined in artwork units, snippets, or GOPs, or information about a maximum depth can be inserted into a header of a bitstream.
.vjl χ
MEXICAN INSTITUTE Dt IA PROPERTY
INDUSTRIAL “** -
Maximum encoding unit divisor 110 and encoding unit determiner 120 correspond to a video encoding layer that determines a reference frame of each of the image frames that form an image sequence by performing prediction and compensation of motion according to encoding units with respect to each image frame, and encodes each image frame using the given reference frame.
Also, as will be discussed later, output unit 130 generates a bit stream by indicating a max_dec_frame_buffering syntax indicating a maximum buffer size required to decode an image frame by a decoder, a num_reorder_frames syntax indicating the number of frames image required to be reordered, and a max_latency_increase syntax indicating latency information of a picture frame that has a larger difference between an encoding order and a display order among the picture frames that make up the picture sequence in a network abstraction layer unit .
In the video encoding apparatus 100 in accordance with an embodiment of the present invention, encoding units corresponding to depths may be encoding units obtained by dividing a height or width of an encoding unit and a
<img file="MX337918B_D0036.tif" />
UTO MEXICANO IA INDUSTRIAL PROPERTY
<img file="MX337918B_D0037.tif" />
depth greater than 2. In other words, when the size of a coding unit of a current depth is 2Nx2N, the size of a coding unit of a lower depth is NxN. Also, the 2Nx2N encoding unit can include four NxN encoding units from the shallowest at best.
Accordingly, the video encoding apparatus 100 can form encoding units having a tree structure by determining encoding units having an optimal shape and size for each maximum encoding unit, based on the size of each encoding unit. maximum and a maximum depth determined considering characteristics of a current illustration. Also, since each maximum encoding unit can be encoded in accordance with any of several prediction modes and transformation methods, an optimal encoding mode can be determined considering characteristics of encoding units of various image sizes.
That way, if an image that has very high resolution or a very large amount of data is encoded in conventional macroblock units, it increases one number of macroblocks per image per illustration excessively. In this way, a quantity of compressed information generated for
<img file="MX337918B_D0038.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY each macroblock increases, and thus it is difficult to transmit the compressed information and the compression of data decreases efficiently. However, the video encoding apparatus 100 is capable of controlling an encoding unit based on image characteristics while increasing a maximum size of the encoding unit in consideration of an image size, thereby increasing compression efficiency. of image.
FIG. 2 is a block diagram of a video decoding apparatus 200 in accordance with an embodiment of the present invention.
The video decoding apparatus 200 includes a receiver 210, an image data and encoding information extractor 220, and an image data decoder 230. Definitions of various terms, such as a coding unit, a depth, a prediction unit, a transformation unit, and information on various coding modes, which are used below to explain various procedures of the video decoding apparatus 200, they are identical to that of the video encoding apparatus 100 described above with reference to Figure 1.
Receiver 210 receives and analyzes a bit stream of encoded video. The image data and encoding information extractor 220 extracts image data
IΜ ΡI
MEXICAN INSTITUTE Λ
OF THE PROPERTY \ *
INDUSTRIAL * «encoded for each of the encoding units that have a tree structure in units of maximum encoding units, from the analyzed bitstream, and then sends the extracted image data to the image data decoder 230. The image data and encoding information extractor 220 can extract information about a maximum size of encoding units from a current artwork, from a header to the current artwork.
Also, the image data and encoding information extractor 220 extracts information about an encoded depth and an encoding mode for encoding units having the tree structure in units of the maximum encoding unit, from the bit stream analyzed. The extracted information about the encoded depth and the encoding mode is sent to the image data decoder 230. In other words, the image data in the bit stream can be divided into the maximum encoding units so that the image data decoder 230 can decode the image data in units of the maximum encoding units.
The information about the encoded depth and the encoding mode for each of the maximum encoding units can be set for at least one encoded depth. Information on the mode of
IMPI
<img file="MX337918B_D0039.tif" />
INSTITUTO MiXICAMO Dt THE PROPERTY coding for each coded depth may include information on a type of division 'of a corresponding unit * corresponding to the coded depth, on a prediction mode, and a size of a transformation unit. Also, depth-conforming division information can be extracted as the encoded depth information.
The information about the encoded depth and the encoding mode for each of the maximum encoding units extracted by the image data and encoding information extractor 220 is information about an encoded depth and a certain encoding mode to generate an error. minimum coding when one coding side, for example, the apparatus d video encoding 100, repeatedly encode each of the encoding units corresponding to depths in units of maximum encoding units. Accordingly, the video decoding apparatus 200 can restore an image by decoding the image data in accordance with the encoded depth and the encoding mode that generates the minimum encoding error.
Since encoding information on the encoded depth and encoding mode can be assigned to data units from among encoding units, prediction units, and a minimum unit
<img file="MX337918B_D0040.tif" />
corresponding, the data extractor of '<sup>N</sup>'<sup>J</sup>Image and encoding information ~ 220 can extract the information about the encoded depth and the encoding mode in units of the data units. If the information about the encoded depth and the encoding mode for each of the maximum encoding units is recorded in units of the data units, data units that include information about the same encoded depth and encoding mode can be inferred to be data units included in the same maximum encoding unit.
The image data decoder 230 restores the current artwork by decoding the image data in each of the maximum encoding units, based on the information about the encoded depth and the encoding mode for each of the maximum encoding units. . In other words, the image data decoder 230 can decode the encoded image data based on a division type, prediction mode, and transformation unit analyzed for each of the encoding units having the included tree structure in each of the maximum encoding units. A decoding procedure may include a prediction procedure that includes intra prediction and motion compensation, and an inverse transformation procedure.
I lt
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0041.tif" />
The image data decoder 230 may perform intra prediction or motion compensation in each of the encoding units in accordance with divisions and a prediction mode thereof, based on information on the type of division and mode. of prediction units of prediction of each of the coding units according to coded depths.
Also, in order to perform inverse transformation on each of the maximum encoding units, the image data decoder 230 performs inverse transformation in accordance with the transformation units of each of the encoding units, based on information from transformation unit size of the deepest encoding unit.
The image data decoder 230 can determine an encoded depth of a current maximum encoding unit, based on information divided according to depths. If the divided information indicates that the image data is no longer divided into the current depth, the current depth is a coded depth. In this way, the image data decoder 230 can decode image data of a current maximum encoding unit by using the information about the prediction unit division type, the mode: • 'Uro MEXICANO /! Λ PROPERTY
INDUSTRIAL -prediction, and the size of the transformation unit of a coding unit corresponding to a current depth.
In other words, data units containing encoding information can be collected including the same information divided by looking at encoding information assigned to a data unit between the encoding unit, the prediction unit, and the minimum unit, and the units The collected data may be considered as a data unit to be decoded in accordance with the same encoding mode by the image data decoder 230.
Also, the receiver 210 and the extractor of image data and encoding information 220 can perform a decoding procedure in a NAL, where a max_dec_frame_buffering syntax indicating a maximum size of a buffer required to decode a picture frame by a decoder, a syntax num_reorder_frames indicating the number of image frames required to be reordered, and a max_latency_increase syntax indicating latency information of an image frame that has a larger difference between a decoding order and a display order among image frames that form an image sequence, are derived from a bit stream and they are sent to the image data decoder 230.
The video decoding apparatus 200 can
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL obtain information about a coding unit that generates a minimum coding error when recursively coding each of the maximum coding units, and can use the information to decode the current illustration. In other words, the encoded image data in the encoding units having the determined tree structure can be decoded to be optimal encoding units in units of the maximum encoding units.
Accordingly, even if image data has high resolution and a very large amount of data, image data can be efficiently decoded to be restored by using an encoding unit size and encoding mode, which are adaptively determined from conformance to image data characteristics, based on information about an optimal encoding mode received from one encoding side.
Hereinafter, methods for determining encoding units in accordance with a tree structure, a prediction unit, and a transformation unit, in accordance with embodiments of the present invention, will be described with reference to Figures 3 to 13.
Figure 3 illustrates a concept of encoding units in accordance with an embodiment of the present invention.
<img file="MX337918B_D0042.tif" />
<img file="MX337918B_D0043.tif" />
An encoding unit size can be expressed in width by height, and can be 64x64, 32x32,
16x16, and 8x8. A 64x64 encoding unit can be divided into 64x64, 64x32, 32x64, or 32x32 divisions, and a 32x32 encoding unit can be divided into 32x32, 32x16, 16x32, or 16x16 divisions, a 16x16 encoding unit can be divided into 16x16, 16x8, 8x16, or 8x8 divisions, and an 8x8 encoding unit can be divided into 8x8, 8x4, 4x8, or 4x4 divisions.
In 310 video data, a resolution is 1920x1080, a maximum size of one encoding unit is 64, and a maximum depth is 2. In 320 video data, a resolution is 1920x1080, a maximum size of one encoding unit is 64 , and a maximum depth is 3. In 330 video data, a resolution is 352x288, a maximum size of one encoding unit is 16, and a maximum depth is 1. The maximum depth shown in Figure 3 denotes a total number of divisions from a maximum encoding unit to a minimum decoding unit.
If a resolution is high or a data amount is large, a maximum size of an encoding unit can be relatively large not only to increase encoding efficiency but also to accurately reflect characteristics of an image. Therefore, the maximum size of the encoding unit
<img file="MX337918B_D0044.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0045.tif" />
of the video data 310 and 320 having the resolution higher than the video data 330 may be 64.
Since the maximum depth of the video data
310 is 2, 315 encoding units of the video data
310 They can include a maximum encoding unit that has a long axis size of 64, and encoding units that have a long axis size of 32 and 16 as depths are intensified to two layers by dividing the maximum encoding unit twice. Meanwhile, since the maximum depth of video data 33 0 is 1, encoding units 335 of video data 330 may include a maximum encoding unit having a long axis size of 16, and encoding units having a long axis size of 8 as depths are intensified to one layer by dividing the maximum encoding unit once.
Since the maximum depth of the video data
320 is 3, 325 encoding units of video data
320 They may include a maximum encoding unit that has a long axis size of 64, and encoding units that have a long axis size of 32, 18, and 8 as depths are intensified to three layers by dividing the maximum encoding unit. three times. As a depth intensifies, detailed information can be accurately expressed.
<img file="MX337918B_D0046.tif" />
Figure 4 is a block diagram of an image encoder 400 based on encoding units, in accordance with the embodiment of the present invention.
Image encoder 400 performs operations of encoding unit determiner 120 of video encoding apparatus 100 to encode image data. Specifically, an intra-predictor 410 performs intra-prediction in encoding units in an intra-mode between a current frame 405, and a motion estimator 420 and a motion compensator 425 perform inter-estimation and motion compensation in encoding units on an inter mode between the current box 405 when using the current box 405 and a reference box 495.
The data output from intra-predictor 410, motion estimator 420, and motion compensator 425 is sent as a quantized transform coefficient through a transformer 430 and a quantizer 440. The quantized transform coefficient is restored as data in a spatial domain through an inverse quantizer 460 and an inverse transformer 470. The data restored in the spatial domain is sent as reference frame 4 95 after being postprocessed through a 48 0 unlock unit and a 490 loop filter unit. The coefficient of í
Quantified transformation can be, ΤΠΌΤΟ MEXICAN INDUSTRIAL PROPERTY sent
<img file="MX337918B_D0047.tif" />
bit stream 455 through an entropy encoder 450. Specifically, the entropy encoder 450 can generate a bit stream by indicating a max_dec_f rame_buf fing syntax that indicates ^. a maximum size of a buffer required to decode an image frame by a decoder, a num_reorder_frames syntax indicating the number of image frames required to be reordered, and a MaxLatencyFrames syntax that indicates a maximum number of a difference value between an encoding order and a display order of image boxes that form an image sequence or a max_latency_increase syntax to determine the MaxLatencyFrames syntax in a NAL unit. Specifically, the entropy encoder 450 can add the max_dec_frame_buffering syntax, the num_reorder_frames syntax, and the max_latency_increase syntax to a sequence parameter group that is header information that includes information related to encoding a general image sequence, as required components.
In order to apply the image encoder 400 to the video encoding apparatus 100, all the elements of the image encoder 4 00, i.e. the intra-predictor 410, the motion estimator 420, the motion compensator 425, the transformer 430, the
A. Z Ό T ¿Vil Jf Jl
INDUSTRIAL PROPERTY INSTITUTE. MEXICO quantizer 440, the entropy encoder 450, the
, ._. ιυ ·· ι -1 -.-.--. -1 ----- v Me Inverse Quantizer 460, Inverse Transformer 470, Unlock Unit 480, and Loop Filtering Unit 490 perform operations based on each encoding unit among encoding units having a structure of tree while considering the maximum depth of each maximum encoding unit.
In particular, the intra-forecaster 410, the motion estimator 420, and the motion compensator
5 they determine divisions and a prediction mode of each encoding unit from among the encoding units having the tree structure while considering the maximum size and maximum depth of a current maximum encoding unit. Transformer 430 determines the size of the transformation unit in each encoding unit from among the encoding units having the tree structure.
FIG. 5 is a block diagram of an image decoder 500 based on encoding units, in accordance with an embodiment of the present invention.
An analyzer 510 analyzes a bit stream 505 to obtain encoded image data to be decoded and encoding information required to decode the encoded image data. Specifically, the 510 parser gets and sends a syntax
<img file="MX337918B_D0048.tif" />
IMPIO .MSTiTUTO MEXICANO
OF THE V * eES PROPERTY
INDUSTRIAL max_dec_frame_buffering indicating a maximum size of a buffer required to decode a picture frame included with a mandatory component in an SPS, a num_reorder_frames syntax indicating the number of picture frames required to be reordered, and a max_latency_increase syntax to determine a syntax
MaxLatencyFrames from a bitstream to an entropy decoder 520. In Figure 5, the parser 510 and the entropy decoder 520 are illustrated to be individual components, but alternatively, procedures can be performed to obtain image data and obtain syntax information related to encoded image data, which is performed by analyzer 510, using entropy decoder 520.
The encoded image data is sent as inversely quantized data through the entropy decoder 520 and an inverse quantizer 530, and the inverse quantized data is restored to image data in a spatial domain through an inverse transformer 540.
With respect to the image data in the spatial domain, an intra predictor 550 performs intra prediction in encoding units in an intramode, and a motion compensator 560 performs motion compensation in encoding units in an intermode using a table. Reference 585.
Image box data restored through the
<img file="MX337918B_D0049.tif" />
Intra-forecaster 550 and motion compensator 560 are post-processed through an unlock unit 570 and sent to a DPB 580. The DPB 580 stores a decoded image frame for storing a reference frame, changing an order of submitting an image box, and submitting an image box. The DPB 580 stores the decoded image frame while setting a maximum buffer size required for normal decoding of an image sequence by using a max_dec_frame_buffering syntax indicating a maximum buffer size required to normally decode an image frame output from the 510 analyzer or the 520 entropy decoder.
Also, the DPB 580 can determine whether to send a pre-decoded and stored reference image frame by using a num_reorder_frames syntax indicating the number of image frames required to be reordered and a max_latency_increase syntax to determine a MaxLatencyFrames syntax. A procedure for sending a reference image frame stored in DPB 580 will be described in detail later.
In order to decode the image data by using the image data decoder 230 of the video decoding apparatus 200, the image decoder
<img file="MX337918B_D0050.tif" />
MEXICAN INSTITUTE ÜE LA PRüflEPAD
INDUSTRIAL
<img file="MX337918B_D0051.tif" />
500 it can perform operations that are performed after a 510 scanner operation.
In order to apply the image decoder 500 to the video decoding apparatus 200, all the elements of the image decoder 500, i.e. analyzer 510, entropy decoder 520, inverse quantizer 530, inverse transformer 540, intra-forecaster 550, motion compensator 560, and unlocking unit 570 can perform decoding operations based on encoding units having a tree structure, in units of maximum encoding units. In particular, intra prediction 550 and motion compensator 560 determine divisions and a prediction mode for each of the encoding units having the tree structure, and inverse transformer 540 determines a size of one transformation unit for each of the encoding units.
Figure 6 is a diagram illustrating coding units corresponding to depths, and divisions, in accordance with one embodiment of the present invention.
The video encoding apparatus 100 and the video decoding apparatus 200 in accordance with one embodiment of the present invention use hierarchical encoding units to consider characteristics of
<img file="MX337918B_D0052.tif" />
INSTITUTO MEXICANO LE LA PROPIEDAD
INDUSTRIAL
<img file="MX337918B_D0053.tif" />
an image. A maximum height, maximum width, and maximum depth of a coding unit can be adaptively determined according to the characteristics of the image, or can be set differently by a user. Coding unit sizes corresponding to depths can be determined according to the predetermined maximum coding unit size.
In a hierarchical structure 600 of encoding units in accordance with an embodiment of the present invention, the maximum height and maximum width of the encoding units are each 64, and the maximum depth is 4. Since a depth is intensified to Along a vertical axis of hierarchical structure 600, each is divided by a height and width of each of the coding units corresponding to depths. Also, a prediction unit and divisions, which are bases for prediction coding of each of the coding units corresponding to depths, are shown along a horizontal axis of the hierarchical structure.
600.
Specifically, hierarchical structure 600, a 610 encoding unit is a maximum encoding unit, and has a depth of 0 and a size of 64x64 (height by width). As the
<img file="MX337918B_D0054.tif" />
depth along the vertical axis, there is a 620 encoding unit having a size of 32x32 and a depth of 1, and a 630 encoding unit having a size of 16x16 and a depth of 2, a 640 encoding unit which has a size of 8x8 and a depth of 3, and a coding unit 650 that has a size of 4x4 and a depth of 4. Coding unit 650 which is 4x4 in size and depth of 4 is a minimal coding unit.
A prediction unit and divisions of each coding unit are arranged along the horizontal axis in accordance with each depth. If the 610 encoding unit having the size 64x64 and the depth 0 is a prediction unit, the prediction unit may be divided into divisions included in the 610 encoding unit, i.e. a 610 division having a size 64x64, 612 divisions that are 64x32 in size, 614 divisions that are 32x64 in size, or 616 divisions that are 32x32 in size.
Similarly, a prediction unit of the 620 encoding unit having the size of 32x32 and the depth of 1 can be divided into divisions included in the 620 encoding unit, i.e. a 620 division having a size of 32x32, 622 divisions that are 32x16 in size, 624 divisions that are 16x32 in size, and 626 divisions that are 16x16 in size.
<img file="MX337918B_D0055.tif" />
INDUSTRIAL ^ * = S_22! _- Similarly, a prediction unit of encoding unit 630 having the size of 16x16 and the depth of 2 can be divided into divisions included in encoding unit 630, i.e. a division 630 that it is 16x16 in size, 632 divisions that are 16x8 in size, 634 divisions that are 8x16 in size, and 636 divisions that are 8x8 in size.
Similarly, a prediction unit of coding unit 640 having the size of 8x8 and depth of 3 can be divided into divisions included in coding unit 640, i.e. a division 640 having a size of 8x8, divisions 642 that are 8x4 in size, 644 divisions that are 4x8 in size, and 646 divisions that are 4x4 in size.
The 650 encoding unit having the size of 4x4 and the depth of 4 is the minimum encoding unit having the lowest depth. A coding unit prediction unit 650 is set only for a division 650 that is 4x4 in size.
In order to determine an encoded depth of the maximum encoding unit 610, the encoding unit determiner 120 of the video encoding apparatus 100 encodes all the encoding units corresponding to each depth, included in the maximum encoding unit 610.
<img file="MX337918B_D0056.tif" />
As the depth intensifies, it increases
<img file="MX337918B_D0057.tif" />
a number of encoding units, which correspond to each depth and include data that have the same range and the same size. For example, four encoding units corresponding to a depth of 2 are required to cover data included in one encoding unit corresponding to a depth of 1. Therefore, in order to compare results for encoding the same data according to depths, the encoding unit corresponding to the depth of 1 and the four encoding units corresponding to the depth of 2 are each encoded.
In order to perform coding in depth units, a minimum coding error for each of the depths can be selected as a representative coding error per prediction coding units in each of the coding units corresponding to the depths, at along the horizontal axis of hierarchical structure 600. Alternatively, a minimum coding error can be sought by coding in units of depths and by comparing minimum coding errors in accordance with depths, as depth is intensified along the vertical axis of hierarchical structure 600. A depth and a division that have the error of
IMPI ^^
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL minimum encoding in the maximum encoding unit 610 can be selected as an encoded depth and a division type of the maximum encoding unit 610.
Figure 7 is a diagram illustrating a relationship between a 710 encoding unit and 7 20 transformation units, in accordance with an embodiment of the present invention.
The video encoding apparatus 100 (or the video decoding apparatus 200) in accordance with an embodiment of the present invention encodes (or decodes) an image in units of maximum encoding units, based on encoding units having sizes less than or equal to the maximum encoding units. During encoding, a size of each transformation unit used to perform transformation can be selected based on a data unit that is not larger than a corresponding encoding unit.
For example, in video encoding apparatus 100 (or video decoding apparatus 200), if a size of encoding unit 710 is 64x64, transformation can be performed using transformation units 720 having a size of 32x32.
Also, data from the 710 encoding unit that is 64x64 in size can be encoded when performing
MEXICAN INSTITUTE DF THE PROPERTY
INDUSTRIAL
<img file="MX337918B_D0058.tif" />
transformation in each of the transformation units having a size of 32x32, 16x16, 8x8, and 4x4, which are smaller than 64x64, and then a transformation unit that has minimal coding error can be selected.
Figure 8 is a diagram illustrating coding information corresponding to depths, in accordance with an embodiment of the present invention.
The output unit 130 of the video encoding apparatus 100 can encode and transmit information 800 on a division type, information 810 on a prediction mode, and information 820 on transformation unit size for each encoding unit corresponding to a depth encoded, as information about an encoding mode.
Information 800 indicates information about a form of a division obtained by dividing a prediction unit from a current encoding unit, such as a data unit for prediction encoding from the current encoding unit. For example, a current CU__0 encoding unit that is 2Nx2N in size can be divided into any one of an 802 division that is 2Nx2N in size, an 804 division that is 2Nx2N in size.
2NxN, an 806 division that is Nx2N in size, and an 808 division that is NxN in size. In this case, information 800 is set to indicate one of the
MEXICAN INSTITUTE
DS THE PROPERTY VtT / tSCáv
INDUSTRIAL division 804 which is 2NxN in size, division 806 which is Nx2N in size, and division 808 which is NxN in size.
Information 810 indicates a prediction mode for each division. For example, information 810 may indicate a prediction encoding mode of the division indicated by information 800, ie, an intra mode 812, an inter mode 814, or a jump mode 816.
Information 820 indicates a transformation unit to be based on when the transformation is performed in a current encoding unit. For example, the transformation unit may be a first intra transformation unit 822, a second intra transformation unit 824, a first inter transformation unit 826, or a second intra transformation unit 828.
The image data and encoding information extractor 220 of the video decoding apparatus 200 can extract and use the information 800, 810, and 820 to decode encoding units corresponding to depths.
Figure 9 is a diagram illustrating coding units corresponding to depths, in accordance with an embodiment of the present invention.
Information divided by being used to indicate a change in depth. The divided information indicates whether a
IMPI
N'TITUTO MF.X1CANO
OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0059.tif" />
Coding unit of a current depth is divided into coding units of a lower depth.
A prediction unit 910 for predicting coding a coding unit 900 having a depth of 0 and a size of 2N_0x2N_0 may include divisions of a division type 912 having a size of 2N_0x2N_0, a division type 914 having a size 2N_0xN_0, a division type 916 that is N_0x2N_0, and a division type 918 that is N_0xN_0. Although Figure 9 illustrates only the types of division 912 to 918 that are obtained by symmetrically dividing the prediction unit 910, one type of division is not limited to these, and the divisions of the prediction unit 910 may include asymmetric divisions, divisions that have an arbitrary shape, and divisions that have a geometric shape.
Prediction encoding is done repeatedly in a division that has a size of 2N_OxN_0 ,. two divisions that are 2N_0x2N_0 in size, two divisions that are N_0x2N_0, and four divisions that are N_0xN_0, according to each type of division. Prediction coding can be performed on divisions having the sizes 2N_0x2N_0, N_0x2N_0, 2N_0xN_0, and N_0xN_0, in accordance with an intra mode and an inter mode. Coding is done by
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0060.tif" />
prediction only in the division that is 2N_0x2N_0 in size, according to a jump mode.
If a coding error is smaller in one of the division types 912 to 916, the prediction unit 910 may not be divided to a lesser depth.
If an encoding error is the smallest in division type 918, a depth is changed from 0 to to divide division type 918 in operation 920, and encoding is repeatedly performed in 930 encoding units that have divisions of one depth of and a size of N_0xN_0 to look for a minimal encoding error.
A prediction unit 940 for predicting coding the coding unit 930 having a depth of 1 and a size of 2N_lx2N_l (= N_0xN_0) can
<td>include divisions of</td><td>a</td><td>type</td><td>of</td><td>division</td><td> 942</td><td>than</td><td>has</td><td>a</td>
<td>size of 2N_lx2N 1,</td><td>a</td><td>type</td><td>of</td><td>division</td><td> 944</td><td>than</td><td>has</td><td>a</td>
<td colspan="2">size of 2N lxN 1, a</td><td>type</td><td>of</td><td>division</td><td> 946</td><td>than</td><td>has</td><td>a</td>
<td>size of N 1χ2Ν 1, and</td><td>a</td><td>type</td><td>of</td><td>division</td><td> 948</td><td>than</td><td>has</td><td>a</td>
size of N_lxN_l.
If an encoding error is the smallest in division type 948 that has a size of N_lxN_l, a depth is changed from 1 to 2 to divide division type 948 in operation 950, and encoding in encoding units is performed repeatedly 960 who have a
<img file="MX337918B_D0061.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0062.tif" />
depth of 2 and a size of N_2xN_2 to look for a minimal encoding error.
When a maximum depth is d, encoding units corresponding to depths can be set when a depth becomes d-1, and divided information can be set to when a depth d-2. In other words, when coding is performed up to when the depth is d-1 after a coding unit corresponding to a depth of d-2 is operationally divided to 170, a prediction unit 990 for predicting coding a coding unit 980 that has a depth of d-1 and a size of 2N_ (d-1) x2N_ (d-1) can include divisions of a division type 992 that has a size of 2N_ (d-1) x2N_ (d-1 ), a division type 994 that has a size of 2N_ (dl) xN_ (d-1), a division type 996 that has a size of N_ (d-1) x2N_ (d1), and a division type 998 that has a size of N_ (dl) xN (d-1).
Prediction coding can be performed repeatedly on one division that is 2N_ (dl) x2N_ (dl) in size, two divisions that are 2N_ (dl) xN_ (dl) in size, two divisions that are N_ (dl) in size x2N_ (dl), and four divisions that are N_ (dl) xN_ (dl) from division types 992 to 998 to find a division type that has minimal coding error.
<img file="MX337918B_D0063.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0064.tif" />
Even when division type 998 has the minimum coding error, since a maximum depth is d, a CU_ (dl) coding unit that has a depth of d-1 is no longer divided to a lower depth, and is determined a depth encoded for a current maximum encoding unit 900 to be d-1 and a division type of encoding unit 900 can be determined to be N_ (d-1) xN_ (d-1). Also, since the maximum depth is d, divided information is not set for an encoding unit 952 that has a depth of (d1).
A data unit 999 may be a 'minimum unit' for the current maximum encoding unit 900. A minimum unit in accordance with an embodiment of the present invention may be a rectangular data unit obtained by dividing a minimum unit having a depth coded lower by 4. By repeatedly encoding as described above, the video encoding apparatus 100 can determine an encoded depth by comparing encoding errors in accordance with the depths of the encoding unit 900 and selecting a depth having the least encoding error, and setting a division type and prediction mode for encoding unit 900 as a encoding depth encoding mode.
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0065.tif" />
As such, minimum coding errors are compared according to depths, i.e. the depths of 0, 1, ..., d-1, and d, with each other, and a depth can be determined that has the minimum coding error as a coded depth. The coded depth, the division type of the prediction unit, and the prediction mode can be coded and transmitted as information about a coding mode.
Also, since one encoding unit is divided from the depth of 0 to the encoded depth, only information divided from the encoded depth is set to 0, and information divided from the other depths excluding the encoded depth is set to 1.
The image data and encoding information extractor 220 of the video decoding apparatus 200 can extract and use the information about the encoded depth and the prediction unit of the encoding unit 900 to decode the division 912. The video decoding apparatus 200 can determine a depth corresponding to divided information '0', such as a coded depth, based on information divided according to depths, and can use information about a coding mode of the coded depth during a decoding procedure.
IMPí ^ js
MEXICAN INSTITUTE
C> £ PROPERTY -M
INDUSTRIAL
Figures 10, 11, and 12 are diagrams illustrating a relationship between encoding units 1010, prediction units 1060, and transformation units 1070, in accordance with one embodiment of the present invention.
The encoding units 1010 are encoding units corresponding to encoded depths for a maximum encoding unit, determined by the video encoding apparatus 100. The prediction units 1060 are divisions of prediction units of the respective encoding units 1010, and transformation units 1070 are transformation units of the respective encoding units 1010.
Among the 1010 encoding units, if a depth of a maximum encoding unit is 0, then encoding units 1012 and 1054 have a depth of 1, encoding units 1014, 1016, 1018,
1028, 1050, and 1052 have a depth of 2, and 1020, 1022, 1024, 1026, 1030, 1032, and 1048 encoding units have a depth of 3, and 1040, 1042, 1044, and 1046 encoding units have a depth of 4.
Among the 1060 prediction units, some divisions 1014, 1016, 1022, 1032, 1048, 1050, 1052, and 1054 are divided into divided divisions of encoding units. In other words, divisions 1014, 1022,
1050, and 1054 are division types 2NxN, divisions 1016,
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX337918B_D0066.tif" />
1048, and 1052 are division types Nx2N, and division 1032
<td colspan="2">it is a type of NxN division.</td><td>Units</td><td>of</td><td>prediction and</td>
<td>unit divisions</td><td>of</td><td colspan="2">coding</td><td>1010 are more</td>
<td>small than or equal</td><td>to</td><td>units</td><td>of</td><td>coding</td>
<td>corresponding to these.</td><td></td><td></td><td></td><td></td>
Among transformation units 1070, reverse transformation or transformation is performed on image data corresponding to encoding unit 1052, based on a data unit that is smaller than encoding unit 1052. Also, transformation units 1014, 1016 , 1022, 1032, 1048, 1050, 1052, and 1054 are data units other than prediction units and corresponding divisions between prediction units 1060, in terms of sizes and shapes. In other words, the video encoding apparatus 100 and the video decoding apparatus 200 in accordance with one embodiment of the present invention can individually perform intra prediction, motion estimation, motion compensation, transformation, and inverse transformation therein. encoding unit, based on different data units.
Accordingly, an optimal encoding unit can be determined by recursively encoding encoding units having a hierarchical structure, in units of regions of each maximum encoding unit, thereby obtaining encoding units having a
<img file="MX337918B_D0067.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0068.tif" />
recursive tree structure. Coding information may include divided information about a coding unit, information about a type of division, information about a prediction mode, and information about a size of a transformation unit. Table 1 shows an example of encoding information that can be set by the video encoding apparatus 100 and the video decoding apparatus 200.
Table 1
<td colspan="5">Divided Information 0 (Coding in Coding Unit having 2Nx2N Size and Current Depth of d)</td><td>Divided Information 1</td>
<td>Mode of Prediction</td><td colspan="2">Division Type</td><td colspan="2">Unit Size Trans formation</td><td>Repeatedly Coding Units Coding they have D + 1 Bottom Depth</td>
<td rowspan="2">Intra Inter Jump (Only 2Nx2N)</td><td>Kind of Division Symmetrical</td><td>Kind of Division Asymmetric</td><td>Information Divided 0 from Unit of Transformation ion</td><td>Divided Information 1 Transformer Unit</td><td></td>
<td>2NX2N 2NxN Nx2N NxN</td><td>2NxnU 2NxnD nLx2N nRx2N</td><td>2Nx2N</td><td>NxN (Kind Symmetrical) N / 2XN / 2 (Kind Asymmetric)</td><td></td>
The output unit 130 of the video encoding apparatus 100 can send the encoding information on the encoding units having a tree structure, and the image data and encoding information extractor 220 of the video decoding apparatus
<img file="MX337918B_D0069.tif" />
MEXICAN INSTITUTE OE THE PROPERTY
INDUSTRIAL
<img file="MX337918B_D0070.tif" />
200 you can extract the encoding information about the encoding units having a tree structure from a received bit stream.
Divided information indicates whether a current encoding unit is divided into encoding units of a lesser depth. If the information divided from a current depth d is 0, a depth, in which the current coding unit is no longer divided into coding units of a lower depth, is a coded depth, and thus information about a division type, a prediction mode, and a transformation unit size for the coded depth. If the current encoding unit is further divided according to the divided information, encoding is independently performed in four divided encoding units of a lower depth.
The prediction mode can be one of an intra mode, an inter mode, and a jump mode. Intramode and intermode can be defined for all division types, and jump mode is defined only for a 2Nx2N division type.
Information about the division type can indicate symmetric division types that have sizes of 2Nx2N, 2NxN, Nx2N, and NxN, which are obtained by symmetrically dividing a height or width of a prediction unit, and asymmetric division types that have
<img file="MX337918B_D0071.tif" />
sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N, which are obtained by asymmetrically dividing the height or width of the prediction unit. Asymmetric division types having the sizes of 2NxnU and 2NxnD can be obtained respectively by dividing the height of the prediction unit by 1: 3 and 3: 1, and asymmetric division types having the sizes of nLx2N and nRx2N respectively by dividing the width of the prediction unit into 1: 3 and 3: 1.
The transformation unit size can be set to be two types in intra mode and two types in inter mode. In other words, the information divided from the transformation unit is 0, the size of the transformation unit can be 2Nx2N equal to the size of the current encoding unit. If the information divided from the transformation unit is 1, transformation units can be obtained by dividing the current encoding unit. Also, a transformation unit size can be NxN when a division type of the current encoding unit that is 2Nx2N is a symmetric division type, and can be N / 2xN / 2 when the division type of the Current encoding unit is an asymmetric type of division.
Coding information on coding units having a tree structure can be assigned to at least one of a coding unit
<img file="MX337918B_D0072.tif" />
TO,
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL _ 'Stó. a unit corresponding to a predicted coded depth, and a minimum unit. The unit3 of<sup>r</sup>Coding ic'acióñ corresponding to the coded depth can include at least one prediction unit and at least one minimum unit that contains the same coding information.
Accordingly, whether adjacent data units are included in encoding units corresponding to the same encoded depth can be determined by comparing encoding information from the adjacent data units. Also, a coding unit corresponding to a coded depth can be determined using coding information from a data unit. In this way, a distribution of coded depths can be determined in a maximum code unit.
Accordingly, if the current encoding unit is predicted based on encoding information from adjacent data units, you can directly indicate and use encoding information from data units in encoding units corresponding to adjacent depths in the current encoding unit.
Alternatively, if the current encoding unit is predicted based on adjacent encoding units, then adjacent encoding units can be indicated when searching for data units
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
O. Ρ I
MEXICAN
ROP1SÜAD OtaiKef :, DUSTIÜAL adjacent to the current encoding unit from among depth units, encoding based on corresponding encoding information from adjacent encoding units corresponding to depths.
Figure 13 is a diagram illustrating a relationship between a coding unit, a prediction unit, and a transformation unit, in accordance with the coding mode information in Table 1.
A maximum 1300 coding unit includes coding units 1302, 1304, 1306, 1312, 1314, 1316, and 1318 of coded depths. Here, since the coding unit 1318 is a coding unit of a coded depth, information divided therefrom can be set to 0. Information about a division type of the 1318 encoding unit that is 2Nx2N in size can be set to be one of a 1322 division type that is 2Nx2N in size, a 1324 division type that is 2NxN in size, a type division type 1326 having a size of Nx2N, a division type 1328 having a size of NxN, a division type 1322 having a size of 2NxnU, a division type 1334 having a size of 2NxnD, a division type 1336 that is nLx2N in size, and a division type 1338 that is nRx2N in size.
<img file="MX337918B_D0073.tif" />
MEXICAN JUDGER OR<sup>c</sup> The property
INDUSTRIAL
<img file="MX337918B_D0074.tif" />
For example, if the division type is set to be a symmetric division type, for example, the division type 1322, 1324, 1326, 1328, then a transformation unit 1342 that has a size of 2Nx2N is set when the information Transformer unit divided (TU size indicator) is '0', and a 1344 transformation unit having NxN size is set when the TU size indicator is '1'.
If the division type is set to be an asymmetric division type, for example the division type
1332, 1334, 1336, 1338, then a 1352 transformation unit is set having a size of 2Nx2N when a TU size indicator is 0, and a 1354 transformation unit is set having a size of N / 2xN / 2 when a TU size indicator is 1.
Figure 14 is a diagram of an image encoding procedure and an image decoding procedure, which are hierarchically classified, in accordance with an embodiment of the present invention.
Coding procedures performed by the video coding apparatus 100 of Figure 1 or the image encoder 400 of Figure 4 can be classified into a coding procedure performed in a video coding layer (VCL). 1410 which handles an image encoding procedure itself
MEXICAN INSTITUTE DS THE PROPERTY
INDUSTRIAL
<img file="MX337918B_D0075.tif" />
itself, and a coding procedure performed on a
NAL 1420 that generates image data and additional encoded information between VCL 1410 and a lower system 1430 that transmits and stores encoded image data, as a bitstream in accordance with a predetermined format as shown in Figure 14. Encoded data 1411 which is an output of encoding procedures of maximum encoding unit divider 110 and encoding unit determiner 120 of video encoding apparatus 100 of Figure 1 is VCL data, and encoded data 1411 is indicated to a VCL NAL unit
1421 through output unit 13 0. Also, information directly related to the encoding procedure of VCL 1410, such as divided information, division type information, prediction mode information, and transformation unit size information On a coding unit used to generate the data encoded 1411 by VCL 1410, the VCL NAL 1421 unit is also indicated. Parameter group information 1412 related to the encoding procedure is indicated to a non-VCL NAL 1422 unit. In particular, in accordance with an embodiment of the present invention, a max_dec_frame_buffering syntax indicating a maximum size of a buffer required to decode a image box by a decoder, a syntax
<img file="MX337918B_D0076.tif" />
A IV1 <sub>u></sub>-___
VX MEXICAN INSTITUTE
FROM PROPERTY V> íxINDUSTRIAL a num_reorder_frames indicating the number of image frames required to be reordered, and max_latency_increase syntax to determine a MaxLatencyFrames are indicated to the non-VCL unit NAL 1422.
Both the VCL NAL 14 21 unit and the non-VCL NAL 1422 unit are NAL units, where the VCL NAL 1421 unit includes image data that is compressed and encoded, and the non-VCL NAL 1422 unit includes parameters corresponding to a sequence of picture and header information of a box.
Similarly, decoding procedures performed by the video decoding apparatus 200 of Figure 2 or the image decoder 500 of Figure 5 can be classified into a decoding procedure performed in VCL 1410 that handles an image decoding procedure per se. same, and a decoding procedure performed on NAL 1420 that obtains encoded image data and additional information from a received and read bitstream between VCL 1410 and lower system 143 0 that receives and reads the encoded image data, as shown in Figure 14. The decoding procedures performed on the receiver 210 and the image data and encoding information extractor 220 of the video decoding apparatus 200 of Figure 2 correspond to the decoding procedures of NAL 1420, and the
IMPI
Decoding of the image data decoder 230 corresponds to the decoding procedures of VCL 1410. In other words, the receiver 210 and the extractor of image data and encoding information 220 obtain, from a stream bit 1431, the VCL NAL drive
1421 including information used to generate encoded image data and encoded data, such as divided information, division type information, prediction mode information, and transformation unit size information of an encoding unit, and the non-VCL unit NAL 1422 that includes parameter group information related to the encoding procedure. In particular, in accordance with an embodiment of the present invention, a max_dec_frame_buffering syntax indicating a maximum size of a buffer required to decode an image frame by a decoder, a num_reorder_frames syntax indicating the number of image frames required to be reordered , and a max_latency_increase syntax to determine a MaxLatencyFrames syntax are included in the non-VCL NAL 1422 unit.
Figure 15 is a diagram of a structure of a NAL 150 0 unit, in accordance with an embodiment of the present invention.
Referring to Figure 15, the NAL 1500 unit includes a NAL 1510 header and a payload of
<img file="MX337918B_D0077.tif" />
<img file="MX337918B_D0078.tif" />
<img file="MX337918B_D0079.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0080.tif" />
unprocessed byte stream (RBSP) 1520.
An RBSP 1530 padding bit is a is a length adjustment bit added at the end of RBSP 1520 to express a length of RBSP 1520 at a multiple of 8 bits. The RBSP 1530 pad bit starts from '1' and includes continuous '0' determined according to the length of RBSP 1520 to have a pattern like '100 ...'. By searching for '1' i.e. an initial bit value, a location of the last bit of the RBSP 1520 can be determined.
The NAL header 1510 includes flag information (nal_ref_idc) 1512 indicating whether a fragment constituting a reference illustration of a corresponding NAL unit is included, and an identifier (nal_unit_type) 1513 indicating this type of NAL unit. '1' 1511 at the beginning of the NAL 1510 header is a fixed bit.
The NAL 1500 unit can be classified into an Instant Decoding Update (IDR) illustration, a Clean Random Access Access (CRA) illustration, an SPS, an illustration parameter group (PPS), Supplemental Improvement Information (SEI), and an Adaptive Parameter Group (APS) in accordance with a value of nal_unit_type 1513. Table 2 shows a unit type of NAL 150 0 nal_unit_type 1513.
in accordance
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL with values
<img file="MX337918B_D0081.tif" />
Table 2
<td>nal_unit_type</td><td>NAL unit type</td>
<td> 0</td><td>Not specified</td>
<td> 1</td><td>Illustration excluding CRA and fragment of illustration excluding IDR</td>
<td> 2-3</td><td>Reserved for future expansion</td>
<td> 4</td><td>CRA Illustration Fragment</td>
<td> 5</td><td>IDR Illustration Fragment</td>
<td> 6</td><td>SEI</td>
<td> 7</td><td>SPS</td>
<td> 8</td><td>PPS</td>
<td> 9</td><td>Access Unit Delimiter (AU)</td>
<td> 10-11</td><td>Reserved for future expansion</td>
<td> 12</td><td>Fill data</td>
<td> 13</td><td>Reserved for future expansion</td>
<td> 14</td><td>APS</td>
<td> 15-23</td><td>Reserved for future expansion</td>
<td> 24-64</td><td>Not specified</td>
As described above, in accordance with one embodiment of the present invention, the max_dec_frame_buff ering syntax, the num_reorder_f rames syntax, and the max_latency_increase syntax are included in the NAL unit, specifically the SPS corresponding to the sequence header information. image, as mandatory components.
Hereinafter, procedures for determining the max_dec_frame_buffering syntax, the num_reorder_frames syntax, and the max_latency_increase syntax, which are included as the mandatory components of the SPS, will be described during the encoding procedure.
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0082.tif" />
An image frame decoded in a VCL is stored in DPB 580 which is an image buffer memory of image decoder 500. DPB 580 marks each stored illustration as a short term reference illustration indicating for a short term, a long-term reference illustration that is indicated for a long term, a non-reference illustration that is not indicated. A decoded artwork is stored in the DPB 580, rearranged according to an output order, and sent from the DPB 580 at an output time or assigned time when the artwork decoded by another picture box is not indicated.
In a general codec, such as an H.264 AVC codec, a maximum size of one DBP required to restore an image box is defined by a profile and a level, or through Video Utility Information (VUI). in English) which is broadcast selectively. For example, the maximum DPB size defined by the H.264 AVC codec is defined as Table 3 below.
Table 3
<td rowspan="2">Resolution</td><td>WQVGA</td><td>Wvga</td><td>HD 720p</td><td>HD 10809</td>
<td>400x240</td><td>800x480</td><td>1280x720</td><td>1920x1080</td>
<td>Minimum level</td><td> 1.3</td><td> 3.1</td><td> 3.1</td><td> 4</td>
<td>MaxDPB</td><td> 891.0</td><td> 6750.0</td><td> 6750.0</td><td> 12288.0</td>
<td>MaxDpbSize</td><td> 13</td><td> 12</td><td> 5</td><td> 5</td>
In Table 3, the maximum size of DPB is defined with respect to a 30 Hz image, and in H.264 AVC codec, the maximum size of DPB is determined by using the max_dec_frame_buffering syntax transmitted selectively through
IMPI
VUI MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY, or in accordance with a predetermined table in accordance with a profile and a level as shown in Table 3 if the max_dec_frame_buffering syntax is not included in the VUI. If a resolution of a decoder is
400x240 (WQVGA) and a frequency of an output image is 30 Hz, a maximum size (MaxDpbSíze) of the DPB is 13, that is, the maximum size of the DPB is set to store 13 decoded illustrations.
In a general video codec, information about a maximum size of a DPB is not necessarily transmitted, but is transmitted selectively. Therefore, in the general video codec, information on a maximum size of a DPB required to decode an image sequence by a decoder cannot always be used.
When such information is not transmitted, the decoder uses a maximum size of a predetermined DPB in accordance with a profile and level, as shown in Table 3 above. However, a DPB size actually required during image sequence encoding and decoding procedures is often smaller than the maximum DPB size in Table 3. That way, if the default maximum size, as shown in Table 3, is used, decoder system resources can be wasted. Also, in accordance with the general video codec, since the size
<img file="MX337918B_D0083.tif" />
IMPI iNsrrfí »fí> ^ .íCaho
Yes THE “INDUSTRIAL OHTOAO of the decoder DPB is smaller than the maximum size
<img file="MX337918B_D0084.tif" />
default of Table 3 but is larger than a size actually required to restore a picture frame, if no information is transmitted about a maximum DPB size required for a decoding procedure even though the decoder is capable of decoding a sequence image size, the default maximum size in Table 3 is
<td>set as</td><td>the size of the</td><td colspan="3">DPB required for</td><td>the</td>
<td>Procedure of</td><td>decoding,</td><td>and of</td><td>that</td><td>shape</td><td>the</td>
<td>Procedure of</td><td>decoding</td><td>can</td><td>to be</td><td>be unable</td><td>of</td>
<td>perform. By</td><td>consequently a</td><td>method</td><td>and</td><td>apparatus</td><td>of</td>
Image encoding in accordance with an embodiment of the present invention transmits a maximum size of one DPB to a decoding apparatus after including the maximum size as a mandatory component of an SPS, and an image decoding method and apparatus can establish a maximum size of a DPB when using a maximum size included in an SPS.
Figures 16A and 16B are reference diagrams for describing maximum size information of a required DPB in accordance with a decoding order during an image sequence encoding procedure.
Referring to Figure 16A, it is assumed that an encoder performs encoding in the order of 10, Pl,
<img file="MX337918B_D0085.tif" />
Ρ2, Ρ3 and Ρ4, and coding is done by indicating ——— r »áai« nm illustrations in directions indicated by arrows. Similar to such an encoding order, decoding is performed in an order of 10, Pl, P2, P3, and P4. In Figure 16A, since An Illustration refers to a reference illustration that is immediately pre-decoded, a maximum DPB size required to normally decode an image sequence is 1.
Referring to Figure 16B, it is assumed that an encoder performs encoding in the order of 10, P2, bl, P4, and b3 by indicating illustrations in directions indicated by arrows. Since a decoding order is the same as the encoding order, decoding is performed in an order of 10, P2, bl, P4, and b3. In an image sequence of Figure 16B, since an illustration P refers to an illustration I that is pre-decoded or a reference illustration of illustration P, and an illustration b refers to illustration I that is pre-decoded or two Reference illustrations in Illustration P, a maximum size of a DPB required to normally decode image sequence 2. Although the maximum DPB size required to normally decode the image sequence has a small value of 1 or 2 as shown in Figures 16A and 16B, if the information about the maximum DPB size is not transmitted
<img file="MX337918B_D0086.tif" />
in:
I
<img file="MX337918B_D0087.tif" />
separately, the decoder has to use information on a maximum size of a predetermined DPB in accordance with profiles and levels of a video codec.
If the decoder's DPB has a maximum value of 3, that is, it is capable of storing a maximum of three decoded image frames, and a maximum DPB size is set to be 13 in accordance with Table 3 as a default value in accordance with a profile or level of a video codec, even though the DPB is large enough to decode an encoded picture frame, the size of the DPB is smaller than the default maximum size of the DPB, and thus the decoder may erroneously determine that the encoded picture frame cannot be decoded.
Accordingly, the video encoding apparatus 10 0 in accordance with one embodiment of the present invention determines a max_dec_frame_buffering syntax indicating a maximum size of a DPB required to decode each picture frame by a decoder, based on an encoding order (or decoding order) of image frames forming an image sequence and an encoding order (or decoding order) of reference frames indicated by the image frames, and inserts and transmits the max_dec_frame_buffering syntax to and with an SPS corresponding to header information of
IMPI
INSTITUTO MSXICANO DS THE PROPERTY, · -l τ-m J INDUSTRIAL the image sequence. The video encoding apparatus 100 includes the max_dec_frame_buffermg syntax in the SPS branch mandatory information instead of selective information.
Meanwhile, when a decoded artwork is stored in the decoder's DPB in a general video codec and a new space is required to store the decoded artwork, a reference artwork having a lower display order is sent (order count of illustration) from the DPB through magnification to obtain an empty space to store a new reference illustration. In the general video codec, the decoder is capable of displaying the decoded artwork only when the decoded artwork is sent from the DPB through such augmentation procedure. However, when the decoded illustration is presented through the magnification procedure as such, the output of a pre-decoded reference illustration is delayed until the magnification procedure.
FIG. 17 is a diagram illustrating a procedure for sending a decoded illustration from a DPB in accordance with an augmentation procedure in a video codec field related to the present invention. In Figure 17, it is assumed that a maximum size (MaxDpbSize) of the DPB is 4, that is, the DPB can store a maximum of four decoded illustrations.
<img file="MX337918B_D0088.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTSLLAL
<img file="MX337918B_D0089.tif" />
Referring to Figure 17, in a general video codec field, if a P4 frame decoded four frames after Illustration 10 is to be stored in a DPB even though Illustration 10 was first decoded in accordance with a Decoding order, Illustration 10 can be sent from the DPB and presented via an augmentation procedure. Accordingly, illustration 10 is sent after being delayed 4 frames from a decoding time.
Therefore, the video decoding apparatus 200 in accordance with an embodiment of the present invention quickly sends a decoded illustration from a DPB without an increase procedure by setting a predetermined latency parameter from a time each decoded illustration is stored in the DPB when using a MaxLatencyFrames syntax that indicates a maximum number of preceding image frames from a default frame in an order-based image sequence display but behind the predetermined frame based on a decoding order, increasing a latency parameter count of the decoded artwork stored in the DPB by 1 any time each artwork in the image sequence is decoded according to the order , and send a decoded artwork whose latency parameter count has reached the MaxLatencyFrames syntax from the DPB. In other words, the apparatus
<img file="MX337918B_D0090.tif" />
IMPI
INSTITUTO MEXICANO DI LA PROPERTY INDUSTRIAL video decoding 200 initially assigns 0 as a latency parameter to a decoded artwork stored in a DPB when the decoded artwork is stored in the DPB, and increases the latency parameter by 1 anytime it is decodes the following illustration lxl in accordance with a decoding order. Also, the video decoding apparatus 200 compares the latency parameter with the MaxLatencyFrames syntax to send a decoded artwork whose latency parameter has the same value as the MaxLatencyFrames syntax from the DPB.
For example, when the MaxLatencyFrames syntax is n, where n is an integer, a decoded artwork is assigned, first decoded based on the decoding order, and stored in the DPB with 0 for a latency parameter. Then, the latency parameter of the first decoded artwork increases by 1 any time the following illustrations are decoded in accordance with the decoding order, and the first decoded and stored artwork is sent from the DPB when the latency parameter reaches n , that is, after an illustration encoded for the (n) th time is encoded based on the decoding order.
Figure 18 is a diagram to describe a procedure for sending a decoded illustration from i
U a DPB when using a syntax
XMPIv ^,
MEXICAN INSTITUTE 'OF PROPERTY C
INDUSTRIAL «» _-MaxLatencyFrames, in accordance with an embodiment of the present invention. In Figure 18, it is assumed that a maximum size (MaxDpbSize) of the DPB is 4, that is, the DPB is capable of storing a maximum of 4 decoded illustrations, and the MaxLatencyFrames syntax is 0.
Referring to Figure 18, since the MaxLatencyFrames syntax has a value of 0, the video decoding apparatus 200 can immediately send a decoded artwork. In Figure 18, the MaxLatencyFrames syntax has the value of 0 in an extreme case, but if the MaxLatencyFrames syntax has a value less than 4, a time point when the decoded artwork is sent from the DPB may move up compared to when the decoded artwork is sent from the DPB after being delayed 4 frames from a decode time through an augmentation procedure.
Meanwhile, a decoded artwork output time may move up as the MaxLatencyFrames syntax has a lower value, but since decoded artwork stored in the DPB must be rendered in accordance with a display order identical to that determined by a decoder, the decoded artwork should not be sent from the DPB until its display order is reached even if the encoded artwork is pre-encoded.
<img file="MX337918B_D0091.tif" />
Accordingly, the ___ video apparatus 100 determines a MaxLatencyFrames syntax that "'Tads a maximum latency frame based on a maximum value of a difference between an encoding order and a display order of each image frame while encoding each one of the image boxes that make up an image sequence inserts the MaxLatencyFrames syntax into a mandatory component of an SPS, and transmits the MaxLatencyFrames syntax to the image decoding device 200.
Alternatively, the video encoding apparatus 100 can insert a syntax to determine the MaxLatencyFrames syntax, and a syntax indicating the number of image frames required to be reordered within the SPS instead of directly inserting the MaxLatencyFrames syntax within the SPS. In detail, the video encoding apparatus 100 can determine a num_reorder_frames syntax indicating a maximum number of image frames required to be reordered since the image frames are first encoded based on an encoding order of between image frames that form an image sequence but are presented after post-encoded image frames based on a display order, and insert a difference value between the syntax
MaxLatencyFrames and the num_reorder_frames syntax, that is, a syntax value MaxLatencyFrames-syntax
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY num_reorder_frames, inside the SPS instead of a max_latency_increase syntax to determine the MaxLatencyFrarnes syntax. When the num_reorder_frames syntax and the max_latency_increase syntax are inserted inside and transmitted with the SPS instead of the MaxLatencyFrarnes syntax, the video decoding apparatus 200 can determine the MaxLatencyFrarnes syntax by using the value of (num_reorder_frames_ + max_latency_inc
Figures 19A through 19D are diagrams for describing a MaxLatencyFrarnes syntax and a num_reorder_frames syntax, in accordance with embodiments of the present invention. In Figures 19A to 19D, a POC denotes a display order, and an encoding order and decoding order of image frames that form an image sequence in an encoder and a decoder are the same. Also, arrows above illustrations FO through F9 in the image sequence indicate reference illustrations.
Referring to Figure 19A, Illustration F8 which is last in the display order and second encoded in the encoding order is an illustration having a larger difference value between the display order and the encoding order . Also, artwork F8 is required to be opened in order since artwork F8 is encoded before artwork F1 through F7 but
<img file="MX337918B_D0092.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY
<img file="MX337918B_D0093.tif" />
behind illustrations F2 to F7 on eT<sup>DUST</sup>order of presentation. Thus, the syntax mim_r eorde r_f ramé s corresponding to the image sequence shown in Figure 19A is 1. The video encoding apparatus 100 can be set to 7 which is the difference value between the display order and the order Illustration F8 encoding as a value in a syntax
MaxLatencyFrames, insert the value of the MaxLatencyFrames syntax as a mandatory component of an SPS, and transmit the value of the MaxLatencyFrames syntax to the 200 video decoding apparatus. Alternatively, the video encoding apparatus 100 can be set to 7 which is a difference value between 8 which is a value of a MaxLatencyFrames syntax and 1 which is a value of a num_reorder_frames syntax, such as a value of a max_latency_increase syntax, insert the num_reorder_frames syntax and max__latency_increase syntax as required components of an SPS instead of syntax
MaxLatencyFrames, and transmit the num_reorder_frames syntax and the max_latency_increase syntax to the 200 video decoding apparatus.
The video decoding device 200 can add the num_reorder_frames syntax and the max_latency_increase syntax transmitted with the SPS to determine the MaxLatencyFrames syntax, and determine a delay time.
<img file="MX337918B_D0094.tif" />
without any output from an alm decoded artwork when using the MaxLatencyFrames syntax augmentation procedure
In an image sequence in Figure 19B, differences between a display order and a coding order of all illustrations that exclude an illustration F0 are 1. Illustrations F2, F4, F6, and F8 are illustrations that have an order of slow encoding but have a fast order of presentation from among illustrations in the image sequence of Figure 19B, and thus require registration. There is only one illustration that has a slow encoding order but has a fast display order based on each of the illustrations F2, F4, F6, and F8. For example, there is only illustration F1 that has a slower encoding order but has a faster display order than illustration F2. Therefore, a value of a num_reorder_frames syntax in the image sequence in Figure 19B is 1. The video encoding apparatus 100 can set 1 as a value of a MaxLatencyFrames syntax, insert the value of the MaxLatencyFrames syntax as a mandatory component of an SPS, and transmit the value of the MaxLatencyFrames syntax to the 200 video decoding apparatus. Alternatively , the video encoding apparatus 100 can be set to 0 which is a value of
IMPI
<img file="MX337918B_D0095.tif" />
MEXICAN INSTITUTE OF THE TÍOPÍEOAD
INIXIJTSLAX___ difference between 1 which is a value of the syntax
MaxLatencyFrames and 1 which is a value of the num_reorder_frames syntax, as a value of a max_latency__increase syntax, insert the num_reorder_frames syntax and the max_latency_increase syntax as mandatory components of a SPS instead of the maxLatencyFrames syntax, and pass the syntax_lantand syntax. to the video decoding apparatus 200.
The video decoding apparatus 200 can add the num_reorder_frames syntax and max_latency_increase syntax transmitted with the SPS to determine the MaxLatencyFrames syntax, and determine an output time of a decoded artwork stored in the DPB by using the MaxLatencyFrames syntax without any augmentation procedure .
In an image sequence of Figure 19C, an illustration F8 that is last in a display order and second encoded in a coding order has a difference value of 7 between the display order and the coding order. . Therefore, a MaxLatencyFrames syntax is 7. Also, illustrations F4 and F8 are required to be rearranged since illustrations F4 and F8 are encoded and stored in the DPB before illustrations F1 through F3 based on the order of
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0096.tif" />
decoding but are presented after illustrations F1 through F3 based on display order, and thus a value of a num_reorder_frames syntax is 2. The video encoding apparatus 100 can set 7 as the value of the MaxLatencyFrames syntax, insert the value of the MaxLatencyFrames syntax as a mandatory component of an SPS, and transmit the value of the MaxLatencyFrames syntax to the 200 video decoding apparatus. Alternatively, the video encoding apparatus 100 can be set to 5 which is a difference value between 7 which is the value of the MaxLatencyFrames syntax and 2 which is the value of the max_latency_increase syntax, as a value of a max_latency_increase syntax, insert the num_reorder_frames syntax and max_latency_increase syntax as mandatory components of the SPS instead of MaxLatencyFrames, and transmit the num_reorder_frames syntax and the max_latency_increase syntax to the video decoding apparatus
200.
The video decoding apparatus 200 can add the num_reorder_frames syntax and max_latency_increase syntax transmitted with the SPS to determine the MaxLatencyFrames syntax, and determine the output time of a decoded artwork stored in the DPB by using the MaxLatencyFrames syntax without any augmentation procedure .
ί
MEXICAN INSTITUTE
OF PROPERTY i%,
INDUSTRIAL
In an image sequence of Figure 19D, illustrations F4 and F8 have a maximum value of 3 of a difference value between a display order and an encoding order. Therefore, a value of a MaxLatencyFrames syntax is 3. Also, F2 and F4 artwork is required to be reordered since F2 and F4 artwork is encoded before an FI artwork but is rendered after the FI artwork based on the order of presentation. Also, illustrations F6 and
F8 are rearranged since illustrations F6 and F8 are coded before an illustration F5 and are presented after the illustration F5 based on the order of presentation. Thus a value of a num_reorder_frames syntax is 2. Video encoding apparatus 100 can set 3 as the value of the MaxLatencyFrames syntax, insert the value of the MaxLatencyFrames syntax as a mandatory component of an SPS, and transmit the value of the MaxLatencyFrames syntax to the 200 video decoding apparatus. Alternatively, the video encoding apparatus 100 can be set to 1 which is a difference value between 3 which is the value of the MaxLatencyFrames syntax and 2 which is the value of the num_reorder_frames syntax, as a value of a max_latency_increase syntax, insert the syntax num__reorder_frames and the max_latency_increase syntax as components
<img file="MX337918B_D0097.tif" />
SPS mandatory instead of MaxLatencyFrames, and transmitting the num_reorder_frames syntax and the max_latency_increase syntax to the 200 video decoding apparatus.
The video decoding apparatus 200 can add the num_reorder_frames syntax and max_latency_increase syntax transmitted with the SPS to determine the MaxLatencyFrames syntax, and determine an output time of a decoded artwork stored in the DPB by using the MaxLatencyFrames syntax without any augmentation procedure .
Figure 2 0 is a flow chart illustrating an image encoding method in accordance with an embodiment of the present invention.
Referring to Figure 20, in operation 2010, the maximum encoding unit divisor 110 and the encoding unit determiner 120 (hereinafter commonly referred to as an encoder), which encode in a VCL of the encoding apparatus of video 100, determine a reference frame of each of the image frames that form an image sequence when performing prediction and compensation in motion, and encode each image frame using the given reference frame.
In operation 2020, the output unit 130 determines a maximum size of a buffer required to decode each image frame by a decoder, and the
<img file="MX337918B_D0098.tif" />
INSTITUTO MtXICANO Ol LA? ROPIi-.3AD INDUSTRIAL number of picture frames required to be reordered, based on a picture box coding order, a reference box decoding order indicated by the picture boxes, a display order of the image boxes, and an order of presentation of the reference tables. In detail, output unit 130 determines a max_dec_frame_buffering syntax indicating a maximum size of a DPB required to decode each picture frame by a decoder based on a picture frame encoding order (or decoding order) and a order of encoding (or order of decoding) of reference boxes indicated by the picture boxes, Inserts the max_dec_frame_buffering syntax into an SPS corresponding to header information for an image sequence, and transmits the max_dec_frame_buffering syntax to an encoder. As described above, output unit 130 includes the max_dec_frame_buffering syntax in the SPS as mandatory information instead of selective information.
In step 2030, output unit 130 determines latency information of an image frame having a larger difference between an encoding order and a display order among the image frames that make up the image sequence, based in the number of image frames required to be reordered.
<img file="MX337918B_D0099.tif" />
IMPI
MEXICAN INSTITUTE OF THE. INDUSTRIAL PROPERTY
In detail, output unit 130 determines a syntax
<td>MaxLatencyFrames</td><td>with</td><td>base</td><td>in</td><td>a value</td><td>Maximum of</td><td>a</td>
<td>difference between</td><td>a</td><td>order</td><td>of</td><td>coding</td><td>and an order</td><td>of</td>
<td colspan="2">presentation of each</td><td>picture</td><td>of</td><td colspan="2">image while encoding</td><td>the</td>
<td>picture boxes</td><td>than</td><td>they form</td><td>the</td><td>sequence of</td><td colspan="2">image. Too,</td>
output unit 130 can determine a num_reorder_frames syntax indicating a maximum number of image frames that are encoded first according to an encoding order based on a predetermined image frame from among the image frames in the image sequence and they are presented after a postcoded picture frame based on a display order, and thus require reordering, and insert a difference value between the MaxLatencyFrames syntax and the num__reorder_frames syntax, that is, a MaxLatencyFrames-num_reorder_frames syntax value, within an SPS as a max_latency_increase syntax to determine the max_latency_increase syntax. If the num_reorder_frames syntax and the max_latency_increase syntax indicating the MaxLatencyFrames-syntax num_reorder_frames syntax are included in and transmitted with the SPS, instead of the MaxLatencyFrames syntax, the MaxLatency value decoder 200 can determine the MaxLatency value syntax. Syntax MaxLatencyFrames-syntax num_reorder_frames.
<img file="MX337918B_D0100.tif" />
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX337918B_D0101.tif" />
At step 2040, output unit 130 generates a bit stream by including the max_dec_frame_buffering syntax, the num_reorder_frames syntax, and the max_latency_increase syntax as mandatory components of the SPS.
FIG. 21 is a flow chart illustrating an image decoding method in accordance with an embodiment of the present invention.
Referring to Figure 21, at step 2110, the image data and encoding information extractor 220 fetches a NAL unit of a NAL from a bit stream, and gets a max_dec_f rame__buf fing syntax indicating a maximum size of a buffer, a syntax num_reorder_frames indicating the number of image frames required to be reordered, and a max_latency_increase syntax to determine a MaxLatencyFrames syntax for the NAL unit that includes an SPS.
At step 2120, the DPB included in the image data decoder 230 sets the maximum buffer size required to decode the image sequence using the max_dec_frame_buffering syntax.
In step 2130, the image data and encoding information extractor 220 obtains encoded data from an image frame included in a VCL NAL unit, and outputs the encoded data obtained to the image data decoder 230. The data decoder of image
<img file="MX337918B_D0102.tif" />
JJqSTnV'O MEXICAN PROPERTY
230 it gets a decoded picture frame by ^ aerWdifrcHT the encoded picture data.
At step 2140, the DPB of the image data decoder 230 stores the decoded image frame.
At step 2150, the DPB determines whether to send the stored decoded image box using the num_reorder_f rames syntax and the max_latency_increase syntax. In detail, the DPB determines the MaxLatencyFrames syntax by adding the num_reorder_f rames syntax and the max_latency_increase syntax. The DPB sets a predetermined latency parameter for each decoded and stored image frame, increases a count of the default latency parameter by one any time an image frame in the image sequence is encoded according to an encoding order, and sends the decoded image box whose default latency parameter count reached syntax
MaxLatencyFrames.
The invention can also be represented as computer readable codes on a computer readable recording medium. The computer-readable recording medium is any data storage device that can store data that can then be read by a computer system. Examples of the computer-readable recording medium include memory-only
<img file="MX337918B_D0103.tif" />
INSTITUTO MEXICANO DE LA fSClPÍEDAD INDUSTRIAL reading (ROM), random access memory (RAM), CD-ROM, magnetic tapes, floppy disks, optical data storage devices, etc. The computer-readable recording medium can also be distributed over network-attached computer systems so that the computer-readable code is stored and executed in a distributed form.
Although the invention has been particularly shown and described with reference to illustrative embodiments thereof, it will be understood by those with basic knowledge in the art that various changes in form and detail may be made there without departing from the spirit and scope of the invention as defined. by the appended claims. Illustrative modalities should be considered in a descriptive sense only and not for purposes of limitation. Therefore, the scope of the invention is defined not by the detailed description of the invention but by the appended claims, and all differences within the scope will be interpreted as being included in the present invention.
It is noted that in relation to this date, the best method known by the applicant to put the aforementioned invention into practice, is the one that is clear from the present description of the invention.
Contents107
123 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111 Sheet 112 Sheet 113 Sheet 114 Sheet 115 Sheet 116 Sheet 117 Sheet 118 Sheet 119 Sheet 120 Sheet 121 Sheet 122 Sheet 123
67 members in 13 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 201161563678 | United States of America | P | |
| 201161563678 | United States of America | P | |
| 61563678 | United States of America | – | |
| 1020120034093 | Republic of Korea | – | |
| 20120034093 | Republic of Korea | A | |
| 20120034093 | Republic of Korea | A | |
| 2012009972 | Republic of Korea | W | |
| 2012009972 | Republic of Korea | W | |
| 1020120034093 | – | – | – |
| 61563678 | – | – | – |
| KR1209972 | – | – | – |
| KR20120034093 | – | – | – |
| US201161563678P | – | – | – |
| WO2012KR09972 | – | – | – |
Members67
| Document | Office | Kind | |
|---|---|---|---|
| CA2856906A1 | Canada | A1 | |
| CA2995095A1 | Canada | A1 | |
| WO2013077665A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20130058584A | Republic of Korea | A | |
| TW201330634A | Taiwan Province of China | A | |
| MX2014006260A | Mexico | A | |
| AU2012341242A1 | Australia | A1 | |
| US2014269899A1 | United States of America | A1 | |
| CN104067618A | China | A | |
| EP2785054A1 | European Patent Office (EPO) | A1 | |
| JP2014533917A | Japan | A | |
| IN1173MUN2014A | India | A | |
| EP2785054A4 | European Patent Office (EPO) | A4 | |
| MX337915B | Mexico | B | |
| MX337916B | Mexico | B | |
| MX337917B | Mexico | B | |
| MX337918BThis record | Mexico | B | |
| TWI542199B | Taiwan Province of China | B | |
| TW201631971A | Taiwan Province of China | A | |
| US9438901B2 | United States of America | B2 | |
| US2016337654A1 | United States of America | A1 | |
| US9560370B2 | United States of America | B2 | |
| US2017105001A1 | United States of America | A1 | |
| TWI587690B | Taiwan Province of China | B | |
| BR112014012549A2 | Brazil | A2 | |
| US9699471B2 | United States of America | B2 | |
| TW201725906A | Taiwan Province of China | A | |
| US2017223362A1 | United States of America | A1 | |
| US9769483B2 | United States of America | B2 | |
| CN104067618B | China | B | |
| US2017347101A1 | United States of America | A1 | |
| CN107465931A | China | A | |
| CN107465932A | China | A | |
| CN107483965A | China | A | |
| CN107517387A | China | A | |
| CN107517388A | China | A | |
| CA2856906C | Canada | C | |
| US9967570B2 | United States of America | B2 | |
| TWI625055B | Taiwan Province of China | B | |
| TW201826790A | Taiwan Province of China | A | |
| US2018249161A1 | United States of America | A1 | |
| US10218984B2 | United States of America | B2 | |
| ZA201706409B | South Africa | B | |
| US2019158848A1 | United States of America | A1 | |
| KR20190065990A | Republic of Korea | A | |
| US10499062B2 | United States of America | B2 | |
| CN107483965B | China | B | |
| CN107465932B | China | B | |
| KR20200037196A | Republic of Korea | A | |
| KR102099915B1 | Republic of Korea | B1 | |
| TWI693819B | Taiwan Province of China | B | |
| KR102135966B1 | Republic of Korea | B1 | |
| CN107517388B | China | B | |
| TW202029756A | Taiwan Province of China | A | |
| CN107465931B | China | B | |
| CN107517387B | China | B | |
| KR20200096179A | Republic of Korea | A | |
| TWI703855B | Taiwan Province of China | B | |
| CA2995095C | Canada | C | |
| KR102221101B1 | Republic of Korea | B1 | |
| KR20210024516A | Republic of Korea | A | |
| KR102321365B1 | Republic of Korea | B1 | |
| KR20210133928A | Republic of Korea | A | |
| ZA201706411B | South Africa | B | |
| ZA201706412B | South Africa | B | |
| KR102400542B1 | Republic of Korea | B1 | |
| ZA201706410B | South Africa | B |
Numbers
- Publication
- 337918
- Publication, DOCDB
- 337918
- Publication, EPODOC
- MX337918
- Application
- 2015006962
- Application, DOCDB
- 2015006962
- Application, EPODOC
- MX20150006962
Titles2
- Spanish
- METODO Y DISPOSITIVO DE CODIFICACION DE IMAGEN PARA MANEJO DE BUFER DE DECODIFICADOR, Y METODO Y DISPOSITIVO DE DECODIFICACION DE IMAGEN.
- English
- IMAGE CODING METHOD AND DEVICE FOR BUFFER MANAGEMENT OF DECODER, AND IMAGE DECODING METHOD AND DEVICE.
Classification
- CPC, 5
- H04N19/152
- H04N19/44
- H04N19/70
- H04N19/105
- H04N19/42
- IPC, 4
- H04N19 10
- H04N19 61
- H04N19 70
- H04N19 152