Method and apparatus for video encoding with arithmectic encoding method and apparatus for video decoding with arithmectic decoding
Abstract
The present invention discloses a video decoding method through symbol decoding. Parsing the symbols of image blocks from the received bitstream, classifying the current symbol into a prefix bit stream and a suffix bit string based on a threshold determined based on the size of the current block, and separately for the prefix bit string and the suffix bit string Disclosed is a video decoding method in which arithmetic decoding is performed according to each determined arithmetic decoding method, and inverse binarization is performed on a prefix bit string and a suffix bit string after arithmetic decoding according to each individually determined binarization method.

Term
8.5 yearsleft in the term
Expires 23 March 2035.
- Priority
- Filed
- Granted
- Today
- Expires
4 claims: 3 independent, 1 dependent
- 1비디오 복호화 방법에 있어서, 현재 블록의 인트라 예측 모드에 대한 정보를 포함하는 비트스트림을 수신하는 단계;상기 비트스트림이 하나의 비트를 포함하는 경우, 상기 하나의 비트에 대해 컨텍스트-기초-산술복호화를 수행하여 인트라 예측 모드 이진열을 획득하는 단계;상기 비트스트림이 복수 개의 비트들을 포함하는 경우, 상기 비트스트림의 첫번째 비트에 대해 컨텍스트-기초-산술복호화를 수행하고, 상기 비트스트림 중 상기 첫번째 비트를 제외한 적어도 하나의 비트에 대해 바이패스 모드 복호화를 수행하여 인트라 예측 모드 이진열을 획득하는 단계;상기 인트라 예측 모드 이진열에 대해 하나의 이진화 방식에 따른 역이진화를 수행하여 상기 현재 블록의 인트라 예측 방향을 나타내는 심볼을 복원하는 단계를 포함하고, 상기 인트라 예측 모드는 상기 현재 블록의 크로마 성분들에 대한 인트라 예측 방향을 나타내는 인트라 크로마 예측 모드를 포함하는 것을 특징으로 하는 비디오 복호화 방법.
- 2삭제
- 3제 1 항에 있어서, 상기 비디오 복호화 방법은, 상기 복원된 심볼을 이용하여 상기 현재 블록의 인트라 예측 방향을 결정하는 단계;및 상기 결정된 인트라 예측 방향에 따라 상기 현재 블록에 대해 인트라 예측을 수행하는 단계를 포함하는 것을 특징으로 하는 비디오 복호화 방법.
- 4비디오 복호화 장치에 있어서, 현재 블록의 인트라 예측 모드에 대한 정보를 포함하는 비트스트림을 수신하는 수신부;상기 비트스트림이 하나의 비트를 포함하는 경우, 상기 하나의 비트에 대해 컨텍스트-기초-산술복호화를 수행하여 인트라 예측 모드 이진열을 획득하고, 상기 비트스트림이 복수 개의 비트들을 포함하는 경우, 상기 비트스트림의 첫번째 비트에 대해 컨텍스트-기초-산술복호화를 수행하고, 상기 비트스트림 중 상기 첫번째 비트를 제외한 적어도 하나의 비트에 대해 바이패스 모드 복호화를 수행하여 인트라 예측 모드 이진열을 획득하는 산술복호화부;상기 인트라 예측 모드 이진열에 대해 하나의 이진화 방식에 따른 역이진화를 수행하여 상기 현재 블록의 인트라 예측 방향을 나타내는 심볼을 복원하는 역이진화부를 포함하고, 상기 인트라 예측 모드는 상기 현재 블록의 크로마 성분들에 대한 인트라 예측 방향을 나타내는 인트라 크로마 예측 모드를 포함하는 것을 특징으로 하는 비디오 복호화 장치.
Independent claims4
277 paragraphs, as filed
BACKGROUND OF THE INVENTION Field of the Invention: A video encoding method accompanied by arithmetic encoding, an apparatus therefor, a video decoding method, and an apparatus thereof TECHNICAL FIELD
The present invention relates to video encoding and decoding involving arithmetic encoding.
With the development and dissemination of hardware capable of reproducing and storing high-resolution or high-definition video content, the need for a video codec for effectively encoding or decoding high-resolution or high-definition video content is increasing. According to the existing video codec, a video is encoded according to a limited encoding method based on a macroblock of a predetermined size.
Image data in the spatial domain is transformed into coefficients in the frequency domain using frequency transform. The video codec divides an image into blocks of a predetermined size for fast frequency transformation and performs DCT transformation for each block to encode frequency coefficients in units of blocks. Compared to image data in the spatial domain, the coefficients in the frequency domain have a form that is easy to compress. In particular, since an image pixel value in a spatial domain is expressed as a prediction error through inter prediction or intra prediction of a video codec, when frequency conversion is performed on the prediction error, a lot of data may be converted to zero. The video codec reduces the amount of data by replacing continuously and repeatedly occurring data with data of a small size.
The present invention proposes a video encoding method and apparatus for classifying symbols into prefix and suffix bit streams and performing arithmetic encoding, and a video decoding method and apparatus thereof.
A video decoding method through symbol decoding according to an embodiment of the present invention includes parsing symbols of image blocks from a received bitstream; classifying the current symbol into a prefix bit stream and a suffix bit stream based on a threshold determined based on the size of the current block; performing arithmetic decoding on the prefix bit string and the suffix bit string according to each individually determined arithmetic decoding method; performing inverse binarization on the prefix bit string and the suffix bit string after the arithmetic decoding according to each separately determined binarization method; and reconstructing image blocks by performing inverse transform and prediction on the current block using the current symbol reconstructed through the arithmetic decoding and the inverse binarization.
According to an embodiment, the performing of the inverse binarization may include performing inverse binarization according to each binarization method individually determined for the prefix bit string and the suffix bit string to obtain a prefix area and a suffix area of the symbol. It may include a step of restoring.
According to an embodiment, the performing of the arithmetic decoding may include: performing arithmetic decoding for determining context modeling for each bit position on the prefix bit string; and performing arithmetic decoding for omitting the context modeling by applying a bypass mode to the suffix bit string.
According to an embodiment, the performing of the arithmetic decoding includes performing the arithmetic decoding using the context of a predetermined index separately allocated in advance for each bit position of the prefix bit string, when the symbol is the final coefficient position information of the transform coefficient. may include the step of
According to an embodiment, the current symbol may include at least one of an intra prediction mode and final coefficient position information of the current block.
According to an embodiment, the binarization method includes a unary binarization method, a truncated unary binarization method, an Exponential Golomb Binarization method, and a fixed length binarization method. It may further include at least any one of.
A video encoding method through symbol encoding according to an embodiment of the present invention includes generating symbols by performing prediction and transformation on blocks of an image; classifying the current symbol into a prefix region and a suffix region based on a threshold determined based on the size of the current block; generating a prefix bit stream and a suffix bit stream by applying each separately determined binarization method to the prefix area and the suffix area; performing symbol encoding by applying each individually determined arithmetic encoding method to the prefix bit string and the suffix bit string; and outputting the bitstreams generated by the symbol encoding in the form of a bitstream.
The performing of the symbol encoding according to an embodiment may include: performing the symbol encoding by applying an arithmetic encoding method of performing context modeling for each bit position with respect to the prefix bit string; and performing the symbol encoding by applying an arithmetic encoding method for omitting the context modeling by applying a bypass mode to the suffix bit string.
According to an embodiment, the performing of the symbol encoding may include performing the arithmetic encoding using the context of a predetermined index separately pre-allocated for each bit position of the prefix bit string when the symbol is the final coefficient position information of the transform coefficient. may include the step of
According to an embodiment, the current symbol may include at least one of an intra prediction mode and final coefficient position information of the current block.
According to an embodiment, the binarization method may further include at least one of a unary binarization method, a truncated single type binarization method, an exponential Gollum coupling type binarization method, and a fixed-length binarization method.
A video decoding apparatus through symbol decoding according to an embodiment of the present invention includes: a parsing unit that parses symbols of image blocks from a received bitstream; Classify the prefix bit string and the suffix bit string of the current symbol based on a threshold determined based on the size of the current block, and perform arithmetic decoding according to each arithmetic decoding method individually determined for the prefix bit string and the suffix bit string a symbol decoding unit that performs inverse binarization on the prefix bit string and the suffix bit string according to each separately determined binarization method; and an image restoration unit that reconstructs image blocks by performing inverse transform and prediction on the current block using the current symbol restored through the arithmetic decoding and inverse binarization.
A video encoding apparatus through symbol encoding according to an embodiment of the present invention includes: an image encoder for generating symbols by performing prediction and transformation on blocks of an image; The prefix bit stream and the suffix bit stream of the current symbol are classified based on a threshold determined based on the size of the current block, and each binarization method determined separately for the prefix area and the suffix area is applied to form the prefix bit string and the a symbol encoding unit generating a fixed bit string and performing symbol encoding by applying each arithmetic encoding method individually determined to the prefix bit string and the suffix bit string; and a bitstream output unit for outputting bitstreams generated by the symbol encoding in the form of a bitstream.
Disclosed is a computer-readable recording medium in which a program for computationally realizing a video decoding method according to an embodiment of the present invention is recorded. Disclosed is a computer-readable recording medium in which a program for computationally realizing a video encoding method according to an embodiment of the present invention is recorded.
1 is a block diagram of a video encoding apparatus according to an embodiment of the present invention. 2 is a block diagram of a video decoding apparatus according to an embodiment of the present invention. 3 and 4 show embodiments of arithmetic coding by classifying symbols into a prefix bit stream and a suffix bit stream according to a predetermined threshold, respectively. 5 is a flowchart of a video encoding method according to an embodiment of the present invention. 6 is a flowchart of a video decoding method according to an embodiment of the present invention. 7 is a block diagram of a video encoding apparatus based on coding units having a tree structure according to an embodiment of the present invention. 8 is a block diagram of a video decoding apparatus based on coding units having a tree structure according to an embodiment of the present invention. 9 illustrates a concept of a coding unit according to an embodiment of the present invention. 10 is a block diagram of an image encoder based on coding units according to an embodiment of the present invention. 11 is a block diagram of an image decoder based on a coding unit according to an embodiment of the present invention. 12 illustrates coding units and partitions for each depth according to an embodiment of the present invention. 13 illustrates a relationship between a coding unit and a transformation unit according to an embodiment of the present invention. 14 illustrates encoding information for each depth according to an embodiment of the present invention. 15 illustrates a coding unit for each depth according to an embodiment of the present invention. 16, 17, and 18 illustrate a relationship between a coding unit, a prediction unit, and a transformation unit according to an embodiment of the present invention. 19 illustrates a relationship between a coding unit, a prediction unit, and a transformation unit according to the coding mode information of Table 1. Referring to FIG.
Hereinafter, a video encoding technique and a video decoding technique accompanying arithmetic encoding according to an embodiment are disclosed with reference to FIGS. 1 to 6 . Also, an embodiment in which arithmetic encoding is performed in a video encoding technique and a video decoding technique based on a coding unit having a tree structure according to an embodiment is disclosed with reference to FIGS. 7 to 19 . Hereinafter, 'image' may represent a still image of a video or a moving image, that is, a video itself.
First, a video encoding technique and a video decoding technique based on a prediction method of an intra prediction mode according to an embodiment are disclosed with reference to FIGS. 1 to 6 .
1 is a block diagram of a video encoding apparatus 10 according to an embodiment of the present invention.
The video encoding apparatus 10 may encode video data in the spatial domain through intra prediction/inter prediction, transformation, quantization, and symbol encoding. Hereinafter, operations occurring while the video encoding apparatus 10 symbolizes symbols generated through intra prediction/inter prediction, transformation, and quantization through arithmetic encoding will be described in detail.
The video encoding apparatus 10 according to an embodiment includes an image encoder 12 , a symbol encoder 14 , and a bitstream output unit 16 .
The video encoding apparatus 10 according to an embodiment may divide image data of a video into a plurality of data units, and may encode each data unit. The shape of the data unit may be a square or a rectangle, and may have any geometric shape. It is not limited to a data unit of a certain size. According to the video encoding method based on the coding unit according to the tree structure, the data unit may be a maximum coding unit, a coding unit, a prediction unit, a transformation unit, or the like. An example in which the arithmetic encoding/decoding method according to an embodiment is applied in the video encoding/decoding method based on coding units according to the tree structure will be described below with reference to FIGS. 7 to 19 .
For convenience of description, a video encoding technique for a 'block', which is a type of data unit, will be described below. However, the video encoding method according to various embodiments of the present invention should not be construed as being limited to the video encoding method for a 'block', but may be applied to various data units.
The image encoder 12 generates symbols by performing operations such as intra prediction/inter prediction, transformation, and quantization on blocks of an image.
The symbol encoder 14 classifies a prefix region and a suffix region of the current symbol based on a threshold determined based on the size of the current block in order to encode the symbol of the current symbol among symbols generated for each block. The symbol encoder 14 may determine a threshold for classifying the prefix region and the suffix region of the current symbol based on at least one of a width and a height of the current block.
The symbol encoder 14 may individually determine a symbol encoding scheme for the prefix region and the suffix region of the symbol, and encode the prefix region and the suffix region according to each symbol encoding scheme.
Symbol encoding can be classified into a binarization process for converting a symbol into a bit string, and an arithmetic encoding process for performing context-based arithmetic encoding on a bit string. The symbol encoder 14 may individually determine each binarization scheme for the prefix region and the suffix region of the symbol, and perform binarization for the prefix region and the suffix region individually according to each binarization scheme. . A prefix bit string may be generated from the prefix area, and a suffix bit string may be generated from the suffix area.
Alternatively, the symbol encoding unit 14 individually determines each arithmetic encoding method for the prefix bit string and the suffix bit string of the symbol, and separately for the prefix bit string and the suffix bit string according to each arithmetic encoding method. Encoding may also be performed.
In addition, the symbol encoding unit 14 performs binarization by individually determining each binarization method for the prefix area and the suffix area of the symbol, and separately performs each arithmetic encoding method for the prefix bit string and the suffix bit string. It can also be determined to perform arithmetic encoding.
The symbol encoder 14 according to an embodiment may individually determine a binarization scheme for the prefix region and the suffix region. The binarization schemes individually determined for the prefix region and the suffix region may be different from each other.
The symbol encoding unit 14 may individually determine each arithmetic encoding method for the prefix bit string and the suffix bit string. The arithmetic encoding schemes individually determined for the prefix bit string and the suffix bit string may be different from each other.
Therefore, in the symbol decoding process of the symbol, the symbol encoder 14 binarizes the prefix region and the suffix region according to different methods only for the binarization process, or converts the prefix bit stream and the suffix bit stream to each other only for the arithmetic encoding process. It may be encoded according to different schemes. Also, the symbol encoder 14 may encode the prefix region (prefix bit stream) and the suffix region (suffix bit stream) in both the binarization process and the arithmetic encoding process according to different schemes.
The binarization method that can be selected according to an embodiment includes, in addition to the general binarization method, a unary binarization method, a truncated unary binarization method, an exponential Golomb binarization method, and It may be at least one of fixed length binarization.
The symbol encoding unit 14 according to an embodiment applies an arithmetic encoding method of performing context modeling for each bit position to a prefix bit string, and applies a bypass mode to a suffix bit string to omit context modeling. By applying the method, symbol encoding of a symbol may be performed.
The symbol encoding unit 14 according to an embodiment performs symbol encoding individually by classifying a prefix region and a suffix region on symbols including at least one of an intra prediction mode and final coefficient position information of a transform coefficient. can do.
The symbol encoder 14 according to an embodiment may perform arithmetic encoding using the context of a predetermined index pre-allocated for the prefix bit string. For example, when the symbol is the final coefficient position information of the transform coefficient, the symbol encoding unit 14 according to an embodiment performs arithmetic encoding using the context of a predetermined index separately allocated in advance for each bit position of the prefix bit string. may be
The bitstream output unit 16 outputs bitstreams generated by symbol encoding in the form of a bitstream.
Therefore, the video encoding apparatus 10 according to an embodiment may arithmetic-encode and output symbols of blocks of a video.
The video encoding apparatus 10 according to an embodiment may include a central processor (not shown) that collectively controls the image encoder 12 , the symbol encoder 14 , and the bitstream output unit 16 . . Alternatively, the image encoding unit 12, the symbol encoding unit 14, and the bitstream output unit 16 are each operated by their own processors (not shown), and as the processors (not shown) operate organically, the video The encoding apparatus 10 may be operated as a whole. Alternatively, the image encoder 12 , the symbol encoder 14 , and the bitstream output unit 16 may be controlled under the control of an external processor (not shown) of the video encoding apparatus 10 according to an embodiment. have.
The video encoding apparatus 10 according to an embodiment includes one or more data storage units (not shown) for storing input/output data of the image encoder 12 , the symbol encoder 14 , and the bitstream output unit 16 . may include The video encoding apparatus 10 may include a memory controller (not shown) that controls data input/output of a data storage unit (not shown).
The video encoding apparatus 10 according to an embodiment performs a video encoding operation including prediction and transformation by operating in conjunction with an internally mounted video encoding processor or an external video encoding processor to output a video encoding result. can The internal video encoding processor of the video encoding apparatus 10 according to an embodiment performs a basic video encoding operation by including a video encoding processing module in the video encoding apparatus 10 or the central processing unit and the graphic processing unit as well as a separate processor. It may include the case of implementing .
2 is a block diagram of a video decoding apparatus according to an embodiment of the present invention.
The video decoding apparatus 20 decodes the video data encoded by the video encoding apparatus 10 through parsing, symbol decoding, inverse quantization, inverse transformation, intra prediction/motion compensation, etc. to obtain a video close to the original video data in the spatial domain. data can be restored. Hereinafter, a process in which the video decoding apparatus 20 performs arithmetic decoding on the symbols parsed from the bitstream to reconstruct the symbols will be described in detail.
The video decoding apparatus 20 according to an embodiment includes a parsing unit 22 , a symbol decoding unit 24 , and an image restoration unit 26 .
The video decoding apparatus 20 may receive a bitstream in which encoded data of a video is recorded. The parsing unit 22 may parse the symbols of the image blocks from the bitstream.
The parsing unit 20 according to an embodiment may parse encoded symbols through arithmetic encoding on video blocks from the bitstream.
The parsing unit 22 may parse symbols including an intra prediction mode of a video block, final coefficient position information of a transform coefficient, and the like from the received bitstream.
The symbol decoding unit 24 determines a threshold for classifying the current symbol into a prefix bit stream and a suffix bit stream. The symbol decoder 24 may determine a threshold for classifying the prefix bit stream and the suffix bit string based on the size of the current block, that is, at least one of a width and a height. The symbol decoding unit 24 individually determines an arithmetic decoding method for the prefix bit string and the suffix bit string. The symbol decoding unit 24 performs symbol decoding by applying each individually determined arithmetic decoding method to the prefix bit string and the suffix bit string.
The arithmetic decoding methods individually determined for the prefix bit string and the suffix bit string may be different from each other.
The symbol decoding unit 24 may individually determine a binarization method for a prefix bit string and a suffix bit string of a symbol. Accordingly, inverse binarization may be performed on each of the symbol prefix bit strings according to the individually determined binarization method. The binarization schemes individually determined for the prefix bit string and the suffix bit string may be different from each other.
In addition, the symbol decoding unit 24 performs arithmetic decoding by applying each individually determined arithmetic decoding method to the prefix bit string and the suffix bit string of the symbol, and the prefix bit string and the suffix bit string generated through the arithmetic decoding. Inverse binarization may also be performed for each column according to an individually determined binarization method.
Accordingly, the symbol decoding unit 24 decodes the prefix bit string and the suffix bit string according to different schemes only for the arithmetic decoding process in the symbol decoding process of the symbol, or the prefix bit sequence and the suffix bit sequence only for the inverse binarization process. Columns can also be de-binarized in different ways. Also, the symbol decoding unit 24 may decode the prefix bit string and the suffix bit string according to different schemes in both the arithmetic decoding process and the inverse binarization process.
* The binarization method, which is individually determined for the prefix bit string and the suffix bit string of the symbol, is at least any one of the unary binarization method, the truncation type single type binarization method, the exponential Gollum combination type binarization method, and the fixed-length binarization method as well as the general binarization method. can be
The symbol decoding unit 24 may apply an arithmetic decoding method of performing context modeling for each bit position with respect to the prefix bit string. The symbol decoder 24 may apply an arithmetic decoding method in which context modeling is omitted by applying a bypass mode to the suffix bit string. Accordingly, the symbol decoding unit 24 may perform symbol decoding through arithmetic decoding separately performed on the prefix bit string and the suffix bit string of the symbol.
The symbol decoder 24 may perform arithmetic decoding by classifying a prefix bit stream and a suffix bit string on symbols including at least one of an intra prediction mode and final coefficient position information of a transform coefficient.
When the symbol is the final coefficient position information of the transform coefficient, the symbol decoding unit 24 may perform arithmetic decoding using the context of a predetermined index separately allocated in advance for each bit position of the prefix bit string.
The image restoration unit 26 may reconstruct a prefix region and a suffix region of a symbol as a result of individually performing arithmetic decoding and inverse binarization on the prefix bit stream and the suffix bit string. The image reconstructor 26 may reconstruct the symbol by synthesizing the prefix region and the suffix region of the symbol.
The image restoration unit 26 performs inverse transformation and prediction on the current block using the current symbol restored through arithmetic decoding and inverse binarization. The image reconstructor 26 may reconstruct image blocks by performing operations such as inverse quantization, inverse transformation, and intra prediction/motion compensation using corresponding symbols for each image block.
The video decoding apparatus 20 according to an embodiment may include a central processor (not shown) that collectively controls the parsing unit 22 , the symbol decoding unit 24 , and the image restoration unit 26 . Alternatively, the parsing unit 22 , the symbol decoding unit 24 , and the image restoration unit 26 are operated by their own processors (not shown), and as the processors (not shown) operate organically, the video decoding apparatus (20) may be operated as a whole. Alternatively, the parsing unit 22 , the symbol decoding unit 24 , and the image restoration unit 26 may be controlled under the control of an external processor (not shown) of the video decoding apparatus 20 according to an embodiment.
The video decoding apparatus 20 according to an embodiment may include one or more data storage units (not shown) for storing input/output data of the parsing unit 22 , the symbol decoding unit 24 , and the image restoration unit 26 . can The video decoding apparatus 20 may include a memory controller (not shown) that controls data input/output of a data storage unit (not shown).
The video decoding apparatus 20 according to an embodiment performs a video decoding operation including inverse transform by operating in conjunction with a video decoding processor mounted therein or an external video decoding processor to reconstruct a video through video decoding. can The internal video decoding processor of the video decoding apparatus 20 according to an embodiment includes a video decoding processing module as well as a separate processor, so that the video decoding apparatus 20 or the central processing unit and the graphic processing unit include a video decoding processing module. It may include the case of implementing .
As a context-based arithmetic encoding decoding method for symbol encoding and decoding, Context-based Adaptive Binary Arithmetic Coding (CABAC) is widely used. According to context-based arithmetic encoding decoding, each bit of a symbol bit string becomes each bin of a context, and each bit position may be mapped to a bin index. The length of the bit string, that is, the length of the bins may vary according to the size of the symbol value. For context-based arithmetic code decoding, context modeling for determining the context of a symbol is required. For context modeling, a complex operation process is required because the context is newly updated for each bit position of the symbol bit string, that is, for each bin index.
According to the video encoding apparatus 10 and the video decoding apparatus 20 described above with reference to FIGS. 1 and 2 , a symbol is divided into a prefix region and a suffix region, and relatively simple binarization is performed for the suffix region compared to the prefix region. method can be applied. In addition, since arithmetic code decoding is performed through context modeling for the prefix bit string and context modeling is omitted for the suffix bit string, the burden of the amount of computation for context-based arithmetic decoding can be reduced. Accordingly, in the video encoding apparatus 10 and the video decoding apparatus 20 according to an embodiment, the computational burden on the suffix region or the suffix bit stream is relatively small during the context-based arithmetic encoding decoding process for symbol encoding and decoding. By performing the binarization method or omitting the context modeling, the efficiency of the symbol encoding/decoding process can be improved.
Hereinafter, various embodiments for arithmetic encoding that can be implemented in the video encoding apparatus 10 and the video decoding apparatus 20 according to an embodiment will be described in detail.
3 and 4 show embodiments of arithmetic coding by classifying symbols into a prefix bit stream and a suffix bit stream according to a predetermined threshold, respectively.
Referring to FIG. 3 , a process of performing symbol encoding on final coefficient position information among symbols according to an embodiment will be described in detail. The final coefficient position information is a symbol indicating the position of the last non-zero coefficient among the transform coefficients of the block. Since the block size is defined by width and height, the final count position information can be expressed as two-dimensional coordinates of x-coordinate values in the width direction and y-coordinate values in the height direction. 3 illustrates a method of symbol encoding an x-coordinate value in a width direction among final coefficient position information when the width of a block is w for convenience of explanation.
Since the range of the x-coordinate value of the final counting position information is within the width of the block, the x-coordinate value of the final counting position information is greater than or equal to 0 and less than or equal to w-1. For arithmetic encoding of a symbol according to an embodiment, a symbol may be classified into a prefix region and a suffix region based on a predetermined threshold th. Accordingly, arithmetic encoding may be performed on the prefix bit string in which the prefix region is binarized based on a context determined through context modeling. In addition, arithmetic encoding may be performed on the suffix bit stream in which the suffix region is binarized according to a bypass mode in which context modeling is omitted.
In this case, the threshold th for classifying the prefix region and the suffix region of the symbol may be determined based on the block width w. As a simple example, the threshold value th may be determined to be (w/2)-1 to halve the bit stream (threshold value determination equation 1). As another example, since the block width w generally has a power of 2, the threshold th may be determined based on the log value of w (threshold value determination equation 2).
<Threshold Determination Formula 1> th = (w/2) - 1;
<Threshold Determination Equation 2> th = (log2w << 1) - 1;
In FIG. 3, when the block width w is 8, the threshold value becomes th = (8/2) - 1 = 3 according to the threshold determination formula 1, so 3 of the x-coordinate values of the final count position information are classified as a prefix area, and , the remaining values except for 3 among the x-coordinate values of the final counting position information may be classified as a suffix area. The prefix region and the suffix region may be binarized according to an individually determined binarization scheme.
When the x-coordinate value N of the current final counting position information is 5, the x-coordinate value of the final counting position information may be classified as N = th + 2 = 3 + 2. That is, among the x-coordinate values of the final counting position information, 3 may be classified as a prefix area, and 2 may be classified as a suffix area.
According to an embodiment, the prefix region and the suffix region may be separately binarized according to different binarization schemes. For example, the prefix region may be binarized according to the unary binarization scheme, and the suffix region may be binarized according to the general binarization scheme.
Accordingly, as a result of binarizing 3 according to the unary binarization method, a prefix bit string 32 '0001' is generated from the prefix area, and as a result of binarizing 2 according to the general binarization method, the suffix bit string 34' is generated from the suffix area. 010' may be generated.
Also, context-based arithmetic encoding may be performed on the prefix bit string 32 '0001' through context modeling. Accordingly, a context index may be determined for each bin of '0001'.
Arithmetic encoding may be performed on the suffix bit string 34 '010' without context modeling according to the bypass mode. In the bypass mode, it is assumed that each bin has an equal probability state, that is, a context of 50%, so that arithmetic encoding may be performed while omitting context modeling.
Accordingly, the context-based arithmetic encoding is individually performed on the prefix bit string 32 '0001' and the suffix bit string 34 '010', thereby completing the symbol encoding of the x-coordinate value N of the current final count position information. can be
In addition, although the embodiment in which symbol encoding is performed through binarization and arithmetic encoding has been described above, symbol decoding may also be performed in the same principle. That is, the parsed symbol bit string is classified into a prefix bit string and a suffix bit string based on the block width w, arithmetic decoding is performed on the prefix bit string 32 through context modeling, and the suffix bit string 34 is performed. For , arithmetic decoding may be performed while omitting context modeling. Inverse binarization is performed on the prefix bit string 32 after arithmetic decoding according to the unary binarization method, the prefix region is restored, and the suffix bit string 34 after arithmetic encoding is inversely binarized according to the general binarization method. The fixed area may be restored. A symbol may be reconstructed by synthesizing the reconstructed prefix area and the suffix area.
Previously, an embodiment in which the unary binarization method is applied to the prefix region (prefix bit string) and the general binarization method is applied to the suffix region (suffix bit string) has been disclosed, but the binarization method is not limited thereto. As another example, the truncation unary binarization method may be applied to the prefix region (prefix bit string), and the fixed-length binarization method may be applied to the suffix region (suffix bit string).
Although only the embodiment of the final counting position information along the block width direction has been described above, the above embodiment can be similarly applied to the final counting position information along the block height direction.
In addition, context modeling is not required for a suffix bit string that is arithmetic-encoded using a context of a fixed probability, but variable context modeling is required for a prefix bit string. According to an embodiment, context modeling for a prefix bit string may be determined according to a block size.
<tables num="0"><table><tgroup cols="2"><colspec align="center" colname="col1" colnum="1" colwidth="1872" /><colspec align="center" colname="col2" colnum="2" colwidth="8919" /><tbody><row><entry align="center" colname="col1"> block size</entry><entry align="center" colname="col2"> Empty index number of the adopted context</entry></row><row><entry align="center" colname="col1">4x4</entry><entry align="justify" colname="col2">0, 1, 2, 2</entry></row><row><entry align="center" colname="col1">8x8</entry><entry align="justify" colname="col2">3, 4, 5, 5</entry></row><row><entry align="center" colname="col1">16x16</entry><entry align="justify" colname="col2">6, 7, 8, 9, 10, 10, 11, 11</entry></row><row><entry align="center" colname="col1">32x32</entry><entry align="justify" colname="col2">12, 13, 14, 15, 16, 16, 16, 16, 17, 17, 17, 17, 18, 18, 18, 18</entry></row></tbody></tgroup></table></tables>
In Table 0 for the context mapping, the position of each number corresponds to an empty index of the prefix bit string, and the number means a context index to be applied to the corresponding bit position. For convenience of explanation, taking a 4x4 block as an example, the prefix bit string consists of a total of 4 bits, and according to the context mapping table, when k is 0, 1, 2, 3, the k-th bin index has context indexes 0 and 1, respectively. , 2, and 2 are determined, so that arithmetic encoding can be performed based on context modeling.
4 illustrates an embodiment of performing arithmetic coding on an intra prediction mode including a luma intra mode and a chroma intra mode indicating intra prediction directions of a luma block and a chroma block according to an embodiment.
When the intra prediction mode is 6, the symbol bit stream 40 '0000001' is generated by the unary binarization method. In this case, the first bit 41 '0' of the symbol bit string 40 of the intra prediction mode is arithmetic-coded through context modeling, and the remaining bits 45 of the bit string 40 '000001' are in the bypass mode. can be arithmetic-coded. That is, the first bit 41 of the symbol bit string 40 corresponds to the prefix bit string, and the remaining bits 45 correspond to the suffix bit string.
How many bits of the symbol bit string 40 are arithmetic-coded through context modeling as a prefix bit string, and how many bits are arithmetic-coded in the bypass mode as a suffix bit string depends on the size of the block or the set of blocks. It can be determined by size. For example, with respect to a block of size 64x64, only the first bit of the bit stream in the intra prediction mode may be arithmetic-coded through context modeling, and the remaining bits may be arithmetic-coded in the bypass mode. For blocks of other sizes, all bits of the bit string of the intra prediction mode may be arithmetic-coded in the bypass mode.
In general, information on bits close to a Most Significant Bit (MSB) in a symbol bit stream is relatively less important to information on bits close to a Least Significant Bit (LSB). Therefore, the video encoding apparatus 10 and the video decoding apparatus 20 adopt an arithmetic encoding method through a binarization method with higher accuracy even if there is a burden on the amount of computation for a prefix bit string close to the MSB, and apply the suffix bit string close to the LSB. An arithmetic encoding method through a binarization method capable of simple operation can be adopted. In addition, the video encoding apparatus 10 and the video decoding apparatus 20 adopt an arithmetic encoding method based on context modeling for a prefix bit string, and an arithmetic encoding method in which context modeling is omitted for a suffix bit string close to the LSB. can be adopted.
An embodiment in which the prefix/suffix bit string of the final coefficient position information of the transform coefficient is binarized in an individually determined manner and arithmetic-encoded in a different manner has been disclosed with reference to FIG. 3 above. Also, with reference to FIG. 4 , an embodiment of performing arithmetic encoding on a prefix/suffix bit string in an intra prediction mode in a different manner is disclosed.
However, according to various embodiments of the present invention, a symbol encoding method in which a binarization/arithmetic encoding method determined individually or a different binarization/arithmetic encoding method is applied to a prefix bit string and a suffix bit string is shown in FIG. 3 It is not limited to the embodiments disclosed with reference to and 4, and can be extended and applied to embodiments in which various binarization/arithmetic encoding schemes are applied to various symbols.
5 is a flowchart of a video encoding method according to an embodiment of the present invention.
In step 51, symbols are generated by performing prediction and transformation on blocks of an image.
In step 53, the current symbol is classified into a prefix area and a suffix area based on a threshold determined based on the size of the current block.
In step 55, a prefix bit stream and a suffix bit stream are generated by applying each separately determined binarization method to the prefix area and the suffix area of the symbol.
In step 57, symbol encoding is performed by applying each individually determined arithmetic encoding method to the prefix bit string and the suffix bit string.
In step 59, bit streams generated by symbol encoding are output in the form of a bitstream.
In step 57, symbol encoding is performed by applying an arithmetic encoding method for performing context modeling for each bit position to the prefix bit string, and applying an arithmetic encoding method for omitting context modeling by applying a bypass mode to the suffix bit string. can
In step 57, when the symbol is the final coefficient position information of the transform coefficient, arithmetic encoding may be performed using the context of a predetermined index separately allocated in advance for each bit position of the prefix bit string.
6 is a flowchart of a video decoding method according to an embodiment of the present invention.
In step 61, symbols of image blocks are parsed from the received bitstream.
In step 63, the current symbol is classified into a prefix bit stream and a suffix bit stream based on a threshold determined based on the size of the current block.
In step 65, arithmetic decoding is performed on the prefix bit stream and the suffix bit stream of the current symbol according to each individually determined arithmetic decoding method.
In step 67, after arithmetic decoding, inverse binarization is performed on the prefix bit string and the suffix bit string according to each separately determined binarization method.
The prefix and suffix regions of the symbol may be reconstructed by performing inverse binarization on the prefix bit stream and the suffix bit stream according to each separately determined binarization scheme.
In step 69, image blocks are reconstructed by performing inverse transform and prediction on the current block using the current symbol reconstructed through arithmetic decoding and inverse binarization.
In operation 65, arithmetic decoding for determining context modeling for each bit position may be performed on the prefix bit string, and arithmetic decoding may be performed on the suffix bit string by applying a bypass mode to omit context modeling.
In step 65, when the symbol is the final coefficient position information of the transform coefficient, arithmetic decoding may be performed using the context of a predetermined index separately allocated in advance for each bit position of the prefix bit string.
In the video encoding apparatus 10 according to an embodiment and the video decoding apparatus 20 according to another embodiment, blocks into which video data is split are split into coding units having a tree structure, and the coding unit is used for intra prediction. As described above, there are cases in which prediction units are used and a transformation unit is used for transformation. Hereinafter, a video encoding method based on a coding unit, a prediction unit, and a transformation unit having a tree structure, an apparatus thereof, a video decoding method, and an apparatus thereof are disclosed with reference to FIGS. 7 to 19 , according to an embodiment.
7 is a block diagram of a video encoding apparatus 100 based on coding units according to a tree structure according to an embodiment of the present invention.
According to an embodiment, the video encoding apparatus 100 accompanying video prediction based on coding units according to a tree structure includes a maximum coding unit divider 110 , a coding unit determiner 120 , and an output unit 130 . . For convenience of description, the video encoding apparatus 100 accompanying video prediction based on coding units according to a tree structure according to an embodiment is abbreviated as 'video encoding apparatus 100'.
The maximum coding unit divider 110 may partition the current picture based on the largest coding unit, which is the largest coding unit for the current picture of the image. If the current picture is larger than the largest coding unit, image data of the current picture may be split into at least one largest coding unit. The maximum coding unit according to an embodiment may be a data unit having a size of 32x32, 64x64, 128x128, 256x256, etc., and may be a square data unit having horizontal and vertical sizes to the power of two. The image data may be output to the coding unit determiner 120 for each at least one largest coding unit.
A coding unit according to an embodiment may be characterized by a maximum size and a depth. The depth represents the number of times a coding unit is spatially split from the largest coding unit, and as the depth increases, the coding unit for each depth may be split from the largest coding unit to the smallest coding unit. The depth of the largest coding unit may be the highest depth, and the smallest coding unit may be defined as the lowest coding unit. Since the size of the coding unit for each depth decreases as the depth of the largest coding unit increases, a coding unit of a higher depth may include a plurality of coding units of a lower depth.
As described above, the image data of the current picture is divided into the largest coding units according to the maximum size of the coding unit, and each largest coding unit may include coding units divided by depth. Since the maximum coding unit according to an embodiment is divided according to depths, image data in a spatial domain included in the maximum coding unit may be hierarchically classified according to depths.
The maximum depth and the maximum size of the coding unit that limit the total number of hierarchically splitting the height and width of the maximum coding unit may be preset.
The coding unit determiner 120 encodes at least one split region in which a region of the largest coding unit is split for each depth, and determines a depth at which a final encoding result is output for each at least one split region. That is, the coding unit determiner 120 encodes image data in coding units for each depth for each maximum coding unit of the current picture, selects a depth at which the smallest coding error occurs, and determines the depth as the coding depth. The determined coded depth and image data for each maximum coding unit are output to the output unit 130 .
Image data in the maximum coding unit is encoded based on the coding units for each depth according to at least one depth less than or equal to the maximum depth, and encoding results based on the coding units for each depth are compared. As a result of comparing encoding errors of coding units for each depth, a depth having the smallest encoding error may be selected. At least one coded depth may be determined for each maximization coding unit.
As the depth of the maximum coding unit increases, the coding units are hierarchically divided and split, and the number of coding units increases. Also, even in coding units of the same depth included in one maximum coding unit, encoding errors for each data are measured and whether to split into lower depths is determined. Accordingly, even for data included in one maximum coding unit, since encoding errors for respective depths vary according to positions, the coded depths may be determined differently according to positions. Accordingly, one or more coding depths may be set for one largest coding unit, and data of the largest coding unit may be partitioned according to coding units of one or more coding depths.
Accordingly, the coding unit determiner 120 according to an embodiment may determine coding units according to a tree structure included in the current largest coding unit. 'Coding units according to a tree structure' according to an embodiment includes coding units having a depth determined as a coding depth among all coding units for each depth included in the current maximum coding unit. A coding unit of a coded depth may be hierarchically determined according to a depth in the same region within the maximum coding unit, and may be determined independently in other regions. Similarly, the coded depth for the current region may be determined independently of the coded depth for other regions.
The maximum depth according to an embodiment is an index related to the number of divisions from the largest coding unit to the smallest coding unit. The first maximum depth according to an embodiment may indicate the total number of splits from the largest coding unit to the smallest coding unit. The second maximum depth according to an embodiment may indicate the total number of depth levels from the largest coding unit to the smallest coding unit. For example, when the depth of the largest coding unit is 0, the depth of the coding unit in which the largest coding unit is split once may be set to 1, and the depth of the coding unit split into two may be set to 2. In this case, if the coding unit divided 4 times from the largest coding unit is the smallest coding unit, the first maximum depth is set to 4 and the second maximum depth is set to 5 because depth levels of depths 0, 1, 2, 3, and 4 exist. can be
Prediction encoding and transformation of the largest coding unit may be performed. Similarly, prediction encoding and transformation are performed based on the coding unit for each depth for each maximum coding unit and for each depth less than or equal to the maximum depth.
Since the number of coding units for each depth increases whenever the maximum coding unit is split for each depth, encoding including prediction encoding and transformation must be performed on all coding units for each depth generated as the depth increases. Hereinafter, for convenience of description, prediction encoding and transformation will be described based on a coding unit of a current depth among at least one maximum coding unit.
The video encoding apparatus 100 according to an embodiment may select various sizes or shapes of data units for encoding image data. In order to encode image data, prediction encoding, transformation, entropy encoding, etc. are performed, and the same data unit may be used in all steps, or the data unit may be changed in each step.
For example, the video encoding apparatus 100 may select a data unit different from a coding unit to perform predictive encoding of image data of a coding unit as well as a coding unit for encoding image data.
For predictive encoding of the largest coding unit, predictive encoding may be performed based on a coding unit having a coded depth, that is, a coding unit that is not further divided. Hereinafter, an odd non-split coding unit that is a basis for predictive encoding is referred to as a 'prediction unit'. The partition in which the prediction unit is divided may include a data unit in which the prediction unit and at least one of a height and a width of the prediction unit are divided. A partition may be a data unit in which a prediction unit of a coding unit is divided, and the prediction unit may be a partition having the same size as a coding unit.
For example, when a coding unit having a size of 2Nx2N (where N is a positive integer) is no longer split, it becomes a prediction unit having a size of 2Nx2N, and the partition size may be 2Nx2N, 2NxN, Nx2N, NxN, or the like. The partition type according to an embodiment includes not only symmetric partitions in which the height or width of a prediction unit is partitioned in a symmetric ratio, but also partitions partitioned in an asymmetric ratio such as 1:n or n:1, in a geometric form. It may optionally include partitioned partitions, partitions of arbitrary shapes, and the like.
The prediction mode of the prediction unit may be at least one of an intra mode, an inter mode, and a skip mode. For example, the intra mode and the inter mode may be performed for partitions having sizes of 2Nx2N, 2NxN, Nx2N, and NxN. Also, the skip mode may be performed only for a partition having a size of 2Nx2N. A prediction mode having the smallest encoding error may be selected because encoding is independently performed for each prediction unit within the coding unit.
Also, the video encoding apparatus 100 according to an embodiment may perform transformation of image data of a coding unit based on a data unit different from the coding unit as well as a coding unit for encoding image data. For transformation of a coding unit, transformation may be performed based on a transformation unit having a size smaller than or equal to that of the coding unit. For example, the transformation unit may include a data unit for an intra mode and a transformation unit for an inter mode.
In a manner similar to a coding unit having a tree structure according to an embodiment, a transformation unit within a coding unit is also recursively divided into smaller transformation units, and residual data of the coding unit is converted into a tree structure according to a transformation depth. It may be partitioned according to the conversion unit.
For the transformation unit according to an embodiment, a transformation depth indicating the number of divisions until the height and width of the coding unit are divided to reach the transformation unit may be set. For example, if the size of the transformation unit of the current coding unit of size 2Nx2N is 2Nx2N, transformation depth 0 is set, transformation depth 1 if the transformation unit size is NxN, and transformation depth 2 if the size of the transformation unit is N/2xN/2. can That is, even for the transformation unit, a transformation unit according to a tree structure may be set according to the transformation depth.
Encoding information for each coded depth requires prediction-related information and transformation-related information as well as coded depth. Accordingly, the coding unit determiner 120 may determine a partition type obtained by dividing a prediction unit into partitions, a prediction mode for each prediction unit, and a size of a transformation unit for transformation, as well as a coding depth at which a minimum encoding error is generated.
A method of determining a coding unit, a prediction unit/partition, and a transformation unit according to a tree structure of the largest coding unit according to an embodiment will be described below in detail with reference to FIGS. 7 to 19 .
The coding unit determiner 120 may measure a coding error of a coding unit for each depth using a rate-distortion optimization technique based on a Lagrangian multiplier.
The output unit 130 outputs, in the form of a bitstream, image data of the largest coding unit encoded based on the at least one coding depth determined by the coding unit determiner 120 and information about the coding mode for each depth.
The encoded image data may be an encoding result of residual data of an image.
The information about the encoding mode for each depth may include encoding depth information, partition type information of a prediction unit, prediction mode information, size information of a transformation unit, and the like.
The coded depth information may be defined using division information for each depth indicating whether encoding is performed in a coding unit of a lower depth instead of encoding the current depth. If the current depth of the current coding unit is the coding depth, since the current coding unit is encoded as the coding unit of the current depth, split information of the current depth may be defined so that it is no longer split into lower depths. Conversely, if the current depth of the current coding unit is not the coding depth, encoding using the coding unit of the lower depth should be attempted. Therefore, the splitting information of the current depth may be defined to be split into coding units of the lower depth.
If the current depth is not the coded depth, encoding is performed on a coding unit divided into coding units of a lower depth. Since one or more coding units of a lower depth exist in a coding unit of the current depth, encoding may be repeatedly performed for each coding unit of each lower depth, and thus recursive encoding may be performed for each coding unit of the same depth.
Since coding units having a tree structure are determined within one largest coding unit, and information about at least one coding mode must be determined for each coding unit of a coding depth, information about at least one coding mode may be determined for one largest coding unit. can In addition, since the data of the largest coding unit is hierarchically partitioned according to the depth and the coding depth may be different for each location, information on the coding depth and the coding mode may be set for the data.
Accordingly, the output unit 130 according to an embodiment may allocate encoding information for a corresponding coded depth and coding mode to at least one of a coding unit, a prediction unit, and a minimum unit included in the largest coding unit. .
The minimum unit according to an embodiment is a square data unit having a size obtained by dividing the minimum coding unit, which is the lowest coding depth, into four. The minimum unit according to an embodiment may be a square data unit having a maximum size that may be included in all coding units, prediction units, partition units, and transformation units included in the maximum coding unit.
For example, the encoding information output through the output unit 130 may be classified into encoding information for each coding unit for each depth and encoding information for each prediction unit. The encoding information for each coding unit for each depth may include prediction mode information and partition size information. The encoding information transmitted for each prediction unit includes information about an estimation direction of the inter mode, information about a reference image index of the inter mode, information about a motion vector, information about a chroma component of an intra mode, and information about an interpolation method of an intra mode. and the like.
Information on the maximum size and maximum depth of a coding unit defined for each picture, slice, or GOP may be inserted into a header of a bitstream, a sequence parameter set, or a picture parameter set.
In addition, information about the maximum size of a transform unit allowed for the current video and information about the minimum size of a transform unit may also be output through a header of a bitstream, a sequence parameter set, a picture parameter set, or the like. The output unit 130 may encode and output the reference information related to the prediction described above with reference to FIGS. 1 to 6 , the prediction information, the unidirectional prediction information, and the slice type information including the fourth slice type.
According to the simplest embodiment of the video encoding apparatus 100 , the coding unit for each depth is a coding unit having a size obtained by dividing the height and width of a coding unit having a depth higher than one layer in half. That is, if the size of the coding unit of the current depth is 2Nx2N, the size of the coding unit of the lower depth is NxN. Also, a current coding unit having a size of 2Nx2N may include up to four lower-depth coding units having a size of NxN.
Accordingly, the video encoding apparatus 100 determines a coding unit having an optimal shape and size for each largest coding unit based on the size and the maximum depth of the largest coding unit determined in consideration of the characteristics of the current picture, thereby Coding units may be configured. In addition, since each largest coding unit may be encoded using various prediction modes and transformation methods, an optimal encoding mode may be determined in consideration of image characteristics of coding units having various image sizes.
Accordingly, if an image having a very high resolution or a very large data amount is encoded in units of existing macroblocks, the number of macroblocks per picture is excessively increased. Accordingly, since the amount of compressed information generated for each macroblock increases, the transmission burden of compressed information tends to increase and data compression efficiency tends to decrease. Accordingly, the video encoding apparatus according to an embodiment may increase the maximum size of the coding unit in consideration of the size of the image and adjust the coding unit in consideration of the image characteristics, so that image compression efficiency may be increased.
The video encoding apparatus 100 of FIG. 7 may perform the operation of the video encoding apparatus 10 described above with reference to FIG. 1 .
The coding unit determiner 120 may perform an operation of the image encoder 12 of the video encoding apparatus 10 . A prediction unit for intra prediction may be determined for each maximum coding unit and coding units according to a tree structure, intra prediction may be performed for each prediction unit, a transformation unit for transformation may be determined, and transformation may be performed for each transformation unit.
The output unit 130 may perform operations of the symbol encoder 14 and the bitstream output unit 16 of the video encoding apparatus 10 . Symbols for various data units such as pictures, slices, maximum coding units, coding units, prediction units, and transformation units are generated, and each symbol is divided into a prefix region and a suffix region based on a threshold determined based on the size of the corresponding data unit. classified. The output unit 130 may generate a prefix bit stream and a suffix bit stream by applying each separately determined binarization method to the prefix area and the suffix area of the symbol. In order to binarize the prefix region and the suffix region, a prefix bit string and a suffix bit string are generated by adopting any one of the general binarization method, the unary binarization method, the truncated single type binarization method, the exponential Gollum associative binarization method, and the fixed-length binarization method, respectively. can be
The output unit 130 may perform symbol encoding by applying each individually determined arithmetic encoding method to the prefix bit string and the suffix bit string. The output unit 130 applies an arithmetic encoding method for performing context modeling for each bit position with respect to the prefix bit string, and applies an arithmetic encoding method for omitting context modeling by applying a bypass mode to the suffix bit string for symbol encoding. can be performed.
For example, when encoding final coefficient position information of a transform coefficient of a transform unit, a threshold for classifying the prefix bit stream and the suffix bit stream may be determined by the size (width or height) of the transform unit. Alternatively, the threshold value may be determined according to the size of a slice including the current transformation unit, a maximum coding unit, a coding unit, a prediction unit, and the like.
As another example, how many bits of the symbol bit stream of the intra prediction mode are arithmetic-coded as a prefix bit stream through context modeling, and how many bits are arithmetic-coded in the bypass mode as a suffix bit stream is determined by the intra prediction mode. It may be determined by the maximum index. For example, a total of 34 intra prediction modes may be used for size 8x8, 16x16, and 32x32 prediction units, 17 intra prediction modes may be used for size 4x4 prediction units, and for a size 64x64 prediction unit Three intra prediction modes may be used. In this case, since the prediction units in which the same number of intra prediction modes can be used can be regarded as having similar statistical characteristics, the first bit among the bit streams of the intra prediction mode for size 8x8, 16x16, and 32x32 prediction units is context modeling. can be arithmetic-coded. In the remaining cases, that is, 4x4 and 64x64 prediction units of size 4x4, mode bits of the bit stream of the intra prediction mode may be arithmetic-coded in the bypass mode.
The output unit 130 may output the bitstreams generated by symbol encoding in the form of a bitstream.
8 is a block diagram of a video decoding apparatus 200 based on coding units according to a tree structure according to an embodiment of the present invention.
According to an embodiment, the video decoding apparatus 200 accompanying video prediction based on coding units according to a tree structure includes a receiver 210 , an image data and encoding information extractor 220 , and an image data decoder 230 . do. For convenience of description, the video decoding apparatus 200 accompanying video prediction based on coding units according to a tree structure according to an embodiment is abbreviated as 'video decoding apparatus 200'.
Definitions of various terms, such as a coding unit, a depth, a prediction unit, a transformation unit, and information on various encoding modes, for a decoding operation of the video decoding apparatus 200 according to an embodiment are shown in FIG. 7 and the video encoding apparatus 100 It is the same as described above with reference.
The receiver 210 receives and parses the bitstream for the encoded video. The image data and encoding information extractor 220 extracts image data encoded for each coding unit according to the coding units according to the tree structure for each largest coding unit from the parsed bitstream, and outputs the extracted image data to the image data decoder 230 . The image data and encoding information extractor 220 may extract information about the maximum size of a coding unit of the current picture from a header, a sequence parameter set, or a picture parameter set for the current picture.
Also, the image data and encoding information extractor 220 extracts information about a coding depth and an encoding mode of coding units according to a tree structure for each largest coding unit from the parsed bitstream. The extracted information about the coded depth and the coded mode is output to the image data decoder 230 . That is, by dividing the image data of the bit string into the largest coding unit, the image data decoder 230 may decode the image data for each largest coding unit.
The information about the coding depth and the coding mode for each maximum coding unit may be set for one or more coding depth information, and the information about the coding mode for each coding depth includes partition type information, prediction mode information, and transformation unit of the corresponding coding unit. may include size information and the like. In addition, as the coded depth information, division information for each depth may be extracted.
The information about the encoding depth and encoding mode for each maximum coding unit extracted by the image data and encoding information extractor 220 is encoded for each depth for each maximum coding unit at the encoding end like the video encoding apparatus 100 according to an embodiment. Information on a coding depth and an encoding mode determined to generate a minimum encoding error by repeatedly performing encoding for each unit. Accordingly, the video decoding apparatus 200 may reconstruct an image by decoding data according to an encoding method that generates a minimum encoding error.
Since the encoding information on the encoding depth and the encoding mode according to an embodiment may be allocated to a predetermined data unit among a corresponding encoding unit, a prediction unit, and a minimum unit, the image data and encoding information extractor 220 may store the predetermined data Information on the coding depth and the coding mode may be extracted for each unit. If information on a coding depth and an encoding mode of a corresponding maximum coding unit is recorded for each predetermined data unit, predetermined data units having information on the same coding depth and encoding mode are inferred as data units included in the same maximum coding unit. can be
The image data decoder 230 reconstructs a current picture by decoding image data of each largest coding unit based on information about a coding depth and an encoding mode for each largest coding unit. That is, the image data decoder 230 may decode the encoded image data based on the read partition type, prediction mode, and transformation unit for each coding unit among the coding units according to the tree structure included in the largest coding unit. can The decoding process may include a prediction process including intra prediction and motion compensation, and an inverse transform process.
The image data decoder 230 may perform intra prediction or motion compensation according to each partition and prediction mode for each coding unit, based on the partition type information and the prediction mode information of the prediction unit of the coding unit for each coding depth. .
In addition, the image data decoder 230 may read transformation unit information according to a tree structure for each coding unit and perform inverse transformation based on the transformation unit for each coding unit for inverse transformation for each maximum coding unit. Through inverse transformation, pixel values of the spatial domain of the coding unit may be reconstructed.
The image data decoder 230 may determine the coding depth of the current largest coding unit by using the division information for each depth. If the split information indicates that the current depth is no longer split, the current depth is the coded depth. Accordingly, the image data decoder 230 may decode the coding unit of the current depth with respect to the image data of the current largest coding unit using information on the partition type, prediction mode, and transformation unit size of the prediction unit.
That is, by observing the encoding information set for a predetermined data unit among the coding unit, the prediction unit, and the minimum unit, the data units having the encoding information including the same segmentation information are gathered, and the image data decoder 230 generates the data. It may be regarded as one data unit to be decoded in the same encoding mode. Decoding of the current coding unit may be performed by acquiring information on the coding mode for each coding unit determined in this way.
Also, the video decoding apparatus 200 of FIG. 8 may perform the operation of the video decoding apparatus 20 described above with reference to FIG. 2 .
The receiving unit 210 and the image data and encoding information extracting unit 220 may perform operations of the parsing unit 22 and the symbol decoding unit 24 of the video decoding apparatus 20 . The image data decoding unit 230 may perform the operation of the image restoration unit 24 of the video decoding apparatus 20 .
The receiver 210 receives a bitstream of an image, and the image data and encoding information extractor 220 parses symbols of image blocks from the received bitstream.
The image data and encoding information extractor 220 may classify the current symbol into a prefix bit stream and a suffix bit stream based on a threshold determined based on the size of the current block. For example, when decoding final coefficient position information of a transform coefficient of a transform unit, a threshold for classifying a prefix bit stream and a suffix bit stream may be determined by the size (width or height) of the transform unit. Alternatively, the threshold value may be determined according to the size of a slice including the current transformation unit, a maximum coding unit, a coding unit, a prediction unit, and the like. As another example, how many bits of the symbol bit stream of the intra prediction mode are arithmetic-coded as a prefix bit stream through context modeling, and how many bits are arithmetic-coded in the bypass mode as a suffix bit stream is determined by the intra prediction mode. It may be determined by the maximum index.
Arithmetic decoding is performed on the prefix bit stream and the suffix bit string of the current symbol according to each individually determined arithmetic decoding method. Arithmetic decoding for determining context modeling for each bit position may be performed on the prefix bit string, and arithmetic decoding in which context modeling is omitted by applying a bypass mode to the suffix bit string may be performed.
After arithmetic decoding, inverse binarization is performed on the prefix bit string and the suffix bit string according to each separately determined binarization method. The prefix and suffix regions of the symbol may be reconstructed by performing inverse binarization on the prefix bit stream and the suffix bit stream according to each separately determined binarization scheme.
The image data decoder 230 may reconstruct image blocks by performing inverse transform and prediction on the current block using the current symbol restored through arithmetic decoding and inverse binarization.
As a result, the video decoding apparatus 200 may obtain information about a coding unit generating a minimum encoding error by recursively encoding each maximum coding unit in an encoding process, and may use it for decoding a current picture. That is, it is possible to decode the encoded image data of the coding units according to the tree structure determined as the optimal coding unit for each largest coding unit.
Therefore, even if a high-resolution image or an image with an excessively large amount of data is used, the information on the optimal encoding mode transmitted from the encoding stage is used to efficiently perform image data according to the size and encoding mode of the coding unit adaptively determined according to the characteristics of the image. can be restored by decrypting it.
9 illustrates a concept of a coding unit according to an embodiment of the present invention.
As an example of the coding unit, the size of the coding unit is expressed as width x height, and may include a coding unit having a size of 64x64, 32x32, 16x16, and 8x8. A coding unit having a size of 64x64 may be divided into partitions having sizes of 64x64, 64x32, 32x64, and 32x32, a coding unit having a size of 32x32 may be partitioned into partitions having a size of 32x32, 32x16, 16x32, and 16x16, and a coding unit having a size of 16x16 may be divided into partitions having a size of 16x16. , 16x8, 8x16, and 8x8 partitions, and a coding unit of size 8x8 may be divided into partitions of size 8x8, 8x4, 4x8, and 4x4.
For the video data 310, the resolution is set to 1920x1080, the maximum size of the coding unit is set to 64, and the maximum depth is set to 2. For the video data 320 , the resolution is set to 1920x1080, the maximum size of the coding unit is set to 64, and the maximum depth is set to 3. For the video data 330 , the resolution is set to 352x288, the maximum size of the coding unit is set to 16, and the maximum depth is set to 1. The maximum depth illustrated in FIG. 9 represents the total number of divisions from the largest coding unit to the smallest coding unit.
When the resolution is high or the amount of data is large, it is preferable that the maximum encoding size be relatively large in order to not only improve encoding efficiency but also accurately reflect image characteristics. Accordingly, in the video data 310 and 320 having a higher resolution than the video data 330 , a maximum encoding size of 64 may be selected.
Since the maximum depth of the video data 310 is 2, the coding unit 315 of the video data 310 is split twice from the largest coding unit having a major axis size of 64, and the depth is two layers deep so that the major axis sizes are 32 and 16. It may include up to phosphorus coding units. On the other hand, since the maximum depth of the video data 330 is 1, the coding unit 335 of the video data 330 is divided once from coding units having a long axis size of 16, and the depth is increased by one layer, so that the long axis size is 8. It may include up to phosphorus coding units.
Since the maximum depth of the video data 320 is 3, the coding unit 325 of the video data 320 is divided three times from the largest coding unit having a major axis size of 64, and the depth is three layers deep, so that the major axis sizes are 32 and 16. , may include up to 8 coding units. As the depth increases, the ability to express detailed information may be improved.
10 is a block diagram of an image encoder 400 based on coding units according to an embodiment of the present invention.
The image encoder 400 according to an embodiment includes operations that the coding unit determiner 120 of the video encoding apparatus 100 undergoes to encode image data. That is, the intra prediction unit 410 performs intra prediction on the intra mode coding unit among the current frame 405 , and the motion estimator 420 and the motion compensator 425 perform the inter mode current frame 405 . and the reference frame 495 to perform inter estimation and motion compensation.
Data output from the intra prediction unit 410 , the motion estimation unit 420 , and the motion compensation unit 425 are output as quantized transform coefficients through the transform unit 430 and the quantization unit 440 . The quantized transform coefficients are reconstructed into spatial domain data through the inverse quantization unit 460 and the inverse transform unit 470 , and the reconstructed spatial domain data passes through the deblocking unit 480 and the loop filtering unit 490 . processed and output as a reference frame 495 . The quantized transform coefficient may be output as a bitstream 455 through the entropy encoder 450 .
In order to be applied to the video encoding apparatus 100 according to an embodiment, the intra predictor 410 , the motion estimator 420 , the motion compensator 425 , and the transform unit ( 430), the quantization unit 440, the entropy encoding unit 450, the inverse quantization unit 460, the inverse transform unit 470, the deblocking unit 480, and the loop filtering unit 490 are all max. An operation based on each coding unit among coding units according to the tree structure should be performed in consideration of the depth.
In particular, the intra predictor 410 , the motion estimator 420 , and the motion compensator 425 partition each coding unit among coding units according to a tree structure in consideration of the maximum size and the maximum depth of the current maximum coding unit. and a prediction mode, and the transform unit 430 must determine the size of a transformation unit in each coding unit among coding units according to the tree structure.
In particular, the entropy encoder 450 classifies the symbol into a prefix region and a suffix region according to a predetermined threshold, and applies different binarization and arithmetic encoding methods to the prefix region and the suffix region, respectively, to form a prefix region and Symbol encoding may be performed individually on the suffix region.
The threshold for classifying the prefix region and the suffix region of the symbol may be determined based on the size of a data unit of the symbol, that is, a slice, a maximum coding unit, a coding unit, a prediction unit, a transformation unit, and the like.
11 is a block diagram of an image decoder 500 based on a coding unit according to an embodiment of the present invention.
The bitstream 505 passes through the parsing unit 510 to parse encoded image data to be decoded and information on encoding required for decoding. The encoded image data is output as inverse quantized data through the entropy decoder 520 and the inverse quantizer 530 , and the image data in the spatial domain is restored through the inverse transform unit 540 .
With respect to the image data in the spatial domain, the intra prediction unit 550 performs intra prediction on the coding unit of the intra mode, and the motion compensation unit 560 uses the reference frame 585 together with the coding unit of the inter mode. motion compensation is performed.
Data in the spatial domain that has passed through the intra prediction unit 550 and the motion compensator 560 may be post-processed through the deblocking unit 570 and the loop filtering unit 580 and output as a restored frame 595 . In addition, data post-processed through the deblocking unit 570 and the loop filtering unit 580 may be output as the reference frame 585 .
In order to decode the image data in the image data decoder 230 of the video decoding apparatus 200, step-by-step operations after the parser 510 of the image decoder 500 according to an embodiment may be performed.
In order to be applied to the video decoding apparatus 200 according to an embodiment, the parsing unit 510 , the entropy decoding unit 520 , the inverse quantization unit 530 , and the inverse transform unit 540 which are components of the image decoding unit 500 . ), the intra prediction unit 550 , the motion compensation unit 560 , the deblocking unit 570 , and the loop filtering unit 580 must perform an operation based on the coding units according to the tree structure for each largest coding unit. do.
In particular, the intra predictor 550 and the motion compensator 560 determine a partition and a prediction mode for each coding unit according to a tree structure, and the inverse transform unit 540 determines a size of a transform unit for each coding unit. .
In particular, the entropy decoding unit 520 classifies the parsed symbol bit string into a prefix bit string and a suffix bit string according to a predetermined threshold, and performs different binarization and arithmetic decoding methods for the prefix bit string and the suffix bit string, respectively. By application, symbol decoding can be individually performed on the prefix bit stream and the suffix bit string.
The threshold for classifying the prefix region and the suffix region of the symbol may be determined based on the size of a data unit of the symbol, that is, a slice, a maximum coding unit, a coding unit, a prediction unit, a transformation unit, and the like.
12 illustrates coding units and partitions for each depth according to an embodiment of the present invention.
The video encoding apparatus 100 according to an embodiment and the video decoding apparatus 200 according to an embodiment use hierarchical coding units in consideration of image characteristics. The maximum height, width, and maximum depth of the coding unit may be adaptively determined according to characteristics of an image, or may be variously set according to a user's request. The size of the coding unit for each depth may be determined according to the preset maximum size of the coding unit.
The hierarchical structure 600 of coding units according to an embodiment illustrates a case where the maximum height and width of the coding unit is 64, and the maximum depth is 4. In this case, the maximum depth represents the total number of divisions from the largest coding unit to the smallest coding unit. Since the depth increases along the vertical axis of the hierarchical structure 600 of the coding units according to an embodiment, the height and the width of the coding units for each depth are respectively divided. Also, along the horizontal axis of the hierarchical structure 600 of coding units, prediction units and partitions that are the basis for predictive encoding of coding units for each depth are illustrated.
That is, the coding unit 610 is the largest coding unit in the hierarchical structure 600 of coding units, and has a depth of 0 and a size of a coding unit, that is, a height and a width of 64x64. The depth increases along the vertical axis, the coding unit 620 having a depth of 1 with a size of 32x32, a coding unit 630 with a depth of 2 having a size of 16x16, a coding unit 640 having a depth of 3 with a size of 8x8, and a coding unit 640 having a size of 4x4 with a depth of 4 A coding unit 650 exists. A coding unit 650 having a size of 4x4 and a depth of 4 is a minimum coding unit.
A prediction unit and partitions of a coding unit are arranged along a horizontal axis for each depth. That is, if the coding unit 610 of size 64x64 with depth 0 is a prediction unit, the prediction unit includes a partition 610 of size 64x64, partitions 612 of size 64x32, and sizes included in the coding unit 610 of size 64x64. It may be divided into partitions 614 of 32x64 and partitions 616 of size 32x32.
Similarly, a prediction unit of the coding unit 620 having a size of 32x32 of depth 1 includes a partition 620 having a size of 32x32, partitions 622 having a size of 32x16, and a partition having a size of 16x32 included in the coding unit 620 having a size of 32x32. 624 , partitions 626 of size 16x16 may be partitioned.
Similarly, the prediction unit of the coding unit 630 of the size 16x16 of the depth 2 includes the partition 630 of the size 16x16, the partitions 632 of the size 16x8, and the partition of the size 8x16 included in the coding unit 630 of the size 16x16. 634, may be divided into partitions 636 of size 8x8.
Similarly, the prediction unit of the coding unit 640 of the size 8x8 of the depth 3 includes the partition 640 of the size 8x8, the partitions 642 of the size 8x4, and the partition of the size 4x8 included in the coding unit 640 of the size 8x8. 644 , may be divided into partitions 646 of size 4x4.
Finally, the coding unit 650 having a size of 4x4 with a depth of 4 is a minimum coding unit and a coding unit with the lowest depth, and a corresponding prediction unit may be set only as a partition 650 having a size of 4x4.
The coding unit determiner 120 of the video encoding apparatus 100 according to an embodiment determines the coding depth of the maximum coding unit 610 , the coding unit of each depth included in the largest coding unit 610 . Encoding must be performed every time.
The number of coding units per depth for including data having the same range and size increases as the depth increases. For example, for data including one coding unit of depth 1, four coding units of depth 2 are required. Accordingly, in order to compare the encoding results of the same data for each depth, encoding must be performed using one coding unit of depth 1 and four coding units of depth 2, respectively.
For encoding for each depth, encoding is performed for each prediction unit of the coding unit for each depth along the horizontal axis of the hierarchical structure 600 of the coding unit, and a representative encoding error that is the smallest encoding error at the corresponding depth may be selected. . In addition, the depth increases along the vertical axis of the hierarchical structure 600 of the coding unit, and encoding is performed for each depth, and the minimum encoding error can be found by comparing the representative encoding error for each depth. A depth and a partition in which a minimum coding error occurs among the largest coding units 610 may be selected as the coding depth and partition type of the largest coding unit 610 .
13 illustrates a relationship between a coding unit and a transformation unit according to an embodiment of the present invention.
The video encoding apparatus 100 according to an embodiment or the video decoding apparatus 200 according to an embodiment encodes or decodes an image in a coding unit having a size smaller than or equal to the maximum coding unit for each largest coding unit. A size of a transformation unit for transformation during an encoding process may be selected based on a data unit that is not larger than each coding unit.
For example, in the video encoding apparatus 100 according to an embodiment or the video decoding apparatus 200 according to an embodiment, when the current coding unit 710 has a size of 64x64, the transformation unit 720 of a size of 32x32 is Transformation can be performed using
In addition, the data of the 64x64 coding unit 710 is transformed into 32x32, 16x16, 8x8, and 4x4 transform units of 64x64 size or less, respectively, and is encoded, and then the transformation unit with the smallest error from the original is selected. can be
14 illustrates encoding information for each depth according to an embodiment of the present invention.
The output unit 130 of the video encoding apparatus 100 according to an embodiment is information about an encoding mode, and includes information 800 about a partition type and information 810 about a prediction mode for each coding unit of each coded depth. , the information 820 on the transform unit size may be encoded and transmitted.
The information 800 on the partition type is a data unit for prediction encoding of the current coding unit, and indicates information on the shape of a partition in which the prediction unit of the current coding unit is divided. For example, a current coding unit CU_0 having a size of 2Nx2N may be one of a size 2Nx2N partition 802, a size 2NxN partition 804, a size Nx2N partition 806, and a size NxN partition 808. It can be divided and used. In this case, the information 800 about the partition type of the current coding unit indicates one of a partition 802 having a size of 2Nx2N, a partition 804 having a size of 2NxN, a partition 806 having a size of Nx2N, and a partition 808 having a size of NxN. is set to
The information 810 on the prediction mode indicates the prediction mode of each partition. For example, through the information 810 on the prediction mode, the partition indicated by the information 800 on the partition type indicates whether prediction encoding is performed in one of the intra mode 812 , the inter mode 814 , and the skip mode 816 . may be set.
In addition, the information 820 on the size of the transformation unit indicates whether the current coding unit is to be transformed based on which transformation unit. For example, the transformation unit may be one of a first intra transformation unit size 822 , a second intra transformation unit size 824 , a first inter transformation unit size 826 , and a second intra transformation unit size 828 . have.
The image data and encoding information extractor 210 of the video decoding apparatus 200 according to an embodiment may include information 800 on a partition type, information on a prediction mode 810, and a transform for each coding unit for each depth. The information 820 on the unit size may be extracted and used for decoding.
15 illustrates a coding unit for each depth according to an embodiment of the present invention.
Segmentation information may be used to indicate a change in depth. The split information indicates whether a coding unit of the current depth is split into a coding unit of a lower depth.
The prediction unit 910 for prediction encoding of the coding unit 900 having a depth of 0 and 2N_0x2N_0 is a partition type 912 having a size of 2N_0x2N_0, a partition type 914 having a size of 2N_0xN_0, a partition type 916 having a size of N_0x2N_0, and N_0xN_0. It may include a partition type 918 of size. Only the partitions 912, 914, 916, and 918 in which the prediction unit is divided at a symmetric ratio are exemplified, but as described above, the partition type is not limited thereto, and an asymmetric partition, an arbitrary-shaped partition, a geometric-shaped partition, etc. may include.
For each partition type, predictive encoding must be repeatedly performed for one partition with a size of 2N_0x2N_0, two partitions with a size of 2N_0xN_0, two partitions with a size of N_0x2N_0, and four partitions with a size of N_0xN_0. Predictive encoding may be performed on a partition having a size of 2N_0x2N_0, N_0x2N_0, and 2N_0xN_0 and N_0xN_0 in an intra mode and an inter mode. In the skip mode, predictive encoding may be performed only on a partition having a size of 2N_0x2N_0.
If the encoding error due to one of the partition types 912, 914, and 916 of sizes 2N_0x2N_0, 2N_0xN_0, and N_0x2N_0 is the smallest, it is no longer necessary to divide into a lower depth.
If the encoding error due to the partition type 918 having the size N_0xN_0 is the smallest, the depth 0 is changed to 1 and divided (920), and the coding units 930 of the partition type 930 having a depth of 2 and a size N_0xN_0 are repeatedly encoded. can be performed to search for the minimum encoding error.
A prediction unit 940 for prediction encoding of a coding unit 930 having a depth of 1 and a size of 2N_1x2N_1 (=N_0xN_0) includes a partition type 942 having a size of 2N_1x2N_1, a partition type 944 having a size of 2N_1xN_1, and a partition type having a size of N_1x2N_1. 946, a partition type 948 of size N_1xN_1.
Also, if the encoding error due to the partition type 948 having the size N_1xN_1 is the smallest, the depth 1 is changed to the depth 2 and divided (950), and the coding units 960 of the depth 2 and the size N_2xN_2 are iteratively performed. A minimum encoding error can be searched for by performing encoding.
When the maximum depth is d, a coding unit for each depth may be set up to a depth d-1, and split information may be set up to a depth d-2. That is, when encoding is performed from depth d-2 to depth d-1 by splitting 970 from depth d-2, prediction encoding of the coding unit 980 having a depth of d-1 and a size of 2N_(d-1)x2N_(d-1) is performed. A prediction unit 990 for , a partition type 992 having a size of 2N_(d-1)x2N_(d-1), a partition type 994 having a size of 2N_(d-1)xN_(d-1), and a size A partition type 996 having a size of N_(d-1)x2N_(d-1) and a partition type 998 having a size of N_(d-1)xN_(d-1) may be included.
Among the partition types, one partition of size 2N_(d-1)x2N_(d-1), two partitions of size 2N_(d-1)xN_(d-1), and two partitions of size N_(d-1)x2N_ Encoding through prediction encoding is iteratively performed for each partition of (d-1) and four partitions of size N_(d-1)xN_(d-1), so that a partition type in which a minimum encoding error occurs can be searched for. .
Even if the encoding error due to the partition type 998 of size N_(d-1)xN_(d-1) is the smallest, since the maximum depth is d, the coding unit CU_(d-1) of depth d-1 is no longer Without the process of dividing into lower depths, the coding depth for the current largest coding unit 900 may be determined to be a depth d-1, and a partition type may be determined to be N_(d-1)xN_(d-1). Also, since the maximum depth is d, split information is not set for the coding unit 952 having a depth of d-1.
The data unit 999 may be referred to as a 'minimum unit' of the current largest coding unit. The minimum unit according to an embodiment may be a square data unit having a size obtained by dividing the minimum coding unit, which is the lowest coding depth, into four. Through this iterative encoding process, the video encoding apparatus 100 according to an embodiment compares the encoding errors for each depth of the coding unit 900, selects a depth at which the smallest encoding error occurs, determines the encoding depth, The corresponding partition type and prediction mode may be set as the encoding mode of the coded depth.
In this way, by comparing the minimum encoding errors for all depths of depths 0, 1, ..., d-1, and d, the depth having the smallest error may be selected and determined as the encoding depth. The coding depth, the partition type of the prediction unit, and the prediction mode may be encoded and transmitted as information about the coding mode. In addition, since the coding unit needs to be split from depth 0 to the coded depth, only split information of the coded depth should be set to '0', and split information for each depth excluding the coded depth should be set to '1'.
The image data and encoding information extractor 220 of the video decoding apparatus 200 according to an embodiment extracts information about a coding depth and a prediction unit for the coding unit 900 to be used to decode the coding unit 912 . can The video decoding apparatus 200 according to an embodiment may determine a depth whose division information is '0' as the coded depth using the division information for each depth, and use the information about the encoding mode for the corresponding depth for decoding. have.
16, 17, and 18 illustrate a relationship between a coding unit, a prediction unit, and a transformation unit according to an embodiment of the present invention.
The coding units 1010 are coding units for each coding depth determined by the video encoding apparatus 100 according to an embodiment with respect to the largest coding unit. The prediction unit 1060 is partitions of prediction units of a coding unit for each coding depth among the coding units 1010 , and the transformation unit 1070 is transformation units of a coding unit for each coding depth.
Assuming that the depth of the maximum coding unit is 0 for the coding units 1010 for each depth, the coding units 1012 and 1054 have a depth of 1, and the coding units 1014 , 1016 , 1018 , 1028 , 1050 , and 1052 have a depth of 1 . 2, coding units 1020, 1022, 1024, 1026, 1030, 1032, and 1048 have a depth of 3, and coding units 1040, 1042, 1044, and 1046 have a depth of 4.
Some partitions 1014 , 1016 , 1022 , 1032 , 1048 , 1050 , 1052 , and 1054 of the prediction units 1060 are in the form of a coding unit divided. That is, the partitions 1014, 1022, 1050, and 1054 have a 2NxN partition type, the partitions 1016, 1048, and 1052 have an Nx2N partition type, and the partition 1032 has an NxN partition type. A prediction unit and partitions of the coding units 1010 for each depth are smaller than or equal to each coding unit.
Transformation or inverse transformation is performed on the image data of some 1052 of the transformation units 1070 in a data unit having a smaller size than that of the coding unit. Also, the transformation units 1014 , 1016 , 1022 , 1032 , 1048 , 1050 , 1052 , and 1054 are data units having different sizes or shapes compared with corresponding prediction units and partitions among the prediction units 1060 . That is, the video encoding apparatus 100 according to an embodiment and the video decoding apparatus 200 according to an embodiment perform an intra prediction/motion estimation/motion compensation operation and a transformation/inverse transformation operation for the same coding unit. Each may be performed based on a separate data unit.
Accordingly, encoding is performed recursively for each largest coding unit and each coding unit having a hierarchical structure for each region to determine an optimal coding unit, so that coding units having a recursive tree structure may be configured. The encoding information may include partitioning information on a coding unit, partition type information, prediction mode information, and transformation unit size information. Table 1 below shows examples that can be set in the video encoding apparatus 100 according to an embodiment and the video decoding apparatus 200 according to an embodiment.
<tables num="1"><table><tgroup cols="6"><colspec align="justify" colname="col1" colnum="1" colwidth="1268" /><colspec align="justify" colname="col2" colnum="2" colwidth="1208" /><colspec align="justify" colname="col3" colnum="3" colwidth="1208" /><colspec align="justify" colname="col4" colnum="4" colwidth="1490" /><colspec align="justify" colname="col5" colnum="5" colwidth="1840" /><colspec align="left" colname="col6" colnum="6" colwidth="0" /><tbody><row><entry align="justify" nameend="col5" namest="col1">Segmentation information 0 (coding for a coding unit having a size of 2Nx2N at a current depth d)</entry><entry align="justify" colname="col6">Split information 1 </entry></row><row><entry align="justify" colname="col1">Prediction mode</entry><entry align="justify" nameend="col3" namest="col2">Partition type</entry><entry align="justify" nameend="col5" namest="col4">Conversion unit size</entry><entry align="justify" colname="col6" morerows="2">Iterative encoding for each coding unit of a lower depth d+1</entry></row><row><entry align="justify" colname="col1" morerows="1">Intra Inter Skip (2Nx2N only)</entry><entry align="justify" colname="col2">symmetrical partition type</entry><entry align="justify" colname="col3">Asymmetric Partition Type</entry><entry align="justify" colname="col4">Transform Unit Split Information 0</entry><entry align="justify" colname="col5">About splitting conversion units 1</entry></row><row><entry align="justify" colname="col2"><b>2</b>Nx2N 2NxN Nx2N NxN</entry><entry align="justify" colname="col3">2NxnU 2NxnD nLx2N nRx2N</entry><entry align="justify" colname="col4">2Nx2N</entry><entry align="justify" colname="col5">NxN (symmetric partition type) N/2xN/2 (asymmetric partition type)</entry></row></tbody></tgroup></table></tables>
The output unit 130 of the video encoding apparatus 100 according to an embodiment outputs encoding information on coding units according to a tree structure, and the encoding information extractor ( 220) may extract encoding information on coding units according to a tree structure from the received bitstream.
The split information indicates whether the current coding unit is split into coding units of a lower depth. If the splitting information of the current depth d is 0, the current coding unit is the coded depth at which the current coding unit is no longer split into lower coding units, so partition type information, prediction mode, and transformation unit size information are defined for the coded depth. can be In the case where the partition is to be further partitioned according to partition information, encoding must be performed independently for each of the four partitioned coding units of the lower depths.
The prediction mode may be expressed as one of an intra mode, an inter mode, and a skip mode. Intra mode and inter mode may be defined in all partition types, and skip mode may be defined only in partition type 2Nx2N.
Partition type information indicates symmetric partition types 2Nx2N, 2NxN, Nx2N, and NxN in which a height or width of a prediction unit is divided at a symmetric ratio, and asymmetric partition types 2NxnU, 2NxnD, nLx2N, and nRx2N in which a height or width of a prediction unit is divided at a symmetric ratio. can The asymmetric partition types 2NxnU and 2NxnD have a height divided by 1:3 and 3:1, respectively, and the asymmetric partition types nLx2N and nRx2N have a width divided by 1:3 and 3:1, respectively.
The transformation unit size may be set to two types of sizes in the intra mode and two types of sizes in the inter mode. That is, if the transformation unit split information is 0, the size of the transformation unit is set to the size of the current coding unit 2Nx2N. If the transformation unit split information is 1, a transformation unit having a size in which the current coding unit is split may be set. In addition, if the partition type of the current coding unit having a size of 2Nx2N is a symmetric partition type, the size of the transformation unit may be set to NxN, and if the partition type is an asymmetric partition type, the size of the transformation unit may be set to N/2xN/2.
The encoding information of the coding units according to the tree structure according to an embodiment may be allocated to at least one of a coding unit, a prediction unit, and a minimum unit of a coded depth. The coding unit of the coded depth may include one or more prediction units and minimum units having the same coding information.
Accordingly, when encoding information possessed by adjacent data units is checked, whether they are included in a coding unit having the same coded depth may be checked. In addition, since the coding unit of the corresponding coding depth can be identified by using the coding information possessed by the data unit, the distribution of the coding depths in the largest coding unit can be inferred.
Accordingly, in this case, when the current coding unit is predicted with reference to a neighboring data unit, encoding information of a data unit in a coding unit for each depth adjacent to the current coding unit may be directly referenced and used.
As another embodiment, when prediction encoding is performed on the current coding unit with reference to a neighboring coding unit, data adjacent to the current coding unit within the coding unit for each depth is stored using encoding information of the coding unit for each depth. By being searched, neighboring coding units may be referred to.
19 illustrates a relationship between a coding unit, a prediction unit, and a transformation unit according to the coding mode information of Table 1. Referring to FIG.
The maximum coding unit 1300 includes coding units 1302 , 1304 , 1306 , 1312 , 1314 , 1316 , and 1318 of a coded depth. Since one coding unit 1318 is a coding unit of the coded depth, split information may be set to 0. Partition type information of the coding unit 1318 having a size of 2Nx2N includes partition types 2Nx2N 1322 , 2NxN 1324 , Nx2N 1326 , NxN 1328 , 2NxnU 1332 , 2NxnD 1334 , and nLx2N 1336 ). and nRx2N (1338).
The transform unit split information (TU size flag) is a type of transform index, and the size of a transform unit corresponding to the transform index may be changed according to the prediction unit type or partition type of the coding unit.
For example, when the partition type information is set to one of the symmetric partition types 2Nx2N (1322), 2NxN (1324), Nx2N (1326), and NxN (1328), if the transformation unit division information is 0, the transformation unit of size 2Nx2N ( 1342) is set, and when the transformation unit division information is 1, the transformation unit 1344 of size NxN may be set.
When the partition type information is set to one of the asymmetric partition types 2NxnU (1332), 2NxnD (1334), nLx2N (1336), and nRx2N (1338), if the transform unit partition information (TU size flag) is 0, the transform unit of size 2Nx2N ( 1352 is set, and when the transformation unit division information is 1, the transformation unit 1354 of size N/2xN/2 may be set.
The transformation unit division information (TU size flag) described above with reference to FIG. 21 is a flag having a value of 0 or 1, but the transformation unit division information according to an embodiment is not limited to a 1-bit flag, and 0 according to the setting. , 1, 2, 3., etc., the transformation unit may be hierarchically divided. The transformation unit division information may be used as an embodiment of the transformation index.
In this case, when the transformation unit division information according to an embodiment is used together with the maximum size of the transformation unit and the minimum size of the transformation unit, the size of the transformation unit actually used may be expressed. The video encoding apparatus 100 according to an embodiment may encode the maximum transformation unit size information, the minimum transformation unit size information, and the maximum transformation unit split information. The encoded maximum transformation unit size information, the minimum transformation unit size information, and the maximum transformation unit split information may be inserted into the SPS. The video decoding apparatus 200 according to an embodiment may use the maximum transformation unit size information, the minimum transformation unit size information, and the maximum transformation unit division information for video decoding.
For example, if (a) the size of the current coding unit is 64x64 and the maximum size of the transformation unit is 32x32, (a-1) when the transformation unit split information is 0, the size of the transformation unit is 32x32, (a-2) the transformation unit When the division information is 1, the size of the transformation unit may be set to 16x16, and (a-3) when the transformation unit division information is 2, the size of the transformation unit may be set to 8x8.
As another example, (b) if the size of the current coding unit is 32x32 and the minimum transformation unit size is 32x32, (b-1) when the transformation unit split information is 0, the size of the transformation unit may be set to 32x32, Since the size cannot be smaller than 32x32, no further transformation unit division information can be set.
As another example, (c) if the size of the current coding unit is 64x64 and the maximum transformation unit splitting information is 1, the transform unit splitting information may be 0 or 1, and other transform unit splitting information cannot be set.
Therefore, when the maximum transform unit split information is defined as 'MaxTransformSizeIndex', the minimum transform unit size is 'MinTransformSize', and the transform unit size when the transform unit split information is 0 is 'RootTuSize', the minimum transform unit possible in the current coding unit is defined as 'RootTuSize'. The size 'CurrMinTuSize' may be defined as in the following relation (1).
CurrMinTuSize
= max(MinTransformSize, RootTuSize/(2^MaxTransformSizeIndex)) ... (1)
Compared with 'CurrMinTuSize', which is the smallest possible transformation unit size in the current coding unit, 'RootTuSize', which is the transformation unit size when transformation unit split information is 0, may indicate the maximum size of a transformation unit that can be adopted in the system. That is, according to the relational expression (1), 'RootTuSize/(2^MaxTransformSizeIndex)' is a transformation obtained by dividing 'RootTuSize', which is the transformation unit size when transformation unit division information is 0, by the number of times corresponding to the maximum transformation unit division information. Since it is a unit size and 'MinTransformSize' is the minimum transformation unit size, a smaller value among them may be 'CurrMinTuSize', the smallest possible transformation unit size in the current coding unit.
The maximum transform unit size RootTuSize according to an embodiment may vary according to a prediction mode.
For example, if the current prediction mode is the inter mode, the RootTuSize may be determined according to the following relation (2). In Relation (2), 'MaxTransformSize' represents the maximum transformation unit size, and 'PUSize' represents the current prediction unit size.
RootTuSize = min(MaxTransformSize, PUSize) ..... (2)
That is, if the current prediction mode is the inter mode, 'RootTuSize', which is the size of the transformation unit when the transformation unit split information is 0, may be set to a smaller value among the maximum transformation unit size and the current prediction unit size.
If the prediction mode of the current partition unit is the intra mode and the prediction mode is the intra mode, 'RootTuSize' may be determined according to the following relation (3). 'PartitionSize' indicates the size of the current partition unit.
RootTuSize = min(MaxTransformSize, PartitionSize) .....(3)
That is, if the current prediction mode is the intra mode, the transformation unit size 'RootTuSize' when the transformation unit partition information is 0 may be set to the smaller of the maximum transformation unit size and the current partition unit size.
However, it should be noted that the current maximum transformation unit size 'RootTuSize' according to an embodiment that varies according to the prediction mode of the partition unit is only an embodiment, and a factor determining the current maximum transformation unit size is not limited thereto.
According to the video encoding technique based on coding units having a tree structure described above with reference to FIGS. 7 to 19 , image data in a spatial domain is encoded for each coding unit having a tree structure, and a video decoding technique based on coding units having a tree structure Accordingly, image data in the spatial domain is reconstructed while decoding is performed for each largest coding unit, so that a video that is a picture and a picture sequence may be reconstructed. The restored video may be played back by a playback device, stored in a storage medium, or transmitted over a network.
Meanwhile, the above-described embodiments of the present invention can be written as a program that can be executed on a computer, and can be implemented in a general-purpose digital computer that operates the program using a computer-readable recording medium. The computer-readable recording medium includes a storage medium such as a magnetic storage medium (eg, a ROM, a floppy disk, a hard disk, etc.) and an optically readable medium (eg, a CD-ROM, a DVD, etc.).
So far, the present invention has been looked at with respect to preferred embodiments thereof. Those of ordinary skill in the art to which the present invention pertains will understand that the present invention can be implemented in a modified form without departing from the essential characteristics of the present invention. Therefore, the disclosed embodiments are to be considered in an illustrative rather than a restrictive sense. The scope of the present invention is indicated in the claims rather than the foregoing description, and all differences within the scope equivalent thereto should be construed as being included in the present invention.
20 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR100718134B1 | Cites | Republic of Korea | Search report |
| JP2011120047A | Cites | Japan | – |
| JP2008113374A | Cites | Japan | – |
| KR1020090129939A | Cites | Republic of Korea | – |
155 members in 27 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 61502038 | United States of America | – | |
| 201161502038 | United States of America | P |
Members155
| Document | Office | Kind | |
|---|---|---|---|
| CA2840481A1 | Canada | A1 | |
| CA2975695A1 | Canada | A1 | |
| WO2013002555A2 | World Intellectual Property Organization (WIPO) | A2 | |
| KR20130002285A | Republic of Korea | A | |
| TW201309031A | Taiwan Province of China | A | |
| WO2013002555A3 | World Intellectual Property Organization (WIPO) | A3 | |
| AU2012276453A1 | Australia | A1 | |
| PH12014500011A1 | Philippines | A1 | |
| MX2014000172A | Mexico | A | |
| CN103782597A | China | A | |
| EP2728866A2 | European Patent Office (EPO) | A2 | |
| KR20140075658A | Republic of Korea | A | |
| JP2014525166A | Japan | A | |
| KR101457399B1 | Republic of Korea | B1 | |
| KR20140146561A | Republic of Korea | A | |
| EP2728866A4 | European Patent Office (EPO) | A4 | |
| EP2849445A1 | European Patent Office (EPO) | A1 | |
| KR20150046772A | Republic of Korea | A | |
| KR20150046773A | Republic of Korea | A | |
| KR20150046774A | Republic of Korea | A | |
| US2015139299A1 | United States of America | A1 | |
| US2015139332A1 | United States of America | A1 | |
| EP2884749A1 | European Patent Office (EPO) | A1 | |
| JP5735710B2 | Japan | B2 | |
| US2015181224A1 | United States of America | A1 | |
| US2015181225A1 | United States of America | A1 | |
| RU2014102581A | Russian Federation | A | |
| JP2015149770A | Japan | A | |
| JP2015149771A | Japan | A | |
| JP2015149772A | Japan | A | |
| JP2015149773A | Japan | A | |
| JP2015149774A | Japan | A | |
| ZA201400647B | South Africa | B | |
| KR101560549B1 | Republic of Korea | B1 | |
| KR101560550B1 | Republic of Korea | B1 | |
| KR101560551B1 | Republic of Korea | B1 | |
| KR101560552B1This record | Republic of Korea | B1 | |
| US9247270B2 | United States of America | B2 | |
| MX336876B | Mexico | B | |
| US9258571B2 | United States of America | B2 | |
| MX337230B | Mexico | B | |
| MX337232B | Mexico | B | |
| CN105357540A | China | A | |
| CN105357541A | China | A | |
| JP5873200B2 | Japan | B2 | |
| JP5873201B2 | Japan | B2 | |
| JP5873202B2 | Japan | B2 | |
| JP5873203B2 | Japan | B2 | |
| CN105516732A | China | A | |
| EP3013054A1 | European Patent Office (EPO) | A1 | |
| CN105554510A | China | A | |
| AU2012276453B2 | Australia | B2 | |
| EP3021591A1 | European Patent Office (EPO) | A1 | |
| US2016156939A1 | United States of America | A1 | |
| RU2586321C2 | Russian Federation | C2 | |
| JP5934413B2 | Japan | B2 | |
| AU2016206258A1 | Australia | A1 | |
| AU2016206259A1 | Australia | A1 | |
| AU2016206260A1 | Australia | A1 | |
| AU2016206261A1 | Australia | A1 | |
| TWI562618B | Taiwan Province of China | B | |
| TW201701675A | Taiwan Province of China | A | |
| US9554157B2 | United States of America | B2 | |
| ZA201502759B | South Africa | B | |
| ZA201502760B | South Africa | B | |
| ZA201502761B | South Africa | B | |
| US9565455B2 | United States of America | B2 | |
| MY160178A | Malaysia | A | |
| MY160179A | Malaysia | A | |
| MY160180A | Malaysia | A | |
| MY160181A | Malaysia | A | |
| MY160326A | Malaysia | A | |
| RU2618511C1 | Russian Federation | C1 | |
| AU2016206258B2 | Australia | B2 | |
| US9668001B2 | United States of America | B2 | |
| BR112013033708A2 | Brazil | A2 | |
| AU2016206259B2 | Australia | B2 | |
| US2017237985A1 | United States of America | A1 | |
| TWI597975B | Taiwan Province of China | B | |
| CA2840481C | Canada | C | |
| PH12017500999A1 | Philippines | A1 | |
| PH12017500999B1 | Philippines | B1 | |
| PH12017501000A1 | Philippines | A1 | |
| PH12017501000B1 | Philippines | B1 | |
| PH12017501001A1 | Philippines | A1 | |
| PH12017501001B1 | Philippines | B1 | |
| PH12017501002A1 | Philippines | A1 | |
| PH12017501002B1 | Philippines | B1 | |
| TW201737713A | Taiwan Province of China | A | |
| AU2016206260B2 | Australia | B2 | |
| AU2016206261B2 | Australia | B2 | |
| EP2884749B1 | European Patent Office (EPO) | B1 | |
| PT2884749T | Portugal | T | |
| DK2884749T3 | Denmark | T3 | |
| AU2018200070A1 | Australia | A1 | |
| LT2884749T | Lithuania | T | |
| HRP20180051T1 | Croatia | T1 | |
| TWI615020B | Taiwan Province of China | B | |
| ES2655917T3 | Spain | T3 | |
| NO3064648T3 | Norway | T3 |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Changes to party contact information recordedST27 STATUS EVENT CODE: A-5-5-R10-R18-OTH-X000 (AS PROVIDED BY THE NATIONAL OFFICE)R18 | R18 | |
| Changes to party contact information recordedST27 STATUS EVENT CODE: A-5-5-R10-R18-OTH-X000 (AS PROVIDED BY THE NATIONAL OFFICE)R18 | R18 | |
| Changes to party contact information recordedST27 STATUS EVENT CODE: A-5-5-R10-R18-OTH-X000 (AS PROVIDED BY THE NATIONAL OFFICE)R18 | R18 | |
| Full renewal or maintenance fee paidU11 | U11 | |
| Annual fee paymentFPAY | FPAY | |
| Written decision to grantGRNT | GRNT | |
| Decision to grant or registration of patent rightE701 | E701 | |
| Notification of reason for refusalE902 | E902 | |
| Request for accelerated examinationA302 | A302 | |
| Divisional application of patentA107 | A107 | |
| Request for examinationA201 | A201 |
Numbers
- Publication
- 10-1560552
- Application
- 100040056
Titles4
- Korean
- 산술부호화를 수반한 비디오 부호화 방법 및 그 장치, 비디오 복호화 방법 및 그 장치
- English
- Method and apparatus for video encoding with arithmectic encoding, method and apparatus for video decoding with arithmectic decoding
- Unlabeled
- 산술부호화를 수반한 비디오 부호화 방법 및 그 장치, 비디오 복호화 방법 및 그 장치{Method and apparatus for video encoding with arithmectic encoding, method and apparatus for video decoding with arithmectic decoding}
- Unlabeled
- BACKGROUND OF THE INVENTION Field of the Invention: A video encoding method accompanied by arithmetic encoding, an apparatus therefor, a video decoding method, and an apparatus thereof TECHNICAL FIELD
Classification
- CPC, 12
- H04N19/13
- H04N19/91
- H04N19/1883
- H04N19/176
- H04N19/186
- H04N19/59
- H04N19/157
- H04N19/44
- H04N19/593
- H04N19/60
- H04N19/50
- H04N19/70
- IPC, 4
- H04N19 13
- H04N19 176
- H04N19 186
- H04N19 59