Image encoding apparatus and control method thereof
Abstract
Problem to be solved.To generate resolution interpolation data by a relatively simple process, and to enable image coding which is visually good and realizes high compression performance by a simple and high-speed process.
Solution.A tile dividing unit 103 extracts tile data of 32 × 32 pixels from an original image data to be encoded and stores it in a tile buffer 104. The resolution conversion unit 105 samples one pixel in a block of 2 × 2 pixels in the stored tile data, generates the reduced tile data constituting the reduced image, and the interpolation data generation unit 110 then sets the original resolution. Generate tile data Generate and output interpolation data. The coding method selection unit 111 outputs a control signal for executing either lossless coding or lossy coding of the reduced tile data as the interpolation data for the tile of interest, and causes the data to be executed. The code sequence forming unit 113 outputs the generated coded data and the interpolated data as coded image data with respect to the original image data. [Selection diagram] Fig. 1

Term
Projected expiry 10 December 2028.
- Priority and filed
- Published
- Today
- Projected expiry
9 claims: 3 independent, 6 dependent
- 1画像データを符号化する画像符号化装置であって、 符号化対象の入力画像データ中のN画素に対応する1画素を出力することで、前記入力画像データに対する縮小画像データを生成する縮小画像生成手段と、 前記縮小画像データから、前記入力画像データと同じ解像度の画像データを生成するため、着目N画素の補間手法を複数種類の補間手法から特定する、解像度補間情報を生成する解像度補間情報生成手段と、 前記縮小画像データにおけるブロックを可逆符号化する可逆符号化手段と、 前記ブロックを非可逆符号化する非可逆符号化手段と、 前記縮小画像データにおける着目ブロックに対応する前記解像度補間情報に基づいて、前記可逆符号化手段、前記非可逆符号化手段のいずれか一方を選択し、前記着目ブロックの符号化処理を実行させる選択手段と、 該選択手段による選択された符号化手段より得られた符号化データと、前記解像度補間情報生成手段で生成された解像度補間情報を、前記入力画像データに対する符号化画像データとして出力する出力手段とを備え、 前記解像度補間情報生成手段が生成する解像度補間情報で示される補間方法の種類には、 前記着目N画素を、当該着目N画素に対応する縮小画像中の1画素、又は、当該縮小画像中の1画素の周囲に位置する少なくとも1画素から復元する第1の補間方法、 前記着目N画素を、当該着目N画素の少なくとも一部を前記縮小画像を参照せずに復元する第2の補間方法 が含まれることを特徴とする画像符号化装置。
- 2前記選択手段は、前記着目ブロックに対する解像度補間情報の補間手法の分布に基づいて、前記可逆符号化手段、前記非可逆符号化手段のいずれか一方を選択することを特徴とする請求項1に記載の画像符号化装置。
- 3前記選択手段は、前記着目ブロックに対応する解像度補間情報の符号量から、前記補間手法の分布を推定することを特徴とする請求項2に記載の画像符号化装置。
- 4前記選択手段は、前記着目ブロックに対応する解像度補間情報の符号量Lを所定の閾値THと比較し、L≦THであれば可逆符号化手段を選択し、L THであれば非可逆符号化手段を選択することを特徴とする請求項3に記載の画像符号化装置。
- 5前記選択手段は、着目ブロックに対応する解像度補間情報のうち、一部の特定の補間手法に関する符号量から補間手法の分布を推定することを特徴とする請求項3に記載の画像符号化装置。
- 6前記縮小画像生成手段は、 前記入力画像データ中の総画素数をMとしたとき、前記入力画像データ中のN画素に対応する1画素を出力する処理をM/N回繰り返してM/N画素数の縮小画像データを生成する処理を、再帰的に所定回数実行し、 前記解像度補間情報生成手段は、前記縮小画像生成手段による1回の縮小画像データの生成を行なう度に、縮小元になる画像データから解像度補間情報を生成することを特徴とする請求項1乃至5のいずれか1項に記載の画像符号化装置。
- 7画像データを符号化する画像符号化装置の制御方法であって、 符号化対象の入力画像データ中のN画素に対応する1画素を出力することで、前記入力画像データに対する縮小画像データを生成する縮小画像生成工程と、 前記縮小画像データから、前記入力画像データと同じ解像度の画像データを生成するため、着目N画素の補間手法を複数種類の補間手法から特定する、解像度補間情報を生成する解像度補間情報生成工程と、 前記縮小画像データにおけるブロックを可逆符号化する可逆符号化工程と、 前記ブロックを非可逆符号化する非可逆符号化工程と、 前記縮小画像データにおける着目ブロックに対応する前記解像度補間情報に基づいて、前記可逆符号化工程、前記非可逆符号化工程のいずれか一方を選択し、前記着目ブロックの符号化処理を実行させる選択工程と、 該選択工程による選択された符号化工程より得られた符号化データと、前記解像度補間情報生成工程で生成された解像度補間情報を、前記入力画像データに対する符号化画像データとして出力する出力工程とを備え、 前記解像度補間情報生成工程が生成する解像度補間情報で示される補間方法の種類は、 前記着目N画素を、当該着目N画素に対応する縮小画像中の1画素、又は、当該縮小画像中の1画素の周囲に位置する少なくとも1画素から復元する第1の補間方法、 前記着目N画素を、当該着目N画素の少なくとも一部を前記縮小画像を参照せずに復元する第2の補間方法 が含まれることを特徴とする画像符号化装置の制御方法。
- 8コンピュータに読み込み込ませ実行させることで、前記コンピュータを請求項1乃至6のいずれか1項に記載の画像符号化装置として機能させるコンピュータプログラム。
- 9請求項8に記載のコンピュータプログラムを格納したことを特徴とするコンピュータ可読記憶媒体。
Independent claims9
132 paragraphs, as filed
The present invention relates to an image coding technique, particularly a technique for encoding high-resolution image data easily and with high efficiency.
Conventionally, an image is divided into blocks composed of m × n pixels, and this block is provided with a lossless coding configuration and a lossy coding configuration, and the coding result of either one is the final block of interest. An image coding technique for outputting as coded data has been proposed.
Such a technology increases the compression rate by selecting and applying a lossy coding method in a natural image region where image quality deterioration is relatively invisible, and conversely, characters, line drawings, CG parts, etc. where image quality deterioration is easily noticeable. The basic principle of this is to suppress visual image quality deterioration by using a lossy coding method.
In order to suppress image quality deterioration and reduce the amount of coding, it is important to appropriately select a coding method for each region. Patent Document 1 is known as a method for selecting lossless or lossy by referring to the amount of information and the number of colors of lossless coded data.
Lossless coding includes predictive coding technology that predicts pixel values from surrounding pixels, obtains the prediction error, and encodes the prediction error, or run-length that converts the prediction error into a continuous number of the same pixel value and encodes it. Coding techniques are often used. On the other hand, as lossy coding, a transform coding technique that converts image data into frequency space data using DCT or wavelet transform and encodes a coefficient value is common. Various devices have been proposed to obtain high compression performance by introducing more advanced operations for each of the lossless and lossy coding technologies.<patcit num="1"><text>Japanese Unexamined Patent Publication No. 08-167030</text></patcit>
<p> In recent years, with the increase in accuracy of image input / output devices, the resolution of image data has been increasing. In order to process high-resolution image data at high speed, more hardware resources and processing time are required than before. For example, if a coding device that handles an image with a resolution of 1200 dpi performs coding processing within the same time as a device that encodes an image of 600 dpi before that, the internal computing power is quadrupled. There is a need to.</p><p> Therefore, high-efficiency lossless and lossy coding processing requires advanced arithmetic processing, and therefore, an image coding technology that satisfies the three requirements of compression performance, image quality, and arithmetic cost in a well-balanced manner is required.</p><p> The present invention has been made in view of the above-mentioned problems, and an object of the present invention is to provide a coding technique that achieves both high compression and high image quality by simple processing.</p>
<p> In order to solve such a problem, the image coding apparatus of the present invention has the following configuration. That is, An image coding device that encodes image data. A reduced image generation means that generates reduced image data for the input image data by outputting one pixel corresponding to N pixels in the input image data to be encoded. In order to generate image data having the same resolution as the input image data from the reduced image data, a resolution interpolation information generation means for generating resolution interpolation information, which specifies an interpolation method for N pixels of interest from a plurality of types of interpolation methods, A lossless coding means for losslessly coding a block in the reduced image data, An irreversible coding means for irreversibly coding the block, A selection means that selects either one of the lossless coding means and the lossy coding means based on the resolution interpolation information corresponding to the block of interest in the reduced image data, and executes the coding process of the block of interest. When, The coding data obtained from the coding means selected by the selection means and the output means for outputting the resolution interpolation information generated by the resolution interpolation information generation means as the coded image data for the input image data. Prepare, The type of interpolation method indicated by the resolution interpolation information generated by the resolution interpolation information generation means is A first interpolation method in which the N pixel of interest is restored from one pixel in a reduced image corresponding to the N pixel of interest, or at least one pixel located around one pixel in the reduced image. A second interpolation method in which at least a part of the N pixels of interest is restored without referring to the reduced image. Is included.</p>
<p> According to the present invention, by configuring the generation of resolution interpolation data by a relatively simple process, it is possible to perform image coding that realizes visually good and high compression performance by simple and high-speed processing. Becomes possible. In particular, image data without sensor noise, such as a PDL rendered image, has a lot of spatial redundancy and can be efficiently used in reduction processing and interpolation data processing, and is therefore suitable for encoding.</p>
Hereinafter, embodiments according to the present invention will be described in detail with reference to the accompanying drawings.
[First Embodiment] FIG. 1 is a block configuration diagram of an image coding device according to an embodiment.
The reduced image generation means of the image encoding device of the present embodiment generates reduced image data by outputting one pixel corresponding to the N pixel of interest (N is an integer of 2 or more) in the input image data from the outside of the device. To do. That is, when the total number of pixels of the input image data is M, the reduced image generation means repeats the process of outputting one pixel corresponding to the N pixels of interest M / N times to reduce the number of M / N pixels. To generate. Then, the resolution interpolation information generation means generates resolution interpolation information that specifies the interpolation method of the N pixels of interest for restoring the image of the original resolution from a plurality of types of interpolation methods. Then, a code string composed of the resolution interpolation data and the coded data of the reduced image data is generated and output. Here, the type of the interpolation method indicated by the resolution interpolation information generated by the resolution interpolation information generation means is that the N pixels of interest are one pixel in the reduced image corresponding to the N pixels of interest, or one in the reduced image. It includes a first interpolation method that restores from at least one pixel located around the pixel, and a second interpolation method that restores the N pixel of interest and at least a part of the N pixel of interest without referring to the reduced image. It is characterized by that.
The input source of the input image data is an image scanner, but it may be a storage medium or the like that stores the image data as a file, or a rendering unit that renders based on information such as print data. It doesn't matter what kind it is. Further, in the embodiment, an example in which the number of pixels represented by the above N is 2 × 2 (= 4 pixels) will be described.
The image data to be encoded by the image coding apparatus according to the present embodiment is RGB color image data, and each component (color component) is 8-bit pixel data expressing a brightness value in the range of 0 to 255. It shall be composed. It is assumed that the image data to be encoded is arranged in point order, that is, each pixel is arranged in the order of raster scan, and each pixel is arranged in the order of R, G, B. The image shall be composed of W pixels in the horizontal direction and H pixels in the vertical direction. For the sake of simplicity, W and H will be described as being an integral multiple of the coding processing unit and the tile size (32 pixels in both the horizontal and vertical directions in this embodiment), which will be described later. However, the input image data may be a color space other than RGB, for example, YCbCr or CMYK, and the type of the color space and the number of components do not matter. Further, the number of bits of one component is not limited to 8 bits, and the number of bits exceeding 8 bits may be used.
Hereinafter, the coding process in the image coding apparatus shown in FIG. 1 will be described. In the figure, 115 is a control unit that controls the operation of the entire image coding apparatus of FIG. Each processing unit operates in cooperation under the control of the control unit 115.
First, the image input unit 101 sequentially inputs the image data to be encoded and outputs the image data to the line buffer 102. The input order of the image data is the raster scan order as described above.
The line buffer 102 has an area for storing a predetermined number of lines (Th) of image data, and sequentially stores the image data input from the image input unit 101. Here, Th is the number of pixels in the vertical direction of the rectangular block (tile) cut out by the tile dividing portion 103, which will be described later, and is 32 in the present embodiment. Hereinafter, the data obtained by dividing the image data to be encoded by the width of the Th line is referred to as a stripe. The capacity required for the line buffer 102, that is, the amount of data for one stripe, is W × Th × 3 (RGB minutes) bytes. As described above, for convenience of explanation, it is assumed that the number of pixels H in the vertical direction is an integral multiple of Th, and incomplete stripes do not occur at the end of the image.
By the way, when the image data of one stripe, that is, the image data of the Th line is stored in the line buffer 102, the tile dividing unit 103 uses the image data of the Th line stored in the line buffer 102 as the horizontal Tw pixels. , It is divided into rectangular blocks composed of vertical Th pixels, read in block units, and stored in the tile buffer 104. As described above, since the number of pixels W in the horizontal direction of the image to be encoded is an integral multiple of Tw, it is assumed that an incomplete block does not occur when the image is divided into rectangular blocks. Further, the rectangular block composed of the horizontal Tw pixels and the vertical Th pixels will be hereinafter referred to as "tiles". In the embodiment, since Tw = Th = 32, the size of one tile is 32 × 32 pixels.
FIG. 8 illustrates the relationship between the image data to be encoded and the stripes and tiles. As shown in the figure, the i-th tile in the horizontal direction and the j-th tile in the vertical direction are indicated as T (i, j) in the image.
The tile buffer 104 has an area for storing pixel data for one tile, and sequentially stores the tile data output from the block division unit 103. The capacity required for the tile buffer 104 is Tw × Th × 3 (RGB minutes) bytes.
The resolution conversion unit 105 performs subsampling to extract one pixel from the area of 2 × 2 pixels for the tile data stored in the tile buffer 104, and generates 1/2 reduced tiles. Here, the 1/2 reduced tile represents a tile in which the tile data stored in the tile buffer 104 is halved in the number of pixels in both the horizontal and vertical directions. Further, the image data in the area of 2 × 2 pixels is hereinafter simply referred to as a pixel block.
Figure 3 shows the tile data and the pixel block B of 2 x 2 pixels.<sub>n</sub>The relationship is illustrated. Also, the pixel block B of interest<sub>n</sub>As shown in the figure, B<sub>n</sub>(0,0), B<sub>n</sub>(0,1), B<sub>n</sub>(1,0), B<sub>n</sub>Expressed as (1,1). Also, attention block B<sub>n</sub>B the pixel block in front of (next to the left)<sub>n-1</sub>, B the pixel block next to Bn (next to the right)<sub>n + 1</sub>And the block under Bn is B<sub>n + b</sub>It is expressed as. Here, "b" is the number of blocks in the horizontal direction in the tile, and when expressed using the number of pixels in the horizontal direction Tw of the tile, b = Tw / 2. In the embodiment, since Tw = 32, b = 16.
The resolution conversion unit 105 of the present embodiment is the pixel block B of interest.<sub>n</sub>Of which, pixel B at the upper left corner<sub>n</sub>Extract (0,0) as one pixel of 1/2 reduced tile. All pixel blocks B in tile data<sub>0</sub>~ B<sub>m</sub>Sub-sampling is performed for (m = Tw / 2 x Th / 2-1 = 16 x 16-1 = 255), and 1/2 reduced tiles with 1/2 the number of pixels in both the horizontal and vertical directions are created with respect to the original tile. Generate. The resolution conversion unit 105 stores the generated 1/2 reduced tile data in the 1/2 reduced tile buffer 106. Therefore, in the case of the embodiment, it functions as a reduced image generation means for generating a reduced image having the number of W × H / 4 pixels from the original image data composed of W × H pixels.
The interpolation data generation unit 110 generates information necessary for restoring the original tile data from the 1/2 reduced tile. That is, the pixel B that is the subsampling target in the pixel block.<sub>n</sub>3 pixels for non-sampling, excluding (0,0) {B<sub>n</sub>(0,1), B<sub>n</sub>(1,0), B<sub>n</sub>It generates interpolated data indicating how to interpolate (1,1)}.
FIG. 12 shows a block configuration diagram inside the interpolation data generation unit 110 according to the present embodiment. The interpolation data generation unit 110 is composed of a flat determination unit 1201 and a non-flat block analysis unit 1202.
The flat determination unit 1201 performs "flat determination" to determine whether or not the image data in the tile can be restored by simple enlargement (pixel repetition) processing from the pixels of the reduced tile data. The minimum unit for flat determination is a pixel block of 2 × 2 pixels. Hereinafter, when this pixel block is a block that can be reproduced by simple enlargement, it is referred to as a flat block, and a block that is not is referred to as a non-flat block. In FIG. 3, the pixel block B of interest<sub>n</sub>Is a flat block when the following equation (1) holds. B<sub>n</sub>(0,0) = B<sub>n</sub>(0,1) = B<sub>n</sub>(1,0) = B<sub>n</sub>(1,1) ...(1)
The role of the flat determination unit 1201 is to generate information indicating whether each pixel block constituting the tile is a flat block or a non-flat block. Hereinafter, the processing procedure of the flat determination unit 1201 will be described with reference to the flowchart of FIG. Note that FIG. 13 is a process for one tile.
If all the pixel blocks included in the tile are flat blocks, the tile is called a flat tile, and if not, it is called a non-flat tile. Similarly, a group of blocks of Tw / 2 (16 in the embodiment) that are continuous in the horizontal direction in the tile is hereinafter referred to as a block line. When all the blocks in the block line are flat blocks, the block line is called a flat block line, and when at least one non-flat block is included, it is called a non-flat block line.
First, the flat determination unit 1201 inputs tile data of Tw × Th pixels (step S1301). Each tile is processed independently, and when referring to a pixel (peripheral pixel) outside the tile, the value of each color component is assumed to be 255.
Next, in step S1302, the flat determination unit 1201 determines flat / non-flat for all the pixel blocks of 2 × 2 pixels constituting the tile of Tw × Th pixels. That is, it is determined whether or not the tile of interest is a flat tile. When it is determined that the tile of interest is a flat tile, the process is moved to step S1303, the 1-bit flag 1 is output to the non-flat block analysis unit 1202, and this process is completed.
On the other hand, when at least one non-flat block exists in the tile of interest (step S1302 is NO), the flat determination unit 1201 shifts the processing to step S1304 and analyzes the 1-bit flag 0 in the non-flat block. Output to unit 1202 and move to step S1305.
In step S1305, we focus on one of the block lines consisting of b (= Tw / 2 = 16) pixel blocks that are continuous in the horizontal direction, and determine whether it is a flat block line or a non-flat block line, as in step S1302. Do. If it is a flat block line (YES), the process is transferred to step S1306, and the 1-bit flag 1 is output to the non-flat block analysis unit 1202. On the other hand, when it is determined that the block line of interest is a non-flat block line (step S1305 is NO), the process is moved to step S1307, and the 1-bit flag "0" is output to the non-flat block analysis unit 1202. , Move the process to step S1308.
When reaching step S1308, there will be at least one non-flat block in the block line of interest. Therefore, for each of the b (= Tw / 2 = 16) blocks constituting the block line of interest, an array of flags indicating the determination result of the flat / non-flat block is required. Here, the flag of the block may be 1 bit, and is set to "1" in the case of a flat block and "0" in the case of a non-flat block. In the case of the embodiment, since the block line is composed of b (= 16) pixel blocks, the flag array of the non-flat block line is b, that is, 16 bits.
In step S1308, the flag array of the block line of interest is compared with the flag array of the previous non-flat block line, and it is determined whether or not they match. If the block line of interest is the first block line of the tile of interest, there is no previous block line. Therefore, the flat determination unit 1201 prepares in advance information indicating that all the blocks of the block line are non-flat blocks in the internal memory prior to the determination of the block line of the tile of interest. That is, when a flat block is defined as "1" and a non-flat block is defined as "0", b (b bits) of "0" are prepared as initial values. Hereinafter, this b bit is referred to as "reference non-flat block line information". Of course, even in the decoding device, when restoring one tile, "reference non-flat block line information" is prepared.
Therefore, in step S1308, it is determined whether or not the arrangement of the determination results of the flat block / non-flat block of each block in the block of interest matches the reference non-flat block line information.
Here, the determination process of step S1308 will be described by taking the tile data of FIG. 10 as an example. In this tile data, the 4th block line from the beginning is a flat block line, and the 5th block line is an example of a non-flat block line. The flat / non-flat determination result of the pixel block with respect to the tile of FIG. 10 is as shown in FIG.
Now, suppose that the fifth block line is the block line of interest. Since the block line of interest is the fifth block line, it is determined that the block line is the first non-flat block line in the tile of interest. Therefore, the determination result of the pixel block included in the block line of interest is compared with the reference non-flat block line information (16-bit 000 ... 000), and it is determined that they match.
In this way, when it is determined that the determination result of the block line of interest and the reference non-flat block line information match, the determination result of each block in the block line of interest is the reference non-flat block line in step S1309. The 1-bit flag 1 indicating that the information is matched is output to the non-flat block analysis unit 1202. In this case, it is not necessary to output the b (= Tw / 2) bit flag indicating the flat / non-flat determination result for each block.
On the other hand, if the array of flags indicating the flat / non-flat determination result of the block line of interest does not match the reference non-non-flat block line (step S1308 is NO), the process is moved to step S1310. Here, the flat determination unit 1201 first outputs a 1-bit flag 0 to the non-flat block analysis unit 1202. Then, following that, a flag array (Tw / 2 bits) indicating the flat / non-flat determination result of the block line of interest is output to the non-flat block analysis unit 1202. After that, the flat determination unit 1201 updates the reference non-flat block line information with the determination result (b bit) of the flat block / non-flat block of the block line of interest in step S1311.
When the flat / non-flat determination for the block line of interest is completed, the process is moved to step S1312, and it is determined whether or not the block line of interest is the final block line in the tile of interest. If it is not the final block line (NO), the block line of interest is moved to the next block line (step S1313). Then, the process is returned to step S1305 and the same process is repeated. If it is the final block line (YES), the flat determination process for the tile data is terminated.
The flat determination unit 1201 performs the above processing on the tile data of interest, and outputs the flat / non-flat determination result to the non-flat block analysis unit 1202.
Next, the processing of the non-flat block analysis unit 1202 will be described.
Similar to the flat determination unit 1201, this non-flat block analysis unit 1202 sets "reference non-flat block line information" indicating that all blocks of one block line are non-flat blocks at the start of processing corresponding to one tile. Keep it. Then, the non-flat block analysis unit 1202 analyzes the determination result from the flat determination unit 1201 and passes it as it is, and outputs it to the interpolation data buffer 112. At this time, if the non-flat block analysis unit 1202 determines in step S1310 of FIG. 13 that there is a block line for which the 1-bit flag 0 is output, the flat block / of each subsequent block The reference non-flat block line information is updated with the judgment result indicating the non-flat block. If the processing of step S1309 of FIG. 13 is performed on the block line of interest, it is possible to determine which block is the non-flat block by referring to the reference non-flat block line at that time. If it is determined that the processing of step S1310 in FIG. 13 is being performed on the block line of interest, the block is non-flat by examining the value of the b bit following the flag 0 of the first 1 bit. It can be determined whether it is a block. That is, the non-flat block analysis unit 1202 can determine the positions and the number of all non-flat blocks in the tile of interest. Then, the non-flat block analysis unit 1202 analyzes the number of colors and the arrangement in the pixel block (non-flat block) composed of 2 × 2 pixels that cannot be reproduced by simply enlarging the pixels of the reduced tile. Then, additional information is generated and output. That is, the non-flat block analysis unit 1202 receives the flat / non-flat determination result output from the flat determination unit 1201, analyzes the non-flat block, and generates information for restoring the non-flat block. ,Output.
The non-flat block analysis unit 1202 performs a non-flat block coding process on a block determined to be a non-flat block according to the flowchart of FIG. Hereinafter, the processing of the non-flat block analysis unit 1202 will be described with reference to the figure.
First, in step S1501, one non-flat block B<sub>n</sub>Is the block of interest, and pixel block data of 2 × 2 pixels is input. Then, in step S1502, the pixel block B of interest<sub>n</sub>However, it is determined whether or not the block satisfies the following equation (2). B<sub>n</sub>(0,1) = B<sub>n + 1</sub>(0,0) and B<sub>n</sub>(1,0) = B<sub>n + b</sub>(0,0) and B<sub>n</sub>(0,1) = B<sub>n + b + 1</sub>(0,0) ...(2)
Block B for which the above equation (2) holds<sub>n</sub>Is hereinafter defined as a surrounding 3-pixel matching block. As shown in the above equation (2), the pair {Bn (0,1), B<sub>n + 1</sub>(0,0)}, {B<sub>n</sub>(1,0), B<sub>n + b</sub>(0,0)} and {B<sub>n</sub>(0,1) and B<sub>n + b + 1</sub>There is a reason to determine a match / mismatch with (0,0)}.
In general, the degree of correlation between the pixel of interest and the pixel adjacent to the pixel of interest is high, and there are many such images. Therefore, when predicting the pixel value of the pixel of interest, the adjacent pixel is often used as a reference pixel for prediction. Pair {Bn (0,1), B under the condition that all four pixels in the block of interest are not the same, that is, they are not flat blocks.<sub>n + 1</sub>(0,0)}, {B<sub>n</sub>(1,0), B<sub>n + b</sub>(0,0)} and {B<sub>n</sub>(0,1) and B<sub>n + b + 1</sub>Experiments have shown that (0,0)} is likely to match.
For this reason, B<sub>n</sub>(0,1) and B<sub>n + 1</sub>Changed to judge whether or not (0,0) matches. Another pair {B<sub>n</sub>(1,0), B<sub>n + W</sub>(0,0)} and {B<sub>n</sub>(0,1) and B<sub>n + W + 1</sub>For the same reason, compare (0,0)}. Then, when the above equation (2) is satisfied (when the three pairs are equal to each other), the amount of information can be reduced by assigning short codewords.
In the embodiment, the pixels at the upper left corner of each pixel block are sub-sampled to generate a reduced image. Therefore, pixel B in Eq. (2)<sub>n + 1</sub>(0,0), B<sub>n + b</sub>(0,0), B<sub>n + b + 1</sub>Note that (0,0) is also a subsampling target pixel in the pixel block adjacent to the pixel block Bn of interest.
Now, when the block of interest Bn is a block for which equation (2) holds, the three pixels at the positions of Bn (0,1), Bn (1,0), and Bn (1,1) are around. It can be judged from the pixels that it can be reproduced. However, when the block is located at the right end of the image or the lower end of the image, the pixels outside the block cannot be referred to. Therefore, it is decided that each component of the outer pixel is virtually set to an arbitrary value in advance, and a match / mismatch determination with the pixel value is performed. In this embodiment, each component value is virtually set to "255". However, the value is not limited to "255", and any other value may be used as long as it is specified that the same value is used on the coding side and the decoding side.
FIG. 4 shows a pixel block of 2 × 2 pixels of interest. Pixel X shown in FIG. 4 indicates a pixel sub-sampled to generate a reduced image, and pixel B in the pixel block Bn of interest.<sub>n</sub>It corresponds to (0,0). In addition, pixels Xa, Xb, and Xc are the blocks of interest B.<sub>n</sub>Pixels in (0,1), B<sub>n</sub>(1,0), B<sub>n</sub>The pixel of (1,1) is shown. Hereinafter, the description will be made using pixels X, Xb, Xc, and Xd.
By the way, when it is determined in step S1502 that the equation (2) holds, the attention block B<sub>n</sub>The pixels Xa, Xb, and Xc in the image can be reproduced as they are from the pixels of the reduced image. Therefore, the process is transferred to step S1503, and a 1-bit flag 1 indicating that the process can be reproduced is output. On the other hand, if it is determined that the block of interest does not satisfy the equation (2), the process is transferred to step S1504 and the 1-bit flag 0 is output.
Next, the process proceeds to step S1505, and it is determined whether the number of colors included in the block of interest is "2" or more (3 or 4 colors). Since the block of interest is a non-flat block, the number of colors cannot be "1".
If it is determined that the number of colors included in the block of interest is greater than "2" (3 or 4 colors), the 1-bit flag "0" is output in step S1512. Then, in step S1513, pixel data of three pixels Xa, Xb, and Xc that are not used as a reduced image is output. In the embodiment, since one pixel is three components of R, G, and B, and one component is 8 bits, the total number of bits of the three pixels Xa, Xb, and Xc is 3 × 8 × 3 = 72 bits.
On the other hand, when it is determined that the number of colors included in the block of interest is "2", a 1-bit flag "1" indicating that the number of appearing colors is "2" is output in step S1506. Then, in step S1507, the pixels in the block of interest indicating which of the patterns 1601 to 1607 of FIG. 16 is the three pixels Xa, Xb, Xc excluding the pixel X of FIG. 4 used for the reduced image. Outputs placement information. As shown in FIG. 16, since there are seven patterns in total, one pattern can be specified if the pixel arrangement information is 3 bits. However, in the embodiment, a pixel having the same color as the pixel used in the reduced image is represented as "1", a pixel having a different color is represented as "0", and each bit corresponding to the pixels Xa, Xb, Xc in FIG. Output in that order (again, 3 bits). For example, when the pattern 1601 in FIG. 16 is matched, the pixels Xa, Xb, and Xc have the same color and are different from the color of the pixel X used in the reduced image. Therefore, the 3 bits of the pixel arrangement information are 000. ". In the case of pattern 1602, it is "100", in the case of pattern 1603, it is "010", and in the case of pattern 1604, it is "001". Others need not be explained.
Then, in step S1508, the attention block B<sub>n</sub>Among the pixels Xa, Xb, and Xc excluding the pixel X used in the reduced image, it is determined whether or not a pixel having the same color as the pixel having a color different from the pixel X exists in the vicinity. In the present embodiment, the pixels in the vicinity to be compared are the block B to the right of the block of interest.<sub>n + 1</sub>Pixel B<sub>n</sub>(0,0), block B directly below the block of interest<sub>n + b</sub>Pixel B<sub>n + b</sub>(0,0), block B adjacent to the lower right of the block of interest<sub>n + b + 1</sub>Pixel B<sub>n + b + 1</sub>Set to (0,0) and compare in this order. Since it is only necessary to indicate which of the three pixels matches, the comparison result is set to 2 bits. And the match is B<sub>n + 1</sub>If it is (0,0), it is a 2-bit 11, B<sub>n + b</sub>If it matches (0,0), 2-bit 01, B<sub>n + b + 1</sub>If (0,0) matches, a 2-bit 10 is output (step S1509). If it does not match any of the three peripheral pixels (NO), the process is moved to step S1510, and the 2-bit flag 00 is output. Then, the pixel value of the second color is output (step S1511), and the attention block B<sub>n</sub>Ends the processing for. In the present embodiment, since the image to be coded is 8-bit RGB data for each color, 24 bits are output in step S1511.
In this way, attention block B<sub>n</sub>When the process of is completed, it is determined in step S1514 whether or not the block of interest is the last non-flat block in the tile of interest. If not, the process proceeds to the next non-flat block process (step S1515). The processing from step S1501 is performed in the same manner for the next non-flat block. If it is determined that the block of interest is the last non-flat block, this process ends.
As described above, the non-flat block analysis unit 1202 outputs the flat / non-flat determination result and the interpolation data generated by the analysis result of the non-flat block. These data are stored in the interpolated data buffer 112.
As can be seen from the above description, the pixel block represented by 2 × 2 pixels can be classified into the following five types (a) to (e). (a) Flat block (b) Non-flat block and surrounding 3-pixel matching block (c) A non-flat block in which the number of colors appearing in the block is "2" and the peripheral pixels have pixels of the same color as the second color. (d) A non-flat block in which the number of colors appearing in the block is "2" and the peripheral pixels do not have pixels of the same color as the second color. (e) A non-flat block with more colors (3 or 4 colors) appearing in the block than "2".
While performing the above processing, the interpolation data generation unit 110 determines which of the above (a) to (e) the pixel block corresponds to, and counts the number of appearances of each block. Hereinafter, the number of blocks of the above types (a) to (e) will be referred to as Na, Nb, Nc, Nd, and Ne.
The interpolation data generation unit 110 initializes Na to Ne with 0 prior to generating the interpolation data for one tile, and counts the number of blocks when the generation and output of the interpolation data for one tile is completed. Na to Ne are output to the coding method selection unit 111.
In the embodiment, since one tile is composed of 32 × 32 pixels, there are 16 × 16 pixel blocks in the tile. Therefore, if the interpolation data generation unit 110 performs the process of step S1303 in FIG. 13, all the pixel blocks in the tile of interest are flat blocks, Na = 256, and the number of blocks other than Na Nb. To Ne can be determined as "0". Further, when the process of step S1306 is performed, "16" may be added to the Na counted up to that point. Then, when the process of step S1503 in FIG. 15 is performed, 1 is added to Na. If the process of step S1512 is performed, "1" is added to Ne. Other than that, you can fully understand from the explanation so far.
The coding method selection unit 111 encodes the 1/2 reduction tile data stored in the 1/2 reduction tile buffer 106 based on the frequency distribution of the number of blocks Na to Ne counted by the interpolation data generation unit 110. Is determined, and the determination result is output as a binary (1 bit) control signal to the lossy coding unit 107, the lossy coding unit 108, and the code sequence forming unit 113. As a result, either one of the lossy coding unit 107 and the lossless coding unit 108 performs the coding process of the 1/2 reduction tile data stored in the 1/2 reduction tile buffer 106. Then, the generated coded data is output to the code string forming unit 113. In the embodiment, when the control signal from the coding method selection unit 111 is 0, the lossy coding unit 107 performs the lossy coding process, and when the control signal is 1, the lossy coding unit 108 performs the lossy coding process. Processing shall be performed.
In general, in characters, line arts, clip art images, etc., there is a tendency that there are many types (a) to (c) of blocks shown above. In other words, in the case of characters, line art, clip art images, etc., the values of the number of blocks Na to Nc become large. On the other hand, in complex CG images and natural images, the values of the number of blocks Nd and Ne tend to be large. Further, in characters, line arts, clip arts, etc., image quality deterioration due to lossy coding is easily noticeable, and high compression performance tends to be obtained by lossless coding. On the other hand, in complicated CG and natural images, on the contrary, the deterioration of image quality due to lossy coding tends to be less noticeable, and high compression performance tends not to be obtained by lossless coding.
Furthermore, the pixel blocks of the block types (a) to (c) restore the original image using the pixel values of the reduced image. Therefore, if a pixel change (deterioration) occurs in the reduced image, it acts in the direction of spreading the deterioration when restoring the image of the original resolution at the time of decoding. Therefore, it is easily affected by image quality deterioration during reduced image coding.
On the contrary, the pixel blocks of the block types (d) and (e) are less affected by the deterioration of image quality because the information of the pixels (colors) not included in the reduced image is directly specified.
Since it is as described above, when the values of Na, Nb, and Nc are large, the coding method selection unit 111 outputs a control signal 1 so that lossless coding is applied to the 1/2 reduction tile data. .. Further, when the values of Nd and Ne are large, the coding method selection unit 111 outputs a control signal 0 so as to apply lossy coding to the 1/2 reduction tile.
Specifically, the coding method selection unit 111 compares the preset threshold value TH1 with the total value of Na + Nb + Nc, and outputs a control signal 1 when the following equation (3) is satisfied. If it is not satisfied, the control signal "0" is output. Na + Nb + Nc> TH1 ... (3) The following equation (4) may be adopted instead of the above equation (3). It is understandable that these equations (3) and (4) are equivalent to each other. Nd + Ne TH1 ... (4) The threshold value TH1 can be appropriately set by the user, but typically, a value indicating 1/2 of the number of blocks included in the tile is preferable. In embodiments, 1 Thailand since the Le includes 16 × 16 pixels blocks, a TH1 = 128.
By the way, when the control signal output from the coding method selection unit 111 is 0, the lossy coding unit 107 does not display the 1/2 reduction tile data stored in the 1/2 reduction tile buffer 106. The coded data is generated by encoding with a lossy coding method, and the generated coded data is stored in the coded data buffer 109.
Various methods can be applied as the lossy coding process in the lossy coding unit 107. Here, the JPEG (ITU-T T.81 | ISO / IEC10918-1) baseline method recommended as the international standard method for still image coding shall be applied. Since there is a detailed explanation of JPEG in the recommendations, etc., the explanation is omitted here. However, the same Huffman table and quantization table used for JPEG coding shall be used for all tiles, and the frame header, scan header, various tables, etc. that are common to all tiles shall be stored in the coded data buffer 109. It is not stored, and only the encoded data part is stored. That is, in the general JPEG baseline coded data configuration shown in FIG. 7, only the entropy-coded data segment from immediately after the scan header to immediately before the EOI marker is stored. For the sake of simplicity, the restart interval is not defined by the DRI and RST markers, and the number of lines is not defined by the DNL marker.
The lossless coding unit 108 losslessly encodes the 1/2 reduction tile data stored in the 1/2 reduction tile buffer 106 when the control signal output from the coding method selection unit 111 is 1. Generate coded data. Then, the lossless coding unit 108 stores the generated coded data in the coded data buffer 109. There are various types of lossless coding performed by the lossless coding unit 108, and here, as an example, JPEG-LS (ITU-T T.87 |) recommended by ISO and ITU-T as international standard methods. ISO / IEC 14495-1) shall be used. However, other lossless coding techniques such as JPEG (ITU-T T.81 | ISO / IEC10918-1) lossless mode may be used. Since the details of JPEG-LS are also described in the recommendations, etc., the description is omitted here. Also, as with JPEG, headers and the like are not stored in the buffer, and only the coded data part is stored.
The code string forming unit 113 includes a control signal output from the coding method selection unit 111, interpolation data stored in the interpolation data buffer 112, and coded data of 1/2 reduction tiles stored in the coding data buffer 109. Is combined, and necessary additional information is added to form a code string that is an output of the image coding apparatus as coded image data corresponding to the original image data. At this time, the code sequence forming unit 113 also adds information indicating which tile is losslessly coded and which tile is lossy coded based on the control signal output from the coding method selection unit 111. And output.
FIG. 9A is a diagram showing a data structure of an output code string of the image coding apparatus. At the beginning of the output code string, information necessary for decoding the image, such as the number of horizontal pixels, the number of vertical pixels, the number of components, the number of bits of each component, the width of the tile, and the height of the image, etc. Additional information is attached as a header. Further, this header portion includes not only information about the image data itself but also information about coding such as a Huffman coding table and a quantization table commonly used for each tile. FIG. 9B is a diagram showing the configuration of the output code string of each tile. At the head of each tile is a tile header containing various information necessary for decoding, such as the tile number and size, and then a control signal output from the coding method selection unit 111 is added. Here, it is shown separately from the tile header for the sake of explanation, but it goes without saying that it may be included in the tile header. Although not shown in FIGS. 9 (a) and 9 (b), the length of the code string of each tile is managed by including it in the header part at the beginning of the tile or the beginning of the coded data. Random access on a tile-by-tile basis may be possible. Alternatively, the same effect can be obtained by setting a special marker by devising so that a predetermined value does not occur in the coded data and placing a marker at the beginning or the end of each tile data.
The code output unit 114 outputs the output code string generated by the code string forming unit 113 to the outside of the device. If the output target is a storage device, it will be output as a file.
As described above, in the image coding apparatus of the present embodiment, the coding process is performed in units of tiles composed of a plurality of pixels (the size of 32 × 32 pixels in the embodiment). For each tile, a reduced image and interpolation data for restoring the original image from the reduced image are generated, and the reduced image is coded by lossless coding or lossy coding. Generation of interpolated data is a simple and simple process, and by limiting the bitmap coding process (coding process by lossy coding unit 107 and lossless coding unit 108) that requires advanced arithmetic processing to reduced images. , The coding process can be realized easily and at high speed as compared with the case of bitmap coding the entire image. Either reversible or lossy coding is selected and applied for coding the reduced image, but it is reversible when the effect on image quality when applying lossy coding based on the structure of the interpolated data is large. Since the coding is applied, it is possible to prevent the deterioration of the image quality from being magnified when the resolution is restored at the time of decoding.
[Modified example of the first embodiment] An example in which the same processing as that of the first embodiment is realized by a computer program will be described below as a modified example of the first embodiment.
FIG. 14 is a block configuration diagram of an information processing device (for example, a personal computer) in this modified example.
In the figure, 1401 is a CPU, which controls the entire apparatus using programs and data stored in RAM 1402 and ROM 1403, and also executes image coding processing and decoding processing described later. The 1402 is RAM and includes an area for storing programs and data downloaded from an external device via an external storage device 1407, a storage medium drive 1408, or an I / F 1409. The RAM 1402 also has a work area used by the CPU 1401 to execute various processes. 1403 is a ROM that stores the boot program and the setting program and data of this device. The 1404 and 1405 are keyboards and mice, respectively, and can input various instructions to the CPU1401.
The 1406 is a display device, which is composed of a CRT and a liquid crystal screen, and can display information such as images and characters. 1407 is a large-capacity external storage device such as a hard disk drive device. The external storage device 1407 stores an OS (operating system), a program for image coding and decoding processing described later, image data to be encoded, coded data of the image to be decoded, and the like as files. In addition, the CPU1401 loads these programs and data into a predetermined area on the RAM1402 and executes them.
The 1408 is a storage medium drive that reads programs and data recorded on a storage medium such as a CD-ROM or DVD-ROM and outputs them to the RAM 1402 or the external storage device 1407. It should be noted that the storage medium may record a program for image coding and decoding processing, which will be described later, image data to be encoded, coded data of the image to be decoded, and the like. In this case, the storage medium drive 1408 loads these programs and data into a predetermined area on the RAM 1402 under the control of the CPU 1401.
The 1409 is an I / F, which connects an external device to the device and enables data communication between the device and the external device. For example, the image data to be encoded and the encoded data of the image to be decoded can be input to the RAM 1402 of the present device, the external storage device 1407, or the storage medium drive 1408. The 1410 is a bus that connects the above-mentioned parts.
In the above configuration, when the power of this device is turned on, the CPU 1401 loads the OS from the external storage device 1407 into the RAM 1402 according to the boot program of the ROM 1403. As a result, keyboard 1404 and mouse 1405 can be input, and the GUI can be displayed on the display device 1406. When the user operates the keyboard 1404 or the mouse 1405 and gives an instruction to start the image coding processing application program stored in the external storage device 1407, the CPU 1401 loads the program into the RAM 1402 and executes it. As a result, this device functions as an image coding device.
Hereinafter, the processing procedure of the application program for image coding executed by the CPU 1401 will be described with reference to the flowchart of FIG. Basically, this program will include functions (or subroutines) corresponding to each component shown in FIG. However, each area of the line buffer 102, the tile buffer 104, the 1/2 reduced tile buffer 106, the encoded data buffer 109, the interpolated data buffer 112, etc. in FIG. 1 is reserved in the RAM 1402 in advance.
First, in step S200, an initialization process before starting coding is performed. Here, header information to be included in the code string is prepared for the image data to be encoded, and various memories, variables, and the like are initialized.
After the initialization process, in step S201, the image data to be encoded is sequentially input from the external device connected by the I / F 1409, and the data for one stripe of the image data to be encoded is stored in the RAM 1402. (Corresponds to the processing of the image input unit 101).
Next, in step S202, the tile data of interest is cut out from the stripe stored in the RAM 1402 and stored at another address position of the RAM 1402 (corresponding to the processing of the tile dividing portion 103).
Subsequently, in step S203, 1/2 reduction tile data is generated by subsampling the pixels in the upper left corner of each pixel block in the tile data of interest and stored in the RAM 1402 (process of the resolution conversion unit 105). Equivalent to).
In step S204, interpolation data for restoring the tile data of the original resolution from the 1/2 reduction tile data is generated (corresponding to the processing of the interpolation data generation unit 110).
In step S205, it is determined from the structural information of the interpolated data whether to apply the reversible or lossy coding method for the 1/2 reduction tile data (corresponding to the processing of the coding method selection unit 111).
When the selected coding method is lossless coding (step S206), the 1/2 reduction tile data is losslessly coded in step S207 to generate coded data and stored in RAM 1402 (lossless coding unit 108). Equivalent to processing).
If the selected coding method is lossy coding (step S206), the 1/2 reduced tile data is irreversibly coded in step S208 to generate coded data and stored in RAM1402 (lossy coding). Corresponds to the processing of part 107).
Subsequently, in step S209, the tile header, the information for identifying lossless / irreversible, the interpolated data, and the 1/2 reduced tile coded data are combined to form the tile coded data (code string). Corresponds to the processing of forming part 113).
In step S210, it is determined whether the coded tile is the last tile of the currently loaded stripe. If it is not the last tile, the process is repeated from step S202 for the next tile. If it is the last tile of the stripe, the process moves to step S211.
In step S211 it is determined whether the currently loaded stripe is the final stripe of the image. If it is not the final stripe, the process returns to step S201, the next stripe data is read, and the process is continued. In the case of the final stripe, the process is moved to step S212.
When the process moves to step S212, the coding process is completed for all the tiles of the image data to be encoded, and the final coded data is generated from the coded data of all the tiles stored in the RAM 1402. Then, the I / F 1409 is solved and output to an external device (corresponding to the processing of the code string forming unit 113 and the code output unit 114).
As described above, it is clear that this modification also makes it possible to obtain the same effects as those of the first embodiment. That is, by decomposing the image to be encoded into a reduced image and an interpolated data and applying the reduced image by switching between lossless coding and lossy coding depending on the structure of the interpolated data, high-resolution image coding can be performed easily and at high speed. It can be carried out.
The processing order of the processing flow shown in FIG. 2 is not necessarily limited to this. For example, in the process of generating 1/2 reduced tile data (step S203), the process of generating resolution interpolation data (step S204), the process of selecting the coding method (step S205), etc., the resolution interpolation information is generated first. You may go to. Alternatively, the processing may be integrated.
[Second Embodiment] In the first embodiment and its modifications, the reduced image is reversibly coded or lossy based on the distribution of the interpolation method applied when restoring the original resolution image using the resolution interpolation data. You have chosen to encode. However, the same effect can be obtained even if the distribution of the interpolation method is estimated from the code amount of the resolution interpolation data and the lossless or lossy coding is determined based on the estimation result. You can also do it.
Therefore, as the second embodiment, a method of selecting a coding method based on the code amount of the resolution interpolation data will be described.
In the second embodiment as well, the target image data is image data composed of 8 bits for each RGB color, but it may be applied to image data in other formats such as CMYK color image data. Further, it is assumed that the image is composed of W pixels in the horizontal direction and H pixels in the vertical direction, and the tile sizes Tw and Th are also "32" as in the first embodiment.
The processing contents of the interpolation data generation unit 110 and the coding method selection unit 111 of the image coding apparatus of the second embodiment are different from those of the first embodiment, and are different from the block configuration diagram of FIG. Since they are the same, they are not shown in the new illustration. Hereinafter, the processing contents of the interpolation data generation unit 110 and the coding method selection unit 111 whose operations are different from those of the first embodiment will be described with respect to the processing of the second embodiment.
In the first embodiment described above, in the process of generating the interpolation data in the interpolation data generation unit 110, the number of blocks Na to Ne indicating the frequency distribution of each pixel block type in the tile is calculated, and this is used as the coding method selection unit. It was decided to provide to 111. On the other hand, in the second embodiment, the interpolation data generation unit 110 obtains the code amount L of the interpolation data for the tile of interest instead of the frequency distribution, and provides it to the coding method selection unit 111. ..
The coding method selection unit 111 compares the code amount L of the interpolation data provided by the interpolation data generation unit 110 with the predetermined threshold value TH2.
Then, if L <TH2, the control signal "1" is output so that lossless coding is applied. If L TH2, the control signal 0 is output so that lossy coding is applied.
Since one tile has 32 x 32 pixels, the amount of data is 32 x 32 x 3 (the number of RGB components) = 3072 bytes. In the second embodiment, the threshold value TH2 is described as 1/8 (= 384 bytes) of 3072 bytes.
Here, the five types (a) to (e) of the 2 × 2 pixel pixel block described in the description of the coding method selection unit 111 of the first embodiment will be considered. As is clear from the explanations so far, in the case of the types (a) to (c) that frequently appear in characters, line arts, clip art images, etc., the code length output as interpolated data is short. On the other hand, in the case of the types (d) and (e) that are often seen in complicated CG and natural images, the code length for including the color information as it is in the interpolation data becomes large. Expressing this in another way, if there are many blocks of types (a) to (c) in the tile of interest, the code amount of the interpolation data of the tile of interest will be small, and conversely, the types (d) and (e). When many blocks of) are included, the code amount of the interpolated data becomes large. Therefore, depending on the amount of code of the interpolated data, it is possible to obtain the same effect as that of referring to the frequency in the first embodiment.
Since it is clear from the modification of the first embodiment described above that the same processing as that of the second embodiment can be realized by the computer program, the description thereof will be omitted.
[Third Embodiment] In the image coding apparatus of the first and second embodiments, the image data to be encoded is reduced to 1/2 in both the horizontal and vertical directions, and this is made a target of reversible or lossy coding. .. However, the reduction is not limited to 1/2, and a smaller image may be generated and used as a target for bitmap coding. In the third embodiment, as an example, the reduction process of 1/2 is recursively executed a predetermined number of times. Then, an example of generating resolution interpolation information from the image of the reduction source each time the reduction process is generated will be described. For the sake of simplicity, an example in which the number of recursive processes is set to 2 will be described below.
FIG. 5 shows a block configuration diagram of the image coding apparatus according to the third embodiment.
The difference from the block diagram in the first and second embodiments is that the resolution conversion unit 501, the 1/4 reduction tile buffer 502, and the interpolation data generation unit 503 are added, and the coding method selection unit 111 is coded. This is a point that has been replaced by the conversion method selection unit 504. Hereinafter, a portion different from the operation of the second embodiment will be described.
The resolution conversion unit 501 further reduces the 1/2 reduction tile data stored in the 1/2 reduction tile buffer 106 to horizontal and vertical 1/2 by the same processing as the resolution conversion unit 105, and 1/4 reduction tile. Store in buffer 502.
The interpolation data generation unit 503 generates the interpolation data necessary for restoring the 1/2 reduced tile data stored in the 1/2 reduced tile buffer from the reduced tile data stored in the 1/4 reduced tile buffer. The processing here is the same as that performed by the interpolation data generation unit 110 on the tile data stored in the tile buffer 104. Hereinafter, in order to distinguish between the interpolation data generated by the interpolation data generation unit 110 and the interpolation data generated by the interpolation data generation unit 503, the former is referred to as level 1 interpolation data and the latter is referred to as level 2 interpolation data.
Both the level 1 interpolation data and the level 2 interpolation data are stored in the interpolation data buffer 112.
Interpolation data generation units 110 and 503 provide the coding method selection unit 504 with a code amount L1 for level 1 interpolation data and a code amount L2 for level 2 interpolation data, respectively. The coding method selection unit 504 compares the sum of the input L1 and L2 with the predetermined threshold value TH3.
That is, if L1 + L2 <TH3, the control signal "1" is output so that lossless coding is applied. On the other hand, if L1 + L2 TH3, the control signal "0" is output so that lossy coding is applied. In the third embodiment, the threshold value TH3 is 1/4 of the original tile data amount of 3072 bytes, that is, 768 (bytes).
FIG. 6 illustrates the structure of the coded data of each tile of the code string generated in the third embodiment. The structure is composed of 1/4 reduced tile-coded data combined with level 2 interpolated data and level 1 interpolated data.
In the third embodiment as well, as in the first and second embodiments, the coding process can be simplified and speeded up. Since the number of pixels to which the bitmap coding process (coding process by the lossy coding unit 107 and the lossless coding unit 108) that requires advanced arithmetic processing is applied is limited to 1/16 of the original image data. The processing load is reduced more than in the first and second embodiments.
Since it is clear that the processing corresponding to the third embodiment can be realized by the computer program as in the modification of the first embodiment described above, the description thereof will be omitted.
[Fourth Embodiment] In the second and third embodiments described above, a method of estimating the configuration of the interpolation method from the code amount of the interpolated data and selecting lossless or lossy coding is shown. The configuration of the interpolation method can be estimated from a part of the code amount of the interpolation data. Such an example will be described as the fourth embodiment.
The fourth embodiment is basically the same as the second embodiment described above, but the operations of the interpolation data generation unit 110 and the coding method selection unit 111 are slightly different. Hereinafter, the parts having different operations will be described.
In the second embodiment described above, the interpolation data generation unit 110 provides the interpolation data amount L of the tile of interest to the coding method selection unit 111. In the fourth embodiment, of the five types of blocks of the types (a) to (e) described in the first embodiment, the code amount Ld of the interpolated data related to the type (d) and the type (e). The sum of the code amounts Le related to the above is provided to the coding method selection unit 111. Here, the interpolated data related to the type (d) refers to the code output for the non-flat block processed in steps S1510 and S1511 in the processing flow of FIG. More specifically, it is a code output in steps S1504, S1506, S1507, S1510, and S1511 for the non-flat blocks leading to steps S1510 and 1511. The interpolated data related to the type (e) refers to the code output for the non-flat block processed in steps S1512 and S1513 in the same figure. This is the code output in steps S1504, S1512, and S1513 for the non-flat blocks leading to steps S1512 and 1513.
The coding method selection unit 111 compares the code amount Ld + Le of the interpolation data provided by the interpolation data generation unit 110 with the predetermined threshold value TH4. If Ld + Le <TH4, the control signal 1 is output so that lossless coding is applied, and if Ld + Le TH4, the control signal 0 is output so that lossy coding is applied.
In complex CG and natural images, the sum of Ld and Le tends to be large, and lossy coding distortion is less noticeable in such images, and even when returning to the original resolution, the reduced image Since it is not easily affected by deterioration, the effects of high image quality, high compression, simple and high speed can be obtained also in this embodiment.
[Other Embodiments] In each of the above embodiments, JPEG-LS is used as lossless coding and JPEG is used as lossy coding. However, as described above, other coding methods such as JPEG2000 and PNG may be applied. ..
Further, in each of the above embodiments, when a 1/2 reduced image is generated from the original image, the pixel B located in the upper left corner of the pixel block of 2 × 2 pixels<sub>n</sub>(0,0) is sampled as a pixel of the reduced image, and other non-sampling target 3 pixels B<sub>n</sub>(1,0), B<sub>n</sub>(0,1), B<sub>n</sub>The interpolation data of (1,1) was generated. However, the pixels used in the reduced image do not necessarily have to be the pixels in the upper left corner of the 2x2 pixel block, B.<sub>n</sub>(1,0), B<sub>n</sub>(0,1), B<sub>n</sub>It does not matter which of (1 and 1).
In short, B<sub>n</sub>(1,0), B<sub>n</sub>(0,1), B<sub>n</sub>Even if any one of (1 and 1) is the pixel X to be sampled, the remaining three pixels can be defined as Xa, Xb, and Xc. In this case, the pixel to be sampled adjacent to the pixel Xa in the adjacent pixel block referred to for restoring the pixel Xa is X1, and the pixel Xb in the adjacent block referenced to restore the pixel Xb is adjacent to the pixel Xb. When the pixel to be sampled is defined as X2, and the pixel to be sampled adjacent to the pixel Xc in the adjacent block referenced to restore the pixel Xc is defined as X3. Xa = X1 and Xb = X2 and Xc = X3 If the condition is satisfied, the pixel block of interest may be determined to be a surrounding 3-pixel matching block.
Further, the size of the pixel block is 3 × 3, and one pixel may be extracted from the pixel block to generate a reduced image having 1/3 of the original number of pixels both horizontally and vertically.
Further, not only one pixel in the pixel block may be extracted, but also a reduced image may be generated by obtaining the average color of the pixel block. In such a case, the method of constructing the interpolated data may be devised so that the image data of the original resolution can be restored according to the method of generating the reduced image.
Further, as a method of constructing the interpolated data, a configuration example in which the image data of the original resolution can be completely restored has been shown, but the present invention is not necessarily limited to this. For example, when the number of colors included in the pixel block is 3 or more, the color may be reduced to return to the original resolution, but the pixel value may be interpolated data that allows a slight change. ..
Further, in the above-described embodiment, the tiles are divided into 32 × 32 pixel tiles for processing, but the tile size is not limited to this, and may be an integral multiple of the pixel block. Therefore. It can be a different block size, such as 16x16, 64x64, 128x128, or it doesn't have to be square.
Further, in the embodiment, the color space of the image is RGB, but since it is clear that it can be applied to various types of image data such as CMYK, Lab, and YCrCb, the number of color components and the color space The present invention is not limited to the type.
In addition, computer programs are usually stored on a computer-readable storage medium such as a CD-ROM, and can be executed by setting it in a computer reader (CD-ROM drive) and copying or installing it on the system. Become. Therefore, it is clear that such computer-readable storage media also fall into the category of the present invention.
<figref num="1">It is a block block diagram of the image coding apparatus which concerns on 1st Embodiment.</figref><figref num="2">It is a flowchart which shows the flow of the process which concerns on the modification of 1st Embodiment.</figref><figref num="3">It is a figure which shows the relationship between the tile data and a pixel block.</figref><figref num="4">It is a figure which shows the position of the sampling target pixel X and the non-sampling target pixel Xa, Xb, Xc in a pixel block.</figref><figref num="5">It is a block block diagram of the image coding apparatus which concerns on 3rd Embodiment.</figref><figref num="6">It is a figure which shows the structure of the coded data of the tile in 3rd Embodiment.</figref><figref num="7">It is a figure which shows the structural example of the JPEG coded data.</figref><figref num="8">It is a figure which shows the relationship between the image data to be encoded, the stripe, and the tile.</figref><figref num="9">It is a figure which shows the structure of the coded data in the 1st and 2nd embodiments.</figref><figref num="10">It is a figure which shows an example of the tile data.</figref><figref num="11">It is a figure which shows the example of the flat, non-flat block determination result.</figref><figref num="12">It is a block diagram which shows the internal structure of the interpolation data generation part.</figref><figref num="13">It is a flowchart which shows the processing procedure of a flat determination part.</figref><figref num="14">It is a block block diagram of the information processing apparatus which concerns on the modification of 1st Embodiment.</figref><figref num="15">It is a flowchart which shows the processing procedure of a non-flat block analysis part.</figref><figref num="16">It is a figure which shows the arrangement pattern of 2 colors in a 2 × 2 pixel block.</figref>
17 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2012054786A | Cited by | Japan | Search report |
| JP2012054842A | Cited by | Japan | Search report |
| JP2011223462A | Cited by | Japan | Examiner |
| JP2008042687A | Cites | Japan | Examiner |
| JPH0981763A | Cites | Japan | Examiner |
| JPH11146394A | Cites | Japan | Examiner |
4 members in 2 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008315030 | Japan | A | |
| JP20080315030 | – | – | – |
Members4
| Document | Office | Kind | |
|---|---|---|---|
| US2010142840A1 | United States of America | A1 | |
| JP2010141533AThis record | Japan | A | |
| JP5116650B2 | Japan | B2 | |
| US8396308B2 | United States of America | B2 |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification of patent or utility model registrationJAPANESE INTERMEDIATE CODE: R151R151 | R151 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 |
Numbers
- Publication
- 2010141533
- Publication, DOCDB
- 2010141533
- Publication, EPODOC
- JP2010141533
- Application
- 315030
- Application, DOCDB
- 2008315030
- Application, EPODOC
- JP20080315030
Titles2
- Japanese
- 画像符号化装置及びその制御方法
- English
- Image coding device and its control method
Classification
- CPC, 9
- H04N19/59
- H04N19/12
- H04N19/14
- H04N19/17
- H04N19/176
- H04N19/46
- H04N19/587
- H04N19/60
- H04N19/70
- IPC, 12
- H04N1 41
- H04N1 413
- H04N7 26
- H04N19 00
- H04N19 12
- H04N19 146
- H04N19 196
- H04N19 59
- H04N19 625
- H04N19 63
- H04N19 90
- H04N19 93