Moving image coding apparatus, moving image decoding apparatus, control method therefor, computer program, and computer-readable storage medium
Abstract
[Subject] The video coding equipment which controlled that the error by the omission of bit plane was gradually accumulated by prediction frame picture like P picture or B picture, and prevented degradation of the picture by choice of the coded data for every bit plane is offered. [Solution means] The block division part 31 which divides the inputted frame into two or more blocks, As it is, in the case of the frame inner code-ized mode, output at the DWT section 33, and in the case of the coding mode between frames, So that it may become below a target code amount in the difference calculation part 32 which calculates difference with the prediction data from the motion compensation part 42, the DWT operation part 33 which codes bit plane, the quantization part 34, and the entropy coding part 35. Only when it has the bit omission part 36 which omits the coded data of bit plane which goes to a higher rank from a least significant bit position and the frame inner code-ized mode is performed, the dequantization part 39 and the reverse DWT section 40 are performed, and the frame memory 41 is updated. [Selection figure] Fig. 1
Term
Term ended
Projected expiry passed 24 January 2025, 1.7 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
20 claims: 10 independent, 10 dependent
- 1A moving image coding device that sequentially inputs and encodes the image data of the frames that make up the moving image, the first coding mode that utilizes the correlation between the frames, and the second code that encodes the frames alone. A mode selection means for adaptively selecting the conversion mode for each frame, a division means for dividing the image data of the input frame into a plurality of blocks, and a local image data encoded according to the output of the mode selection means. In the decoding means for decoding and the first coding mode, prediction data is extracted from the previously locally decoded frame based on the block image obtained by dividing by the dividing means, and the divided block image and the said A calculation means that outputs a block that is different from the prediction data, and outputs a block divided by the division means in the second coding mode, and a conversion means that converts the block obtained by the calculation means into spatial frequency component data. And a coded data generation means that generates intermediate coded data in bit plane units composed of bit information of each bit position representing each frequency component value obtained by conversion. The adjustment means for adjusting the amount of coded data and the adjustment means for adjusting the amount of coded data by truncating from the least significant bit position to the coded data of the desired bit plane toward the upper bit position in the generated coded data. A moving image coding device including an output means for outputting the coded data. 動画像を構成するフレームの画像データを順次入力し、符号化する動画像符号化装置であって、 フレーム間の相関を利用した第1の符号化モード、フレーム単独で符号化する第2の符号化モードをフレーム単位に適応的に選択するモード選択手段と、 入力したフレームの画像データを複数のブロックに分割する分割手段と、 前記モード選択手段の出力に応じて符号化された画像データをローカルデコードする復号手段と、 前記第1の符号化モードでは、前記分割手段で分割して得られたブロック画像に基づき従前のローカルデコードされたフレームから予測データを抽出し、前記分割したブロック画像と前記予測データとを差分したブロックを出力し、第2の符号化モードでは前記分割手段で分割したブロックを出力する演算手段と、 前記演算手段で得られたブロックを空間周波数成分データに変換する変換手段と、 変換して得られた各周波数成分値を表わす各ビット位置のビット情報で構成されるビットプレーン単位に中間的な符号化データを生成する符号化データ生成手段と、 生成された符号化データ中の最下位ビット位置から上位ビット位置に向かう所望とするビットプレーンの符号化データまでを切り捨てることで、符号化データ量を調整する調整手段と、 前記調整手段により調整された符号化データを出力する出力手段と を備えることを特徴とする動画像符号化装置。
- 8This is a control method of a moving image coding device that sequentially inputs and encodes image data of frames constituting a moving image. The first coding mode utilizing the correlation between frames, the first coding by a single frame. A mode selection step of adaptively selecting the coding mode of 2 on a frame-by-frame basis, a division step of dividing the input frame image data into a plurality of blocks, and an image encoded according to the output of the mode selection means. In the decoding step of locally decoding the data and the first coding mode, the predicted data is extracted from the conventional locally decoded frame based on the block image obtained by dividing in the dividing step, and the divided block. In the second coding mode, a calculation step of outputting a block obtained by differentiating the image and the prediction data and outputting the block divided in the division step, and a block obtained in the calculation step are used as spatial frequency component data. A conversion step of conversion, a coded data generation step of generating intermediate coded data in a bit plane unit composed of bit information of each bit position representing each frequency component value obtained by the conversion, and a coded data generation step. An adjustment step for adjusting the amount of coded data by truncating from the least significant bit position to the coded data of the desired bit plane toward the upper bit position in the generated coded data, and an adjustment step for adjusting the amount of coded data. A control method for a moving image encoding device, which comprises an output process for outputting the encoded data. 動画像を構成するフレームの画像データを順次入力し、符号化する動画像符号化装置の制御方法であって、 フレーム間の相関を利用した第1の符号化モード、フレーム単独で符号化する第2の符号化モードをフレーム単位に適応的に選択するモード選択工程と、 入力したフレームの画像データを複数のブロックに分割する分割工程と、 前記モード選択手段の出力に応じて符号化された画像データをローカルデコードする復号工程と、 前記第1の符号化モードでは、前記分割工程で分割して得られたブロック画像に基づき従前のローカルデコードされたフレームから予測データを抽出し、前記分割したブロック画像と前記予測データとを差分したブロックを出力し、前記第2の符号化モードでは前記分割工程で分割したブロックを出力する演算工程と、 前記演算工程で得られたブロックを空間周波数成分データに変換する変換工程と、 前記変換して得られた各周波数成分値を表わす各ビット位置のビット情報で構成されるビットプレーン単位に中間的な符号化データを生成する符号化データ生成工程と、 前記生成された符号化データ中の最下位ビット位置から上位ビット位置に向かう所望とするビットプレーンの符号化データまでを切り捨てることで、符号化データ量を調整する調整工程と、 前記調整工程で調整された符号化データを出力する出力工程と を備えることを特徴とする動画像符号化装置の制御方法。
- 9A computer program that functions as a moving image coding device that sequentially inputs and encodes frames that make up a moving image by being read and executed by a computer, and is a first coding mode that utilizes the correlation between frames. For the mode selection means that adaptively selects the second coding mode that encodes the frame alone for each frame, the division means that divides the image data of the input frame into a plurality of blocks, and the output of the mode selection means. Decoding means that locally decodes the image data encoded accordingly, and in the first coding mode, prediction data is obtained from the conventional locally decoded frame based on the block image obtained by dividing by the dividing means. A calculation means that outputs a block obtained by extracting and differentiating the divided block image and the prediction data, and outputs a block divided by the dividing means in the second coding mode, and a calculation means obtained by the calculation means. A code that generates intermediate coded data in bit plane units consisting of a conversion means that converts blocks into spatial frequency component data and bit information at each bit position that represents each frequency component value obtained by the conversion. Computerized data generation means and An adjustment means for adjusting the amount of coded data by truncating from the least significant bit position to the coded data of the desired bit plane toward the upper bit position in the generated coded data, and adjustment by the adjustment means. A computer program characterized in that it functions as an output means for outputting the encoded data. コンピュータが読み込み実行することで、動画像を構成するフレームを順次入力し、符号化する動画像符号化装置として機能するコンピュータプログラムであって、 フレーム間の相関を利用した第1の符号化モード、フレーム単独で符号化する第2の符号化モードをフレーム単位に適応的に選択するモード選択手段と、 入力したフレームの画像データを複数のブロックに分割する分割手段と、 前記モード選択手段の出力に応じて符号化された画像データをローカルデコードする復号手段と、 前記第1の符号化モードでは、前記分割手段で分割して得られたブロック画像に基づき従前のローカルデコードされたフレームから予測データを抽出し、前記分割したブロック画像と前記予測データとを差分したブロックを出力し、前記第2の符号化モードでは前記分割手段で分割したブロックを出力する演算手段と、 前記演算手段で得られたブロックを空間周波数成分データに変換する変換手段と、 前記変換して得られた各周波数成分値を表わす各ビット位置のビット情報で構成されるビットプレーン単位に中間的な符号化データを生成する符号化データ生成手段と、 前記生成された符号化データ中の最下位ビット位置から上位ビット位置に向かう所望とするビットプレーンの符号化データまでを切り捨てることで、符号化データ量を調整する調整手段と、 前記調整手段により調整された符号化データを出力する出力手段 として機能することを特徴とするコンピュータプログラム。
- 11A moving image coding device that sequentially inputs and encodes the image data of the frames that make up the moving image, the first coding mode that utilizes the correlation between the frames, and the second code that encodes the frames alone. The mode selection means for adaptively selecting the conversion mode for each frame, the division means for dividing the image data of the input frame into a plurality of blocks, the storage means for storing the image data for at least one frame, and the mode selection. When the first coding mode is selected by the means, prediction data is extracted from the image data stored in the storage means based on the block image obtained by dividing by the dividing means, and extracted. A calculation means that outputs the difference between the prediction data and the block image, and outputs the block image divided by the division means when the second coding mode is selected by the mode selection means. It is encoded in a bit plane unit composed of a conversion means for converting a block output from the calculation means into spatial frequency component data and bit information at each bit position representing each frequency component value obtained by the conversion means. Coding means and When the second coding mode is selected by the mode selection means, the coding data generated by the coding means is locally decoded, and the storage means is updated with the image data obtained by the decoding. A moving image encoding device, which comprises. 動画像を構成するフレームの画像データを順次入力し、符号化する動画像符号化装置であって、 フレーム間の相関を利用した第1の符号化モード、フレーム単独で符号化する第2の符号化モードをフレーム単位に適応的に選択するモード選択手段と、 入力したフレームの画像データを複数のブロックに分割する分割手段と、 少なくとも1フレーム分の画像データを記憶する記憶手段と、 前記モード選択手段で前記第1の符号化モードが選択された場合には、前記分割手段で分割して得られたブロック画像に基づき、前記記憶手段に記憶された画像データから予測データを抽出し、抽出した予測データと前記ブロック画像との差分を出力し、前記記モード選択手段で前記第2の符号化モードが選択された場合には、前記分割手段で分割されたブロック画像を出力する演算手段と、 該演算手段より出力されたブロックを空間周波数成分データに変換する変換手段と、 該変換手段で得られた各周波数成分値を表わす各ビット位置のビット情報で構成されるビットプレーン単位に符号化する符号化手段と、 前記モード選択手段によって前記第2の符号化モードを選択した場合、前記符号化手段で生成された符号化データをローカルデコードし、デコードして得られた画像データで前記記憶手段を更新する更新手段と を備えることを特徴とする動画像符号化装置。
- 12In the coding means, the bit position of the uppermost bit plane is Nmax, the coded data of the n (0 n Nmax) bit plane is C (i), and the amount of the coded data is L (C (i). )), Coded data up to the maximum k that satisfies ΣL (C (Nmax-k)) T, where T is the threshold value indicating the allowable code amount for one frame C (Nmax), C (Nmax-1), ..., C (Nmax-k) is output as valid coded data, and the coded data C (0), ..., C (Nmax-k-1) is discarded. 11. The moving image encoding device according to 11. 前記符号化手段は、 最上位のビットプレーンのビット位置をNmax、n(0≦n≦Nmax)目のビットプレーンの符号化データをC(i)、その符号化データ量をL(C(i))、1フレームの許容符号量を示す閾値をTとしたとき、 ΣL(C(Nmax-k))≦Tを満たす最大kまでの符号化データC(Nmax),C(Nmax-1)、...、C(Nmax-k)を有効な符号化データとして出力し、符号化データC(0)、...、C(Nmax-k-1)まで破棄することを特徴とする請求項11に記載の動画像符号化装置。
- 14It is a control method of a moving image coding device that is provided with a storage means for storing image data for at least one frame, sequentially inputs and encodes image data of frames constituting a moving image, and uses correlation between frames. A mode selection step of adaptively selecting the first coding mode and the second coding mode of coding the frame alone for each frame, and a dividing step of dividing the image data of the input frame into a plurality of blocks. When the first coding mode is selected in the mode selection step, prediction data is extracted from the image data stored in the storage means based on the block image obtained by dividing in the division step. Then, the difference between the extracted prediction data and the block image is output, and when the second coding mode is selected in the mode selection step, the block image divided in the division step is output. A bit plane composed of a calculation step to be performed, a conversion step of converting a block output from the calculation step into spatial frequency component data, and bit information of each bit position representing each frequency component value obtained in the conversion step. A coding process that encodes into units, When the second coding mode is selected by the mode selection step, an update step of locally decoding the coded data generated in the coding step and updating the storage means with the image data obtained by decoding. A method for controlling a moving image encoding device, which comprises. 少なくとも1フレーム分の画像データを記憶する記憶手段を備え、動画像を構成するフレームの画像データを順次入力し、符号化する動画像符号化装置の制御方法であって、 フレーム間の相関を利用した第1の符号化モード、フレーム単独で符号化する第2の符号化モードをフレーム単位に適応的に選択するモード選択工程と、 入力したフレームの画像データを複数のブロックに分割する分割工程と、 前記モード選択工程で前記第1の符号化モードが選択された場合には、前記分割工程で分割して得られたブロック画像に基づき、前記記憶手段に記憶された画像データから予測データを抽出し、抽出した予測データと前記ブロック画像との差分を出力すると共に、前記記モード選択工程で前記第2の符号化モードが選択された場合には、前記分割工程で分割されたブロック画像を出力する演算工程と、 該演算工程より出力されたブロックを空間周波数成分データに変換する変換工程と、 該変換工程で得られた各周波数成分値を表わす各ビット位置のビット情報で構成されるビットプレーン単位に符号化する符号化工程と、 前記モード選択工程によって前記第2の符号化モードを選択した場合、前記符号化工程で生成された符号化データをローカルデコードし、デコードして得られた画像データで前記記憶手段を更新する更新工程と を備えることを特徴とする動画像符号化装置の制御方法。
- 15A computer program for a moving image coding device that is provided with a storage means for storing image data for at least one frame, sequentially inputs and encodes image data of frames constituting a moving image, and correlates between frames. A mode selection means for adaptively selecting the first coding mode used and the second coding mode for encoding the frame alone for each frame, and a dividing means for dividing the image data of the input frame into a plurality of blocks. When the first coding mode is selected by the mode selection means, prediction data is obtained from the image data stored in the storage means based on the block image obtained by dividing by the division means. It extracts and outputs the difference between the extracted prediction data and the block image, and when the second coding mode is selected by the mode selection means, the block image divided by the division means is output. A bit composed of a calculation means for output, a conversion means for converting a block output from the calculation means into spatial frequency component data, and bit information at each bit position representing each frequency component value obtained by the conversion means. Coding means that encodes in plane units and When the second coding mode is selected by the mode selection means, the coding data generated by the coding means is locally decoded, and the storage means is updated with the image data obtained by the decoding. A computer program characterized by functioning as. 少なくとも1フレーム分の画像データを記憶する記憶手段を備え、動画像を構成するフレームの画像データを順次入力し、符号化する動画像符号化装置用のコンピュータプログラムであって、 フレーム間の相関を利用した第1の符号化モード、フレーム単独で符号化する第2の符号化モードをフレーム単位に適応的に選択するモード選択手段と、 入力したフレームの画像データを複数のブロックに分割する分割手段と、 前記モード選択手段で前記第1の符号化モードが選択された場合には、前記分割手段で分割して得られたブロック画像に基づき、前記記憶手段に記憶された画像データから予測データを抽出し、抽出した予測データと前記ブロック画像との差分を出力すると共に、前記記モード選択手段で前記第2の符号化モードが選択された場合には、前記分割手段で分割されたブロック画像を出力する演算手段と、 該演算手段より出力されたブロックを空間周波数成分データに変換する変換手段と、 該変換手段で得られた各周波数成分値を表わす各ビット位置のビット情報で構成されるビットプレーン単位に符号化する符号化手段と、 前記モード選択手段によって前記第2の符号化モードを選択した場合、前記符号化手段で生成された符号化データをローカルデコードし、デコードして得られた画像データで前記記憶手段を更新する更新手段 として機能することを特徴とするコンピュータプログラム。
- 17It is a moving image decoding device that decodes the encoded moving image data, and the frame of interest uses the correlation between the frames based on the storage means that stores at least one frame of image data and the input encoded data. Determining means for determining whether the data is encoded by the first coding mode or the coded data by the second coding mode, which is encoded by the frame alone, and decoding to decode the encoded data of the frame of interest. When the means and the determination means determine that the frame of interest is the coded data in the first coding mode, the decoding result by the decoding means is assumed to be the difference image data and is stored in the storage means. A frame image is generated by adding the combined image data and the difference image data, and when the determination means determines that the frame of interest is the encoded data in the second coding mode, it is decoded. When the addition means that outputs the result as a frame image and the determination means determine that the frame of interest is the coded data in the second coding mode, the storage means is the frame image output from the addition means. A moving image decoding device, which comprises an updating means for updating a data. 符号化された動画像データを復号する動画像復号装置であって、 少なくとも1フレーム分の画像データを記憶する記憶手段と、 入力した符号化データに基づき、注目フレームがフレーム間の相関を利用した第1の符号化モードによる符号化データであるか、フレーム単独で符号化する第2の符号化モードによる符号化データであるのかを判定する判定手段と、 注目フレームの符号化データを復号する復号手段と、 前記判定手段で注目フレームが前記第1の符号化モードによる符号化データであると判定された場合、前記復号手段による復号結果は差分画像データであるものとし、前記記憶手段に記憶された画像データと前記差分画像データとを加算することでフレーム画像を生成すると共に、前記判定手段で注目フレームが前記第2の符号化モードによる符号化データであると判定された場合には、復号結果をフレーム画像として出力する加算手段と、 前記判定手段で注目フレームが前記第2の符号化モードによる符号化データであると判定された場合、前記加算手段より出力されたフレーム画像で前記記憶手段を更新する更新手段と を備えることを特徴とする動画像復号装置。
- 18It is a control method of a moving image decoding device that has a storage means for storing image data for at least one frame and decodes the encoded moving image data, and the frame of interest is between frames based on the input encoded data. The determination step of determining whether the data is encoded by the first coding mode using the correlation of the above, or the encoded data by the second coding mode that encodes the frame alone, and the coding of the frame of interest. When the decoding step of decoding the data and the determination step determine that the frame of interest is the encoded data by the first coding mode, the decoding result by the decoding step is assumed to be the difference image data. A frame image is generated by adding the image data stored in the storage means and the difference image data, and in the determination step, it is determined that the frame of interest is the encoded data in the second coding mode. In this case, the addition step of outputting the decoding result as a frame image, and the frame output from the addition step when the frame of interest is determined to be the coded data by the second coding mode in the determination step. With the update process of updating the storage means with an image A method for controlling a moving image decoding device, which comprises. 少なくとも1フレーム分の画像データを記憶する記憶手段を有し、符号化された動画像データを復号する動画像復号装置の制御方法であって、 入力した符号化データに基づき、注目フレームがフレーム間の相関を利用した第1の符号化モードによる符号化データであるか、フレーム単独で符号化する第2の符号化モードによる符号化データであるのかを判定する判定工程と、 注目フレームの符号化データを復号する復号工程と、 前記判定工程で注目フレームが前記第1の符号化モードによる符号化データであると判定された場合、前記復号工程による復号結果は差分画像データであるものとし、前記記憶手段に記憶された画像データと前記差分画像データとを加算することでフレーム画像を生成すると共に、前記判定工程で注目フレームが前記第2の符号化モードによる符号化データであると判定された場合には、復号結果をフレーム画像として出力する加算工程と、 前記判定工程で注目フレームが前記第2の符号化モードによる符号化データであると判定された場合、前記加算工程より出力されたフレーム画像で前記記憶手段を更新する更新工程と を備えることを特徴とする動画像復号装置の制御方法。
- 19A computer program for controlling a moving image decoding device that has a storage means for storing at least one frame of image data and decodes the encoded moving image data, and is of interest based on the input encoded data. Attention is paid to the determination means for determining whether the frame is the coded data in the first coding mode using the correlation between the frames or the coded data in the second coding mode in which the frame is encoded by itself. When the decoding means for decoding the coded data of the frame and the determination means determine that the frame of interest is the coded data in the first coding mode, the decoding result by the decoding means is the difference image data. A frame image is generated by adding the image data stored in the storage means and the difference image data, and the frame of interest in the determination means is the coded data in the second coding mode. If it is determined by the addition means that outputs the decoding result as a frame image, and if the determination means determines that the frame of interest is the coded data by the second coding mode, the addition means Update means for updating the storage means with the output frame image A computer program characterized by functioning as. 少なくとも1フレーム分の画像データを記憶する記憶手段を有し、符号化された動画像データを復号する動画像復号装置を制御するためのコンピュータプログラムであって、 入力した符号化データに基づき、注目フレームがフレーム間の相関を利用した第1の符号化モードによる符号化データであるか、フレーム単独で符号化する第2の符号化モードによる符号化データであるのかを判定する判定手段と、 注目フレームの符号化データを復号する復号手段と、 前記判定手段で注目フレームが前記第1の符号化モードによる符号化データであると判定された場合、前記復号手段による復号結果は差分画像データであるものとし、前記記憶手段に記憶された画像データと前記差分画像データとを加算することでフレーム画像を生成すると共に、前記判定手段で注目フレームが前記第2の符号化モードによる符号化データであると判定された場合には、復号結果をフレーム画像として出力する加算手段と、 前記判定手段で注目フレームが前記第2の符号化モードによる符号化データであると判定された場合、前記加算手段より出力されたフレーム画像で前記記憶手段を更新する更新手段 として機能することを特徴とするコンピュータプログラム。
Independent claims10
109 paragraphs, as filed
The present invention relates to a technique for encoding moving image data.
In recent years, the contents flowing on networks, especially on the Internet, have been increased in capacity and diversified from text information to still image information and further to moving image information. Along with this, the development of coding technology that compresses the amount of information has also progressed, and the developed coding technology has become widespread due to international standardization.
On the other hand, the capacity and diversification of the network itself is progressing, and one content has to pass through various environments before it reaches the receiving side. In addition, the processing performance of the transmitting / receiving side equipment is also diversifying. General-purpose information processing devices (hereinafter referred to as PCs) such as personal computers, which are mainly used as transmission / reception devices, have significantly improved performance such as CPU performance and graphics performance, while processing performance such as PDAs, mobile phones, TVs, and hard disk recorders has been improved. Various different devices have come to have a network connection function. For this reason, attention is being paid to a function called scalability that can respond to changing communication line capacities and processing performance of receiving devices with a single piece of data.
The JPEG2000 coding method is widely known as a still image coding method having this scalability function. This method has been internationally standardized and is described in detail in ISO / IEC15444-1 (Information technology --JPEG 2000 image coding system --Part 1: Core coding system). Its feature is that the input image data is subjected to discrete wavelet transform (DWT) and separated into multiple frequency bands. These coefficients are quantized and their values are arithmetically coded for each bit plane. By encoding and decoding the required number of bit planes, fine-grained hierarchy control is possible.
In addition, the JPEG2000 coding method also realizes a technology such as ROI (Region Of Interest) that relatively improves the image quality of the region of interest in the image, which is not found in the conventional coding technology.
Figure 8 shows the coding procedure of the JPEG2000 coding method. The tile dividing unit 9001 divides the input image into a plurality of areas (tiles). This feature is optional. The DWT section 9002 performs discrete wavelet transform and separates it into frequency bands. Each coefficient is quantized by the quantization unit 9003. However, this feature is optional. The ROI section 9007 is an option, the area of interest is set, and the quantization section 9003 shifts up. The entropy coding unit 9004 performs entropy coding by the EBCOT (Embeded Block Coding with Optimized Truncation) method, and the bit plane truncation unit 9005 truncates the lower bit plane to control the rate of the encoded data. Header information is added by the code forming unit 9006, various scalability functions are selected, and coded data is output.
Figure 9 shows the decoding procedure of the JPEG2000 coding method. The code analysis unit 9020 analyzes the header and obtains information for constructing the hierarchy. The bit plane truncation unit 9021 truncates the lower bit plane of the input coded data according to the capacity of the internal buffer and the decoding processing capacity. The entropy decoding unit 9022 decodes the coded data of the EBCOT coding method and obtains the quantized wavelet transform coefficient. The inverse quantization unit 9023 applies inverse quantization to this, and the inverse DWT unit 9024 performs inverse discrete wavelet transform to reproduce the image data. The tile composition unit 9025 synthesizes a plurality of tiles and reproduces the image data.
The Motion JPEG 2000 method (ISO / IEC15444-3 (Information technology --JPEG 2000 image coding system Part 3: Motion JPEG 2000)) that encodes moving images by making this JPEG 2000 coding method correspond to each frame of the moving image is also available. It is recommended. In this method, coding processing is performed independently for each frame, and coding is not performed using time correlation, so that redundancy remains between frames. Therefore, there is a problem that it is difficult to effectively reduce the amount of coding as compared with the moving image coding method using time correlation.
On the other hand, in the MPEG coding method, motion compensation is performed to improve the coding efficiency (Non-Patent Document 1). Figure 10 shows the coding procedure. The block division unit 9031 divides the pixel block into 8 × 8, the difference unit 9032 draws the prediction data by motion compensation, the DCT unit 9033 performs the discrete cosine transform, and the quantization unit 9034 performs the quantization. The result is encoded by the entropy coding unit 9035, header information is added by the code forming unit 9036, and the coded data is output.
At the same time, the inverse quantization unit 9037 performs inverse quantization, the inverse DCT unit 9038 performs the inverse transform of the discrete cosine transform, and the addition unit 9039 adds the prediction data and stores it in the frame memory 9040. The motion compensation unit 9041 obtains a motion vector by referring to the input image and the reference frame stored in the frame memory 9040, and generates prediction data.<patcit num="1"><text>"Latest MPEG Textbook", page 76, etc. ASCII Publishing Bureau 1994</text></patcit>
<p> In the above-mentioned MPEG coding method, in order to apply it to a method that performs bit plane coding so as to realize scalability such as JPEG2000, an error in motion compensation accumulates due to the termination of bit plane coding, and the error of the image is accumulated. Problems such as deterioration occur. That is, the DCT section 9033 and the inverse DCT section 9038 in FIG. 10 are replaced with the discrete wavelet conversion and the inverse discrete wavelet conversion, the entropy coding section 9035 performs bit plane coding, and the code formation 9036 adds the bit plane truncation section 9005. When bit plane truncation is performed, the number of bit planes reproduced in each frame differs depending on the bit plane truncation performed by the bit plane truncation section 9021 in FIG. Also, if the truncated lower bit plane is filled with 0, the value will be different from the original error, so if a further P picture or B picture is generated from the P picture in MPEG, the error will be accumulated. As a result, the image quality of the moving image deteriorates.</p><p> The present invention has been made in view of the above problems, and suppresses the error generated by the termination of bit plane coding from being gradually accumulated in a predicted frame image such as a P picture or a B picture. It is intended to provide a technique that makes it possible to prevent image deterioration.</p>
<p> In order to solve this problem, for example, the image coding apparatus of the present invention has the following configuration. That is, it is a moving image coding device that sequentially inputs and encodes the image data of the frames constituting the moving image, the first coding mode utilizing the correlation between the frames, and the second encoding by the frame alone. A mode selection means for adaptively selecting the coding mode of the above in frame units, a division means for dividing the input frame image data into a plurality of blocks, and image data encoded according to the output of the mode selection means. In the first coding mode, the prediction data is extracted from the conventional locally decoded frame based on the block image obtained by dividing by the dividing means, and the divided block image is extracted. In the second coding mode, a calculation means for outputting the block divided by the division means and a block obtained by the calculation means are converted into spatial frequency component data. Conversion means and A coded data generating means for generating intermediate coded data in a bit plane unit composed of bit information of each bit position representing each frequency component value obtained by conversion, and a coded data generating means in the generated coded data. An adjustment means for adjusting the amount of coded data by truncating up to the coded data of the desired bit plane from the lowest bit position to the upper bit position, and an output for outputting the coded data adjusted by the adjusting means. Provide means.</p>
<p> According to the present invention, even when the final coded data is generated by selecting the coded data for each bit plane, the bit plane gradually becomes a predicted frame image such as a P picture or a B picture. It is possible to suppress the accumulation of errors due to truncation and prevent image deterioration.</p>
Hereinafter, embodiments according to the present invention will be described in detail with reference to the accompanying drawings.
<First Embodiment> FIG. 1 is a block configuration diagram of a moving image coding device according to the first embodiment. In the first embodiment, the MotionJPEG2000 coding method will be described as an example of the image coding method used by the moving image coding device, but the present invention is not limited to this.
In FIG. 1, 31 is a block division unit that divides the input image data into block units, and 32 is a difference calculation unit that obtains a difference between the image data and the prediction data obtained by motion compensation described later. 33 is a DWT part that performs a discrete wavelet transform on the divided blocks. 34 is a quantization unit that quantizes the conversion coefficient obtained by discrete wavelet conversion, 35 is an entropy coding unit that performs EBCOT coding of the JPEG2000 coding method for each bit plane, and 36 is an entropy coding unit that performs EBCOT coding of the JPEG2000 coding method for each bit plane, and 36 is from the coded data. It is a bitplane truncation section that selects valid upper bitplane encoded data and truncates the lower bitplane encoded data, 37 generates the required headers and cuts the encoded data from the output of the bitplane truncation section 36. It is a code forming part to be formed.
Reference numeral 43 denotes a mode determination unit that determines the coding mode on a frame-by-frame basis, and determines whether to use the intra-frame coding (intra-frame coding) mode or the inter-frame coding (inter-frame coding) mode. To do. 39 is an inverse quantization unit that performs the inverse quantization of the quantization unit 34, and 40 is an inverse DWT unit that performs the inverse conversion of the DWT unit 33. The inverse quantization unit 39 and the inverse DWT unit 40 are executed only when the mode determination unit 43 determines to encode in the in-frame coding mode. Therefore, in response to the determination result by the mode determination unit 43, a switch 38 for giving execution permission to the inverse quantization unit 39 and the inverse DWT unit 40 is provided.
41 is a frame memory for storing a decoded image (an image locally decoded by the inverse quantization unit 39 and the inverse DWT unit 04) for reference of motion compensation. As described above, since the inverse quantization unit 39 and the inverse DWT unit 40 are executed only in the in-frame coding mode, only the decoded image result of the in-frame coded image is stored in the frame memory 41. It will be stored. Reference numeral 42 denotes a motion compensation that performs motion prediction from the frame memory 41 and the input image and calculates the motion vector and the prediction data.
The operation of the moving image coding device configured as described above will be described below. In the present embodiment, a case where a GOP (Group Of Pictures) is composed of only an I picture that performs in-frame coding and a P picture that performs inter-frame coding by forward prediction will be described. One GOP shall consist of 15 frames. Normally, during playback, playback is performed at a frame rate of 30 frames / second, so 1 GOP is about 0.5 seconds of moving image data. Further, in the embodiment, one I picture (in-frame coded data) is set in 1GOP, the remaining 14 frames are set as P pictures (inter-frame coded data), and the timing of generating the I picture is fixed. The number of I pictures may be 2 or more. As the number of I pictures in 1 GOP increases, the image quality as a moving image improves, but the amount of coded data increases at the cost. Assuming that the playback is performed at 30 frames / second and 1 GOP consists of 15 frames, about 2 I pictures in 1 GOP will be sufficient.
In the block division unit 31, one frame of the input moving image is divided into blocks of N × N (N is a natural number) (each block has a size of 32 × 32 pixels), and each block image is divided into the difference calculation unit 32. Send to motion compensation unit 42. The motion compensation unit 42 calculates a motion vector from the frame memory 41 for the input block image and obtains the block image data which is the prediction data thereof. When the mode determination unit 43 selects the inter-frame coding mode (interframe coding mode), the difference calculation unit 32 subtracts the prediction data from the current frame. Further, when the mode determination unit 43 selects the intra-frame coding (intra-frame coding) mode, the difference is not taken (the coefficients of the prediction data may be all 0), and the input frame information is used as it is in the DWT. Output to unit 33.
The DWT unit 33 performs discrete wavelet transform and outputs it to the quantization unit 34. The quantization unit 34 quantizes the coefficient after the discrete wavelet transform and outputs it to the entropy encoding unit 35 and the inverse quantization unit 39. The entropy coding unit 35 encodes the quantized coefficient for each bit plane and outputs it to the bit plane truncation unit 36. The bit plane truncation unit 36 truncates the bit plane so that the code amount of 1 GOP falls within a predetermined code amount, and outputs the bit plane to the code forming unit 37.
When the threshold of the coded data amount of I picture (in-frame coded data) is defined as Ti and the threshold value of the code amount of P picture (interframe coded data) is defined as Tp, the allowable amount of data amount of 1 GOP is Ti ×. It can be expressed by n + Tp × m (in the embodiment, n = 1, m = 14). It is defined as the amount of data D of one frame received by the bit plane truncation unit 36 from the entropy encoding unit 35.
Now, when the in-frame coding mode is selected and there is a relationship of D Ti, the bit plane truncation unit 36 does not perform truncation. When there is a relationship of D> Ti, the bit plane truncation unit 36 truncates the coded data of the bit plane going upward from the least significant bit plane input from the entropy coding unit 35 until the relationship becomes D Ti. I will go.
For example, if the coded data of the nth bit plane is C (n), the amount of coded data is L ((Cn)), and the most significant bit is Nmax, then L (C (Nmax)) + L Find the maximum value of k that satisfies (C (Nmax-1)) + ... + L (C (Nmax-k)) Ti, and find C (Nmax), C (Nmax-1), ..., Outputs C (Nmax-k) as valid encoded data, and discards the encoded data C (Nmax-k-1), C (Nmax-k-2), ..., C (0).
The above is the same for the P picture. However, the threshold value in the case of P picture is different in that it is Tp. The I picture is a reference when generating the P picture, and it is desired that the image quality is high. Further, since the I picture is a picture to be coded in the frame, the relationship between the threshold value Ti and Tp is Ti> Tp. As a result of the above, the amount of data of 1 GOP can be maintained below the allowable amount of data. The threshold values Ti and Tp may be appropriately determined.
The coding unit 37 adds header information to the code in the code forming unit 37, and outputs the coded data.
As described above, the inverse quantization unit 39 and the inverse DCT unit 40 function only when the switch 38 is turned on based on the information indicating the in-frame coding mode from the mode determination unit 43. Therefore, the data that has passed through the inverse quantization unit 39 and the inverse DWT unit 40 becomes the restored image data (meaning that it is not difference data). This restored image data will be stored in the frame memory at 41. In the case of the embodiment, since there is one I picture in 1 GOP, the frame memory 41 is updated at intervals of 15 frames. Of course, if there are two or three I pictures in one GOP, the frame memory 41 will be updated at each interval.
The motion compensation unit 42 obtains a motion vector by referring to the input image and the reference frame stored in the frame memory 41 only when the frame to be currently encoded is encoded between frames, and generates prediction data.
The simple flow of the above moving image coding process will be described with reference to the flowchart of FIG. The figure is a flowchart showing a processing procedure in the moving image coding apparatus according to the first embodiment.
First, in step S100, when coding starts, the picture type flag PicType indicating the coding mode is set to 0, and the counter cnt is set to 0. When this picture type flag PicType is 0, it indicates an inter-frame coding mode, and when it is 1, it indicates an in-frame coding mode. The counter cnt counts up each time a frame is input, and is reset to "0" again when the value exceeds 14. That is, the range from 0 to 14 is repeatedly counted. This is because, in the embodiment, 1GOP = 15 frames, and 1GOP has one I-picture explaining an example.
Next, in step S116, it is determined whether or not the frame input is completed, and if not, the processing of step S101 and subsequent steps is repeated. If the device of the embodiment is a video camera, the determination of the end of frame input is determined by whether or not the recording button (not shown) is turned off. Further, it may be determined whether or not the set number of frames (or time) has been reached.
Proceeding to step S101, one frame of the image is input and divided into blocks for wavelet transform. At this time, the counter cnt is increased by 1. Next, in step S102, it is determined whether or not it is the timing to encode the input frame as an I picture. This determination is made based on whether or not the counter cnt = 1.
If it is determined that the counter cnt = 1, the process proceeds to step S104, the flag PicType is set to 1, and the in-frame coding mode coding process is set for the input frame. If the counter cnt is other than "1", the input frame performs the inter-frame coding mode, so the flag PicType is set to "0" in step S103.
When any of the processes of steps S103 and S104 is performed, the flag PicType is set to either 0 or 1, and the mode determination unit 43 in FIG. 1 makes this determination. The mode determination unit 43 supplies the value set in the flag PicType as a signal to each of the difference calculation unit 32, the inverse quantization unit 39, the inverse DWT unit 40, and the bit plane truncation unit 36 in FIG. When the supplied signal is 1, the difference calculation unit 32 does not use the signal from the motion compensation unit 42, but supplies each input pixel block as it is to the DWT unit 32, and when it is 0. Calculates the difference between the pixel block and the input block from the motion compensation unit 42, and supplies the result to the DWT unit 33.
The bit plane truncation unit 36 selects either the threshold value Ti or Tp according to the signal from the mode determination unit 43, and the code data of the bit plane going from the lowest to the upper level so that the code amount is equal to or less than the selection threshold value. Will be truncated.
Next, the process proceeds to step S105, DWT conversion is performed by the DWT unit 32 for each block from the difference calculation unit 32, and quantization processing is performed by the quantization unit 34 in step S106.
In the next step S107, it is determined whether or not the flag PicType is 1, that is, whether or not it is in the in-frame coding mode. When it is determined that the flag PicType is "1", the switch 38 is turned ON in step S108, and the inverse quantization unit 39 and the inverse DWT unit 40 are set to function. If the flag PicType is 0, the switch 38 is turned off and the processes of steps S109 to S111 are not performed.
When the process proceeds to step S109, the inverse quantization unit 39 performs the inverse quantization process, the inverse DWT transform is performed in step S110, and the image data of the conversion result is stored in the frame memory 41 in step S111. Then, update the frame memory 41.
In step S113, the entropy coding unit 35 performs entropy coding. This entropy coding is also a process of generating coded data for each bit plane.
Next, in step S114, the bit plane truncation unit 36 causes the bit plane truncation unit 36 to perform truncation processing of the coded data of the bit plane going from the least significant bit plane to the upper side so that the coded data is within the set threshold value. Then, in step S115, a predetermined header (including information indicating whether it is an I or P picture) is added to the coded data for one frame to generate the coded data. Output. After that, the process returns to step S116, and the above process is repeated.
As described above, according to the present embodiment, in the moving image coding in which bit plane coding is performed and the code amount is controlled by truncating the bit plane, the frame to be coded between frames is a frame to be coded in frame. By performing motion compensation by referring only to the image, the cumulative error due to motion compensation on the decoding side, that is, the error not accumulated as in the case of generating a P picture from a P picture, is not accumulated, so that image deterioration is suppressed. It is possible to generate video-encoded data.
In the embodiment, 1 GOP is composed of 15 frames, one I picture is generated in 1 GOP, and the remaining 14 frames are generated by P pictures. However, the P picture near the end is temporally from the I picture. There is a high possibility that the accuracy of motion compensation will deteriorate. In such a case, by setting the number of I pictures to about 2 or 3 and allocating the number of P pictures to be inserted between them on average, it is possible to deal with the case where there is a relatively large moving object. Yeah.
Further, in the embodiment, whether or not to generate an I picture is determined by counting the number of frames and according to the count value, but the amount of coded data in a predetermined time unit (or a predetermined number of GOPs) It may be determined whether or not to generate an I picture according to the size of. In this case, it should be avoided that the I-pictures are solidified and generated. Therefore, it is desirable to generate the coded data on the condition that at least one P-picture is generated after the I-pictures are generated.
In the present embodiment, only the I picture and the P picture have been described, but the present invention is not limited to this, and a B picture for bidirectional prediction may be introduced. In the case of a B picture, two frames are referenced, so this can be achieved by increasing the capacity of the frame memory and storing the referenced two-frame image in the frame memory.
Further, each process in FIG. 1 in the embodiment may be realized by software executed by a personal computer or the like. In this case, the input of moving image data can be dealt with by installing a video capture card or the like. In addition, a computer program can usually be executed by setting a computer-readable storage medium such as a CD-ROM containing the computer program in the computer and copying or installing it in the system. Therefore, of course, such a computer-readable storage medium Is also included in the scope of the present invention.
Further, the coding method is not limited to the JPEG2000 coding method, and the extension layer coding method in the FGS coding of the MPEG-4 coding method may be adopted.
Further, a B picture that performs two-way prediction may be introduced. In this case, it can be realized by performing motion compensation by referring to the previous and next I pictures.
Further, in the present embodiment, the coefficient after quantization is inversely quantized to obtain a decoded image, but the present invention is not limited to this. FIG. 11 is a configuration diagram in the case where the coefficient after bit truncation is inversely quantized to obtain a decoded image. In the figure, the dequantization unit 60 that receives the result of decoding by the entropy decoding unit 61 has a function of performing dequantization by shifting by the number of truncated bits. As a result, it is possible to realize a moving image coding device in consideration of bit truncation.
<Second Embodiment> FIG. 3 is a block diagram showing a configuration of a moving image coding device according to a second embodiment of the present invention. In FIG. 3, the same numbers are assigned to the parts that perform the same functions as those in FIG. 1 of the first embodiment, and the description thereof will be omitted.
In the figure, 1414 is a frame memory for storing input image data, and 142 is a motion compensation unit that performs motion prediction from the frame memory 1414 and the input image and calculates a motion vector and prediction data. Reference numeral 143 is a switch that controls the output by the output of the mode determination unit 43, and when a signal indicating the in-frame coding mode is received, each input block is overwritten on the frame memory 1414. Then, in the inter-frame coding mode, the switch is turned off and no writing is performed in the frame memory 1414. By doing so, the image of the frame encoded in the frame before the input current frame is stored and held in the frame memory 1414, which is the same as that of the first embodiment.
Reference numeral 135 denotes an entropy coding unit that encodes the conversion coefficient generated by the DWT unit 33 for each bit plane. 144 and 145 are selectors that select input / output with a lossless selection signal given from the external indicator 150.
The lossless coding operation in the moving image coding device configured as described above will be described below. An example in which the GOP is composed of only the I picture and the P picture will be described as in the first embodiment. Further, in the embodiment, JPEG2000 will be described as an example, but the description is not limited to this.
Similar to the first embodiment, the block division unit 31 divides the input frame into blocks and sends the input frame to the difference calculation unit 32, the motion compensation unit 142, and the switch 143. The mode selection 43 generates a signal indicating either the in-frame coding mode or the inter-frame coding mode as the coding mode of the input frame, and outputs the signal to the differential calculation unit 32, the motion compensation unit 142, and the switch 143.
When the mode determination unit 43 receives a signal indicating the inter-frame coding mode for the current frame, the difference calculation unit 32 subtracts the prediction data by motion compensation from each block divided by the block division unit 31. When a signal indicating the in-frame coding mode is received, the difference calculation is not performed and the information of the input frame is output to the DWT unit 33 as it is. The DWT unit 33 performs the discrete wavelet transform, encodes the coefficient after the discrete wavelet transform to be output to the entropy coding unit 135 in bit plane units, and outputs it to the code forming unit 37. The entropy encoding unit 135 encodes the quantized coefficient and outputs it to the selector 144. When the selector 144 is instructed to perform lossless coding from the outside, the data encoded by the entropy coding unit 135 is directly output to the code forming unit 37 without interposing the bit plane truncation unit 36. When instructed to perform lossy coding, the coded data generated by the entropy coding unit 135 is supplied to the bit plane truncation unit 36, and the result is supplied to the code forming unit 37.
The bit plane truncation unit 36 performs the same processing as in the first embodiment. That is, the coded data of the bit plane is truncated so that the coded data amount falls within a predetermined code amount. The code forming unit 37 adds header information to the code and outputs the coded data.
On the other hand, when the switch 143 receives a signal indicating the in-frame coding mode from the mode determination unit 43, the switch 143 is turned ON and sends and writes an input frame to the frame memory 1414. At this time, the motion compensation unit 142 does not operate, and 0 is output to the difference unit 32 as prediction data. On the other hand, when a signal indicating the inter-frame coding mode is received, the switch 143 is turned off and the input frame is not sent to the frame memory 1414 (the frame memory 1414 is not updated). The motion compensation 142 obtains a motion vector by referring to the input image and the reference frame stored in the frame memory 1414 for the frame to be encoded at present, and generates prediction data. That is, the frame memory 141 holds the information until it is overwritten.
As described above, a simple flow of the moving image coding process in the second embodiment will be described with reference to the flowchart of FIG.
First, in step S200, each parameter is initialized. Then, in step S212, the following processes after step S201 are repeated until it is determined that the coding process is completed. The processing of steps S200 and S212 is the same as that of steps S100 and S116 of the first embodiment.
Proceeding to step S201, one frame of image is input and divided into blocks for wavelet transform. At this time, the counter cnt is increased by 1. Next, in step S202, it is determined whether or not it is the timing to encode the input frame as an I picture. This determination is made based on whether or not the counter cnt = 1.
When it is determined that the coded data for the I picture is to be created, the switch 143 is turned ON in step S203. Then, the frame image input in step S204 is stored in the frame memory 1414 to update it, and the flag PicType is set to 1.
On the other hand, when the input frame is encoded as a P picture, the switch 143 is turned off in step S205 so that the frame memory 1414 is not updated. Next, in step S206, motion compensation is performed between the image stored in the frame memory 1414 and the input image data, the result is output to the difference calculation unit 32, and the flag PicType is set to 0. Set.
In step S207, the input image data or the difference image data is subjected to the discrete wavelet transform by the DWT unit 33. Then, in step S208, entropy encoding is performed for each bit plane.
Next, in step S209, it is determined whether or not lossless coding is instructed, and if lossy coding is instructed, bit plane truncation processing is performed in step S210, and lossless coding is instructed. If so, the process of step S210 is not performed.
In step S211, the coded data is input, necessary headers and the like are added to the coded data, the code is formed, and the data is output. After that, the process returns to step S212, and the process of step S201 and subsequent steps is repeated until it is determined that the frame is the final frame.
As described above, according to the second embodiment, in the moving image lossless coding in which bit plane coding is performed, when performing interframe coding, in-frame coding has been performed in the past. By performing motion compensation with reference only to the frame image, it is possible to obtain the same effect as that of the first embodiment. Moreover, in the second embodiment, the inverse quantization unit and the inverse DWT unit shown in the first embodiment are not required, and when it is realized by hardware, the circuit scale can be reduced and it is realized by software. In that case, it is possible to reduce the load on the CPU. Further, according to the second embodiment, since it is possible to appropriately select whether or not to perform the bit plane truncation process, it is possible to deal with irreversible coding.
In this embodiment, only the I picture and the P picture have been described, but the present invention is not limited to this, and even if the B picture of bidirectional prediction is introduced, the frame memory is increased and the I picture is similarly referred to. It is possible.
Further, each process in FIG. 3 in the embodiment may be realized by software executed by a personal computer or the like. In this case, the input of moving image data can be dealt with by installing a video capture card or the like. In addition, a computer program can usually be executed by setting a computer-readable storage medium such as a CD-ROM containing it in the computer and copying or installing it in the system. Therefore, of course, such a computer-readable storage medium Is also included in the scope of the present invention.
Further, the coding method is not limited to the JPEG2000 coding method, and the extension layer coding method in the FGS coding of the MPEG-4 coding method may be adopted.
<Third Embodiment> Next, the third embodiment will be described. FIG. 5 is a block configuration diagram showing a moving image coding device according to the third embodiment.
In the figure, 300 is a central processing unit (CPU) that controls the entire device and performs various processes, 301 is an operating system (OS) required to control this device, a computer program for image compression processing, and calculations. It is a memory that provides the storage area required for the computer. 302 is a bus that connects various devices and exchanges data and control signals.
The 303 is an input unit composed of a pointing device such as a switch, a keyboard, and a mouse (registered trademark) for starting a device, setting various conditions, and instructing playback. Reference numeral 304 denotes a storage device (for example, a hard disk) for storing the above OS and various software. Reference numeral 305 is a storage device for storing a stream in a storage medium, and the storage medium includes a writable CD disc, a DVD disc, a magnetic tape, or the like. Reference numeral 306 is a camera that captures moving images. 307 is a monitor for displaying images, and 309 is a communication circuit, which is composed of LAN, public line, wireless line, broadcast radio wave, and the like. Reference numeral 308 is a communication interface for transmitting and receiving a stream via the communication circuit 309.
The memory 301 controls the entire device, stores the OS for operating various software and the software to be operated, an image area for storing image data, a code area for storing generated coded data, various operations and coding. There is a working area to store the parameters and data related to the watermark.
The moving image coding process will be described in such a configuration. A case where the image data input from the camera 306 is encoded and output to the communication circuit 309 will be described as an example.
The memory usage and storage status of the memory 301 is as shown in FIG. The memory 301 contains an OS for controlling the entire device and operating various software, moving image coding software for moving image coding, object extraction software for extracting objects from moving images, communication software for communication, and moving images from camera 305. Contains image input software that inputs images in frame units. The moving image coding software will be described by taking the one based on the Motion JPEG 2000 coding method as an example, but the description is not limited to this.
Prior to the process, the input unit 303 instructs the entire device to start, and each unit is initialized. An instruction as to whether or not to maintain compatibility with the Motion JPEG 2000 encoding method is input from the input unit 303, the software stored in the storage device 304 is expanded to the memory 301 via the bus 302, and the software is started. To.
In such a configuration, the code area and working area on the memory 301 are cleared to 0 prior to processing. To maintain compatibility with JPEG2000 coding, image area 2 is not used and is open.
The image input software stores the image data captured by the camera 305 frame by frame in the image area on the memory 301. After that, the object extraction software extracts the object from the image in the image area and stores the shape information in the image area.
Next, the operation of coding by the moving image coding software by the CPU 300 will be described according to the flowchart shown in FIG.
First, in step S301, the header required by the MotionJPEG2000 coding method is generated and stored in the code area secured on the memory 301. When the coded data is stored in the code area, the communication software sends the coded data to the communication line 309 via the communication interface 308, and after sending, clears the corresponding area of the code area. Hereinafter, the transmission of the coded data in the coded area will not be particularly mentioned.
In step S302, the end determination of the coding process is performed. When the end of the coding process is input from the input unit 303, all the processes are ended. If not, the process proceeds to step S303.
When the process proceeds to step S303, the image data is read from the image area on the memory 301. In step S304, it is determined whether the frame to be coded is intra-frame coded or inter-frame correlation code coded. When the input unit 303 instructs to maintain compatibility with Motion JPEG 2000, the coding mode is determined to perform in-frame coding. The determined result is stored in the working area on the memory 301. In addition, information on whether or not to maintain compatibility with Motion JPEG 2000 is also stored in the working area.
In step S305, it is determined whether or not the processing of all blocks is completed. When the coding processing of all the blocks is completed, the process returns to step S302 and the coding processing of the next frame is performed. Otherwise, proceed to step S306.
In step S306, the block to be encoded is extracted from the image area on the memory 301 and stored in the working area. Then, in step S307, the coding mode of the working area on the memory 301 is referred to, and if it is in-frame coding (I picture), the process proceeds to step S308. Otherwise, proceed to step S314.
In step S308, the discrete wavelet transform is performed on the block data stored in the working area, and the obtained conversion coefficient is re-stored in the portion where the block data stored in the working area is stored. Then, in step S309, the conversion coefficient stored in the working area is quantized, and the obtained quantization result is stored in the area where the conversion coefficient is stored in the previous processing of the working area.
In step S310, the compatibility information of the working area on the memory 301 with Motion JPEG2000 is referred to, and if the compatibility is maintained, the process proceeds to step S317. If not, the process proceeds to step S311.
In step S311, the quantization result stored in the working area on the memory 301 is inversely quantized, and the obtained conversion coefficient is stored in the portion of the working area where the quantization result is stored. Then, in step S312, the conversion coefficient stored in the working area is subjected to inverse discrete wavelet transform, and the image data obtained in step S313 is stored in the image area 2. The image stored in the image area 2 corresponds to the frame memory 41 in the first embodiment.
On the other hand, if it is determined in step S307 that the image is other than the I picture, motion compensation is performed between the decoded image stored in the image area 2 and the block extracted from the input image data in step S314. To calculate the motion vector and prediction error data. Then, the motion vector data is encoded in the same manner as the motion vector coding in MPEG-4 coding, and is stored in the code area on the memory 301. The prediction error data is stored in the working area on the memory 301.
Next, in step S315, the prediction error data stored in the working area on the memory 301 is subjected to discrete wavelet transform, and the obtained conversion coefficient is stored in the portion of the working area where the block data is stored. .. In step S316, the conversion coefficient stored in the working area is quantized, and the obtained quantization result is stored in the part where the conversion coefficient is stored in the working area.
In step S317, the quantization result obtained in step S309 or step S316 is encoded in bit plane units and stored in the working area on the memory 301. Then, in step S318, the coded data stored in the working area is selected and stored in the coded area on the frame memory 301 by selecting the coded data that can be transmitted by rate control. Next, in step S319, the coded data on the coded area is multiplexed and transmitted. After that, clear the working area and the code area. After that, the process returns to step S305.
Through such a series of operations, it becomes possible to select a method capable of coding that is highly compatible with the conventional still image coding method and an inter-frame coding method.
The moving image coding processing of the first embodiment and the second embodiment may be realized by software, or the moving image coding device of the third embodiment may be realized by hardware. Absent.
<Fourth Embodiment> FIG. 12 is a block configuration diagram of the moving image decoding apparatus in the fourth embodiment.
In the fourth embodiment, the MotionJPEG2000 coding method will be described as an example of the image coding method used by the moving image decoding apparatus, but the present invention is not limited to this.
In FIG. 12, 71 is a code separator that separates the code string into header information and an image code data string, and 72 is an entropy decoding unit that performs the EBCOT decoding process of the JPEG2000 coding method for each bit plane. , 73 is an inverse quantization unit that dequantizes the decoded post-quantization coefficient, and 74 is an inverse DWT unit that performs an inverse DWT conversion on the DWT coefficient.
Reference numeral 77 denotes a mode determination unit that determines the coding mode on a frame-by-frame basis, and determines either an intra-frame coding (intra-frame coding) mode or an inter-frame coding (inter-frame coding) mode. Reference numeral 75 denotes an addition unit for obtaining addition of the image data with the prediction data obtained by motion compensation described later. 79 is a frame memory for storing the decoded image for motion compensation reference. The frame memory 79 stores the decoded image only when the mode determination unit 77 determines to decode in the in-frame coding mode. Therefore, a switch 80 for storing the output of the addition calculation unit 75 in the frame memory 79 is provided.
78 is motion compensation that calculates motion vector and prediction data by performing motion prediction from the frame memory 79 and the input image.
The code separation unit 71 inputs the coded data in GOP units, and separates the coded data into a header, a code related to the DCT coefficient, a motion vector code, and the like. The entropy decoding unit 72 entropy-decodes the separated coding and outputs it to the inverse quantization unit 73. The inverse quantization unit 73 dequantizes the information about the DC coefficient and outputs it to the inverse DWT. In the reverse DWT 74, the reverse DWT is applied and output to the addition unit 75.
The mode determination unit 77 detects that the image being decoded is an I picture or any other image based on the header of the picture layer in the input encoded data in the information obtained from the code separation unit 71, and determines the determination. The result information is sent to the addition unit 75, the motion compensation unit 78, and the switch 80. When the result of the mode determination is the I picture, the switch 80 writes the output from the adder 75 to the frame 79. That is, the contents of the frame memory 79 are updated only when the I picture is decoded.
The motion compensation unit 78 performs motion compensation using the motion vector information output from the entropy decoding unit 72 and the information of the frame memory 79, and sends the compensation image to the addition unit 75. If the result of the mode determination is an I picture, the addition unit 75 sends the result of the inverse DWT74 as it is to the block coupling 76, and if it is not an I picture, the image output from the motion compensation 78 and the image of the inverse DWT74 (difference image). ) Is added to the corresponding pixels and sent to the block join 76.
The simple flow of the above moving image decoding process will be described with reference to the flowchart of FIG. The figure is a flowchart which shows the decoding processing procedure in the moving image decoding apparatus which concerns on 4th Embodiment. Briefly explaining the outline, the processing from step S500 to step S511 is performed in units of small areas that divide the image into small areas, and in step S512, the small areas are combined to generate and output a video frame.
First, when the decoding is started (step S500), the presence / absence of the input code is detected (step S501), and when the decoding of all frames is completed, the decoding is completed (step S502). If all decoding is not completed, the code is separated into a header, an image signal code, a motion vector code, and the like (step S503).
Next, the entropy decoding process is performed (step S504). Then, the picture type of the image to be decoded is detected from the header information (step S505). In the case of I-picture, dequantization and inverse DWT are applied to the entropy-decoded result (steps S509 to S510). After that, the decoded image is stored in the frame memory (step S511), and the result is output as a video frame (step S512).
If the picture type is not I-picture in step S505, inverse quantization and inverse DWT are applied (steps S506 to S507), motion compensation is performed using the decoded image stored in step S511 (step S508), and the reverse of the compensation image. It is added to the result of DWT (step S507) and output as a video frame (step S512).
As described above, according to the fourth embodiment, even when the moving image coded data in which the bit plane coding is performed and the code amount is controlled by the bit plane truncation is input, the P picture is used. There is no decoding process that accumulates errors that generate P-pictures, and it is possible to reproduce good moving images.
In the fourth embodiment, only the I picture and the P picture have been described, but the present invention is not limited to this, and the frame memory is increased even for the code data in which the B picture of the bidirectional prediction is introduced, and the I picture is similarly used. It can be realized by referring to it.
Further, each process in FIG. 12 in the embodiment may be realized by software executed by a personal computer or the like. In this case, the input of moving image data can be dealt with by installing a video capture card or the like. In addition, a computer program can usually be executed by setting a computer-readable storage medium such as a CD-ROM containing the computer program in the computer and copying or installing it in the system. Therefore, of course, such a computer-readable storage medium Is also included in the scope of the present invention.
Further, the coding method is not limited to the JPEG2000 coding method, and the extension layer coding method in the FGS coding of the MPEG-4 coding method may be adopted.
Although the first to fourth embodiments have been described above, since it is clear that the functions corresponding to each processing unit can be realized by a computer program, the present invention also includes a computer program in its category. In addition, a computer program is usually stored in a computer-readable storage medium such as a CD-ROM, and can be executed by setting it in a computer and copying or installing it in the system. Such computer-readable storage media also fall within the scope of the present invention.
<figref num="1">It is a block block diagram of the moving image coding apparatus in 1st Embodiment.</figref><figref num="2">It is a flowchart which shows the moving image coding processing procedure in 1st Embodiment.</figref><figref num="3">It is a block block diagram of the moving image coding apparatus in 2nd Embodiment.</figref><figref num="4">It is a flowchart which shows the moving image coding processing procedure in 2nd Embodiment.</figref><figref num="5">It is a block block diagram of the moving image coding apparatus in 3rd Embodiment.</figref><figref num="6">It is a flowchart which shows the moving image coding processing procedure in 3rd Embodiment.</figref><figref num="7">It is a figure which shows the memory map during processing in 3rd Embodiment.</figref><figref num="8">It is a block block diagram of the image coding apparatus of JPEG2000.</figref><figref num="9">It is a block block diagram of the decoding apparatus of JPEG2000.</figref><figref num="10">It is a block block diagram of the conventional moving image coding apparatus.</figref><figref num="11">It is another block block diagram of the moving image coding apparatus in 1st Embodiment.</figref><figref num="12">It is a block block diagram of the moving image decoding apparatus in 4th Embodiment.</figref><figref num="13">It is a flowchart which shows the processing procedure in 4th Embodiment.</figref>
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9106915B2 | Cited by | United States of America | Applicant |
| WO2017115483A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2007318752A | Cited by | Japan | Examiner |
| JP2017120979A | Cited by | Japan | Search report |
| JP2002034043A | Cites | Japan | Examiner |
| JPH0795571A | Cites | Japan | Examiner |
6 priority claims, no other members on record
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004071399 | Japan | A | |
| 2004071399 | Japan | – | |
| 2005015847 | Japan | A | |
| 2004200471399 | – | – | – |
| JP20040071399 | – | – | – |
| JP20050015847 | – | – | – |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2005295505
- Publication, DOCDB
- 2005295505
- Publication, EPODOC
- JP2005295505
- Application
- 15847
- Application, DOCDB
- 2005015847
- Application, EPODOC
- JP20050015847
Titles3
- Japanese
- 動画像符号化装置及び動画像復号装置及びそれらの制御方法、並びに、コンピュータプログラム及びコンピュータ可読記憶媒体
- English
- Video coding device, video decoding device, their control method, computer program, and computer-readable storage medium.
- English
- MOVING IMAGE CODING APPARATUS, MOVING IMAGE DECODING APPARATUS, CONTROL METHOD THEREFOR, COMPUTER PROGRAM, AND COMPUTER-READABLE STORAGE MEDIUM
Classification
- CPC, 1
- H04N19/63
- IPC, 15
- H03M7 36
- H04N19 102
- H04N19 134
- H04N19 146
- H04N19 50
- H04N19 159
- H04N19 177
- H04N19 34
- H04N19 503
- H04N19 513
- H04N19 60
- H04N19 61
- H04N19 63
- H04N19 85
- H04N19 91