Device and method for generating stream, device and method for transmitting stream, device and method for encoding and recording medium
14 claims: 8 independent, 6 dependent
- 1In a coded stream converter that converts a coded stream into a recoded stream An input means for inputting the history coding parameters used in the past coding process or decoding process for the coded stream together with the coded stream. A conversion means for converting the coded stream input by the input means into the recoded stream, and A combination information generating means for generating combination information indicating a selective combination of the history coding parameters, and An output means that outputs the history coding parameter corresponding to the combination information generated by the combination information generation means and the combination of the combination information together with the recoded stream. A coded stream converter comprising. 符号化ストリームを再符号化ストリームに変換処理する符号化ストリーム変換装置において、 前記符号化ストリームに対する過去の符号化処理または復号処理において利用された履歴符号化パラメータを、前記符号化ストリームとともに入力する入力手段と、 前記入力手段により入力された前記符号化ストリームを前記再符号化ストリームに変換処理する変換手段と、 前記履歴符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成手段と、 前記組み合わせ情報生成手段により生成された前記組み合わせ情報及び前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記再符号化ストリームとともに出力する出力手段と を備える符号化ストリーム変換装置。
- 8In the coded stream conversion method of the coded stream converter that converts the coded stream into a recoded stream. An input step for inputting the history coding parameters used in the past coding process or decoding process for the coded stream together with the coded stream. A conversion step of converting the coded stream input by the processing of the input step into the recoded stream, and a conversion step. A combination information generation step that generates combination information indicating a selective combination of the history coding parameters, and With the output step that outputs the history coding parameter corresponding to the combination information generated by the processing of the combination information generation step and the combination of the combination information together with the recoded stream. A coded stream conversion method that includes. 符号化ストリームを再符号化ストリームに変換処理する符号化ストリーム変換装置の符号化ストリーム変換方法において、 前記符号化ストリームに対する過去の符号化処理または復号処理において利用された履歴符号化パラメータを、前記符号化ストリームとともに入力する入力ステップと、 前記入力ステップの処理により入力された前記符号化ストリームを前記再符号化ストリームに変換処理する変換ステップと、 前記履歴符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成ステップと、 前記組み合わせ情報生成ステップの処理により生成された前記組み合わせ情報および前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記再符号化ストリームとともに出力する出力ステップと を含む符号化ストリーム変換方法。
- 9In a coded stream converter that converts a coded stream into a recoded stream An input means for inputting the history coding parameters used in the past coding process or decoding process for the coded stream together with the coded stream. A conversion means for converting the coded stream input by the input means into the recoded stream, and A coding parameter calculation means for calculating the current coding parameter, and A combination information generating means that generates combination information indicating a selective combination of the history coding parameter and the current coding parameter calculated by the coding parameter calculating means, and a combination information generating means. An output means that outputs the history coding parameter corresponding to the combination information generated by the combination information generation means and the combination of the combination information together with the recoded stream. A coded stream converter comprising. 符号化ストリームを再符号化ストリームに変換処理する符号化ストリーム変換装置において、 前記符号化ストリームに対する過去の符号化処理または復号処理において利用された履歴符号化パラメータを、前記符号化ストリームとともに入力する入力手段と、 前記入力手段により入力された前記符号化ストリームを前記再符号化ストリームに変換処理する変換手段と、 現在の符号化パラメータを算出する符号化パラメータ算出手段と、 前記履歴符号化パラメータ、および、前記符号化パラメータ算出手段により算出された前記現在の符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成手段と、 前記組み合わせ情報生成手段により生成された前記組み合わせ情報および前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記再符号化ストリームとともに出力する出力手段と を備える符号化ストリーム変換装置。
- 10In the coded stream conversion method of the coded stream converter that converts the coded stream into a recoded stream. An input step for inputting the history coding parameters used in the past coding process or decoding process for the coded stream together with the coded stream. A conversion step of converting the coded stream input by the processing of the input step into the recoded stream, and a conversion step. A coding parameter calculation step to calculate the current coding parameter, and A combination information generation step that generates combination information indicating a selective combination of the history coding parameter and the current coding parameter calculated by the processing of the coding parameter calculation step, and a combination information generation step. With the output step that outputs the history coding parameter corresponding to the combination information generated by the processing of the combination information generation step and the combination of the combination information together with the recoded stream. A coded stream conversion method that includes. 符号化ストリームを再符号化ストリームに変換処理する符号化ストリーム変換装置の符号化ストリーム変換方法において、 前記符号化ストリームに対する過去の符号化処理または復号処理において利用された履歴符号化パラメータを、前記符号化ストリームとともに入力する入力ステップと、 前記入力ステップの処理により入力された前記符号化ストリームを前記再符号化ストリームに変換処理する変換ステップと、 現在の符号化パラメータを算出する符号化パラメータ算出ステップと、 前記履歴符号化パラメータ、および、前記符号化パラメータ算出ステップの処理により算出された前記現在の符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成ステップと、 前記組み合わせ情報生成ステップの処理により生成された前記組み合わせ情報および前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記再符号化ストリームとともに出力する出力ステップと を含む符号化ストリーム変換方法。
- 11In a stream output device that outputs a coded stream, An input means for inputting the history coding parameters used in the past coding process or decoding process for the coded stream, and A combination information generating means for generating combination information indicating a selective combination of the history coding parameters according to an application using the coded stream, and a combination information generating means. An output means that outputs the history coding parameter corresponding to the combination information generated by the combination information generation means and the combination of the combination information together with the coded stream. A stream output device equipped with. 符号化ストリームを出力するストリーム出力装置において、 前記符号化ストリームに対する過去の符号化処理または復号処理において使用された履歴符号化パラメータを入力する入力手段と、 前記符号化ストリームを利用するアプリケーションに応じた前記履歴符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成手段と、 前記組み合わせ情報生成手段により生成された前記組み合わせ情報及び前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記符号化ストリームとともに出力する出力手段と を備えるストリーム出力装置。
- 12In the stream output method of the stream output device that outputs the coded stream, An input step for inputting the history coding parameters used in the past coding or decoding process for the coded stream. A combination information generation step that generates combination information indicating a selective combination of the history coding parameters according to the application that uses the coded stream, and a combination information generation step. With the output step that outputs the history coding parameter corresponding to the combination information generated by the processing of the combination information generation step and the combination of the combination information together with the coded stream. Stream output method including. 符号化ストリームを出力するストリーム出力装置のストリーム出力方法において、 前記符号化ストリームに対する過去の符号化処理または復号処理において使用された履歴符号化パラメータを入力する入力ステップと、 前記符号化ストリームを利用するアプリケーションに応じた前記履歴符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成ステップと、 前記組み合わせ情報生成ステップの処理により生成された前記組み合わせ情報及び前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記符号化ストリームとともに出力する出力ステップと を含むストリーム出力方法。
- 13In a stream output device that outputs a coded stream, An input means for inputting the history coding parameters used in the past coding process or decoding process for the coded stream, and Combination information indicating a selective combination of the history coding parameters according to the capacity of the transmission line for transmitting the coded stream or the recording medium for recording the coded stream.Combination information generation means to be generated and An output means that outputs the history coding parameter corresponding to the combination information generated by the combination information generation means and the combination of the combination information together with the coded stream. A stream output device equipped with. 符号化ストリームを出力するストリーム出力装置において、 前記符号化ストリームに対する過去の符号化処理または復号処理において使用された履歴符号化パラメータを入力する入力手段と、 前記符号化ストリームを伝送する伝送路又は前記符号化ストリームを記録する記録媒体の容量に応じた前記履歴符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成手段と、 前記組み合わせ情報生成手段により生成された前記組み合わせ情報及び前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記符号化ストリームとともに出力する出力手段と を備えるストリーム出力装置。
- 14In the stream output method of the stream output device that outputs the coded stream, An input step for inputting the history coding parameters used in the past coding or decoding process for the coded stream. A combination information generation step for generating combination information indicating a selective combination of the history coding parameters according to the capacity of the transmission line for transmitting the coded stream or the recording medium for recording the coded stream. With the output step that outputs the history coding parameter corresponding to the combination information generated by the processing of the combination information generation step and the combination of the combination information together with the coded stream. Stream output method including. 符号化ストリームを出力するストリーム出力装置のストリーム出力方法において、 前記符号化ストリームに対する過去の符号化処理または復号処理において使用された履歴符号化パラメータを入力する入力ステップと、 前記符号化ストリームを伝送する伝送路又は前記符号化ストリームを記録する記録媒体の容量に応じた前記履歴符号化パラメータの選択的な組み合わせを示す組み合わせ情報を生成する組み合わせ情報生成ステップと、 前記組み合わせ情報生成ステップの処理により生成された前記組み合わせ情報及び前記組み合わせ情報の組み合わせに対応する前記履歴符号化パラメータを、前記符号化ストリームとともに出力する出力ステップと を含むストリーム出力方法。
Independent claims8
514 paragraphs, as filed
[Technical field to which the invention belongs] The present invention<u style="single">A coded stream conversion device, a coded stream conversion method, a stream output device, and a stream output method.</u>In particular, it is suitable for use in a transcoding device for changing the GOP (Group of Pictures) structure of a coded bit stream encoded based on the MPEG standard or changing the bit rate of a coded bit stream. Nana<u style="single">A coded stream conversion device, a coded stream conversion method, a stream output device, and a stream output method.</u>Regarding.
[0002] In recent years, in broadcasting stations that produce and broadcast television programs, MPEG (Moving Picture Experts Group) technology has been generally used for compressing / encoding video data. Has become. In particular, when recording video data on a randomly accessible recording medium material such as tape, and when transmitting video data via a cable or satellite, this MPEG technology is becoming the de facto standard.
[0003] An example of processing in a broadcasting station until a video program produced in the broadcasting station is transmitted to each home will be briefly described. First, the source video data is encoded and recorded on a magnetic tape by an encoder provided in a camcorder in which a video camera and a VTR (Video Tape Recorder) are integrated. At this time, the camcorder encoder encodes the source video data so as to be suitable for the recording format of the VTR tape. For example, the GOP structure of the MPEG bitstream recorded on this magnetic tape is a structure consisting of 1 GOP from 2 frames (for example, I, B, I, B, I, B, ...). Will be done. The bit rate of the MPEG bit stream recorded on the magnetic tape is 18 Mbps.
Next, the main broadcasting station performs an editing process for editing the video bit stream recorded on the magnetic tape. Therefore, the GOP structure of the video bitstream recorded on the magnetic tape is converted into a GOP structure suitable for editing processing. The GOP structure suitable for editing processing is a GOP structure in which 1 GOP is composed of one frame and all pictures are I pictures. This is because I-pictures, which do not correlate with other pictures, are the most suitable for editing on a frame-by-frame basis. In the actual operation, the video stream recorded on the magnetic tape is once decoded and returned to the baseband video data. Then, the baseband video signal is re-encoded so that all the pictures are I pictures. By performing the decoding process and the re-encoding process in this way, it is possible to generate a bit stream having a GOP structure suitable for the editing process.
Next, in order to transmit the edited video program generated by the above-mentioned editing process from the main station to the local station, the bit stream of the edited video program is converted into a GOP structure and a bit rate suitable for the transmission process. To do. The GOP structure suitable for transmission between broadcasting stations is, for example, a GOP structure in which 1 GOP is composed of 15 frames (for example, I, B, B, P, B, B, P ...). In addition, the bit rate suitable for transmission between broadcasting stations is generally a high bit rate of 50 Mbps or more because a dedicated line having a high transmission capacity such as an optical fiber is provided between broadcasting stations. Is desirable. Specifically, the bitstream of the edited video program is once decoded and returned to the baseband video data. Then, the baseband video data is re-encoded so as to have a GOP structure and a bit rate suitable for transmission between the above-mentioned broadcasting stations.
[0006] In the local station, an editing process is performed in order to insert a commercial peculiar to the local area into the video program transmitted from the main station. That is, in the same manner as the editing process described above, the video stream transmitted from the main station is once decoded and returned to the baseband video data. Then, by re-encoding the baseband video signal so that all the pictures are I pictures, a bitstream having a GOP structure suitable for editing processing can be generated.
[0007] Subsequently, the video program edited at this local station is converted into a GOP structure and bit rate suitable for this transmission process in order to be transmitted to each home via a cable or satellite. For example, a GOP structure suitable for transmission processing for transmission to each home is a GOP structure in which 1 GOP consists of 15 frames (for example, I, B, B, P, B, B, P ...). Therefore, the bit rate suitable for the transmission process for transmission to each home is a low bit rate of about 5 Mbps. Specifically, the bitstream of the edited video program is once decoded and returned to the baseband video data. Then, the baseband video data is re-encoded so as to have a GOP structure and a bit rate suitable for the above-mentioned transmission process.
[Problem to be Solved by the Invention] As can be understood from the above description, a plurality of decoding processes and encoding processes are repeated while a video program is transmitted from a broadcasting station to each home. There is. Actually, the processing in the broadcasting station requires various signal processing other than the above-mentioned signal processing, and the decoding processing and the coding processing must be repeated each time.
[0009] However, it is well known that the coding process and the decoding process based on the MPEG standard are not 100% reversible processes. That is, the baseband video data before being encoded and the video data after being decoded are not 100% the same, and the image quality is deteriorated by this coding process and the decoding process. That is, as described above, when the decoding process and the encoding process are repeated, there is a problem that the image quality deteriorates each time the processing is performed. In other words, the deterioration of image quality accumulates every time the decoding / encoding process is repeated.
[0010] The present invention has been made in view of such a situation, and decoding and coding for changing the structure of the GOP (Group of Pictures) of a coded bit stream encoded based on the MPEG standard. This is intended to realize a transcoding system in which image quality deterioration does not occur even if the conversion process is repeated.
[Means for Solving Problems] [Means for Solving Problems]<u style="single"> The coded stream conversion device according to the first aspect of the present invention is a coded stream conversion device that converts a coded stream into a recoded stream, and in a past coding process or decoding process for the coded stream. An input means for inputting the used history coding parameter together with the coded stream, a conversion means for converting the coded stream input by the input means into the recoded stream, and the history coding parameter. The combination information generation means for generating the combination information indicating the selective combination of the above, and the history coding parameter corresponding to the combination of the combination information and the combination information generated by the combination information generation means are re-encoded. It is provided with an output means for outputting together with a stream.</u>【0012】<u style="single"> The combination information can be information that is distinguished according to the degree of image quality deterioration due to the conversion process.</u>【0013】<u style="single"> The combination information can be information that is distinguished according to the application that uses the recoded stream, and the combination information generating means is provided with the combination information that satisfies the requirements of the application. Can be generated.</u>【0014】<u style="single"> The combination information may be information that is distinguished according to the transmission line through which the recoded stream is transmitted or the capacity of the recording medium on which the recoded stream is recorded.</u>【0015】<u style="single"> The output means may be made to describe and output the combination information and the history coding parameter corresponding to the combination of the combination information in the recoding stream.</u>【0016】<u style="single"> The conversion means can be made to reuse the history coding parameter to generate the recoded stream.</u>【0017】<u style="single">The conversion means can reuse the history coding parameter only when the picture type included in the physical history coding parameter matches the picture type used in the conversion process.</u>【0018】<u style="single"> The coded stream conversion method of the first aspect of the present invention includes an input step of inputting a history coding parameter used in a past coding process or a decoding process for the coded stream together with the coded stream, and the above-mentioned. A conversion step of converting the coded stream input by the processing of the input step into the recoded stream, a combination information generation step of generating combination information indicating a selective combination of the history coding parameters, and the above. It includes an output step that outputs the combination information generated by the processing of the combination information generation step and the history coding parameter corresponding to the combination of the combination information together with the recoded stream.</u>【0019】<u style="single">In the first aspect of the present invention, the historical coding parameters used in the past coding or decoding process for the coded stream are input together with the coded stream, and the coded stream is converted into a recoded stream. , Combination information indicating the selective combination of the history coding parameters is generated, and the history coding parameters corresponding to the generated combination information and the combination of the combination information are output together with the recoded stream.</u>【0020】<u style="single"> The coded stream conversion device according to the second aspect of the present invention is a coded stream conversion device that converts a coded stream into a recoded stream, and in a past coding process or decoding process for the coded stream. An input means for inputting the used history coding parameter together with the coded stream, a conversion means for converting the coded stream input by the input means into the recoded stream, and a current coding parameter. A combination information generating means that generates combination information indicating a selective combination of the coding parameter calculating means for calculating the above, the history coding parameter, and the current coding parameter calculated by the coding parameter calculating means. And an output means that outputs the history coding parameter corresponding to the combination information generated by the combination information generation means and the combination of the combination information together with the recoded stream.</u>【0021】<u style="single"> The coded stream conversion method of the second aspect of the present invention includes an input step of inputting a history coding parameter used in a past coding process or a decoding process for the coded stream together with the coded stream, and the above-mentioned. A conversion step for converting the coded stream input by the processing of the input step into the recoded stream, a coding parameter calculation step for calculating the current coding parameter, the history coding parameter, and the above. A combination information generation step that generates combination information indicating a selective combination of the current coding parameters calculated by the processing of the coding parameter calculation step, the combination information generated by the processing of the combination information generation step, and the combination information. It includes an output step that outputs the history coding parameter corresponding to the combination of the combination information together with the recoded stream.</u>【0022】<u style="single">In the second aspect of the present invention, the historical coding parameters used in the past coding or decoding process for the coded stream are input together with the coded stream, and the input coded stream is the recoded stream. Is converted to, the current coding parameters are calculated, the history coding parameters, and the combination information indicating the selective combination of the calculated current coding parameters is generated, and the generated combination information and the combination information The history coding parameters corresponding to the combination are output with the recoded stream.</u>【0023】<u style="single"> The stream output device of the third aspect of the present invention is a stream output device that outputs a coded stream, and inputs the history coding parameters used in the past coding process or decoding process for the coded stream. A combination information generating means that generates combination information indicating a selective combination of the input means and the history coding parameter according to the application that uses the coding stream, and the combination information generated by the combination information generating means. And an output means for outputting the history coding parameter corresponding to the combination of the combination information together with the coded stream.</u>【0024】<u style="single"> The stream output method of the third aspect of the present invention is for an input step for inputting a history coding parameter used in a past coding process or a decoding process for the coded stream, and an application using the coded stream. The history corresponding to the combination information generation step that generates the combination information indicating the selective combination of the history coding parameters according to the history, and the combination information generated by the processing of the combination information generation step and the combination of the combination information. It includes an output step that outputs the coding parameters together with the coding stream.</u>【0025】<u style="single">In the third aspect of the present invention, the history coding parameters used in the past coding process or decoding process for the coded stream are input, and the history coding parameters are selected according to the application using the coded stream. Combination information indicating a specific combination is generated, and the generated combination information and the history coding parameters corresponding to the combination of the combination information are output together with the coded stream.</u>【0026】<u style="single"> The stream output device of the fourth aspect of the present invention is a stream output device that outputs a coded stream, and inputs historical coding parameters used in the past coding process or decoding process for the coded stream. An input means and a combination information generating means for generating combination information indicating a selective combination of the history coding parameters according to the capacity of the transmission line for transmitting the coded stream or the recording medium for recording the coded stream. An output means that outputs the history coding parameter corresponding to the combination information generated by the combination information generation means and the combination of the combination information together with the coded stream.</u>【0027】<u style="single"> The stream output method of the fourth aspect of the present invention includes an input step for inputting history coding parameters used in the past coding process or decoding process for the coded stream, and a transmission line for transmitting the coded stream. Alternatively, the combination information generation step for generating combination information indicating a selective combination of the history coding parameters according to the capacity of the recording medium for recording the coded stream, and the combination information generation step generated by the processing of the combination information generation step. It includes an output step that outputs the combination information and the history coding parameter corresponding to the combination of the combination information together with the coded stream.</u>【0028】<u style="single">In the fourth aspect of the present invention, the historical coding parameters used in the past coding process or decoding process for the coded stream are input, and the transmission line or the coded stream for transmitting the coded stream is recorded. Combination information indicating a selective combination of history coding parameters according to the capacity of the medium is generated, and the generated combination information and the history coding parameters corresponding to the combination of the combination information are output together with the coded stream.</u>[Embodiments of the Invention] Hereinafter, a transcoder to which the present invention is applied will be described, but before that, compression coding of a moving image signal will be described. In addition, the term of the system in this specification means an overall apparatus composed of a plurality of apparatus, means and the like.
[0030] For example, in a system for transmitting a moving image signal to a remote location such as a video conferencing system or a videophone system, in order to efficiently use a transmission line, line correlation or interframe correlation of video signals is used. Then, the image signal is compressed and encoded.
By using the line correlation, the image signal can be compressed by, for example, DCT (discrete cosine transform) processing.
Further, by using the inter-frame correlation, the image signal can be further compressed and encoded. For example, as shown in FIG. 1, when frame images PC1 to PC3 are generated at times t1 to t3, the difference between the image signals of the frame images PC1 and PC2 is calculated to generate PC12, and the frame is also generated. The difference between the images PC2 and PC3 is calculated to generate PC23. Normally, the images of frames adjacent in time do not have such a large change, so when the difference between the two is calculated, the difference signal becomes a small value. Therefore, if this difference signal is encoded, the code amount can be compressed.
[0033] However, the original image cannot be restored by transmitting only the difference signal. Therefore, the image of each frame is set to one of three types of picture types, I picture, P picture, and B picture, and the image signal is compressed and coded.
That is, as shown in FIG. 2, for example, the image signals of 17 frames from frames F1 to F17 are designated as a group of pictures (GOP) and are used as one unit of processing. Then, the image signal of the first frame F1 is encoded as an I picture, the second frame F2 is processed as a B picture, and the third frame F3 is processed as a P picture. Hereinafter, the fourth and subsequent frames F4 to F17 are alternately processed as B pictures or P pictures.
[0035] As the image signal of the I picture, the image signal for one frame is transmitted as it is. On the other hand, as the image signal of the P picture, basically, as shown in FIG. 2, the difference from the image signal of the I picture or the P picture that precedes it in time is transmitted. Further, as the image signal of the B picture, basically, as shown in FIG. 3, the difference from the average value of both the frame preceding in time and the frame following in time is obtained, and the difference is encoded.
[0036] FIG. 4 shows the principle of the method of encoding the moving image signal in this way. As shown in the figure, since the first frame F1 is processed as an I picture, it is directly transmitted to the transmission line as transmission data F1X (in-image coding). On the other hand, since the second frame F2 is processed as a B picture, the difference between the time-preceding frame F1 and the time-lagging frame F3 is calculated, and the difference is calculated. It is transmitted as transmission data F2X.
[0037] However, there are four types of processing as the B picture, to be described in more detail. The first process is to transmit the data of the original frame F2 as it is as the transmission data F2X (SP1) (intra-coding), which is the same process as in the case of the I picture. The second process is to calculate the difference from the later frame F3 in time and transmit the difference (SP2) (backward prediction coding). The third process is to transmit the difference (SP3) from the time-preceding frame F1 (forward prediction coding). Further, the fourth process generates a difference (SP4) between the average value of the frame F1 that precedes in time and the frame F3 that follows, and transmits this as transmission data F2X (bidirectional prediction coding). ..
[0038] Actually, the method with the least transmission data among the above-mentioned four methods is adopted.
[0039] When transmitting the difference data, the motion vector x1 between the image (reference image) of the frame for which the difference is calculated (the motion vector between the frames F1 and F2) (in the case of forward prediction). , Or x2 (motion vector between frames F3 and F2) (for backward prediction), or both x1 and x2 (for bidirectional prediction) are transmitted along with the difference data.
[0040] Further, in the frame F3 of the P picture, the difference signal (SP3) from this frame and the motion vector x3 are calculated with the frame F1 preceding in time as the reference image, and this is transmitted as the transmission data F3X. (Forward prediction coding). Alternatively, the data in the original frame F3 is transmitted as is as data F3X (SP1) (intra-coding). Of these methods, the method with less transmitted data is selected, as in the case of the B picture.
FIG. 5 shows a configuration example of a device that encodes and transmits a moving image signal and decodes the moving image signal based on the above-mentioned principle. The coding device 1 encodes the input video signal and transmits it to the recording medium 3 as a transmission line. Then, the decoding device 2 reproduces the signal recorded on the recording medium 3, decodes the signal, and outputs the signal.
In the coding device 1, the input video signal is input to the preprocessing circuit 11, where the brightness signal and the color signal (in the case of the present embodiment, the color difference signal) are separated and A / D converted respectively. The analog signal is converted to a digital signal by the devices 12 and 13. The video signal converted into a digital signal by the A / D converters 12 and 13 is supplied to the frame memory 14 and stored. The frame memory 14 stores the luminance signal in the luminance signal frame memory 15 and the luminance signal in the luminance signal frame memory 16.
[0043] The format conversion circuit 17 converts the frame format signal stored in the frame memory 14 into a block format signal. That is, as shown in FIG. 6, the video signal stored in the frame memory 14 is regarded as frame format data as shown in FIG. 6 (A) in which V lines of H dots are collected per line. There is. The format conversion circuit 17 divides this one-frame signal into M slices in units of 16 lines, as shown in FIG. 6 (B). Then, each slice is divided into M macroblocks. As shown in FIG. 6C, the macroblock is composed of a luminance signal corresponding to 16 × 16 pixels (dots), and this luminance signal is further divided into blocks Y [1] in units of 8 × 8 dots. ] To Y [4]. The 16 × 16 dot luminance signal corresponds to an 8 × 8 dot Cb signal and an 8 × 8 dot Cr signal.
[0044] As described above, the data converted into the block format is supplied from the format conversion circuit 17 to the encoder 18, where encoding is performed. The details will be described later with reference to FIG. 7.
The signal encoded by the encoder 18 is output to the transmission line as a bit stream. For example, it is supplied to the recording circuit 19 and recorded on the recording medium 3 as a digital signal.
The data reproduced from the recording medium 3 by the reproduction circuit 30 of the decoding device 2 is supplied to the decoder 31 and decoded. Details of the decoder 31 will be described later with reference to FIG.
The data decoded by the decoder 31 is input to the format conversion circuit 32, and is converted from the block format to the frame format. Then, the luminance signal of the frame format is supplied to and stored in the luminance signal frame memory 34 of the frame memory 33, and the color difference signal is supplied to and stored in the luminance signal frame memory 35. The luminance signal and the luminance signal read from the luminance signal frame memory 34 and the color difference signal frame memory 35 are converted into analog signals by the D / A converters 36 and 37, respectively, and supplied to the post-processing circuit 38. The post-processing circuit 38 synthesizes and outputs a luminance signal and a color difference signal.
Next, the configuration of the encoder 18 will be described with reference to FIG. 7. The encoded image data is input to the motion vector detection circuit 50 in macroblock units. The motion vector detection circuit 50 processes the image data of each frame as an I picture, a P picture, or a B picture according to a predetermined sequence set in advance. Whether to process the images of each frame sequentially input as a picture of I, P, or B is predetermined (for example, as shown in FIGS. 2 and 3, frames F1 to F17). The group of pictures composed of is processed as I, B, P, B, P, ... B, P).
The image data of the frame (for example, frame F1) processed as an I picture is transferred and stored from the motion vector detection circuit 50 to the front original image unit 51a of the frame memory 51, and is processed as a B picture (for example). For example, the image data of the frame F2) is transferred and stored in the original image unit 51b, and the image data of the frame processed as a P picture (for example, frame F3) is transferred and stored in the rear original image unit 51c.
[0050] Further, at the next timing, when an image of a frame to be processed as a B picture (frame F4) or a P picture (frame F5) is further input, the first image stored in the rear original image unit 51c until then. The image data of the P picture (frame F3) of the above is transferred to the front original image section 51a, the image data of the next B picture (frame F4) is stored (overwritten) in the reference original image section 51b, and the next P picture is stored. The image data of (frame F5) is stored (overwritten) in the rear original image unit 51c. Such an operation is repeated in sequence.
The signal of each picture stored in the frame memory 51 is read out from the signal, and the frame prediction mode processing or the field prediction mode processing is performed in the prediction mode switching circuit 52.
[0052] Furthermore, under the control of the prediction determination circuit 54, the arithmetic unit 53 performs in-image prediction, forward prediction, backward prediction, or bidirectional prediction. Which of these processes is performed is determined according to the prediction error signal (difference between the reference image to be processed and the predicted image with respect to the reference image). Therefore, the motion vector detection circuit 50 generates the sum of absolute values (or the sum of squares) of the prediction error signals used for this determination.
[0053] Here, the frame prediction mode and the field prediction mode in the prediction mode switching circuit 52 will be described.
When the frame prediction mode is set, the prediction mode switching circuit 52 directly calculates the four luminance blocks Y [1] to Y [4] supplied by the motion vector detection circuit 50 in the subsequent stage. Output to device 53. That is, in this case, as shown in FIG. 8, the data of the odd-numbered field lines and the data of the even-numbered field lines are mixed in each luminance block. In this frame prediction mode, prediction is performed in units of four luminance blocks (macroblocks), and one motion vector corresponds to the four luminance blocks.
On the other hand, in the field prediction mode, the prediction mode switching circuit 52 displays the signals input from the motion vector detection circuit 50 in the configuration shown in FIG. 8 with four luminances as shown in FIG. Of the blocks, the luminance blocks Y [1] and Y [2] are composed of, for example, only the dots of the lines of the odd field, and the other two luminance blocks Y [3] and Y [4] are of the even field. It is composed of only line dots and output to the arithmetic unit 53. In this case, one motion vector corresponds to the two luminance blocks Y [1] and Y [2], and for the other two luminance blocks Y [3] and Y [4]. And one other motion vector is corresponded.
The motion vector detection circuit 50 outputs the sum of the absolute values of the prediction errors in the frame prediction mode and the sum of the absolute values of the prediction errors in the field prediction mode to the prediction mode switching circuit 52. The prediction mode switching circuit 52 compares the sum of the absolute values of the prediction errors in the frame prediction mode and the field prediction mode, performs processing corresponding to the prediction mode having a small value, and outputs the data to the calculator 53.
However, such processing is actually performed by the motion vector detection circuit 50. That is, the motion vector detection circuit 50 outputs a signal having a configuration corresponding to the determined mode to the prediction mode switching circuit 52, and the prediction mode switching circuit 52 outputs the signal as it is to the arithmetic unit 53 in the subsequent stage.
[0058] In the frame prediction mode, the color difference signal is supplied to the arithmetic unit 53 in a state where the data of the line of the odd field and the data of the line of the even field are mixed as shown in FIG. In the field prediction mode, as shown in FIG. 9, the upper half (4 lines) of the color difference blocks Cb and Cr is the color difference signal of the odd field corresponding to the luminance blocks Y [1] and Y [2]. The lower half (4 lines) is used as the color difference signal of the even field corresponding to the luminance blocks Y [3] and Y [4].
[0059] Further, the motion vector detection circuit 50 determines whether to perform in-image prediction, forward prediction, backward prediction, or bidirectional prediction in the prediction determination circuit 54 as shown below. Generate the absolute sum of prediction errors.
That is, as the absolute value sum of the prediction errors of the prediction in the image, the sum sum ΣAij | of the signal Aij of the macroblock of the reference image and the absolute value | Aij | of the signal Aij of the macroblock Σ Find the difference between | Aij |. Also, as the absolute value sum of the prediction errors of the forward prediction, the difference between the macroblock signal Aij of the reference image and the macroblock signal Bij of the predicted image Aij-Bij absolute value | Aij-Bij | sum total Σ | Aij- Ask for Bij |. Further, the sum of the absolute values of the prediction errors of the backward prediction and the bidirectional prediction is also obtained in the same manner as in the case of the forward prediction (changing the predicted image to a prediction image different from the case of the forward prediction).
The sum of these absolute values is supplied to the prediction determination circuit 54. The prediction determination circuit 54 selects the smallest absolute sum of the prediction errors of the forward prediction, the backward prediction, and the bidirectional prediction as the absolute sum of the prediction errors of the inter-prediction. Furthermore, the absolute sum of the prediction errors of this inter-prediction is compared with the absolute sum of the prediction errors of the in-image prediction, the smaller one is selected, and the mode corresponding to this selected absolute sum is set as the prediction mode. select. That is, if the sum of the absolute values of the prediction errors of the in-image prediction is smaller, the in-image prediction mode is set. If the absolute sum of the prediction errors of the inter-prediction is smaller, the mode in which the corresponding absolute sum of the forward prediction, the backward prediction, or the bidirectional prediction mode is the smallest is set.
[0062] As described above, the motion vector detection circuit 50 switches the prediction mode by switching the macroblock signal of the reference image to the mode selected by the prediction mode switching circuit 52 among the frame or field prediction modes. While supplying to the arithmetic unit 53 via the circuit 52, the motion vector between the prediction image and the reference image corresponding to the prediction mode selected by the prediction determination circuit 54 among the four prediction modes is detected, and the variable length code is used. Output to the conversion circuit 58 and the motion compensation circuit 64. As described above, as this motion vector, the one that minimizes the sum of the absolute values of the corresponding prediction errors is selected.
[0063] When the motion vector detection circuit 50 reads out the image data of the I picture from the front original image unit 51a, the prediction determination circuit 54 performs a frame or field (image) prediction mode (motion compensation) as the prediction mode. (No mode) is set, and the switch 53d of the arithmetic unit 53 is switched to the contact a side. As a result, the image data of the I picture is input to the DCT mode switching circuit 55.
[0064] As shown in FIG. 10 or 11, the DCT mode switching circuit 55 uses data of four luminance blocks in a state in which odd-numbered field lines and even-numbered field lines are mixed (frame DCT mode), or It is output to the DCT circuit 56 in either the separated state (field DCT mode).
That is, the DCT mode switching circuit 55 compares the coding efficiency when the odd-numbered field and the even-numbered field data are mixed and DCT processed with the coding efficiency when the DCT processing is performed in the separated state. Select a mode with good coding efficiency.
[0066] For example, as shown in FIG. 10, the input signal has a configuration in which odd-numbered field and even-numbered field lines are mixed, and the difference between the signal of the odd-numbered field line adjacent to the top and bottom and the signal of the even-numbered field line. Is calculated, and the sum of the absolute values (or the sum of squares) is obtained.
Further, as shown in FIG. 11, the input signal has a configuration in which the lines of the odd field and the even field are separated, and the difference between the signals of the lines of the odd field adjacent to the top and bottom and the line of the even field Calculate the difference between the signals and find the sum (or square sum) of the absolute values of each.
[0068] Further, the two (absolute value sum) are compared, and the DCT mode corresponding to the smaller value is set. That is, if the former is smaller, the frame DCT mode is set, and if the latter is smaller, the field DCT mode is set.
[0069] Then, the data of the configuration corresponding to the selected DCT mode is output to the DCT circuit 56, and the DCT flag indicating the selected DCT mode is output to the variable length coding circuit 58 and the motion compensation circuit 64.
[0070] As is clear from a comparison between the prediction mode (FIGS. 8 and 9) in the prediction mode switching circuit 52 and the DCT mode (FIGS. 10 and 11) in the DCT mode switching circuit 55, the luminance block is described. The data structure in each mode of both is substantially the same.
When the frame prediction mode (mode in which odd-numbered lines and even-numbered lines are mixed) is selected in the prediction mode switching circuit 52, the frame DCT mode (odd-numbered lines and even-numbered lines are mixed) is also selected in the DCT mode switching circuit 55. Mode) is likely to be selected, and if the field prediction mode (mode in which the odd and even field data is separated) is selected in the prediction mode switching circuit 52, the field in the DCT mode switching circuit 55. DCT mode (mode in which odd and even field data is separated) is likely to be selected.
However, the mode is not always selected in this way, and in the prediction mode switching circuit 52, the mode is determined so that the sum of the absolute values of the prediction errors becomes small, and in the DCT mode switching circuit 55, the mode is determined. , The mode is determined so that the coding efficiency is good.
[0073] The image data of the I picture output from the DCT mode switching circuit 55 is input to the DCT circuit 56, DCT processed, and converted into a DCT coefficient. This DCT coefficient is input to the quantization circuit 57, quantized on a quantization scale corresponding to the data storage amount (buffer storage amount) of the transmission buffer 59, and then input to the variable length coding circuit 58.
[0074] The variable-length coding circuit 58 corresponds to the quantization scale (scale) supplied from the quantization circuit 57, and the image data supplied from the quantization circuit 57 (in this case, the data of the I picture). Is converted into a variable length code such as a Huffman code and output to the transmission buffer 59.
[0075] The variable-length coding circuit 58 is also set to a quantization scale (scale) from the quantization circuit 57 and a prediction mode (in-image prediction, forward prediction, backward prediction, or bidirectional prediction) from the prediction determination circuit 54. The motion vector detection circuit 50 indicates the motion vector, the prediction mode switching circuit 52 indicates the prediction flag (the flag indicating whether the frame prediction mode or the field prediction mode is set), and the DCT mode switching circuit 55. The output DCT flag (a flag indicating whether the frame DCT mode or the field DCT mode is set) is input, and these are also variable-length encoded.
[0076] The transmission buffer 59 temporarily stores the input data, and outputs the data corresponding to the accumulated amount to the quantization circuit 57. When the remaining amount of data of the transmission buffer 59 is increased to the allowable upper limit value, the data amount of the quantization data is reduced by increasing the quantization scale of the quantization circuit 57 by the quantization control signal. On the contrary, when the remaining amount of data decreases to the allowable lower limit value, the transmission buffer 59 reduces the quantization scale of the quantization circuit 57 by the quantization control signal, thereby reducing the amount of data of the quantization data. Increase. In this way, overflow or underflow of transmit buffer 59 is prevented.
Then, the data stored in the transmission buffer 59 is read out at a predetermined timing, output to the transmission line, and recorded on the recording medium 3 via, for example, the recording circuit 19.
On the other hand, the I-picture data output from the quantization circuit 57 is input to the inverse quantization circuit 60 and inversely quantized according to the quantization scale supplied from the quantization circuit 57. The output of the inverse quantization circuit 60 is input to the IDCT (inverse discrete cosine transform) circuit 61, processed for inverse discrete cosine transform, and then supplied to and stored in the forward prediction image unit 63a of the frame memory 63 via the arithmetic unit 62. Will be done.
[0079] When the motion vector detection circuit 50 processes the image data of each frame sequentially input as, for example, a picture of I, B, P, B, P, B ..., the motion vector detection circuit 50 is input first. After processing the image data of the frame as an I picture, before processing the image of the next input frame as a B picture, the image data of the next input frame is processed as a P picture. This is because the B picture involves backward prediction, and therefore cannot be decoded unless the P picture as the backward prediction image is prepared in advance.
[0080] Therefore, the motion vector detection circuit 50 starts processing the image data of the P picture stored in the rear original image unit 51c after the processing of the I picture. Then, as in the case described above, the absolute value sum of the inter-frame difference (prediction error) in macroblock units is supplied from the motion vector detection circuit 50 to the prediction mode switching circuit 52 and the prediction determination circuit 54. The prediction mode switching circuit 52 and the prediction judgment circuit 54 correspond to the sum of the absolute values of the prediction errors of the macroblock of this P picture in the frame / field prediction mode, or in-image prediction, forward prediction, backward prediction, or bidirectional prediction. Set the prediction mode of.
[0081] When the in-image prediction mode is set, the arithmetic unit 53 switches the switch 53d to the contact a side as described above. Therefore, this data is transmitted to the transmission line via the DCT mode switching circuit 55, the DCT circuit 56, the quantization circuit 57, the variable length coding circuit 58, and the transmission buffer 59, similarly to the data of the I picture. Further, this data is supplied to and stored in the rear prediction image unit 63b of the frame memory 63 via the inverse quantization circuit 60, the IDCT circuit 61, and the arithmetic unit 62.
Further, when the forward prediction mode is set, the switch 53d is switched to the contact b, and the image (in this case, the image of the I picture) data stored in the forward prediction image unit 63a of the frame memory 63. Is read out, and the motion compensation circuit 64 compensates for the motion corresponding to the motion vector output by the motion vector detection circuit 50. That is, when the motion compensation circuit 64 is instructed by the prediction determination circuit 54 to set the forward prediction mode, the motion vector detection circuit 50 is currently outputting the read address of the forward prediction image unit 63a of the macroblock. Data is read out by shifting from the position corresponding to the position by the amount corresponding to the motion vector, and the predicted image data is generated.
[0083] The predicted image data output from the motion compensation circuit 64 is supplied to the arithmetic unit 53a. The arithmetic unit 53a subtracts the predicted image data corresponding to this macroblock supplied from the motion compensation circuit 65 from the data of the macro block of the reference image supplied from the prediction mode switching circuit 52, and the difference (prediction error). ) Is output. This difference data is transmitted to the transmission line via the DCT mode switching circuit 55, the DCT circuit 56, the quantization circuit 57, the variable length coding circuit 58, and the transmission buffer 59. Further, this difference data is locally decoded by the inverse quantization circuit 60 and the IDCT circuit 61, and input to the arithmetic unit 62.
[0084] The arithmetic unit 62 is also supplied with the same data as the predicted image data supplied to the arithmetic unit 53a. The arithmetic unit 62 adds the predicted image data output by the motion compensation circuit 64 to the difference data output by the IDCT circuit 61. As a result, the image data of the original (decoded) P picture is obtained. The image data of this P picture is supplied to and stored in the rear prediction image unit 63b of the frame memory 63.
[0085] In this way, the motion vector detection circuit 50 executes the processing of the B picture after the data of the I picture and the P picture are stored in the front prediction image unit 63a and the rear prediction image unit 63b, respectively. The prediction mode switching circuit 52 and the prediction judgment circuit 54 set the frame / field mode according to the magnitude of the absolute sum of the differences between frames in macroblock units, and set the prediction mode to the in-image prediction mode. Set to either forward prediction mode, backward prediction mode, or bidirectional prediction mode.
[0086] As described above, in the in-image prediction mode or the forward prediction mode, the switch 53d is switched to the contact a or b. At this time, the same processing as in the case of the P picture is performed, and the data is transmitted.
[0087] On the other hand, when the backward prediction mode or the bidirectional prediction mode is set, the switch 53d is switched to the contact points c or d, respectively.
[0088] In the backward prediction mode in which the switch 53d is switched to the contact c, the image (in this case, the image of the P picture) data stored in the backward prediction image unit 63b is read out, and the motion compensation circuit 64 As a result, motion compensation is performed corresponding to the motion vector output by the motion vector detection circuit 50. That is, when the motion compensation circuit 64 is instructed by the prediction determination circuit 54 to set the backward prediction mode, the motion vector detection circuit 50 currently outputs the read address of the backward prediction image unit 63b of the macroblock. Data is read out by shifting from the position corresponding to the position by the amount corresponding to the motion vector, and the predicted image data is generated.
The predicted image data output from the motion compensation circuit 64 is supplied to the arithmetic unit 53b. The arithmetic unit 53b subtracts the predicted image data supplied from the motion compensation circuit 64 from the macroblock data of the reference image supplied from the prediction mode switching circuit 52, and outputs the difference. This difference data is transmitted to the transmission line via the DCT mode switching circuit 55, the DCT circuit 56, the quantization circuit 57, the variable length coding circuit 58, and the transmission buffer 59.
[0090] In the bidirectional prediction mode in which the switch 53d is switched to the contact d, the image (in this case, the image of the I picture) data stored in the front prediction image unit 63a and the rear prediction image unit 63b are stored. The image data (in this case, the image of the P picture) is read out, and the motion compensation circuit 64 compensates for the motion corresponding to the motion vector output by the motion vector detection circuit 50.
That is, in the motion compensation circuit 64, when the prediction determination circuit 54 commands the setting of the bidirectional prediction mode, the motion vector detection circuit 50 now sets the read addresses of the front prediction image unit 63a and the rear prediction image unit 63b. Data is read out by shifting from the position corresponding to the position of the output macroblock by the amount corresponding to the motion vector (in this case, there are two motion vectors, one for the forward predicted image and the other for the backward predicted image), and the predicted image. Generate data.
The predicted image data output from the motion compensation circuit 64 is supplied to the arithmetic unit 53c. The arithmetic unit 53c subtracts the average value of the predicted image data supplied from the motion compensation circuit 64 from the macroblock data of the reference image supplied by the motion vector detection circuit 50, and outputs the difference. This difference data is transmitted to the transmission line via the DCT mode switching circuit 55, the DCT circuit 56, the quantization circuit 57, the variable length coding circuit 58, and the transmission buffer 59.
Since the image of the B picture is not regarded as a predicted image of another image, it is not stored in the frame memory 63.
[0094] In the frame memory 63, the front prediction image unit 63a and the rear prediction image unit 63b are bank-switched as necessary and are stored in one or the other with respect to a predetermined reference image. Can be switched and output as a forward prediction image or a backward prediction image.
[0095] In the above description, the luminance block has been mainly described, but the color difference block is similarly processed and transmitted in units of the macroblocks shown in FIGS. 8 to 11. As the motion vector when processing the color difference block, the motion vector of the corresponding luminance block is halved in the vertical direction and the horizontal direction, respectively.
[0096] FIG. 12 is a block diagram showing a configuration of the decoder 31 of FIG. The encoded image data transmitted via the transmission line (recording medium 3) is received by a reception circuit (not shown), reproduced by a playback device, temporarily stored in the reception buffer 81, and then decrypted by the decoding circuit 90. It is supplied to the variable length decoding circuit 82 of. The variable length decoding circuit 82 performs variable length decoding of the data supplied from the receive buffer 81, outputs the motion vector, the prediction mode, the prediction flag, and the DCT flag to the motion compensation circuit 87, and reverse-quantizes the quantization scale. In addition to outputting to 83, the decoded image data is output to the inverse quantization circuit 83.
[0097] The inverse quantization circuit 83 dequantizes the image data supplied from the variable length decoding circuit 82 according to the quantization scale also supplied from the variable length decoding circuit 82, and outputs the image data to the IDCT circuit 84. The data (DCT coefficient) output from the inverse quantization circuit 83 is subjected to the inverse discrete cosine transform process by the IDCT circuit 84 and supplied to the arithmetic unit 85.
[0098] When the image data supplied from the IDCT circuit 84 to the arithmetic unit 85 is I-picture data, the data is output from the arithmetic unit 85 and is input to the arithmetic unit 85 later (P or B). It is supplied to and stored in the forward prediction image unit 86a of the frame memory 86 for generating the prediction image data (picture data). Further, this data is output to the format conversion circuit 32 (FIG. 5).
[0099] When the image data supplied from the IDCT circuit 84 is P-picture data whose predictive image data is the image data one frame before the IDCT circuit 84 and is the forward prediction mode data, the forward prediction image of the frame memory 86 The image data (I picture data) one frame before, which is stored in the part 86a, is read out, and the motion compensation circuit 87 applies motion compensation corresponding to the motion vector output from the variable length decoding circuit 82. .. Then, in the arithmetic unit 85, the image data (difference data) supplied from the IDCT circuit 84 is added and output. This added data, that is, the decoded P-picture data is behind the frame memory 86 for predictive image data generation of the image data (B-picture or P-picture data) that is later input to the arithmetic unit 85. It is supplied to the prediction image unit 86b and stored.
[0100] Even if it is the data of the P picture, the data of the in-image prediction mode is not processed by the arithmetic unit 85 like the data of the I picture, and is stored as it is in the backward prediction image unit 86b.
Since this P picture is an image to be displayed next to the next B picture, it is not yet output to the format conversion circuit 32 at this point (as described above, the P input after the B picture). The picture is processed and transmitted before the B picture).
[0102] When the image data supplied from the IDCT circuit 84 is B picture data, it is stored in the forward prediction image unit 86a of the frame memory 86 corresponding to the prediction mode supplied from the variable length decoding circuit 82. Image data of the I picture (in the case of forward prediction mode), image data of the P picture stored in the backward prediction image unit 86b (in the case of backward prediction mode), or both image data (in the bidirectional prediction mode) Case) is read, and in the motion compensation circuit 87, motion compensation corresponding to the motion vector output from the variable length decoding circuit 82 is applied, and a predicted image is generated. However, if motion compensation is not required (in-image prediction mode), no prediction image is generated.
[0103] In this way, the data for which motion compensation has been applied by the motion compensation circuit 87 is added to the output of the IDCT circuit 84 in the arithmetic unit 85. This addition output is output to the format conversion circuit 32.
However, since this additional output is B picture data and is not used for generating a predicted image of another image, it is not stored in the frame memory 86.
[0105] After the image of the B picture is output, the image data of the P picture stored in the rear prediction image unit 86b is read out and supplied to the arithmetic unit 85 via the motion compensation circuit 87. However, at this time, motion compensation is not performed.
Although the decoder 31 does not show circuits corresponding to the prediction mode switching circuit 52 and the DCT mode switching circuit 55 in the encoder 18 of FIG. 5, processing corresponding to these circuits, that is, an odd number. The process of returning the configuration in which the signals of the field and the even field lines are separated to the original configuration as needed is executed by the motion compensation circuit 87.
[0107] Further, in the above description, the processing of the luminance signal has been described, but the processing of the color difference signal is also performed in the same manner. However, as the motion vector in this case, the motion vector for the luminance signal is halved in the vertical direction and the horizontal direction.
FIG. 13 shows the quality of the encoded image. Image quality (SNR: Signal to Noise Ratio) is controlled according to the picture type, I picture and P picture are considered to be high quality, and B picture is considered to be inferior to I and P picture. Be transmitted. This is a method that utilizes human visual characteristics, and the visual image quality is better when the quality is vibrated than when all the image qualities are averaged. The image quality control corresponding to this picture type is executed by the quantization circuit 57 of FIG.
[0109] FIGS. 14 and 15 show the configuration of the transcoder 101 to which the present invention is applied, and FIG. 15 shows a more detailed configuration of FIG. 14. The transcoder 101 converts the GOP structure and bit rate of the encoded video bit stream input to the decoding device 102 into the GOP structure and bit rate desired by the operator. In order to explain the function of the transcoder 101, although not shown in FIG. 15, three transcoders having almost the same functions as the transcoder 101 are connected to the front stage of the transcoder 101. It is assumed that there is. That is, in order to change the GOP structure and bit rate of the bitstream in various ways, the first transcoder, the second transcoder, and the third transcoder are connected in series in order, and the third transcoder It is assumed that the fourth transcoder shown in FIG. 15 is connected to the back.
[0110] In the following description of the present invention, the coding process performed in this first transcoder is defined as the first generation coding process, and the second transcoder is connected behind the first transcoder. The coding process performed in the transcoder is defined as the second generation coding process, and the coding process performed in the third transcoder connected behind the second transcoder is defined as the third generation code. The coding process performed in the fourth transcoder (transcoder 101 shown in FIG. 15), which is defined as the conversion process and is connected behind the third transcoder, is the fourth generation coding process or the current coding process. We will define it as coding processing.
[0111] Further, the coding parameter generated in the first generation coding process is referred to as the first generation coding parameter, and the coding parameter generated in the second generation coding process is referred to as the second generation code. The coding parameters generated in the 3rd generation coding process are called the 3rd generation coding parameters, and the coding parameters generated in the 4th generation coding process are called the 4th generation coding. We will call it the coding parameter or the current coding parameter.
First, the coded video stream ST (3rd) supplied to the transcoder 101 shown in FIG. 15 will be described. ST (3rd) indicates that it is a 3rd generation coded stream generated in the 3rd generation coding process in the 3rd transcoder provided in front of the transcoder 101. In the coded video stream ST (3rd) generated in this 3rd generation coding process, the 3rd generation coding parameters generated in the 3rd generation coding process are included in this coded coded video stream ST (3rd). (3rd) sequence layer, GOP layer, picture layer, slice layer, and macro block layer, sequence_header () function, sequence_extension () function, group_of_pictures_header () function, picture_header () function, picture_coding_extension () function, picture_data () Function, slice () It is described as a function and a macroblock () function. It is defined in the MPEG2 standard that the third coding parameter used in the third coding process is described in the third coded stream generated by the third coding process in this way. There is nothing new.
[0113] The unique point of the transcoder 101 of the present invention is that not only the third coding parameter is described in the third coding stream ST (3rd), but also the first generation and the second generation. The point is that the first-generation and second-generation coding parameters generated in the coding process of are also described.
Specifically, the first-generation and second-generation coding parameters are described as a history stream history_stream () in the user data area of the picture layer of the third-generation coded video stream ST (3rd). Has been done. In the present invention, the history stream described in the user data area of the picture layer of the 3rd generation encoded video stream ST (3rd) is referred to as "history information" or "history information", and is referred to as this history stream. The described coding parameters are called "history parameters" or "history parameters".
[0115] As another name, when the third-generation coding parameter described in the third-generation coding stream ST (3rd) is called the "current coding parameter", the third generation is called. Since the 1st and 2nd generation coding processes are the coding processes performed in the past from the viewpoint of the generation coding process, the user data area of the picture layer of the 3rd generation coded stream ST (3rd). The coding parameter described as the history stream described in is also called "past coding parameter".
[0116] As described above, not only the third coding parameter is described in the third coding stream ST (3rd), but also the first and second generation coding processes are generated. The reason for describing the first-generation and second-generation coding parameters is that image quality deterioration can be prevented even if the GOP structure and bit rate of the coded stream are repeatedly changed by the transcoding process.
[0117] For example, a picture is encoded as a P picture in the first generation coding process, and the picture is B in the second generation coding process in order to change the GOP structure of the first generation coded stream. In order to encode as a picture and further change the GOP structure of the second generation coded stream, it is conceivable to code the picture again as a P picture in the third generation coding process. Since the coding process and the decoding process based on the MPEG standard are not 100% reversible processes, it is known that the image quality deteriorates every time the coding and decoding processes are repeated.
[0118] In such a case, in the third generation coding process, instead of recalculating the coding parameters such as the quantization scale, the motion vector, and the prediction mode, in the first generation coding process. Reuse the generated coding parameters such as quantization scale, motion vector, and prediction mode. Quantization scale, motion vector, prediction mode newly generated by 1st generation coding process rather than coding parameters such as quantization scale, motion vector, prediction mode newly generated by 3rd generation coding process Since the coding parameters such as the above are clearly more accurate, the image quality deterioration can be reduced even if the coding and decoding processes are repeated by reusing the first-generation parameters.
[0119] In order to explain the above-described processing according to the present invention, the processing of the 4th generation transcoder 101 shown in FIG. 15 will be described in more detail by taking as an example.
[0120] The decoding device 102 decodes the coded video contained in the third-generation coded bitstream ST (3rd) using the third-generation coded parameters, and decodes the baseband digital. It is a device for generating video data. Further, the decoding device 102 is for decoding the first-generation and second-generation coding parameters described as the history stream in the user data area of the picture layer of the third-generation coded bitstream ST (3rd). It is also a device.
Specifically, as shown in FIG. 16, the decoding device 102 has basically the same configuration as the decoder 31 (FIG. 12) of the decoding device 2 of FIG. 5, and is supplied with bits. Receive buffer 81 for buffering the stream, variable length decoding circuit 112 for variable length decoding of the coded bit stream, inverse quantum according to the quantization scale supplied from the variable length decoding circuit 112 for the variable length decoded data. It is equipped with an inverse quantized circuit 83 to be converted, an IDCT circuit 84 to perform inverse discrete cosine transform of the inverse quantized DCT coefficient, an arithmetic unit 85 for performing motion compensation processing, a frame memory 86, and a motion compensation circuit 87.
[0122] The variable-length decoding circuit 112 decodes the third-generation coded bitstream ST (3rd) by decoding the picture layer, slice layer, and macro of the third-generation coded bitstream ST (3rd). Extract the 3rd generation coding parameters described in the block layer. For example, the third-generation coding parameters extracted in this variable-length decoding circuit 112 are picture_coding_type indicating the picture type, quantizer_scale_code indicating the quantization scale step size, macroblock_type indicating the prediction mode, motion_vector indicating the motion vector, and Frame prediction. Frame / field_motion_type indicating mode or field prediction mode, and Frame DCT mode or Field It is dct_type etc. indicating whether it is DCT mode. The quatntiser_scale_code extracted in the variable length decoding circuit 112 is supplied to the inverse quantization circuit 83, and parameters such as picture_coding_type, quatntiser_scale_code, macroblock_type, motion_vector, frame / field_motion_type, and dct_type are supplied to the motion compensation circuit 87.
[0123] The variable-length decoding circuit 112 has not only these coding parameters required for decoding the third-generation coded bitstream ST (3rd), but also a third-generation transcoder in the subsequent stage. The coding parameters to be transmitted as generation history information are extracted from the sequence layer, GOP layer, picture layer, slice layer, and macroblock layer of the third generation coded bitstream ST (3rd). Of course, the 3rd generation coding parameters such as picture_coding_type, quatntiser_scale_code, macroblock_type, motion_vector, frame / field_motion_type, and dct_type used in the 3rd generation decoding process are included in this 3rd generation history information. What kind of coding parameter is extracted as history information is preset by the operator or the host computer according to the transmission capacity or the like.
[0124] Further, the variable length decoding circuit 112 extracts the user data described in the user data area of the picture layer of the third generation coded bit stream ST (3rd), and extracts the user data into the history decoding device. Supply to 104.
[0125] This history decoding device 104 uses the user data described in the picture layer of the third generation coded bitstream ST (3rd) to describe the first generation coding parameters and history information. This is a circuit for extracting the second-generation coding parameters (coding parameters of generations earlier than the immediately preceding generation). Specifically, the history decoding device 104 detects the unique History_Data_Id described in the user data by analyzing the syntax of the received user data, and thereby extracts the converted_history_stream (). be able to. Further, the history decoding device 104 obtains history_stream () by taking 1-bit marker bits (marker_bit) inserted at predetermined intervals in converted_history_stream (), and the history_stream () is thinned. By analyzing the tax, the 1st and 2nd generation coding parameters described in history_stream () can be obtained. The detailed operation of this history decoding device 104 will be described later.
The history information multiplexing device 103 supplies the first-generation, second-generation, and third-generation coding parameters to the decoding device 106 that performs the fourth-generation coding process, so that the decoding device 102 It is a circuit for multiplexing the encoding parameters of the first generation, the second generation, and the third generation on the video data of the base band decoded in. Specifically, the history information multiplexing device 103 includes baseband video data output from the arithmetic unit 85 of the decoding device 102, and a third-generation coding parameter output from the variable-length decoding device 112 of the decoding device 102. , And the 1st generation coding parameter and the 2nd generation coding parameter output from the history decoding device 104 are received, and these 1st generation, 2nd generation and 1st generation are added to the video data of this baseband. Multiplex 3 generations of coding parameters. The baseband video data in which the first-generation, second-generation, and third-generation coding parameters are multiplexed is supplied to the history information separator 105 via a transmission cable.
Next, a method of multiplexing these first-generation, second-generation, and third-generation coding parameters into baseband video data will be described with reference to FIGS. 17 and 18. , Figure 17 shows one macroblock consisting of 16 pixels x 16 pixels as defined in the MPEG standard. This 16-pixel x 16-pixel macroblock consists of four 8-pixel x 8-pixel subblocks (Y [0], [1], [2] and Y [3]) for luminance signals and color difference signals. Consists of four 8 pixel x 8 pixel subblocks (Cr [0], r [1], b [0], and Cb [1]).
FIG. 18 represents a format with video data. This format is defined in ITU Recommendation-RDT601 and represents the so-called "D1 format" used in the broadcasting industry. This D1 format has been standardized as a format for transmitting 10-bit video data, so that one pixel of video data can be represented by 10 bits.
Since the baseband video data decoded by the MPEG standard is 8 bits, in the transcoder of the present invention, as shown in FIG. 18, the upper 8 bits (D9 to D9) of the 10 bits of the D1 format are shown. D2) is used to transmit baseband video data decoded according to the MPEG standard. When the decoded 8-bit video data is written to the D1 format in this way, the lower 2 bits (D1 and D0) become unallocated bits. In the transcoder of the present invention, history information is transmitted by using this unallocated area.
The data blocks described in FIG. 18 are subblocks (Y [0], Y [1], Y [2], Y [3], Cr [0], Cr [1], Cb [ Since it is a data block for transmitting 1 pixel in 0], Cb [1]), 64 data blocks shown in FIG. 18 are transmitted in order to transmit 1 macroblock of data. .. By using the lower 2 bits (D1 and D0), a total of 1024 (= 16 × 64) bits of history information can be transmitted to one macroblock of video data. Therefore, since the history information for one generation is generated so as to be 256 bits, the history information for the past 4 (= 1024/256) generations can be superimposed on the video data of one macroblock. it can. In the example shown in FIG. 18, the first-generation history information, the second-generation history information, and the third-generation history information are superimposed.
[0131] The history information separator 105 is a circuit for extracting baseband video data from the upper 8 bits of data transmitted as the D1 format and extracting history information from the lower 2 bits. In the example shown in FIG. 15, the history information separator 105 extracts the baseband video data from the transmitted data, supplies the video data to the coding device 106, and supplies the first generation and second generation data from the transmitted data. The history information of the generation and the third generation is extracted and supplied to the encoding device 106 and the history encoding device 107, respectively.
[0132] The coding device 106 encodes the baseband video data supplied from the history information separating device 105 into a bit stream having a GOP structure and a bit rate specified by the operator or the host computer. It is a device of. Note that changing the GOP structure means, for example, the number of pictures included in the GOP, the number of P pictures existing between I pictures and I pictures, and the number of P pictures existing between I pictures and P pictures (or I pictures). B Means to change the number of pictures.
In the example shown in FIG. 15, the supplied baseband video data is superposed with first-generation, second-generation, and third-generation history information, so that the encoding device 106 is , The 4th generation coding process is performed by selectively reusing these history information so that the image quality deterioration due to the recoding process is reduced.
[0134] FIG. 19 is a diagram showing a specific configuration of the encoder 121 provided in the coding device 106. This encoder 121 is basically configured in the same manner as the encoder 18 shown in FIG. 7, and has a motion vector detection circuit 50, a frame / field prediction mode switching circuit 52, an arithmetic unit 53, a DCT mode switching circuit 55, and a DCT circuit. It is equipped with 56, a quantization circuit 57, a variable length coding circuit 58, a transmission buffer 59, an inverse quantization circuit 60, an inverse DCT circuit 61, an arithmetic unit 62, a frame memory 63, and a motion compensation circuit 64. Since the functions of these circuits are almost the same as the functions of the encoder 18 described with reference to FIG. 7, the description thereof will be omitted. Hereinafter, the differences between the encoder 121 and the encoder 18 described with reference to FIG. 7 will be mainly described.
[0135] This encoder 121 has a controller 70 for controlling the operation and function of each of the circuits described above. The controller 70 receives instructions regarding the GOP structure from the operator or host computer and determines the picture type of each picture corresponding to the GOP structure. In addition, the controller 70 receives the target bit rate information from the operator or the host computer, and controls the quantization circuit 57 so that the bit rate output from the encoder 121 becomes the specified target bit rate. To do.
[0136] Further, the controller 70 receives the history information of a plurality of generations output from the history information separation device 105, and reuses the history information to perform the coding process of the reference picture. This will be described in detail below.
First, the controller 70 determines whether or not the picture type of the reference picture determined from the GOP structure specified by the operator matches the picture type included in the history information. That is, it is determined whether or not this reference picture has been encoded in the past with the same picture type as the specified picture type.
[0138] To explain more clearly with the example shown in FIG. 15, in this controller 70, the picture type assigned to this reference picture as the fourth generation coding process is the first generation. Whether it matches the picture type of this reference picture in the coding process, the picture type of this reference picture in the 2nd generation coding process, or the picture type of this reference picture in the 3rd generation coding process. to decide.
[0139] If the picture type specified for this reference picture as the 4th generation coding process does not match any picture type in the past coding process, the controller 70 may "normally code". "I do. That is, in this case, in any of the 1st generation, 2nd generation, or 3rd generation coding processing, this reference picture is the coding processing with the picture type assigned as the 4th generation coding processing. It means that it has never been done. On the other hand, if the picture type specified for this reference picture as the 4th generation coding process matches any of the picture types in the past coding process, the controller 70 will perform "parameter reuse". Encoding process "is performed. That is, in this case, in the coding process of either the 1st generation, the 2nd generation, or the 3rd generation, the reference picture is encoded with the picture type assigned as the 4th generation coding process. It means that it has been processed.
[0140] First, the normal coding process of the controller 70 will be described.
[0141] The motion vector detection circuit 50 detects a prediction error in the frame prediction mode and a prediction error in the field prediction mode, respectively, in order to determine whether the frame prediction mode or the field prediction mode should be selected, and predicts the prediction error. The error value is supplied to the controller 70. The controller 70 compares the values of those prediction errors and selects the prediction mode having the smaller prediction error value. The prediction mode switching circuit 52 performs signal processing so as to correspond to the prediction mode selected by the controller 70, and supplies it to the arithmetic unit 53.
[0142] Specifically, when the frame prediction mode is selected, the prediction mode switching circuit 52 is an arithmetic unit with respect to the luminance signal as it is input, as described with reference to FIG. The signal is processed so as to be output to 53, and the color difference signal is signal-processed so that the odd field line and the even field line are mixed. On the other hand, when the field prediction mode is selected, as described with reference to FIG. 9, for the luminance signal, the luminance blocks Y [1] and Y [2] are composed of odd field lines, and the luminance block is formed. Signal processing Y [3] and Y [4] so that they are composed of even field lines, and for color difference signals, the upper 4 lines are composed of odd field lines and the lower 4 lines are composed of even field lines. Signal processing.
[0143] Further, in each prediction mode, the motion vector detection circuit 50 determines which of the in-image prediction mode, the forward prediction mode, the backward prediction mode, and the bidirectional prediction mode is selected. A prediction error is generated, and the prediction error in each prediction mode is supplied to the controller 70, respectively. The controller 70 selects the smallest prediction error of the forward prediction, the backward prediction, and the bidirectional prediction as the prediction error of the inter prediction. Further, the prediction error of the inter-prediction is compared with the prediction error of the in-image prediction, the smaller one is selected, and the mode corresponding to the selected prediction error is selected as the prediction mode. That is, if the prediction error of the in-image prediction is smaller, the in-image prediction mode is set. If the prediction error of the inter-prediction is smaller, the mode with the smallest corresponding prediction error of the forward prediction, the backward prediction, or the bidirectional prediction mode is set. The controller 70 controls the arithmetic unit 53 and the motion compensation circuit 64 so as to correspond to the selected prediction mode.
[0144] In order to select either the frame DCT mode or the field DCT mode, the DCT mode switching circuit 55 displays data of four brightness blocks in a signal form in which odd-numbered field lines and even-numbered field lines are mixed. The frame DCT mode) is converted, and the odd field line and the even field line are converted into a separated signal form (field DCT mode), and each signal is supplied to the DCT circuit 56. The DCT circuit 56 calculates the coding efficiency when the odd-numbered field and the even-numbered field are mixed and DCT-processed, and the coding efficiency when the odd-numbered field and the even-numbered field are separated and DCT-processed, and the result is a controller. Supply to 70. The controller 70 compares the coding efficiencies supplied from the DCT circuit 56, selects the DCT mode having the better coding efficiency, and controls the DCT mode switching circuit 55 so as to be in the selected DCT mode. ..
The controller 70 receives a target bit rate indicating a target bit rate supplied from the operator or the host computer and a signal indicating the amount of bits buffered in the transmission buffer 59, that is, a signal indicating the remaining amount of the buffer. , Generates feedback_q_scale_code to control the quantization step size of the quantization circuit 57 based on this target bit rate and the remaining buffer capacity. This feedback_q_scale_code is a control signal generated according to the remaining buffer level of the transmit buffer 59 so that the transmit buffer 59 does not overflow or underflow, and is a bit of the bit stream output from the transmit buffer 59. It is also a signal that controls the rate to be the target bit rate.
[0146] Specifically, for example, when the amount of bits buffered in the transmission buffer 59 becomes small, the quantization step size is increased so that the amount of generated bits of the picture to be encoded next increases. On the other hand, if the amount of bits buffered in the transmission buffer 59 becomes large, the quantization step size is increased so that the amount of generated bits of the next encoded picture is small. .. Note that feedback_q_scale_code and the quantization step size are proportional, and when feedback_q_scale_code is increased, the quantization step size is increased, and when feedback_q_scale_code is decreased, the quantization step size is decreased.
Next, the parameter reuse coding process, which is one of the features of the transcoder 101, will be described. To explain this process more clearly, the reference picture is encoded as a P-picture in the first-generation coding process, encoded as an I-picture in the second-generation coding process, and is a third-generation coding process. It is assumed that it was encoded as a B picture in the coding process, and that this reference picture must be encoded as a P picture in the 4th generation coding process this time.
[0148] In this case, the controller has the same picture type (I picture) as the picture type assigned as the 4th generation picture type, and this reference picture is encoded in the 1st generation coding process. The 70 does not create a new coding parameter from the supplied video data, but uses the first generation coding parameter to perform the coding process. Typical coding parameters to be reused in this fourth coding process are quantizer_scale_code indicating the quantization scale step size, macroblock_type indicating the prediction direction mode, motion_vector indicating the motion vector, and Frame prediction mode or Field. Frame / field_motion_type indicating whether it is the prediction mode, dct_type indicating whether it is the Frame DCT mode or the Field DCT mode, and the like.
[0149] The controller 70 does not reuse all the coding parameters transmitted as the history information, but reuses and reuses the coding parameters as described above, which are assumed to be desirable to be reused. Code parameters that should not be used are newly generated.
[0150] Next, the coding parameter reuse coding process will be described focusing on the differences from the above-described normal coding process.
[0151] The motion vector detection circuit 50 detects the motion vector of the reference picture in the above-mentioned normal coding process, but does not detect the motion vector motion_vector in this parameter reuse coding process. In addition, the motion vector motion_vector supplied as the history information of the first generation is reused. The reason will be explained.
[0152] Since the baseband video data obtained by decoding the third-generation coded stream has undergone decoding and coding processing at least three times, the image quality is clearly deteriorated as compared with the original video data. There is. Even if the motion vector is detected from the video data whose image quality is deteriorated, the accurate motion vector cannot be detected. That is, the motion vector supplied as the first-generation history information is clearly a more accurate motion vector than the motion vector detected in the fourth-generation coding process. That is, by reusing the motion vector transmitted as the first-generation coding parameter, the image quality does not deteriorate even if the fourth-generation coding process is performed. The controller 70 uses the motion vector motion_vector supplied as the 1st generation history information as the motion vector information of the reference picture encoded in the 4th generation coding process by the motion compensation circuit 64 and the variable length coding. Supply to circuit 58.
[0153] Further, the motion vector detection circuit 50 detects a prediction error in the frame prediction mode and a prediction error in the field prediction mode in order to determine whether the frame prediction mode or the field prediction mode is selected. In this parameter reuse coding process, the Frame prediction mode or Field prediction supplied as the first generation history information is not performed without detecting the prediction error in this frame prediction mode and the prediction error in the field prediction mode. Reuse frame / field_motion_type to indicate the mode. This is because the prediction error in each prediction mode detected in the 1st generation is more accurate than the prediction error in each prediction mode detected in the 4th generation coding process, so it is determined by the highly accurate prediction error. This is because more optimal coding processing can be performed by selecting the predicted prediction mode.
[0154] Specifically, the controller 70 supplies the control signal corresponding to the frame / field_motion_type supplied as the history information of the first generation to the prediction mode switching circuit 52, and the prediction mode switching circuit 52 is this. Performs signal processing corresponding to the reused frame / field_motion_type.
Further, in the normal coding process, the motion vector detection circuit 50 further predicts any of an in-image prediction mode, a forward prediction mode, a backward prediction mode, and a bidirectional prediction mode (hereinafter, this prediction mode). In order to determine whether to select the prediction direction mode), the prediction error in each prediction direction mode was calculated, but in this parameter reuse coding process, the prediction error in each prediction direction mode is calculated. No calculation is performed, and the prediction direction mode is determined based on the macroblock_type supplied as the first generation history information. This is because the prediction error in each prediction direction mode in the 1st generation coding process is more accurate than the prediction error in each prediction direction mode in the 4th generation coding process, so that the prediction error is more accurate. This is because more efficient coding processing can be performed by selecting the prediction direction mode determined by. Specifically, the controller 70 selects the prediction direction mode indicated by macroblock_type included in the first-generation history information, and the arithmetic unit 53 and the motion compensation circuit correspond to the selected prediction direction mode. Control 64.
[0156] In the normal coding process, the DCT mode switching circuit 55 uses the signal converted into the signal form of the frame DCT mode in order to compare the coding efficiency of the frame DCT mode with the coding efficiency of the field DCT mode. , Both the signal converted to the signal form of the field DCT mode was supplied to the DCT circuit 56, but in this parameter reuse coding process, the signal converted to the signal form of the frame DCT mode and the signal of the field DCT mode are supplied. The process of generating both the signals converted into forms is not performed, and only the process corresponding to the DCT mode indicated by dct_type included in the history information of the first generation is performed. Specifically, the controller 70 reuses the dct_type contained in the first-generation history information, and the DCT mode switching circuit 55 performs signal processing corresponding to the DCT mode indicated by this dct_type. Controls the mode switching circuit 55.
[0157] In the normal coding process, the controller 70 controls the quantization step size of the quantization circuit 57 based on the target bit rate and the remaining amount of the transmission buffer specified by the operator. However, this parameter reuse In the coding process, the quantization step size of the quantization circuit 57 is controlled based on the target bit rate, the remaining amount of the transmission buffer, and the past quantization scale included in the history information. In the following explanation, the past quantization scale included in the history information will be described as history_q_scale_code. Further, in the history stream described later, this quantization scale is described as quantizer_scale_code.
[0158] First, the controller 70 uses the current quantization scale feedback_q_scale_code as in the normal coding process. To generate. This feedback_q_scale_code is a value determined according to the remaining buffer amount of the transmission buffer 59 so that the transmission buffer 59 does not overflow or underflow. Then, the value of the past quantization scale history_q_scale_code contained in the 1st generation history stream is compared with the value of this current quantization scale feedback_q_scale_code to determine which quantization scale is larger. .. A large quantization scale means a large quantization step. If the current quantization scale feedback_q_scale_code is larger than the past quantization scale history_q_scale_code, the controller 70 supplies this current quantization scale feedback_q_scale_code to the quantization circuit 57. On the other hand, if the past quantization scale history_q_scale_code is larger than the current quantization scale feedback_q_scale_code, the controller 70 supplies this past quantization scale history_q_scale_code to the quantization circuit 57.
That is, the controller 70 has the largest quantization scale code among the plurality of past quantization scales contained in the history information and the current quantization scale calculated from the remaining amount of the transmission buffer. select. In other words, the controller 70 is used in the quantization step in the past (1st, 2nd and 3rd generation) coding process or in the present (4th generation) coding process. The quantization circuit 57 is controlled so as to perform the quantization using the largest quantization step among the quantization steps. The reason for this will be explained below.
[0160] For example, the bit rate of the stream generated in the third-generation coding process is 4 [Mbps], and the target bit rate set for the encoder 121 that performs the fourth-generation coding process. Is 15 [Mbps]. At this time, the target bit rate is increasing, so if we simply reduce the quantization step, that is not the case. Even if a picture encoded in a large quantization step in the past coding process is encoded in the current coding process by reducing the quantization step, the image quality of this picture is improved. There is no. That is, coding in a quantization step smaller than the quantization step in the past coding process merely increases the amount of bits and does not improve the image quality. Therefore, among the quantization steps in the past (1st, 2nd and 3rd generation) coding processes or the present (4th generation) coding processes, the largest quantization step is selected. The most efficient coding process can be performed by using it for quantization.
Next, the history decoding device 104 and the history encoding device 107 in FIG. 15 will be further described. As shown in the figure, the history decoding device 104 has a history from the output of the user data decoder 201 that decodes the user data supplied from the decoding device 102, the converter 202 that converts the output of the user data decoder 201, and the output of the converter 202. It is composed of history VLD203 that reproduces information.
Further, the history encoding device 107 uses the history VLC211 for formatting the coding parameters for three generations supplied from the history information separator 105, the converter 212 for converting the output of the history VLC211, and the output of the converter 212. It consists of a user data formatter 213 that formats the data into a format.
[0163] The user data decoder 201 decodes the user data supplied from the decoding device 102 and outputs the user data to the converter 202. The details will be described later with reference to FIG. 51, but the user data (user_data ()) consists of user_data_start_code and user_data. In the MPEG standard, user_data contains consecutive 23-bit "0" (the same code as start_code). ) Is prohibited. This is to prevent the data from being falsely detected as start_code. History information (history_stream ()) is described in the user data area (as a type of MPEG standard user_data), and there may be such continuous 23-bit or more "0" s in it. , It is necessary to insert "1" at a predetermined timing and convert it to converted_history_stream () (Fig. 38, which will be described later) so that continuous "0" of 23 bits or more does not occur. It is the converter 212 of the history encoding device 107 that performs this conversion. The converter 202 of the history decoding device 104 performs a conversion process opposite to that of the converter 212 (removes the inserted 1 so as not to generate 0 of 23 consecutive bits or more).
[0164] The history VLD 203 generates history information (in this case, a first-generation coding parameter and a second-generation coding parameter) from the output of the converter 202, and outputs the history information to the history information multiplexing device 103.
On the other hand, in the history encoding device 107, the history VLC211 sets the coding parameters (first generation, second generation, and third generation) for the three generations supplied from the history information separator 105 as the history information. Convert to format. This format includes a fixed length format (FIGS. 40 to 46 described later) and a variable length format (FIG. 47 described later). Details of these will be described later.
[0166] The history information formatted by the history VLC211 is converted into converted_history_stream () in the converter 212. This is a process to prevent false detection of start_code of user_data () as described above. That is, although there are continuous 23-bit or more "0" in the history information, continuous 23-bit or more "0" cannot be placed in user_data, so do not touch this prohibited item. Data is converted by the converter 212 (1 is inserted at a predetermined timing).
[0167] The user data formatter 213 is an MPEG standard that can be inserted into the video stream by adding History_Data_ID to converted_history_stream () supplied from the converter 212 and further adding user_data_stream_code based on FIG. 38 described later. Generate user_data and output it to the encoding device 106.
[0168] FIG. 20 shows a configuration example of the history VLC211. The code word converter 301 and the code length converter 305 have a coding parameter (a coding parameter transmitted as history information this time) (item data) and information for specifying a stream in which the coding parameter is arranged (for example,). , The name of the syntax (for example, the name of sequence_header described later)) (item NO.) Is supplied from the history information separator 105. The codeword converter 301 converts the input coding parameters into codewords corresponding to the instructed syntax and outputs them to the barrel shifter 302. The barrel shifter 302 shifts the codeword input from the codeword converter 301 by the amount corresponding to the shift amount supplied from the address generation circuit 306, and outputs the codeword as a codeword in byte units to the switch 303. The switch 303, which is switched by the bit select signal output from the address generation circuit 306, is provided for each bit, and supplies the codeword supplied from the barrel shifter 302 to the RAM 304 and stores it. The write address at this time is specified by the address generation circuit 306. When a read address is specified from the address generation circuit 306, the data (codeword) stored in the RAM 304 is read and supplied to the converter 212 in the subsequent stage, and if necessary, via the switch 303. Is supplied to RAM 304 again and stored.
[0169] The code length converter 305 determines the code length of the coded parameter from the input syntax and the coded parameter, and outputs the code length to the address generation circuit 306. The address generation circuit 306 generates the shift amount, bit select signal, write address, or read address described above according to the input code length, and supplies them to the barrel shifter 302, the switch 303, or the RAM 304, respectively. ..
As described above, the history VLC211 is configured as a so-called variable length encoder, and outputs the input coding parameter in variable length coding.
[0171] FIG. 21 shows a configuration example of the history VLD 203 that decodes the data formatted as described above. In this history VLD203, the data of the coding parameters supplied from the converter 202 is supplied to the RAM311 and stored. The write address at this time is supplied from the address generation circuit 315. The address generation circuit 315 also generates a read address at a predetermined timing and supplies it to the RAM 311. At this time, the RAM 311 reads the data stored in the read address and outputs the data to the barrel shifter 312. The barrel shifter 312 shifts the input data by the amount corresponding to the shift amount output by the address generation circuit 315, and outputs the data to the inverse code length converter 313 and the inverse code word converter 314.
[0172] The inverse code length converter 313 is also supplied with the syntax name (item No.) of the stream in which the coding parameters are arranged from the converter 202. The inverse code length converter 313 obtains the code length from the input data (code word) based on the syntax, and outputs the obtained code length to the address generation circuit 315.
Further, the inverse codeword converter 314 decodes the data supplied from the barrel shifter 312 based on the syntax (inverse codeword conversion), and outputs the data to the history information multiplexing device 103.
[0174] Further, the inverse codeword converter 314 extracts the information necessary for identifying what kind of codeword is included (information necessary for determining the codeword delimiter), and then extracts the information necessary for determining the codeword delimiter. Output to the address generation circuit 315. The address generation circuit 315 generates a write address and a read address based on this information and the code length input from the inverse code length converter 313, outputs the write address to the RAM311 and outputs the shift amount to the barrel shifter 312. To do.
FIG. 22 shows a configuration example of the converter 212. In this example, 8-bit data is read from the read address output by the controller 326 of the buffer memory 320 located between the history VLC211 and the converter 212, and the D-type flip-flop (D-FF) 321 is used. It is designed to be supplied and retained. Then, the data read from the D-type flip-flop 321 is supplied to the staff circuit 323 and also supplied to and held in the 8-bit D-type flip-flop 322. The 8-bit data read from the D-type flip-flop 322 is combined with the 8-bit data read from the D-type flip-flop 321 and supplied to the staff circuit 323 as 16-bit parallel data.
[0176] The stuff circuit 323 inserts (stuffs) the code "1" at the position of the signal (stuff position) indicating the stuff position supplied from the controller 326, and outputs the data as a total of 17 bits to the barrel shifter 324. ..
[0177] The barrel shifter 324 shifts the input data based on the signal (shift) indicating the shift amount supplied from the controller 326, extracts the 8-bit data, and puts it into the 8-bit D-type flip-flop 325. Output. The data held in the D-type flip-flop 325 is read from the data and supplied to the user data formatter 213 in the subsequent stage via the buffer memory 327. At this time, the controller 326 generates a write address together with the output data, and supplies the write address to the buffer memory 327 interposed between the converter 212 and the user data formatter 213.
FIG. 23 shows a configuration example of the staff circuit 323. The 16-bit data input from the D-type flip-flops 322 and 321 are input to the contacts a of switches 331-16 to 331-1, respectively. The data of the switch adjacent to the MSB side (upper in the figure) is supplied to the contact c of the switch 331-i (i = 0 to 15). For example, the contact c of the switch 331-12 is supplied with the 13th data from the LSB supplied to the contact a of the switch 331-13 adjacent to the MSB side, and the contact c of the switch 331-13 is supplied with the 13th data. , The 14th data from the LSB side supplied to the contact a of the switch 331-14 adjacent to the MSB side is supplied.
[0179] However, the contact a of the switch 331-0 further below the switch 331-1 corresponding to the LSB is open. In addition, the contact c of the switch 331-16 corresponding to the MSB is open because there is no switch higher than that.
[0180] Data "1" is supplied to the contact b of each switch 331-0 to 331-16.
[0181] The decoder 332 puts one of the switches 331-0 to 331-16 on the contact b side corresponding to the signal stuff position indicating the position where the data "1" supplied from the controller 326 is inserted. Switching, the switch on the LSB side is switched to the contact c side, and the switch on the MSB side is switched to the contact a side.
[0182] FIG. 23 shows an example in which the data 1 is inserted 13th from the LSB side. Therefore, in this case, the switches 331-0 to 331-12 are all switched to the contact c side, the switch 331-13 is switched to the contact b side, and the switches 331-14 to 331-16 are the contacts. It has been switched to the a side.
[0183] With the above configuration, the converter 212 of FIG. 22 converts the 22-bit code into 23 bits and outputs the code.
FIG. 24 shows the timing of the output data of each part of the converter 212 of FIG. When the controller 326 of the converter 212 generates a read address (FIG. 24 (A)) in synchronization with the clock in bytes, the corresponding data is read in bytes from the buffer memory 320, and a D-type flip-flop is used. It is temporarily held at 321. Then, the data (FIG. 24 (B)) read from the D-type flip-flop 321 is supplied to the staff circuit 323 and is supplied to and held in the D-type flip-flop 322. The data held in the D-type flip-flop 322 is further read from the data (FIG. 24 (C)) and supplied to the staff circuit 323.
Therefore, the input of the staff circuit 323 (FIG. 24 (D)) is the first 1-byte data D0 at the timing of the read address A1, and the 1-byte data D0 at the timing of the next read address A2. And 1 byte of data D1 becomes 2 bytes of data, and at the timing of read address A3, it becomes 2 bytes of data composed of data D1 and data D2.
[0186] The stuff position (FIG. 24 (E)) indicating the position where the data "1" is inserted is supplied to the stuff circuit 323 from the controller 326. The decoder 332 of the stuff circuit 323 switches the switch corresponding to this signal stuff position to the contact b among the switches 331-16 to 331-0, switches the switch on the LSB side to the contact c side, and further MSB. Switch the switch on the side to the contact a side. As a result, the data "1" is inserted, and the staff circuit 323 outputs the data (FIG. 24 (F)) in which the data "1" is inserted at the position indicated by the signal stuff position.
[0187] The barrel shifter 324 barrel-shifts the input data by the amount indicated by the signal shift (FIG. 24 (G)) supplied from the controller 326, and outputs the input data (FIG. 24 (H)). This output is further held by the D-type flip-flop 325 and then output to the subsequent stage (Fig. 24 (I)).
[0188] In the data output from the D-type flip-flop 325, the data "1" is inserted next to the 22-bit data. Therefore, between the data "1" and the next data "1", even if all the bits between them are 0, the number of consecutive 0 data is 22.
[0189] FIG. 25 shows a configuration example of the converter 202. The configuration of the converter 202 including the D-type flip-flops 341 to the controller 346 is basically the same as that of the D-type flip-flops 321 to the controller 326 of the converter 212 shown in FIG. Instead, the delay circuit 343 is inserted, which is different from the case of the converter 212. Other configurations are the same as in the case of the converter 212 of FIG.
That is, in this converter 202, according to the signal delete position indicating the position of the bit to be deleted output by the controller 346, the delay circuit 343 performs the bit (data 1 inserted in the staff circuit 323 of FIG. 22). ) Is deleted.
Other operations are the same as in the case of the converter 212 of FIG.
FIG. 26 shows a configuration example of the delay circuit 343. In this configuration example, of the 16-bit data input from the D-type flip-flops 342 and 341, 15 bits on the LSB side are supplied to the contacts a of the corresponding switches 351-0 to 351-14, respectively. Only one bit of data on the MSB side is supplied to the contact b of each switch. The decoder 352 deletes the bit specified by the signal delete position supplied from the controller 346 and outputs it as 15-bit data.
FIG. 26 shows a state in which the thirteenth bit from the LSB is delayed. Therefore, in this case, the switch 351-0 to the switch 351-11 are switched to the contact a side, and the 12 bits from the LSB to the 12th are selected and output as they are. Further, since the switches 351-12 to 351-14 are switched to the contact b side, the 14th to 16th data are selected and output as the data of the 13th to 15th bits. To.
The inputs of the staff circuit 323 of FIG. 23 and the delay circuit 343 of FIG. 26 are 16 bits because the inputs of the staff circuit 323 of the converter 212 of FIG. 22 are supplied from the D-type flip-flops 322 and 321, respectively. This is because the input of the delay circuit 343 is also 16 bits by the D-type flip-flops 342 and 341 in the converter 202 of FIG. In FIG. 22, by barrel-shifting the 17 bits output by the staff circuit 323 with the barrel shifter 324, for example, 8 bits are finally selected and output. The 15-bit data output by the 343 is converted into 8-bit data by barrel-shifting the data by a predetermined amount with the barrel shifter 344.
FIG. 27 represents another configuration example of the converter 212. In this configuration example, the counter 361 counts the number of consecutive 0 bits in the input data and outputs the count result to the controller 326. The controller 326 outputs the signal stuff position to the stuff circuit 323, for example, when the counter 361 counts 22 consecutive 0 bits. At this time, the controller 326 resets the counter 361 and causes the counter 361 to count the number of consecutive 0 bits again.
Other configurations and operations are the same as in FIG. 22.
FIG. 28 represents another configuration example of the converter 202. In this configuration example, the counter 371 counts the number of consecutive 0s in the input data, and the count result is output to the controller 346. When the count value of the counter 371 reaches 22, the controller 346 outputs the signal delete position to the delay circuit 343, resets the counter 371, and causes the counter 371 to count the number of new consecutive 0 bits again. .. Other configurations are the same as in FIG. 25.
[0198] As described above, in this configuration example, the data "1" as the marker bit is inserted and deleted based on a predetermined pattern (consecutive number of data "0"). ..
[0199] The configurations shown in FIGS. 27 and 28 enable more efficient processing than the configurations shown in FIGS. 22 and 25. However, the length after conversion depends on the original history information.
FIG. 29 shows a configuration example of the user data formatter 213. In this example, when the controller 383 outputs a read address to a buffer memory (not shown) arranged between the converter 212 and the user data formatter 213, the data read from the read address is the user data formatter 213. It is supplied to the contact a side of the switch 382. The ROM 381 stores data necessary for generating user_data () such as a user data start code and a data ID. The controller 313 switches the switch 382 to the contact a side or the contact b side at a predetermined timing, appropriately selects and outputs the data stored in the ROM 381 or the data supplied from the converter 212. As a result, the data in the format of user_data () is output to the encoding device 106.
Although not shown, the user data decoder 201 is realized by outputting input data via a switch that is read from ROM 381 in FIG. 29 and deletes the inserted data. be able to.
[0202] FIG. 30 shows a state in which a plurality of transcoders 101-1 to 101-N are connected and used in series, for example, in a video editing studio. The history information multiplexing device 103-i of each transcoder 101-i (i = 1 to N) was used by itself in the section in which the oldest coding parameter of the above-mentioned coding parameter area was recorded. Overwrite the latest coding parameters. As a result, the coding parameters (generation history information) for the latest four generations corresponding to the same macroblock are recorded in the baseband image data (Fig. 18).
[0203] The encoder 121-i (FIG. 19) of each encoding device 106-i is based on the coding parameters used this time supplied from the history information separator 105-i in the variable length coding circuit 58. Encode the video data supplied by the quantization circuit 57. The current coding parameters are multiplexed in the bitstream thus generated (eg, picture_header ()).
[0204] The variable length coding circuit 58 also multiplexes the user data (including generation history information) supplied by the history encoding device 107-i into the output bitstream (embedding as shown in FIG. 18). Multiplex in bitstream, not processing). Then, the bit stream output by the encoding device 106-i is input to the transcoder 101- (i + 1) in the subsequent stage via the SDTI (Serial Data Transfer Interface) 108-i.
[0205] The transcoder 101-i and the transcoder 101- (i + 1) are respectively configured as shown in FIG. Therefore, the process is the same as that described with reference to FIG.
[0206] When it is desired to change what was currently encoded as an I picture to a P or B picture as the coding using the history of the actual coding parameters, look at the history of the past coding parameters and see the past. Look for cases that are P or B pictures, and if these histories exist, change the picture type using parameters such as the motion vector. On the other hand, if there is no history in the past, change the picture type that does not detect motion is abandoned. Of course, even if there is no history, the picture type can be changed by performing motion detection.
[0207] In the case of the format shown in FIG. 18, the coding parameters for four generations are embedded, but the parameters of the picture types I, P, and B can also be embedded. FIG. 31 shows an example of the format in this case. In this example, when the same macroblock is encoded with a change in the picture type in the past, the coding parameters (picture history information) for one generation are recorded for each picture type. Therefore, the decoder 111 shown in FIG. 16 and the encoder 121 shown in FIG. 19 replace the current (latest), 3rd, 2nd, and 1st generation coding parameters with I-pictures, P-pictures. , And one generation of coding parameters corresponding to the B picture will be input and output.
[0208] Further, in the case of this example, since the free area of Cb [1] [x] and Cr [1] [x] is not used, the area of Cb [1] [x] and Cr [1] [x] is not used. The present invention can also be applied to image data in 4: 2: 0 format that does not have.
[0209] In the case of this example, the decoding device 102 extracts the coding parameter at the same time as decoding, determines the picture type, and writes (multiplexes) the coding parameter at a location corresponding to the picture type of the image signal. Output to the history information separator 105. The history information separation device 105 can separate the coding parameters and perform recoding while changing the picture type in consideration of the picture type to be encoded and the past coding parameters input.
Next, in each transcoder 101, a process of determining a picture type that can be changed will be described with reference to the flowchart of FIG. Since the change of the picture type in the transcoder 101 uses the past motion vector, it is premised that this process is executed without motion detection. Further, the process described below is executed by the history information separation device 105.
[0211] In step S1, the coding parameters (picture history information) for one generation for each picture type are input to the history information separator 105.
[0212] In step S2, the history information separation device 105 determines whether or not the coding parameter when the picture is changed to the B picture exists in the picture history information. If it is determined that the picture history information has the coding parameter when the picture is changed to the B picture, the process proceeds to step S3.
[0213] In step S3, the history information separation device 105 determines whether or not the coding parameter when the picture is changed to the P picture exists in the picture history information. If it is determined that the picture history information has the coding parameter when the picture is changed to the P picture, the process proceeds to step S4.
[0214] In step S4, the history information separator 105 determines that the changeable picture types are I picture, P picture, and B picture.
[0215] If it is determined in step S3 that the picture history information does not have the coding parameter when the picture is changed to the P picture, the process proceeds to step S5.
[0216] In step S5, the history information separator 105 determines that the changeable picture types are I picture and B picture. Furthermore, the history information separator 105 can be changed to a P picture in a pseudo manner by performing special processing (using only the forward prediction vector without using the backward prediction vector included in the history information of the B picture). to decide.
[0217] If it is determined in step S2 that the picture history information does not have the coding parameter when the picture is changed to the B picture, the process proceeds to step S6.
[0218] In step S6, the history information separator 105 determines whether or not the picture history information has a coding parameter when the picture is changed to a P picture. If it is determined that the picture history information has the coding parameter when the picture is changed to the P picture, the process proceeds to step S7.
[0219] In step S7, the history information separator 105 determines that the changeable picture types are I picture and P picture. Further, the history information separation device 105 determines that the P picture can be changed to the B picture by performing a special process (only the forward prediction vector included in the history information is used for the P picture).
[0220] If it is determined in step S6 that the picture history information does not have the coding parameter when the picture is changed to the P picture, the process proceeds to step S8. In step S8, since the motion vector does not exist, the history information separator 105 determines that the only picture type that can be changed is the I picture (because it is an I picture, it cannot be changed other than the I picture).
[0221] In step S9 following the processing of steps S4, S5, S7, and S8, the history information separator 105 displays a changeable picture type on a display device (not shown) and notifies the user.
[0222] FIG. 33 shows an example of changing the picture type. When changing the picture type, the number of frames that make up the GOP is changed. That is, in the case of this example, a 4 Mbps Long GOP (4 Mbps Long GOP) composed of frames with N = 15 (number of frames of GOP N = 15) and M = 3 (I or P picture appearance period M = 3 in GOP). 1st generation) is converted to 50Mbps Short GOP (2nd generation) consisting of N = 1, M = 1 frames, and again 4Mbps Long composed of N = 15, M = 3 frames. It has been converted to GOP (3rd generation). In the figure, the broken line indicates the boundary of GOP.
[0223] When the picture type is changed from the first generation to the second generation, as is clear from the description of the changeable picture type determination process described above, all frames should change the picture type to I picture. Is possible. At the time of this picture type change, all the motion vectors calculated when the moving image (0th generation) is converted to the 1st generation are stored (remained) in the picture history information. Next, when it is converted to Long GOP again (the picture type is changed from the 2nd generation to the 3rd generation), the motion vector for each picture type when it is converted from the 0th generation to the 1st generation is saved. Therefore, by reusing this, it is possible to suppress the deterioration of image quality and convert it to Long GOP again.
FIG. 34 shows another example of changing the picture type. In the case of this example, the 4 Mbps Long GOP (1st generation) with N = 14, M = 2 is converted to the 18 Mbps Short GOP (2nd generation) with N = 2, M = 2, and then N. It is converted to a 50 Mbps Short GOP (3rd generation) with = 1, M = 1 and 1 frame, and converted to a 1 Mbps, random GOP (4th generation) with N frames.
[0225] Also in this example, the motion vector for each picture type when converted from the 0th generation to the 1st generation is stored until the conversion from the 3rd generation to the 4th generation. Therefore, as shown in FIG. 34, even if the picture type is changed in a complicated manner, the image quality deterioration can be suppressed to be small by reusing the stored coding parameters. Furthermore, if the quantization scale of the stored coding parameters is effectively used, coding with less deterioration in image quality can be realized.
The reuse of this quantization scale will be described with reference to FIG. FIG. 35 shows that a given frame is always converted to an I picture from the 1st generation to the 4th generation, and only the bit rate is changed to 4Mbps, 18Mbps, or 50Mbps.
[0227] For example, when converting from the first generation (4 Mbps) to the second generation (18 Mbps), the image quality does not improve even if re-encoded on a fine quantization scale as the bit rate increases. This is because the data quantized in the rough quantization step in the past is not restored. Therefore, as shown in FIG. 35, even if the bit rate is increased in the middle, the quantization in the fine quantization step accordingly only increases the amount of information and does not lead to the improvement of the image quality. Therefore, if controlled so as to maintain the coarsest (larger) quantization scale in the past, the least wasteful and efficient coding becomes possible.
[0228] When changing from the 3rd generation to the 4th generation, the bit rate is reduced from 50 Mbps to 4 Mbps, but even in this case, the coarsest (larger) quantization scale in the past is maintained. ..
[0229] As described above, when the bit rate is changed, it is very effective to use the history of the past quantization scale for coding.
[0230] This quantization control process will be described with reference to the flowchart of FIG. In step S11, the history information separator 105 determines whether or not the input picture history information has a coding parameter of the picture type to be converted. If it is determined that the coding parameter of the picture type to be converted exists, the process proceeds to step S12.
[0231] In step S12, the history information separator 105 extracts history_q_scale_code from the coding parameters that are the targets of the picture history information.
[0232] In step S13, the history information separator 105 calculates feedback_q_scale_code based on the remaining buffer amount fed back from the transmission buffer 59 to the quantization circuit 57.
[0233] In step S14, the history information separator 105 determines whether or not the history_q_scale_code is larger (coarse) than the feedback_q_scale_code. If it is determined that history_q_scale_code is larger than feedback_q_scale_code, the process proceeds to step S15.
[0234] In step S15, the history information separator 105 outputs history_q_scale_code as a quantization scale to the quantization circuit 57. The quantization circuit 57 performs quantization using history_q_scale_code.
[0235] In step S16, it is determined whether or not all the macroblocks included in the frame have been quantized. If it is determined that all the macroblocks have not been quantized yet, the process returns to step S12, and the processing of steps S12 to S16 is repeated until all the macroblocks are quantized.
[0236] If it is determined in step S14 that history_q_scale_code is not larger (fine) than feedback_q_scale_code, the process proceeds to step S17.
[0237] In step S17, the history information separator 105 outputs feedback_q_scale_code as a quantization scale to the quantization circuit 57. The quantization circuit 57 executes quantization using feedback_q_scale_code.
[0238] If it is determined in step S11 that the coding parameter of the picture type to be converted does not exist in the history information, the process proceeds to step S18.
[0239] In step S18, the history information separator 105 calculates feedback_q_scale_code based on the remaining buffer amount fed back from the transmission buffer 59 to the quantization circuit 57.
[0240] In step S19, the quantization circuit 57 performs quantization using Feedback_q_scale_code.
[0241] In step S20, it is determined whether or not all the macroblocks included in the frame have been quantized. If it is determined that all the macroblocks have not been quantized yet, the process returns to step S18, and the processing of steps S18 to S20 is repeated until all the macroblocks are quantized.
[0242] In the transcoder 101 according to the present embodiment, as described above, the decoding side and the coding side are roughly coupled, and the coding parameters are multiplexed and transmitted to the image data. As shown in FIG. 37, the decoding device 102 and the coding device 106 may be directly connected (tightly coupled).
[0243] The transcoder 101 described with reference to FIG. 15 multiplexes the past coding parameters into the baseband video data in order to supply the first to third generation past coding parameters to the coding apparatus 106. I was trying to transmit. However, in the present invention, a technique for multiplexing past coding parameters into baseband video data is not essential, and as shown in FIG. 37, a transmission line different from that of baseband video data (for example, a data transfer bus) is required. ) May be used to transmit past coding parameters.
That is, the decoding device 102, the history decoding device 104, the coding device 106, and the history encoding device 107 shown in FIG. 37 are the decoding device 102, the history decoding device 104, and the coding device described in FIG. It has exactly the same functions and configurations as the 106 and the history encoding device 107.
[0245] The variable-length decoding circuit 112 of the decoding device 102 is composed of the sequence layer, the GOP layer, the picture layer, the slice layer, and the macroblock layer of the third-generation coded stream ST (3rd), and the third-generation coding parameters. Is extracted and supplied to the controller 70 of the history encoding device 107 and the coding device 106, respectively. The history encoding device 107 converts the received 3rd generation coding parameter into converted_history_stream () so that it can be described in the user data area of the picture layer, and the converted_history_stream () is used as user data for variable length coding of the coding device 106. Supply to circuit 58.
[0246] Further, the variable length decoding circuit 112 extracts the user data user_data including the first generation coding parameter and the second coding parameter from the user data area of the picture layer of the third generation coded stream. Then, it is supplied to the variable length coding circuit 58 of the history decoding device 104 and the coding device 106. The history decoding device 104 extracts the first-generation coding parameter and the second-generation coding parameter from the history stream described as converted_history_stream () in the user data area, and outputs the first-generation coding parameter and the second-generation coding parameter to the controller of the coding device 106. Supply.
The controller 70 of the coding device 106 is based on the first and second generation coding parameters received from the history decoding device 104 and the third generation coding parameters received from the coding device 102. It controls the coding process of the coding device 106.
[0248] The variable length coding circuit 58 of the coding device 106 receives the user data user_data including the first generation coding parameters and the second coding parameters from the decoding device 102, and also receives the user data user_data from the history encoding device 107. It receives user data user_data containing 3rd generation coding parameters, and describes those user data as history information in the user data area of the picture layer of the 4th generation coding stream.
[0249] FIG. 38 is a diagram showing the syntax for decoding an MPEG video stream. The decoder extracts a plurality of meaningful data items (data elements) from the bitstream by decoding the MPEG bitstream according to this syntax. In the syntax described below, the functions and conditional statements are represented in fine print, and the data elements are represented in bold type. Data items are described in Mnemonic, which indicates their name, bit length, and their type and transmission order.
First, the functions used in the syntax shown in FIG. 38 will be described.
[0251] The next_start_code () function is a function for searching for a start code described in a bit stream. In the syntax shown in FIG. 38, the sequence_header () function and the sequence_extension () function are arranged in order after this next_start_code () function, so this bit stream has this sequence_header () function and Describes the data elements defined by the sequence_extension () function. Therefore, when decoding a bitstream, this next_start_code () function finds the start code (a type of data element) described at the beginning of the sequence_header () and sequence_extension () functions in the bitstream, and uses that as a reference. It finds more sequence_header () and sequence_extension () functions and decodes each data element defined by them.
The sequence_header () function is a function for defining the header data of the sequence layer of the MPEG bitstream, and the sequence_extension () function is for defining the extension data of the sequence layer of the MPEG bitstream. It is a function.
The do {} while syntax placed next to the sequence_extension () function is data described based on the function in the {} of the do statement while the condition defined by the while statement is true. A syntax for extracting elements from a data stream. That is, according to the do {} while syntax, while the condition defined by the while statement is true, the decoding process is performed to extract the data element described based on the function in the do statement from the bitstream.
The nextbits () function used in this while statement is a function for comparing a bit or a bit string appearing in a bitstream with a data element to be decoded next. In this example syntax in Figure 38, the nextbits () function compares the bitstring in the bitstream with the sequence_end_code, which marks the end of the video sequence, and when the bitstring in the bitstream and sequence_end_code do not match, this while The condition of the statement is true. Therefore, in the do {} while syntax placed after the sequence_extension () function, the data element defined by the function in the do statement is bitstreamed while the sequence_end_code indicating the end of the video sequence does not appear in the bitstream. Indicates that it is described in.
[0255] In the bitstream, the data element defined by the extension_and_user_data (0) function is described next to each data element defined by the sequence_extension () function. This extension_and_user_data (0) function is a function for defining the extension data and user data of the sequence layer of the MPEG bitstream.
The do {} while syntax placed next to this extension_and_user_data (0) function is described based on the function in the {} of the do statement while the condition defined by the while statement is true. This is a function for extracting the data element from the bitstream. The nextbits () function used in this while statement is a function for determining the match between the bit or bit string appearing in the bitstream and picture_start_code or group_start_code, and the bit or bit string appearing in the bitstream. If the picture_start_code or group_start_code matches, the condition defined by the while statement is true. Therefore, this do { } In the while syntax, when picture_start_code or group_start_code appears in the bitstream, the code of the data element defined by the function in the do statement is described after the start code, so this picture_start_code or group_start_code By finding the start code indicated by, the data elements defined in the do statement can be extracted from the bitstream.
[0257] The if statement described at the beginning of this do statement indicates the condition that group_start_code appears in the bitstream. If the condition by this if statement is true, the data elements defined by the group_of_picture_header (1) function and the extension_and_user_data (1) function are described in order after this group_start_code in the bitstream.
[0258] This group_of_picture_header (1) function is a function for defining the header data of the GOP layer of the MPEG bit stream, and the extension_and_user_data (1) function is the extension data (extension_data) of the GOP layer of the MPEG bit stream. It is a function to define user data (user_data).
Further, in this bitstream, the data elements defined by the group_of_picture_header (1) and extension_and_user_data (1) functions are followed by the data elements defined by the picture_header () and picture_coding_extension () functions. It has been described. Of course, if the condition of the if statement explained above is not true, the data element defined by the group_of_picture_header (1) function and the extension_and_user_data (1) function is not described, so it is defined by the extension_and_user_data (0) function. Next to the data elements that are set, the data elements defined by the picture_header () and picture_coding_extension () functions are described.
[0260] This picture_header () function is a function for defining the header data of the picture layer of the MPEG bitstream, and the picture_coding_extension () function defines the first extension data of the picture layer of the MPEG bitstream. Function for.
The following while statement is a function for determining the condition of the next if statement while the condition defined by this while statement is true. The nextbits () function used in this while statement is a function for determining the match between the bit string appearing in the bitstream and the extension_start_code or user_data_start_code, and the bit string appearing in the bitstream and the extension_start_code or user_data_start_code If they match, the condition defined by this while statement is true.
[0262] The first if statement is a function for determining the match between the bit string appearing in the bit stream and the extension_start_code. If the bit string appearing in the bitstream matches the 32-bit extension_start_code, the data element defined by the extension_data (2) function is described next to the extension_start_code in the bitstream.
[0263] The second if statement is a syntax for determining a match between the bit string appearing in the bit stream and user_data_start_code, and when the bit string appearing in the bit stream and the 32-bit user_data_start_code match, the second if statement is used. The condition of the third if statement is judged. This user_data_start_code is a start code for indicating the start of the user data area of the picture layer of the MPEG bitstream.
[0264] The third if statement is a syntax for determining a match between a bit string appearing in a bit stream and History_Data_ID. If the bit string appearing in the bitstream matches this 32-bit History_Data_ID, then in the user data area of the picture layer of this MPEG bitstream, the code indicated by this 32-bit History_Data_ID is followed by the converted_history_stream () function. Describes the data elements defined by.
The converted_history_stream () function is a function for describing history information and history data for transmitting all the coding parameters used at the time of MPEG coding. Details of the data elements defined by this converted_history_stream () function will be described later as history_stream () with reference to FIGS. 40 to 47. Further, this History_Data_ID is a start code for indicating the beginning of the history information and the history data described in the user data area of the picture layer of the MPEG bit stream.
[0266] The else statement is a syntax for indicating that the condition is non-true in the third if statement. Therefore, in the user data area of the picture layer of this MPEG bitstream, if the data element defined by the converted_history_stream () function is not described, the data element defined by the user_data () function is described.
[0267] In FIG. 38, the history information is described in converted_history_stream () and not in user_data (), but this converted_history_stream () is described as a kind of user_data of the MPEG standard. Therefore, in the present specification, it is also explained that the history information is described in user_data in some cases, but it means that it is described as a kind of user_data of the MPEG standard.
[0268] The picture_data () function is a function for describing data elements related to the slice layer and the macroblock layer next to the user data of the picture layer of the MPEG bit stream. Normally, the data element indicated by this picture_data () function is the data element defined by the converted_history_stream () function or the data element defined by the user_data () function described in the user data area of the picture layer of the bitstream. As described below, if extension_start_code or user_data_start_code does not exist in the bitstream indicating the data element of the picture layer, the data element indicated by this picture_data () function is defined by the picture_coding_extension () function. It is described next to the data element.
[0269] Next to the data element indicated by this picture_data () function, the data elements defined by the sequence_header () function and the sequence_extension () function are arranged in order. The data elements described by the sequence_header () and sequence_extension () functions are exactly the same as the data elements described by the sequence_header () and sequence_extension () functions described at the beginning of the sequence in the video stream. The reason for describing the same data in the stream in this way is that when reception is started from the middle of the data stream (for example, the bitstream part corresponding to the picture layer) on the bitstream receiver side, the data in the sequence layer is received. This is to prevent the stream from being unable to be decoded.
[0270] Next to the data element defined by this last sequence_header () function and sequence_extension () function, that is, at the end of the data stream, a 32-bit sequence_end_code indicating the end of the sequence is described.
[0271] The outline of the basic configuration of the above syntax is shown in FIG. 39.
Next, the history stream defined by the converted_history_stream () function will be described.
[0273] This converted_history_stream () is a function for inserting a history stream indicating history information into the user data area of the MPEG picture layer. The meaning of "converted" is that in order to prevent start emulation, a conversion process is performed in which a marker bit (1 bit) is inserted at least every 22 bits of a history stream composed of history data to be inserted into the user area. It means that it is a stream.
[0274] This converted_history_stream () is described in the form of either a fixed-length history stream (FIGS. 40 to 46) or a variable-length history stream (FIG. 47) described below. When a fixed-length history stream is selected on the encoder side, there is an advantage that the circuit and software for decoding each data element from the history stream on the decoder side are simplified. On the other hand, when a variable-length history stream is selected on the encoder side, the history information (data element) described in the user area of the picture layer can be arbitrarily selected on the encoder as needed, so that the history stream can be selected. The amount of data in the data can be reduced, and as a result, the data rate of the entire encoded bit stream can be reduced.
[0275] The "history stream", "history stream", "history information", "history information", "history data", "history data", "history parameter", and "history parameter" described in the present invention are It means the coding parameter (or data element) used in the past coding process, and does not mean the coding parameter used in the current (final stage) coding process. For example, in the first generation coding process, a certain picture is encoded by an I picture and transmitted, and in the next second generation coding process, this picture is encoded and transmitted as a P picture, and further. In the third generation coding process, an example of encoding this picture with a B picture and transmitting it will be described.
[0276] The coding parameters used in the third-generation coding process are the sequence layer, GOP layer, picture layer, slice layer, and macroblock layer of the coded bitstream generated in the third-generation coding process. It is described in a predetermined position. On the other hand, the coding parameters used in the 1st generation and 2nd generation coding processing, which are the past coding processing, are the sequence layer and GOP layer in which the coding parameters used in the 3rd generation coding processing are described. It is described in the user data area of the picture layer as the history information of the coding parameter according to the syntax already described.
First, a fixed-length history stream syntax will be described with reference to FIGS. 40 to 46.
[0278] In the user data area of the picture layer of the bitstream generated in the final stage (for example, third generation) coding process, first of all, the past (for example, first generation and second generation) coding processing The coding parameters included in the sequence header of the sequence layer used in are inserted as a history stream. Note that the history information such as the sequence header of the sequence layer of the bitstream generated in the past coding process is not inserted into the sequence header of the sequence layer of the bitstream generated in the final coding process. It should be noted that.
[0279] The data elements included in the sequence header (sequence_header) used in the past coding process are sequence_header_code, sequence_header_present_flag, horizontal_size_value, marker_bit, vertical_size_value, aspect_ratio_information, frame_rate_code, bit_rate_value, VBV_buffer_size_value, const. It is composed of non_intra_quantiser_matrix and so on.
[0280] The sequence_header_code is data representing the start synchronization code of the sequence layer. sequence_header_present_flag is data indicating whether the data in sequence_header is valid or invalid. horizontal_size_value is data consisting of the lower 12 bits of the number of pixels in the horizontal direction of the image. marker_bit is bit data inserted to prevent start code emulation. vertical_size_value is data consisting of the lower 12 bits of the number of vertical lines in the image. The aspect_ratio_information is data representing the aspect ratio (aspect ratio) of the pixel or the display screen aspect ratio. frame_rate_code is data representing the display cycle of the image.
[0281] bit_rate_value is the lower 18 bits (rounded up in 400 bsp units) data of the bit rate for limiting the amount of generated bits. VBV_buffer_size_value is the lower 10-bit data of the value that determines the size of the virtual buffer (video buffer verifier) for controlling the generated code amount. constrained_parameter_flag is data indicating that each parameter is within the limit. load_intra_quantiser_matrix is data indicating the existence of quantization matrix data for intra MB. load_non_intra_quantiser_matrix is data indicating the existence of quantization matrix data for non-intra MB. intra_quantiser_matrix is data indicating the value of the quantization matrix for intra MB. non_intra_quantiser_matrix is data representing the value of the quantization matrix for non-intra MB.
[0282] In the user data area of the picture layer of the bitstream generated in the final-stage coding process, a data element representing the sequence extension of the sequence layer used in the past coding process is described as a history stream. To.
[0283] The data elements representing the sequence extension (sequence_extension) used in this past coding process are extension_start_code, extension_start_code_identifier, sequence_extension_present_flag, profile_and_level_indication, progressive_sequence, chroma_format, horizontal_size_extension, vertical_size_extension, bit_rate_extension, vertical_size_extension, bit_rate_extension, etc. Data element of.
[0284] The extension_start_code is data representing the start synchronization code of the extension data. The extension_start_code_identifier is data indicating which extension data is sent. sequence_extension_present_flag is data indicating whether the data in the sequence extension is valid or invalid. profile_and_level_indication is data for specifying the profile and level of video data. The progressive_sequence is data indicating that the video data is a sequential scan. chroma_format is data for specifying the color difference format of video data.
[0285] horizontal_size_extension is the upper 2 bits of data to be added to the horizontal_size_value of the sequence header. vertical_size_extension is the upper 2 bits of data to be added to the vertical_size_value of the sequence header. bit_rate_extension is the data of the upper 12 bits to be added to bit_rate_value of the sequence header. vbv_buffer_size_extension is the upper 8 bits of data to be added to vbv_buffer_size_value in the sequence header. low_delay is data indicating that the B picture is not included. frame_rate_extension_n is data for obtaining the frame rate in combination with the frame_rate_code of the sequence header. frame_rate_extension_d is data for obtaining the frame rate in combination with frame_rate_code in the sequence header.
Subsequently, in the user area of the picture layer of the bitstream, a data element representing the sequence display extension of the sequence layer used in the past coding process is described as a history stream.
The data element described as this sequence display extension (sequence_display_extension) is composed of extension_start_code, extension_start_code_identifier, sequence_display_extension_present_flag, video_format, colour_description, colour_primaries, transfer_characteristics, matrix_coeffients, display_horizontal_size, and display_vertical.
[0288] The extension_start_code is data representing the start synchronization code of the extension data. extension_start_code_identifier is a code that indicates which extension data is sent. sequence_display_extension_present_flag is data indicating whether the data element in the sequence display extension is valid or invalid. video_format is data representing the video format of the original signal. color_description is data indicating that there is detailed data of the color space. The color_primaries are data showing the details of the color characteristics of the original signal. The transfer_characteristics are data showing details of how the photoelectric conversion was performed. matrix_coeffients are data showing details of how the primary signals are transformed from the three primary colors of light. display_horizontal_size is data representing the active area (horizontal size) of the intended display. display_vertical_size is data representing the active area (vertical size) of the intended display.
Subsequently, in the user area of the picture layer of the bitstream generated in the final-stage coding process, macroblock assignment data (macroblock_assignment_in_user_data) indicating the phase information of the macroblock generated in the past coding process. ) Is described as a history stream.
[0290] The macroblock_assignment_in_user_data indicating the phase information of this macroblock is composed of data elements such as macroblock_assignment_present_flag, v_phase, and h_phase.
[0291] This macroblock_assignment_present_flag is data indicating whether the data element in macroblock_assignment_in_user_data is valid or invalid. v_phase is data indicating the phase information in the vertical direction when a macroblock is cut out from the image data. h_phase is data indicating the phase information in the horizontal direction when a macroblock is cut out from the image data.
Subsequently, in the user area of the picture layer of the bitstream generated by the final-stage coding process, a data element representing the GOP header of the GOP layer used in the past coding process is used as a history stream. It has been described.
[0293] The data element representing this GOP header (group_of_picture_header) is composed of group_start_code, group_of_picture_header_present_flag, time_code, closed_gop, and broken_link.
[0294] The group_start_code is data indicating the start synchronization code of the GOP layer. group_of_picture_header_present_flag is data indicating whether the data element in group_of_picture_header is valid or invalid. time_code is a time code indicating the time from the beginning of the sequence of the first picture of the GOP. closed_gop is flag data indicating that the image in the GOP can be played independently from other GOPs. broken_link is flag data indicating that the first B picture in the GOP cannot be reproduced accurately due to editing or the like.
Subsequently, in the user area of the picture layer of the bitstream generated by the final-stage coding process, a data element representing the picture header of the picture layer used in the past coding process is used as a history stream. It has been described.
[0296] The data element related to this picture header (picture_header) is composed of picture_start_code, temporary_reference, picture_coding_type, vbv_delay, full_pel_forward_vector, forward_f_code, full_pel_backward_vector, and backward_f_code.
[0297] Specifically, picture_start_code is data representing the start synchronization code of the picture layer. temporal_reference is the data that is reset at the beginning of the GOP with a number indicating the display order of the pictures. picture_coding_type is data indicating the picture type. vbv_delay is data indicating the initial state of the virtual buffer at the time of random access. The full_pel_forward_vector is data indicating whether the accuracy of the forward motion vector is in integer units or half pixel units. The forward_f_code is data representing the forward motion vector search range. The full_pel_backward_vector is data indicating whether the accuracy of the reverse motion vector is in integer units or half pixel units. The backward_f_code is data representing the reverse motion vector search range.
Subsequently, in the user area of the picture layer of the bitstream generated by the final coding process, the picture coding extension of the picture layer used in the past coding process is described as a history stream. There is.
The data elements related to this picture coding extension (picture_coding_extension) are extension_start_code, extension_start_code_identifier, f_code [0] [0], f_code [0] [1], f_code [1] [0], f_code [1] [1]. , Intra_dc_precision, picture_structure, top_field_first, frame_predictive_frame_dct, concealment_motion_vectors, q_scale_type, intra_vlc_format, alternate_scan, repeat_firt_field, chroma_420_type, progressive_frame, composite_display_flag, v_axis,
[0300] The extension_start_code is a start code indicating the start of the extension data of the picture layer. extension_start_code_identifier is a code that indicates which extension data is sent. f_code [0] [0] is data representing the horizontal motion vector search range in the forward direction. f_code [0] [1] is data representing the vertical motion vector search range in the forward direction. f_code [1] [0] is data representing the horizontal motion vector search range in the backward direction. f_code [1] [1] is data representing the vertical motion vector search range in the backward direction.
[0301] intra_dc_precision is data representing the accuracy of the DC coefficient. The picture_structure is data indicating whether it is a frame structure or a field structure. In the case of a field structure, it is data that also indicates the upper field or the lower field. In the case of a frame structure, top_field_first is data indicating whether the first field is high or low. frame_predictive_frame_dct is data indicating that in the case of a frame structure, the prediction of the frame mode DCT is only in the frame mode. The concealment_motion_vectors are data indicating that the intra macroblock has a motion vector for concealing the transmission error.
[0302] q_scale_type is data indicating whether to use a linear quantization scale or a non-linear quantization scale. intra_vlc_format is data indicating whether to use another 2D VLC for the intra macroblock. alternate_scan is data that represents the choice between using a zigzag scan or an alternate scan. repeat_firt_field is the data used for 2: 3 pulldown. chroma_420_type is the data that represents the same value as the next progressive_frame if the signal format is 4: 2: 0, or 0 otherwise. The progressive_frame is data indicating whether or not this picture can be scanned sequentially. The composite_display_flag is data indicating whether or not the source signal is a composite signal.
[0303] v_axis is data used when the source signal is PAL. field_sequence is the data used when the source signal is PAL. sub_carrier is the data used when the source signal is PAL. burst_amplitude is the data used when the source signal is PAL. sub_carrier_phase is the data used when the source signal is PAL.
Subsequently, in the user area of the picture layer of the bitstream generated by the final-stage coding process, the quantization matrix extension used in the past coding process is described as a history stream.
[0305] Data elements related to the quantization matrix extension (quant_matrix_extension), extension_start_code, extension_start_code_identifier, quant_matrix_extension_present_flag, load_intra_quantiser_matrix, intra_quantiser_matrix [64], load_non_intra_quantiser_matrix, non_intra_quantiser_matrix [64], load_chroma_intra_quantiser_matrix, chroma_intra_quantiser_matrix [64], load_chroma_non_intra_quantiser_matrix, and chroma_non_intra_quantiser_matrix [64] Consists of.
[0306] extension_start_code is a start code indicating the start of this quantization matrix extension. extension_start_code_identifier is a code that indicates which extension data is sent. quant_matrix_extension_present_flag is data to indicate whether the data element in this quantized matrix extension is valid or invalid. load_intra_quantiser_matrix is data indicating the existence of quantization matrix data for intra macroblocks. intra_quantiser_matrix is data indicating the value of the quantization matrix for the intra macroblock.
[0307] load_non_intra_quantiser_matrix is data indicating the existence of quantization matrix data for non-intra macroblocks. non_intra_quantiser_matrix is data representing the value of the quantization matrix for non-intra macroblocks. load_chroma_intra_quantiser_matrix is data indicating the existence of quantization matrix data for the color difference intra macroblock. chroma_intra_quantiser_matrix is data indicating the value of the quantization matrix for the color difference intra macroblock. load_chroma_non_intra_quantiser_matrix is data indicating the existence of quantization matrix data for color difference non-intra macroblocks. chroma_non_intra_quantiser_matrix is data indicating the value of the quantization matrix for the color difference non-intra macroblock.
Subsequently, in the user area of the picture layer of the bitstream generated by the final-stage coding process, the copyright extension used in the past coding process is described as a history stream.
[0309] The data element related to this copyright extension (copyright_extension) is composed of extension_start_code, extension_start_code_itentifier, copyright_extension_present_flag, copyright_flag, copyright_identifier, original_or_copy, copyright_number_1, copyright_number_2, and copyright_number_3.
[0310] The extension_start_code is a start code indicating the start of the copyright extension. This is the code that indicates which extension data of extension_start_code_itentifier is sent. copyright_extension_present_flag is data to indicate whether the data element in this copyright extension is valid or invalid. copyright_flag indicates whether or not copy rights are granted to the encoded video data until the next copyright extension or sequence end.
[0311] The copyright_identifier is data for identifying the registration authority of the copy right specified by ISO / IEC JTC / SC29. The original_or_copy is data indicating whether the data in the bitstream is original data or copy data. copyright_number_1 is data representing bits 44 to 63 of the copyright number. copyright_number_2 is data representing bits 22 to 43 of the copyright number. copyright_number_3 is data representing bits 0 to 21 of the copyright number.
Subsequently, in the user area of the picture layer of the bitstream generated by the final-stage coding process, the picture display extension (picture_display_extension) used in the past coding process is described as a history stream. There is.
[0313] The data element representing this picture display extension is composed of extension_start_code, extension_start_code_identifier, picture_display_extension_present_flag, frame_center_horizontal_offset_1, frame_center_vertical_offset_1, frame_center_horizontal_offset_2, frame_center_vertical_offset_2, frame_center_offset_2, frame_center_offset_2.
[0314] The extension_start_code is a start code for indicating the start of the picture display extension. extension_start_code_identifier is a code that indicates which extension data is sent. The picture_display_extension_present_flag is data indicating whether the data element in the picture display extension is valid or invalid. frame_center_horizontal_offset is data indicating the horizontal offset of the display area, and up to three offset values can be defined. frame_center_vertical_offset is data indicating the vertical offset of the display area, and up to three offset values can be defined.
[0315] In the user area of the picture layer of the bitstream generated in the final-stage coding process, the user data used in the past coding process (next to the history information representing the picture display extension already described) ( user_data) is described as a history stream.
[0316] Next to this user data, information about the macroblock layer used in the past coding process is described as a history stream.
[0317] Information on this macroblock layer includes data elements related to the position of macroblocks (macroblock) such as macroblock_address_h, macroblock_address_v, slice_header_present_flag, skipped_macroblock_flag, macroblock_quant, macroblock_motion_forward, macroblock_motion_backward, mocroblock_pattern, macroblock_tempor, macroblock_pattern, macroblock_in Data elements related to macroblock modes (macroblock_modes []), data elements related to quantization step control such as quantizer_scale_code, PMV [0] [0] [0], PMV [0] [0] [1], motion_vertical_field_select [0] ] [0], PMV [0] [1] [0], PMV [0] [1] [1], motion_vertical_field_select [0] [1], PMV [1] [0] [0], PMV [1] Data related to motion compensation such as [0] [1], motion_vertical_field_select [1] [0], PMV [1] [1] [0], PMV [1] [1] [1], motion_vertical_field_select [1] [1] It is composed of elements, data elements related to macroblock patterns such as coded_block_pattern, and data elements related to generated code amounts such as num_mv_bits, num_coef_bits, and num_other_bits.
[0318] The data elements related to the macroblock layer will be described in detail below.
[0319] macroblock_address_h is data for defining the absolute horizontal position of the current macroblock. macroblock_address_v is the data for defining the absolute position of the current macroblock in the vertical direction. The slice_header_present_flag is data indicating whether or not this macroblock is the head of the slice layer and is accompanied by a slice header. skipped_macroblock_flag is data indicating whether or not to skip this macroblock in the decoding process.
[0320] The macroblock_quant is data derived from the macroblock type (macroblock_type) shown in FIGS. 63 and 64, which will be described later, and is data indicating whether or not the quantizer_scale_code appears in the bit stream. The macroblock_motion_forward is the data derived from the macroblock types shown in FIGS. 63 and 64, and is the data used in the decoding process. The macroblock_motion_backward is data derived from the macroblock types shown in FIGS. 63 and 64, and is used in the decoding process. The mocroblock_pattern is data derived from the macroblock types shown in FIGS. 63 and 64, and is data indicating whether or not coded_block_pattern appears in the bitstream.
[0321] macroblock_intra is data derived from the macroblock types shown in FIGS. 63 and 64, and is used in the decoding process. spatial_temporal_weight_code_flag is the data derived from the macroblock types shown in Figures 63 and 64, and spatial_temporal_weight_code, which indicates how to upsample the lower layer image with time scalability, is the data that indicates whether or not it exists in the bitstream. Is.
[0322] frame_motion_type is a 2-bit code indicating the prediction type of the macroblock of the frame. If there are two prediction vectors and a field-based prediction type, it is "00", if there is one prediction vector and it is a field-based prediction type, it is "01", and there is one prediction vector and it is frame-based. If it is a prediction type of, it is "10", and if it is a prediction type of dial prime with one prediction vector, it is "11". field_motion_type is a 2-bit code that indicates the motion prediction of the field macroblock. If there is one prediction vector and the field-based prediction type is "01", if there are two prediction vectors and the 18x8 macroblock-based prediction type is "10", the prediction vector is 1. If the number is a dial prime prediction type, it is "11". dct_type is data indicating whether the DCT is in frame DCT mode or field DCT mode. quantiser_scale_code is data indicating the quantization step size of the macroblock.
Next, data elements related to motion vectors will be described. The motion vector is encoded as a difference with respect to the previously encoded vector in order to reduce the motion vector required during decoding. To perform motion vector decoding, the decoder must maintain four motion vector predictions (with horizontal and vertical components, respectively). This predicted motion vector is expressed as PMV [r] [s] [v]. [r] is a flag indicating whether the motion vector in the macroblock is the first vector or the second vector, and is "0" when the vector in the macroblock is the first vector. Therefore, when the vector in the macroblock is the second vector, it becomes "1". [s] is a flag indicating whether the direction of the motion vector in the macroblock is the forward direction or the backward direction. In the case of the forward motion vector, it becomes "0" and the backward motion vector. In the case of, it becomes "1". [v] is a flag indicating whether the vector component in the macroblock is horizontal or vertical. In the case of the horizontal component, it is "0", and in the case of the vertical component, it is "0". Is "1".
Therefore, PMV [0] [0] [0] represents the data of the horizontal component of the forward motion vector of the first vector, and PMV [0] [0] [1] is the first. Represents the data of the vertical component of the forward motion vector of the vector, PMV [0] [1] [0] represents the data of the horizontal component of the posterior motion vector of the first vector, PMV [0] 0] [1] [1] represents the data of the vertical component of the backward movement vector of the first vector, and PMV [1] [0] [0] represents the forward movement of the second vector. Representing the data of the horizontal component of the vector, PMV [1] [0] [1] represents the data of the vertical component of the forward motion vector of the second vector, PMV [1] [1] [0] ] Represents the data of the horizontal component of the posterior motion vector of the second vector, and PMV [1] [1] [1] is the data of the vertical component of the posterior motion vector of the second vector. Represents.
[0325] motion_vertical_field_select [r] [s] is data indicating which reference field is used in the prediction format. When this motion_vertical_field_select [r] [s] is "0", the top reference field is used, and when it is "1", the bottom reference field is used.
Therefore, motion_vertical_field_select [0] [0] indicates a reference field when generating a forward motion vector of the first vector, and motion_vertical_field_select [0] [1] indicates a backward motion of the first vector. Indicates the reference field when generating the motion vector of, motion_vertical_field_select [1] [0] indicates the reference field when generating the forward motion vector of the second vector, and motion_vertical_field_select [1] [1] indicates the reference field. , Shows the reference field when generating the backward motion vector of the second vector.
[0327] The coded_block_pattern is variable-length data indicating which DCT block has a significance coefficient (non-zero coefficient) among a plurality of DCT blocks storing the DCT coefficient. num_mv_bits is data indicating the code amount of the motion vector in the macroblock. num_coef_bits is data indicating the code amount of the DCT coefficient in the macroblock. num_other_bits is the code amount of the macroblock, and is the data indicating the code amount other than the motion vector and the DCT coefficient.
Next, the syntax for decoding each data element from the variable length history stream will be described with reference to FIGS. 47 to 67.
[0329] This variable-length history stream includes next_start_code () function, sequence_header () function, sequence_extension () function, extension_and_user_data (0) function, group_of_picture_header () function, extension_and_user_data (1) function, picture_header () function, picture_coding_extension ( ) Function, re_coding_stream_info () function, extension_and_user_data (2) function, and data elements defined by the picture_data () function.
[0330] Since the next_start_code () function is a function for searching for a start code existing in the bit stream, the head of the history stream was used in the past coding process as shown in FIG. 48. A data element that is defined by the sequence_header () function is described.
[0331] The data elements defined by the sequence_header () function are sequence_header_code, sequence_header_present_flag, horizontal_size_value, vertical_size_value, aspect_ratio_information, frame_rate_code, bit_rate_value, marker_bit, VBV_buffer_size_value, constrained_parameter_flag, load_in.
[0332] The sequence_header_code is data representing the start synchronization code of the sequence layer. sequence_header_present_flag is data indicating whether the data in sequence_header is valid or invalid. horizontal_size_value is data consisting of the lower 12 bits of the number of pixels in the horizontal direction of the image. vertical_size_value is data consisting of the lower 12 bits of the number of vertical lines in the image. The aspect_ratio_information is data representing the aspect ratio (aspect ratio) of the pixel or the display screen aspect ratio. frame_rate_code is data representing the display cycle of the image. bit_rate_value is the lower 18 bits (rounded up in 400bsp) data of the bit rate to limit the amount of generated bits.
[0333] The marker_bit is bit data inserted to prevent start code emulation. VBV_buffer_size_value is the lower 10-bit data of the value that determines the size of the virtual buffer (video buffer verifier) for controlling the generated code amount. constrained_parameter_flag is data indicating that each parameter is within the limit. load_intra_quantiser_matrix is data indicating the existence of quantization matrix data for intra MB. intra_quantiser_matrix is data indicating the value of the quantization matrix for intra MB. load_non_intra_quantiser_matrix is data indicating the existence of quantization matrix data for non-intra MB. non_intra_quantiser_matrix is data representing the value of the quantization matrix for non-intra MB.
[0334] Next to the data element defined by the sequence_header () function, the data element defined by the sequence_extension () function as shown in FIG. 49 is described as a history stream.
[0335] The data elements defined by the sequence_extension () function are extension_start_code, extension_start_code_identifier, sequence_extension_present_flag, profile_and_level_indication, progressive_sequence, chroma_format, horizontal_size_extension, vertical_size_extension, bit_rate_extension, vertical_size_extension, bit_rate_extension, vbv_buffer.
[0336] The extension_start_code is data representing the start synchronization code of the extension data. The extension_start_code_identifier is data indicating which extension data is sent. sequence_extension_present_flag is data that indicates whether the data in the sequence extension is valid or invalid. profile_and_level_indication is data for specifying the profile and level of video data. The progressive_sequence is data indicating that the video data is a sequential scan. chroma_format is data for specifying the color difference format of video data. horizontal_size_extension is the upper 2 bits of data to be added to the horizontal_size_value of the sequence header. vertical_size_extension is the upper 2 bits of data to be added to the vertical_size_value of the sequence header. bit_rate_extension is the data of the upper 12 bits to be added to bit_rate_value of the sequence header. vbv_buffer_size_extension is the upper 8 bits of data to be added to vbv_buffer_size_value in the sequence header.
[0337] low_delay is data indicating that the B picture is not included. frame_rate_extension_n is data for obtaining the frame rate in combination with the frame_rate_code of the sequence header. frame_rate_extension_d is data for obtaining the frame rate in combination with frame_rate_code in the sequence header.
[0338] Next to the data element defined by the sequence_extension () function, the data element defined by the extension_and_user_data (0) function as shown in FIG. 50 is described as a history stream. When "i" is other than 1, the extension_and_user_data (i) function does not describe the data element defined by the extension_data () function, but describes only the data element defined by the user_data () function as a history stream. .. Therefore, the extension_and_user_data (0) function describes only the data elements defined by the user_data () function as a history stream.
[0339] The user_data () function describes user data as a history stream based on the syntax as shown in FIG.
Next to the data element defined by the extension_and_user_data (0) function, the data element defined by the group_of_picture_header () function as shown in FIG. 52 and the data element defined by the extension_and_user_data (1) function are Described as a history stream. However, only when the group_start_code indicating the start code of the GOP layer is described in the history stream, the data element defined by the group_of_picture_header () function and the data element defined by the extension_and_user_data (1) function are described. ing.
The data element defined by the group_of_picture_header () function is composed of group_start_code, group_of_picture_header_present_flag, time_code, closed_gop, and broken_link.
[0342] The group_start_code is data indicating the start synchronization code of the GOP layer. group_of_picture_header_present_flag is data indicating whether the data element in group_of_picture_header is valid or invalid. time_code is a time code indicating the time from the beginning of the sequence of the first picture of the GOP. closed_gop is flag data indicating that the image in the GOP can be played independently from other GOPs. broken_link is flag data indicating that the first B picture in the GOP cannot be reproduced accurately due to editing or the like.
[0343] The extension_and_user_data (1) function, like the extension_and_user_data (0) function, describes only the data elements defined by the user_data () function as a history stream.
If there is no group_start_code indicating the start code of the GOP layer in the history stream, the data elements defined by these group_of_picture_header () and extension_and_user_data (1) functions will be included in the history stream. Not described. In that case, the data element defined by the picture_headr () function is described as a history stream next to the data element defined by the extension_and_user_data (0) function.
The data elements defined by the picture_headr () function are picture_start_code, temporary_reference, picture_coding_type, vbv_delay, full_pel_forward_vector, forward_f_code, full_pel_backward_vector, backward_f_code, extra_bit_picture, and extra_information_picture, as shown in FIG.
[0346] Specifically, picture_start_code is data representing the start synchronization code of the picture layer. temporal_reference is the data that is reset at the beginning of the GOP with a number indicating the display order of the pictures. picture_coding_type is data indicating the picture type. vbv_delay is data indicating the initial state of the virtual buffer at the time of random access. The full_pel_forward_vector is data indicating whether the accuracy of the forward motion vector is in integer units or half pixel units. The forward_f_code is data representing the forward motion vector search range. The full_pel_backward_vector is data indicating whether the accuracy of the reverse motion vector is in integer units or half pixel units. The backward_f_code is data representing the reverse motion vector search range. extra_bit_picture is a flag indicating the existence of subsequent additional information. When this extra_bit_picture is "1", extra_information_picture exists next, and when extra_bit_picture is "0", it means that there is no data following it. extra_information_picture is the information reserved in the standard.
[0347] Next to the data element defined by the picture_headr () function, the data element defined by the picture_coding_extension () function as shown in FIG. 54 is described as a history stream.
The data elements defined by this picture_coding_extension () function are extension_start_code, extension_start_code_identifier, f_code [0] [0], f_code [0] [1], f_code [1] [0], f_code [1] [ 1], intra_dc_precision, picture_structure, top_field_first, frame_predictive_frame_dct, concealment_motion_vectors, q_scale_type, intra_vlc_format, alternate_scan, repeat_firt_field, chroma_420_type, progressive_frame, composite_display_flag, v_axis, field_sequence,
[0349] The extension_start_code is a start code indicating the start of the extension data of the picture layer. extension_start_code_identifier is a code that indicates which extension data is sent. f_code [0] [0] is data representing the horizontal motion vector search range in the forward direction. f_code [0] [1] is data representing the vertical motion vector search range in the forward direction. f_code [1] [0] is data representing the horizontal motion vector search range in the backward direction. f_code [1] [1] is data representing the vertical motion vector search range in the backward direction. intra_dc_precision is data representing the precision of the DC coefficient.
[0350] The picture_structure is data indicating whether it is a frame structure or a field structure. In the case of a field structure, it is data that also indicates the upper field or the lower field. In the case of a frame structure, top_field_first is data indicating whether the first field is high or low. frame_predictive_frame_dct is data indicating that in the case of a frame structure, the prediction of the frame mode DCT is only in the frame mode. The concealment_motion_vectors are data indicating that the intra macroblock has a motion vector for concealing the transmission error. q_scale_type is data indicating whether to use a linear quantization scale or a non-linear quantization scale. intra_vlc_format is data indicating whether to use another 2D VLC for the intra macroblock.
[0351] The alternate_scan is data indicating the selection of whether to use the zigzag scan or the alternate scan. repeat_firt_field is the data used for 2: 3 pulldown. chroma_420_type is the data that represents the same value as the next progressive_frame if the signal format is 4: 2: 0, or 0 otherwise. The progressive_frame is data indicating whether or not this picture can be scanned sequentially. The composite_display_flag is data indicating whether or not the source signal is a composite signal. v_axis is the data used when the source signal is PAL. field_sequence is the data used when the source signal is PAL. sub_carrier is the data used when the source signal is PAL. burst_amplitude is the data used when the source signal is PAL. sub_carrier_phase is the data used when the source signal is PAL.
[0352] Next to the data element defined by the picture_coding_extension () function, the data element defined by the re_coding_stream_info () function is described as a history stream. This re_coding_stream_info () function is mainly used when describing a combination of historical information, and the details thereof will be described later with reference to FIG. 71.
[0353] Next to the data element defined by the re_coding_stream_info () function, the data element defined by extensions_and_user_data (2) is described as a history stream. As shown in Fig. 50, this extension_and_user_data (2) function describes the data element defined by the extension_data () function when the extension start code (extension_start_code) exists in the bitstream. Next to this data element, if the user data start code (user_data_start_code) exists in the bitstream, the data element defined by the user_data () function is described. However, if the extension start code and user data start code do not exist in the bitstream, the data elements defined by the extension_data () and user_data () functions are not described in the bitstream.
[0354] As shown in FIG. 55, the extension_data () function is a bit stream of a data element indicating extension_start_code and a data element defined by the quant_matrix_extension () function, the copyright_extension () function, and the picture_display_extension () function. It is a function to describe as a history stream inside.
[0355] Data elements defined by the quant_matrix_extension () function, as shown in FIG. 56, extension_start_code, extension_start_code_identifier, quant_matrix_extension_present_flag, load_intra_quantiser_matrix, intra_quantiser_matrix [64], load_non_intra_quantiser_matrix, non_intra_quantiser_matrix [64], load_chroma_intra_quantiser_matrix, chroma_intra_quantiser_matrix [64], load_chroma_non_intra_quantiser_matrix , And chroma_non_intra_quantiser_matrix [64].
[0356] extension_start_code is a start code indicating the start of this quantization matrix extension. extension_start_code_identifier is a code that indicates which extension data is sent. quant_matrix_extension_present_flag is data to indicate whether the data element in this quantized matrix extension is valid or invalid. load_intra_quantiser_matrix is data indicating the existence of quantization matrix data for intra macroblocks. intra_quantiser_matrix is data indicating the value of the quantization matrix for the intra macroblock.
[0357] load_non_intra_quantiser_matrix is data indicating the existence of quantization matrix data for non-intra macroblocks. non_intra_quantiser_matrix is data representing the value of the quantization matrix for non-intra macroblocks. load_chroma_intra_quantiser_matrix is data indicating the existence of quantization matrix data for the color difference intra macroblock. chroma_intra_quantiser_matrix is data indicating the value of the quantization matrix for the color difference intra macroblock. load_chroma_non_intra_quantiser_matrix is data indicating the existence of quantization matrix data for color difference non-intra macroblocks. chroma_non_intra_quantiser_matrix is data indicating the value of the quantization matrix for the color difference non-intra macroblock.
[0358] As shown in FIG. 57, the data element defined by the copyright_extension () function is composed of extension_start_code, extension_start_code_itentifier, copyright_extension_present_flag, copyright_flag, copyright_identifier, original_or_copy, copyright_number_1, copyright_number_2, and copyright_number_3.
[0359] extension_start_code is a start code indicating the start of the copyright extension. extension_start_code_itentifier Code that indicates which extension data will be sent. copyright_extension_present_flag is data to indicate whether the data element in this copyright extension is valid or invalid.
[0360] copyright_flag indicates whether or not copy rights are granted to the encoded video data until the next copyright extension or sequence end. The copyright_identifier is data for identifying the registration authority of the copy right specified by ISO / IEC JTC / SC29. The original_or_copy is data indicating whether the data in the bitstream is original data or copy data. copyright_number_1 is data representing bits 44 to 63 of the copyright number. copyright_number_2 is data representing bits 22 to 43 of the copyright number. copyright_number_3 is data representing bits 0 to 21 of the copyright number.
[0361] As shown in FIG. 58, the data elements defined by the picture_display_extension () function are extension_start_code_identifier, frame_center_horizontal_offset, frame_center_vertical_offset, and the like.
[0362] The extension_start_code_identifier is a code indicating which extension data is sent. frame_center_horizontal_offset is data indicating the horizontal offset of the display area, and can define the number of offset values defined by number_of_frame_center_offsets. frame_center_vertical_offset is data indicating the vertical offset of the display area, and can define the number of offset values defined by number_of_frame_center_offsets.
[0363] Returning to FIG. 47 again, the data element defined by the picture_data () function is described as a history stream next to the data element defined by the extension_and_user_data (2) function. However, this picture_data () function exists when red_bw_flag is not 1 or red_bw_indicator is 2 or less. The red_bw_flag and red_bw_indicator are described in the re_coding_stream_info () function, which will be described later with reference to FIGS. 71 and 72.
[0364] The data element defined by the picture_data () function is the data element defined by the slice () function, as shown in FIG. 59. At least one data element defined by this slice () function is described in the bitstream.
The slice () function is defined by a data element such as slice_start_code, slice_quantiser_scale_code, intra_slice_flag, intra_slice, reserved_bits, extra_bit_slice, extra_information_slice, and extra_bit_slice, and a data element such as extra_bit_slice, as shown in FIG. , A function to describe as a history stream.
[0366] slice_start_code is a start code indicating the start of a data element defined by the slice () function. The slice_quantiser_scale_code is data indicating the quantization step size set for the macroblock existing in this slice layer. However, if quantizer_scale_code is set for each macroblock, the data of macroblock_quantiser_scale_code set for each macroblock is preferentially used.
[0367] The intra_slice_flag is a flag indicating whether or not intra_slice and reserved_bits are present in the bit stream. intra_slice is data indicating whether or not a non-intra macroblock exists in the slice layer. If any of the macroblocks in the slice layer is a non-intra macroblock, intra_slice will be "0", and if all the macroblocks in the slice layer are non-intra macroblocks, intra_slice will be "1". Become. reserved_bits is 7-bit data and takes a value of "0". extra_bit_slice is a flag indicating that additional information exists as a history stream, and is set to "1" if extra_information_slice exists next. Set to "0" if no additional information exists.
[0368] Next to these data elements, the data elements defined by the macroblock () function are described as a history stream.
[0369] As shown in FIG. 61, the macroblock () function is defined by data elements such as macroblock_escape, macroblock_address_increment, macroblock_quantiser_scale_code, and marker_bit, and the macroblock_modes () function, motion_vectors (s) function, and code_block_pattern () function. It is a function to describe the data element.
[0370] macroblock_escape is a fixed bit string indicating whether or not the horizontal difference between the reference macroblock and the previous macroblock is 34 or more. If the horizontal difference between the reference macroblock and the previous macroblock is 34 or more, add 33 to the macroblock_address_increment value. macroblock_address_increment is data showing the horizontal difference between the reference macroblock and the previous macroblock. If there is one macroblock_escape before this macroblock_address_increment, the value obtained by adding 33 to the value of this macroblock_address_increment is the data showing the horizontal difference between the actual reference macroblock and the previous macroblock. ..
[0371] macroblock_quantiser_scale_code is a quantization step size set for each macroblock, and exists only when macroblock_quant is "1". Slice_quantiser_scale_code indicating the quantization step size of the slice layer is set in each slice layer, but when macroblock_quantiser_scale_code is set for the reference macroblock, this quantization step size is selected.
[0372] Next to macroblock_address_increment, data elements defined by the macroblock_modes () function are described. As shown in FIG. 62, the macroblock_modes () function is a function for describing data elements such as macroblock_type, frame_motion_type, field_motion_type, and dct_type as a history stream.
[0373] macroblock_type is data indicating the coding type of macroblock. The details will be described later with reference to FIGS. 65 to 67.
If macroblock_motion_forward or macroblock_motion_backward is "1", the picture structure is a frame, and frame_pred_frame_dct is "0", a data element representing frame_motion_type is described next to the data element representing macroblock_type. ing. Note that this frame_pred_frame_dct is a flag indicating whether or not frame_motion_type exists in the bitstream.
[0375] frame_motion_type is a 2-bit code indicating the prediction type of the macroblock of the frame. If there are two prediction vectors and a field-based prediction type, it is "00", if there is one prediction vector and it is a field-based prediction type, it is "01", and there is one prediction vector and it is frame-based. If it is a prediction type of, it is "10", and if it is a prediction type of dial prime with one prediction vector, it is "11".
[0376] If the condition for describing the frame_motion_type is not satisfied, the data element representing the field_motion_type is described next to the data element representing the macroblock_type.
[0377] field_motion_type is a 2-bit code indicating the motion prediction of the macroblock of the field. If there is one prediction vector and the field-based prediction type is "01", if there are two prediction vectors and the 18x8 macroblock-based prediction type is "10", the prediction vector is 1. If the number is a dial prime prediction type, it is "11".
[0378] If the picture structure is a frame, frame_pred_frame_dct indicates that frame_motion_type exists in the bitstream, and frame_pred_frame_dct indicates that dct_type exists in the bitstream, data representing macroblock_type. Next to the element, the data element representing dct_type is described. The dct_type is data indicating whether the DCT is in the frame DCT mode or the field DCT mode.
Returning to FIG. 61 again, if the reference macroblock is either a forward prediction macroblock, or the reference macroblock is an intramacroblock and is either a concealed macroblock, motion_vectors. (0) The data element defined by the function is described. If the reference macroblock is a backward prediction macroblock, the data element defined by the motion_vectors (1) function is described. The motion_vectors (0) function is a function for describing the data element related to the first motion vector, and the motion_vectors (1) function is a function for describing the data element related to the second motion vector. Is.
[0380] The motion_vectors (s) function is a function for describing data elements related to motion vectors, as shown in FIG. 63.
[0381] If there is only one motion vector and the dial prime prediction mode is not used, the data elements defined by motion_vertical_field_select [0] [s] and motion_vector (0, s) are described.
[0382] In this motion_vertical_field_select [r] [s], whether the first motion vector (which may be either the forward or backward vector) is a vector created by referring to the bottom field or the top field. It is a flag indicating whether the vector is created by referring to. This index "r" is an index indicating whether the vector is the first vector or the second vector, and "s" is whether the prediction direction is forward or backward prediction. It is an index indicating whether or not.
[0383] As shown in FIG. 64, the motion_vector (r, s) function includes a data string related to motion_code [r] [s] [t], a data string related to motion_residual [r] [s] [t], and a data string related to motion_residual [r] [s] [t]. It is a function to describe the data representing dmvector [t].
[0384] motion_code [r] [s] [t] is variable length data representing the magnitude of the motion vector in the range of -16 to +16. motion_residual [r] [s] [t] is variable-length data representing the residuals of the motion vector. Therefore, a detailed motion vector can be described by the values of this motion_code [r] [s] [t] and motion_residual [r] [s] [t]. dmvector [t] already exists depending on the time distance to generate a motion vector in one field (for example, the top field is one field relative to the bottom field) in dual prime prediction mode. This is data in which the motion vector of is scaled and correction is performed in the vertical direction in order to reflect the vertical deviation between the lines of the top field and the bottom field. This index "r" is an index indicating whether the vector is the first vector or the second vector, and "s" is whether the prediction direction is forward or backward prediction. It is an index indicating whether or not. S is data indicating whether the motion vector is a vertical component or a horizontal component.
[0385] The motion_vector (r, s) function shown in FIG. 64 first describes a data string representing the horizontal motion_coder [r] [s] [0] as a history stream. Since the number of bits of both motion_residual [0] [s] [t] and motion_residual [1] [s] [t] is indicated by f_code [s] [t], f_code [s] [t] is not 1. In some cases, it indicates that motion_residual [r] [s] [t] exists in the bitstream. The fact that the horizontal component motion_residual [r] [s] [0] is not "1" and the horizontal component motion_code [r] [s] [0] is not "0" means that motion_residual in the bit stream. It means that there is a data element representing [r] [s] [0] and there is a horizontal component of the motion vector. In that case, the horizontal component motion_residual [r] [s] ] A data element representing [0] is described.
Subsequently, a data string representing a vertical motion_coder [r] [s] [1] is described as a history stream. Similarly, the number of bits of both motion_residual [0] [s] [t] and motion_residual [1] [s] [t] is indicated by f_code [s] [t], so f_code [s] [t] If is not 1, it means that motion_residual [r] [s] [t] exists in the bitstream. The fact that motion_residual [r] [s] [1] is not "1" and motion_code [r] [s] [1] is not "0" means that motion_residual [r] [s] [1] is not in the bitstream. ], Which means that there is a vertical component of the motion vector. In that case, the data element that represents the motion_residual [r] [s] [1] of the vertical component. Is described.
Next, macroblock_type will be described with reference to FIGS. 65 to 67. macroblock_type is variable length data generated from flags such as macroblock_quant, dct_type_flag, macroblock_motion_forward, and macroblock_motion_backward. macroblock_quant has a flag indicating whether macroblock_quantiser_scale_code for setting the quantization step size is set for the macroblock, and if macroblock_quantiser_scale_code exists in the bitstream, macroblock_quant is a value of "1". I take the.
[0388] The dct_type_flag is a flag (in other words, a flag indicating whether or not DCT is performed) for indicating whether or not there is a dct_type indicating whether or not the reference macroblock is encoded by the frame DCT or the field DCT. If dct_type exists in the bit stream, this dct_type_flag takes a value of "1". macroblock_motion_forward is a flag indicating whether or not the reference macroblock is forward-predicted, and if it is forward-predicted, it takes a value of "1". macroblock_motion_backward is a flag indicating whether or not the reference macroblock is backward predicted, and if it is backward predicted, it takes a value of "1".
[0389] In the variable length format, the history information can be reduced in order to reduce the bit rate to be transmitted.
[0390] That is, when macroblock_type and motion_vectors () are transferred but quantizer_scale_code is not transferred, the bit rate can be reduced by setting slice_quantiser_scale_code to "00000".
[0391] Further, when only macroblock_type is transferred and motion_vectors (), quantizer_scale_code, and dct_type are not transferred, the bit rate can be reduced by using "not coded" as macroblock_type.
[0392] Furthermore, when only the picture_coding_type is transferred and all the information below slice () is not transferred, the bit rate can be reduced by using picture_data () which does not have slice_start_code.
[0393] In the above, in order to prevent the continuous "0" of 23 bits in user_data from appearing, "1" is inserted every 22 bits, but it does not have to be every 22 bits. .. It is also possible to check Byte_allign and insert it instead of counting the number of consecutive "0" and inserting "1".
[0394] Further, in MPEG, the generation of 23-bit continuous "0" is prohibited, but in reality, the problem is only when 23 bits are continuous from the beginning of the byte, not the beginning of the byte. , If 0 is 23 bits in a row from the middle, it does not matter. Therefore, for example, "1" may be inserted at a position other than the LSB every 24 bits.
[0395] Further, in the above, the history information is in a format close to the video elementary stream, but may be in a format close to the packetized elementary stream or the transport stream. Also, the location of user_data in Elementary Stream is before picture_data, but it can be in another location.
[0396] In the transcoder 101 of FIG. 15, the coding parameters for four generations are output as history information in the subsequent stage, but in reality, not all of the history information is required, and each application. The history information required for the application will be different. In addition, the actual transmission line or recording medium (transmission medium) has a limited capacity, and even if it is compressed, if all the history information is transmitted, the capacity becomes a burden, and as a result, it becomes a burden. The bit rate of the image bit stream is suppressed, and the effectiveness of history information transmission is impaired.
[0397] Therefore, a descriptor describing a combination of items to be transmitted as history information is incorporated into the history information and transmitted to the subsequent stage, and information corresponding to various applications is provided instead of transmitting all the history information. It can be transmitted. FIG. 68 shows a configuration example of the transcoder 101 in such a case.
[0398] In FIG. 68, the parts corresponding to the case in FIG. 15 are designated by the same reference numerals, and the description thereof will be omitted as appropriate. In the configuration example of FIG. 68, the coding parameter selection circuit 501 is inserted between the history information separating device 105 and the coding device 106, and between the history encoding device 107 and the coding device 106.
The coding parameter selection circuit 501 outputs the coding parameter calculation unit 512 that calculates the coding parameter from the baseband video signal output by the history information separating device 105, and the transcoder 101 output by the history information separating device 105. From the information on the coding parameters determined to be optimal for coding (eg, second generation coding parameters), the coding parameters and descriptors (red_bw_flag, red_bw_indicator) (see FIG. 72), which will be described later. ) Is separated from the combination descriptor separation unit 511, the coding parameter output by the coding parameter calculation unit 512, and the coding parameter output by the combination descriptor separation unit 511. It has a switch 513 that selects corresponding to the descriptor separated in part 511 and outputs it to the encoding device 106. Other configurations are the same as in FIG. 15.
[0400] Here, a combination of items to be transmitted as history information will be described. The history information can be classified into picture unit information and macroblock unit information. Information in slice units can be obtained by collecting information in macroblocks contained therein, and information in units of GOP can be obtained by collecting information in picture units included in it.
[0401] Since the information in the picture unit is transmitted only once for each frame, the bit rate in the information transmission is not so large. On the other hand, since information in macroblock units is transmitted for each macroblock, for example, in the case of a video system with 525 scanning lines per frame and a field rate of 60 fields / second, the number of pixels per frame Assuming that is 720 x 480, information in macroblock units needs to be transmitted 1350 (= (720/16) x (480/16)) times per frame. Therefore, a considerable part of the history information is occupied by the information for each macroblock. Therefore, as the history information, at least the information in the picture unit is always transmitted, but the information in the macroblock unit is selected and transmitted according to the application, so that the amount of information to be transmitted can be suppressed.
[0402] The macroblock unit information transferred as history information includes, for example, num_coef_bits, num_mv_bits, num_other_bits, q_scale_code, q_scale_type, motion_type, mv_vert_field_sel [] [], mv [] [] [], mb_mfwd, mb_mbwd, mb_pattern, coded. , mb_intra, slice_start, dct_type, mb_quant, skipped_mb, etc. These are expressed using the elements of macroblock rate information.
[0403] num_coef_bits represents the code amount required for the DCT coefficient among the code amounts of the macroblock. num_mv_bits represents the code amount required for the motion vector among the code amounts of the macroblock. num_other_bits represents the code amount of macroblock other than num_coef_bits and num_mv_bits.
[0404] q_scale_code represents the q_scale_code applied to the macroblock. motion_type represents the type of motion vector applied to the macroblock. mv_vert_field_sel [] [] represents the field select of the motion vector applied to the macroblock.
[0405] mv [] [] [] represents the motion vector applied to the macroblock. mb_mfwd is a flag indicating that the prediction mode of macroblock is forward prediction. mb_mbwd is a flag indicating that the prediction mode of macroblock is backward prediction. mb_pattern is a flag indicating the presence or absence of a non-zero DCT coefficient of macroblock.
[0406] The coded_block_pattern is a flag indicating for each DCT block whether or not the macroblock has a non-zero DCT coefficient. mb_intra is a flag that indicates whether the macroblock is intra_macro or not. slice_start is a flag that indicates whether macroblock is the beginning of slice. dct_type is a flag that indicates whether the macroblock is field_dct or flame_dct.
[0407] mb_quant is a flag indicating whether or not macroblock transmits quantizer_scale_code. skipped_mb is a flag indicating whether or not the macroblock is a skipped macroblock.
[0408] Not all of these items are always required, and the required items vary depending on the application. For example, items such as num_coef_bits and slice_start are needed in applications that have a transparent requirement to restore the re-encoded bitstream to its original form as much as possible. In other words, these items are not necessary for applications that change the bit rate. In addition, there are applications in which it is only necessary to know the coding type of each picture when the transmission line is very restricted. From such a situation, as an example of the combination of items for transmitting history information, for example, the combination shown in FIG. 69 can be considered.
[0409] In FIG. 69, the value "2" corresponding to the item in each combination means that the information exists and is available, and "0" means that the information does not exist. To do. "1" indicates that the information itself has no meaning, for the purpose of assisting the existence of other information, or for the purpose of syntactically existing but not related to the original bitstream information. .. For example, slice_start is "1" in the macroblock at the beginning of slice when transmitting history information, but if slice is not necessarily in the same positional relationship with the original bit stream, history It becomes meaningless as information.
[0410] In the example of FIG. 69, (num_coef_bits, num_mv_bits, num_other_bits), (q_scale_code, q_scale_type), (motion_type, mv_vert_field_sel [] [], mv [] [] []), (mb_mfwd, mb_mbwd), (mb_pattern) ), (Coded_block_pattern), (mb_intra), (slice_start), (dct_type), (mb_quant), (skipped_mb) 5 combinations of combinations 1 to 5 are prepared depending on the presence or absence of each item.
Combination 1 is a combination intended to reconstruct a completely transparent bitstream. According to this combination, highly accurate transcoding can be realized by using the generated code amount information. Combination 2 is also a combination aimed at reconstructing a completely transparent bitstream. Combination 3 is a combination that allows a completely transparent bitstream to be reconstructed, but a visually nearly transparent bitstream to be reconstructed. Combination 4 is inferior to Combination 3 in terms of transparent, but is a combination that can reconstruct a bitstream that does not cause any visual problems. Combination 5 is inferior to Combination 4 in terms of transparency, but it is a combination that can reconstruct the bitstream incompletely with less historical information.
[0412] Of these combinations, the smaller the number of the combination number, the higher the functionality, but the larger the capacity required for transferring the history. Therefore, it is necessary to determine the combination to be transmitted by considering the assumed application and the capacity that can be used for the history.
Next, the operation of the transcoder 101 of FIG. 68 will be described with reference to the flowchart of FIG. 70. In step S41, the decoding device 102 of the transcoder 101 decodes the input bit stream, extracts the coding parameter (4th) used in encoding the bit stream, and extracts the coding parameter (4th). ) Is output to the history information multiplexing device 103, and the decoded video data is also output to the history information multiplexing device 103. In step S42, the decoding device 102 also extracts user_data from the input bitstream and outputs it to the history decoding device 104. The history decoding device 104 extracts the combination information (descriptor) from the input user_data in step S43, and further extracts the coding parameters (1st, 2nd, 3rd) as the history information by using the combination information (descriptor). , Output to the history information multiplexing device 103.
[0414] In step S44, the history information multiplexing device 103 includes the current coding parameter (4th) supplied from the decoding device 102 retrieved in step S41 and the past output by the history decoding device 104 in step S43. The coding parameters (1st, 2nd, 3rd) of the above are multiplexed with the baseband video data supplied from the decoding device 102 according to the format shown in FIG. 18 or FIG. 31, and output to the history information separator 105. ..
[0415] In step S45, the history information separator 105 extracts a coding parameter from the baseband video data supplied from the history information multiplexing device 103, and the code most suitable for the current coding is extracted from the coding parameters. The conversion parameter (for example, the second generation coding parameter) is selected and output to the combination descriptor separation unit 511 together with the descriptor. Further, when the history information separator 105 determines that a coding parameter other than the coding parameter determined to be optimal for the current coding (for example, the optimum coding parameter is the second generation coding parameter). Other 1st generation, 3rd generation, and 4th generation coding parameters) are output to the history encoding device 107. The history encoding device 107 describes the coding parameter input from the history information separating device 105 in user_data in step S46, and outputs the user_data (converted_history_stream ()) to the coding device 106.
[0416] The combination descriptor separator 511 of the coding parameter selection circuit 501 separates the coding parameter and the descriptor from the data supplied by the history information separator 105, and sets the coding parameter (2nd) of the switch 513. Supply to one contact. To the other contact of the switch 513, the coding parameter calculation unit 512 calculates and supplies the coding parameter from the baseband video data output by the history information separator 105. The switch 513 corresponds to the descriptor output by the combination descriptor separation unit 511 in step S48, and the coding parameter output by the combination descriptor separation unit 511 or the coding parameter output by the coding parameter calculation unit 512. Select one of the above and output it to the encoding device 106. That is, in the switch 513, when the coding parameter supplied from the combination descriptor separation unit 511 is valid, the coding parameter output by the combination descriptor separation unit 511 is selected, but the combination descriptor separation unit 511 When it is determined that the coding parameter output by 511 is invalid, the coding parameter calculated by the coding parameter calculation unit 512 processing the baseband video is selected. This selection is made in response to the capacity of the transmission medium.
[0417] In step S49, the coding device 106 encodes the baseband video signal supplied from the history information separator 105 based on the coding parameters supplied from the switch 513. Further, in step S50, the coding device 106 multiplexes and outputs the user_data supplied from the history encoding device 107 to the coded bit stream.
[0418] In this way, even when the combination of coding parameters obtained by each history is different, transcoding can be performed without any problem.
[0419] As shown in FIG. 38, the history information is transmitted by history_stream () (more accurately, converted_history_stream ()) as a kind of user_data () function of the video stream. The syntax of the history_stream () is as shown in Figure 47. Descriptors (red_bw_flag, red_bw_indicator) representing combinations of historical information items, and items not supported by MPEG streams (num_other_bits, num_mv_bits, num_coef_bits) are transmitted by the re_coding_stream_info () function in FIG. 47.
[0420] As shown in FIG. 71, the re_coding_stream_info () function is composed of data elements such as user_data_start_code, re_coding_stream_info_ID, red_bw_flag, red_bw_indicator, marker_bit, num_other_bits, num_mv_bits, num_coef_bits.
[0421] user_data_start_code is a start code indicating that user_data starts. re_coding_stream_info_ID is a 16-bit integer used to identify the re_coding_stream_info () function. Specifically, the value is set to "1001 0001 1110 1100" (0x91ec).
[0422] red_bw_flag is a 1-bit flag, which is set to 0 when the history information transmits all items, and when the value of this flag is 1, by examining the red_bw_indicator following this flag, the figure is shown. It is possible to determine which of the five combinations shown in 69 is used to send the item.
[0423] red_bw_indicator is a 2-bit integer, and the combination of items is described as shown in FIG. 72.
That is, of the five combinations shown in FIG. 69, red_bw_flag is set to 0 in the case of combination 1, and red_bw_flag is set to 1 in the case of combinations 2 to 5. On the other hand, red_bw_indicator is 0 for combination 2, 1 for combination 3, 2 for combination 4, and 3 for combination 5.
[0425] Therefore, red_bw_indicator is specified when red_bw_flag is 1 (in the case of combination 2 to 5).
Further, as shown in FIG. 71, when red_bw_flag is 0 (in the case of combination 1), marker_bit, num_other_bits, num_mv_bits, num_coef_bits are described for each macroblock. These four data elements are not specified for combination 2 to 5 (when red_bw_flag is 1).
As shown in FIG. 59, the picture_data () function is composed of one or more slice () functions. However, in the case of combination 5, syntax elements below it, including the picture_data () function, are not transmitted (Fig. 69). In this case, the history information is intended to transmit information in units of pictures such as picture_type.
[0428] In the case of combination 1 to 4, the slice () function shown in FIG. 60 exists. However, the position information of slice determined by this slice () function and the position information of slice of the original bitstream depend on the combination of the items of the history information. In the case of combination 1 or combination 2, the slice position information of the bitstream that is the source of the history information and the slice position information determined by the slice () function must be the same.
[0429] The syntax element of the macroblock () function shown in FIG. 61 depends on the combination of items in the history information. The macroblock_escape, macroblock_address_increment, macroblock_modes () functions are always present. However, the informative validity of macroblock_escape and macroblock_address_increment is determined by the combination. If the combination of historical information items is combination 1 or combination 2, the same skipped_mb information of the original bitstream must be transmitted.
[0430] In the case of combination 4, the motion_vectors () function does not exist. In the case of combination 1 to 3, the existence of the motion_vectors () function is determined by the macroblock_type of the macroblock_modes () function. For combination 3 or 4, the coded_block_pattern () function does not exist. For combination 1 and combination 2, the macroblock_type of the macroblock_modes () function determines the existence of the coded_block_pattern () function.
[0431] The syntax element of the macroblock_modes () function shown in FIG. 62 depends on the combination of items in the history information. macroblock_type is always present. If the combination is combination 4, flame_motion_type, field_motion_type, dct_type do not exist.
[0432] The effectiveness of the parameter obtained from macroblock_type as information is determined by the combination of items in the history information.
If the combination of historical information items is combination 1 or combination 2, macroblock_quant must be the same as the original bitstream. For combination 3 or 4, macroblock_quant represents the presence of quantizer_scale_code in the macroblock () function and does not have to be the same as the original bitstream.
[0434] If the combination is combination 1 to combination 3, macroblock_motion_forward and macroblock_motion_backward must be the same as the original bitstream. If the combination is combination 4 or combination 5, it is not necessary.
If the combination is combination 1 or combination 2, macroblock_pattern must be identical to the original bitstream. For combination 3, macroblock_pattern is used to indicate the presence of dct_type . When the combination is combination 4, the relationship as in the case of combination 1 to 3 does not hold.
[0436] When the combination of the items of the history information is combination 1 to combination 3, macroblock_intra needs to be the same as the original bit stream. In the case of combination 4, this is not the case.
[0437] History_stream () in FIG. 47 is a syntax when the history information has a variable length, but as shown in FIGS. 40 to 46, when the history information has a fixed length syntax, the fixed length history information. The descriptors (red_bw_flag and red_bw_indicator) as information indicating which of the items to be transmitted are valid are superimposed on the baseband image and transmitted. As a result, by examining this descriptor, it is possible to determine that the field exists but its contents are invalid.
Therefore, as shown in FIG. 44, user_data_start_code, re_coding_stream_info_ID, red_bw_flag, red_bw_indicator, marker_bit are arranged as re_coding_stream_information. The meaning of each is the same as in FIG. 71.
[0439] By transmitting the elements of the coding parameters to be transmitted as the history in a combination according to the application, it is possible to transmit the history according to the application with an appropriate amount of data.
[0440] As described above, when the history information is transmitted as a variable length code, the re_coding_stream_info () function is configured as shown in FIG. 71 and is transmitted as a part of the history_stream () function as shown in FIG. 47. Will be done. On the other hand, when the history information is transmitted as a fixed length code, re_coding_stream_information () is transmitted as a part of the history_stream () function as shown in FIG. 44. In the example of FIG. 44, user_data_start_code, re_coding_stream_info_ID, red_bw_flag, red_bw_indicator are transmitted as re_coding_stream_information.
Further, the Re_Coding information Bus macroblock format as shown in FIG. 73 is defined for the transmission of the history information in the baseband signal output by the history information multiplexing device 103 of FIG. 68. This macroblock consists of 16 x 16 (= 256) bits. Then, in FIG. 73, the 32 bits shown in the third and fourth lines from the top are designated as picrate_element. Picture rate elements shown in FIGS. 74 to 76 are described in this picrate_element. The 1-bit red_bw_flag is specified in the second line from the top of FIG. 74, and the 3-bit red_bw_indicator is specified in the third line. That is, these flags red_bw_flag and red_bw_indicator are transmitted as picrate_element in FIG. 73.
Explaining the other data in FIG. 73, SRIB_sync_code is a code indicating that the first line of the macroblock of this format is left-justified, specifically set to "11111". Will be done. fr_fl_SRIB is set to 1 if picture_structure is a frame picture structure (its value is "11"), indicating that the Re_Coding Information Bus macroblock is transmitted over 16 lines, and picture_structure is not a frame structure. If set to 0, it means that the Re_Coding Information Bus will be transmitted over 16 lines. This mechanism locks the Re_Coding Information Bus to the corresponding pixels of the spatially and temporally decoded video frame or field.
SRIB_top_field_first is set to the same value as top_field_first held in the original bitstream and represents the temporal alignment of the Re_Coding Information Bus of the associated video with repeat_first_field. SRIB_repeat_first_field is set to the same value as repeat_first_field held in the original bitstream. The contents of the Re_Coding Information Bus in the first field need to be repeated as shown by this flag.
422_420_chroma represents whether the original bitstream is 4: 2: 2 or 4: 2: 0. A value of 0 indicates that the bitstream is 4: 2: 0 and the color difference signal is upsampled so that a 4: 2: 2 video is output. A value of 0 indicates that the color difference signal filtering process has not been executed.
Rolling_SRIB_mb_ref represents a 16-bit modulo 65521, which value is incremented for each macroblock. This value must be continuous across frames in the frame picture structure. Otherwise, this value must be contiguous across the field. This value is initialized to a given value between 0 and 65520. This allows the recorder system to incorporate a unique Re_Coding Information Bus identifier.
[0446] Since the meanings of the other data of the Re_Coding Information Bus macroblock are as described above, they are omitted here.
As shown in FIG. 77, the 256-bit Re_Coding Information Bus data in FIG. 73 is the LSB of the color difference data, one bit at a time, Cb [0] [0], Cr [0] [0], Cb. It is placed in [1] [0], Cr [1] [0]. Since 4-bit data can be transmitted by the format shown in FIG. 77, the 256-bit data in FIG. 73 can be transmitted by sending 64 (= 256/4) formats in FIG. 77.
[0448] According to the transcoder of the present invention, since the coding parameters generated in the past coding processing are reused in the current coding processing, the decoding processing and the coding processing are repeated. However, the image quality does not deteriorate. That is, it is possible to reduce the accumulation of image quality deterioration due to the repetition of the decoding process and the coding process.
[0449] According to the transcoder of the present invention, the coding parameters generated in the past coding process are described in the user data area of the coded stream generated in the current coding process, and are generated. Since the bitstream is a coded stream conforming to the MPEG standard, any existing decoder can perform decoding processing. Furthermore, according to the transcoder of the present invention, it is not necessary to provide a dedicated line for transmitting the coding parameters in the past coding processing, so that the conventional data stream transmission environment can be used as it is. , Past coding parameters can be transmitted.
[0450] According to the transcoder of the present invention, the coding parameters generated in the past coding process are selectively described in the coding stream generated in the current coding process. , The past coding parameters can be transmitted without extremely increasing the bit rate of the output bitstream.
[0451] According to the transcoder of the present invention, the optimum coding parameter for the current coding process is selected from the past coding parameters and the current coding parameter to perform the coding process. Therefore, even if the decoding process and the coding process are repeated, the deterioration of the image quality is not accumulated.
[0452] According to the transcoder of the present invention, the optimum coding parameter for the current coding process is selected from the past coding parameters according to the picture type to perform the coding process. Therefore, even if the decoding process and the coding process are repeated, the deterioration of the image quality is not accumulated.
[0453] According to the transcoder of the present invention, it is determined whether or not to reuse the past coding parameter based on the picture type included in the past coding parameter, so that the optimum coding process is performed. It can be performed.
[0454] The computer program that performs each of the above processes is provided by recording on a recording medium such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, and is transmitted via a network such as the Internet or a digital satellite. It can be provided by recording on a user's recording medium.
[Effect of the Invention]<u style="single"> According to the first aspect of the present invention, a coded stream can be converted into a recoded stream, in particular a combination information indicating a selective combination of history coding parameters and a history corresponding to a combination of combination information. Since the coding parameters are output together with the re-encoding stream, it is possible to transmit a stream that can suppress the deterioration of the image due to the re-encoding.</u>【0456】<u style="single">According to the second aspect of the present invention, the coded stream can be converted into a recoded stream, in particular the current coded parameters are calculated, the history coded parameters, and the calculated current coded. Since the combination information indicating the selective combination of parameters and the history coding parameter corresponding to the combination of the combination information are output together with the recoded stream, the stream capable of suppressing the deterioration of the image due to the recoding can be displayed. It becomes possible to transmit.</u>【0457】<u style="single">According to the third aspect of the present invention, the coded stream can be output, and in particular, the combination information indicating the selective combination of the history coding parameters according to the application using the coded stream, and the combination information. Since the history coding parameters corresponding to the combination of are output together with the coded stream, it is possible to transmit a stream capable of suppressing image deterioration due to recoding.</u><u style="single">According to the fourth aspect of the present invention, the coded stream can be output, and in particular, the history coding parameter according to the capacity of the transmission line for transmitting the coded stream or the recording medium for recording the coded stream. Since the combination information indicating the selective combination and the history coding parameter corresponding to the combination of the combination information are output together with the coded stream, the stream capable of suppressing the deterioration of the image due to the recoding is transmitted. Is possible.</u>BRIEF DESCRIPTION OF THE DRAWINGS [FIG. 1] FIG. 1 is a diagram illustrating a principle of high-efficiency coding.
FIG. 2 is a diagram illustrating a picture type in the case of compressing image data.
FIG. 3 is a diagram illustrating a picture type in the case of compressing image data.
FIG. 4 is a diagram illustrating a principle of encoding a moving image signal.
FIG. 5 is a block diagram showing a configuration of a device that encodes and decodes a moving image signal.
FIG. 6 is a diagram illustrating a structure of image data.
FIG. 7 is a block diagram showing a configuration of an encoder 18 of FIG.
8 is a diagram illustrating the operation of the prediction mode switching circuit 52 of FIG. 7. FIG.
9 is a diagram illustrating the operation of the prediction mode switching circuit 52 of FIG. 7. FIG.
10 is a diagram illustrating the operation of the prediction mode switching circuit 52 of FIG. 7. FIG.
11 is a diagram illustrating the operation of the prediction mode switching circuit 52 of FIG. 7. FIG.
12 is a block diagram showing a configuration of the decoder 31 of FIG. 5. FIG.
FIG. 13 is a diagram illustrating SNR control corresponding to a picture type.
FIG. 14 is a block diagram showing a configuration of a transcoder 101 to which the present invention is applied.
15 is a block diagram showing a more detailed configuration of the transcoder 101 of FIG. 14. FIG.
16 is a block diagram showing a configuration of a decoder 111 built in the decoding device 102 of FIG. 14. FIG.
FIG. 17 is a diagram illustrating pixels of a macroblock.
FIG. 18 is a diagram illustrating an area in which coding parameters are recorded.
FIG. 19 is a block diagram showing a configuration of an encoder 121 incorporated in the coding device 106 of FIG.
FIG. 20 is a block diagram showing a configuration example of the history VLC211 of FIG.
FIG. 21 is a block diagram showing a configuration example of the history VLD 203 of FIG.
22 is a block diagram showing a configuration example of the converter 212 of FIG. 15. FIG.
FIG. 23 is a block diagram showing a configuration example of the staff circuit 323 of FIG. 22.
FIG. 24 is a timing chart illustrating the operation of the converter 212 of FIG.
FIG. 25 is a block diagram showing a configuration example of the converter 202 of FIG.
FIG. 26 is a block diagram showing a configuration example of the delay circuit 343 of FIG. 25.
FIG. 27 is a block diagram showing another configuration example of the converter 212 of FIG.
FIG. 28 is a block diagram showing another configuration example of the converter 202 of FIG.
FIG. 29 is a block diagram showing a configuration example of the user data formatter 213 of FIG.
FIG. 30 is a diagram showing a state in which the transcoder 101 of FIG. 14 is actually used.
FIG. 31 is a diagram illustrating an area in which coding parameters are recorded.
32 is a flowchart illustrating a changeable picture type determination process of the coding device 106 of FIG. 14. FIG.
FIG. 33 is a diagram showing an example in which the picture type is changed.
FIG. 34 is a diagram showing another example in which the picture type is changed.
FIG. 35 is a diagram illustrating a quantization control process of the coding device 106 of FIG.
FIG. 36 is a flowchart illustrating a quantization control process of the coding device 106 of FIG.
FIG. 37 is a block diagram showing a configuration of a tightly coupled transcoder 101.
FIG. 38 is a diagram illustrating a stream syntax of a video sequence.
FIG. 39 is a diagram illustrating a configuration of the syntax of FIG. 38.
FIG. 40 is a diagram illustrating the syntax of history_stream () that records fixed-length history information.
FIG. 41 is a diagram illustrating the syntax of history_stream () that records fixed-length history information.
FIG. 42 is a diagram illustrating the syntax of history_stream () that records fixed-length history information.
FIG. 43 is a diagram illustrating the syntax of history_stream () that records fixed-length history information.
FIG. 44 is a diagram illustrating the syntax of history_stream () that records fixed-length history information.
FIG. 45 is a diagram illustrating the syntax of history_stream () that records fixed-length history information.
FIG. 46 is a diagram illustrating the syntax of history_stream () that records fixed-length history information.
FIG. 47 is a diagram illustrating the syntax of history_stream () that records variable-length history information.
FIG. 48 is a diagram illustrating the syntax of sequence_header ().
FIG. 49 is a diagram illustrating the syntax of sequence_extension ().
FIG. 50 is a diagram illustrating the syntax of extension_and_user_data ().
FIG. 51 is a diagram illustrating the syntax of user_data ().
FIG. 52 is a diagram illustrating the syntax of group_of_pictures_header ().
FIG. 53 is a diagram illustrating the syntax of picture_header ().
FIG. 54 is a diagram illustrating the syntax of picture_coding_extension ().
FIG. 55 is a diagram illustrating the syntax of extension_data ().
FIG. 56 is a diagram illustrating the syntax of quant_matrix_extension ().
FIG. 57 is a diagram illustrating the syntax of copyright_extension ().
FIG. 58 is a diagram illustrating the syntax of picture_display_extension ().
FIG. 59 is a diagram illustrating the syntax of picture_data ().
FIG. 60 is a diagram illustrating the syntax of slice ().
FIG. 61 is a diagram illustrating the syntax of macroblock ().
FIG. 62 is a diagram illustrating the syntax of macroblock_modes ().
FIG. 63 is a diagram illustrating the syntax of motion_vectors (s).
FIG. 64 is a diagram illustrating the syntax of motion_vector (r, s).
FIG. 65 is a diagram illustrating a variable length code of macroblock_type for an I picture.
FIG. 66 is a diagram illustrating a variable length code of macroblock_type for a P picture.
FIG. 67 is a diagram illustrating a variable length code of macroblock_type for a B picture.
FIG. 68 is a block diagram showing another configuration of the transcoder 101 to which the present invention is applied.
FIG. 69 is a diagram illustrating a combination of items of history information.
70 is a flowchart illustrating the operation of the transcoder 101 of FIG. 68. FIG.
FIG. 71 is a diagram illustrating the syntax of re_coding_stream_info ().
FIG. 72 is a diagram illustrating red_bw_flag and red_bw_indicator.
FIG. 73 is a diagram illustrating Re_Coding Information Bus macroblock formation.
FIG. 74 is a diagram illustrating Picture rate elements.
FIG. 75 is a diagram illustrating Picture rate elements.
FIG. 76 is a diagram illustrating Picture rate elements.
FIG. 77 is a diagram illustrating an area in which a Re_Coding Information Bus is recorded.
[Code description] 1 Encoding device, 2 Decoding device, 3 Recording medium, 12,13 A / D converter, 14 frame memory, 15 Bright signal frame memory, 16 Color difference signal frame memory, 17 Format conversion circuit, 18 Encoder , 31 Decoder, 32 Format conversion circuit, 33 Frame memory, 34 Brightness signal frame memory, 35 Color difference signal frame memory, 36, 37 D / A converter, 50 Motion vector detection circuit, 51 Frame memory, 52 Prediction mode switching circuit, 53 Calculator, 54 Prediction Judgment Circuit, 55 DCT Mode Switching Circuit, 56 DCT Circuit, 57 Quantization Circuit, 58 Variable Length Coding Circuit, 59 Transmission Buffer, 60 Inverse Quantization Circuit, 61 IDCT Circuit, 62 Calculator, 63 Frame memory, 64 motion compensation circuit, 81 receive buffer, 82 variable length decoding circuit, 83 inverse quantization circuit, 84 IDCT circuit, 85 arithmetic unit, 86 Frame memory, 87 Motion compensation circuit, 101 Transcoder, 102 Decoder, 103 History information multiplexing device, 105 History information separator, 106 Encoding device, 111 Decoder, 112 Variable length decoding circuit, 121 Encoder,
77 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP08331555A | Cites | Japan |
| JP08111870A | Cites | Japan |
| JP10032830A | Cites | Japan |
| JP10271496A | Cites | Japan |
| JP07107461A | Cites | Japan |
| JP10070729A | Cites | Japan |
| JP08130712A | Cites | Japan |
| JP07154602A | Cites | Japan |
| JP11262005A | Cites | Japan |
| JP943338A | Cites | Japan |
| JP11164330A | Cites | Japan |
26 members in 6 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 1999031944 | Japan | – | |
| 3194499 | Japan | A | |
| 34315699 | Japan | A | |
| 199931944 | – | – | – |
| JP19990031944 | – | – | – |
| JP19990343156 | – | – | – |
Members26
| Document | Office | Kind | |
|---|---|---|---|
| WO0048402A1 | World Intellectual Property Organization (WIPO) | A1 | |
| JP2000299856A | Japan | A | |
| JP2000299857A | Japan | A | |
| EP1069779A1 | European Patent Office (EPO) | A1 | |
| CN1294820A | China | A | |
| KR20010042575A | Republic of Korea | A | |
| JP3672185B2 | Japan | B2 | |
| JP2005245002A | Japan | A | |
| JP2005253092A | Japan | A | |
| KR20050109629A | Republic of Korea | A | |
| CN1241416C | China | C | |
| KR100571307B1 | Republic of Korea | B1 | |
| KR100571687B1 | Republic of Korea | B1 | |
| JP3890838B2 | Japan | B2 | |
| JP2007060708A | Japan | A | |
| US7236526B1 | United States of America | B1 | |
| EP1069779A4 | European Patent Office (EPO) | A4 | |
| US2007253488A1 | United States of America | A1 | |
| US2008043839A1 | United States of America | A1 | |
| JP4139983B2This record | Japan | B2 | |
| US7680187B2 | United States of America | B2 | |
| JP4482811B2 | Japan | B2 | |
| JP4539637B2 | Japan | B2 | |
| JP4543321B2 | Japan | B2 | |
| US8681868B2 | United States of America | B2 | |
| EP1069779B1 | European Patent Office (EPO) | B1 |
18 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4139983
- Publication, DOCDB
- 4139983
- Publication, EPODOC
- JP4139983B
- Application
- 34315699
- Application, DOCDB
- 34315699
- Application, EPODOC
- JP19990343156
Titles2
- English
- A coded stream conversion device, a coded stream conversion method, a stream output device, and a stream output method.
- Japanese
- 符号化ストリーム変換装置、および、符号化ストリーム変換方法、並びに、ストリーム出力装置、および、ストリーム出力方法
Classification
- IPC, 18
- H04N19 114
- H04N7 08
- H04N7 081
- H04N7 24
- H04N19 00
- H04N19 126
- H04N19 134
- H04N19 172
- H04N19 186
- H04N19 189
- H04N19 423
- H04N19 46
- H04N19 51
- H04N19 625
- H04N19 70
- H04N19 85
- H04N19 91
- H04N7 26
