Editing device and method, re-coding device and method, and splicing device and method
Abstract
[Subject] The splicing equipment and stream edit equipment which realize carrying out splicing of two or more coding streams seamlessly are offered. [Solution means] The coding streams STA and STB in the re-encoding section are decoded by MPEG decoders 14A and 14B. The splice controller 13 calculates target bit quantity, when performing re-encoding processing in MPEG encoder 15. And the new quantization characteristic is set up using this calculated target code amount and the quantization characteristic included in the coding streams STA and STBS. MPEG encoder 15 re-encodes a video data based on this new quantization characteristic, and outputs it as the re-encoding stream STRE. The splice controller 13 controls the switch 17, and switches and outputs coding stream STA, the re-encoding stream STRE, and the coding stream STB. [Selection figure] Fig. 19

Term
Term ended
Projected expiry passed 16 May 2025, 1.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
69 claims: 16 independent, 53 dependent
- 1It is an editing device for inputting a plurality of encoded bitstreams obtained by encoding a plurality of video materials and connecting the plurality of encoded bitstreams, and inputting a plurality of encoded bitstreams. A decoding means for decoding each coded bitstream before and after the connection point including the connection point of the plurality of coded bitstreams and outputting the coded image data, and a new target code amount for the coded image data are set. The target code amount setting means to be used and the coded image data output from the decoding means are encoded according to a new target code amount set by the target code amount setting means to generate a new coded bit stream. The bitstream edited by switching between the original coded bitstream supplied and the new coded bitstream output by the coding means at the connection point. An editing device characterized in that it is equipped with an output means for outputting. 複数の映像素材を符号化して得られる複数の符号化ビットストリームを入力して、この複数の符号化ビットストリームを接続するための編集装置であって、複数の符号化ビットストリームを入力して、この複数の符号化ビットストリームの接続点を含む接続点前後の各符号化ビットストリームを復号化して、符号化画像データを出力する復号化手段と、符号化画像データに対する新たな目標符号量を設定する目標符号量設定手段と、上記復号化手段から出力された符号化画像データを、上記目標符号量設定手段によって設定される新たな目標符号量に従って符号化して、新たな符号化ビットストリームを生成する符号化手段と、上記供給された元の符号化ビットストリームと上記符号化手段によって出力される新たな符号化ビットストリームとを、上記接続点において切り換えて出力することによって、編集されたビットストリームを出力する出力手段とを備えたことを特徴とする編集装置。
- 7It is an editing method for inputting a plurality of encoded bitstreams obtained by encoding a plurality of video materials and connecting the plurality of encoded bitstreams, and inputting a plurality of encoded bitstreams. A decoding step of decoding each coded bitstream before and after the connection point including the connection point of the plurality of coded bitstreams and outputting the coded image data, and setting a new target code amount for the coded image data. The target code amount setting step to be performed and the coded image data output from the decoding step are encoded according to the new target code amount set by the target code amount setting step to generate a new coded bit stream. The bitstream edited by switching between the original coded bitstream supplied and the new coded bitstream output by the coding step at the connection point. An editing method consisting of an output process that outputs. 複数の映像素材を符号化して得られる複数の符号化ビットストリームを入力して、この複数の符号化ビットストリームを接続するための編集方法であって、複数の符号化ビットストリームを入力して、この複数の符号化ビットストリームの接続点を含む接続点前後の各符号化ビットストリームを復号化して、符号化画像データを出力する復号化工程と、符号化画像データに対する新たな目標符号量を設定する目標符号量設定工程と、上記復号化工程から出力された符号化画像データを、上記目標符号量設定工程によって設定される新たな目標符号量に従って符号化して、新たな符号化ビットストリームを生成する符号化工程と、上記供給された元の符号化ビットストリームと上記符号化工程によって出力される新たな符号化ビットストリームとを、上記接続点において切り換えて出力することによって、編集されたビットストリームを出力する出力工程とから構成される編集方法。
- 13In an editing device that edits a plurality of source-encoded streams, an edit point setting means for setting an edit point for each of the plurality of source-encoded streams, and a picture of the plurality of source-encoded streams in the vicinity of the edit point. Decoding means that decodes and outputs the decoded video data, re-encoding means that re-encodes the decoded video data and outputs the re-encoded stream, and the source-encoded stream and the re-encoded stream. An edit stream generation means that generates an edited edit stream by switching and outputting, and an edit control that controls the re-encoding means and the edit stream generation means so that the edit stream does not become discontinuous at the time of decoding. An editing device characterized by having means. 複数のソース符号化ストリームを編集する編集装置において、上記複数のソース符号化ストリームに対して、それぞれ編集ポイントを設定する編集ポイント設定手段と、上記複数のソース符号化ストリームの上記編集ポイント付近のピクチャをそれぞれデコードし、デコードされたビデオデータを出力するデコード手段と、上記デコードされたビデオデータを再エンコードし、再エンコードストリームを出力する再エンコード手段と、上記ソース符号化ストリームと上記再エンコードストリームとを切り換えて出力することによって、編集された編集ストリームを生成する編集ストリーム生成手段と、上記上記編集ストリームがデコード時に不連続とならないように上記再エンコード手段及び上記編集ストリーム生成手段を制御する編集制御手段とを備えたことを特徴とする編集装置。
- 15The editing control means is the re-encoding means so that the trajectory of the data occupancy of the VBV buffer does not become discontinuous at the switching point between the source encoding stream and the re-encoding stream or the editing point of the re-encoding stream. The target bit amount for re-encoding in the above is calculated, and the re-encoding means encodes the source-encoded stream based on the target bit amount supplied from the editing control means. The editing device described in the section. 上記編集制御手段は、上記ソース符号化ストリームと上記再エンコードストリームの切換えポイント又は上記再エンコードストリームの上記編集ポイントにおいて、VBVバッファのデータ占有量の軌跡が不連続にならないように、上記再エンコード手段における再エンコードの目標ビット量を演算し、上記再エンコード手段は、上記編集制御手段から供給された上記目標ビット量に基いて、上記ソース符号化ストリームを符号化することを特徴とする請求項13項記載の編集装置。
- 19The editing control means extracts information on the quantization characteristics included in the source coding stream for each picture, and each in the re-encoding process in the re-encoding means so that overflow and underflow do not occur in the VBV buffer. The picture target bit amount is calculated, information about the quantization characteristic included in the source coding stream is extracted for each picture, and the quantization characteristic extracted from the source coding stream and the calculated target bit amount are extracted. 13. The re-encoding means is controlled so as to calculate a new quantization characteristic and perform the re-encoding process based on the calculated new quantization characteristic. The editing device described. 上記編集制御手段は、上記ソース符号化ストリームに含まれる量子化特性に関する情報を各ピクチャ毎に抽出し、VBVバッファにオーバーフロー及びアンダーフローが発生しないように、上記再エンコード手段における再エンコード処理における各ピクチャ目標ビット量を演算し、上記ソース符号化ストリームに含まれる量子化特性に関する情報を各ピクチャ毎に抽出し、上記ソース符号化ストリームから抽出された量子化特性と、上記演算された目標ビット量とに基いて、新たな量子化特性を演算し、上記演算された新たな量子化特性に基いて上記再エンコード処理を行なうように上記再エンコード手段を制御することを特徴とする請求項13項記載の編集装置。
- 23The re-encoding means includes a motion detecting means for detecting the motion of each picture and generating a motion vector, and a motion compensating means for performing motion compensation based on the motion vector detected by the motion detecting means. The control means determines whether or not to reuse the motion vector information extracted from the source coding stream during the re-encoding process in the re-encoding means, and if it is determined to reuse the motion vector information, the motion detecting means. 13. The re-encoding means is controlled so as to supply the motion vector extracted from the stream instead of the motion vector detected in the above stream to the motion compensation circuit of the re-encoding means. Editing device. 上記再エンコード手段は、各ピクチャの動きを検出して、動きベクトルを生成する動き検出手段と、上記動き検出手段によって検出された動きベクトルに基いて動き補償を行なう動き補償手段を備え、上記編集制御手段は、上記ソース符号化ストリームから抽出された動きベクトル情報を、再エンコード手段における再エンコード処理時に再利用するか否かを判断し、再利用すると判断された場合には、上記動き検出手段において検出された動きベクトルではなくて、上記ストリームから抽出された動きベクトルを、上記再エンコード手段の動き補償回路に供給するように上記再エンコード手段を制御することを特徴とする請求項13項記載の編集装置。
- 27The plurality of source coded streams include at least a first coded stream and a second coded stream, and the editing control means is the first of the plurality of pictures constituting the first coded stream. For the first edit point set in the coded stream 1, the output picture is selected so that the future picture is not output as the edit stream on the presentation time axis, and the second coded stream is set. A claim characterized in that the output picture is selected so that the past picture is not output on the presentation time axis for the second edit point set in the second coded stream among the plurality of constituent pictures. Item 13. The editing device according to item 13. 上記複数のソース符号化ストリームは、少なくとも第1の符号化ストリームと第2の符号化ストリームを含み、上記編集制御手段は、上記第1の符号化ストリームを構成する複数のピクチャのうち、上記第1の符号化ストリームに設定された第1の編集ポイントに対して、プレゼンテーション時間軸において未来のピクチャを上記編集ストリームとして出力しないように出力ピクチャを選択し、且つ、上記第2の符号化ストリームを構成する複数のピクチャのうち、第2の符号化ストリームに設定された第2の編集ポイントに対して、プレゼンテーション時間軸において過去のピクチャを出力しないように出力ピクチャを選択することを特徴とする請求項13項記載の編集装置。
- 29In the editing method of editing a plurality of source-encoded streams to generate an edited stream, an edit point setting step of setting edit points for each of the plurality of source-encoded streams and the above-mentioned plurality of source codes. A decoding process that decodes the pictures near the edit point of the conversion stream and outputs the decoded video data, a re-encoding process that re-encodes the decoded video data and outputs the re-encoded stream, and the above source. The edit stream generation step of generating the edited edit stream by switching between the encoded stream and the re-encoded stream, and the re-encoded step and the above-mentioned re-encoding step so that the above-mentioned edit stream is not discontinuous at the time of decoding. An editing method characterized by including an editing control process for controlling an editing stream generation process. 複数のソース符号化ストリームを編集して、編集されたストリームを生成する編集方法において、上記複数のソース符号化ストリームに対して、それぞれ編集ポイントを設定する編集ポイント設定工程と、上記複数のソース符号化ストリームの上記編集ポイント付近のピクチャをそれぞれデコードし、デコードされたビデオデータを出力するデコード工程と、上記デコードされたビデオデータを再エンコードし、再エンコードストリームを出力する再エンコード工程と、上記ソース符号化ストリームと上記再エンコードストリームとを切り換えて出力することによって、編集された編集ストリームを生成する編集ストリーム生成工程と、上記上記編集ストリームがデコード時に不連続とならないように上記再エンコード工程及び上記編集ストリーム生成工程を制御する編集制御工程とを備えたことを特徴とする編集方法。
- 30In a splicing device that splics a plurality of source-encoded streams, a splicing point setting means for setting a splicing point for each of the plurality of source-encoded streams, and a picture of the plurality of source-encoded streams in the vicinity of the splicing point. Decoding means that decodes and outputs the decoded video data, re-encoding means that re-encodes the decoded video data and outputs the re-encoded stream, and the source-encoded stream and the re-encoded stream. The splicing stream generating means for generating the spliced splicing stream by switching and outputting, and the splicing control for controlling the re-encoding means and the splicing stream generating means so that the splicing stream does not become discontinuous at the time of decoding. A splicing device characterized by being equipped with means. 複数のソース符号化ストリームをスプライシングするスプライシング装置において、上記複数のソース符号化ストリームに対して、それぞれスプライシングポイントを設定するスプライシングポイント設定手段と、上記複数のソース符号化ストリームの上記スプライシングポイント付近のピクチャをそれぞれデコードし、デコードされたビデオデータを出力するデコード手段と、上記デコードされたビデオデータを再エンコードし、再エンコードストリームを出力する再エンコード手段と、上記ソース符号化ストリームと上記再エンコードストリームとを切り換えて出力することによって、スプライシングされたスプライシングストリームを生成するスプライシングストリーム生成手段と、上記上記スプライシングストリームがデコード時に不連続とならないように上記再エンコード手段及び上記スプライシングストリーム生成手段を制御するスプライス制御手段とを備えたことを特徴とするスプライシング装置。
- 32The splice control means is a re-encoding means so that the trajectory of the data occupancy of the VBV buffer is not discontinuous at the switching point between the source encoding stream and the re-encoding stream or the splice point of the re-encoding stream. 30. The re-encoding means calculates the target bit amount of the re-encoding in the above, and the re-encoding means encodes the source coded stream based on the target bit amount supplied from the splice control means. The splicing device according to the section. 上記スプライス制御手段は、上記ソース符号化ストリームと上記再エンコードストリームの切換えポイント又は上記再エンコードストリームの上記スプライスポイントにおいて、VBVバッファのデータ占有量の軌跡が不連続にならないように、上記再エンコード手段における再エンコードの目標ビット量を演算し、上記再エンコード手段は、上記スプライス制御手段から供給された上記目標ビット量に基いて、上記ソース符号化ストリームを符号化することを特徴とする請求項30項記載のスプライシング装置。
- 36The splice control means extracts information on the quantization characteristics included in the source coded stream for each picture, and each in the re-encoding process in the re-encoding means so that overflow and underflow do not occur in the VBV buffer. The picture target bit amount is calculated, information about the quantization characteristic included in the source coding stream is extracted for each picture, and the quantization characteristic extracted from the source coding stream and the calculated target bit amount are extracted. 30. The present invention is characterized in that a new quantization characteristic is calculated based on the above, and the re-encoding means is controlled so as to perform the re-encoding process based on the calculated new quantization characteristic. The splicing device described. 上記スプライス制御手段は、上記ソース符号化ストリームに含まれる量子化特性に関する情報を各ピクチャ毎に抽出し、VBVバッファにオーバーフロー及びアンダーフローが発生しないように、上記再エンコード手段における再エンコード処理における各ピクチャ目標ビット量を演算し、上記ソース符号化ストリームに含まれる量子化特性に関する情報を各ピクチャ毎に抽出し、上記ソース符号化ストリームから抽出された量子化特性と、上記演算された目標ビット量とに基いて、新たな量子化特性を演算し、上記演算された新たな量子化特性に基いて上記再エンコード処理を行なうように上記再エンコード手段を制御することを特徴とする請求項30項記載のスプライシング装置。
- 44The plurality of source coded streams include at least a first coded stream and a second coded stream, and the splice control means is the first of the plurality of pictures constituting the first coded stream. For the first splice point set in the coded stream 1, the output picture is selected so that the future picture is not output as the splicing stream on the presentation time axis, and the second coded stream is set. A claim characterized in that an output picture is selected so as not to output a past picture on the presentation time axis with respect to a second splicing point set in the second coded stream among a plurality of constituent pictures. Item 30. The splicing device according to item 30. 上記複数のソース符号化ストリームは、少なくとも第1の符号化ストリームと第2の符号化ストリームを含み、上記スプライス制御手段は、上記第1の符号化ストリームを構成する複数のピクチャのうち、上記第1の符号化ストリームに設定された第1のスプライスポイントに対して、プレゼンテーション時間軸において未来のピクチャを上記スプライシングストリームとして出力しないように出力ピクチャを選択し、且つ、上記第2の符号化ストリームを構成する複数のピクチャのうち、第2の符号化ストリームに設定された第2のスプライシングポイントに対して、プレゼンテーション時間軸において過去のピクチャを出力しないように出力ピクチャを選択することを特徴とする請求項30項記載のスプライシング装置。
- 46In a splicing method for splicing a plurality of source-encoded streams, a splicing point setting step for setting a splicing point for each of the plurality of source-encoded streams, and a picture of the plurality of source-encoded streams near the splicing point. The decoding process that decodes and outputs the decoded video data, the re-encoding process that re-encodes the decoded video data and outputs the re-encoded stream, and the source-encoded stream and the re-encoded stream. Splice control that controls the splicing stream generation step of generating the spliced splicing stream and the re-encoding step and the splicing stream generation step so that the splicing stream does not become discontinuous at the time of decoding by switching and outputting. A splicing device characterized by having a process. 複数のソース符号化ストリームをスプライシングするスプライシング方法において、上記複数のソース符号化ストリームに対して、それぞれスプライシングポイントを設定するスプライシングポイント設定工程と、上記複数のソース符号化ストリームの上記スプライシングポイント付近のピクチャをそれぞれデコードし、デコードされたビデオデータを出力するデコード工程と、上記デコードされたビデオデータを再エンコードし、再エンコードストリームを出力する再エンコード工程と、上記ソース符号化ストリームと上記再エンコードストリームとを切り換えて出力することによって、スプライシングされたスプライシングストリームを生成するスプライシングストリーム生成工程と、上記上記スプライシングストリームがデコード時に不連続とならないように上記再エンコード工程及び上記スプライシングストリーム生成工程を制御するスプライス制御工程とを備えたことを特徴とするスプライシング装置。
- 48The splice control step is a re-encoding step so that the trajectory of the data occupancy of the VBV buffer does not become discontinuous at the switching point between the source encoding stream and the re-encoding stream or the splice point of the re-encoding stream. 46. The target bit amount for re-encoding in the above is calculated, and the re-encoding step encodes the source coded stream based on the target bit amount supplied from the splice control step. The splicing method described in the section. 上記スプライス制御工程は、上記ソース符号化ストリームと上記再エンコードストリームの切換えポイント又は上記再エンコードストリームの上記スプライスポイントにおいて、VBVバッファのデータ占有量の軌跡が不連続にならないように、上記再エンコード工程における再エンコードの目標ビット量を演算し、上記再エンコード工程は、上記スプライス制御工程から供給された上記目標ビット量に基いて、上記ソース符号化ストリームを符号化することを特徴とする請求項46項記載のスプライシング方法。
- 52In an encoding device that encodes the supplied source video data, the source video data is encoded and encoded by the first encoding means that outputs the first encoded stream and the first encoding means. By selectively using the decoding means for decoding the encoded stream and outputting the decoded video data and the encoding parameters included in the first encoded stream and used at the time of the first encoding, the encoding parameters are selectively used. The second encoding means that re-encodes the video data decoded by the decoding means and outputs it as a re-encoded stream, and the second encoding means so that overflow and underflow do not occur in the VBV buffer corresponding to the re-encoded stream. An encoding device including an encoding control means for controlling the encoding means of the data. 供給されたソースビデオデータを符号化する符号化装置において、上記ソースビデオデータを符号化し、第1の符号化ストリームを出力する第1のエンコード手段と、上記第1のエンコード手段によって符号化された符号化ストリームをデコードし、デコードされたビデオデータを出力するデコード手段と、上記第1の符号化ストリームに含まれる上記第1の符号化時において使用した符号化パラメータを選択的に使用して、上記デコード手段によってデコードされたビデオデータを再エンコードして再エンコードストリームとして出力する第2のエンコード手段と、上記再エンコードストリームに対応するVBVバッファにオーバーフロー及びアンダーフローが発生しないように、上記第2のエンコード手段を制御するエンコード制御手段とを備えたことを特徴とする符号化装置。
- 61In the encoding method for encoding the supplied source video data, the source video data is encoded and encoded by the first encoding step of outputting the first encoded stream and the first encoding step. By selectively using the decoding process of decoding the encoded stream and outputting the decoded video data, and the encoding parameters included in the first encoded stream and used at the time of the first encoding, the encoding parameters are selectively used. The second encoding step of re-encoding the video data decoded by the decoding step and outputting it as a re-encoded stream, and the second encoding step so as not to cause overflow and underflow in the VBV buffer corresponding to the re-encoded stream. An encoding method including an encoding control process for controlling the encoding process of the above. 供給されたソースビデオデータを符号化する符号化方法において、上記ソースビデオデータを符号化し、第1の符号化ストリームを出力する第1のエンコード工程と、上記第1のエンコード工程によって符号化された符号化ストリームをデコードし、デコードされたビデオデータを出力するデコード工程と、上記第1の符号化ストリームに含まれる上記第1の符号化時において使用した符号化パラメータを選択的に使用して、上記デコード工程によってデコードされたビデオデータを再エンコードして再エンコードストリームとして出力する第2のエンコード工程と、上記再エンコードストリームに対応するVBVバッファにオーバーフロー及びアンダーフローが発生しないように、上記第2のエンコード工程を制御するエンコード制御工程とを備えたことを特徴とする符号化方法。
Independent claims16
172 paragraphs, as filed
The present invention comprises an editing device and method for generating an edited video material by editing a plurality of video materials, a bitstream splicing device and method for generating a seamless splicing stream by splicing a plurality of bitstreams, and a bitstream splicing device and method. , A coding device and a method for encoding video data.
In recent years, using a storage medium such as a DVD (digital versatile disk or digital video disk), which is an optical disk capable of recording a large amount of digital data, compression-encoded image data is recorded or reproduced. Reproduction systems and multiplex transmission systems that multiplex and transmit a plurality of compressed and encoded broadcast materials (programs) have been proposed. These systems utilize compression coding technology for image data based on the MPEG (Moving Picture Experts Group) standard.
In this MPEG standard, a bidirectional predictive coding method is adopted as a coding method. In this bidirectional predictive coding method, three types of coding are performed: intraframe coding, interframe forward predictive coding, and bidirectional predictive coding. They are called intra coded picture), P picture (predictive coded picture) and B picture (bidirectionally predictive coded picture). In addition, GOP (Group of Picture), which is a unit of random access, is constructed by appropriately combining each picture of I, P, and B. In general, the amount of code generated in each picture is the largest in the I picture, the second largest in the P picture, and the smallest in the B picture.
In a coding method such as the MPEG standard in which the amount of bits generated differs for each picture, in order to appropriately transmit and decode the obtained coded bit stream (hereinafter, also simply referred to as a stream) to obtain an image, image decoding is performed. The image encoding device must know the amount of data occupied in the input buffer of the computer. Therefore, in the MPEG standard, a VBV (Video Buffering Verifier) buffer, which is a virtual buffer corresponding to the input buffer in the image decoding device, is assumed, and on the image coding device side, the VBV buffer is broken, that is, underflow or overflow. You have to generate a stream so that it doesn't let you.
Here, with reference to FIG. 1, an outline of a transmission system and a recording / playback system according to the MPEG standard will be described. Note that FIG. 1 shows a transmission system configured to realize ISO / IEC13818-1 (MPEG1) and ISO / IEC13818-2 (MPEG2).
This transmission system, as a configuration on the encoding device 110 side, encodes (encodes) the input video data DV and outputs a video elementary stream (video ES) which is an encoded bit stream and a video encoder 111. , The packetizer 112 that outputs the video packaged elementary stream (video PES) by adding a header etc. to the video elementary stream output from this video encoder 111 and packetizing it, and the input audio data DA The audio encoder 113 that encodes and outputs the audio elementary stream (audio ES) that is the encoded bit stream, and the audio elementary stream that is output from this audio encoder 113 are packetized by adding a header or the like to the audio. The packetizer 114 that outputs the packetized elemental stream (audio PES), the video packetized elemental stream output from the packetizer 112, and the audio packaged elemental stream output from the packetizer 114 are multiplexed and 188. It is equipped with a transport stream multiplexer (referred to as TSMUX in the figure) 115 that creates a byte transport stream packet and outputs it as a transport stream (TS).
Further, in the transmission system shown in FIG. 1, as a configuration on the decoding device 120 side, a transport stream output from the transport stream multiplexer 115 and transmitted via the transmission medium 116 is input, and the video packaged elementary is input. Output from the transport stream demultiplexer (referred to as TSDEMUX in the figure) 121 that outputs separately to the stream (video PES) and the audio packaged elementary stream (audio PES), and from this transport stream demultiplexer 121. The depacketizer 122 that depackets the video packaged elemental stream to be output and outputs the video elemental stream (video ES) and the video elemental stream output from this depacketizer 122 are decoded and the video data DV The video decoder 124 that outputs the data, the depacketizer 124 that depackets the audio packaged elemental stream output from the transport stream demultiplexer 121, and outputs the audio elementary stream (audio ES), and the depacketizer 124 that outputs the data. It is equipped with an audio decoder 125 that decodes the audio elementary stream and outputs the audio data DA.
The decoding device 120 in FIG. 1 is commonly referred to as an intelligent receiver decoder (IRD).
In the case of a recording / playback system that records encoded data on a storage medium, instead of the transport stream multiplexer 115 in FIG. 1, a video packaged elementary stream and an audio packaged elementary stream are used. A program stream multiplexer is provided to multiplex and output the program stream (PS), a storage medium for recording the program stream is used instead of the transmission media 116, and the program stream is videotaped instead of the transport stream demultiplexer 121. A program stream demultiplexer that separates the packaged elemental stream and the audio packaged elemental stream is provided.
Next, the operation of the system shown in FIG. 1 will be described by taking as an example the case where the video data is coded, transmitted, and decoded.
First, on the encoder 110 side, the input video data DV in which each picture has the same bit amount is encoded by the video encoder 111, and each picture is converted and compressed into a different bit amount according to its redundancy. Output as a video elementary stream. The packetizer 112 inputs a video elementary stream, packets it to absorb (average) the fluctuation of the bit amount on the bit stream time axis, and outputs it as a video packaged elementary stream (PES). The PES packet packetized by the packetizer 112 is composed of one access unit or a plurality of access units. Generally, this one access unit is composed of one frame. The transport stream multiplexer 115 multiplexes the video packaged elemental stream and the audio packaged elemental stream output from the packetizer 114 to create a transport stream packet, and transmits it as a transport stream (TS). It is sent to the decoding device 120 via the media 116.
On the decoding device 120 side, the transport stream demultiplexer 121 separates the transport stream into a video packaged elementary stream and an audio packaged elementary stream. The depacketizer 122 depackets the video packaged elementary stream and outputs the video elementary stream, and the video decoder 123 decodes the video elementary stream and outputs the video data DV.
The decoding device 120 buffers the stream transmitted at a constant transmission rate in the VBV buffer, and extracts data for each picture from the VBV buffer based on a predefined time stamp (DTS) set for each picture in advance. .. The capacity of this VBV buffer is determined according to the standard of the transmitted signal, and in the case of the standard video signal of the main profile main level (MP @ ML), it has a capacity of 1.75 Mbits. There is. On the encoder 110 side, it is necessary to control the bit generation amount of each picture so as not to overflow or underflow this VBV buffer.
Next, the VBV buffer will be described with reference to FIG. In FIG. 2, the fold line represents the change in the amount of data occupied in the VBV buffer, the slope 131 represents the transmission bit rate, and the vertically falling portion 132 represents the VBV of the video decoder 123 for reproduction of each picture. Represents the amount of bits to pull out of the buffer. The timing extracted by the video decoder 123 is specified by information called a presentation time stamp (PTS) or a coded time stamp (DTS). The interval between PTS and DTS is generally one frame period. In FIG. 2, I, P, and B represent I picture, P picture, and B picture, respectively. This also applies to other figures. Also, in Fig. 2, vbv_delay is the time from when the occupancy of the VBV buffer is zero to when it is fully filled, and is TP. Represents the presentation time cycle. On the decoder side, as shown in FIG. 2, the transmitted stream fills the VBV buffer at a constant bit rate 131, and data is extracted from the VBV buffer for each picture at a timing according to the presentation time.
Next, with reference to FIG. 3, the rearrangement of pictures in the bidirectional predictive coding method of the MPEG standard will be described. In FIG. 3, (a) represents the picture order of the input video data supplied to the encoder, (b) shows the order of the pictures after the sorting is performed in the encoder, and (c) shows the picture order from the encoder. Represents the picture order of the output video stream. As shown in FIG. 3, in the encoder, the input video frames are sorted according to the picture type (I, P, B) at the time of encoding, and are encoded in the sorted order. Specifically, the B picture is predictively coded from the I picture or P picture, so that it is located after the I picture or P picture used for the predictive coding, as shown in FIG. 3 (b). Sorted to. The encoder performs coding processing in the order of the rearranged pictures, and outputs each encoded picture as a video stream in the order shown in FIGS. 3 (b) and 3 (c). The output coded stream is supplied to the decoder or the storage medium via the transmission line. In FIG. 3, it is assumed that the GOP is composed of 15 pictures.
FIG. 4 is a diagram for showing which picture is used for predictive coding processing in each picture in the predictive coding method described with reference to FIG. 3, and is a diagram for rearranging the pictures and predicting code in the encoder. It shows the relationship with the conversion process. In FIG. 4, (a) represents the order of the pictures in the input video data to the encoder, (b) represents the order of the pictures after the sorting is performed in the encoder, and (c) and (d) are the encoders, respectively. It represents the picture held in the two frame memories FM1 and FM2, and (e) represents the elementary stream (ES) output from the encoder. In the figure, the numbers attached to I, P, and B indicate the order of the pictures. In the encoder, the input video data as shown in FIG. 4 (a) is sorted in the order of the pictures as shown in FIG. 4 (b). Further, the two frame memories FM1 and FM2 hold pictures as shown in FIGS. 4 (c) and 4 (d), respectively. Then, when the input video data is an I picture, the encoder performs encoding processing based only on the input video data (I picture), and when the input video data is a P picture, it is held in the input video data and the frame memory FM1. Prediction coding processing is performed based on the I picture or P picture, and if the input video data is B picture, the prediction code is based on the input video data and the two pictures held in the frame memories FM1 and FM2. Perform the conversion process. The reference numerals in FIG. 4 (e) represent the pictures used in the coding process.
Next, with reference to FIG. 5, the rearrangement of pictures in the bidirectional predictive decoding method of the MPEG standard will be described. In FIG. 5, (a) represents the order of the pictures of the coded video stream supplied from the encoder to the decoder via the transmission line, and (b) is the coded video supplied from the encoder to the decoder via the transmission line. The stream is represented, and (c) represents the picture order of the video data output from the decoder. As shown in FIG. 5, in this decoder, since the B picture is predictively coded using the I picture or the P picture at the time of encoding, the same picture as the picture used for this predictive coding process is used. Decryption processing is performed using this. As a result, as shown in FIG. 5 (c), the B picture is output earlier than the I picture and the P picture in the order of the pictures of the video data output from this decoder.
FIG. 6 is a diagram for explaining the decoding process described in FIG. 5 in more detail. In FIG. 6, (a) represents the order of the encoded video elementary stream (ES) pictures supplied to the decoder, and (b) and (c) are the two frame memories FM1 in the decoder, respectively. , FM2 represents the picture, and (d) represents the output video data output from the decoder. Similar to FIG. 4, the numbers attached to I, P, and B in the figure indicate the order of the pictures. When the input elementary stream (ES) as shown in Fig. 6 (a) is input, the decoder must store the two pictures used for the predictive coding process, so the two frame memories FM1, The FM2 holds the pictures shown in FIGS. 6 (b) and 6 (c), respectively. Then, when the input elementary stream is an I picture, the decoder performs decoding processing based only on the input elementary stream (I picture), and when the input elementary stream is a P picture, the input elementary stream and the frame memory. Decoding processing is performed based on the I picture or P picture held in FM1, and if the input elementary stream is B picture, the input elementary stream and the two pictures held in the frame memories FM1 and FM2 Decoding process is performed based on the above to generate output video data as shown in FIG. 6 (d).
Next, motion detection and motion compensation in the bidirectional predictive coding method of the MPEG standard will be described with reference to FIG. 7. FIG. 7 shows the prediction directions (directions in which differences are taken) between pictures arranged in the order of input video frames by arrows. The MPEG standard employs motion compensation that enables higher compression. In order to perform motion compensation, the encoder performs motion detection in accordance with the prediction direction shown in FIG. 7 at the time of encoding to obtain a motion vector. The P picture and the B picture are composed of a difference value between this motion vector and a search image or a predicted image obtained according to the motion vector. At the time of decoding, the P picture and the B picture are reconstructed based on the motion vector and the difference value.
In general, as shown in FIG. 7, an I-picture is a picture encoded from the information of the I-picture, and is a picture generated without using interframe prediction. A P-picture is a picture generated by making a prediction from a past I-picture or P-picture. The B picture is a picture predicted from both directions of the past I or P picture and the future I or P picture, a picture predicted from the forward direction of the past I or P picture, or a picture predicted from the opposite direction of the I or P picture. It is one of the pictures made.
By the way, in the editing system of a broadcasting station, the video material recorded by using a camcorder or a digital VTR at the interview site, the video material transmitted from a local station or a program supply company, etc. are edited and used for one on-air. I am trying to generate a video program of. In recent years, video materials recorded using camcorders and digital VTRs and video materials supplied by local stations, program suppliers, etc. are supplied as compressed and coded encoded streams using the above-mentioned MPEG technology. Mostly done. The reason is that the recording area of the recording medium can be effectively utilized by compressing and encoding using MPEG technology rather than recording the uncompressed baseband video data as it is on the recording medium. This is because the transmission line can be effectively utilized by compressing and encoding using MPEG technology rather than transmitting the unrecorded baseband video data as it is.
In this conventional editing system, for example, in order to edit two encoded video streams to generate one on-air video program, all the data of the encoded stream is once decoded before the editing process. , I have to go back to the baseband video data. This is because the prediction direction of each picture included in the coded stream according to the MPEG standard is interrelated with the prediction direction of the preceding and following pictures, so that the coded stream can be connected at an arbitrary position on the stream. Because it cannot be done. If you forcibly connect two coded streams, the connection will not be discontinuous and you will not be able to decode accurately.
Therefore, when trying to edit two coded streams supplied in a conventional editing system, a decoding process that decodes the two source coded streams once to generate two baseband video data and two basebands. The editing process of editing the video data of the above to generate the edited video data for on-air and the encoding process of re-encoding the edited video data to generate an encoded video stream must be performed. should not.
However, since MPEG coding is not 100% reversible coding, there is a problem that the image quality deteriorates when the decoding process and the coding process are performed only for the editing process in this way. That is, in order to perform editing processing in the conventional editing system, decoding processing and coding processing are always required, so that there is a problem that the image quality is deteriorated by that amount.
As a result, in recent years, there has been a demand for a technique that enables editing in the state of the coded stream without decoding all of the supplied coded stream. At the bitstream level encoded in this way, concatenating two different encoded bitstreams to generate a concatenated bitstream is called "splicing". In other words, splicing means editing a plurality of streams in the state of the encoded stream.
However, there are three major problems in realizing this splicing.
The first problem is that it comes from the viewpoint of the order of picture presentation. The picture presentation order is the display order of each video frame. This first problem will be described with reference to FIGS. 8 and 9.
Both FIGS. 8 and 9 show the relationship between the order of pictures in the stream before and after the splice and the order of the picture presentation after the splice when the stream is simply spliced. FIG. 8 shows the case where the problem does not occur in the order of the picture presentations, and FIG. 9 shows the case where the problem occurs in the order of the picture presentations. Also, in FIGS. 8 and 9, (a) shows one video stream A to be spliced, (b) shows the other video stream B to be spliced, and (c) shows the stream after splicing. (d) shows the order of presentations. Also, SPA indicates the splice point in stream A, SPB indicates the splice point in stream B, STA indicates stream A, STB indicates stream B, and STSP indicates the spliced stream.
Since the picture is rearranged in the encoder, the order of the pictures in the encoded stream after encoding and the order of the presentation after decoding are different. However, the splice processing shown in FIG. 8 is a diagram showing an example in which the splice point SPA is set after the B picture of the stream A and the splice point SPB is set before the P picture of the stream B. As shown in FIG. 8 (d), the picture of stream A is displayed after the picture of stream B, or the picture of stream B is the picture of stream A at the boundary SP between stream A and stream B. There is no phenomenon that it is displayed before. That is, in the case of splicing at the splice points SPA and SPB shown in FIG. 8, there is no problem in the order of presentation.
On the other hand, the splice processing shown in FIG. 9 is a diagram showing an example in which the splice point SPA is set after the P picture of the stream A and the splice point SPB is set before the B picture of the stream B. As shown in FIG. 8 (d), the last picture of stream A is displayed after the picture of stream B, or two pictures of stream B are streamed at the boundary SP between stream A and stream B. There is a phenomenon that it is displayed before the last picture of A. In other words, if the pictures are displayed in this order, the video image of stream A is switched to the video image of stream B near the splice point, and two frames later, the video image of stream A is only one frame again. It becomes a strange image that it is displayed. That is, when splicing at the splice points SPA and SPB shown in FIG. 9, problems arise in the order of presentation.
Therefore, in order to realize the splice at an arbitrary picture position, such a problem regarding the presentation order may occur.
The second problem is that it comes from the viewpoint of motion compensation. This problem will be described with reference to FIGS. 10 to 12. Each of these figures shows the relationship between the order of pictures in the stream after splicing and the order of picture presentations when a simple splice of the stream is performed. FIG. 10 shows an example when there is no problem with motion compensation, and FIGS. 11 and 12 show an example when there is a problem with motion compensation. Note that FIGS. 11 and 12 show the case where the splices shown in FIGS. 8 and 9 are performed, respectively. In each figure, (a) shows the stream after splicing, and (b) shows the order of presentation.
In the example shown in FIG. 10, since stream B is a closed GOP (a closed GOP whose prediction does not depend on the previous GOP) and the splice is performed at the break of the GOP, the motion compensation is performed without excess or deficiency. The picture is decoded without any problem.
On the other hand, in FIGS. 11 and 12, motion compensation for referencing pictures in different streams is performed at the time of decoding, which causes a problem in motion compensation. Specifically, the motion compensation by the motion vector in the prediction direction shown by the broken line in the figure is invalid because the B picture and P picture of the stream B cannot be created by referring to the P picture of the stream A (in the figure). It is written as NG.). Therefore, in the examples shown in FIGS. 11 and 12, the picture remains broken until the next I picture.
When splicing in arbitrary picture units, splicing on a simple stream cannot solve this problem.
The third problem comes from the perspective of the VBV buffer. This problem will be described with reference to FIGS. 13 to 18. FIG. 13 shows an example of ideal stream splicing that satisfies the conditions of picture presentation order, motion compensation, and VBV buffer. In the figure, STA, STB, and STC indicate stream A, stream B, and stream C, respectively. In FIG. 13, (a) shows the state of the VBV buffer on the decoder side, (b) shows the stream after splicing, (c) shows the generation timing of each picture after sorting in the encoder, and (d). Indicates the order of the pictures after decoding. In FIG. 13, SPV indicates the splice point in the VBV buffer, VOC indicates the occupancy of the VBV buffer in the splice point SPV, and SPS indicates the splice point in the stream. In the example shown in FIG. 13, from the viewpoint of the VBV buffer, the VBV buffer does not collapse due to the overflow or underflow of the VBV buffer due to the splice.
However, in general, when splicing in arbitrary picture units, the condition that the VBV buffer does not collapse is not always satisfied. This will be described with reference to FIGS. 14 to 18. 14 and 15 show normal streams A and B that satisfy the constraints of the VBV buffer, respectively, (a) shows the state of the VBV buffer on the decoder side, and (b) shows streams A and B, respectively. There is. Figures 16 to 18 show three examples when such streams A and B are simply spliced at arbitrary positions. In these figures, (a) shows the state of the VBV buffer on the decoder side, and (b) shows the stream after splicing. As shown in these figures, the state of the VBV buffer differs depending on where the streams A and B are spliced. In the example shown in FIG. 16, the stream after splicing also satisfies the constraint of the VBV buffer, but in the example shown in FIG. 17, the overflow indicated by reference numeral 141 occurs, and the constraint of the VBV buffer is satisfied. Absent. Further, in the example shown in FIG. 18, the underflow indicated by reference numeral 142 occurs, and the restriction of the VBV buffer is not satisfied. On the decoder (IRD) side, if underflow or overflow of the VBV buffer occurs, picture decoding fails due to the failure of the VBV buffer, and seamless image reproduction cannot be realized. The state of image reproduction at that time varies depending on the function of the encoder (IRD), such as image skip, freeze, and processing stop. When splicing in arbitrary picture units, splicing on a simple stream cannot solve this problem.
As described above, conventionally, there has been a problem that there is no effective means for realizing seamless splice on the stream.
The following prior art documents can be mentioned as related to the above description (Patent Documents 1 to 3).
<patcit num="1"><text>Japanese Unexamined Patent Publication No. 08-205079</text></patcit><patcit num="2"><text>Japanese Unexamined Patent Publication No. 08-37640</text></patcit><patcit num="3"><text>Japanese Patent Application Laid-Open No. 06-253331</text></patcit>
<p> The present invention has been made in view of the above problems, and the first object thereof is that the virtual buffer corresponding to the input buffer on the decoding device side is broken or the data occupancy in the virtual buffer is discontinuous. It is an object of the present invention to provide a splicing device and a stream editing device that can seamlessly splice a plurality of coded streams so as not to do so.</p><p> A second object of the present invention provides, in addition to the above object, a splicing device, a stream editing device, and a coding device that reduce image quality deterioration in the vicinity of a splicing point and prevent image quality deterioration in re-encoding processing. To do.</p>
<p> The stream editing device of the present invention is a stream editing device for inputting a plurality of coded bit streams obtained by encoding a plurality of video materials and connecting the plurality of coded bit streams, and is a plurality of stream editing devices. Decoding that inputs a coded bitstream, decodes each coded bitstream in a predetermined section before and after the connection point including the connection points of the plurality of coded bitstreams, and outputs image data in the predetermined section. The means, the target code amount setting means for setting a new target code amount for the image data in the predetermined section output by the decoding means, and the image data in the predetermined section output by the decoding means are set as target codes. A coding means that encodes according to a new target code amount set by the quantity setting means and outputs a new coded bit stream in a predetermined section, and a coding means that outputs the original coded bit stream in the predetermined section. With a coded bitstream output means for connecting and outputting the original coded bitstream and the new coded bitstream before and after the predetermined section by replacing with a new coded bitstream in the predetermined section output by It is equipped with.</p><p> In the stream editing device of the present invention, each coded bit stream in a predetermined section before and after the connection point including the connection point of the plurality of coded bit streams is decoded by the decoding means, and the image data in the predetermined section is obtained. Output, a new target code amount is set for the image data in the predetermined section by the target code amount setting means, and the image data in the predetermined section is encoded according to the new target code amount by the coding means, and is predetermined. A new coded bitstream within the interval is output. Then, the coded bitstream output means replaces the original coded bitstream in the predetermined section with a new coded bitstream in the predetermined section, and the original coded bitstream before and after the predetermined section is replaced with a new coded bitstream. The coded bit stream is concatenated and output.</p><p> The splicing device of the present invention decodes a splicing point setting means for setting a splicing point for each of a plurality of source-encoded streams and a picture near the splicing point of the plurality of source-encoded streams, and decodes the video. Spliced splicing by switching between a decoding means that outputs data, a re-encoding means that re-encodes the decoded video data and outputs a re-encoded stream, and a source-encoded stream and a re-encoded stream. The apparatus is characterized by including a splicing stream generating means for generating a stream, a re-encoding means for controlling the splicing stream so as not to be discontinuous at the time of decoding, and a splice control means for controlling the splicing stream generating means.</p><p> Further, the splice control means of the splicing device of the present invention has a function of calculating the target bit amount of re-encoding in the re-encoding means so that overflow and underflow do not occur in the VBV buffer.</p><p> Further, the splice control means of the splicing device of the present invention re-uses the data occupancy locus of the VBV buffer so as not to be discontinuous at the switching point between the source coded stream and the re-encoded stream or the splice point of the re-encoded stream. It has a function of calculating the target bit amount for re-encoding in the encoding means.</p><p> Further, in the splice control means of the splicing device of the present invention, it can be assumed that the locus of the data occupancy of the VBV buffer corresponding to the re-encoded stream originally had the data occupancy of the VBV buffer. It has a function to control the re-encoding means so as to be close to the trajectory of.</p><p> Further, the splice control means of the splicing apparatus of the present invention extracts a coding parameter included in the source coding stream and selectively reuses the extracted code parameters during the re-encoding process in the re-encoding means. , It has a function of preventing deterioration of the image quality of the splicing stream.</p><p> Further, the splice control means of the splicing apparatus of the present invention extracts information on the quantization characteristics included in the source coding stream, and re-encodes the re-encoding means based on the extracted quantization characteristics. It has a function of controlling the encoding means.</p><p> Further, the splice control means of the splicing device of the present invention extracts information on the quantization characteristics included in the source coded stream for each picture, and in the re-encoding means so that overflow and underflow do not occur in the VBV buffer. The amount of each picture target bit in the re-encoding process is calculated, the information about the quantization characteristic contained in the source coding stream is extracted for each picture, and the quantization characteristic extracted from the source coding stream and the calculated target are calculated. It has a function of calculating a new quantization characteristic based on the amount of bits and controlling the re-encoding means so as to perform a re-encoding process based on the calculated new quantization characteristic.</p><p> Further, the splice control means of the splicing device of the present invention calculates the target bit amount in the re-encoding process in the re-encoding means so that overflow and underflow do not occur in the VBV buffer, and the past included in the source coded stream. By referring to the quantization characteristics generated in the coding process, the target bit amount is assigned to each picture to be re-encoded, and the target bit amount assigned to each picture is re-encoded for each picture. It has a function of controlling the re-encoding means so as to perform processing.</p><p> Further, the splice control means of the splicing device of the present invention calculates the target bit amount in the re-encoding process in the re-encoding means so that overflow and underflow do not occur in the VBV buffer, and the source coded stream is encoded in the past. A target bit amount is assigned to each picture to be re-encoded so as to be close to the amount of bits generated for each picture in the processing, and re-encoding processing is performed for each picture according to the target bit amount assigned to each picture. It has a function to control the re-encoding means so as to perform.</p><p> Further, the splice control means of the splicing device of the present invention prevents deterioration of the image quality of the splicing stream by selectively reusing the motion vector information extracted from the source coding stream during the re-encoding process in the re-encoding means. It has the function of</p><p> Further, the splice control means of the splicing device of the present invention determines whether or not to reuse the motion vector information extracted from the source coding stream during the re-encoding process in the re-encoding means, and it is determined that the motion vector information is reused. In some cases, it has a function of controlling the re-encoding means so as to supply the motion vector extracted from the stream to the motion compensation circuit of the re-encoding means instead of the motion vector detected by the motion detecting means. There is.</p><p> Further, the splice control means of the splicing device of the present invention is located in the vicinity of the splice point where the re-encoding process of the re-encoding means is performed so as not to be predicted from the pictures of different source coding streams across the splicing point. It has a function of setting the prediction direction for the picture.</p><p> Further, the splice control means of the splicing apparatus of the present invention selectively changes the picture type of the picture near the splice point re-encoded by the re-encoding means, thereby deteriorating the image quality of the picture near the splice point of the splicing stream. It has a function to prevent.</p><p> Further, the splice control means of the splicing apparatus of the present invention is a picture of a picture in the vicinity of the splice point re-encoded by the re-encoding means so as not to be predicted from the pictures of different source coding streams across the splice points. It has a function to selectively change the type.</p>
<p> According to the splicing device and the editing method of the present invention, each coded bitstream in a predetermined section before and after the connection point including the connection points of the plurality of coded bitstreams is decoded and the image data in the predetermined section is output. The image data in the predetermined section is encoded according to a new target code amount, a new coded bit stream in the predetermined section is output, and the original coded bit stream in the predetermined section is replaced with a new coded bit stream in the predetermined section. Since the original coded bitstream before and after the predetermined section and the new coded bitstream are concatenated and output by replacing with the coded bitstream, the virtual corresponding to the input buffer on the decoding device side is output. It has the effect of being able to connect a plurality of video materials in arbitrary picture units on the encoded bitstream without causing buffer failure or image failure.</p><p> Further, according to the splicing device and the editing device of the present invention, the coding is performed by using the information included in the original coded bit stream and used at the time of decoding. It has the effect of reducing image quality deterioration in the vicinity.</p><p> Further, according to the splicing device and the editing device of the present invention, when the original coded bit stream is rearranged for pictures to be coded by the bidirectional predictive coding method, the image within a predetermined section is used. When coding the data, the picture type is reconstructed and the coding is performed so that the predictive coding process using the pictures belonging to different coding bit streams is not performed. Even when coding by the predictive coding method, there is an effect that the image is not corrupted.</p><p> Further, according to the splicing device and the editing device of the present invention, a new target code amount is set so as to reduce the deviation of the data occupancy locus of the virtual buffer corresponding to the input buffer on the decoding device side before and after the connection point. Since it is set, it has the effect of more reliably preventing the virtual buffer from collapsing.</p>
Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
FIG. 19 is a block diagram showing a configuration of a splicing device and an editing device according to an embodiment of the present invention. This splicing device and editing device are, for example, a plurality of coded bit streams obtained by encoding video data VDA and VDB of a plurality of video materials by encoders 1A and 1B according to a bidirectional predictive coding method according to the MPEG standard (hereinafter referred to as ). It's just called a stream.) STA and STB are input. The stream in this embodiment may be an elemental stream, a packaged elementary stream, or a transport stream.
The splicing device and the editing device according to the present embodiment include a buffer memory 10 for inputting streams STA and STB and temporarily storing them, a stream counter 11 for counting the number of bits of the streams STA and STB, and a stream counter 11. It includes a stream analysis unit 12 that analyzes the syntax of streams STA and STB, and a splice controller 13 that controls each block described later in order to perform splicing processing.
The splicing device and the editing device further decode the streams STA and STB output from the buffer memory 10 according to the MPEG standard, and output the baseband video data, and the MPEG decoders 14A and 14B and these MPEG decoders 14A, respectively. Switch 15 that switches the output video data from 14B, MPEG encoder 16 that re-encodes the output video data output from this switch 15, and streams STA, STB and MPEG encoder 16 that are output from the buffer memory 10. It is equipped with a switch 17 that outputs a spliced stream STSP by switching and outputting the re-encoded stream STRE.
Each of the above-mentioned blocks will be described in detail below.
The buffer memory 10 temporarily stores the two supplied streams STA and STB in response to a write command from the splice controller 13, which will be described later, and is stored in response to a read command from the splice controller 13. Read streams STA and STB respectively. As a result, the phase and timing of the splicing points of the streams STA and STB can be matched in order to perform splicing at the splicing points set for each of the streams STA and STB.
The stream counter 11 receives the streams STA and STB, counts the number of bits of each of these streams, and supplies the count values to the splice controller 13, respectively. The reason for counting the number of bits of the bitstreams STA and STB supplied in this way is to virtually grasp the trajectory of the data occupancy of the VBV buffer corresponding to the streams STA and STB.
The stream analysis unit 12 extracts appropriate information from the sequence layer, the GOP layer, the picture layer, and the macroblock layer by analyzing the syntax of the streams STA and STB. For example, encoding information such as a picture type indicating a picture type (I, B or P), a motion vector, a quantization step, and a quantization matrix is extracted, and the information is output to the splice controller 13.
These encoding information are encoding information generated in the past encoding processing in the encoders 1A and 1B, and in the splicing apparatus of the present invention, these past encoding information are selectively used in the re-encoding processing. To do.
This splice controller 13 sets the count value output from the stream counter 11, the encoding information output from the stream analysis unit 12, the parameters n0 and m0 for setting the re-encoding interval, and the parameter p0 for indicating the splice point. It receives and controls the switch 15, the MPEG encoder 16 and the switch 17 based on the information. Specifically, the splice controller 13 controls the switching timing of the switch 15 according to the input parameter p0, and controls the switching timing of the switch 17 according to the parameters n0, m0, and p0.
Further, the splice controller 13 does not overflow / underflow the VBV buffer due to the spliced stream based on the count value supplied from the stream counter 11 and the stream analysis unit 12 and the encoding information supplied from the stream analysis unit 12. In addition, a new target code amount is calculated for each picture in the re-encoded interval so that the trajectory of the data occupancy of the VBV buffer is not discontinuous due to the spliced stream.
Further, the price controller 13 adjusts the delay amount of each stream STA, STB in the buffer memory 10 by controlling the write address and the read address of the buffer memory 10, for example, and each stream is based on the presentation time. The phase of the splice points of STA and STB is adjusted.
FIG. 20 is a block diagram showing the configurations of the MPEG decoders 14A and 14B and the MPEG encoder 16 in FIG. 19. In this figure, the MPEG decoders 14A and 14B are represented as the MPEG decoder 14, and the streams STA and STB are represented as the stream ST.
The MPEG decoder 14 is a variable-length decoding circuit (referred to as VLD in the figure) 21 that inputs a stream ST and performs variable-length decoding, and an inverse quantization that dequantizes the output data of the variable-length decoding circuit 21. A circuit (denoted as IQ in the figure) 22 and an inverse DCT circuit (denoted as IDCT in the figure) 23 that performs inverse DCT (inverse discrete cosine conversion) on the output data of this inverse quantization circuit 22 and vice versa. A switch 25 that selectively outputs one of the output data of the DCT circuit 23 and the output data of the inverse DCT circuit 23 and the output data of the adder circuit 24 as the output data of the MPEG decoder 14 and the adder circuit 24 that adds the output data of the DCT circuit 23 and the predicted image data. And two frame memories (denoted as FM1 and FM2 in the figure) 26,27 for holding the output data of the adder circuit 24, and the data held in the frame memory 26,27 and the movement included in the stream ST. It is equipped with a motion compensation unit (denoted as MC in the figure) 28 that performs motion compensation based on vector information to generate predicted image data and outputs the predicted image data to the addition circuit 24.
The MPEG encoder 16 includes an encoder preprocessing unit 30 that performs preprocessing for encoding and the like on the output video data supplied from the MPEG decoder 14. The encoder pre-processing unit 30 performs processing such as rearranging pictures for coding by a bidirectional predictive coding method, macroblocking 16 × 16 pixels, and calculating the coding difficulty of each picture. It has become like.
Further, the MPEG encoder 16 selectively selects one of the subtraction circuit 31 that takes the difference between the output data of the encoder preprocessing unit 30 and the predicted image data, and the output data of the encoder preprocessing unit 30 and the output data of the subtraction circuit 31. A DCT circuit (denoted as DCT in the figure) 33 that performs DCT on the output switch 32 and the output data of this switch 32 in units of DCT (discrete cosine transform) blocks and outputs the DCT coefficient, and this DCT circuit. A quantization circuit (indicated as Q in the figure) 34 that quantizes the output data of 33, and a variable length coding circuit that encodes the output data of this quantization circuit 34 into a variable length code and outputs it as a re-encoded stream STRE. In the figure, it is referred to as VLC.) It is equipped with 35.
The MPEG encoder 16 further performs an inverse DCT on the dequantization circuit (referred to as IQ in the figure) 36 that dequantizes the output data of the quantization circuit 34 and the output data of the dequantization circuit 36. To hold the reverse DCT circuit (denoted as IDCT in the figure) 37, the adder circuit 38 that adds and outputs the output data of the reverse DCT circuit 37 and the predicted image data, and the output data of this adder circuit 38. Based on the two frame memories (denoted as FM1 and FM2 in the figure) 39,40 and the data held in the frame memory 39,40 and the motion vector information, motion compensation is performed to generate predicted image data. , A motion compensating unit (denoted as MC in the figure) 41 that outputs this predicted image data to the subtraction circuit 31 and the addition circuit 38 is provided.
The MPEG encoder 16 further detects a motion vector based on the data held in the frame memories 39 and 40 and the output data of the encoder preprocessing unit 30, and outputs the motion vector information (with ME in the figure). (Note.) 42, the encoding information supplied from the splice controller 13, and the target code amount, and based on this information, the encoding that controls the quantization circuit 34, the inverse quantization circuit 36, and the frame memory 39, 40. Controlled by the controller 43 and the encoding controller 43, one of the motion vector information output from the encoding controller 43 and the motion vector information output from the motion detection circuit 42 is selectively output to the motion compensation unit 41. It is equipped with a switch 44 to be used.
Next, the outline of the operation of the MPEG decoder 14 and the MPEG encoder 16 shown in FIG. 20 will be described.
First, in the MPEG decoder 14, the stream ST is variable-length decoded by the variable-length decoding circuit 21, dequantized by the dequantization circuit 22, reverse DCT is performed by the reverse DCT circuit 23, and the reverse DCT circuit 23 is performed. The output data of is input to the adder circuit 24 and the switch 25. In the case of the I picture, the output data of the inverse DCT circuit 23 is output as the output data of the MPEG decoder 14 via the switch 25. In the case of a P picture or a B picture, the adder circuit 24 adds the output data of the inverse DCT circuit 23 and the predicted image data output from the motion compensation unit 28 to reproduce the P picture or the B picture, and the adder circuit 24 The output data of is output as the output data of the MPEG decoder 14 via the switch 25. Further, the I picture or the P picture is appropriately held in the frame memories 26 and 27 and used for the motion compensation unit 28 to generate the predicted image data.
On the other hand, in the MPEG encoder 16, the output data of the MPEG decoder 14 is input to the encoder preprocessing unit 30, and the encoder preprocessing unit 30 rearranges the pictures, macroblocks, and the like. Here, the encoder preprocessing unit 30 rearranges the pictures based on the picture type information from the splice controller 13.
The output data of the encoder preprocessing unit 30 is input to the subtraction circuit 31 and the switch 32. In the case of an I picture, the switch 32 selectively outputs the output data of the encoder preprocessing unit 30. In the case of P picture or B picture, the subtraction circuit 31 subtracts the predicted image data output from the motion compensation unit 41 from the output data of the encoder preprocessing unit 30, and the switch 32 selects the output data of the subtraction circuit 31. Output.
The output data of the switch 32 is DCT performed by the DCT circuit 33, the output data of the DCT circuit 33 is quantized by the quantization circuit 34, variable length coded by the variable length coding circuit 35, and output as a stream STRE. To.
Further, in the case of the I picture, the output data of the quantization circuit 34 is inversely quantized by the inverse quantization circuit 36, the inverse DCT is performed by the inverse DCT circuit 37, and the output data of the inverse DCT circuit 37 is the frame. It is held in memory 39 or frame memory 40. In the case of the P picture, the output data of the quantization circuit 34 is inversely quantized by the inverse quantization circuit 36, the inverse DCT is performed by the inverse DCT circuit 37, and the output data of the inverse DCT circuit 37 is performed by the adder circuit 38. The predicted image data from the motion compensation unit 41 is added and held in the frame memory 39 or the frame memory 40. The I picture or P picture held in the frame memory 39 or the frame memory 40 is appropriately used by the motion compensation unit 41 to generate predicted image data. Further, the motion detection circuit 42 detects the motion vector based on the data held in the frame memories 39 and 40 and the output data of the encoder preprocessing unit 30, and outputs the motion vector information.
The encoding controller 43 receives the encoding information supplied from the splice controller 13 and the target code amount for each picture, and based on the information, the quantization circuit 34, the inverse quantization circuit 36, the frame memory 39, 40 and the switch. Control 44.
Specifically, when the encoding controller 43 reuses the motion vector included in the encoding information supplied from the splice controller 13, the motion vector information is input to the motion compensation unit 41 via the switch 44. When the motion vector included in the encoding information supplied from the splice controller 13 is not reused, the switch 44 is turned on so that the motion vector information newly generated in the motion detection circuit 42 is input to the motion compensation unit 41. Control. Further, the encoding controller 43 uses the frame memory 39, 40 to hold the pictures necessary for generating the predicted image data based on the picture type included in the encoding information supplied from the splice controller 13. Control 40.
Further, the encode controller 43 monitors the generated code amount of the variable length coding circuit 35 and controls the variable length coding circuit 35. Then, when the generated code amount is insufficient for the set target code amount and the VBV buffer is likely to overflow, the encoding controller 43 adds dummy data to compensate for the shortage of the generated code amount for the target code amount. That is, it is designed to perform stuffing. Further, when the generated code amount exceeds the set target code amount and the VBV buffer is likely to underflow, the encode controller 43 performs skipped macroblock processing (skip macroblock processing) which is a stop processing of coding processing in macroblock units. ISO / IEC 13818-2 7.6.6) is to be performed.
Next, with reference to FIGS. 21 to 26, the picture rearrangement control and the motion compensation control according to the embodiment of the present invention will be described in detail.
FIG. 21 is an explanatory diagram showing an example of a splice point and a re-encoded section in the video data (hereinafter referred to as presentation video data) obtained by decoding with the MPEG decoders 14A and 14B. FIG. 21 (a) shows the presentation video data corresponding to the stream A (STA), and FIG. 21 (b) shows the presentation video data corresponding to the stream B (STB). First, when determining the splice point, a picture is specified on the presentation video data. The splice point is specified by the parameter p0. In addition, a re-encoding section is set as a predetermined section before and after the splice point including the splice point. The re-encoding interval is set by the parameters n0 and m0.
In the following description, as shown in FIG. 21, when the picture of the splice point in the presentation video data corresponding to the stream STA is expressed as An-P0 by using the parameter p0, it is more future than the picture of the splice point An-P0. The picture of is expressed as A (n-P0) +1, A (n-P0) +2, A (n-P0) +3, A (n-P0) +4, ... Yes, the pictures older than the picture at the splice point An-P0 are A (n-P0) -1, A (n-P0) -2, A (n-P0) -3, A (n-P0)- It can be expressed as 4, ..... Similarly, if the picture of the splice point in the presentation video data corresponding to the stream STB is represented as Bm-P0, the pictures in the future than the picture B0 of the splice point are B (m-P0) + 1, B (m-P0). ) +2, B (m-P0) +3, B (m-P0) +4, ... It can be expressed as P0) -1, B (m-P0) -2, B (m-P0) -3, B (m-P0) -4, ...
The re-encoding section shall be n0 sheets before the splice point for the presentation video data corresponding to the stream STA, and m0 sheets after the splice point for the presentation video data corresponding to the stream STB. Therefore, the re-encoding section is picture A (n-P0) + n0 ~ picture An-P0 picture Bm-P0 ~ picture B (m-P0) -m0.
In the present embodiment, the re-encoding process is performed on the re-encoding section set in this way. This re-encoding process returns the baseband video data by decoding the supplied source coded streams STA and STB, connects the two decoded video data at the splice point, and then re-encodes the video data. It is the process of encoding and creating a new stream STRE.
This re-encoding process eliminates the problem of picture sorting and motion compensation. This will be described below.
FIG. 22 shows the arrangement of the pictures before and after decoding in the example shown in FIG. 21. In FIG. 22, (a) shows the stream STA near the re-encoded section, (b) shows the presentation video data corresponding to the stream STA near the re-encoded section, and (c) shows the stream STB near the re-encoded section. The corresponding presentation video data is shown, and (d) shows the stream STB near the re-encoded interval. In the figure, REPA indicates the picture to be re-encoded in the presentation video data corresponding to the stream STA, and REBP indicates the picture to be re-encoded in the presentation video data corresponding to the stream STB. Further, in the figure, the curved arrow indicates the prediction direction.
FIG. 23 shows the state after the stream STA and the stream STB shown in FIGS. 21 and 22 are spliced, and FIG. 23 (a) shows the presentation video data after splicing the two streams. , Figure 23 (b) shows the stream STSP after splicing the two streams. The stream STSP shown in FIG. 23 (b) re-encodes the image data shown in FIG. 23 (a) for the re-encoded section to generate a new stream STRE, and further, the original stream STA before the re-encoded section. It is formed by concatenating (hereinafter referred to as OSTA), a new stream STRE in the re-encoded section, and the original stream STB (hereinafter referred to as OSTB) after the re-encoded section. In the figure, TRE indicates the re-encoding period.
In the examples shown in FIGS. 21 to 23, since the predictive coding process using the pictures belonging to different streams is not performed in the vicinity of the splice point, there is no problem with the presentation order of the pictures.
Next, other examples having different splice points from the examples shown in FIGS. 21 to 23 are shown in FIGS. 24 to 26.
FIG. 24 is an explanatory diagram showing another example of the splice point and the re-encoded interval in the presentation video data. In FIG. 24, (a) shows the presentation video data corresponding to the stream STA, and (b) shows the presentation video data corresponding to the stream STB.
FIG. 25 shows the arrangement of the pictures before and after decoding in the example shown in FIG. 24. In FIG. 25, (a) shows the stream STA near the re-encoded section, (b) shows the presentation video data corresponding to the stream STA near the re-encoded section, and (c) shows the stream STB near the re-encoded section. The corresponding presentation video data is shown, and (d) shows the stream STB near the re-encoded interval.
FIG. 26 shows the arrangement of pictures after splicing in the example shown in FIG. 24. In FIG. 26, (a) shows the image data after concatenating the presentation video data shown in FIG. 25 (b) and the presentation video data shown in FIG. 25 (c), and (b) is spliced. Shows stream STSP. The stream STSP shown in (b) re-encodes the image data shown in (a) for the re-encoded section to generate a new stream STRE, and further, the original stream STA (OSTA) before the re-encoded section. And the new stream STRE in the re-encoded interval and the original stream STB (OSTB) after the re-encoded interval are concatenated.
In the examples shown in FIGS. 23 to 26, since the picture Bm-P0 at the splice point is a P picture, the presentation video data shown in FIG. 25 (b) and FIG. 25 ( When the presentation video data shown in c) is concatenated, when the picture Bm-P0 is re-encoded, the predictive coding process using the pictures belonging to different stream STAs is performed, resulting in image quality deterioration. Therefore, in the present embodiment, the picture type is reconstructed so that the predictive coding process using the pictures belonging to different streams is not performed in the vicinity of the splice point. FIG. 26 (a) shows the state after the picture Bm-P0 is changed from the P picture to the I picture as a result of the reconstruction of this picture type.
Further, in FIG. 26 (a), the picture An-P0 is a B picture, and the picture An-P0 is originally a picture in which the prediction coding process is performed from both directions. -When re-encoding P0, predictive coding processing using pictures belonging to different stream STBs is performed, resulting in image quality deterioration. Therefore, in the present embodiment, even if the B picture is re-encoded, the prediction coding process is performed without using the prediction from the picture side belonging to a different stream. Therefore, in the example shown in FIG. 26 (a), for the picture An-P0, the prediction coding process using only the P picture (A (n-P0) +1) before the picture An-P0 is performed.
The picture type reconstruction setting as described above is performed by the splice controller 13, and the information of the picture type reconstruction setting is given to the encoding controller 43 of the MPEG encoding 17. The encoding controller 43 performs the encoding process according to the setting of the picture type reconstruction. Reuse of the encoding information generated in the past coding process such as motion vector is also performed according to the setting of the picture type reconstruction.
Since the splice processing according to the present embodiment is different from the splice on a simple stream, B that existed after (past) the picture Bm-P0 in the stream STB as shown in FIG. 25 (d). The pictures (B (m-P0) +2 and B (m-P0) + 1) do not exist in the re-encoded picture string because they are discarded after decoding as shown in Fig. 25 (c).
Next, a method of calculating a new target code amount for the image data in the re-encoded section in the present embodiment will be described with reference to FIGS. 27 to 30.
Simply splicing two streams will cause the VBV buffer of the spliced stream to underflow or overflow after the splicing point, or the data occupancy trajectory of the VBV buffer of the spliced stream will be discontinuous. Things happen. The re-encoding process of the splicing apparatus of the present invention for solving these problems will be described with reference to FIGS. 27 to 30.
First, with reference to FIG. 27, the problem that the VBV buffer of the splice stream underflows and the data occupancy of the VBV buffer becomes discontinuous will be described.
FIG. 27 shows an example in which the simple splice processing corresponding to FIG. 23 described above is performed, and FIG. 27 (a) shows the locus of the data occupancy of the VBV buffer of the stream to be re-encoded STRE'. FIG. 27 (b) is a diagram showing the stream to be re-encoded STRE'. In FIG. 27, TRE indicates the re-encoding control period, OSTA indicates the original stream A, and STRE'indicates the re-encoding target stream to be re-encoded. Note that this re-encoded stream STRE'is different from the actually re-encoded re-encoded stream STRE, and indicates a stream that is expected to become such a stream STRE'when simple splice processing is performed. .. Also, OSTB indicates the original stream B, SPVBV indicates the splice point in the VBV buffer, and SP indicates the splice point in the stream.
As shown in FIG. 27 (a), the locus of the VBV buffer of the stream STRE'that is the target of the splice processing becomes the locus of the data occupancy of the VBV buffer of the stream A (STA) before the splice point SP. After the splice point SP, it becomes the trajectory of the data occupancy of the VBV buffer of stream B (STB). Simply splicing stream A and stream B, what is the level at the splice point of the data occupancy of the VBV buffer of stream A (STA) and the level of the data occupancy of the VBV buffer of stream B (STB) at the splice point? Because they are different, the trajectory of the data occupancy of the VBV buffer becomes discontinuous.
In order to realize seamless splicing in which the locus of VBV data occupancy at the splice point of stream A and the locus of data occupancy of VBV of stream B at the splice point are continuous, Fig. 27 (a) ), The start level of VBV data occupancy of stream B at the splice point must match the end level of VBV data occupancy at the splice point of stream A. That is, in order to match those levels, in the example shown in FIG. 27 (a), the trajectory of the data occupancy of the VBV buffer of stream B would have originally been possessed in the re-encoding control period TRE. The level must be lower than the trajectory. The locus that the locus of the data occupancy would originally have is the locus of the data occupancy of the VBV buffer related to the stream B when it is assumed that the supplied stream B was not spliced. , Shown by the extension locus of VBVOST_B in FIG. 27 (a).
As a result, as shown in FIG. 27 (a), this VBV buffer underflows at the extraction timing of the I picture having the largest amount of extraction bits from the VBV buffer.
In the embodiment of the present invention, as shown in FIG. 27 (a), the re-encoding period is such that the locus of the data occupancy of the VBV buffer is continuous at the splice point and no underflow occurs after the splice point. A new target code amount is set for each picture of.
Also, simply lowering the trajectory of the data occupancy of the VBV buffer of stream B so that the trajectory of the data occupancy of the VBV buffer of the spliced stream is continuous at the splash point will only cause underflow. Instead, at the switching point between the re-encoded stream STRE'and the original stream OSTB, the trajectory of the data occupancy of the VBV buffer becomes discontinuous.
Further, in the embodiment of the present invention, as shown in FIG. 27 (a), the trajectory of the data occupancy of the VBV buffer is continuous at the switching point between the re-encoding target stream STRE'and the original stream OSTB. A new target code amount is set for each picture during the re-encoding period. The reason why the locus VBVOST_B of the data occupancy of the VBV buffer corresponding to the original stream OSTB is not controlled is that the locus VBVOST_B is the locus that the locus of the data occupancy of the VBV buffer of the stream B would originally have. This is because it is not possible to control this trajectory. This is because this locus VBVOST_B is the optimum locus determined so that the original stream OSTB does not overflow or underflow, and if the level of this optimum locus is controlled, overflow or underflow may occur. Is.
Next, the problem of overflowing the VBV buffer of the splice stream will be described with reference to FIG. 29, as in the case of the VBV underflow problem. FIG. 29 is an example in which the splice processing corresponding to FIG. 26 described above is performed, and FIG. 29 (a) is a diagram showing the trajectory of the data occupancy of the VBV buffer of the splicing stream STSP. FIG. 29 (b) is a diagram showing the splicing stream STSP. In FIG. 29, TRE indicates the splicing controlled splice period, OSTA indicates the original stream A, STRE indicates the stream to be re-encoded, OSTB indicates the original stream B, and SPVBV indicates the splice in the VBV buffer. Indicates the point, and SP indicates the splice point in the stream.
As shown in FIG. 29 (a), the trajectory of the VBV buffer of the spliced re-encoded stream STRE'is the trajectory of the data occupancy of the VBV buffer of stream A (STA) before the splice point SP. After the splice point SP, it becomes the trajectory of the data occupancy of the VBV buffer of stream B (STB). Simply splicing stream A and stream B, what is the level at the splice point of the data occupancy of the VBV buffer of stream A (STA) and the level of the data occupancy of the VBV buffer of stream B (STB) at the splice point? Because they are different, the trajectory of the data occupancy of the VBV buffer becomes discontinuous.
In order to realize seamless splicing in which the locus of VBV data occupancy at the splice point of stream A and the locus of data occupancy of VBV of stream B at the splice point are continuous, Fig. 29 (a) ), The start level of VBV data occupancy of stream B at the splice point must match the end level of VBV data occupancy at the splice point of stream A. That is, in order to match those levels, in the example shown in FIG. 29 (a), the trajectory of the data occupancy of the VBV buffer of stream B originally had in the re-encoding processing control period TRE. The level must be higher than the deaf trajectory. The locus that the locus of the data occupancy would originally have is the locus of the data occupancy of the VBV buffer related to the stream B when it is assumed that the supplied stream B was not spliced. , Shown by the extension locus of VBVOST_B in FIG. 29 (a).
As a result, as shown in FIG. 29 (a), this VBV buffer overflows after some B pictures and P pictures with a small amount of extraction bits from the VBV buffer are continuously drawn from the VBV buffer. It ends up.
In the embodiment of the present invention, as shown in FIG. 29 (a), the trajectory of the data occupancy of the VBV buffer is continuous at the splice point, and the re-encoding period is set so that overflow does not occur after the splice point. A new target code amount is set for each picture.
Also, simply increasing the trajectory of the data occupancy of the VBV buffer of stream B so that the trajectory of the data occupancy of the VBV buffer of the spliced stream is continuous at the splash point will only cause an overflow. Instead, the trajectory of the data occupancy of the VBV buffer becomes discontinuous at the switching point between the re-encoded stream STRE'and the original stream OSTB.
Further, in the embodiment of the present invention, as shown in FIG. 29 (a), the trajectory of the data occupancy of the VBV buffer is continuous at the switching point between the re-encoding target stream STRE'and the original stream OSTB. A new target code amount is set for each picture during the re-encoding period. The reason why the locus VBVOST_B of the data occupancy of the VBV buffer corresponding to the original stream OSTB is not controlled is that the locus VBVOST_B is the locus that the locus of the data occupancy of the VBV buffer of the stream B would originally have. This is because it is not possible to control this trajectory. This is because this locus VBVOST_B is the optimum locus determined so that the original stream OSTB does not overflow or underflow, and if the level of this optimum locus is controlled, overflow or underflow may occur. Is.
Next, the splice control method of the present invention for avoiding the underflow or overflow of the VBV buffer described above, and the splice control method of the present invention for which the data occupancy of the VBV buffer does not become discontinuous will be described.
In FIGS. 27 to 30, vbv_under indicates the amount of underflow of the VBV buffer, vbv_over indicates the amount of overflow of the VBV buffer, and vbv_gap indicates the amount of VBV buffer at the switching point between the stream to be re-encoded STRE'and the original stream OSTB. It is the data which shows the gap value of.
First, the splice controller 13 determines the trajectory of the data occupancy of the VBV buffer of the original stream OSTA and the VBV buffer of the original stream OSTB based on the bit count value of stream A and the bit count value of stream B supplied from the stream counter 11. And the trajectory of the data occupancy of the VBV buffer of the re-encoded stream STRE'when stream A and stream B are simply spliced. The calculation of the trajectory of the data occupancy of each VBV buffer is performed by subtracting the bit amount output from the VBV buffer according to the presentation time from the bit count value supplied from the stream counter 11 for each presentation time. It can be calculated easily. Therefore, the splice controller 13 has a locus of data occupancy in the VBV buffer of the original stream OSTA, a locus of data occupancy in the VBV buffer of the original stream OSTB, and a re-encoding target when stream A and stream B are simply spliced. It is possible to virtually grasp the trajectory of the data occupancy of the VBV buffer of stream STRE'.
Next, the splice controller 13 refers to the trajectory of the data occupancy of the VBV buffer of the virtually obtained re-encoding target stream STRE', so that the underflow amount (vbv_under) or overflow of the re-encoding target stream STRE' Calculate the quantity (vbv_over). Further, the splice controller 13 refers to the trajectory of the data occupancy of the VBV buffer of the virtually obtained re-encoded stream STRE'and the trajectory of the data occupancy of the VBV buffer of the original stream OSTB (VBVOST_B). Calculates the gap value (vbv_gap) of the VBV buffer at the switching point between the re-encoded stream STRE'and the original stream OSTB.
Then, the splice controller 13 obtains the offset amount vbv_off of the target code amount by the following equations (1) and (2). vbv_off =-(vbv_under --vbv_gap) (1) vbv_off = + (vbv_over --vbv_gap) (2)
When the VBV buffer underflows as shown in Fig. 27 (a), the offset amount vbv_off is calculated using Eq. (1), and as shown in Fig. 29 (a). If the VBV buffer overflows, use Eq. (2) to calculate the offset amount vbv_off.
The splice controller 13 then obtains the target code amount (target bit amount) TBP0 by the following equation (3) using the offset amount vbv_off obtained by the equation (1) or the equation (2).
<img file="JP2005295587A_D0001.tif" />
The target bit amount TBP0 is a value indicating the target bit amount allocated to the picture to be re-encoded. In (3), GB_A is a value indicating the bit generation amount of any of the pictures from picture An-P0 to picture A (n-P0) + n0 in stream A, and is ΣGB_A (n-P0) + i. Is the sum of the generated bits of each picture from picture An-P0 to picture A (n-P0) + n0. Similarly, in Eq. (3), GB_B is a value indicating the amount of generated bits of any of the pictures from picture Bm-P0 to picture B (m-P0) -m0 in stream B, and is ΣGB_B (m-). P0) -i is the sum of the bit generation amounts of each picture from picture Bm-P0 to picture B (m-P0) -m0.
That is, the target code amount TBP0 represented by the equation (3) is a value obtained by adding the VBV offset value vbv_off to the total generated bit amount of picture A (n-P0) + n0 to picture B (m-P0) -m0. Is. By adding the offset value vbv_off and correcting the target bit amount TBP0 in this way, the gap between the loci of the data occupancy at the switching point between the re-encoded stream STSP and the original stream OSTB can be set to 0. Therefore, seamless splicing without joints can be realized.
Next, the splice controller 13 allocates the target bit amount TBP0 obtained based on the equation (3) to the picture A (n-P0) + n0 to the picture B (m-P0) -m0. Normally, the quantization characteristic of each picture is determined so that the target bit amount TBP0 is simply distributed so that the ratio of I picture: P picture: B picture is 4: 2: 1.
However, the splicing apparatus of the present invention does not simply use the quantization characteristic of distributing the target bit amount TBP0 to the I picture: P picture: B picture at a fixed ratio of 4: 2: 1. The new quantization characteristics are determined by referring to the past quantization steps of picture A (n-P0) + n0 to picture B (m-P0) -m0 and the quantization characteristics such as the quantization matrix. Specifically, the encode controller 43 refers to the information of the quantization step and the quantization matrix contained in the stream A and the stream B, and greatly resembles the quantization characteristics of the encoders 1A and 1B during the past encoding process. Determine the quantization characteristics during re-encoding so that they do not differ. However, for a picture whose picture type has been changed due to the reconstruction of the picture, the quantization characteristic is newly calculated at the time of re-encoding without referring to the information of the quantization step and the quantization matrix.
FIG. 28 shows the data occupancy of the VBV buffer when the re-encoding process is performed by the target bit amount TBP0 calculated by the splice controller 13 in order to solve the problem of VBV buffer underflow described in FIG. 27. It is a figure for. Further, FIG. 30 shows the data occupancy of the VBV buffer when the re-encoding process is performed by the target bit amount TBP0 calculated by the splice controller 13 in order to solve the problem of the overflow of the VBV buffer described in FIG. It is a figure to show
Therefore, the re-encoded stream STRE after re-encoding is the locus of the data occupancy of the VBV buffer of the re-encoded stream STRE'in FIG. 27 (a) and FIG. 28, as shown in FIGS. 28 and 30. The trajectory is similar to the trajectory of the data occupancy of the VBV buffer of the re-encoded stream STRE in (a), and the trajectory of the data occupancy of the VBV buffer of the re-encoded stream STRE'in FIG. 29 (a) and FIG. 30 The trajectory is similar to the trajectory of the data occupancy of the VBV buffer of the re-encoded stream STRE in (a).
Next, the operations of the splicing device and the editing device according to the present embodiment will be described with reference to FIGS. 31 and 32. In addition, this embodiment satisfies the Annex C provisions of ISO13818-2 and ISO11172-2 and the AnnexL provisions of ISO13818-1.
First, in step S10, the splice controller 13 receives the splice points p0 for splicing the streams STA and STB at arbitrary picture positions and the re-encoded intervals n0 and m0 in the splicing process. Actually, the operator inputs these parameters from the outside, but the re-encoding sections n0 and m0 may be automatically set according to the GOP configuration of the stream and the like. In the following description, the case of switching from stream STA to stream STB at the splice point will be described as an example, but of course the reverse is also possible.
In step S11, the splice controller 13 controls the write operation of the buffer memory 10 so that the stream STA and the stream STB are temporarily stored in the buffer memory 10, respectively, and the stream STA and the stream STB are based on the presentation time. The read operation of the buffer memory 10 is controlled so that the phases of the splicing points of the stream STB are synchronized.
In step S12, the splice controller 13 selects a picture of the stream STA so that it does not output a picture of the future than the picture An-P0 of the splice point set in the stream STA, and of the splice point set in the stream STB. Select a picture in the stream STB so that no pictures older than picture Bm-P0 are output. For example, in the examples shown in FIGS. 25 (a) and 25 (b), the P picture of picture A (n-P0) -2 is older than the splice point picture An-P0 on the stream STA. However, in the presentation order, it is a future picture rather than picture An-P0. Therefore, the P picture which is this picture A (n-P0) -2 is not output. Further, in the examples shown in FIGS. 25 (c) and 25 (d), the B picture having picture B (m-P0) +2 and picture B (m-P0) + 1 is a splice point on the stream STB. It is a future picture than the picture Bm-P0, but in the presentation order, it is a past picture than the picture Bm-P0. Therefore, the B picture which is picture B (m-P0) +2 and picture B (m-P0) + 1 is not output. Since the splice controller 13 controls the decoders 14A and 14B, the pictures not selected in this step are not supplied to the encoder 16.
In this way, since the pictures to be output are selected based on the presentation order, even if the splicing process is performed, the problem related to the presentation order as described in FIG. 9 does not occur.
In step S13, the splice controller 13 starts a process for setting coding parameters necessary for the picture reconstruction process when performing the re-encoding process. This picture reconstruction process means the process from step S14 to step S30 below, and the parameters set in this process are the picture type, the prediction direction, the motion vector, and the like.
In step S14, the splice controller 13 determines whether or not the picture to be subjected to the picture reconstruction process is the picture An-P0 of the splice point. If the picture to be reconstructed is the splice point picture An-P0, the process proceeds to the next step S15. On the other hand, if this is not the case, that is, if the picture to be subjected to the picture reconstruction process is from picture A (n-P0) + n0 to picture A (n-P0) + 1, the process proceeds to step S20.
In step S15, the splice controller 13 determines whether the picture to be subjected to the picture reconstruction process is a B picture, a P picture, or an I picture. If the picture that is the target of the picture reconstruction process is a B picture, the process proceeds to step S17, and if the picture that is the target of the picture reconstruction process is a P or I picture, the step is performed. Proceed to S18.
In step S16, the splice controller 13 determines in the spliced splice stream STSP whether or not there are two or more B pictures before the picture An-P0. For example, as shown in Figure 26 (b), when there are two B pictures (picture A (n-P0) +2 and picture A (n-P0) +3) before picture An-P0. Goes to step S18. If not, the process proceeds to step S17.
In step S17, the splice controller 13 determines that it is not necessary to change the picture type of the picture An-P0, and sets it as the picture type in the re-encoding process of the picture An-P0 in the past encoding process in the encoder 1A. Set the same picture type as the picture type (B picture). Therefore, the picture An-P0 is encoded again as a B picture at the time of the re-encoding process described later.
In step S18, the splice controller 13 changes the picture type of picture An-P0 from B picture to P picture. The reason for changing the picture type in this way will be described. Reaching the step of this step S18 means that two B pictures (picture A (n-P0) + 2 and picture A (n-P0) + in FIG. 8) are preceded by the B picture (picture An-P0). It means that 3) exists. That is, three B pictures are lined up in the re-encoded stream STRE'. A normal MPEG decoder only has two frame memories to temporarily store the predicted picture, so if three B pictures are arranged consecutively on the stream , The last B picture cannot be decoded. Therefore, as described with reference to FIG. 26, the picture An-P0 can be reliably decoded by changing the picture type of the picture An-P0 from the B picture to the P picture.
In step S19, the splice controller 13 determines that it is not necessary to change the picture type of the picture An-P0, and sets it as the picture type in the re-encoding process of the picture An-P0 in the past encoding process in the encoder 1A. Set the same picture type as the picture type (I picture or P picture).
In step S20, the splice controller 13 determines that it is not necessary to change the picture type of the picture An-P0, and sets it as the picture type in the re-encoding process of the picture An-P0 in the past encoding process in the encoder 1A. Set the same picture type as the picture type (I picture, P picture or B picture).
In step S21, the splice controller 13 sets the prediction direction and the parameters related to the motion vector for each picture. For example, as shown in the examples of FIGS. 25 and 26, when the picture An-P0 that is the target of the picture reconstruction process is a B picture in the original stream OSTA, the picture reconstruction process is performed. The target picture An-P0 is a picture that is bidirectionally predicted from both the P picture of A (n-P0) + 1 and the P picture of A (n-P0) -2. That is, in the past encoding process in the encoder 1A, the picture An-P0 is bidirectionally predicted from both the P picture of A (n-P0) + 1 and the P picture of A (n-P0) -2. It is a generated picture. As described in step S12, the P picture of A (n-P0) -2 is not output as a splicing stream, so A is used as the reverse prediction picture of the picture An-P0 that is the target of the picture reconstruction process. (n-P0) -2 P picture cannot be specified.
Therefore, for the picture An-P0 (B picture) in which the picture type is set to be unchanged in step S17, a forward one-sided prediction such that only the P picture of A (n-P0) + 1 is predicted is performed. Must be done. Therefore, in this case, the splice controller 13 sets forward one-sided prediction such that only the P picture of A (n-P0) + 1 is predicted for the picture An-P0. Similarly, for the picture An-P0 changed from the B picture to the P picture in step S18, a one-sided prediction parameter for predicting only the P picture of A (n-P0) + 1 is set.
In the picture An-P0 (P picture) in which the picture type is set to be unchanged in step S19, the prediction direction is not changed. That is, in this case, the splice controller 13 sets forward one-sided prediction for the picture An-P0 so as to predict only the same picture as the picture predicted at the time of the past encoding process in the encoder 1A.
It is not necessary to change the prediction direction for the pictures from picture A (n-P0) + n0 to picture A (n-P0) + 1 which are set to have no change in the picture type in step S20. That is, in this case, the splice controller 13 predicts the same picture as the picture predicted during the past encoding process in the encoder 1A for the picture A (n-P0) + n0 to the picture A (n-P0) +1. Set the prediction direction so that it does. However, both the picture A (n-P0) +1 and the picture An-P0 are predicted from the bidirectional picture of the P picture or I picture in the forward direction and the I picture or P picture in the reverse direction. In the case of a picture, not only the picture An-P0 but also the picture A (n-P0) +1 must be changed to one-sided prediction such that the prediction is made only from the forward picture.
Further, in this step S21, whether the splice controller 13 reuses the motion vector set by the past encoding process in the encoder 1A for each picture in the re-encoding process based on the newly set prediction direction. Decide whether or not.
As described above, for the P picture and B picture whose prediction direction has not changed, the motion vector used in the past encoding process in the encoder 1A is used as it is at the time of the re-encoding process. For example, in the examples shown in FIGS. 23 and 26, for pictures A (n-P0) + n0 to picture A (n-P0) + 1, the motion vectors used in the past encoding processing in the encoder 1A are used. Reuse when re-encoding.
Further, when picture A (n-P0) +1 and picture An-P0 are a P picture or I picture in the forward direction and a B picture predicted from both directions of the I picture or P picture in the reverse direction. Has been changed to one-sided prediction such that prediction is made from only the forward picture, so it is necessary to use only the motion vector corresponding to the forward picture. That is, in step S21, when the picture A (n-P0) +1 and the picture An-P0 are B pictures, the splice controller 13 sets a motion vector regarding the forward picture for these pictures. Use it and set it so that the motion vector of the picture in the opposite direction is not used.
If the picture An-P0 is a picture predicted on one side in the opposite direction only from the future picture A (n-P0) -2 in the past encoding process in the encoder 1A, the re-encoding process is performed. In, a new motion vector corresponding to A (n-P0) +1 is generated without using any motion vector generated during the past encoding process in the encoder 1A. That is, the splice controller 13 makes a setting in step S21 that the past motion vector is not used at all.
Next, in step S22, has the splice controller 13 set parameters for the picture type, prediction direction, and past motion vector for all the pictures from picture A (n-P0) + n0 to picture An-P0? Judge whether or not.
In step S23, the splice controller 13 determines whether or not the picture to be subjected to the picture reconstruction process is the picture Bm-P0 of the splice point. If the picture to be reconstructed is the splice point picture Bm-P0, the process proceeds to the next step S24. On the other hand, if this is not the case, that is, if the picture to be subjected to the picture reconstruction process is picture B (m-P0) -1 to picture B (m-P0) + m0, the process proceeds to step S28.
In step S24, the splice controller 13 determines whether the picture to be subjected to the picture reconstruction process is a B picture, a P picture, or an I picture. If the picture that is the target of the picture reconstruction process is a B picture, the process proceeds to step S25. If the picture that is the target of the picture reconstruction process is a P picture, the process proceeds to step S26. If the picture to be processed for the picture reconstruction process is an I picture, the process proceeds to step S27.
In step S25, the splice controller 13 determines that it is not necessary to change the picture type of the picture Bm-P0 during the re-encoding process, as in the examples shown in FIGS. 22 and 23, and re-encodes the picture Bm-P0. As the picture type at the time of encoding processing, the same picture type as the picture type (B picture) set in the past encoding processing in the encoder 1B is set.
In step S26, the splice controller 13 changes the picture type of picture Bm-P0 from P-picture to I-picture, as in the examples shown in FIGS. 25 and 26. The reason for changing the picture type in this way will be described. Since the P picture is a one-sided prediction picture predicted from the forward I picture or the P picture, it is a picture that always exists at a position behind those predicted pictures on the stream. If the first picture Bm-P0 of the splice point in the stream STB is a P picture, it must be predicted from the forward picture of the stream STA existing before this picture Bm-P0. Since stream STA and stream STB are completely different, it is clear that if the picture type of the first picture Bm-P0 is set to P picture, even if this picture is decoded, the picture quality will be considerably deteriorated.
Therefore, when the picture type of the first picture Bm-P0 of the splice point in the stream STB is P picture, the splice controller 13 changes the picture type of this picture Bm-P0 to I picture.
In step S27, the splice controller 13 determines that it is not necessary to change the picture type of the picture Bm-P0, and sets it as the picture type in the re-encoding process of the picture Bm-P0 in the past encoding process in the encoder 1B. Set the same picture type as the picture type (I picture).
In step S28, the splice controller 13 determines that there is no need to change the picture types from picture B (m-P0) -1 to picture B (m-P0) -m0, and during the re-encoding process of those pictures. As the picture type, the same picture type as the picture type (I picture, P picture or B picture) set in the past encoding process in the encoder 1B is set.
In step S29, the splice controller 13 sets the prediction direction and the motion vector for each picture. For example, as in the examples shown in FIGS. 22 and 23, when the picture Bm-P0 that is the target of the picture reconstruction process is a B picture in the original stream OSTB, it is the target of the picture reconstruction process. The picture Bm-P0 is a picture that is bidirectionally predicted from both the P picture of B (m-P0) + 1 and the I picture of B (m-P0) -2. That is, in the past encoding process in the encoder 1B, the picture Bm-P0 is bidirectionally predicted from both the P picture of B (m-P0) + 1 and the I picture of B (m-P0) -2. It is a generated picture. As described in step S12, the P picture of B (m-P0) + 1 is not output as a splicing stream, so B is used as the forward prediction picture of the picture Bm-P0 that is the target of the picture reconstruction process. (m-P0) + 1 P picture cannot be specified.
Therefore, for the picture Bm-P0 (B picture) in which the picture type is set to be unchanged in step S25, a reverse one-sided prediction such that only the I picture of B (m-P0) -2 is predicted is performed. Must be done. Therefore, in this case, the splice controller 13 sets the prediction direction for the picture Bm-P0 so as to perform one-sided prediction in the opposite direction so as to predict only the I picture of B (m-P0) -2. To do.
It is not necessary to change the prediction direction for the pictures from picture B (m-P0) + m0 to picture B (m-P0) + 1 which are set to have no change in the picture type in step S28. That is, in this case, the splice controller 13 predicts the same picture as the picture predicted during the past encoding process in the encoder 1B for the picture B (m-P0) + m0 to the picture B (m-P0) +1. Set the prediction direction so that it does. However, when B (m-P0) -1 is a B picture, B (m-P0) -2 is opposed to picture B (m-P0) -1 as in the case of picture Bm-P0. A prediction direction is set such that one-sided prediction in the opposite direction is performed so as to predict only the I picture of.
Further, in this step S29, whether the splice controller 13 reuses the motion vector set by the past encoding process in the encoder 1B for each picture in the re-encoding process based on the newly set prediction direction. Decide whether or not.
As described above, for the P picture and B picture whose prediction direction has not changed, the motion vector used in the past encoding process in the encoder 1B is used as it is at the time of the re-encoding process. For example, in FIGS. 22 and 23, for each picture from the I picture of B (m-P0) -2 to the P picture of B (m-P0) -m0, the motion used in the past encoding. Use the vector as it is.
In the past encoding process in encoder 1B, the pictures Bm-P0 and picture B that were bidirectionally predicted from both the P picture of B (m-P0) + 1 and the I picture of B (m-P0) -2. For (m-P0) -1, the prediction direction has been changed to one-sided prediction that predicts only the I picture of B (m-P0) -2, so picture B (m-P0) + 1 It is necessary to use only the motion vector corresponding to picture B (m-P0) -2 for picture Bm-P0 without using the motion vector corresponding to. That is, in this step S29, the splice controller 13 reuses the past motion vector in only one direction for the picture Bm-P0 and the picture B (m-P0) -1, and the other past motion vector. Set not to use.
Next, in step S30, the splice controller 13 determines whether or not the parameters related to the picture type, the prediction direction, and the motion vector are set for all the pictures from the picture Bm-P0 to the picture B (m-P0) -m0. to decide.
In step S31, the splice controller 13 calculates the target bit amount (TBP0) to be generated during the re-encoding period based on the equation (3) already described. This will be described in detail below. First, the splice controller 13 determines the trajectory of the data occupancy of the VBV buffer of the original stream OSTA and the VBV buffer of the original stream OSTB based on the bit count value of stream A and the bit count value of stream B supplied from the stream counter 11. And the trajectory of the data occupancy of the VBV buffer of the re-encoded stream STRE'when stream A and stream B are simply spliced.
Next, the splice controller 13 analyzes the trajectory of the data occupancy of the VBV buffer of the virtually obtained re-encoding target stream STRE', so that the underflow amount (vbv_under) or overflow of the re-encoding target stream STRE' Calculate the quantity (vbv_over). Further, the splice controller 13 compares the trajectory of the data occupancy of the VBV buffer of the virtually obtained re-encoded stream STRE'with the trajectory of the data occupancy of the VBV buffer of the original stream OSTB (VBVOST_B). Calculates the gap value (vbv_gap) of the VBV buffer at the switching point between the re-encoded stream STRE'and the original stream OSTB. Subsequently, the splice controller 13 obtains the offset amount vbv_off of the target code amount by the equations (1) and (2) already described, and further uses the offset amount vbv_off obtained by the equation (1) or the equation (2). Then, the target code amount (target bit amount) TBP0 is obtained by the equation (3) already described.
Next, in step S32, the splice controller 13 applies the target bit amount TBP0 obtained based on the equation (3) to picture A (n-P0) + n0 to picture B (m-P0) -m0. Based on the allocation, the quantization characteristics set for each picture are determined. In the splicing apparatus of the present invention, the past quantization steps in the encoders 1A and 1B of each picture A (n-P0) + n0 to picture B (m-P0) -m0 and the quantization characteristics such as the quantization matrix are referred to. To determine new quantization characteristics. Specifically, the splicing controller 13 first inputs the coding parameter information generated in the past coding processing in the encoders 1A and 1B such as the quantization step and the quantization matrix included in the stream A and the stream B. Received from the stream analysis unit 12.
Then, the splicing controller 13 assigns the target bit amount TBP0 obtained based on the equation (3) to the code amount assigned to the picture A (n-P0) + n0 to the picture B (m-P0) -m0. Rather than determining the quantization characteristics solely from the target bit amount, referring to the code amount assigned from this target bit amount TBP0 and these past coding parameter information, the quantization characteristics at the time of encoding in the encoders 1A and 1B are large. Determine the quantization characteristics during re-encoding so that they do not differ. However, as described in step S18 and step S26, for a picture whose picture type has been changed by the picture reconstruction process, a new picture is newly created during the re-encoding process without referring to the information in the quantization step and the quantization matrix. Calculate the quantization characteristics.
Next, in step S33, the splice controller 13 decodes picture A (n-P0) + n0 to picture B (m-P0) -m0 included in the re-encoding period.
Next, in step S34, it occurs using the quantization characteristics set for each picture A (n-P0) + n0 to picture B (m-P0) -m0 in step S34 and step S32. Re-encode picture A (n-P0) + n0 ~ picture B (m-P0) -m0 while controlling the amount of bits.
In this re-encoding process, the splice controller 13 tells the encoding controller to supply the motion vector to the motion compensation unit 41 via the switch 44 when reusing the motion vector used in the past encoding process in the encoders 1A and 1B. When a control signal is given and the motion vector used in the past encoding processing in the encoders 1A and 1B is not used, the motion vector newly generated by the motion detection unit 42 is supplied to the motion compensation unit 41 via the switch 41. Control the encoding controller 43 so as to. At that time, the encoding controller 43 controls the frame memory 39,40 so that the picture required for generating the predicted image data is held in the frame memory 39,40 based on the picture type information from the splice controller 13. .. Further, the encode controller 43 sets the quantization characteristics set for each picture in the re-encoded section supplied from the splice controller 13 for the quantization circuit 34 and the inverse quantization circuit 36.
In step S35, the splice controller 13 controls the switch 17 to select one of the streams STA, STB output from the buffer memory 10 and one of the new streams STRE in the re-encoded section output from the MPEG encoder 16. By selectively outputting, the stream STA before the re-encoded section, the new stream STRE in the re-encoded section, and the stream STB after the re-encoded section are concatenated and output as a spliced stream STSP.
In the present embodiment, a new stream STRE in the re-encoded section obtained by re-encoding the MPEG encoder 17 while controlling the rate according to the target bit amount TBP0 is generated by the switch 17 as a picture in the original stream. Fit in the position of A (n-P0) + n0 ~ Picture B (m-P0) -m0. This provides a seamless splice.
When splicing a stream STB to a stream STA, adjust the state of the VBV buffer after splicing of the stream STB to the state before splicing in order to prevent overflow and underflow of the VBV buffer in the stream STB after splicing. However, it is a condition for buffer control to guarantee continuous picture presentation. In the present embodiment, as a general rule, this condition is satisfied by setting a new target code amount (target bit amount) as described above.
As described above, according to the present embodiment, each stream in the re-encoded section including the splice points of a plurality of streams is decoded, and the obtained image data in the re-encoded section is used as a new target code amount. Re-encoded according to, generated a new stream in the re-encoded section, and connected the original stream and the new stream before and after the re-encoded section for output, so the decoder (IRD) side It is possible to seamlessly splice a plurality of video materials in any picture unit on the stream without breaking the VBV buffer of the above and without causing the picture presentation to continuously break the image before and after the splice point.
Further, in the present embodiment, at the time of re-encoding, the information used for decoding the motion vector or the like is reused. That is, in the present embodiment, the motion detection information such as the motion vector detected at the time of the previous encoding is reused for the pictures other than the pictures for which the motion detection is invalid in the vicinity of the splice point. The image quality is not deteriorated by the encoding process.
Furthermore, when determining the quantization characteristics at the time of re-encoding, the information of the quantization step and the quantization matrix used at the time of decoding is referred to. Therefore, according to the present embodiment, deterioration of image quality due to repeated decoding and re-encoding can be suppressed as much as possible, and the accuracy of image reconstruction can be guaranteed.
In the compression coding process according to the MPEG standard, the information used for decoding is reused from the calculation accuracy of orthogonal transform and non-linear operations such as mismatch processing (processing to insert an error in the high frequency range of DCT coefficient). Reconstruction errors that cannot be suppressed by themselves are introduced. Therefore, at the time of re-encoding, even if the information used at the time of decoding is reused, the complete image cannot be reconstructed. Therefore, considering the existence of image quality deterioration, decoding and re-encoding should be performed only on the pictures in a certain section in the vicinity of the splice point including the splice point. Therefore, in the present embodiment, instead of decoding and re-encoding the entire stream, a re-encoding section is set, and only that section is decoded and re-encoded. This also makes it possible to prevent deterioration of image quality. The re-encoding section may be automatically set according to the degree of deterioration of image quality, the GOP length, and the GOP structure. Can be set arbitrarily in consideration of.
The present invention is not limited to the above embodiment. For example, the method of calculating the new target code amount for the re-encoded section is not limited to the methods shown in the equations (1) to (3), and can be appropriately set. Is.
<figref num="1">It is a block diagram which shows the schematic structure of the transmission system according to the MPEG standard.</figref><figref num="2">It is explanatory drawing for demonstrating the VBV buffer.</figref><figref num="3">It is explanatory drawing for demonstrating the rearrangement of a picture in an encoder required in the bidirectional predictive coding system of an MPEG standard.</figref><figref num="4">It is explanatory drawing which shows the relationship between the rearrangement of a picture in an encoder, and the coding process.</figref><figref num="5">It is explanatory drawing for demonstrating the rearrangement of a picture in a decoder.</figref><figref num="6">It is explanatory drawing which shows the relationship between the rearrangement of a picture in a decoder and the decoding process.</figref><figref num="7">It is explanatory drawing for demonstrating motion detection and motion compensation in the bidirectional predictive coding system of the MPEG standard.</figref><figref num="8">It is explanatory drawing which shows an example of the relationship between the order of the picture in the stream before and after splicing, and the order of picture presentation after splicing when the stream is simply spliced.</figref><figref num="9">It is explanatory drawing which shows another example of the relationship between the order of the picture in the stream before and after splicing, and the order of picture presentation after splicing when the stream is simply spliced.</figref><figref num="10">It is explanatory drawing which shows an example of the relationship between the order of a picture in a stream after splicing, and the order of a picture presentation when the stream is simply spliced.</figref><figref num="11">It is explanatory drawing which shows another example of the relationship between the order of a picture in a stream after splicing and the order of a picture presentation when the stream is simply spliced.</figref><figref num="12">It is explanatory drawing which shows still another example of the relationship between the order of a picture in a stream after splicing and the order of a picture presentation when the stream is simply spliced.</figref><figref num="13">It is explanatory drawing which shows the example of the ideal stream splice which satisfies the condition of a picture presentation order, motion compensation and a VBV buffer.</figref><figref num="14">It is explanatory drawing which shows the normal stream which satisfies the constraint of a VBV buffer.</figref><figref num="15">It is explanatory drawing which shows the other normal stream which satisfies the constraint of a VBV buffer.</figref><figref num="16">It is explanatory drawing for demonstrating an example of the case where two streams are simply spliced at an arbitrary position.</figref><figref num="17">It is explanatory drawing for demonstrating another example of the case where two streams are simply spliced at an arbitrary position.</figref><figref num="18">It is explanatory drawing for demonstrating still another example in the case where two streams are simply spliced at an arbitrary position.</figref><figref num="19">It is a block diagram which shows the structure of the splicing apparatus and stream editing apparatus which concerns on one Embodiment of this invention.</figref><figref num="20">It is a block diagram which shows the structure of the MPEG decoder and the MPEG encoder in FIG.</figref><figref num="21">It is explanatory drawing which shows an example of the splice point and the re-encoding section in the presentation video data obtained by decoding by the MPEG decoder in FIG.</figref><figref num="22">It is explanatory drawing which shows the arrangement of the picture before and after decoding of two streams in the example shown in FIG.</figref><figref num="23">It is explanatory drawing which shows the arrangement of the picture of the splice stream after splice in the example shown in FIG.</figref><figref num="24">It is explanatory drawing which shows another example of the splice point and the re-encoded section in the presentation video data obtained by decoding by the MPEG decoder in FIG.</figref><figref num="25">It is explanatory drawing which shows the arrangement of the picture of two streams before and after decoding in the example shown in FIG.</figref><figref num="26">It is explanatory drawing which shows the arrangement of the picture of the splice stream after splice in the example shown in FIG.</figref><figref num="27">It is explanatory drawing which shows the example which underflow occurs in the data occupancy of the VBV buffer.</figref><figref num="28">It is explanatory drawing which shows the example which improved the underflow explained in FIG. 27 by the splicing apparatus of this invention.</figref><figref num="29">It is explanatory drawing which shows the example which overflow occurs in the data occupancy of the VBV buffer.</figref><figref num="30">It is explanatory drawing which shows the example which improved the overflow described in FIG. 29 by the splicing apparatus of this invention.</figref><figref num="31">It is a flowchart for demonstrating operation of the splicing apparatus and stream editing apparatus of this invention.</figref><figref num="32">It is a flowchart for demonstrating operation of the splicing apparatus and stream editing apparatus of this invention.</figref>
Code description
1A, 1B ... encoder, 10 ... buffer memory, 11 ... stream counter, 12 ... stream analyzer, 13 ... splice controller, 14A, 14B ... MPEG decoder, 15 ... Switch, 16 ... MPEG encoder, 17 ... switch, VDA, VDB ... video data, STA, STB ... encoded stream, STRE ... re-encoded stream, STSP ... spliced stream , N0, m0, p0 ... parameters.
1 sheet
Sheet 1
6 priority claims, no other members on record
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 1997199923 | Japan | – | |
| 19992397 | Japan | A | |
| 2005143130 | Japan | A | |
| 1997199923 | – | – | – |
| JP19970199923 | – | – | – |
| JP20050143130 | – | – | – |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Decision of refusalA02 | A02 | |
| Notification of reasons for refusalA131 | A131 |
Numbers
- Publication
- 2005295587
- Publication, DOCDB
- 2005295587
- Publication, EPODOC
- JP2005295587
- Application
- 143130
- Application, DOCDB
- 2005143130
- Application, EPODOC
- JP20050143130
Titles3
- Japanese
- 編集装置、編集方法、再符号化装置、再符号化方法、スプライシング装置及びスプライシング方法
- English
- Editing device, editing method, recoding device, recoding method, splicing device and splicing method
- English
- EDITING DEVICE AND METHOD, RE-CODING DEVICE AND METHOD, AND SPLICING DEVICE AND METHOD
Classification
- CPC, 26
- H04N21/23424
- H04N5/92
- G11B27/034
- G11B27/036
- G11B2220/2562
- H04N7/52
- H04N21/23406
- H04N21/2343
- H04N21/44004
- H04N21/44016
- H04N21/440254
- H04N19/114
- H04N19/124
- H04N19/146
- H04N19/149
- H04N19/15
- H04N19/159
- H04N19/172
- H04N19/177
- H04N19/40
- H04N19/577
- H04N19/61
- H04N19/142
- H04N19/50
- H04N19/513
- H04N19/52
- IPC, 39
- H04N5 91
- G06T9 00
- G11B20 10
- G11B27 034
- G11B27 036
- H03M7 30
- H03M7 36
- H04N5 92
- H04N7 24
- H04N7 52
- H04N19 114
- H04N19 12
- H04N19 124
- H04N19 134
- H04N19 136
- H04N19 139
- H04N19 146
- H04N19 152
- H04N19 156
- H04N19 159
- H04N19 172
- H04N19 189
- H04N19 40
- H04N19 423
- H04N19 44
- H04N19 50
- H04N19 503
- H04N19 51
- H04N19 513
- H04N19 577
- H04N19 61
- H04N19 625
- H04N19 70
- H04N19 85
- H04N19 91
- H04N21 234
- H04N21 2343
- H04N21 44
- H04N21 4402