Moving image encoding method and apparatus
Abstract
This record has no abstract on file.
Term
Projected expiry 15 July 2028.
- Priority
- Filed
- Granted
- Today
- Projected expiry
12 claims: 12 independent, 0 dependent
- 1The moving image is coded by switching between a variable-length coding method that performs adaptive coding according to the surrounding situation and a binary arithmetic coding method that performs adaptive coding according to the surrounding situation. A determination step for determining a continuous section to be seamlessly continuously reproduced, and an in-screen coded picture in which the continuous section is configured and the first picture in the decoding order is an in-screen coded picture. In the reproduction section, a generation step of generating a moving image stream by encoding a moving image using either the variable length coding method or the arithmetic coding method, and a continuous reproduction section. Management information that generates management information including a first flag information indicating that the coding method is fixed to either the variable length coding method or the arithmetic coding method for each reproduction section. A moving image coding method characterized by having a generation step. 周囲の状況に応じて、適応的に符号化を行う可変長符号化方式と、周囲の状況に応じて、適応的に符号化を行う2値の算術符号化方式とを切り替えて動画像を符号化する動画像符号化方法であって、 シームレスな連続再生の対象となる連続区間を決定する決定ステップと、 前記連続区間を構成し、復号順で先頭のピクチャが画面内符号化ピクチャである各再生区間において、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方を用いて、動画像を符号化することにより動画像ストリームを生成する生成ステップと、 連続する前記再生区間における符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方に固定されていることを、前記再生区間ごとに示す第1フラグ情報を含む管理情報を生成する管理情報生成ステップと を有することを特徴とする動画像符号化方法。
- 2The moving image is coded by switching between a variable-length coding method that performs adaptive coding according to the surrounding situation and a binary arithmetic coding method that performs adaptive coding according to the surrounding situation. It is a moving image coding device to be converted, and a determination means for determining a continuous section to be seamlessly continuously reproduced, and each of the continuous sections being configured and the first picture in the decoding order being an in-screen coded picture. In the reproduction section, a coding means for generating a moving image stream by encoding a moving image by using either the variable length coding method or the arithmetic coding method, and the continuous reproduction section. Generates management information including a first flag information indicating that the coding method in the above is fixed to either the variable length coding method or the arithmetic coding method for each reproduction section. A moving image coding device comprising means. 周囲の状況に応じて、適応的に符号化を行う可変長符号化方式と、周囲の状況に応じて、適応的に符号化を行う2値の算術符号化方式とを切り替えて動画像を符号化する動画像符号化装置であって、 シームレスな連続再生の対象となる連続区間を決定する決定手段と、 前記連続区間を構成し、復号順で先頭のピクチャが画面内符号化ピクチャである各再生区間において、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方を用いて、動画像を符号化することにより動画像ストリームを生成する符号化手段と、 連続する前記再生区間における符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方に固定されていることを、前記再生区間ごとに示す第1フラグ情報を含む管理情報を生成する生成手段と を備えることを特徴とする動画像符号化装置。
- 3It was encoded by switching between a variable-length coding method that performs adaptive coding according to the surrounding situation and a binary arithmetic coding method that performs adaptive coding according to the surrounding situation. It is a moving image decoding method that decodes a moving image from a moving image stream including a moving image and other information, and management information, and constitutes a continuous section to be seamlessly continuously reproduced in the order of decoding. In the reproduction section in which the first picture is an in-screen coded picture, the coding method in the continuous reproduction section is fixed to either the variable length coding method or the arithmetic coding method. The first extraction step of extracting the first flag information indicating the above for each reproduction section from the management information, and when the first flag information is extracted, decoding is performed using the same coding method. A determination step for determining and a second extraction step for extracting second flag information indicating whether the coding method is the variable length coding method or the arithmetic coding method from the moving image stream. When, When it is determined in the determination step that decoding is performed using the same coding method, the decoding step in which decoding is seamlessly performed at the connection points of the continuous reproduction sections using the coding method indicated by the second flag. A moving image decoding method characterized by having. 周囲の状況に応じて、適応的に符号化を行う可変長符号化方式と、周囲の状況に応じて、適応的に符号化を行う2値の算術符号化方式とを切り換えて符号化された動画像と、その他の情報とを含む動画像ストリームと、管理情報とから動画像を復号する動画像復号化方法であって、 シームレスな連続再生の対象となる連続区間を構成し、復号順で先頭のピクチャが画面内符号化ピクチャである再生区間において、連続する前記再生区間における符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方に固定されていることを示す第1フラグ情報を、前記管理情報から前記再生区間ごとに抽出する第1の抽出ステップと、 前記第1フラグ情報を抽出した場合に、同一の符号化方式を用いて復号を行うことを決定する決定ステップと、 前記動画像ストリームから、符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれであるかを示す第2フラグ情報を抽出する第2の抽出ステップと、 前記決定ステップにおいて同一の符号化方式を用いて復号すると決定した場合に、前記第2フラグが示す符号化方式を用いて、連続する前記再生区間の接続点においてシームレスに復号を行う復号ステップと を有することを特徴とする動画像復号化方法。
- 4It was encoded by switching between a variable-length coding method that performs adaptive coding according to the surrounding situation and a binary arithmetic coding method that performs adaptive coding according to the surrounding situation. A moving image decoding device that decodes moving images from moving images, a moving image stream containing other information, and management information, and constitutes a continuous section to be seamlessly played back, in the order of decoding. In the reproduction section in which the first picture is an in-screen coded picture, the coding method in the continuous reproduction section is fixed to either the variable length coding method or the arithmetic coding method. The first extraction means for extracting the first flag information indicating the above for each reproduction section from the management information, and when the first flag information is extracted, decoding is performed using the same coding method. A second extraction means for extracting from the determination means for determining and the second flag information indicating whether the coding method is the variable length coding method or the arithmetic coding method from the moving image stream. When, When the determination means determines to decode using the same coding method, the decoding means that seamlessly decodes at the connection point of the continuous reproduction section using the coding method indicated by the second flag. A moving image decoding device characterized by having. 周囲の状況に応じて、適応的に符号化を行う可変長符号化方式と、周囲の状況に応じて、適応的に符号化を行う2値の算術符号化方式とを切り換えて符号化された動画像と、その他の情報とを含む動画像ストリームと、管理情報とから動画像を復号する動画像復号化装置であって、 シームレスな連続再生の対象となる連続区間を構成し、復号順で先頭のピクチャが画面内符号化ピクチャである再生区間において、連続する前記再生区間における符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方に固定されていることを示す第1フラグ情報を、前記管理情報から前記再生区間ごとに抽出する第1の抽出手段と、 前記第1フラグ情報を抽出した場合に、同一の符号化方式を用いて復号を行うことを決定する決定手段と、 前記動画像ストリームから、符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれであるかを示す第2フラグ情報を抽出する第2の抽出手段と、 前記決定手段において同一の符号化方式を用いて復号すると決定した場合に、前記第2フラグが示す符号化方式を用いて、連続する前記再生区間の接続点においてシームレスに復号を行う復号手段と を有することを特徴とする動画像復号化装置。
- 5A moving image encoded by switching between a variable-length coding method that adaptively encodes according to the surrounding situation and a binary arithmetic coding method that adaptively encodes according to the surrounding situation. It is a recording method for recording a moving image stream including an image and management information on a computer-readable recording medium, and constitutes a determination step for determining a continuous section to be seamlessly continuously reproduced and the continuous section. By encoding the moving image using either the variable length coding method or the arithmetic coding method in each reproduction section in which the first picture in the decoding order is the in-screen coded picture. The reproduction section indicates that the generation step for generating the moving image stream and the coding method in the continuous reproduction section are fixed to either the variable length coding method or the arithmetic coding method. Recording on a recording medium, which comprises a management information generation step for generating management information including the first flag information shown for each, and a recording step for recording the moving image stream and the management information on a recording medium. Method. 周囲の状況に応じて、適応的に符号化を行う可変長符号化方式と、周囲の状況に応じて、適応的に符号化を行う2値の算術符号化方式とを切り替えて符号化した動画像を含む動画像ストリームと、管理情報とをコンピュータ読み取り可能な記録媒体に記録する記録方法であって、 シームレスな連続再生の対象となる連続区間を決定する決定ステップと、 前記連続区間を構成し、復号順で先頭のピクチャが画面内符号化ピクチャである各再生区間において、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方を用いて、動画像を符号化することにより動画像ストリームを生成する生成ステップと、 連続する前記再生区間における符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方に固定されていることを、前記再生区間ごとに示す第1フラグ情報を含む管理情報を生成する管理情報生成ステップと、 前記動画像ストリームと管理情報とを記録媒体に記録する記録ステップと、 を有することを特徴とする記録媒体への記録方法。
- 6It was encoded by switching between a variable-length coding method that performs adaptive coding according to the surrounding situation and a binary arithmetic coding method that performs adaptive coding according to the surrounding situation. A computer-readable recording medium in which data including moving images and other information is recorded, a computer-readable recording medium in which data including management information is recorded, and a moving image decoding device that reads and decodes the data from the recording medium. A moving image decoding system composed of the above, the data recorded on the recording medium constitutes a continuous section to be seamlessly continuously reproduced, and the first picture in the decoding order is an in-screen encoded picture. In each reproduction section, the moving image is encoded by using either the variable length coding method or the arithmetic coding method, and the coding method in the continuous reproduction section is the said. Management information including the first flag information indicating for each reproduction section that it is fixed to either the variable length coding method or the arithmetic coding method, and A second indicating whether the coding method is the variable length coding method or the arithmetic coding method for each predetermined unit of the coded moving image and the coded moving image. The moving image decoding device has a moving image stream including flag information, and the moving image decoding device has a first extraction means for extracting the first flag information from the management information for each reproduction section, and the first flag information. Is the same in the determination means as the determination means for deciding to perform decoding using the same coding method and the second extraction means for extracting the second flag information from the moving image stream. When it is determined to decode using the coding method of the above, it is characterized by having a decoding means that seamlessly decodes at the connection point of the continuous reproduction section by using the coding method indicated by the second flag. Video decoding system. 周囲の状況に応じて、適応的に符号化を行う可変長符号化方式と、周囲の状況に応じて、適応的に符号化を行う2値の算術符号化方式とを切り替えて符号化された動画像と、その他の情報とを含む動画像ストリームと、管理情報とを含むデータが記録されたコンピュータ読み取り可能な記録媒体と、 前記記録媒体から、前記データを読み取り復号を行う動画像復号化装置とから構成される動画像復号化システムであって、 前記記録媒体に記録されたデータは、 シームレスな連続再生の対象となる連続区間を構成し、復号順で先頭のピクチャが画面内符号化ピクチャである各再生区間において、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方を用いて、動画像を符号化されており、 連続する前記再生区間における符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれか一方に固定されていることを、前記再生区間ごとに示す第1フラグ情報を含む管理情報と、 符号化された動画像、及び、前記符号化された動画像の所定単位毎に、符号化方式が、前記可変長符号化方式、または、前記算術符号化方式のいずれであるかを示す第2フラグ情報を含む動画像ストリームとを有し、 前記動画像復号化装置は、 前記第1フラグ情報を、前記管理情報から前記再生区間ごとに抽出する第1の抽出手段と、 前記第1フラグ情報を抽出した場合に、同一の符号化方式を用いて復号を行うことを決定する決定手段と、 前記動画像ストリームから前記第2フラグ情報を抽出する第2の抽出手段と、 前記決定手段において同一の符号化方式を用いて復号すると決定した場合に、前記第2フラグが示す符号化方式を用いて、連続する前記再生区間の接続点においてシームレスに復号を行う復号手段と を有することを特徴とする動画像復号化システム。
Independent claims6
271 paragraphs, as filed
The present invention relates to a moving image coding method and apparatus for switching a variable length coding method to encode a moving image, a moving image decoding method and apparatus, a recording method, and a moving image decoding system.
A DVD-Video disc (hereinafter simply referred to as DVD), which is a conventional technology, will be described.
FIG. 1 is a diagram showing the structure of a DVD. As shown in the lower part of Fig. 1, a logical address space is provided on the DVD disc from read-in to read-out, and the volume information of the file system is recorded from the beginning of the logical address space, followed by the video. Application data such as voice is recorded.
A file system is ISO9660 or UDF (Universal Disc Format), and is a mechanism for expressing data on an optical disc in units called directories or files. Even in the case of a PC (personal computer) that is used on a daily basis, the data recorded on the hard disk in the structure of directories and files is expressed on the computer by passing through a file system called FAT or NTFS, improving usability.
In the case of DVD, both UDF and ISO9660 are used (both are sometimes collectively called "UDF bridge"), and data can be read by either UDF or ISO9660 file system driver. Of course, DVD-RAM / R / RW, which is a rewritable DVD disc, can physically read, write, and delete data via these file systems.
The data recorded on the DVD can be viewed as a directory or file as shown in the upper left of Fig. 1 through the UDF bridge. A directory called "VIDEO_TS" is placed directly under the root directory ("ROOT" in the figure), and the application data of the DVD is recorded here. Application data is recorded as multiple files, and the main files are as follows.
VIDEO_TS.IFO Disk playback control information file VTS_01_0.IFO Video title set # 1 Playback control information file VTS_01_0.VOB Video Title Set # 1 Stream File .....
Two types of extensions are specified. "IFO" is a file in which playback control information is recorded, and "VOB" is a file in which an MPEG stream, which is AV data, is recorded. Playback control information is information for realizing the interactivity (technology that dynamically changes playback according to the user's operation) adopted in DVD, and information attached to titles and AV streams such as metadata. And so on. Further, in DVD, the playback control information is generally called navigation information.
The playback control information file is "VIDEO_TS.IFO" that manages the entire disc, and individual video title sets (multiple titles on DVD, in other words, different movies and different versions of movies can be recorded on one disc. There is "VTS_01_0.IFO" which is the playback control information for each. Here, "01" in the file name body indicates the number of the video title set. For example, in the case of the video title set # 2, it is "VTS_02_0.IFO".
The upper right part of FIG. 1 is a DVD navigation space in the DVD application layer, and is a logical structure space in which the above-mentioned playback control information is expanded. The information in "VIDEO_TS.IFO" is as VMGI (Video Manager Information), and the playback control information that exists for each other video title set is as "VTS_01_0.IFO" in the DVD navigation space as VTSI (Video Title Set Information). Be expanded.
In VTSI, PGCI (Program Chain Information), which is information on a playback sequence called PGC (Program Chain), is described. PGCI is composed of a set of cells and a kind of programming information called a command. Cell itself is a set of some or all sections of a VOB (an abbreviation for Video Object, which refers to an MPEG stream), and cell playback means playing the section specified by the Cell of the VOB. ing.
The commands are processed by the DVD virtual machine and are similar to Java (registered trademark) scripts that are executed on the browser. However, while Java® scripts control windows and browsers (for example, open a new browser window) in addition to logical operations, DVD commands play AV titles in addition to logical operations. It differs in that it only executes control, such as specifying the chapter to play.
The Cell has the VOB start and end addresses (logical recording addresses on the disk) recorded on the disk as its internal information, and the player receives the VOB start and end address information described in the Cell. Use to read and play data.
FIG. 2 is a schematic diagram illustrating navigation information embedded in an AV stream. The interactivity that is a feature of DVD is not realized only by the navigation information recorded in "VIDEO_TS.IFO" and "VTS_01_0.IFO" mentioned above, but some important information is the navigation pack (navigation). It is multiplexed with video and audio data in the VOB using a dedicated carrier called a pack or NV_PCK).
Here, the menu is explained as an example of simple interactivity. Several buttons appear on the menu screen, and each button defines the processing when the button is selected and executed. In addition, one button is selected on the menu (a semi-transparent color is overlaid on the selection button by highlighting to indicate to the user that the button is in the selected state), and the user can use the remote control up / down / left / right. You can use the keys to move the selected button to either the up, down, left, or right button. Use the up / down / left / right keys on the remote control to move the highlight to the button you want to select and execute, and confirm (press the enter key) to execute the program of the corresponding command. Generally, the corresponding title or chapter is played back by a command.
The upper left part of Fig. 2 shows the outline of the control information stored in NV_PCK.
NV_PCK contains highlight color information and individual button information. Color palette information is described in the highlight color information, and the semi-transparent color of the highlight to be overlaid is specified. The button information includes rectangular area information that is the position information of each button, movement information from the button to another button (designation of the destination button corresponding to each user's up / down / left / right key operation), and button command information. (The command to be executed when the button is determined) is described.
The highlights on the menu are created as overlay images, as shown in the upper right center of Figure 2. The overlay image is the rectangular area information of the button information with the color of the color palette information added. This overlay image is combined with the background image shown on the right side of FIG. 2 and displayed on the screen.
As mentioned above, the DVD realizes the menu. Also, the reason why part of the navigation data is embedded in the stream using NV_PCK is that the menu information is dynamically updated in synchronization with the stream (for example, 5 to 10 minutes during movie playback). This is because even in the case of an application where synchronization timing is likely to be a problem, it can be realized without problems (for example, the menu is displayed only in between). Another major reason is that NV_PCK stores information to support special playback, and AV data is smoothly decoded and played even during abnormal playback such as fast-forwarding and rewinding during DVD playback. This is to improve the operability of the user.
Figure 3 is an image of a VOB, which is a DVD stream. As shown in the figure, data such as video, audio, and subtitles (stage A) is packetized and packed (stage B) based on the MPEG system standard (ISO / IEC13818-1), and each is multiplexed into one. It is made into the MPEG program stream of (C stage). Also, as mentioned above, NV_PCK, which includes button commands to realize interactive, is also multiplexed.
A feature of MPEG system multiplexing is that each piece of data to be multiplexed is a bit string based on its decoding order, but between the multiplexed data, that is, between video, audio, and subtitles, it is not necessarily based on the playback order. The bit string is not formed. It has a decoder model of the multiplexed MPEG system stream (commonly called System Target Decoder, or STD (stage D in Figure 3)) that has a decoder buffer corresponding to each elemental stream after demultiplexing and decodes it. It is derived from the fact that data is temporarily accumulated by the timing. For example, the decoder buffer specified by DVD-Video has a different size for each elementary stream, and has 232KB for video, 4KB for audio, and 52KB for subtitles. ..
That is, the subtitle data multiplexed along with the video data is not necessarily decoded or reproduced at the same timing.
On the other hand, there is BD (Blu-ray Disc) as a next-generation DVD standard.
DVDs have been aimed at package distribution (DVD-Video standard) and analog broadcast recording (DVD Video Recording standard) for standard definition image quality, but BDs have high precision image quality (High Definition image quality). It is possible to record the digital broadcast of (Blu-ray Disc Rewritable standard, hereinafter BD-RE) as it is.
However, the BD-RE standard is special because it is widely intended for recording digital broadcasts. Regeneration support information is not optimized. In the future, considering that high-precision video will be packaged and distributed at a higher rate than digital broadcasting (BD-ROM standard), a mechanism that does not stress the user even during abnormal playback will be required.
In addition, MPEG-4 AVC (Advanced Video Coding) is adopted as one of the video coding methods in BD. MPEG-4 AVC is the next generation of high compression ratio jointly formulated by ISO / IEC (International Electrotechnical Commission) JTC1 / SC29 / WG11 and ITU-T (International Telecommunication Union Telecommunication Standardization Division). It is a coding method.
Generally, in video coding, the amount of information is compressed by reducing redundancy in the temporal and spatial directions. Therefore, in inter-screen prediction coding for the purpose of reducing temporal redundancy, motion is detected and a prediction image is created in block units by referring to the front or rear picture, and the obtained prediction image and coding are performed. Encoding is performed on the difference value from the target picture. Here, a picture is a term representing one screen, and means a frame in a progressive image and a frame or a field in an interlaced image. Here, the interlaced image is an image in which one frame is composed of two fields having different times. In the coding and decoding processing of an interlaced image, one frame can be processed as a frame, it can be processed as two fields, or each block in the frame can be processed as a frame structure or a field structure. it can.
An image that does not have a reference image and performs in-screen predictive coding is called an I picture. Further, a picture that refers to only one picture and performs inter-screen prediction coding is called a P picture. Further, a picture capable of performing inter-screen prediction coding by referring to two pictures at the same time is called a B picture. As for the B picture, it is possible to refer to two pictures as any combination of the display time from the front or the back. The reference image (reference picture) can be specified for each block, which is the basic unit of encoding and decoding, but the reference picture described first in the encoded bitstream is the first reference picture. , The one described later is distinguished as the second reference picture. However, as a condition for encoding and decoding these pictures, the referenced picture must already be encoded and decoded.
The residual signal obtained by subtracting the prediction signal obtained from the in-screen prediction or the inter-screen prediction from the encoded image is frequency-converted and quantized, then variable-length coded and output as a coded stream. .. For MPEG-4 AVC, CAVLC (Context-Adaptive Variable Length Coding) and CABAC (Context-Adaptive Binary Arithmetic Coding) are used as variable length coding methods. There are two types, which can be switched on a picture-by-picture basis. Here, the context-adaptive type is a method of adaptively selecting an efficient coding method according to the surrounding situation.
FIG. 4 is an example showing a variable length coding method applied to a picture constituting a randomly accessible unit in an MPEG-4 AVC stream. Here, in MPEG-4 AVC, there is no concept equivalent to GOP (Group of Pictures) of MPEG-2 video, but if data is divided into special picture units that can be decoded without depending on other pictures, it becomes GOP. Since a corresponding randomly accessible unit can be constructed, this is called a random access unit (RAU). As shown in FIG. 4, whether CABAC or CAVLC is applied as the variable length coding method is switched for each picture.
Next, since the processing at the time of variable length decoding differs between CABAC and CAVLC, each variable length decoding processing will be described with reference to FIGS. 5A to 5C. Figure 5A shows CABAD (Context-Adaptive Binary Arithmetic Decoding), which is the decoding process of variable-length coded data by CABAC, and CABAD (Context-Adaptive Binary Arithmetic Decoding), which is the decoding process of variable-length coded data by CAVLC. A block diagram of an image decoding device that performs a certain CAVLD (Context-Adaptive Variable Length Decoding) is shown.
The image decoding process accompanied by CABAD is performed as follows. First, the coded data Vin to which CABAC is applied is input to the stream buffer 5001. Subsequently, the arithmetic decoding unit 5002 reads the encoded data Vr from the stream buffer, performs arithmetic decoding, and inputs the binary data Bin1 into the binary data buffer 5003. The binary data decoding processing unit 5004 acquires binary data Bin2 from the binary data buffer 5003, decodes the binary data, and inputs the decoded binary data Din1 to the pixel restoration unit 5005. The pixel restoration unit 5005 performs inverse quantization, inverse conversion, motion compensation, and the like on the binary decoding data Din1, restores the pixels, and outputs the decoding data Vout.
FIG. 5B is a flowchart showing the operation from the start of decoding the coded data to which CABAC is applied to the execution of the pixel restoration process. First, in step 5001, the coded data Vin to which CABAC is applied is arithmetically decoded to generate binary data. Next, in step 5002, it is determined whether or not the binary data for a predetermined data unit such as one or more pictures are aligned, and if they are aligned, the process proceeds to step S5003, and if they are not aligned, the process of step S5002 is performed. repeat. Here, the reason why the binary data is buffered is that in CABAC, the code amount of the binary data per picture or macroblock may be remarkably large, and the processing load of arithmetic decoding may be remarkably increased accordingly. This is because it is necessary to perform a certain amount of arithmetic decoding processing in advance in order to realize uninterrupted reproduction even in the worst case. In step S5003, the binary data is decoded, and in step S5004, the pixel restoration process is performed. As described above, in CABAD, since the pixel restoration process cannot be started until the binary data for a predetermined data unit is prepared in step S5001 and step S5002, a delay occurs at the start of decoding.
The image decoding process involving CAVLD is performed as follows. First, the coded data Vin to which CAVLC is applied is input to the stream buffer 5001. Subsequently, the CAVLD unit 5006 performs a variable length decoding process, and inputs the VLD decoding data Din2 to the pixel restoration unit 5005. The pixel restoration unit 5005 performs inverse quantization, inverse conversion, motion compensation, and the like, restores pixels, and outputs decoded data Vout. FIG. 5C is a flowchart showing the operation from the start of decoding the coded data to which CAVLC is applied to the execution of the pixel restoration process. First, CAVLD is performed in step S5005, and then pixel restoration processing is performed in step S5004. In this way, in CAVLD, unlike CABAD, it is not necessary to wait until the data for a predetermined data unit is prepared before starting the pixel restoration processing, and variable length decoding such as the binary data buffer 5003 is performed. It is not necessary to have an intermediate buffer in processing.
FIG. 6 is a flowchart showing the operation of a conventional decoding device that decodes a stream in which the variable length coding method is switched in the middle of the stream, as in the example of FIG. First, in step S5101, information indicating the variable-length coding method applied to the picture is acquired, and the process proceeds to step S5102. In step S5102, it is determined whether or not the variable length coding method is switched between the immediately preceding picture and the current picture in the decoding order. Since the buffer management method in the variable-length decoding process differs between CABAD and CAVLD, when the variable-length coding method is switched, the process proceeds to step S5103 to switch the buffer management to perform variable-length coding. If the method has not been switched, the process proceeds to step S5104. In step S5104, it is determined whether or not the variable length coding method is CAVLC. If it is CAVLC, the process proceeds to step S5105 to perform CAVLD processing, and if it is CABAC, the process proceeds to step S5106. In step S5106, it is determined whether or not the variable length coding method has been switched between the immediately preceding picture and the current picture in the decoding order, and when the switching is made, the process proceeds to step S5107, as shown in steps S5001 and S5002 of FIG. , Arithmetic decoding is performed until the binary data for a predetermined data unit is prepared, and then the binary data is decoded. When it is determined in step S5106 that the variable length coding method has not been switched, the process proceeds to step S5108, and normal CABAD processing is performed. Here, the normal CABAD process is a process that does not buffer the binary data required when switching from CAVLC to CABAC or starting decoding of the stream to which CABAC is applied. Finally, the pixel restoration process is performed in step S5109.<patcit num="1"><text>Japanese Unexamined Patent Publication No. 2000-228656</text></patcit><nplcit num="1"><text>Proposed SMPTE Standard for Television: VC-1 Compressed Video Bitstream Format and Decoding Process, Final Committee Draft 1 Revision 6, 2005.7.13</text></nplcit>
<p> In this way, when playing back a conventional information recording medium in which a conventional MPEG-4 AVC stream is multiplexed, the variable length coding method is switched for each picture, so when switching from CAVLC to CABAC, , There was the first problem that a delay occurred due to the buffering of binary data. In particular, if the frequency of switching is high, delays may accumulate and playback may be interrupted. In addition, since the buffer management method differs between CABAC and CAVLC, it is necessary to switch the buffer management method each time the variable-length coding method is switched, which increases the processing load at the time of decoding. there were.</p><p> An object of the present invention is to provide a moving image coding method and device, a moving image decoding method and device, a recording method, and a moving image decoding system that do not cause playback interruption without increasing the processing load at the time of decoding. To do.</p>
<p> The image coding method that achieves the above object is a moving image coding method that encodes a moving image by switching a variable length coding method, determines a continuous section to be continuously reproduced, and determines the continuous section. The management information including the first flag information indicating that the variable-length coding method in the continuous section is fixed by generating a stream by encoding the moving image without switching the variable-length coding method in To generate.</p><p> According to this configuration, since the variable-length coding method is fixed in the section to be continuously reproduced, it is possible to eliminate the delay at the time of decoding due to the switching of the variable-length coding method, and at the time of decoding. Reproduction quality can be improved. In addition, the processing load associated with switching the buffer management method can be reduced.</p><p> Here, the continuous section corresponds to one stream and may be a section specified by the packet identifier of the transport stream.</p><p> Here, the streams constituting the continuous section may be specified by the packet identifier of the transport stream.</p><p> Here, the continuous section may be a section including a stream to be seamlessly connected.</p><p> Here, the continuous section may be a section including a stream corresponding to each angle constituting the seamless multi-angle.</p><p> Here, the continuous section may be a section including a stream corresponding to each angle constituting the non-seamless multi-angle.</p><p> Here, the moving image coding method may further insert the second flag information indicating the variable length coding method into the stream for each predetermined unit of the coded moving image.</p><p> Here, the management information includes a playlist indicating the playback order of one or more playback sections, with all or part of the stream as a playback section, and the first flag information is for each playback section shown in the playlist. The predetermined unit may be a unit to which a picture parameter set in an MPEG-4 AVC-encoded stream is added.</p><p> According to this configuration, it is possible to prevent playback interruption without increasing the processing load at the time of decoding by using coding such as MPEG-4 AVC that can switch the variable length coding method in the stream.</p><p> Here, the management information includes a playlist indicating the playback order of one or more playback sections, with all or a part of the stream as a playback section, and the first flag information is for each playback section shown in the playlist. The second flag information is generated in, and indicates a variable length coding method for information in macro block units, the predetermined unit is a picture unit in a stream, and the information in the macro block unit is variable length coding. When the method is the first method in bit plane coding, it is added in units of macro blocks in the stream, and when the variable length coding method is the second method in bit plane coding, the picture in the stream. It is added in units, and in the second method, the second flag information corresponding to all the macro blocks in the picture may be inserted at the beginning of the picture in the stream.</p><p> According to this configuration, it is possible to prevent playback interruption without increasing the processing load at the time of decoding by using coding such as VC-1 in which the variable length coding method can be switched in the stream.</p><p> Here, the first flag information may indicate that the variable length coding method in the continuous section is fixed and that the stream is a seamless target.</p><p> According to this configuration, the amount of management information data can be reduced.</p><p> Further, since the image coding apparatus of the present invention includes the same means as the above-mentioned image coding method, the details will be omitted.</p><p> Further, in the present invention, the data having a computer-readable structure includes management information and a stream showing a coded moving image, and the management information is a variable length code in a continuous section to be continuously reproduced. The stream includes first flag information indicating that the coding method is fixed, and the stream includes second flag information indicating the variable length coding method for each predetermined unit of the encoded moving image.</p><p> According to this configuration, since the variable-length coding method is fixed in the section to be continuously reproduced, it is possible to eliminate the delay at the time of decoding due to the switching of the variable-length coding method, and at the time of decoding. Reproduction quality can be improved. In addition, the processing load associated with switching the buffer management method can be reduced.</p><p> Here, the management information includes a playlist indicating the playback order of one or more playback sections, with all or a part of the stream as a playback section, and the first flag information is provided in each stream shown in the playlist. Correspondingly, the predetermined unit may be a unit to which a picture parameter set in the stream is added.</p><p> Here, the management information includes a playlist indicating the playback order of one or more playback sections, with all or a part of the stream as a playback section, and the first flag information is provided in each stream shown in the playlist. Correspondingly, the second flag information indicates a variable length coding method for each macro block, and the predetermined unit is a macro block unit in the stream when the variable length coding method is the first method. , When the variable-length coding method is the second method, it is a picture unit in the stream, and when the variable-length coding method is the second method, it corresponds to all the macro blocks in the picture. 2 The flag information may be inserted at the beginning of the picture in the stream.</p>
<p> As described above, according to the moving image coding method of the present invention, the variable length coding method applied to the coded data of the moving image is fixed in the continuous section, so that the variable length coding method can be used. It is possible to eliminate the delay during decoding due to switching, and it is possible to reduce the processing load associated with switching the buffer management method. For example, it is possible to improve the playback quality of packaged media in which streams such as MPEG-4 AVC and VC-1 that can switch the variable length coding method within the stream are multiplexed, and reduce the processing load of the playback device. It can be done and its practical value is high.</p>
Hereinafter, embodiments of the present invention will be described with reference to the drawings.
(Embodiment 1) First, the first embodiment of the present invention will be described.
In the present embodiment, when decoding the coded data of a moving image in a package medium such as BD-ROM, the delay of the decoding operation due to the switching of the variable length coding method and the buffer management required at the same time. An information recording medium capable of suppressing an increase in processing load due to method switching, and a playback device thereof will be described. Here, the video coding method is MPEG-4 AVC, but another coding method capable of switching the variable length coding method in the middle of the stream may be used.
In the MPEG-4 AVC stream stored in the information recording medium of the present embodiment, the unit in which the variable length coding method can be switched is restricted, and the switching unit is restricted or restricted. Information indicating the switching unit is stored in the management information.
Figure 7 shows MPEG-4 An example of restrictions on the switching unit of the variable-length coding method in the AVC stream is shown. In package media such as BD-ROM, a play list or the like indicates a unit for continuously reproducing encoded data of a moving image (hereinafter referred to as a continuous reproduction unit), so that variable length coding is performed in the continuous reproduction unit. If the method is fixed, the delay of the decoding operation due to the switching of the variable-length coding method and the switching of the buffer management method do not occur in the continuously reproduced section. Therefore, in the present embodiment, the variable length coding method is fixed in the continuous reproduction unit. Figures 7 (a) and 7 (b) show examples in which the variable-length coding method is limited to CAVLC only and CABAC only in the continuous playback unit, respectively. Furthermore, there are two types of connection conditions for clips that are played continuously: seamless connection and non-seamless connection. The connection here includes the case of connecting a plurality of sections in the same clip. In a non-seamless connection, for example, a gap may occur in the decoding operation as in the case of connecting to an open GOP. Therefore, it is decided to allow switching of the variable length coding method, and the continuous playback unit is seamlessly connected. In, the variable length coding method may be fixed.
The variable length coding method may be fixed in a unit different from the continuous playback unit such as a clip or a random access unit (RAU). 7 (c) and 7 (d) show an example of fixing in clip units, and FIG. 7 (e) shows an example of fixing in random access units.
Next, in the management information, flag information indicating that the switching unit of the variable length coding method is restricted is stored in the MPEG-4 AVC stream. Here, the identification information of the coding method is used as a flag. FIG. 8 shows an example of storing flags in a BD-ROM. In BD-ROM, the coding method of each clip referenced from the playlist is stored in an area called StreamCodingInfo in the management information, so it is shown here that the coding method is MPEG-4 AVC. In this case, it is assumed that the variable length coding method is fixed in the continuous playback unit. It should be noted that whether the variable length coding method is CABAC or CAVLC may be indicated separately.
A flag indicating that the switching unit of the variable-length coding method is restricted may be separately specified and stored, or information indicating the switching unit may be stored. In addition, this information may be stored in the MPEG-4 AVC stream. For example, information indicating that the variable-length coding method is fixed for a picture in a random access unit can be obtained as SEI (Supplemental Enhancement Information) in the first picture of the random access unit, or NAL (Network Abstraction Layer) having an Unspecified type. ) Can be stored in the unit.
In MPEG-4 AVC, the entropy_coding_mode_flag in the PPS (Picture Parameter Set) showing the initialization information for each picture indicates whether the variable length coding method is CAVLC or CABAC. Therefore, when the variable-length coding method is fixed in a certain section, the field value of entropy_coding_mode_flag is fixed in all PPSs referenced by the picture in the section. MPEG-4 In AVC, it is permissible to store PPS that is not referenced from pictures existing in a predetermined section in the decoding order in the predetermined section, but the field value of entropy_coding_mode_flag in PPS that is not referenced from pictures in the section is particularly limited. No need. For example, all PPS referenced by a picture in a random access unit RAU is guaranteed to be in the random access unit RAU, but there is a PPS in the random access unit that is not referenced by a picture in the random access unit RAU. You may. At this time, since the unreferenced PPS does not affect the decoding, it is not necessary to limit the field value of entropy_coding_mode_flag. However, since it is easier to handle by uniquely defining the field value of entropy_coding_mode_flag in the PPS included in the predetermined interval, the field value may be fixed including the unreferenced PPS.
FIG. 9 is a flowchart showing a decoding operation of a continuous reproduction unit in the information recording medium of the present embodiment. Since the variable-length coding method is fixed in the continuous playback unit, unlike the conventional decoding operation of FIG. 6, it is not necessary to buffer the binary data during decoding and switch the buffer management method. Since the operation of each step is the same as the step with the same reference numeral in FIG. 6, the description thereof will be omitted.
Furthermore, as a new coding method, the VC-1 (Non-Patent Document 1) standard is currently being formulated by SMPTE (The Society of Motion Picture and Television Engineers). In VC-1, various flags indicating the coding method of a macroblock (a unit having a size of 16 × 16 pixels) are defined. Examples of the flag include whether or not it is a skip macroblock, whether or not it is a field mode / frame mode, and whether or not it is a direct mode macroblock.
One of the extended coding tools is bit plane coding. Bitplane coding is used to indicate a flag indicating how to code the macroblock described above. In bit-plane coding, these flags can be grouped together for one picture and indicated in the picture header. Generally, adjacent macroblocks have a high correlation, so the flag also has a high correlation. Therefore, by encoding the flags of a plurality of adjacent macroblocks together, the amount of code representing the flags can be reduced.
In bit plane coding, seven kinds of coding methods are specified. One of them is a method of encoding each flag with a macroblock header, which is called RAW MODE and is similar to the MPEG-2 video method and MPEG-4 visual method. The remaining six methods are methods that encode the flags for one picture together, and different methods are defined depending on how the flags of adjacent macroblocks are encoded together. These six methods include, for example, encoding the flags of two adjacent macroblocks on the left and right together, and if all the flags of a row of macroblocks arranged in the horizontal direction are "0", set it to 1 bit. There is a method of encoding each flag as it is if it is represented by "0" and there is at least one "1" in the flags of a macroblock in a row.
Which of these seven methods is used for bit plane coding can be changed independently for each flag and for each picture.
Here, in bit plane coding, mode 1 is when only the method of encoding each flag in the macroblock header is used, and mode 2 is when only the method of encoding the flags for one picture is used together. Since the operation at the time of decoding differs between mode 1 and mode 2, the processing load increases at the mode switching portion, and a delay may occur. Therefore, in the same way as the switching unit of the variable length coding is restricted, the switching unit of the mode 1 and the mode 2 may be restricted for the bit plane coding. For example, the mode is fixed in the continuous playback unit or the continuous playback unit that is seamlessly connected. Further, the management information may include flag information indicating that the bit plane coding mode is fixed in a predetermined unit. For example, the coding method indicated by StreamCodingInfo is used as flag information, and when the coding method is indicated to be VC -1, the bit plane coding mode can be fixed in a predetermined unit.
Furthermore, if mode 3 is used when both the method of encoding each flag in the macroblock header and the method of encoding the flags for one picture at once can be used, the environment in which VC-1 is used. There are cases where mode 1 and mode 3 are used properly depending on the situation. For example, mode 1 can be used for terminals with low processing power, and mode 3 can be used for terminals with high processing power. In such a case, it is effective to fix to either mode 1 or mode 3 in a predetermined reproduction unit. Further, flag information indicating that the mode is fixed to either mode 1 or mode 3 or information indicating which mode is fixed can be stored in the management information or the coded stream. In addition, mode 2 and mode 3 may be used properly.
FIG. 10 is a block diagram showing a configuration of a multiplexing device 5100 that realizes the multiplexing method of the present embodiment. The multiplexing device 5100 includes a switching unit determination unit 5101, a switching information generation unit 5102, a coding unit 5103, a system multiplexing unit 5104, a management information creation unit 5105, and a coupling unit 5106. The operation of each part will be described below.
The switching unit determination unit 5101 determines a unit for which the variable length coding method can be switched, and inputs the determined switching unit Unit to the switching information generation unit 5102 and the coding unit 5103. The switching unit is predetermined, but may be set from the outside. The switching information generation unit 5102 generates switching information SwInf indicating a unit in which variable length coding can be switched based on the switching unit Unit, and inputs the switching information SwInf to the management information creation unit 5105. The coding unit 5103 encodes the data of each clip so as to satisfy the constraint of the switching unit Unit, and inputs the coded data Cdata1 to the system multiplexing unit 5104. The system multiplexing unit 5104 system-multiplexes the encoded data Cdata1, inputs the stream information StrInf1 to the management information creation unit 5105, and inputs the multiplexing data Mdata1 to the coupling unit 5106. In BD-ROM, as a system multiplexing method, a method called a source packet, in which a 4-byte header is added to an MPEG-2 transport stream, is used. In addition, the stream information StrInf1 includes information for generating management information about the multiplexed data Mdata1 such as a time map. The management information creation unit 5105 generates the management information CtrInf1 including the time map generated based on the stream information StrInf1 and the switching information SwInf, and inputs the management information CtrInf1 to the connection unit 5106. The combining unit 5106 combines the management information CtrlInf1 and the multiplexed data Mdata1 and outputs the recorded data Dout1.
When creating data with an authoring tool, etc., the coded data may be generated and the system multiplexing or management information may be created by separate devices. Even in such a case, the operation of each device is performed. May be the same as each part in the multiplexing device 5100.
FIG. 11 is a flowchart showing the operation of the multiplexing method for creating the multiplexed data stored in the information recording medium according to the present embodiment. The multiplexing method of the present embodiment includes a step of determining a unit for which the variable-length coding method can be switched (step S5201), a step of encoding a clip based on the determined unit (step S5202), and a variable-length code. It differs from the conventional multiplexing method in that it includes a step (step S5204) of generating flag information indicating a switching unit of conversion.
First, in step S5201, the unit for which the variable length coding method can be switched is determined. That is, it is determined whether the switching is possible in a continuous playback unit, a clip, or a random access unit. Subsequently, in step S5202, the MPEG-4 AVC clip data is encoded based on the switching unit determined in step S5201. In step S5203, it is determined whether or not the coding of the final clip is completed. If it is determined that the coding of the final clip is completed, the process proceeds to step S5204. If it is determined that the coding of the final clip is not completed, the process returns to step S5202 to encode the clip. repeat. In step S5204, flag information indicating the switching unit of variable length coding is generated, and the process proceeds to step S5205. In step S5205, management information including the flag information generated in step S5204 is created, and the management information and the clip data are multiplexed and output.
FIG. 12 is a flowchart showing a specific example of the step (S5201) of determining the unit in which the variable length code method in FIG. 11 can be switched. In the figure, the minimum unit in which the variable length code method can be switched is the clip shown in FIGS. 7 (c) and 7 (d). Here, the clip is stored as a file of AV data on the recording medium, and refers to, for example, one stream containing one stream of MPEG-4 AVC and one file storing one stream of VC-1. Also, in the transport stream, the clip refers to the stream identified by the identifier of the TS packet.
In FIG. 12, the switching unit determination unit 5101 determines whether or not the picture to be encoded is the start picture of the clip (S5201a), and if it is not the start picture, that is, if it is a picture in the middle of the clip, the corresponding picture is concerned. In clip coding, it is determined that the variable length code method cannot be switched (S5201f).
If it is a start picture, the switching unit determination unit 5101 determines whether or not the clip of the start picture is seamlessly connected to the immediately preceding clip that has been encoded (S5201b), and if it is seamlessly connected, the start picture is started. It is determined that the variable length code method cannot be switched in the encoding of the clip of the picture (S5201f).
If the connection is not seamless, the switching unit determination unit 5101 determines whether or not the clip of the start picture is a clip corresponding to the angles constituting the multi-angle (S5201c), and if the clip corresponds to the angle. Determines that the variable-length code method cannot be switched between the angles that make up the multi-angle when encoding the clip of the start picture (S5201f). Here, in the seamless multi-angle that can be seamlessly connected to each angle, the variable length coding method of each angle is determined to be the same method as the clip immediately before the multi-angle section. On the other hand, in the non-seamless multi-angle where seamless connection to each angle is not guaranteed, if the variable length coding method is the same for each angle, even if the method is different from the clip immediately before the multi-angle section. Good.
Further, when the picture to be encoded is the start picture of the clip and does not correspond to any of S5101b to S5101c (when no), the switching unit determination unit 5101 encodes the clip of the start picture with a variable length. Determines that the method is switchable for the immediately preceding clip that has been encoded (S5201e).
As described above, in the flowchart of FIG. 12, the clips determined not to be switched by the switching unit determination unit 5101 are (a) clips specified by the packet identifier of the transport stream, and (b) a plurality of clips to be seamlessly connected. Clips, (c) Multiple clips corresponding to each angle constituting the multi-angle are determined. The determinations of S5201a to S5201c may be performed in any order. Also in the case of multi-angle, the variable-length coding method may not be switchable only in seamless multi-angle. Further, the clip may be identified by information different from the packet identifier such as a file name. Further, in FIG. 12, the case where the minimum unit for switching the variable length code method is the clip shown in FIGS. 7 (c) and 7 (d) has been described, but the RAU as shown in FIG. 7 (e) is used as the minimum unit. May be good. In that case, the process of replacing "clip" in the figure with "RAU" may be performed.
FIG. 13 is a flowchart showing a specific example of the clip coding step (S5202) in FIG. FIG. 13 shows a case where MPEG-4 AVC is encoded. In the figure, the coding unit 5103 determines whether or not the variable length code method of the clip can be switched prior to the start of coding of the clip (S5202a). This determination follows the decision in FIG. The coding unit 5103 arbitrarily determines the variable length coding method of the clip when it is determined that the clip can be switched (S5202b), and when it is determined that the clip cannot be switched, the variable length coding method of the clip is determined. Are determined to be in the same manner as other clips that make up the same multi-angle, just before they are seamlessly connected to each other (S5202c). Further, the coding unit 5103 sets a flag indicating the determined variable-length coding method in the picture parameter set PPS (S5202d), and encodes the clip according to the determined variable-length coding method (S5202e). ). This flag is called entropy_coding_mode_flag in MPEG4-AVC.
In this way, the coding unit 5103 generates the coded data Cdata1 by coding the moving image without switching the variable length coding method for the clips in the continuous section determined to be non-switchable. ..
FIG. 14 is a flowchart showing a specific example of the flag information generation step (S5204) and the management information generation step (S5205) in FIG.
In the figure, the switching information generation unit 5102 determines whether or not the clip encoded by the coding unit 5103 is a clip that is determined to be switchable in the variable-length coding method (S5204a), and is said to be switchable. If the clip is determined, flag information indicating that the variable-length coding method is fixed is generated, and the flag information is stored in the work area of the memory in association with the clip (S5204b) and switched. If the clip is not determined to be possible, flag information indicating that the variable-length coding method is not fixed is generated, and the flag information is stored in the work area of the memory in association with the clip (S5204b). Further, the switching information generation unit 5102 determines whether or not the clip is the last clip encoded by the coding unit 5103 (S5204d), and if it is not the last clip, the above S5204a to S5204c are repeated. If it is the last clip, the flag information accumulated in the work area of the memory is output as switching information SwInf to the management information creation unit 5105.
Furthermore, the management information creation unit 5105 generates management information including the playlist (S5205a), refers to the switching information SwInf, and determines that the variable length coding method is fixed for the play items included in the playlist. Add the indicated flag information (S5205b). The flag information may indicate whether or not the playback section referred to by the immediately preceding play item and the variable length coding method are the same. Here, the playlist indicates the playback order of one or more play items. A play item is information that points to a clip to be played, and points to all or part of one clip as a play section. Further, the above flag information may be used in combination with other parameters added to the play item. In that case, for example, a parameter that means that the clips are seamlessly connected (eg "" connection_condition = 5 ) can also be used as the above flag information. Because, in FIG. 12, the continuous interval (the interval in which the variable length coding method is fixed) determined to be non-switchable is (a) transformer. Clips specified by the packet identifier of the port stream, (b) multiple clips to be seamlessly connected, and (c) multiple clips corresponding to each angle constituting the multi-angle, of which (c) is seamless. This is because connection is a prerequisite. Also, since it can be indicated by a flag called "is_multi_angle" whether it is a multi-angle section, this flag is also used as a flag indicating that the variable length coding method is fixed. This may reduce the amount of management information data.
FIG. 15 is a flowchart showing another specific example of the clip coding step (S5202) in FIG. FIG. 15 shows a case where VC-1 is coded. In the figure, the coding unit 5103 determines whether or not the variable-length code method of the clip can be switched between the low mode and the other modes prior to the start of coding of the clip (S5202a). .. This determination follows the decision in FIG. The coding unit 5103 arbitrarily determines the bit plane coding method of the clip when it is determined that the clip can be switched (S5202f), and when it is determined that the clip cannot be switched, the bit plane coding of the clip is performed. Determine the method to be the same as the previous clip (S5202g). The coding unit 5103 determines whether the determined bit plane coding method is the raw mode (RAW MODE) or other method (S5202h). The coding unit 5103 adds information indicating the mode to each picture, and when it is determined to be in RAW MODE, encodes predetermined information for each macroblock MB in each macroblock (S5202i). Low mode (RAW If it is determined that it is not MODE), the predetermined information for each macroblock MB is set at the beginning of the picture, and the clip is encoded (S5202j). Information indicating the mode is indicated by a field called IMODE in VC-1.
In this way, the coding unit 5103 generates the coded data Cdata1 by coding the moving image without switching the bit plane coding method for the clips in the continuous section determined to be non-switchable. ..
The playlist is not limited to use on an optical disk, and when receiving a stream via a network, the playlist is first received and analyzed, the stream to be received is determined, and then the actual stream is received. It can also be used to start. Also, when a stream is packetized into RTP (Real-time Transport Protocol) packets or TS packets and then transmitted over an IP (Internet Protocol) network, the playback control information is, for example, SDP (Session Description Protocol). , It may indicate whether the variable length coding method is fixed in the reproduction section.
The data structure of the BD-ROM disc storing the data generated by the image coding method according to the present embodiment and the configuration of the player for playing the disc are shown below.
(Logical data structure on disk) FIG. 16 is a diagram showing a configuration of a BD-ROM, particularly a BD disc (104) which is a disc medium and a configuration of data (101, 102, 103) recorded on the disc. The data recorded on the BD disc (104) is AV data (103), BD management information (102) such as management information related to AV data and AV playback sequence, and a BD playback program (101) that realizes interactive. .. In the present embodiment, for convenience of explanation, the BD disc will be described mainly for an AV application for playing back AV contents of a movie, but of course, it is the same even if it is used for other purposes.
FIG. 17 is a diagram showing a directory / file structure of logical data recorded on the above-mentioned BD disc. A BD disc, like other optical discs such as DVDs and CDs, has a recording area spirally from the inner circumference to the outer circumference, and logical data is stored between the read-in on the inner circumference and the read-out on the outer circumference. It has a recordable logical address space. Also, inside the read-in, there is a special area called BCA (Burst Cutting Area) that can only be read by a drive. Since this area cannot be read from the application, it may be used for copyright protection technology, for example.
Application data such as video data is recorded in the logical address space, starting with file system information (volume). As explained in the conventional technology, the file system is UDF, ISO9660, etc., and it is possible to read the logical data recorded in the same way as a normal PC using a directory and a file structure. ..
In the case of this embodiment, the BDVIDEO directory is placed directly under the root directory (ROOT) as the directory and file structure on the BD disc. This directory is a directory in which data such as AV contents and management information handled by BD (101, 102, 103 described in FIG. 16) are stored.
The following 7 types of files are recorded under the BDVIDEO directory.
BD.INFO (fixed file name) It is one of the "BD management information" and is a file that records information about the entire BD disc. The BD player first reads this file.
BD.PROG (fixed file name) It is one of the "BD playback programs" and is a file that records playback control information related to the entire BD disc.
XXX.PL ("XXX" is variable, extension "PL" is fixed) It is one of the "BD management information" and is a file that records playlist information that is a scenario (playback sequence). I have one file for each playlist.
XXX.PROG ("XXX" is variable, extension "PL" is fixed) It is one of the "BD playback programs" and is a file that records the playback control information for each playlist described above. The correspondence with the playlist is identified by the file body name (matching "XXX").
YYY.VOB ("YYY" is variable, extension "VOB" is fixed) It is one of the "AV data" and is a file that records VOB (same as VOB explained in the conventional example). I have one file for each VOB.
YYY.VOBI ("YYY" is variable, extension "VOBI" is fixed) It is one of the "BD management information" and is a file that records stream management information related to VOB, which is AV data. Correspondence with VOB is identified by the file body name (matching "YYY").
ZZZ.PNG ("ZZZ" is variable, extension "PNG" is fixed) It is one of the "AV data" and is a file that records image data PNG (an image format standardized by W3C and read as "ping") for composing subtitles and menus. It has one file for each PNG image.
(Player configuration) Next, the configuration of the player that reproduces the BD disc described above will be described with reference to FIGS. 18 and 19.
FIG. 18 is a block diagram showing a rough functional configuration of the player.
The data on the BD disc (201) is read through the optical pickup (202). The read data is transferred to a dedicated memory according to each type of data. The BD playback program (the contents of the "BD.PROG" or "XXX.PROG" file) is stored in the program recording memory (203) with BD management information ("BD.INFO", "XXX.PL" or "YYY.VOBI"). Is transferred to the management information recording memory (204), and the AV data (YYY.VOB or ZZZ.PNG) is transferred to the AV recording memory (205), respectively.
The BD playback program recorded in the program recording memory (203) is recorded by the program processing unit (206), and the BD management information recorded in the management information recording memory (204) is recorded by the management information processing unit (207). The AV data recorded in the memory (205) is processed by the presentation processing unit (208), respectively.
The program processing unit (206) receives event information such as playlist information to be played back and program execution timing from the management information processing unit (207) and processes the program. In addition, the program can dynamically change the playlist to be played, and in this case, it is realized by sending a playlist playback command to the management information processing unit (207). The program processing unit (206) receives an event from the user, that is, a request from the remote control key, and executes a program corresponding to the user event, if any.
The management information processing unit (207) receives instructions from the program processing unit (206), analyzes the corresponding playlist and the VOB management information corresponding to the playlist, and sends the presentation processing unit (208) the target AV data. Instruct to play. In addition, the management information processing unit (207) receives reference time information from the presentation processing unit (208), instructs the presentation processing unit (208) to stop AV data playback based on the time information, and also provides a program processing unit. Generates an event indicating the program execution timing for (206).
The presentation processing unit (208) has decoders corresponding to video, audio, and subtitles / images (still images), and decodes and outputs AV data according to instructions from the management information processing unit (207). In the case of video data and subtitles / images, after decoding, they are drawn on their respective dedicated planes, video planes (210) and image planes (209), and the video composition processing is performed by the composition processing unit (211) to display TV, etc. Output to the device.
As described above, as shown in FIG. 18, the BD player has a device configuration based on the data configuration recorded on the BD disc shown in FIG.
FIG. 19 is a block diagram detailing the above-mentioned player configuration. In FIG. 19, the AV recording memory (205) is the image memory (308) and the track buffer (309), the program processing unit (206) is the program processor (302) and the UOP manager (303), and the management information processing unit (207). ) Is the scenario processor (305) and the presentation controller (306), and the presentation processor (208) is the clock (307), demultiplexer (310), image processor (311), video processor (312) and sound processor (313). It corresponds / deploys to each.
The VOB data (MPEG stream) read from the BD disc (201) is recorded in the track buffer (309), and the image data (PNG) is recorded in the image memory (308). The demultiplexer (310) extracts the VOB data recorded in the track buffer (309) based on the time of the clock (307), and sends the video data to the video processor (312) and the audio data to the sound processor (313), respectively. The video processor (312) and sound processor (313) each consist of a decoder buffer and a decoder, as defined by the MPEG system standard. That is, the video and audio data sent from the demultiplexer (310) are temporarily recorded in the respective decoder buffers, and are decoded by the individual decoders according to the clock (307).
PNG recorded in the image memory (308) has the following two processing methods.
If the image data is for subtitles, the presentation controller (306) indicates the decoding timing. When the subtitle display time (start and end) is reached, the subtitles are displayed to the presentation controller (306) so that the scenario processor (305) can once receive the time information from the clock (307) and display the subtitles appropriately. Give instructions to hide. Upon receiving the decoding / display instruction from the presentation controller (306), the image processor (311) extracts the corresponding PNG data from the image memory (308), decodes it, and draws it on the image plane (314).
Next, when the image data is for a menu, the decoding timing is instructed by the program processor (302). When the program processor (302) instructs the decoding of the image depends on the BD program processed by the program processor (302) and is not unconditionally determined.
As described in FIG. 18, the image data and the video data are output to the image plane (314) and the video plane (315) after decoding, respectively, and are output after being synthesized by the compositing processing unit (316).
The management information (scenario, AV management information) read from the BD disc (201) is stored in the management information recording memory (304), but the scenario information ("BD.INFO" and "XXX.PL") is stored. It is read into the scenario processor (305) and processed. Further, the AV management information (YYY.VOBI) is read and processed by the presentation controller (306).
The scenario processor (305) analyzes the playlist information, indicates the VOB referenced by the playlist and its playback position to the presentation controller (306), and the presentation controller (306) manages the target VOB. Analyze ("YYY.VOBI") and instruct the drive controller (317) to read the target VOB.
The drive controller (317) moves the optical pickup according to the instruction of the presentation controller (306) and reads out the target AV data. The read AV data is read into the image memory (308) or the track buffer (309) as described above.
In addition, the scenario processor (305) monitors the time of the clock (307) and throws an event to the program processor (302) at the timing set in the management information.
The BD program (BD.PROG or XXX.PROG) recorded in the program recording memory (301) is executed and processed by the program processor 302. The program processor (302) processes the BD program when an event is sent from the scenario processor (305) or when an event is sent from the UOP manager (303). The UOP manager (303) generates an event for the program processor (302) when a request is sent from the user by the remote control key.
(Application space) FIG. 20 is a diagram showing a BD application space.
In the BD application space, a playlist (PlayList) is one playback unit. A playlist is a concatenation of cells, and has a static scenario, which is a playback sequence determined by the concatenation order, and a dynamic scenario, which is described by a program. Unless there is a dynamic scenario by the program, the playlist only plays the individual cells in order, and the playlist ends when all the cells have finished playing. On the other hand, the program can dynamically change the playback target beyond the playlist and the playback target depending on the user selection or the state of the player. A typical example is a menu. In the case of BD, a menu can be defined as a scenario to be played by the user's selection, and a playlist is dynamically selected by a program.
The program referred to here is an event handler executed by a time event or a user event.
A time event is an event generated based on time information embedded in a playlist. This corresponds to the event sent from the scenario processor (305) to the program processor (302) described in FIG. When a time event is issued, the program processor (302) executes the event handler associated with the ID. As described above, the program to be executed can instruct the playback of other playlists, in which case the playback of the currently playing playlist is stopped and the specified playlist is played. And transition.
A user event is an event generated by a user's remote control key operation. User events can be broadly divided into two types. The first is a menu selection event generated by operating the cursor keys ("up", "down", "left", "right" keys) or the "enter" key. The event handler corresponding to the menu selection event is valid only for a limited period in the playlist (the validity period of each event handler is set as the playlist information), and the "top" and "up" of the remote control The event handler that is valid when the "Down", "Left", "Right" key or "Enter" key is pressed is searched, and if there is a valid event handler, the event handler is executed. In other cases, the menu selection event will be ignored.
The second user event is a menu call event generated by operating the "menu" key. When the menu call event is generated, the global event handler is called. Global event handlers are playlist-independent and always valid event handlers. By using this function, it is possible to implement a DVD menu call (a function of calling an audio, subtitle menu, etc. during title playback and executing title playback from the interrupted point after changing the audio or subtitle).
A cell, which is a unit that constitutes a static scenario in a playlist, refers to a playback section of all or part of a VOB (MPEG stream). The cell has the playback section in the VOB as information on the start and end times. The VOB management information (VOBI) paired with each VOB has a time map (Time Map or TMAP), which is table information of the recording address corresponding to the playback time of the data, inside the VOB management information (VOBI). It is possible to derive the read start address and end address in the VOB (that is, in the target file "YYY.VOB") from the playback and end times of the VOB described above by the map. The details of the time map will be described later.
(Details of VOB) FIG. 21 is a configuration diagram of an MPEG stream (VOB) used in this embodiment.
As shown in FIG. 21, a VOB is composed of a plurality of VOBUs (Video Object Units). VOBU is GOP (Gr) in MPEG video stream. It is a playback unit as a multiplexed stream that also includes audio data, based on oup Of Pictures). VOBU has a video playback time of 1.0 seconds or less, and usually has a playback time of about 0.5 seconds.
The TS packet (MPEG-2 Transport Stream Packet) at the beginning of the VOBU contains the sequence header, the GOP header that follows it, and the I picture (Intra-coded), and decoding from this I picture can be started. There is. Also, the address of the TS packet including the beginning of the I picture at the beginning of this VOBU (start address), the address from this start address to the TS packet including the end of the I picture (end address), and the playback start time of this I picture. (PTS) is managed by the time map. Therefore, a time map entry is given for each TS packet at the beginning of the VOBU.
VOBU has a video packet (V_PKT) and an audio packet (A_PKT) inside. Each packet is 188 bytes, and although not shown in FIG. 21, an ATS (Arrival Time Stamp), which is the relative decoder supply start time of the TS packet, is added immediately before each TS packet. ..
ATS is assigned to each TS packet because the system rate of this TS stream is not a fixed rate but a variable rate. Generally, when fixing the system rate, a dummy TS packet called a NULL packet is inserted, but in order to record with high image quality in a limited recording capacity, a variable rate is suitable. In BD, it is recorded as a TS stream with ATS.
FIG. 22 is a diagram showing the configuration of the TS packet.
As shown in FIG. 22, the TS packet is composed of a TS packet header, an applicable field, and a payload part. A PID (Packet Identifier) is stored in the TS packet header, which identifies what kind of information the TS packet stores. PCR (Program Clock Reference) is stored in the applicable field. PCR is a reference value of the reference clock (system time clock, STC) of the device that decodes the stream. The instrument typically demultiplexes the system stream at the timing of PCR and reconstructs various streams such as video streams. PES packets are stored in the payload.
DTS (Decoding Time Stamp) and PTS (Presentation Time Stamp) are stored in the PES packet header. DTS indicates the decoding timing of the picture / audio frame stored in the PES packet, and PTS indicates the presentation timing such as video / audio output. Elementary data such as video data and audio data are sequentially put into a data storage area of a packet (PES Packet) called a PES packet payload (PES Packet Payload) from the beginning. The PES packet header also records an ID (stream_id) for identifying which stream the data stored in the payload is.
The details of the TS stream are specified in ISO / IEC 13818-1, and the characteristic of BD is that ATS is added to each TS packet.
(VOB interleave record) Next, the interleave recording of the VOB file will be described with reference to FIGS. 23 and 24.
The upper part of FIG. 23 is a part of the above-mentioned player configuration diagram. As shown in the figure, the data on the BD disc is input to the track buffer if it is a VOB or MPEG stream through the optical pickup, and is input to the image memory if it is PNG or image data.
The track buffer is a FIFO, and the input VOB data is sent to the demultiplexer in the order in which they were input. At this time, each TS packet is extracted from the track buffer according to the above-mentioned ATS, and data is delivered to the video processor or sound processor via the demultiplexer. On the other hand, in the case of image data, which image is drawn is instructed by the presentation controller. Further, the image data used for drawing is deleted from the image memory at the same time in the case of the image data for subtitles, but in the case of the image data for the menu, it is left as it is in the image memory during the drawing of the menu. This is because the drawing of the menu depends on the user operation, and a part of the menu may be redisplayed or replaced with a different image according to the user operation, and the image data of the redisplayed part is decoded at that time. This is to make it easier.
The lower part of FIG. 23 is a diagram showing interleave recording of VOB files and PNG files on a BD disc. Generally, in the case of ROM, for example, CD-ROM or DVD-ROM, AV data, which is a series of continuous playback units, is continuously recorded. This means that the drive only needs to read the data sequentially and deliver it to the decoder as long as it is continuously recorded, but if the continuous data is divided and discretely arranged on the disk, seek between individual continuous sections. This is because an operation is performed, data reading is stopped during this period, and data supply may be stopped. Similarly, in the case of BD, it is desirable that the VOB file can be recorded in a continuous area, but there is data that is played back in synchronization with the video data recorded in the VOB, such as subtitle data, and the VOB file. Similarly, it is necessary to read the subtitle data from the BD disc by some method.
As one means of reading the subtitle data, there is a method of reading the image data (PNG file) for the subtitles in a lump before starting the playback of the VOB. However, in this case, a large amount of memory is required, which is unrealistic.
Therefore, a method is used in which the VOB file is divided into several blocks and interleaved with the image data. The lower part of FIG. 23 is a diagram explaining the interleave record.
By appropriately interleaving the VOB file and the image data, it is possible to store the image data in the image memory at the required timing without the large amount of temporary recording memory as described above. However, when reading the image data, the reading of the VOB data will naturally stop.
FIG. 24 is a diagram illustrating a VOB data continuous supply model using a track buffer that solves this problem.
As described above, the VOB data is temporarily stored in the track buffer. If there is a difference (Va> Vb) between the data input rate (Va) to the track buffer and the data output rate (Vb) from the track buffer, the data in the track buffer will be accumulated as long as the data is continuously read from the BD disk. The amount will increase.
As shown in the upper part of FIG. 24, it is assumed that the continuous recording area of the VOB continues from the logical address "a1" to "a2". It is assumed that the section between "a2" and "a3" is a section in which image data is recorded and VOB data cannot be read.
The lower part of FIG. 24 is a diagram showing the inside of the track buffer. The horizontal axis shows time and the vertical axis shows the amount of data stored in the track buffer. The time "t1" indicates the time when the reading of "a1", which is the start point of the continuous recording area of the VOB, is started. After this time, data will be accumulated in the track buffer at a rate of Va-Vb. Needless to say, this rate is the difference between the input / output rates of the track buffers. The time "t2" is the time to read the data of "a2" which is the end point of the continuous recording area. That is, the amount of data in the track buffer increases at the rate Va-Vb between the time "t1" and "t2", and the amount of data accumulated B (t2) at the time "t2" can be calculated by the following equation.
B (t2) = (Va-Vb) × (t2-t1) (Equation 1)
After this, since the image data continues up to the address "a3" on the BD disc, the input to the track buffer becomes 0, and the amount of data in the track buffer decreases at the output rate "-Vb". Become. This is up to the read position "a3" and up to "t3" in time.
The important thing here is that if the amount of data stored in the track buffer before the time "t3" becomes 0, the VOB data supplied to the decoder will be exhausted, and VOB playback may stop. There is. However, if data remains in the track buffer at time "t3", it means that VOB playback can be continued without stopping.
This condition can be expressed by the following equation.
B (t2) -Vb × (t3-t2) (Equation 2)
That is, the arrangement of the image data (non-VOB data) should be determined so as to satisfy Equation 2.
(Navigation data structure) The BD navigation data (BD management information) structure will be described with reference to FIGS. 25 to 31.
FIG. 25 is a diagram showing the internal structure of the VOB management information file (YYY.VOBI).
The VOB management information has the stream attribute information (Attribute) of the VOB and the time map. The stream attributes are configured to have individual video attributes (Video) and audio attributes (Audio # 0 to Audio # m). Especially in the case of audio streams, since the VOB can have multiple audio streams at the same time, the presence or absence of a data field is indicated by the number of audio streams (Number).
The following are the fields that the video attribute (Video) has and the values that each can have.
Compression method (Coding): MPEG1 MPEG2 MPEG4 MPEG4-AVC (Advanced Video Coding) Resolution: 1920x1080 1440x1080 1280x720 720x480 720x565 Aspect: 4: 3 16: 9 Frame rate: 60 59.94 (60 / 1.001) 50 30 29.97 (30 / 1.001) twenty five twenty four 23.976 (24 / 1.001) The following are the fields that the audio attribute (Audio) has and the values that each can have.
Compression method (Coding): AC3 MPEG1 MPEG2 LPCM Number of channels (Ch): 1 ~ 8 Language attribute:
The time map (TMAP) is a table having information for each VOBU, and has the number of VOBUs (Number) possessed by the VOB and each VOBU information (VOBU # 1 to VOBU # n). Each VOBU information is composed of the address I_start of the VOBU first TS packet (I picture start), the offset address (I_end) to the end address of the I picture, and the playback start time (PTS) of the I picture.
The value of I_end may have an offset value, that is, an actual end address of the I picture, instead of having the size of the I picture.
FIG. 26 is a diagram illustrating details of VOBU information.
As is widely known, MPEG video streams may be variable bit rate compressed for high quality recording, and there is no simple correlation between their playback time and data size. On the contrary, since AC3, which is an audio compression standard, compresses at a fixed bit rate, the relationship between time and address can be obtained by a linear equation. However, in the case of MPEG video data, each frame has a fixed display time, for example, in the case of NTSC, one frame has a display time of 1 / 29.97 seconds, but the data size after compression of each frame is a characteristic of the picture. The data size varies greatly depending on the picture type used for compression, the so-called I / P / B picture. Therefore, in the case of MPEG video, the relationship between time and address cannot be expressed in the form of a linear expression.
As a matter of course, the MPEG system stream that multiplexes the MPEG video data, that is, VOB, cannot express the time and the data size in the form of a linear expression. Therefore, it is the time map (TMAP) that connects the relationship between the time and the address in the VOB.
In this way, when a certain time information is given, first search which VOBU the time belongs to (follow the PTS for each VOBU), and make the VOBU that has the PTS immediately before the time in TMAP. Dive (address specified by I_start), decryption starts from the I picture at the beginning of VOBU, and display starts from the picture at that time.
Next, the internal structure of the playlist information (XXX.PL) will be described with reference to FIG. 27.
The playlist information is composed of a cell list (CellList) and an event list (EventList).
A cell list is a playback cell sequence in a playlist, and cells are played in the order described in this list. The contents of the cell list (Cell List) are the number of cells (Number) and each cell information (Cell # 1 to Cell # n).
The cell information (Cell #) has a VOB file name (VOBName), a start time (In) and an end time (Out) in the VOB, and a subtitle table (SubtitleTable). The start time (In) and end time (Out) are represented by frame numbers in the VOB, respectively, and the address of VOB data required for playback can be obtained by using the time map described above.
The subtitle table (Subtitle Table) is a table having subtitle information to be played back in synchronization with the VOB. Subtitles can have multiple languages as well as audio, and the first information in the subtitle table is also composed of the number of languages (Number) and the subsequent tables for each language (Language # 1 ~ Language # k). There is.
The table for each language (Language #) includes language information (Lang), the number of subtitle information displayed individually (Number), and subtitle information for individually displayed subtitles (Speech # 1 to Speech # j). The subtitle information (Speech #) is composed of the corresponding image data file name (Name), subtitle display start time (In), subtitle display end time (Out), and subtitle display position (Position). ..
The event list (EventList) is a table that defines the events that occur in the playlist. The event list consists of individual events (Event # 1 to Event # m) following the number of events (Number), and each event (Event #) is the event type (Type) and event ID (ID). , It consists of the event occurrence time (Time) and the validity period (Duration).
FIG. 28 is an event handler table (XXX.PROG) having event handlers (time events and user events for menu selection) for each playlist.
The event handler table has a defined event handler / number of programs (Number) and individual event handlers / programs (Program # 1 to Program # n). The description in each event handler / program (Program #) has the event handler start definition (<event_handler> tag) and the event handler ID (ID) that is paired with the event ID described above, and then the program also Describe between the parentheses "{" and "}" following the Function. The events (Event # 1 to Event # m) stored in the above-mentioned event list (EventList) of "XXX.PL" are specified by using the ID (ID) of the event handler of "XXX.PROG".
Next, the internal structure of information (BD.INFO) regarding the entire BD disc will be described with reference to FIG. 29.
The entire BD disc information is composed of a title list (Title List) and an event table (Event List) for global events.
The title list (Title List) is composed of the number of titles (Number) in the disk and the title information (Title # 1 to Title # n) following it. The individual title information (Title #) includes a playlist table (PLTable) included in the title and a chapter list (ChapterList) in the title. The playlist table (PL Table) has the number of playlists in the title (Number) and the playlist name (Name), that is, the file name of the playlist.
The chapter list is composed of the number of chapters (Number) included in the title and individual chapter information (Chapter # 1 to Chapter # n), and the individual chapter information (Chapter #) is the cell contained in the chapter. It has a table (Cell Table), and the cell table (Cell Table) is composed of the number of cells (Number) and the entry information (CellEntry # 1 to CellEntry # k) of each cell. The cell entry information (CellEntry #) is described by the playlist name including the cell and the cell number in the playlist.
The event list has the number of global events (Number) and information on individual global events. It should be noted here that the first global event defined is called the First Event, which is the first event to be called when a BD disc is inserted into the player. The event information for global events has only the event type (Type) and the event ID (ID).
FIG. 30 is a table (BD.PROG) of the program of the global event handler.
This table has the same contents as the event handler table described with reference to FIG. 28.
(Mechanism of event occurrence) The mechanism of event occurrence will be described with reference to FIGS. 31 to 33.
Figure 31 is an example of a time event.
As described above, the time event is defined in the event list (EventList) of the playlist information (XXX.PL). When the event defined as a time event, that is, the event type (Type) is "TimeEvent", the time event with the ID "Ex1" is changed from the scenario processor to the program processor when the event generation time ("t1") is reached. Can be given to. The program processor searches for an event handler with the event ID "Ex1" and executes and processes the target event handler. For example, in the case of this embodiment, it is possible to draw two button images.
FIG. 32 is an example of a user event that performs a menu operation.
As described above, the user event for performing the menu operation is also defined in the event list (EventList) of the playlist information (XXX.PL). When the event defined as a user event, that is, the event type (Type) is "UserEvent", the user event becomes ready when the event generation time ("t1") is reached. At this time, the event itself has not yet been generated. The event is ready for the period indicated in the Duration information (Duration).
As shown in FIG. 32, when the user presses the "up", "down", "left", "right" or "enter" keys of the remote control keys, a UOP event is first generated by the UOP manager and raised to the program processor. The program processor sends a UOP event to the scenario processor, and the scenario processor searches for a valid user event at the time when the UOP event is received, and if there is a target user event, the user event is sent. Generate and lift to program processor. The program processor searches for an event handler with the event ID "Ev1" and executes the target event handler. For example, in the case of this embodiment, the playback of playlist # 2 is started.
The generated user event does not contain information about which remote control key was pressed by the user. The information of the selected remote control key is transmitted to the program processor by the UOP event, and is recorded and held in the register SPRM (8) of the virtual player. The event handler program can check the value of this register and execute branch processing.
Figure 33 is an example of a global event.
As mentioned above, global events are defined in the event list (EventList) of information about the entire BD disc (BD.INFO). When the event defined as a global event, that is, the event type (Type) is "GlobalEvent", the event is generated only when the user operates the remote control key.
When the user presses "Menu", a UOP event is first generated by the UOP manager and raised to the program processor. The program processor sends a UOP event to the scenario processor, and the scenario processor generates the corresponding global event and sends it to the program processor. The program processor searches for an event handler with the event ID "menu" and executes the target event handler. For example, in the case of this embodiment, the playback of playlist # 3 is started.
In this embodiment, it is simply called a "menu" key, but there may be a plurality of menu keys such as a DVD. It is possible to correspond by defining the ID corresponding to each menu key.
(Virtual player machine) The functional configuration of the program processor will be described with reference to FIG. 34.
A program processor is a processing module that has a virtual player machine inside. The virtual player machine is a functional model defined as BD and does not depend on the implementation of each BD player. That is, it is guaranteed that the same function can be executed in any BD player.
The virtual player machine has two main functions. Programming functions and player variables (registers). The programming function is based on Java (registered trademark) Script and defines the following two functions as BD-specific functions.
Link function: Stops current playback and starts playback from the specified playlist, cell, and time Link (PL #, Cell #, time) PL #: Playlist name Cell #: Cell number time: Playback start time in the cell PNG drawing function: Draws the specified PNG data on the image plane Draw (File, X, Y) File: PNG file name X: X coordinate position Y: Y coordinate position Image plane clear function: Clears the specified area of the image plane Clear (X, Y, W, H) X: X coordinate position Y: Y coordinate position W: X direction width H: Y direction width Player variables include system parameters (SPRM) that indicate the state of the player and general parameters (GPRM) that can be used for general purposes.
Figure 35 is a list of system parameters (SPRM).
SPRM (0): Language code SPRM (1): Audio stream number SPRM (2): Subtitle stream number SPRM (3): Angle number SPRM (4): Title number SPRM (5): Chapter number SPRM (6): Program number SPRM (7): Cell number SPRM (8): Select key information SPRM (9): Navigation timer SPRM (10): Playback time information SPRM (11): Mixing mode for karaoke SPRM (12): Country information for parrenting SPRM (13): Parent level SPRM (14): Player settings (video) SPRM (15): Player settings (audio) SPRM (16): Language code for audio stream SPRM (17): Language code for audio stream (extended) SPRM (18): Language code for subtitle stream SPRM (19): Language code for subtitle stream (extended) SPRM (20): Player region code SPRM (21): Spare SPRM (22): Spare SPRM (23): Playback status SPRM (24): Spare SPRM (25): Spare SPRM (26): Spare SPRM (27): Spare SPRM (28): Spare SPRM (29): Spare SPRM (30): Spare SPRM (31): Reserve In this embodiment, the programming function of the virtual player is based on Java (registered trademark) Script, but instead of Java (registered trademark) Script, B-Shell used in UNIX (registered trademark) OS etc. Other programming functions such as Perl Script may be used, in other words, the present invention is not limited to Java® Script.
(Program example) 36 and 37 are examples of programs in the event handler.
FIG. 36 is an example of a menu with two select buttons.
The program on the left side of Fig. 36 is executed using a time event at the beginning of the cell (PlayList # 1.Cell # 1). Here, "1" is set to one of the general parameters GPRM (0) first. GPRM (0) is used to identify the selected button in the program. In the initial state, the button 1 to be placed on the left side is selected as the initial value.
Next, PNG drawing is done for button 1 and button 2 using the drawing function Draw. Button 1 draws a PNG image "1black.png" starting from the coordinates (10, 200) (far left). Button 2 draws a PNG image "2white.png" starting from the coordinates (330,200) (far left).
At the end of this cell, the program on the right side of Fig. 36 is executed using the time event. Here, it is specified to play again from the beginning of the cell by using the Link function.
FIG. 37 is an example of an event handler for a menu selection user event.
When any of the "left" key, "right" key, and "enter" key is pressed, the corresponding program is written in the event handler. When the user presses the remote control key, as described in FIG. 32, a user event is generated and the event handler of FIG. 37 is activated. In this event handler, branch processing is performed using the value of GPRM (0) that identifies the select button and SPRM (8) that identifies the selected remote control key.
Condition 1) When button 1 is selected and the selection key is the "right" key Reset GPRM (0) to 2 and change the selected button to right button 2.
Rewrite the images of Button 1 and Button 2 respectively.
Condition 2) When the selection key is "OK" and button 1 is selected Start playing playlist # 2
Condition 3) When the selection key is "OK" and button 2 is selected Start playing playlist # 3 Execution processing is performed as described above.
(Player processing flow) Next, the processing flow in the player will be described with reference to FIGS. 38 to 41.
FIG. 38 shows a basic processing flow up to AV playback.
When a BD disc is inserted (S101), the BD player reads and parses the BD.INFO file (S102) and reads BD.PROG (S103). Both BD.INFO and BD.PROG are temporarily stored in the management information recording memory and analyzed by the scenario processor.
The scenario processor then generates the first event according to the First Event information in the BD.INFO file (S104). The generated first event is received by the program processor, and the event handler corresponding to the event is executed (S105).
It is expected that the event handler corresponding to the first event records the playlist information to be played first. If the playlist playback is not instructed, the player does not play anything and just keeps waiting for the user event. (S201). When the BD player accepts the remote control operation from the user, the UOP manager launches a UOP event for the program manager (S202).
The program manager determines whether the UOP event is a menu key (S203), and if it is a menu key, sends the UOP event to the scenario processor, and the scenario processor generates a user event (S204). The program processor executes and processes the event handler corresponding to the generated user event (S205).
FIG. 39 is a processing flow from the start of PL reproduction to the start of VOB reproduction.
As mentioned above, the playlist playback is started by the first event handler or the global event handler (S301). The scenario processor reads and analyzes the playlist information "XXX.PL" (S302) and reads the program information "XXX.PROG" corresponding to the playlist as information necessary for playing the playlist to be played (S303). ). Subsequently, the scenario processor instructs the cell to be regenerated based on the cell information registered in the playlist (S304). Cell playback means that the scenario processor makes a request to the presentation controller, and the presentation controller starts AV playback (S305).
When the start of AV playback (S401) is started, the presentation controller reads and analyzes the VOB information file (XXX.VOBI) corresponding to the cell to be played (S402). The presentation controller uses the time map to identify the VOBU to start playback and its address, instructs the drive controller to read the address, the drive controller reads the target VOB data (S403), and the VOB data is sent to the decoder. Playback begins (S404).
VOB reproduction is continued until the reproduction section of the VOB ends (S405), and when it ends, the process shifts to the next cell reproduction (S304). If there are no cells next, playback will stop (S406).
FIG. 40 shows an event processing flow after the start of AV playback.
The BD player is an event-driven player model. When the playlist playback is started, the time event system, user event system, and subtitle display system event processing processes are started respectively, and the event processing is executed in parallel.
The processing of the S500 system is a processing flow of the time event system.
After the playlist playback starts (S501), the scenario processor checks whether the time event occurrence time has come through the step (S502) of checking whether the playlist playback has ended (S503). When the time event occurrence time is reached, the scenario processor generates a time event (S504), and the program processor receives the time event and executes the event handler (S505).
If the time event occurrence time has not been reached in step S503, or after the event handler execution process is performed in step S504, the process returns to step S502 and the above process is repeated. Further, when it is confirmed in step S502 that the playlist playback is completed, the time event processing is forcibly terminated.
The S600 system processing is a user event system processing flow.
After the playlist playback starts (S601), the process proceeds to the UOP reception confirmation step (S603) through the playlist playback end confirmation step (S602). When the UOP is accepted, the UOP manager generates a UOP event (S604), and the program processor that receives the UOP event checks whether the UOP event is a menu call (S605), and if it is a menu call, , The program processor causes the scenario processor to generate an event (S607), and the program processor executes and processes the event handler (S608).
If it is determined in step S605 that the UOP event is not a menu call, it indicates that the UOP event is a cursor key or "OK" key event. this In this case, the scenario processor determines whether the current time is within the valid period of the user event (S606), and if it is within the valid period, the scenario processor generates the user event (S607), and the program processor is the target event. Execute the handler (S608).
If there is no UOP reception in step S603, if the current time is not in the user event valid period in step S606, or after the event handler execution process in step S608, the process returns to step S602 and the above process is repeated. Further, when it is confirmed in step S602 that the playlist playback is completed, the user event processing is forcibly terminated.
FIG. 41 shows the flow of subtitle processing.
After the playlist playback starts (S701), the process proceeds to the subtitle drawing start time confirmation step (S703) through the playlist playback end confirmation step (S702). For the subtitle drawing start time, the scenario processor instructs the presentation controller to draw the subtitle, and the presentation controller instructs the image processor to draw the subtitle (S704). If it is determined in step S703 that it is not the subtitle drawing start time, it is confirmed whether it is the subtitle display end time (S705). When it is determined that the subtitle display end time is reached, the presentation controller instructs the image processor to erase the subtitles and erases the drawn subtitles from the image plane (S706).
After the subtitle drawing step S704 is completed, after the subtitle erasing step S706 is completed, or when it is determined in the subtitle display end time confirmation step S705 that the time is not the relevant time, the process returns to step S702 and the above-described processing is repeated. Further, when it is confirmed in step S702 that the playlist playback is completed, the subtitle display system processing is forcibly terminated.
(Embodiment 2) Next, a second embodiment of the present invention will be described.
The second embodiment is a description for realizing a slide show using a still image by applying the above application. Basically, the content is based on the first embodiment, and the explanation will focus on the extension or the different part.
(See I picture) Figure 42 shows the relationship between a slide show (still image application) and a time map. A slide show usually consists of only still images (I pictures). The time map has the position and size information of the still image data, and when a certain still image is selected, one still image is displayed by extracting the necessary data and sending it to the decoder. Normally, slide shows are not always displayed in order like a video, but because the display order is not fixed due to user interaction, it is intra-encoded so that it can be displayed from anywhere. I am using the I picture.
However, in order to reduce the amount of data, it is possible to realize a slide show even with a P picture that is compressed by referring to an I picture or a B picture that is compressed by referring to two or more pictures before and after.
However, the P picture and B picture cannot be decoded without the referenced picture. Therefore, even if an attempt is made to start playback from a P picture or B picture in the middle due to user interaction, decoding cannot be performed. Therefore, as shown in FIG. 43, Prepare a flag to indicate that the picture pointed to by the time map is an I picture and that it does not refer to any other image. By referring to this flag, if the reference image is not needed, that is, if it can be decoded independently, it is possible to decode and display from that image regardless of the previous and next display, but if the reference image is needed. Indicates that the image cannot be displayed depending on the display order because the related image cannot be displayed unless it has been decoded so far.
As a whole time map, a flag indicating that the picture referenced from the time map is always an I picture, that is, any picture can be decoded independently is set in the time map or as shown in FIG. It may be recorded in one place in the relevant navigation information. If this flag is not set, the timemap entry does not always point to an I picture, and there is no guarantee that the referenced picture can be decoded.
In the explanation so far, it was explained as an I picture based on the MPEG2 video stream, but in the case of MPEG4-AVC (also called H.264 or JVT), IDR (Instantaneous Decoder refresh) picture or IDR picture. It may be an I-picture other than the above, and even in the case of an image of another format, if it is an image that can be decoded independently, it can be easily applied.
(Guarantee of reference to all I pictures) Figure 45 shows the difference between a video application and a still image application (slide show). As shown in Fig. 45 (a), in the case of a video application, once playback is started, subsequent pictures are continuously decoded, so it is not necessary to set a reference from the time map for all I pictures. , At least the time map entry needs to be set only at the point where you want to start playback.
FIG. 45 (b) is an example of a slide show. In the case of a slide show, it is necessary to display still images regardless of the order such as skip operation without displaying the previous and next images by the user's operation. Therefore, unless the time map entries are registered for all I pictures, the data of the I pictures to be displayed cannot be sent to the decoder unless all the streams are actually analyzed, which is inefficient. If each I picture has a time map entry, only the necessary I picture data can be directly accessed, the data can be read and sent to the decoder, so access efficiency is good and the time to display is short, which is efficient. Is good.
If it can be identified that an entry exists for all I pictures, the range of data to be read can be known by referring to the time map entry when accessing any I picture, so the preceding and following streams are analyzed extra. You don't have to.
If it is not guaranteed that entries exist for all I pictures, and if you specify to display I pictures that are not registered in the time map, the necessary data is extracted while analyzing the streams before and after that. It is inefficient because it must be accessed, access efficiency is poor, and it takes time to display.
Therefore, as shown in FIG. 46, it is necessary to analyze the previous and next streams by providing a flag indicating whether or not all I pictures are guaranteed to be referenced from the time map. Such flags are useful because they are not needed or can be identified simply by parsing static data.
Note that this flag is effective not only in still image applications such as slide shows, but also in moving image applications, and is a flag that guarantees that playback can be started from any I picture.
(Embodiment 3) In the second embodiment, it was described that MPEG-4 AVC can be used as a coding method for realizing a still image application. The still image of MPEG-4 AVC is defined as AVC Still Picture in the extended standard for MPEG-4 AVC (ISO / IEC 13818-1 Amendment 3) of the MPEG-2 system, not the MPEG-4 AVC standard itself. However, the MPEG-2 system standard does not specify the playback method of still images, and it is necessary to separately specify the playback method in order to use it in a still image application. In this embodiment, a still image data structure and a display method for applying MPEG-4 AVC to a still image application will be described.
The AVC Still Picture in the MPEG-2 system standard is defined to include an IDR picture, an SPS (Sequence Parameter Set) referenced from the IDR picture, and a (Picture Parameter Set). FIG. 47 shows the data structure of the MPEG-4 AVC still image (hereinafter referred to as AVC still image) in the present embodiment. The boxes in the figure indicate NAL units (Network Abstraction Units). Be sure to include the End of Sequence NAL unit in the AVC still image. Since End of Sequence is identification information that indicates the end of the sequence in MPEG-4 AVC, by arranging the NAL unit of End of Sequence and ending the sequence, the MPEG-4 AVC can be used to display the AVC still image. It can be defined independently outside the standard. Here, the order of appearance of each NAL unit shall be in accordance with the regulations stipulated in the MPEG-4 AVC standard.
Next, a method of displaying an AVC still image will be described with reference to FIG. 48. In the still image application, it is necessary to specify the display time of the still image and the display time length of the still image. The display time (PTS: Presentation Time Stamp) of the AVC still image is acquired from the time map or the header of the PES (Packetized Elemantary Stream) packet. Here, when the display time of all the still images is indicated by the time map, the display time can be acquired by referring only to the time map. From the display time of the Nth AVC still image to the display time of the N + 1th AVC still image, the display of the Nth AVC still image is frozen, that is, the Nth AVC still image is repeatedly displayed. I will do it.
When playing back an AVC still image, it is desirable to be able to obtain the frame rate from the AVC still image data. In MPEG-4 AVC, the display rate of a video stream can be indicated by VUI (Video Usability Information) in SPS. Specifically, it refers to the three fields num_units_in_tick, time_scale, and fixed_frame_rate_flag. Here, time_scale indicates a time scale, and for example, the time_scale of a clock operating at 30000 Hz can be 30000. num_units_in_tick is a basic unit indicating the time when the clock operates. For example, if the num_units_in_tick of the clock whose time_scale is 30000 is 1001, it can be indicated that the basic period when the clock operates is 29.97 Hz. Also, by setting fixed_frame_rate_flag, it can be shown that the frame rate is fixed. MPEG-4 In AVC, these fields can be used to indicate the difference value between the display times of two consecutive pictures, but in the present embodiment, these fields are used to repeatedly display the AVC still image. I will show the frame rate at the time. First, by setting fixed_frame_rate_flag to 1, we show that the frame rate is fixed. Next, when setting the frame rate to 23.976Hz, for example, set num_units_in_tick to 1001 and time_scale to 24000. That is, frame rate = time_scale / num_units_in_tick Set both fields so that In addition, set both vui_parameters_present_flag in SPS and timing_info_present_flag in VUI to 1 to ensure that the above three fields in VUI and VUI exist. When the Nth AVC still image is the final AVC still image, the display is frozen until there is a user action or until the next action or the like predetermined by the program starts. The frame rate setting method is not limited to time_scale / num_units_in_tick. For example, in an MPEG-4 AVC video stream, time_scale / num_units_in_tick indicates the field rate (a parameter indicating the field display interval), so the frame rate is time_scale / num_units_in_tick / 2. Therefore, the frame rate may be set to time_scale / num_units_in_tic / 2 even for still images.
The frame rate indicated by the above method shall match the frame rate value indicated in the BD management information. Specifically, it matches the value indicated by the frame_rate field in StreamCodingInfo.
The display cycle when the AVC still image is repeatedly displayed can be obtained from the frame rate indicated by the above method. This display cycle may be an integral multiple of the frame grid or the field grid. By doing so, it is possible to guarantee synchronous playback with other video sources such as video and graphics. Here, the frame grid or field grid is generated based on the frame rate of a specific stream such as video. Further, the difference value of the display time of the Nth and N + 1th AVC still images may be an integral multiple of the frame grid or the field grid.
As the time map to be referred to when reproducing the AVC still image, the time map in the second embodiment is used.
Note that these fields may be omitted by specifying the default values of num_units_in_tick, time_scale, and fixed_frame_rate_flag in the BD ROM standard and the like.
In the case of a video stream, it is prohibited to change the resolution in the stream, but in the still image stream, even if the resolution is converted, buffer management in the decoding operation can be realized without failure, so in the stream. The resolution may be changeable. Here, the resolution is indicated by a field in the SPS.
Even if the encoding method is other than MPGE-4 AVC, if it has the same data structure, the data structure and the reproduction method of the present embodiment can be applied.
(Embodiment 4) In the second and third embodiments, it has been described that MPEG-4 AVC can be used as a coding method for realizing a still image application. In the fourth embodiment, an information recording medium capable of encoding a still image with high image quality while suppressing the amount of processing during moving image reproduction in a package medium such as a BD-ROM, and a reproduction device thereof will be described.
First, a conventional information recording medium will be described. In MPEG-4 AVC, the maximum value of the code amount of a picture is specified. In application standards such as BD, the specified value in MPEG-4 AVC or the value set independently in the application is set as the upper limit of the code amount of the picture. The upper limit can be limited by a parameter called MinCR (Minimum Compression Ratio) specified in the MPEG-4 AVC standard. MinCR is a parameter that indicates the lower limit of the compression ratio of the coded picture with respect to the original image. For example, if MinCR is 2, it means that the code amount of the coded picture is less than half the size of the original image.
In the conventional information recording medium, the same value is used as MinCR in the moving image application and the still image application. In moving images, the amount of processing required to decode the coded data is large, so operation is guaranteed even in the worst case where the amount of calculation required to decode one picture is the upper limit specified by the standard. MinCR is determined so that it can be done. On the other hand, in a still image, since the display interval is longer than that of a moving image, the image quality is more important than the processing amount at the time of decoding. However, since the amount of coding increases when a still image is encoded with high image quality, a conventional information recording medium in which the MinCR is the same for a still image and a moving image allocates a sufficient number of bits for a picture, especially at the time of intra-coding. There was a problem that it could not be done.
In the information recording medium according to the fourth embodiment, by applying different MinCRs to the moving image and the still image, the MinCR value is increased for the moving image in consideration of the processing amount at the time of decoding, and the still image has high image quality. Make the MinCR value smaller than the video so that you can guarantee a sufficient picture size to encode to.
FIG. 49 shows an example of the data structure of the information recording medium of the fourth embodiment. The stream management information in the BD management information indicates the attribute information of the clip in the data object called ClipInfo. The clip refers to an AV data file, and for example, one file containing an MPEG-4 AVC still image stream is one clip. In order to show that different MinCR is applied to the moving image and the still image, information indicating the MinCR value for each clip is required. Therefore, information indicating the MinCR value applied in the referenced clip is added to ClipInfo. Here, assuming that the MinCR value applied to the clip of the still image and the clip of the moving image is predetermined, by storing the flag information indicating whether the clip to be referred to is a moving image or a still image. , I will show the MinCR value applied to the clip. In the example of FIG. 49, at least still image and video clips are stored in the disc, which are referred to by ClipInfo # 1 and ClipInfo # 2, respectively. Here, ClipInfo # 1 is clicked Flag information indicating that the clip is a still image, ClipInfo # 2 is a video clip Flag information indicating that is stored. By referring to this flag information, the MinCR value in the pictures constituting the clip can be acquired. In the example of FIG. 49, the MinCR of the clip of the still image is set to 2, and the MinCR of the clip of the moving image is set to 4, so that the image quality of the still image is improved and the processing amount at the time of decoding the moving image is suppressed. The MinCR value here is an example, and other combinations may be used. In an application in which the processing amount of the playback device is sufficient, the MinCR value of the still image and the moving image may be the same. Further, a plurality of combinations of MinCR values for still images and moving images may be determined in advance, and the MinCR values may be indicated by introducing a parameter indicating a specific combination. Further, when it is shown that the clip is a still image, it may be guaranteed that the interval between decoding or displaying two consecutive pictures is equal to or larger than a predetermined value. For example, in the case of a still image, it is assumed that the display interval of two consecutive pictures is 0.5 seconds or more. As a result, even when the MinCR value is 2 and the code amount per picture is large, the display interval is sufficiently long, 0.5 seconds or more, so that the decoding of each picture can be guaranteed.
Note that ClipInfo has a field called application_type that indicates the type of application that plays the clip. In this field, it is possible to indicate whether the application is a moving image or a still image, and if it is a still image, whether it is time-based or browserable. Here, the time base is for displaying still images at predetermined intervals, and the browserable is an application in which the user can determine the display timing of the still images. Therefore, when the field value of application_type refers to a time-based or browserable still image application, the MinCR value for still images is applied, and when referring to a video application, the MinCR value for video is applied. May be done.
The MinCR value can be switched between clips of different videos as well as between moving images and still images. For example, when the main video and the sub video are included, the MinCR value can be set small for the main video and encoded with high image quality, and the MinCR value can be set large for the sub video in consideration of the processing amount. .. At this time, as the information indicating the MinCR value, the information indicating the MinCR value for each clip is used instead of the flag information indicating whether the image is a still image or a moving image.
The parameter indicating the upper limit of the code amount of the moving image or the still image is not limited to MinCR, and may be another parameter such as directly indicating the upper limit value of the code amount as the data size.
Information indicating the upper limit of the code amount of the picture in the clip may be stored in BD management information other than ClipInfo, or may be stored in the coded data. When storing in coded data, it can be stored for each random access unit such as GOP (Group Of Picture). For example, in MPEG-4 AVC, a data unit for storing user data can be used. The data unit for storing user data includes a NAL (Network Abstraction Layer) unit having a specific type, an SEI (Supplemental Enhancement Information) message for storing user data, and the like. Further, the upper limit of the code amount of the picture may be switched in a unit different from the clip, such as a random access unit.
Further, depending on the data playback device, when decoding a moving image, if it is determined that the time required to decode the coded data of one picture exceeds a predetermined time or the display interval of the picture, the picture is determined. Decoding may be skipped and decoding of the next picture may start. Alternatively, even if the worst case can be dealt with when decoding a moving image, when reproducing a still image in the information recording medium of the present embodiment, the upper limit of the code amount of the still image becomes larger than that of the moving image, and the code amount becomes larger. As the size increases, the time required for decoding also increases, and as a result, decoding of the still image may be skipped. Here, since the display interval of a still image is generally longer than that of a moving image, even if the decoding is not completed by the preset display start time, if the image is displayed after the decoding is completed, the reproduction quality is deteriorated. It is minor. Therefore, when decoding a still image, even if the decoding is not completed by the preset display start time, the decoding may be displayed after the decoding is completed without skipping the decoding.
Although BD has been described above, the same method can be used as long as it is an information recording medium capable of storing still images and moving images. Also, the coding method is not limited to MPEG-4 AVC, and can be applied to other coding methods such as MPEG-2 Video.
(Embodiment 5) FIG. 50 is a flowchart showing a multiplexing method for creating data stored in the information recording medium according to the fourth embodiment. In the multiplexing method of the present embodiment, a step of switching the MinCR value according to the type of clip (step S2001, step S2002, and step S2003) and a flag information for specifying the MinCR value are generated and management information is generated. It differs from the conventional multiplexing method in that it includes steps (step S2004 and step S2005) to be included in.
First, in step S2001, it is determined whether the generated clip is a moving image or a still image. If the clip is a still image, proceed to step S2002 to set a predetermined MinCR value for the still image clip, and if the clip is a moving image, proceed to step S2003 to set a predetermined MinCR value for the moving image clip. Set the MinCR value of. Next, in step S1001, the pictures constituting the clip are encoded so as to satisfy the MinCR value set in step S2002 or step S2003, and the process proceeds to step S1002. In step S1002, the data encoded in step S1001 is system-multiplexed. BD uses an MPEG-2 transport stream as the system multiplexing method. Next, in step S2004, flag information for specifying the MinCR value applied to the pictures constituting the clip is generated, and in step S2005, management information including the flag information generated in step S2004 is generated. Finally, in step S1003, the management information and the system-multiplexed coded data are combined and output. Further, the information for specifying the MinCR value may be other than the flag information, such as directly storing the maximum value of the code amount of the picture.
Data such as audio and graphics can also be multiplexed together with moving images or still images, but the description thereof will be omitted here.
FIG. 51 is a block diagram showing a configuration of a multiplexing device 2000 that realizes the multiplexing method of the fifth embodiment. The multiplexing device 2000 includes a MinCR determination unit 2001, a MinCR information generation unit 2002, a coding unit 1001, a system multiplexing unit 1002, a management information creation unit 2003, and a coupling unit 1003, and includes a MinCR determination unit 2001 and a MinCR information generation unit 2002. It differs from the conventional multiplexing device in that it is provided with the above and that the management information creation unit 2003 generates management information including flag information for specifying the MinCR value.
The operation of each part will be described below. The MinCR determination unit determines the MinCR value to be applied to the pictures constituting the clip based on the clip attribute ClipChar indicating whether the clip is a moving image or a still image, and sets the determined MinCR value cr as the encoding unit 1001. Input to MinCR information generator 2002. The coding unit 1001 encodes the input moving image or image data Vin based on the MinCR value cr determined by the MinCR determining unit, and inputs the coded data Cdata to the system multiplexing unit 1002. The system multiplexing unit 1002 system-multiplexes the encoded data Cdata and inputs the multiplexed data Mdata to the coupling unit 1003. On the other hand, the MinCR information generation unit generates MinCR information crInf, which is flag information for specifying the MinCR value applied to the pictures constituting the clip, based on the MinCR value cr, and inputs it to the management information creation unit 2003. .. The management information generation unit acquires the stream information StrInf for generating management information about the multiplexed data Mdata such as a time map from the system multiplexing unit 1002, generates the management information CtrInf including the MinCR information crInf, and combines them. Output to unit 1003. The coupling unit 1003 combines the management information CtrInf and the multiplexed data Mdata and outputs them as recorded data Dout. Here, the coding unit 1001 may decode two consecutive pictures or set a lower limit of the display interval based on the clip type or the MinCR value.
When creating data with an authoring tool, etc., the coded data may be generated and the system multiplexing or management information may be created by separate devices. Even in such a case, the operation of each device may be performed. May be the same as each part in the multiplexing device 2000.
(Embodiment 6) Further, the information recording medium shown in each of the above embodiments, a reproduction method thereof, and a program for realizing the recording method are shown in each of the above embodiments by recording on a recording medium such as a flexible disk. The processing can be easily performed in an independent computer system.
52A to 52C are explanatory views when the reproduction method and the recording method of each of the above-described embodiments are carried out by a computer system using a program recorded on a recording medium such as a flexible disk.
FIG. 52B shows the front view, cross-sectional structure, and flexible disc of the flexible disc, and FIG. 52A shows an example of the physical format of the flexible disc, which is the main body of the recording medium. The flexible disk FD is built in the case F, and a plurality of track Trs are concentrically formed on the surface of the disk from the outer circumference toward the inner circumference, and each track is divided into 16 sectors Se in the angular direction. ing. Therefore, in the flexible disk in which the program is stored, the program is recorded in the area allocated on the flexible disk FD.
Further, FIG. 52C shows a configuration for recording / reproducing the above program on the flexible disk FD. When recording the above program that realizes the reproduction method and the recording method on the flexible disk FD, the above program is written from the computer system Cs via the flexible disk drive. Further, when a reproduction method and a recording method for realizing a reproduction method and a recording method by a program in a flexible disk are constructed in a computer system, the program is read from the flexible disk by a flexible disk drive and transferred to the computer system.
In the above description, a flexible disk is used as the recording medium, but an optical disk can also be used in the same manner. The recording medium is not limited to this, and any recording medium such as an IC card or a ROM cassette that can record a program can be used in the same manner.
Each functional block in the block diagram shown in FIGS. 10, 18, 19, 23, 51, etc. is typically realized as an LSI which is an integrated circuit device. This LSI may be integrated into a single chip or a plurality of chips. (For example, functional blocks other than memory may be integrated into a single chip.) Here, LSI is used, but it may also be called IC, system LSI, super LSI, or ultra LSI depending on the degree of integration.
The method of making an integrated circuit is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after the LSI is manufactured, or a reconfigurable processor that can reconfigure the connection and settings of circuit cells inside the LSI may be used.
Furthermore, if an integrated circuit technology that replaces an LSI appears due to advances in semiconductor technology or another technology derived from it, it is naturally possible to integrate functional blocks using that technology. There is a possibility of adaptation of biotechnology.
Further, among each functional block, the unit for storing data may not be integrated into one chip, but may have a separate configuration like the recording medium of the present embodiment.
In addition, in each functional block of the block diagram shown in FIGS. 10, 18, 19, 23, 51, etc. and the flowchart shown in FIGS. 9, 11 to 15, 38 to 41, 50, the central part depends on the processor and the program. Is also realized.
As described above, the image coding method or the image decoding method shown in the above-described embodiment can be used for any of the above-mentioned devices / systems, and by doing so, the effects described in the above-described embodiment can be obtained. Obtainable.
The image coding method according to the present invention is a variable-length coding method in which the variable-length coding method applied to the coded data of the moving image is fixed in the continuous reproduction unit indicated by the management information. Since it is possible to eliminate the delay during decoding due to switching and reduce the processing load associated with switching the buffer management method, multiple streams such as MPEG-4 AVC that can switch the variable length coding method within the stream are multiplexed. Suitable for packaged media, etc.
<figref num="1">FIG. 1 is a block diagram of a DVD.</figref><figref num="2">FIG. 2 is a block diagram of highlights.</figref><figref num="3">FIG. 3 is a diagram showing an example of multiplexing on a DVD.</figref><figref num="4">FIG. 4 is a diagram showing an example of a variable length coding method applied to each picture in a conventional MPEG-4 AVC stream.</figref><figref num="5A">FIG. 5A is a block diagram showing a configuration of a decoding device that decodes a coded stream to which CABAC and CAVLC are applied.</figref><figref num="5B">FIG. 5B is a flowchart showing the operation of decoding the coded stream to which CABAC is applied.</figref><figref num="5C">FIG. 5C is a flowchart showing the operation of decoding the coded stream to which CAVLC is applied.</figref><figref num="6">FIG. 6 is a flowchart showing the operation of the conventional decoding device.</figref><figref num="7">FIG. 7 is a diagram showing an example of a variable length coding method applied to each picture in the MPEG-4 AVC stream stored in the information recording medium of the first embodiment.</figref><figref num="8">FIG. 8 is a diagram showing an example of storing flag information indicating a unit in which the variable length coding method is fixed in the information recording medium.</figref><figref num="9">FIG. 9 is a flowchart showing the operation of the decoding device that reproduces the information recording medium.</figref><figref num="10">FIG. 10 is a block diagram showing the configuration of the multiplexing device.</figref><figref num="11">FIG. 11 is a flowchart showing the operation of the multiplexing device.</figref><figref num="12">FIG. 12 is a flowchart showing a specific example of S5201 in FIG.</figref><figref num="13">FIG. 13 is a flowchart showing a specific example of S5202 in FIG.</figref><figref num="14">FIG. 14 is a flowchart showing another specific example of S5203 in FIG.</figref><figref num="15">FIG. 15 is a flowchart showing a specific example of S5204 in FIG.</figref><figref num="16">FIG. 16 is a data hierarchy diagram of HD-DVD.</figref><figref num="17">FIG. 17 is a block diagram of the logical space on the HD-DVD.</figref><figref num="18">FIG. 18 is a schematic block diagram of an HD-DVD player.</figref><figref num="19">FIG. 19 is a block diagram of an HD-DVD player.</figref><figref num="20">FIG. 20 is an explanatory diagram of the application space of HD-DVD.</figref><figref num="21">FIG. 21 is a block diagram of an MPEG stream (VOB).</figref><figref num="22">FIG. 22 is a block diagram of the pack.</figref><figref num="23">FIG. 23 is a diagram illustrating the relationship between the AV stream and the player configuration.</figref><figref num="24">FIG. 24 is a model diagram of continuous supply of AV data to the track buffer.</figref><figref num="25">FIG. 25 is a VOB information file configuration diagram.</figref><figref num="26">FIG. 26 is an explanatory diagram of the time map.</figref><figref num="27">FIG. 27 is a configuration diagram of a playlist file.</figref><figref num="28">FIG. 28 is a configuration diagram of a program file corresponding to a playlist.</figref><figref num="29">FIG. 29 is a configuration diagram of a BD disc overall management information file.</figref><figref num="30">FIG. 30 is a configuration diagram of a file that records a global event handler.</figref><figref num="31">FIG. 31 is a diagram illustrating an example of a time event.</figref><figref num="32">FIG. 32 is a diagram illustrating an example of a user event.</figref><figref num="33">FIG. 33 is a diagram illustrating an example of a global event handler.</figref><figref num="34">FIG. 34 is a configuration diagram of a virtual machine.</figref><figref num="35">FIG. 35 is a diagram of the player variable table.</figref><figref num="36">FIG. 36 is a diagram showing an example of an event handler (time event).</figref><figref num="37">FIG. 37 is a diagram showing an example of an event handler (user event).</figref><figref num="38">FIG. 38 is a flowchart of the basic processing of the player.</figref><figref num="39">FIG. 39 is a flowchart of the playlist reproduction process.</figref><figref num="40">FIG. 40 is a flowchart of event processing.</figref><figref num="41">FIG. 41 is a flowchart of subtitle processing.</figref><figref num="42">FIG. 42 is a diagram illustrating the relationship between the time map of the second embodiment and the still image.</figref><figref num="43">FIG. 43 is a diagram illustrating a flag indicating whether or not the referenced picture can be decoded.</figref><figref num="44">FIG. 44 is a diagram illustrating a flag indicating that all entries refer to an I picture.</figref><figref num="45">FIG. 45 is a diagram illustrating the difference between a video application and a slide show.</figref><figref num="46">FIG. 46 is a diagram illustrating a flag that guarantees that all I-pictures are referenced.</figref><figref num="47">FIG. 47 is a diagram showing a still image data structure in the MPEG-4 AVC of the third embodiment.</figref><figref num="48">FIG. 48 is a diagram illustrating a method of reproducing a still image in MPEG-4 AVC.</figref><figref num="49">FIG. 49 is a diagram illustrating a flag indicating that a specific MinCR value is applied to a clip, and a data structure.</figref><figref num="50">FIG. 50 is a flowchart showing the operation of the multiplexing method according to the fifth embodiment.</figref><figref num="51">FIG. 51 is a block diagram showing the configuration of the multiplexing device.</figref><figref num="52A">FIG. 52A shows an example of the physical format of the flexible disk which is the main body of the recording medium according to the sixth embodiment.</figref><figref num="52B">FIG. 52B shows the front view, cross-sectional structure, and flexible disc of the flexible disc.</figref><figref num="52C">FIG. 52C shows a configuration for recording / reproducing the above program on the flexible disk FD.</figref>
Code description
201 BD disc 202 optical pickup 203 Program recording memory 204 Management information recording memory 205 AV recording memory 206 Program processing unit 207 Management Information Processing Department 208 Presentation processing department 209 image plane 210 video plane 211 Synthesis section 301 Program recording memory 302 program processor 303 UOP Manager 304 Management information recording memory 305 scenario processor 306 Presentation controller 307 clock 308 image memory 309 track buffer 310 demultiplexer 311 image processor 312 video processor 313 sound processor 314 Image plane 315 video plane 316 Synthesis processing unit 317 drive controller 3207 Video Down Converter 3215 Subtitle Down Converter 3223 Still image down converter 3228 Audio down converter S101 disk insertion step S102 BD.INFO reading step S103 BD.PRO G reading step S104 First event generation step S105 event handler execution step S201 UOP reception step S202 UOP event generation step S203 Menu call judgment step S204 event generation step S205 Event Handler Execution Step S301 Playlist playback start step S302 Playlist information (XXX.PL) loading step S303 Playlist program (XXX.PROG) loading step S304 cell regeneration start step S305 AV playback start step S401 AV playback start step S402 VOB Information (YYY.VOBI) Reading Step S403 VOB (YYY.VOB) reading step S404 VOB playback start step S405 VOB playback end step S406 Next cell existence determination step S501 Playlist playback start step S502 Playlist playback end determination step S503 time event time determination step S504 event generation step S505 Event Handler Execution Step S601 Playlist playback start step S602 Playlist playback end determination step S603 UOP reception judgment step S604 UOP event generation step S605 Menu call judgment step S606 User event validity period determination step S607 event generation step S608 Event handler execution step S701 Playlist playback start step S702 Playlist playback end determination step S703 Subtitle drawing start judgment step S704 Subtitle drawing step S705 Subtitle display end judgment step S706 Subtitle Erase Step
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2003204550A | Cites | Japan |
| WO2004049710A1 | Cites | World Intellectual Property Organization (WIPO) |
| WO2004030351A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP20036979A | Cites | Japan |
48 members in 9 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004251870 | Japan | A | |
| 2004251870 | Japan | A | |
| 2004251870 | Japan | – | |
| 2008183938 | Japan | A | |
| 20042004251870 | – | – | – |
| JP20040251870 | – | – | – |
| JP20080183938 | – | – | – |
Members48
| Document | Office | Kind | |
|---|---|---|---|
| WO2006025388A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1791358A1 | European Patent Office (EPO) | A1 | |
| KR20070056059A | Republic of Korea | A | |
| CN101010951A | China | A | |
| US2007274393A1 | United States of America | A1 | |
| JPWO2006025388A1 | Japan | A1 | |
| JP4099512B2 | Japan | B2 | |
| JP2008148337A | Japan | A | |
| EP1791358A4 | European Patent Office (EPO) | A4 | |
| JP2008301512A | Japan | A | |
| JP4201213B2This record | Japan | B2 | |
| JP2010028861A | Japan | A | |
| JP2010028862A | Japan | A | |
| JP2010028863A | Japan | A | |
| US2010027971A1 | United States of America | A1 | |
| CN101010951B | China | B | |
| US7756205B2 | United States of America | B2 | |
| JP4516109B2 | Japan | B2 | |
| CN101800900A | China | A | |
| KR20100092524A | Republic of Korea | A | |
| KR20100092525A | Republic of Korea | A | |
| KR20100092526A | Republic of Korea | A | |
| KR20100092527A | Republic of Korea | A | |
| CN101820544A | China | A | |
| CN101835046A | China | A | |
| CN101841714A | China | A | |
| CN101848389A | China | A | |
| US2010290523A1 | United States of America | A1 | |
| US2010290758A1 | United States of America | A1 | |
| EP1791358B1 | European Patent Office (EPO) | B1 | |
| ATE511314T1 | Austria | T1 | |
| ES2362787T3 | Spain | T3 | |
| EP2346243A1 | European Patent Office (EPO) | A1 | |
| PL1791358T3 | Poland | T3 | |
| JP4813591B2 | Japan | B2 | |
| JP4813592B2 | Japan | B2 | |
| JP4813593B2 | Japan | B2 | |
| US8085851B2 | United States of America | B2 | |
| PL1791358T4 | Poland | T4 | |
| KR101116965B1 | Republic of Korea | B1 | |
| KR101138047B1 | Republic of Korea | B1 | |
| KR101138093B1 | Republic of Korea | B1 | |
| KR101148701B1 | Republic of Korea | B1 | |
| CN101820544B | China | B | |
| CN101841714B | China | B | |
| CN101835046B | China | B | |
| EP2346243B1 | European Patent Office (EPO) | B1 | |
| US8660189B2 | United States of America | B2 |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD |
Numbers
- Publication
- 4201213
- Publication, DOCDB
- 4201213
- Publication, EPODOC
- JP4201213B
- Application
- 183938
- Application, DOCDB
- 2008183938
- Application, EPODOC
- JP20080183938
Titles2
- Japanese
- 動画像符号化方法および装置、動画像復号化方法および装置、記録方法並びに動画像復号化システム
- English
- Moving image coding method and device, moving image decoding method and device, recording method and moving image decoding system
Classification
- CPC, 14
- H04N9/8042
- H04N5/92
- G11B27/3027
- G11B27/329
- G11B2220/2541
- H04N5/85
- H04N9/8063
- H04N9/8205
- H04N9/8227
- H04N19/13
- H04N19/61
- H04N19/179
- H04N19/103
- G11B20/10
- IPC, 18
- H04N7 32
- H04N5 92
- G11B20 12
- G11B27 00
- G11B20 10
- H04N5 85
- H04N5 91
- H04N5 93
- H04N19 00
- H04N19 12
- H04N19 13
- H04N19 169
- H04N19 34
- H04N19 423
- H04N19 44
- H04N19 46
- H04N19 70
- H04N19 91