Moving picture stream generation apparatus, moving picture coding apparatus, moving picture multiplexing apparatus and moving picture decoding apparatus
Abstract
This record has no abstract on file.
Term
Term ended
Projected expiry passed 25 April 2025, 1.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1Patent claims Zastrzeżenia patentowe 1. An apparatus for generating a moving image stream, generating a stream containing images, having at least a supplementary information memory unit and a pixel information memory unit, the device comprising:1. Urządzenie do generowania strumienia ruchomego obrazu, generujące strumień zawierający obrazy, dysponujące co najmniej jednostką pamięci informacji uzupełniających oraz jednostką pamięci informacji pikselowych, przy czym urządzenie to tworzą: an auxiliary information generating unit, generating auxiliary information regarding all images contained in each random access unit, in which the internally encoded image with the possibility of decoding independently of other images is placed at the top in the decoding order, the supplementary information containing a lot of pieces of information about the type images indicating the types of all images in the random access unit, wherein many pieces of information about the type of image are placed in the order of image decoding as identification information to identify the images to be decoded during image reproduction using the trickplay function;and a stream generating unit generating the stream by writing the supplementary information to the supplementary information memory unit of the image placed on top in the respective random access unit. jednostka generowania informacji pomocniczych, generująca informacje pomocnicze dotyczące wszystkich obrazów zawartych w każdej jednostce o swobodnym dostępie, w której obraz kodowany wewnętrznie z możliwością dekodowania niezależnie od innych obrazów jest umieszczony u góry w kolejności dekodowania, przy czym informacje uzupełniające zawierają wiele fragmentów informacji na temat typu obrazów wskazujących typy wszystkich obrazów, jakie znajdują się w jednostce o swobodnym dostępie, przy czym wiele fragmentów informacji na temat typu obrazu jest umieszczanych w kolejności dekodowania obrazów jako informacje identyfikacyjne do identyfikacji obrazów do zdekodowania w trakcie odtwarzania obrazów z zastosowaniem funkcji trickplay;oraz jednostka generowania strumienia generuj ąca strumień poprzez zapisanie informacji uzupełniających do jednostki pamięci informacji uzupełniających obrazu umieszczonego na górze w odpowiedniej jednostce o swobodnym dostępie. 2. A method of generating a moving image stream that allows generating a stream containing images having at least a supplementary information memory unit and a pixel information memory unit, the method comprising: 2. Sposób generowania strumienia ruchomego obrazu umożliwiaj ący generowanie strumienia zawierającego obrazy dysponujące co najmniej jednostką pamięci informacji uzupełniających oraz jednostką pamięci informacji pikselowych, przy czym sposób ten obejmuje: etap generowania informacji pomocniczych, podczas którego generuje się informacje pomocnicze dotyczące wszystkich obrazów zawartych w każdej jednostce o swobodnym dostępie, w której obraz kodowany wewnętrznie z możliwością dekodowania niezależnie od innych obrazów umieszcza się na górze w kolejności dekodowania, informacje uzupełniające zawierające wiele fragmentów informacji na temat typu obrazów wskazujących typy wszystkich obrazów, jakie znajdują się w jednostce o swobodnym dostępie, przy czym wiele fragmentów informacji na temat typu obrazu umieszcza się w kolejności dekodowania obrazów jako informacje identyfikacyjne do identyfikacji obrazów do zdekodowania w trakcie odtwarzania obrazów z zastosowaniem funkcji trick-play;oraz etap generowania strumienia, podczas którego generuje się strumień poprzez zapisanie informacji uzupełniających do jednostki pamięci informacji uzupełniających obrazu znajdującego się na górze w odpowiedniej jednostce o swobodnym dostępie. the stage of generating auxiliary information, during which auxiliary information is generated regarding all images contained in each random access unit, in which an internally encoded image with the possibility of decoding independently of other images is placed on top in the decoding order, supplementary information containing many pieces of information about image type indicating the types of all images in the random access unit, wherein many pieces of information about the type of image are placed in the order of image decoding as identification information for identifying the images to be decoded during image reproduction using the trick-play function;and a stream generating step during which the stream is generated by writing the supplemental information to the image supplemental information storage unit located at the top in the respective random access unit. 3. A moving image decoding device that decodes a stream containing encoded images that form a moving image, and reproduces a decoded stream, the moving image decoding device comprising: 3. Urządzenie do dekodowania ruchomego obrazu, które dekoduje strumień zawierający zakodowane obrazy, które tworzą ruchomy obraz, oraz odtwarza zdekodowany strumień, przy czym urządzenie do dekodowania ruchomego obrazu obejmuje: an instruction obtaining unit enabling instructions to be obtained that the trick-play function should be performed;jednostkę uzyskiwania instrukcji umożliwiającą uzyskiwanie instrukcji wskazujących, że powinna zostać przeprowadzona funkcja trick-play;an analysis unit that allows to obtain and analyze supplementary information based on a given random access unit, the random access unit forming a stream;jednostkę analizy pozwalająca uzyskać i przeanalizować informacje uzupełniające na podstawie danej jednostki o swobodnym dostępie, przy czym jednostka o swobodnym dostępie tworzy strumień;a reproduction image determining unit enabling the determination of images from among the images contained in the random access unit that are necessary to perform the trick-play function indicated by the instruction obtained by the instruction obtaining unit based on the results of the analysis carried out by the analysis unit;and a decoding unit enabling the decoding of images determined by the image determining unit for reproduction and playback of the decoded images, wherein in the random access unit the image internally encoded with the possibility of decoding independently of other images is placed on top, and the supplementary information contains a lot of pieces of information about the type of images indicating the types of all images in the unit with free access, wherein many pieces of information about the type of image are placed in the order of image decoding as identification information to identify the images to be decoded during image reproduction using the trick-play function. jednostkę określania obrazów do odtwarzania umożliwiającą określenie obrazów, spośród obrazów zawartych w jednostce o swobodnym dostępie, które są niezbędne do przeprowadzenia funkcji trick-play wskazanej przez instrukcję uzyskaną przez jednostkę uzyskiwania instrukcji na podstawie wyników analizy przeprowadzonej przez jednostkę analizy;oraz jednostkę dekodującą umożliwiającą dekodowanie obrazów określonych przez jednostkę określania obrazów do odtwarzania oraz odtwarzanie zdekodowanych obrazów, przy czym w jednostce swobodnego dostępu obraz kodowany wewnętrznie z możliwością dekodowania niezależnie od innych obrazów jest umieszczony na górze, a informacje uzupełniające zawierają wiele fragmentów informacji na temat typu obrazów wskazujących typy wszystkich obrazów, jakie znajdują się w jednostce o swobodnym dostępie, przy czym wiele fragmentów informacji na temat typu obrazu umieszczanych jest w kolejności dekodowania obrazów jako informacje identyfikacyjne do identyfikacji obrazów do zdekodowania w trakcie odtwarzania obrazów z zastosowaniem funkcji trick-play. 4. A method of decoding a moving image, enabling decoding a stream containing images that form a moving image, and reproducing a decoded stream, the method for decoding a moving image includes: 4. Sposób dekodowania ruchomego obrazu, umożliwiający dekodowanie strumienia zawierającego obrazy, które tworzą ruchomy obraz, oraz odtwarzanie zdekodowanego strumienia, przy czym sposób do dekodowania ruchomego obrazu obejmuje: a step of obtaining instructions to obtain instructions indicating that the trick-play function should be performed;etap uzyskiwania instrukcji umożliwiający uzyskiwanie instrukcji wskazujących, że należy przeprowadzić funkcję trick-play;an analysis step to obtain and analyze supplementary information based on a given random access entity, the random access entity forming a stream;etap analizy pozwalający uzyskać i przeanalizować informacje uzupełniające na podstawie danej jednostki o swobodnym dostępie, przy czym jednostka o swobodnym dostępie tworzy strumień;etap określania obrazów do odtwarzania umożliwiający określenie obrazów, spośród obrazów zawartych w jednostce o swobodnym dostępie, które są niezbędne do przeprowadzenia funkcji trick-play wskazanej przez instrukcję uzyskaną we wspomnianym etapie uzyskiwania instrukcji na podstawie wyników analizy przeprowadzonej w etapie analizy;oraz etap dekodowania umożliwiający dekodowanie obrazów określonych w etapie określania obrazów do odtwarzania oraz odtwarzanie zdekodowanych obrazów, przy czym w jednostce swobodnego dostępu obraz kodowany wewnętrznie z możliwością dekodowania niezależnie od innych obrazów umieszcza się na górze, a informacje uzupełniające zawierają wiele fragmentów informacji na temat typu obrazów wskazujących typy wszystkich obrazów, jakie znajdują się w jednostce o swobodnym dostępie, przy czym wiele fragmentów informacji na temat typu obrazu umieszcza się w kolejności dekodowania obrazów jako informacje identyfikacyjne do identyfikacji obrazów do zdekodowania w trakcie odtwarzania obrazów z zastosowaniem funkcji trick-play. the step of determining the images to be reproduced enabling the images to be determined from among the images contained in the random access unit that are necessary to carry out the trick-play function indicated by the instruction obtained in said step of obtaining instructions based on the results of the analysis carried out in the analysis step;and a decoding step enabling the decoding of the images specified in the step of determining the images to be reproduced and reproducing the decoded images, wherein in the random access unit the image encoded with the possibility of decoding independently of other images is placed on top, and the supplementary information contains a lot of pieces of information about the type of images indicating the types of all images in the unit with free access, wherein many pieces of information about the type of image are placed in the order of image decoding as identification information to identify the images to be decoded during image reproduction using the trick-play function. 5. A recording medium with the possibility of reading by means of a computer on which the stream is recorded, the stream containing images that form a moving image, the stream being constructed in such a way that supplementary information is added to the image placed on top of the unit with free access, the supplementary information contains many pieces of information about the type of images indicating the types of all images, which are found in the random access unit, with a plurality of pieces of information about the type of image being placed in the order of image decoding as identification information to identify the images to be decoded during image reproduction using the trick-play function. 5. Nośnik zapisujący z możliwością odczytu za pomocą komputera, na którym zapisywany jest strumień, przy czym strumień zawiera obrazy, które tworzą ruchomy obraz, przy czym strumień jest skonstruowany w taki sposób, że informacje uzupełniające są dodawane do obrazu umieszczonego na górze jednostki o swobodnym dostępie, przy czym informacje uzupełniające zawierają wiele fragmentów informacji na temat typu obrazów wskazujących typy wszystkich obrazów, jakie znajdują się w jednostce o swobodnym dostępie, przy czym wiele fragmentów informacji na temat typu obrazu jest umieszczanych w kolejności dekodowania obrazów jako informacje identyfikacyjne do identyfikacji obrazów do zdekodowania w trakcie odtwarzania obrazów z zastosowaniem funkcji trick-play. 6. A recording method that allows recording a stream containing images that form a moving image on a recording medium, the recording method comprising the step of recording a stream generated using the method of generating the image stream according to claim 1. 2. 6. Sposób zapisu umożliwiający zapisywanie strumienia zawierającego obrazy, które tworzą ruchomy obraz na nośniku zapisującym, przy czym sposób zapisu składa się z etapu zapisywania strumienia generowanego przy użyciu sposobu generowania strumienia obrazu zgodnie z zastrz. 2. 7. A moving image decoding system consisting of a recording medium according to claim 1. And a moving image decoding device according to claim 1. 3, which reads the stream containing the images that form the moving image, from the recording medium, decodes and plays the stream. 7. System dekodowania ruchomego obrazu składający się z nośnika zapisującego zgodnie z zastrz. 5 oraz urządzenia dekodującego ruchomy obraz zgodnie z zastrz. 3, który odczytuje strumień zawierający obrazy, które tworzą ruchomy obraz, z nośnika zapisującego, dekoduje i odtwarza strumień. Panasonic Corporation Panasonic Corporation Pełnomocnik: Proxy: FIG. ΙΑ <- FIG. ΙΑ <— GOP GOP GOP //////////////// GOP //////////////// FIG, IB FIG, IB Stream Strumień GOP obraz GOP picture FIG. 3A stream FIG. 3A strumień RAU RAU RAU RAU RAU RAU AU AU AU AU AU AU Slice Slice WALU frame information start code frame (0x000001) frame WALU ramki informacje kod startowy ramki (0x000001) ramki Slice Slice NALU frame —T frame information NALU ramki —T informacje ramki FIG. 3B FIG. 3B NALU NALU NALU 1 heading NALU 1 nagłówek 7 "NAL unit type data block 7“ typ jednostki NAL blok danych FIG. 8A FIG. 8A Γ " Γ” RAU RAU A AND RAU RAU AU AU NALU trick-play information informacje trick-play NALU AUD AUD SHAFT WALU NAL unit jednostka NAL FIG. 8B informacje trick-play komunikatu SEI FIG. 8B trick-play information of the SEI message \ \ plaster \ \ plaster NALU patch NALU plaster NALU patch NALU plaster NALU informacje trick-play komunikatu SEI NALU trick-play information of the SEI message FIG. 9A display order FIG. 9A kolejność wyświetlania FIG. 9B ..... hau_;, FIG. 9B ..... hau_;, FIG. 10A display order FIG. 10A kolejność wyświetlania FIG. 10B _rau__ FIG. 10B _rau__ <-num_picJn ~ RAU nu m_speed <-num_picJn~RAU n u m_speed Variable Speed Piay { nufn_picJn_RAU;num_speed;Variable Speed Piay {nufn_picJn_RAU;num_speed;For (i = 0;i <num_speed;1 ++) {play_speed;num "dec_pic;For (i=0;i < num_speed;1++) { play_speed;num„dec_pic;for (j ~ 0;j <num "dec_pic;j ++) {dec_pic;for (j~0;j < num„dec_pic;j++) { dec_pic;> > > > > > przykład składni syntax example FIG, 13A danych FIG, 13A data FIG. 13B FIG. 13B FIG. 14 FIG. 14 Variable Speed Play { num_pic_in_RAU;num_speed;Variable Speed Play {num_pic_in_RAU;num_speed;for 0 = 0;and <num__speed;i ++) {play_speed;num_dec_pic;pts_dts_flag;for 0=0;i < num__speed;i++) { play_speed;num_dec_pic;pts_dts_flag;for (j = 0;j <num_dec_pic;j ++) {dec_pic;for (j=0;j < num_dec_pic;j++) { dec_pic;if (pts_dts_flag) diplay_order;if (pts_dts_flag) diplay_order;> > } } FIG. 15A display order FIG. 15A kolejność wyświetlania FIG. 15B FIG. 15B decoding order kolejność dekodowania FIG. 17C FIG. 17C FIG. 17A display order FIG. 17A kolejność wyświetlania FIG. 17B FIG. 17B RAU RAU RAU map < nurtt_AUJn_RAU;RAU map <nurtt_AUJn_RAU;for (1 = 0;I <num_AU_ln_RAU;1 ++) {frame_fleld_flag;plcjype;for (1=0;I < num_AU_ln_RAU;1++) { frame_fleld_flag;plcjype;> > > > FIG, 17F FIG, 17F FIG. 17D FIG. 17D RAU map < num_frame Jn_RAU;for (1=0;I < num_frame_ln_RAU;I++) { ftame_flag;RAU map <num_frame Jn_RAU;for (1 = 0;I <num_frame_ln_RAU;I ++) {ftame_flag;if (frame_fleld_flag) frame_type;else field_pair_type;if (frame_fleld_flag) frame_type;else field_pair_type;> > > > <-I4 <-P5 <-Bl «-Β2 Hi 3 <-I4 <-P5 <-Bl «-Β2 Hi 3 Ή8 ^ B6 ^ -B7 «-Ρ13« -Ρ14 k-B9 f-BlD <-BU «-Β12 Ή8 ^B6 ^-B7 «-Ρ13 «-Ρ14 k-B9 f-BlD <-BU «-Β12 H> 19 «-Ρ20 H>19 «-Ρ20 -B9, B1O -B9,B1O FIG. 17E FIG. 17E FIG. 18A FIG. 18A RAU map { num_AU_ln_RAU;RAU map {num_AU_ln_RAU;for 0 = 0;and <num_AU_in_RAU;i ++) <picture_structure;picture_type;for 0=0;i < num_AU_in_RAU;i++) < picture_structure;picture_type;> > > > FIG. 18B box or frame FIG. 18B pole lub ramka FIG. 18C picture_type: obraz 1 lub obraz odniesienia B lub obraz B bez odniesienia lub obraz P kolejność wyświetlania FIG. 18C picture_type:picture 1 or reference image B or image B without reference or image P display order FIG. 20A FIG. 20A FIG. 22 FIG. 22 Nie No JL JL END KONIEC FIG. 24A HLP support information FIG. 24A informacje pomocnicze HLP - ^ presence / no structure with free access. presence / no prediction structure> presence / no special information presence / no information -> indicating units -^obecność/brak struktury o swobodnym dostępie .obecność/brak struktury predykcji >obecność/brak informacji specjalnych obecność/brak informacji -> wskazujących jednostki AU do zdekodowania obecność/brak informacji -> wskazujących jednostki AU to be decoded presence / lack of information -> indicating units AU do wyświetlenia AU to display FIG. 24B HLP support information FIG. 24B informacje pomocnicze HLP -> Have the BD-ROM requirements been met? ->Czy spełniono wymagania BD-ROM? Czy spełniono wymagania HD-DVD? Have you met the HD-DVD requirements? kontrola zarządzania management control FIG. 25 FIG. 25 JJ stream is a random access unit JJ strumień jednostka swobodnego dostępu FIG. 26 free access unit ζ START FIG. 26 jednostka swobodnego dostępu ζ START S51 określanie informacji dotyczących atrybutów TYPE .S52 kodowanie strumienia ,S53 generowanie informacji pomocniczych HLP .S54 generowanie informacji zarządzania INFO ,S5$ multipleksowanie zapisywanie S51 specifying information about TYPE attributes. S52 stream coding, S53 generating auxiliary information HLP. S54 generating management information INFO, S5 $ multiplexing saving -SS6 and -SS6 i ( KONIEC j (END j FT (7 77 FT(7 77 FIG. 28 FIG. 28 Nie r * KONIEC J No r * END J FIG. 29 startJ) FIG. 29 startJ) Nie wyszukiwanie prefiksu kodu startowego No start code prefix search S31 S31 Czy jednostka NAL przechowuje informacje trickplay? Does the NAL unit store trickplay information? S30 S30 Tak Yes Nie No FIG. 31 lll FIG. 31 th 112 112 113 113 114 114 FIG. 32 FIG. 32 BD.IMFO BD.IMFO BD.PROG PROG XXX. PROG XXX. EN XXX. PROG XXX. PL ΥΫΥ.νΟΒΙ ΥΫΥ.νΟΒΙ ΥΎΥ.νΟΒ ΥΎΥ.νΟΒ ZZZ.PNG ZZZ.PNG FIG. 33 FIG. 33 Coding Coding Resolution Resolution Video video Number Number Aspect Aspect Framerate framerate Audio # 0 Audio#0 ΥΥΥ.νΟΒΙ ΥΥΥ.νΟΒΙ Γ. Γ. VOB YYY.VOB Attribute Attribute ΤΜΑΡ ΤΜΑΡ Audio # rrt »\ Audio#rrt » \ AND I «I Coding Coding Ch. Ch. Lang. Lang. FIG. 34 FIG. 34 MPEG-TS (VOB) MPEG-TS(VOB) PTS # k PTS#k I end # k I end#k And "start # k I„start#k FIG. 36 function Eventl () { FIG. 36 function Eventl() { 4 > 4 > FIG. 37 FIG. 37 ± ± Type Type ID ID BD.PROG _Z BD.PROG _Z FIG. 38 FIG. 38 FIG. 39 FIG. 39
294 paragraphs in 1 section, as filed
Technical field The present invention relates to a device and the like which generates an encoded moving image stream, and in particular to a device and the like which generates a stream on which trickplay mode can be used, including jump-in playback, variable playback speed, reverse play and more.
Background Art [0002] Not so long ago, the multimedia era has begun, which means that nowadays sound, image and other pixel values are integrated into obtaining one medium, and traditional information media used as communication tools such as newspapers, magazines, television, radio and the phone are treated as multimedia targets. Basically, multimedia is a form of simultaneous presentation of not only characters, but also graphics, sound, and especially photos. To process the traditional information media described above to obtain a multimedia form, information should be provided in digital form.
[0003] However, it is impossible to directly digitally process the massive amount of information using the traditional information media described above, because, if you convert the amount of data for each of the information media described above with the amount of digital data, the amount of data per character is from 1 to 2 bytes, while when a second of sound takes at least 64 kb (telephone sound quality), and a second of moving image takes at least 100 Mb (the quality of the current TV signal). For example, TV telephones have become available in practice thanks to ISDN (Integrated Services Digital Network) technology, providing transmission from 64 kbps to
1.5 Mb / s, however it is not possible to transfer moving images from the television camera in its present form using
ISDN.
[0004] That is why it becomes necessary to use compression techniques. For example, the H.261 or H.263 moving image compression technique, which is recommended by the ITU-T (International Telecommunication Union-Telecommunication) standardization sector, is used for TV telephones. Moreover, the information compression technique in the MPEG-1 standard allows you to store image information along with audio information on a standard CD (Compact Disc) for recording music.
[0005] MPEG (Moving Picture Experts Group) is an international standard for digital compression of moving image signals that has been developed by the ISO / IEC (International Standardization Organization / International Engineering Consortium). MPEG-1 is a standard for compressing moving image signals up to 1.5 Mb / s, i.e. compressing the TV signal by about 100 times. The quality, which meets the requirements of the MPEG-1 standard, is at an average level, which can be achieved at a speed of about 1.5 Mb / s. Therefore, the MPEG-2 standard was developed to provide higher image quality and compression of moving image signals from 2 to 15 Mb / s. Currently, the ISO / IEC JTC1 / SC29 / WG11 working group, which has developed the MPEG-1 and MPEG-2 standards, has created the MPEG-4 standard with a higher compression ratio. The MPEG-4 (i) standard achieves a higher compression ratio than that offered by the MPEG-1 and MPEG-2 standards, ii) enables coding, decoding and object-for-object operations, and (iii performs the new functions required for the multimedia age. The initial goal of creating the MPEG-4 standard was to standardize the method of encoding images with a low data rate, but later the goal was extended to the general-purpose coding method for interlaced image with high transmission speeds. Then, ISO / IEC and ITU-T jointly developed the MPEG-4 AVC (Advanced Video Coding) standard as a way of coding a new generation image with a high compression ratio. It was intended to be created for devices related to new generation optical discs or for transmitting on portable terminals.
[0006] Basically, when coding a moving image, information is compressed by reducing time and spatial redundancy. During cross-picture predictive coding, which aims to reduce time redundancy, motion estimation, and predictive image creation, block-by-block is performed taking into account the next and previous image, and coding is performed at a different value between the obtained predictive image and the image to be encoded. In this case, the term "image" defines one image. For a progressive image, the image means a frame, but for interlaced image it can be a frame or field. The term "interlaced image" means a frame composed of two fields with a slight time delay. During the process of encoding and decoding interlaced images, it is possible to process the frame as such, as two fields or frame-by-frame or field-by-field for each block in a frame.
[0007] An image for performing internal predictive coding without taking into account any reference image is called an I (Intra Coded Picture) image. In turn, an image that is used to perform inter-image predictive coding relating to only one image is called a P (Predictive Coded Picture) image ). In turn, the image that is used to perform inter-picture predictive coding, referring to two reference images at the same time is called the image B (Bi-predictive Coded Picture). Image B can refer to two images selected as any combination of the next and previous image in the display order. These two reference images can be determined block-by-block, with the block being the basic unit of coding and decoding. These reference images are distinguished from others in the following ways: the reference image previously described in the coded bit stream is called the first reference image, and the other reference image described later is called the second reference image. Note that such reference images must already be encoded or decoded to encode or decode P images and B images.
[0008] Motion compensation in inter-picture coding is used to encode P and B images. Intra-picture predictive coding with motion compensation is a method of intra-picture predictive coding that uses motion compensation. Motion compensation is a method of increasing the precision of prediction and reducing the amount of data by estimating the amount of motion (hereinafter referred to as the motion vector) of each image block and by performing predictive coding taking into account the motion vector. For example, the amount of data is reduced by estimating the motion vectors of the images to be encoded and by encoding each predictive residue. residual) between each predictive value that changes by the size of each motion vector and each current image to be encoded. With this method, as traffic vector information is required in the decoding process, traffic vectors are also encoded and saved or transmitted.
[0009] Motion vectors are estimated by the macroblock-macroblock method. In particular, motion vectors are estimated by determining the image macroblock to be encoded, shifting the reference image macroblock within the search range, and finding the location of the reference block that is closest to the standard block.
[0010] Figures 1A and 1B show structural diagrams of typical MPEG-2 streams. As shown in Fig. 1B, the MPEG-2 stream has a hierarchical structure, which will be described below. The stream consists of a group of images (GOP - Group of Pictures). Using GOP as the basic unit in the coding process allows you to edit a moving image or gain free access. GOP consists of I, P and B images. The stream, GOP and image further comprise a synchronization signal (sync) indicating a frame of units and a header indicating data in units, with the units being stream, GOP and image, respectively.
[0011] Figs. 2A and 2B are respectively examples showing how inter-picture predictive coding is performed which is used in the MPEG-2 standard. Diagonally hatched images in the drawing mean those images to which other images refer. As shown in fig. 2A, during MPEG-2 predictive coding, P images (P0, P6, P9, P12 and P15) can only refer to a single image selected as the I or P image immediately preceding it in the display order. In addition, B images (B1 .
B2, B4, B5, B7, B8, B10, B11, B13, B14, B16, B17, B19 and
B20) can refer to two images selected as a combination of the I image or the P image immediately before the image and the P image immediately after the image. The order of the images in the stream is determined. The images I and the image P are placed in the order of display, and each image B is placed immediately after the image I, which is placed immediately after the image B or directly after the image P. In shown in Fig. 2B, GOP structural example, I3-B14 images are grouped into a single GOP block.
<td> [0012]</td><td>FIG.</td><td>3A</td>
<td>stream</td><td>MPEG-4</td><td>AVC</td>
<td>standard</td><td>MPEG-4</td><td>AVC.</td>
There is no GOP equivalent in constructing a random access unit that is the equivalent of a GOP by segmenting data based on a special image that can be decoded independently of other images, the unit will now be called RAU (Random Access Unit). In other words, a random access unit RAU is an encoded group of images that begins with an internally encoded image that can be decoded independently of any other image.
[0013] The access unit, which is the basic unit in serving the stream (hereinafter referred to as AU), will be described below. AU is a unit for storing coded data equivalent to one image and includes a set of PS parameters, slice data and the like. There are two types of PS parameter sets. One of them is a set of PPS image parameters (henceforth called PPS), which is equivalent to the header data of each image. Another is the set of SPS sequence parameters (henceforth called SPS), which is equivalent to the header that is part of the GOP unit or more images in the MPEG-2 standard. The SPS contains the maximum number of reference images, image size and the like. On the other hand, PPS includes a variable length coding type, initial quantization step value, number of reference images, and the like. Each image is assigned an identifier indicating which of the sets described above (PPS and SPS) is associated with it. In addition, the frame number (FN), which is an identification number used to identify the image contained in the patch data. Note that the sequence begins with a special image, in which all statuses necessary for decoding are zeroed as described below, where it consists of a group of images that begins with a special image and ends with an image that is placed immediately before the next special picture.
[0014] There are two types of I pictures in the MPEG-4 AVC standard. They are IDR (Instantaneous Decoder Refresh) and other images. An IDR image is an I image that can decode all images placed after an IDR image in decoding order without being associated with images placed before an IDR image in decoding order. In other words, it is an I image for which the decoding statuses are zeroed. The IDR image corresponds to the upper image and closed GOP in MPEG-2. The MPEG-4 AVC sequence begins with the IDR image. In the case of an I image that is not an IDR image, the image placed after the image I in the decoding order may refer to the image placed before the image in the decoding order. The respective image types will be defined below. The IDR image and the I image are images that consist exclusively of I slices. The P image is an image that can consist of P slices and I slices. Image B is an image that can consist of B slices, P slices and I slices. Note that IDR image slices are stored in a NAL unit whose type is different from that of the NAL unit in which the image slices other than IDR. In this case, the NAL unit is a sub-picture unit.
[0015] The AU unit in the MPEG-4 AVC standard may contain not only data necessary for decoding, but also supplementary information and frame information.
AU.
This supplementary information is called
SEI (Supplemental Enhancement Information) and they are not necessary for decoding slice data. All data such as PS parameter set, patch data, SEI are stored in the NAL (Network Abstraction Layer) unit, which is called NALU. The NAL unit consists of a header and a data block. The header contains a field denoting the type of data to be stored (hereinafter referred to as NAL unit type). NAL unit type values are defined for data types such as slice or SEI, respectively. By referencing such a NAL unit type value, you can identify the type of data to be stored in the NAL unit. The NAL unit header contains a field named nal_ref_idc. The nal_ref_idc field is a 2-bit field and takes the value 0, 1 or higher depending on the NAL unit types. For example, the NAL unit of the SPS or PPS set corresponds to a value of 1 or greater. For the NAL unit of the patch, the slice to which other segments refer is set to 1 or higher, while the slice to which other slices do not apply is set to 0. In addition, the NI unit of the SEI information always sets to 0.
[0016] One or more SEI messages may be stored in the NAL unit of SEI information. The SEI message consists of a header and a data block, and the type of information to be stored in the data block is identified based on the type of SEI message indicated in the header. AU decoding now means decoding slice data in an AU and AU display means displaying the result of decoding slice data in a unit
AU.
[0017] Since the NAL unit does not contain frame identification information for the NAL unit, it is possible to add frame information at the top of each NAL unit when saving the NAL unit as an AU. When processing the MPEG-4 AVC stream in the MPEG-2 transport stream (TS) or in the MPEG-2 program stream (PS), the start code prefix shown as 3 bytes 0x000001 is added to the top of the NAL unit. It was further defined that the NAL unit indicating the AU frame should be placed on top of the AU in the TS or PS MPEG-2 stream, this AU being referred to as the Access Unit Delimiter.
[0018] Various techniques have been proposed to date for coding a moving image similar to this (for example in Patent Document 1).
[0019] Patent Document 1: Japanese Patent Publication Open to Public, No. 2003-18549.
[0020] Fig. 4 is a block diagram of a typical moving image coding apparatus.
[0021] The device for the coding of a moving image 1 is a device that outputs an encoded stream Str obtained by conversion, by coding combined with compression, wherein the video signal Vin at the input is input into a bit stream with a variable length coding type or the like . The moving image coding device has a PTYPE prediction structure determination unit, ME motion vector estimation unit, MC motion compensation unit, Sub subtraction unit, Orthogonal transformation unit T, Q quantization unit, IQ inverse quantization unit, IT inverse orthogonal transformation unit, Add unit , PicMem image memory, switch, and VLC variable length coding unit.
[0022] The video input signal Vin is input to the subtraction unit Sub and to the motion vector estimation unit ME. The Sub subtraction unit calculates the difference between the input video input signal Vin and the predictive image and outputs the result to the input of the orthogonal transformation unit. The orthogonal transformation unit T converts the difference value into a frequency factor and sends it to the quantization unit Q. The Q quantization unit performs quantization on the entered frequency factor and generates the Qcoef quantization value for the variable length coding unit.
[0023] The inverse quantization unit IQ performs an inverse quantization on the Qcoef quantization value to reconstruct the frequency factor and introduces the result to the inverse IT orthogonal transformation unit. The inverse orthogonal IT transformation unit performs the frequency inverse frequency transformation of the frequency ratio into a pixel differential value and enters the result into the addition unit Add. The Add unit adds a pixel differential value to the predictive image that is generated by the MC motion compensation unit to create a decoded image. The SW switch is in the ON position when recording a coded image is requested. The encoded image is then saved to the PicMem image memory.
[0024] On the other hand, the ME motion vector estimation unit into which the video input signal is input
Vin macroblock-by-macroblock, looks for a decoded image stored in the PicMem image memory and estimates the image area that is closest to the image input signal, and consequently determines the MV motion vector indicating the position. The estimation of the motion vector is carried out block-block, the block being part of a macroblock. As many images can now be used as reference images, it is necessary to use identification numbers for specific images (relative indexes) block-by-block. This makes it possible to specify reference images by calculating the image numbers indicated by relative indexes, with the image numbers assigned to the respective images in the PicMem image memory.
[0025] The motion compensation unit MC selects the image area that is optimal as a predictive image from the decoded images stored in the PicMem image memory.
[0026] The PTYPE prediction structure determination unit instructs the ME motion vector estimation unit and the MC motion compensation unit to perform in-image coding on the target image as a freely available special image using its Ptype image type, in the event that the initial image of the random access unit RAUin indicates that the random access unit RAU begins with the current image, and further instructs the VLC variable coding unit to encode the Ptype image type.
[0027] The VLC variable coding unit performs variable length coding on the Qcoef quantization value, relative index, Ptype image type, and MV motion vector to obtain the coded stream Str.
[0028] Fig. 5 shows a block diagram of a typical moving image decoding device 2. This moving image decoding device 2 has a variable length VLD decoding unit, PicMem image memory, MC motion compensation unit, Add addition unit, inverse IT orthogonal transformation unit and the inverse quantization unit IQ. It should be noted that in the drawing, those processing units that perform the same operations as the processing units in the typical moving image coding apparatus shown in Fig. 4 have been assigned the same reference numbers and their descriptions will be omitted.
[0029] The variable length decoding unit VLD decodes the encoded stream Str and generates a Qcoef quantization value, relative index, Ptype image type, and MV motion vector. The Qcoef quantization value, Relative Index index, and MV motion vector are entered into the PicMem image memory, the MC motion compensation unit and the inverse IQ quantization unit, respectively, and then the decoding process is performed on them. Such operations of a typical moving image coding apparatus have already been described using the block diagram shown in Fig. 4.
[0030] The random access unit RAU indicates that decoding can be performed starting from the top AU in the random access unit. However, because a typical MPEG-4 AVC stream allows very flexible prediction structures, a storage device having an optical disk or hard disk cannot obtain information to determine AU access units to be decoded or displayed during variable speed playback or reverse playback.
[0031] Figs. 6A and 6B show examples of AU access unit prediction structures. The image is stored in each AU. Fig. 6A shows the prediction structure of AUs used in the MPEG-2 stream. Diagonally hatched images in the drawing are images that other AUs may reference. For the MPEG-2 stream, the P image AUs (P4 and P7) may perform predictive coding by referring only to the single AU selected as the AU of the image immediately before the I image or the P image in display order. In addition, AUs of B images (B1, B2, B3, B5 and B6) can perform predictive coding only on two AUs selected as a combination of AUs of the image I or the P image immediately before the image and the image I or the P image immediately after the image in order of display. The order of the images in the stream is determined as follows: AUs of the I image and the P images are arranged in order of display; each of the AUs of the B pictures is placed immediately after the AUs of the I picture or one of the P pictures that is placed immediately after the AU of each B picture. As a result, decoding can be performed in three ways: (1) all images are decoded; (2) only AUs of the I picture and P pictures are decoded and displayed; (3) only the AU of the I picture is decoded and displayed. Therefore, the following three types of playback can easily be performed using: (1) standard playback, (2) medium speed playback and (3) high speed playback.
[0032] For the MPEG-4 AVC stream, it is possible to perform a prediction in which the AU of the B image refers to the AU of the B image. Fig. 6B shows an example of the prediction structure in the MPEG-4 AVC stream and the AU of the B images (B1 and B3) refer to unit AU (B2) of image B. In this example, the following four types of decoding can be implemented: (1) all images are decoded; (2) only AUs of the I picture, P pictures and B pictures referenced are decoded and displayed; (3) only AUs of the I picture and P pictures are decoded and displayed; (4) only the AU of picture I is decoded and displayed.
[0033] Also, for the MPEG-4 AVC stream, the P image AU may refer to the B image AU. As shown in Fig. 7, the P image P AU (P7) may refer to the B (B2) image AU. In this case, the AU of the P picture (P7) can be decoded only after the AU of the B (B2) picture has been decoded. Therefore, it is possible to implement the following three types of decoding or display: (1) all images are decoded; (2) only AUs of the I picture, P pictures and B pictures referenced are decoded and displayed; (3) only the AU of the I picture is decoded and displayed.
[0034] In this way, as different prediction structures are allowed in the MPEG-4 AVC standard, data segment analysis and prediction structure assessment must be performed to know the reference relationship between AUs. This has the problem that the AUs to be decoded or displayed cannot be determined based on a rule that is predetermined depending on the playback speed during jump-in, variable-speed playback and reverse playback , unlike the MPEG-2 format. D1 discloses a device for generating a moving image stream in which information about the position of I and P frames required to use trick-play modes is recorded in the input sector created at the top of each GOP.
<td>AVC</td><td>which</td>
<td>and</td><td>(Ii)</td>
<td>and</td><td>ago</td>
Disclosure of the Invention [0035] The object of the present invention is to provide (i) a device for generating a moving picture stream and a method for generating a moving picture stream that allows performing trick-play effects such as jump-in playback, variable speed playback and reverse playback, even in the case of a coding method similar to that offered by the MPEG-4 AVC standard, which allows flexible prediction structures, apparatus for decoding a moving picture and the like which decodes such a moving picture stream.
[0036] The invention is a device for generating a moving picture stream according to claim 1, a method of generating a moving image stream according to claim 1. 2, a moving image decoding device according to claim 1. 3, the method of decoding a moving image according to claim 1. 4, a recordable medium with the ability to be read by a computer according to claim 5, the recording method according to claim And a moving image decoding system according to claim 7.
[0037] As shown here, in this invention, AUs to be decoded during trick-play effects, such as variable speed playback and reverse playback, can be determined by reference to a specific NAL unit in the upper AU of the random access unit RAU . Therefore, a device for decoding a moving image with excellent trick-play function can be easily implemented, and therefore the present invention is highly practical.
Brief Description of the Drawings [0038] These and other objects, advantages and features of the invention will become apparent from the following description in conjunction with the accompanying drawings, which illustrate a specific embodiment of the invention. In the drawings:
Figures 1A and 1B are schematics respectively showing MPEG-2 stream structures known in the art;
Figures 2A and 2B are schematics respectively showing MPEG-2 GOP structures known in the art;
Figs. 3A and 3B are schematics respectively showing MPEG-4 stream structures known in the art;
Fig. 4 is a block diagram showing the structure of a typical coding device;
Fig. 5 is a block diagram showing the structure of a typical decoding device;
Figures 6A and 6B are schematics respectively showing examples of prediction structure in a typical MPEG-4 AVC stream;
Fig. 7 is a diagram showing another example of the prediction structure in a typical MPEG-4 stream
AVC;
Figures 8A and 8B are schematics respectively showing the MPEG-4 AVC stream structures of the present invention;
<td>FIG.</td><td>9A to 9D are diagrams of the first</td>
<td>example</td><td>presenting AUs to be decoded in</td>
<td>unit</td><td>free RAU access;</td>
<td>FIG.</td><td>10A to 10D are diagrams of the second</td>
<td>example</td><td>presenting AUs to be decoded in</td>
<td>unit</td><td>free RAU access;</td>
<td>FIG.</td><td>11A to 11C are diagrams of the third</td>
<td>example</td><td>presenting AUs to be decoded in</td>
<td>unit</td><td>free RAU access;</td>
<td>FIG.</td><td>12A to 12F are diagrams of an example</td>
showing a method of determining AUs to be decoded in a random access unit RAU;
Fig. 13A is a diagram showing an example of a table syntax storing information related to variable speed playback and Fig. 13B is a diagram showing a data storage unit;
Fig. 14 is a diagram showing an example of a table syntax storing information related to variable speed playback;
Figures 15A to 15C are diagrams of an example showing AUs of picture I and pictures P in the random access unit RAU as information related to variable speed playback;
Figures 16A to 16C show diagrams of an example in which the buffer time is used as a priority indicator when using AU priorities as variable speed playback information;
Figs. 17A and 17B are diagrams showing respectively examples in which the AU unit frame structure and AU unit field structure coexist in respective RAUs;
<td>FIG.</td><td>17C presents</td><td>scheme</td><td>showing</td><td>example</td>
<td>composition</td><td>first map</td><td colspan="3">(RAU map1) depicting</td>
<td>structure</td><td>each unit</td><td colspan="2">AU in the RAU unit;</td><td></td>
<td>FIG.</td><td>17D presents</td><td>scheme</td><td>showing</td><td>RAU_map1</td>
<td>units</td><td>RAU visible on</td><td>Fig. 17B;</td><td></td><td></td>
<td>FIG.</td><td>17E presents</td><td>scheme</td><td>showing</td><td>RAU_map</td>
<td colspan="2">as a free unit</td><td>access</td><td>RAU visible</td><td>in fig</td>
<td>17B;</td><td></td><td></td><td></td><td></td>
<td>FIG.</td><td>17F presents</td><td>scheme</td><td>showing</td><td>example</td>
<td>syntax</td><td colspan="2">second map (RAU map2)</td><td colspan="2">representing the type</td>
<td>coding</td><td>each frame or</td><td colspan="3">each image of a pair of fields;</td>
<td>FIG.</td><td>18A to Fig.</td><td>18C</td><td>presents</td><td>diagrams</td>
presenting another example map as information
<td>connected with</td><td colspan="3">playback;</td>
<td>FIG.</td><td> 19</td><td>presents</td><td>diagram of the way of indicating</td>
<td>information</td><td>on</td><td>frame theme</td><td>in the free access unit</td>
<td>RAU;</td><td></td><td></td><td></td>
<td>FIG.</td><td>20A</td><td>and fig.</td><td>20B are diagrams</td>
presenting examples of image prediction structures in
<td>unit</td><td>free RAU access;</td>
<td>FIG.</td><td>21 is a block diagram showing</td>
<td>structure</td><td>moving image coding devices</td>
object of the present invention;
<td>FIG.</td><td>22 is a flow diagram of a method</td>
<td>coding</td><td>moving image;</td>
<td>FIG.</td><td>23 is a block diagram showing</td>
<td>structure</td><td>moving image multiplexing devices</td>
according to the present invention;
Figs. 24A and 24B are diagrams showing example content of HLP support information;
Fig. 25 is a diagram showing an example of a NAL unit where trick-play information is stored in HLP support information;
Fig. 26 is a flow chart showing
<td>action</td><td>mobile multiplexing devices</td>
<td>image;</td><td></td>
<td>FIG.</td><td>27 is a block diagram showing</td>
<td>structure</td><td>moving image decoding devices</td>
according to the present invention;
Fig. 28 is a flowchart of a typical image decoding method;
<td>FIG.</td><td>29 is a flow diagram for determining</td>
<td>units</td><td>AU for decoding as part of a decoding method</td>
<td>movable</td><td>image of the present invention;</td>
<td>FIG.</td><td>30 is a flow chart showing</td>
processing carried out when the AUs to be decoded do not correspond to the AUs to be displayed as part of the method for decoding a moving image in accordance with the present invention;
Fig. 31 is a diagram showing the data hierarchy of the HD-DVD format;
Fig. 32 is a structural diagram of the logical space of the HD-DVD;
Fig. 33 is a structural diagram of the file with VOB information;
Fig. 34 is a time map diagram;
Fig. 35 is a structural diagram of a play list file;
Fig. 36 is a structural diagram of a program file corresponding to a play list;
Fig. 37 is a structural diagram of the management information file for the entire BD disk;
Fig. 38 is a structural diagram of a file for recording the global event handler;
Fig. 39 is a block diagram showing the operation of an HD-DVD player;
and Figures 4A to 40C are diagrams showing a program storage medium for implementing the method of encoding a moving image and the method of decoding a moving image according to the present invention.
Preferred Mode for Carrying Out the Invention [0039] The method for carrying out the present invention will be described based on the drawings.
(AVC stream structure) [0040] First, the structure of the AVC stream generated by the moving image generating apparatus, the moving image coding apparatus and the moving image multiplexing apparatus according to the present invention will be described. In other words, the structure of the AVC stream for insertion into the moving picture decoding apparatus of the present invention.
[0041] Figs. 8A and f8B respectively show the structures of the AVC streams of the present invention. Note that the frame information that will be added at the top of the NAL unit is not visible in the drawing. The AVC stream differs from a typical AVC stream in that trick-play information has been added, with the trickplay information indicating AUs to be decoded during trick-play functions such as jump-in, variable speed playback and reverse playback. The trick-play information is stored in the NAL unit for storing playback information (Fig. 8A). For the MPEG-4 AVC standard, the relationship between storage information and the NAL unit type of a specific NAL unit can be determined by the application. More specifically, the values 0 and 24-31 can be used, with these NAL unit types being referred to as user-defined NAL unit types. As a result, trick-play information is stored in a NAL unit having these types of user-defined NAL units. In the event that specific NAL unit types are reserved to store information other than trick-play information, NAL unit types that are other than NAL unit types are allocated to trick-play information. NAL units with trick-play information are stored on top of the AU of the random access unit RAU. Such a NAL unit is placed immediately after the PPS NAL unit (if present) in the AU unit, but may be placed in a different position as long as the order obtained meets the requirements of the MPEG-4 AVC standard or another standard. Also, in case it is not possible to interpret the NAL unit with trick-play information, the NAL unit data may be skipped and the decoding is restarted from the top of the next NAL unit. Therefore, even a terminal that cannot interpret the NAL unit with trick-play information can successfully perform decoding.
[0042] Note that such a NAL with trick-play information may be located not at the top of the AU of the random access unit RAU, but in another AU, such as the last AU. In addition, such a NAL unit with trick-play information may be located in any AU that is a random access unit RAU.
[0043] Figures 9 to 11 show examples of AUs to be decoded during variable speed playback. Fig. 9A shows the order in which AU units are displayed. In this case, oblique dashed AUs are AUs to which other AUs refer, and the arrows represent the images to which they refer. Negative reference numbers are associated with AUs displayed before I0, and positive reference numbers are associated with AUs displayed after B15. Fig. 9B shows the decoding order of the AUs shown in Fig. 9A, with I0 to B11 forming a random access unit RAU. At this time I0, -B14,
P4, B2, P8, P6, P12 and B10 are decoded to perform double speed playback (Fig. 9C), while I0, P4, P8 and P12 are decoded to perform playback at four times the speed (Fig. 9D) . Figures 9C and 9D show that AUs with the * sign must be decoded during double-speed playback and quad-speed playback and these pieces of information are stored in the NAL unit with trick-play information. In the example shown in Figures 10A to 10D, the images I0 to B11 in the decoding order form a random access unit RAU. In this case, I0, B13, P3, B1, P6, B4, P9, B7, P12 and B10 are decoded to perform playback at 1.5x speed, while I0, P3, P6, P9 and P12 are decoded to perform triple speed playback. In addition, in the example shown in Fig. 11A to 11C, I0, P3, P6, P9 and P12 are decoded to perform triple-speed playback.
[0044] In this case, the playback speeds need not be accurate because they are presented as guidelines of the playback speed. In the example shown in Fig. 11C, in the event that all AUs represented as AUs to be decoded during triple-speed playback are decoded, a 3.2x speed is obtained based on the expression: 16 ^ 5, in other words, this is not triple speed exactly. In addition, the playing time at the speed Mx in the case where the smallest value above M is N of the playback speeds presented as trick-play information, it is possible to decode the units
AUs to be decoded at N-speed playback and determining how the rest of the AUs should be decoded depending on the implementation of the decoding device. In addition, it is possible to set high priorities for AUs that need to be decoded when the playback speed is high, and also specify AUs to be decoded based on priorities.
[0045] Note that some AUs among the AUs to be decoded during variable speed playback may not be displayed. For example, the AU unit number N is displayed during double speed playback, but the unit number M is no longer displayed. At this time, in case there is no need to decode the AU with the M number to decode the AU with the N number, the AU with the M number is decoded but is not displayed during double speed playback.
[0046] Next, the method for determining AUs to be decoded during variable speed playback based on Figs. 12A to 12F will be described. Figures 12A to 12F show examples of determining AUs for decoding in the same random access unit RAU as in Figure 9. As shown in Figure 12D, images I0, -B14, P4, B2, P8, P6, P12, B10 are decoded during double speed playback. These AUs are units numbered 1, 2, 5, 6, 9, 10, 13 and 14, assuming that the counting begins with the AU unit at the top of the random access unit RAU. In this way, it is possible to uniquely specify AUs to be decoded during variable speed playback by presenting the ordinal numbers of AUs in the random access unit RAU. The access unit limiter is definitely located at the top of the AU when multiplexing the AVC stream by the MPEG-2 (TS) transport stream. When obtaining AU data to be decoded during variable speed playback, the access unit delimiters are searched in order to recognize AU frames. This search method eliminates the need to analyze a block of NAL unit data, such as slice data, and is therefore easier.
[0047] Note that it is possible to specify AUs to be decoded by determining that AUs to which other AUs refer, such as AUs of the I picture and P pictures (such AUs that are referenced are known as AU reference units) are decoded during variable speed playback, and by specifying ordinal reference numbers of AU reference units in a random access unit RAU. In the random access unit RAU shown in Fig. 12B, as shown in Fig. 12C, the images I0, -B14, P4, B2, P8, P6, P12, B10 are AU reference units. In addition, during double-speed playback, I0, -B14, P4, B2, P8, P6, P12, B10 are decoded, but when indicating these AUs in the order of AU reference units they correspond to the first, second, third, fourth, fifth, sixth , seventh and eighth reference units AU as shown in Fig. 12F. Whether or not the AU is an AU reference unit can be assessed by reference to a specific field in the header of the NAL unit in the patch. In particular, if the nal_ref_idc field value is not 0, the AU is the reference unit of AU. Note that the AU reference unit to be decoded can be determined based on the frame number because the AU reference unit can be identified based on the frame number. <
> [0048] Furthermore, it is possible to specify AUs to be decoded by determining the equivalent of the offset value in length in bytes from the start position of the upper AU of the random access unit RAU to the start position of the AU to be decoded. For example in fig. 12A to 12F, when I0 starts from a position distant from the top of the stream by 10,000 bytes, and P4 starts from a position distant from P4 by 20,000 bytes, the offset value to P4 is 10,000 bytes. It is obtained from the expression: 20,000-10000. If the multiplexed stream is used in the MPEG-2 (TS) transport stream, it is possible to specify the offset value taking into account the additional header of the TS packet or PES (Packetized Elementary Stream) packet or it is possible to specify the offset value taking into account the above when performing data complementation (padding) by the application. In addition, it is possible to specify an AU based on the FN frame number.
[0049] Note that when using a multiplexed stream in an MPEG-2 TS transport stream, it is possible to specify AUs based on the number of TS packets from (i) the TS packet for storing the index number and address information for identifying the TS packet including data on top of AUs to be decoded or data on top of random access units RAU to (ii) the current TS packet. You can use information about the source package used for the Blu-ray Disc (BD) recording format instead of the TS package. The source package is obtained by adding to the TS packet a 4-byte header containing time information for the TS packet, copy protection information and the like.
[0050] Fig. 13A shows an example of a table syntax storing information related to variable speed playback. In the syntax num_pic_in_RAU presents the number of AUs that make up a random access unit RAU, num_speed presents the number of playback speeds at which AUs will be decoded, play_speed presents the playback speed, num_dec_pic presents the number of AUs to be decoded during playback at the playback speed shown in play_speed , dec_pic presents the order numbers of AUs to be decoded when AUs are counted, starting from the top of the AU, in the unit with random access RAU. Fig. 13B shows an example of storing information in AUs to be decoded in the random access unit RAU of Figs. 9A to 9D during double-speed playback and quad-speed playback. Note that the num_pic_in_RAU variable is used when calculating the exact playback speed based on the number of AUs to be decoded and the total number of AUs in a random access unit RAU or when skipping based on a random access unit RAU in sequence.
However, the num_pic_in_RAU variable can be omitted because the same information can be obtained by searching the top AUs of a random access unit RAU. In addition, a field indicating the size of the table can be added to the array. It should be noted that in the syntax example presented in Fig. 13A, the order number of the AU to be decoded (when counting the random access unit RAU) is presented directly, but whether or not it is necessary to decode each AU can be represented by enabling or disabling the bits corresponding to each AU. For example, a random access unit RAU consists of 16 AUs in the example shown in Figures 9A to 9D. If you assign 1 bit to each AU, 16 bits are needed. During quadruple playback, it was shown that the first, fifth, ninth and thirteenth AUs are decoded by allocating 16-bit information, which is represented as 0b1000100010001000 (0b is a binary number). In this case, the highest bit and the last bit correspond respectively to the upper AU and the last AU of the random access unit RAU.
[0051] Note that the size of the table is variable in the example of the syntax shown in Fig. 13A. The maximum value of the array size is set when the maximum value of the number of AUs that make up the random access unit RAU is specified, and the maximum value of the variable num_speed. As a result, it is possible to set the size of the table at a certain maximum value and, in the event that the amount of information for variable speed playback does not reach the maximum value, it is possible to perform padding. Determining the size of an array in this way always allows you to get the data of the set size when obtaining information about playback at variable speed, which allows you to speed up the acquisition of information. Note that the array size or NAL unit size for the storage of the array is presented as management information. In addition, it is possible to specify the size of the NAL unit in advance for storing trickplay information, and (in the event that the information cannot be stored in a single NAL unit), it is possible to save information for variable speed playback separately in multiple NAL units. At this point, padding is performed on the data block of the last NAL unit, so that the NAL unit size has a predetermined size. In addition, some predefined values are referred to as array size values, and an index number indicating the specified array size value can be displayed in the array or using application management information.
[0052] Furthermore, it is possible to present differential information instead of listing all AUs to be decoded at each playback speed. If the information will be played back at the speed M (where M <N), only AUs to be decoded will be shown except for the units to be decoded at playback speed N. In the example in Fig. 13B, since during double speed playback the second, sixth, tenth and fourteenth AU units are decoded - except for AU units decoded at four speeds - only the second, sixth, tenth and fourteenth AU units can be presented as information for double playback speed.
[0053] It should be noted that the AUs needed to be decoded during variable speed playback have been presented in the above description, but it is possible to show information indicating the order in which AUs needed to be decoded are displayed. For example, information during double speed and quadruple speed playback is shown in the example in Fig. 9A to 9D, but here is an example of playing a random access unit RAU at triple speed. By displaying some of the AUs to be displayed during double-speed playback, in addition to the AUs to be displayed during 4-speed playback, you can perform triple-speed playback. When we consider the case where one more AU is displayed between I0 and P4 to be displayed during quadruple playback, the information for the purposes of double speed playback shows that the images are -B14, B2, B6 and B10. However, the display order of these four AUs can only be obtained when slice header information is being analyzed. Since, according to the display order information, only the -B14 image is displayed between I0 and P4, it is possible to determine that the -B14 image is decoded. Fig. 14 is an example of syntax indicating information about the display order. It was obtained by adding information about the display order to the syntax shown in Fig. 13A. In this case, the variable pts_dts_flag shows whether the order of decoding AUs to be decoded at the playback speed matches the order in which AUs are played, only if the decoding order does not match the display order, information about the display order is presented in the display_order field.
[0054] It should be noted that for playback at a playback speed that is not represented by the variable speed playback information, it is possible to specify AUs to be decoded and AUs to be displayed based on a rule that is predetermined in the terminal. For example, when playing at triple speed in the example shown in Fig. 9 it is possible to display I0, B3, B6, B9 and P12 images in addition to AU units for display during quadruple speed instead of displaying some AU units for display during double speed playback. In this case, B pictures in AU reference units may be decoded or displayed in a privileged manner.
[0055] In addition, there is a case in which trick-play functions, such as variable speed playback, are implemented by playing only the image units AU of the image I or only the image units AU of the image I and the images P, therefore the list containing the image I and the images P may be saved as trick-play information. Figures 15A to 15C show another example. In this case, images from I0 to B14 are part of the RAU unit, which is shown in Fig. 15B, among them AU units of picture I and pictures P are I0, P3, P6, P9, P12 and P15, as shown in Fig. 15C. Therefore, information is stored for identification purposes I0, P3, P6, P9, P12 and P15. At this point, it is possible to add information to distinguish the AU of the I picture from the AU of the P picture. In addition, it is possible to provide information for distinguishing the following images from each other: image I, images P, images B referenced (hereinafter referred to as reference images B), and images B not referenced (hereinafter referred to as non-reference images B).
[0056] In addition, it is possible to save priority information of respective AUs as trick-play information and to decode or display AUs by priority during playback with a variable. It is possible to use image types as priority information. For example, AUs can be prioritized in the following order:
(i) picture I; (ii) P images; (iii) reference B images and (iv) non-reference B images. In addition, you can set priority information as follows: the longer the time between decoding an AU and its display, the higher the priorities. Figures 16 A to 16C show an example of prioritizing according to the time spent in the buffer. Fig. 16A shows the prediction structure of AU units, and B7 and P9 also refer to P3. At this point, in the case where the random access unit RAU consists of AUs I0 to B11 (Fig. 16B), the buffer time for each AU is as in Fig. 16C. The residence time in the buffer is here according to the speed.
presented based on the number of frames. For example, a P3 image is required before P9 decoding, and the buffer time must be equivalent to six images. Therefore, decoding AUs for which the buffer time is 3 or more, means decoding all images - I image and P images, and performing triple-speed playback. In this case, the residence time in the P3 image buffer is longer than in the case of I0, but it is possible to add the offset value to the AU of the I image to set the highest priority for the AU of the I image. In addition, it is possible to set high priorities for AUs needed to be decoded during high speed playback and to use N priority in AUs to be decoded during N-time playback. It should be noted that if the AU refers to other AUs after it has been decoded or displayed, it is possible to present the time interval during which reference is made to the AU.
[0057] Note that trick-play information may be stored in the SEI message (Fig. 8B). In this case, the SEI message type is defined for trick-play information, the trick-play information is SEI of the defined type.
stored in the message The SEI message for trick-play information is stored in the SEI NAL unit separately or together with other SEI messages. Note that it is possible to save the trick-play information in the SEI user data registered itu t t35 message or in the SEI user_data_unregistered message. These are SEI messages to store user-defined information. When using these SEI messages, it is possible to demonstrate that the trick-play information is retained, or to the type of trick-play information in the SEI data block by adding identification information of the information to be retained.
[0058] Note that it is possible to keep trick-play information in AUs other than the top AU in the random access unit RAU. In addition, it is possible to set values in advance for the identification of AUs necessary to be decoded during playback at a specific playback speed and to add values specified for each AU. For example, for AUs to be decoded at a playback speed of N or less, N is given as information about the playback speed. In addition, it is possible to present the following information in the variable nal_ref_idc and the like regarding the NAL unit of the patch: image structure in AU, structure constituting the frame structure or field structure, and in addition, if the image has a field structure, it is possible to present the field type, there is an upper or lower field. For example, since there is a need for alternating display of upper and lower fields in interlaced display, it is desirable that whether the next field to be decoded is to be the upper or lower field can be easily determined when decoding fields by bypassing some fields in during high speed playback. In the case where the field type can be determined based on the NAL unit header, there is no need to analyze the slice header, and the amount of resources needed for such determination can be reduced.
[0059] Note that information indicating whether each AU that forms the random access unit RAU is a field or frame can be stored in the upper AU of the random access unit RAU. In addition, it is possible to easily specify AUs to be decoded when using the trick-play function even if the field structure and frame structure coexist, storing this information in the upper AU of the random access unit. FIG. 17A and 17B show examples in which ANUS having a frame structure and AUs having a field structure coexist in a random access unit RAU, wherein they display the order of displaying AUs and the order of decoding AUs, respectively. The following images are coded as field pairs, respectively: B2 and B3; I4 and P5; B9 and B10; B11 and B12; P13 and P14;
B15 and B16; B17 and B18; as well as P19 and P20. In addition, other AUs are encoded as AUs having a frame structure. When playing only the AUs of I picture and P pictures, the following pictures can be decoded and played back in the following order: pair of fields I4 and P5; P8 frame; pair of fields P13 and P14; and also a pair of fields P19 and P20. However, adding this information is effective because it is necessary to determine if each AU is one of the fields that are a pair of fields, or whether each AU is a frame when determining AUs to be decoded.
[0060] Fig. 17C is an example of the first map syntax (RAU_map1) indicating whether an AU in a random access unit RAU is a frame or a field. The number of AUs that make up the random access unit is presented in the variable num_AU_in_RAU, and information about each AU is presented in the following loop in the decoding order frame_field_flag shows
In this case, the variable to write in whether the image of the AU is a frame or a field. In addition, the variable pic_type presents information about the type of image encoding. The types of coding that can be represented include: picture I;
IDR image; image P; reference image B; picture B without reference and the like. Therefore, it is possible to specify images to be decoded during the trick-play function by referring to this map. Note that it is possible to indicate whether or not each I image and P image are referenced. It is also possible to indicate information for the purpose of determining whether a predefined requirement is applied to prediction structures.
[0061] Fig. 17D is a map of RAU_map1 associated with the random access unit RAU of Fig. 17B. In this case, the pic_type of the I image, P images, B reference images, and B referenced images will be 0, 1, 2, 3, respectively In this case, it is possible to store information indicating the types of image coding according to the above method, because the images are played back frame-by-frame or pair-by-pair of fields during the trick-play function.
[0062] Fig. 17F is an example of the second map syntax (RAU_map2) indicating the types of frame-by-frame or frame-by-field pair coding. In this case, the variable num_frame_in_RAU represents the number of frames that make up the RAU unit and the number of field pairs. In addition, the variable frame_flag indicates whether the image is a frame or not, and if it is a frame, it is set to 1. When frame_flag is 1, information about the frame coding type is presented in the frame_type variable. In the case where frame_flag is 0 (i.e. the image is one of the field pairs), the coding type of each field that makes up the field pair is presented in the field_pair_type variable.
[0063] Fig. 17E is a map of RAU_map2 associated with the random access unit RAU of Fig. 17B. In Fig. 17E, the values indicating the frame_type of the I picture, P pictures, reference B pictures, and B non-reference pictures are 0, 1, 2, 3, respectively. In addition, field_pair_type presents the type of each field in decoding order. The following field types are available: I for image I, P for image P; Br for reference images B and Bn for reference images. For example, when the first field is image I and the second field is image P, the IP designation will be used. If the first and second fields are B images without reference, the BnBn designation will be used. In this case, the values indicating the combinations of IP, PI, PP, BrBr, BnBn and the like are determined in advance. Note that the following information may be used as information indicating the coding types of a field pair: information about whether the field pair contains an I image or one or more P images; information on whether the field pair contains one or more reference B images; information on whether a field pair contains one or more B images without reference.
[0064] For example, the trick-play information may be a map of the random access unit RAU, as in the syntax example shown in Fig. 18A. This map contains the picture_structure variable denoting the structure of each image included in the RAU unit and the picture_type variable denoting the image type. As shown in Fig. 18B, the variable picture_structure presents the structure of each image, namely field structure or frame structure and the like. In addition, as shown in Fig. 18C, the variable picture_type represents the image type of each image, namely image I, reference image B, reference image B and image P. In this way, the moving image decoding device that receives this map can easily identify AUs on which the trick-play function is performed by reference to this map. For example, it is possible to decode and play (high speed playback) only the I image and P or reference B images next to the I image and P images.
[0065] It should be noted that in the case where information indicating the structure of the image, such as 3-2 pull down correction, is contained in the AU, which is a random access unit RAU, it is possible to attach information about the structure of the image to the described before the first or second map. For example, it is possible to show whether each image has display fields corresponding to three images, or each image has display fields corresponding to two images. Further, in the event that it has display fields corresponding to three images, it is possible to provide information indicating whether the first field is displayed repeatedly, or information indicating whether the first field is the top field. In turn, if it has display fields corresponding to two images, it is possible to provide information indicating whether the first field is the top field. In the case of the MPEG-4 AVC standard, whether the image has an image structure such as 3-2 pull down correction can be represented by (i) the pic_struct_present_flag variable of the SPS sequence parameter set or (ii) the picture_to_display_conversion_flag variable and similar in the synchronization descriptor AVC and HRD, defined in the MPEG-2 system standard. In addition, the structure of each image is presented using the pic_struct field of SEI information on image synchronization. Therefore, it is possible to present the image structure by setting the flag only if the pic_struct field has a specific value, for example, if the image has display fields corresponding to three images. In other words, the indication of the following three types of information for each image is effective (i) when jump-in playback is performed in the center of a random access unit RAU and (ii) when specifying a field to display at a specified time or frame, in which the field is stored. The same applies to specifying images to be displayed during variable speed playback. Three types of information are available:
(i) field (ii) a frame (which is used when 3-2 pull down correction is not used, or which is also used when using 3-2 pull down correction. In the second case, the frame has display fields corresponding to two images.) (iii) a frame having displayed fields corresponding to three images when using 3-2 pull down correction. Note that these types of information can be signaled in the picture_structure variable of the RAU map shown in Fig. 18A.
[0066] Indicating information on the types of respective images that make up the RAU, thus allows easily specifying the images to be decoded or displayed while performing trick-play functions such as variable speed playback, jump-in playback, and reverse playback. This is especially effective in the following cases:
(i) when only the I picture and P pictures are played back;
(ii) when high-speed playback of I image, P images and reference B images is performed, and (iii) when images for which there are requirements for prediction structures are identified based on image types, the images needed to be decoded while the trick-play function is being performed, the selected images are played back as part of the trick-play function.
[0067] In addition, it is possible to save the default value of trick-play information in a zone that is different from the AVC stream, such as application-level management information, as well as attach trick-play information to the random access unit RAU only if the information trick-play is different from the trick-play information presented using the default value.
[0068] Trick-play information related to variable speed playback has been described above, but it is possible to use similar information as complementary information during reverse playback. You can finish decoding during reverse playback if all the images to be displayed can be stored in memory. Resources needed for decoding can be reduced in this case. If you consider the reverse playback case in the order listed P12, P8, P4 and I0 in the example shown in Figures 9A to 9D, provided that all decoding results of these four AUs are preserved, it is possible to decode the images I0, P4, P8 and P12 in this order at a time and perform reverse playback. Therefore, it is possible to assess whether all decoded AU data can be saved based on the number of AUs to be decoded or displayed during N-speed playback, and to specify AUs to display when performing reverse playback based on the result of this evaluation.
[0069] Similarly, trick-play information may be used as supplementary information during jump-in playback. It is assumed here that jump43 in playback means fast forward the moving image and perform a standard playback of the moving image starting from a random position. Specifying images for fast forward using such supplementary information even during jump-in playback allows you to specify the image at which jump-in playback begins.
[0070] It should be noted that the AU to which reference is made of each AU that forms a random access unit may be directly represented in the trick-play information. In the case where there are multiple AU reference units, all are presented. Here it is assumed that in case the AU reference unit belongs to a random access unit other than the random access unit containing the AU which refers to the AU reference unit, the AU may be determined in the following particular way: AU with number M random access units, which is placed before or after the N random access units, or the AU may be determined as follows: AU unit belonging to the random access unit, which is placed before or after the N number of random access units. Note that it is possible to display the order number (in decoding order) of an AU when counting from an AU that refers to the AU reference unit. Then AU units are counted according to one of the following rules: all AU units; AU reference units; AUs of a particular type of image, for example I, P and B. In addition, it is possible to indicate that each AU may only refer to AUs to the number N of AUs before and after in decoding order. Note that when referring to an AU that is not included in AUs located up to the number N of AUs before and after in decoding order, it is possible to add information indicating this fact.
[0071] Note that it is also possible to use the trick-play information described above in a similar manner also in a multiplexing format, such as MP4, where the NAL unit size is used instead of using the start code prefix as the frame information of the NAL unit.
[0072] It should be noted that when receiving and saving an encoded stream that is packaged using the MPEG-2 TS transport stream packet or RTP (Real Time Transmission Protocol) packet loss occurs. In this way, when writing data received in an environment where packet loss occurs, it is possible to store in the coded stream in the form of supplementary information or management information information indicating that the data in the stream has been lost due to packet loss. You can present data loss due to packet loss by entering signaling information indicating whether the stream data has been lost or not, or a special error notification code to report the lost fragment. It should be noted that in the event of an error concealment process where data has been lost, it is possible to retain identification information indicating the presence / absence or concealment of the error.
[0073] The trick-play information for determining the AUs to be decoded or displayed during the trick-play function has been described previously. The data structure for frame detection of the random access unit RAU will be described with reference to Fig. 19.
[0074] In the upper AU of the random access unit RAU, there is always written the NAL unit of the sequence parameter set (SPS) to which the AU refers which forms the random access unit RAU. On the other hand, in the MPEG-4 AVC standard, it is possible to preserve the NAL unit of the parameter sequence set (SPS), to which the AU with the N number refers in order of decoding, in the AU that is freely selected from among the AU with the N number or units AU placed before AU with number N in decoding order. Such a NAL unit is maintained so that the NAL unit of the SPS sequence parameter set can be repeatedly transmitted pending the case where the NAL unit of the SPS sequence parameter set is lost due to packet loss during stream transmission during communication or broadcasting. However, for the purposes of using storage applications, the following rule is effective. Only the single NAL unit of the SPS sequence parameter set, to which all AUs of the random access unit RAU refer, is stored in the upper AU of the random access unit RAU, the NAL unit of the SPS sequence set parameter is not stored in subsequent AUs in unit with free access. This ensures that the AU is the upper AU of the random access unit RAU if it contains the NAL unit of the SPS sequence parameter set. The beginning of a random access unit RAU can be found by searching the NAL unit of the SPS sequence parameter set. Stream management information, such as a time map, does not guarantee the provision of access information for all units with random access RAU. Therefore, it is particularly important that the start position of each random access unit RAU can be obtained by searching the NAL unit of the SPS sequence parameter set in the stream, for example when performing jump-in playback on an image placed in the center of a random access unit RAU access information.
[0075] In the case where the upper AU of the random access unit RAU is the AU of the IDR image, the AU of the random access unit of the RAU does not refer to the AU in the random access unit RAU that is placed in the order of decoding first. This type of random access unit RAU is called closed type random access unit RAU. On the other hand, in the case where the upper AU of the random access unit RAU is the AU of the I image which is not the IDR image, the AU of the random access unit of the RAU may refer to the AU of the random access unit of the RAU which is previously placed in decoding order. This type of random access unit RAU is called open type random access unit RAU. When angles in an optical disk or the like are switched during playback, the switching takes place from the closed-type random access unit RAU. Therefore, determining whether the random access unit RAU is of the open or closed type can be carried out at the top of the random access unit RAU. For example, it is possible to present signaling information for type determination (open or closed type unit) in the nal_ref_idc field of the NAL unit of the SPS sequence parameter set. Since the nal_ref_idc value has been defined to be 1 or more in the SPS NAL unit, the high bit is always set to 1 and the signaling information is represented by the low bit. Note that the AU in the random access unit RAU may not refer to the AU in the random access unit RAU that is placed in the decoding order before, even if the upper AU is an AU of the picture I, which is not an IDR image. This type of random access unit RAU can be considered as a closed type random access unit RAU. Note that signaling information can be presented using the nal_ ref_idc field.
[0076] Note that it is possible to determine the start position of the random access unit RAU based on a NAL unit other than the SPS sequence parameter set for storage only in the upper AU of the random access unit RAU. In addition, it is possible to show the type (open type or closed type) of each unit with random access RAU using the nal_ref_idc field of each unit with random access
RAU.
[0077] Finally, Figs. 20A and 20B show examples of AU prediction structures that form a random access unit RAU. Fig. 20A shows AU positions in order of display, and Fig. 20B shows AU positions in decoding order. As shown in the drawings, B1 and B2, which are located in front of I3, which is the top AU of the random access unit RAU, may refer to AUs for display after I3. In Figure B1, refers to P6. To ensure that AUs of the I3 image and subsequent images in the display order can be correctly decoded, AUs of the I3 image and subsequent images in the display order are not allowed to refer to AUs prior to I3 in the display order.
(Moving image coding apparatus) [0078] Fig. 21 is a block diagram of a moving image coding apparatus 100 that implements the moving image coding method of the present invention. The present device for encoding a moving image 100 generates an encoded stream (shown in Figs. 8 to 20) of a moving image that can be played using trick-play functions such as jump-in playback, variable speed playback and reverse playback. The moving picture coding apparatus 100 has a trick-play information generation unit (TrickPlay) in addition to the units of the typical moving picture coding device 1 of Fig. 4. It should be noted that those processing units that perform the same operations as the processing units in a typical the moving image coding apparatus shown in Fig. 4, the same reference numbers have been assigned and their descriptions will be omitted.
[0079] The trick-play information generation unit (TrickPlay) is an example of a unit that generates, based on a random access unit containing one or more images, complementary information that is referenced during playback of random access units. The trickplay information generation unit (TrickPlay) generates trick-play information based on Ptype image types and passes trick-play information to a variable length VLC coding unit.
[0080] The VLC variable length coding unit is an example of a stream generation unit that generates a stream containing supplementary information and images by adding generated supplementary information to each respective random access unit. The VLC variable coding unit encodes and places a NAL unit for storing trickplay information in the upper AU of the random access unit
RAU.
[0081] Fig. 22 is a flowchart showing how the moving image coding apparatus 100 (mainly the trick-play information generation unit (TrickPlay)) shown in Fig. 21 performs the procedure for generating the coded stream containing the image 100 generating the trick- information play.
[0082] First, during step 10, the moving image coding apparatus 100 determines whether the AU to be coded is the upper AU of the random access unit RAU. If it is an upper AU, then it goes to step 11. And if it is not an upper AU, it goes to step
12. In step 11, the moving image coding apparatus 100 performs pre-processing to generate trick-play information of the random access unit RAU, and also secures the area to write trickplay information in the upper AU of the random access unit RAU. In step 12, the moving picture coding apparatus 100 encodes the AU data followed by step 13. In step 13, the mobile coding apparatus obtains the information necessary during trick-play information. This information is as follows: AU image types, namely I image, P image, reference B image, or non-reference B image; or whether there is a need to decode the AU at the time of performing N-fold playback. Then, the moving image coding apparatus 100 proceeds to step 14. In step 14, the moving image coding apparatus 100 determines whether the AU is the last AU of the random access unit RAU. In the case where this is the last AU unit, the moving image coding apparatus 100 proceeds to step 15. In turn, if it is not the last AU unit, then proceeds to step 16. At step 15, the moving image coding apparatus 100 determines the trick-play information, generates a NAL unit for storing trick-play information, and also saves the generated NAL unit to the security area in step 11. After completing step 15, the moving coding device 100 passes to step 16. In step 16, the moving image coding apparatus 100 determines whether an AU exists for subsequent coding. In the case where an AU exists for coding, step 10 and subsequent steps are repeated. Conversely, when there are no more AUs to encode, processing is complete. In the event that the moving picture coding apparatus 100 determines that there are no AUs to be encoded in step 16, the apparatus retains the trick-play information of the last random access unit RAU, after which processing ends.
[0083] For example, when the moving image coding apparatus 100 generates trick-play information as shown in Fig. 18A, obtains in step 13: image type; information about whether the image has a field structure or a frame structure, and / or information indicating whether the displayed image field is the equivalent of two images or the equivalent of three images in the event that information about 3-2 pull down correction is contained in the encoded stream. In step 15, the moving image coding apparatus 100 sets the picture_structure and picture_type variables of all images in the random access unit RAU in decoding order.
[0084] Note that if the size of the NAL unit for storing trick-play information is not known at the time the encoding of the upper AU of the random access unit RAU is started, processing to secure the area for storing trick-play information will be skipped in step 11. In this case, the generated trick-play information storage unit is placed in the upper AU in step 15.
[0085] Also, saving or not saving trick-play information can be switched based on the encoded stream. Particularly in the case where the prediction structure between AUs that form a random access unit is determined by the application, it is possible to specify that trick-play information is not saved. For example, if the encoded stream has the same prediction structure as MPEG-2 stream, there is no need to save trick-play information. This is because it is possible to determine AUs needed to be decoded during the trick-play function without trick-play information. Note that switching can be done on the basis of a random access unit
RAU.
(Moving image multiplexing device) [0086]
FIG.
is a block diagram showing the structure of a moving image multiplexing apparatus 108 of the present invention; This moving image multiplexing apparatus 108 introduces moving image data, encodes the moving image data to create an MPEG-4 AVC stream, multiplexes the stream with access information to AUs that make up the stream, and management information containing supplementary information to determine actions performed during execution trick-play function and saves the multiplexed stream. The moving image multiplexing apparatus 108 includes the stream attribute determining unit 101, the coding unit 102, the management information generating unit 103, the multiplexing unit 106 and the memory unit 107. The coding unit 102 has the function of adding trickplay information to the moving image coding device 100 shown in FIG. 21.
[0087] The stream attribute determining unit 101 specifies the requirements for trick-play functions performed during the coding of the MPEG-4 AVC stream, and then passes the results to the coding unit 102 and the reproduction assistance information generating unit 105 as TYPE attribute information. Requirements related to the trick-play function include information indicating: whether the requirement to create a random access unit is applied to the MPEG-4 AVC stream; whether information indicating AUs to be decoded or displayed during variable speed playback or reverse playback are included in the stream; or whether the prediction structure requirement between AUs is set. The attribute determination unit 101 then sends to the general management information generation unit 104 general management information, which is the information necessary to generate management information, such as compression format or resolution. The coding unit 102 encodes the input video data into the MPEG-4 AVC stream based on the TYPE attribute information, transmits the encoded data to the multiplexing unit 106, as well as access information in the stream to the general management information generation unit 104. In the event that the TYPE attribute information shows that information indicating AUs to be decoded or displayed during variable speed playback or reverse playback is not in the stream, the trick-play information is not placed in the encoded stream. Note that the access information indicates information about the access unit, which is the basic unit when accessing the stream, including the start address, display time and the like regarding the upper AU in the access unit. The generic management information generation unit 104 generates array data that is referenced when accessing the stream, as well as array data that stores attribute information, such as compression format, based on access information and general management information, and sends array data. to multiplexing unit 106 as INFO management information. The reproduction assistance information generating unit 105 generates HLP assistance information indicating whether the stream has a random access structure based on the entered information regarding TYPE attributes, and sends the HLP assistance information to multiplexing unit 106. The multiplexing unit 106 generates encoded data input by the coding unit 102, INFO management information and multiplexing data by multiplexing the HLP auxiliary information, and then sends the results to the memory unit 107. The memory unit 107 writes the multiplexing data introduced by the multiplexing unit 106 onto a storage medium such as optical disk, hard disk or memory. Note that the coding unit 102 may packet the MPEG-4 AVC stream by creating, for example, MPEG-2 TS transport streams or MPEG-2 PS program streams, and then generate MPEG-2 TS or PS packet streams. In addition, the coding unit 102 may packet the stream using a format established by the application, such as BD.
[0088] Note that the content of management information need not depend on whether trickplay information is stored in the encoded stream or not. HLP support information may then be omitted. Furthermore, the moving image multiplexing apparatus 108 may have a structure without a reproduction assistance information generating unit 105.
[0089] Figs. 24A and 24B show examples of information presented by HLP assistance information. HLP support information includes a method directly indicating information about the stream as shown in Fig. 24A, as well as a method indicating whether the stream meets the requirements imposed by a particular application standard as shown in Fig. 24.
[0090] Fig. 24A shows the following information regarding the stream: information about whether the stream has a random access structure; information on whether there is a requirement for a prediction structure between the images stored in the AU; and information on whether there is information indicating the AUs to be decoded or displayed during the trick-play function.
[0091] Information regarding AUs to be decoded or displayed while performing the trick-play function may directly indicate AUs to be decoded or displayed or indicate priorities during decoding or displaying. For example, it can be pointed out that information indicating AUs to be decoded or displayed based on a random access unit is stored in a NAL unit having a special type of NAL unit, specified by the application, SEI message and the like. Note that it is possible to determine if there is information indicating the structure of the prediction between AUs that make up the random access entity. In addition, information about AUs to be decoded or displayed during the trick-play function may be added based on one or more random access units or to each AU that creates a random access unit.
[0092] Furthermore, in the case where information indicating the AU information to be decoded or displayed is stored in a NAL unit that has a special type, it is possible to show the NAL unit type. In the example shown in Fig. 25, in the HLP support information, information regarding AUs to be decoded or displayed during the trick-play function is contained in a type 0 NAL unit. At this time, it is possible to obtain information about the trick-play function by demultiplexing the Type 0 NAL unit based on the stream AU data. In case information related to the trick-play function is stored using a SEI message, it is possible to indicate information enabling identification of the SEI message.
[0093] In addition, due to the requirements for prediction structures, it is possible to indicate whether one or more of the predetermined requirements have been met or it is possible to indicate the following requirements met independently:
(i) as regards the AUs of the I picture and P pictures, the decoding order should correspond to the display order;
(ii) the P image AU may not refer to the B image AU;
(iii) AUs following the upper AU in the display order of the random access unit may only refer to AUs contained in the random access unit; and (iv) each AU may only refer to AUs placed max. N numbers before and after in decoding order. In this case, all AUs are counted together or AUs are counted based on the AU reference unit and the N value can be represented in the HLP support information.
[0094] Note that for MPEG-4 AVC it is possible to use as reference images images on which, after decoding, filtering (block separation) is carried out to remove block distortion to improve image quality, and it is also possible to use images before cutting blocks as display images. In this case, the moving image decoding device should store the image data before and after the blocks are split. Therefore, it is possible to store information in the HLP support information indicating whether there is a need to save the images before cutting the blocks for display purposes. The MPEG-4 AVC standard defines the maximum buffer size (DPB: Decoded Picture Buffer) necessary to store reference images or images for display as decoding results. Therefore, thanks to the DPB buffer with the maximum size or the buffer with the maximum size specified by the application, it is possible to determine whether decoding can be performed without failure even when storing images for displaying reference images. Note that to preserve images before cutting up blocks of reference images, it is possible to indicate the buffer size to be protected, in addition to the size required as DPB using the number of bytes or the number of frames. Whether or not block extraction is performed on each image can be obtained from information in the stream or non-stream information, such as management information. When obtaining information in a stream, it can for example be obtained from SEI. In addition, when decoding an MPEG-4 AVC stream, it is possible to determine whether the images before cutting the blocks of reference images can be used for display purposes or not, based on the buffer size that can be used in the decoding unit and the information described above, and then exists the ability to specify how images are displayed.
[0095] Note that all information or some information may be included as HLP support information. In addition, it is possible to include the necessary information based on a predetermined condition, for example to include information about the presence or absence of trick-play information only if there is no requirement for a prediction structure. In addition, information other than those described above may be included in the HLP support information.
[0096] Fig. 24B does not directly indicate stream structure information, but determines whether the stream meets the stream structure requirements set by the Blu-ray Disc (BD-ROM) or High Definition (HD) DVD standard, which is a standard for storing high definition images on DVD discs. In addition, in the event that multiple modes are defined as stream requirements in an application standard, such as BD-ROM or similar, information indicating the mode used may be retained. For example, the following modes are used:
Mode 1 indicating that there are no requirements;
mode 2 indicating that the stream has a random access structure and contains information for determining AUs to be decoded during the trick-play function; and the like. Note that it is possible to determine whether a stream meets the requirements of a communication service, such as downloading or streaming, or a broadcasting standard.
[0097] It should be noted that it is possible to indicate both the information shown in Fig. 24A and in Fig. 24B. In addition, if the stream is known to meet the requirements of a particular application standard, it is possible to maintain the requirements of the application standard by converting the stream structure to a format for direct description, as shown in Fig. 24A, instead of indicating whether the stream meets the given application standard.
[0098] Note that it is possible to save information indicating AUs to be decoded or displayed during trick-play as management information. Also, in the event that the content of HLP ancillary information is changed in the stream, HLP ancillary information may be determined section-by-section.
[0099] Fig. 26 is a flowchart showing the operation of a moving image multiplexing apparatus 108. In step 51, the stream attribute determining unit 101 determines information about TYPE attributes based on user settings or predetermined conditions. At step 52, the coding unit 102 encodes the stream based on the TYPE attribute information. In step 53, the reproduction assistance information generation unit 105 generates HLP assistance information based on the TYPE attribute information. Consequently, in step 54, the coding unit 102 generates access information based on the access unit of the encoded stream, and the generic management information generating unit 104 generates INFO management information by adding access information to other necessary information (general management information). At step 55, the multiplexing unit 106 multiplexes the stream, HLP support information, and INFO management information. At step 56, the memory unit 107 keeps multiplexed
<td>data. Belongs</td><td>note,</td><td>that</td><td>step 53</td><td>maybe</td><td>be</td>
<td>executed</td><td>before step 52 or</td><td>after</td><td>step 54.</td><td></td><td></td>
<td colspan="2">[0100] Note that</td><td>that</td><td>unit</td><td>coding</td><td> 102</td>
can store in the stream the information presented in the HLP support information. In this case, the information presented in the HLP support information is stored in the NAL unit for trick-play data storage. For example, if P pictures do not refer to B pictures, it is possible to decode only the I picture and P pictures during variable speed playback. Therefore, signaling information is stored indicating whether only the I picture and P pictures can be decoded and displayed. In addition, there is a case in which some AUs to be decoded during variable speed playback cannot obtain SPS or PPS from AUs to which the corresponding AUs should refer. This is the case in which the PPS referred to in the P picture is stored only in the B picture AU when decoding only the I picture and P pictures. In this case it is necessary to obtain the PPS necessary to decode the P picture from the B picture AU. Therefore, it is possible to include signaling information indicating whether the SPS or PPS to which each AU relates to be decoded during variable speed playback can certainly be obtained from one of the other AUs to be decoded during variable speed playback . In this way, it is possible to perform operations such as SPS or PPS detection also from the AU of the image not to be decoded during variable speed playback62 only if the marker is not set. In addition, when it turns out that only the I picture and P pictures can be decoded and displayed, it is possible to adjust the playback speed by decoding also the B pictures, in particular the reference B pictures to which other pictures refer.
[0101] Furthermore, it is possible to store the signaling information in the header of another NAL unit, such as SPS or PPS or a patch, instead of using any NAL unit for storing trick-play information. For example, in the case where the SPS to which the AU refers to forms the random access unit RAU is stored in the upper AU in the random access unit RAU, the nal_ref_idc field of the NAL SPS may indicate the signaling information. Since the nal_ref_idc value has been defined to be 1 or more in the SPS NAL unit, the high bit can always be set to 1 and the signaling information to be indicated by the low bit.
[0102] It should be noted that the content of the HLP support information may be recorded either in the stream or in administrative information, or in both places. For example, the content may be presented in the management information in the event that the content of the HLP ancillary information is fixed in the stream, while the content may be presented in the stream in the event that the content is variable. In addition, it is possible to keep signaling information indicating whether or not HLP support information is set in management information. Also, in the event that HLP ancillary information is predetermined in an application standard such as BD-ROM or RAM, or where HLP ancillary information is separately provided by communication or broadcasting, HLP ancillary information may not be recorded.
(Moving image decoding device) [0103] Fig. 27 shows a block diagram of a moving image decoding device 200 that implements the moving image decoding method of the present invention. This moving image decoding apparatus 200 reproduces the coded stream shown in Figs. 8A and 8B to Fig. 20. The device can not only perform standard playback, but also perform trickplay functions such as jump-in playback, variable speed playback and reverse playback. The moving image decoding apparatus 200 further includes an EXT stream extraction unit and AU units selection unit for decoding AUsel in addition to the units of the typical decoding apparatus 2 shown in Fig. 5. It should be noted that the processing units that perform the same operations as the corresponding processing units of the typical decoding device 2 shown in the block diagram (Fig. 5) are assigned the same reference numbers and their descriptions will be omitted.
[0104] AU units selection for decoding AUsel specifies AUs necessary for decoding based on the trick-play information GrpInf decoded in the variable length VLD decoding unit according to the trick-play instruction externally implemented. The trick-play instruction indicating trick-play information is introduced from the AU selection unit for AUSel decoding. In addition, the AU selection unit to be decoded AUsel notifies the EXT stream extraction unit of the DecAU, i.e., information indicating the AUs defined as AUs necessary for decoding. The EXT stream extraction unit only extracts the stream corresponding to AUs that are specified as AUs necessary to be decoded by the AUs unit selection for decoding AUsel, and then sends the stream to the VLD variable length decoding unit.
[0105] Fig. 28 is a flowchart showing the manner in which the moving image decoding apparatus 200 (mainly AU unit selection for decoding AUsel) shown in Fig. 27 performs the procedure of decoding a stream containing trick-play information while performing the trick- function play.
[0106] First, in step 20, the AU entity selection for decoding AUsel determines whether the AU is the top AU of the random access unit RAU by detecting SPS or the like in the stream. In the case where the AU is an upper AU, then proceeds to step 21; in turn, if the AU is not the upper AU, then proceed to step 22. The starting position of the random access unit RAU can be obtained from management information such as a time map. Especially when the start position of playback is specified during jump-in playback or only the upper image of the random access unit RAU is selected and high speed playback is performed on the selected upper image, it is possible to specify the start position of the random access unit RAU to the time map. In step 21, the AU entity selection unit to decode AUsel obtains trick-play information from the AU entity data, analyzes the AU entity data and determines the AU units to be decoded before moving to step 22. In step 22, the AU entity selection unit to decode AUsel determines whether the unit AU is an AU that is specified in step 21 as the AU to be decoded. In the case where the AU is determined in this way, the moving image decoding apparatus 200 decodes the AU in step 23; in turn, if the AU is not determined in this way, then proceeds to step 24. In step 24, the moving image decoding apparatus 200 determines if the AUs still need to be decoded. In the case where an AU exists, the moving image decoding apparatus 200 repeats the processing of step 20 and subsequent steps; in turn, if the AU does not exist, the process ends. Note that it is possible to skip processing of steps 21 and 22 or skip processing for determining in step 21, as well as sending information indicating that all AUs are decoded during standard playback when all AUs are decoded and displayed in order.
[0107] Fig. 29 is a flowchart showing the processing (processing by the AU of the AUSel decode unit) in step 21. First, the AU of the AUSel decode unit selection detects the start position of the NAL unit that creates the AU by searching the data AU for the start code prefix, starting from the top byte in step 30, and then proceeds to step 31. Note that the search of the start code prefix may start not from the upper byte of AU data, but from another position, such as the end position of the access units limiter. At step 31, the AU selection unit for AUSel decoding obtains the NAL unit type, followed by step 32. At step 32, the AU entity selection unit for AUSel decoding determines whether the NAL unit type obtained in step 31 is the NAL unit type for storing trick-play information. In case trick-play information is saved, proceed to step 33; in turn, if the trick-play information is not saved, the processing of step 30 and subsequent steps is repeated. Here, it is assumed that in the case where the trick-play information is stored in the SEI message, the AU entity selection unit for AUSel decoding obtains the NAL first, and then determines whether the SEI message for trick-play information storage is included in the NAL unit or not. At step 33, the AU selection unit to decode AUSel obtains trick-play information, followed by step 34. At step 34, the AU selection unit for AUSel decoding determines the images needed to be decoded during a particular trick-play operation. For example, when dual speed playback is specified. In the event that trick-play information indicates that it is possible to perform double-speed playback by decoding and reproducing only the I picture, P pictures and reference B pictures, it is determined that these three types of pictures will be decoded and played back. Note that if trick-play information is not detected in the upper image of a random access unit RAU during processing from step 30 to 32, the images needed to be decoded to perform a specific trick-play operation are determined in accordance with the determined in advance. For example, it is possible to determine if an image is a reference image or not by reference to the field indicating the image type in the access unit delimiter or by checking nal_ref_idc of the NAL unit header. For example, it is possible to distinguish reference B images from non-reference B images by referring to both the field indicating the image types and nal_ref_idc.
[0108] Fig. 30 is a flowchart showing the processing (processing by the AU of the AUSel decode units) in case all AUs to be decoded are not always displayed. The steps associated with the same processing as for the stages shown in Fig. 28 were assigned the same reference numbers and their descriptions will be omitted. In step 41, the AU unit selection for decoding, AUSel obtains and analyzes trickplay information, specifies the AUs to be decoded, and AUs to be displayed during a particular trick-play operation, followed by step 42. In step 42, the AU unit selection for AUSel decode determines whether the AUs to be decoded completely match the AUs to be displayed. If there is total agreement, then proceed to step 22; in turn, if there is no complete agreement, then proceeds to step 43. In step 43, the AU unit selection unit to decode AUSel sends the information in the form of a list of AU units to display, followed by the step 22. Information in the form of a list of sent units AUs are used in the step of determining AUs of decoded AUs.
[0109] Note that for MPEG-4 AVC it is possible to use as reference images images on which after decoding filtering (block separation) is carried out to remove block distortion to improve image quality, and it is also possible to use images from pre-cutting blocks as displayable images. In this case, the moving image decoding device 200 should store image data before and after the blocks are split. It is assumed here that, provided that the moving image decoding apparatus 200 has a memory that can save the data after decoding, which is equivalent to four images, in the case when it writes the image data before and after the blocks are divided into memory, the memory needs to be saved data, equivalent to two images, to preserve the images before cutting up the blocks of reference images. However, as described above, it is recommended that as many images as possible be saved in memory during reverse playback. While the moving image decoding apparatus 200 also uses block-split images also for display purposes, it can store four image data in memory because there is no need to save images before (not shown) for display with block cut. Therefore, displaying images before cutting blocks to improve image quality during playback in the standard direction, and displaying images after block cutting during reverse playback allows you to store more images in memory and reduce the amount of resources needed during reverse playback. For example, in the case shown in Fig. 15A to 15C, which present a list of AUs of I picture and P pictures as trick-play information, all data of the four pictures can be stored in memory during reverse playback, while the following sets of two pictures, which are freely selected from I0, P3, P6 and P9 can be stored in memory at the same time during playback in the normal direction: I0 and P3; P3 and
P6; and P6 and P9.
(Example of a trick-play recording format on an optical disc) [0110] The trick-play function is particularly important on optical disc devices that play packet media. The following describes an example of saving the trick-play information described above to a Blu-Ray (BD) disc, which is a new generation optical disc.
[0111] First, the BD-ROM recording format will be described.
[0112] Fig. 31 is a diagram showing the structure of the BD-ROM standard, in particular the structure of the BD 114 disk, which is a carrier, as well as data 111, 112 and 113 stored on the disk. The data stored on BD 114 includes AV 113 data, BD 112 management information, such as management information regarding AV data and AV playback sequence, as well as BD 111 playback program that implements interactivity. For the sake of convenience, the description of the BD disc will be presented with emphasis on the AV application for playing audio and visual content of films, but a similar description can be given with emphasis on another application.
[0113] Fig. 32 is a diagram showing the structure of logical data files stored on the BD disc described above. The BD disc has a recording area from its inner area along the radius to the outer area along the radius, such as DVDs or CDs, and also has a logical address space for storing data between read-in in the inner area along the radius and read-out on the outer area along the radius. In addition, there is a special area inside the read-in that can only be read by the BCA (Burst Cutting Area) drive. As this area cannot be read by the application, it can be used, for example, for copyright protection techniques.
[0114] File system information (volume) is stored at the top of the logical address space, and application data, such as video data, is also stored there. As shown in the prior art, the file system can be, for example, UDF or ISO9660, so that it is possible to read logical data stored using the directory structure or file structure as in the case of a typical computer.
[0115] In the present embodiment, as the directory structure and file structure on the BD disk, the BDVIDEO directory is located directly under the root directory (ROOT). This directory is a directory that stores data such as AV content or management information (101, 102 and 103, which are described in Fig. 32), which are supported by BD.
[0116] Below the BDVIDEO directory, the following 7 files are recorded.
(i) BD. INFO (the file name is immutable), which is part of the "BD management information" and is a file that stores information about the entire BD disc. The BD player reads this file first.
(ii) BD. PROG (the file name is immutable), which is one of the "BD playback programs" and is a file that stores playback control information about the entire BD disc.
(iii) XXX.PL (where the element "XXX" is variable and the extension "PL" is constant), which is part of the "BD management information" and is a file storing information about the playlist that makes up the scenario (playback sequence). Each playlist has a file (iv) XXX.PROG (where "XXX" is variable and "PROG" is fixed), which is one of
BD "of" playback programs is a playback file and storing control information prepared from a playlist. The corresponding playlist is identified by the file name (based on matching "XXX").
(<sub>V</sub>) YYY.VOB (where the "VYY" element is variable and the "VOB" extension is constant), which is part of the "AV data" and is a VOB storage file (the same as the VOB described in the prior art). Each VOB has a file ( vi) YYY.VOBI (where the element "YYY" is variable and the extension "VOBI" is constant), which is part of the "BD administrative information" and is a file that stores stream management information regarding VOB, i.e. AV data. The corresponding playlist is identified by the file name (based on matching "YYY").
(vii) ZZZ.PNG (where the "ZZZ" element is variable and the "PNG" extension is fixed), which is part of the "AV data" and is a file that stores PNG image data (image format developed by the W3C organization and called "ping ") for subtitles and menus. Each PNG image has its own file.
[0117] The structure of BD navigation data (BD management information) will be described based on Fig. 33 to Fig.
38.
[0118] Fig. 33 is a diagram showing the internal structure of a VOB management information file ("YYY.VO-BI"). The VOB management information has stream attribute information (Attribute) and a time map (TMAP). The stream attribute has a video attribute (Video) ) and separately with the audio attribute (Audio # 0 to Audio # m). Especially in the case of an audio stream, since the VOB has multiple audio streams at the same time, the presence or absence of a data field is signaled by the number (Number) of audio streams.
[0119] The following are video attributes (Video) stored in fields respectively, and values that the respective fields may have.
(i) compression format (Coding): MPEG-1; MPEG-2; MPEG4; and MPEG-4 AVC (Advanced Video Coding).
(ii) resolution: 1920 x 1080; 1440 x1080; 1280 x 720; 720x480; and 720 x 565.
(iii) Aspect ratio: 4 to 3 and 16 to 9.
(iv) frame rate (Framerate): 60;
59, 94 (60/1, 001); 50; thirty; 29, 97 (30 / 1,001); 25; 24 and
23, 976 (24/1,001).
[0120] The following are audio attributes (Audio) stored in fields respectively, and values that the respective fields may have.
(i) compression format (Coding): AC3; MPEG-1; MPEG-2 and LPCM.
(ii) number of channels (Ch): 1 to 8 (iii) language attribute (Language):
[0121] A time map (TMAP) is a table storing information based on VOBU, and includes a number of VOBUs that VOB has, and corresponding fragments of VOBU information (VOBU # 1 to VOBU # n). Corresponding portions of the VOBU information include I_start, i.e. the address (start address of the image I) of the upper TS VOBU packet, and the offset address (I_end) to the end address of the image I, as well as the start time (PTS) of the image I.
[0122] Fig. 34 is a diagram illustrating details of VOBU information. As is well known, as MPEG video stream can be compressed with variable bit rate to save the video stream in high quality, lack of proportionality between playback time and data size. On the other hand, in the case of constant-speed compression carried out in the AC3 standard, which is the audio compression standard, the relationship between time and address can be obtained from the original expression. However, for MPEG video data, each frame has a fixed display time, for example, the frame has a display time of 1 / 29.97 seconds in the NTSC standard, but the data size after compressing each frame changes significantly depending on the image content or type of image used in the compression (image I, image P or image B). Therefore, for an MPEG video stream, it is not possible to express the relationship between time and address using the original expression.
[0123] As would be expected, it is not possible to show the relationship between time and size of data using the original expression in the case of an MPEG system stream, where MPEG video data is multiplexed, i.e., VOB. Therefore, a time map (TMAP) associates time with an address in VOB.
[0124] In this way, when time information is specified, the VOBU to which the time belongs is searched first (in the order of the PTS VOBU time stamps), the PTS immediately before the specified time is entered into the VOBU located at TMAP map (address specified by I_start), decoding starts from the top picture and VOBU, and the display starts from the picture that corresponds to this time.
[0125] Next, the internal structure of the playlist information ("XXX.PL") will be described with reference to Fig. 35. The playlist information includes a cell list (CellList) and an event list (EventList).
[0126] The cell list (CellList) is a sequence of reproduction cells in a playlist, and the cells are reproduced in the order of description indicated in that list. The contents of the cell list (CellList) is the number of cells (Number) and information about each cell (Cell # 1 to Cell # n).
[0127] Cell information (Cell #) consists of the VOB file name (VOBName), start time (In) and end time (Out) in VOB, as well as subtitles (SubtitleTable). Start time (In) and end time (Out) are presented as frame numbers in each VOB. You can obtain the VOB data address necessary for playback using the time map (TMAP) described above.
[0128] The subtitle table (SubtitleTable) is a table storing subtitle information that is played back synchronously with the VOB. As with sound, subtitles have many language versions. The first information of the subtitle table (SubtitleTable) contains the number of languages (Number) and the following tables (Language # 1 to Language # k) based on the language.
[0129] Each language table (Language #) includes language information (Lang), number (Number) of subtitle information items to be displayed separately, and subtitle information (Speech # 1 to Speech # j) of subtitles to be displayed separately. Subtitle information (Speech #) includes the name of the image data file (Name), subtitle display start time (In), subtitle display end time (Out) and subtitle display position (Position).
[0130] The event list (EventList) is a table defining each event that occurs in the play list. The event list includes the number of events (Number) and the corresponding events (Event # 1 to Event # m). Each event (Event #) includes the event type (Type), event ID (ID), time of event occurrence (Time) and duration of event (Duration).
[0131] Fig. 36 shows an event handler table ("XXX. PROG") having an event handler (i.e. a time event and a user-related event selectable from a menu) prepared based on a play list. The event handler table includes the number of defined event handlers / programs (Number) and relevant event handlers / programs (Program # 1 to Program # n). The content of each event handler / program (Program #) is the definition of the beginning of the event handler (<event_handler> tag) and the event handler ID (ID) that matches the previously described event ID, as well as the program between the characters and and after the word Function. The event (Event # 1 to Event # m) stored in the event list (EventList) in the file "XXX. PL "is determined using the identifier (ID) of the event handler" XXX. PROG ".
[0132] Next, the internal structure of information for the entire BD disk ("BD. INFO") will be described with reference to Fig. 37. The information for the entire BD disk includes the title list (TitleList) and the event table regarding global events (EventList).
and fragments Each [0133] Title list (TitleList) includes the number of disk titles (Number) and fragments of title information (Title # 1 to Title # n) that follow the number of titles. The relevant parts of the title information (Title #) include the playlist table contained in the title (PLTable) and the chapter list in the title (ChapterList). The playlist table (PLTable) includes the number of playlists in the title (Number) and the names of the playlists (Name), which are the names of the playlist files.
[0134] The chapter list (ChapterList) includes the number of chapters contained in the title (Number) of chapter information (Chapter # 1 to Chapter # n) the fragment of chapter information (Chapter #) includes the cell table (CellTable) contained in the chapter, and the cell table (CellTable) containing the number of cells (Number), as well as fragments of information about cell entries (CellEntry # 1 to CellEntry # k). Information about cell entries (CellEntry #) includes the playlist name containing the cell and the cell number in the playlist.
[0135] The event list (EventList) includes the number of global events (Number) and pieces of information about global events. Note that the global event defined first is called the first event (FirstEvent), but the event dispatched first when the BD disc is placed in the player. Event information for global events includes only the event type (Type) and the event ID (ID).
[0136] Fig. 38 shows a table ("BD. PROG") of the global event handler program. The content of this table is the same as the content of the event handler table in Fig. 36.
[0137] When storing the trick-play information described above in the BD-ROM format described so far, it is assumed that the VOBU contains one or more units with random access RAU, and the trick-play information is in the upper AU of the unit VOBU. Note that the MPEG-4 AVC standard includes a NAL unit in which trick-play information is stored.
[0138] Note that trick-play information may be stored in BD management information. For example, it is possible to save trick-play information prepared on the basis of VOBU by expanding the time map of VOB management information. In addition, you can define a new map for storing trick-play information.
[0139] Furthermore, it is possible to save trick-play information either in VOBU or in BD management information.
[0140] In addition, it is possible to save only the default value of trick-play information in BD administrative information, and only if the trick-play information regarding VOBU is other than the default value, it is possible to save trick-play information in VOBU.
[0141] Furthermore, it is possible to save a set of one or more pieces of trick-play information in BD management information as information that is common to streams. VOBU may refer to one piece of trick-play information among the pieces of trick-play information stored in BD management information. In this case, the index information of the trick-play information to which the VOBU refers is stored in the management information of the VOBU or in this VOBU.
(Player for playing optical discs) [0142] Fig. 39 is a block diagram showing an outline of the functional structure of a player that plays the BD discs shown in Fig. 31 and the like. The data on the BD 201 disk is read by means of the optical head 202. The read data is sent to one of the memories depending on the type of data. The BD playback program (content "BD. PROG" or "XXX.PROG") is transferred to the memory of program 203. In addition, BD management information ("BD. INFO "," XXX. PL "or" YYY. VOBI ") are transferred to the management information memory 204. In addition, AV data (" YYY. VOB "or" ZZZ. PNG ") is transferred to the AV 205 memory.
[0143] The BD playback program stored in the program memory 203 is processed by the program processing unit 206. In addition, the BD management information stored in the management information memory 204 is processed by the management information processing unit 207. In addition, the AV data stored in the AV memory 205 is processed by the unit processing presentations 208.
[0144] The program processing unit 206 receives information about play lists to be reproduced by the management information processing unit 207, as well as event information, such as synchronization of program execution, and performs program processing. In addition, you can dynamically change the playlists to be played by the program. This can be accomplished by sending a play list reproduction instruction to the management information processing unit 207. The program processing unit 206 receives an event from a user, i.e. receives a request by a pilot, and in the event that a program suitable for the user event occurs, the program is executed.
[0145] The management information processing unit 207 receives the instruction from the program processing unit 206, analyzes the play lists and management information of the VOBs associated with the play lists, and also instructs the presentation processing unit 208 to play the AV target data. In addition, the management information processing unit 207 receives standard time information from the presentation processing unit 208, instructs the presentation processing unit 208 to stop playing AV data based on the time information. In addition, the management information processing unit 207 generates an event to inform the program processing unit 206 about program synchronization.
[0146] The presentation processing unit 208 has a decoder that can process image, sounds, subtitles / pictures (still images), respectively. The unit decodes and outputs AV data according to the processing unit
In the case of video data, decoding is also performed, followed by rendering in one of the planes: video plane 210 and image plane 209. Then the unit instructions information derived management 207.
subtitle / image, fusion processing 211 performs fusion processing on the video signal and sends the video signal to a display device such as a television.
[0147] When performing trick-play functions such as jump-in playback, variable-speed playback and reverse playback, the presentation processing unit 208 interprets the trick-play operation that is required by the user, and notifies the management information processing unit 207 by providing information such as playback speed. The management information processing unit 207 analyzes the trick-play information stored in the upper AU of the VOBU and determines the AU to be decoded and displayed to ensure that the user-defined trick-play operation is performed correctly. Note that the management information processing unit 207 may obtain trick-play information, send it to the presentation processing unit 208 and specify AUs to be decoded, as well as AUs to be displayed in the presentation processing unit 208.
[0148] It should be noted that the autonomous computer system can easily perform the processing described in this solution by saving the program implementing the moving image coding method and the moving image decoding method presented in the present embodiment on a recording medium such as a diskette.
[0149] Figs. 40A to 40C illustrate how a computer system implements a method of encoding a moving image and a method of decoding a moving image according to the present solution using a program recorded on a recording medium such as a floppy disk.
[0150] Fig. 40A shows an example of a physical diskette format as a recording medium. Fig. 40B shows the diskette, as well as the front view and cross section of the diskette. The diskette (FD) is located in the F housing, the tracks (Tr) are arranged concentrically on the disk surface from the outer area along the radius to the inner area along the radius of the disk, and each track is divided into 16 sectors (Se) in the angular direction. Therefore, for a floppy disk that stores the previously described program is saved in the area allocated on the floppy disk (FD).
[151] In addition, Fig. 40C shows a structure for recording and playing a program on a diskette. In the case of writing said program implementing a method of coding a moving image and a method of decoding a moving image on a FD diskette, the computer system Cs saves the program on a diskette using a disk drive. In addition, when constructing the described moving image coding device and the moving image decoding device implementing the moving image coding method and the method of moving image decoding using a diskette program, the program is read from a diskette using a disk drive and sent to a computer system.
[0152] Note that the above description was created using a floppy disk as a recording medium, but the program can be written to an optical disk. Furthermore, the recording medium is not limited to the above examples. Other recording media, such as an IC card or ROM cassette, can be used as long as they can save the program.
[0153] The above-mentioned apparatus for generating a moving picture stream, moving picture coding device, moving picture multiplexing device and moving picture decoding device according to the present invention have been described in the selected embodiment, but the present invention is not limited to this embodiment. The present invention includes variations that one skilled in the art could create based on this embodiment, these variations not departing from the scope of the present invention.
[0154] For example, the present invention includes the following elements in accordance with this embodiment: (i) a device for generating a moving image; an optical disc recording device having a moving image coding device or a moving image decoding device; a device for transmitting a moving image; digital television signal transmitting device; web server; communication device; a mobile information terminal and the like; and (ii) a moving image receiving device having a moving image decoding device; device receiving the digital television signal; communication device; mobile information terminal and the like.
[0155] It should be noted that the respective functional blocks shown in Fig. 21, Fig. 23, Fig. 27 and Fig. 39 are usually implemented as LSI, which is a large scale integration circuit. Each of the functional blocks may be made in the form of a single integrated circuit or a part or all of the functional blocks may be integrated in a single integrated circuit (for example, functional blocks except memory may be made in the form of a single integrated circuit). The integrated circuit, here referred to as LSI, may be called IC, system LSI, super LSI or ultra LSI, depending on the level of integration. In addition, the method of producing them in an integrated circuit is not limited to the method of production as LSI systems. Blocks can be implemented as a separate system or a typical processor. In addition, it is possible to use (i) a configurable processor in which the connection or settings of the circuit cells can be reconfigured, or (ii) the programmable FPGA (Field Programmable Gate Array) after conversion to LSI. In addition, when the technique for producing blocks in an integrated circuit instead of producing them in the form of LSIs occurs when the semiconductor technique or derivative technique is still developed at the right time, functional blocks can be produced in the form of an integrated circuit using the new technique. It is also possible to use biotechnics. In addition, among the respective functional blocks, the memory unit (image memory) in which the image data to be encoded or decoded is stored can be configured separately instead of being placed in a separate integrated circuit.
[0156] Although only the exemplary embodiment of the present invention has been described above, those skilled in the art will appreciate that numerous modifications of the exemplary embodiment are possible without significantly departing from the novel solutions and the benefits of the present invention. Therefore, all such modifications fall within the scope of the present invention.
Application in industry [0157] The present invention finds use as: a device for generating a moving picture stream that generates a moving picture for reproduction using the trick-play function; a moving picture coding apparatus that generates (using coding) a moving picture for reproduction using the trick-play function; a moving image multiplexing apparatus that generates a moving image for reproduction using trick-play functions using packet multiplexing; and a moving image decoding device that reproduces a moving image using a trick-play function, and in particular as a device for creating an MPEG-4 AVC stream playback system using trick-play mode, such as variable speed playback and reverse playback, what such a device is, for example, a device associated with an optical disk, to which the trick-play function refers essentially.
Panasonic Corporation
Proxy:
80 members in 13 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004134212 | Japan | A | |
| 2004134212 | Japan | A | |
| 2004165005 | Japan | A | |
| 2004165005 | Japan | A | |
| 2004251871 | Japan | A | |
| 2004251871 | Japan | A | |
| 05736693 | European Patent Office (EPO) | A | |
| 2005008319 | Japan | W | |
| 2005008319 | Japan | W | |
| EP20050736693 | – | – | – |
| JP20040134212 | – | – | – |
| JP20040165005 | – | – | – |
| JP20040251871 | – | – | – |
| WO2005JP08319 | – | – | – |
Members80
| Document | Office | Kind | |
|---|---|---|---|
| CA2542266A1 | Canada | A1 | |
| CA2811897A1 | Canada | A1 | |
| WO2005106875A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200601281A | Taiwan Province of China | A | |
| KR20070007028A | Republic of Korea | A | |
| EP1743338A1 | European Patent Office (EPO) | A1 | |
| CN1950907A | China | A | |
| JP2007535187A | Japan | A | |
| JP2007325289A | Japan | A | |
| JP4071811B2 | Japan | B2 | |
| US2008117988A1 | United States of America | A1 | |
| US2008118218A1 | United States of America | A1 | |
| US2008118224A1 | United States of America | A1 | |
| US2008131079A1 | United States of America | A1 | |
| EP1968063A1 | European Patent Office (EPO) | A1 | |
| US2008219393A1 | United States of America | A1 | |
| JP4185567B1 | Japan | B1 | |
| JP2008295055A | Japan | A | |
| EP1743338B1 | European Patent Office (EPO) | B1 | |
| ATE443327T1 | Austria | T1 | |
| DE602005016663D1 | Germany | D1 | |
| TW200949824A | Taiwan Province of China | A | |
| ES2330864T3 | Spain | T3 | |
| JP2010028860A | Japan | A | |
| CN101697575A | China | A | |
| CN101697576A | China | A | |
| PL1743338T3This record | Poland | T3 | |
| TWI324338B | Taiwan Province of China | B | |
| EP2182519A1 | European Patent Office (EPO) | A1 | |
| EP2182520A1 | European Patent Office (EPO) | A1 | |
| EP2182521A1 | European Patent Office (EPO) | A1 | |
| CN101707723A | China | A | |
| EP1968063B1 | European Patent Office (EPO) | B1 | |
| CN101778235A | China | A | |
| EP2207181A1 | European Patent Office (EPO) | A1 | |
| EP2207182A1 | European Patent Office (EPO) | A1 | |
| EP2207183A1 | European Patent Office (EPO) | A1 | |
| ATE471562T1 | Austria | T1 | |
| DE602005021926D1 | Germany | D1 | |
| US7809060B2 | United States of America | B2 | |
| ES2347385T3 | Spain | T3 | |
| US7843994B2 | United States of America | B2 | |
| JP4614991B2 | Japan | B2 | |
| JP2011019270A | Japan | A | |
| CN1950907B | China | B | |
| JP4763824B2 | Japan | B2 | |
| MY144441A | Malaysia | A | |
| MY145551A | Malaysia | A | |
| US8130843B2 | United States of America | B2 | |
| JP4910065B2 | Japan | B2 | |
| EP2182519B1 | European Patent Office (EPO) | B1 | |
| EP2182520B1 | European Patent Office (EPO) | B1 | |
| EP2207181B1 | European Patent Office (EPO) | B1 | |
| EP2207182B1 | European Patent Office (EPO) | B1 | |
| EP2207183B1 | European Patent Office (EPO) | B1 | |
| ATE555473T1 | Austria | T1 | |
| ATE555474T1 | Austria | T1 | |
| ATE555475T1 | Austria | T1 | |
| ATE555476T1 | Austria | T1 | |
| ATE555477T1 | Austria | T1 | |
| KR101148765B1 | Republic of Korea | B1 | |
| CN101697575B | China | B | |
| ES2383652T3 | Spain | T3 | |
| ES2383654T3 | Spain | T3 | |
| ES2383655T3 | Spain | T3 | |
| ES2383656T3 | Spain | T3 | |
| CN101697576B | China | B | |
| CA2542266C | Canada | C | |
| EP2207183B8 | European Patent Office (EPO) | B8 | |
| US8254446B2 | United States of America | B2 | |
| US8254447B2 | United States of America | B2 | |
| PL2182519T3 | Poland | T3 | |
| PL2182520T3 | Poland | T3 | |
| PL2207181T3 | Poland | T3 | |
| PL2207182T3 | Poland | T3 | |
| PL2207183T3 | Poland | T3 | |
| ES2388397T3 | Spain | T3 | |
| TWI395207B | Taiwan Province of China | B | |
| CN101778235B | China | B | |
| CA2811897C | Canada | C |
Numbers
- Publication, DOCDB
- 1743338
- Publication, EPODOC
- PL1743338T
- Application
- 736693
- Application, DOCDB
- 05736693
- Application, EPODOC
- PL20050736693T
Titles2
- English
- MOVING PICTURE STREAM GENERATION APPARATUS, MOVING PICTURE CODING APPARATUS, MOVING PICTURE MULTIPLEXING APPARATUS AND MOVING PICTURE DECODING APPARATUS
- Polish
- Urządzenie do generowania strumieni ruchomego obrazu, urządzenie do kodowania ruchomego obrazu, urządzenie do multipleksowania ruchomego obrazu oraz urządzenie do dekodowania ruchomego obrazu
Classification
- CPC, 18
- H04N9/8042
- G11B27/00
- G11B27/005
- G11B2220/2541
- H04N5/783
- H04N5/85
- H04N9/8063
- H04N9/8205
- H04N9/8227
- H04N21/42646
- H04N21/4325
- H04N21/4332
- H04N21/4334
- H04N21/44008
- H04N21/8451
- H04N21/85406
- H04N9/804
- G11B20/10
- IPC, 19
- G11B27 00
- G11B20 12
- H04N5 76
- H04N5 783
- H04N5 91
- H04N5 92
- H04N7 173
- H04N9 804
- H04N19 159
- H04N19 16
- H04N19 172
- H04N19 177
- H04N19 46
- H04N19 50
- H04N19 51
- H04N19 70
- H04N19 82
- H04N19 85
- H04N19 91