Image encoder, image encoding method, image decoder, image decoding method, and distribution media
Abstract
An image decoder for decoding an encoded bit string, produced by the encoding of an image formed by a sequence of objects, an object being encoded by intra-coding, an intra-video object plane (I-VOP), being a object encoded by intra-coding or by prediction coding forward, a prediction VOP (P-VOP), and being an object encoded by intra-coding, prediction coding forward, backward prediction coding or bidirectional prediction coding, a bidirectional prediction VOP (B-VOP), where said VOPs have been grouped into one or more groups (GOV), each group having an order of presentation associated, according to the which presents a plurality of decoded VOPs of the corresponding group when reproducing the image, and each or more groups comprise a group time code representing the absolute time corresponding to a synchronization point associated with a first object in the order of presentation of the corresponding group (GOV), the group time code comprising a value of hours_time_code representing a time unit of time, a value of minutes_time_code representing a unit of time in minutes, and a value of seconds_time_code representing a unit of time in seconds of the synchronization point, and each VOP of the group comprising time information with precision of seconds (base_time_ module) indicative of a time value in units of a second, and information Detailed time (increase_VOP_time) indicative of a time value in units with finer precision than a second, as information representative of a presentation time of said VOP, the image decoder comprising: receiving means for receiving said encoded bit string; a presentation time computer (36, 102) for calculating said presentation time of said VOP, adding said time information with precision of one second (module_time_base) and a detailed time information (VOP_time_ increment) of each VOP, with said code of group time, of the corresponding group; and means for decoding (72N) said VOP according to the corresponding presentation time calculated.

Term
Term ended
Projected expiry passed 31 March 2018, 8.5 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
3 claims: 2 independent, 1 dependent
- 1ES 2 323 482 T3 REIVINDICACIONES 1. Un descodificador de imágenes para descodificar una cadena de bits codificados, producida por la codificación de una imagen formada por una secuencia de objetos, siendo un objeto codificado por intra-codificación, un plano de objeto intra-vídeo (I-VOP), siendo un objeto codificado por intra-codificación o bien por codificación de predicción hacia delante, un VOP de predicción (P-VOP), y siendo un objeto codificado por intra-codificación, codificación de predicción hacia delante, codificación de predicción hacia atrás o codificación de predicción bidireccional, un VOP de predicción bidireccional (B-VOP), donde dichos VOP han sido agrupados en uno o más grupos (GOV), teniendo asociado cada grupo un orden de presentación, de acuerdo con el cual se presenta una pluralidad de VOP descodificados del correspondiente grupo al reproducir la imagen, y cada uno o más grupos comprende un código de tiempo del grupo que representa el tiempo absoluto correspondiente a un punto de sincronización asociado con un primer objeto en el orden de presentación del correspondiente grupo (GOV), comprendiendo el código de tiempo del grupo un valor de horas_código_tiempo que representa una unidad horaria de tiempo, un valor de minutos_código_tiempo que representa una unidad de tiempo en minutos, y un valor de segundos_código_tiempo que representa una unidad de tiempo en segundos del punto de sincronización, y comprendiendo cada VOP del grupo una información del tiempo con precisión de segundos (base_tiempos_ módulo) indicativa de un valor del tiempo en unidades de un segundo, e información detallada del tiempo (incremento_tiempo_VOP) indicativa de un valor del tiempo en unidades con precisión más fina que un segundo, como información representativa de un tiempo de presentación de dicho VOP, comprendiendo el descodificador de imágenes:medios receptores para recibir dicha cadena de bits codificados;un computador (36, 102) de tiempo de presentación para calcular dicho tiempo de presentación de dichos VOP, sumando dicha información de tiempo con precisión de un segundo (base_tiempos_módulo) y una información detallada del tiempo (incremento_tiempo_VOP) de cada VOP, con dicho código de tiempo de grupo, del grupo correspondiente;y medios para descodificar (72N) dichos VOP de acuerdo con el tiempo de presentación correspondiente calculado.
- 2Un descodificador de imágenes como se reivindica en la reivindicación 1, en el que dicho código de tiempo del grupo corresponde a un tiempo absoluto cuando se inicia la codificación de un grupo correspondiente de VOP.
- 3Un método de descodificación de imágenes para descodificar una imagen formada por una secuencia de objetos, siendo un objeto codificado mediante intra-codificación, un plano de objeto intra-vídeo (I-VOP), siendo un objeto codificado por intra-codificación o bien por codificación de predicción hacia delante, un VOP de predicción (P-VOP), y siendo un objeto codificado por intra-codificación, codificación de predicción hacia delante, codificación de predicción hacia atrás o codificación de predicción bidireccional, un VOP de predicción bidireccional (B-VOP), donde dichos VOP han sido agrupados en uno o más grupos (GOV), teniendo asociado cada grupo un orden de presentación, de acuerdo con el cual se presenta una pluralidad de VOP descodificados del correspondiente grupo al reproducir la imagen, y cada uno o más grupos comprende un código de tiempo del grupo que representa el tiempo absoluto correspondiente a un punto de sincronización asociado con un primer objeto en el orden de presentación del correspondiente grupo (GOV), comprendiendo el código de tiempo del grupo un valor de horas_código_tiempo que representa una unidad horaria de tiempo, un valor de minutos_código_tiempo que representa una unidad de tiempo en minutos, y un valor de segundos_código_tiempo que representa una unidad de tiempo en segundos del punto de sincronización, y comprendiendo cada VOP del grupo una información del tiempo con precisión de segundos (base_tiempos_ módulo) indicativa de un valor del tiempo en unidades de un segundo, e información detallada del tiempo (incremento_tiempo_VOP) indicativa de un valor del tiempo en unidades con precisión más fina que un segundo, como información representativa de un tiempo de presentación de dicho VOP, comprendiendo dicho método de descodificación los pasos de:recibir dicha cadena de bits codificados ES 2 323 482 T3 calcular dicho tiempo de presentación de dichos VOP en un computador (36, 102) de tiempo de presentación, sumando dicha información de tiempo con precisión de un segundo (base_tiempos_módulo) y una información detallada del tiempo (incremento_tiempo_VOP) de cada VOP, con dicho código de tiempo de grupo, del grupo correspondiente;y descodificar (72N) dichos VOP de acuerdo con el tiempo de presentación correspondiente calculado.
Independent claims3
376 paragraphs in 19 sections, as filed
ES 2 323 482 T3
DESCRIPTION
Image encoding device, image encoding method, image decoding device, image decoding method, and supplying medium.
Technical field
The present invention relates to an image encoder, an image encoding method, an image decoder, an image decoding method, and distribution means. More particularly, the invention relates to an image encoder, an image encoding method, an image decoder, an image decoding method and distribution means suitable for use, for example, in the case where record dynamic image data on storage media, such as magneto-optical disk, magnetic tape, etc., and also regenerate and display the recorded data on a screen, or in the case where the dynamic image data is transmitted from a transmitting side to a receiving side, through a transmission path, and, on the receiving side, the received dynamic image data is presented, edited and record, as in videoconferencing systems, videophone systems, broadcast equipment, and multimedia database retrieval systems.
Previous technique
For example, as in video conferencing systems and videophone systems, in systems that transmit dynamic image data to a remote location, the image data is compressed and encoded taking advantage of inter-line correlation and inter-frame correlation, in order to efficiently take advantage of transmission paths.
As a representative high-efficiency dynamic image coding system, there is a dynamic image coding system for storage media, based on the Moving Picture Expert Group (MPEG) standard. This MPEG standard has been studied by the International Organization for Standardization (ISO) - IEC / JTC1 / SC2 / WG11 and has been proposed as a proposed standard. The MPEG standard has adopted a hybrid system that uses a combination of motion compensation prediction coding and discrete cosine transform (DCT) coding.
The MPEG standard defines some profiles and levels in order to support a wide range of applications and functions. The MPEG standard is mainly based on the Main Profile and Main Level (MP @ ML).
Figure 1 illustrates the constitution example of an MP @ ML encoder in the standard MPEG system.
The image data to be encoded is input into a frame memory 31 and temporarily stored. A motion vector detector 32 reads the image data stored in the frame memory 31, for example, in a macroblock unit consisting of 16 (16 pixels and detects the motion vectors.
In this case, the moving vector detector 32 processes the image data for each frame as either an intracoded image (I-image), a forward prediction encoding image (P-image), or a bidirectional prediction encoding image. (image-B). Note that the way in which I, -P, and -B images of input frames are processed has been predetermined (that is, images are processed as I-image, B-image, P-image, B-image, B-image. P, ...., B-image, and P-image, in the order mentioned).
That is, in the motion vector detector 32, reference is made to a predetermined reference frame in the image data stored in the frame memory 31, and a small block of 16 pixels (16 lines (macroblock)) in the current frame to be encoded, is matched with a set of blocks of the same size in the reference frame. With block adaptation, the motion vector of the macroblock is detected.
In this case, in the MPEG standard, the prediction modes for an image include four classes: intracoding, forward prediction coding, backward prediction coding, and bidirectional prediction coding. An I-image is encoded by intracoding. A P-picture is encoded by intracoding or by prediction forward coding. A B-picture is encoded by intracoding, forward prediction coding, backward prediction coding, or bidirectional prediction coding.
That is, the motion vector detector 32 sets the intracoding mode to an I-picture as a prediction mode. In this case, the motion vector detector 32 delivers the prediction mode (intracoding mode) to a variable word length (VLC) encoding unit 36 and compensator 42, without detecting the motion vector.
Motion vector detector 32 also performs forward prediction of a P-image and detects the motion vector. Furthermore, at motion vector detector 32, a prediction error caused by making a forward prediction is compared with the spread, for example, of macroblocks to be encoded (macroblocks
ES 2 323 482 T3 in the P-image). As a result of the comparison, when the spread of the macroblocks is less than the prediction error, the motion vector detector 32 sets an intracoding mode as the prediction mode and delivers it to the VLC unit 36 and the VLC compensator 42. movement. In addition, if the prediction error caused by performing the forward prediction is smaller, the motion detector 32 sets a forward prediction coding mode as the prediction mode. The forward prediction encoding mode, along with the detected motion vector, is delivered to VLC unit 36 and motion compensator 42.
Motion vector detector 32 further performs forward prediction, backward prediction, and bidirectional prediction for a B-image, and detects the respective motion vectors. Then, the motion vector detector 32 detects the minimum error between the prediction errors in the forward prediction, the backward prediction, and the bidirectional prediction (hereinafter referred to as the minimum prediction error when needed), and compares the minimum prediction error with the spread, for example, of macroblocks to be encoded (macroblocks in B-image). As a result of the comparison, when the spread of the macroblocks is less than the minimum prediction error, the motion vector detector 32 sets an intracoding mode as the prediction mode, and delivers it to the VLC unit 36 and the compensator. 42 movement. Furthermore, if the minimum prediction error is smaller, the motion vector detector 32 sets as a prediction mode a prediction mode in which the minimum prediction error was obtained. The prediction mode, along with the corresponding motion vector, is delivered to VLC unit 36 and motion compensator 42.
If motion compensator 42 receives the prediction mode and motion vector from motion vector detector 32, motion compensator 42 will read the locally encoded and previously decoded image data stored in frame memory 41 from according to the received prediction mode and the motion vector. This read image data is supplied to arithmetic units 33 and 40 as predicted image data.
The arithmetic unit 33 reads from the memory 31 the same macroblock as the image data read from the frame memory 31 by the motion vector detector 32, and calculates the difference between the macroblock and the predicted image that was supplied from the compensator. 42 movement. This differential value is supplied to the DCT unit 34.
On the other hand, in the case where only one prediction mode is received from the motion vector detector 32, that is, the case where the prediction mode is an intracoding mode, the motion compensator 42 does not deliver a predicted image. In this case, the arithmetic unit 33 (as well as the arithmetic unit 40) delivers to the DCT unit 34 the macroblock read from the frame memory 31, without processing it.
In the DCT unit 34, the DCT is applied to the output data of the arithmetic unit 33, and the resulting coefficients of the DCT are supplied to the quantizer 35. In the quantizer 35, a quantization step is set (quantization scale ), corresponding to the amount of data stored in buffer 37 (which is the amount of data stored in buffer 37) (buffer feedback). In the quantization step, the DCT coefficients are quantized from the DCT unit 34. The quantized DCT coefficients, (hereinafter referred to as quantized coefficients when needed), along with the set quantization step, are supplied to the VLC unit 36.
In VLC unit 36, the quantized coefficients supplied by quantizer 35 are transformed into variable length word codes, such as Huffman codes, and delivered to buffer 37. In addition, in the VLC unit 36, the quantization step from the quantizer 35 is encoded by variable-length word encoding, and likewise, the prediction mode (which indicates whether it is an intracoding (prediction intracoding) is encoded. image), forward prediction coding, backward prediction coding, or bidirectional prediction coding), and the motion vector of the motion vector detector 32. The resulting encoded data is delivered to buffer 37.
Buffer 37 temporarily stores encoded data supplied from VLC unit 36, thereby smoothing the amount of stored data. For example, smoothed data is delivered to a transmission path or recorded on a storage medium, as a string of coded bits.
The buffer 37 also outputs the amount of stored data to the quantizer 35. The quantizer 35 sets a quantization step in correspondence with the amount of stored data that is delivered by this buffer 37. That is, when there is a possibility that the capacity of the buffer 37 overflows, the quantizer 35 increases the size of the quantization step, thereby reducing the amount of quantized coefficient data. When there is a possibility that the capacity of the buffer 37 goes into an underflow state, the quantizer 35 reduces the size of the quantization step, thereby increasing the amount of data of the quantized coefficients. In this way, overflow and underflow of buffer 37 is prevented.
The quantized coefficients and the quantization step, delivered by the quantizer 35, are supplied not only to the VLC unit 36, but also to an inverse quantizer 38. In the inverse quantizer 38, the quantized coefficients from the quantizer 35 are inversely quantized. according to the quantization step provided by the quantizer 35, whereby the quantized coefficients are transformed into
ES 2 323 482 T3 DCT coefficients. The DCT coefficients are provided to an inverse DCT unit (IDCT unit) 39. In the IDCT
39, an inverse DCT is applied to the DCT coefficients and the resulting data is supplied to the arithmetic unit
40.
In addition to the output data from IDCT unit 39, the same predicted image data supplied to arithmetic unit 33 is supplied from motion compensator 42 to arithmetic unit 40, as described above. The arithmetic unit 40 adds the output data (residual prediction (differential data)) from the IDCT unit 39 and the predicted image data from the motion compensator 42, thereby decoding the original image data locally. The locally decoded image data is delivered to the output. (However, in the case where the prediction mode is an intracoding mode, the output data of the IDCT 39 is transferred through the arithmetic unit 40 and supplied to the frame memory 41, as decoded image data. locally, without being processed). Note that this decoded image data is consistent with the decoded image data obtained at the receiver side.
Decoded image data obtained in arithmetic unit 40 (locally decoded image data) is supplied to and stored in frame memory 41. Thereafter, the decoded image data is used as reference image data (reference frame) with respect to an image to which intracoding is applied (forward prediction coding, backward prediction coding, or backward prediction coding). bidirectional prediction).
Next, Figure 2 illustrates the constitution example of an MP @ ML decoder in the standard MPEG system that decodes the encoded data output from the encoder of Figure 1.
The string of encoded bits (encoded data) transmitted through a transmission path is received by a receiver (not illustrated), or the string of encoded bits (encoded data) recorded on a storage medium is regenerated by a regenerator ( not illustrated). The string of received or regenerated bits is supplied and stored in a buffer 101.
A reverse VLC unit (VLC (Variable-Length Word Decoder) Unit) 102, reads the encoded data from the buffer 101 and performs variable-length word decoding, thereby separating the encoded data into the motion vector, prediction mode, quantization step, and quantization coefficients in a macroblock unit. Among them, the motion vector and the prediction mode are supplied to a motion compensator 107, while the quantization step and the quantized coefficients of the macroblock are supplied to an inverse quantizer 103.
In the inverse quantizer 103, the quantized coefficients of the macroblock, supplied from the IVLC unit 102, are inversely quantized according to the quantization step supplied from the same IVLC unit 102. The resulting DCT coefficients are fed to an IDCT unit 104. In IDCT 104, an inverse DCT is applied to the macroblock DCT coefficients, supplied from the inverse quantizer 103, and the resulting data is supplied to an arithmetic unit 105.
In addition to the output data from the IDCT unit 104, the output data from the motion compensator 107 is also supplied to the arithmetic unit 105. That is, in the motion compensator 107, as in the case of the motion compensator 42 of FIG. 1, the previously decoded image data, stored in the frame memory 106, is read in accordance with the motion vector and the prediction mode supplied from IVLC unit 102, and supplied to arithmetic unit 105 as predicted image data. The arithmetic unit 105 sums the output data (residual prediction (differential value)) from the IDCT unit 104 and the predicted image data from the motion compensator 107, thereby decoding the original image data. This decoded image data is supplied to and stored in frame memory 106. Note that in the case where the output data from the IDCT unit 104 is intracoded data, the output data is transferred through the arithmetic unit 105 and supplied to the frame memory 106 as decoded image data, without being processed.
The decoded image data stored in the frame memory 106 is used as reference image data for the next image data to be decoded. Furthermore, the decoded image data is supplied, for example, to a screen (not illustrated) and displayed as a reproduced output image.
Note that in the MPEG-1 standard and in the MPEG-2 standard, a B-image is not stored in the decoder frame memory 41 (figure 1) and the decoder frame memory 106 (figure 2), because no it is used as reference image data.
The aforementioned encoder and decoder illustrated in Figures 1 and 2 are based on the MPEG-1/2 standard. Currently, a system for encoding video in a video object unit (VO) of a sequence of objects constituting an image is being standardized as an MPEG-4 standard by ISOIEC / JTC1 / SC29 / WG11.
By the way, as the MPEG-4 standard is being standardized under the assumption that it is mainly used in the field of communications, it does not prescribe the group of pictures (GOP) prescribed in the MPEG standard.
ES 2 323 482 T3
1/2. Therefore, in the case where the MPEG-4 standard is used in storage media, efficient random access will be difficult.
According to one aspect, the present invention provides an image decoder for decoding a string of coded bits, produced by encoding an image formed by a sequence of objects, with an object coded by intracoding, this being an intracoding plane of objects. video (I-VOP), an intracoding or forward prediction encoding encoded object that is a prediction VOP (P-VOP), and an intracoding encoded object, forward prediction coding, backward prediction coding, or bidirectional prediction coding that is a bidirectional prediction VOP (B-VOP), where said VOPs have been grouped into one or more groups (GOV), each group having a screen associated according to which a plurality of decoded VOPs of the corresponding group is displayed, when reproducing the image, and each of the one or more groups comprises a group time code representing an absolute time corresponding to a synchronism point associated with a first object in the corresponding group presentation order (GOV), comprising the time code of the group a time_code_hours value that represents an hourly unit of time, a timecode_minutes value that represents a unit in minutes of time, and a value of seconds_time_code that represents a unit in seconds of time of the synchronism point, and each VOP of the group comprising time information with precision of seconds (time_base_module) indicative of a value of time in units of a second, and detailed information of time (VOP_time_increment) indicative of a time value in units of precision finer than one second, as information representing a presentation time of said VOP, comprising the image decoder:
receiving means for receiving said encoded bit stream;
a presentation time computer to calculate said presentation time of said VOPs, adding said time information with precision of seconds (time_base_module) and detailed time information (VOP_time_increment) of each VOP to said group time code of the corresponding group ; and means for decoding said VOPs in accordance with the corresponding calculated presentation time.
Various other aspects of the invention are detailed in the appended claims.
The present invention has been made in view of such circumstances, and the embodiments of the invention aim to make efficient random access possible.
Figure 1 is a block diagram illustrating a constitution example of a conventional encoder;
Figure 2 is a block diagram illustrating a constitution example of a conventional decoder;
Figure 3 is a block diagram illustrating a constitution example of an embodiment of an encoder to which the present invention is applied;
Fig. 4 is a diagram for explaining that the position and size of a video object (VO) vary with time;
Figure 5 is a block diagram illustrating an example of constitution of the VOP coding sections 31 to 3N of Figure 3;
Figure 6 is a diagram for explaining spatial scalability;
Figure 7 is a diagram for explaining spatial scalability;
Figure 8 is a diagram for explaining spatial scalability;
Figure 9 is a diagram for explaining spatial scalability;
Fig. 10 is a diagram for explaining a method for determining the size data and compensation data of a video object plane (VOP);
FIG. 11 is a block diagram illustrating the constitution example of the base layer encoding section 25 of FIG. 5;
ES 2 323 482 T3
FIG. 12 is a block diagram illustrating the constitution example of the reinforcing layer coding section 23 of FIG. 5;
Figure 13 is a diagram for explaining spatial scalability;
Fig. 14 is a diagram for explaining time scalability;
Figure 15 is a block diagram illustrating the constitution example of an embodiment of a decoder to which the present invention is applied;
Figure 16 is a block diagram illustrating another example of constitution of sections 72<sub>2</sub> to 72<sub>N</sub> VOP decoding box of FIG. 15;
FIG. 17 is a block diagram illustrating the constitution example of the base layer decoding section 95 of FIG. 16;
Fig. 18 is a block diagram illustrating the constitution example of the decoding section 93 of the reinforcing layer of Fig. 16;
Figure 19 is a diagram illustrating the syntax of a bit string obtained by scalable encoding;
Figure 20 is a diagram illustrating the syntax of VS;
Figure 21 is a diagram illustrating the syntax of VO;
Figure 22 is a diagram illustrating the syntax of VOL;
Figure 23 is a diagram illustrating the syntax of VOP;
FIG. 24 is a diagram illustrating the relationship between module_time_base and VOP_time_increment time;
Figure 25 is a diagram illustrating the syntax of a bit string, in accordance with the present invention;
Figure 26 is a diagram illustrating the syntax of a GOV;
Figure 27 is a block diagram illustrating the constitution of a time_code;
FIG. 28 is a diagram illustrating a method of encoding the GOV layer time_code and the module_time_base and VOP_time_increment of the first I-VOP of the GOV;
FIG. 29 is a diagram illustrating a method of encoding the GOV layer time_code and also the module_time_base and VOP_time_increment of the B-VOP located before the first I-VOP of the GOV;
Fig. 30 is a diagram illustrating the relationship between modulo_time_base and VOP_time_increment, when their definitions do not change;
Fig. 31 is a diagram illustrating a process of encoding the modulo_time_base and VOP_time_increment of the B-VOP, based on a first method;
Fig. 32 is a flowchart illustrating a process of encoding the I / P-VOP_time_base and VOP_time_increment, based on a first method and a second method;
FIG. 33 is a flowchart illustrating a process of encoding the module_time_base and VOP_time_increment of the B-VOP, based on a first method;
Fig. 34 is a flowchart illustrating a decoding process of the I / P-VOP_time_base_time_base and I / P-VOP_time_increment, encoded by the first and second methods;
FIG. 35 is a flowchart illustrating a decoding process of the module_time_base and VOP_time_increment of the B-VOP, encoded by the first method;
FIG. 36 is a diagram illustrating a process of encoding the modulo_time_base and VOP_time_increment of the B-VOP, based on a second method;
Fig. 37 is a flowchart illustrating the encoding process of the module_time_base and VOP_time_increment of the B-VOP, based on the second method;
ES 2 323 482 T3
Fig. 38 is a flowchart illustrating a decoding process of the module_time_base and VOP_time_increment of the B-VOP, encoded by the second method;
Fig. 39 is a diagram for explaining the module_time_base; Y
Fig. 40 is a block diagram illustrating the constitution example of another embodiment of an encoder and a decoder, to which the present invention is applied.
Best mode of carrying out the invention
The embodiments of the present invention will now be described in detail with reference to the drawings.
Figure 3 shows the constitution example of an embodiment of an encoder, to which the present invention is applied.
The image data (dynamic data) to be encoded is input into a video object constitution (VO) section 1. In VO constitution section 1, the image is constituted, for each object, by a VO sequence. The sequence of the VOs is delivered to sections 21 to 2N of VOP constitution. That is, in section 1 of VO constitution, in the case in which N video objects are generated (VO # 1 to VO # N), the VO # 1 to VO # N are delivered to sections 21 to 2N of constitution of VOP, respectively.
More specifically, for example, when the image data to be encoded is constituted by a sequence of an independent second plane F1 and a first plane F2, the VO constitution section 1 delivers the first plane F2, for example, to section 21 VOP constitution as VO # 1 and also enter the second plane F1 to VOP constitution section 22 as VO # 2.
Note that, in the case where the image data to be encoded is, for example, an image previously synthesized by the second plane F1 and the first plane F2, the VO constitution section 1 divides the image in the second plane F1 and the background F2, according to a predetermined algorithm. The second plane F1 and the first plane F2 are delivered to the corresponding VOP constitution sections 2n (where n = 1, 2, ... and N). The VOP building sections 2n generate VO planes (VOP) from the outputs of the VO building section 1. That is, for example, an object is extracted from each frame. For example, the minimum rectangle surrounding the object (hereinafter referred to as the minimum rectangle, when needed), is considered to be the VOP. Note that, at this moment, the VOP constitution sections 2n generate the VOP, so that the number of horizontal pixels and the number of vertical pixels are multiples of 16. If the VO constitution sections 2n generate VOP, the VOPs are delivered to VOP coding sections 3n, respectively.
Furthermore, VOP constitution sections 2n detect size data (VOP size) indicating the size of a VOP (e.g. horizontal and vertical lengths) and offset data (VOP offset) indicating the position of the VOP. VOP in a frame (for example, coordinates where the upper left of a frame is the origin). The size data and the offset data are also supplied to the VOP coding sections 3n.
The VOP coding sections 3n encode the outputs of the VOP constitution sections 2n, for example, with a method based on the MPEG standard or the H.263 standard. The resulting bit strings are delivered to a multiplexing section 4 which multiplexes the bit strings obtained from VOP encoding sections 31 to 3N. The resulting multiplexed data is transmitted over ground waves or through a transmission path, such as a satellite line, a CATV network, etc. Alternatively, the multiplexed data is recorded on a storage medium 6, such as a magnetic disk, a magneto-optical disk, an optical disk, a magnetic tape, etc.
In this case, a description will be made of the video object (VO) and the video object plane (VOP).
In the case of a synthesized image, each of the images that make up the synthesized image is called VO, while VOP means a VO at a given moment.
That is to say, for example, in the case of a synthesized image F3, constituted by the images F1 and F2, when the image F1 and F2 are arranged in the form of a time series, they are VO. Image F1 or F2 at any given time is a VOP. Therefore, it can be said that VO is a set of VOPs of the same object, at different times.
For example, if it is assumed that the image F1 is the background and also that the image F2 is the foreground, the synthesized image F3 will be obtained by synthesizing the images F1 and F2 with a key signal to extract the image F2. The VOP of the F2 image in this case is assumed to include the key signal in addition to the image data (luminance signal and color difference signal) that make up the F2 image.
ES 2 323 482 T3
A frame of an image does not vary in size or position, but there are cases where the size or position of a VO changes. That is, even in the case where a VOP constitutes the same VO, there are cases where the size and position vary with time.
Specifically, Figure 4 illustrates a synthesized image made up of image F1 (background) and image F2 (foreground).
For example, suppose that image F1 is an image obtained by photographing a certain natural scene, and that the entire image is a single VO (eg, VO # 0). Suppose also that image F2 is an image obtained by photographing a person who is walking and that the minimum rectangle surrounding the person is a single VO (eg, VO # 1).
In this case, since VO # 0 is the image of a scene, basically both the position and the size do not change as in a normal image frame. On the other hand, since VO # 1 is the image of a person, the position or size will change if the person scrolls left and right or moves to this side or depth side of Figure 4. Thus, although Figure 4 shows VO # 0 and VO # 1 at the same time, there are cases where the position or size of the VO varies over time.
Therefore, the output bit string of the VOP coding sections 3n of Figure 3 includes information on the position (coordinates) and the size of a VOP in a predetermined absolute coordinate system, in addition to the data indicating a coded VOP. Note in figure 4, that the vector that indicates the position of the VOP of VO # 0 (image F1) at a certain instant, is represented by OST0 and also a vector that indicates the position of the VOP of VO # 1 (image F2) at a certain instant, it is represented by OST1.
Next, Figure 5 illustrates the constitution example of the VOP coding sections 3n of Figure 3, which perform scalability. That is, the MPEG standard introduces a scalable encoding method that effects scalability, comprising different image sizes and frame rates. The VOP coding sections 3n illustrated in Figure 5 are constructed so that such scalability can be realized.
The VOP (image data), size data (VOP size), and compensation data (VOP compensation of the VOP constitution sections 2n are all supplied to an image layering section 21.
The image layering section 21 generates one or more layers of image data from the VOP (VOP layering is performed). That is, for example, in the case of performing spatial scalability coding, the image data input into the image layering section 21, as is, is delivered as a reinforcing layer of the image data. At the same time, the number of pixels constituting the image data is reduced (the resolution is lowered) by making the pixels smaller, and the image data reduced in the number of pixels is delivered as a base layer of the image data. .
Note that an input VOP can be used as a database layer and also the VOP increased in number of pixels (resolution) by some other method can be used as a data reinforcement layer.
Also, although the number of layers can be made equal to 1, this case cannot effect scalability. In this case, the VOP coding sections 3n are constituted, for example, only by a base layer coding section 25.
Also, the number of layers can be made equal to 3 or more. But in this embodiment, the two-layer case will be described for simplicity.
For example, in the case of performing time scalability encoding, the image layering section 21 alternately delivers image data, for example, as base layer data or backing layer data, in correspondence with time. That is, for example, when it is assumed that the VOPs that constitute a certain VO are entered in order, VOP0, VOP1, VOP2, VOP3, ..., the image layering section 21 outputs VOP0, VOP2, VOP4, VOP6, ..., as base layer data and VOP1, VOP3, VOP5, VOP7, ..., as backing layer data. Note that, in the case of temporal scalability, the VOPs thus dwarfed are delivered merely as base layer data and backing layer data, and increasing or reducing the image data (resolution conversion) is not performed. (but it is possible to make the increase or decrease).
Also, for example, in the case of performing signal-to-noise ratio (SNR) scalability encoding, the image data input into the image layering section 21, as is, is delivered as data from the reinforcement layer or base layer data. That is, in this case, the base layer data and the backing layer data are consistent with each other.
In this case, for spatial scalability in the case where an encoding operation is performed for each VOP, there are, for example, the following three classes.
ES 2 323 482 T3
That is, for example, if it is now assumed that a synthesized image consisting of images F1 and F2, such as the one illustrated in figure 4, is input as a VOP, in the first spatial scalability, the full input VOP (figure 6 (A)) is considered to be a backing layer, as illustrated in Figure 6, and the entire reduced VOP (Figure 6 (B)) is considered to be a base layer.
Also, in the second spatial scalability, as illustrated in Figure 7, an object constituting part of an input VOP (Figure 7 (A) (corresponding to image F2)) is extracted. The extracted object is considered to be in a reinforcing layer, while all the reduced VOP (Figure 7 (B)) is considered to be a base layer. (Such extraction is carried out, for example, in the same way as in the case of sections 2n of constitution of the VOPs. Therefore, the extracted object is also a single VOP).
Furthermore, in the third scalability, as illustrated in Figures 8 and 9, the objects (VOP) constituting an input VOP are extracted, and a backing layer and a base layer are generated for each object. Note that figure 8 shows a reinforcement layer and a base layer generated from the second plane (image F1) that constitutes the VOP illustrated in figure 4, while figure 9 shows a reinforcement layer and a base layer generated by starting from the first plane (image F2) that constitutes the VOP illustrated in figure 4.
It has been predetermined which of the aforementioned scalabilities is used. The image layering section 21 performs layering of a VOP, so that encoding can be performed according to a predetermined scalability.
Furthermore, the image layering section 21 calculates (or determines) the size data and the compensation data of the generated base and reinforcement layers, from the size data and the compensation data of the input VOP ( hereinafter referred to as initial size data and initial offset data, respectively, when required). Offset data indicates the position of a base or backing layer in a predetermined absolute coordinate system of the VOP, while size data indicates the size of the base or backing layer.
In this case, a method for determining the compensation data (position information) and the size data of the VOPs in the base and reinforcement layers will be described, for example, in the case where the aforementioned second scalability is performed. (figure 7).
In this case, for example, the compensation data of a base layer, FPOS_B, as illustrated in Fig. 10 (A), is determined such that, when the image data of the base layer is scaled up (oversampling) based on in the difference between the resolution of the reinforcing layer, that is, when the image of the base layer is enlarged with a magnification ratio such that the size is consistent with that of the image of the reinforcement layer (which is the reciprocal of the demagnification ratio because the image of the base layer is generated by reducing the image of the reinforcement layer) (hereinafter referred to as FR magnification, when needed), the compensation data of the enlarged image in the absolute coordinate system is consistent with the initial compensation data. The data of the size of the base layer, FSZ_B, is also determined such that the data of the size of an enlarged image, obtained when the image of the base layer is enlarged with the magnification FR, is consistent with the initial data of the size. That is, the compensation data FPOS_B is determined to be FR times themselves, or consistent with the initial compensation data. Also, the FSZ_B size data is determined in the same way.
On the other hand, for the FPOS_E compensation data of a reinforcement layer, the coordinates of the upper left corner of the minimum rectangle (VOP) surrounding an object extracted from an input VOP, for example, are computed based on the initial data offset, as illustrated in Fig. 10 (B), and this value is determined as offset data FPOS_E. Also, the FPOS_E size data of the backing layer is determined with the horizontal and vertical lengths, for example, of the minimum rectangle surrounding an object extracted from an input VOP.
Therefore, in this case, the compensation data FPOS_B and the size data FPOS_B of the base layer are first transformed according to the magnification FR. (The FPOS_B offset data and the FPOS_B size data after transformation are called the transformed FPOS_B offset data and the size FPOS_B size data, respectively.) Then, at a position corresponding to the FPOS_B offset data transformed in the absolute coordinate system, consider a frame of an image of the size corresponding to the transformed size data FSZ_B. If an enlarged image is arranged, obtained by enlarging the image data on the FR base layer times, in the aforementioned corresponding position (Figure 10 (A)) and furthermore, if the image of the reinforcing layer is also arranged in the system of absolute coordinates, according to the FPOS-E compensation data and the FPOS_E size data of the reinforcement layer (figure 10 (B)), the pixels constituting the enlarged image and the pixels constituting the image in the backing layer will be arranged so that the mutually corresponding pixels are located in the same position. That is, for example, in Figure 10, the person in the reinforcing layer and the person in the enlarged image will be arranged in the same position.
Even in the case of the first scalability and the third scalability, the FPOS_B offset data, FPOS_E offset data, FSZ_B size data, and FSZ_E size data are determined equally.
ES 2 323 482 T3 so that the mutually corresponding pixels constituting an enlarged image in the base layer and an image in a reinforcing layer are located in the same position in the absolute coordinate system.
Returning to Fig. 5, the image data, the compensation data FPOS_E and the size data FSZ_E of the reinforcing layer, generated in the image layering section 21, are delayed by a delay circuit 22 in the period of process of a base layer coding section 25, which will be described later, and are supplied to a backing layer coding section 23. Also the image data, the offset data FPOS_B and the size data, FSZ_B, of the base layer are supplied to the coding section 25 of the base layer. In addition, the FR magnification is provided to the backing layer encoding section 23 and the resolution transformation section 24, through the delay circuit 25.
In the base layer encoding section 25, the base layer image data is encoded. The resulting encoded data (bit string) includes the offset data FPOS_B and the size data FSZ_B, and is supplied to the multiplexing section 26.
Furthermore, the base layer encoding section 25 decodes the locally encoded data and delivers the locally decoded image data in the base layer to the resolution transformation section 24. In the resolution transformation section 24, the base layer image data from the base layer encoding section 25 is returned to the original size by enlarging (or reducing) the image data according to the FR magnification. The resulting enlarged image is delivered to the backing layer encoding section 23.
On the other hand, in the backing layer encoding section 23, the image data of the backing layer is encoded. The resulting encoded data (bit string) includes the offset data FPOS_E and the size data FSZ_E, and is supplied to the multiplexing section 26. Note that in the backing layer encoding section 23, the encoding of the backing layer image data is performed using, as a reference image, the enlarged image supplied from the resolution transformation section 24.
The multiplexing section 26 multiplexes the outputs of the reinforcement layer decoding section 23 and the base layer encoding section 25, and delivers the multiplexed bitstream.
Note that the size data FSZ_B, the offset data FPOS_B, the motion vector (MV), the signaling COD, etc., of the base layer, are supplied from the base layer coding section 25 to the section Backing layer encoding 23, and that backing layer encoding section 23 is constructed to perform the process, referencing the supplied data that is needed. Details will be described later.
Next, figure 11 shows the detailed constitution example of the coding section 25 of the base layer, of figure 5. In figure 11, the same numerical references are applied to the parts that correspond to figure 1. It is that is, basically the base layer encoding section 25 is constituted as the encoder of figure 1.
The image data from the image layering section 21 (FIG. 5), that is, the VOP of the base layer, as in FIG. 1, is supplied to the frame memory 31 and stored therein. In a motion vector detector 32, the motion vector is detected in a macroblock unit.
But the size data FSZ_B and the offset data FPOS_B of the VOP of a base layer are supplied to the motion vector detector 32 of the coding section 25 of the base layer, which in turn detects the motion vector of a macroblock. , based on the supplied size data FSZ_B and offset data FPOS_B.
That is, as described above, the size and position of a VOP varies with time (frame). Therefore, when detecting the motion vector, there is a need to set a reference coordinate system for detection and to detect the motion in the coordinate system. Therefore, in the motion vector detector 32 in this case, the above-mentioned absolute coordinate system is used as the reference coordinate system, and a VOP to be encoded and a reference VOP are arranged in the absolute coordinate system, according to the size data FSZ_B and offset data FPOS_B, so the motion vector is detected.
Note that the detected motion vector (MV), along with the prediction mode, are supplied to a VLC unit 36 and motion compensator 42, and are also supplied to reinforcement layer encoding section 23 (Figure 5).
Even in the case of performing motion compensation, there is also a need to detect motion in a reference coordinate system, as described above. Therefore, the size data FSZ_B and the compensation data FPOS_B are supplied to the motion compensator 42.
ES 2 323 482 T3
A VOP whose motion vector was detected is quantized as in the case of FIG. 1, and the quantized coefficients are supplied to the VLC unit 36. Furthermore, as in the case of FIG. 1, the size data FSZ_B and the compensation data FPOS_B from the image layering section 21 are supplied to the VLC unit 36 in addition to the quantized coefficients, the quantization step, the motion vector and the prediction mode. In the VLC unit 36, the supplied data is encoded by means of variable length word encoding.
In addition to the aforementioned coding, the VOP whose motion vector was detected is decoded locally as in the case of figure 1 and stored in the frame memory 41. This decoded image is used as a reference image, as previously described, and is further delivered to the resolution transformation section 24 (FIG. 5).
Note that, unlike the MPEG-1 standard and the MPEG-2 standard, the MPEG-4 standard also uses a B-picture (B-VOP) as the reference picture. For this reason, a B-picture is also decoded locally and stored in the frame memory 41. (However, a B-image is currently used only in a backing layer, as a reference image.)
On the other hand, as described in FIG. 1, the VLC unit 36 determines whether the macroblock of an I-image, P-image, or B-image (I-VOP, P-VOP or B-VOP) is made as a macroblock. omitted. The VLC unit 36 sets COD and MODB flags indicating the result of the determination. The COD and MODB flags are also encoded by variable length word encoding and transmitted. In addition, the COD flag is supplied to the backing layer encoding section 23.
Next, figure 12 shows the example of constitution of the coding section 23 of the reinforcing layer of figure 5. In figure 12, the same reference numerals are applied to the parts corresponding to figure 11 or 1. That is, basically the backing layer encoding section 23 is constituted as the encoding section 25 of the base layer of Figure 11 or of the encoder of Figure 1, except that the frame memory 52 is again provided.
The image data from the image layering section 21 (FIG. 5), that is, the VOP of the backing layer, as in the case of FIG. 1, is supplied to the frame memory 31 and stored therein. In motion vector detector 32, the motion vector is detected as a macroblock unit. Even in this case, as in the case of FIG. 11, the size data FSZ_E and the offset data POS_E are supplied to the motion vector detector 32, in addition to the VOP of the backing layer, etc. In the motion vector detector 32, as in the aforementioned case, the position taken by the VOP of the reinforcing layer in the absolute coordinate system is recognized based on the size data FSZ_E and the compensation data FPOS_E , and the motion vector of the macroblock is detected.
In this case, in the motion vector detectors 32 of the reinforcement layer coding section 23 and the base layer coding section 25, the VOPs are processed according to a predetermined sequence, as described in the Figure 1. For example, the sequence is set as follows.
That is, in the case of spatial scalability, as illustrated in figure 13 (A) or 13 (B), the VOPs are processed in a reinforcing layer or in a base layer, for example, in the order P, B , B, B, ... or I, P, P, P, ...
And in this case, the first P-image (P-VOP) of the reinforcing layer is encoded, for example using as reference image the VOP of the base layer present at the same time as the P-image (in this case, image-I (IVOP)). The second B-image (B-VOP) is also encoded in the backing layer, for example, using as reference images the image of the backing layer immediately before that and also the VOP of the base layer present at the same time than image-B. That is, in this example, the B-image of the backing layer, like the P-image of the base layer, is used as a reference image to encode another VOP.
For the base layer, encoding is performed, for example, as in the case of the MPEG-1 standard, the MPEG-2 standard, or the H.263 standard.
The scalability of the SNR is processed in the same way as the aforementioned spatial scalability, because it is the same as the spatial scalability when the FR magnification of the spatial scalability is 1.
In the case of temporal scalability, that is, for example, in the case in which a VO is made up of VOP0, VOP1, VOP2, VOP3, ..., and it is also considered that VOP1, VOP3, VOP5, vOp7, .. . are in a reinforcing layer (figure 14 (A)) and VOP0, VOP2, VOP4, VOP6, ... are in a base layer (figure 14 (B)), as described above, the VOPs of the layers Reinforcement and base are processed, respectively, in the order B, B, B, ... and in the order I, P, P, P, ..., as illustrated in Figure 14.
And in this case, the first VOP1 (B-image) is encoded in the backing layer, for example using the VOP0 (I-image) and the VOP2 (P-image) of the base layer, as reference images. The second VOP3 (B-image) of the reinforcement layer is coded, for example, using as reference images the first coded VOP1 (B-image) of the reinforcement layer, immediately before that, and of the VOP4 (B-image). P) of the base layer present in
ES 2 323 482 T3 the instant (frame) close to VOP3. The third VOP5 (B-image) of the backing layer, as in the VOP3 encoding, is encoded, for example, using as reference images the second encoded VOP3 (B-image) of the backing layer immediately before it, and of the VOP6 (P-image) of the base layer, which is an image present at the instant (frame) following VOP5.
As described above, for VOPs in one layer (in this case, the backing layer), the VOPs of another layer (scalable layer) (in this case, base layer) can be used as reference images to encode a P-image and a B-image. In the case in which an OP is encoded in this way in one layer, using a VOP of another layer as the reference image, that is, as in this embodiment, in the case in which a VOP of the layer is used as the reference image base as a reference image by encoding a backing layer VOP in a predictable way, the motion vector detector 32 of the backing layer encoding section 23 (FIG. 12) is constructed to set and deliver the id_layer_ref flag indicating that a base layer VOP is used to encode a base layer VOP. base coat predictably. (In the case of 3 or more layers, the flag id_capa_ref represents a layer to which a VOP belongs, used as a reference image).
In addition, the motion vector detector 32 of the backing layer coding section 23 is constructed to set and deliver the flag_selecc_ref (reference image information) in accordance with the flag id_layer_ref for a VOP. The flag_select_ref (reference picture information) indicates which layer and which VOP of the layer are used as the reference picture when performing forward prediction coding or backward prediction coding.
More specifically, for example, in the case where a P-image is encoded in a reinforcing layer, using as a reference image a VOP belonging to the same layer as a decoded (locally decoded) image immediately before the image -P, the flag_selecc_ref is set to 00. Furthermore, in the case in which a P-image is encoded using as a reference image a VOP that belongs to a layer (in this case a base layer (reference layer)), different from an image presented immediately before the image- P, the flag_selecc_ref is set to 01. Furthermore, in the case where the P-image is encoded using as a reference image a VOP belonging to a different layer from the image to be displayed immediately after the P-image, the flag code_selecc_ref is set to 10. Furthermore, in the case where the P-image is encoded using as a reference image a VOP belonging to a different layer that is present at the same time as a P-image, the flag code_selecc_ref is set to 11.
On the other hand, for example, in the case where a P-image in the reinforcement layer is encoded using as a reference image for forward prediction, a VOP belonging to a different layer, which is present at the same time as B-image, and also using as a reference image for backward prediction, a VOP belonging to the same layer as a decoded image immediately before the B-image, the flag code_selecc_ref is set to 00. In addition, in the case in which an image of the reinforcement layer is encoded, using as a reference image for forward prediction, a VOP belonging to the same layer as the B-image and also using as a reference image for the forward prediction. backward prediction of a VOP that belongs to a different layer than the displayed image immediately before the B-image, the flag_selecc_ref is set to 01. Furthermore, in the case where the B-image of the reinforcement layer is encoded using as a reference image for forward prediction, a VOP belonging to the same layer as a decoded image immediately before the B-image and also using as a reference image for the backward prediction, a VOP belonging to a different layer of an image to be displayed immediately after the B-image, the flag code_selecc_ref is set to 10. Furthermore, in the case where the B-image of the reinforcing layer is encoded using as a reference image for forward prediction, a VOP belonging to a different layer from the image presented immediately before the B-image, and also using as a reference image for backward prediction, a VOP belonging to a different layer of an image to be displayed immediately after the B-image, the flag code_selecc_ref is set to 11.
In this case, the prediction coding illustrated in Figures 13 and 14 is merely a single example. Therefore, it is possible within the aforementioned range, to freely set which layer and which VOP is used as the reference picture for forward prediction coding, backward prediction coding or bidirectional prediction coding.
In the aforementioned case, although the terms spatial scalability, temporal scalability, and SNR scalability have been used for the convenience of explanation, it becomes difficult to discriminate the spatial scalability, temporal scalability, and SNR scalability from each other, in the case in that a reference picture for a prediction encoding is set by the flag_selecc_ref. That is, conversely speaking, the use of the flag_selecc_ref makes the aforementioned discrimination between scalabilities unnecessary.
In this case, if a correlation is made between the aforementioned scalability and the flag_selecc_ref, the correlation will be, for example, as follows. That is, with respect to the P-image, as in the case where the flag_selecc_ref is 11, it is a case where the VOP is used at the same time in the layer indicated by the flag id_layer_ref as a reference image (for forward prediction), this case corresponds to spatial scalability or SNR scalability. And the cases other than the case in which the flag_selecc_ref is 11, corresponds to the temporal scalability.
ES 2 323 482 T3
In addition, with respect to a B-image, the case where the flag_selecc_ref is 00 is also the case where a VOP is used at the same time in the layer indicated by the flag id_layer_ref as a reference image for forward prediction , so this case corresponds to spatial scalability or SNR scalability. And the cases other than the case in which the flag_selecc_ref is 00, correspond to the temporal scalability.
Note that, in the case where in order to encode a VOP in a backing layer in a predictable way, a VOP is employed at the same time in a layer (in this case, base layer) other than the backing layer, such as reference image, there is no movement between them, so the movement vector is always 0 ((0,0)).
Returning to Figure 12, the aforementioned id_layer_ref flag and code_selecc_ref flag are set to the motion vector detector 32 of the reinforcement layer encoding section 23, and supplied to the motion compensator 42 and the unit. 36 of VLC.
Furthermore, the motion vector detector 32 detects a motion vector not only by referring to the frame memory 31, in accordance with the id_capa_ref flag and the code_selecc_ref flag, but also by referring to the frame memory 52, when needed. .
In this case, a locally decoded enlarged image in the base layer is supplied from the resolution transformation section 24 (FIG. 5) to the frame memory 52. That is, in the resolution transformation section 24, the locally decoded VOP in the base layer is amplified, for example, by a so-called interpolation filter, etc. With this, an enlarged image is generated which is FR times the size of the VOP, that is, an enlarged image of the same size as the VOP of the reinforcing layer, corresponding to the VOP of the base layer. The generated image is supplied to the backing layer encoding section 23. The frame memory 52 stores the enlarged image supplied from the resolution transformation section 24 in this manner.
Therefore, when the FR magnification is 1, the resolution transform section 24 does not process the locally decoded VOP from the base layer encoding section 25. The locally decoded VOP from the base layer encoding section 25, as is, is supplied to the backing layer encoding section 23.
The FSZ_B size data and the FPOS_B offset data are supplied from the base layer coding section 25 to the motion vector detector 32, and the FR magnification is also supplied from the delay circuit 22 (Figure 5) to the motion vector detector 32. In the case where the enlarged image stored in the frame memory 52 is used as the reference image, that is, in the case where, in order to encode a VOP in a backing layer in a predictable manner, a VOP of the base layer at the same time as the VOP of the reinforcement layer, as a reference image (in this case, the flag code_selecc_ref is set equal to 11 for a P-image and 00 for a B-image), the motion vector detector 32 multiplies the size data FSZ_B and the offset data FPOS_B corresponding to the enlarged image, by the magnification FR. And, based on the result of the multiplication, the motion vector detector 32 recognizes the position of the enlarged image in the absolute coordinate system, thereby detecting the motion vector.
Note that the motion vector and prediction mode are supplied in a base layer, to motion vector detector 32. These data are used in the following case. That is, in the case in which the flag_selecc_ref for a B-image of a reinforcement layer is 00, when the FR magnification is 1, that is, in the case of SNR scalability (in this case, as a VOP of the reinforcement layer to encode the reinforcement layer in a predictable way, the SNR scalability used in this case differs in this respect from that prescribed in the MPEG-2 standard), the images of the reinforcement layer and the layer base are the same. Thus, when B-image prediction coding is performed on a reinforcing layer, the motion vector detector 32 can employ the motion vector and prediction mode on a base layer that are present at the same time as the B-image, as they are. Therefore, in this case, the motion vector detector 32 does not process the B-image of the backing layer, but adopts the motion vector and base layer prediction mode as is.
In this case, in the backing layer coding section 23, neither a motion vector nor a prediction mode is delivered from the motion vector detector 32 to the VLC unit 36. (Therefore, they are not transmitted). This is because the receiving side can recognize the motion vector and prediction mode of a backing layer, from the result of decoding a base layer.
As described above, the motion vector detector 32 detects a motion vector using a VOP in a backing layer and in an enlarged image as reference images. Furthermore, as illustrated in FIG. 1, motion vector detector 32 sets a prediction mode that minimizes prediction error (or spread). Furthermore, the motion vector detector 32 sets and delivers necessary information, such as the select_code_ref flag, the layer_id_id flag, and so on.
In FIG. 12, the COD flag indicates whether a macroblock constituting an I-image or a P-image in the base layer is a skip macroblock, and the COD flag is supplied from the base layer coding section 25 to the motion vector detector 32, VLC unit 36, and motion compensator 42.
ES 2 323 482 T3
The macroblock, whose motion vector was detected, is encoded in the same way as in the aforementioned case. As a result of the encoding, the variable length codes are delivered from the VLC unit 36.
The VLC unit 36 of the backing layer coding section 23, as in the case of the base layer coding section 25, is constructed to fix and deliver COD and MODB flags. In this case, the COD flag, as described above, indicates whether a macroblock in an I-image or a P-image is a skip macroblock, while the MOD-B flag indicates whether a macroblock in a B-image it is a macroblock bypass.
The quantized coefficients, the quantization step, the motion vector, the prediction mode, the FR magnification, the code_selecc_ref flag, the id_layer_ref flag, the size data FSZ_E and the compensation data FPOS_E are also supplied to the unit 36 of VLC. In VLC unit 36, these are encoded by variable length word encoding and output.
On the other hand, after a macroblock whose motion vector has been detected has been encoded, it is also decoded locally as described above, and stored in the frame memory 41. And in the motion compensator 42, as in the case of the motion vector detector 32, motion compensation is performed using as reference images a locally decoded VOP in a reinforcement layer, stored in the frame memory 41, and a VOP locally decoded and expanded on a base layer, stored in frame memory 52. With this compensation, a predicted image is generated.
That is, in addition to the motion vector and prediction mode, the code_selecc_ref flag, the id_capa_ref flag, the FR magnification, the FSZ_B size data, the FSZ_E size data, the FPOS_B offset data, and the FPOS_E offset data are supplied to motion compensator 42. The motion compensator 42 recognizes a reference image to be motion compensated based on the flags code_selecc_ref and id_layer_ref. Furthermore, in the case where a locally decoded VOP in a backing layer is used as the reference image, or an enlarged image is used as the reference image, the motion compensator 42 recognizes the position and size of the reference image. in the absolute coordinate system, based on size data FSZ_E and FPOS_E, or size data FSZ_B and offset data FPOS_B. Motion compensator 42 generates a predicted image using FR magnification, when needed.
Next, figure 15 shows the constitution example of an embodiment of a decoder that decodes the bit stream delivered from the encoder of figure 3.
This decoder receives the bit stream supplied by the encoder of figure 3, through the transmission path 5 or the storage medium 6. That is, the bit string, delivered from the encoder of Figure 3 and transmitted through the transmission path 5, is received by the receiver (not illustrated).
Alternatively, the bit string recorded on the storage medium 6 is regenerated by a regenerator (not illustrated). The received or regenerated string of bits is supplied to a reverse multiplexing section 71.
The reverse multiplexing section 71 receives the bit stream (video stream (VS) described below) input into it. Furthermore, in the reverse multiplexing section 71, the input bit string is separated into bit strings VO # 1, VO # 2, ... The bit strings are supplied to the corresponding VOP decoding sections 72n, respectively. In VOP decoding sections 72n; the VOP (image data) constituting a VO, the size data (VOP size), and the offset data (VOP offset), are decoded from the bit string supplied from the reverse multiplexing section 71 . The decoded data is supplied to an image reconstruction section 73.
Image reconstruction section 73 reconstructs the original message, based on the respective outputs of sections 72<sub>1</sub> to 72<sub>N</sub> decoding of VOP. This reconstructed image is supplied, for example, to a monitor 74 and displayed on it.
Next, figure 16 shows the constitution example of section 72<sub>N</sub> VOP decoding method of FIG. 15, which effects scalability.
The bit string supplied from the reverse multiplexing section 71 (Figure 15) is input into a reverse multiplexing section 91, in which the input bit string is separated into a VOP bit string in a reinforcement layer. , and a string of bits from a VOP on a base layer. The bitstream of a VOP in a backing layer is delayed by a delay circuit 92, in a process period in the base layer decoding section 95, and is supplied to the base layer decoding section 93. reinforcement. In addition, the bitstream of a VOP in a base layer is supplied to the decoding section 95 of the base layer.
In the base layer decoding section 95, the base layer bitstream is decoded, and the resulting decoded image from a base layer is supplied to the resolution transformation section 94. What's more,
ES 2 323 482 T3 in the base layer decoding section 95, the information necessary to decode a VOP in a reinforcement layer, obtained by decoding the bit stream of a base layer, is supplied to the decoding section 93 of the reinforcing layer. The necessary information includes size data FSZ_B, compensation data FPOS_B, motion vector (MV), prediction mode, COD flag, etc.
In the reinforcement layer decoding section 93, the string of bits in a reinforcing layer, supplied through the delay circuit 92, is decoded by referring to the outputs of the base layer decoding section 95, and of section 94 of resolution transformation, as needed. The decoded image resulting from a reinforcement layer, the FSZ-E size data, and the FPOS_E offset data are output. Furthermore, in the reinforcement layer decoding section 93, the FR magnification, obtained by decoding the bit stream of a reinforcing layer, is output to the resolution transformation section 94. In the resolution transform section 94, as in the case of the resolution transform section 24 of FIG. 5, the decoded image on a base layer is transformed using the FR magnification supplied from the image decoding section 93. reinforcing layer. An enlarged image, obtained with this transformation, is supplied to the decoding section 93 of the reinforcing layer. As described above, the enlarged image is used to decode the bit stream of a reinforcing layer.
Next, Figure 17 shows the constitution example of the decoding section 95 of the base layer of Figure 16. In Figure 17, the same reference numerals are applied to the parts corresponding to the case of the decoder of Figure 2. That is, basically the decoding section 95 of the base layer is constituted in the same way as the decoder of Figure 2.
The base layer bit string from the reverse multiplexing section 91 is supplied to a buffer 101 and temporarily stored. An IVLC unit 102 reads the bit string from buffer 101, corresponding to a next stage block processing state, as needed, and the bit string is decoded by variable length word decoding, and is separated into quantized coefficients, a motion vector, a prediction mode, a quantization step, size data FSZ_B, offset data FPOS_B, and the COD flag. The quantized coefficients and the quantization step are supplied to an inverse quantizer 103. The motion vector and prediction mode are supplied to a motion compensator 107, and to the decoding section 93 of the reinforcement layer (FIG. 16). Furthermore, the size data FSZ_B and the compensation data FPOS_B are supplied to the motion compensator 107, the image reconstruction section 73 (Fig. 15), and the reinforcement layer decoding section 93, while the flag COD is supplied to the backing layer decoding section 93.
The inverse quantizer 103, the IDCT unit 104, the arithmetic unit 105, the frame memory 106, and the motion compensator 107 perform similar processes corresponding to the inverse quantizer 38, the IDCT unit 39, the arithmetic unit 40, the memory 41 of frames and the motion compensator 42 of the coding section 25 of the base layer of FIG. 11, respectively. With this, the VOP of a base layer is decoded. The decoded VOP is supplied to the image reconstruction section 73, the reinforcement layer decoding section 93 and the resolution transformation section 94 (FIG. 16).
Next, figure 18 shows the example of constitution of the decoding section 93 of the reinforcing layer of figure 16. In figure 18, the same numerical references are applied to the parts corresponding to the case of figure 2. It is That is, basically, the backing layer decoding section 93 is constituted in the same way as the decoder of FIG. 2, except that a new frame memory 112 is provided.
The bit string of a reinforcement layer from the reverse multiplexing section 91 is supplied to an IVLC 102 through a buffer 101. The IVLC unit 102 decodes the bit stream of a reinforcement layer, by decoding variable length words, thereby separating the bit stream into quantized coefficients, a motion vector, a prediction mode, a quantization step , size data FSZ_E, compensation data FPOS_E, magnification FR, flag id_layer_ref, flag code_selecc_ref, flag COD and flag MODB. The quantized coefficients and the quantization step, as in the case of FIG. 17, are supplied to an inverse quantizer 103. The motion vector and prediction mode are supplied to the motion compensator 107. In addition, the FSZ_E size data and FPOSE compensation data are supplied to the motion compensator 107 and the image reconstruction section 73 (FIG. 15). The COD flag, MOD flag, layer_id flag, and code_selection_ref flag are supplied to motion compensator 107. In addition, the FR magnification is supplied to motion compensator 107 and resolution transformation section 94 (FIG. 16).
Note that the motion vector, the COD flag, the FSZ_B size data, and the FPOS_B offset data of a base layer are supplied from the base layer decoding section 95 (FIG. 16) to the motion compensator 107, in addition to the aforementioned data. Furthermore, an enlarged image is supplied from the resolution transformation section 94 to the frame memory 112.
The inverse quantizer 103, the IDCT unit 104, the arithmetic unit 105, the frame memory 106, the motion compensator 107, and the frame memory 112 perform similar processes corresponding to the inverse quantizer 38, to the IDCT unit 39, to the arithmetic unit 40, the frame memory 41, the motion compensator 42 and the frame memory 52 of the layer coding section 23
ES 2 323 482 T3 reinforcement of Figure 12, respectively. With this, the VOP of a backing layer is decoded. The decoded VOP is supplied to the image reconstruction section 73.
In this case, in the VOP decoding sections 72n having the reinforcing layer decoding section 93 and the base layer decoding section 95, constituted as described above, both the decoded image, the FSZ_E size data and FPOS_E offset data of a backing layer (hereinafter referred to as backing layer data, when needed), such as the decoded image, the size data FSZ_B and the compensation data FPOS_B of the base layer (hereinafter referred to as the base layer data, when required). In the image reconstruction section 73, an image is reconstructed from the backing layer data or the base layer data, for example, as follows.
That is, for example, in the case where the first spatial scalability is realized (Figure 6) (that is, in the case where all the input VOP is made a backing layer and all the reduced VOP is made a layer base), when both the base layer data and the backing layer data are decoded, the image reconstruction section 73 sets the decoded image (VOP) of the backing layer, of size corresponding to the size data FSZ_E at the position indicated by the offset data FPOS_E, based only on the data of the backing layer. Also, for example, when an error occurs in the bit string of a backing layer, or when the monitor 74 processes only a low resolution image and thus only the base layer data is decoded, the section Image reconstruction 73 arranges the decoded image (VOP) of a reinforcing layer of size corresponding to the size data FSZ_B, at the position indicated by the compensation data FPOS_B, based only on base layer data.
Also, for example, in the case that the second spatial scalability is performed (figure 7), (that is, in the case where part of an input VOP is made a reinforcing layer and all the reduced VOP is made a base layer), when both the base layer data and the backing layer data are decoded, the image reconstruction section 73 enlarges the base layer decoded data to the size corresponding to the FSZ_B size data, according to the FR magnification and generates the magnified image. Furthermore, the image reconstruction section 73 magnifies FR times the compensation data FPOS_B and places the magnified image at the position corresponding to the resulting value. And the image reconstruction section 73 sets the decoded image of the reinforcement layer with the size corresponding to the size of the data FSZ_E at the position indicated by the compensation data FPOS_E.
In this case, the part of the decoded image of a reinforcing layer is presented with a higher resolution than the remaining part.
Note that in the case where the decoded image is provided with a reinforcing layer, the decoded image and an enlarged image are mutually synthesized.
Furthermore, although not illustrated in Fig. 16 (Fig. 15), the FR magnification from the backing layer decoding section 93 (VOP decoding sections 72n) is supplied to the image reconstruction section 73, in addition to the aforementioned data. The image reconstruction section 73 generates an enlarged image using the supplied FR magnification.
On the other hand, in the case where the second spatial scalability is performed, when only the base layer data is decoded, an image is reconstructed in the same way as the aforementioned case, where the first spatial scalability is performed.
In addition, in the case in which the third spatial scalability is carried out, (Figures 8 and 9), (that is, in the case in which each of the objects that constitute an input VOP is made a reinforcing layer and the VOP that excludes the objects a base layer is made), an image is reconstructed in the same way as the aforementioned case in which the second spatial scalability is performed.
As described above, the compensation data FPOS_B and the compensation data FPOS_E are constructed so that mutually corresponding pixels, which constitute the enlarged image of a base layer and an image of a reinforcement layer, are arranged therein. position in the absolute coordinate system. Therefore, by reconstructing an image in the aforementioned manner, an accurate image (without position compensation) can be obtained.
Next, the syntax of the output of the string of bits encoded by the encoder of figure 3 will be described, for example, with the video verification model (version 6.0) of the MPeG-4 standard (hereinafter called VM-6.0 when needed), as an example.
Figure 19 shows the syntax of a string of encoded bits in VM-6.0.
The encoded bit stream is made up of Video Session Classes (VS). Each VS is made up of one or more classes of video objects (VO). Each VO is made up of one or more classes of object layers (VOL). (When an image is not layered, it is made up of a single VOL. In the case where the image is layered, it is made up of the VOLs corresponding to the number of layers). Each VOL is made up of Video Object Plane (VOP) classes.
ES 2 323 482 T3
Note that VS are a sequence of images and equivalents, for example, with a single program or movie.
Figures 20 and 21 show the syntax of a VS and the syntax of a VO. The VO is a string of bits corresponding to a complete image or to a sequence of objects that make up an image. Therefore, VS are constituted by a set of such sequences. (Therefore, the vS are equivalent, for example, to a single program).
Figure 22 shows the syntax of a VOL.
VOL is a class for the aforementioned scalability and is identified by a number indicated with a video_object_layer_id. For example, the video_object_layer_id for a base layer VOL becomes 0, while the video_object_layer_id for a reinforcement layer VOL becomes 1. Note that, as described above, the number of scalable layers is not limited to 2, but can be an arbitrary number that includes 1, 3, or more.
Also, if a VOL is a complete image or part of an image, it is identified by the form_layer_object_ video. This video_object_layer_form is a flag to indicate the shape of a VOL and is set as follows.
When the shape of a VOL is rectangular, the video_object_layer_form is given, for example, the value 00. Furthermore, when a VOL is in the form of a zone cut by means of a hard key (a binary signal that takes the value 0 or 1), the video_object_layer_form is given, for example, the value 01. Also, when a VOL is in the form of zone cut by means of a soft key (a signal that can take on a continuous value (grayscale) in the range of 0 to 1) (when synthesized with a soft key), the video_object_layer_form is given a value of, for example, 10.
In this case, when the video_object_layer_form is given the value 00, the shape of a VOP is rectangular and also the position and size of a VOL in the absolute coordinate system do not vary with time, that is, they are constant. In this case, the sizes (horizontal length and vertical length) are indicated by video_object_layer_width and video_object_layer_height. Video_object_layer_width and video_object_layer_height are both 10-bit fixed-length flags. In the case where the video_object_layer_form is 00, it is transmitted first only once. (This is because, in the case where the video_object_layer_form is 00, as described above, the size of a VOL in the absolute coordinate system is constant).
Also, if a VOL is a base layer or a backing layer, it is indicated by the scalability which is a 1-bit flag. When a VOL is a base layer, scalability is given, for example, the value 1. In any case other than that, scalability is given, for example, the value 0.
Furthermore, in the case where a VOL uses an image in a VOL other than itself as a reference image, the VOL to which the reference image belongs is represented by id_capa_ref, as described above. Note that the id_layer_ref is transmitted only when a VOL is a backing layer.
In Fig. 22, the see_sample factor and the view_sample_factor indicate a value corresponding to the horizontal length of a VOP in a base layer and a value corresponding to the horizontal length of a VOP in a backing layer, respectively. The horizontal length from a reinforcement layer to a base layer (horizontal resolution magnification), is given by the following equation:
factor_n_sample_hor / factor_m_sample_hor.
In FIG. 22, the hor_sample_factor and hor_m_sample_factor indicate a value corresponding to the vertical length of a VOP in a base layer and a value corresponding to the vertical length of a VOP in a reinforcing layer, respectively. The vertical length from a reinforcement layer to a base layer (vertical resolution magnification) is given by the following equation:
factor_n_sample_ver / factor_m_sample_ver.
Next, Figure 23 shows the syntax of a VOP.
The sizes (horizontal length and vertical length) of a VOP are indicated, for example, by VOP_width and VOP_height, which have a fixed length of 10 bits. Furthermore, the positions of a VOP in the absolute coordinate system are indicated, for example, by ref_mc_espacial_horizontal_VOP and by ref_mc_vertical_VOP of fixed length of 10 bits. VOP_width and VOP_height represent the horizontal length and vertical length of a VOP, respectively. These are equivalent to the FSZ_B size data and the FSZ_E size data described above. The ref_mc_espacial_horizontal_VOP and the ref_mc_vertical_VOP represent the horizontal and vertical coordinates (x and y coordinates) of a VOP, respectively. These are equivalent to the compensation data FPOS_B and the compensation data FPOS_E described above.
ES 2 323 482 T3
The VOP_width, the VOP_height, the mc_horizontal_VOP_ref, and the mc_vertical_VOP_ref are transmitted only when the video_object_layer_form is not 00. That is, when the video_object_layer_form is 00, as described above, the size and position of a VOP are both constant, so that there is no need to transmit VOP_width, VOP_height, mc_space_ref_horizontal_VOP and mc_vertical_VOP_ref. In this case, on the receiver side, a VOP is arranged so that the upper left corner is consistent, for example, with the origin of the absolute coordinate system. Furthermore, the sizes are recognized from the video_object_layer_width and video_object_layer_height described in figure 22.
In Figure 23, the code_selecc_ref, as described in Figure 19, represents an image that is used as a reference image, and is prescribed by the syntax of a VOP.
By the way, in VM-6.0, the presentation time of each VOP (equivalent to a conventional frame) is determined by the modulo_time_base and by the modulo_time_increment (figure 23), as follows:
That is, the modulo_time_base represents the encoder time in the local timebase, with a precision of one second (1000 milliseconds). The module_time_base is represented as a marker transmitted in the VOP header, and is made up of a required number of 1s and 0s. The number of consecutive “1s” that make up the modulo_time_base followed by a “0” is the cumulative period from the synchronization point (time to the precision of one second) marked by the last encoded / decoded modulo_time_base. For example, when modulo_time_base indicates 0, the cumulative period from the sync point marked by the last encoded / decoded modulo_timebase is 0 seconds. Also, when the modulo_timebase indicates 10, the cumulative period from the synchronization point marked by the last encoded / decoded modulo_timebase is 1 second. Also, when the modulo_time_base indicates 110, the cumulative period from the synchronization point marked by the last encoded / decoded modulo_timebase is 2 seconds. Therefore, the number of ones in the modulo_time_base is the number of seconds from the sync point marked by the last encoded / decoded modulo_time_base.
Note that, for the module_time_base, VM-6.0 specifies that:
This value represents the local time base in the resolution unit of one second (1000 milliseconds). It is represented as a marker transmitted in the header of the VOP. The number of consecutive 1s followed by a 0 indicates that the number of seconds has elapsed since the synchronization point marked by the last encoded / decoded module_time_base.
The VOP_time_increment represents the encoder time in the local time base within the precision of 1 ms. In VM-6.0, for I-VOP and P-VOP, the VOP_time_increment is the time from the sync point marked by the last encoded / decoded module_time_base. For B-VOPs, the VOP_time_increment is the relative time since the last encoded / decoded I- or P-VOP.
Note that, for the VOP_time_increment, the VM-6.0 specifies that:
This value represents the local time base in units of milliseconds. For the I- and P-VOPs, this value is the absolute VOP_time_increment from the synchronization point marked by the last module_time_base. For B-VOPs, this value is the relative VOP_time_increment since the last encoded / decoded I- or P-VOP.
And the VM-6.0 specifies that:
In the encoder, the following formulas are used to determine the absolute and relative VOP_time_steps for the I / P-VOP and B-VOP, respectively.
That is, VM-6.0 prescribes that, in the encoder, the presentation times for the I / P-VOP and B-VOP are, respectively, encoded by the following formulas:
tGTB (n) = nx 1000 ms + tEST tAVT1 = tETB (I / P) = tGTB (n) tRVTI = tETB (B) - tETB (I / P) ... (1) where tGTB (n) represents the time of the synchronization point (as described above, with precision of one second) marked by the nth encoded modulo_time_base, tEST represents the encoder time at the start of VO encoding ( the absolute time in which the VO coding started), tAVT1 represents the VOP_time_increment for the I or P-VOP, tETB (I / P) represents the encoder time at the start of the I or P-VOP encoding (the absolute time at which the VOP encoding started), tRVTI represents the VOP_time_increment for the B-VOP, and tETB (B) represents the encoder time at the start of the B-VOP encoding.
ES 2 323 482 T3
Note that, for the tGTB (n), tEST, tAVT1, tETB (I / P), tRVTI and tETB (B) of formulas (1), the VM-6.0 specifies that:
tGTB (n) is the encoder time base marked by the nth encoded module_time_base, tEST is the start time of the encoder timebase, tAVTI is the absolute VOP_time_increment for the I or P-VOP, tETB (I / P) is the encoder time base at the start of the I or P-VOP encoding, tRVTI is the relative VOP_time_increment for the B-VOP, and tETB (B) is the encoder time base at the start of the B-VOP encoding.
Additionally, VM-6.0 specifies that:
In the decoder, the following formulas are used to determine the recovered time base of the I / PVOP and B-VOP, respectively.
That is, VM-6.0 prescribes that on the decoder side, the presentation times for the I / P-VOPs and the B-VOPs are decoded, respectively, by the following formulas:
tGTB (n) = nx 1000 ms + tDST tDTB (I / P) = tAVT1 + tGTB (n) tDTB (B) = tRVTI + tDTB (I / P) ... (2) where tGTB (n) represents the time of the synchronization point marked by the decoded nth module_time_base, tDST represents the decoder time at the start of VO decoding (the absolute time at which VO decoding started) , tDTB (I / P) represents the decoder time at the start of decoding the I-VOP or P-VOP, tAVTI represents the VOP_time_increment for the I-VOP or P-VOP, tDTB (B) represents the decoder time at the start of decoding the B-VOP (the absolute time at which the decoding of the VOP started), tRVTI represents the VOP_time_increment for the B_VOP.
Note that, for the tGTB (N), tDST, tDTB (I / P), tAVTI, tDTB (B) and tRVTI of formulas (2), the VM-6.0 specifies that:
tGTB (n) is the encoding time base marked by the nth decoded module_time_base, tDST is the start time of the decoding time base, tDTB (I / P) is the decoding time base at the start of decoding of I or P-VOP, tAVTI is the absolute decoding VOP_time_increment for the P-VOP, tDTB (B) is the decoding time base at the beginning of the B-VOP decoding, and tRTVI is the decoded relative VOP_time_increment for the B-VOP.
Figure 24 shows the relationship between modulo_time_base and VOP_time_increment, based on the above definition.
In the figure, a VO is made up of a sequence of VOPs, such as I1 (I-VOP), B2 (B-VOP), B3, P4 (P-VOP), B5, P6, etc. Now suppose that the encode-decode start time (absolute time) of the VO is t0, the module_time_base will represent time (sync point), such as t0 + 1 sec. t0 + 2 sec. etc., because the elapsed time from the start time t0 is accurately represented within one second. In Figure 24, although the display order is I1, B2, B3, p4, B5, P6, etc., the encoding / decoding order is I1, P4, B2, B3, P6, etc.
In Figure 24 (same as Figures 28 to 31 and Figure 36 to be described later), the VOP_time_increment for each VOP is indicated by a number (in units of milliseconds) enclosed within a square. Synchronization point switching indicated by the module_time_base is indicated by a ▼ mark. In figure 24, therefore, the VOP_time_steps for I1, B2, B3, P4, B5 and P6 are 350 ms, 400 ms, 800 ms, 550 ms, 400 ms and 350 ms, and in P4 and P6, it is switched the synchronization point.
Now, in Figure 24, the VOP_time_increment for I1 is 350 ms. The encoding / decoding time of I1, therefore, is the time 350 ms after the synchronization point marked by the last encoded / decoded modulo_time_base. Note that immediately after the start of I1 encoding / decoding, the start time t0 (encoding / decoding start time) becomes a synchronization point. The encoding / decoding time of I1 will therefore be time t0 + 350 ms after 350 ms from the start time t0 (encoding / decoding start time).
And the encode / decode time of the B2 or B3 is the time of the VOP_time_increment that has elapsed since the last encoded / decoded I-VOP or P-VOP. In this case, as the encoding / decoding time of the last encoded / decoded I1 is t0 + 350 ms, the encoding / decoding time of B2 or B3 is time t0 + 750 ms or t0 + 1200 ms after 400 ms or 800 ms.
ES 2 323 482 T3
Then, for P4, the synchronization point indicated by the module_time_base is switched at P4. Therefore, the synchronization point is time t0 + 1 sec. As a result, the encode / decode time of P4 is time (t0 + 1) sec + 550 ms after 350 ms from time t0 + 1 sec.
The B5 encode / decode time is the VOP_time_increment time that has elapsed since the last encoded / decoded I-VOP or P-VOP. In this case, since the encoding / decoding time of the last encoded / decoded P4 is (t0 + 1) sec. + 550 ms, the encoding / decoding time of the B5 is time (t0 + 1) sec. + 950 ms after 400 ms.
Then, for P6, the synchronization point of P6 indicated by the modulo_time_base is switched. Therefore, the synchronization point is time t0 + 2 sec. As a result, the encoding / decoding time of P6 is time (t0 + 2) sec. + 350 ms after 350 ms from time t0 + 2 sec.
Note that in VM-6.0, the switching of the synchronization points indicated by the modulo_time_base is allowed only for I-VOPs and P-VOPs and is not allowed for B-VOPs.
Furthermore, VM-6.0 specifies that for I-VOPs and P-VOPs, the VOP_time_increment is the time from the synchronization point marked by the last encoded / decoded module_time_base, while for B-VOPs, the VOP_time_increment is the Relative time from the sync point marked by the last encoded / decoded I-VOP or P-VOP. This is mainly for the following reason. That is, a B-VOP is predictably coded using the I-VOP or P-VOP arranged through the BVOP in order of appearance as a reference image. Therefore, the temporal distance to the I-VOP or P-VOP is set at the VOP_time_increment for the B-VOP, so that the weight, in relation to the I-VOP or P-VOP that is used as a reference image to effect predictable coding, it is determined from the B-VOP on the basis of the time distance to the IVOP or P-VOP disposed across the B-VOP. This is the main reason.
By the way, the definition of the VOP_time_increment of the aforementioned VM-6.0 has a drawback. That is, in Figure 24, the VOP_time_increment for a B-VOP is not the relative time since the last IVOP or P-VOP encoded / decoded, immediately before the B-VOP, but the relative time since the last I-VOP or P-VOP presented. This is for the following reason. For example, consider B2 or B3. The I-VOP or PVOP that is encoded / decoded immediately before B2 or B3 is P4 from the point of view of the above-mentioned encoding / decoding order. Thus, when the VOP_time_increment for a B-VOP is assumed to be the relative time from the encoded / decoded I-VOP or P-VOP immediately before the B-VOP, the VOP_time_increment for B2 or B3 is the relative time from the moment encode / decode number of P4, and becomes a negative value.
On the other hand, in the MPEG-4 standard, the VOP_time_increment is 10 bits. If the VOP_time_increment has only a value equal to or greater than 0, it can express a value in the range of 0 to 1023. Therefore, the position between contiguous synchronization points can be represented in units of milliseconds with the previous time synchronization point ( in the left direction of Figure 24) for reference.
However, if the VOP_time_increment is allowed to have not only a value equal to or greater than 0, but also a negative value, the position between contiguous synchronization points will be represented with the previous time synchronization point as a reference, or it will be represented with the next time sync point for reference. For this reason, the process of calculating the encoding time or decoding time of a VOP becomes complicated.
Therefore, as described above, for VOP_time_increment, VM-6.0 specifies that:
This value represents the local time base in units of milliseconds. For I- and P-VOP, this value is the absolute VOP_time_increment from the synchronization point marked by the last module_time_base. For the B-VOP, this value is the relative VOP_time_increment since the last encoded / decoded I- or P-VOP.
However, the last sentence “For B-VOPs this value is the relative VOP_time_increment since the last encoded / decoded P-VOP Io”, should be changed to “For B-VOPs, this value is the relative VOP_time_increment since the last I - or P-VOP filed ”. With this, the VOP_time_increment should not be defined as the relative time since the last encoded / decoded I- or P-VOP, but should be defined as the relative time since the last presented I- or P-VOP.
Defining the VOP_time_increment in this way, the basis for calculating the encoding / decoding time for a B-VOP is the presentation time of the I / P-VOP (I-VOP or P-VOP) that have a presentation time before the B-VOP. Therefore, the VOP_time_increment for a B-VOP always has a positive value, as long as no I-VOP reference image is presented for the B-VOP before the B-VOP. Therefore, the VOP_time_increases for the I / P-VOPs also have a positive value at all times.
Furthermore, in figure 24, the definition of the VM-6.0 is also changed so that the time represented by the module_time_base and by the VOP_time_increment is not the encoding / decoding time of a VOP, but rather the presentation time of a VOP. VOP. That is, in figure 24, when absolute time is considered
ES 2 323 482 T3 in a sequence of VOP, the tEST (I / P) of formulas (1) and tDTB (I / P) of formulas (2) represent absolute times present in a sequence of I-VOP or of P-VOP, respectively, and the tEST (B) of formulas (1) and tDTB (B) of formulas (2) represent the absolute times present in a sequence of B-VOP, respectively.
Then, in the VM-6.0, the tEST start time of the encoder time base of formulas (1) is not encoded, but the module_time_base and the VOP_time_increment are encoded as differential information between the tEST start time of the encoder time base and the presentation time of each VOP (absolute time representing the position of a VOP present in a sequence of VOPs). For this reason, on the decoder side, the relative time between the VOPs can be determined using the module_time_base and the VOP_time_increment, but the absolute presentation time of each VOP, that is, the position of each VOP in a VOP sequence, cannot can be determined. Therefore, only modulo_time_base and VOP_time_increment cannot access a bit string, that is, random access.
On the other hand, if the encoder time base start time tEST is merely encoded, the decoder can decode the absolute time of each VOP, using the encoded tEST. However, when decoding in the header of the encoded bit string, the start time tEST of the encoder time base and also the module_time_base and the VOP_time_increment which are the relative time information of each VOP, there is a need to control the cumulative absolute time. This is troublesome, so that efficient random access cannot be performed.
Therefore, in the embodiment of the present invention, a layer to encode the absolute time that is present in a VOP sequence is introduced in the hierarchical constitution of the encoded bit stream of the VM-6.0, to easily perform an access effective random. (This layer is not a layer that performs scalability (the aforementioned base layer or backing layer) but is a layer of the encoded bitstream.) This layer is an encoded bitstream layer that can be inserted at an appropriate position as well as the header of the encoded bitstream.
Like this layer, this embodiment introduces, for example, a prescribed layer in the same way as a GOP (group of pictures) layer used in the MPeG-1/2 standard. With this, the compatibility between the MPEG-4 standard and the MPEG-1 standard can be strengthened compared to the case where an original layer of encoded bitstream is used in the MPEG-4 standard. This entered layer is now called GOV (or group of video object planes (GVOP)).
Figure 25 shows the constitution of the encoded bit stream in which a GOV layer is inserted to encode the absolute times present in a VOP sequence.
The GOV layer is prescribed between a VOL layer and a VOP layer, so that it can be inserted at the arbitrary position of a string of coded bits, as well as at the head of the string of coded bits.
With this, in the case where a certain VOL # 0 is constituted by a sequence of VOP, such as VOP # 0, VOP # 1, ... VOP # n, VOP # (n + 1), ... and VOP # m, the GOV layer can be inserted, for example, directly before VOP # (n + 1), as well as directly before the VOP # 0 header. Thus, in the encoder, the GOV layer can be inserted, for example, at the position of a string of coded bits where random access is performed. Thus, upon inserting the GOV layer, a sequence of VOP constituting a certain VOL is separated into a plurality of groups (hereinafter referred to as GOV, when needed) and is encoded.
The GOV layer syntax is defined, for example, as illustrated in Figure 26.
As illustrated in the figure, the GOV layer is made up of a group_start_code, a time_code, a closed_gop, a broken_link and a next_start_code OR arranged in sequence.
Next, a description of the semantics of the GOV layer will be made. The semantics of the GOV layer are basically the same as the GOP layer in the MPEG-2 standard. Therefore, for the parts that are not described here, see the MPEG-2 video standard (ISO / IEC-13818-2).
The group_start_code is 000001B8 (hexadecimal) and indicates the start position of a GOV.
The time_code, as illustrated in Figure 27, consists of a 1-bit frame_down_flag, 5-bit time_code_hours, 6-bit time_code_minutes, a 1-bit marker bit, 6-bit time_code_seconds, and 6-bit time_code_images. Therefore, the time_code is made up of 25 bits in total.
The time_code is equivalent to the “time and control codes for videotape recorders”, prescribed in publication 461 of the IEC standard. In this case, the MPEG-4 standard does not have the concept of video frame rate. (Therefore, a VOP can be represented at an arbitrary moment). Therefore, this embodiment does not take advantage of the frame_down_flag indicating whether the time_code is described in the frame_down_mode, and the value is set, for example, to 0. Furthermore, this embodiment does not take advantage of time_code_images for the same reason, and the value is set, for example, to 0. Therefore, the time_code used in this case represents the time of the head of a gOv times the time_code hours that represent the
ES 2 323 482 T3 unit in hours of the time, the minutes_time_code that represent the unit in minutes of the time, and the seconds_time_code that represent the unit in seconds of the time. As a result, the time_code (absolute time accurate to one second from the start of encoding) in a GOV layer expresses the head time of the GOV layer, that is, the absolute time in a VOP sequence when encoding starts of the GOV layer, accurate to within one second. For this reason, this embodiment of the present invention sets the time with a precision finer than one second (in this case, milliseconds) for each VOP.
Note that the marker bit in the timecode is set equal to 1 so that 23 or more zeros do not continue in a string of coded bits.
The gop_cerrado means one in which the I-, P- and B- images in the definition of the gop_cerrado of the MPEG-2 video standard (ISO / IEC 13818-2) have been replaced by an I-VOP, a P-VOP and a B-VOP, respectively. Thus, the B-VOP of one VOP represents not only a VOP constituting the GOV, but represents whether the VOP has been encoded with a VOP in another GOV as a reference image. In this case, for the definition of the closed gop in the MPEG-2 video standard (ISO / IEC 13818-29), the phrases that effect the above-mentioned substitution are illustrated as follows:
This is a one-bit flag indicating the nature of the predictions used in the first consecutive B-VOPs (if any) that immediately follow the first coded I-VOP that follows the head-group of the plane. The gop_cerrado is set to 1 to indicate that these B-VOPs have been encoded using only backward prediction or intra-encoding. This bit is provided to be used during any editing that takes place after encoding. If the previous images are removed for editing, the broken_link can be set to 1, so that a decoder can prevent the presentation of these B-VOPs that follow the first I-VOP that follows the group of the head of the plane. However, if the gop_close bit is set to 1, then the publisher can choose not to set the link_broken bit, as these B-VOPs can be decoded correctly.
The broken_link also means one in which the same substitution has been made as in the case of the closed gop in the definition of the broken_link of the MPEG-2 video standard (ISO / IEC 13818-29). The broken_link therefore represents whether the header B-VOP of a GOV can be regenerated correctly. In this case, for the definition of the broken_link in the MPEG-2 video standard (ISO / IEC 13818-2), the phrases that effect the above-mentioned substitution are illustrated as follows:
This is a one-bit flag that will be set to 0 during encoding. It is set to 1 to indicate that the first consecutive B-VOPs (if any), immediately following the first encoded I-VOP, which follows the group at the head of the shot, cannot be decoded correctly because the reference frame that used for prediction is not available (due to edit action). A decoder can use this flag to avoid displaying frames that cannot be decoded correctly.
The code () _ start_ next gives the position of the head of the next GOV.
The above-mentioned absolute time in a GOV sequence that introduces the GOV layer and also initiates the GOV layer encoding (hereinafter referred to as the absolute decoding start time, when required), is set to the GOV time_code. Furthermore, as described above, since the time_code in the GOV layer has a precision within one second, this embodiment sets a finer precision part in the absolute time of each VOP that is present in a VOP sequence to each VOP.
Figure 28 shows the relationship between time_code, module_time_base and VOP_time_increment in the case where the GOV layer of Figure 26 has been introduced.
In the figure, the GOV is made up of I1, B2, B3, P4, B5 and P6 arranged in order of presentation from the header.
Now assuming, for example, that the absolute time of the GOV encoding start is 0h: 12m: 35sec: 350msec (0 hours, 12 minutes, 35 seconds 350 milliseconds), the GOV timecode will be set to 0h: 12m: 35sec because it is accurate to within one second, as described above. (The time_code_hours, time_code_minutes_, and time_code_seconds that make up the time_code will be set to 0, 12, and 35, respectively). On the other hand, in the case where the absolute time of I1 in a VOP sequence (absolute time of a VOP sequence before coding (or after coding) of a VS that includes the GOV of figure 28 ) (since this is equivalent to the presentation time of I1, when a VOP sequence is presented, it will be referred to as presentation time, when needed) is, for example, 0h: 12m: 35sec: 350msec, the semantics of the VOP_time_increment changes such that 350 ms, which is a precision finer than the precision of one second, is set to the VOP_time_increment of the I-VOP of I1 and is encoded (that is, so that the encoding is done with the VOP_time_increment of I1 = 350).
That is, in FIG. 28, the I-VOP header_time_time_increment (I1) of a GOV in display order has a differential value between the GOV timecode and the I-VOP display time. Therefore, the time accurately within one second that the time_code represents is the first synchronization point of the GOV (in this case, a point that represents the time accurately within one second).
ES 2 323 482 T3
Note that, in figure 28, the semantics of the VOP_time_increments for B2, B3, P4, B5 and P6 of the GOV, which are arranged as if the VOP were the second or later, is the same as the one in which the definition changes of the VM-6.0, as described in figure 24.
Therefore, in FIG. 28, the presentation time of B2 or B3 is the time in which the VOP_time_increment has elapsed since the last I-VOP or P-VOP presented. In this case, since the presentation time of the last I1 presented is 0h: 12m: 35sec: 350msec, the presentation time of B2 or B3 is 0h: 12m: 35sec: 750msec or 0h: 12m: 36sec: 200msec after 400 msec or 800 msec.
Then, for P4, the synchronization point indicated by the module_time_base is switched at P4. Therefore, the synchronization point time is 0h: 12m: 36sec after 1 second from 0h: 12m: 35sec. As a result, the P4 display time is 0h: 12m: 36sec: 550msec after 550msec from 0h: 12m: 36sec.
The B5 display time is the time in which the VOP_time_increment has elapsed since the last I-VOP or P-VOP presented. In this case, the presentation time of the B5 is 0h: 12m: 36sec: 950msec after the 400msec from the presentation time 0h: 12m: 36sec: 550msec of the last P4 presented.
Then, for P6, the synchronization point indicated by the module_time_base is switched at P6. Therefore, the synchronization point time is 0h: 12m: 35sec + 2sec, that is, 0h: 12m: 37sec. As a result, the P6 display time is 0h: 12m: 37sec: 350msec after 350msec. from 0h: 12m: 37sec.
Next, Figure 29 shows the relationship between time_code, module_time_base and VOP_time_increment in the case where the VOP of a GOV header is a B-VOP in order of presentation.
In the figure, the GOV is made up of B0, I1, B2, B3, P4, B5 and P6, arranged in order of presentation from the header. That is, in figure 29, the GOV is made up of B0 added before I1 in figure 28.
In this case, if it is assumed that the VOP_time_increment for the GOV header B0 is determined with the GOV I / P-VOP presentation time as standard, that is, for example, if it is assumed to be determined with the presentation time of the I1 as standard, the value will be a negative value, which is disadvantageous as described above.
Therefore, the semantics of the VOP_time_increment for the B-VOP that occurs before the I-VOP of the GOV (the B-VOP that occurs before the I-VOP of the GOV that occurs first) changes as follows:
That is, the VOP_time_increment for such a B-VOP has a differential value between the GOV time_code and the B-VOP presentation time. In this case, when the B0 display time is, for example, 0h: 12m: 35sec: 200msec and when the GOV timecode is, for example, 0h: 12m: 35sec, as illustrated in figure 29, the VOP_time_increment for B0 is 350 ms (= 0h: 12m: 35sec: 200msec - 0h: 12m: 35sec). If done this way, the VOP_time_increment will always have a positive value.
With the two aforementioned changes in the semantics of the VOP_time_increment, a correlation can be made between the time_code of a GOV and the module_time_base and the VOP_time_increment of a VOP. Furthermore, with this, the absolute time (presentation time) of each VOP can be specified.
Next, Figure 30 shows the relationship between the time_code of a GOV and the module_time_base and the VOP_time_increment of a VOP in the case where the interval between the I-VOP presentation time and the B-VOP presentation time predicted from the I-VOP is equal to greater than 1 second (exactly speaking, 1.023 seconds).
In figure 30, the GOV is made up of I1, B2, B3, B4 and P6 arranged in order of presentation. The B4 is presented at the instant after 1 second from the presentation time of the last presented I1 (I-VOP).
In this case, when the B4 presentation time is encoded with the above-mentioned VOP_time_increment, whose semantics have been changed, the VOP_time_increment is 10 bits as described above, and can only express the time up to 1023. For this reason, it cannot express a time greater than 1.023 seconds. Therefore, the semantics of the VOP_time_increment is further changed and the semantics of the module_time_base is also changed in order to cope with that case.
In this embodiment, such changes are made, for example, either by the first method or by the second method.
That is, in the first method, the elapsed time between the presentation time of an I / P-VOP and the presentation time of a B-VOP predicted from the I / P-VOP is detected, with an accuracy within one second. For time, the unit of a second is expressed as the base_module_time, while the unit of a millisecond is expressed as the increment_time_VOP.
ES 2 323 482 T3
Figure 31 shows the relationship between the time_code for a GOV and the module_time_base and the VOP_time_increment for a VOP, in the case in which the module_time_base and the VOP_time_increment have been coded in the case illustrated in figure 30, according to the first method.
That is, in the first method, the addition of the modulo_time_base is allowed not only for an I-VOP and a P-VOP, but also for a B-VOP. And the modulo_time_base added to a B-VOP does not represent the switching of sync points, but represents the units carried from a second unit obtained from the presentation time of the last presented I / P-VOP.
Furthermore, in the first method, the time after the units carried from a second unit from the presentation time of the last presented I / P-VOP, indicated by the module_time_base added to a B-VOP, is subtracted from the B-VOP presentation time, and the resulting value is set as VOP_time_increment.
Therefore, according to the first method, in figure 30, it is assumed that the presentation time of I1 is 0h: 12m: 35sec: 350msec and also the presentation time of B4 is 0h: 12m: 36sec: 550msec , then the difference between the presentation times of I1 and B4 is 1200 ms over 1 second, and therefore the module_time_base (illustrated by a ▼ mark in figure 31) that indicates the units that are carried from a second unit from the presentation time of the last I1 presented, is added to B4 as illustrated in the figure 31. More specifically, the module_time_base to be added to B4 is 10, which represents the units that are carried with the value of 1 sec. which is the value of the 1 second digit in 1200 ms. And the VOP_time_increment for B4 is 200, which is the value less than 1 second, obtained from the difference between the presentation times between I1 and B4 (the value is obtained by subtracting from the presentation time of B4, the time after the units of a second that have been obtained since the presentation time of the last I / P-VOP presented, indicated by the base_time_module for B4).
The aforementioned process for the module_time_base and for the VOP_time_increment, according to the first method, is carried out in the encoder by means of the VLC unit 36 illustrated in Figures 11 and 12, and in the encoder by means of the unit 102 IVLS illustrated in Figures 17 and 18.
Therefore, the process for the module_time_base and for the VOP_time_increment that is performed by the VLC unit 36 will be described first, with reference to the flow chart of FIG. 32.
The VLC unit 36 splits a sequence of VOPs into the GOVs and performs the process for each of the GOVs. Note that the GOV is so constituted that it includes at least one VOP that is encoded by intracoding.
If a GOV is received, the VLC unit 36 will set the received time as the absolute encoding start time of the GOV, and the GOV will be encoded to within one second of the absolute encoding start time, as time_code ( encodes the absolute encoding start time up to the one second digit). The encoded time_code is contained in a string of encoded bits. Each time an I / P-VOP constituting the GOV is received, the VLC unit 36 sets the I / P-VOP to an attention I / P-VOP, calculates the module_time_base and the I / P-VOP_VOP_time_increment attention, according to the flow chart of figure 32 and performs the encoding.
That is, in VLC unit 36, firstly, in step S1, 0B is set (where B represents a binary number) in the module_time_base and 0 is also set for the VOP_time_increment, so the module_time_base and VOP_time_increment are reset.
And in step S2, it is judged whether the attention I / P-VOP is the first I-VOP of a GOV to be processed (hereinafter referred to as the process GOV object). In step S2, in the case where the I / P-VOP is judged to be the first I-VOP of the process object GOB, step S2 advances to step S4. In step S4, the difference between the time_code of the process GOV object and the one second precision of the process I / P-VOP (in this case, the first I-VOP of the process GOV object), that is, the The difference between the time_code and the second digit of the attention I / P-VOP display time is calculated and set to a variable D. Then, step S4 advances to step S5.
Furthermore, in step S2, in the case where the attention I / P-VOP is judged not to be the first I-VOP of the process object GOV, step S2 advances to step S3. In step S3, the differential value between the seconds digit of the presentation time of the last I / P-VOP of attention, and the seconds digit of the presentation time of the last I / P-VOP presented (which is displayed immediately before the attention I / P-VOP of the VOP constituting the process object GOV) and the differential value is set in variable D. Then, step S3 continues to step S5.
In step S5, it is judged whether the variable D is equal to 0. That is, it is judged whether the difference between the time_code and the seconds digit of the attention I / P-VOP presentation time is equal to 0, or it is judged if the differential value between the seconds digit of the attention I / P-VOP presentation time and the seconds digit of the presentation time of the last I / P presented is equal to 0. In step S5, in the case where the variable D is judged not to be equal to 0, that is, in the case where the variable D is equal to or greater than 1, the step S5 advances to the step S6, in the which 1 is added as the most significant bit (MSB) of the module_time_base. That is, in this case, when the
ES 2 323 482 T3 modulo_time_base is, for example, 0B immediately before its reset, it is set to 10B. Also, when the modulo_time_base is, for example, 10B, it is set to 110B.
And step S6 continues to step S7, in which variable D is incremented by 1. Then, step S7 returns to step S5. Thereafter, steps S5 to S7 are repeated until it is judged, in step S5, that the variable D is equal to 0. That is, the number of consecutive ones in the module_time_base is the same as the number of seconds corresponding to the difference between the time_code and the second digit of the attention I / P-VOP presentation time, or the differential value between the seconds digit of the attention I / P-VOP presentation time and the seconds digit of the presentation time of the last I / P-VOP presented. And the module_time_base has a 0 in the least significant digit (LSD) of it.
And in step S5, in the case where the variable D is judged equal to 0, step S5 advances to step S8, in which a time finer than the one-second precision of the display time is set. of the attention I / P-VOP, that is, the time in which the units of milliseconds are set to the VOP_time_increment, and the process ends.
In the VLC circuit 36, the modulo_time_base and the VOP_time_increment of an attention I / P-VOP calculated in the aforementioned manner are added to the attention I / P-VOP. With this, it is included in a string of encoded bits.
Note that the module_time_base, the VOP_time_increment, and the time_code are encoded in the VLC circuit 36 by encoding variable length words.
Each time a B-VOP constituting a process GOV object is received, the VLC unit 36 sets the B-VOP in a B-VOP for attention, calculates the base_time_module and the increment_time_VOP of the B-VOP for attention, according to with a flow chart of Figure 33, and performs the encoding.
That is, in the VLC unit 36, in step S11, as in the case of step S1 of FIG. 32, a reset of the module_time_base and the VOP_time_increment is first made.
And step S11 advances to step S12, in which it is judged whether the attention B-VOP occurs before the first I-VOP of the process object GOV. In step S12, in the case where the attention B-VOP is judged to be one that occurs before the first I-VOP of the process object GOV, step S12 advances to step S14. In step S14, the difference between the time_code of the process GOV object and the presentation time of the attention B-VOP (in this case, the B-VOP that is presented before the first I-VOP of the GOV object is calculated. process) and is set to a variable D. Then, step S13 proceeds to step S15. Therefore, in figure 33, a time with the precision of one millisecond (the time to the digit of milliseconds) is set in the variable D (on the other hand, the time with the precision of one second is set in the variable of Figure 32, as described above).
Furthermore, in step S12, in the case where the attention B-VOP is judged to be one that occurs after the first I-VOP of the process object GOV, step S12 advances to step S14. In step S14, the differential value between the presentation time of the attention B-VOP and the presentation time of the last I / P-VOP presented (which is presented immediately before the attention B-VOP of the VOP that constitutes the process object GOV) and the differential value is set to variable D. Then, step S13 advances to step S15.
In step S15 it is judged whether the variable D is greater than 1. That is, it is judged whether the value of the difference between the time_code and the presentation time of the B-VOP for attention is greater than 1, or it is judged whether the The value of the difference between the presentation time of the attention B-VOP and the presentation time of the last I / P-VOP presented is greater than 1. In step S15, in the case where the variable D is judged to be greater than 1, step S15 proceeds to step S17, in which 1 is added as the most significant bit (MSB) of the module_time_base. In step S17, variable D is decremented by 1. Then, step S17 returns to step S15. And until it is judged in step S15 that the variable D is not greater than 1, steps S15 to S17 are repeated. That is to say, with this, the number of consecutive ones in the module_time_base is the same as the number of seconds corresponding to the difference between the time_code and the presentation time of the attention B-VOP or the differential value between the presentation time of the B-vOp of attention and the presentation time of the last I / P-VOP presented. And the base_time_module has a 0 as the least significant digit (LSD) of it.
And in step S15, in the case where the variable D is judged not to be greater than 1, step S15 proceeds to step S18, in which the value of the current variable D, that is, the differential value between the time_code and presentation time of the attention B-VOP, or the millisecond digit to the right of the seconds digit of the differential between the attention B-VOP presentation time and the last I / P presentation time -VOP presented, set to VOP_time_increment, and the process ends.
In the VLC circuit 36, the module_time_base and the VOP_time_increment of an attention B-VOP calculated in the aforementioned manner, are added to the attention B-VOP. With this, it is included in a string of encoded bits.
Next, each time the encoded data for each VOP is received, the IVLC unit 102 processes the VOP as a care VOP. With this process, the IVLC unit 102 recognizes the presentation time of
ES 2 323 482 T3 is a VOP included in an encoded string that is outputted by the VLC unit 36 by dividing a sequence of VOPs into GOVs, and also processing each GOV in the aforementioned manner. Then, the IVLC unit 102 encodes the variable length words so that the VOP is displayed at the recognized presentation time. That is, if a GOV is received, the IVLC unit 102 will recognize the GOV's time_code. Each time an I / P-VOP constituting the GOV is received, the IVLC unit 102 sets the I / P-VOP to an attention I / P-VOP, and calculates the presentation time of the I / P-VOP. based on the module_time_base and the VOP_time_increment of the attention I / P-VOP, according to a flowchart in figure 34.
That is, in the IVLC unit 102, it is first judged in step S21 whether the attention I / P-VOP is the first I-VOP of the process object GOV. In step S21, in the case where it is judged whether the attention I / P-VOP is the first I-VOP of the process object GOV, step S21 proceeds to step S23. In step S23, the time_code of the process GOV object is set to a variable T, and step S23 continues in step S24.
Also, in step S21, in the case where it is judged that the attention I / P-VOP is not the first I-VOP of the process object GOV, step S21 proceeds to step S22. In step S22, a value of up to the second digit of the presentation time of the last I / P-VOP presented (which is one of the VOPs that constitute the process object GOV) presented immediately before is set in the variable T of the I / P-VOP of attention. Then, step S22 continues to step S24.
In step S24, it is judged whether the module_time_base added to the attention I / P-VOP is equal to 0B. In step S24, in the case where it is judged that the module_time_base added to the attention I / P-VOP is not equal to 0B, that is, in the case where the module_time_base added to the attention I / P-VOP includes a 1, step S24 continues to step S25, in which the 1 is removed from the MSB of the module_time_base. Step S25 continues to step S26, in which the variable T is incremented by 1. Then, step S26 returns to step S24. Thereafter, until in step S24 it is judged that the module_time_base added to the attention I / P-VOP is equal to 0B, steps S24 to S26 are repeated. With this, the variable T is increased in the number of seconds that corresponds to the number of ones of the first base_time_module added to the I / P-VOP of attention.
And in step S24, in the case where the module_time_base added to the attention I / P-VOP is equal to 0B, step S24 continues in step S27, in which the time with precision of one millisecond, indicated by the VOP_time_increment. The added value is recognized as the time of presentation of the I / P-VOP of attention and the process ends.
Then when a B-VOP constituting the process GOV object is received, the IVLC unit 102 sets the B-VOP to a care B-VOP and calculates the presentation time of the care B-VOP, based on the base_time_module and in the increment_time_VOP of the B-VOP of attention, according to a flow diagram of figure 35.
That is, in the IVLC unit 102, first, in step S31, it is judged whether the attention B-VOP is one that occurs before the first I-VOP of the process object GOV. In step S31, in the case where the attention BVOP is judged to be one that occurs before the first I-VOP of the process object GOV, step S31 proceeds to step S33. Thereafter, in steps S33 to S37, as in the case of steps S23 to S27 of FIG. 34, a similar process is performed, whereby the attention B-VOP presentation time is calculated.
On the other hand, in step S31, in the case where the attention B-VOP is judged to be one that occurs after the first I-VOP of the process object GOV, step S31 proceeds to step S32. Thereafter, in steps S32 and S34 to S37, as in the case of steps S22 and S24 to S27 of FIG. 34, a similar process is carried out, thereby calculating the presentation time of the B-VOP of attention.
Next, in the second method, the time between the presentation time of an I-VOP and the presentation time of a B-VOP predicted from the I-VOP is calculated, down to the second digit. The value is expressed with the modulo_time_base, while the millisecond precision of the B-VOP presentation time is expressed with the VOP_time_increment. That is, the VM-6.0, as described above, the temporal distance to an I-VOP or P-VOP is set in the increment_time_VOP for a B-VOP, so that the weight, in relation to the I-VOP or P-VOP, which is used as a reference image to carry out the prediction coding of the B-VOP, is determined from the B-VOP on the basis of the temporal distance to the I-VOP or P-VOP arranged through the B- VOP. For this reason, the VOP_time_increment for the I-VOP or P-VOP is different from the time from the sync point marked by the last encoded / decoded module_time_base. However, if the presentation time of a B-VOP and also the I-VOP or P-VOP arranged through the B-VOP is calculated, the time distance between them can be calculated by the difference between them. Therefore, there is very little need to handle only the VOP_time_increment for the B-VOP independently of the I-VOP and P-VOP_VOP_time increments. On the contrary, from the point of view of the efficiency of the process, it is preferable that all the increments_time_VOP (detailed information of the time) for the I-, B- and P-VOP and also the base_time_module (information of time with seconds) are handled in the same way.
Therefore, in the second method, the modulo_time_base and VOP_time_increment for the B-VOP are handled in the same way as for the I / P-VOPs.
ES 2 323 482 T3
Figure 36 shows the relationship between the time_code for a GOV and the module_time_base and the VOP_time_increment in the case where the module_time_base and theVOP_time_increment have been encoded according to the second method, for example, in the case illustrated in Figure 30.
That is, even in the second method, the addition of the modulo_time_base is allowed not only for an I-VOP and a P-VOP, but also for a B-VOP. And the modulo_time_base added to a B-VOP, like the modulo_time_base added to an I / P-VOP, represents the switching of the sync points.
Furthermore, in the second method, the synchronization time marked by the module_time_base added to a B-VOP is subtracted from the B-VOP presentation time, and the resulting value is set as VOP_time_increment.
Therefore, according to the second method, in figure 30, the modulo_time_bases for I1 and B2, presented between the first synchronization point of a GOV (which is the time represented by the GOV time_code) and the marked synchronization point by the time_code + 1 second, they are both 0B. And the values of the unit of milliseconds less than the unit of seconds of the presentation times of I1 and B2, are set in the VOP_time_increments for I1 and B2, respectively. Furthermore, the modulo_time_bases for B3 and B4, presented between the sync point marked by the time_code + 1 second and the synchronization point marked by the time_code + 2 seconds, are both 10B. And the values of the unit of milliseconds less than the unit of seconds of the presentation times of the B3 and B4 are set in the VOP_time_steps for B3 and B4, respectively. Furthermore, the module_time_base for P5, presented between the synchronization point marked by the time_code + 2 seconds and the synchronization point marked by the time_code + 3 seconds, is 110B. And the value of the unit of milliseconds less than the unit of seconds of the display time of P5 is set to VOP_time_increment for P5.
For example, in figure 30, if it is assumed that the presentation time of I1 is 0h: 12m: 35s: 350ms and also that the presentation time of B4 is 0h: 12m: 36s: 550ms, as described above, the modulo_time_bases for I1 and B4 are 0B and 10B, respectively. Also, the VOP_time_steps for I1 and B4 are 0B, 350 ms, and 550 ms (which are the unit of milliseconds of display time), respectively.
The aforementioned process for the module_time_base and the VOP_time_increment, according to the second method, as in the case of the first method, is carried out by the VLC unit 36 illustrated in Figures 11 and 12, and also by the illustrated IVLC unit 102 in Figures 17 and 18.
That is, the VLC unit 36 calculates the module_time_base and the VOP_time_increment for an I / P-VOP, in the same way as in the case of FIG. 32.
Furthermore, for a B-VOP, each time the B-VOP constituting a GOV is received, the VLC unit 36 sets the B-VOP to an attention B-VOP and calculates the module_time_base and the VOP_time_increment of the B-VOP. attention, according to a flow chart in Figure 37.
That is, in the VLC unit 36, firstly, in step S41 a reset is made of the module_time_base and the VOP_time_increment, in the same way as in the case of step S1 of FIG. 32.
And step S41 continues to step S42, in which it is judged whether the attention B-VOP is the one that occurs before the first I-VOP of a GOV to be processed (a process GOV object). In step S42, in the case where it is judged whether the attention B-VOP is one that occurs before the first I-VOP of the process object GOV, step S42 proceeds to step S44. In step S44, the difference between the time_code of the process GOV object and the precision in seconds of the attention B_VOP, that is, the difference between the time_code and the seconds digit of the display time of the B-VOP of attention, and set to a variable D. Then, step S44 continues to step S45.
Furthermore, in step S42, in the case where the attention B-VOP is judged to be the one that occurs after the first I-VOP of the process object GOV, step S42 continues in step S43. In step S43, the differential value between the seconds digit and the presentation time of the attention B-VOP and the seconds digit of the presentation time of the last presented I / P-VOP (which is one of the VOPs constituting the process object GOV, presented immediately before the attention B-VOP) and the differential value is set in variable D. Then, step S43 continues to step S45.
In step S45 it is judged whether the variable D is equal to 0. That is, it is judged whether the difference between the time_code and the seconds digit of the presentation time of the attention B-VOP is equal to 0, or it is judged if the differential value between the seconds digit of the attention B-VOP presentation time and the seconds digit of the presentation time of the last I / P-VOP presented is equal to 0 seconds. In step S45, in the case where it is judged whether the variable D is not equal to 0, that is, in the case where the variable D is equal to or greater than 1, the step S45 continues in the step S46, at which adds 1 to the MSB of the module_time_base.
And step S46 advances to step S47, in which variable D is incremented by 1. Then, step S47 returns to step S45. Thereafter, until step S45 in which it is judged whether the variable D is equal to 0, steps S45 to S47 are repeated. That is, with this, the number of consecutive 1s in the module_time_base is the same as
ES 2 323 482 T3 the number of seconds corresponding to the difference between the time_code and the seconds digit of the attention B-VOP presentation time or the differential value between the seconds digit of the B-VOP presentation time attention and the seconds digit of the presentation time of the last I / P-VOP presented. And the base_time_module has a 0 in the LSD of it.
And in step S45, in the case where the variable D is judged equal to zero, step 45 continues in step S48, in which the time with precision finer than the seconds of the display time of the B- Attention VOP, that is, the time in milliseconds, is set to VOP_time_increment, and the process ends.
On the other hand, for an I / P-VOP, the IVLC unit 102 calculates the presentation time of the I / P-VOP, based on the module_time_base and the VOP_time_increment, in the same way as the aforementioned case of the figure 3. 4.
In addition, for a B-VOP, each time the B-VOP constituting a GOV is received, the IVLC unit 102 sets the B-VOP to a care B-VOP and calculates the presentation time of the B-VOP of Attention, based on the module_time_base and on the VOP_time_increment of the Attention B-VOP, according to a flow chart of figure 38.
That is, in the IVLC unit 102, first, in step S51 it is judged whether the attention B-VOP is one that occurs before the first I-VOP of the process object GOV. In step S51, in the case where the attention B-VOP is judged to be the one presented before the first I-VOP of the process object GOV. In step S51, in the case where the attention B-VOP is judged to be one that occurs before the first I-VOP of the process object GOV, step S51 proceeds to step S52. In step S52, the time_code of the process object GOV is set to a variable T, and step S52 continues in step S54.
Also, in step S51, in the case where the attention B-VOP is judged to be one that occurs after the first I-VOP of the process object GOV, step S51 continues in step S53. In step S53, a value up to the second digit of the presentation time of the last I / P-VOP presented is set in the variable T (which is one of the VOPs that constitutes the process object GOV, presented immediately before the B-VOP of attention). Then, step S53 continues to step S54.
In step S54, it is judged whether the modulus_time_base added to the attention B-VOP is equal to 0B. In step S54, in the case where it is judged that the modulo_time_base added to the attention B-VOP is not equal to 0B, that is, in the case where the module_time_base added to the attention B-VOP includes 1, the step S54 proceeds to step S55, in which the 1 of the MSB is removed from the module_time_base. Step S55 continues to step S56, in which the variable T is incremented by 1. Then, step S56 returns to step S54.
Thereafter, until it is judged in step S54 that the modulo_time_base added to the attention B-VOP is equal to 0B, steps S54 to S56 are repeated. With this, the variable T is increased in the number of seconds that corresponds to the number of ones in the first base_time_module added to the B-VOP of attention.
And in step S54, in the case that the module_time_base added to the B-VOP of attention is equal to 0B, step S54 advances to step S57, in which a time with the precision of one millisecond, indicated by the VOP_time_increment. The added value is recognized as time of presentation of the B-VOP of attention, and the process ends.
Thus, in the embodiment of the present invention, the GOV layer for encoding the absolute encoding start time is introduced into the hierarchical constitution of a string of encoded bits. This GOV layer can be inserted at an appropriate position in the encoded bit stream, as well as at the head of the encoded bit stream. Also, the definitions of the module_time_base and VOP_time_increment prescribed in VM-6.0 have been changed as described above. Thus, it becomes possible in all cases to calculate the display time (absolute time) of each VOP, regardless of the arrangement of the image types of the VOPs and the time interval between neighboring VOPs.
Thus, in the encoder, the absolute encoding start time is encoded in a GOV unit, and the module_time_base and VOP_time_increment are also encoded. The encoded data is included in a string of encoded bits. With this, in the decoder, the absolute encoding start time can be decoded in the GOV unit and the module_time_base and VOP_time_increment of each VOP can also be decoded. And the display time of each VOP can be decoded, so that it becomes possible to perform random access efficiently in a GOV unit.
Note that if the number of ones that are added to the modulo_time_base merely increases when a sync point is switched, it will reach a huge number of bits. For example, if an hour (3600 seconds) has elapsed since the time marked by time_code (in the case that a gOv is made up of the VOPs equivalent to that time), the module_time_base will reach 3601 bits, because it is made up of a 1 3600 bits and a 1-bit 0.
ES 2 323 482 T3
Therefore, in MPEG-4, the modulo_time_base is prescribed such that it is reset on an I / P-VOP that is the first presented after a sync point has been switched.
Therefore, for example, as illustrated in figure 39, in the case where a GOV is made up of I1 and B2, presented between the first synchronization point of the GOV (which is the time represented by the GOV time_code) and the synchronization point marked by the time_code + 1 second, B3 and B4 presented between the synchronization point marked by the time_code + 1 second and the synchronization point marked by the time_code + 2 seconds, P5 and P6 presented between the synchronization point marked by the time_code + 2 seconds and the synchronization point marked by the time_code + 3 seconds, B7 presented between the synchronization point marked by the time_code + 3 seconds and the synchronization point marked by the time_code + 4 seconds, and B8 presented between the synchronization point marked by the time_code + 4 seconds and the synchronization point marked by the time_code + 5 seconds, the module_time_bases for I1 and B2, presented between the first synchronization point of the GOV and the synchronization point marked by the time_code + 1 second, are set to 0B.
Furthermore, the module_time_bases for B3 and B4, presented between the synchronization point marked by the time_code + 1 second and the synchronization point marked by the time_code + 2 seconds, are set to 10B. In addition, the module_time_base for P5, displayed between the synchronization point marked by the time_code + 2 seconds and the synchronization point marked by the time_code + 3 seconds, is set to 110B.
Since P5 is a P-VOP that occurs first after the first sync point of a GOV has been switched, to the sync point marked by type_code + 1 second, the modulo_time_base for P5 is set to 0B . The module_time_base for B6, which is presented after B5, is set under the assumption that a reference synchronization point, used to calculate the presentation time of P5, that is, the synchronization point marked by the time_code + 2 seconds in this case, is the first synchronization point of the GOV. Therefore, the modulo_time_base for B6 is set to 0B.
Thereafter, the modulo_time_base for the B7, presented between the sync point marked by the time_code + 3 seconds and the synchronization point marked by the time_code + 4 seconds, is set to 10B. The module_time_base for the B8, presented between the sync point marked by the time_code + 4 seconds and the synchronization point marked by the time_code + 5 seconds, is set to 110B.
The process in the encoder (VLC unit 36) described in Figures 32, 33 and 37 is performed in such a way as to set the modulo_time_base in the aforementioned manner.
Furthermore, in this case, when the first I / P-VOP presented after the synchronization point switching is detected, in the decoder (IVLC unit 102) there is the need to add the number of seconds indicated by the module_time_base for the I / P-VOP, with the time code and calculate the presentation time. For example, in the case illustrated in FIG. 39, the display times from I1 to P5 can be calculated by adding the number of seconds corresponding to the module_time_base of each VOP and the VOP_time_increment with the time_code. However, the presentation times from B6 to B8, presented after P5, which occurs first after a synchronization point switching, need to be calculated by adding the number of seconds corresponding to the module_time_base of each VOP and the VOP_time_increment with the time_code and, furthermore, adding 2 seconds which is the number of seconds corresponding to the module_time_base of P5. For this reason, the process described in Figures 34, 35 and 38 is carried out to calculate the display time in the aforementioned manner.
Then, the aforementioned encoder and decoder can also be realized by specific hardware or by having a computer execute a program that performs the aforementioned process.
Figure 40 shows an example of constitution of an embodiment of a computer that functions as the encoder of Figure 3 or the decoder of Figure 15.
A read-only memory (ROM) 201 stores a boot program, etc. A central processing unit 202 performs various processes by executing a program stored on a hard disk (HD) 206 in a random access memory (RAM) 203. RAM 203 temporarily stores programs that are executed by CPU 202 or the data necessary to be processed by the CPU. An input section 204 is made up of a keyboard or a mouse. The input section 204 is actuated when a necessary command or data is entered. An output section 205 is constituted, for example, by a screen and displays data in accordance with the control of the CPU 202. The HD 206 stores programs to be executed by the CPU 202, image data to be encoded, encoded data (encoded bit string), decoded image data, etc. A communication interface (I / F) 207 receives the image data of an encoding object from external equipment, or transmits a string of encoded bits to external equipment, controlling communication between it and the external equipment. In addition, the communication I / F 207 receives a string of encoded bits from an external unit or transmits decoded image data to an external unit.
By causing the CPU 202 of the computer thus constituted to execute a program that performs the aforementioned process, this computer functions as an encoder of Figure 3 or the decoder of Figure 15.
ES 2 323 482 T3
In the embodiment of the present invention, although the VOP_time_increment represents the display time of a VOP in the unit of milliseconds, the VOP_time_increment can also be made as follows. That is, the time between a sync point and the next sync point is divided into N points, and the VOP_time_increment can be set to a value representing the nth position of the divided point corresponding to the presentation time of a VOP. In the case where the VOP_time_increment is defined like this, if N = 1000, it will represent the presentation time of a VOP in the unit of milliseconds. In this case, although the information on the number of divided points between contiguous synchronization points is required, the number of divided points can be predetermined or the number of divided points included in a layer higher than the GOV layer can be transmitted to a decoder. .
According to the image encoder and the exposed image coding method, one or more layers of each sequence of objects constituting an image are divided into a plurality of groups, and the groups are encoded. Thus, it becomes possible to have random access to the coded result in a group unit.
According to the image decoder and the image decoding method set forth, a string of coded bits is decoded, obtained by dividing one or more layers of each sequence of objects that make up the image into a plurality of groups, and also encoding the groups. Thus, it becomes possible to have random access to the encoded bit string in a group unit and to decode the bit string.
According to the disclosed distribution means, a string of coded bits, obtained by dividing one or more layers of each sequence of objects that make up the image, is distributed into a plurality of groups and also coding the groups. Thus, it becomes possible to have random access to the encoded bit string in a unit of groups.
According to the exposed image encoder and coding method, the time information is generated with precision of one second, which indicates the time accurately within one second, and the detailed information of time is generated, which indicates a time period between the time information accurate to one second, directly before the I-VOP, P-VOP or B-VOP display time, and the display time within a precision finer than one second. Thus, it becomes possible to recognize the presentation times of the I-VOP, P-VOP and B-VOP on the basis of the time information with precision of one second and the detailed time information, and to perform a random access on the basis of the recognition result.
According to the picture decoder and decoding method set forth, the I-VOP, P-VOP and B-VOP display times are calculated based on one-second precision timing information and detailed information from time. Therefore, it becomes possible to perform a random access, based on the presentation time.
According to the exposed distribution means, a string of coded bits is distributed that is obtained by generating time information with precision of one second, which indicates the time with precision within one second, also generating detailed information of the time that indicates a period. time between time information accurately to one second, directly before the I-VOP display time, P-VOP or B-VOP and the display time with a precision finer than one second precision, and further adding the time information with one second precision and the detailed time information to a corresponding I-VOP, P -VOP or B-VOP as information that indicates the presentation time of said I-VOP, P-VOP or B-VOP. Thus, it becomes possible to recognize the presentation times of the I-VOP, P-VOP and B-VOP on the basis of the time information with precision of one second and the detailed time information, and to perform a random access on the basis of the recognition result.
Industrial applicability
The present invention can be used in image information recording-regeneration units, in which dynamic image data is recorded on storage media, such as a magneto-optical disk, magnetic tape, etc., and also recorded data are regenerated and presented on a screen. The invention can also be used in video conferencing systems, videophone systems, broadcasting equipment and multimedia database retrieval systems, in which dynamic image data is transmitted from a transmitting side to a receiving side through a transmission path and, on the receiving side, the received dynamic data is displayed or edited and recorded.
Contents19
40 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40
50 members in 14 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 19970099683 | Japan | – | |
| 9968397 | Japan | A | |
| 9968397 | Japan | A | |
| 996839798911112 | – | – | – |
| JP19970099683 | – | – | – |
Members50
| Document | Office | Kind | |
|---|---|---|---|
| CA2255923A1 | Canada | A1 | |
| CA2421090A1 | Canada | A1 | |
| WO9844742A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU6520298A | Australia | A | |
| JPH10336669A | Japan | A | |
| ID20680A | Indonesia | A | |
| EP0914007A1 | European Patent Office (EPO) | A1 | |
| CN1220804A | China | A | |
| IL127274A0 | Israel | A0 | |
| KR20000016220A | Republic of Korea | A | |
| TW398150B | Taiwan Province of China | B | |
| JP2001036911A | Japan | A | |
| AU732452B2 | Australia | B2 | |
| CN1312655A | China | A | |
| EP0914007A4 | European Patent Office (EPO) | A4 | |
| EP1152622A1 | European Patent Office (EPO) | A1 | |
| US6414991B1 | United States of America | B1 | |
| US2002114391A1 | United States of America | A1 | |
| US2002118750A1 | United States of America | A1 | |
| US2002122486A1 | United States of America | A1 | |
| HK1043461A1 | Hong Kong, China | A1 | |
| HK1043707A1 | Hong Kong, China | A1 | |
| JP3380980B2 | Japan | B2 | |
| JP3380983B2 | Japan | B2 | |
| US6535559B2 | United States of America | B2 | |
| US2003133502A1 | United States of America | A1 | |
| US6643328B2 | United States of America | B2 | |
| KR100417932B1 | Republic of Korea | B1 | |
| KR100418141B1 | Republic of Korea | B1 | |
| CN1185876C | China | C | |
| CN1186944C | China | C | |
| CA2421090C | Canada | C | |
| CA2255923C | Canada | C | |
| CN1630375A | China | A | |
| HK1043461B | Hong Kong, China | B | |
| HK1080650A1 | Hong Kong, China | A1 | |
| IL127274A | Israel | A | |
| US7302002B2 | United States of America | B2 | |
| EP0914007B1 | European Patent Office (EPO) | B1 | |
| EP1152622B1 | European Patent Office (EPO) | B1 | |
| AT425637T | Austria | T | |
| AT425638T | Austria | T | |
| ATE425637T1 | Austria | T1 | |
| ATE425638T1 | Austria | T1 | |
| ES2323358T3 | Spain | T3 | |
| ES2323482T3This record | Spain | T3 | |
| EP0914007B9 | European Patent Office (EPO) | B9 | |
| EP1152622B9 | European Patent Office (EPO) | B9 | |
| CN100579230C | China | C | |
| IL167288A | Israel | A |
Numbers
- Publication
- 2323482
- Publication, DOCDB
- 2323482
- Publication, EPODOC
- ES2323482T
- Application
- 98911112
- Application, DOCDB
- 98911112
- Application, EPODOC
- ES19980911112T
Titles2
- Spanish
- DISPOSITIVO DE CODIFICACION DE IMAGENES, METODO DE CODIFICACION DE IMAGENES, DISPOSITIVO DE DESCODIFICACION DE IMAGENES, METODO DE DESCODIFICACION DE IMAGENES, Y MEDIO SUMINISTRADOR.
- English
- IMAGE CODING DEVICE, IMAGE CODING METHOD, IMAGE DECODING DEVICE, IMAGE DECODING METHOD, AND SUPPLY MEDIUM.
Classification
- CPC, 8
- H04N19/00
- H04N19/177
- H04N19/70
- H04N19/61
- H04N19/29
- H04N19/33
- H04N19/31
- H04N5/93
- IPC, 22
- H04N5 91
- G06T9 00
- G11B20 12
- G11B27 10
- H03M7 30
- H03M7 40
- H04N5 92
- H04N5 93
- H04N7 24
- H04N19 20
- H04N19 31
- H04N19 33
- H04N19 423
- H04N19 50
- H04N19 503
- H04N19 51
- H04N19 593
- H04N19 60
- H04N19 61
- H04N19 625
- H04N19 70
- H04N19 91