Moving picture decoding method
Abstract
An image coding method, comprising: generating commands indicating a correspondence between reference images and reference indexes that designate the reference images, in which a reference index is assigned to a reference image and a plurality of indexes of reference are assigned to each of at least one reference image in said commands; determine a maximum value of the benchmarks; determine a plurality of sets of weighting coefficients for the reference indices by determining a set of weighting coefficients for each reference index; select a reference index and a reference image corresponding to the reference index for a current block to be encoded, the reference image corresponding to the reference index to which reference is made when a motion compensation is carried out in the current block that it will be coded; generate a predictive image of the current block by carrying out a linear prediction in pixel values of a reference block, applying the set of weighting coefficients corresponding to the reference index selected in said selection, obtaining the reference block from the reference image when movement compensation is carried out in the current block to be encoded; generate a prediction error that is a difference between the current block to be encoded and the predictive image of the current block; and provide an encoded image signal obtained by encoding the reference index selected in dichaselection, information indicating the maximum value of the reference indices, the commands, the plurality of sets of weighting coefficients and the prediction error, in which the information indicating the maximum value of the reference indices is located in a common image information area included in the encoded image signal.

Term
Term ended
Projected expiry passed 22 July 2023, 3.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
2 claims: 2 independent, 0 dependent
- 1ES 2 401 774 T3 REIVINDICACIONES 1. - Un procedimiento de codificación de imágenes, que comprende:generar comandos que indican una correspondencia entre imágenes de referencia e índices de referencia que designan a las imágenes de referencia, en el que un índice de referencia se asigna a una imagen de referencia y una pluralidad de índices de referencia se asignan a cada una de al menos una imagen de referencia en dichos comandos;determinar un valor máximo de los índices de referencia;determinar una pluralidad de conjuntos de coeficientes de ponderación para los índices de referencia determinando un conjunto de coeficientes de ponderación para cada índice de referencia;seleccionar un índice de referencia y una imagen de referencia correspondiente al índice de referencia para un bloque actual que va a codificarse, correspondiendo la imagen de referencia al índice de referencia al que se hace referencia cuando se lleva a cabo una compensación de movimiento en el bloque actual que va a codificarse;generar una imagen predictiva del bloque actual llevando a cabo una predicción lineal en valores de píxel de un bloque de referencia, aplicando el conjunto de coeficientes de ponderación correspondiente al índice de referencia seleccionado en dicha selección, obteniéndose el bloque de referencia a partir de la imagen de referencia cuando se lleva a cabo la compensación de movimiento en el bloque actual que va a codificarse;generar un error de predicción que es una diferencia entre el bloque actual que va a codificarse y la imagen predictiva del bloque actual;y proporcionar una señal de imagen codificada obtenida codificando el índice de referencia seleccionado en dicha selección, información que indica el valor máximo de los índices de referencia, los comandos, la pluralidad de conjuntos de coeficientes de ponderación y el error de predicción, en el que la información que indica el valor máximo de los índices de referencia está ubicada en un área de información común de imagen incluida en la señal de imagen codificada.
- 2- Un aparato de codificación de imágenes, que comprende:una unidad de generación de comandos que puede hacerse funcionar para generar comandos que indican una correspondencia entre imágenes de referencia e índices de referencia que designan a las imágenes de referencia, en el que un índice de referencia se asigna a una imagen de referencia y una pluralidad de índices de referencia se asignan a cada una de al menos una imagen de referencia en dichos comandos;una unidad de determinación de valor máximo que puede hacerse funcionar para determinar un valor máximo de los índices de referencia;una pluralidad de conjuntos de coeficientes de ponderación para los índices de referencia determinando un conjunto de coeficientes de ponderación para cada índice de referencia;una unidad de selección que puede hacerse funcionar para seleccionar un índice de referencia y una imagen de referencia correspondiente al índice de referencia para un bloque actual que va a codificarse, correspondiendo la imagen de referencia al índice de referencia al que se hace referencia cuando se lleva a cabo una compensación de movimiento en el bloque actual que va a codificarse;una unidad de generación de imágenes predictivas que puede hacerse funcionar para generar una imagen predictiva del bloque actual llevando a cabo una predicción lineal en valores de píxel de un bloque de referencia, aplicando el conjunto de coeficientes de ponderación correspondiente al índice de referencia seleccionado en dicha unidad de selección, obteniéndose el bloque de referencia a partir de la imagen de referencia cuando se lleva a cabo la compensación de movimiento en el bloque actual que va a codificarse;una unidad de generación de errores de predicción que puede hacerse funcionar para generar un error de predicción que es una diferencia entre el bloque actual que va a codificarse y la imagen predictiva del bloque actual;y una unidad de provisión de señales de imágenes codificadas que puede hacerse funcionar para proporcionar una señal de imagen codificada obtenida codificando el índice de referencia seleccionado en dicha unidad de selección, información que indica el valor máximo de los índices de referencia, los comandos, la pluralidad de conjuntos de coeficientes de ponderación y el error de predicción. ES 2 401 774 T3 en el que la información que indica el valor máximo de los índices de referencia está ubicada en un área de información común de imagen incluida en la señal de imagen codificada.
Independent claims2
595 paragraphs in 18 sections, as filed
ES 2 401 774 T3
DESCRIPTION
Encoding procedure and decoding procedure for moving images
Technical field
The present invention relates to a moving image coding method and a moving image decoding method and, in particular, to a coding method and a decoding method using inter-image prediction with reference to coded images. previously.
Previous technique
With the development of multimedia applications, fully handling all kinds of multimedia information such as video, audio and text has become commonplace. To that end, the digitization of all this multimedia information allows it to be treated in its entirety. However, since digitized images have a huge amount of data, image information compression techniques are absolutely necessary to store and transmit such information. It is also important to standardize such compression techniques for the interworking of compressed image data. There are international standards for image compression techniques, such as H.261 and H.263 standardized by the International Telecommunication Union - Telecommunication Standardization Sector (ITU-T) and MPEG-1, MPEG-4 and other standards. standardized by the International Organization for Standardization (ISO). The ITU is currently working on standardizing H.26L as the latest standard for encoding images.
In general, in the coding of moving images, the amount of information is compressed reducing redundancies in both temporal and spatial directions. Therefore, in interimage prediction coding, which aims to reduce temporal redundancy, the movement of a current image is estimated by each block with reference to previous or subsequent images to create a predictive image, subsequently encoding differential values between the Predictive images obtained and the current image.
In this case, the term "image" represents a single layer of an image and represents a frame when used in the context of a progressive image, while it represents a frame or a field in the context of an interlaced image. In this case, the interlaced image is a single frame that is made up of two fields that have different times, respectively. In the interlaced image encoding and decoding process, a single frame can be treated as one frame, as two fields, as a frame structure, or as a field structure in each frame block.
The following description is provided assuming that an image is a frame in a progressive image, but the same description can be given even assuming that an image is a frame or a field in an interlaced image.
FIG. 30 is a diagram explaining types of images and the reference relationships between them.
An image such as image I1, which is a coded intra-image prediction without reference to any image, is referred to as an I-image. An image such as P10 image, which is a coded inter-image prediction with reference to only an image is called a P-image and an image, which can be a coded inter-image prediction with reference to two images at the same time, is called a B-image.
B images, like B6, B12, and B18 images, can reference two images located at arbitrary temporal directions. The reference images can be designated in each block, with respect to which the movement is estimated, and are discriminated between a first reference image described above in a coded stream obtained by coding images and a second reference image described later in the stream encoded.
However, to encode and decode previous images it is necessary that the reference images are already encoded and decoded. FIG. 31A and 31B show examples of B-image encoding and decoding order. FIG. 31A shows a display order of the images, and FIG. 31B shows a reordered encoding and decoding order from the display order shown in FIG. 31A. These diagrams show that the pictures are rearranged so that the pictures referenced by pictures B3 and B6 are already encoded and decoded.
A method for creating a predictive image in case the above-mentioned B-image is encoded with reference to two images at the same time will be explained in detail using FIG. 32. It should be noted that a predictive image is created by decoding in exactly the same way.
ES 2 401 774 T3
Image B4 is a current B-image to be encoded, and blocks BL01 and BL02 are current blocks to be encoded belonging to the current B-image. By referring to a block BL11 belonging to image P2 as a first reference image and to a block BL21 belonging to image P3 as a second reference image, a predictive image is created for block BL01. Also, by referring to a block BL12 belonging to image P2 as a first reference image and to a block BL22 belonging to image P1 as a second reference image, a predictive image is created for block BL02 (see document 1, which is not a patent).
FIG. 33 is a diagram explaining a procedure for creating a predictive image for the current block BL01 to be encoded using the two referenced blocks BL11 and BL21. The following explanation will assume in this case that the size of each block is 4 by 4 pixels. Assuming that Q1 (i) is a pixel value of BL11, that Q2 (i) is a pixel value of BL21, and that P (i) is a pixel value of the predictive image for the target block BL01, the value of pixel P (i) can be calculated using a linear prediction equation such as equation 1 below. "i" indicates the position of a pixel, and in this example, "i" has values between 0 and 15.
P (¡) = (wl x Ql (i) + w2 x Q2 (i)) / pow (2, d) + c ... Equation 1 (where pow (2, d) indicates the “d” -th power of 2) "w1", "w2", "c" and "d" are the coefficients for carrying out a linear prediction, and these four coefficients are treated as a set of weighting coefficients. This set of weighting coefficients is determined by a reference index that designates an image to which each block refers. For example, four values of w1_1, w2_1, c_1, and d_1 are used for BL01, and w1_2, w2_2, c_2, and d_2 are used for BL02, respectively.
Next, reference indices designating reference images will be explained with reference to FIG. 34 and FIG. 35. A value called an image number, which increases by values of one each time an image is stored in memory, is assigned to each image. In other words, an image number with a value resulting from adding one to the maximum value of existing image numbers is assigned to a newly stored image. However, a reference image is not actually designated using this image number, but rather by using a value called the reference index, which is defined separately. The indices that indicate first reference images are called first reference indices, and the indices that indicate second reference images are called second reference indices, respectively.
FIG. 34 is a diagram explaining a procedure for assigning two reference indices to picture numbers. When there is a sequence of images arranged in the order of display, the image numbers are assigned in the order of encoding. The commands to assign the reference indices to the image numbers are described in a heading of a section that is a subdivision of an image, as the encoding unit, and therefore the allocation of the same is updated every time a section is encoded. The command indicates the differential value between an image number that is assigned a current benchmark and an image number that is assigned a benchmark immediately before the current allocation, serially by the number of benchmarks.
Taking the first benchmark of FIG. 34 as an example, since "-1" is first provided as a command, 1 is subtracted from image number 16 from the current image to be encoded, and thus reference index 0 is assigned to image number 15. Next, since "-4" is provided, 4 is subtracted from image number 15, and thus reference index 1 is assigned to image number 11. Subsequent benchmarks are assigned to respective image numbers in the same processing. The same applies to the second benchmarks.
FIG. 35 shows the result of the assignment of the benchmarks. The first benchmarks and the second benchmarks are assigned to respective image numbers separately, but by examining each benchmark, it becomes apparent that a benchmark is assigned to a picture number.
Next, it will be explained, with reference to FIG. 36 and FIG. 37, a procedure for determining the sets of weights to be used.
An encoded stream of an image is made up of a common image information area and a plurality of section data areas. FIG. 36 shows a structure of a section data area thereof. The section data area is made up of a section header area and a plurality of block data areas. As an example of a block data area, in this case block areas corresponding to BL01 and BL02 are shown in FIG. 32.
"Ref1" and "ref2" included in block BL01 indicate the first reference index and the second index indicating two reference images for this block, respectively. In the section header area, data (pset0,
ES 2 401 774 T3 psetol, pset2, pset3 and pset4) to determine the sets of weighting coefficients for linear prediction are described for ref1 and ref2, respectively. FIG. 37 shows tables of the aforementioned data included in the section header area, by way of example.
Each piece of data indicated by an identifier "pset" has four values, w1, w2, c and d, and is structured to be directly referenced by the values of ref1 and ref1. Additionally, the section header area describes a script idx_cmd1 and idx_cmd2 for assigning reference indices to image numbers.
Using ref1 and ref2 described in BL01 in FIG. 36, one set of weighting coefficients is selected from the table for ref1 and another set thereof is selected from the table for ref2. By carrying out a linear prediction of equation 1 using sets of respective weighting coefficients, two predictive images are generated. A desired predictive image can be obtained by calculating the average of these two predictive images for each pixel.
Furthermore, there is another method for obtaining a predictive image using a predetermined fixed equation, different from the above-mentioned method for generating a predictive image using a prediction equation obtained by sets of weighting coefficients of linear prediction coefficients. In the first procedure, in case an image designated by a first benchmark appears, in the display order, after an image designated by a second benchmark, the following equation 2a is selected, which is a fixed equation composed of fixed coefficients, and, in other cases, the following equation 2b, which is a fixed equation composed of fixed coefficients, is selected to generate a predictive image.
P (¡) - 2x Ql (i) -Q2 (i) ... Equation 2a
P (i) - (Ql (i) + Q2 (i)) / 2 ... Equation 2b
As is evident from the above, this method has the advantage that it is not necessary to encode or transmit the sets of weighting coefficients to obtain the predictive image since the prediction equation is fixed. This method has another advantage in that it is not necessary to encode or transmit a flag to designate the sets of weighting coefficients of linear prediction coefficients since the fixed equation is selected based on the positional relationship between the images. Furthermore, this procedure allows a significant reduction in the amount of processing for linear prediction thanks to a simple linear prediction formula.
(Document 1, which is not a patent)
ITU-T Rec. H.264 | ISO / IEC 14496-10 AVC
Joint Committee Draft (CD) (05/10/2002) (P.34 8.4.3 Re-Mapping of frame numbers indicator,
P.105 11.5 Prediction signal generation procedure)
In the procedure to create a predictive image using sets of weights according to equation 1, since the number of commands to assign reference indices to reference images is the same as the number of reference images, only one index is assigned reference to a reference image and therefore sets of weighting coefficients used for linear prediction of blocks that refer to the same reference image have exactly the same values. There is no problem if the images change uniformly in an image as a whole, but there is a high probability that the optimal predictive image cannot be generated if the respective images change differently. Also, there is the additional problem that the amount of processing for linear prediction increases because the equation includes multiplications.
The final committee draft corresponding to the above document 1, which is not a patent, was published as Text of final committee draft of joint video specification, by Wiegand, T, (ITU-T Rec.H.264 / ISO / IEC 14496- 10 AVC) MPEG02 / N4920).
Document WO 01/86960 A2 discloses a system for video encoding and decoding using short-term and long-term buffers. The reconstruction of each block of an image can be carried out with reference to one of the buffers, so that different parts of an image, or different images of a sequence, can be reconstructed using different buffers. Also provided herein is systems for signaling, between an encoder and a decoder, the use of the above buffers and related address information. For example, the encoder can transmit information identifying video data corresponding to a particular memory of the buffers; and the decoder can transmit information related to the size of the memory buffer.
ES 2 401 774 T3 short term and long term buffer. Buffer sizes can be modified during transmission of video data by including buffer allocation information in the video data. Procedures and apparatus as discussed above are also described in this document.
Description of the invention
In view of the foregoing, an object of the present invention is to provide an image coding method and an image decoding method and apparatus and programs for executing these procedures to allow the assignment of a plurality of reference indices to an image. reference and, therefore, improve the decoding efficiency of benchmarks in case a plurality of benchmarks are assigned and in case a benchmark is assigned.
To achieve this object, the image coding method according to the present invention is structured as follows. The image coding method according to the present invention includes: a reference image storage step that stores an encoded image identified by an image number, such as a reference image, in a storage unit; a command generation step that generates commands indicating a correspondence between reference indices and image numbers, said reference indices designating reference images and coefficients used for the generation of predictive images; a reference image designation step designating a reference image by a reference index, said reference image being used when performing motion compensation on a current block of a current image to be encoded; a stage of generation of predictive images that generates a predictive image by carrying out a linear production in a block by using a coefficient corresponding to the reference index, said block being obtained by estimating the movement in the reference image designated in the stage reference image designation; and a stage of provision of coded signals that provides a coded image signal including a coded signal obtained by coding a prediction error, the commands, the reference index and the coefficient, said prediction error being a difference between the current block of the current image to be encoded and the predictive image, where in the stage of provision of encoded signals, information indicating a maximum reference index value is encoded and input into the encoded image signal.
In this case, it may be structured so that the information indicating the maximum reference index value is input into a common image information area included in the encoded image signal.
According to this structure, when the picture numbers and the benchmarks are matched to each other according to the commands of the decoding apparatus, the information indicating the maximum benchmark value is included in the encoded signal. Therefore, by matching the image numbers and the reference indices according to the commands until the number of the reference indices reaches the maximum value, all the reference indices and the image numbers can be easily mapped to each other. As a result, it is not only possible to assign a plurality of benchmarks to a reference picture, but also to carry out decoding of benchmarks efficiently in case of assigning a plurality of benchmarks or in case of assigning a benchmark index.
In this case, it can be structured so that in the command generation stage, the commands are generated so that at least one reference image, among the reference images stored in the storage unit, has an image number which is assigned a plurality of benchmarks.
Furthermore, it may be structured so that in the reference image designation step, when the plurality of reference indices are matched with the image number of the reference image, one of the reference indices is selected based on coefficients corresponding respectively to said plurality of reference indices, and in the stage of generation of predictive images, linear prediction is carried out using a coefficient corresponding to the reference index selected in the reference image designation step.
According to this structure, a plurality of reference indices are mapped to one image. Therefore, when a linear prediction is carried out using the coefficient corresponding to the reference index designating the reference image, that coefficient can be selected from a plurality of coefficients. In other words, the optimal coefficient can be selected as a coefficient used for linear prediction. As a result, the coding efficiency can be improved.
In this case, it may be structured so that in the predictive imaging stage, linear prediction is carried out using only a bit shift operation, an addition and a subtraction.
According to this structure, no multiplications or divisions that require a large processing load are used, but only bit shifting operations, additions and subtractions that require less processing load are used, so that the amount of processing of linear prediction.
ES 2 401 774 T3
In this case, it can be structured so that at the predictive imaging stage, such as the coefficient used in linear prediction, only a value indicating a direct current component in a linear prediction equation is mapped to the index reference.
According to this structure, it is not necessary to encode coefficient values other than the value indicated by a CC component, so that the encoding efficiency can be effectively improved. Also, since multiplications and divisions that require a large processing load are not used, but rather addition and subtractions that require less processing load are used, the amount of processing for linear prediction can be limited.
In this case, it can be structured so that the reference index has a first reference index indicating a first reference image and a second reference index indicating a second reference image, and in the predictive image generation stage , in case a coefficient generation procedure according to display order information of each of the reference images is used as a procedure to carry out linear prediction, when a reference image designated by the first reference index and a reference image designated by the second reference index have the same display order information, linear prediction is carried out using a predetermined coefficient for each of the first and second benchmarks instead of the coefficient.
Furthermore, it can be structured so that the predetermined coefficients have the same weight.
According to this structure, it is possible to determine the coefficients to carry out linear prediction even in case that two reference pictures have the same display order information, and therefore, the efficiency of coding can be improved.
Furthermore, the image decoding method, the image encoding apparatus, the image decoding apparatus, the image encoding program, the image decoding program, and the encoded image data of the present invention have the same structures. , functions and effects than the image encoding procedure described above.
Furthermore, the image coding method of the present invention may be structured in any of the following ways (1) to (14).
(1) The image coding method according to the present invention includes: a reference image storage step that stores an encoded image identified by an image number in a storage unit; a command generation stage that generates commands to allow a plurality of reference indices to refer to the same image, said commands indicating a correspondence between reference indices and image numbers, said reference indices designating reference images and coefficients used for predictive imaging, and referring to said reference images when motion compensation is performed on a current block of a current image to be encoded and which are arbitrarily selected from a plurality of encoded images stored in the storage unit; a reference image designation step designating a reference image by a reference index, said reference image being referenced when motion compensation is performed on the current block of the current image to be encoded; a stage of generation of predictive images that generates a predictive image by carrying out a linear prediction in a block by using a coefficient corresponding to the reference index that designates the reference image, said block being obtained by estimating movement in the reference image designated in the reference image designation step; and a stage of provision of coded signals that provides a coded image signal that includes a coded signal obtained by coding a prediction error, the commands, the reference index and the coefficient, said prediction error being the difference between the current block of the current frame to be encoded and the predictive image.
(2) According to another image coding method of the present invention, it is possible to assign a plurality of reference indices to the image number of the reference image, and it is also possible to select a reference index from one or more indexes of reference corresponding to respective coded images in the reference image designation step and therefore determine the coefficient used for linear prediction in the predictive imaging stage.
(3) According to another image coding method of the present invention, at least one reference image, out of a plurality of reference images to which a section refers, has an image number that is assigned a plurality of indices reference.
(4) According to another image coding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from among
ES 2 401 774 T3 the plurality of coded images and a second reference index that indicates a second reference frame that is arbitrarily designated from among the plurality of coded frames, where in the stage of generation of predictive images, the linear prediction is carried out carried out in the block using a coefficient corresponding to the first benchmark, and the linear prediction is carried out in the block using a coefficient corresponding to the second reference index, and then a final predictive image for the block is generated by calculating the average of the pixel values in the two predictive images obtained respectively by the predictions linear.
(5) According to another image coding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from the plurality of encoded images and a second reference index that indicates a second reference frame that is arbitrarily designated from among the plurality of coded frames, where in the stage of generation of predictive images, the coefficient used for linear prediction is determined by calculating the average of the coefficients designated by the first selected benchmark and the second selected benchmark, respectively.
(6) According to another image coding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from among the plurality of encoded images and a second reference index that indicates a second reference frame that is arbitrarily designated from the plurality of coded frames, where the sets of coefficients are matched with the first reference index and the second reference index, and in the predictive image generation stage, the predictive image is generated using a part of the set of coefficients corresponding to one of the first and the second benchmark and a part of the set of coefficients corresponding to the other benchmark.
(7) According to another image coding method of the present invention, in an equation used for linear prediction in the stage of generation of predictive images, neither a multiplication nor a division is used, but only an operation of bit shift, an addition and a subtraction. Therefore, linear prediction can be performed by carrying out processing with fewer calculations.
(8) According to another image coding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from among the plurality of encoded images and a second reference index that indicates a second reference frame that is arbitrarily designated from among the plurality of coded frames, where in the stage of generation of predictive images, the predictive image is generated using the coefficient corresponding to any one of the first reference index and the second reference index that is selected, as a coefficient used for the bit-shifting operation, from among the sets of coefficients corresponding to the first index benchmark and second benchmark, and using the average of the coefficients corresponding to the first benchmark and the second benchmark, respectively, as a coefficient used for other operations.
(9) According to another image coding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from among the plurality of encoded images and a second reference index that indicates a second reference frame that is arbitrarily designated from among the plurality of coded frames, where in the stage of generation of predictive images, Only one value indicating a direct current component in a linear prediction equation is used as a coefficient used for linear prediction, and one coefficient is mapped to each of the first benchmark and second benchmark.
(10) Another image encoding method of the present invention includes: a reference image storage step that stores an encoded image identified by an image number in a storage unit; a command generation stage that generates commands to allow a plurality of reference indices to refer to the same image, said commands indicating a correspondence between reference indices and image numbers, said reference indices indicating reference images to which reference is made when a motion compensation is carried out on a current block of a current image to be encoded and which are arbitrarily selected from among a plurality of encoded images stored in the storage unit; a reference image designation step designating a reference image by a reference index, said reference image being referenced when motion compensation is performed on the current block of the current image to be encoded; a stage of generation of predictive images that generates a coefficient from information on the display order of each reference image and that generates a predictive image by carrying out a linear prediction in a block by using the generated coefficient, said block being obtained by an estimate of movement in the designated reference image in the reference image designation step; and a stage of provision of coded signals that provides a coded image signal that includes a coded signal obtained by coding a prediction error, the commands, the reference index and the coefficient, said prediction error being the difference between
ES 2 401 774 T3 the current block of the current frame to be encoded and the predictive image.
(11) According to another image coding method of the present invention, in the predictive image generation stage, as a method to carry out linear prediction, a procedure that uses the coefficient generated according to the display order information and a procedure that uses a predetermined fixed equation toggle to be used based on the display order information indicating a temporal relationship between the reference image designated by the first index benchmark and the benchmark image designated by the second benchmark.
(12) According to another image coding method of the present invention, in the predictive image generation stage, in the event that a method that uses the coefficient generated according to display order information of each of the reference images is use as a procedure to carry out linear prediction, When a reference image designated by the first benchmark and a reference image designated by the second benchmark have the same display order information, linear prediction is carried out using a predetermined coefficient for each of the first and of the second benchmark instead of the coefficient.
(13) According to another image coding method of the present invention, in the predictive image generation stage, when the coefficient is generated using the display order information, the coefficient approaches a power of 2, so that linear prediction can be carried out without using multiplication or division, but using only a bit-shifting operation, an addition, and a subtraction.
(14) According to another image coding method of the present invention, when the approximation is carried out, a rounding up approximation and a rounding down approximation switch to be used based on the display order information indicating a temporal relationship between the reference image designated by the first benchmark and the reference image designated by the second benchmark.
(15) A program of the present invention may be structured to make a computer execute the image coding procedure described in any of the above-mentioned points (1) to (14).
Furthermore, a computer-readable recording medium of the present invention may be structured as described in the following points (16) to (25).
(16) A computer-readable recording medium in which an encoded signal that is a signal of an encoded motion picture is recorded, wherein the encoded signal includes data obtained by encoding the following: a coefficient used to generate a predictive image ; commands to allow a plurality of reference indices to refer to the same image, said commands indicating a correspondence between reference indices and image numbers, said reference indices designating reference images and coefficients used for the generation of predictive images, and referring to said reference images when motion compensation is carried out on a current block of a current image to be encoded and which are arbitrarily selected from a plurality of encoded images stored in a storage unit for storing the encoded images identified by the image numbers; a reference index for designating the coefficient used for generating the predictive image and the reference image used when performing motion compensation on the current block of the current image; the predictive image that is generated by carrying out a linear prediction on a block by using the coefficient corresponding to the reference index that designates the reference image, said block being obtained by estimating movement in the selected reference image.
(17) The encoded signal includes the maximum reference index value.
(18) The maximum value is located in a common image information area included in the encoded signal.
(19) A header of a section that includes a plurality of blocks, a common image information area, or a header of each block, which is included in the encoded signal, includes an indicator that indicates whether or not a coefficient has been encoded. used to generate a predictive image of the block using linear prediction.
(20) A header of a section that includes a plurality of blocks, a common image information area, or a header of each block, which is included in the encoded signal, includes an indicator that indicates whether to generate the predictive image of the block without use a coefficient, but rather using a predetermined fixed equation or whether to generate the predictive image using a predetermined fixed equation using only a coefficient indicating a direct current component.
(21) A header of a section including a plurality of blocks, a common image information area
ES 2 401 774 T3 or a header of each block, which is included in the coded signal, includes an indicator that indicates whether to generate the predictive image of the block using two predetermined equations by switching between them or whether to generate the predictive image using any of the two equations, when the predictive image is generated using a fixed equation consisting of the two equations.
(22) A header of a section that includes a plurality of blocks, a common image information area, or a header of each block, which is included in the encoded signal, includes an indicator that indicates whether or not to generate a coefficient used to generating the predictive image of the block by linear prediction using display order information from a reference image.
(23) A header of a section that includes a plurality of blocks, a common image information area, or a header of each block, which is included in the coded signal, includes an indicator that indicates whether or not to approach a power of 2 a coefficient used to generate the predictive image of the block by linear prediction.
(24) The coded signal includes a flag indicating that the calculation for linear prediction can be carried out without using multiplication or division, but using only a bit shift operation, an addition and a subtraction.
(25) The coded signal includes an indicator that indicates that the calculation for linear prediction can be carried out using only a value that indicates a direct current component.
Furthermore, the image decoding method of the present invention may be structured as described in the following points (26) to (39).
(26) The image decoding method of the present invention includes: a stage of obtaining information from coded images that decodes a coded image signal that includes a coded signal obtained by encoding coefficients used for the generation of predictive images, commands that allow a plurality of reference indices to refer to the same image, indices reference and a prediction error, said commands indicating a correspondence between the reference indices and image numbers, by designating said reference indices reference images and the coefficients, and referring to said reference images when motion compensation is performed on a current block of a current image to be encoded and which are arbitrarily selected from a plurality of images encoded stored in a storage unit; a reference image designation step that designates a reference image according to decoded commands and a decoded reference index, said reference image being used when a motion compensation is carried out on a current block of a current image to be decoded ; a stage of generation of predictive images that generates a predictive image by carrying out a linear prediction in a block by using a coefficient corresponding to the reference index that designates the reference image, said block being obtained by estimating movement in the designated reference image; and a decoded image generation step for generating a decoded image from the predictive image and a decoded prediction error.
(27) According to another image decoding method of the present invention, it is possible to assign a plurality of reference indices to the image number of the reference image and determine the coefficient used for linear prediction in the predictive image generation stage using the decoded reference index in the reference image designation stage.
(28) According to another image decoding method of the present invention, at least one reference image, out of a plurality of reference images to which a section refers, has an image number assigned to a plurality of indices reference.
(29) According to another image decoding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from among a plurality of decoded images and a second reference index that indicates a second reference frame that is arbitrarily designated from among a plurality of decoded frames, where in the predictive image generation stage, linear prediction is carried out in the block by a coefficient corresponding to the first reference index, and linear prediction is carried out in the block by a coefficient corresponding to the second reference index, and then a final predictive image for the block it is generated by calculating the average of the pixel values in the two predictive images respectively obtained by linear prediction.
(30) According to another image decoding method of the present invention, the reference index includes a first reference index indicating a first reference image that is arbitrarily designated by a plurality of decoded images and a second reference index indicating a second reference frame that is arbitrarily designated from a plurality of decoded frames, where in the
ES 2 401 774 T3 predictive imaging stage, the coefficient used for linear prediction is determined by calculating the average of the coefficients designated by the first selected benchmark and by the second selected benchmark, respectively.
(31) According to another image decoding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from a plurality of decoded images and a second reference index that indicates a second reference frame that is arbitrarily designated from a plurality of decoded frames, where sets of coefficients are matched with the first reference index and with the second reference index, and in the stage of generation of predictive images, the predictive image is generated using a part of the set of coefficients corresponding to one of the first and the second benchmark and a part of the set of coefficients corresponding to the other benchmark.
(32) According to another image decoding method of the present invention, in an equation used for linear prediction in the stage of generation of predictive images, neither multiplication nor division are used, but only a shift operation of bits, an addition and a subtraction. Therefore, linear prediction can be performed by carrying out processing with fewer calculations.
(33) According to another image decoding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from among a plurality of decoded images and a second reference index that indicates a second reference frame that is arbitrarily designated from among a plurality of decoded frames, where in the predictive image generation stage, a predictive image is generated using the coefficient corresponding to any one of the first reference index and the second reference index that is selected, as a coefficient used for the bit shift operation, from among the sets of coefficients corresponding to the first index benchmark and second benchmark, and using the average of the coefficients corresponding to the first benchmark and the second benchmark, respectively, as a coefficient used for other operations.
(34) According to another image decoding method of the present invention, the reference index includes a first reference index that indicates a first reference image that is arbitrarily designated from among a plurality of decoded images and a second reference index that indicates a second frame that is arbitrarily designated from among a plurality of decoded frames, where in the predictive image generation stage, Only one value indicating a direct current component in a linear prediction equation is used as a coefficient used for linear prediction, and one coefficient is mapped to each of the first benchmark and second benchmark.
(35) The image decoding method of the present invention includes: a first stage that decodes an encoded image signal that includes an encoded signal obtained by encoding commands that allow a plurality of reference indices to refer to the same image, reference indices and a prediction error, said commands indicating a correspondence between reference indices and image numbers, said reference indices designating reference images referred to when motion compensation is carried out on a current block of a current image to be encoded and which are arbitrarily selected from a plurality of encoded images stored in a storage unit; a reference image designation step that designates a reference image according to decoded commands and a decoded reference index, said reference image being used when a motion compensation is carried out on a current block of a current image to be decoded ; a stage of generation of predictive images that generates a coefficient from information on the display order of each reference image and that generates a predictive image by carrying out a linear prediction in a block by using the generated coefficient, said block being obtained by an estimate of movement in the designated reference image; and a decoded image generation step that generates a decoded image from the predictive image and a decoded prediction error.
(36) In the predictive imaging stage, as a procedure to carry out linear prediction, a procedure that uses a coefficient generated according to the display order information and a procedure that uses a predetermined fixed equation switch to use based on the display order information indicating a temporal relationship between the reference image designated by the first index benchmark and the benchmark image designated by the second benchmark.
(37) According to another image decoding method of the present invention, in the predictive image generation stage, in the event that a method that uses the coefficient generated according to display order information of each of the reference images is use as a procedure to carry out linear prediction, when a reference image designated by the first benchmark and a reference image designated by the second benchmark have the same order information
ES 2 401 774 T3 display, linear prediction is carried out using a predetermined coefficient for each of the first and second benchmarks instead of the coefficient.
(38) According to another image decoding method of the present invention, in the predictive image generation stage, when the coefficient is generated using the display order information, the coefficient approaches a power of 2, such that linear prediction can be carried out without using multiplication or division, but using a bit-shifting operation, an addition, and a subtraction.
(39) According to another image decoding method of the present invention, when approximation is carried out, a rounding-up approximation and a rounding-down approximation switch to be used based on the display order information indicating a temporal relationship between the reference image designated by the first benchmark and the reference image designated by the second benchmark.
(40) A program of the present invention may be structured to cause a computer to execute the image decoding procedure described in any of the points (26) to (39) mentioned above.
As described above, the moving image encoding method and the moving image decoding method of the present invention allow the creation of a plurality of candidates for sets of weighting coefficients used in a linear prediction to generate a predictive image and, therefore, allow the selection of the optimal set for each block. As a result, the benchmarks can be decoded more efficiently in any case where a plurality of benchmarks are assigned or in case one benchmark is assigned. Furthermore, since the present invention enables a significant improvement in coding efficiency, it is highly efficient for coding and decoding of moving pictures.
Brief description of the drawings
FIG. 1 is a block diagram showing the structure of a coding apparatus in a first embodiment of the present invention.
FIG. 2 is a block diagram showing the structure of a decoding apparatus in a sixth embodiment of the present invention.
FIG. 3 is a schematic diagram explaining a procedure for assigning reference indices to image numbers.
FIG. 4 is a schematic diagram showing an example of the relationship between benchmarks and image numbers.
FIG. 5 is a schematic diagram explaining motion compensation operations.
FIG. 6 is a schematic diagram explaining the structure of an encoded stream.
FIG. 7 is a schematic diagram showing an example of weight coefficient sets of linear prediction coefficients.
FIG. 8 is a functional block diagram showing the generation of a predictive image in a coding apparatus.
FIG. 9 is another functional block diagram showing the generation of a predictive image in the encoding apparatus.
FIG. 10A and 10B are still further functional block diagrams showing the generation of a predictive image in the encoding apparatus.
FIG. 11 is yet another functional block diagram showing the generation of a predictive image in the encoding apparatus.
FIG. 12 is a schematic diagram explaining the structure of an encoded stream.
FIG. 13 is a schematic diagram showing an example of weight coefficient sets of linear prediction coefficients.
FIG. 14 is a schematic diagram explaining the structure of an encoded stream.
ES 2 401 774 T3
FIG. 15 is a schematic diagram showing an example of weight coefficient sets of linear prediction coefficients.
FIG. 16 is a functional block diagram showing the generation of a predictive image in the encoding apparatus.
FIG. 17A and 17B are schematic diagrams explaining the structure of a coded stream and an example of flags.
FIG. 18 is a functional block diagram showing the generation of a predictive image in a decoding apparatus.
FIG. 19 is another functional block diagram showing the generation of a predictive image in the decoding apparatus.
FIG. 20A and 20B are still further functional block diagrams showing the generation of a predictive image in the decoding apparatus.
FIG. 21 is yet another functional block diagram showing the generation of a predictive image in the decoding apparatus.
FIG. 22 is yet another functional block diagram showing the generation of a predictive image in the decoding apparatus.
FIG. 23 is a schematic diagram explaining the structure of an encoded stream.
FIG. 24 is a schematic diagram explaining the structure of an encoded stream.
FIG. 25A, B and C are illustrations of a recording medium for storing a program for carrying out the motion picture encoding procedure and the motion picture decoding procedure of each of the aforementioned embodiments using a computer system.
FIG. 26 is a block diagram showing an overall configuration of a content delivery system.
FIG. 27 is an external view of a mobile phone.
FIG. 28 is a block diagram showing the structure of the mobile phone.
FIG. 29 is a diagram showing an example of a digital broadcasting system.
FIG. 30 is a schematic diagram explaining reference relationships between images in the prior art.
FIG. 31A and B are schematic diagrams explaining image rearrangement in the prior art.
FIG. 32 is a schematic diagram explaining the operations of motion compensation in the prior art.
FIG. 33 is a schematic diagram explaining the operations of linear prediction processing in the prior art.
FIG. 34 is a schematic diagram explaining a procedure for assigning reference indices to image numbers in the prior art.
FIG. 35 is a schematic diagram showing an example of relationship between reference indices and image numbers in the prior art.
FIG. 36 is a schematic diagram explaining the structure of a coded stream in the prior art.
FIG. 37 is a schematic diagram showing an example of sets of weighting coefficients of linear prediction coefficients in the prior art.
FIG. 38 is a schematic diagram explaining the structure of a coded stream in the prior art.
ES 2 401 774 T3
FIG. 39 is a schematic diagram explaining the relationship between image numbers and display order information.
FIG. 40A and 40B are schematic diagrams explaining the structure of a coded stream and an example of flags.
FIG. 41A and 41B are schematic diagrams explaining the structure of a coded stream and an example of flags.
Best Mode of Carrying Out the Invention (First Embodiment)
FIG. 1 is a block diagram showing the structure of a moving picture coding apparatus in the first embodiment of the present invention. The moving picture coding procedure executed in this moving picture coding apparatus, specifically, (1) an overview of the coding, (2) a procedure for assigning reference indices and (3) a procedure for generating a predictive image, will be explained in this order using the block diagram shown in FIG. 1.
(1) Encoding overview
A moving picture to be encoded is input into a picture memory 101, picture by picture in the order of display, and the input pictures are reordered in the order of coding. FIG. 31A and 31B are diagrams showing an example of image rearrangement. FIG. 31A shows an example of images in the order of display and FIG. 31B shows an example of the images rearranged in the encoding order. In this case, since images B3 and B6 refer to images before and after in time, the reference images need to be encoded before encoding these current images, and therefore the images are rearranged in FIG. 31B so that pictures P4 and P7 have been coded previously. Each of the images is divided into blocks, each of which is referred to as a macroblock of 16 horizontal pixels by 16 vertical pixels, for example, the following process being carried out in each block.
An input image signal read from image memory 101 is input to a differential computing unit 112. Differential computing unit 112 computes the difference between the input image signal and the predictive image signal provided by a unit of motion compensation coding 107 and provides the obtained differential image signal (prediction error signal) to a prediction error coding unit 102. The prediction error coding unit 102 performs image coding processing, such as frequency transformation and quantization, and provides a coded error signal. The encoded error signal is input to a prediction error decoding unit 104, which performs image decoding processing, such as inverse quantization and inverse frequency transformation, and provides a decoded error signal. A summing unit 113 sums the decoded error signal and the predictive image signal to generate a reconstructed image signal and stores, in an image memory 105, the reconstructed image signals that can be referred to in inter-prediction. posterior image from the reconstructed image signals obtained.
On the one hand, the input image signal read per macroblock from the image memory 101 is also input into a motion vector estimation unit 106. Here, the reconstructed image signals stored in the image memory 105 are they seek to estimate an image area that is most similar to the input image signal and thus determine a motion vector pointing to the position of the image area. The motion vector estimation is carried out for each block, which is a subdivision of a macroblock, and the obtained motion vectors are stored in a motion vector storage unit 108.
At this time, since a plurality of images can be used as reference in the H.26L standard, which is currently being considered for standardization, identification numbers are needed in each block to designate reference images. The identification numbers are called reference indices, and a reference index / image number conversion unit 111 maps the reference indices and the image numbers of the images stored in the image memory 105 to allow the designation of reference images. The operation of the image number / reference index conversion unit 111 will be explained in detail in section (2).
The motion compensation coding unit 107 extracts the most suitable image area for the predictive image from among the reconstructed image signals stored in the image memory 105 using the motion vectors estimated by the aforementioned processing and the reference indices. . The pixel value conversion processing, such as the interpolation processing by linear prediction, is carried out on the pixel values of the image area obtained to obtain the
ES 2 401 774 T3 final predictive image. The linear prediction coefficients used for that purpose are generated by the linear prediction coefficient generating unit 110 and stored in the linear prediction coefficient storage unit 109. This predictive image generation procedure will be explained in detail in the section (3).
The encoded stream generation unit 103 performs variable-length encoding for the encoded information, such as linear prediction coefficients, reference indices, motion vectors, and encoded error signals provided as a result of the above. serial processing to obtain an encoded stream to be provided from this encoding apparatus.
The flow of operations in the case of inter-picture prediction coding has been described above, and a switch 114 and a switch 115 switch between an inter-picture prediction coding and an intra-picture prediction coding. In the case of intra-image prediction coding, a predictive image is not generated by motion compensation, but a differential image signal is generated by calculating the difference between a current area and a predictive image of the current area that is generated at starting from a coded area in the same image. The prediction error encoding unit 102 converts the differential image signal into the encoded error signal in the same way as the inter-image prediction encoding, and the encoded stream generation unit 103 performs length encoding variable for the signal to obtain an encoded stream to be provided.
(2) Procedure for assigning benchmarks
Next, it will be explained, using FIG. 3 and FIG. 4, the method by which the reference index / image number conversion unit 111 shown in FIG. 1 assigns benchmarks.
FIG. 3 is a diagram explaining the procedure for assigning two reference indices to picture numbers. Assuming there is a sequence of images arranged in the order of display, as shown in this diagram, image numbers are assigned to the images in the order of encoding. The commands for assigning reference indices to image numbers are described in a header of each section, which is a subdivision of an image, as the encoding unit, and, therefore, their assignment is updated every time. encode a section. The command indicates, serially by the number of benchmarks, the differential value between an image number that is assigned a current benchmark and an image number that is assigned a benchmark immediately before the current allocation.
Taking the first benchmarks of FIG. 3 as an example, since “-1” is provided as a command first, an image with image number 15 is assigned to reference index number 0 by subtracting 1 from current image number 16. Then, -4 is provided as a command, so an image with image number 11 is assigned to reference index number 1 by subtracting 4 from image number 15 that has been assigned together before it. Each subsequent image number is assigned in the same way. The same applies to the second benchmarks.
According to the conventional benchmark assignment procedure shown in FIG. 34, all reference indices are mapped to respective image numbers. On the other hand, in the example of FIG. 3, although exactly the same allocation procedure is used, a plurality of reference indices are matched to the same image number by modifying the values of the commands.
FIG. 4 shows the result of the assignment of the benchmarks. This diagram shows that the first benchmark and the second benchmark are assigned to each image number separately, but a plurality of benchmarks are assigned to one image number in some cases. In the encoding method of the present invention, it is assumed that a plurality of reference indices are assigned to at least one picture number, like this example.
If the reference indices are used only to determine reference images, the conventional method of one-to-one assignment of reference indices to image numbers is the most efficient coding procedure. However, in case a set of linear prediction coefficient weights is selected to generate a predictive image using reference indices, the same linear prediction coefficients have to be used for all blocks having the same reference images , so there is an extremely high possibility that the optimal predictive image cannot be generated.
Therefore, if it is possible to assign a plurality of reference indices to an image number as in the case of the present invention, the set of optimal weighting coefficients of linear prediction coefficients can be selected for each block from a plurality of Candidate sets even if all the blocks have the same reference image and thus the predictive image can be generated with higher coding efficiency.
ES 2 401 774 T3
It should be noted that the above description shows the case where the image numbers are provided assuming that all reference images are stored in a reference memory. However, a current image is given an image number that is one greater than the number of an image that has been encoded immediately before the current image, only when the current image that has been encoded in is stored. last place, whereby the continuity of the image numbers is kept in the reference memory even if some images are not stored and therefore the above mentioned procedure can be used without changes.
(3) Procedure to generate predictive images
Next, it will be explained, using FIG. 5, the predictive imaging method in the motion compensation encoding unit 107 shown in FIG. 1. Although the method of generating images by linear prediction is exactly the same as the conventional method, the flexibility in the selection of linear prediction coefficients is increased because a plurality of reference index numbers can be matched to the same image.
Image B16 is a current B image to be encoded, and blocks BL01 and BL02 are current blocks to be encoded belonging to image B. Image P11 and image B15 are used as the first reference image and as the second reference image for BL01, and the predictive image is generated with reference to blocks BL11 and BL21 belonging to images P11 and B15, respectively. In the same way, the image P11 and the image B15 are used as the first reference image and as the second reference image for BL02, the predictive image is generated with reference to the blocks BL12 and BL22 that belong to those reference images. , respectively.
Although BL01 and BL02 refer to the same images as their first reference image and second reference image, it is possible to assign different values to the first reference index ref1 and the second reference index ref2 for BL01 and BL02 using the assignment procedure of benchmarks explained in section (2). Taking FIG. 4 As an example, 1 and 3 are assigned to the first benchmark corresponding to image number 11, while 1 and 6 are assigned to the second benchmark corresponding to image number 15.
As a result, it is assumed that there are four combinations of these benchmarks, (ref1, ref2) = (1, 1), (1, 6), (3, 1) and (3, 6) and thus it is It is possible to select the combination to obtain the optimal set of weighting coefficients for each block among these combinations. In FIG. 5, ref1 = 1 and ref2 = 1 are assigned to BL01, and ref1 = 3 and ref2 = 6 are assigned to BL02, for example.
According to the conventional homing procedure shown in FIG. 35, only a combination of (ref1, ref2) = (1, 1) can be selected for BL01 and BL02 in the case of FIG. 5, and thus only one set of weights of linear prediction coefficients can be selected. On the other hand, according to the present invention, there are four options available and it can be said that the possibility of selecting the optimal set of weighting coefficients increases.
An encoded stream of an image is made up of a common image information area and a plurality of section data areas. FIG. 6 shows the structure of the section data area in the encoded stream. The section data area is further formed by a section header area and a plurality of block data areas. This diagram shows each of the block areas corresponding to BL01 and BL02 in FIG. 5 as an example of the block data area. ref1 and ref2 included in BL01 designate the first reference index and the second reference index, respectively, indicating two images to which the block BL01 refers.
Furthermore, in the section header area, data (pset0, pset1, pset2 ...) to provide the sets of weighting coefficients to carry out the above-mentioned linear prediction are described for ref1 and ref2, respectively. In this area, "psets" can be set to a number equivalent to the number of benchmarks explained in section (2). More specifically, in case ten benchmarks, ranging from 0 to 9, are used as each of the first benchmark and second benchmark, ten sets from 0 to 9 can also be set for ref1 and ref2. .
FIG. 7 shows an example of tables of the sets of weights included in the section header area. Each piece of data indicated by an identifier set has four values w1, w2, c and d, and these tables are structured so that the values of ref1 and ref2 can refer to the data directly. Also, the idx_cmd1 and idx_cmd2 scripts for assigning reference indices to image numbers are described in the section header area.
Using ref1 and ref2 described in BL01 in FIG. 6, a set of weighting coefficients is selected from each of the tables for ref1 and ref2 of FIG. 7. By performing a linear prediction on the pixel values of the reference images using these two sets of weighting coefficients, an image is generated
ES 2 401 774 T3 predictive.
As described above, using the encoding method in which a plurality of reference indices are assigned to an image number, a plurality of candidates can be generated for the weight coefficient sets of linear prediction coefficients and, therefore, the best of them can be selected. For example, in case two first benchmarks and two second benchmarks are assigned, there are four sets of weights available as candidates for selection, and in case three first benchmarks and three second benchmarks are assigned For reference, there are nine sets of weights available as candidates for selection.
In particular, this linear prediction procedure has an important effect in case the brightness of the whole image or part of it changes significantly, such as fading or flickering. In many cases, the degree of change in brightness is different between parts of an image. Therefore, the structure of the present invention in which the best set for each block can be selected from a plurality of sets of weighting coefficients is very efficient in image coding.
Next, the processing flow from determining sets of weighting coefficients to generating a predictive image will be explained in detail.
FIG. 8 is a functional block diagram showing the functional structure for generating a predictive image in the linear prediction coefficient generation unit 110, the linear prediction coefficient storage unit 109, and the motion compensation coding unit 107.
A predictive image is generated through the linear prediction coefficient generating unit 110, the linear prediction coefficient storage unit 109a, the linear prediction coefficient storage unit 109b, the averaging unit 107a, and the linear prediction operations unit 107b.
The sets of weighting coefficients generated by the linear prediction coefficient generating unit 110 are stored in the linear prediction coefficient storage unit 109a and in the linear coefficient storage unit 109b. The averaging unit 107a obtains, from the linear prediction coefficient storage unit 109a, a set of weighting coefficients (w1_1, w2_1, c_1, d_1) selected by the first reference index ref1 determined by the processing estimation of motion and, in addition, obtains, from the linear coefficient storage unit 109b, a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected by the second reference index ref2.
Then, the averaging unit 107a calculates the average, for respective parameters, of the sets of weighting coefficients from the storage units of linear prediction coefficients 109a and 109b to consider it as the set of weighting coefficients (w1 , w2, c, d) to be used for the real linear prediction, and provides it to the linear prediction operations unit 107b. The linear prediction operation unit 107b calculates the predictive image using equation 1 as a function of the obtained set of weighting coefficients (w1, w2, c, d) to be provided.
FIG. 9 is a functional block diagram showing another functional structure for generating a predictive image. A predictive image is generated through the linear prediction coefficient generating unit 110, the linear prediction coefficient storage unit 109a, the linear prediction coefficient storage unit 109b, the linear prediction operating unit 107c, the linear prediction operating unit 107d and the averaging unit 107e.
The sets of weighting coefficients generated by the linear prediction coefficient generating unit 110 are stored in the linear prediction coefficient storage unit 109a and in the linear prediction coefficient storage unit 109b. The linear prediction operations unit 107c obtains, from the linear prediction coefficient storage unit 109a, a set of weighting coefficients (w1_1, w2_1, c_1, d_1) selected by the first reference index ref1 determined by the motion estimation processing, and calculates the predictive image using equation 1 based on the set of weighting coefficients to be provided to the averaging unit 107e.
In the same way, the linear prediction operations unit 107d obtains, from the linear prediction coefficients storage unit 109b, a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected by the second index of reference determined by motion estimate processing, and calculates the predictive image using equation 1 based on the set of weighting coefficients to be provided to the averaging unit 107e.
The averaging unit 107e calculates the average of respective pixel values of images
ES 2 401 774 T3 provided by the linear prediction operations unit 107c and by the linear prediction operations unit 107d respectively to generate the final predictive image to be provided.
FIG. 10A is a functional block diagram showing another functional structure for generating a predictive image. A predictive image is generated through the linear prediction coefficient generating unit 110, the linear prediction coefficient storage unit 109c, the linear prediction storage unit 109d, the averaging unit 107f, and the measurement unit. linear prediction operations 107g.
The sets of weighting coefficients generated by the linear prediction coefficient generating unit 110 are stored in the linear prediction coefficient storage unit 109c and in the linear prediction coefficient storage unit 109d. The averaging unit 107f obtains, from the linear prediction coefficient storage unit 109c, the parameters of c_1 and d_1 in a set of weighting coefficients (w1_1, w2_1, c_1, d_1) selected by the first index reference ref1 determined by the motion estimation processing and also obtains, from the linear prediction coefficient storage unit 109d, the parameters of c_2 and d_2 in a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected by the second reference index ref2. The averaging unit 107f calculates the average of c_1 and c_2 and the average of d_1 and d_2 obtained from the linear prediction coefficient storage unit 109c and the linear prediction coefficient storage unit 109d to obtain c and d to be provided to the linear prediction operation unit 107g.
Furthermore, the linear prediction operations unit 107g obtains the parameter of w1_1 from the set of weighting coefficients mentioned above (w1_1, w2_1, c_1, d_1) from the linear prediction coefficient storage unit 109c, it obtains the parameter of w2_2 of the set of weighting coefficients mentioned above (w1_2, w2_2, c_2, d_2) from the storage unit of linear prediction coefficients 109d, and obtains c and d which are the averages calculated by the averaging unit 107f, and then calculates the predictive image using equation 1 to be provided.
More specifically, when determining the set of weighting coefficients (w1, w2, c, d), which is actually used for linear prediction, from among the set of weighting coefficients (w1_1, w2_1, c_1, d_1) obtained from the linear prediction coefficient storage unit and the set of weighting coefficients (w1_2, w2_2, c_2, d_2) obtained from the linear prediction coefficient storage unit 109d, the linear prediction operating unit 107g uses the following rule:
w1 = w1_1 w2 = w2_2 c = (average of c_1 and c_2) d = (average of d_1 and d_2)
As previously described, in the generation of the predictive image as explained in FIG. 10A, the linear prediction coefficient storage unit 109c does not need w2_1 of the set of weighting coefficients. Therefore, w2 is not required for the set of weighting coefficients for ref1, and thus the amount of data in an encoded stream can be reduced.
Furthermore, the linear prediction coefficient storage unit 109d does not need w1_2 of the set of weighting coefficients. Therefore, w1 is not required for the set of weighting coefficients for ref2, and thus the amount of data in an encoded stream can be reduced.
FIG. 10B is a functional block diagram showing another functional structure for generating a predictive image. A predictive image is generated through the linear prediction coefficient generating unit 110, the linear prediction coefficient storage unit 109e, the linear prediction coefficient storage unit 109f, and the linear prediction operations unit 107h.
The sets of weighting coefficients generated by the linear prediction coefficient generating unit 110 are stored in the linear prediction coefficient storage unit 109e and in the linear prediction coefficient storage unit 109f. The linear prediction operations unit 107h obtains, from the linear prediction coefficients storage unit 109e, the parameters of w1_1, c_1 and d_1 that are part of a set of weighting coefficients (w1_1, w2_1, c_1, d_1 ) selected by the first reference index ref1 determined by the motion estimation processing and also obtains, from the linear prediction coefficient storage unit 109f, the parameter of w2_2 which is part of a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected based on the second reference index ref2. Linear prediction operations unit 107h calculates a predictive image using equation 1 based on w1_1, c_1, d_1, w2_2 obtained from linear prediction coefficient storage unit 109e and prediction coefficient storage unit 109f to be provided.
ES 2 401 774 T3
More specifically, when determining the set of weighting coefficients (w1, w2, c, d), which is actually used for linear prediction, from among the set of weighting coefficients (w1_1, w2_1, c_1, d_1) obtained from the linear prediction coefficient storage unit 109e and the set of weighting coefficients (w1_2, w2_2, c_2, d_2) obtained from the linear prediction coefficient storage unit 109f, the linear prediction operation unit 107h uses the following rule.
w1 = w1_1 w2 = w2_2 c = c_1 d = d_1
In the generation of the predictive image as explained in FIG. 10B, the linear prediction coefficient storage unit 109e does not need w2_1 of the set of weighting coefficients. Therefore, w2 is not required for the set of weighting coefficients of ref1, and thus the amount of data in an encoded stream can be reduced.
Furthermore, the linear prediction coefficient storage unit 109f does not need w1_2, c_2 and d_2 from the set of weighting coefficients. Therefore w1, c and d are not required for the set of weighting coefficients for ref2 and thus the amount of data in an encoded stream can be reduced.
Furthermore, it is also possible to use one or more parameters among w1, w2, c and d as fixed values. FIG. 11 is a functional block diagram in the case where only d is used as a fixed value for the functional structure of FIG. 10A. A predictive image is generated through the linear prediction coefficient generating unit 110, the linear prediction coefficient storage unit 109i, the linear prediction coefficient storage unit 109j, the averaging unit 107j and the linear prediction operations unit 107k.
The coefficients selected by the first reference index ref1 from the linear prediction coefficient storage unit 109i are only (w1_1, c_1), and the coefficients selected by the second reference index ref2 from the storage unit of linear prediction coefficients 109j are only (w2_2, c_2). The averaging unit 107j calculates the average of c_1 and c_2 obtained from the prediction coefficient storage unit 109i and the linear prediction coefficient storage unit 109j to obtain c and provides it to the operations unit of linear prediction 107k.
The linear prediction operations unit 107k obtains the parameter of w1_1 from the storage unit of linear prediction coefficients 109i, the parameter of w2_2 from the storage unit of linear prediction coefficients 109j, and the parameter of ca starting from averaging unit 107j and then calculating the predictive image based on equation 1 using a predetermined fixed value as a parameter of d, and provides the predictive image. More specifically, the following values are entered as the coefficients (w1, w2, c, d) of equation 1.
w1 = w1_1 w2 = w2_2 c = (average of c_1 and c_2) d = (fixed value)
Assigning the above values in equation 1 provides the following equation 1a.
P (i) = (wl_l x Ql (i) + w2_2 x Q2 (i)) / pow (2, d) + (c_l + c_2) / 2 ...... Equation 1a (where pow (2, d ) indicates the “d” -th power of 2)
Furthermore, modifying equation 1a gives the following equation 1b. It is possible that the linear prediction operation unit 107k treats the linear prediction operation procedure exactly the same in the format of equation 1b or equation 1.
P (i) = (wl_l xQl (i) / pow (2, dl) + c_l + w2_2 x Q2 (i) / pow (2, d-1) 4-c_2) / 2 ...... Equation 1b (where pow (2, d-1) indicates the “d-1” -th power of 2)
ES 2 401 774 T3
Although pow (2, d-1) is used in equation 1b, the system can be structured using pow (2, d ') by entering d' (assuming that d '= d-1) in the linear prediction operations unit 107k , since d is a fixed value.
Furthermore, in the generation of the predictive image as explained in FIG. 11, the linear prediction coefficient storage unit 109i only needs w1_1 and c_1 from among the sets of weighting coefficients for this, and the linear prediction coefficient storage unit 109j only needs w2_2 and c_2 from among the sets of weighting coefficients for it. Therefore, it is not necessary to encode parameters other than the parameters required above, and therefore the amount of data in the encoded stream can be reduced.
It should be noted that it is possible to use a predetermined fixed value as a value of d in either case, but the fixed value can toggle for each section with the fixed value described in the section header. Also, the fixed value can be switched for each image or for each sequence being described in the image common information area or in the sequence common information area.
FIG. 12 shows an example of a structure of a section data area in case the above-mentioned linear prediction procedure is used. This is different from FIG. 6 where only d is described in the section header area and only w1_1 and c_1 are described as pset for ref1. FIG. 13 shows tables showing an example of the previous sets of weights included in the section header area. Each piece of data indicated by the identifier "pset" has two values of w1_1 and c_1 or w2_2 and c_2, and is structured so that the values of ref1 and ref2 refer to them directly.
It should be noted that the linear prediction coefficient generating unit 110 generates the sets of weighting coefficients by examining the characteristics of an image, and the motion compensation coding unit 107 creates a predictive image using any of the methods explained in FIG. . 8, FIG. 9, FIG. 10 and FIG. 11, and determines the combination of two reference indices ref1 and ref2 to minimize the prediction error. In case any of the procedures of FIG. 10A, FIG. 10B and FIG. 11 not requiring all the parameters, it is possible to skip the processing of creating unnecessary parameters in the phase where the linear prediction coefficient generating unit 110 of the coding apparatus creates the sets of weighting coefficients.
In the procedures of FIG. 10A, FIG. 10B and FIG. 11, the linear prediction coefficient generating unit 110 may search for and create the optimal sets of weighting coefficients for ref1 and ref2, w1_1 and w2_2, for example, separately, when such sets of weighting coefficients are created. In other words, using these procedures, it is possible to reduce the amount of processing performed by the encoding apparatus to create sets of weighting coefficients.
It should be noted that the coding procedures mentioned above refer to a B-picture having two reference pictures, but it is also possible to perform the same processing in the coding mode of a reference picture for a P-picture or a B-picture having just a reference image. In this case, using only one of the first reference index and the second reference index, pset and idx_cmd for ref1 or ref2 are described only in the section header area included in the encoded stream of FIG. 6, according to the reference index described in the block data area.
Also, as a linear prediction procedure, the following equation 3 is used instead of the equation 1 explained in the conventional procedure. In this case, Q1 (i) is a pixel value of a referenced block, P (i) is a pixel value of a predictive image of a current block to be encoded, and w1, w2, c and d are linear prediction coefficients provided by the selected set of weighting coefficients.
P (¡) = (wl x Ql (i) + w2 x Ql (í)) / pow (2, d) + c ...... Equation 3 (where pow (2, d) denotes the d-th power of 2)
It should be noted that it is possible to use equation 4 as a linear prediction equation, instead of equation 3. In this case, Q1 (i) is a pixel value of a referenced block, P (i) is a pixel value of a predictive image of a current block to be encoded, and w1, c and d are linear prediction coefficients provided by the selected set of weighting coefficients.
<img file="ES2401774T3_D0001.tif" />
(where pow (2, d) indicates the “d” -th power of 2)
Using equation 1 and equation 3 requires four parameters w1, w2, c and d, while using equation 4 requires only three parameters w1, c and d for linear prediction. In other words, in case any one of the first benchmark and the second benchmark is used
ES 2 401 774 T3 for an image as a whole, such as a P-image, it is possible to reduce to three the number of data items of each set of weighting coefficients to be described in the section header area.
When using equation 3, it is possible to make a linear prediction available for B-images and P-images adaptively without changes in structure. When equation 4 is used, the amount of data to be described in the header area of an image P can be reduced, and therefore, the reduction of the amount of processing can be achieved by a simplified calculation. However, since the reference indexing method suggested by the present invention can be applied directly to any of the above methods, a predictive image with high coding efficiency can be created, which is extremely efficient in image coding.
On the other hand, the images referred to in a motion compensation are determined by designating the reference indices assigned to the respective images. In that case, the maximum number of images that are available for reference has been described in the image common information area of the encoded stream.
FIG. 38 is a schematic diagram of a coded stream describing the maximum number of images that are available for reference. As this diagram shows, the maximum number of images for Max_pic1 of ref1 and the maximum number of images for Max_pic2 of ref2 are described in the image common information of the encoded stream.
The information required for encoding is not the maximum number of actual images, but the maximum reference index value available to designate images.
Since in the conventional method a reference index is assigned to an image, the above-mentioned description of the maximum number of images does not generate any contradiction. However, the different number of reference indices and images has a significant influence in case a plurality of reference indices are assigned to a picture number, such as the present invention.
As described above, the idx_cmd1 and idx_cmd2 scripts are described in a coded stream for the purpose of assigning reference indices to image numbers. Image numbers and reference indices are mapped to each other based on each command in these idx_cmd1 and idx_cmd2 scripts. To that end, knowing the maximum benchmark value shows that all benchmarks and image numbers have been mapped to each other, specifically the end of the idx_cmd1 and idx_cmd2 script commands.
Therefore, in the present embodiment, the maximum number of available reference indices, instead of the maximum number of images in the prior art, is described in the common image information area, which is the header of the image. Alternatively, both the maximum number of images and the maximum number of reference indices are described.
FIG. 23 shows a common image information area in a coded stream of an image in which the maximum number of reference indices is described. The maximum number of reference indices available for Max_idx1 of ref1 and the maximum number of reference indices available for Max_idx2 of ref2 are described in the image common information area.
In FIG. 23, the maximum number of reference indices is described in the common image information, but may be structured so that the maximum number of reference indices is described in the section data area as well as in the common image information. For example, the maximum number of benchmarks required for each section can be clearly described, in case the maximum number of benchmarks required for each section is significantly different from the maximum number of benchmarks described in the common information area of image, from section to section; For example, the maximum number of benchmarks in an image is 8, the maximum number of benchmarks required for section 1 of the image is 8, and the maximum number of benchmarks required for section 2 is 4.
In other words, it can be structured so that the maximum number of benchmarks described in the common image information is set as a default value common to all sections of the image and that the maximum number of benchmarks required for a section, which is different from the default, is described in the section header.
Although FIG. 23 and FIG. 38 show examples in which a coded stream is made up of a common image information area and section data areas, the common image information area and the section data areas can be treated as coded streams different from exactly the one. same way as an encoded stream.
(Second realization)
ES 2 401 774 T3
Next, the moving picture coding method of the second embodiment of the present invention will be explained. Since the structure of the coding apparatus, the coding processing flow and the reference index assignment procedure are exactly identical to those of the first embodiment, the explanation thereof will not be repeated.
In the first embodiment, linear prediction is carried out on each pixel to generate a predictive image in a motion compensation using equation 1, equation 3 or equation 4. However, all of these equations include multiplications, which causes a significant increase in the amount of processing considering that these multiplications are carried out on all pixels.
Therefore, it is possible to use Equation 5 in place of Equation 1, Equation 6 in place of Equation 3, and Equation 7 in place of Equation 4. These equations allow calculations using only bit-shifting operations without use multiplication and thus reduce the amount of processing. In the following equations, Q1 (i) and Q2 (i) are pixel values of referenced blocks, P (i) is a pixel value of a predictive image of a current block to be encoded, and 'm', 'n' and 'c' are linear prediction coefficients provided by a selected set of weighting coefficients.
P (i) = ± pow (2, m) x Ql (i) ± pow (2, n) xQ2 (i) + c ......
Equation 5
P (i) = ± pow (2, m) x Ql (i) ± pow (2, n) x.Ql (i) + c ......
Equation 6
P (i) = ± pow (2, m) xQl (i) + c ...... Equation 7 (where pow (2, m) indicates the “m” -th power of 2, and pow (2, n) indicates the n-th power of 2)
As in the first embodiment, equation 5 is used to generate a predictive image with reference to two images at the same time, and equation 6 or equation 7 is used to generate a predictive image with reference to only one image. Since these equations require identifiers that indicate plus or minus signs, sets of weights required for prediction operations are (sign1, m, sign2, n, c) for equations 5 and 6, and (sign1, m, c ) for equation 7. sign1 and sign2 are parameters that identify the first and second plus and minus signs, respectively. Although the number of parameters is larger than in the first embodiment, the amount of processing increases little since sign1 and sign2 can be represented by 1 bit.
Next, the processing flow from determining sets of weighting coefficients to generating a predictive image with reference to two images at a time will be explained in detail using equation 5.
First, the case where the functional structure for generating a predictive image is as shown in FIG. 8. The averaging unit 107a obtains the set of weighting coefficients (sign1_1, m_1, sign2_1, n_1, c_1) from the linear prediction coefficient storage unit 109a. Furthermore, the averaging unit 107a obtains the set of weighting coefficients (sign1_2, m_2, sign2_2, n_2, c_2) from the linear prediction coefficient storage unit 109b.
The averaging unit 107a calculates, for respective parameters, the average of the sets of weighting coefficients obtained from the linear prediction coefficient storage unit 109a and the linear prediction coefficient storage unit 109b to consider the average as the set of weights (sign1, m, sign2, n, c). The linear prediction operations unit 107b calculates the predictive image using equation 5 based on the set of weighting coefficients (sign1, m, sign2, n, c) provided by the averaging unit 107a.
It should be noted that FIG. 8 shows the set of weighting coefficients (w1_1, w2_1, c_1, d_1), and the like obtained from the storage unit of linear prediction coefficients 109a and similar, which are calculated in case of using equation 1 explained in the first embodiment, and does not show the parameters used in case the predictive image is obtained using equation 5, since the parameters used in the first case can be replaced by the parameters of the second case. This also applies to the cases of FIG. 9 and FIG. 10 described below.
Now the case where the functional structure for generating a predictive image is as shown in FIG. 9. The linear prediction operations unit 107c computes a predictive image 1
ES 2 401 774 T3 based on the set of weighting coefficients (sign1_1, m_1, sign2_1, n_1, c_1) obtained from the linear prediction coefficient storage unit 109a. The linear prediction operations unit 107d calculates a predictive image 2 based on the set of weighting coefficients (sign1_2, m_2, sign2_2, n_2, c_2) obtained from the linear prediction coefficient storage unit 109b. Furthermore, the averaging unit 107e calculates, for respective pixels, the average of the predictive images computed by the linear prediction operating units 107c and 107d to obtain a predictive image.
In this case, the linear prediction operations unit 107c first calculates the predictive image using equation 5 based on the set of weighting coefficients (sign1_1, m_1, sign2_1, n_1, c_1), whereby it is possible to calculate the Predictive image using bit shift operations without using multiplication. This also applies to the linear prediction operating unit 107d. On the other hand, in the case of FIG. 8, since the average of the set of weighting coefficients (sign1_1, m_1, sign2_1, n_1, c_1) and of the set of weighting coefficients (sign1_2, m_2, sign2_2, n_2, c_2), the average of m_1 and m_2 or the average of n_1 and n_2, specifically the exponents of 2, may not be whole numbers and therefore there is a possibility that the amount of processing may increase. Also, if the exponents of 2 are rounded to whole numbers, there is a chance that the errors will increase.
Next, the case where a predictive image is generated in the functional structure shown in FIG. 10A. The linear prediction operations unit 107g calculates a predictive image using equation 5, based on the parameters sign1_1 and m_1 that are obtained from the linear prediction coefficient storage unit 109c and used for bit shift operations, the sign2_2 and n_2 parameters obtained from linear prediction coefficient storage unit 109d and used for bit shift operations, and the average c calculated by the averaging unit 107f from the parameters c_1 and c_2 which are obtained from the linear prediction coefficient storage units 109c and 109d.
In this case, since the coefficients used for the bit-shifting operations are the values that are obtained directly from the linear prediction coefficient storage unit 109c or the linear prediction coefficient storage unit 109d, the exponents of 2 in equation 5 are whole numbers. Therefore, calculations can be carried out using bit shift operations, and thus the amount of processing can be reduced.
Next, the case where a predictive image is generated in the functional structure shown in FIG. 10B. The linear prediction operations unit 107h calculates a predictive image using equation 5 based on the parameters sign1_1, m_1 and c_1, which are obtained from the storage unit of linear prediction coefficients 109e, and on the parameters sign2_2 and n_2 , which are obtained from the linear prediction coefficient storage unit 109f.
In this case, since the coefficients used for the bit-shifting operations are values that are obtained directly from the linear prediction coefficient storage unit 109e or the linear prediction coefficient storage unit 109f, the exponents of 2 in equation 5 are whole numbers. Therefore, the predictive image can be calculated using bit shift operations, and therefore the amount of processing can be reduced.
In the cases of FIGS. 10A and 10B, there are parameters that do not need to be added to the encoded stream for transmission, as is the case in FIGS. 10A and 10B in the first embodiment, and the amount of data in the encoded stream can be reduced.
Using the linear prediction equations explained in the second embodiment, the calculations can be performed using bit shift operations without using multiplication, whereby the amount of processing can be significantly reduced with respect to the first embodiment.
In the present embodiment, linear prediction is carried out using equations 5, 6 and 7 instead of equations 1, 3 and 4, and using the set of parameters to be coded (sign1, m, sign2, n, c) instead of (w1, w2, c, d), so that calculations can be performed using only bit shift operations and thus a reduction in the amount of processing is achieved. However, it is also possible, as another approach, to use equations 1, 3, and 4 and (w1, w2, c, d) as is, limiting the selectable values of w1 and w2 to only the values available for shift operations of bits, so that calculations can be performed using only bit shift operations, thereby achieving a reduction in the amount of processing.
When the values of w1 and w2 are determined, the linear prediction coefficient generating unit 110 of FIG. 1 selects only, as one of the options, the values available for bit-shifting operations and describes the selected values directly in the encoded streams of FIG. 6 and FIG. 12, as the values of w1 and w2 in them. As a result, it is possible to reduce the amount of
ES 2 401 774 T3 processing for linear prediction even in exactly the same structure as that of the first embodiment. It is also possible to determine the coefficients easily since the options for the coefficients are limited.
Furthermore, as a procedure for such a limitation, it is possible to limit the values of w1 and w2 so that 1 is always selected for such values and generate the optimal values of only c1 and c2, which are DC components, in the generation unit. of prediction coefficients 110. Taking the structure of FIG. 11 as an example, (1, c_1) for ref1 and (1, c_2) for ref2 are encoded as parameter sets. In this case, the pixel value P (i) of the predictive image is calculated by the following equation where w1_1 and w2_2 from equation 1a are replaced by 1.
P (i) = (Ql (i) + Q2 (i)) / pow (2, d) + (c_l + c_2) / 2 (where pow (2, d) indicates the “d” -th power of 2)
Accordingly, it is possible to significantly reduce the amount of processing for linear prediction even in exactly the same structure as that of the first embodiment. It is also possible to significantly simplify the procedure for determining coefficients since the necessary coefficients are only c_1 and c_2.
FIG. 24 shows an example in which a flag sft_flg, which indicates whether or not it is possible to carry out a linear prediction using only bit shift operations, and a flag cc_flg, which indicates whether or not it is possible to carry out a linear prediction using only c, which is a component of cc, they are described in common image information in a coded stream of one image. A decoding apparatus can decode the picture without referring to these flags. However, referring to these flags, it is possible to perform a decoding in the frame suitable for a linear prediction using only bit shift operations, or a decoding in the frame suitable for a linear prediction using only a DC component, so these flags can be very important information depending on the structure of the decoding apparatus.
Although FIG. 24 shows an example in which a coded stream is made up of an image common information area and section data areas, the image common information area and the section data areas can be treated as exactly different encoded streams of the same way as an encoded stream. Furthermore, in the example of FIG. 24, sft_flg and cc_flg are described in the image common information area, but they can be treated in exactly the same way even if they are described in the sequence common information area or another independent common information area. Furthermore, it is not only possible to use the two flags sft_flg and cc_flg at the same time, but one of them can also be used, and they can be treated in the same way in the second case.
(Third embodiment)
Next, the moving picture coding method of the third embodiment of the present invention will be explained. Since the structure of the coding apparatus, the coding processing flow and the reference index assignment procedure are exactly identical to those of the first embodiment, the explanation thereof will not be repeated.
As explained in the prior art section, there is a procedure to generate a predictive image using a predetermined fixed equation, such as equation 2a and equation 2b, unlike the first and second embodiments in which an image Predictive is generated using a prediction equation obtained from sets of weights of linear prediction coefficients. This conventional method has the advantage that the amount of data for encoding can be reduced since it is not necessary to encode or transmit the sets of weighting coefficients used to generate the predictive image. Also, the amount of processing for linear prediction can be significantly reduced since the equations for linear prediction are simple. However, the procedure using the fixed equations has the problem that the prediction precision is lower because only two linear prediction equations, 2a and 2b, can be selected.
Therefore, in the present embodiment, equations 8a and 8b are used instead of equations 2a and 2b. These equations 8a and 8b are obtained by adding C1 and C2 to equations 2a and 2b. Since only the number of sums in the operation increases, the amount of processing increases very little compared to the original equations 2a and 2b. In the following equations, Q1 (i) and Q2 (i) are pixel values of referenced blocks, P (i) is a pixel value of a predictive image of a current block to be encoded, and C1 and C2 are linear prediction coefficients provided by a selected set of weighting coefficients.
<img file="ES2401774T3_D0002.tif" />
ES 2 401 774 T3
P (¡) = (Ql (i) + Cl + Q2 (i) + C2) / 2 ...... Equation 8b
Equations 8a and 8b are prediction equations for generating a predictive image with reference to two images at the same time, but when generating a predictive image with reference to only one image, equation 9 is used instead of equation 3 or equation Equation 4 explained in the previous embodiments.
P (i) = Ql (i) + Cl ...... Equation 9
The sets of weights to use this procedure are only (C1) for ref1 and (C2) for ref2. Therefore, an example of a coded stream of an image obtained using this procedure is as shown in FIG. 14. In the section header area, the sets of weights for linear prediction (pset0, pset1, pset2 ...) are described for ref1 and ref2 separately, and each of the sets of weighting coefficients includes only C. Also, FIG. 15 shows an example of sets of weights included in the section header area. Unlike FIG. 7, each of the sets of weighting coefficients of FIG. fifteen includes only C.
FIG. 16 is a block diagram showing the functional structure for generating a predictive image through the linear prediction coefficient generation unit 110, the linear prediction coefficient storage unit 109, and the motion compensation coding unit 107 of FIG. 1.
A predictive image is generated through the linear prediction coefficient generating unit 110, the linear prediction coefficient storage unit 109g, the linear prediction coefficient storage unit 109h, and the linear prediction operation unit 107i.
The sets of weighting coefficients generated by the linear prediction coefficient generating unit 110 are stored in the linear prediction coefficient storage unit 109g and in the linear prediction coefficient storage unit 109h. Using the first benchmark ref1 and the second benchmark ref2 determined by the motion estimation processing, the sets of weights (C1) and (C2), which have one element respectively, are obtained from the units storage of linear prediction coefficients 109g and 109h. These values are input to the linear prediction operating unit 107i, where a linear prediction is performed on them using equations 8a and 8b, the predictive image then being generated.
Also, when linear prediction is carried out with reference to one image only, any one of the sets of weighting coefficients (C1) and (C2) is obtained using only one of between ref1 and ref2 of FIG. 16, linear prediction is carried out using equation 9, and then the predictive image is generated.
It should be noted that the linear prediction coefficient generation unit 110 generates the sets of weighting coefficients (C1) and (C2) by examining the characteristics of an image, creates a predictive image using the method explained in FIG. 16 and then determines a combination of the two reference indices ref1 and ref2 to minimize the prediction error.
Since the present embodiment requires only one parameter to be used for ref1 and ref2, the encoding apparatus can easily determine the values of the parameters, and furthermore, the amount of data to be described in the encoded stream can be reduced. Furthermore, since linear prediction equations do not require complicated operations such as multiplication, the number of operations can also be minimized. Furthermore, the use of the coefficients C1 and C2 allows a radical improvement of the low prediction precision, considered as a disadvantage of the conventional procedure that uses a fixed equation.
It should be noted that it is possible to use the linear prediction procedure explained in the present embodiment, regardless of whether or not a plurality of reference indices can refer to the same image.
(Fourth embodiment)
Next, the moving picture coding method of the fourth embodiment of the present invention will be explained. Since the structure of the coding apparatus, the coding processing flow and the reference index assignment procedure are exactly identical to those of the first embodiment, the explanation thereof will not be repeated.
Display order information indicating the display time, or an alternative to it, as well as an image number are assigned to each image. FIG. 39 is a diagram showing an example of image numbers and the corresponding display order information. Certain values are assigned to the display order information based on the display order. This example uses the value that increases one by one for each image. In the fourth embodiment, a procedure for generating
ES 2 401 774 T3 coefficient values used in an equation for a linear prediction using this display order information.
In the first embodiment, linear prediction is carried out for each pixel using equation 1, equation 3 or equation 4 when generating a predictive image in motion compensation. However, since linear prediction requires coefficient data, such coefficient data is described in the section head areas in a coded stream as sets of weighting coefficients to be used for the creation of the predictive image. Although this method achieves high coding efficiency, it requires additional processing to create data from the sets of weighting coefficients and causes an increase in the number of bits as the sets of weighting coefficients are described in the encoded stream.
Therefore, it is also possible to carry out a linear prediction using Equation 10, Equation 11a, and Equation 12a instead of Equation 1. Using these equations, the weighting coefficients can be determined based on only the order information display of each reference image, so it is not necessary to code the sets of weighting coefficients separately.
In the following equations, Q1 (i) and Q2 (i) are referenced block pixel values, P (i) is a pixel value of a predictive image of a current block to be encoded, V0 and V1 are weighting coefficients, T0 is the display order information of the current image to be encoded, T1 is the display order information of the image designated by the first benchmark and T2 is the display order information of the image designated by the second benchmark.
P (i) = Vl x Q1 (i) 4-V2 x Q2 (i) Equation 10
Vl - (T2-T0) / (T2-TI) ...... Equation 11a
V2 = (T0-T1) / (T2-TI) Equation 12a
When it is assumed, for example, that the current image to be encoded at # 16, that the image designated by the first benchmark is # 11, and that the image designated by the second benchmark is # n 10, the display order information of the respective images is 15, 13 and 10, and therefore the following linear prediction equations are determined.
Vl = (10-15) / (10-13) = 5/3
V2 = (15-13) / (10-13) = -2/3
<img file="ES2401774T3_D0003.tif" />
Compared to the procedure that performs a linear prediction using the sets of weighting coefficients from Equation 1, the above equations have less flexibility with regard to coefficient values and therefore it can be said that it is impossible to create the optimal predictive image. However, compared to the procedure of switching the two fixed equations 2a and 2b depending on the positional relationship between two reference images, the above equations are more efficient linear prediction equations.
When the first benchmark and second benchmark refer to the same image, equation 11a and equation 12a do not hold because T1 equals T2. Therefore, when the two reference images have the same display order information, a linear prediction will be performed using 1/2 as the value of V1 and V2. In that case, the linear prediction equations are as follows.
Vl = 1/2
<img file="ES2401774T3_D0004.tif" />
<img file="ES2401774T3_D0005.tif" />
Furthermore, equations 11a and 12a are not satisfied because T1 equals T2 when the first reference index and the second reference index refer to different images but these images have the same display order information. As mentioned above, when the two reference images have the same display order information, a linear prediction will be performed using 1/2 as the
ES 2 401 774 T3 value of V1 and V2.
As described above, when two reference images have the same display order information, it is possible to use a predetermined value as the coefficient. Such a predetermined coefficient can be one having the same weight as 1/2 as shown in the example above.
On the other hand, the use of equation 10 in the above embodiments requires multiplication and division for linear prediction. Since the linear prediction operation using equation 10 is the operation for all pixels in a current block to be encoded, the addition of multiplications causes a significant increase in the amount of processing.
Therefore, the approximation of V1 and V2 to the powers of 2 as the case in the second embodiment makes it possible to carry out a linear prediction operation using only shift operations and therefore reduce the amount of processing. Equations 11b and 12b are used as linear prediction equations for that case instead of equations 11a and 12a. In the following equations, v1 and v2 are whole numbers.
Vl = ± pow (2, vl) = apx ((T2 — T0) / (T2 — TI)) ......
Equation 11b
V2 = ± pow (2, v2) = apx ((T0 — T1) / (T2 — TI)) ......
Equation 12b (where pow (2, v1) indicates the “v1” -th power of 2 and pow (2, v2) indicates the “v2” -th power of 2) (where = apx () indicates that the value within () will approximate the value on the left)
It should be noted that it is also possible to use equations 11c and 12c instead of equations 11a and 12a, where v1 is an integer.
Vl = ± pow (2, vl) - apx ((T2 — T0) / (T2 — TI)) ...... Equation 11c
V2 = 1 — Vl ...... Equation 12c (where pow (2, v1) indicates the “v1” -th power of 2) (where = apx () indicates that the value within () will approximate the value left)
It should be noted that it is also possible to use equations 11d and 12d instead of equations 11a and 12a, where v1 is an integer.
Vl = 1 —V2 ...... Equation 11d
V2 = ± pow (2, v2) = apx ((T0 — T1) / (T2 — TI)) ......
Equation 12d (where pow (2, v2) indicates the “v2” -th power of 2) (where = apx () indicates that the value inside () will approximate the value on the left)
It should be noted that the value of V1 and V2 approximated to the power of 2 will be, taking equation 11b as an example, the value of ± pow (2, v1) obtained when the values of ± pow (2, v1) and (T2 - T0) / (T2 - T1) get very close to each other as the value of v1 changes by values of one.
For example, in FIG. 39, when the current image to be encoded is No. 16, the image designated by the first benchmark is No. 11 and the image designated by the second benchmark is No. 10, the order information display of the respective images is 15, 13 and 10, so that (T2 -T0) / (T2 T1) and ± pow (2, v1) are determined as follows.
(T2-T0) / (T2 - TI) = (10-15) / (10-13) = 5/3 + pow (2, 0) = 1 + pow (2, 1) = 2
ES 2 401 774 T3
Since 5/3 has a value closer to 2 than to 1, as a result of the approximation it is obtained that V1 = 2.
As another approximation procedure, it is also possible to switch between a round-up approximation and a round-down approximation depending on the relationship between two display order information values T1 and T2.
In that case, the round-up approximation is carried out at V1 and V2 when T1 is after T2, and the round-down approximation is carried out at V1 and V2 when T1 is before T2. It is also possible, conversely, to carry out a round-down approximation of V1 and V2 when T1 is after T2, and a round-down approximation of V1 and V2 when T1 is before T2.
As another approximation procedure that uses display order information, the round-up approximation is carried out on an equation for V1 and the round-down approximation is carried out on an equation for V2 when T1 is later than T2 . As a result, the difference between the values of the two coefficients increases and it is possible to obtain suitable values for an extrapolation. Conversely, when T1 is earlier than T2, the value in the equation for V1 and the value in the equation for V2 are compared and then a round-up approximation is carried out on the lower value and a rounding down approximation to the upper value. As a result, the difference between the values of the two coefficients decreases, making it possible to obtain suitable values for an interpolation.
For example, in FIG. 39, when the current image to be encoded is No. 16, the image designated by the first benchmark is No. 11 and the image designated by the second benchmark is No. 10, the order information display of the respective images is 15, 13 and 10. Since T1 is later than T2, a round-up approximation is carried out in the equation for V1 and a round-down approximation is carried out in the equation for V2. As a result, equations 11b and 12b are calculated as follows.
(1) Equation 11b (T2 — T0) / (T2 —TI) = (10-15) / (10-13) = 5/3 + pow (2, 0) = 1 + pow (2, 1) = 2
As a result of the rounding-up approximation, we obtain that V1 = 2.
(2) Equation 12b (T0 — T1) / (T2 — TI) = (15-13) / (10-13) = -2/3
-pow (2, 0) = -1
-pow (2, -1) = -1/2
As a result of the rounding down approximation, it is obtained that V2 = -1. It should be noted that although equation 10 is only an equation for a linear prediction in the above embodiments, it is also possible to combine this procedure with the linear prediction procedure by using the two fixed equations 2a and 2b explained in the technique section previous. In that case, equation 10 is used in place of equation 2a, and equation 2b is used as is. In other words, equation 10 is used when the image designated by the first benchmark appears behind the image designated by the second benchmark in the display order, while equation 2b is used in other cases.
On the contrary, it is also possible to use equation 10 instead of equation 2b and use equation 2a as is. In other words, equation 2a is used when the image designated by the first benchmark appears behind the image designated by the second benchmark, and equation 10 is used in other cases. However, when the two reference images have the same display order information, a linear prediction is performed using 1/2 as the value of V1 and V2.
It is also possible to describe only the coefficient C in the section head areas to be used for linear prediction, in the same way as the concept of the third embodiment. In that case, equation 13 is used instead of equation 10. V1 and V2 are obtained in the same way as the previous embodiments.
P (i) = Vlx (Ql (i) + Cl) + V2x (Q2 (i) + C2) ...... Equation 13
ES 2 401 774 T3
The processing to generate C coefficients is necessary, and in addition, the C coefficient needs to be encoded in the section head area, but the use of C1 and C2 allows more accurate linear prediction even if the precision of V1 and V2 is low. This is particularly effective when V1 and V2 approach the powers of 2 to perform a linear prediction.
It should be noted that using equation 13, linear prediction can be carried out in the same way in case a reference index is assigned to an image and in case a plurality of reference indices are assigned to an image.
In calculating the values of each of equations 11a, 12a, 11b, 12b, 11c, 12c, 11d, and 12d, the combinations of available values are somewhat limited in each section. Therefore, only one operation is required to encode a section, as opposed to Equation 10 or Equation 13 where the operation needs to be carried out for all pixels in a current block to be encoded, and thus it looks like which does not greatly affect the total processing amount.
It should be noted that the display order information In the present embodiment, it is not limited to the display order, but can be the actual display time or the order of respective images starting from a predetermined image, the value of which increases as time passes. display.
(Fifth realization)
Next, the moving picture coding method of the fifth embodiment of the present invention will be explained. Since the structure of the coding apparatus, the coding processing flow and the reference index assignment procedure are exactly identical to those of the first embodiment, the explanation thereof will not be repeated.
In the conventional procedure, it is possible to switch between generating a predictive image by using fixed equations and generating a predictive image by using sets of weighting coefficients of linear prediction coefficients, using indicators described in an area of common image information in an encoded stream, if necessary.
In the present embodiment, another method for switching the various explained linear prediction procedures from the first to the fourth above embodiments will be explained.
FIG. 17A shows the structure used for the case where five flags (p_flag, c_flag, d_flag, t_flag, s_flag) to control the above switching are described in the section header area of the encoded stream.
As shown in FIG. 17B, p_flag is an indicator indicating whether the weights have been encoded or not. c_flag is an indicator that indicates whether or not only the data for parameter C (C1 and C2), among the parameters for ref1 and ref2, has been encoded. t_flag is an indicator that indicates whether or not the weights for linear prediction are to be encoded using the display order information from the reference images. Finally, s_flag is an indicator that indicates whether or not the weights for linear prediction will approach the powers of 2 for computation using shift operations.
Furthermore, d_flag is an indicator that indicates whether or not to switch two predetermined fixed equations, such as equations 2a and 2b, depending on the temporal positional relationship between the image designated by ref1 and the image designated by ref2, when a prediction is carried out. linear using such two fixed equations. More specifically, when this flag indicates the switching of equations, equation 2a is used in case the image designated by ref1 is later than the image designated by ref2 in the display order, and equation 2b is used in other cases to carry out a linear prediction, as is the case of the conventional procedure. On the other hand, when this indicator indicates not to switch the equations, equation 2b is always used to carry out the linear prediction, regardless of the positional relationship between the image designated by ref1 and the image designated by ref2.
It should be noted that even if Equation 2a is used instead of Equation 2b as an equation to be used without commutation, Equation 2a can be treated in the same way as Equation 2b.
In the encoding apparatus shown in FIG. 1, the motion compensation encoding unit 107 determines whether or not to encode the data related to the sets of weighting coefficients for each section, provides the information of the indicator p_flag to the encoded stream generation unit 103 as a function of the determination and describes the information in the encoded stream shown in FIG. 17A. As a result, it is possible to use the sets of weighting coefficients in a higher performance apparatus to carry out linear prediction and not to use the sets of weighting coefficients in a lower performance apparatus to carry out linear prediction.
ES 2 401 774 T3
Also, in the encoding apparatus shown in FIG. 1, the motion compensation encoding unit 107 determines, for each section, whether or not encoding only the data related to the parameter C (C1 and C2) corresponding to the DC components of the image data, provides the information of the flag c_flag to the coded stream generation unit 103 as a function of the determination and describes the information in the coded stream shown in FIG. 17A. As a result, it is possible to use all sets of weighting coefficients in a higher performance apparatus to carry out linear prediction and to use only the DC components in a lower performance apparatus to carry out linear prediction.
Also, in the encoding apparatus shown in FIG. 1, when linear prediction is carried out using fixed equations, the motion compensation coding unit 107 determines, for each section, whether or not to carry out a coding by switching two equations, provides the information of the flag d_flag to the unit encoded stream generation generator 103 as a function of the determination and describes the information in the encoded stream shown in FIG. 17A. As a result, it is possible to use any one of the fixed equations for linear prediction in case there is a slight temporal change in the brightness of an image and to switch between the two fixed equations for linear prediction in case the brightness of the image image changes as time goes by.
Also, in the encoding apparatus shown in FIG. 1, the motion compensation coding unit 107 determines, for each section, whether or not to generate coefficients for linear prediction using the display order information of the reference images, provides the information of the t_flag flag to the generation unit of coded streams 103 as a function of the determination and describes the information in the coded stream shown in FIG. 17A. As a result, it is possible to encode the sets of weighting coefficients for linear prediction in case there is enough space to encode in the encoded stream and generate the coefficients from the display order information for linear prediction in case there is more space to code.
Also, in the encoding apparatus shown in FIG. 1, the motion compensation coding unit 107 determines, for each section, whether or not to approximate the coefficients for linear prediction to the powers of 2 to allow the calculation of these coefficients through shift operations, provides the information of the flag s_flag to the encoded stream generating unit 103 as a function of the determination and describes the information in the encoded stream shown in FIG. 17A. As a result, it is possible to use the weighting coefficients without approaching in a higher performance apparatus to carry out a linear prediction and to use the weighting coefficients approaching the powers of 2 in a lower performance apparatus to carry out the linear prediction. which can be done only by shift operations.
For example, the case of (1) (p, c, d, t, s_flag) = (1, 0, 0, 0, 1) shows that all sets of weighting coefficients are coded and that a linear prediction only by shift operations with the coefficients represented as the powers of 2 explained in the second embodiment to generate a predictive image.
The case of (2) (p, c, d, t, s_flag) = (1, 1, 1, 0, 0) shows that only the data related to the parameter C (C1 and C2) is encoded, the method of generating a predictive image by adding the coefficient C to the fixed equations as explained in the third embodiment, and furthermore, the two fixed equations are switched for use.
In the case of (3) (p, c, d, t, s_flag) = (0, 0, 0, 0, 0), no set of weights is encoded. In other words, it shows that the procedure is used to generate a predictive image using only equation 2b from among the fixed equations of the conventional procedure.
The case of (4) (p, c, d, t, s_flag) = (0, 0, 1, 1, 1) shows that no set of weights is encoded, but rather that the weights are generated at From the information of the display order of the reference images and the linear prediction is carried out only by means of shifting operations, the generated coefficients approaching the powers of 2 and then two fixed equations switch for use in generating a predictive image.
It should be noted that in the present embodiment, the determinations are made using five flags (p_flag, c_flag, d_flag, t_flag, s_flag), each of which is 1 bit, but it is also possible to represent the determinations only with a 5-bit flag. instead of these five indicators. In that case an encoding using variable length encoding is possible, not through 5-bit representation.
It should be noted that In the present embodiment, all five flags (p_flag, c_flag, d_flag, t_flag, s_flag) are used, each of which is 1 bit, but this can also be applied to the case where the linear prediction procedure toggles using only some of these indicators. In that case, only the indicators necessary for linear prediction among the indicators shown in FIG. 17A.
ES 2 401 774 T3
The conventional procedure allows switching, in each image, between the predictive imaging procedure using fixed equations and the predictive imaging procedure using sets of weighting coefficients of linear prediction coefficients, providing an indicator to switch these procedures in a common image information area in an encoded stream. However, in the conventional method, the predictive imaging method can only be switched for each image.
As explained above, the present embodiment allows the predictive imaging procedure to be switched in each section, which is a subdivision of an image, by providing this switching flag in a section header of a coded stream. Therefore, it is possible to generate a predictive image using the weighting coefficients in a section that has complex images and to generate a predictive image using fixed equations in a section that has simple images, and thus it is possible to improve the image quality while minimizing the increase in the amount of processing.
It should be noted that in the present embodiment, five flags (p_flag, c_flag, d_flag, t_flag, s_flag) are described in a section header area to determine whether or not the coefficients are encoded for each section, but it is also possible to toggle the determination for each image describing these indicators in the image common information area. Furthermore, it is also possible to generate a predictive image using the optimal procedure in each block by providing the toggle flag in each block, which is a subdivision of a section.
It should be noted that the display order information In the present embodiment, it is not limited to the display order, but can be the actual display time or the order of respective images starting from a predetermined image, the value of which increases as it passes. the viewing time.
(Sixth embodiment)
FIG. 2 is a block diagram showing the structure of a moving picture decoding apparatus of the sixth embodiment of the present invention. Using this block diagram shown in FIG. 2, the moving picture decoding procedure executed by this moving picture decoding apparatus will be explained in the following order: (1) a decoding overview, (2) a procedure for assigning benchmarks, and (3) a procedure for generating a predictive image. In this case, it is assumed that an encoded stream that was generated by the moving picture encoding method of the first embodiment is input to the present decoding apparatus.
(1) Decoding overview
First, the coded stream analysis unit 201 extracts from the section header area the data sequence of sets of weighting coefficients for linear prediction and the command sequence for the assignment of reference indices, and from the area of coded block information various information such as reference index information, motion vector information, and coded prediction error data. FIG. 6 shows the various encoded information mentioned above of the encoded stream.
The data sequence of the sets of weighting coefficients for the linear prediction extracted by the coded stream analysis unit 201 is provided to a storage unit of linear prediction coefficients 206, the command sequence for the assignment of reference indices is provided at a 207 benchmark / image number conversion unit, the reference indices are provided to a motion compensation decoding unit 204, the motion vector information is provided to a motion vector storage unit 205, and the encoded prediction error signal is provided to a decoding unit prediction errors 202, respectively.
The prediction error decoding unit 202 performs image decoding processing, such as inverse quantization and inverse frequency transform for the input encoded prediction error signal, and provides a decoded error signal. The summing unit 208 sums the decoded prediction error signal and the predictive image signal provided by the motion compensation decoding unit 204 to generate a reconstructed image signal. The reconstructed image signal obtained is stored in an image memory 203 for use as a reference in subsequent inter-image prediction and provided for display.
The motion compensation decoding unit 204 extracts the most suitable image area as a predictive image from among the reconstructed image signals stored in the image memory 203, using the motion vectors input from the motion vector storage unit. 205 and the benchmarks entered from the encoded stream analysis unit 201. At this time, the reference index / image number conversion unit 207 designates the images of
ES 2 401 774 T3 is referenced in the image memory 203 based on the correspondence between the reference indices provided by the encoded stream analysis unit 201 and the image numbers.
The operation of the image number / reference index conversion unit 207 will be explained in detail in section (2). In addition, the motion compensation decoding unit 204 performs pixel value conversion processing, such as interpolation processing by linear prediction on pixel values in the extracted image area, to generate the final predictive image. . The linear prediction coefficients used for that processing are obtained from the data stored in the linear prediction coefficient storage unit 206 using the reference indices as search keys.
This predictive imaging procedure will be explained in detail in section (3).
The decoded image generated through the aforementioned series of processes is stored in the image memory 203 and provided as an image signal for display according to the display time.
The flow of operations in the case of inter-picture prediction decoding has been described above, but a switch 209 switches between inter-picture prediction decoding and intra-picture prediction decoding. In the case of intra-image decoding, a predictive image is not generated by motion compensation, but a decoded image is generated by generating a predictive image of a current area to be decoded from a decoded area of the same image. and adding the predictive image. The decoded image is stored in the image memory 203, as in the case of inter-image prediction decoding, and is provided as an image signal for display according to the display time.
(2) Procedure for assigning benchmarks
Next, it will be explained, using FIG. 3 and FIG. 4, a benchmark assignment procedure in the benchmark / picture number conversion unit 207 of FIG. 2.
FIG. 3 is a diagram explaining a procedure for assigning two reference indices to picture numbers. When there is a sequence of images arranged in the order of display, the image numbers are assigned in the order of decoding. The commands for assigning the reference indices to the image numbers are described in a section header, which is a subdivision of an image, as the decoding unit and, therefore, the allocation of the same is updated every time a section is decoded. The command indicates, serially by the number of benchmarks, the differential value between an image number that is assigned a current benchmark and an image number that is assigned a benchmark immediately before the current allocation .
Taking the first benchmark of FIG. 3 as an example, since "-1" is first provided as a command, 1 is subtracted from image number 16 of the current image to be decoded, and thus reference index 0 is assigned to image number 15. Next, since -4 is provided, 4 is subtracted from image number 15, and thus benchmark 1 is assigned to image number 11. Subsequent benchmarks are assigned to respective image numbers in the same way. The same applies to the second benchmarks.
According to the conventional benchmark assignment procedure shown in FIG. 34, all reference indices are mapped to respective image numbers. On the other hand, in the example of FIG. 3, the allocation procedure is exactly the same as the conventional procedure, but a plurality of reference indices are matched to the same image number by modifying the command values.
FIG. 4 shows the result of the assignment of the benchmarks. This diagram shows that the first benchmarks and the second benchmarks are assigned to respective image numbers separately, but a plurality of benchmarks are assigned to a picture number in some cases. In the decoding method of the present invention, it is assumed that a plurality of reference indices are assigned to at least one picture number, like this example.
If the benchmarks are used only to determine reference images, the conventional procedure of one-to-one assignment of benchmarks to image numbers is the most efficient procedure. However, in case a set of linear prediction coefficient weights is selected for generating a predictive image using reference indices, the same linear prediction coefficients have to be used for all blocks having the same images. reference, so there is an extremely high probability that the optimal predictive image cannot be generated. Therefore, if it is possible to assign a plurality of reference indices to an image number as in the case of the
ES 2 401 774 T3 present invention, the set of optimal weighting coefficients of linear prediction coefficients can be selected from a plurality of candidate sets for each block even if all blocks have the same reference image and therefore can be generated the predictive image with higher prediction accuracy.
It should be noted that the above description shows the case where the image numbers are provided assuming that all reference images are stored in a reference memory. However, since a current image is given an image number that is one greater than the number of an image that was encoded immediately before the current image, only when the current image that was encoded last is stored Instead, the continuity of the image numbers is kept in the reference memory even if some images are not stored, and therefore the above-mentioned procedure can be used without changes.
(3) Predictive imaging procedure
Next, it will be explained, using FIG. 5, the predictive image generation method in the motion compensation decoding unit 204 of FIG. 2. Although the method of generating images by linear prediction is exactly identical to the conventional method, the flexibility in the selection of linear prediction coefficients is increased because a plurality of numbers of reference indices can be matched to the same image.
Picture B16 is a current B picture to be decoded, and blocks BL01 and BL02 are current blocks to be decoded belonging to picture B. Picture P11 and picture B15 are used as the first reference picture and the second reference image for BL01, and the predictive image is generated with reference to blocks BL11 and BL21 belonging to images P11 and B15, respectively. In the same way, the image P11 and the image B15 are used as the first reference image and the second reference image for BL02, and the predictive image is generated with reference to the blocks BL12 and BL22, respectively.
Although BL01 and BL02 refer to the same images as their first reference image and their second reference image, it is possible to assign different values to the first reference index ref1 and the second reference index ref2 for BL01 and BL02, respectively, using the benchmark assignment procedure explained in section (2). Taking FIG. 4 As an example, 1 and 3 are assigned to the first benchmark corresponding to image number 11, while 1 and 6 are assigned to the second benchmark corresponding to image number 15. As a result, four combinations of these indexes are assumed reference (ref1, ref2) = (1, 1), (1, 6), (3, 1) and (3, 6) and therefore it is possible to select the combination to obtain the optimal set of weighting coefficients for each block between these combinations. In FIG. 5, ref1 = 1 and ref2 = 1 are assigned for BL01, and ref1 = 3 and ref2 = 6 are assigned for BL02, for example.
According to the conventional homing procedure shown in FIG. 35, only a combination of (ref1, ref2) = (1, 1) can be selected for BL01 and BL02 in the case of FIG. 5 and thus only one set of weights of linear prediction coefficients can be selected. On the other hand, according to the present invention, there are four options available and it can be said that the possibility of selection of the optimal set of weighting coefficients increases.
An encoded stream of an image is made up of a common image information area and a plurality of section data areas. FIG. 6 shows the structure of the section data area in the encoded stream. The section data area is further formed by a section header area and a plurality of block data areas. This diagram shows each block area corresponding to BL01 and BL02 of FIG. 5 as an example of the block data area.
"Ref1" and "ref2" included in BL01 designate the first reference index and the second reference index, respectively, indicating two images referenced by block BL01. Furthermore, in the section header area, data to provide the sets of weighting coefficients for the aforementioned linear prediction (pset0, pset1, pset2 ...), are described for ref1 and ref2, respectively. In this area, a number of "p-sets" equivalent to the number of benchmarks explained in section (2) can be set. More specifically, in case ten benchmarks, ranging from 0 to 9, are used as each of the first benchmark and the second benchmark, ten sets ranging from 0 to 9 can be set for ref1 and ref2 .
FIG. 7 shows an example of tables of the sets of weights included in the section header area. Each piece of data indicated by an identifier pset has four values w1, w2, c and d, and these tables are structured in such a way that the values of ref1 and ref2 can refer directly to those values. Also, the idx_cmd1 and idx_cmd2 scripts for assigning reference indices to image numbers are described in the section header area.
Using ref1 and ref2 described in BL01 in FIG. 6, a set of weights is selected to
ES 2 401 774 T3 from each of the tables for ref1 and ref2 in FIG. 7. By performing a linear prediction on the pixel values of the reference images using these two sets of weighting coefficients, a predictive image is generated.
Next, the processing flow from determining sets of weighting coefficients to generating a predictive image will be explained.
FIG. 18 is a functional block diagram showing a functional structure for generating a predictive image in the linear prediction coefficient storage unit 206 and in the motion compensation decoding unit 204.
A predictive image is generated through the linear prediction coefficient storage unit 206a, the linear prediction coefficient storage unit 206b, the averaging unit 204a, and the linear prediction operation unit 204b.
The averaging unit 204a obtains, from the linear prediction coefficient storage unit 206a, a set of weighting coefficients (w1_1, w2_1, c_1, d_1) selected by ref1 provided by the coded stream analysis unit 201 and obtains, from the linear prediction coefficient storage unit 206b, a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected by ref2 provided by the encoded stream analysis unit 201.
The averaging unit 204a calculates the average, for respective parameters, of the sets of weighting coefficients obtained from the linear prediction coefficient storage units 206a and 206b, respectively, to provide the averages to the operations unit. linear prediction 204b as the set of weighting coefficients (w1, w2, c, d) that is actually used for linear prediction. The linear prediction operating unit 204b calculates the predictive image using equation 1 based on the set of weighting coefficients (w1, w2, c, d) to be provided.
FIG. 19 is a functional block diagram showing another functional structure for generating a predictive image. A predictive image is generated through the linear prediction coefficient storage unit 206a, the linear prediction coefficient storage unit 206b, the linear prediction operations unit 204c, the linear prediction operations unit 204d and the unit Averaging 204e.
The linear prediction operations unit 204c obtains, from the linear prediction coefficients storage unit 206a, a set of weighting coefficients (w1_1, w2_1, c_1, d_1) selected by ref1 provided by the flow analysis unit encoded 201, and calculates the predictive image using equation 1 based on the set of weighting coefficients to be provided to the averaging unit 204e.
In the same way, the linear prediction operations unit 204d obtains, from the linear prediction coefficients storage unit 206b, a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected by ref2 provided by the coded stream analysis unit 201, and calculates the predictive image using equation 1 based on the set of weighting coefficients to be provided to the averaging unit 204e.
The averaging unit 204e calculates the average, for respective pixel values, of the predictive images provided by the linear prediction operation unit 204c and the linear prediction operation unit 204d, respectively, to generate the final predictive image at provide.
FIG. 20A is a functional block diagram showing another functional structure for generating a predictive image. A predictive image is generated through the linear prediction coefficient storage unit 206c, the linear prediction storage unit 206d, the averaging unit 204f, and the linear prediction operation unit 204g.
The averaging unit 204f obtains, from the linear prediction coefficient storage unit 206c, the parameters of c_1 and d_1 of a set of weighting coefficients (w1_1, w2_1, c_1, d_1) selected by ref1 provided by the coded stream analysis unit 201 and, likewise, obtains, from the linear prediction coefficient storage unit 206d, the parameters of c_2 and d_2 of a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected by ref2 provided by the coded stream analysis unit 201. The averaging unit 204f calculates the average of c_1 and c_2 and the average of d_1 and d_2 obtained from the linear prediction coefficient storage unit 206c and the linear prediction coefficient storage unit 206d to obtain 'c 'and' d 'and provided to linear prediction operations unit 204g.
Furthermore, the linear prediction operations unit 204g obtains the parameter of w1_1 from the set of weighting coefficients mentioned above (w1_1, w2_1, c_1, d_1) from the unit of
ES 2 401 774 T3 storage of linear prediction coefficients 206c, obtains the parameter of w2_2 from the set of weighting coefficients mentioned above (w1_2, w2_2, c_2, d_2) from the storage unit of linear prediction coefficients 206d, obtains c and d which are the averages calculated by the averaging unit 204f and then calculates the predictive image using equation 1 to provide itself.
More specifically, when determining the set of weighting coefficients (w1, w2, c, d), which is actually used for linear prediction, from among the set of weighting coefficients (w1_1, w2_1, c_1, d_1) obtained from the linear prediction coefficient storage unit 206c and the set of weighting coefficients (w1_2, w2_2, c_2, d_2) obtained from the linear prediction coefficient storage unit 206d, the linear prediction operating unit 204g uses the following rule.
w1 = w1_1 w2 = w2_2 c = (average of c_1 and c_2) d = (average of d_1 and d_2)
FIG. 20B is a functional block diagram showing another functional structure for generating a predictive image. A predictive image is generated through the linear prediction coefficient storage unit 206e, the linear prediction coefficient storage unit 206f, and the linear prediction operation unit 204h.
The linear prediction operations unit 204h obtains, from the linear prediction coefficients storage unit 206e, the parameters of w1_1, c_1 and d_1 that are part of a set of weighting coefficients (w1_1, w2_1, c_1, d_1 ) selected by ref1 provided by the coded stream analysis unit 201 and, likewise, obtains, from the linear prediction coefficient storage unit 206f, the parameter of w2_2 that is part of a set of weighting coefficients (w1_2, w2_2, c_2, d_2) selected by ref2 provided by the coded stream analysis unit 201. The linear prediction operations unit 204h calculates the predictive image using equation 1 based on w1_1, c_1, d_1, w2_2 obtained from the linear prediction coefficient storage unit 206e and the linear prediction coefficient storage unit 206f to be provided.
More specifically, when determining the set of weighting coefficients (w1, w2, c, d), which is actually used for linear prediction, from among the set of weighting coefficients (w1_1, w2_1, c_1, d_1) obtained from the linear prediction coefficient storage unit 206e and the set of weighting coefficients (w1_2, w2_2, c_2, d_2) obtained from the linear prediction coefficient storage unit 206f, the linear prediction operation unit 204h uses the following rule.
w1 = w1_1 w2 = w2_2 c = c_1 d = d_1
Furthermore, it is possible to use one or more parameters among w1, w2, c and d as fixed values. FIG. 21 is a functional block diagram in the case where only "d" is used as a fixed value for the functional structure of FIG. 20 A. A predictive image is generated by the linear prediction coefficient storage unit 206g, the linear prediction coefficient storage unit 206h, the averaging unit 204i, and the linear prediction operation unit 204j.
The coefficients selected by the first reference index ref1 from the linear prediction coefficient storage unit 206g are only (w1_1, c_1), and the coefficients selected by the second reference index ref2 from the storage unit of linear prediction coefficients 206h are only (w2_2, c_2). The averaging unit 204i calculates the average of c_1 and c_2 obtained from the linear prediction coefficient storage unit 206g and the linear prediction coefficient storage unit 206h to obtain "c", and provides it to the linear prediction operations unit 204j.
The linear prediction operations unit 204j obtains the parameter of w1_1 from the linear prediction coefficient storage unit 206g, obtains the parameter of w2_2 from the linear prediction coefficient storage unit 206h, obtains the parameter of c from the averaging unit 204i, and then computes the predictive image using a predetermined fixed value as a parameter of d and equation 1 to be provided. In this case, it is also possible to transform equation 1 into equation 1b for use as explained in the first embodiment.
It is possible to use a predetermined fixed value as a value of d in many cases, but in case the encoding apparatus describes the previous fixed value in the section header, that fixed value can be switched for each
ES 2 401 774 T3 section by extracting it in the encoded stream analysis unit 201. Also, the fixed value can be switched for each image or for each sequence being described in the image common information area or in the sequence common information area.
It should be noted that the decoding procedures mentioned above are related to a B picture having two reference pictures, but it is also possible to perform the same processing in the decoding mode of a reference picture for a P picture or a B picture having just a reference image. In this case, using only one of the first reference index and the second reference index, pset and idx_cmd for ref1 or ref2 are only described in the section header area included in the encoded stream of FIG. 6, according to the reference index described in the block data area. Also, as a linear prediction procedure, the following equation 3 or equation 4 is used instead of the equation 1 explained in the conventional procedure.
Using equation 1 and equation 3 requires four parameters w1, w2, c and d, but using equation 4 allows a linear prediction using only three parameters w1, c and d. In other words, it is possible to reduce the number of data items in the set of weights to be described in each section header area to three, when the first benchmark or the second benchmark is used for a full image as P.
Using equation 3, it is possible to make a linear prediction available for the B images and P images adaptively without changes in structure. If equation 4 is used, the amount of data to be described in the header area of an image P can be reduced, and therefore, it is possible to reduce the amount of processing thanks to a simplified calculation. However, since the reference index assignment procedure suggested by the present invention can be applied directly to any of the above procedures, a predictive image can be created with higher prediction accuracy, which is extremely efficient in image decoding. .
On the other hand, the images referred to in a motion compensation are determined by designating the reference indices assigned to respective images. In that case, the maximum number of images that are available for reference has been described in the image common information area of the encoded stream.
FIG. 38 is a schematic diagram of a coded stream in which the maximum number of images that are available for reference is described. As this diagram shows, the maximum number of images for "Max_pic1" of ref1 and the maximum number of images for "Max_pic2" of ref2 are described in the common image information in the encoded stream.
The information required for decoding is not the maximum number of actual images, but the maximum reference index value available to designate images.
Since in the conventional method a reference index is assigned to an image, the above-mentioned description of the maximum number of images does not generate any contradiction. However, the different number of reference indices and images has a significant influence in case a plurality of reference indices are assigned to a picture number, such as the present invention.
As described above, the idx_cmd1 and idx_cmd2 scripts are described in a coded stream for the purpose of assigning reference indices to image numbers. Image numbers and reference indices are mapped to each other based on each command in these idx_cmd1 and idx_cmd2 scripts. To that end, knowing the maximum benchmark value shows that all benchmarks and image numbers have been mapped to each other, specifically the end of the idx_cmd1 and idx_cmd2 script commands.
Therefore, in the present embodiment, the maximum number of reference indices available, instead of the maximum number of images in the prior art, is described in the common image information area, which is an image header. Alternatively, both the maximum number of images and the maximum number of reference indices are described.
FIG. 23 shows a common image information area of a coded stream of an image in which the maximum number of reference indices is described. The maximum number of reference indices available for "Max_idx1" of ref1 and the maximum number of reference indices available for "Max_idx2" of ref2 are described in the image common information area.
In FIG. 23, the maximum number of reference indices is described in the image common information, but may be structured so that the maximum number of reference indices is described in the section data area as well as in the image common information. . For example, the maximum number of benchmarks required for each section can be clearly described, in case the maximum number of benchmarks
ES 2 401 774 T3 required for each section is significantly different from the maximum number of them described in the common image information area, from section to section; For example, the maximum number of benchmarks in an image is 8, the maximum number of benchmarks required for section 1 in the image is 8, and the maximum number of benchmarks required for section 2 is 4.
In other words, it may be structured so that the maximum number of reference indices described in the common image information is set to be the default value common to all sections of the image, and the maximum number of reference indices The required reference for each section, which is different from the default, is described in the section header.
Although FIG. 23 and FIG. 38 show examples where a coded stream is made up of a common image information area and section data areas, the common image information area and the section data areas can be treated as coded streams other than exactly the one. same way as an encoded stream.
(Seventh realization)
Next, the moving picture decoding method of the seventh embodiment of the present invention will be explained. Since the structure of the decoding apparatus, the decoding processing flow and the reference index assignment procedure are exactly identical to those of the sixth embodiment, the explanation thereof will not be repeated.
In the sixth embodiment, linear prediction is carried out on each pixel using equation 1, equation 3 or equation 4 to generate a predictive image in a motion compensation. However, all these equations include multiplications, which cause a significant increase in the amount of processing, considering that these multiplications are carried out on all pixels.
Therefore, it is possible to use Equation 5 in place of Equation 1, Equation 6 in place of Equation 3, and Equation 7 in place of Equation 4. These equations only allow calculations by bit-shifting operations without using multiplications and therefore reduce the amount of processing.
As in the case of the sixth embodiment, equation 5 is used to generate a predictive image with reference to two images at the same time, and equation 6 or equation 7 is used to generate a predictive image with reference to one image. only. Since these equations require identifiers that indicate plus and minus signs, the sets of weights required for prediction operations are (sign1, m, sign2, n, c) for equations 5 and 6, and (sign1, m, c) for equation 7. sign1 and sign2 are parameters that identify the first and second plus and minus signs, respectively. Although the number of parameters is greater than in the third embodiment, there is only a slight increase in the amount of processing since both sign1 and sign2 can be represented by 1 bit.
Next, the processing flow from determining sets of weighting coefficients to generating a predictive image with reference to two images at the same time will be explained in detail using equation 5.
In the first place, the case in which the functional structure for generating a predictive image is as shown in FIG. 18. The averaging unit 204a obtains the set of weighting coefficients (sign1, m_1, sign2_1, n_1, c_1) from the linear prediction coefficient storage unit 206a. Furthermore, the averaging unit 204a obtains the set of weighting coefficients (sign1_2, m_2, sign2_2, n_2, c_2) from the linear prediction coefficient storage unit 206b.
The averaging unit 204a calculates, for respective parameters, the average of the sets of weighting coefficients obtained from the linear prediction coefficient storage unit 206a and the linear prediction coefficient storage unit 206b to consider the average as the set of weights (sign1, m, sign2, n, c). The linear prediction operations unit 204b calculates the predictive image using equation 5 based on the set of weighting coefficients (sign1, m, sign2, n, c) provided by the averaging unit 204a.
It should be noted that FIG. 18 shows the set of weighting coefficients (w1_1, w2_1, c_1, d_1) and the like obtained from the storage unit of linear prediction coefficients 206a and the like, which are calculated in the case of using equation 1 explained in sixth embodiment, and does not show the parameters used in case the predictive image is obtained using equation 5, since the parameters used in the first case can be replaced by the parameters of the second case. The same applies for the cases of FIG. 19 and FIG. 20, described below.
Now we will explain the case in which the functional structure for the generation of a predictive image is like the
ES 2 401 774 T3 shown in FIG. 19. The linear prediction operations unit 204c calculates a predictive image 1 based on the set of weighting coefficients (sign1_1, m_1, sign2_1, n_1, c_1) obtained from the linear prediction coefficient storage unit 206a. The linear prediction operations unit 204d calculates a predictive image 2 based on the set of weighting coefficients (sign1_2, m_2, sign2_2, n_2, c_2) obtained from the linear prediction coefficient storage unit 206b. Furthermore, the averaging unit 204e calculates, for respective pixels, the average of the predictive images computed by the linear prediction operating units 204c and 204d, respectively, to obtain a predictive image.
In this case, the linear prediction operations unit 204c first calculates the predictive image using equation 5 based on the set of weighting coefficients (sign1_1, m_1, sign2_1, n_1, c_1), so that it is possible to calculate the Predictive image using bit shift operations without using multiplication. The same applies to the linear prediction operating unit 204d. On the other hand, in the case of FIG. 18, since the average of the set of weighting coefficients (sign1_1, m_1, sign2_1, n_1, c_1) and of the set of weighting coefficients (sign1_2, m_2, sign2_2, n_2, c_2) is calculated first, the average of m_1 and m_2 or the average of n_1 and n_2, specifically the exponents of 2, may not be whole numbers and therefore there is a possibility that the amount of processing may increase. Also, if the exponents of 2 are rounded to whole numbers, there is a possibility that the errors will increase.
Next, the case where a predictive image is generated in the functional structure shown in FIG. 20 A. The linear prediction operations unit 204g calculates a predictive image using equation 5, based on the parameters sign1_1 and m_1 that are obtained from the linear prediction coefficient storage unit 206c and used for bit shift operations, the sign2_2 and n_2 parameters that are obtained from the linear prediction coefficient storage unit 206d and used for bit shift operations, and the average c calculated by the averaging unit 204f of the parameters c_1 and c_2 which are obtained from the linear prediction coefficient storage units 206c and 206d.
In this case, since the coefficients used for the bit-shifting operations are the values obtained directly from the linear prediction coefficient storage unit 206c or the linear prediction coefficient storage unit 206d, the exponents of 2 of Equation 5 are whole numbers. Therefore, calculations can be carried out using bit shift operations, and thus the amount of processing can be reduced.
Next, the case where a predictive image is generated in the functional structure shown in FIG. 20B. The linear prediction operations unit 204h calculates a predictive image using equation 5 based on the parameters sign1_1, m_1 and c_1 that are obtained from the storage unit of linear prediction coefficients 206e and the parameters sign2_2 and n_2 that are obtained from the linear prediction coefficient storage unit 206f.
In this case, since the coefficients used for the bit-shifting operations are values that were obtained directly from the linear prediction coefficient storage unit 206e or the linear prediction coefficient storage unit 206f, the exponents of 2 in Equation 5 are whole numbers. Therefore, the predictive image can be calculated using bit shift operations, and therefore the amount of processing can be reduced.
In the cases of FIGS. 20A and 20B, there are parameters that do not need to be added to the encoded stream for transmission, as is the case in FIGS. 10A and 10B in the sixth embodiment, and the amount of data in the encoded stream can be reduced.
Using the linear prediction equations explained in the seventh embodiment, the calculations can be carried out using bit shift operations without using multiplication, so that the amount of processing can be significantly reduced with respect to the sixth embodiment.
In the present embodiment, linear prediction is carried out using equations 5, 6 and 7 instead of equations 1, 3 and 4, and using the set of parameters to be coded (sign1, m, sign2, n, c) instead of (w1, w2, c, d), so that calculations can be performed using only bit-shifting operations and thus the amount of processing is reduced. However, it is also possible, as another approach, to use equations 1, 3 and 4 and (w1, w2, c, d) as is, limiting the selectable values of w1 and w2 to only the values available for shift operations of bits, so that the calculations can be performed using only bit shift operations, and thus the reduction of the amount of processing is achieved in exactly the same structure as that of the sixth embodiment.
Furthermore, as a procedure for such a limitation, it is possible to limit the values of w1 and w2 so that 1 is always selected for such values and an encoded stream is entered that has arbitrary values of only c1 and 2, which are components of cC . Taking the structure of FIG. 21 as an example, (1, c_1) for
ES 2 401 774 T3 ref1 and (1, c_2) for ref2 will be encoded as parameter sets. In this case, the pixel value P (i) of the predictive image is calculated by the following equation where w1_1 and w2_2 in equation 1a are replaced by 1.
P (i) = (Ql (i) + Q2 (i)) / pow (2, d) + (c_l + c_2) / 2 (where pow (2, d) indicates the “d” -th power of 2)
Accordingly, it is possible to significantly reduce the amount of processing for linear prediction even in exactly the same structure as that of the sixth embodiment.
Furthermore, as shown in FIG. 24, in case a flag sft_flg, which indicates whether or not it is possible to perform a linear prediction using only bit shift operations, and a flag cc_flg, which indicates whether or not it is possible to perform a linear prediction using only c , which is a CC component, are described in the common image information of an encoded stream of an image, the decoding apparatus can carry out decoding, referring to these flags, in the structure suitable for linear prediction using only bit shift operations or a decoding in the structure suitable for linear prediction using only a DC component. Accordingly, the amount of processing can be significantly reduced depending on the structure of the decoding apparatus.
(Eighth embodiment)
Next, the moving picture decoding method of the eighth embodiment of the present invention will be explained. Since the structure of the decoding apparatus, the decoding processing flow and the reference index assignment procedure are exactly identical to those of the sixth embodiment, the explanation thereof will not be repeated.
As explained in the prior art section, there is a method for generating a predictive image using a predetermined fixed equation, such as equation 2a and equation 2b, as opposed to the sixth and seventh embodiments in which an image Predictive is generated using a prediction equation obtained from sets of weights of linear prediction coefficients. This conventional method has the advantage that the amount of encoding data can be reduced since it is not necessary to encode or transmit the set of weighting coefficients used to generate the predictive image. Also, the amount of processing for linear prediction can be significantly reduced since the linear prediction equations are simple. However, this procedure using the fixed equations has the problem that the prediction precision is lower because only two linear prediction equations, 2a and 2b, can be selected.
Therefore, in the present embodiment, equations 8a and 8b are used instead of equations 2a and 2b. Equations 8a and 8b are obtained by adding C1 and C2 to equations 2a and 2b. Since only the number of sums for the operation increases, the amount of processing increases slightly compared to the original equations 2a and 2b.
Equations 8a and 8b are prediction equations for generating a predictive image with reference to two images at the same time, but when generating a predictive image with reference to only one image, equation 9 is used in place of equation 3 or equation Equation 4, explained in the previous embodiments.
The sets of weights to use this procedure are only (C1) for ref1 and (C2) for ref2. Therefore, an example of a coded stream of an image obtained using this procedure is as shown in FIG. 14. In the section header area, the sets of weighting coefficients for linear prediction (pset0, pset1, pset2 ...) are described separately for ref1 and ref2, and each of the weighting coefficient sets includes only C. Also, FIG. 15 shows an example of sets of weights included in the section header area. Unlike FIG. 7, each of the sets of weighting coefficients of FIG. fifteen includes only C.
FIG. 22 is a block diagram showing the functional structure for generating a predictive image through the linear prediction coefficient storage unit 206 and the motion compensation decoding unit 204 of FIG. 2.
A predictive image is generated through the linear prediction coefficient storage unit 206a, the linear prediction coefficient storage unit 206b, and the linear prediction operation unit 204a.
The sets of weighting coefficients (C1) and (C2), which have one element respectively, are obtained at
ES 2 401 774 T3 starting from the linear prediction coefficient storage units 206a and 206b by means of the first reference index ref1 and the second reference index ref2 provided by the coded stream analysis unit 201. These values are entered in the linear prediction operations unit 204a, where a linear prediction is carried out on them using equations 8a and 8b, then the predictive image is generated.
Also, when performing a linear prediction with reference to only one image, any of the sets of weighting coefficients (C1) and (C2) are obtained using only one of ref1 and ref2 of FIG. 22, a linear prediction is carried out using equation 9 and then a predictive image is generated.
Since the present embodiment only requires to use one parameter for each of ref1 and ref2, the amount of data to be described in the encoded stream can be reduced. Furthermore, since linear prediction equations do not require complex operations such as multiplication, the number of operations can also be minimized. Furthermore, the use of the coefficients C1 and C2 allows a radical improvement in the low prediction precision, considered as a disadvantage of the conventional procedure of using fixed equations.
It should be noted that it is possible to use the linear prediction procedure explained in the present embodiment, regardless of whether or not a plurality of reference indices can refer to the same image.
(Ninth accomplishment)
Next, the moving picture decoding method of the ninth embodiment of the present invention will be explained. Since the structure of the decoding apparatus, the decoding processing flow and the reference index assignment procedure are exactly identical to those of the sixth embodiment, the explanation thereof will not be repeated.
Display order information indicating the display time, or an alternative to it, as well as an image number are assigned to each image. FIG. 39 is a diagram showing an example of image numbers and the corresponding display order information. Certain values are assigned to the display order information based on the display order. This example uses a value that increases one by one for each image. In the ninth embodiment, a method for generating coefficient values used in a linear prediction equation using this display order information will be explained.
In the sixth embodiment, linear prediction is carried out for each pixel using equation 1, equation 3 or equation 4 during the generation of a predictive image in a motion compensation. However, since linear prediction requires coefficient data, such coefficient data is described in section head areas in a coded stream as sets of weighting coefficients that are used for the creation of the predictive image. Although this method achieves high coding efficiency, it requires additional processing to create data from the sets of weighting coefficients and causes an increase in the number of bits as the sets of weighting coefficients are described in the encoded stream.
Therefore, it is also possible to carry out linear prediction using Equation 10, Equation 11a, or Equation 12a instead of Equation 1. Using these equations, the weighting coefficients can be determined only based on the order information display of each reference image, so there is no need to code the sets of weighting coefficients separately.
When it is assumed, for example, that the current image to be decoded is number 16, that the image designated by the first reference index is number 11 and that the image designated by the second reference index is n 10, the respective image display order information is 15, 13 and 10, and therefore, the following linear prediction educations are determined.
<img file="ES2401774T3_D0006.tif" />
V2 = (15-13) / (10-13) = -2/3
<img file="ES2401774T3_D0007.tif" />
Compared to the procedure for carrying out linear prediction using the sets of weighting coefficients from equation 1, the above equations have less flexibility with regard to the coefficient values and hence it can be said that it is impossible create the optimal predictive image. However, compared to the procedure of switching the two fixed equations 2a and 2b depending on the relation
ES 2 401 774 T3 positional between two reference images, the above equations are more efficient linear prediction equations.
When the first benchmark and second benchmark refer to the same image, equation 11a and equation 12a do not hold because T1 equals T2. Therefore, when the two reference images have the same display order information, a linear prediction can be carried out using 1/2 as the value of V1 and V2. In that case, the linear prediction equations are as follows.
<img file="ES2401774T3_D0008.tif" />
<img file="ES2401774T3_D0009.tif" />
<img file="ES2401774T3_D0010.tif" />
Also, equations 11a and 12a are not satisfied because T1 equals T2 when the first reference index and the second reference index refer to different images but these images have the same display order information. When the two reference images have the same display order information as mentioned above, the linear prediction will be carried out using 1/2 as the value of V1 and V2.
As described above, when two reference images have the same display order information, it is possible to use a predetermined value as the coefficient. Such a predetermined coefficient can be one that has the same weight as 1/2, as shown in the example above.
On the other hand, the use of equation 10 in the present embodiment requires multiplication and division for linear prediction. Since the linear prediction operation using equation 10 is an operation on all pixels in a current block to be decoded, the addition of multiplications causes a significant increase in the amount of processing.
Therefore, the approximation of V1 and V2 to the powers of 2, as in the case of the seventh embodiment, makes it possible to carry out a linear prediction operation only by shifting operations and thus reduce the amount processing. Equations 11b and 12b are used as linear prediction equations for that case instead of equations 11a and 12a.
It should be noted that it is also possible to use equations 11c and 12c in place of equations 11a and 12a.
It should be noted that it is also possible to use equations 11d and 12d in place of equations 11a and 12a.
It should be noted that the value of V1 and V2 approximated to the power of 2 will be, taking equation 11b as an example, the value of ± pow (2, v1) obtained when the values of ± pow (2, v1) and (T2 - T0) / (T2 - T1) get very close to each other as the value of v1 changes by values of 1.
For example, in FIG. 39, when a current image to be decoded is No. 16, the image designated by the first benchmark is No. 11 and the image designated by the second benchmark is No. 10, the order information display of the respective images is 15, 13 and 10, so (T2 - T0) / (T2 - T1) and ± pow (2, v1) are determined as follows.
(T2-T0) / (T2 - TI) = (LO-15) / (10- 13) = 5/3 + pow (2, 0) = 1 + pow (2, 1) = 2
Since 5/3 has a value closer to 2 than to 1, as a result of the approximation it is obtained that V1 = 2.
As another approximation procedure, it is also possible to toggle between a round-up approximation and a round-down approximation, depending on the relationship between two display order information values T1 and T2.
In that case, the round-up approximation is carried out at V1 and V2 when T1 is after T2, and the round-down approximation is carried out at V1 and V2 when T1 is before T2. It is also possible, conversely, to carry out a rounding down approximation on V1 and V2 when T1 is after T2 and a rounding down approximation on V1 and V2 when T1 is before T2.
ES 2 401 774 T3
As another approximation procedure that uses display order information, a round-up approximation is carried out on an equation for V1 and a round-down approximation is carried out on an equation for V2 when T1 is later than T2 . As a result, the difference between the values of the two coefficients increases, making it possible to obtain suitable values for an extrapolation. Conversely, when T1 is earlier than T2, the value in the equation for V1 and the value in the equation for V2 are compared and then a round-up approximation is carried out on the lower value and carried out to perform a rounding down approximation on the top value. As a result, the difference between the values of the two coefficients decreases, making it possible to obtain suitable values for an interpolation.
For example, in FIG. 39, when a current image to be decoded is No. 16, the image designated by the first benchmark is No. 11 and the image designated by the second benchmark is No. 10, the order information display of the respective images is 15, 13 and 10. Since T1 is later than T2, a round-up approximation is carried out in the equation for V1 and a round-down approximation is carried out in the equation for V2. As a result, equations 11b and 12b are calculated as follows.
(1) Equation 11b (T2-T0) / (T2-TI) = (10-15) / (10-13) = 5/3 + pow (2, 0) = 1 + pow (2, 1) = 2
As a result of the rounding-up approximation, we obtain that V1 = 2.
(2) Equation 12b (TO — T1) / (T2 —TI) = (15-13) / (10-13) = -2/3 - pow (2, 0) = —1
-pow (2, -1) = -1/2
As a result of the rounding down approximation, it is obtained that V2 = -1.
It should be noted that although equation 10 is only an equation for linear prediction in the above embodiments, it is also possible to combine this procedure with the linear prediction procedure by using the two fixed equations 2a and 2b explained in the technique section previous. In that case, equation 10 is used in place of equation 2a and equation 2b is used as is. More specifically, equation 10 is used when the image designated by the first benchmark appears behind the image designated by the second benchmark in the display order, while equation 2b is used in other cases.
On the contrary, it is also possible to use equation 10 instead of equation 2b and use equation 2a as is. More specifically, equation 2a is used when the image designated by the first benchmark appears behind the image designated by the second benchmark, and equation 10 is used in other cases. However, when the two reference images have the same display order information, linear prediction is carried out using 1/2 as the value of V1 and V2.
It is also possible to describe only the coefficient C in the section head areas to be used for linear prediction, in the same way as the concept of the eighth embodiment. In that case, equation 13 is used instead of equation 10. V1 and V2 are obtained in the same way as the previous embodiments.
The processing to generate coefficients is necessary and, in addition, the coefficient data needs to be decoded in the section head area, but the use of C1 and C2 allows a more accurate linear prediction even if the precision of V1 and V2 is low . This is particularly effective where V1 and V2 approach the powers of 2 for linear prediction.
It should be noted that using equation 13, linear prediction can be carried out in the same way in case of assigning a reference index to an image and in case of assigning a plurality of reference indices to an image.
In calculating the values of each of equations 11a, 12a, 11b, 12b, 11c, 12c, 11d, and 12d, the combinations of allowed values are limited to some extent in each section. Therefore, a single operation is sufficient to decode a section, as opposed to Equation 10 or Equation 13 where the operation needs to be carried out for all pixels of a current block to be decoded and thus Therefore, it appears to have little influence on the amount of overall processing.
ES 2 401 774 T3
It should be noted that the display order information In the present embodiment, it is not limited to the display order, but may be the actual display time or the order of the respective images starting from a predetermined image, the value of which increases as display time elapses.
(Tenth accomplishment)
Next, the moving picture decoding method of the tenth embodiment of the present invention will be explained. Since the structure of the decoding apparatus, the decoding processing flow and the reference index assignment procedure are exactly identical to those of the sixth embodiment, the explanation thereof will not be repeated.
In the conventional procedure it is possible to switch, if necessary, between the generation of a predictive image by using fixed equations and the generation of a predictive image by using sets of weighting coefficients of linear prediction coefficients, using described indicators. in a common image information area of an encoded stream.
In the present embodiment, another method for switching the various explained linear prediction procedures from the sixth to the ninth above embodiments using pointers will be explained.
FIG. 17A shows the structure used for the case where five flags (p_flag, c_flag, d_flag, t_flag, s_flag) to control the above switching are described in the section header area of the encoded stream.
As shown in FIG. 17B, p_flag is an indicator indicating whether the weights have been encoded or not. c_flag is an indicator that indicates whether or not only the data for parameter C (C1 and C2), among the parameters for ref1 and ref2, has been encoded. t_flag is an indicator that indicates whether or not the weights for linear prediction are to be generated using the display order information from the reference images. Finally, s_flag is an indicator that indicates whether or not the weights for linear prediction will approach the powers of 2 for computation using shift operations.
Furthermore, d_flag is an indicator that indicates whether or not to switch two predetermined fixed equations, such as equations 2a and 2b, depending on the temporal positional relationship between the image designated by ref1 and the image designated by ref2, when a prediction is carried out. linear using such two fixed equations. More specifically, when this flag indicates the switching of equations, equation 2a is used in case the image designated by ref1 is later than the image designated by ref2 in the display order, and equation 2b is used in other cases to carry out a linear prediction, as is the case of the conventional procedure. On the other hand, when this indicator indicates not to switch the equations, equation 2b is always used for linear prediction, regardless of the positional relationship between the image designated by ref1 and the image designated by ref2.
It should be noted that even if Equation 2a is used instead of Equation 2b as an equation to be used without commutation, Equation 2a can be treated in the same way as Equation 2b.
In the decoding apparatus shown in FIG. 2, the encoded stream analysis unit 201 analyzes the value of the flag p_flag and provides the motion compensation decoding unit 204 with an instruction indicating whether or not to decode the data related to the sets of weighting coefficients to generate an image. predictive based on the result of the analysis, and then motion compensation decoding unit 204 performs motion compensation by linear prediction. As a result, it is possible to use the sets of weighting coefficients in a higher performance apparatus to carry out linear prediction and not to use the sets of weighting coefficients in a lower performance apparatus to carry out linear prediction.
Also, in the decoding apparatus shown in FIG. 2, the encoded stream analysis unit 201 analyzes the value of the flag c_flag and provides the motion compensation decoding unit 204 with an instruction indicating whether or not to decode only the data related to the corresponding parameter C (C1 and C2) to the DC components of the image data to generate a predictive image using fixed equations, based on the results of the analysis, and then the motion compensation decoding unit 204 performs motion compensation by linear prediction. As a result, it is possible to use the sets of weighting coefficients in a higher performance apparatus to carry out linear prediction and to use only the DC components in a lower performance apparatus to carry out linear prediction.
Also, in the decoding apparatus shown in FIG. 2, when the coded stream analysis unit 201 analyzes the value of the flag d_flag and a linear prediction is carried out using fixed equations based on the result of the analysis, the coded stream analysis unit 201 provides an instruction, indicating whether
ES 2 401 774 T3 whether or not to switch two decoding equations, to the motion compensation decoding unit 204, in which motion compensation is performed. As a result, it is possible to toggle the procedure so that either of the fixed equations is used for linear prediction in the event of a slight temporal change in the brightness of an image, and the two fixed equations toggle for linear prediction at case the brightness of the image changes as time passes.
Also, in the decoding apparatus shown in FIG. 2, the coded stream analysis unit 201 analyzes the value of the flag t_flag and, based on the result of the analysis, provides the motion compensation decoding unit 204 with an instruction indicating whether or not to generate the coefficients for a linear prediction using the display order information of the reference images, and the motion compensation decoding unit 204 performs motion compensation. As a result, it is possible for the encoding apparatus to encode the sets of weighting coefficients for linear prediction in case further encoding can be carried out and generate the coefficients from the display order information for linear prediction in case no further encoding can be carried out.
Also, in the encoding apparatus shown in FIG. 2, the coded stream analysis unit 201 analyzes the value of the flag s_flag and, depending on the result of the analysis, provides the motion compensation decoding unit 204 with an instruction indicating whether or not to approximate the linear prediction coefficients to the powers of 2 to allow the calculation of these coefficients by means of displacement operations, and the motion compensation decoding unit 204 performs motion compensation. As a result, it is possible to use the weights approaching the powers of 2 in a higher-performance apparatus to carry out the linear prediction, and to use the weights approaching the powers of 2 in a lower-performance apparatus to carry out the prediction. linear that can be done by shift operations.
For example, the case of (1) (p, c, d, t, s_flag) = (1, 0, 0, 0, 1) shows that all sets of weights are decoded and that a linear prediction only by shift operations with the coefficients represented as the powers of 2 explained in the seventh embodiment to generate a predictive image.
The case of (2) (p, c, d, t, s_flag) = (1, 1, 1, 0, 0) shows that only the data related to parameter C (C1 and C2) is decoded, the method for generating a predictive image by adding the coefficient C to the fixed equations as explained in the eighth embodiment, and further, the two fixed equations are switched for use.
In the case of (3) (p, c, d, t, s_flag) = (0, 0, 0, 0, 0), no set of weights is decoded. In other words, it shows that the procedure is used to generate a predictive image using only equation 2b from among the fixed equations of the conventional procedure.
The case of (4) (p, c, d, t, s_flag) = (0, 0, 1, 1, 1) shows that no set of weights is decoded, but linear prediction is carried out only by means of shift operations generating the weighting coefficients from display order information of the reference images and additionally approximating the coefficients to the powers of 2, as explained in the ninth embodiment, and then the two fixed equations are switched to use and generate a predictive image.
It should be noted that in the present embodiment, the determinations are made using five flags (p_flag, c_flag, d_flag, t_flag, s_flag), each of which is 1 bit, but it is also possible to represent the determinations only with a 5-bit flag. instead of these five indicators. In that case a decoding is possible using variable length decoding, not through 5-bit representation.
It should be noted that In the present embodiment, all five flags (p_flag, c_flag, d_flag, t_flag, s_flag) are used, each of which is 1 bit, but this can also be applied to the case where the linear prediction procedure toggles using only some of these indicators. In that case, only the indicators necessary for linear prediction among the indicators shown in FIG. 17A.
In the conventional method an indicator is provided in a common image information area of a coded stream to switch between generating a predictive image using fixed equations and generating a predictive image using sets of coefficients of weighting of linear prediction coefficients, to allow switching between them in each image. However, in this procedure, the predictive imaging procedure can only be switched for each image.
On the contrary, in the present embodiment, as mentioned above, it is possible to switch the method of generating a predictive image for each section, which is a subdivision of an image, by providing the switch flag in a section header of a encoded stream. Therefore it is
ES 2 401 774 T3 possible, for example, to generate the predictive image using the sets of weighting coefficients in a section containing complex images and to generate the predictive image using the fixed equations in a section containing simple images. As a result, the image quality can be improved while minimizing the increase in the amount of processing.
It should be noted that in the present embodiment, five flags (p_flag, c_flag, d_flag, t_flag, s_flag) are described in a section header area for the determination of the procedure in each section, but it is also possible to switch the determination for each image describing these indicators in the image common information area. Furthermore, it is also possible to generate a predictive image using the optimal procedure in each block by providing the toggle flag in each block, which is a subdivision of a section.
It should be noted that the display order information In the present embodiment, it is not limited to the display order, but can be the actual display time or the order of respective images starting from a predetermined image, the value of which increases as it passes. the viewing time.
(Eleventh realization)
Next, the moving picture coding method and the moving picture decoding method of the eleventh embodiment of the present invention will be explained. Since the structures of the encoding apparatus and the decoding apparatus, the encoding processing and decoding processing flows and the reference index assignment procedure are exactly identical to those of the first and sixth embodiments, the following will not be repeated. explanation of them.
In the present embodiment, a technology similar to that explained in the fifth and tenth embodiments will be explained.
The p_flag indicator indicates whether or not a set of parameters is encoded and the c_flag indicator indicates whether or not only the data related to parameter C (C1 and C2) are encoded, among the parameters for ref1 and ref2, described in each section .
In the encoding apparatus shown in FIG. 1, the motion compensation encoding unit 107 determines, for each section or for each block, whether or not to encode the data related to the parameter set and provides, based on the determination, the information of the flag p_flag to the unit of coded stream generation 103, wherein the information is described in the coded stream, as shown in FIG. 40A.
Also, in the encoding apparatus shown in FIG. 1, the motion compensation coding unit 107 determines, for each section or for each block, whether or not to encode only the data related to the parameter C (C1 and C2) corresponding to the DC components of the image data and provides, based on the determination, the information of the c_flag flag to the encoded stream generation unit 103, wherein the information is described in the encoded stream, as shown in FIG. 40A.
On the other hand, in the decoding apparatus shown in FIG. 2, the encoded stream analysis unit 201 analyzes the values of the switching flags p_flag and c_flag and, based on the analysis, provides the motion compensation decoding unit 204 with an instruction indicating whether to generate a predictive image by means of the use of downloaded parameter sets or whether to generate a predictive image using fixed equations, for example, and the motion compensation decoding unit 204 performs motion compensation by linear prediction.
For example, as shown in FIG. 40B, (1) when the p_flag flag is 1 and the c_flag flag is 0, the encoding apparatus encodes all parameter sets, (2) when the p_flag flag is 1 and the c_flag flag is 1, the encoding apparatus encodes only the data related to the parameter C (C1 and C2) and furthermore (3) when the flag p_flag is 0 and the flag c_flag is 0, the encoding apparatus does not encode any parameter set. It should be noted that by determining the indicator values shown in FIG. 40B it can be determined whether or not the DC component of the image data has been encoded using the value of the p_flag flag.
The encoding apparatus processes the parameters as explained in FIG. 8 to FIG. 10, for example, in the previous case (1). Process the parameters as explained in FIG. 16, for example, in the previous case (2). Process the parameters using fixed equations, for example, in the previous case (3).
The decoding apparatus processes the parameters as explained in FIG. 18 to FIG. 20, for example, in the previous case (1). Process the parameters as explained in FIG. 22, for example, in the previous case (2). Process the parameters using fixed equations, for example, in the previous case (3).
Next, another example of a combination of the above cases (1) to (3) will be specifically explained.
ES 2 401 774 T3
In the above example, the parameter encoding (whether or not the decoder receives the parameters) toggles explicitly using the p_flag and c_flag flags, but it is also possible to use a variable-length encoding table (VLC table) in instead of the previous indicators.
As shown in FIGS. 41A and 41B, it is also possible to explicitly select whether to switch between fixed equation 2a and fixed equation 2b.
In this case, the non-commutation of equation 2 means the following. In the prior art section it was explained that, for example, to generate a predictive image, the fixed equation 2a including the fixed coefficients is selected when the image designated by the first reference index appears behind, in the display order, of the image designated by the second reference index, and the equation 2b that includes fixed coefficients is selected in other cases. On the other hand, when the equation is not commanded to switch, as shown in the example of FIG. 41B, this means that the fixed equation 2b that includes fixed coefficients is selected even when the image designated by the first reference index appears behind, in the coding order, the image designated by the second reference index to generate a predictive image .
The information of the v_flag flag for explicitly selecting whether to switch between the fixed equation 2a and the fixed equation 2b is provided by the encoded stream generation unit 103 and described in the encoded stream shown in FIG. 41A.
FIG. 41B shows an example of processing using the v_flag flag. As shown in FIG. 41B, when the v_flag flag is 1, the parameters are not encoded (the parameters are not downloaded to the decoding apparatus. This also applies to what is described below) and the fixed equation 2 does not switch. When the v_flag flag is 01, the parameters are not encoded and fixed equation 2 toggles. When the v_flag flag is 0000, only parameter C is encoded and equation 2 does not switch.
Also, when the v_flag flag is 0001, only the C parameter is encoded and equation 2 toggles. When the v_flag flag is 0010, all parameters are encoded and equation 2 does not switch. When the v_flag flag is 0011, all parameters are encoded and fixed equation 2 toggles.
It should be noted that since all the parameters are coded when v_flag is 0010 and 0011, it is possible to carry out a linear prediction using weighting parameters, without using fixed equations, and in that case, determining whether or not to switch the fixed equation it is ignored.
It should be noted that the v_flag can be switched by the motion compensation encoding unit 107 of the encoding apparatus shown in FIG. 1 and by the motion compensation decoding unit 204 of the decoding apparatus shown in FIG. 2. It is also possible to use the d_flag flag that indicates whether or not to toggle the fixed equation, instead of the v_flag flag, in addition to the previous flags p_flag and c_flag.
As described above, the use of flags allows the switching of whether or not the decoding apparatus receives (download) encoded parameters after the encoding apparatus encodes the parameters. As a result, the parameters to be encoded (or received) can be explicitly switched according to the characteristics of the application and the performance of the decoding apparatus.
In addition, since the fixed equation can explicitly switch, the variety of means for improving the image quality increases, and therefore the coding efficiency also improves. Furthermore, even if the decoding apparatus has no fixed equation, it can switch to the fixed equation explicitly and thus generate a predictive image using the explicitly selected fixed equation.
It should be noted that the location of the indicators is not limited to that shown in FIG. 40. Furthermore, the values of the indicators are not limited to those explained above. Also, since four types of parameter usages can be explicitly displayed using two types of indicators, the parameters can be assigned differently than explained above. Also, all parameters are transmitted in the above example, but all necessary parameter sets can be transmitted, as explained in FIG. 10 and in FIG. 20, for example.
(Twelfth realization)
If a program for realizing the structures of the image encoding method or the image decoding method shown in the above embodiments is recorded on a memory medium such as a floppy disk, then it is possible to easily carry out the processing shown in each one of the embodiments in a separate computer system.
FIG. 25A, 25B and 25C are illustrations showing the case where the processing is carried out in a
ES 2 401 774 T3 computer system using a floppy disk that stores the image encoding procedure or the image decoding procedure from the first to the eleventh embodiments above.
FIG. 25B shows a front view and a cross-sectional view of the appearance of a floppy disk, and the floppy disk itself, and FIG. 25A shows an example of a physical format of a floppy disk as a recording medium body. The flexible disk FD is contained in a sleeve F, and a plurality of tracks Tr are concentrically formed on the surface of the disk in the direction of the radius from the periphery, and each track is divided into 16 sectors Se in the angular direction. Therefore, as far as the floppy disk storing the above-mentioned program is concerned, the image encoding procedure as a program is recorded in an allocated area on the floppy disk FD.
FIG. 25C shows the structure for recording and reproducing the program on and from the FD floppy disk. When the program is recorded on the floppy disk FD, the image encoding procedure or the picture decoding procedure as a program is written to the floppy disk from the computer system Cs through a floppy disk drive. When the image encoding process is generated in the computer system by the program on the floppy disk, the program is read from the floppy disk using the floppy disk drive and transferred to the computer system.
The above explanation is made under the assumption that a recording medium is a floppy disk, but the same processing can also be carried out using an optical disk. Furthermore, the recording medium is not limited to a floppy disk and an optical disk, but any other medium, such as an IC card and a ROM cassette, that can record a program can be used.
(Thirteenth realization)
FIG. 26 to 29 are illustrations of devices for carrying out the encoding processing or decoding processing described in the above embodiments and a system using them.
FIG. 26 is a block diagram showing the global configuration of a content delivery system ex100 to carry out a content delivery service. The area for providing the communication service is divided into cells of the desired size and base stations ex107 to ex110, which are fixed wireless stations, are located in respective cells.
In this content delivery system ex100, devices such as a computer ex111, a personal digital assistant (PDA) ex112, a camera ex113, a mobile phone ex114 and a mobile phone equipped with a camera ex115 are connected to the Internet ex101 through a Internet service provider ex102, a telephone network ex104, and base stations ex107 to ex110.
However, the content delivery system ex100 is not limited to the configuration shown in FIG. 26, and a combination of any of the devices can be connected. Furthermore, each device can be connected directly to the telephone network ex104, not through the base stations ex107 to ex110.
The ex113 camera is a device such as a digital video camera that can capture moving images. The mobile phone may be a mobile phone of a personal digital communication system (PDC), a code division multiple access system (CDMA), a wideband code division multiple access system (W-CDMA ), a global system for mobile communications (GSM), a personal portable telephone system (PHS), or the like.
An ex103 streaming server is connected to the ex113 camera through the ex109 base station and the ex104 telephone network, which allows live distribution or similar using the ex113 camera based on the encoded data transmitted from a user . The camera ex113 or the server to transmit the data can encode the data.
Furthermore, the moving image data captured by a camera ex116 can be transmitted to the streaming server ex103 through the computer ex111. The ex116 camera is a device such as a digital camera that can capture still and moving images. The camera ex116 or the computer ex111 can encode the moving image data. An LSI ex117 included in the computer ex111 or in the camera ex116 actually performs the encoding processing.
The software for encoding and decoding moving images can be integrated into any type of storage medium (such as a CD-ROM, a floppy disk and a hard disk) that is a recording medium that can be read by the computer ex111 or similar. . In addition, a mobile phone equipped with an ex115 camera can transmit the moving image data. This moving picture data is the data encoded by the LSI included in the mobile phone ex115.
The content delivery system ex100 encodes content (such as video of musical performances on
ES 2 401 774 T3 direct) captured by users using the camera ex113, the camera ex116, or the like, in the same way as the previous embodiments and transmits them to the streaming server ex103, while the streaming server ex103 performs a flow distribution of content data to clients upon request. Clients include computer ex111, PDA ex112, camera ex113, mobile phone ex114, and the like, capable of decoding the encoded data mentioned above. In the content delivery system ex100, clients can thus receive and reproduce the encoded data, and furthermore, clients can receive, decode and reproduce the data in real time for personal broadcasting.
When each device of this system performs encoding or decoding, the moving picture encoding apparatus or the moving picture decoding apparatus shown in each of the above-mentioned embodiments can be used.
Next, a mobile phone will be explained as an example of the device.
FIG. 27 is a diagram showing the mobile phone ex115 using the moving picture coding procedure and the moving picture decoding procedure explained in the above embodiments. The mobile phone ex115 has an antenna ex201 to send and receive radio waves to and from the base station ex110, a camera unit ex203, such as a CCD camera that can capture still and video images, a display unit ex202, such such as a liquid crystal display to display the data obtained by decoding video, and the like, captured by the camera unit ex203 and received through the antenna ex201, a body unit including a set of operating keys ex204, a voice output unit ex208, such as a speaker for emitting voice, a voice input unit 205, such as a microphone for entering voice, a storage medium ex207 to store encoded or decoded data, such as still or motion image data captured by the camera, text data, and still or motion image data from received emails, and a slot unit ex206 for coupling the storage medium ex207 to the mobile phone ex115. The storage medium ex207 includes a flash memory element, a type of erasable electrically programmable read-only memory (EEPROM), which is a non-volatile memory that can be electrically rewritten and erased, in a plastic coating such as an SD card. .
The mobile phone ex115 will be further explained with reference to FIG. 28. In the mobile phone ex115, a power supply circuit unit ex310, an operation input control unit ex304, an image coding unit ex312, a camera interface unit ex303, a glass screen control unit liquid (LCD) ex302, an image decoding unit ex309, a multiplexing / demultiplexing unit ex308, a recording / reproducing unit ex307, a modem circuit unit ex306 and a voice processing unit ex305 are connected to a main control unit ex311 structured to globally control the display unit ex202 and the body unit including the operation keys ex204, and are connected to each other via a synchronous bus ex313.
When a user presses the call end key or the power key, the power supply circuit unit ex310 supplies power to the respective units from a battery pack to activate the digital camera equipped mobile phone ex115 to make it pass to a ready state.
In the mobile phone ex115, the voice processing unit ex305 converts the voice signals received by the voice input unit ex205, in talk mode, into digital voice data under the control of the main control unit ex311 which includes a CPU, ROM, RAM and other devices, the modem circuit unit ex306 performs spread spectrum processing of the digital voice data and the send / receive circuit unit ex301 performs a digital-to-analog conversion and frequency transform of the data to transmit it through the antenna ex201. Furthermore, in the mobile phone ex115, after the data received by the antenna ex201 in the talk mode is amplified and subjected to a frequency transform and analog-to-digital conversion, the modem circuit unit ex306 leads to performs reverse spread spectrum processing of the data, and the voice processing unit ex305 converts it to analog voice data to provide it through the voice output unit 208.
In addition, when e-mails are transmitted in the data communication mode, the e-mail text data entered by operation of the body unit operation keys ex204 is sent to the main control unit ex311 via the operations input control unit ex304. In main control unit ex311, after modem circuit unit ex306 performs spread spectrum processing of text data and send / receive circuit unit ex301 performs digital to analog conversion and a frequency transform therein, the data is transmitted to the base station ex110 through the antenna ex201.
When the image data is transmitted in the data communication mode, the image data captured by the camera unit ex203 is supplied to the image coding unit ex312 through the camera interface unit ex303. When no image data is transmitted, it is also possible to display the image data captured by the camera unit ex203 directly on the display unit 202 through the
ES 2 401 774 T3 camera interface unit ex303 and LCD control unit ex302.
The image coding unit ex312, including the image coding apparatus explained in the present invention, compresses and encodes the image data supplied by the camera unit ex203 by the coding method used for the image coding apparatus shown. in the above embodiments to transform it into encoded image data and send it to the multiplexing / demultiplexing unit ex308. At this time, the mobile phone ex115 sends the voices received by the voice input unit ex205 during capture by means of the camera unit ex203 to the multiplexing / demultiplexing unit ex308 as digital voice data through the recording unit. speech processing ex305.
The multiplexing / demultiplexing unit ex308 multiplexes the coded image data supplied by the image coding unit ex312 and the voice data supplied by the voice processing unit ex305 by a predetermined procedure, the modem circuit unit ex306 leads to perform spread spectrum processing on the multiplexed data obtained as a result of multiplexing, and the send / receive circuit unit ex301 performs a digital to analog conversion and a frequency transform on the data for transmission through the antenna ex201.
Regarding the reception of data from a moving picture file that is linked to a web page or the like in the data communication mode, the modem circuit unit ex306 performs reverse spread spectrum processing of the signal received from the base station ex110 through the antenna ex201 and sends the multiplexed data obtained as a result of the processing to the multiplexing / demultiplexing unit ex308.
In order to decode the multiplexed data received through the antenna ex201, the multiplexing / demultiplexing unit ex308 separates the multiplexed data into an encoded image data bitstream and an encoded voice data bitstream and supplies the data. image data encoded to image decoding unit ex309 and voice data to voice processing unit ex305, respectively, via synchronous bus ex313.
Then, the image decoding unit ex309, including the image decoding apparatus explained in the present invention, decodes the encoded bit stream of image data by the decoding method corresponding to the encoding method shown in the above-mentioned embodiments to generate reproduced moving image data and supplies this data to the display unit ex202 through the display unit. LCD control ex302 and therefore Movie clip data contained in a movie clip file linked to a web page, for example, is displayed. At the same time, the voice processing unit ex305 converts the voice data into analog voice data and then supplies this data to the voice output unit ex208 and thus the voice data included in an audio file is reproduced. moving images linked to a web page, for example.
The present invention is not limited to the above-mentioned system, and at least any of the image coding apparatus and the image decoding apparatus of the above-mentioned embodiments can be incorporated into a digital broadcasting system shown in FIG. 29. This digital satellite or terrestrial broadcasting has recently appeared in the media. More specifically, an encoded bit stream of video information is transmitted from a broadcast station ex409 to a communications or broadcast satellite ex410 via radio waves. Upon reception of the same, the broadcasting satellite ex410 transmits radio waves for broadcasting, a domestic antenna ex406 with a satellite broadcast reception function receives the radio waves, and a television (receiver) ex401 or a radio unit. set-top box (STB) ex407 decodes the encoded bitstream for playback.
The image decoding apparatus shown in the above-mentioned embodiments can be implemented in the playback apparatus ex403 to read and decode the encoded bit stream recorded on a storage medium ex402, that is, a recording medium such as a CD and a DVD. In this case, the reproduced video signals are displayed on an ex404 monitor. It is also conceived to implement the image decoding apparatus in the set-top box ex407 connected to a cable ex405 for a cable television or the antenna ex406 for satellite and / or terrestrial broadcasting for its reproduction on the television monitor ex408 ex401. The image decoding apparatus may be incorporated in the television, not in the set-top box. Alternatively, a car ex412 having an antenna ex411 can receive signals from satellite ex410, base station ex107 or the like to reproduce moving images on a display device such as a car navigation system ex413.
Furthermore, the image encoding apparatus shown in the above-mentioned embodiments can encode image signals for recording on a recording medium. As a concrete example, there is an ex420 recorder, such as a DVD recorder for recording image signals on a DVD ex421 disc and a disc recorder for recording on a hard disk. They can also be recorded on an ex422 SD card. If the ex420 recorder includes the image decoding apparatus shown in the mentioned embodiments
ES 2 401 774 T3 above, the image signals recorded on the DVD disc ex421 or the SD card ex422 can be played back for display on the monitor ex408.
The structure without the camera unit ex203, the camera interface unit ex303 and the image coding unit ex312, among the units shown in FIG. 28, can be thought of as the structure of the car navigation system ex413. The same applies to the computer ex111, the television (receiver) ex401 and other devices.
Furthermore, three types of implementations can be conceived for a terminal such as the mobile phone ex114 mentioned above: a sending / receiving terminal that includes both an encoder and a decoder, a sending terminal that includes only one encoder, and a receiving terminal that includes only one decoder.
As described above, it is possible to use the moving picture coding method or the moving picture decoding process in the aforementioned embodiments in any of the aforementioned apparatus and systems and, using this method, they can be obtained the effects described in the previous embodiments.
Industrial applicability
The present invention is suitable for an image coding apparatus to carry out inter-image coding to generate a predictive image by generating commands that indicate a correspondence between image numbers and reference indices to designate reference images and coefficients used for the predictive imaging, designating a reference image that is referenced using a reference index, when performing motion compensation on a block of a current image to be encoded, and performing a linear prediction using a coefficient corresponding to the reference index on a block obtained by motion compensation on the designated reference image . Furthermore, the present invention is suitable for an image decoding apparatus to decode an encoded signal obtained as a result of the encoding performed by the image encoding apparatus.
Contents18
51 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51
121 members in 13 offices
Priority claims34
| Document | Office | Kind | Date |
|---|---|---|---|
| 2002232160 | Japan | A | |
| 2002232160 | Japan | A | |
| 2002232160 | Japan | – | |
| 2002273992 | Japan | A | |
| 2002273992 | Japan | A | |
| 2002273992 | Japan | – | |
| 2002289294 | Japan | A | |
| 2002289294 | Japan | A | |
| 2002289294 | Japan | – | |
| 2002296726 | Japan | A | |
| 2002296726 | Japan | A | |
| 2002296726 | Japan | – | |
| 2002370722 | Japan | A | |
| 2002370722 | Japan | A | |
| 2002370722 | Japan | – | |
| 2003008751 | Japan | A | |
| 2003008751 | Japan | A | |
| 2003008751 | Japan | – | |
| 0309228 | Japan | W | |
| 0309228 | Japan | W | |
| 2002232160 | – | – | – |
| 2002273992 | – | – | – |
| 2002289294 | – | – | – |
| 2002296726 | – | – | – |
| 2002370722 | – | – | – |
| 2003008751 | – | – | – |
| JP20020232160 | – | – | – |
| JP20020273992 | – | – | – |
| JP20020289294 | – | – | – |
| JP20020296726 | – | – | – |
| JP20020370722 | – | – | – |
| JP20030008751 | – | – | – |
| PCTJP2003009228 | – | – | – |
| WO2003JP09228 | – | – | – |
Members121
| Document | Office | Kind | |
|---|---|---|---|
| WO03031023A1 | World Intellectual Property Organization (WIPO) | A1 | |
| JP2003117320A | Japan | A | |
| WO2004015999A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20040054710A | Republic of Korea | A | |
| EP1440722A1 | European Patent Office (EPO) | A1 | |
| JP2004242276A | Japan | A | |
| US2004244344A1 | United States of America | A1 | |
| CN1568622A | China | A | |
| EP1440722A4 | European Patent Office (EPO) | A4 | |
| KR20050031451A | Republic of Korea | A | |
| EP1530374A1 | European Patent Office (EPO) | A1 | |
| US2005105809A1 | United States of America | A1 | |
| PL373781A1 | Poland | A1 | |
| KR100604116B1 | Republic of Korea | B1 | |
| JP2006325238A | Japan | A | |
| US7179516B2 | United States of America | B2 | |
| EP1530374A4 | European Patent Office (EPO) | A4 | |
| CN101083770A | China | A | |
| US7308145B2 | United States of America | B2 | |
| US2008063288A1 | United States of America | A1 | |
| US2008063291A1 | United States of America | A1 | |
| US2008069461A1 | United States of America | A1 | |
| US2008069462A1 | United States of America | A1 | |
| EP1906671A2 | European Patent Office (EPO) | A2 | |
| KR20080073368A | Republic of Korea | A | |
| EP1440722B1 | European Patent Office (EPO) | B1 | |
| DE60228472D1 | Germany | D1 | |
| US7492952B2 | United States of America | B2 | |
| JP4367683B2 | Japan | B2 | |
| EP1906671A3 | European Patent Office (EPO) | A3 | |
| CN1568622B | China | B | |
| CN101083770B | China | B | |
| JP4485157B2 | Japan | B2 | |
| JP4485494B2 | Japan | B2 | |
| KR100969057B1 | Republic of Korea | B1 | |
| KR100969057B1 | Republic of Korea | B1 | |
| KR20100082015A | Republic of Korea | A | |
| KR20100082016A | Republic of Korea | A | |
| KR20100082017A | Republic of Korea | A | |
| KR100976017B1 | Republic of Korea | B1 | |
| US7813568B2 | United States of America | B2 | |
| US7817867B2 | United States of America | B2 | |
| US7817868B2 | United States of America | B2 | |
| CN101867823A | China | A | |
| CN101873491A | China | A | |
| CN101873492A | China | A | |
| CN101873501A | China | A | |
| CN101883279A | China | A | |
| KR101000609B1 | Republic of Korea | B1 | |
| KR101000635B1 | Republic of Korea | B1 | |
| US2010329350A1 | United States of America | A1 | |
| KR101011561B1 | Republic of Korea | B1 | |
| EP2302931A2 | European Patent Office (EPO) | A2 | |
| EP2320659A2 | European Patent Office (EPO) | A2 | |
| US8023753B2 | United States of America | B2 | |
| CN101873491B | China | B | |
| US2011293007A1 | United States of America | A1 | |
| EP2302931A3 | European Patent Office (EPO) | A3 | |
| EP2320659A3 | European Patent Office (EPO) | A3 | |
| US8150180B2 | United States of America | B2 | |
| CN101867823B | China | B | |
| CN101873492B | China | B | |
| US2012195514A1 | United States of America | A1 | |
| CN101873501B | China | B | |
| CN101883279B | China | B | |
| EP1530374B1 | European Patent Office (EPO) | B1 | |
| US8355588B2 | United States of America | B2 | |
| DK1530374T3 | Denmark | T3 | |
| PT1530374E | Portugal | E | |
| US2013094583A1 | United States of America | A1 | |
| ES2401774T3This record | Spain | T3 | |
| US2013101045A1 | United States of America | A1 | |
| US8606027B2 | United States of America | B2 | |
| US2014056351A1 | United States of America | A1 | |
| US2014056359A1 | United States of America | A1 | |
| EP2320659B1 | European Patent Office (EPO) | B1 | |
| PT2320659E | Portugal | E | |
| ES2524117T3 | Spain | T3 | |
| EP2824927A2 | European Patent Office (EPO) | A2 | |
| EP2824928A2 | European Patent Office (EPO) | A2 | |
| US9002124B2 | United States of America | B2 | |
| EP2903272A1 | European Patent Office (EPO) | A1 | |
| EP2903277A1 | European Patent Office (EPO) | A1 | |
| EP2903278A1 | European Patent Office (EPO) | A1 | |
| US9113149B2 | United States of America | B2 | |
| EP2824927A3 | European Patent Office (EPO) | A3 | |
| EP2928185A1 | European Patent Office (EPO) | A1 | |
| EP2824928A3 | European Patent Office (EPO) | A3 | |
| US9456218B2 | United States of America | B2 | |
| US2016360197A1 | United States of America | A1 | |
| US2016360233A1 | United States of America | A1 | |
| EP2903278B1 | European Patent Office (EPO) | B1 | |
| EP2903277B1 | European Patent Office (EPO) | B1 | |
| DK2903277T3 | Denmark | T3 | |
| EP2928185B1 | European Patent Office (EPO) | B1 | |
| PT2903278T | Portugal | T | |
| EP2824927B1 | European Patent Office (EPO) | B1 | |
| PT2903277T | Portugal | T | |
| ES2636947T3 | Spain | T3 | |
| ES2638189T3 | Spain | T3 |
Numbers
- Publication
- 2401774
- Publication, DOCDB
- 2401774
- Publication, EPODOC
- ES2401774T
- Application
- 3741520
- Application, DOCDB
- 03741520
- Application, EPODOC
- ES20030741520T
Titles2
- Spanish
- Procedimiento de codificación y procedimiento de descodificación de imágenes en movimiento
- English
- Coding procedure and decoding procedure of moving images
Classification
- CPC, 23
- H04N19/105
- H04N19/117
- H04N19/134
- H04N19/147
- H04N19/172
- H04N19/174
- H04N19/176
- H04N19/18
- H04N19/42
- H04N19/46
- H04N19/573
- H04N19/577
- H04N19/58
- H04N19/61
- H04N19/70
- H04N19/51
- H04N19/103
- H04N19/137
- H04N19/139
- H04N19/159
- H04N19/177
- H04N19/182
- H04N19/593
- IPC, 7
- G06K9 36
- G06T9 00
- H04N19 89
- H04N7 26
- H04N7 36
- H04N7 46
- H04N7 50