High dynamic range codecs
Abstract
A method for encoding an image of high dynamic range (12), the method comprising the steps of: - obtaining an image of lower dynamic range (14) corresponding to the image of high dynamic range (12), being able to obtain the image of smaller dynamic range (14) from the high dynamic range image (12) by a dynamic range reduction process and containing the same scene as the high dynamic range image (12); - generate a prediction function (19) to predict a high dynamic range image (12) from the lower dynamic range image (14), where the generation of the prediction function implies: for each pixel value represented in the image with the lowest dynamic range, identify the set of these pixels in the image with the lowest dynamic range (14) that have said pixel value; and for each of these sets, identify the pixels in the high dynamic range image (12) that correspond to the pixels in the respective set, thereby identifying the pixel values in the high dynamic range image ( 12) that correspond to each pixel value represented in the image with the lowest dynamic range (14), and where the prediction function (19) is based, at least in part, in the pixel values of the pixels in the high dynamic range image (12), for which the corresponding pixels in the lower dynamic range image (14) all have the same pixel value, and use the statistical relationships between the pixel values of the pixels in the lower dynamic range image and the corresponding pixel values in the high dynamic range image (12), thereby determining for each set a pixel value predicted in the high dynamic range image (12); - apply the prediction function (19) to the image with the lowest dynamic range (14) to obtain an image of high expected dynamic range (29); - obtaining a residual image (32) from the expected high dynamic range image (29) and the high dynamic range image (12); and - encode and store the data representing the image with the lowest dynamic range (14), the prediction function (19) and the residual image (32) in a video stream, where the high dynamic range image (12) and the image with the lowest dynamic range (14), each, comprises a frame in a video sequence, and the prediction function (19) is updated for each frame in the video sequence.
Term
Term ended
Projected expiry passed 7 September 2026, 0 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
12 claims: 5 independent, 7 dependent
- 1ES 2 551 561 T3 REIVINDICACIONES 1. Un método para codificar una imagen de alto rango dinámico (12), comprendiendo el método los pasos de:- obtener una imagen de menor rango dinámico (14) correspondiente a la imagen de alto rango dinámico (12), pudiéndose obtener la imagen de menor rango dinámico (14) a partir de la imagen de alto rango dinámico (12) mediante un proceso de reducción del rango dinámico y que contiene la misma escena que la imagen de alto rango dinámico (12);- generar una función de predicción (19) para predecir una imagen de alto rango dinámico (12) a partir de la imagen de menor rango dinámico (14), en donde la generación de la función de predicción supone: para cada valor de píxel representado en la imagen de menor rango dinámico, identificar el conjunto de estos píxeles en la imagen de menor rango dinámico (14) que tienen dicho valor de píxel;y para cada uno de estos conjuntos, identificar los píxeles en la imagen de alto rango dinámico (12) que se corresponden con los píxeles en el conjunto respectivo, identificando, de ese modo, los valores de píxel en la imagen de alto rango dinámico (12) que se corresponden con cada valor de píxel representado en la imagen de menor rango dinámico (14), y en donde la función de predicción (19) se basa, al menos en parte, en los valores de píxel de los píxeles en la imagen de alto rango dinámico (12), para los cuales los píxeles correspondientes en la imagen de menor rango dinámico (14) tienen todos el mismo valor de píxel, y utiliza las relaciones estadísticas entre los valores de píxel de los píxeles en la imagen de menor rango dinámico y los valores de píxel correspondientes en la imagen de alto rango dinámico (12), determinando, de ese modo, para cada conjunto un valor de píxel previsto en la imagen de alto rango dinámico (12);- aplicar la función de predicción (19) a la imagen de menor rango dinámico (14) para obtener una imagen de alto rango dinámico prevista (29);- obtener una imagen residual (32) a partir de la imagen de alto rango dinámico prevista (29) y de la imagen de alto rango dinámico (12);y - codificar y almacenar los datos que representan la imagen de menor rango dinámico (14), la función de predicción (19) y la imagen residual (32) en un flujo de video, en donde la imagen de alto rango dinámico (12) y la imagen de menor rango dinámico (14), cada una, comprende una trama en una secuencia de video, y la función de predicción (19) se actualiza para cada trama en la secuencia de video.
- 2Un método de acuerdo con cualquiera de la reivindicación 1, en donde generar la función de predicción (19) comprende uno o más de:(a) generar una función que es no lineal en el dominio logarítmico;(b) calcular una media aritmética de los valores de píxel de los píxeles en el alto rango dinámico (12) que pertenecen a cada uno de la pluralidad de los grupos;(c) calcular una mediana de los valores de píxel de los píxeles en la imagen de alto rango dinámico (12) que pertenecen a cada uno de la pluralidad de los grupos;(d) calcular un promedio de los valores de píxel más alto y más bajo de los píxeles en la imagen de alto rango dinámico (12) que pertenecen a cada uno de la pluralidad de los grupos;y (f) combinaciones de estos.
- 3Un método de acuerdo con cualquiera de las reivindicaciones 1 a 2, en donde la función de predicción (19) comprende una relación de uno a muchos entre los píxeles de la imagen de menor rango dinámico (14), que tiene un mismo valor de píxel y píxeles correspondientes que la imagen prevista.
- 4Un método de acuerdo con una cualquiera de las reivindicaciones 1 a 3 que comprende filtrar la imagen residual (32) con el fin de eliminar el ruido antes de almacenar los datos que representan la imagen residual (32).
- 5Un método de acuerdo con la reivindicación 4 en el que el filtrado de la imagen residual (32) comprende:aplicar una transformada wavelet (de ondículas) discreta a la imagen residual (32) y a la imagen de alto rango ES 2 551 561 T3 dinámico (12) con el fin de obtener una imagen residual transformada y una imagen de alto rango dinámico transformada;asignar el valor cero a los valores umbral para los coeficientes en la imagen residual transformada basados en valores de coeficientes en la imagen residual transformada si los coeficientes tienen valores que no exceden los umbrales correspondientes.
- 6Un método de acuerdo con la reivindicación 5 en el que el establecimiento de los valores umbral comprende aplicar una función de evaluación de umbral a los coeficientes en la imagen de alto rango dinámico transformada.
- 7Un método de acuerdo con la reivindicación 6 en el que la función de elevación de umbral comprende uno o más de los siguientes:(a) elevar los coeficientes a una potencia constante predeterminada;(b) multiplicar los coeficientes por un número constante predeterminado;(c) una función dada por en caso contrario o un equivalente matemático de los mismos.
- 8Un método de acuerdo con una cualquiera de las reivindicaciones 6 a 7 que comprende aplicar factores de ponderación de una función de sensibilidad al contraste predeterminada a los coeficientes de la imagen de alto rango dinámico transformada antes de aplicar la función de elevación de umbral a los coeficientes.
- 9Un método de acuerdo con una cualquiera de las reivindicaciones 6 a 8 que comprende aplicar una función de incertidumbre de fase a los coeficientes de la imagen de alto rango dinámico transformada antes de aplicar la función de elevación de umbral a los coeficientes.
- 10Un codificador de imágenes para codificar una imagen de alto rango dinámico (12), comprendiendo un procesador configurado para ejecutar instrucciones que hagan que el procesador:- obtenga una imagen de menor rango dinámico (14) correspondiente a la imagen de alto rango dinámico (12), pudiéndose obtener la imagen de menor rango dinámico (14) a partir de la imagen de alto rango dinámico (12) mediante un proceso de reducción del rango dinámico y que contiene la misma escena que la imagen de alto rango dinámico (12);- genere una función de predicción (19) para predecir una imagen de alto rango dinámico (12) a partir de la imagen de menor rango dinámico (14), en donde la generación de la función de predicción (19) supone: para cada valor de píxel representado en la imagen de menor rango dinámico, identificar el conjunto de estos píxeles en la imagen de menor rango dinámico (14) que tienen dicho valor de píxel;y para cada uno de estos conjuntos, identificar los píxeles en la imagen de alto rango dinámico (12) que se corresponden con los píxeles en el conjunto respectivo, identificando, de ese modo, los valores de píxel en la imagen de alto rango dinámico (12) que se corresponden con cada valor de píxel representado en la imagen de menor rango dinámico (14), y en donde la función de predicción (19) se basa, al menos en parte, en los valores de píxel de los píxeles en la imagen de alto rango dinámico, para los cuales los píxeles correspondientes en la imagen de menor rango dinámico (14) tienen todos el mismo valor de píxel, y utiliza las relaciones estadísticas entre los valores de píxel de los píxeles en la imagen de menor rango dinámico y los valores de píxel correspondientes en la imagen de alto rango dinámico (12), determinando, de ese modo, para cada conjunto un valor de píxel previsto en la imagen de alto rango dinámico (12);- aplique la función de predicción (19) a la imagen de menor rango dinámico (14) con el fin de obtener una imagen de alto rango dinámico prevista (29);- obtenga una imagen residual (32) a partir de la imagen de alto rango dinámico prevista (29) y la imagen de alto rango dinámico (12);y - codifique y almacene los datos que representan la imagen de menor rango dinámico (14), la función de predicción (19) y la imagen residual (32) en un flujo de video, ES 2 551 561 T3 en donde la imagen de alto rango dinámico (12) y la imagen de menor rango dinámico (14), cada una, comprende una trama en una secuencia de video, y la función de predicción (19) se actualiza para cada trama en la secuencia de video.
- 11Un equipo para decodificar una imagen de alto rango dinámico, comprendiendo el equipo:- medios para recuperar los datos que representan una imagen de menor rango dinámico (22) correspondiente a la imagen de alto rango dinámico y a una imagen residual (35);- medios para recuperar los datos que representan una función de predicción (37), siendo transmitidos los datos desde un codificador;- medios para aplicar la función de predicción a la imagen de menor rango dinámico con el fin de obtener una imagen de alto rango dinámico prevista;y - medios para combinar la imagen residual con la imagen de alto rango dinámico prevista con el fin de obtener la imagen de alto rango dinámico, en donde la función de predicción se basa, al menos en parte, en los valores de píxel de los píxeles en la imagen de alto rango dinámico, para los cuales los píxeles correspondientes en la imagen de menor rango dinámico (22) tienen todos el mismo valor de píxel, y utiliza las relaciones estadísticas entre los valores de píxel de los píxeles en la imagen de menor rango dinámico (22) y los valores de píxel correspondientes en la imagen de alto rango dinámico;en donde para cada valor de píxel representado en la imagen de menor rango dinámico (22), los valores de píxel correspondientes en la imagen de alto rango dinámico son los valores de píxel de aquellos píxeles en la imagen de alto rango dinámico que se corresponden con los píxeles en un conjunto respectivo de píxeles en la imagen de menor rango dinámico (22) en el que tienen todos el valor de píxel respectivo en la imagen de menor rango dinámico;en donde la imagen de menor rango dinámico (22) se puede obtener a partir de la imagen de alto rango dinámico mediante un proceso de reducción del rango dinámico y contiene la misma escena que la imagen de alto rango dinámico;y en donde la imagen de alto rango dinámico (12) y la imagen de menor rango dinámico (14), cada una, comprende una trama en una secuencia de video, y la función de predicción (19) se actualiza para cada trama en la secuencia de video.
- 12Un método para decodificar una imagen de alto rango dinámico en un decodificador, comprendiendo el método:- recuperar los datos que representan una imagen de menor rango dinámico (22) correspondiente a la imagen de alto rango dinámico y una imagen residual (35);- recuperar los datos que representan una función de predicción (37), donde los datos se transmiten desde un codificador;- aplicar la función de predicción a la imagen de menor rango dinámico con el fin de obtener una imagen de alto rango dinámico prevista;y - combinar la imagen residual con la imagen de alto rango dinámico prevista con el fin de obtener la imagen de alto rango dinámico, en donde la función de predicción se basa, al menos en parte, en los valores de píxel de los píxeles en la imagen de alto rango dinámico, para los cuales los píxeles correspondientes en la imagen de menor rango dinámico (22) tienen todos el mismo valor de píxel, y utiliza las relaciones estadísticas entre los valores de píxel de los píxeles en la imagen de menor rango dinámico (22) y los valores de píxel correspondientes en la imagen de alto rango dinámico;en donde para cada valor de píxel representado en la imagen de menor rango dinámico (22), los valores de píxel correspondientes en la imagen de alto rango dinámico son los valores de píxel de aquellos píxeles en la imagen de alto rango dinámico que se corresponden con los píxeles en un conjunto respectivo de píxeles en la imagen de menor rango dinámico (22) en el que tienen todos el valor de píxel respectivo en la imagen de menor rango dinámico;en donde la imagen de menor rango dinámico (22) se puede obtener a partir de la imagen de alto rango dinámico mediante un proceso de reducción del rango dinámico y contiene la misma escena que la imagen de alto ES 2 551 561 T3 rango dinámico;y - en donde la imagen de alto rango dinámico (12) y la imagen de menor rango dinámico (14), cada una, comprende una trama en una secuencia de video, y la función de predicción (19) se actualiza para cada trama en la secuencia de video.
Independent claims12
224 paragraphs in 9 sections, as filed
ES 2 551 561 T3
DESCRIPTION
High dynamic range codees.
Technical field
The invention is related to the encoding of image data. The invention has a specific application for encoding images or for encoding video data sequences.
Background
Dynamic range is a measure of the relative brightness of the lightest and darkest parts of an image. Until recently, most televisions, computer monitors, and other display devices have been capable of reproducing dynamic ranges of only a few hundred to one. This is much less than the dynamic range that can be appreciated by the human eye. Display devices with higher dynamic ranges are becoming available. Such high dynamic range displays can provide images that are much more natural and realistic than images produced by "low dynamic range" displays.
High dynamic range displays are suitable for a wide spectrum of applications. For example, high dynamic range displays can be used to display realistic video images ranging from movies and game visuals to visual effects displays in simulators such as flight simulators. High dynamic range displays can also be used in demanding image processing applications such as medical imaging.
Many current image data formats specify the values of a pixel using 24 or fewer bits per pixel. These bits specify both the brightness and the color of the pixel. 24 bits are too few to specify both a full range of colors and a brightness that can vary smoothly throughout the range that a high dynamic range display is capable of reproducing. In order to fully benefit from a high dynamic range display it is necessary to provide image data capable of specifying a wide range of pixel values. Various high dynamic range data formats have been developed or proposed that provide a higher number of bits per pixel. Such high dynamic range data formats are typically not backward compatible with earlier low dynamic range data formats.
For example, the "Perception-motivated HDR Video Encoding" HDRV as described in R. Mantiuk, G. Krawczyk, K. Myszkowski and HP. Seidel. Perceptionmotivated high dynamic range video encoding. ACM Transactions on Graphics (SIGGRAPH Proceedings 2004), 23 (3): 730-38, 2004 is a lossy HDR video compression method which is not backward compatible. The method encodes HDR pixels using 11 bits for luminance and 2 by 8 bits for chrominance. The resulting video stream does not contain any information about the LDR frames.
HDR JPEG is described in Greg Ward and Mariann Simmons. Subband encoding of high dynamic range imagery. In APGV '04: Proceedings of the 1<sup>st</sup> Symposium on Applied perception in graphics and visualization, pages 83-90, New York, NY, USA, 2004. ACM Press. This method involves subsampling a subband layer, which results in the loss of high frequencies. In order to prevent this, the method suggests three techniques: a pre-correction of the LDR layer, in order to encode within this layer the high frequencies that can be lost due to subsampling; a post-correction that attempts to restore lost high frequencies instead of modifying the LDR image and performing a full sampling, which means no subsampling is performed.
Therefore there remains a need for practical methods and equipment for encoding and decoding HDR images, especially HDR video images. There is still a specific need for such methods and equipment that provide backward compatibility with existing hardware for reproducing lower dynamic range images.
US Patent Application US 2005/0259729 A1 discloses a method for encoding a quality scalable video stream. An N-bit input frame is converted to an M-bit input frame, where M is an integer between 1 and N. In order to be backward compatible in existing 8-bit video systems, it can be chosen that M let 8. The M-bit input frame is encoded to produce a base layer output bit stream. An M-bit output frame is reconstructed from the base layer output bit stream and converted to an N-bit output frame. The N-bit output frame is compared to the N-bit input frame in order to obtain a difference from the N-bit image that is encoded to produce an enhanced layer bit stream.
ES 2 551 561 T3
Summary of the invention
This invention provides methods and equipment for encoding high dynamic range image data and for decoding the data to provide both lower dynamic range image data and higher dynamic range image data. The methods and equipment can be applied to encoding video data. In some embodiments of the invention the data with a lower dynamic range is encoded in a standard format such as, for example, an MPEG (Moving Picture Experts Group) format.
One aspect of the invention provides a method for encoding a high dynamic range image. The method comprises obtaining a lower dynamic range image corresponding to the high dynamic range image; identifying groups of pixels in the high dynamic range image for which the corresponding pixels in the lower dynamic range image all have the same pixel value; generating a prediction function based at least in part on the pixel values of the pixels in the high dynamic range image belonging to each of a plurality of the groups; apply the prediction function to the image of lower dynamic range in order to obtain a predicted image; calculating a residual image representing the differences between the pixel values in the predicted image and the corresponding pixel values in the high dynamic range image; and, encoding and storing the data representing the image with the lowest dynamic range, the prediction function, and the residual image.
Other aspects of the invention provide methods for decoding high dynamic range images that have been encoded in accordance with the invention and equipment for encoding and / or decoding high dynamic range images.
Additional aspects of the invention and characteristics of specific embodiments of the invention are described below.
Brief description of the drawings
In drawings illustrating non-limiting embodiments of the invention, Figure 1 is a flow chart illustrating a coding method in accordance with one embodiment of the invention;
Figure 1A is a histogram of pixel values of a high dynamic range image for which all corresponding pixels in a lower dynamic range version of the image have the same pixel value;
Figure 2 is a flow chart illustrating a decoding method in accordance with the invention;
Figure 3 is a flow chart illustrating an MPEG encoding method according to a specific embodiment;
Figures 4A to 4F show the relationship between the luma (image brightness) values in the corresponding HDR and LDR images for various tone mapping algorithms;
Figure 5 shows a method for filtering residual image data in accordance with one embodiment of the invention; and, Figure 6 is a diagram illustrating bit rate as a function of image quality parameter for a prototype coding system.
Description
Specific details are set forth throughout the following description in order to provide a thorough understanding of the invention. However, the invention can be practiced without these details. In other examples, well-known elements are not shown or described in detail in order to avoid unnecessary masking of the invention. Accordingly, the specification and drawings are to be considered illustrative rather than restrictive.
Figure 1 shows a method 10 for encoding an image data frame in accordance with a basic embodiment of the invention. The method 10 encodes both high dynamic range (HDR) data 12 and lower dynamic range (LDR) data 14 into encoded image data 38. As described below, encoded image data 38 can be decoded to reconstruct both LDR data as HDR data.
By way of example only, HDR 12 data can be represented in a color space such as the CIE XYZ (2E standard observer) color space in which the color and brightness of each pixel is specified using three floating point numbers. The LDR 14 data can be represented in a color space such as the sRGB color space in which the color and brightness of each pixel is specified by three bytes. In some embodiments, the LDR data 14 is obtained from the
ES 2 551 561 T3 HDR data 12 (or a precursor to HDR data 12) by an appropriate dynamic range reduction process 16.
Dynamic range reduction may comprise tone mapping and / or color gamut mapping, for example. Any tone mapping or color gamut mapping operator can be used. For example, a tone mapping operator can be selected in order to saturate both luminance and color, change color values, and improve local contrast. Such changes may result in a lower compression rate, but both LDR and HDR frames will be preserved in the resulting video stream.
At block 18, method 10 applies a prediction function 19. Prediction function 19 outputs a predicted pixel value for a pixel in the HDR data 12 based on the pixel value for the corresponding pixel in the LDR data 14. Since the goal is to be able to reproduce the HDR data 12 and the LDR data 14 from the encoded image data 38, it is preferable to base the prediction function 19 on a version of the LDR data 14 that can be reconstructed from the data. encoded image images 38.
When using a lossy algorithm to encode and compress the LDR 14 data, it is not possible to guarantee that the reconstructed version of the LDR 14 data is identical to the original LDR 14 data. For this reason, Figure 1 shows that block 19 receives as input the reconstructed LDR data 26. The reconstructed LDR data 26 is obtained by encoding / compressing the LDR data 14 in block 20 in order to provide compressed encoded LDR data 22 and then decoding / decompressing the compressed encoded LDR data 22 in block 24. The LDR data Compressed encoders 22 are included in the encoded image data 38. Line 15 illustrates a less precise alternative in which block 18 directly uses the LDR data 14 to apply the prediction function 19.
The prediction function 19 preferably uses the statistical relationships between the pixel values in the reconstructed LDR data 26 and the corresponding pixel values in the HDR data 12. In general, if all the pixels in the reconstructed LDR image 26 are taken for the Since all pixels have the same specific pixel value, the corresponding pixels in the HDR data image 12 will not all have the same pixel value. That is, in general, there is a one-to-many relationship between LDR pixel values and HDR pixel values.
Figure 1A is a histogram in which the horizontal axis represents all possible HDR pixel values and the vertical axis indicates the number of pixels in which the image represented by the HDR image data 12 has such a value. There can be a significant number of pixel values for which the image does not have any pixels that have that value. The shaded bars in Figure 1A represent pixel values in the HDR image data 12 for which all corresponding pixels in the reconstructed LDR image data 26 have the same Xldr pixel value. correspond to the range of LDR pixel value Xldr vary between A and B. All HDR pixel values for pixels that correspond to the same pixel value in the reconstructed LDR image data 26 can be called a container. It is normal, but not mandatory, that the different containers do not overlap.
A prediction function 19 for an image can be obtained from the HDR image data 12 and the reconstructed LDR image data 26 by grouping the HDR pixel values into containers and statistically analyzing each of the containers. Grouping the HDR pixel values into containers can comprise:
• taking the data from the reconstructed LDR image 26, and for each of the pixel values represented in the data from the reconstructed LDR image 26 identifying the set of all pixels having said pixel value;
• for each of the pixel sets identify the corresponding pixels in the HDR data 12 and determine the pixel values of those corresponding pixels to generate a set of all HDR pixel values that correspond to each of the pixel values LDR.
Prediction function 19 can be obtained by any of the following operations:
• calculate the arithmetic mean of the HDR pixel values in each of the bins;
• calculate the median of the HDR pixel values in each of the bins;
• calculate the average of the A and B values that delimit the container;
• some combination of the above; or • similar.
The arithmetic mean is considered to provide a good combination of computational precision and efficiency for many applications.
ES 2 551 561 T3
Given a prediction function 19, it is only necessary to encode the differences between the values predicted by the prediction function 19 and the actual values of the HDR image data 12. Said differences are normally close to zero and therefore can be efficiently compress into residual frames.
The prediction function 19 needs to be defined only for the possible pixel values in the LDR data 14 (256 values in the case where the pixel values are represented by an 8-bit number). The prediction function 19 may comprise a lookup table indexed with the valid values for the LDR pixels. Prediction function 19 can be implemented as a lookup table that has a corresponding output value for each of the index values. For example, when the LDR pixels have 8-bit values, the lookup table can comprise 256 different values indexed by integers in the range 1 to 256. Prediction function 19 does not need to be continuous as its main function is to do the Residual frame values as small as possible. Alternatively, the prediction function 19 may be partially or completely represented by an appropriately parameterized solid curve.
In block 28 the method 10 obtains a predicted HDR image by applying the prediction function 19 to the reconstructed LDR data 26. The pixel value for each of the pixels in the reconstructed LDR data 26 is used as an input to the prediction function 19 and the pixel value is substituted for the resulting output of the prediction function 19 to generate a predicted HDR image 29.
Block 30 calculates a difference between the predicted HDR image 29 and the HDR data image 12 in order to provide a residual image 32. The residual image 32 is encoded / compressed in block 34 and results in the data of the residual image 35 for inclusion in the encoded image data 38. Block 34 may comprise filtering and quantization of residual image 32 in order to remove information that will not have an observable effect (or, with more aggressive filtering and / or quantization, an excessively detrimental effect) on the fidelity of an HDR image. reconstructed from the encoded image data 38.
Figure 2 shows a method 40 for decoding the encoded image data 38. The LDR data 22 can be extracted from the encoded image data 38 and decoded / decompressed in block 32 in order to generate the LDR data 43 that they are obtained as LDR 44 data output. If LDR 44 data output is all that is needed, then no additional processing is required.
If an HDR data output 56 is also required, then in block 46 a prediction function 37 is transformed in order to generate a prediction function 47 and in block 50 residual image data 35 is decoded / decompressed with in order to generate the residual image 52.
In block 48 the prediction function 47 is applied to the LDR data 43 in order to generate a predicted HDR image 49. In block 54 the predicted HDR image 49 is combined with the residual image 52 in order to generate the output HDR 56 data. A decoder that operates as shown in Figure 2 can be backward compatible with systems and devices that need LDR 44 data output while providing high quality HDR data on HDR 56 data output.
Methods 10 and 40 can be implemented by:
• programmed data processors, which may comprise one or more of the following: microprocessors, digital signal processors, some combination thereof, or the like running software that causes the data processors to implement the methods;
• hardware circuits, for example circuits that include functional blocks that cooperate to implement the method - the circuits can comprise, for example, field-programmable gate arrays (“FPGAs”) or configured application-specific integrated circuits (“ASICs”) appropriately; or, • carry out some parts of the methods on programmed data processors and other parts of the methods on appropriate hardware circuits.
Figure 3 shows a method 70 according to a more specific exemplary embodiment. Method 70 encodes video frames in a way that meets the standards set by the Moving Picture Experts Group (MPEG) standards. Method 70 receives two video data input streams. At input 72 a stream containing HDR frames 74 is received. At input 78 a stream containing LDR frames 76 is received. LDR frames 76 can be obtained from HDR frames 74 or some earlier precursor of HDR frames 74 from input 78.
An encoder operating as shown in Figure 3 produces three compressed streams: an LDR 80 stream, which can be fully MPEG-compliant; a residual stream 82, containing the differences between the LDR frames 76 and the corresponding HDR frames 74; and an auxiliary stream 84 containing auxiliary data for the reconstruction of the HDR frames 74. The best performance can be achieved when the residual stream 82 and the auxiliary stream 84 do not duplicate the information encoded in the LDR stream 80.
ES 2 551 561 T3
The LDR frames 76 are encoded in block 88 using an appropriate encoder. For example, block 88 may use an ISO / IEC 14496-2 standard compliant MPEG video encoder. Alternatively other video encoders can be used. The resulting video stream can be encapsulated in an appropriate media container format, such as Audio Video Interlaced (AVI) or QuickTime ™, so that it can be recognized and played back by existing software.
In block 90 the MPEG encoded LDR frames are decoded. In order to minimize computation, decoding of block 90 can be performed by the MPEG encoder used in block 88. MPEG encoders typically decode frames internally for use in motion estimation vectors. Block 90 may comprise access to the decoded frames generated by the MPEG encoder. Alternatively, block 90 can be implemented independently of block 88.
The output of block 90 will generally be different from the input of block 88 because MPEG is a lossy compression method. The LDR frames that are encoded with MPEG and then decoded are not exactly the same as the original LDR frames but contain unwanted effects of compression.
In blocks 92A and 92B, if necessary, the color spaces of one or both LDR frames 76 and HDR frames 74 are transformed in order to provide LDR frames and HDR frames that are represented in mutually compatible color spaces. The type of transformations performed on blocks 92A and 92B, if performed, depends on the color spaces of LDR frames 76 and HDR frames 74. In some cases, blocks 92A and 92B are not necessary. In other cases, only block 92A or 92B is necessary.
HDR and LDR color spaces are supported when the color channels of both LDR and HDR color spaces represent approximately the same information. It is also desirable that the HDR and LDR color spaces are perceptually uniform. Perceptual uniformity makes it easy to estimate color differences based on perceivable rather than arithmetic differences. It is also desirable that the HDR color space preserve a wide color gamut, ideally the full gamut of viewable colors, even though the full gamut of viewable colors cannot be displayed on existing displays.
A good color space for use in rendering HDR image data is considered by the inventors to be a combination of the 1976 CIE Uniform Chromaticity Scales (uo, vo) with the gamma correction of the sRGB color space. Other color spaces could also be used. In one example, the input LDR 76 frames are rendered in the sRGB color space while the input HDR 74 frames are rendered in the CIE XYX (2E standard observer) color space. In this case, block 92A comprises converting LDR pixels from sRGB color space to lldruldrVdr space. This can be done by calculating the CIE XYZ color coordinates and then calculating the luma and u 'and v' color coordinates from the XYZ values. XYZ values can be determined using the sRGb conversion formulas provided in IEC 61966-2-1: 1999. Multimedia systems and equipment - Color measurement and management - Part 2-1: Color management - Default RGB color space - sRGB ( Multimedia systems and equipment - Color measurement and management - Part 2-1: Color management - RGB color space by default - sRGB). International Electrotechnical Commission, 1999. For example, for R<sub>8</sub>-bit the 8-bit color coordinate is:
<sup>(1)</sup>
255 (2)
<img file="ES2551561T3_D0001.tif" />
G8-bit and B8-bit color coordinates can also be converted to floating point values, then X, Y, and Z can be determined from:
(3)
<img file="ES2551561T3_D0002.tif" />
The example matrix in Equation (3) assumes the white pixel D65. The luma can be calculated for each of the LDR pixels using the appropriate corrected color values. For example, the luma can be calculated by:
(4) lldr = 0.2126xRe-¿it + 0.7152xG8-bit + 0.0722xB8-bit where: l »is the luma value for an LDR pixel. The luma is the weighted sum of the non-linear components R 'G' B 'after applying the gamma correction.
The chromaticities u 'and v' can be obtained from:
ES 2 551 561 T3 (5) and
(6)
A '+ 15i' + 32 gy
X + 15? + 32
The 8-bit numbers uidr and vidr can then be obtained by multiplying each of the u 'and v' by an appropriate scale factor such as (7)
Uidr = U'x410 <sup>Y</sup> (8) Vidr = V'x410
In the transformed color space, each of the pixels in the LDR data is represented by the pixel values lidr, Vdr, u<sub>Ur</sub>.
Block 92B may transform the color values of the HDR frames 74 in substantially the same way as described above for the values of the LDR pixels. Normal gamma correction cannot normally be used for the range of luminance values that can be specified in an HDR frame. Therefore, some embodiments use a perceptually uniform luminance representation that has been obtained from contrast detection measurements for human observers. This space has properties similar to a space in which gamma correction is performed to the LDR pixel values but the entire visible range of luminance can be encoded (using, for example, 11-12 bits).
In an exemplary embodiment, the HDR luminance, y, is transformed into the 12-bit HDR luma, hdr, by the formula:
(9)
A- = I <sup>b</sup> Y<sup>c</sup> + <sup>d yes</sup> »= £ and <rt
V loeCO + f if y> y<sub>h</sub> where the constants are listed in Table 1 below. The inverse transformation is obtained by:
<sup>(10)</sup>
EyCÍlhdrJ = i'tW + ¿'A if 6 <W <íjí
U'-espCf'-ijur) sil<sub>k</sub>¿<sub>r</sub>> l<sub>h</sub> where the different constants used in Equations (9) and (10) are listed in Table 1 below.
<td colspan="6">TABLE I - Example Constants for Equations (9) and 10)</td>
<td>to</td><td>b</td><td>c</td><td>d</td><td>and</td><td>F</td>
<td> 17,554</td><td> 826,81</td><td> 0,10013</td><td> -884,17</td><td> 209,16</td><td> -731,28</td>
<td>yi</td><td>yh</td><td></td><td></td><td></td><td></td>
<td> 5,6046</td><td> 10469</td><td></td><td></td><td></td><td></td>
<td>to'</td><td>b '</td><td>c '</td><td>d '</td><td>and'</td><td>F</td>
<td> 0,056968</td><td>7.3014e-30</td><td> 9,9872</td><td> 884,17</td><td> 32994</td><td> 0,00478</td>
<td>ii</td><td>ih</td><td></td><td></td><td></td><td></td>
<td> 98,381</td><td> 1204,7</td><td></td><td></td><td></td><td></td>
ES 2 551 561 T3
Block 94 generates a prediction function for the HDR image data. The prediction function attempts to predict a pixel value for a pixel in the HDR image data based on a corresponding pixel value for the corresponding pixel in the LDR image data. Ideally the prediction function is selected to minimize the number of pixels in the HDR image data that have values that differ significantly from the values predicted by the prediction function. The prediction function is preferably non-linear in the logarithmic domain.
In cases where the pixel values representing the chromaticity in the HDR data are practically the same as the corresponding pixel values in the LDR image data, it is not necessary to calculate a prediction function for the chromaticity pixel values (eg u 'and vj. In such cases, it is only necessary to provide a prediction function for the brightness values (eg luma, luminance or the like).
Since LDR frames 76 and HDR frames 74 contain similar information, these frames are strongly correlated. When LDR frames 76 are obtained by applying a tone mapping algorithm to HDR frames 74, the precise nature of the correlation depends on the tone mapping algorithm that has been used.
Figures 4A to 4F show how the luma values of an LDR frame are related to the luma values of the corresponding HDR frame. Each of these Figures is applied to a different tone mapping function to obtain an LDR image from an example HDR image. These tone mapping functions provide, in general, a linear relationship between h<sub>dr</sub> and lhdr for small values. There is greater variation between the tone mapping functions for higher luminance values. In each of Figures 4A to 4D, the LDR luma values are plotted on the horizontal axis and the HDR luma values are plotted on the vertical axis. The points marked with an X indicate the pixel values of the corresponding pixels in the LDR and HDR images.
Figures 4A to 4F correspond respectively to the tone mapping functions disclosed in:
• S. Pattanaik, JE Tumblin, H. Yee and DP Greenberg. Time dependent visual adaptation for realistic image display. In ACM SIGGRAPH 2000 Proceedings, Computer Graphics Proceedings, Annual Conference Series, pages 47-54, July 2000.
• Erik Reinhard, Michael Stark, Peter Shirley and Jim Ferwerda. Photographic tone reproduction for digital images. ACM Publications on Graphics, 21 (3): 267266, 2002.
• Frédo Durand and Julie Dorsey. Fast bilateral filtering for the display of high-dynamic-range images. ACM Publications on Graphics, 21 (3): 257-266, 2002.
• Raanan Fattal, Dani Lischinski and Michael Werman. Gradient domain high dynamic range compression. ACM Publications on Graphics, 21 (3): 249-256, 2002.
• Frédéric Drago, Karol Myszkowski, Thomas Annen and Norishige Chiba. Adaptive logarithmic mapping for displaying high contrast scenes. Computer Graphics Forum, Eurographics Publications 2003, 22 (3): 419-426, 2003.
• Rafal Mantiuk, Karol Myszkowski and Hans-Peter Seidel. A perceptual framework for contrast processing of high dynamic range images. In APGV '05: Publications of the 2nd Symposium on Applied Perception in Graphics and Visualization, pages 87-94, New York, NY, USA, 2005. AcM Press.
The prediction function can be generated as described above. In cases where the prediction function is defined as the arithmetic mean of the values of all the HDR pixels that are in a certain container, the prediction can be performed by:
<img file="ES2551561T3_D0003.tif" />
where Z¡ = {i = 1 ... N * Mi) = l}, l = 0 ... 255;
N is the number of pixels in a frame and lldr (i) and lhdr (I) are the luma values for the i-th pixel of the frames
ES 2 551 561 T3
LDR and HDR, respectively. The prediction function is preferably updated for each of the frames.
In Figures 4A to 4F, the prediction functions are shown as solid lines. The prediction functions will depend on the content of the image as well as the tone mapping function used. Figures 4A to 4F show prediction functions for typical HDR images. Figures 4A to 4F show that typical prediction functions tend to change slowly with a slope increasing over significant parts of their range. Therefore, in some embodiments, instead of encoding the prediction function values for each of the containers, the differences of the prediction function values for two consecutive containers are encoded. These differences can be compressed in order to further reduce the number of bits, for example, using an adaptive Huffman algorithm as indicated in block 95. In some embodiments the size of the auxiliary data stream 84 is 1% or less of the total stream size. Thus, the storage overhead of a prediction function can be practically negligible. Prediction functions or parts of prediction functions can also be represented in other ways, for example, as parameterized polynomial curves, spline curves (basic smoothed polynomials), or other parameterized functions.
In block 96 the residual frames are calculated. Each of the pixel values in the residual frame represents the difference between the pixel value for the corresponding pixel of the HDR frame and the pixel value for the predicted pixel by applying the prediction function to the pixel value of the pixel. corresponding part of the LDR frame. Block 96 can be executed independently for each of the pixel values (l, u, and v in this example). For the luminance values, each of the pixels r (i) of the residual raster can be calculated by (12) Γι (ί) = lhdr (i) -RF (lidr (i)) for the chromatic values, the function Prediction can be an identity function, in which case:
(13) ru (i) = Uhdr (i) -U | d<sub>r</sub>(i) <sup>Y</sup> (14) r<sub>v</sub>(i) = Vhdr (i) -Vldi ^ i)
An appropriately selected prediction function can significantly reduce the amount of data that HDR frames encode. Regardless of these savings, residual frames can still contain a significant amount of noise that does not visibly improve the quality of reconstructed HDR images. The compression ratio can be improved without causing an appreciable reduction in image quality by filtering out residual frames in order to reduce or eliminate this noise. Block 98 filters the residual frames. The signal in residual frames is often relatively close to the visibility threshold. Therefore, filtering can result in significant data reduction without significant degradation in the quality of HDR images reconstructed from the data.
An output from block 98 is a residual frame in which the high frequencies have been attenuated in those regions where they are not visible. Figure 5 shows a method 110 that can be applied to filter out residual frames. The method 110 can be put into practice in the context of a coding method according to the invention but it can also be applied in other contexts in which it is desired to reduce the amount of data that represents an image without introducing visible effects in it. not wanted.
The description made below describes the processing carried out on a luma channel. The same processing can also be applied to chroma channels. In order to reduce processing, chroma channels can be subsampled, for example, to half their original resolution. This reduction roughly explains the differences in luminance and chrominance CSF.
The method 110 receives a residual frame 112 and an HDR frame 114 that masks the residual frame. In blocks 116 and 118 a Discrete Wavelet Transform (DWT) is applied to divide the masking frame 114 and the residual frame 112 into several frequency and orientation selective channels. Instead of DWT other appropriate transforms can be used, such as, for example, the cortex transform described in AB Watson. The cortex transform: rapid computation of simulated neural images. Computer Vision Graphics and Image Processing, 39: 311-327, 1987. Cortex transform can be very computationally intensive so it is only practical if sufficient computational resources are available.
One embodiment of a prototype is based on the CDF 9/7 discrete wavelet (which is also used for lossy image compression according to the JPEG-2000 standard). The wavelet base offers a good relationship between smoothness and computational efficiency. In the prototype, only the three finer scales of wavelet decomposition are used since the filtering of smaller spatial frequencies in coarser scales
ES 2 551 561 T3 could cause noticeable unwanted effects.
At block 120 a function such as a contrast sensitivity function (CSF) is applied to respond to the lower sensitivity of the human visual system for high spatial frequencies. The application of the CSF involves weighting with a constant value each of the bands of the wavelet coefficients. Table 2 provides examples of weighting factors for an observation distance of 1700 pixels.
<td colspan="4">TABLE 2 - CSF coefficients</td>
<td>Scale</td><td>LH</td><td>HL</td><td>H H</td>
<td> 1</td><td> 0,275783</td><td> 0,275783</td><td> 0,090078</td>
<td> 2</td><td> 0,837755</td><td> 0,837755</td><td> 0,701837</td>
<td> 3</td><td> 0,999994</td><td> 0,999994</td><td> 0,999988</td>
Human vision channels have limited phase sensitivity. This provides an additional opportunity to discard information without obtaining noticeable degradation of the reconstructed images. A masking signal does not only affect the regions in which the wavelet coefficient values are the highest, but it also affects the neighboring regions. A phase uncertainty also reduces the masking effect at borders, unlike textures that show higher amounts of masking.
Phase uncertainty can be modeled with the L0.2 standard, which is also used in JPEG-2000 image compression. The L0.2 norm is given by:
(15)
<img file="ES2551561T3_D0004.tif" />
and its mathematical equivalents where l represents the environment of a coefficient (in the prototype implementation a 13H13 box is used as the environment), Lcsf is a wavelet coefficient that has been weighted by applying a CSF factor and - is the CSF-weighted wavelet coefficient after taking phase uncertainty into account.
Block 124 predicts how the threshold contrast changes in the presence of the masking signal from the original HDR frame 114. In order to model the masking of the contrast, a threshold elevation function can be used. The threshold elevation function can, for example, take the form:
(16)
<img file="ES2551561T3_D0005.tif" />
In the prototype embodiment, the constants of Equation (16) are a = 0.093071, b = 1.0299 and c = 11.535.
Each of the CSF-weighted coefficients for the residual frame, RCSF, is compared to the corresponding threshold rise value Te calculated from the original HDR frame 114. If Rcsf is smaller than the threshold elevation Te of Equation (16), the coefficient can be assigned the value zero without introducing any perceptible changes in the eventual reconstructed image. This can be expressed by:
(17) p _ _ _ ίθ YES 1 R csf <sup>jT:</sup>'<sup>L</sup> L? otherwise
Finally, the filtered wavelet coefficients, Rfilt, are transformed back to the image domain. The pre-filtering method presented above can substantially reduce the size of the waste stream. Filtering is a reasonable balance between computational efficiency and visual model accuracy. Filtering such as that described in the present application typically increases encoding time by no more than about 80%. Filtering during encoding does not increase decoding times.
Returning to Figure 3, block 100 quantizes the filtered residual frames. Although the magnitudes of the differences encoded in the residual frames are normally small, they can take on values in the range of! 4095 to 4095 (for a 12-bit HDR luma encoding). Obviously, such values cannot be encoded using an 8-bit MPEG encoder. Although the MPEG standard provides an extension to encode luma values in 12 bits, such an extension is infrequently implemented, especially in hardware.
The quantization block 100 makes it possible to reduce the magnitude of the residual values, preferably in a
ES 2 551 561 T3 sufficient so that said values can be encoded using a standard MPEG 8-bit encoder. Various quantification schemes can be used. For example, some embodiments apply non-linear quantization, in which large absolute values of the residual are strongly quantized, while small values are held with the highest precision. Since there are very few pixels that contain a residue with a large magnitude, most of the pixels are not affected by the strong quantization.
Strong quantization can cause some images to have poor visual quality. This is because even if few pixels have large quantization errors, they can protrude in a way that decreases perceived image quality.
Simply narrowing down residual values (for example to an 8-bit range) can produce visually better results at the cost of losing detail in very bright or dark regions. Furthermore, in typical images, with appropriately chosen prediction functions, only a few pixels have residual values that exceed an 8-bit range.
In some embodiments, in order to reduce the cost limitation of stronger quantization, the residual values are divided by a constant quantization factor. The factor can be chosen based on a balance between errors due to dimensioning and errors due to quantization. Said quantization factors can be established separately for each of the containers, depending on the maximum magnitude of the residue of all the pixels belonging to said container. Therefore, residual values after quantification can be calculated by:
<img file="ES2551561T3_D0006.tif" />
where:
• The operator [·]<sup>-127+127</sup> rounds the value inside the brackets to the nearest whole number and then limits the value if it is greater than 127 or less than -127;
• q (l) is a quantification factor that is selected separately for each of the Σκ containers.
The quantization factor is given by (19) q (l) = max q<sub>min</sub>, <sup>max</sup>^, (k / (Otr
127 where what<sub>m</sub>¡N is a minimum quantization factor that can be, for example, 1 or 2.
The quantization factors q (l) can be stored together with the prediction function in the auxiliary data stream 84. This data can be compressed first as in block 95. In most cases, most of the quantification factors q (l) will have the value q<sub>m</sub>¡N · Thus, string length encoding followed by Huffman encoding is an effective way to compress the data representing quantization factors.
In block 102 the residual values are encoded. When the residual values are 8-bit values, they can be encoded using normal MPEG compression (eg MPEG-4 compression). In one embodiment of the prototype, the quantized residuals, Π, and the residuals of chroma ru and r<sub>v</sub> they are encoded with MPEG after rounding to the nearest integer value. Bear in mind that the operations applied in order to obtain the residual values are approximately linear in the cases in which the prediction function is almost linear and the effect of the adaptive quantization of Equation (18) is minimal. In such cases, the visual information of a residual frame is in the same frequency bands as the original HDR frame, and the DCT quantization of the residual has a similar effect as for the original HDR pixel values. Therefore, a standard DCT quantization matrix can be used to encode the residual frames.
Since the MPEG encoding in blocks 88 and 102 is independent, it is possible to separately configure the MPEG quality parameters for each of blocks 88 and 102. In most applications, it is neither intuitive nor appropriate to configure two sets of MPEG quality parameters. In preferred embodiments, a single quality control sets the quality parameters for both blocks 88 and 102. It has been found that, in general, setting the quality parameters in blocks 88 and 102 to be equal to each other provides satisfactory results.
Some quality settings for blocks 88 and 102 provide better compression results than
ES 2 551 561 T3 others. In order to achieve the best quality HDR images, block 102 should comprise an encoding that uses the best quality. The quality settings in block 88 mainly affect the quality of LDR images reconstructed from stream 80 but may also have some impact on HDR images.
Some embodiments of the invention use the fact that both LDR and HDR frames contain the same scenes. In this way the optical flow should be the same for both. In these embodiments, the same motion vectors are used for the residual frames as those calculated for the LDR frames. Data structure 38 can include only one set of motion vectors. In alternative embodiments of the invention, the motion vectors are calculated separately for the LDR and residual frames and both sets of motion vectors are stored in the encoded image data 38.
The software for practicing the methods according to the invention can be implemented in various ways. In one embodiment of the prototype, the software is implemented as a dynamic library in order to simplify integration with external software. A separate set of command line tools enables encoding and decoding of video streams to and from HDR image files.
Since HDR video playback involves decoding two MPEG 80 and 82 streams, achieving an acceptable frame rate is more challenging than normal LDR video playback. The playback frame rate can be increased by executing some parts of the decoding process using graphics hardware. For example, both color space conversion and color channel oversampling can be costly in computational power when running on a CPU and yet can be run extremely efficiently on a graphics processor (GPU) such as modules. Program. Additionally, some color conversion functions can be significantly sped up by using fixed-point arithmetic and lookup tables.
Figure 6 illustrates the performance of the prototype embodiment as a function of quality settings. The lower points correspond to the LDR 80 flow while the upper points correspond to the sum of the LDR 80 flow and the residual flow 82. It can be seen that for the lower values of the quality parameter qscale (that is, for higher quality images) the percentage of the overall data stream composed of the residual stream 82 is lower than for the higher values of the quality parameter (corresponding to lower quality LDR images).
Codecs as described in the present application can be used to encode and decode both individual images and video sequences. Such codecs can be used to encode and decode movies to be stored on media such as DVD or other storage media that may be common in the future for storing movies.
Some aspects of the invention provide media players that include an output for HDR images to which an HDR display device is or can be connected. Media players include hardware, software, or a combination of hardware and software that implement decoding methods such as those shown, for example, in Figure 2.
Some implementations of the invention comprise computer processors that execute software instructions that cause the processors to implement a method of the invention. For example, one or more processors in a data processing system may implement the encoding methods of Figures 1 or 3 or the decoding method of Figure 2 by executing software instructions stored in memory accessible to the processors. . The invention can also be provided in the form of a product in the form of a program. The product in program form may comprise any medium that includes a set of computer-readable signals comprising instructions, which, when executed by a data processor, result in the data processor executing a method of the invention. . Products in program form according to the invention can be in any of a wide variety of forms. The product in program form may comprise, for example, a physical medium such as, for example, a magnetic data storage medium including floppy disks, hard drives, optical data storage media including CD ROM, DVD, electronic storage media. data including ROM, flash, RAM, or the like. The computer-readable signals of the product in program form may alternatively be compressed or encrypted.
Unless otherwise indicated, when a component (for example a software module, a processor, an assembly, a device, a circuit, etc.) is mentioned above, the reference to said component (including a reference to "media ”) Should be interpreted as including as equivalents of said component any component that performs the function of the described component (that is, it is functionally equivalent), including components that are not structurally equivalent to the disclosed structure that performs the function in the exemplary embodiments illustrated in the invention.
ES 2 551 561 T3
Additional embodiments of the invention:
1. A method for encoding a high dynamic range image, comprising the method:
obtain an image with a lower dynamic range that corresponds to the image with a high dynamic range;
identifying groups of pixels in the high dynamic range image for which all corresponding pixels in the lower dynamic range image have the same pixel value;
generating a prediction function based at least partially on pixel values of the pixels in the high dynamic range image belonging to each of a plurality of the groups;
apply the prediction function to the image of lower dynamic range in order to obtain a predicted image;
calculating a residual image representing the differences between the pixel values in the predicted image and the corresponding pixel values in the high dynamic range image; and, encoding and storing the data representing the image with the lowest dynamic range, the prediction function, and the residual image.
two. One method in which obtaining the lower dynamic range image comprises encoding the lower dynamic range image and decoding the lower dynamic range image.
3. A method comprising transforming the high dynamic range image, the lower dynamic range image, or both the high dynamic range image and the lower dynamic range image between color spaces before setting the prediction function.
Four. A method in which immediately before generating the prediction function, both the lower dynamic range image and the high dynamic range image are expressed in color spaces that include a luma or luminance value of the pixels and two or more values chromaticity of the pixels.
5. A method in which the lower dynamic range image is represented in a color space comprising one pixel intensity value and two or more pixel chroma values.
6. A method in which the prediction function is nonlinear in the logarithmic domain.
7. A method in which the generation of the prediction function comprises calculating an arithmetic mean of the pixel values of the pixels in the high dynamic range image belonging to each of the plurality of groups.
8. A method in which the generation of the prediction function comprises calculating an average of the pixel values of the pixels in the high dynamic range image belonging to each of the plurality of groups.
9. A method generating the prediction function comprising calculating an average of the highest and lowest pixel values of the pixel values of the pixels in the high dynamic range image belonging to each of the plurality of groups.
10. A method in which the generation of the prediction function comprises one or more of the following:
calculating an arithmetic mean of the pixel values of the pixels in the high dynamic range image belonging to each of the plurality of groups;
calculating a median of the pixel values of the pixels in the high dynamic range image belonging to each of the plurality of groups;
calculating an average of the highest and lowest pixel values of the pixel values of the pixels in the high dynamic range image belonging to each of the plurality of groups;
calculating a centroid of a subset of the pixel values of the pixels in the high dynamic range image belonging to each of the plurality of groups; and, combinations thereof.
eleven. A method in which the lower dynamic range image and the high dynamic range image comprise
ES 2 551 561 T3 each frame in a video sequence.
12. A method comprising generating a new prediction function for each of the frames in the video sequence.
13. A method comprising monitoring a difference between successive frames in the video sequence and generating a new prediction function each time the difference indicates that the current frame is significantly different from a previous frame.
14. A method in which the pixel values are pixel intensity values.
fifteen. A method in which the intensity values of the pixels comprise luminance values, luma values or radiance values.
16. A method comprising generating a chroma prediction function for each one or more chroma values and for each of the one or more chroma values:
applying the chroma prediction function corresponding to the corresponding chroma values for the pixels in the lower dynamic range image in order to obtain a predicted image;
calculating a residual chroma image that represents the differences between the chroma values for the pixels in the predicted and high dynamic range images; and, encoding and storing the data representing the chroma prediction functions and the chroma residual images.
17. A method comprising subsampling the residual image before storing the data representing the residual image.
18. A method comprising filtering the residual image in order to remove the noise before storing the data representing the residual image.
19. A method in which the filtering of the residual image comprises:
applying a discrete wavelet transform to the residual image and the high dynamic range image to obtain a transformed residual image and a transformed high dynamic range image;
setting threshold values for the coefficients in the transformed residual image based on the values of the coefficients in the transformed high dynamic range image; and, assigning the value zero to the coefficients in the transformed residual image if the coefficients have values that do not exceed the corresponding thresholds.
twenty. A method in which setting the threshold values comprises applying a threshold raising function to the coefficients in the transformed high dynamic range image.
twenty-one. A method in which the threshold raising function comprises raising the coefficients to a predetermined constant power.
22. A method in which the threshold raising function comprises multiplying the coefficients by a predetermined constant number.
2. 3. A method according to any one of claims 20 to 22 in which the lift function is given by:
yes <ίΐ otherwise or a mathematical equivalent of it.
24. A method comprising applying a predetermined contrast sensitivity function that weights the coefficients of the transformed high dynamic range image factors before applying the threshold raising function to the coefficients.
25. A method comprising applying a phase uncertainty function to the coefficients of the transformed high dynamic range image before applying the threshold raising function to the coefficients.
26. A method in which the phase uncertainty function is given by:
<img file="ES2551561T3_D0007.tif" />
ES 2 551 561 T3
<img file="ES2551561T3_D0008.tif" />
or a mathematical equivalent thereof, where l represents the neighborhood of a coefficient, Lcsf is a wavelet coefficient and is the wavelet coefficient after applying the phase uncertainty function.
27. A device for encoding a high dynamic range image, the device comprising a data processor that executes instructions that make the data processor:
Obtain an image with a lower dynamic range than that of a high dynamic range image;
identify groups of pixels in the high dynamic range image for which all corresponding pixels in the lower dynamic range image have the same pixel value;
generate a prediction function based at least partially on pixel values of the pixels in the high dynamic range image belonging to each of a plurality of the groups;
apply the prediction function to the image of lower dynamic range in order to get a predicted image;
calculate a residual image that represents the differences between pixel values in the predicted image and corresponding pixel values in the high dynamic range image; and, encode and store the data representing the lower dynamic range image, the prediction function, and the residual image.
28. A device for encoding a high dynamic range image, the device comprising:
means for obtaining a lower dynamic range image corresponding to the high dynamic range image;
means for identifying groups of pixels in the high dynamic range image for which all corresponding pixels in the lower dynamic range image have the same pixel value;
means for generating a prediction function based at least partially on pixel values of the pixels in the high dynamic range image belonging to each of a plurality of the groups;
means for applying the prediction function to the image of lower dynamic range in order to obtain a predicted image;
means for calculating a residual image representing the differences between pixel values in the predicted image and corresponding pixel values in the high dynamic range image; and, means for encoding and storing the data representing the lower dynamic range image, the prediction function and the residual image.
29. A device for decoding a high dynamic range image, the device comprising a data processor that executes instructions that make the data processor:
retrieve the data representing a lower dynamic range image corresponding to the high dynamic range image, a prediction function and a residual image;
Apply the prediction function to the lower dynamic range image in order to obtain a predicted high dynamic range image; and, combine the residual image with the intended high dynamic range image to obtain the high dynamic range image.
30. A device for decoding a high dynamic range image, the device comprising:
means for retrieving the data representing a lower dynamic range image corresponding to the high dynamic range image, a prediction function, and a residual image;
ES 2 551 561 T3 means for applying the prediction function to the lower dynamic range image in order to obtain a predicted high dynamic range image; and, means for combining the residual image with the intended high dynamic range image in order to obtain the high dynamic range image.
Contents9
42 members in 8 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 761510P | United States of America | – | |
| 76151006 | United States of America | P |
Members42
| Document | Office | Kind | |
|---|---|---|---|
| WO2007082562A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2007082562A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1989882A2 | European Patent Office (EPO) | A2 | |
| KR20080107389A | Republic of Korea | A | |
| CN101371583A | China | A | |
| JP2009524371A | Japan | A | |
| HK1129181A1 | Hong Kong, China | A1 | |
| CN101742306A | China | A | |
| US2010172411A1 | United States of America | A1 | |
| EP2290983A2 | European Patent Office (EPO) | A2 | |
| EP2320653A2 | European Patent Office (EPO) | A2 | |
| CN101371583B | China | B | |
| EP2290983A3 | European Patent Office (EPO) | A3 | |
| EP2320653A3 | European Patent Office (EPO) | A3 | |
| JP5249784B2 | Japan | B2 | |
| JP2013153508A | Japan | A | |
| US8537893B2 | United States of America | B2 | |
| US2013322532A1 | United States of America | A1 | |
| US8611421B1 | United States of America | B1 | |
| KR101356548B1 | Republic of Korea | B1 | |
| US2014086321A1 | United States of America | A1 | |
| JP5558603B2 | Japan | B2 | |
| US8989267B2 | United States of America | B2 | |
| US2015156506A1 | United States of America | A1 | |
| EP2290983B1 | European Patent Office (EPO) | B1 | |
| EP2320653B1 | European Patent Office (EPO) | B1 | |
| EP1989882B1 | European Patent Office (EPO) | B1 | |
| ES2551561T3This record | Spain | T3 | |
| ES2551562T3 | Spain | T3 | |
| US9210439B2 | United States of America | B2 | |
| EP2988499A1 | European Patent Office (EPO) | A1 | |
| US2016119638A1 | United States of America | A1 | |
| US9544610B2 | United States of America | B2 | |
| US2017041626A1 | United States of America | A1 | |
| EP2988499B1 | European Patent Office (EPO) | B1 | |
| EP3197157A1 | European Patent Office (EPO) | A1 | |
| US9894374B2 | United States of America | B2 | |
| US2018103263A1 | United States of America | A1 | |
| US10165297B2 | United States of America | B2 | |
| US2019052892A1 | United States of America | A1 | |
| US10931961B2 | United States of America | B2 | |
| EP3197157B1 | European Patent Office (EPO) | B1 |
Numbers
- Publication
- 2551561
- Application
- 10185996
Titles2
- Spanish
- Codecs de alto rango dinámico
- English
- High dynamic range codecs
Classification
- CPC, 12
- H04N19/105
- H04N19/85
- H04N19/50
- H04N19/63
- H04N19/186
- H04N19/184
- H04N19/187
- H04N19/33
- H04N19/98
- H04N19/59
- H04N19/136
- H04N19/124
- IPC, 3
- H04N19 105
- H04N19 50
- H04N19 63