Piecewise cross color channel predictor
Abstract
A sequence of visual dynamic range (VDR) images may be encoded using a standard dynamic range (SDR) base layer and one or more enhancement layers. A prediction image is generated by using piecewise cross-color channel prediction (PCCC), wherein a color channel in the SDR input may be segmented into two or more color channel segments and each segment is assigned its own cross-color channel predictor to derive a predicted output VDR image. PCCC prediction models may include first order, second order, or higher order parameters. Using a minimum mean-square error criterion, a closed form solution is presented for the prediction parameters for a second-order PCCC model. Algorithms for segmenting the color channels into multiple color channel segments are also presented.

Term
6.3 yearsto projected expiry
Projected expiry 23 January 2033, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
14 claims: 10 independent, 4 dependent
- 1Patent claims Zastrzeżenia patentowe 1. A method comprising:1. Sposób obejmujący: accessing a first image and a second image representing the same scene, each image containing one or more color channels, each image containing a plurality of pixels, each pixel having a corresponding pixel value for each of the one or more color channels the second image having a dynamic range greater than the dynamic range of the first image;uzyskiwanie dostępu do pierwszego obrazu i drugiego obrazu przedstawiających tę samą scenę, przy czym każdy z obrazów zawiera jeden lub większą liczbę kanałów koloru, każdy z obrazów zawiera wiele pikseli, przy czym każdy piksel ma odpowiednią wartość piksela dla każdego z jednego lub większej liczby kanałów koloru, przy czym drugi obraz ma zakres dynamiki większy niż zakres dynamiki pierwszego obrazu;segmenting at least one color channel of the first image into two or more non-overlapping color channel segments by a set of cutoff points, where each color channel segment corresponds to two consecutive cutoff points, and where the pixel values of the color channel that are between two consecutive cutoff points. are assigned to the corresponding segment of the color channel;and for the first image color channel segment: segmentowanie co najmniej jednego kanału koloru pierwszego obrazu na dwa lub większą liczbę nienakładających się segmentów kanału koloru za pomocą zbioru punktów granicznych, gdzie każdy segment kanału koloru odpowiada dwóm kolejnym punktom granicznym, i gdzie wartości pikseli kanału koloru, które są pomiędzy dwoma kolejnymi punktami granicznymi, są przypisywane odpowiedniemu segmentowi kanału koloru;i dla segmentu kanału koloru pierwszego obrazu: Determining a first-order color transmission prediction model that expresses a predicted pixel value for a pixel of a second image in one color channel as a combination of at least the corresponding pixel values for all pixel color channels within the first image having the same pixel coordinates as the pixel of the second image, with which the prediction model of color penetration contains the matrix of the prediction parameter, which converts the input vector to an output channel segment, the input vector including, for each pixel of the first image color channel segment, an input vector component of at least a first order, the first order input vector components comprising the products of the respective pixel values for two or more color channels a pixel within the first image, and wherein the output channel segment includes, for each pixel of the first image color channel segment, a predicted pixel value for the second image pixel in one color channel having the same pixel coordinates as the pixel of the first image;określanie modelu predykcji przenikania kolorów co najmniej pierwszego rzędu, który wyraża prognozowaną wartość piksela dla piksela drugiego obrazu w jednym kanale koloru jako kombinację co najmniej odpowiednich wartości pikseli dla wszystkich kanałów koloru piksela w obrębie pierwszego obrazu mającego te same współrzędne piksela co piksel drugiego obrazu, przy czym model predykcji przenikania kolorów zawiera macierz parametru predykcji, która przekształca wektor wejściowy na wyjściowy segment kanału, przy czym wektor wejściowy zawiera, dla każdego piksela segmentu kanału koloru pierwszego obrazu, składową wektora wejściowego co najmniej pierwszego rzędu, gdzie składowe wektora wejściowego pierwszego rzędu zawierają iloczyny odpowiednich wartości pikseli dla dwóch lub większej liczby kanałów koloru piksela w obrębie pierwszego obrazu, i przy czym wyjściowy segment kanału zawiera, dla każdego piksela segmentu kanału koloru pierwszego obrazu, prognozowaną wartość piksela dla piksela drugiego obrazu w jednym kanale koloru, mającego te same współrzędne piksela co piksel pierwszego obrazu;computing the parameters of the prediction parameter matrix by minimizing the mean square error between the predicted pixel values of the output color channel segment and corresponding pixel values of the second image;obliczanie parametrów macierzy parametru predykcji przez zminimalizowanie błędu średniokwadratowego między prognozowanymi wartościami pikseli wyjściowego segmentu kanału koloru a odpowiednimi wartościami pikseli drugiego obrazu;computing an output color channel segment by transformation obliczanie wyjściowego segmentu kanału koloru przez przekształcenie -17wektora wejściowego za pomocą macierzy parametru predykcji;i podawanie parametrów macierzy parametru predykcji do użycia przez dekoder. Input vector using a prediction parameter matrix;and supplying matrix parameters of the prediction parameter to be used by the decoder.
- 4A method according to any of the claims 1-3, wherein the first image is a standard dynamic range (SDR) image and the second image has a dynamic range extending at least 14-15 orders of magnitude. 4. Sposób według dowolnego z zastrz. 1-3, w którym pierwszy obraz jest obrazem o standardowym zakresie dynamiki (SDR), a drugi obraz ma zakres dynamiki rozciągający się na co najmniej 14-15 rzędów wielkości.
- 5A method according to any of the claims Wherein the first image is a first SDR image in a sequence of SDR images including a second, different SDR image, the method further comprising:5. Sposób według dowolnego z zastrz. 1-3, w którym pierwszy obraz jest pierwszym obrazem SDR w sekwencji obrazów SDR, zawierającej drugi, inny obraz SDR, który to sposób ponadto obejmuje: performing a two-step search algorithm to identify a first cutoff point for a color channel segment in the first SDR image;and using the first cutoff point as the starting point in the second step of the two-step search algorithm to identify the cutoff point of the color channel segment in the second SDR image. wykonywanie dwuetapowego algorytmu wyszukiwania dla zidentyfikowania pierwszego punktu granicznego dla segmentu kanału koloru w pierwszym obrazie SDR;i stosowanie pierwszego punktu granicznego jako punktu początkowego w drugim etapie dwuetapowego algorytmu wyszukiwania dla zidentyfikowania punktu granicznego segmentu kanału koloru w drugim obrazie SDR.
- 9The method according to p. 8, including:9. Sposób według zastrz. 8, obejmujący ponadto: compressing the first image into an encoded base layer signal;and compressing the image derived from the second image and the projected image into one or more encoded enhancement layer signals;and kompresowanie pierwszego obrazu do zakodowanego sygnału warstwy bazowej;i kompresowanie obrazu wyprowadzonego z drugiego obrazu i obrazu prognozowanego do jednego lub większej liczby zakodowanych sygnałów warstwy wzmocnienia;i -18gdzie parametry macierzy parametru predykcji są przesyłane do dekodera jako metadane. Where the parameters of the prediction parameter matrix are transmitted to the decoder as metadata.
- 10A method according to any of the claims 1-9, in which minimizing the mean square error also includes:10. Sposób według dowolnego z zastrz. 1-9, w którym minimalizowanie błędu średniokwadratowego obejmuje ponadto: applying numerical methods that minimize the mean square error between the predicted pixel values of the output color channel segment and the corresponding pixel values of the second image. zastosowanie metod numerycznych, które minimalizują błąd średniokwadratowy między prognozowanymi wartościami pikseli wyjściowego segmentu kanału koloru a odpowiednimi wartościami pikseli drugiego obrazu.
- 11An image decoding method, comprising:11. Sposób dekodowania obrazu, obejmujący: accessing metadata containing data for a prediction model, wherein the metadata is generated and communicated with the method of claim 9;uzyskiwanie dostępu do metadanych zawierających dane dla modelu predykcji, gdzie metadane są generowane i przesyłane sposobem według zastrz. 9;decompressing the base layer signal to obtain a decompressed image;and generating an output color channel segment from the decompressed image data and data for the specified color diffusion prediction model. dekompresję sygnału warstwy bazowej dla otrzymania obrazu zdekompresowanego;i generowanie wyjściowego segmentu kanału koloru na podstawie danych obrazu zdekompresowanego i danych dla określonego modelu predykcji przenikania kolorów.
- 12The method according to p. 11, further comprising computing an output forecast image including the output color channel segment; and further including:12. Sposób według zastrz. 11, obejmujący ponadto obliczanie wyjściowego obrazu prognozowanego zawierającego wyjściowy segment kanału koloru;i obejmujący ponadto: accessing the residual image;uzyskiwanie dostępu do obrazu resztkowego;combining the residual image and the output forecast image to generate a decoded image, wherein the decoded image has a greater dynamic range than the dynamic range of the first image. łączenie obrazu resztkowego i wyjściowego obrazu prognozowanego dla generowania obrazu zdekodowanego, gdzie obraz zdekodowany ma większy zakres dynamiki niż zakres dynamiki pierwszego obrazu.
Independent claims10
128 paragraphs in 11 sections, as filed
Description
REFERENCE TO RELATED NOTIFICATIONS
[0001] The present disclosure may also be related to U.S. Provisional Application No. 61 / 475,359, filed April 14, 2011, entitled "Multiple color channel multiple regression predictor" which was also filed as PCT Application No. PCT / US2012 / 033605 April 13, 2012 year. This application claims priority to U.S. Provisional Patent Application No. 61 / 590,175, filed January 24, 2012.
TECHNOLOGY
The present invention relates generally to images. More specifically, an embodiment of the invention relates to a fragmentary color cross-pass channel predictor of high dynamic range images by means of standard dynamic range images.
BACKGROUND
[0003] As used herein, the term "dynamic range" (DR) can refer to the ability of the human psychovisual system (HVS) to perceive a certain range of intensity in an image (e.g., luminance, luma). e.g. from the darkest darkness to the brightest brightness. In this sense, DR refers to "scene related" intensity. DR may also relate to the ability of the display device to adequately or approximate to visualize an intensity range with a certain width. In this sense, DR refers to "display related" intensity. Until the meanings of a particular meaning are explicitly stated anywhere in this specification, it is inferred that the term may be used with any of these meanings, e.g., interchangeably.
[0004] As used herein, the term "high dynamic range" (HDR) refers to a DR width spanning approximately 14-15 orders of magnitude of the human visual system (HVS). For example, well-adjusted humans of substantially normal (e.g., one or more of a statistical, biometric, or ophthalmic sense) have an intensity range that extends over about 15 orders of magnitude. Adapted people can perceive dark light sources that emit only a handful of photons. But these same people can perceive the almost painful intensity of solar noon brightness in the desert, at sea, or in the snow (or even gaze at the sun, albeit briefly to avoid damage). However, this range is available to "adjusted" people, such as those for whom HVS has had time to reset and adjust.
[0005] In contrast, a DR in which a human can simultaneously perceive a large range of intensity may be somewhat trimmed relative to HDR. In meaning
As used herein, the terms "visual dynamie range" or "variable dynamie range" (VDR) may separately or alternately refer to a DR that HVS can perceive simultaneously. As used herein, a VDR can refer to a DR that extends over 5-6 orders of magnitude. So, while maybe the VDR is a bit narrower than the real HDR scene, it still constitutes a wide DR. As used herein, the term simultaneous dynamie range may refer to VDR.
[0006] Until relatively recently, images had a DR much narrower than HDR or VDR. Television (TV) and computer monitor equipment using a conventional eathode ray tube (CRT), liquid erystal display (LCD) with solid white fluorescent backlight, or plasma screen technology may be limited in terms of DR visualization capabilities up to about three orders of magnitude. Thus, such conventional displays typically have a low dynamic range (LDR), also referred to as standard dynamic range (SDR), relative to VDR and HDR.
[0007] However, advances in the underlying technology make it possible for displays with a more modern design to visualize the content of images and videos with significant improvements in various quality characteristics of the same content than when visualized on less modern displays. For example, more modern display devices may be capable of visualizing high definition content. high definition, HD) and / or content that can be scaled down according to various display capabilities such as an image ratio. In addition, some modern displays are able to visualize content with a DR greater than that of conventional displays.
[0008] For example, some modern LCD displays have a baeklight unit (BLU) having a light emitting diode (LED) array. The BLU LEDs may be modulated separately from the modulation of the polarization states of the active LCD elements. This dual modulation approach is extendable (e.g. on N modulation layers where N comprises an integer greater than two), as in the case of controllable intermediate layers between BLU and LCD screen elements. leh BLU based on the LED chip and the dual (or N-fold) module effectively increases the DR of the display of LCD monitors having such eeehy.
[0009] Such "HDR displays" as they are often called (although in fact their capabilities may be closer to the VDR range) and the DR extension that they can perform over conventional SDR displays represent a significant advance in the ability to display images. , video content and other visual information. The range of colors that such an HDR display can visualize can also extend well beyond the color gamut of more conventional displays, even to being able to visualize a wide eolor gamut (WCG). With "HDR" displays (henceforth referred to as "HDR displays"), it is now possible to more faithfully and more effectively display HDR or VDR scene-related and WCG image content such as that which can be generated by "next generation" movie and TV cameras. As in case of
-3 scaled video coding and HDTV technology, the DR widening of the image usually involves the bifurcation approach. For example, scene-related HDR content captured with a modern HDR capable camera can be used to generate an SDR version of that content that can be displayed on conventional SDR displays. In one approach, generating an SDR version from a registered VDR version may involve applying a global tone mapping operator (TMO) to the intensity (e.g., luminance, luma) of the associated pixel values in the HDR content. In the second approach, described in patent application PCT / US2011 / 048861 "Extending Image Dynamic Range" by W. Gish et al., Generating an SDR image may involve applying a reversible operator (or predictor) to the VDR data. From the standpoint of saving bandwidth or otherwise, transmitting both the actual recorded VDR content and the corresponding SDR version may not be the best approach.
[0010] Thus, an inverse color mapping generator (iTMO) inverse of the original TMO or an inverse operator of the original predictor may be applied to the generated SDR content, which enables the prediction of the version of the VDR content. The projected version of the VDR content can be compared to the originally recorded HDR content. For example, subtracting the projected VDR version from the original VDR version can generate a residual image. The encoder may transmit the generated SDR content as a base layer (BL), and package the generated SDR content version, an optional residual image and iTMO or other predictors as an enhancement layer (EL) or metadata.
[0011] Sending EL and metadata, along with SDR content, residual and predictors, into bitstreams typically takes less bandwidth than that would be occupied if both HDR and SDR content was sent directly to the bitstream. Compatible decoders that will receive the bitstream sent by the encoder can decode and visualize SDR on conventional displays. However, compatible decoders may also use residual image, iTMO predictors, or metadata to compute a predicted version of HDR content from this, for use with higher capacity displays. It is an object of the invention to provide new predictor generation methods that enable efficient encoding, transmission and decoding of VDR data using the corresponding SDR data.
[0012] In WO 2010/105036 the HDR pixel luma values are derived as a function of the respective LDR pixel luma values, the function parameters depending on the LDR pixel luma value. However, WO 2010/102036 does not consider the use of other color channels when deriving the HDR pixel values of the luma component.
[0013] The approaches described in this section are those that could be pursued, but need not be those that have been conceived or have been pursued previously. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section are prior art solely for inclusion in this section. Likewise, it should not be assumed on the basis of this section that the problems identified for one or more of the approaches have already been identified
Recognized in any art, unless otherwise stated.
[0014] The invention is set out in the appended set of claims.
BRIEF DESCRIPTION OF THE DRAWING FIGURES
[0015] An embodiment of the invention is shown by way of translation, and not as a limitation, in the figures of the accompanying drawing, in which like numerals refer to like elements and wherein:
Fig. 1 shows an exemplary data flow for a VDR-SDR system according to an embodiment of the invention;
Fig. 2 shows a VDR encoding system according to an embodiment of the invention;
Fig. 3 shows an exemplary process for fragment color penetration channel prediction according to an embodiment of the present invention;
Fig. 4 shows an exemplary image decoder with a predictor operating in accordance with embodiments of the present invention.
DESCRIPTION OF DEMO EXAMPLES
[0016] Fragment channel color prediction is described herein. Assuming a given pair of corresponding VDR and SDR images, that is, images representing the same scene but at different dynamic range levels, this section describes methods for allowing the encoder to approximate a VDR image as an SDR image and the piecewise cross channel predictor. color channel, PCCC). In the following description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the invention. In other instances, well-known structures and devices are not described exhaustively in detail to avoid unnecessarily confusing, blurring or obscuring the present invention.
GENERAL DESCRIPTION
[0017] The exemplary embodiments described herein relate to encoding high dynamic range pictures. In one embodiment, a sequence of visual dynamic range (VDR) images may be encoded using a standard dynamic range (SDR) base layer and one or more enhancement layers. The predicted image is generated by means of fragmentary color penetration channel prediction (PCCC), where the color channel in the SDR input data can be split into two or more color channel segments and each segment is assigned its own color crossing channel predictor to obtain the predicted output. VDR image. The PCCC prediction models for individual segments may include first order, second order, or higher parameters. Using the minimum mean square error criterion, the solution was presented in the form of an explicit formula for the prediction parameters of the second order PCCC model. Algorithms for segmenting color channels into multiple color channel segments are also presented. Prediction parameters can be transmitted to the decoder with ancillary data such as metadata.
[0018] In another embodiment, the decoder accesses the underlying SDR layer, the residual layer, and the PCCC prediction modeling metadata. The decoder generates an output predicted image using a base layer and a PCCC prediction parameter that can be used with the residual layer to generate an output VDR image.
SAMPLE VDR-SDR SYSTEM
[0019] Fig. 1 illustrates an exemplary data flow in a VDR-SDR system 100 according to one embodiment of the invention. The HDR image or video sequence is captured by HDR camera 110 or other similar means. After recording, the captured image or video is (s) processed during mastering to produce the target 125 VDR image. Mastering can include a variety of processing steps such as editing, major and secondary color correction, color conversion, and noise filtering. The 125 VDR output from this method typically represents the director's intentions as to how the captured image will be displayed on the target VDR display.
[0020] Mastering may also produce a corresponding SDR image 145 showing the director's intentions as to how a captured image will be displayed on an inherited SDR display. The output of the 145 SDR may be provided directly from the mastering circuit 120 or may be generated by a separate 140 VDRSDR converter.
[0021] In this exemplary embodiment, the VDR 125 and SDR 145 signals represent the input to the encoder 130. The purpose of the encoder 130 is to create an encoded bitstream that reduces the bandwidth required for the transmission of VDR and SDR signals, but also allows the corresponding decoder 150 to decode and decode and visualizing either SDR or VDR signal. In an exemplary embodiment, encoder 130 may be a layered encoder, such as one of the encoders defined by the MPEG-2 and H.264 encoding standards, presenting its output as a base layer, an optional gain layer, and metadata. As used herein, the term "metadata" refers to any auxiliary information transmitted as part of an encoded bitstream and assisting the decoder in visualizing the decoded image. Such metadata may include, but are not limited to, data such as color range or gamut information, dynamic range information, hue mapping information, or predictor operators such as those described herein.
[0022] At the receiver, the decoder 150 applies the received coded bitstreams and metadata to a visualized either an SDR image 157 or a VDR image 155, according to the capabilities of the target display. For example, an SDR display may only use the base layer and metadata to visualize the SDR image. On the other hand, a VDR display can use information from all input layers and metadata to visualize the VDR signal.
[0023] Fig. 2 shows in more detail an exemplary embodiment of an encoder 130 using methods of the invention. In Fig. 2, an optional SDR '207 signal is enhanced
-6 SDR signal. Typically SDR video today is 8-bit, 4: 2: 0, ITU Rec. 709 data. SDR 'may have the same color range (primary color and white point) as SDR, but may employ high precision, say 12 bits per pixel , with all color components in full spatial resolution (e.g. 4: 4: 4 RGB). From Fig. 2, one can derive the SDR from the SDR 'signal using a set of forward transforms. forward transforms), which can include quantization from, say, 12 bits per pixel to 8 bits per pixel, a color transform, say, from RGB to YUV, and color subsampling, say, from 4: 4: 4 to 4: 2: 0. The SDR output of transducer 210 is fed to compression system 220. Depending on the application of the compression system 220, it can either be lossy, such as H.264 or MPEG-2, or lossless, such as JPEG2000. The output from the compression system 220 can be transferred as the base layer 225. To reduce drift between the encoded and decoded signals, the encoder 130 often follows the compression process 220 with a corresponding decompression process 230 and inverse transforms 240 corresponding to forward transforms 210. Therefore, the predictor 250 may have the following inputs: VDR input 205 and either a compressed-decompressed SDR signal '(or SDR) 245, corresponding to the SDR signal' (or SDR) that will be received by the corresponding decoder 150, or the original SDR input data ' 207. The predictor 250, using the VDR and SDR '(or SDR) inputs, will produce a signal 257 which is an approximation or estimate of the input 205 VDR. The adder 260 subtracts the predicted VDR 257 from the original VDR 205 to produce a residual output 265. Then (not shown), residual 265 also may be encoded by another lossy or lossless encoder and may be sent to the decoder as a gain layer. In some examples, compression system 220 may receive SDR input 215 directly. In such examples, the units of the forward transform 210 and the inverse transform 240 may be optional.
[0024] The predictor 250 may also provide prediction parameters used in prediction as metadata 255. The prediction parameters may change during the encoding, for example between frames or between scenes, so this metadata may be transmitted to the decoder as part of the data also including the base layer and reinforcement layer.
[0025] Both VDR 205 and SDR '207 (or SDR 215) show the same scene but are intended for different displays with different characteristics such as dynamic range and color gamut, so it is expected that there will be a very tight correlation between these two signals. Co-owned U.S. Provisional Application No. 61 / 475,359, filed April 14, 2011 (now PCT Application No. PCT / US2012 / 033605, filed April 13, 2012 and entitled "Multiple color channel multiple regression predictor," henceforth referred to as Application '359, disclosed new multivariate and multi-regression multivariate, multi-regression, MMR) a prediction model that enables the prediction of a VDR input signal using its respective SDR signal (or SDR) and the MMR operator.
[0026] The MMR predictor of the '359 application can be considered "global" multicolor as it can be applied to all pixels of a frame independent of their individual color values. However, there are several when translating a VDR video sequence into an SDR video sequence
-7 operational factors that can degrade the performance of global predictors, such as color clipping and secondary color sorting.
[0027] In color clipping, the values of some pixels in one channel or a color component (e.g., the red channel) may be clipping more than the values of the same pixels in other channels (say, the green or blue channel). Clipping operations are non-linear operations, so the predicted values of these pixels may not match the global mapping assumptions, thus resulting in large prediction errors.
[0028] Another factor that can affect the SDR prediction on the VDR is secondary color sorting. In secondary color sorting, the coloring system can further divide each color channel into segments such as highlights, midtones, and shadows. These color boundaries can be controlled and adjusted when sorting colors. Estimating these color boundaries can improve overall prediction and reduce color artifacts in decoded video.
EXAMPLES OF PREDICTION MODELS
An example of notation and naming
[0029] Without loss of generality, an embodiment of a fragmentary color penetration channel predictor (PCCC) with two inputs: an input SDR (or SDR ') and an input signal v VDR is contemplated. Each of these inputs contains multiple color channels, also commonly referred to as color components (e.g., RGB, YCbCr, XYZ, and the like). Without loss of generality, regardless of bit depth, pixel values in each color component can be normalized to [0,1).
Assuming that all input and output signals are expressed using three color components, let us denote the three color components of the i-th pixel in the SDR image as:
<sup>S.</sup>, = kl Y γξ (1)
Let us denote the three components of the color of the i-th pixel in the VDR input data as:
v<sub>f</sub> = [ry v<sub>f2</sub> vj.
and the predicted three color components of the i-th pixel in the projected VDR should be marked as: νγίύι bil.
[0032] Each color channel, say c-th, can be divided into a set of multiple non-overlapping color segments by a set of cut-off points (e.g., Uci, u, 2, ..., u, u) such that within two consecutive segments (e.g., uiu +1) 0 <Ucu <Uc (u + i) <1. For example, in one embodiment, each color channel can be divided into three shadow, midtones, and highlights segments using two breakpoints, uci, and uc2. The shadows will then be defined in the range [0, uci), the midtones will be defined in the range [uc 1, uc 2), and the highlights will be defined in the range [uc 2, 1).
Let us denote the set of pixels having values within the u-th segment in the c-th channel
-8 color as Φu. Let us denote the number of pixels in Φ “as Pu. For ease of discussion and to simplify the notation, the procedure for the th segment in the-th color channel is described and can be repeated for all segments in all color channels. The proposed PCCC modeling can be combined with other color diffusion based models such as the models described in the '359 application. The second order PCCC model is described as an example and without loss of generality; however, these methods can easily be extended to other prediction models as well.
Exemplary second order PCCC model
Optimization of the prediction for the color channel segment
[0034] For the SDR signal let us denote the three color components of the i-th pixel in ° C as
8<sub>;</sub>= Ια ι AA (4)
[0035] For each SDR pixel in Φu a corresponding VDR co-located pixel can be found, denoted as <= kl · (5)
[0036] As used herein, the term "corresponding SDR and VDR co-located pixels" means two pixels, one in an SDR image and one in a VDR image, which may have different dynamic ranges but have the same pixel coordinates in the individual images. For example, for a 5 (10.20) SDR pixel, the corresponding pixel co-located VDR pixel is v (10.20).
Let us denote the predicted value of the c-th color component for this VDR pixel as:
= kl (6)
[0038] By collecting all Pu pixels in Φu, the following vector expressions can be generated:
<img file="PL2807823T3_D0001.tif" />
<img file="PL2807823T3_D0002.tif" />
and original VDR data V “-
<img file="PL2807823T3_D0003.tif" />
[0039] For a given SDR input signal, a prediction model may be defined having first and second (or higher) order SDR inputs, such as:
SC “. = Fr.<sub>j2</sub><sup>S.</sup>loam'<sup>S.</sup>i3 A Ά ΑΆάΙ '(8)
4], (9)
<img file="PL2807823T3_D0004.tif" />
(10)
[0040] These data vectors can be combined to form an input vector for a second order PCCC model:
<1· (11)
[0041] Based on equations (4) to (11), the VDR prediction problem can be expressed as:
(12) where M.<sup>at</sup><sub>c</sub> denotes the matrix of the prediction parameter for the u-th segment within the c-th color component. Note that this is a color gradient channel prediction model. In equation (12), the c-th color component of the predicted output signal is expressed as the combination of all the color components in the input signal. In other words, unlike other single-channel color predictors, where each color channel is processed separately and independently of the others, this model can account for all the color components of a given pixel and therefore can take full advantage of any inter-color correlations and redundancy.
[0042] By collecting all pu pixels, an appropriate data matrix can be created:
<img file="PL2807823T3_D0005.tif" />
[0043] The prediction operation may then be expressed in matrix form as:
V "= SCA -M". (14)
[0044] In one embodiment, a predictor of Mj solutions can be obtained using least squared optimization techniques, where the elements of Ms are selected to minimize the mean square error (MSE) between the original VDR and the predicted VDR.
min || v -ν<sup>Η</sup>| Γ. (15) μ; η U
Using the MSE criterion, the optimal solution of equation (15) can be expressed as:
(2) 7 (2) and (2) 7
M " <sub>=</sub>(SC “SC“) SC “V“. (16)
[0045] The above formulation derives a predictor for a specific segment within one of the color channels, assuming that the boundaries of those segments within the color channel are known. However, in practice, specific breakpoints of individual channel segments
-10may be inaccessible and may need to be outputted during the encoding process.
[0046] Fig. 3 shows an example prediction process according to one embodiment of the invention. In step 310, the predictor accesses the input signals VDR and SDR. In step 320, each color channel in the SDR input signal may be divided into two or more non-overlapping segments. The boundaries of these segments may be obtained as part of an input, say from a VDR to SDR color sorting process, or may be derived from the input using techniques such as described in the next section. In step 330, a prediction parameter matrix (e.g., M)) and an optimization criterion such as minimizing the MSE of prediction for each color segment in each of the color channels using a multi-color prediction model, such as the second order PCCC model of equations (4) to (14). In step 340, the output of the predicted VDR is calculated. In addition to computing the predicted VDR image, the prediction parameter matrix may be provided to the decoder by means of ancillary data such as metadata.
Optimization of the prediction for the entire color channel
[0047] Consider the problem of optimizing the prediction for all pixels in all segments within the c-th color channel. For all p pixels within this channel let us denote the projected VDR as:
<img file="PL2807823T3_D0006.tif" />
(Π)
<img file="PL2807823T3_D0007.tif" />
and the original VDR data as:
<img file="PL2807823T3_D0008.tif" />
(18)
[0048] The optimization problem for the c-th color channel can be formulated as the MSE minimization problem for finding:
<img file="PL2807823T3_D0009.tif" />
(19)
Given the set of cutoff points Uci, one can decompose the parameter optimization problem for the whole channel into several sub-problems, one for each segment of the color channel cth, and can derive a solution for each sub-problem by equation (16). More specifically, having given the set U of color segments, one can express equation (19) as:
<img file="PL2807823T3_D0010.tif" />
[0050] Let
<img file="PL2807823T3_D0011.tif" />
[0051] Given the set of cutoff points Uci, the total distortion for a set of prediction parameters can be given by the formula:
<img file="PL2807823T3_D0012.tif" />
[0052] When the value of any cutoff point changes, the above overall distortion also changes. Therefore, the aim is to identify those breakpoints for which the total distortion in the c-th channel is minimized.
<img file="PL2807823T3_D0013.tif" />
Sample solutions
A color channel with only two segments
[0053] Most of the SDR content is usually limited to 8 bits, so if you exclude 0 and 255 the total number of breakpoints is limited to 2<sup>8</sup>-2 = 254. If a color channel has only two color segments, only one boundary (Uc 1) needs to be identified in the interval [1.255). In one example, the complete search may compute J ({Uc 1}) for all possible 254 cutoff points and then select as cutoff point Uc 1, that cutoff point for which J ({Uc 1}) is minimal.
[0054] In one embodiment, the best cutoff point can be derived using a heuristic iterative search technique, which may improve the search time but not necessarily yield optimal cutoffs. For example, in one embodiment, the original SDR range may be further divided into K segments (e.g., K = 8). Then, assuming that the limit Uc 1 is approximately in the middle of each of these segments, one can calculate equation (22) K times. Let kc denote the segment with the minimum prediction error among all K segments. Either a full search or similar hierarchical searches can then be performed within segment kc to identify the locally optimal breakpoint. The steps of this two-step search algorithm are summarized in pseudo-code form in Table 1.
Table 1. Two-step search algorithm
Dividing the color range interval into K segments // First step (a) Computing the prediction error Jk ({Uc 1}) for each k segment, assuming that the cutoff point Uc 1 is approximately in the middle of the kth segment.
(b) Finding a segment, say kc, for which Jk ({Uc 1}) is minimal.
// Second step (a) Within segment kc - use a full search or repeat this two-step algorithm to find Uc 1 that minimizes prediction error.
[0055] This two-step search algorithm can be easily modified for the alternative examples. For example, instead of assuming the cutoff point is approximately in the middle of the kth segment, it can be assumed that the cutoff point is at the beginning, end, or anywhere else of the segment.
[0056] For color ranges having more than two segments, similar heuristics and iterative search techniques may also be used. For example, for 8-bit SDR data, after identifying the first cut-off point Uc1 in the interval (1,255), an attempt can be made to identify two candidates for the second cut-off point: one candidate in the sub-interval (0, Uc 1) and the other in the sub-interval (Uc 1, 255) ). By computing the total distortion J ({Uci}) for each of these two candidates, the second cut-off point (Uc 2) can then be determined as the one that gives the smallest prediction error (e.g. by equation (22)) of the two candidate solutions.
[0057] The color sorting of video frames is highly correlated, especially for all frames within the same scene, so the search for breakpoints for the nth frame may also consider known results from previous frames within the same scene. Alternatively, breakpoints can only be calculated once for the entire scene. An example of a scene-based search algorithm is described in pseudo-code form in Table 2. In this embodiment, after the breakpoint for the first frame has been identified using the full dynamic range of the color channel, subsequent frames use it as the starting point to define the breakpoint within a much smaller segment of the color range.
Table 2. Scene-based search algorithm
For the first frame in a given scene (1), the execution of a two-step algorithm to identify the cut-off point within the color channel.
For the remaining frames in the same scene (2), use the boundary point from the previous frame to determine the segment that will be used as the starting point of the second step in the two-step search (see table 1).
[0058] It should be appreciated that the steps of the algorithm can be performed in a variety of alternative ways. For example, in step (1), a full search or any other type of search algorithm may be used instead of using a two-step search algorithm to identify a breakpoint. As another example in step (2), to have a given starting point, you can take that starting point as an approximate midpoint
-13segment with a predetermined length. Alternatively, it can be considered as the segment start point, segment end point, or any other position on the segment.
[0059] The methodology described herein can also be used in deriving other PCCC models. For example, a first-order PCCC model can be derived by using only the first three terms of equation (11) using the equations sA = [and<sub>S.</sub>“. <sub>SC</sub>"J, (24) i
/ ΑΧ (25)
[0060] Similarly, the data vectors in equations (8) - (11) can be extended to define third order or higher order PCCC models.
IMAGE DECODING
[0061] Embodiments of the invention may be performed in either an image encoder or an image decoder. Fig. 4 shows an exemplary embodiment of a decoder 150 according to one embodiment of the invention.
The decoding system 400 receives an encoded bitstream that may combine the base layer 490, an optional gain (or residual) layer 465, and metadata 445, which are obtained after decompressing 430 and various optional inverse transforms 440. For example, in a VDR-SDR system, base layer 490 may be the SDR representation of the encoded signal and the metadata 445 may include information about the PCCC prediction model that has been used in the encoder predictor 250 and the corresponding prediction parameters. In one exemplary embodiment, when the encoder uses a PCCC predictor according to the methods of the invention, the metadata may include cutoff values that identify individual color segments within each color channel, identification of the model used (e.g., first order PCCC, second order PCCC, and the like), and all. the coefficients of the matrix of the prediction parameter associated with that particular model. Given the underlying layer 490 s and the prediction parameters derived from metadata 445, the predictor 450 can compute predicted v 480 using any of the appropriate equations described herein (e.g., equation (14)). If there is no residual or the residual is negligible then the predicted value 480 can be output directly as the final VDR image. Otherwise, the output from the predictor (480) is added in the adder 460 to remainder 465, giving an output of 470 VDR.
EXAMPLE DESIGN OF A COMPUTER SYSTEM
Embodiments of the invention may be implemented with a computer system, systems configured in circuits and electronic components, an integrated circuit (IC) device such as a microcontroller, a directly programmable gate array, FPGA) or other configurable or programmable logic
-14 (programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or apparatus including one or more such systems, devices or components. The computer and / or IC may execute, control or execute PCCC based prediction commands such as described herein. The computer and / or IC may calculate any of a variety of parameters or values that relate to PCCC prediction as described herein. Embodiments for image and video dynamic range extension may be implemented with hardware, software, firmware, and various combinations thereof.
[0063] Some embodiments of the invention include computer processors executing program commands that cause the processors to perform the inventive method. For example, one or more processors in a display, encoder, set top box (STB), transcoder, or the like may perform PCCC based prediction methods as described above by executing program commands in program memory accessible to the processors. The invention may also be provided as a program product. The program product may include any medium that contains a set of computer-readable signals containing commands that, when executed by the data processor, cause the data processor to perform the method of the invention. The products of the inventive program can be in any of a wide variety of forms. For example, the program product may contain physical media such as magnetic media including flexible disks, hard drives, optical media including CD ROMs, DVDs, electronic media including ROMs, flash RAM or the like. The computer readable signals in the software product may optionally be compressed or encrypted.
[0064] Where reference is made to some item (e.g., software module, processor, assembly, device, circuit, etc.) above, reference to that item (including reference to "means") is to be construed as including equivalents of this element, being any element that performs the function of the described element (e.g. is functionally equivalent), including elements that are not structurally equivalent to the disclosed structure that performs a given function in the illustrative embodiments of the invention shown.
BALANCES, EXTENSIONS, ALTERNATIVES AND OTHER
[0065] Illustrative embodiments are therefore described that relate to using PCCC prediction in coding VDR and SDR pictures. In the foregoing specification, embodiments of the invention have been described with reference to a number of specific details that may vary in the various embodiments. Thus, one and sole indicator of what the invention is, and what is intended as an invention by the applicants, is the set of patent claims which are derived from the present application, in the specific form in which such claims originate, including any subsequent amendments. Any definitions of terms expressly set forth herein in such claims will determine the meaning of such terms when used in the claims. Therefore
Any limitation, element, feature, feature, advantage or attribute not expressly stated in the claim will in no way limit the scope of such claim. Accordingly, the specifications and drawings are to be considered in an illustrative rather than restrictive sense.
Contents11
17 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17
69 members in 11 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 201261590175 | United States of America | P | |
| 13703222 | European Patent Office (EPO) | A | |
| 2013022673 | United States of America | W | |
| 137032223 | – | – | – |
| 201261590175P | – | – | – |
| EP20130703222 | – | – | – |
| US201261590175P | – | – | – |
| WO2013US22673 | – | – | – |
Members69
| Document | Office | Kind | |
|---|---|---|---|
| WO2012142471A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2012142506A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2013112532A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2013112532A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2012142471A8 | World Intellectual Property Organization (WIPO) | A8 | |
| CN103493489A | China | A | |
| US2014029675A1 | United States of America | A1 | |
| CN103563372A | China | A | |
| US2014037205A1 | United States of America | A1 | |
| EP2697971A1 | European Patent Office (EPO) | A1 | |
| EP2697972A1 | European Patent Office (EPO) | A1 | |
| US8731287B2 | United States of America | B2 | |
| US2014185930A1 | United States of America | A1 | |
| US8811490B2 | United States of America | B2 | |
| JP2014520414A | Japan | A | |
| US8837825B2 | United States of America | B2 | |
| EP2782348A1 | European Patent Office (EPO) | A1 | |
| HK1193688A1 | Hong Kong, China | A1 | |
| US2014307796A1 | United States of America | A1 | |
| EP2807823A2 | European Patent Office (EPO) | A2 | |
| US2014369409A1 | United States of America | A1 | |
| EP2697972B1 | European Patent Office (EPO) | B1 | |
| US8971408B2 | United States of America | B2 | |
| US2015092850A1 | United States of America | A1 | |
| EP2697971B1 | European Patent Office (EPO) | B1 | |
| JP5744318B2 | Japan | B2 | |
| US2015222916A1 | United States of America | A1 | |
| JP2015165665A | Japan | A | |
| EP2945377A1 | European Patent Office (EPO) | A1 | |
| HK1204741A1 | Hong Kong, China | A1 | |
| JP5921741B2 | Japan | B2 | |
| US9386313B2 | United States of America | B2 | |
| HK1214049A1 | Hong Kong, China | A1 | |
| US9420302B2 | United States of America | B2 | |
| JP2016167834A | Japan | A | |
| US2016269756A1 | United States of America | A1 | |
| CN103493489B | China | B | |
| US9497475B2 | United States of America | B2 | |
| US2017034521A1 | United States of America | A1 | |
| CN103563372B | China | B | |
| CN106878707A | China | A | |
| US9699483B2 | United States of America | B2 | |
| CN107105229A | China | A | |
| US2017264898A1 | United States of America | A1 | |
| JP6246255B2 | Japan | B2 | |
| EP2782348B1 | European Patent Office (EPO) | B1 | |
| US9877032B2 | United States of America | B2 | |
| ES2659961T3 | Spain | T3 | |
| TR201802291T4 | Türkiye | T4 | |
| EP2807823B1 | European Patent Office (EPO) | B1 | |
| JP2018057019A | Japan | A | |
| PL2782348T3 | Poland | T3 | |
| CN106878707B | China | B | |
| EP3324622A1 | European Patent Office (EPO) | A1 | |
| ES2670504T3 | Spain | T3 | |
| US10021390B2 | United States of America | B2 | |
| PL2807823T3This record | Poland | T3 | |
| US2018278930A1 | United States of America | A1 | |
| EP2945377B1 | European Patent Office (EPO) | B1 | |
| US10237552B2 | United States of America | B2 | |
| JP6490178B2 | Japan | B2 | |
| PL2945377T3 | Poland | T3 | |
| EP3324622B1 | European Patent Office (EPO) | B1 | |
| DK3324622T3 | Denmark | T3 | |
| CN107105229B | China | B | |
| PL3324622T3 | Poland | T3 | |
| HUE046186T2 | Hungary | T2 | |
| ES2750234T3 | Spain | T3 | |
| CN107105229B9 | China | B9 |
Numbers
- Publication
- 2807823
- Publication, DOCDB
- 2807823
- Publication, EPODOC
- PL2807823T
- Application
- 13703222
- Application, DOCDB
- 13703222
- Application, EPODOC
- PL13703222T
Titles2
- English
- PIECEWISE CROSS COLOR CHANNEL PREDICTOR
- Polish
- Predyktor kanału fragmentowego przenikania kolorów
Classification
- CPC, 7
- H04N19/182
- H04N19/105
- H04N19/137
- H04N19/186
- H04N19/192
- H04N19/194
- H04N19/30
- IPC, 5
- H04N19 105
- H04N19 137
- H04N19 182
- H04N19 186
- H04N19 30