Multiple color channel multiple regression predictor
11 claims: 3 independent, 8 dependent
- 1Patent claims Zastrzeżenia patentowe 1. In a video coding or decoding system, a method of prediction of an image using a processor, the method comprising:1. W systemie kodowania lub dekodowania wideo, sposób predykcji obrazu za pomocą procesora, przy czym sposób obejmuje: odbieranie pierwszego obrazu;receiving the first image;odbieranie parametrów predykcji dla modelu predykcji wielokanałowej regresji wielokrotnej (MMR), który to model MMR służy do określania predykcyjnego drugiego obrazu z pierwszego obrazu, przy czym drugi obraz reprezentuje tę samą scenę co pierwszy obraz i gdzie zakres dynamiki pierwszego obrazu jest mniejszy niż zakres dynamiki drugiego obrazu;i zastosowanie pierwszego obrazu i parametrów predykcji do modelu predykcji MMR w celu wygenerowania obrazu wyjściowego zbliżonego do drugiego obrazu, w którym wartości pikseli co najmniej jednej składowej barwy obrazu wyjściowego oblicza się na podstawie wartości pikseli co najmniej dwóch składowych barwy w pierwszym obrazie, przy czym do generowania określanej predykcyjnie wartości (v i) pikseli wyjściowych dla składowej luma i/albo chroma, model MMR obejmuje model MMR pierwszego rzędu z iloczynami wektorowymi zawierającymi: stałą wartość (n), co najmniej liniową kombinację wartości pikseli dwóch różnych składowych barwy i co najmniej jedno wyrażenie iloczynu wartości pikseli dwóch różnych składowych barwy. receiving prediction parameters for a multi-channel multiple regression prediction model (MMR), which MMR model is used to determine the predictive second image from the first image, the second image representing the same scene as the first image and where the dynamic range of the first image is less than the dynamic range of the second image;and applying the first image and prediction parameters to the MMR prediction model to generate an output image similar to the second image in which the pixel values of at least one color component of the output image are calculated based on the pixel values of at least two color components in the first image, wherein generating a predetermined value (vi) of the output pixels for the luma and / or chroma component, the MMR model includes the first-order MMR model with vector products containing: constant value (n), at least a linear combination of pixel values of two different color components and at least one expression of the product of the pixel values of two different color components.
- 6A video system that includes:6. System wideo, który obejmuje: dane wejściowe do odbioru pierwszego obrazu i parametry predykcji dla modelu predykcji wielokanałowej regresji wielokrotnej (MMR), w którym model MMR służy do predykcji drugiego obrazu z pierwszego obrazu, przy czym drugi obraz reprezentuje to samo ujęcie co pierwszy obraz i gdzie zakres dynamiki pierwszego obrazu jest mniejszy niż zakres dynamiki drugiego obrazu;input data for receiving the first image and prediction parameters for the multichannel multiple regression prediction model (MMR), in which the MMR model is used to predict the second image from the first image, the second image representing the same image as the first image and where the dynamic range of the first image is smaller than the dynamic range of the second image;a processor for applying the first image and prediction parameters to the MMR prediction model to generate an output image approximating a second image in which the pixel values of at least one color component of the output image are calculated based on the pixel values of at least two color components in the first image, where generating a predetermined value (vi) of the output pixels for the luma and / or chroma component, the MMR model includes the first-order MMR model with vector products containing: constant value (n), at least a linear combination of pixel values of two different color components and at least one expression of the product of the pixel values of two different color components. procesor do zastosowania pierwszego obrazu i parametrów predykcji do modelu predykcji MMR w celu wygenerowania obrazu wyjściowego aproksymującego drugi obraz, w którym wartości pikselowe co najmniej jednej składowej barwy obrazu wyjściowego są obliczane na podstawie wartości pikseli co najmniej dwóch składowych barwy w pierwszym obrazie, gdzie w celu wygenerowania określonej predykcyjnie wartości (vi) pikseli wyjściowych dla składowej luma i/albo chroma, model MMR obejmuje model MMR pierwszego rzędu z iloczynami wektorowymi zawierającymi: stałą wartość (n), co najmniej liniowe połączenie wartości pikseli dwóch różnych składowych barwy i co najmniej jedno wyrażenie iloczynu wartości pikseli dwóch różnych składowych barwy.
- 11A computer readable durable storage medium on which instructions computer executes are stored for performing, using one or more processors, a method according to any of claims 1-5. 11. Trwały nośnik danych czytelny dla komputera, na którym są przechowywane instrukcje wykonywane przez komputer dla wykonywania, za pomocą jednego albo większej liczby procesorów, sposobu według dowolnego z zastrzeżeń 1-5. FIG. 3 FIG. 3 400 400 FIG. 4 FIG. 4 420> * FIG. 5 420>* FIG. 5 FIG. 6 FIG. 6 Base layer 6 & 0 Warstwa bazowa 6&0
Independent claims3
163 paragraphs, as filed
Description
REFERENCE TO RELATED APPLICATIONS [0001] This application is a European application separated from the European patent application EP 14171538.3 (reference: D11020EP02), for which the EPO 1001 form was filed on June 6, 2014. EP 14171538.3 is a European application separated from the Euro-PCT patent application EP 12717990.1 (reference: D11020EP01), submitted on April 13, 2012 and granted as EP 2 697 971 on July 8, 2015.
TECHNICAL FIELD [0002] The present invention generally relates to images. In particular, an embodiment of the present invention relates to a multi-color channel, a multiple regression predictor between high dynamic range images and standard dynamic range images.
BACKGROUND OF THE INVENTION [0003] As used herein, the term 'dynamic range' (DR) may refer to the ability of the human psycho-visual system (HVS) to perceive the range of intensity ( for example luminance, lumens) in the image, for example from the darkest darkness to the brightest brightness. In this sense, the scope of DR relates to the intensity 'relative to the shot'. The DR range may also refer to the ability of the display device to appropriately or approximately render the intensity range of a certain width. In this sense, the DR range refers to the intensity 'relative to the display'. Unless it is explicitly stated in any place in this description that a given meaning is particularly important, it should be concluded that the expression may be used in both meanings, for example interchangeably.
[0004] As used herein, the expression large dynamic range (HDR) refers to the width of the DR range, which includes about 14-15 orders of magnitude of the human visual system (HVS). For example, well-adapted individuals with a generally normal condition (for example, in one or more of statistical, biometric or ophthalmological meanings) have an intensity range that includes about 15 orders of magnitude. Adapted people can see dimmed light sources as low as just a few photons. The same people, however, can notice the painfully bright intensity of the midday sun in the desert, sea or snow (and even look at the sun, but for a moment to avoid damage). However, this coverage is available to 'adapted' people, for example, whose HVS system has a period of time during which they can re-adjust and adapt.
[0005] In contrast, the DR range, in which a human can simultaneously see a large width in the intensity range, relative to the HDR range, may be somewhat truncated. As used herein, the terms 'visual dynamic range' or 'variable dynamic range' (VDR) may individually or interchangeably refer to the DR range, which is also able to be seen by the HVS system. As used herein, the VDR range may refer to the DR range, which includes 5-6 orders of magnitude. Thus, while it is slightly narrower compared to the true HDR range in terms of shot, the VDR range represents a wide DR range. As used herein, the expression "simultaneous dynamic range" may refer to the VDR range.
[0006] Until quite recently, the displays had a much narrower DR range than HDR or VDR. The ability of a television and computer display device that uses a cathode ray tube (CRT), a liquid crystal display (LCD) with constant fluorescent white background lighting or plasma screen technology to render the DR range be limited to approximately three orders of magnitude. Thus, such traditional displays are characteristic of the low dynamic range (LDR), also called standard dynamic range (SDR), compared to the VDR or HDR range.
[0007] However, advances in the underlying technology allow that in more modern display designs the rendering of image and video content is significantly improved in terms of different quality characteristics with the same content as compared to content rendered on less modern displays. For example, more modern display devices may be able to render high definition (HD) content high definition) and / or content that can be scaled according to various display capabilities, such as the image scaling element. In addition, some modern displays are able to render content with a DR range that is larger than the SDR range of traditional displays.
[0008] For example, some modern LCD displays have a backlight unit (BLU) that includes an arrangement of light emitting diodes (LEDs). BLU LEDs can be modulated separately from the polarization modulation of the active LCD elements. This double modulation approach is extensible (for example, to N modulation layers, where N contains an integer greater than two), as with controlled intermediate layers between the BLU system and LCD screen elements. Their BLU units, based on a LED system and double modulation (or N modulation), effectively increase the DR range with respect to LCD monitor displays that have such characteristics.
[0009] Such 'HDR displays' as they are often called (although, in fact, their capabilities may be closer to the VDR range), and the DR extension they are capable of, compared to traditional SDR displays, are significant progress in the ability to display images, video content and other visual information. The range of colors that such an HDR display can render also can significantly exceed the range of colors of more traditional displays, even to the degree of rendering ability of a wide range of colors (WCG). The HDR or VDR range in relation to the shot and WCG image content that can be generated by 'next generation' film and television cameras now, using 'HDR' displays (hereinafter 'HDR displays'), can be displayed more accurately and efficiently.
[0010] As with scalable video coding and HDTV technology, extending the DR range of an image usually involves a forked approach. For example, the content of the HDR range for a shot that has been recorded with a modern HDR camera can be used to generate the SDR version of the shot that can be displayed on traditional SDR displays. In one approach, generating an SDR version from a registered VDR version may involve applying a global tone mapping operator (TMO) to pixel intensity values (e.g., luminance, lumens) in HDR content. In the second approach, as described in International Patent Application No. PCT / US2011 / 048861 filed on August 23, 2011, incorporated herein by reference to all purposes, generating an SDR image may involve using a reversible operator (or predictor) on VDR data. Uploading actual registered VDR content may not be the best approach to conserve bandwidth or for other reasons.
[0011] Thus, for the version of SDR content that has been generated, a reverse tone mapping operator (iTMO) can be applied, inverted relative to the original TMO, or an inverse operator relative to the original predictor, which allows predictive determination of the VDR content version. The prediction version of VDR content can be compared with the originally recorded HDR content. For example, subtracting the predictive VDR version from the original VDR version may generate a residual image. The encoder can send the generated SDR content as a base layer (BL) and package the generated version of the SDR content, any residual image and iTMO or other predictors as an enhancement layer (EL) or as metadata.
[0012] Sending an EL layer and metadata, with SDR content, debris and predictors, in a bit stream typically consumes less bandwidth than would consume sending both HDR and SDR content directly to the bit stream. Compatible decoders that receive the bit stream sent by the encoder can decode and render the SDR range on traditional displays. However, compatible decoders can also use residual image, iTMO predictors or metadata to calculate the prediction-specific version of HDR content from them for use on more suitable displays. The object of the invention is to provide modern predictor generation methods enabling efficient encoding, transmission and decoding of VDR data using corresponding SDR data.
[0013] The approaches described in this section are approaches that could be implemented, but are not necessarily approaches that were previously developed or that were previously implemented. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section merely by virtue of their inclusion in this section qualifies as prior art. Similarly, this section should not assume that any state of the art has identified problems identified in connection with one or more approaches, unless otherwise indicated.
[0014] EP 2 009 921 A2 discloses methods of coding and decoding image sequences with scalable bit depth. In particular, the methods perform inverse mapping of image tones.
[0015] EP 2 144 444 A1 discloses methods of compressing a video frame data stream in which tone mapping functions are determined. These tone mapping functions are different for frames that refer to different shots, and the tone mapping functions are changed for stream frames that refer to the same shot.
[0016] US 2009/097561 A1 discloses methods associated with bit depth gain for scalable video coding.
[0017] YU Y ET AL: "Bit depth", 26 JVT MEETING; 83 MPEG MEETING; 131-2008 - 18-1-2008, ANTALYA, No. JVT-Z045, January 22, 2008 describes research work with SVC bit depth. In particular, using a filter for the reconstructed image from the bottom layer, an average of 6.0% BDBR or an average of 0.20 dB BDPSNR can be achieved.
[0018] WO 2008/128898 A1 discloses methods of coding / decoding video data in the case where two or more video versions with different color bit depths use different color coding.
BRIEF DESCRIPTION OF THE FIGURES [0019] An embodiment of the present invention is shown by way of example and not by way of limitation, in the figures of the accompanying drawing and in which the same reference numerals apply to similar elements and in which:
FIG. 1 depicts an example data flow for a VDR-SDR system, according to an embodiment of the present invention;
FIG. 2 depicts an example VDR coding system according to an embodiment of the present invention;
FIG. 3 illustrates the input and output interfaces of a multiple variable multiple regression predictor according to an embodiment of the present invention;
FIG. 4 depicts an example multi-variable regression predictive determination according to an embodiment of the present invention;
FIG. 5 illustrates an exemplary process of selecting a multiple variable multiple regression predictor model according to an embodiment of the invention;
FIG. 6 shows an example image decoder with a predictor operating according to an embodiment of the present invention.
DESCRIPTION OF EXEMPLARY EMBODIMENTS [0020] This document describes inter-color prediction of an image based on multi-variable multiple regression modeling. Having a pair of corresponding VDR and SDR images, i.e. images that represent the same scene but at different levels of dynamic range, this section describes ways that allow the encoder to zoom in on the VDR image in terms of the SDR image and the multiple-regression multiple predictor ( MMR) ( multivariate multi-regression). The following description, for clarification purposes, to provide a thorough understanding of the invention, contains many specific details. However, it will be apparent that the invention can be used without these specific details. In other cases, well-known constructions and devices have not been exhaustively described to avoid unnecessarily obstructing or obfuscating the present invention.
OVERVIEW [0021] The exemplary embodiments described in this document relate to the coding of images with a high dynamic range. In an embodiment, the MMR predictor has been created that allows the VDR image to be expressed relative to its corresponding SDR representation.
EXAMPLE VDR-SDR SYSTEM [0022] FIG. 1 depicts an example data flow in a 100 VDR-SDR system, according to an embodiment of the present invention. The video sequence or HDR image was recorded using a 110 HDR camera. After recording, the recorded image or video is processed during the mastering process to create the target 125 VDR image. The mastering process can include many processing steps, such as assembly, primary and secondary color correction, color conversion and noise filtering. The output of the 125 VDR of this process represents the director's intention as to how the recorded image will be displayed on the target VDR display.
[0023] The mastering process may also output the corresponding 145 SDR image, showing the director's intention as to how the recorded image will be displayed on the older SDR display. The 145 SDR output can be obtained directly from the 120 mastering system or can be generated by a separate 140 VDR-to-SDR converter.
[0024] In this embodiment, the VDR 125 and SDR 145 signals are input to the encoder 130. The purpose of the encoder 130 is to create an encoded bit stream that will reduce the bandwidth required to transmit the VDR and SDR signals, but also allows the corresponding decoder 150 to decode and rendering either SDR or VDR signals. In an exemplary embodiment, the encoder 130 may be a layer encoder, such as one of those defined in the MPEG-2 and H.264 encoding standards, which presents its result as a base layer, optional gain layer, and metadata. As used herein, the expression "metadata" refers to any auxiliary information that is transmitted as part of an encoded bit stream and helps the decoder render the decoded image. Such metadata may include, but are not limited to, data such as space or color gamut information, dynamic range information, tone mapping information, and MMR predictor operators, such as those described in this document.
[0025] At the receiver, the decoder 150 uses received encoded bit streams and metadata to render either the SDR image or the VDR image in accordance with the capabilities of the target display. For example, to render an SDR image, the SDR display can only use the base layer and metadata. However, to render the VDR signal, the VDR display can use information from all input layers and metadata.
[0026] FIG. 2 shows a more detailed embodiment of the encoder 130 including the methods of the present invention. In FIG. 2, SDR 'means amplified SDR signal. Today, SDR video is 8-bit data, 4: 2: 0, ITU Rec. 709. SDR 'may have the same color space (primary colors and white point) as SDR, but it can use high precision, say, 12 bits pixel, with all color components at full spatial resolution (for example, 4: 4: 4 RGB). From FIG. 2, from the SDR 'signal easily using a set of ordinary transforms, which may include quantization from, say, 12 bits to 8 bits per pixel, color conversion, say, from RGB to YUV, and color sub-sampling, say, from 4: 4: 4 by 4: 2: 0, you can get an SDR. Converter SDR 210 output is applied to compression system 220. Depending on the application, the 220 compression system can be either loss-making, such as H.264 or MPEG-2, or lossless. The output of the compression system 220 can be sent as the base layer 225. To reduce the drift between coded and decoded signals, it is not uncommon for the encoder 130 to perform the compression process 220 with the corresponding decompression process 230 and the inverse transforms 240 corresponding to the normal transforms 210. Thus, the predictor 250 may have the following input data: 205 VDR input data and either the 245 SDR 'signal that corresponds to the SDR signal that will be received by the corresponding decoder or the SDR' 207 input data. The predictor 250 using the VDR and SDR input data ', creates a signal 257 that is an approximation or estimation of the VDR 205 input data. The add element 260 to create the output residual signal 265 subtracts the predetermined VDR 257 from the original VDR 205. Later (which is not shown), the residue 265 can also be encoded either with a loss or lossless encoder, and sent to the decoder as a gain layer.
[0027] The predictor 250 can also pass prediction parameters used during the prediction process as metadata 255. Since the prediction parameters may change during the coding process, for example, from frame to frame or from shot to shot, these metadata can be sent to the decoder as part of the data, which also includes the base layer and the enhancement layer.
[0028] Since both VDR 205 and SDR '207 present the same shot, but are directed to different displays with different characteristics, such as dynamic range and color range, it is expected that there is a very close correlation between these two signals. In an exemplary embodiment of the present invention, a new multi-variable multiple regression predictor 250 (MMR) has been developed that can predict the input signal with the corresponding SDR 'signal and the multi-variable MMR operator
VDR.
EXAMPLES OF PREDICTION MODELS [0029] FIG. 3 shows the 300 MMR predictor input and output interfaces according to an exemplary embodiment of the present invention. From FIG. 3, the predictor 330 receives v 310 and 320 input vectors representing VDR and SDR image data, respectively, and outputs the v 340 vector representing the predicted input v value.
Example notation and nomenclature [0030] Let us denote the three color components of the z-pixel in the 320 SDR image as <sup>S</sup>r = [\ l [0031] Let's denote the three color components of the z-pixel in the input VDR 310 as = HI [0032] Let's denote the three predetermined color components of the z-pixel in the predictively determined VDR 340 as u = [Ai Aj AJ (3 )
Let's denote the total number of pixels in one color component as p.
[0033] In equations (1-3), the color pixels can be represented in the RGB, YUV, YCbCr, modelu model or in any other color representation. While in equations (1-3) for each pixel in the image or video frame the representation of three colors is assumed, which is also shown later, the methods described in this document can easily be extended to represent the image and video with more than three color components per pixel or image representation in which one of the input can have pixels with a different number of color representations than the other input.
First order model (MMR-1) [0034] Using a multiple variable multiple regression model (MMR), the first order prediction model can be expressed as:
v<sub>f</sub> - -ι-n, (4) where is a 3x3 matrix, an is a 1x3 vector defined as:
(5) [0035] Please note that this is a multi-color channel prediction model. In Vj from equation (4), each color component is expressed as a linear combination of all color components in the input data. In other words, unlike other single channel color predictors, in which each color channel is processed independently and independently of each other pixel, in this model all components of the applause of the i-th pixel are taken into account, thereby fully applying all correlation and redundancy between colors.
[0036] Equation (4) can be simplified using an expression based on a single matrix:
v. = (6) where
M<sup>(1)</sup> =
<img file="PL3324622T3_D0001.tif" />
(7) [0037] By gathering together all the pixel of the frame (or other suitable segment or other relevant part of the input data), the following matrix expression can be obtained,
V = S.<sup>r</sup>M<sup>loam]</sup> (H) where p<sub>0 </sub>si is the input data and predictive output data, S 'is the data matrix / 2x4, V is the matrix / 7x3, and M<sup>(1)</sup> is a 4x3 matrix. As used herein, M<sup>(1)</sup> can be called, interchangeably, a multi-variable operator or a prediction matrix.
[0038] Based on this linear system of equations (8), this MMR system can be expressed in two different problems: (a) least squares problem or (b) least squares problem; each of which can be solved using well-known numerical methods. For example, using the least squares approach, the solution problem for M can be expressed as minimizing the residual or error of the mean squares of the prediction, or (10) where V is a matrix / 7x3 formed with the corresponding VDR input.
[0039] Having equations (8) and (10), the optimal solution for M<sup>(1)</sup> gives
M<sup>b</sup> = (S<sup>r</sup>S ') -<sup>1</sup>S '<sup>r</sup>V where, S '<sup>T</sup> indicates the transposition of S 'and S'<sup>T</sup>S 'is a 4x4 matrix.
[0040] If S 'has a full columnar row, for example, <sup>RZC? C /</sup>(S ') = 4 <p, then M<sup>(1)</sup> can also be solved by using many different numerical techniques, including SVD, QR or LU distributions.
Second order model (MMR-2) [0041] Equation (4) shows the first order MMR prediction model. You may also consider adopting a higher order of prediction, as described later.
[0042] The second order MMR model can be expressed as:
v, + «<sub>Γ</sub>Μ<sup>(Ι)</sup>+ n (12) where M ^<sup>2</sup>) is a 3x3 matrix,
<img file="PL3324622T3_D0002.tif" />
(13) [0043] Equation (12) can be simplified using a single prediction matrix,
ν. = Λΐ<sup>(2</sup>>, IL '(14) where
<td></td><td>«ΙΪ</td><td>and;</td>
<td></td><td></td><td></td>
<td></td><td></td><td></td>
<td></td><td> '4’</td><td>to me·'</td>
<td></td><td></td><td></td>
<td></td><td></td><td></td>
<td></td><td></td><td></td>
(15) <sup>s</sup>and<sup>2t</sup> = b% ν V], (16) [0044] By collecting all p pixels together, the following matrix expression can be determined:
V = S.<sup><3)</sup>M<sup>i2></sup> (17) where
S<sup>(from}</sup> =
si
(] «) In
% -iJ [0045] Equation (14) can be solved using the same optimization and the same solutions as described in the previous section. The optimal solution for M.<sup>(2)</sup> for least squares problems is
M<sup>t2</sup>> = (S<sup>(1> T</sup>S<sup>tIL</sup>) 'S ^ Y, (19) where S<sup>(2) T</sup>S<sup>(2)</sup> it's now a 7x7 matrix.
[0046] Third or higher order MMR models can also be built in a similar manner.
First order model with vector product (MMR-1C) [0047] In the alternative MMR model, the first order prediction model from equation (4) can be extended by a vector product between the color components of each pixel as follows:
v, = sc<sub>J</sub>Æ<sup>loam)</sup> + n (20) where is the 3x3 matrix, an is the 1x3 vector, both as defined in equation (5), and.
<td></td><td></td><td></td><td>'""at?</td>
<td>c<sup>(L</sup>> =</td><td></td><td></td><td>dregs</td>
<td></td><td>-and?</td><td></td><td>WTY</td>
<td></td><td></td><td></td><td>^ ty_</td>
(21) [0048] Using the same approach as before, the MMR-1C model of equation (20) can be simplified by using a single MC prediction matrix as in:
v. = scj<sup>,)</sup> -MC<sup>LL)</sup>where (22)
MC ™
<img file="PL3324622T3_D0003.tif" />
<td>«II</td><td>fl |<sub>2</sub></td><td></td>
<td></td><td></td><td>'and?</td>
<td></td><td></td><td></td>
<td></td><td></td><td></td>
<td></td><td></td><td></td>
<td></td><td>Hlity</td><td>dregs</td>
<td></td><td>dregs</td><td>dregs</td>
<td>Toilets'<sup>1</sup>/</td><td>dregs</td><td>dregs</td>
(23) a
80 ^ = (1 p<sub>7</sub> scj = [l 5,., ó '<sub>loam</sub> Sn '\<sub>3</sub> Ai '%] [0049] By collecting all p pixels together, a simplified matrix expression can be obtained, as in
V = SC MC<sup>(AND)</sup> ,
U5) where
<img file="PL3324622T3_D0004.tif" />
(26) [0050] SC is the matrix / x x (1 + 7), and equation (25) can be solved using the same least squares equation described earlier.
Second-order model with vector product (MMR-2C) [0051] The first-order MMR-1C model can be extended to also include second-order data. For example, sc<sup>2</sup>C<sup>Background></sup> + s; M<sup><2j</sup> 4-sc, C<sup>(L)</sup> + p<sub>r</sub>M<sup>(L}</sup> + n, (27) where
C<sup><!)</sup> =
<td>jmc<sub>L</sub>|</td><td>mc /</td><td><sup>HC</sup>l?</td>
<td></td><td></td><td>JHcJ<sup>2</sup>’</td>
<td></td><td>/ NCJ?</td><td></td>
<td></td><td></td><td></td>
<sup>Λ</sup> 2 2 22 2 > ^2^
SC<sub>;</sub><sup>=</sup> p'| '^ 2 \ l (28) (29) while the other components of equation (27) are the same as those previously specified in equations (5-26).
As before, equation (27) can be simplified using a single MC matrix<sup>(2) </sup>prediction, v, = sc}<sup>2)</sup>MC<sup><?></sup> where
MC<sup><2></sup> = n
C<sup>(L) </sup>ii<sup>(2) </sup>c<sup>(I></sup> scl<sup>21</sup> - [1 s. sc, p<sup>2</sup> sc<sup>2</sup> (31) [0053] By collecting all p pixels together, a simplified matrix expression can be obtained
V = SC<sup>i2></sup> · MC<sup>|2)</sup> , (32) where
<td>V =</td><td>* 0 V1</td><td>SC<sup>!and</sup>> =</td><td>Γ Ρΐ Ί</td><td> (33)</td>
<td></td><td></td><td></td><td></td><td></td>
<td></td><td></td><td></td><td>Stick</td><td></td>
and SC<sup>(2)</sup> is the matrix /? x (l + 2 * 7) and the same least squares solutions as described earlier can be used.
[0054] Similarly, you can also build third-order or higher-order models using vector product parameters. Alternatively, as described in chapter 5.4.3 "Digital Color Imaging Handbook", CRC Press, 2002, edited by Gaurava Sharma, the representation of the K order of the MMR vector product model can also be described by the following wording:
κ a · Jf jf-0, a0 _'_ 0 £ KR
<img file="PL3324622T3_D0005.tif" />
Χ = ΰ 1 “V 2 = 0 a
κ r κ
J, _O v = o 2 = 0 (34) (35) (36) where K is the highest order of the MMR predictor.
Spatial extension based on MMR (MMR-CS) [0055] In all MMR models described so far, the value of the predicted pixel Vj depends only on the corresponding, usually occurring together, input values of Si. Based on MMR prediction, you can also benefit from taking into account data from neighboring pixels. This approach corresponds to the inclusion in the MMR model of any linear input data processing in the spatial domain, such as FIR filtering.
[0056] If all eight possible neighboring pixels are included in the image, this approach can add on our matrix M predictions to eight more first-order variables per color component. However, in practice, it is usually enough to add prediction variables corresponding to two horizontal and two vertical adjacent pixels and do not consider diagonal neighbors. On the prediction matrix, you add up to four variables for the color component, i.e. those corresponding to the top, left, bottom and right pixels. Similarly, you can also add parameters corresponding to a higher order of values of an adjacent pixel.
[0057] To simplify the complexity and computational requirements of such an MMR spatial model, you can consider adding spatial extensions to traditional models only for a single color component, such as the luminance component (as in the luma-chroma color representation) or the green component (as in the RGB representation) . For example, assuming the addition of spatial pixel prediction for only the green color component, from equations (34-36), the general expression for the prediction of the value of the green output pixel would be λ r λ ν + '- ΣΣΣ' »,. ™
First order model with spatial extension (MMR-1S) [0058] In another embodiment, the first order MMR model (MMR-1) of equation (4) can be reconsidered, but now extended to include spatial extensions in one or more of color components; For example, if you use up to four adjacent pixels of each pixel in the first color component:
and, = sti D<sup>(l</sup>'+ 5, ™ ^' + n, (38) where is the 3x3 matrix, an is the 1x3 vector, both as specified in equation (5),
D<sup>(L></sup> =
0-1) 1 G 1) 1 ^ ίιν)!
(39) where m in equation (39) means the number of columns in the input frame with columns m and rows n or the total number of pixels mxn = p. Equation (39) can easily be expanded to apply these methods to both other color components and variant configurations of adjacent pixels.
[0059] Using the same approaches as before, equation (38) can easily be expressed as a system of linear equations,
V = SD'MD<sup>(L)</sup> , (40) which can be solved as described earlier.
Application to VDR signals with more than three primary colors [0060] All of the proposed MMR prediction models can easily be extended to signal spaces with more than three primary colors. As an example, you can think about the case in which the SDR signal has three primary colors, say RGB, but the VDR signal is defined in the P6 color space, with six basic colors. In this case equations (1 - 3) can be rewritten as<sup>s</sup>i = k (41) <sup>v</sup>f = [<sup>v</sup>n hi h «heb (<sup>42</sup>>
and <sup>C</sup>'i2 <sup>L</sup>’14 (<sup>43</sup> [0061] As before, the number of pixels in one jakop color component should be determined. Given now the first order MMR prediction model (MMR-1) from equation (4), (44) is now a 3x6 matrix, an is a 1x6 vector defined by
<td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td> =</td><td></td><td></td><td></td><td></td><td></td><td></td><td> (45)</td>
<td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>and</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>n =</td><td>[li</td><td>L2 O</td><td> «14</td><td>l</td><td>"id-</td><td></td><td> (46)</td>
[0062] Equation (41) can be expressed using a single matrix M<sup>(1)</sup> prediction as (47) where
M<sup>(1)</sup>
<td> «11</td><td> «12</td><td> «13</td><td> «14</td><td> «15</td><td> «16</td>
<td></td><td></td><td></td><td>«1P *</td><td></td><td><sup>λ</sup>16</td>
<td></td><td><sup>m</sup>22</td><td><sup>m</sup>23</td><td> (1)</td><td><sup>m</sup>25</td><td> «*26</td>
<td> (1) <sup>m</sup>3I</td><td><sup>m</sup>32</td><td></td><td><sup>m</sup>34</td><td> (1) <sup>m</sup>35</td><td>«* Ίβ</td>
a <= [1 M ^ 2 sj. (48) [0063] By collecting all p pixels together, this prediction problem can be described as
V = SM<sup>loam</sup>* (49)
<img file="PL3324622T3_D0006.tif" />
is the px6 matrix, «here is the pxĄ matrix, and M<sup>(1)</sup> is a matrix [0064] Similarly, higher order MMR prediction models can be extended, and solutions for the prediction matrix can be obtained using the methods outlined above.
SAMPLE PROCESS OF MULTI-CHANNEL REGRESSION PREDICTION [0065] FIG. 4 depicts an example multi-channel multiple regression prediction process according to an exemplary embodiment of the present invention.
[0066] The process begins at step 410, in which a predictor, such as the predictor 250, receives the VDR and SDR input signals. Having these two input signals, in step 420, the predictor decides which MMR model to choose. As described earlier, the predictor can choose from many MMR models, including (among others): first order (MMR-1), second order (MMR-2), third or higher order, first order with vector product (MMR-1C) , second order with vector product (MMR-2C), third or higher order with vector product, or any of the above models with added spatial extensions.
[0067] The MMR model can be selected using a variety of methods that take into account a number of criteria, including: prior knowledge of input SDRs and VDRs, available computational and memory resources, and target coding efficiency. FIG. 5 illustrates an exemplary implementation of step 420 based on the requirement for a lower balance than a predetermined threshold.
[0068] As described earlier, each MMR model can be represented as a set of linear equations in the form
V = SM, (50) where M is a prediction matrix.
[0069] In step 430, M can be solved using a number of numerical methods. For example, with the constraint imposing a minimization of the mean square of the remainder between V and its V estimation,
M = (S<sup>r</sup>S) 'S<sup>r</sup>V. (51) [0070] Finally, in step 440, using equation (50), the predictor outputs to output V and M.
[0071] FIG. 5 shows an example process 420 in terms of selecting the MMR model during prediction. The predictor 250 may start at step 510 with the initial MMR model, such as the one that was used in the previous frame or shot, for example the second order model (MMR-2) or the simplest possible model such as MMR1. When resolved relative to M, at step 520, the predictor calculates a prediction error between the input V and its predictively determined value. At step 530, if the prediction error is less than a given threshold, then the predictor selects the existing model and stops the dialing process (540), otherwise, at step 550, it examines whether to use a more complex model. For example, if the current model is MMR-2, the predictor may decide to use MMR-2-C or MMR-2-CS. As described earlier, this decision may depend on a number of criteria, including prediction error values, processing power requirements, and target coding efficiency. If a more complex model is possible, the new model is selected in step 560, and the method returns to step 520. Otherwise, the predictor will use the existing model (540).
[0072] The prediction process 400 can be repeated in as many intervals as necessary to maintain coding efficiency while using available computing resources. For example, when encoding video signals, the process 400 may be repeated based on a predetermined video slice size, for each frame, group of frames, or whenever a prediction residual exceeds a specified threshold.
[0073] During the prediction process 400, all available input pixels or a sub-sample of these pixels can also be used. In one embodiment, pixels can only be used from each row of pixels k and each column of k k input data, where k is an integer equal to or greater than two. In another exemplary implementation, you may decide to omit input pixels that are below a specific clipping threshold (e.g., very close to zero), or pixels that are above a specific saturation threshold (e.g., for 'bit data, pixel values that are very close to 2<sup>n</sup>-1.) In yet another implementation, a combination of such subsampling and thresholding techniques can be used to reduce the pixel sample size and adjust the computational constraints of a particular implementation.
IMAGE DECODING [0074] Embodiments of the invention may be implemented either on an image encoder or on an image decoder. FIG. 6 shows an embodiment of a decoder 150 according to an embodiment of the present invention.
The decoding system 600 receives an encoded bit stream that can combine the base layer 690, the optional extension layer (or residual) 665 and metadata 645, which are extracted after decompressing 630 and various inverse transformations 640. For example, in a VDR-SDR system, the base layer 690 may represent an SDR representation of an encoded signal, and metadata 645 may include information about an MMR prediction model that uses an encoder predictor 250, and corresponding prediction parameters. In one embodiment, while the encoder uses the MMR predictor according to the methods of the invention, the metadata may contain information identifying the model used (e.g., MMR-1, MMR-2, MMR-2C and the like) and all matrix coefficients associated with that particular model . Having a 690 s base layer and MMR-related color parameters extracted from 645 metadata, the predictor 650 can calculate the predictive V 680 using any of the appropriate equations described in this document. For example, if the specified model is MMR-2C, then V 680 can be calculated using equation (32). In the absence of a remainder or when the remainder is negligible, the predetermined value of 680 can be output directly as the final VDR image. Otherwise, in the summation element 660, to output the VDR 670 signal, the prediction result (680) is added to the rest 665.
EXAMPLE IMPLEMENTATION OF THE COMPUTER SYSTEM [0075] Embodiments of the present invention can be implemented using a computer system, systems configured in an electrical circuits and components, devices with an integrated circuit (IC) such as a microcontroller, directly programmable gate matrix (FPGA) ) (or field programmable gate array) or another configurable or programmable electronic system (PLD) programmable logic device), a digital signal processor (DSP) or a discrete time processor, specialized IC (ASIC) or devices that contain one or more of such systems, devices or components. The computer or IC can operate, control commands, or execute prediction commands based on MMR, such as those described in this document. The computer and / or IC system may calculate any of many parameters or any of many values that relate to MMR prediction as described in this document. Embodiments of expanding the range of image and video dynamics can be implemented in hardware, software, firmware and their various combinations.
[0076] Some embodiments of the invention include computer processors that execute program instructions that cause the processors to perform the method of the invention. For example, one or more processors in the display, encoder, STB decoder, transcoder or the like may implement MMR prediction methods as described above by executing program instructions in program memory to which the processors have access. The invention may also be in the form of a product in the form of a program. The program product may contain any medium that stores a set of computer readable signals containing instructions that, when the processor executes them, will cause the processor to perform the method of the invention. Products in the form of a program according to the invention may take the form of any of the many forms available. The program product may contain, for example, physical media such as magnetic storage media, including floppy disks, hard disk drives, optical storage media, including CD ROMs, DVDs, electronic data storage media, including memory only for reading ROM, flash direct RAM memory and the like. Computer-readable signals appearing in the product in the form of a program can be optionally compressed or encrypted.
[0077] Where reference is made above to a subassembly (e.g., software module, processor, assembly, device, circuit, and the like), unless otherwise indicated, reference to that subassembly (including reference to "means") shall be interpreted as including the equivalents of this sub-assembly, i.e. any sub-assembly that performs the function of the described sub-assembly (for example, which is functionally equivalent), including sub-assemblies, which in construction do not correspond to the disclosed structure which performs the function in the exemplary embodiments of the invention.
EQUIVALENTS, EXTENSIONS, VARIANT AND MISCELLANEOUS SOLUTIONS [0078] Thus, exemplary embodiments that relate to the use of MMR prediction in the encoding of VDR and SDR images are described. In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Thus, the sole and exclusive indicator of what an invention is and what the applicants intend to be an invention is a set of claims that result from this application, in a particular form in which such reservations occur, including any subsequent correction. Any definitions expressly set out in this document regarding the expressions contained in such claims determine the meaning of such expressions used in the claims. Thus, any restriction, element, property, feature, advantage or attribute that is not explicitly specified in the claim should in no way limit the scope of such claim. Accordingly, the nature of the description and the figure should be regarded as demonstrative rather than limiting.
14 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14
69 members in 11 offices
Priority claims16
| Document | Office | Kind | Date |
|---|---|---|---|
| 201161475359 | United States of America | P | |
| 201161475359 | United States of America | P | |
| 12717990 | European Patent Office (EPO) | A | |
| 12717990 | European Patent Office (EPO) | A | |
| 14171538 | European Patent Office (EPO) | A | |
| 14171538 | European Patent Office (EPO) | A | |
| 17204552 | European Patent Office (EPO) | A | |
| 2012033605 | United States of America | W | |
| 2012033605 | United States of America | W | |
| 172045528 | – | – | – |
| 201161475359P | – | – | – |
| EP20120717990 | – | – | – |
| EP20140171538 | – | – | – |
| EP20170204552 | – | – | – |
| US201161475359P | – | – | – |
| WO2012US33605 | – | – | – |
Members69
| Document | Office | Kind | |
|---|---|---|---|
| WO2012142471A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2012142506A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2013112532A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2013112532A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2012142471A8 | World Intellectual Property Organization (WIPO) | A8 | |
| CN103493489A | China | A | |
| US2014029675A1 | United States of America | A1 | |
| CN103563372A | China | A | |
| US2014037205A1 | United States of America | A1 | |
| EP2697971A1 | European Patent Office (EPO) | A1 | |
| EP2697972A1 | European Patent Office (EPO) | A1 | |
| US8731287B2 | United States of America | B2 | |
| US2014185930A1 | United States of America | A1 | |
| US8811490B2 | United States of America | B2 | |
| JP2014520414A | Japan | A | |
| US8837825B2 | United States of America | B2 | |
| EP2782348A1 | European Patent Office (EPO) | A1 | |
| HK1193688A1 | Hong Kong, China | A1 | |
| US2014307796A1 | United States of America | A1 | |
| EP2807823A2 | European Patent Office (EPO) | A2 | |
| US2014369409A1 | United States of America | A1 | |
| EP2697972B1 | European Patent Office (EPO) | B1 | |
| US8971408B2 | United States of America | B2 | |
| US2015092850A1 | United States of America | A1 | |
| EP2697971B1 | European Patent Office (EPO) | B1 | |
| JP5744318B2 | Japan | B2 | |
| US2015222916A1 | United States of America | A1 | |
| JP2015165665A | Japan | A | |
| EP2945377A1 | European Patent Office (EPO) | A1 | |
| HK1204741A1 | Hong Kong, China | A1 | |
| JP5921741B2 | Japan | B2 | |
| US9386313B2 | United States of America | B2 | |
| HK1214049A1 | Hong Kong, China | A1 | |
| US9420302B2 | United States of America | B2 | |
| JP2016167834A | Japan | A | |
| US2016269756A1 | United States of America | A1 | |
| CN103493489B | China | B | |
| US9497475B2 | United States of America | B2 | |
| US2017034521A1 | United States of America | A1 | |
| CN103563372B | China | B | |
| CN106878707A | China | A | |
| US9699483B2 | United States of America | B2 | |
| CN107105229A | China | A | |
| US2017264898A1 | United States of America | A1 | |
| JP6246255B2 | Japan | B2 | |
| EP2782348B1 | European Patent Office (EPO) | B1 | |
| US9877032B2 | United States of America | B2 | |
| ES2659961T3 | Spain | T3 | |
| TR201802291T4 | Türkiye | T4 | |
| EP2807823B1 | European Patent Office (EPO) | B1 | |
| JP2018057019A | Japan | A | |
| PL2782348T3 | Poland | T3 | |
| CN106878707B | China | B | |
| EP3324622A1 | European Patent Office (EPO) | A1 | |
| ES2670504T3 | Spain | T3 | |
| US10021390B2 | United States of America | B2 | |
| PL2807823T3 | Poland | T3 | |
| US2018278930A1 | United States of America | A1 | |
| EP2945377B1 | European Patent Office (EPO) | B1 | |
| US10237552B2 | United States of America | B2 | |
| JP6490178B2 | Japan | B2 | |
| PL2945377T3 | Poland | T3 | |
| EP3324622B1 | European Patent Office (EPO) | B1 | |
| DK3324622T3 | Denmark | T3 | |
| CN107105229B | China | B | |
| PL3324622T3This record | Poland | T3 | |
| HUE046186T2 | Hungary | T2 | |
| ES2750234T3 | Spain | T3 | |
| CN107105229B9 | China | B9 |
Numbers
- Publication
- 3324622
- Publication, DOCDB
- 3324622
- Publication, EPODOC
- PL3324622T
- Application
- 17204552
- Application, DOCDB
- 17204552
- Application, EPODOC
- PL20170204552T
Titles2
- English
- MULTIPLE COLOR CHANNEL MULTIPLE REGRESSION PREDICTOR
- Polish
- PREDYKTOR REGRESJI WIELOKROTNEJ KANAŁU WIELU BARW
Classification
- CPC, 8
- H04N19/103
- H04N19/105
- H04N19/147
- H04N19/192
- H04N19/30
- H04N19/16
- H04N19/98
- G06F17/18
- IPC, 4
- H04N19 105
- H04N19 147
- H04N19 192
- H04N19 30
