Bit-depth scalability
Abstract
This record has no abstract on file.
Term
1.6 yearsto projected expiry
Projected expiry 16 April 2028, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1Zastrzeżenia patentowe 1. Koder do kodowania danych źródłowych (160) obrazu lub wideo do strumienia danych (112) o skalowalnej jakości, zawierający:środki (102) kodowania bazowego, do kodowania danych źródłowych (160) obrazu lub wideo, do strumienia danych kodowania bazowego reprezentującego reprezentację danych źródłowych obrazu lub wideo z pierwszą głębią bitową próbki obrazu;środki (104) mapowania, do mapowania próbek reprezentacji źródłowych danych (160) obrazu lub wideo z pierwszą głębią bitową próbki obrazu, z pierwszego zakresu dynamicznego, odpowiadającego pierwszej głębi bitowej próbki obrazu, do drugiego zakresu dynamicznego, większego od pierwszego zakresu dynamicznego i odpowiadającego drugiej głębi bitowej próbki obrazu, wyższej od pierwszej głębi bitowej próbki obrazu, poprzez użycie jednej lub większej liczby globalnych funkcji mapowania, stałych w ramach źródłowych danych (160) obrazu lub wideo lub zmiennych przy pierwszej granulacji, i lokalnej funkcji mapowania, lokalnie modyfikującej jedną lub większą liczbę globalnych funkcji mapowania, przy drugiej granulacji, drobniejszej od pierwszej granulacji, dla uzyskania predykcji danych źródłowych obrazu lub wideo, mających drugą głębię bitową próbki obrazu;środki (106) kodowania resztkowego, do kodowania resztki predykcji, predykcji, do strumienia danych warstwy poprawy jakości głębi bitowej;oraz środki (108) łączenia, do formowania strumienia danych o skalowalnej jakości w oparciu o strumień danych kodowania bazowego, lokalną funkcję mapowania i strumień danych wzbogacenia głębi bitowej tak, że lokalna funkcja mapowania może być pozyskiwana ze strumienia danych o skalowalnej jakości. 2. Dekoder do dekodowania strumienia danych o skalowalnej jakości, w którym zakodowane są źródłowe dane obrazu lub wideo, przy czym źródłowe dane obrazu lub wideo zawierają strumień danych warstwy bazowej reprezentujący źródłowe dane obrazu lub wideo o pierwszej głębi bitowej próbki obrazu, strumień danych warstwy poprawy jakości głębi bitowej reprezentujący resztkę predykcji o drugiej głębi bitowej próbki obrazu, wyższej od pierwszej głębi bitowej próbki obrazu i lokalną funkcję mapowania zdefiniowaną z drugą granulacją, przy czym dekoder zawiera: środki (204) do dekodowania strumienia danych warstwy bazowej do zrekonstruowanych danych obrazu lub wideo o niższej głębi bitowej;środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej do resztki predykcji;środki (206) do mapowania próbek zrekonstruowanych danych obrazu lub wideo o niższej głębi bitowej z pierwszą głębią bitową próbki obrazu, z pierwszego zakresu dynamicznego, odpowiadającego pierwszej głębi bitowej próbki obrazu, do drugiego zakresu dynamicznego, większego od pierwszego zakresu dynamicznego i odpowiadającego drugiej głębi bitowej próbki obrazu, poprzez użycie jednej lub większej liczby globalnych funkcji mapowania, stałych w wideo lub zmiennych przy pierwszej granulacji, oraz lokalnej funkcji mapowania, lokalnie modyfikującej jedną lub większą liczbę globalnych funkcji mapowania przy drugiej granulacji, mniejszej od pierwszej granulacji, dla uzyskania predykcji danych źródłowych obrazu lub wideo, mających drugą głębię bitową próbki obrazu;oraz środki (210) do rekonstrukcji obrazu z drugą głębią bitową próbki obrazu w oparciu o predykcję i resztkę predykcji. 3. Dekoder według zastrz. 2, w którym środki (206) mapowania są przystosowane do mapowania próbek zrekonstruowanych danych obrazu lub wideo o niższej głębi bitowej z pierwszą głębią bitową próbki obrazu, poprzez użycie kombinowanej funkcji mapowania, będącej arytmetyczną kombinacją jednej z jednej lub większej liczby globalnych funkcji mapowania i lokalnej funkcji mapowania. 4. Dekoder według zastrz. 3, w którym więcej niż jedna globalna funkcja mapowania jest użyta przez środki (206) mapowania i środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej są przystosowane do identyfikacji jednej, z więcej niż jednej globalnej funkcji mapowania, ze strumienia danych wzbogacenia głębi bitowej. 5. Dekoder według zastrz. 3 albo 4, w którym arytmetyczna kombinacja obejmuje operację dodawania. 6. Dekoder według dowolnego z zastrz. 2 do 5, w którym środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej i środki (104) mapowania są przystosowane tak, że druga granulacja dzieli źródłowe dane (160) obrazu lub wideo na wiele bloków (170) obrazu, a środki (206) mapowania są przystosowane tak, że lokalną funkcją mapowania jest m-s+n, z m i n zmieniającymi się przy drugiej granulacji, przy czym środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej są przystosowane do pozyskiwania m i n ze strumienia danych wzbogacenia głębi bitowej dla każdego bloku (170) obrazu, źródłowych danych (160) obrazu lub wideo, tak że m i n mogą się różnić między wieloma blokami obrazu. 7. Dekoder według dowolnego z zastrz. 2 do 6, w którym środki (206) mapowania są przystosowane tak, że druga granulacja zmienia się w źródłowych danych (160) obrazu lub wideo, a środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej są przystosowane do pozyskiwania drugiej granulacji ze strumienia danych wzbogacenia głębi bitowej. 8. Dekoder według dowolnego z zastrz. 2 do 7, w którym środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej i środki (206) mapowania są przystosowane tak, że druga granulacja dzieli źródłowe dane (160) obrazu lub wideo na wiele bloków (170) obrazu, a środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej są przystosowane do pozyskiwania resztki (ńm, ńn) lokalnej funkcji mapowania ze strumienia danych wzbogacenia głębi bitowej dla każdego bloku (170) obrazu i pozyskiwania lokalnej funkcji mapowania ustalonego bloku obrazu źródłowych danych (160) obrazu lub wideo, poprzez użycie predykcji przestrzennej i/lub czasowej z jednego lub większej liczby sąsiednich bloków obrazu lub odpowiedniego bloku obrazu, obrazu, danych źródłowych obrazu lub wideo poprzedzającego obraz, do którego ustalony blok obrazu należy i resztki lokalnej funkcji mapowania ustalonego bloku obrazu. 9. Dekoder według dowolnego z zastrz. 2 do 8, w którym środki (206) mapowania są przystosowane tak, że co najmniej jedna z jednej lub większej liczby globalnych funkcji mapowania jest nieliniowa. 10. Dekoder według dowolnego z zastrz. 2 do 9, w którym środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej są przystosowane do pozyskiwania co najmniej jednej z jednej lub większej liczby globalnych funkcji mapowania ze strumienia danych wzbogacenia głębi bitowej. 11. Dekoder według dowolnego z zastrz. 2 do 10, w którym środki (206) mapowania są przystosowane tak, że co najmniej jedna z globalnych funkcji mapowania jest zdefiniowana jako 2 M-N-K x + 2 M-1 - 2 N-1-K , gdzie x jest próbką reprezentacji danych źródłowych obrazu lub wideo z pierwszą głębią bitową próbki obrazu, N jest pierwszą głębią bitową próbki obrazu, M jest drugą głębią bitową próbki obrazu, a K jest parametrem mapowania, 2 M-N-K x + D, gdzie N jest pierwszą głębią bitową próbki obrazu, M jest drugą głębią bitową próbki obrazu, a K i D są parametrami mapowania, floor (2 M-N-K x + 2 M-2N-K x + D), gdzie floor (a) zaokrągla w dół do najbliższej liczby całkowitej, N jest pierwszą głębią bitową próbki obrazu, M jest drugą głębią bitową próbki obrazu, a K i D są parametrami mapowania, liniowa funkcja dla części do mapowania próbek z pierwszego zakresu dynamicznego do drugiego zakresu dynamicznego z informacją punktu interpolacji definiującą liniowe mapowanie dla części, albo tablica przeglądowa do zindeksowania poprzez użycie próbek z pierwszego zakresu dynamicznego i wyprowadzanie próbek drugiego zakresu dynamicznego. 12. Dekoder według zastrz. 11, w którym środki (208) do dekodowania strumienia danych wzbogacenia głębi bitowej są przystosowane do pozyskiwania parametru (parametrów) mapowania, informacji punktu interpolacji lub tablicy przeglądowej, ze strumienia danych wzbogacenia głębi bitowej. 13. Sposób kodowania źródłowych danych (160) obrazu lub wideo do strumienia danych (112) o skalowalnej jakości, obejmujący: kodowanie źródłowych danych (160) obrazu lub wideo do strumienia danych kodowania bazowego, reprezentującego reprezentację danych źródłowych obrazu lub wideo z pierwszą głębią bitową próbki obrazu;mapowanie próbek reprezentacji źródłowych danych (160) obrazu lub wideo z pierwszą głębią bitową próbki obrazu z pierwszego zakresu dynamicznego odpowiadającego pierwszej głębi bitowej próbki obrazu, do drugiego zakresu dynamicznego, większego od pierwszego zakresu dynamicznego i odpowiadającego drugiej głębi bitowej próbki obrazu, wyższej od pierwszej głębi bitowej próbki obrazu, poprzez użycie jednej lub większej liczby globalnych funkcji mapowania, stałych w ramach źródłowych danych (160) obrazu lub wideo, lub zmiennych przy pierwszej granulacji i lokalnej funkcji mapowania, lokalnie modyfikującej jedną lub większą liczbę globalnych funkcji mapowania, przy drugiej granulacji, drobniejszej od pierwszej granulacji, dla uzyskania predykcji danych źródłowych obrazu lub wideo, mających drugą głębię bitową próbki obrazu;kodowanie resztki predykcji, predykcji, do strumienia danych warstwy poprawy jakości głębi bitowej;formowanie strumienia danych o skalowalnej jakości w oparciu o strumień danych kodowania bazowego, lokalną funkcję mapowania i strumień danych wzbogacenia głębi bitowej tak, że lokalna funkcja mapowania może być pozyskiwana ze strumienia danych o skalowalnej jakości. 14. Sposób dekodowania strumienia danych o skalowalnej jakości, w którym zakodowane są źródłowe dane obrazu lub wideo, przy czym strumień danych o skalowalnej jakości zawiera strumień danych warstwy bazowej reprezentujący źródłowe dane obrazu lub wideo z pierwszą głębią bitową próbki obrazu, strumień danych warstwy poprawy jakości głębi bitowej reprezentujący resztkę predykcji z drugą głębią bitową próbki obrazu, wyższą od pierwszej głębi bitowej próbki obrazu, lokalną funkcję mapowania zdefiniowaną przy drugiej granulacji, przy czym sposób obejmuje: dekodowanie strumienia danych warstwy bazowej do zrekonstruowanych danych obrazu lub wideo o niższej głębi bitowej;dekodowanie strumienia danych wzbogacenia głębi bitowej do resztki predykcji;mapowanie próbek zrekonstruowanych danych obrazu lub wideo o niższej głębi bitowej z pierwszą głębią bitową próbki obrazu, z pierwszego zakresu dynamicznego odpowiadającego pierwszej głębi bitowej próbki obrazu do drugiego zakresu dynamicznego, większego od pierwszego zakresu dynamicznego i odpowiadającego drugiej głębi bitowej próbki obrazu, poprzez użycie jednej lub większej liczby globalnych funkcji mapowania, stałych w wideo lub zmiennych przy pierwszej granulacji, oraz lokalnej funkcji mapowania, lokalnie modyfikującej jedną lub większą liczbę globalnych funkcji mapowania przy drugiej granulacji, mniejszej od pierwszej granulacji, dla uzyskania predykcji danych źródłowych obrazu lub wideo, mających drugą głębię bitową próbki obrazu;oraz rekonstrukcję obrazu z pierwszą głębią bitową próbki obrazu w oparciu o predykcję i resztkę predykcji. 15. Strumień danych o skalowalnej jakości, w którym zakodowane są źródłowe dane obrazu lub wideo, przy czym strumień danych o skalowalnej jakości zawiera strumień danych warstwy bazowej, reprezentujący strumień danych warstwy bazowej reprezentujący zrekonstruowane dane źródłowe obrazu lub wideo z pierwszą głębią bitową próbki obrazu, strumień danych warstwy poprawy jakości głębi bitowej, reprezentujący resztkę predykcji z drugą głębią bitową próbki obrazu, większą od pierwszej głębi bitowej próbki obrazu i lokalną funkcję mapowania zdefiniowaną przy drugiej granulacji, przy czym rekonstrukcja obrazu z drugą głębią bitową próbki obrazu może być pozyskiwana z resztki predykcji i predykcji uzyskanych przez mapowanie próbek zrekonstruowanych danych obrazu lub wideo o niższej głębi bitowej, z pierwszą głębią bitową próbki obrazu, z pierwszego zakresu dynamicznego, odpowiadającego pierwszej głębi bitowej próbki obrazu, do drugiego zakresu dynamicznego, większego od pierwszego zakresu dynamicznego i odpowiadającego drugiej głębi bitowej próbki obrazu, poprzez użycie jednej lub większej liczby globalnych funkcji mapowania, stałych w wideo lub zmiennych przy pierwszej granulacji i lokalnej funkcji mapowania, lokalnie modyfikującej jedną lub większą liczbę globalnych funkcji mapowania przy drugiej granulacji, mniejszej od pierwszej granulacji. 16. Program komputerowy zawierający kod programu do realizacji, gdy jest uruchomiony w komputerze, sposobu określonego w zastrz. 13 albo 14. Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e.V., Niemcy Pełnomocnik EP 2 279 622 B1 Z-12825 EP 2 279 622 B1 Z-12825 FIG 2 EP 2 279 622 B1 Z-12825 154> próbkowania w gorę Γ FIG3 FIG 4 EP 2 279 622 B1 Z-12825 162 FIG 5 EP 2 279 622 B1 Z-12825 i ustawianie indeksu k dla każdego bloku drobnej granulacji FIG 6 EP 2 279 622 B1 Z-12825 282 284 286 288 280 FIG 7 EP 2 279 622 B1 Z-12825 918b
114 paragraphs in 4 sections, as filed
[0001] The present invention relates to image and / or video coding, and in particular to scalable quality coding that provides bit depth scalability using scalable quality data streams.
[0002] The Joint Video Team (JVT) from the ISO / IEC Moving Pictures Experts Group (MPEG) and ITU-T Video Coding Experts Group (VCEG) has recently finalized the scalable extension of the prior art H.264 / AVC video coding standard, called Scalable Video Coding (SVC). SVC supports time, spatial and scalable SNR (signal-to-noise ratio) coding of video sequences or any combination thereof.
[0003] H.264 / AVC described in ITU-T Rec. & ISO / IEC 14496-10 AVC, "Advanced Video Coding for Generic Audiovisual Services" version 3, 2005, defines a hybrid video codec in which the macroblock prediction signals are either generated in the time domain by motion-compensated prediction, or in the spatial domain by Intra prediction, and both are followed by residual coding. A bit-to-distortion ratio comparable to a single-layer H.264 / AVC means that the same visual playback quality is typically achieved at 10% bit rate. Considering the above, scalability is considered to be a function for removing part of the bit stream while obtaining a ratio of bit rate to distortion at every supported spatial, temporal or SNR resolution that is comparable to single-layer H.264 / AVC coding at a given resolution. [0004] The basic construction of scalable video coding (SVC) can be classified as a layered video codec. In each layer used basic prediction concepts with motion compensation and Intra prediction, as in H.264 / AVC. However, additional inter-layer prediction mechanisms have been integrated to utilize redundancy between several spatial layers or SNRs. SNR scalability is basically obtained by residual quantization, while for spatial scalability a combination of motion prediction and compensation is used for decomposition of the oversampled pyramid. The temporal scalability approach with H.264 / AVC is maintained.
[0005] Generally, the encoder structure depends on the scalability space that is required by the application. For example, Fig. 8 shows an encoder structure 900 with two spatial layers 902a, 902b. Each layer uses an independent hierarchical structure 904a, b motion compensation prediction, with parameters 906a, b layer-specific motion. The redundancy between successive layers 902a, b is used by the interlayer prediction concepts 908, which include prediction mechanisms for parameters 906a, motion b, as well as texture data 910a, b. The base representation of 912a, b input images 914a, b of each layer 902a, b is obtained by transform coding 916a, b, similar to H.264 / AVC encoding, the corresponding NAL (Network Abstraction Layer) units contain traffic information and data textures, NALs of the base representation of the lowest layer, i.e. 912a are compatible with single-layer H.264 / AVC.
[0006] The resulting bit streams output by coding 916a, b of the base layer and progressive coding 918a, respectively, b textures improving the SNR of the respective layers 902a, b are multiplexed by the multiplexer 920 to obtain the scalable bit stream 922. This scalable bit stream 922 is scalable in SNR time, space and quality.
[0007] In summary, according to the above scalable extension of Video Coding Standard H.264 / AVC, temporal scalability is provided by using a hierarchical prediction structure. For this hierarchical prediction structure, one of the single-layer H.264 / AVC standards can be used without any changes. For spatial scalability or SNR, additional tools must be added to the single-layer H.264 / MPEG4.AVC, as described in the SVC extension of the H.264 / AVC codec. All three types of scalability are combined to generate a bit stream that supports a large degree of combined scalability.
[0008] Problems arise when the video source signal has a different dynamic range than that required by the decoder or player, respectively. In the current SVC standard above, the scalability tools are only specified for cases where both the base layer and the enrichment layer represent the given video source at the same bit depth of the respective luma (brightness) and / or chroma (color) sample arrangements. Therefore, taking into account different decoders and players requiring different bit depths, there would have to be several coding streams separately for each bit depth. However, in terms of bit rate / distortion, this means an increase in overhead and performance, respectively.
[0009] It has already been proposed to add bit depth scalability to the SVC standard. For example, Shan Liu et al. describes in the document introduced to JVT, namely JVT-X075 - the possibility of obtaining inter-layer prediction from a representation with a lower bit depth of the base layer, by using inverted tone mapping, according to which the inter-layer predicted or inverse mapped p 'pixel value is calculated from the p value<sub>b</sub> base layer pixels from the formula p '= p<sub>b</sub> scale + offset with the statement that inter-layer prediction is implemented on macroblocks or on smaller block sizes. In JVTY067, Shan Liu presents the results for this method of inter-layer prediction. Similarly, Andrew Segall et al. proposes in the JVT-X071 document inter-layer prediction for bit depth scalability, according to which the gain and shift operation for inverted tone mapping is used. The gain parameters are indexed and transmitted in the bit stream of the quality improvement layer, block by block. Signaling of scalable coefficients and displacement coefficients is provided by a combination of prediction and purification. In addition, it has been described that high-level syntax supports more coarse granulation than block-by-block transfer. We also refer to the Andrew Segall document "Scalable Coding of High Dynamic Range Video" in ICIP 2007, from I-1 to I-4 and JVT, JVT-X067 and JVT-W113 documents also from Andrew Segall.
[0010] Although the aforementioned suggestions for using inverse tone mapping to obtain a prediction from a lower bit depth base layer remove some of the redundancy between lower bit depth information and higher bit depth information, it would be beneficial to achieve even better delivery performance for such a bit stream, scalable relative to bit depth, in particular in terms of bit rate to distortion.
[0011] It is an object of the present invention to provide a coding method that provides a more efficient method of providing image or video coding suitable for different bit depths.
[0012] This object is achieved by the encoder as defined in claim 1, the decoder as defined in claim 11, the method as defined in claim 22 or 23, or the scalable quality data stream as defined in claim 24.
[0013] The present invention is based on the finding that the performance of a scalable bit depth data stream can be increased when interlayer prediction is obtained by mapping samples of an image or video source data representation with a first bit depth of the image sample from the first dynamic range corresponding to the first bit depth of the image sample, to the second dynamic range, greater than the first dynamic range and corresponding to the second bit depth of the image sample, which is greater than the first bit depth of the image sample, by using one or more global mapping functions that are fixed within the image or video source data or variables with first granulation and local a mapping function that locally modifies one or more global mapping functions and variables with a second granulation, smaller than the first granulation, with forming a scalable quality data stream based on the local mapping function such that the local mapping function can be obtained from a scalable quality data stream. Although providing one or more global mapping functions in addition to the local mapping function, which in turn modifies one or more global mapping functions locally, at first glance increases the amount of additional information in the scalable data stream, this increase is more than offset by the fact that this division into a global mapping function on the one hand and a local mapping function on the other hand allows that the local mapping function and its parameters for its parameterization can be small and hence can be coded in a very efficient way. The global mapping function can be encoded in a scalable quality data stream and because it is constant within the image or video source data or varies with greater granulation, the overhead or flexibility for defining this global mapping function can be increased so that this global function the mapping could exactly match the average statistics of the image or video source data, thereby further reducing the local mapping function module.
[0014] Preferred embodiments of the present invention are described below with reference to the figures. In particular
Fig. 1 is a schematic diagram of a video encoder according to an embodiment of the present invention;
Fig. 2 is a diagram of a video decoder according to an embodiment of the present invention;
Fig. 3 is a flowchart of potential implementation of the mode of operation of the prediction module 134 in Fig. 1, according to an embodiment;
Fig. 4 is a diagram of the video and its division into image sequences, images, macroblock pairs, macroblocks and transformation blocks according to an embodiment;
Fig. 5 is a diagram of a portion of an image divided into blocks according to the fine granulation underlying the local mapping / adaptation function while illustrating a method of predictive coding for encoding the parameters of the local mapping / adaptation function according to an embodiment;
Fig. 6 is a flowchart to illustrate a method of inverted tone mapping in an encoder according to an embodiment;
Fig. 7 is a flowchart of a reverse tone mapping method in a decoder corresponding to the decoder in Fig. 6, according to an embodiment; and Fig. 8 is a block diagram of a conventional encoder structure for scalable video coding.
[0015] Fig. 1 shows an encoder 100 comprising base coding means 102, prediction means 104, residual coding means 106 and combining means 108 as well as input 110 and output 112. The encoder 100 of Fig. 1 is a video encoder receiving a high quality video signal on input 110 and outputs a bit stream with scalable quality on output 112. Base coding means 102 encode data at input 110 in the base coding data stream representing the content of this video signal at input 110 with reduced bit depth of the image sample and optionally reduced spatial resolution compared to the input signal at input 110. The prediction means 104 are adapted, based on the base coding data stream output by the base coding means 102, to provide a prediction signal with full or increased bit depth of the image sample and optionally, full or increased spatial resolution for the video signal at input 110. The subtraction module 114, also included in the encoder 100, creates a prediction residual from the prediction signal provided by the means 104 relative to the high quality of the input signal at the input 110, the residual signal being encoded by the residual coding means 106 in the data stream of the quality improvement layer. The merging means 108 combine the base coding data stream from the base coding means 102 and the quality improvement layer data stream output by the residual encoding means 106 to form a data stream 112 of scalable quality at the output 112. The quality scalability means that the data stream at the output 112 consists from the part that is self-sufficient, i.e. allows reconstruction of video signal 110 with reduced bit depth and optionally reduced spatial resolution, without any additional information and ignoring the remainder of data stream 112 on one side and the distal part, which, in combination with the first part, allows reconstruction of the video signal at input 110 in the original bit depth and the original spatial resolution, greater than the bit depth and / or the spatial resolution of the first part.
[0016] After a fairly general description of the structure and functionality of the encoder 100, its internal structure is described in more detail below. In particular, base coding means 102 include down conversion module 116, subtraction module 118, transformation module 120 and quantization module 122, connected in series in said order between input 110 and connection means 108 and prediction means 104, respectively. Down conversion module 116 serves to reduce the bit depth of the image sample and optionally the spatial resolution of the video signal images at input 110. In other words, down conversion module 116 irreversibly converts the high quality video input signal 110 to the base quality video signal. As will be described in more detail below, this downward conversion may include reducing the bit depth of the signal samples, i.e. pixel values in the video signal at input 110 using any tone mapping method, such as rounding sample values, subsampling of chroma components when the video signal is provided in the form of luma and chroma components, filtering the input signal at input 110, as with conversion RGB to YCbCr or any combination thereof. More details about possible prediction mechanisms are provided below. In particular, it is possible that down conversion module 116 uses different down conversion methods for each video signal image or image sequence input 110, or uses the same method for all images. This is also discussed in more detail below.
[0017] The subtraction module 118, transformation module 120 and quantization module 122 cooperate to encode the base quality signal output by the down conversion module 116 by using, for example, a non-scalable video coding method such as H.264 / AVC. According to the example of Fig. 1, subtraction module 118, transformation module 120 and quantization module 122 interact with the optional predictive loop filter 124, prediction module 126, inverse transformation module 128 and addition module 130, typically included in base coding means 102 and prediction means 104, to form the part reducing the insignificance of the hybrid encoder, which encodes the base quality video signal output by downward conversion module 116 by motion compensation based prediction and subsequent compression of the prediction residuals. In particular, the subtraction module 118 subtracts from the current image or base quality video macroblock a predicted image or predicted portion of the macroblock reconstructed from pre-coded base quality video signals using for example motion compensation. Transformation module 120 applies transformation on residual prediction, such as DCT, FFT or elemental wave transformation. The transformed residual signal may represent a spectral representation, and its transformation coefficients are irreversibly quantized in the quantization module 122. The resulting quantized residual signal represents the residual base coding data stream output by the base coding means 102.
[0018] In addition to the optional predictive loop filter 124 and prediction module 126, inverse transformation module 128 and addition module 130, the prediction means 104 include an optional filter for reducing coding artifacts 132 and prediction module 134. Inverted transformation module 128, addition module 130, optional predictive loop filter 124 and prediction module 126 cooperate to reconstruct the video signal with reduced bit depth and optionally reduced spatial resolution as defined by down conversion module 116. In other words, they create a video signal with low bit depth and optionally low spatial resolution to the optional filter 132, which represents a low quality representation of the source signal at input 110, which can also be reconstructed on the decoder side. In particular, the reverse transformation module 128 and the addition module 130 are connected in series between quantization module 122 and the optional filter 132, while the optional predictive loop filter 124 and the prediction module 126 are connected in series in said order between the output of the add module 130 and the next module input 130 addition. The output of the prediction module 126 is also connected to the inverting input of the subtraction module 118. The optional filter 132 is connected between the output of the add module 130 and the prediction module 134, which in turn is connected between the output of the optional filter 132 and the inverting input of the subtraction module 114.
[0019] The inverted transformation module 128 performs the inverted transformation of the base coded residual images output by the base coding means 102 to obtain residual images with low bit depth and optionally low spatial resolution. Accordingly, the inverted transformation module 128 implements the inverted transformation, which is the inversion of the transformation and quantization carried out by modules 120 and 122. Alternatively, the dequantization module can be provided separately on the input side of the inverted transformation module 128. The addition module 130 adds a prediction to the reconstructed residual images, with the prediction based on previously reconstructed video signal images. In particular, the addition module 130 outputs a reconstructed video signal with reduced bit depth and optionally reduced spatial resolution. These reconstructed images, for example, are filtered through a loop filter 124 to reduce artifacts and then used by prediction module 126 to predict the image to be reconstructed, using, for example, motion compensation, from previously reconstructed images. The resulting base quality signal at the output of the addition module 130 is used by serially connecting the optional filter 132 and the prediction module 134 to obtain a high-quality prediction of the input signal at the input 110, the latter prediction being used to form a high-quality enrichment signal at output residual encoding means 106. This is described in more detail below. [0020] In particular, the low quality signal obtained from the addition module 130 is optionally filtered through an optional filter 132 to reduce coding artifacts. Filters 124 and 132 can even operate in the same way and thus, although filters 124 and 132 are shown separately in Fig. 1, both can be replaced by only one filter located between the output of the add-module 130 and the corresponding input prediction module 126 and prediction module 134. Then, a low quality video signal is used by the prediction module 134 to create a prediction signal for the high quality video signal received at the non-inverting input of the addition module 114 connected to the input 110. This high-quality prediction forming operation may include mapping of sample images of a decoded base quality signal by using the combined mapping function described in more detail below, using the appropriate value of the base quality signal samples to index the lookup table, which contains the appropriate high quality sample values using the quality base signal sample value for interpolation to obtain the appropriate high quality sample value by upsampling the chroma components, filtering the base quality signal by using, for example, conversion of YCbCr to RGB or any combination thereof. Other examples are described below.
[0021] For example, prediction module 134 may map a basic quality video signal sample from a first dynamic range to a second dynamic range that is higher than the first dynamic range and optionally, by using a spatial interpolation filter, spatially interpolate a basic quality video signal sample for increasing the spatial resolution to match the spatial resolution of the video signal at input 110. In a similar manner to the description of the down-conversion module 116, it is possible to use a different prediction method for different images of basic video signal sequences as well as to use the same prediction method for all images.
[0022] The subtraction module 114 subtracts the high quality prediction received from the prediction module 134 from the high quality video signal received from the input 110 to derive the high quality prediction residual signal, i.e. with original bit depth and optionally spatial resolution, to the residual coding means 106. In the residual encoding means 106, the difference between the original high-quality input signal and the prediction obtained from the decoded basic signal quality is encoded, for example using a compression coding method such as, for example, the H.264 / AVC standard. To this end, the residual coding means 106 of Fig. 1 include, for example, transformation module 136, quantization module 138, and entropy coding module 140 connected in series between the output of the subtraction module 114 and the combining means 108 in said order. Transformation module 136 transforms a residual signal or its images, respectively, into a transformation domain or spectral domain respectively, where spectral components are quantized by quantization module 138, wherein the quantized transformation values are entropy coded by entropy coding module 140. The entropy coding result represents the high quality data stream of the quality improvement layer output by the residual coding means 106. If modules 136 to 140 implement H.264 / AVC encoding that supports 4x4 or 8x8 sample size transformations, for luma encoding, the transformation size for the luma transformation of the residual signal from the subtraction module 114, in the transformation module 136, can be arbitrarily selected for each macroblock and may not necessarily be the same as used for coding the base quality signal in the transformation module 120. For encoding chroma components, the H.264 / AVC standard does not offer a choice. When quantizing transformation coefficients in the quantization module 138, the same quantization method as in H.264 / AVC can be used, which means that the size of the quantizer step can be controlled by the quantization QP parameter, which can take values from -6 * (bit depth of the component) high quality video signal - 8) up to 51. The QP used to encode the basic quality macroblock representation of the quantization module 122 and the QP used to encode the high quality macroblock enrichment in the quantization module 138 need not be the same.
[0023] The merging means 108 includes entropy coding module 142 and multiplexer 144. The entropy coding module 142 is connected between the output of quantization module 122 and the first input of multiplexer 144, while the second input of multiplexer 144 is connected to the output of entropy coding module 140. The output of multiplexer 144 represents output 112 of the encoder 100.
[0124] The entropy coding module 142 encodes the entropy quantized transformation values output by the quantization module 122 to form a quality base layer data stream from the base coding data stream output by the quantization module 122. Hence, as mentioned above, modules 118, 120, 122, 124, 126, 128, 130 and 142 can be adapted to cooperate in accordance with H.264 / AVC and jointly represent a hybrid encoder with entropy encoder 142, realizing lossless compression of quantized prediction residuals .
[0025] Multiplexer 144 receives both the quality base layer data stream and the high quality layer data stream and combines them with each other to form a scalable quality data stream.
[0026] As already described above and as shown in Fig. 3, the way prediction module 134 performs prediction from a reconstructed base quality signal to a high quality signal domain may include sample bit depth extension, also called inverse 150 tone mapping and optional operation up-spatial sampling, i.e. up-sampling filtering operation 152, when the base quality and high quality signals have different spatial resolution. The order in which the prediction module 134 implements inverted 150-tone mapping and optional upward spatial sampling operation 152 can be fixed and can accordingly be known in advance on both the encoder and decoder side, or it can be selected adaptively for each block or for each image or some other granulation, in which case the prediction module 134 signals the order information used for steps 150 and 152 to some object, such as encoder 106 to enter it as additional information into bit stream 112, to be signaled on the decoder side as part of the additional information. The adaptation of the order of steps 150 and 152 is shown in Fig. 3 by using the dashed line 154 with two arrows, and the granulation at which the order can be selected adaptively can be signaled also and may even change within the video.
[0027] In implementing the 150 tone inverse mapping, the prediction module 134 uses two parts, namely one or more global functions of the reverse tone mapping and its local adaptation. Generally, one or more of the global inverse tone mapping functions are dedicated to taking into account the overall average properties of the video image sequence and accordingly the tone mapping that was initially used for the high quality video input signal to obtain the base quality video signal in module 116 down conversion. In comparison, local adaptation should take into account individual deviations from the global inverted tone mapping model for individual video blocks.
[0028] To illustrate this, Fig. 4 shows a video portion 160, for example consisting of four consecutive images 162a to 162d, video 160. In other words, video 160 consists of a sequence of images 162, of which four are shown, for example, in Fig. 4. Video 160 may be divided into non-overlapping sequences of subsequent images for which global parameters or syntax elements are sent in the data stream 112. For illustrative purposes only, it is assumed that the four consecutive images 162a to 162d shown in Fig. 4 should form a sequence of 164 images. Each image in turn is divided into many macroblocks 166, as shown in the lower left corner of image 162d. A macroblock is a container in which transformation coefficients are sent along with other control syntax elements related to the type of macroblock coding in bit stream 112. A pair of 168 macroblocks 166 cover the continuous portion of the corresponding image 162d. Depending on the mode of the macroblock pair, the corresponding pair of macroblocks, the upper macroblock 162, of this pair 168, covers either the samples of the upper half of the macroblock pair 168, or the samples of each odd line in the pair of macroblocks, with the lower macroblock respectively for its other samples. Each macroblock 160, in turn, can be divided into transformation blocks, as shown under the designation 170, wherein these transformation blocks form block bases on which the transformation module 120 performs the transformation, and the reverse transformation module 128 performs the inverted transformation.
[0029] Referring again to the aforementioned inverse tone mapping function, the prediction module 134 can be configured to use one or more of such global inverse tone mapping functions continuously for all video 160 or, alternatively, for sub-parts thereof such as a sequence of 164 consecutive images or image alone 162. The latter option implies that prediction module 134 changes the global function of inverted tone mapping at granulation corresponding to the size of the image sequence or image size. Examples of global reverse tone mapping functions are given below. If the prediction module 134 adapts the global reverse tone mapping function to video statistics 160, the prediction module 134 outputs information to the entropy coding module 140 or the multiplexer 144, so that the bit stream 112 includes information about the global reverse tone mapping function (s) and their changes in video framework 160. In the event that one or more of the global inverse tone mapping functions used by the prediction module 134 are constantly applicable to the entire video 160, they may be known in advance to the decoder or may be transmitted as additional information in bit stream 112.
[0030] With even less granulation, the local inverse tone mapping function used by the prediction module 134 changes within the video 160. For example, such a local inverse tone mapping function changes with granulation smaller than the image size, such as, for example, the macroblock size or macroblock pairs or transformation block size.
[0031] For both the global (inverted) tone mapping function and the local inverse tone mapping function, the granulation at which the respective function changes, or at which functions are defined in the bit stream 112, may change within the video 160. Change granulation in turn can be signaled in bit stream 112.
[0032] During inverse 150 tone mapping, the prediction module 134 maps a predetermined video sample 160, from a base quality bit depth to a high quality bit depth, by using a combination of one global reverse tone mapping function applicable to the corresponding image and local reverse tone mapping function defined in the corresponding block to which the determined sample belongs.
[0033] For example, the combination may be an arithmetic combination, and in particular an addition. The prediction module can be configured to obtain a predicted shigh value with a high quality bit depth from the corresponding reconstructed s value<sub>low</sub> samples of low quality bit depth using formula s<sub>high</sub> = f<sub>k</sub>(s<sub>low</sub>) + m · p<sub>low</sub> + n.
[0034] In this formula, function f<sub>k</sub> represents the global reverse tone mapping operator in which the index k decides which global reverse tone mapping operator is selected when more than one method or more than one global reverse tone mapping function is used. The remainder of this formula is a local adaptation or local function of inverted tone mapping, with z being the shift value and m as the scale factor. The k, min values can be specified block by block in bit stream 112. In other words, bit stream 112 will allow the disclosure of triples (k, m, n) for all 160 video blocks, with the block size of these blocks dependent on the global adaptation of the global function reverse tone mapping, and this granulation in turn potentially changes within the 160 video.
[0035] The following mapping mechanisms can be used for the prediction method in the scope of the global function f (x) of inverted time mapping. For example, linear part mapping can be used where an arbitrary number of interpolation points can be specified. For example, for a base quality sample with the value x and the two given points (xn, yn) and (xn + 1, yn + 1) interpolation, the corresponding prediction sample y is obtained by module 134 according to the following formula
<img file="PL2279622T3_D0001.tif" />
χ —χ »[0036] Linear interpolation can be performed with low computational complexity by using only bit shift instead of division operations if x<sub>n +</sub>and - x<sub>n</sub> is limited to the power of two.
[0037] A further potential global mapping mechanism represents a lookup table mapping in which, using the value of a base quality sample, table lookup is implemented in a lookup table in which for each potential value of the quality base sample, as regards the global inverse tone mapping function, determined the global (x) value of the prediction sample is appropriate. A lookup table may be provided to the decoder page as additional information or may be known on the decoder side by default.
[0038] In addition, fixed shift scaling can be used for global mapping. According to this alternative approach, to obtain a suitable sample of high-quality global prediction with a greater bit depth, module 134 multiplies the sample x base quality by a constant factor of 2<sup>MNK</sup>and then move it permanently 2<sup>M-1</sup>-2<sup>M-K-1</sup> is added in accordance with, for example, one of the following formulas:
<img file="PL2279622T3_D0002.tif" />
or respectively f (x) = min (2
MNK
<img file="PL2279622T3_D0003.tif" />
where M is the bit depth of the high quality signal and N is the bit depth of the basic quality signal.
[0039] In this way, the dynamic range [0; 2<sup>N-1</sup>] of low quality is mapped to the second dynamic range [0; 2<sup>M-1</sup>] in such a way that the mapped x values are distributed centrally to the potential dynamic range [0; 2<sup>M-1</sup>] higher quality as part of the extension, which is determined by K. The K value can be an integer value or an actual value and can be sent as additional information to the decoder, as part of, for example, a scalable quality data stream, so that some prediction means in the decoder they can operate in the same way as the prediction module 134, as will be described below. To obtain integer f (x) values, the rounding function can be used.
[0040] Another option for global scaling is variable shift scaling: xo base quality samples are multiplied by a fixed factor and then the variable shift is added according to, for example, one of the following formulas:
<img file="PL2279622T3_D0004.tif" />
or f (x) = min + D, 2 * - 1) [0041] In this way, the low quality dynamic range is globally mapped to the second dynamic range in a way that the mapped values of x are distributed over part of the potential dynamic range of samples with high quality, whose extension is determined by K and whose displacement within the lower limit is determined by D. D can be an integer or a real number. The resulting f (x) represents the globally mapped value of the high-quality prediction signal sample image. K and D values can be transmitted as additional information to the decoder, as part of, for example, a scalable quality data stream. Again, the rounding operation can be used to obtain integer f (x) values, the latter being also true for the other examples given in this application, for global bit depth mappings without re-affirming this explicitly.
[0042] Yet another possibility of global mapping is overlay scaling: globally mapped f (x) high bit depth prediction samples are obtained from the corresponding xo sample of base quality according to, for example, one of the following formulas, where floor (a) rounds "a "Down to the nearest integer:
f (x) = floor (2 "'<sup>ν</sup>χ + 2<sup>M-2w</sup>x) or f (x) = min (floor (2<sup>M</sup>‘<sup>N</sup>x + 2<sup>M</sup>~<sup>2N</sup>x), 2 "- 1) [0043] The abovementioned possibilities can be combined. For example, global overlap and fixed shift scaling can be used: globally mapped f (x) high bit depth prediction samples are obtained according to for example one from the following formulas, where floor (a) rounds "a" down to the nearest integer:
f (x) = floor (2
NMK x + 2<sup>m</sup>‘<sup>2w</sup>'<sup>k</sup>x + 2 "- * - 2, Μ-1-Κ
<img file="PL2279622T3_D0005.tif" />
[0044] The value of K may be determined as additional information for the decoder.
[0045] Similarly, global overlap, variable shift scaling can be used: globally mapped high bit depth prediction samples (x) are obtained according to the following formula, where floor (a) rounds "a" down to the nearest integer :
f (x) = floor (2
NMK x + 2<sup>M</sup>'<sup>2N</sup>~<sup>K</sup>x + D) f (x) = min (floor (2
MNK
M-2N-K x + D), 2 "- 1) [0046] The values of D and K can be specified as additional information to a decoder. [0047] By transferring the examples just mentioned for the global inverted tone mapping function to Fig. 4, the parameters mentioned therein for defining the global inverted tone mapping function, namely (x1, y1), K and D can be known to the decoder, can be transmitted in bit stream 112 for all video 160, in case the global inverted tone mapping function is constant within video 160, or these parameters are transmitted in a bit stream for its various parts, such as for example a sequence of 164 images, or images 162, depending on coarse granulation, underlying global mapping function. In the case where the prediction module 134 uses more than one global function of inverted tone mapping, said parameters (x1, y1), K and D can be considered as supplied with the index k, with these parameters (x<sub>1z</sub> s<sub>1</sub>) K<sub>k</sub> and D<sub>k</sub> defining the k-th global function f<sub>k</sub> inverted tone mapping.
[0048] Thus, it is possible to determine the use of different global inverse tone mapping mechanisms for each block, each image, by signaling the corresponding k value of bit stream 112, as well as, alternatively, by using the same mechanism for the complete video sequence 160.
[0049] In addition, it is possible to specify, by various global mapping mechanisms for the luma and chroma components, of a signal of base quality that the statistics, such as the probability density function, may be different.
[0050] As noted above, the prediction module 134 uses a combination of a global inverse tone mapping function and a local inverse tone mapping function for local adaptation of the global function. In other words, local adaptation is implemented on every globally mapped sample f (x) of high bit depth prediction obtained in this way. Local adaptation of the global reverse tone mapping function is accomplished by using a local reverse tone mapping function that is locally adapted, for example, by locally adapting some of its parameters. In the example above, these parameters are the scaling factor m and the offset n value. The scaling factors m and the value of n shift can be determined for each block, where the block may correspond to the size of the transformation block, which is for example in the case of H.264 / AVC 4x4 or 8x8 samples, or the size of the macroblock which is in the case of H.264 / AVC for example 16x16 samples. Which block meaning is actually used can be either immutable, i.e. known in advance to both the encoder and the decoder, or can be selected for the image or for the sequence by the prediction module 134, in which case it should be signaled to the decoder as part of the additional information in bit stream 112. For H.264 / AVC, a set of sequence parameters and / or a set of image parameters can be used for this purpose. In addition, it can be specified in the additional information that either the scaling factor m, or the offset value n, or both are set to zero, for the complete video sequence or for a well-defined set of images in the video sequence 160. In the event that either the scaling factor m, or the offset value n, or both are specified for each block for a given image, in order to reduce the required bit rate for encoding these values, only the Am values, An differences for the respective predicted m values<sub>pred</sub>, n<sub>pred</sub> they can be coded so that the actual values for m, n can be obtained as follows: m = e.g.<sub>r</sub>ed + Am, n = e.g.<sub>r</sub>ed + An.
[0051] In other words, as shown in Fig. 5, which shows a portion of the image 162 divided into blocks 170 forming the basis for the granulation of local adaptation of the global reverse tone mapping function, prediction module 134, entropy coding module 140 and multiplexer 144 are configured such that the parameters for local adaptation of the global reverse tone mapping function, namely the scaling factor m and the shift n value used for individual blocks 170 are not directly coded into bit stream 112, but only as a residual prediction, for the prediction obtained from the scaling factors and offset values of adjacent blocks 170. So {Am, An, k} are transmitted for each block 170 when more than one global mapping function is used, as shown in Fig. 5. For example, suppose the prediction module 134 uses the parameters m<sub>and</sub>, jin<sub>and</sub>, j for the reverse tone mapping of samples of a certain block i, j, namely the middle block of Fig. 5. Next, the prediction module 134 is configured to calculate the values of mi, jpred and ni, jpred prediction from scaling factors and the offset values of adjacent blocks 170, for example, m<sub>and</sub>, j_i et al<sub>and</sub>,<sub>j-1</sub>. In this case, the prediction module 134 will cause the difference between the actual parameters m<sub>and</sub>, jin<sub>and</sub>, j and predicted m<sub>and</sub>, jn<sub>and</sub>j<sub>pred</sub> in<sub>and</sub>, jn<sub>and</sub>j<sub>pred</sub> will be introduced to bit stream 112. These differences are designated in Am. 5 as Am<sub>and</sub>, ji An<sub>and</sub>, j, wherein the indices i, j are for example the jth block from the top and the ith block from the left of image 162. Alternatively, the prediction of the scaling factor m and the offset n values can be obtained from already transferred blocks, rather than from adjacent blocks of the same image. For example, the prediction values can be obtained from blocks 170 of the previous image, located in the same or the corresponding spatial location. In particular, the predicted value can be: • a constant value that is either transmitted in the notes, or is already known to both the encoder and the decoder, • the value of the corresponding variable and the preceding block, • the median value of the corresponding variables in adjacent blocks, • the average value of variables in adjacent blocks, • linearly interpolated or extrapolated value, obtained from the values of the corresponding variables in adjacent blocks.
[0052] Which of these prediction mechanisms is actually used for a given block may be known to both the encoder and the decoder, or may depend on the values of m, n adjacent blocks, if any.
[0053] In the encoded high quality enrichment signal output by entropy coding module 140, the following information may be sent for each macroblock in the case where modules 136, 138 and 140 implement encoding in accordance with H.264 / AVC. Coded block pattern (CBP) information may be included, indicating which of the four 8x8 luma transformation blocks in the macroblock and which of the associated macroblock chroma transformation blocks may contain nonzero transformation coefficients. If there are no non-zero transformation coefficients, no further information is sent for the given macroblock. Further information may relate to the transformation size used to encode the luma component, i.e. the size of transformation blocks in which a macroblock consisting of 16x16 luma samples is sent, in transformation module 136, i.e. in 4x4 or 8x8 transformation blocks. In addition, the high quality improvement layer data stream may include the QP quantization parameter used in the quantization module 138 to control the size of the quantization step. In addition, quantized transformation coefficients, i.e. levels of transformation coefficients may be included for each macroblock in the high quality enhancement layer data stream output by entropy coding module 140.
[0054] In addition to the above information, the following information should be included in data stream 112. For example, if more than one single global inverted tone mapping method is used for the current image, the corresponding k index value must also be transmitted for each enrichment signal block high-quality. In terms of local adaptation, the variables ńm and ńn can be signaled for each block of these images, where the transformation of the respective values is indicated in the additional information by, for example, a set of sequence parameters and / or a set of image parameters in H.264 / AVC. For all three new variables k, ńm and ńn, new syntax elements with appropriate binaryization methods must be introduced. Simple unary binarization can be used to prepare three variables for the binary arithmetic coding method used in H.264 / AVC. Because all three variables typically have small modules, a simple unary binaryization method is very suitable. Because ńn and ńn are signed integers, they can be converted to unsigned values, as described in the following table:
<td>Am value, Δη with a sign</td><td>unsigned value (for binarization)</td>
<td> 0</td><td> 0</td>
<td> 1</td><td> 1</td>
<td> 2</td><td> -1</td>
<td> 3</td><td> 2</td>
<td> 4</td><td> -2</td>
<td> 5</td><td> 3</td>
<td> 6</td><td> -3</td>
<td>k</td><td> (-1<sup>) k</sup>+<sup>1</sup> Ceil (k + 2)</td>
[0055] Thus, by summarizing some of the above embodiments, the prediction means 134 may perform the following steps during their mode of operation. In particular, as shown in Fig. 6, the prediction means 134, in step 180, set one or more global (global) m (f) mapping functions, or, in the case of more than one global fk (x) mapping functions, z denoting the appropriate index value, indicating the appropriate global mapping function. As described above, one or more global (global) mapping functions can be set or defined as fixed throughout the entire video 160, or can be defined or set in coarse granulation, such as granulation of image size 162 or sequence of 164 images. One or more global (global) mapping functions may be defined as described above. In general, global mapping functions can be non-trivial functions, i.e. unequal f (x) = fixed quantity, and in particular can be nonlinear functions. Whenever any of the above examples of the global mapping function is used, step 180 leads to a corresponding parameter of the global mapping function, such as K, D or (x<sub>n</sub>, y<sub>n</sub>) from n G (1, ... N) set for each of the sections into which the video is divided, according to coarse granulation.
[0056] In addition, the prediction module 134 sets the local mapping / adaptation function such as indicated above, namely the linear function according to m-x + n. However, another local mapping / adaptation function is also possible, such as a constant function, only parameterized by using the offset value. Setting 182 is accomplished with finer granulation, such as, for example, granulation size smaller than the image, e.g., macroblock, transformation block or pair of macroblocks, or even a segment in the image, where the segment is a subset of macroblocks or pairs of image macroblocks. Thus, step 182 leads, in the case of the above embodiment for the local mapping / adaptation function, to a pair of ńm and nn values, set or defined for each of the blocks, into which video images 162 are divided according to finer granulation.
[0057] Optionally, namely when more than one global mapping function is used in step 180, the prediction means 134, in step 184, sets the index k for each fine granulation block used in step 182.
[0058] Although it would be possible for the prediction module 134 to set one or more global (global) mapping functions at step 180 only depending on the mapping function used by the down conversion module 116 to reduce the bit depth of the sample of the original image samples as per example by using the reverse mapping function as a global mapping function in step 180, it should be noted that it is also possible that the prediction means 134 implement all settings in step 180, 182 and 184 so that a certain optimization criterion, such as the ratio of bit rate to distortion of the resulting bit stream 112, is extreme, such as maximized or minimized. Thanks to this, both the global (global) mapping function (s) and its (their) local adaptation, namely the local mapping / adaptation function are adapted to the tone statistics of the video sample value, determining the best compromise between coding overhead of the necessary additional information for global mapping on one side and obtaining the best fit of the global mapping function to tone statistics, imposing only a small local mapping / adaptation function on the other hand.
[0059] When implementing the actual reverse tone mapping in step 186, the prediction module 134 uses, for each sample of the reconstructed low quality signal, a combination of a global mapping function in the case where there is only one global mapping function, or one of the global mapping functions in the case, when there is more than one function on the one hand and a local mapping / adaptation function on the other hand, both as defined for the block or section, to which the current sample belongs. In the above example, the combination is addition. However, an arithmetic combination is also possible. For example, the combination can be the serial application of both functions.
[0060] Furthermore, in order to inform the decoder side of the combined mapping function used for the inverse tone mapping, the prediction module 134 at least causes the local mapping / adaptation function information to be provided to the decoder page by the bit stream 112. For example, the module The prediction 134 causes, at step 188, that the min values are encoded in bit stream 112 for each fine granulation block. As described above, the coding in step 188 may be a predictive coding according to which the remains of the min prediction are encoded in a bit stream rather than their actual values, with obtaining the prediction of the values that are obtained from these min values, in adjacent blocks 170, or in the corresponding block in the previous image. In other words, local or temporal prediction along with residual coding can be used to encode the parameter mi n.
[0061] Similarly, in the case where more than one global mapping function is used in step 180, the prediction module 134 may cause the index k to be encoded in the bit stream for each block in step 190. In addition, the prediction module 134 may cause the information of (x) or fk (x) is encoded in the bit stream when the same information is not known to the decoder a priori. This is done in step 192. In addition, prediction module 134 may cause video granulation and granulation change information for the local mapping / adaptation function and / or global mapping function (global functions) to be encoded in the bit stream at step 194. [0062] It should be noted that all steps 180 to 194 need not be implemented in the order mentioned. It is not even necessary to rigidly implement these steps in sequence. Rather, steps 180 to 194 are shown in sequential order for illustrative purposes only and these steps will preferably be implemented overlapping.
[0063] Although not explicitly stated in the above description, it should be noted that the additional information generated in steps 190 to 194 may be input into the high quality quality layer signal or the high quality part of the bit stream 112 rather than the base part quality from entropy coding module 142.
[0064] After describing the embodiment of the encoder with reference to Fig. 2, an embodiment of the decoder is described. The decoder of Fig. 2 is marked with reference number 200 and includes demultiplex means 202, base decoding means 204, prediction means 206, means
208 residual decoding and reconstruction means 210 as well as input 212, first output 214 and second output 216. The decoder 200 receives, at its input 212, a scalable quality data stream, which, for example, was output by the encoder 100 of Fig. 1. described above, quality scalability can refer to bit depth and optionally to spatial reduction. In other words, the bit stream at input 212 may include a self-sufficient part that can be used in isolation for the reconstruction of a video signal with reduced bit depth and optionally reduced spatial resolution, as well as an additional part which, in combination with the first part, allows the reconstruction of the video signal higher bit depth and optional spatial resolution. A lower quality reconstruction video signal is output at output 216 while a higher quality reconstruction video signal is output at output 214.
[0065] The demultiplex means 202 divide the incoming quality scalable data stream at the input 212 into the base coding data stream and the high quality quality stream of the quality improvement layer which both were mentioned with reference to Fig. 1. These base decoding means 204 are used for decoding a base coding data stream to represent the basic quality of the video signal, which may be directly as in the example of Fig. 2, or indirectly via an artifact reduction filter (not shown), optionally output at output 216. Based on the video signal of the basic quality representation, the prediction means 206 create a prediction signal having increased bit depth of the image sample and / or increased chroma sampling resolution. The decoding means 208 decode the data stream of the quality improvement layer to obtain a residual prediction having increased bit depth and optionally increased spatial resolution. The reconstruction means 210 obtain a high quality video signal from the prediction and the remains of the prediction and output it at output 214, via an optional artifact reduction filter.
[0066] Inside, the demultiplex means 202 include the demultiplexer 218 and entropy decoding module 220. The demultiplexer 218 input is connected to input 212 and the first demultiplexer 218 output is connected to the residual decoding means 208. The entropy decoding module 220 is connected between the other demultiplexer output 218 and the base decoding means 204. The demultiplexer 218 divides the scalable quality data stream into a base layer data stream and a quality improvement layer data stream that has been separately introduced into multiplexer 144, as described above. The entropy decoding module 220 implements, for example, a Huffman decoding algorithm or arithmetic decoding to obtain levels of transformation coefficients, motion vectors, information about the size of the transformation and other syntax elements necessary to obtain a base representation of the video signal. The base coding data stream is formed at the output of the entropy decoding module 220.
[0067] The base decoding means 204 includes inverse transformation module 222, addition module 224, optional loop filter 226 and prediction module 228. Modules 222 to 228 of the base decoding means 204 correspond in terms of functionality and connections between them, elements 124 to 130 of Fig. 1. More specifically, the reverse transformation module 222 and addition module 224 are connected in series in said order between the demultiplex centers 202 on the one hand and the prediction means 206 and the base quality output on the other, respectively, and the optional loop filter 226 and the prediction module 228 are connected in series in said order between the output of the add module 224 and another input of the add module 224. In this way, the addition module 224 outputs a video signal of the base representation with reduced bit depth and optionally with reduced spatial resolution that can be received from outside from output 216.
[0068] Prediction means 206 comprise an optional artifact reduction filter 230 and prediction information module 232, both modules operating synchronously with respect to elements 132 and 134 of Fig. 1. In other words, the optional artifact reduction filter 230 optionally filters the base quality video signal to reduce the artifacts contained therein, and the prediction forming module 232 acquires predicted images with increased bit depths and optionally increased spatial resolution, as already described above with respect to the module 134 predictions. That is, the predictor information module 232 can, using additional information contained in the scalable quality data stream, map incoming image samples to a higher dynamic range and optionally apply a spatial interpolation filter on the image content to increase spatial resolution.
[0069] The decoding means 208 include an entropy decoding module 234 and a module
236 inverted transformations that are connected in series between the demultiplexer
218 and reconstruction means 210 in said order. The entropy decoding module 234 and the inverse transformation module 236 cooperate to invert coding implemented by modules 136, 138 and 140 of Fig. 1. In particular, entropy decoding module 234 implements, for example, Huffman decoding or arithmetic decoding algorithms to provide syntax elements, including, but not limited to, levels of transformation coefficients, which are inversely transformed by inverse transformation module 236 to obtain prediction residual signals or sequential residual image sequences. In addition, entropy decoding module 234 discloses additional information generated in steps 190 to 194 on the encoder side, such that the prediction forming module 232 is able to emulate the inverse mapping procedure implemented on the encoder side by the prediction module 134, as already noted above.
[0070] As in Fig. 6, which refers to the encoder, Fig. 7 shows in more detail the operating mode of the prediction forming module 232 and, in part, the entropy decoding module 234. As shown, the method of obtaining a prediction from the reconstructed base layer signal on the decoder side begins with the interaction of the demultiplexer 218 and entropy decoder 234, to obtain the granulation information encoded in step 194, in step 280, to obtain information about global function (global functions) mapping encoded in step 192, in step 282, to obtain the k value of the index for each fine granulation block that was encoded in step 190, in step 284 and obtaining the min parameter of the local mapping / adaptation function for the finer granulation block that was encoded in step 188, in step 286. As shown by the dashed lines, steps 280 to 284 are optional, the use of which depends on the currently used embodiment . [0071] In step 288, prediction forming module 232 performs an inverse tone mapping based on the information obtained in steps 280 to 286, thereby accurately emulating the inverse tone mapping implemented on the encoder side in step 186. As in the description of step 188, the acquisition of parameters m, n may include predictive decoding, in which the value of the prediction residual is obtained from a part of the high quality data stream introduced into the demultiplexer 218, by using, for example, entropy decoding implemented by entropy decoder 234 and acquisition actual min values by adding these residual prediction values to the prediction value, acquired by local and / or time prediction. [0072] Reconstruction means 210 comprise an add module 238 whose inputs are connected to the output of the prediction information module 232 and the output of the module respectively
236 inverted transformation. Module 238 adds a prediction residual and a prediction signal to obtain a high quality video signal with reduced bit depth and optionally with reduced spatial resolution, which is provided by the optional artifact reduction filter 240, to output 214.
[0073] Thus, as can be seen from Fig. 2, the basic quality decoder can reconstruct the basic quality video signal from the scalable quality data stream at input 212 and may, for this purpose, not include elements 130, 232, 238, 234, 236 and 240. On the other hand, the high quality decoder may not include outputs 216.
[0074] In other words, in the decoding method, decoding the base quality representation is simple. To decode a high quality signal, a basic quality signal must first be decoded, which is implemented by modules 218 to 228. Then, the prediction method described above for module 232 and optional module 230 is used using the decoded base representation. The quantized transformation coefficients of the high quality enrichment signal are scaled and inversely transformed by the inverted transformation module 236, for example as specified in H.264 / AVC, to obtain residual signal or difference samples that are added to the prediction obtained from the decoded base samples representation by the prediction module 232. In the final step of decoding the high-quality video signal to be output at output 214, an optional filter can be used to remove or reduce visually disturbing coding artifacts. It should be noted that the motion compensation prediction loop, including modules 226 and 228, is fully self-sufficient, using only a representation of base quality. Hence, the decoding complexity is moderate and there is no need for an interpolation filter that works on image data with a higher bit depth and optionally with a higher spatial resolution in the method of prediction with motion compensation, the prediction module 228.
[0075] In the context of the above embodiments, it should be noted that the artifact reduction filters 132 and 230 are optional and can be removed. The same applies to loop filters 124 and 226 and filter 240, respectively. For example, referring to Fig. 1 it should be noted that filters 124 and 132 can be replaced by only one common filter, such as, for example, an unblocking filter whose output is connected both to the input of the motion prediction module 126 and the input of the inverse tone mapping module 134. Similarly, filters 226 and 230 may be replaced by one common filter, such as an unblocking filter, whose output is connected to output 216, input of prediction forming module 232, and input of prediction module 228. In addition, to ensure completeness, it should be noted that prediction modules 228 and 128, respectively, do not necessarily temporarily determine samples in the macroblocks of the current image. Rather, spatial prediction or Intra prediction may also be used, using samples of the same image. In particular, the type of prediction can be selected, for example, for each macroblock or some other granulation. In addition, the present invention is not limited to video coding. Rather, the above description also applies to the encoding of still images. Accordingly, the motion compensation prediction loop comprising elements 118, 128, 130, 126 and 124 and elements 224, 228 and 226, respectively, can also be removed. Similarly, said entropy coding need not necessarily be implemented.
[0076] More specifically, in the above embodiments, encoding 118130, 142 of the base layer was based on motion compensation prediction based on the reconstruction of already loss-coded images. In this case, the base coding reconstruction method can also be seen as part of the high quality prediction method that was implemented in the above description. However, in the case of lossless coding of the underlying representation, no reconstruction would be necessary and the down-converted signal could be sent directly to means 132, 134. In the case of prediction without motion compensation in lossy base layer coding, reconstruction, for reconstruction of the base quality signal on the encoder side would be particularly dedicated to the formation of high-quality prediction in 104. In other words, the above association of items 116-134 and 142 with measures 102, 104 and 108, respectively, can be implemented in a different manner. In particular, entropy coding module 142 can be seen as part of the base coding means 102, with prediction means comprising only modules 132 and 134 and combining means 108 only containing multiplexer 144. This approach is correlated with the combination of modules / means used in Fig. 2 such that the prediction means 206 do not include motion-based prediction. Additionally, however, demultiplex means 202 may be viewed as not having entropy module 220, so that the base decoding means also include entropy decoding module 220. However, both approaches lead to the same results, because the prediction in 104 is based on a representation of the source material with reduced bit depth and optionally a reduced spatial resolution, which is losslessly coded for losslessly obtained from a bit stream with scalable quality and a layer data stream base. According to the approach underlying Fig. 1, prediction 134 is based on reconstruction of the base coding data stream, while in the alternative approach, reconstruction would start with an intermediate coded version or half coded version of the quality base signal in which lossless coding by module 142 is omitted for complete coding to base layer data stream. In this regard, it should further be noted that the downward conversion in module 116 need not be performed by the encoder 100. Rather, the encoder 100 may have two inputs, one for receiving a high quality signal and the other for receiving a down-converted version from the outside.
[0077] In the embodiments described above, the quality scalability concerned only bit depth and optionally spatial resolution. However, the above embodiments can easily be extended to include time scalability, chroma format scalability, and fine granularity scalability.
[0078] Accordingly, the above embodiments of the present invention form a coding concept for scalable coding of image or video content with different granulation in terms of bit depth of the sample and optionally spatial resolution, by using locally adaptive inverted tone mapping. In accordance with embodiments of the present invention, both the temporal and spatial prediction method specified in the scalable H.264 / AVC video coding extension are extended to also include mapping from lower fidelity to higher bit depth of the sample, optionally from lower to higher resolution spatial. The SVC extension described above for scalability in the sample bit depth and optional spatial resolution, allows the coder to store a representation of the base quality of the video sequence, which can be decoded by any older video decoder, together with the enrichment signal for higher bit depth and optionally higher spatial resolution that are ignored by older video decoders. For example, the base quality representation may contain an 8 bit version of the video sequence in CIF resolution, namely 352x288 samples, while the high quality enrichment signal contains "processing" up to 10 bit version, in 4CIF resolution, i.e. 704x476 samples of the same sequence. In another configuration, it is also possible to use the same spatial resolution for both representations, base quality and enriched quality, so that the high quality enrichment signal only contains processing of the bit depth of the sample, e.g. from 8 to 10 bits.
In other words, the above-mentioned embodiments allow for the creation of a video encoder, i.e. encoder or decoder, for encoding, i.e. encoding or decoding, a layered representation of a video signal, including a standard video coding method, for coding a base quality layer, a prediction method for performing a prediction of a high quality signal of a quality improvement layer, by using a reconstructed base signal quality and a residual coding method, for coding residual prediction High quality signal layer improving quality. In this range, the prediction can be realized by applying a mapping function from the dynamic range associated with the base quality layer to the dynamic range associated with the high quality enrichment layer. In addition, the mapping function can be built as the sum of the global mapping function that is compatible with the reverse tone mapping method described above and local adaptation. Local adaptation, in turn, can be accomplished by scaling the x value of the baseline quality samples and adding the offset value according to m ^ x n. In any case, residual coding can be performed according to H.264 / AVC.
[0080] Depending on the actual implementation, the coding method according to the invention can be implemented in hardware or in software. Hence, the present invention also relates to a computer program that can be stored on a computer-readable medium, such as a CD, disk, or any other data medium. The present invention is therefore also a computer program comprising a program code which, when executed on a computer, implements the method according to the invention described with reference to the above figures. In particular, the implementation of the means and modules in Fig. 1 and Fig. 2 may include standard subprograms implemented, for example, in the CPU, parts of the ASIC (Application Specific Integrated Circuit), etc.
[0081] Thus, among other things, the above embodiments describe an encoder for encoding image or video source data (160), to a scalable quality data stream (112), including base coding means (102), for encoding source data (160) image or video, to a base coding data stream representing a representation of the image data of the video or video, with a first bit depth of the image sample; mapping means (104) for mapping samples of image or video data source representation (160) with a first bit depth of the image sample from the first dynamic range corresponding to the first bit depth of the image sample to the second dynamic range greater than the first dynamic range and the corresponding a second bit depth of the image sample, higher than the first bit depth of the image sample, by using one or more global mapping functions, constants within the image or video source data (160), or variables at the first granulation, and local mapping function that locally modifies one or more global mapping functions, at the second granulation, finer than the first granulation, in order to predict the source image data or video having a second bit depth of the image sample; residual coding means (106) for coding residual prediction, prediction, to the bit stream quality improvement layer data stream; and merging means (108) for forming a scalable quality data stream based on a base coding data stream, a local mapping function and a bit depth enrichment data stream such that the local mapping function can be obtained from the scalable quality data stream.
[0082] The mapping means may include mapping means (124, 126, 128, 130, 132) for image or video reconstruction of low bit depth reconstruction as a representation of the image or video source data with the first bit depth of the image sample based on a base coding data stream wherein the low bit depth reconstruction image or video has a first bit depth of the image sample.
[0083] The mapping means (104) may be adapted to map samples of the image or video source data representation (160) with the first bit depth of the image sample, by using a combined mapping function being an arithmetic combination of one of one or more global mapping functions and local mapping function. More than one global mapping function may be used by the mapping means (104), and the merging means (108) are adapted to form a data stream (112) of scalable quality such that one or more global mapping functions can be obtained from the data stream with scalable quality. The arithmetic combination may include an addition operation.
[0084] The merging means (108) and mapping means (104) can be adapted such that the second granulation divides the image or video source data (160) into a plurality of image blocks (170), and the mapping means can be adapted such that a local function mapping is m-s + n, variables change at the second granulation, with the means of joining adapted so that min are defined in the data stream of scalable quality for each block (170) of image source data (160) of image or video, so that min may differ between multiple image blocks.
[0085] The mapping means (134) may be adapted such that the second granulation varies in the image or video source data (160), and the combining means (108) can be adapted so that the second granulation can also be obtained from the data stream scalable quality.
[0086] The merging means (108) and mapping means (134) can be adapted such that the second granulation divides the image or video source data (160) into a plurality of image blocks (170), and the merging means (108) can be adapted to such the way that the remainder (ńm, ńn) of the local mapping function is introduced into the data stream (112) of scalable quality for each block (170) of the image, and the local mapping function of a fixed block of image source data (160) of image or video can be obtained from a scalable quality data stream by using spatial and / or time prediction, one or more adjacent image blocks or a corresponding image block, image, source data image or video that precedes the image to which the fixed image block belongs, and remains of the local mapping function of the fixed image block. [0087] The combining means (108) can be adapted such that at least one of the one or more global mapping functions can also be obtained from a scalable quality data stream.
[0088] The base coding means (102) may include means (116) for mapping image samples with a second bit depth of the image sample from the second dynamic range to the first dynamic range corresponding to the first bit depth of the image sample to obtain a reduced quality image ; and means (118, 120, 122, 124, 126, 128, 130) for encoding the reduced quality image to obtain a base coding data stream.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany
Proxy:
EP 2 279 622 B1
Z-12825
Contents4
29 members in 10 offices
Priority claims1
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008003047 | European Patent Office (EPO) | W |
Members29
| Document | Office | Kind | |
|---|---|---|---|
| WO2009127231A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2279622A1 | European Patent Office (EPO) | A1 | |
| CN102007768A | China | A | |
| US2011090959A1 | United States of America | A1 | |
| JP2011517245A | Japan | A | |
| CN102007768B | China | B | |
| JP5203503B2 | Japan | B2 | |
| EP2279622B1 | European Patent Office (EPO) | B1 | |
| PT2279622E | Portugal | E | |
| DK2279622T3 | Denmark | T3 | |
| ES2527932T3 | Spain | T3 | |
| EP2835976A2 | European Patent Office (EPO) | A2 | |
| PL2279622T3This record | Poland | T3 | |
| US8995525B2 | United States of America | B2 | |
| EP2835976A3 | European Patent Office (EPO) | A3 | |
| US2015172710A1 | United States of America | A1 | |
| HUE024173T2 | Hungary | T2 | |
| EP2835976B1 | European Patent Office (EPO) | B1 | |
| PT2835976T | Portugal | T | |
| DK2835976T3 | Denmark | T3 | |
| ES2602100T3 | Spain | T3 | |
| PL2835976T3 | Poland | T3 | |
| HUE031487T2 | Hungary | T2 | |
| US2019289323A1 | United States of America | A1 | |
| US10958936B2 | United States of America | B2 | |
| US2021211720A1 | United States of America | A1 | |
| US11711542B2 | United States of America | B2 | |
| US2023421806A1 | United States of America | A1 | |
| US12457361B2 | United States of America | B2 |
Numbers
- Application
- 8735287
Titles2
- English
- BIT-DEPTH SCALABILITY
- Polish
- Skalowalność głębi bitowej
Classification
- CPC, 7
- H04N19/117
- H04N19/593
- H04N19/184
- H04N19/80
- H04N19/33
- H04N19/36
- H04N19/82
- IPC, 1
- H04N19 00