Bit-depth scalability
Abstract
This record has no abstract on file.
Term
No projected expiry on record.
- Priority and filed
- Published
- Today
16 claims: 11 independent, 5 dependent
- 1Claims 1. Encoder fór encoding a picture or videó source data (160) intő a quality-scalable data stream (112), comprising:base encoding means (102) fór encoding the picture or videó source data (160) intő a base encoding data stream representing a representation ofthe picture or videó source data with a first picture sample bit depth;mapping means (104) fór mapping samples ofthe representation ofthe picture or videó source data (160) with the first picture sample bit depth from a first dynamic rangé corresponding to the first picture sample bit depth ΕΡ 2 279 622 Β1 to a second dynamic rangé greater than the first dynamic rangé and corresponding to a second picture sample bit depth being higher than the first picture sample bit depth, by use of one or more global mapping functions being constant within the picture or videó source data (160) or varying at a first granularity, and a local mapping function locally modifying the one or more global mapping functions at a second granularity finer than the first granularity to obtain a prediction of the picture or videó source data having the second picture sample bit depth;residual encoding means (106) fór encoding a prediction residual of the prediction intő a bit-depth enhancement layerdata stream;and combining means (108) fór forming the quality-scalable data stream based on the base encoding data stream, the local mapping function and the bit-depth enhancement layer data stream so that the local mapping function is derivable from the quality-scalable data stream.
- 2Decoderfordecoding a quality-scalable data stream intő which picture or videó source data is encoded, the qualityscalable data stream comprising a base layerdata stream representing the picture or videó source data with a first picture sample bit depth, a bit-depth enhancement layer data stream representing a prediction residual with a second picture sample bit depth being higher than the first picture sample bit depth, and a local mapping function defined at a second granularity, the decoder comprising:means (204) fór decoding the base layer data stream intő a lower bit-depth reconstructed picture or videó data;means (208) fór decoding the bit-depth enhancement data stream intő the prediction residual;means (206) fór mapping samples of the lower bit-depth reconstructed picture or videó data with the first picture sample bit depth from a first dynamic rangé corresponding to the first picture sample bit depth to a second dynamic rangé greater than the first dynamic rangé and corresponding to the second picture sample bit depth, by use of one or more global mapping functions being constant within the videó or varying at a first granularity, and a local mapping function locally modifying the one or more global mapping functions at the second granularity being smaller than the first granularity, to obtain a prediction of the picture or videó source data having the second picture sample bit depth;and means (210) fór reconstructing the picture with the second picture sample bit depth based on the prediction and the prediction residual.
- 6Decoder according to any of claims 2 to 5, wherein the means (208) fór decoding the bit-depth enhancement data stream and the mapping means (104) are adapted such that second granularity subdivides the picture or videó source data (160) intő a plurality of picture blocks (170) and the mapping means (206) is adapted such that the local mapping function is m s + n with m and n varying at the second granularity with the means (208) fór decoding the bit-depth enhancement data stream being adapted to dérivé m and n from the bit-depth enhancement data stream fór each picture block (170) of the picture or videó source data (160) so that m and n may differ among the plurality of picture blocks.
- 7Decoder according to any of claims 2 to 6, wherein the mapping means (206) is adapted such that the second granularity varies within the picture or videó source data (160), and the means (208) fór decoding the bit-depth enhancement data stream is adapted to dérivé the second granularity from the bit-depth enhancement data stream.
- 8Decoder according to any of claims 2 to 7, wherein the means (208) fór decoding the bit-depth enhancement data stream and the mapping means (206) are adapted such that the second granularity divides-up the picture or videó source data (160) intő a plurality of picture blocks (170), and the means (208)fordecoding the bit-depth enhancement data stream is adapted to dérivé a local mapping function residual (Διτι, Δη) from the bit-depth enhancement data stream fór each picture block (170), and dérivé the local mapping function of a predetermined picture block of the picture or videó source data (160) by use of a spatial and/or temporal prediction from one or more neighboring picture blocks or a corresponding picture block of a picture of the picture or videó source data preceding a picture ΕΡ 2 279 622 Β1 to which the predetermined picture block belongs, and the local mapping function residual of the predetermined picture block.
- 11Decoder according to any of claims 2 to 10, wherein the mapping means (206) is adapted such that the at least one ofthe global mapping functions is defined as 2M-N-K x + 2M-1.2N-1-K w herein x is a sample of the representation of the picture or videó source data with the first picture sample bit-depth, N is the first picture sample bit-depth, M is the second picture sample bit-depth, and K is a mapping paraméter, 2M-n-k x + o, where N is the first picture sample bit depth, M is the second picture sample bit depth, and K und D are a mapping parameters, floor (2 m ~ n ~ k x + 2 m ~ 2N ~ k x + D), where floor (a) rounds a down to the nearest integer, N is the first picture sample bit depth, M is the second picture sample bit depth, and K und D are a mapping parameters, a piece-wise linear function fór mapping the samples from the first dynamic rangé to the second dynamic rangé with interpolation point Information defining the piece-wise linear mapping, or a look-up table fór being indexed by use of samples from the first dynamic rangé und outputting samples ofthe second dynamic rangé thereupon.
- 13Method fór encoding a picture or videó source data (160) intő a quality-scalable data stream (112), comprising:encoding the picture or videó source data (160) intő a base encoding data stream representing a representation ofthe picture or videó source data with a first picture sample bit depth;mapping samples of the representation of the picture or videó source data (160) with the first picture sample bit depth from a first dynamic rangé corresponding to the first picture sample bit depth to a second dynamic rangé greaterthan the first dynamic rangé and corresponding to a second picture sample bit depth being higher than the first picture sample bit depth, by use of one or more global mapping functions being constant within the picture or videó source data (160) or varying at a first granularity, and a local mapping function locally modifying the one or more global mapping functions at a second granularity finer than the first granularity to obtain a prediction ofthe picture or videó source data having the second picture sample bit depth;encoding a prediction residual ofthe prediction intő a bit-depth enhancement layerdata stream;and forming the quality-scalable data stream based on the base encoding data stream, the local mapping function and the bit-depth enhancement layerdata stream so that the local mapping function isderivablefrom the qualityscalable data stream.
- 14Method fór decoding a quality-scalable data stream intő which picture or videó source data is encoded, the qualityscalable data stream comprising a base layerdata stream representing the picture or videó source data with a first picture sample bit depth, a bit-depth enhancement layerdata stream representing a prediction residual with a second picture sample bit depth being higher than the first picture sample bit depth, and a local mapping function defined at a second granularity, the method comprising:decoding the base layerdata stream intő a lower bit-depth reconstructed picture or videó data;decoding the bit-depth enhancement data stream intő the prediction residual;mapping samples ofthe lower bit-depth reconstructed picture or videó data with the first picture sample bit depth from a first dynamic rangé corresponding to the first picture sample bit depth to a second dynamic rangé greater than the first dynamic rangé and corresponding to the second picture sample bit depth, by use of one or more global mapping functions being constant within the videó or varying at a first granularity, and a local mapping function locally modifying the one or more global mapping functions at the second granularity being smaller than the first granularity, to obtain a prediction ofthe picture or videó source data having the second picture ΕΡ 2 279 622 Β1 sample bit depth;and reconstructing the picture with the second picture sample bit depth based on the prediction and the prediction residual.
- 15Quality-scalable data stream intő which picture or videó source data is encoded, the quality-scalable data stream comprising a base layer data stream representing the picture or videó source data with a first picture sample bit depth, a bit-depth enhancement layer data stream representing a prediction residual with a second picture sample bit depth being higher than the first picture sample bit depth, and a local mapping function defined at a second granularity, wherein a reconstruction of the picture with the second picture sample bit depth is derivable from the prediction residual and a prediction obtained by mapping samples of the lower bit-depth reconstructed picture or videó data with the first picture sample bit depth from a first dynamic rangé corresponding to the first picture sample bit depth to a second dynamic rangé greater than the first dynamic rangé and corresponding to the second picture sample bit depth, by use of one or more global mapping functions being constant within the videó or varying at a first granularity, and the local mapping function locally modifying the one or more global mapping functions at the second granularity being smaller than the first granularity.
Independent claims11
118 paragraphs, as filed
(56) References cited:
• MARTIN WINKEN ETAL: Bit-Depth Scalable Videó Coding IMAGE PROCESSING, 2007. ICIP 2007. IEEE INTERNATIONAL CONFERENCE ON, IEEE, Pl, 1 September 2007 (2007-09-01), pages I5, XP031157664 ISBN: 978-1-4244-1436-9 • LIU S ET AL: SVC inter-layer pred fór SVC bitdepth scalability VIDEÓ STANDARDS AND DRAFTS, XX, XX, no. JVT-X075, 30 June 2007 (2007-06-30), XP030007182 cited in the application • SEGALL A ET AL: CE1: Inter-layer prediction fór SVC bit-depth scalability 25. JVT MEETING; 82. MPEG MEETING; 21-10-2007 - 26-10-2007; SHENZHEN, CN; (JOINT VIDEÓ TEAM OFISO/IEC JTC1/SC29/WG11 AND ITU-T SG.16 )„ no. JVTY071, 24 October 2007 (2007-10-24), XP030007275 • SEGALL A ET AL: System fór bit-depth scalable coding VIDEÓ STANDARDS AND DRAFTS, XX, XX, no. JVT-W113, 25 April 2007 (2007-04-25), XP030007073 cited in the application
Note: Within nine months ofthe publication ofthe mention ofthe grant ofthe European patent in the European Patent Bulletin, any person may give notice to the European Patent Office of opposition to that patent, in accordance with the Implementing Regulations. Notice of opposition shall nőt be deemed to have been filed until the opposition fee has been paid. (Art. 99(1) European Patent Convention).
Printed by Jouve, 75001 PARIS (FR)
ΕΡ 2 279 622 Β1
Description [0001] The present invention is concerned with picture and/or videó coding, and in particular, quality-scalable coding enabling bit-depth scalability using quality-scalable data streams.
[0002] The Joint Videó Team (JVT) of the ISO/IEC Moving Pictures Experts Group (MPEG) and the ITU-T Videó Coding Experts Group (VCEG) have recently finalized a scalable extension of the state-of-the-art videó coding standard H.264/AVC called Scalable Videó Coding (SVC). SVC supports temporal, spatial and SNR scalable coding of videó sequences or any combination thereof.
[0003] H.264/AVC as described in ITU-T Rec. & ISO/IEC 14496-10 AVC, Advanced Videó Coding fór Generic Audiovisual Services, version 3, 2005, specifies a hybrid videó codec in which macroblock prediction signals are either generated in the temporal domain by motion-compensated prediction, or in the spatial domain by intra prediction, and both predictions are followed by residual coding. H.264/AVC coding without the scalability extension is referred to as single-layer H.264/AVC coding. Rate-distortion performance comparable to single-layer H.264/AVC means that the same visual reproduction quality is typically achieved at 10% bit-rate. Given the above, scalability is considered as a functionality fór removal of parts of the bit-stream while achieving an R-D performance at any supported spatial, temporal or SNR resolution that is comparable to single-layer H.264/AVC coding at that particular resolution.
[0004] The basic design of the scalable videó coding (SVC) can be classified as a layered videó codec. In each layer, the basic concepts of motion-compensated prediction and intra prediction are employed as in H.264/AVC. However, additional inter-layer prediction mechanisms have been integrated in orderto exploit the redundancy between several spatial or SNR layers. SNR scalability is basically achieved by residual quantization, while fór spatial scalability, a combination of motion-compensated prediction and oversampled pyramid decomposition is employed. The temporal scalability approach of H.264/AVC is maintained.
[0005] In generál, the coderstructuredependson the scalability space that is required by an application. Fór illustration, Fig. 8 shows a typical coderstructure 900 with two spatial layers 902a, 902b. In each layer, an independent hierarchical motion-compensated prediction structure 904a,b with layer-specific motion parameters 906a, b is employed. The redundancy between consecutive layers 902a,b is exploited by inter-layer prediction concepts 908 that include prediction mechanisms fór motion parameters 906a,b as well as texture data 910a,b. A base representation 912a,b of the input pictures 914a,b of each layer 902a,b is obtained by transform coding 916a,b similar to that of H.264/AVC, the corresponding NAL units (NAL - Network Abstraction Layer) contain motion Information and texture data; the NAL units of the base representation of the lowest layer, i.e. 912a, are compatible with single-layer H.264/AVC.
[0006] The resulting bit-streams output by the base layer coding 916a,b and the Progressive SNR refinement texture coding 918a,b ofthe respective layers 902a,b, respectively, are multiplexed by a multiplexer 920 in orderto result in the scalable bit-stream 922. This bit-stream 922 is scalable in time, space and SNR quality.
[0007] Summarizing, in accordance with the above scalable extension ofthe Videó Coding Standard H.264/AVC, the temporal scalability is provided by using a hierarchical prediction structure. Fór this hierarchical prediction structure, the one of single-layer H.264/AVC standards may be used without any changes. Fór spatial and SNR scalability, additional tools have to be added to the single-layer H.264/MPEG4.AVC as described in the SVC extension of H.264/AVC. All three scalability types can be combined in order to generate a bit-stream that supports a large degree on combined scalability.
[0008] Problems arise when a videó source _signal has a different dynamic rangé than required by the decoder or player, respectively. In the above current SVC standard, the scalability tools are only specified fór the case that both the base layer and enhancement layer represent a given videó source with the same bit depth of the corresponding arrays of luma and/or chroma samples. Hence, considering different decoders and players, respectively, requiring different bit depths , several coding streams dedicated fór each ofthe bit depths would have to be provided separately. However, in rate/distortion sense, this means an increased overhead and reduced efficiency, respectively.
[0009] There have already been proposals to add a scalability in terms of bit-depth to the SVC Standard. Fór example, Shan Liu et al. describe in the input document to the JVT - namely JVT-X075 - the possibility to dérivé a an inter-layer prediction from a lower bit-depth representation of a base layer by use of an inverse tone mapping according to which an inter-layer predicted or inversely tone-mapped pixel value p’ is calculated from a base layer pixel value p<sub>b</sub> by p’ = p<sub>b </sub> scale + offset with stating that the inter-layer prediction would be performed on macro blocks or smaller block sizes. In JVT-Y067 Shan Liu, presents results fór this inter-layer prediction scheme. Símilarly, Andrew Segall et al. propose in JVT-X071 an inter-layer prediction fór bit-depth scalability according to which a gain plus offset operation is used fór the inverse tone mapping. The gain parameters are indexed and transmitted in the enhancement layer bit-stream on a blockby-block hasis. The signaling of the scale factors and offset factors is accomplished by a combination of prediction and refinement. Further, it is described that high level syntax supports coarsergranularities than the transmission on a blockby-block hasis. Reference is alsó made to Andrew Segall Scalable Coding of High Dynamic Rangé Videó in ICIP 2007, 1-1 to I-4 and the JVT document, JVT-X067 and JVT-W113, alsó stemming from Andrew Segall.
[0010] Although the above-mentioned proposals fór using an inverse tone-mapping in orderto obtain a prediction from
ΕΡ 2 279 622 Β1 a lower bit-depth base layer, remove somé of the redundancy between the lower bit-depth information and the higher bit-depth information, it would be favorable to achieve an even better efficiency in providing such a bit-depth scalable bit-stream, especially in the sense of rate/distortion performance.
[0011] It is the object of the present invention to provide a coding scheme that enables a more efficient way of providing a coding of a picture or videó being suitable fór different bit depths.
[0012] This object is achieved by an encoderaccording to claim 1, adecoderaccording to claim 11, a method according to claim 22 or 23, or a quality-scalable data stream according to claim 24.
[0013] The present invention is based on the finding that the efficiency of a bit-depth scalable data-stream may be increased when an inter-layer prediction is obtained by mapping samples of the representation of the picture or videó source data with a first picture sample bit-depth from a first dynamic rangé corresponding to the first picture sample bitdepth to a second dynamic rangé greater than the first dynamic rangé and corresponding to a second picture sample bit-depth being higher than the first picture sample bit-depth by use of one or more global mapping functions being constant within the picture or videó source data or varying at a first granularity, and a local mapping function locally modifying the one or more global mapping functions and varying at a second granularity smallerthan the first granularity, with forming the quality-scalable data-stream based on the local mapping function such that the local mapping function is derivable from the quality-scalable data-stream. Although the provision of one or more global mapping functions in addition to a local mapping function which, in turn, locally modifies the one or more global mapping functions, primefacie increases the amount of side information within the scalable data-stream, this increase is more than recompensated by the fact that this sub-division intő a global mapping function on the one hand and a local mapping function on the other hand enables that the local mapping function and its parameters fór its parameterization may be small and are, thus, codeable in a highly efficient way. The global mapping function may be coded within the quality-scalable datastream and, since same is constant within the picture or videó source data orvaries at a greater granularity, the overhead or flexibility fór defining this global mapping function may be increased so that this global mapping function may be precisely fit to the average statistics of the picture or videó source data, thereby further decreasing the magnitude of the local mapping function.
[0014] In the following, preferred embodiments of the present application are described with reference to the Figs. In particular, as it is shown in
Fig. 1 a block diagram of a videó encoder according to an embodiment of the present invention;
Fig. 2 a block diagram of a videó decoder according to an embodiment of the present invention;
Fig. 3 a flow diagram fór a possible implementation of a mode of operation of the prediction modulé 134 of Fig. 1 according to an embodiment;
Fig. 4 a schematic of a videó and its sub-division intő picture sequences, pictures, macroblock pairs, macroblocks and transform blocks according to an embodiment;
Fig. 5 a schematic of a portion of a picture, sub-divided intő blocks according to the fine granularity underlying the local mapping/adaptation function with concurrently illustrating a predictive coding scheme fór coding the local mapping/adaptation function parameters according to an embodiment;
Fig. 6 a flow diagram fór illustrating an inverse tone mapping process at the encoder according to an embodiment;
Fig. 7 a flow diagram of an inverse tone mapping process at the decoder corresponding to that of Fig. 6, according to an embodiment; and
Fig. 8 a block diagram of a conventional coder structure fór scalable videó coding.
[0015] Fig. 1 shows an encoder 100 comprising a base encoding means 102, a prediction means 104, a residual encoding means 106 and a combining means 108 as well as an input 110 and an output 112. The encoder 100 of Fig. 1 is a videó encoder receiving a high quality videó signal at input 110 and outputting a quality-scalable bit stream at output 112. The base encoding means 102 encodes the data at input 110 intő a base encoding data stream representing the content of this videó signal at input 110 with a reduced picture sample bit-depth and, optionally, a decreased spatial resolution compared to the input signal at input 110. The prediction means 104 is adapted to, based on the base encoding data stream output by base encoding means 102, provide a prediction signal with full or increased picture sample bitdepth and, optionally, full or increased spatial resolution fór the videó signal at input 110. Asubtractor 114 alsó comprised by the encoder 100 forms a prediction residual of the prediction signal provided by means 104 relatíve to the high quality
ΕΡ 2 279 622 Β1 inputsignal at input 110, theresidualsignal being encoded bytheresidualencoding means 106 intoaqualityenhancement layerdata stream. The combining means 108 combines the base encoding data stream írom the base encoding means 102 and the quality enhancement layer data stream output by residual encoding means 106 to form a quality scalable data stream 112 at the output 112. The quality-scalability means that the data stream at the output 112 is composed of a part that is self-contained in that it enables a reconstruction of the videó signal 110 with the reduced bit-depth and, optionally, the reduced spatial resolution without any further Information and with neglecting the remainder of the data stream 112, on the one hand and a further part which enables, in combination with the first part, a reconstruction of the videó signal at input 110 in the original bit-depth and original spatial resolution being higher than the bit depth and/or spatial resolution of the first part.
[0016] After having rather generally described the structure and the functionality of encoder 100, its internál structure is described in more detail below. In particular, the base encoding means 102 comprises a down conversion modulé 116, a subtractor 118, a transform modulé 120 and a quantization modulé 122 serially connected, in the order mentioned, between the input 110, and the combining means 108 and the prediction means 104, respectively. The down conversion modulé 116 is fór reducing the bit-depth of the picture samples of and, optionally, the spatial resolution of the pictures of the videó signal at input 110. In other words, the down conversion modulé 116 irreversibly down-converts the high quality input videó signal at input 110 to a base quality videó signal. As will be described in more detail below, this downconversion may include reducing the bit-depth of the signal samples, i.e. pixel values, in the videó signal at input 110 using any tone-mapping scheme, such as rounding of the sample values, sub-sampling of the chroma components in case the videó signal is given in the form of luma plus chroma components, filtering of the inputsignal at input 110, such as by a RGB to YCbCr conversion, or any combination thereof. More details on possible prediction mechanisms are presented in thefollowing. In particular, it is possible that the down-conversion modulé 116 usesdifferentdown-conversion schemes fór each picture of the videó signal or picture sequence input at input 110 or uses the same scheme fór all pictures. This is alsó discussed in more detail below.
[0017] The subtractor 118, the transform modulé 120 and the quantization modulé 122 co-operate to encode the base quality signal output by down-conversion modulé 116 by the use of, fór example, a non-scalable videó coding scheme, such as H.264/AVC. According to the example of Fig. 1, the subtractor 118, the transform modulé 120 and the quantization modulé 122 co-operate with an optional prediction loop filter 124, a predictor modulé 126, an inverse transform modulé 128, and an adder 130 commonly comprised by the base encoding means 102 and the prediction means 104 to form the irrelevance reduction partof a hybrid encoder which encodes the base quality videó signal output by down-conversion modulé 116 by motion-compensation based prediction and following compression ofthe prediction residual. In particular, the subtractor 118 subtracts from a current picture or macroblock ofthe base quality videó signal a predicted picture or predicted macroblock portion reconstructed from previously encoded pictures of the base quality videó signal by, fór example, use of motion compensation. The transform modulé 120 applies a transform on the prediction residual, such as a DCT, FFT or wavelet transform. The transformed residual signal may represent a spectral representation and its transform coefficients are irreversibly quantized in the quantization modulé 122. The resulting quantized residual signal represents the residual ofthe base-encoding data stream output by the base-encoding means 102.
[0018] Apart from the optional prediction loop filter 124 and the predictor modulé 126, the inverse transform modulé 128, and the adder 130, the prediction means 104 comprises an optional filter fór reducing coding artifacts 132 and a prediction modulé 134. The inverse transform modulé 128, the adder 130, the optional prediction loop filter 124 and the predictor modulé 126 co-operate to reconstruct the videó signal with a reduced bit-depth and, optionally, a reduced spatial resolution, as defined by the down-conversion modulé 116. In other words, they create a low bit-depth and, optionally, low spatial resolution videó signal to the optional filter 132 which represents a low quality representation of the source signal at input 110 alsó being reconstructable at decoder side. In particular, the inverse transform modulé 128 and the adder 130 are serially connected between the quantization modulé 122 and the optional filter 132, whereas the optional prediction loop filter 124 and the prediction modulé 126 are serially connected, in the order mentioned, between an output of the adder 130 as well as a further input of the adder 130. The output of the predictor modulé 126 is alsó connected to an inverting input of the subtractor 118. The optional filter 132 is connected between the output of adder 130 and the prediction modulé 134, which, in turn, is connected between the output of optional filter 132 and the inverting input of subtractor 114.
[0019] The inverse transform modulé 128 inversely transforms the base-encoded residual pictures output by baseencoding means 102 to achieve low bit-depth and, optional, low spatial resolution residual pictures. Accordingly, inverse transform modulé 128 performs an inverse transform being an inversion ofthe transformation and quantization performed by modules 120 and 122. Alternatively, a de-quantization modulé may be separately provided at the input side ofthe inverse transform modulé 128. The adder 130 adds a prediction to the reconstructed residual pictures, with the prediction being based on previously reconstructed pictures ofthe videó signal. In particular, the adder 130 outputs a reconstructed videó signal with a reduced bit-depth and, optionally, reduced spatial resolution. These reconstructed pictures are filtered by the loop filer 124 fór reducing artifacts, fór example, and used thereafter by the predictor modulé 126 to predict the picture currently to be reconstructed by means of, fór example, motion compensation, from previously reconstructed
ΕΡ 2 279 622 Β1 pictures. The base quality signal thus obtained at the output of adder 130 is used by the serial connection of the optional filter 132 and prediction modulé 134 to get a prediction of the high quality input signal at input 110, the latter prediction to be used fór forming the high quality enhancement signal at the output of the residual encoding means 106. This is described in more detail below.
[0020] In particular, the low quality signal obtained from adder 130 is optionally filtered by optional filter 132forreducing coding artifacts. Filters 124 and 132 may even operate thesameway and, thus, althoughfilters 124 and 132areseparately shown in Fig. 1, both may be replaced by just one filter arranged between the output of adder 130 and the input of prediction modulé 126 and prediction modulé 134, respectively. Thereafter, the low quality videó signal is used by prediction modulé 134 to form a prediction signal fór the high quality videó signal received at the non-inverting input of adder 114 being connected to the input 110. This process of forming the high quality prediction may include mapping the decoded base quality signal picture samples by use of a combined mapping function as described in more detail below, using the respective value of the base quality signal samples fór indexing a look-up table which contains the corresponding high quality sample values, using the value of the base quality signal sample fór an interpolation process to obtain the corresponding high quality sample value, up-sampling of the chroma components, filtering of the base quality signal by use of, fór example, YCbCr to RGB conversion, or any combination thereof. Other examples are described in the following.
[0021] Fór example, the prediction modulé 134 may map the samples of the base quality videó signal from a first dynamic rangé to a second dynamic rangé being higher than the first dynamic rangé and, optionally, by use of a spatial interpolation filter, spatially interpolate samples of the base quality videó signal to increase the spatial resolution to correspond with the spatial resolution of the videó signal at the input 110. Ina way similar to the above description of the down-conversion modulé 116, it is possible to use a different prediction process fór different pictures of the base quality videó signal sequence as well as using the same prediction process fór all the pictures.
[0022] The subtractor 114 subtracts the high quality prediction received from the prediction modulé 134 from the high quality videó signal received from input 110 to output a prediction residual signal of high quality, i.e. with the original bitdepth and, optionally, spatial resolution to the residual encoding means 106. At the residual encoding means 106, the difference between the original high quality input signal and the prediction derived from the decoded base quality signal is encoded exemplarily using a compression coding scheme such as, fór example, specified in H.264/AVC. To this end, the residual encoding means 106 of Fig. 1 comprises exemplarily a transform modulé 136, a quantization modulé 138 and an entropy coding modulé 140 connected in series between an output of the subtractor 114 and the combining means 108 in the mentioned order. The transform modulé 136 transforms the residual signal or the pictures thereof, respectively, intő a transformation domain orspectral domain, respectively, where the spectral components are quantized by the quantization modulé 138 and with the quantized transform values being entropy coded by the entropy-coding modulé 140. The result of the entropy coding represents the high quality enhancement layerdata stream output by the residual encoding means 106. If modules 136 to 140 implement an H.-264-/AVC coding, which supports transforms with a size of 4x4 or 8x8 samples fór coding the luma content, the transform size fór transforming the luma component of the residual signal from the subtractor 114 in the transform modulé 136 may arbitrarily be chosen fór each macroblock and does nőt necessarily have to be the same as used fór coding the base quality signal in the transform modulé 120. Fór coding the chroma components, the H.264/AVC standard, provides no choice. When quantizing the transform coefficients in the quantization modulé 138, the same quantization scheme as in the H.264/AVC may be used, which means that the quantizerstep-size may be controlled by a quantization paraméter QP, which can take values from -6*(bit depth of high quality videó signal component- 8) to 51. The QP used fór coding the base quality representation macroblock in the quantization modulé 122 and the QP used fór coding the high quality enhancement macroblock in the quantization modulé 138 do nőt have to be the same.
[0023] Combining means 108 comprises an entropy coding modulé 142 and the multiplexer 144. The entropy-coding modulé 142 is connected between an output of the quantization modulé 122 and a first input of the multiplexer 144, whereas a second input of the multiplexer 144 is connected to an output of entropy coding modulé 140. The output of the multiplexer 144 represents output 112 of encoder 100.
[0024] The entropy encoding modulé 142 entropy encodes the quantized transform values output by quantization modulé 122 to form a base quality layerdata stream from the base encoding data stream output by quantization modulé 122. Therefore, as mentioned above, modules 118,120,122,124,126,128,130 and 142 may be designed to co-operate in accordance with the H.264/AVC, and represent together a hybrid coder with the entropy coder 142 performing a lossless compression of the quantized prediction residual.
[0025] The multiplexer 144 receives both the base quality layer data stream and the high quality layer data stream and puts them together to form the quality-scalable data stream.
[0026] As already described above, and as shown in Fig. 3, the way in which the prediction modulé 134 performs the prediction from the reconstructed base quality signal to the high quality signal domain may include an expansion of sample bit-depth, alsó called an inverse tone mapping 150 and, optionally, a spatial upsampling operation, i.e. an upsampling filtering operation 152 in case base and high quality signals are of different spatial resolution. The order in
ΕΡ 2 279 622 Β1 which prediction modulé 134 performs the inverse tone mapping 150 and the optional spatial up-sampling operation 152 may be fixed and, accordingly, ex ante be known to both encoder and decodersides, or it can be adaptively chosen on a block-by-blockorpicture-by-picture basis orsome othergranularity, in which casethe prediction modulé 134signals information on the order between steps 150 and 152 used, to somé entity, such as the coder 106, to be introduced as side information intő the bit stream 112 so as to be signaled to the decoder side as part of the side information. The adaptation of the order among steps 150 and 152 is i11ustrated in Fig. 3 by use of a dotted double-headed arrow 154 and the granularity at which the order may adaptively be chosen may be signaled as well and even varied within the videó. [0027] In performing the inverse tone mapping 150, the prediction modulé 134 uses two parts, namely one or more global inverse tone mapping functions and a local adaptation thereof. Generally, the one or more global inverse tone mapping functions are dedicated fór accounting fór the generál, average characteristics of the sequence of pictures of the videó and, accordingly, of the tone mapping which has initially been applied to the high quality input videó signal to obtain the base quality videó signal at the down conversion modulé 116. Compared thereto, the local adaptation shall account fór the individual deviations from the global inverse tone mapping model fór the individual blocks of the pictures of the videó.
[0028] In order to illustrate this, Fig. 4 shows a portion of a videó 160 exemplary consisting offour consecutive pictures 162a to 162d of the videó 160. In other words, the videó 160 comprises a sequence of pictures 162 among which four are exemplary shown in Fig. 4. The videó 160 may be divided-up intő non-overlapping sequences of consecutive pictures fór which global parameters or syntax elements are transmitted within the data stream 112. Fór illustration purposes only, it is assumed that the four consecutive pictures 162a to 162d shown in Fig. 4 shall form such sequence 164 of pictures. Each picture, in turn, is sub-divided intő a plurality of macroblocks 166 as illustrated in the bottom left-hand corner of picture 162d. A macroblock is a Container within which transformation coefficients along with other control syntax elements pertaining the kind of coding of the macroblock is transmitted within the bit-stream 112. A pair 168 of macroblocks 166 covers a continuous portion of the respective picture 162d. Depending on a macroblock pair mode of the respective macroblock pair 168, the top macroblock 162 of this pair 168 covers either the samples of the upper half of the macroblock pair 168 or the samples of every odd-numbered line within the macroblock pair 168, with the bottom macroblock relating to the other samples therein, respectively. Each macroblock 160, in turn, may be sub-divided intő transform blocks as illustrated at 170, with these transform blocks forming the block basis at which the transform modulé 120 performs the transformation and the inverse transformation modulé 128 performs the inverse transformation. [0029] Referring back to the just-mentioned global inverse tone mapping function, the prediction modulé 134 may be configured to use one or more such global inverse tone mapping functions constantly fór the whole videó 160 or, alternatively, fór a sub-portion thereof, such as the sequence 164 of consecutive pictures or a picture 162 itself. The latter options would imply that the prediction modulé 134 varies the global inverse tone mapping function at a granularity corresponding to a picture sequence size or a picture size. Examples fór global inverse tone mapping functions are given in the following. If the prediction modulé 134 adapts the global inverse tone mapping function to the statistics of the videó 160, the prediction modulé 134 outputs information to the entropy coding modulé 140 orthe multiplexer 144 so that the bit stream 112 contains information on the global inverse tone mapping function(s) and the variation thereof within the videó 160. In case the one or more global inverse tone mapping functions used by the prediction modulé 134 constantly apply to the whole videó 160, same may be ex ante known to the decoder or transmitted as side information within the bit stream 112.
[0030] At an even smaller granularity, a local inverse tone mapping function used by the prediction modulé 134 varies within the videó 160. Fór example, such local inverse tone mapping function varies at a granularity smaller than a picture size such as, fór example, the size of a macroblock, a macroblock pair or a transformation block size.
[0031] Fór both the global inverse tone mapping function(s) and the local inverse tone mapping function, the granularity at which the respective function varies or at which the functions are defined in bit stream 112 may be varied within the videó 160. The variance of the granularity, in turn, may be signaled within the bit stream 112.
[0032] During the inverse tone mapping 150, the prediction modulé 134 maps a predetermined sample of the videó 160 from the base quality bit-depth to the high quality bit-depth by use of a combination of one global inverse tone mapping function applying to the respective picture and the local inverse tone mapping function as defined in the respective block to which the predetermined sample belongs.
[0033] Fór example, the combination may be an arithmetic combination and, in particular, an addition. The prediction modulé may be configured to obtain a predicted high bit-depth sample value s<sub>high</sub> from the corresponding reconstructed low bit-depth sample value s,<sub>ow</sub> by use of s<sub>high</sub> = f<sub>k</sub>(s,<sub>ow</sub>) + m s,<sub>ow</sub> + n.
[0034] In this formula, the function f<sub>k</sub> represents a global inverse tone mapping operator wherein the index k selects which global inverse tone mapping operator is chosen in case more than one single scheme or more than one global inverse tone mapping function is used. The remaining part of this formula constitutes the local adaptation or the local inverse tone mapping function with n being an offset value and m being a scaling factor. The values of k, m and n can be specified on a block-by-block basis within the bit stream 112. In other words, the bit stream 112 would enable revealing the triplets {k, m, n) fór all blocks of the videó 160 with a block size of these blocks depending on the granularity of the
EP 2 279 622 Β1 local adaptation of the global inverse tone mapping function with this granularity, in turn, possibly varying within the videó 160.
[0035] The following mapping mechanisms may be used fór the prediction process as far as the global inverse time mapping function f(x) is concerned. Fór example, piece-wise linear mapping may be used where an arbitrary number of interpolation points can be specified. Fór example, fór a base quality sample with value x and two given interpolation points (x<sub>n</sub>, y<sub>n</sub>) and (x<sub>n+</sub>-|, y<sub>n+1</sub>) the corresponding prediction sample y is obtained by the modulé 134 according to the following formula f(x) = y<sub>n</sub> +* .-<sup>X</sup>_ (y<sub>n</sub>+i—y<sub>n</sub>)
Xn+\~ Xn [0036] This linear interpolation can be performed with little computational complexity by using only bit shift instead of division operations if x<sub>n+1</sub> - x<sub>n</sub> is restricted to be a power of two.
[0037] Afurther possible global mapping mechanism represents a look-up table mapping in which, by means of the base quality sample values, a table look-up is performed in a look-up table in which fór each possible base quality sample value as far as the global inverse tone mapping function is concerned the corresponding global prediction sample value (x) is specified. The look-up table may be provided to the decoder side as side Information or may be known to the decoder side by default.
[0038] Further, scaling with a constant offset may be used fór the global mapping. According to this alternative, in ordertoachieve the corresponding high quality global prediction sample (x) having higher bit-depth, modulé 134 multiplies the base quality samples x by a constant factor 2<sup>M_N_K</sup>, and afterwards a constant offset 2<sup>M_1</sup>-2<sup>M</sup><sup>1</sup><sup>K</sup> is added, according to, fór example, one of the following formuláé:
f (x)=2<sup>w</sup>-<sup>n</sup>-<sup>k</sup>x + 2<sup>M_1</sup> - 2-<sup>1</sup>-* or r(x)=min(2 x + 2 -2 , 2-1), respectively, wherein M is the bit-depth of the high quality signal and N is the bit-depth of the base quality signal.
[0039] By this measure, the low quality dynamic rangé [0;2<sup>N</sup>-1] is mapped to the second dynamic rangé [0;2<sup>M</sup>-1] in a manner according to which the mapped values of x are distributed in a centralised manner with respect to the possible dynamic rangé [0;2<sup>M</sup>-1] of the higher quality within a extension which is determined by K. The value of K could be an integer value or reál value, and could be transmitted as side Information to the decoder within, fór example, the qualityscalable data stream so that at the decoder somé predicting means may act the same way as the prediction modulé 134 as will be described in the following. A round operation may be used to get integer valued f(x) values.
[0040] Another possibility forthe global scaling is scaling with variable offset: the base quality samples x are multiplied by a constant factor, and afterwards a variable offset is added, according to, fór example, one of the following formuláé:
f(x)=2<sup>M-N</sup><sup>K</sup>x + D or f (x) =min (2<sup>m_w_x</sup>x + D, 2<sup>m</sup> - 1 ) [0041] By this measure, the low quality dynamic rangé is globally mapped to the second dynamic rangé in a manner according to which the mapped values of x are distributed within a portion of the possible dynamic rangé of the highquality samples, the extension of which is determined by K, and the offset of which with respect to the lower boundary is determined by D. D may be integer or reál. The result f(x) represents a globally mapped picture sample value of the high bit-depth prediction signal. The values of K and D could be transmitted as side information to the decoder within,
ΕΡ 2 279 622 Β1 fór example, the quality-scalable data stream. Again, a round operation may be used to get integer valued f(x) values, the latter being true alsó fór the otherexamples given in the present application fór the global bit-depth mappings without explicitly stating it repeatedly.
[0042] An even further possibiIity fór global mapping is scaling with superposition: the globally mapped high bit depth prediction samples f(x) are obtained from the respective base quality sample x according to, fór example, one of the following formuláé, where floor (a) rounds a down to the nearest integer:
f (x)= floor (2<sup>M_N</sup>x + 2<sup>M-2w</sup>x)
ΟΓ f (x) = min( floor (2<sup>m-n</sup>x + 2'<sup>2N</sup>x) , 2 - 1 ) [0043] Thejust mentioned possibilities may be combined. Fór example, global scaling with superposition and constant offset may be used: the globally mapped high bit depth prediction samples f(x) are obtained according to, fór example, one of the following formuláé, where floor(a) rounds a down to the nearest integer:
f (x)=floor (2<sup>M</sup>'<sup>N</sup>~<sup>K</sup>x + 2<sup>m</sup>~<sup>2n</sup>'<sup>k</sup>x + 2<sup>W_1</sup> - 2<sup>1</sup>'*) f (x)=min( floor (2-<sup>n</sup>-<sup>k</sup>x + 2<sup>M</sup>~<sup>2N</sup>~<sup>K</sup>x + 2’<sup>1</sup> - 2^^),2^1) [0044] The value of K may be specified as sídé Information to the decoder.
[0045] Similarly, global scaling with superposition and variable offset may be used: the globally mapped high bit depth prediction samples (x) are obtained according to the following formula, where floor(a) rounds a down to the nearest integer:
f (x)=floor (2<sup>m</sup>~<sup>n</sup>~<sup>k</sup>x + 2<sup>m</sup>~<sup>2N</sup>~<sup>k</sup>x + D) f(x)=min( floor (2<sup>M</sup>~<sup>N_K</sup>x + 2<sup>M_2N</sup>'<sup>K</sup>x + D) , 2<sup>M</sup> - 1 ) [0046] The values of D and K may be specified as sídé Information to the decoder.
[0047] Transferring thejust-mentioned examplesforthe global inverse tone mapping function to Fig. 4, the parameters mentioned there fór defining the global inverse tone mapping function, namely (x^ y^, K and D, may be known to the decoder, may be transmitted within the bit stream 112 with respect to the whole videó 160 in case the global inverse tone mapping function is constant within the videó 160 or these parameters are transmitted within the bit stream fór different portions thereof, such as, fór example, fór picture sequences 164 or pictures 162 depending on the coarswe granularity underlying the global mapping function. In case the prediction modulé 134 uses more than one global inverse tone mapping function, the aforementioned parameters (x^ y^, K and D may be thought of as being provided with an index k with these parameters (x-|, y^, K<sub>k</sub> and D<sub>k</sub> defining the k<sup>th</sup> global inverse tone mapping function f<sub>k</sub>.
[0048] Thus, itwould be possible to specify usage of different global inverse tone mapping mechanismsforeach block of each picture by signaling a corresponding value k within the bit stream 112 as well as, alternatively, using the same mechanism fór the complete sequence of the videó 160.
[0049] Further, it is possible to specify different global mapping mechanisms fór the luma and the chroma components ofthe base quality signal to take intő accountthat thestatistics, such as their probability density function, may be different. [0050] As already denoted above, the prediction modulé 134 uses a combination ofthe global inverse tone mapping function and a local inverse tone mapping function fór locally adapting the global one. In other words, a local adaptation is performed on each globally mapped high bit-depth prediction sample f(x) thus obtained. The local adaptation ofthe
EP 2 279 622 Β1 global inverse tone mapping function is performed by use of a local inverse tone mapping function which is locally adapted by, fór example, locally adapting somé parameters thereof. In the example given above, these parameters are the scaling factor m and the offset value n. The scaling factors m and the offset value n may be specified on a blockby-block hasis where a block may correspond to a transform block size which is, fór example, in case of H.264/AVC 4x4 or 8x8 samples, or a macroblock size which is, in case of H.264/AVC, fór example, 16x16 samples. Which meaning of block is actually used may either be fixed and, therefore, ex ante known to both encoder and decoder or it can be adaptively chosen per picture or per sequence by the prediction modulé 134, in which case it has to be signaled to the decoder as part of sídé Information within the bit stream 112. In case of H.264/AVC, the sequence paraméter set and/or the picture paraméter set could be used to this end. Furthermore, it could be specified in the sídé Information that either the scaling factor m or the offset value n or both values are set equal to zero fór a complete videó sequence or fór a well-defined set of pictures within the videó sequence 160. In case either the scaling factor m or the offset value n or both of them are specified on a block-by-block hasis fór a given picture, in orderto reduce the required bit-rate fór coding of these values, only the difference values Am, Δη to corresponding predicted values m<sub>pred</sub>, n<sub>pred</sub> may be coded such that the actual values fór m, n are obtainable as follows: m=n<sub>pred</sub> + Am, n=n<sub>pred</sub> + Δη.
[0051] In otherwords and as illustrated in Fig. 5, which shows a portion of a picture 162 divided-up intő blocks 170 forming the basis of the granularity of the local adaptation of the global inverse tone mapping function, the prediction modulé 134, the entropy coding modulé 140 and the multiplexer 144 are configured such that the parameters fór locally adapting the global inverse tone mapping function, namely the scaling factor m and the offset value n as used fór the individual blocks 170 are notdirectly coded intő the bit stream 112, bút merely as a prediction residual to a prediction obtained from scaling factors and offset values of neighboring blocks 170. Thus, {Am, Δη, k} are transmitted fór each block 170 in case more than one global mapping function is used as is shown in Fig. 5. Assume, fór example, that the prediction modulé 134 used parameters ιτηj and nj j fór inverse tone mapping the samples of a certain block ij, namely the middle one of Fig. 5. Then, the prediction modulé 134 is configured to compute prediction values m, j <sub>pred</sub> and n, j <sub>pred </sub>from the scaling factors and offset values of neighboring blocks 170, such as, fór example, mjj^ and n, In this case, the prediction modulé 134 would cause the difference between the actual parameters ττη j and nj j and the predicted ones <sup>m</sup>i,F<sup>n</sup>i,j,pred <sup>ar|</sup>d <sup>n</sup>i,f<sup>n</sup>i,j,pred t° inserted intő the bit-stream 112. These differences are denoted in Fig. 5 as Διττήj and Anjj with the indicies ij indicating the j<sup>th</sup> block from the top and the i<sup>th</sup> block from the left-hand sídé of the picture 162, fór example. Alternatively, the prediction of the scaling factor m and the offset value n may be derived from the already transmitted blocks ratherthan neighboring blocks ofthe same picture. Fór example, the prediction values may be derived from the block 170 of the preceding picture lying at the same or a corresponding spatial location. In particular, the predicted value can be a fixed value, which is either transmitted in the sídé Information or already known to both encoder and decoder, the value ofthe corresponding variable and the preceding block, the médián value ofthe corresponding variables in the neighboring blocks, the mean value ofthe corresponding variables in the neighboring blocks, a linear interpolated or extrapolated value derived from the values of corresponding variables in the neighboring blocks.
[0052] Which one of these prediction mechanisms is actually used fór a particular block may be known to both encoder and decoder or may depend on values of m,n in the neighboring blocks itself, if there are any.
[0053] Within the coded high quality enhancement signal output by entropy coding modulé 140, the following Information could be transmitted fór each macroblock in case modules 136, 138 and 140 implement an H.264/AVC conforming encoding. A coded block pattern (CBP) Information could be included indicating as to which of the four 8x8 luma transformation blocks within the macroblock and which ofthe associated chroma transformation blocks ofthe macroblock may contain non-zero transform coefficients. If there are no non-zero transform coefficients, no further Information is transmitted fór the particular macroblock. Further Information could relate to the transform size used fór coding the luma component, i.e. the size ofthe transformation blocks in which the macroblock consisting of 16x16 luma samples is transformed in the transform modulé 136, i.e. in 4x4 or 8x8 transform blocks. Further, the high quality enhancement layer data stream could include the quantization paraméter QP used in the quantization modulé 138 fór controlling the quantizerstep-size. Further, the quantized transform coefficients, i.e. the transform coefficient levels, could be included fór each macroblock in the high quality enhancement layer data stream output by entropy coding modulé 140.
[0054] Besides the above Information, the following Information should be contained within the data stream 112. Fór example, in case more than one single global inverse tone mapping scheme is used fór the current picture, alsó the corresponding index value k has to be transmitted fór each block of the high quality enhancement signal. As far as the local adaptation is concerned, the variables Am and Δη may be signaled fór each block of those pictures where the transmission ofthe corresponding values is indicated in the sídé Information by means of, fór example, the sequence paraméter set and/or the picture paraméter set in H.264/AVC. Fór all three new variables k, Am and Δη, new syntax
EP 2 279 622 Β1 elements with a corresponding binarization schemes have to be introduced. A simple unary binarization may be used in order to prepare the three variables fór the binary arithmetic coding scheme used in H.264/AVC. Since all the three variables have typically small magnitudes, the simple unary binarization scheme is very well suited. Since Δη and Δη are signed integer values, they could be converted to unsigned values as described in the following Table:
<td> signed value of Am, An</td><td> unsigned value (to be binarized)</td>
<td> 0</td><td> 0</td>
<td> 1</td><td> 1</td>
<td> 2</td><td> -1</td>
<td> 3</td><td> 2</td>
<td> 4</td><td> -2</td>
<td> 5</td><td> 3</td>
<td> 6</td><td> -3</td>
<td> k</td><td> (-1)<sup>k+1</sup> Ceil(k+2)</td>
[0055] Thus, summarizing somé of the above embodiments, the prediction means 134 may perform the following steps during its mode of operation. In particular, as shown in Fig. 6, the prediction means 134 sets, in step 180, one or more global mapping function(s) f(x) or, in case of more than one global mapping function, f<sub>k</sub>(x) with k denoting the corresponding index value pointing to the respective global mapping function. As described above, the one or more global mapping function(s) may be set or defined to be constant across the videó 160 or may be defined or set in a coarse granularity such as, fór example, a granularity of the size of a picture 162 or a sequence 164 of pictures. The one or more global mapping function(s) may be defined as described above. In generál, the global mapping function(s) may be a non-trivial function, i.e. unequal to f(x)=constant, and may especially be a non-linear function. In any case, when using any of the above examples fór a global mapping function, step 180 results in a corresponding global mapping function paraméter, such as K,D or (x<sub>n</sub>,y<sub>n</sub>) with n e (1, ... N) being set fór each section intő which the videó is sub-divided according to the coarse granularity.
[0056] Further, the prediction modulé 134 sets a local mapping/adaptation function, such as that indicated above, namely a linear function according to m x + n. However, another local mapping/adaptation function is alsó feasible such as a constant function merely being parametrized by use of an offset value. The setting 182 is performed at a finer granularity such as, fór example, the granularity of a size smaller than a picture such as a macroblock, transform block or macroblock pair or even a slice within a picture wherein a slice is a sub-set of macroblocks or macroblock pairs of a picture. Thus, step 182 results, in case of the above embodiment fór a local mapping/adaptation function, in a pair of values Am and Δη being set or defined fór each of the blocks intő which the pictures 162 of the videó are sub-divided according to the finer granularity.
[0057] Optionally, namely in case more than one global mapping function is used in step 180, the prediction means 134 sets, in step 184, the index k fór each block of the fine granularity as used in step 182.
[0058] Although it would be possible fór the prediction modulé 134 to set the one or more global mapping function (s) in step 180 solely depending on the mapping function used by the down conversion modulé 116 in order to reduce the sample bit-depth of the original picture samples, such as, fór example, by using the inverse mapping function thereto as a global mapping function in step 180, it should be noted that it is alsó possible that the prediction means 134 performs all settings within steps 180, 182 and 184 such that a certain optimization criterion, such as the rate/distortion ratio of the resulting bit-stream 112 is extremized such as maximized or minimized. By this measure, both the global mapping function(s) and the local adaptation thereto, namely the local mapping/adaptation function are adapted to the sample value tone statistics of the videó with determining the best compromise between the overhead of coding the necessary side Information fór the global mapping on the one hand and achieving the best fit of the global mapping function to the tone statistics necessitating merely a small local mapping/adaptation function on the other hand.
[0059] In performing the actual inverse tone mapping in step 186, the prediction modulé 134 uses, fór each sample of the reconstructed low-quality signal, a combination of the one global mapping function in case there is just one global mapping function orone of the global mapping functions in case there are more than one on the one hand, and the local mapping/adaptation function on the other hand, both as defined at the block or the section to which the current sample belongs. In the example above, the combination has been an addition. However, any arithmetic combination would alsó be possible. Fór example, the combination could be a serial application of both functions.
[0060] Further, in order to inform the decoder side about the used combined mapping function used fór the inverse
ΕΡ 2 279 622 Β1 tone mapping, the prediction modulé 134 causes at least Information on on the local mapping/adaptation function to be provided to the decoder sídé via the bit-stream 112. Fór example, the prediction modulé 134 causes, in step 188, the values of m and n to be coded intő the bit-stream 112 fór each block of the fine granularity. As described above, the coding in step 188 may be a predictive coding according to which prediction residuals of m and n are coded intő the bitstream ratherthan the actual values itself, with deriving the prediction of the values being derived from the values of m and n at neighboring blocks 170 or of a corresponding block in a previous picture. In otherwords, a local ortemporal prediction along with residual coding may be used in order to code paraméter m and n.
[0061] Similarly, fór the case that more than one global mapping function is used in step 180, the prediction modulé 134 may cause the index k to be coded intő the bit-stream fór each block in step 190. Further, the prediction modulé 134 may cause the Information on f(x) or f<sub>k</sub>(x) to be coded intő the bit-stream in case same Information is nőt a priori known to the decoder. This is done in step 192. In addition, the prediction modulé 134 may cause Information on the granularity and the change of the granularity within the videó fór the local mapping/adaptation function and/or the global mapping function(s) to be coded intő the bit-stream in step 194.
[0062] It should be noted that all steps 180 to 194 do nőt need to be performed in the order mentioned. Evén a strici sequential performance of these steps is nőt necessary. Rather, the steps 180 to 194 are shown in a sequential order merely fór illustrating purposes and these steps will preferably be performed in an overlapping manner.
[0063] Although nőt explicitly stated in the above description, it is noted that the side Information generated in steps 190 to 194 may be introduced intő the high quality enhancement layer signal or the high quality portion of bit-stream 112 rather than the base quality portion stemming from the entropy coding modulé 142.
[0064] After having described an embodiment fór an encoder, with respect to Fig. 2, an embodiment of a decoder is described. The decoder of Fig. 2 is indicated by reference sign 200 and comprises a de-multiplexing means 202, a base decoding means 204, a prediction means 206, a residual decoding means 208 and a reconstruction means 210 as well as an input 212, a first output 214 and a second output 216. The decoder 200 receives, at its input 212, the qualityscalable data stream, which has, fór example, been output by encoder 100 of Fig. 1. As described above, the quality scalability may relate to the bit-depth and, optionally, to the spatial reduction. In otherwords, the data stream atthe input 212 may have a self-contained part which is isolatedly usable to reconstruct the videó signal with a reduced bit-depth and, optionally, reduced spatial resolution, as well as an additional part which, in combination with the first part, enables reconstructing the videó signal with a higher bit-depth and, optionally, higher spatial resolution. The lower quality reconstruction videó signal is output at output 216, whereasthe higher quality reconstruction videó signal is output at output 214. [0065] The demultiplexing means 202 divides up the incoming quality-scalable data stream at input 212 intő the base encoding data stream and the high quality enhancement layer data stream, both of which have been mentioned with respect to Fig. 1. The base decoding means 204 is fór decoding the base encoding data stream intő the base quality representation of the videó signal, which is directly, as it is the case in the example of Fig. 2, or indirectly via an artifact reduction filter (nőt shown), optionally outputable at output 216. Based on the base quality representation videó signal, the prediction means 206 forms a prediction signal having the increased picture sample bit depth and/or the increased chroma sampling resolution. The decoding means 208 decodes the enhancement layer data stream to obtain the prediction residual having the increased bit-depth and, optionally, increased spatial resolution. The reconstruction means 210 obtains the high quality videó signal from the prediction and the prediction residual and outputs same at output 214 via an optional artifact reducing filter.
[0066] Internally, the demultiplexing means 202 comprises a demultiplexer 218 and an entropy decoding modulé 220. An input of the demultiplexer 218 is connected to input 212 and a first output of the demultiplexer 218 is connected to the residual decoding means 208. The entropy-decoding modulé 220 is connected between another output of the demultiplexer 218 and the base decoding means 204. The demultiplexer 218 divides the quality-scalable data stream intő the base layer data stream and the enhancement layer data stream as having been separately input intő the multiplexer 144, as described above. The entropy decoding modulé 220 performs, fór example, a Huffman decoding or arithmetic decoding algorithm in order to obtain the transform coefficient levels, motion vectors, transform size Information and other syntax elements necessary in order to dérivé the base representation of the videó signal therefrom. At the output of the entropy-decoding modulé 220, the base encoding data stream results.
[0067] The base decoding means 204 comprises an inverse transform modulé 222, an adder 224, an optional loop filter 226 and a predictor modulé 228. The modules 222 to 228 of the base decoding means 204 correspond, with respect to functionality and inter-connection, to the elements 124 to 130 of Fig. 1. To be more precise, the inverse transform modulé 222 and the adder 224 are connected in series in the order mentioned between the demultiplexing means 202 on the one hand and the prediction means 206 and the base quality output, respectively, on the other hand, and the optional loop filter 226 and the predictor modulé 228 are connected in series in the order mentioned between the output of the adder 224 and another input of the adder 224. By this measure, the adder 224 outputs the base representation videó signal with the reduced bit-depth and, optionally, the reduced spatial resolution which is receivable from the outside at output 216.
[0068] The prediction means 206 comprises an optional artifact reduction filter 230 and a prediction Information modulé
ΕΡ 2 279 622 Β1
232, both modules functioning in a synchronous manner relatíve to the elements 132 and 134 of Fig. 1. In other words, the optional artifact reduction filter 230 optionally filters the base quality videó signal in order to reduce artifacts therein and the prediction formation modulé 232 retrieves predicted pictures with increased bit depths and, optionally, increased spatial resolution in a manner already described above with respect to the prediction modulé 134. That is, the prediction information modulé 232 may, by means of side information contained in the quality-scalable data stream, map the incoming picture samples to a higher dynamic rangé and, optionally, apply a spatial interpolation filter to the content of the pictures in order to increase the spatial resolution.
[0069] The residual decoding means 208 comprises an entropy decoding modulé 234 and an inverse transform modulé 236, which are serially connected between the demultiplexer 218 and the reconstruction means 210 in the order just mentioned. The entropy decoding modulé 234 and the inverse transform modulé 236 cooperate to reverse the encoding performed by modules 136,138, and 140 of Fig. 1. In particular, the entropy-decoding modulé 234 performs, forexample, a Huffman decoding or arithmetic decoding algorithms to obtain syntax elements comprising, among others, transform coefficient levels, which are, by the inverse transform modulé 236, inversely transformed to obtain a prediction residual signal or a sequenceof residual pictures. Further, the entropy decoding modulé 234 reveals the side information generated in steps 190 to 194 at the encoder side so that the prediction formation modulé 232 is able to emulate the inverse mapping procedure performed at the encoder side by the prediction modulé 134, as already denoted above.
[0070] Similarto Fig. 6, which referred to the encoder, Fig. 7 shows the mode of operation of the prediction formation modulé 232 and, partially, the entropy decoding modulé 234 in more detail. As shown therein, the process of gaining the prediction from the reconstructed base layer signal at the decoder side starts with a co-operation of demultiplexer 218 and entropy decoder 234 to dérivé the granularity information coded in step 194 in step 280, dérivé the information on the global mapping function(s) having been coded in step 192 in step 282, dérivé the index value k fór each block of the fine granularity as having been coded in step 190 in step 284, and dérivé the local mapping/adaptation function paraméter m and n fór each finegranularity block as having been coded in step 188 in step 286. As illustrated by way of the dotted lines, steps 280 to 284 are optional with its appliance depending on the embodiment currently used. [0071] In step 288, the prediction formation modulé 232 performs the inverse tone mapping based on the information gained in steps 280 to 286, thereby exactly emulating the inverse tone mapping having been performed at the encoder side at step 186. Similar to the description with respect to step 188, the derivation of parameters m, n may comprise a predictive decoding where a prediction residual value is derived from the high quality part of the data-stream entering demultiplexer 218 by use of, fór example, entropy decoding as performed by entropy decoder 234, and obtaining the actual values of m and n by adding these prediction residual values to a prediction value derived by local and/ortemporai prediction.
[0072] The reconstruction means 210 comprises an adder 238 the inputs of which are connected to the output of the prediction information modulé 232, and the output of the inverse transform modulé 236, respectively. The adder 238 adds the prediction residual and the prediction signal in order to obtain the high quality videó signal having the increased bit depth and, optionally, increased spatial resolution which is fed via an optional artifact reducing filter 240 to output 214. [0073] Thus, as is derivable from Fig. 2, a base quality decoder may reconstruct a base quality videó signal from the quality-scalable data stream at the input 212 and may, in order to do so, nőt include elements 230, 232, 238, 234, 236, and 240. On the other hand, a high quality decoder may nőt include the output 216.
[0074] In other words, in the decoding process, the decoding of the base quality representation is straightforward. Fór the decoding of the high quality signal, first the base quality signal has to be decoded, which is performed by modules 218 to 228. Thereafter, the prediction process described above with respect to modulé 232 and optional modulé 230 is employed using the decoded base representation. The quantized transform coefficients of the high quality enhancement signal arescaled and inversely transformed bythe inverse transform modulé 236, fór example, asspecified in H.264/AVC in order to obtain the residual or difference signal samples, which are added to the prediction derived from the decoded base representation samples by the prediction modulé 232. As a final step in the decoding process of the high quality videó signal to be output at output 214, optional a filter can be employed in order to remove or reduce visually disturbing coding artifacts. It is to be noted that the motion-compensated prediction loop involving modules 226 and 228 is fully self-contained using only the base quality representation. Therefore, the decoding complexity is moderate and there is no need fór an interpolation filter, which operates on the high bit depth and, optionally, high spatial resolution image data in the motion-compensated prediction process of predictor modulé 228.
[0075] Regarding the above embodiments, it should be mentioned that the artifact reduction filters 132 and 230 are optional and could be removed. The same applies fór the loop filters 124 and 226, respectively, and filter 240. Fór example, with respect to Fig. 1, it had been noted that the filters 124 and 132 may be replaced byjust one common filter such as, fór example, a de-blocking filter, the output of which is connected to both the input of the motion prediction modulé 126 as well as the input of the inverse tone mapping modulé 134. Similarly, filters 226 and 230 may be replaced by one common filter such as, a de-blocking filter, the output of which is connected to output 216, the input of prediction formation modulé 232 and the input of predictor 228. Further, fór the sake of completeness, it is noted that the prediction modules 228 and 126, respectively, nőt necessarily temporally predict samples within macroblocks of a current picture.
ΕΡ 2 279 622 Β1
Rather, a spatial prediction or intra-prediction using samples of the same picture may alsó be used. In particular, the prediction type may be chosen on a macroblock-by-macroblock basis, fór example, or somé other granularity. Further, the present invention is nőt restricted to videó coding. Rather, the above description is alsó applicable to still image coding. Accordingly, the motion-compensated prediction loop involving elements 118, 128, 130, 126, and 124 and the elements 224, 228, and 226, respectively, may be removed alsó. Similarly, the entropy coding mentioned needs nőt necessarily to be performed.
[0076] Evén more precise, in the above embodiments, the base layer encoding 118-130, 142 was based on motioncompensated prediction based on a reconstruction of already lossy coded pictures. In this case, the reconstruction of the base encoding process may alsó be viewed as a part of the high-quality prediction forming process as has been done in the above description. However, in case of a lossless encoding of the base representation, a reconstruction would nőt be necessary and the down-converted signal could be directly forwarded to means 132, 134, respectively. In the case of no motion-compensation based prediction in a lossy base layer encoding, the reconstruction fór reconstructing the base quality signal at the encoder side would be especially dedicated fór the high-quality prediction formation in 104. In other words, the above association of the elements 116-134 and 142 to means 102,104 and 108, respectively, could be performed in another way. In particular, the entropy coding modulé 142 could be viewed as a part of base encoding means 102, with the prediction means merely comprising modules 132 and 134 and the combining means 108 merely comprising the multiplexer 144. This view correlates with the module/means association used in Fig. 2 in that the prediction means 206 does nőt comprise the motion compensation based prediction. Additionally, however, demultiplexing means 202 could be viewed as nőt including entropy modulé 220 so that base decoding means alsó comprises entropy decoding modulé 220. However, both views lead to the same result in that the prediction in 104 is performed based on a representation of the source matéria! with the reduced bit-depth and, optionally, the reduced spatial resolution which is losslessly coded intő and losslessly derivable from the quality-scalable bit stream and base layer data stream, respectively. According to the view underlying Fig. 1, the prediction 134 is based on a reconstruction ofthe base encoding data stream, whereas in case ofthe alternative view, the reconstruction would start from an intermediate encoded version or halfway encoded version ofthe base quality signal which misses the lossless encoding according to modulé 142 fór being completely coded intő the base layer data stream. In this regard, itshould be further noted that the down-conversion in modulé 116 does nőt have to be performed by the encoder 100. Rather, encoder 100 may have two inputs, one fór receiving the high-quality signal and the other fór receiving the down-converted version, from the outside.
[0077] In the above-described embodiments, the quality-scalability did merely relate to the bit depth and, optionally, the spatial resolution. However, the above embodiments may easily be extended to include temporal scalability, chroma formát scalability, and fine granular quality scalability.
[0078] Accordingly, the above embodiments of the present invention form a concept fór scalable coding of picture or videó content with different granularities in terms of sample bit-depth and, optionally, spatial resolution by use of locally adaptive inverse tone mapping. In accordance with embodiments ofthe present invention, both the temporal and spatial prediction processes as specified in the H.264/AVC scalable videó coding extension are extended in a way that they include mappings from lower to higher sample bit-depth fidelity as well as, optionally, from lower to higher spatial resolution. The above-described extension of SVC towards scalability in terms of sample bit-depth and, optionally, spatial resolution enables the encoder to store a base quality representation of a videó sequence, which can be decoded by any legacy videó decoder together with an enhancement signal fór higher bit-depth and, optionally, higher spatial resolution, which is ignored by legacy videó decoders. Fór example, the base quality representation could contain an 8-bit version of the videó sequence in CIF resolution, namely 352x288 samples, while the high quality enhancement signal contains a refinement to a 10-bit version in 4CIF resolution, i.e. 704x476 samples ofthe same sequence. In a different configuration, it is alsó possible to use the same spatial resolution fór both base and enhancement quality representations, such that the high quality enhancement signal only contains a refinement ofthe sample bit-depth, e.g. from 8 to 10 bits. [0079] In other words, the above-outlined embodiments enable to form a videó coder, i.e. encoder or decoder, fór coding, i.e. encoding or decoding, a layered representation of a videó signal comprising a standardized videó coding method fór coding a base-quality layer, a prediction method fór performing a prediction ofthe high-quality enhancement layer signal by using the reconstructed base-quality signal and a residual coding method fór coding of the prediction residual ofthe high-quality enhancement layer signal. In this regard, the prediction may be performed by using a mapping function from the dynamic rangé associated with the base-quality layer to the dynamic rangé associated with the highquality enhancement layer. Further, the mapping function may be built as the sum of a global mapping function, which follows the inverse tone mapping schemes described above and a local adaptation. The local adaptation, in turn, may be performed by scaling the sample values x of the base quality layer and adding an offset value according to m x n. In any case, the residual coding may be performed according to H.264/AVC.
[0080] Depending on an actual implementation, the inventive coding scheme can be implemented in hardware or in software. Therefore, the present invention alsó relates to a computer program, which can be stored on a computerreadable médium such as a CD, a disc or any other data carrier. The present invention is, therefore, alsó a computer program having a program code which, when executed on a computer, performs the inventive method described in
ΕΡ 2 279 622 Β1 connection with the above figures. In particular, the implementations ofthe means and modules in Fig. 1 and 2 may comprise sub-routines running on a CPU, Circuit parts of an ASIC or the like, fór example.
[0081] Thus, inter alias, above embodiments describe an encoder fór encoding a picture or videó source data (160) intő a quality-scalable data stream (112), comprising base encoding means (102) fór encoding the picture or videó source data (160) intő a base encoding data stream representing a representation ofthe picture or videó source data with a first picture sample bit depth; mapping means (104) fór mapping samples ofthe representation ofthe picture or videó source data (160) with the first picture sample bit depth from a first dynamic rangé corresponding to the first picture sample bit depth to a second dynamic rangé greater than the first dynamic rangé and corresponding to a second picture sample bit depth being higher than the first picture sample bit depth, by use of one or more global mapping functions being constant within the picture or videó source data (160) or varying at a first granularity, and a local mapping function locally modifying the one or more global mapping functions at a second granularity finer than the first granularity to obtain a prediction ofthe picture or videó source data having the second picture sample bit depth; residual encoding means (106) fór encoding a prediction residual ofthe prediction intő a bit-depth enhancement layerdata stream; and combining means (108) forforming the quality-scalable data stream based on the base encoding data stream, the local mapping function and the bit-depth enhancement layerdata stream so that the local mapping function is derivablefrom the qualityscalable data stream.
[0082] The mapping means may comprise means (124, 126, 128, 130, 132) fór reconstructing a low bit-depth reconstruction picture or videó as the representation ofthe picture or videó source data with the first picture sample bit depth being based on the base encoding data stream, the low bit-depth reconstruction picture or videó having the first picture sample bit depth.
[0083] The mapping means (104) may be adapted to map the samples ofthe representation ofthe picture or videó source data (160) with the first picture sample bit-depth by use of a combined mapping function being an arithmetic combination of one ofthe one or more global mapping functions and the local mapping function. More than one global mapping functions may be used by the mapping means (104) and the combining means (108) is adapted to form the quality-scalable data-stream (112) such that the one ofthe more than one global mapping functions is derivable from the quality-scalable data-stream. The arithmetic combination may comprise an add operation.
[0084] The combining means (108) and the mapping means (104) may be adapted such that second granularity subdivides the picture or videó source data (160) intő a plurality of picture blocks (170) and the mapping means may be adapted such that the local mapping function is m s + n with m and n varying at the second granularity with the combining means being adapted such that m and n are defined within the quality scalable data-stream fór each picture block (170) of the picture or videó source data (160) so that m and n may differ among the plurality of picture blocks. [0085] The mapping means (134) may be adapted such that the second granularity varies within the picture or videó source data (160), and the combining means (108) may be adapted such that the second granularity is alsó derivable from the quality-scalable data-stream.
[0086] The combining means (108) and the mapping means (134) may be adapted such that the second granularity divides-up the picture or videó source data (160) intő a plurality of picture blocks (170), and the combining means (108) may be adapted such that a local mapping function residual (Am, Δη) is incorporated intő the quality-scalable datastream (112) fór each picture block (170), and the local mapping function of a predetermined picture block ofthe picture or videó source data (160) is derivable from the quality-scalable data-stream by use of a spatial and/or temporal prediction from one or more neighboring picture blocks or a corresponding picture block of a picture of the picture or videó source data preceding a picture to which the predetermined picture block belongs, and the local mapping function residual of the predetermined picture block.
[0087] The combining means (108) may be adapted such that at least one ofthe one or more global mapping functions is alsó derivable from the quality-scalable data-stream.
[0088] The base encoding means (102) may comprise means (116) fór mapping samples representing the picture with the second picture sample bit depth from the second dynamic rangé to the first dynamic rangé corresponding to the first picture sample bit depth to obtain a quality-reduced picture; and means (118, 120, 122, 124, 126, 128, 130) fór encoding the quality-reduced picture to obtain the base encoding data stream.
29 members in 10 offices
Priority claims1
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008003047 | European Patent Office (EPO) | W |
Members29
| Document | Office | Kind | |
|---|---|---|---|
| WO2009127231A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2279622A1 | European Patent Office (EPO) | A1 | |
| CN102007768A | China | A | |
| US2011090959A1 | United States of America | A1 | |
| JP2011517245A | Japan | A | |
| CN102007768B | China | B | |
| JP5203503B2 | Japan | B2 | |
| EP2279622B1 | European Patent Office (EPO) | B1 | |
| PT2279622E | Portugal | E | |
| DK2279622T3 | Denmark | T3 | |
| ES2527932T3 | Spain | T3 | |
| EP2835976A2 | European Patent Office (EPO) | A2 | |
| PL2279622T3 | Poland | T3 | |
| US8995525B2 | United States of America | B2 | |
| EP2835976A3 | European Patent Office (EPO) | A3 | |
| US2015172710A1 | United States of America | A1 | |
| HUE024173T2This record | Hungary | T2 | |
| EP2835976B1 | European Patent Office (EPO) | B1 | |
| PT2835976T | Portugal | T | |
| DK2835976T3 | Denmark | T3 | |
| ES2602100T3 | Spain | T3 | |
| PL2835976T3 | Poland | T3 | |
| HUE031487T2 | Hungary | T2 | |
| US2019289323A1 | United States of America | A1 | |
| US10958936B2 | United States of America | B2 | |
| US2021211720A1 | United States of America | A1 | |
| US11711542B2 | United States of America | B2 | |
| US2023421806A1 | United States of America | A1 | |
| US12457361B2 | United States of America | B2 |
Numbers
- Publication
- E024173
- Application
- 8735287
Titles2
- English
- BIT-DEPTH SCALABILITY
- Hungarian
- Bitmélység skálázhatóság
Classification
- CPC, 7
- H04N19/117
- H04N19/593
- H04N19/184
- H04N19/80
- H04N19/33
- H04N19/36
- H04N19/82
- IPC, 1
- H04N19 00