Inter-layer prediction using sample-adaptive adjustments for bit depth scalable video coding
30 claims: 6 independent, 24 dependent
- 1ビデオデータをコーディングするように構成された装置であって、前記装置は、 第1のレイヤと第2のレイヤとに関連付けられたビデオデータを記憶するように構成されたメモリと、前記第1のレイヤが、第1のビット深度を有する第1のレイヤサンプルを含む、 前記メモリと通信しているプロセッサと、前記プロセッサは、 予備予測サンプルを生成するために前記第1のレイヤサンプルに予備マッピング関数を適用することと、 前記第1のレイヤサンプルに関連付けられた1つまたは複数の値に基づいて前記第1のレイヤサンプルの第1のカテゴリーを決定することと、ここにおいて、前記第1のカテゴリーは、前記第1のレイヤサンプルに関連付けられた前記1つまたは複数の値と、少なくとも複数の帯域内の隣接する帯域間の1つまたは複数の境界点によって定義された前記複数の帯域との間の比較に基づく、 前記第1のレイヤサンプルの前記決定された第1のカテゴリーに基づいて1つまたは複数の調整パラメータを決定することと、 改良予測サンプルを決定するために、前記1つまたは複数の調整パラメータを使用して区分的調整演算を前記予備予測サンプルに対して行うことと、前記区分的調整演算は、 前記1つまたは複数の調整パラメータによる、前記予備予測サンプルの 乗算、除算、累乗、または対数のうちの1つまたは複数を備え、前記改良予測サンプルが、前記第1のビット深度よりも大きい第2のビット深度を有する、 を行うように構成された、 を備える、装置。
- 2前記第1のビット深度が8ビットであり、前記第2のビット深度が10ビット、12ビット、および14ビットのうちの1つである、請求項1に記載の装置。
- 3前記第1のレイヤサンプルの前記第1のカテゴリーが、前記第1のレイヤサンプルの1つまたは複数のルミナンス値またはクロミナンス値から決定される、請求項1に記載の装置。
- 4前記プロセッサは、(i)前記第1のレイヤサンプルに関連付けられた前記1つまたは複数の値、および(ii)前記第1のレイヤ内の隣接サンプルに関連付けられた1つまたは複数の値に基づいて、前記第1のレイヤサンプルの第2のカテゴリーを決定することと、前記第1のレイヤサンプルの前記決定された第2のカテゴリーに基づいて1つまたは複数の追加の調整パラメータを決定することとを行うようにさらに構成され、前記区分的調整演算は、前記1つまたは複数の追加の調整パラメータにさらに基づく、請求項1に記載の装置。
- 5前記第1のレイヤサンプルが、可能なルミナンス値のスケール上のルミナンス値を表し、 可能なルミナンス値の前記スケールが複数のルミナンス帯域に分割され、 前記第1のレイヤサンプルによって表される前記ルミナンス値が前記ルミナンス帯域のうちの1つ内にあり、 前記第1のレイヤサンプルの前記第1のカテゴリーは、前記第1のレイヤサンプルがその内にある前記ルミナンス帯域に対応する、請求項1に記載の装置。
- 6前記第1のレイヤサンプルが、可能なクロミナンス値のスケール上のクロミナンス値を表し、 可能なクロミナンス値の前記スケールが複数のクロミナンス帯域に分割され、 前記第1のレイヤサンプルによって表される前記クロミナンス値が前記クロミナンス帯域のうちの1つ内にあり、 前記第1のレイヤサンプルの前記第1のカテゴリーは、前記第1のレイヤサンプルがその内にある前記クロミナンス帯域に対応する、請求項1に記載の装置。
- 7前記第1のレイヤサンプルの前記第2のカテゴリーが、前記第1のレイヤサンプルと、前記ビデオデータ中の前記第1のレイヤサンプルに空間的に隣接する他のサンプルとの間の複数の比較の結果に依存する、請求項4に記載の装置。
- 8前記予備マッピング関数が少なくとも1つの対数演算または指数演算を備える、請求項1に記載の装置。
- 9前記予備マッピング関数が、左ビットシフト、または2よりも大きいかまたはそれに等しい数による乗算を備える、請求項1に記載の装置。
- 10前記予備マッピング関数が、各可能な第1のレイヤサンプル値を対応する第2のレイヤサンプル値にマッピングするルックアップテーブルを備える、請求項1に記載の装置。
- 11前記予備予測サンプルが、前記第2のビット深度に等しいビット深度を有する、請求項1に記載の装置。
- 12前記1つまたは複数の調整パラメータが、比、係数、指数、または対数の底を備える、請求項1に記載の装置。
- 13前記区分的調整演算が、加算または減算をさらに備える、請求項1に記載の装置。
- 14前記プロセッサが、第2のレイヤサンプルを決定するために、前記改良予測サンプルに残差値を加算するようにさらに構成された、請求項1に記載の装置。
- 15ビデオデータをコーディングする方法であって、前記方法は、 第1のビット深度を有する第1のレイヤサンプルを備える前記ビデオデータを受信することと、 予備予測サンプルを生成するために前記第1のレイヤサンプルに予備マッピング関数を適用することと、 前記第1のレイヤサンプルに関連付けられた1つまたは複数の値に基づいて前記第1のレイヤサンプルの第1のカテゴリーを決定することと、ここにおいて、前記第1のカテゴリーは、前記第1のレイヤサンプルに関連付けられた前記1つまたは複数の値と、少なくとも複数の帯域内の隣接する帯域間の1つまたは複数の境界点によって定義された前記複数の帯域との間の比較に基づく、 前記第1のレイヤサンプルの前記決定された第1のカテゴリーに基づいて1つまたは複数の調整パラメータを決定することと、 改良予測サンプルを決定するために、前記1つまたは複数の調整パラメータを使用して区分的調整演算を前記予備予測サンプルに対して行うことと、前記区分的調整演算は、 前記1つまたは複数の調整パラメータによる、前記予備予測サンプルの 乗算、除算、累乗、または対数のうちの1つまたは複数を備え、前記改良予測サンプルが、前記第1のビット深度よりも大きい第2のビット深度を有する、 を備える、方法。
- 16前記第1のビット深度が8ビットであり、前記第2のビット深度が10ビット、12ビット、および14ビットのうちの1つである、請求項15に記載の方法。
- 17前記第1のレイヤサンプルの前記第1のカテゴリーが、前記第1のレイヤサンプルの1つまたは複数のルミナンス値またはクロミナンス値から決定される、請求項15に記載の方法。
- 18(i)前記第1のレイヤサンプルに関連付けられた前記1つまたは複数の値、および(ii)第1のレイヤ内の隣接サンプルに関連付けられた1つまたは複数の値に基づいて、前記第1のレイヤサンプルの第2のカテゴリーを決定することと、前記第1のレイヤサンプルの前記決定された第2のカテゴリーに基づいて1つまたは複数の追加の調整パラメータを決定することとをさらに備え、前記区分的調整演算は、前記1つまたは複数の追加の調整パラメータにさらに基づく、請求項15に記載の方法。
- 19前記第1のレイヤサンプルが、可能なルミナンス値のスケール1ルミナンス値を表し、 可能なルミナンス値の前記スケールが複数のルミナンス帯域に分割され、 前記第1のレイヤサンプルによって表される前記ルミナンス値が前記ルミナンス帯域のうちの1つ内にあり、 前記第1のレイヤサンプルの前記第1のカテゴリーは、前記第1のレイヤサンプルがその内にある前記ルミナンス帯域に対応する、請求項15に記載の方法。
- 20前記第1のレイヤサンプルが、可能なクロミナンス値のスケール上のクロミナンス値を表し、 可能なクロミナンス値の前記スケールが複数のクロミナンス帯域に分割され、 前記第1のレイヤサンプルによって表される前記クロミナンス値が前記クロミナンス帯域のうちの1つ内にあり、 前記第1のレイヤサンプルの前記第1のカテゴリーは、前記第1のレイヤサンプルがその内にある前記クロミナンス帯域に対応する、請求項15に記載の方法。
- 21前記第1のレイヤサンプルの前記第2のカテゴリーが、前記第1のレイヤサンプルと、前記ビデオデータ中の前記第1のレイヤサンプルに空間的に隣接する他のサンプルとの間の複数の比較の結果に依存する、請求項18に記載の方法。
- 22前記予備マッピング関数が少なくとも1つの対数演算または指数演算を備える、請求項15に記載の方法。
- 23前記予備マッピング関数が、左ビットシフト、または2よりも大きいかまたはそれに等しい数による乗算を備える、請求項15に記載の方法。
- 24前記予備マッピング関数が、各可能な第1のレイヤサンプル値を対応する第2のレイヤサンプル値にマッピングするルックアップテーブルを備える、請求項15に記載の方法。
- 25前記予備予測サンプルが、前記第2のビット深度に等しいビット深度を有する、請求項15に記載の方法。
- 26前1つまたは複数の記調整パラメータが、比、係数、指数、または対数の底を備える、請求項15に記載の方法。
- 27前記区分的調整演算が、加算または減算をさらに備える、請求項15に記載の方法。
- 28プロセッサが、第2のレイヤサンプルを決定するために、前記改良予測サンプルに残差値を加算するようにさらに構成された、請求項15に記載の方法。
- 29実行されたとき、 第1のビット深度を有する第1のレイヤサンプルを備えるビデオデータを受信することと、 予備予測サンプルを生成するために前記第1のレイヤサンプルに予備マッピング関数を適用することと、 前記第1のレイヤサンプルに関連付けられた1つまたは複数の値に基づいて前記第1のレイヤサンプルの第1のカテゴリーを決定することと、ここにおいて、前記第1のカテゴリーは、前記第1のレイヤサンプルに関連付けられた前記1つまたは複数の値と、少なくとも複数の帯域内の隣接する帯域間の1つまたは複数の境界点によって定義された前記複数の帯域との間の比較に基づく、 前記第1のレイヤサンプルの前記決定された第1のカテゴリーに基づいて1つまたは複数の調整パラメータを決定することと、 改良予測サンプルを決定するために、前記1つまたは複数の調整パラメータを使用して区分的調整演算を前記予備予測サンプルに対して行うことと、前記区分的調整演算は、 前記1つまたは複数の調整パラメータによる、前記予備予測サンプルの 乗算、除算、累乗、または対数のうちの1つまたは複数を備え、前記改良予測サンプルが、前記第1のビット深度よりも大きい第2のビット深度を有する、 を装置に行わせるコードを備える非一時的コンピュータ可読媒体。
- 30ビデオデータをコーディングするように構成されたビデオコーディングデバイスであって、前記ビデオコーディングデバイスは、 第1のビット深度を有する第1のレイヤサンプルを備える前記ビデオデータを受信するための手段と、 予備予測サンプルを生成するために前記第1のレイヤサンプルに予備マッピング関数を適用するための手段と、 前記第1のレイヤサンプルに関連付けられた1つまたは複数の値に基づいて前記第1のレイヤサンプルの第1のカテゴリーを決定するための手段と、ここにおいて、前記第1のカテゴリーは、前記第1のレイヤサンプルに関連付けられた前記1つまたは複数の値と、少なくとも複数の帯域内の隣接する帯域間の1つまたは複数の境界点によって定義された前記複数の帯域との間の比較に基づく、 前記第1のレイヤサンプルの前記決定された第1のカテゴリーに基づいて1つまたは複数の調整パラメータを決定するための手段と、 改良予測サンプルを決定するために、前記1つまたは複数の調整パラメータを使用して区分的調整演算を前記予備予測サンプルに対して行うための手段と、前記区分的調整演算は、 前記1つまたは複数の調整パラメータによる、前記予備予測サンプルの 乗算、除算、累乗、または対数のうちの1つまたは複数を備え、前記改良予測サンプルが、前記第1のビット深度よりも大きい第2のビット深度を有する、 を備える、ビデオコーディングデバイス。
Independent claims30
161 paragraphs, as filed
[0001] The present disclosure relates generally to the field of video coding and compression, and more specifically to techniques for inter-layer prediction in scalable video coding (SVC).
[0002] Digital video capabilities include digital television, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video. It can be incorporated into a wide range of devices, including game consoles, cellular or satellite radio phones, video teleconferencing devices, and more. Digital video devices include MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part10, Advanced Video Coding (AVC), high efficiency currently under development. Video Coding (HEVC: High Efficiency Video) Coding) Implement video compression techniques, such as the standards defined by the standard and the video compression techniques described in extensions to such standards. Video devices may more efficiently transmit, receive, encode, decode, and / or store digital video information by implementing such video coding techniques.
[0003] Video compression techniques make spatial (intra-picture) and / or temporal (inter-picture) predictions to reduce or eliminate the redundancy inherent in video sequences. For block-based video coding, video slices (eg, video frames, parts of video frames, etc.) are divided into tree blocks, coding units (CUs) and / or video blocks, sometimes called coding nodes. obtain. Intra-coding a picture (I) The video blocks in a slice are encoded using spatial prediction for reference samples in adjacent blocks in the same picture. A video block in an intercoded (P or B) slice of a picture makes a spatial prediction for a reference sample in an adjacent block in the same picture, or a temporal prediction for a reference sample in another reference picture. Can be used. Pictures are sometimes referred to as frames, and reference pictures are sometimes referred to as reference frames.
[0004] Spatial or temporal prediction yields predictive blocks of blocks to be coded. The residual data represents the pixel difference between the original block to be coded and the predicted block. The intercoding block is encoded according to a motion vector pointing to a block of reference samples forming the prediction block and residual data showing the difference between the coding block and the prediction block. The intracoding block is encoded according to the intracoding mode and the residual data. For further compression, the residual data can be converted from the pixel area to the conversion area to obtain a residual conversion factor, which in turn can be quantized. The quantization transform coefficients are initially composed of a two-dimensional array, which can be scanned to generate a one-dimensional vector of the transform coefficients, and entropy coding can be applied to achieve further compression.
[0005] Certain block-based video coding and compression may use scalable techniques. Scalable video coding (SVC) uses a base layer (BL), sometimes referred to as a reference layer (RL), and one or more scalable enhancement layers (EL). Refers to video coding. In the case of SVC, the base layer can carry video data with base level quality. One or more enhancement layers can carry additional video data to support higher spatial, temporal, and / or signal-to-noise (SNR) levels. Enhancement layers can be defined for previously encoded layers. For example, the lowest layer can act as BL and the highest layer can act as EL. The intermediate layer can act as either EL or RL, or both. For example, the layer in the middle is the base layer or the intervening enhancement layer. It is an EL for the layer below it, such as layer), and at the same time can act as an RL for one or more enhancement layers above it. Similarly, in HEVC standard multi-view or 3D extensions, there can be multiple views, where the information in one view is the information in another view (eg motion estimation, motion vector prediction) and /. Or it can be used to code (eg, encode or decode) other redundancy. In some cases, the base layer may be transmitted in a more reliable way than the enhancement layer is transmitted. Techniques for SVC are also inter-layer to reduce or eliminate redundancy between the base layer and the enhancement layer. prediction) can be used. Inter-layer prediction produces a prediction enhancement layer block from the corresponding base layer block. The enhancement layer block can be coded using the prediction block generated from the base layer, along with residual data showing the difference between the prediction block and the block to be coded. This residual data can be transformed, quantized, and entropy-encoded, similar to the residual data associated with spatial and temporal predictions.
[0006] High dynamic range (HDR) range) Sequences are used in professional generation environments and high quality displays are available that can play 10-bit or higher content. One way to represent and distribute such HDR content is to generate a bitstream using a single layer encoder. For example, 10-bit content can be encoded with a single layer encoder such as HEVC or H.264 / AVC (eg, in a high 10 profile). In such cases, only 10-bit displays can play the decrypted content, and legacy 8-bit displays require down-conversion of 10-bit content to 8-bit, which is still 10-bit. Requires a compatible decoder. Legacy 8-bit decoders are not capable of decoding 10-bit bitstreams. In this example, if both the 8-bit and 10-bit displays require access to the same HDR video content, the HDR video content will have separate bitstreams for the two displays (eg, 8-bit bitstream and It can be simulcast in a 10-bit bitstream. However, such techniques have high bandwidth requirements because there can be a lot of redundant information in the two bitstreams.
[0007] Alternatively, a scalable bitstream can be generated by the scalable encoder. A scalable decoder may be able to decode 10-bit video content from a scalable bit stream, and an 8-bit decoder is used to advance the information contained in the enhancement layer (eg, 8 bits to 10 bits). The 8-bit base layer can be decoded, ignoring the information that is given. Alternatively, a bitstream extractor located on the server side or in the network can extract, for example, an 8-bit base layer from a scalable bitstream.
[0008] Therefore, a base layer that can be decoded by a legacy decoder (eg, 8 bits) to produce video content with a lower bit depth (eg, 8 bits) and a higher bid depth video content (eg, 10 bits). Back compatibility with legacy decoders is provided by using the SVC to generate a scalable bitstream that contains one or more enhancement layers that can be decoded by the scalable decoder to generate (bits). Obtaining, bandwidth requirements can be reduced compared to simulcasting separate bitstreams, which improves coding efficiency and performance. Thus, the techniques described in the present disclosure may reduce the computational complexity associated with the method of coding video information, improve coding efficiency, and / or improve overall coding performance.
[0009] Each of the systems, methods and devices of the present disclosure has several invention aspects, of which a single aspect may not be solely responsible for the desired attributes disclosed herein. ..
[0010] One aspect of the present disclosure provides an apparatus configured to code video data. The present device includes a memory unit configured to store video data. The video data may have a base layer and an enhancement layer. The base layer comprises a video sample (also known as a pixel) with a certain bit depth. The enhancement layer comprises a sample having a higher bit depth than the video sample in the base layer. Both samples in the base layer and samples in the enhancement layer can be grouped into video blocks, and the dimensions of the blocks can vary within each layer and between different layers, but video blocks in the base layer are generally Corresponds to one or more video blocks in the enhancement layer.
[0011] The device further comprises a processor communicating with a memory unit, which determines predicted video samples for the enhancement layer based on the video samples associated with the base layer. It is configured as follows. The processor first applies a preliminary mapping function to the video sample from the base layer to determine the preliminary prediction, and then to determine the refined prediction. Predictive video samples for the enhancement layer can be determined by applying adaptive adjustments to the preliminary predictions. Processors may apply different adaptive adjustments for different categories of base layer samples.
[0012] In some embodiments, the preliminary mapping function may comprise a non-linear mathematical function that maps the base layer sample to the predictive enhancement layer sample, for example by calculating the logarithm or power of the base layer sample. In other embodiments, the preliminary mapping function may not be used at all, or the base layer sample may simply be used as a preliminary prediction for the enhancement layer sample. In some embodiments, the adaptive adjustment may comprise a ratio or coefficient by which the preliminary prediction is multiplied to determine the improved prediction. As an addition or alternative, the adjustment may include an offset added to the preliminary forecast to determine the improved forecast. In some embodiments, adaptive adjustment may depend on categories such as the intensity range of individual samples or the pattern of adjacent samples.
[0013] In some embodiments, the base layer sample can have a bit depth of 8 bits and the enhancement layer sample can have a bit depth of 10 bits. Base layer samples can be assigned one or more categories based on one or more luminance or chrominance values of the base layer sample and / or other samples in the video data.
[0014] Another aspect of the present disclosure provides a method for coding video data. The method comprises determining predicted samples for the enhancement layer based on the samples associated with the base layer of the video data. The prediction video sample for the enhancement layer first applies a preliminary mapping function to the video sample from the base layer to determine the preliminary prediction, and then adapts to the preliminary prediction to determine the improvement prediction. Can be determined by applying. Different adaptive adjustments can be applied for different categories of base layer samples.
[0015] Another aspect of the disclosure is a non-transient computer comprising code that, when executed, causes the device to determine a predictive sample for the enhancement layer based on the sample associated with the base layer of video data. Provide a readable medium. The device first applies a preliminary mapping function to the video sample from the base layer to determine the preliminary prediction, and then applies adaptive adjustments to the preliminary prediction to determine the improved prediction. It can be programmed to determine a predictive video sample for the enhancement layer. The device may be configured to apply different adaptive adjustments for different categories of base layer samples.
[0016] Another aspect of the present disclosure provides a video coding device for coding video data. The device includes means for determining a predictive video sample for the enhancement layer based on the video sample associated with the base layer of video data. The device includes means for applying a preliminary mapping function to a video sample from the base layer to determine a preliminary prediction and means for applying adaptive adjustments to the preliminary prediction to determine an improved prediction. obtain. The device may apply different adaptive adjustments for different categories of base layer samples.
The disclosed devices, methods, computer-readable media, and devices also minimize one or more measures of error or strain associated with the predicted sample produced by applying adaptive adjustments to the base layer sample. To limit, calculations may include components, steps, modules, or functions for determining adaptive adjustments for various categories of base layer samples. The error measure may include, for example, an average error, an average squared error, or a computationally efficient approximation of either.
[0018] Details of one or more examples are given in the accompanying drawings and in the description below. Other features, objectives, and advantages will become apparent from its description and drawings, as well as the claims.
<figref num="1">It is a block diagram which shows an example of the video coding and decoding system which can utilize the technique by the aspect described in this disclosure.</figref><figref num="2">It is a block diagram which shows an example of the video encoder which can implement the technique by the aspect described in this disclosure.</figref><figref num="3">It is a block diagram which shows an example of the video decoder which can implement the technique by the aspect described in this disclosure.</figref><figref num="4">FIG. 3 is a block diagram illustrating an exemplary scalable video encoder that may utilize the techniques according to the embodiments described in the present disclosure.</figref><figref num="5">It is a block diagram which shows an example of the scalable video decoder which can utilize the technique by the aspect described in this disclosure.</figref><figref num="6">FIG. 5 is a flow chart illustrating an exemplary method for determining the prediction of an enhancement layer sample with a higher bit depth than the corresponding base layer sample according to aspects of the present disclosure.</figref>
[0025] Some embodiments described herein relate to interlayer prediction for scalable video coding in the context of advanced video codecs, such as HEVC (High Efficiency Video Coding). More specifically, the present disclosure relates to systems and methods for improving the performance of interlayer prediction in HEVC's Scalable Video Coding (SVC) extension.
[0026] The following description describes H.264 / AVC techniques related to some embodiments, as well as HEVC standards and related techniques. Some embodiments will be described herein in the context of HEVC and / or H.264 standards, but the systems and methods disclosed herein may be applicable to any suitable video coding standard. The person skilled in the art will understand. For example, the embodiments disclosed herein include the following standards: ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU- Includes T H.263, ISO / IEC MPEG-4 Visual, and its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions (also as ISO / IEC MPEG-4 AVC). Applicable to one or more of (known) ITU-T H.264.
[0027] HEVC generally follows the framework of previous video coding standards in many respects. The units of prediction in HEVC are different from the units in some previous video coding standards (eg macroblocks). In fact, the concept of macroblocks does not exist in HEVC, as understood in some previous video coding standards. Macroblocks can be replaced with a quadtree-based hierarchy that can provide high flexibility, among other possible benefits. For example, within the HEVC method, a coding unit (CU), a prediction unit (PU: Prediction Unit), and a transformation unit (TU: Transform). Three types of blocks called Unit) are defined. CU may refer to the basic unit of area division. CUs can be considered similar to the concept of macroblocks, but they do not limit the maximum size and can allow recursive splitting into four equally sized CUs to improve content adaptability. The PU can be considered the basic unit of inter / intra prediction, which may contain multiple arbitrary shape divisions in a single PU in order to effectively code irregular image patterns. The TU can be considered the basic unit of conversion. It can be defined independently of the PU, but its size can be limited to the CU to which the TU belongs. This separation of block structures into three different concepts can allow each to be optimized according to its role, which can improve coding efficiency.
[0028] For purposes of illustration only, some embodiments disclosed herein include examples that include only two layers (eg, a lower level layer such as a base layer and an upper level layer such as an enhancement layer). It will be described using. It should be understood that such an example may be applicable to configurations involving multiple base and / or enhancement layers. Further, for the sake of brevity, the following disclosure includes the term "frame" or "block" for some embodiments. However, these terms are not limited. For example, the techniques described below can be used with any suitable video unit, such as blocks (eg, CU, PU, TU, macroblocks, etc.), slices, frames, and so on.
Video coding standard [0029] A digital image, such as a video image, TV image, still image, or image generated by a video recorder or computer, can consist of pixels or samples composed of horizontal and vertical lines. The number of pixels in a single image is generally tens of thousands. Each pixel generally contains luminance information and chrominance information. The amount of information to be carried from the image encoder to the image decoder without compression is so great that it makes real-time image transmission impossible. Several different compression methods have been developed, such as the JPEG, MPEG and H.263 standards, to reduce the amount of information to be transmitted.
[0030] Video coding standards are ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG- 4 Includes ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including Visual and its Scalable Video Coding (SVC) and Multiview Video Coding (MVC) extensions, all of which Is incorporated by reference in its entirety.
In addition, new video coding standards, namely High Efficiency Video Coding (HEVC), are the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Developed by Joint Collaboration Team on Video Coding (JCT-VC) with Group).
A recent draft of HEVC is available from http://phenix.it-sudparis.eu/jct/doc_end_user/documents/12_Geneva/wg11/JCTVC-L1003-v34.zip as of November 22, 2013. The whole is incorporated by reference. For a complete quote of HEVC Draft 10, see Documents JCTVC-L1003, Bross et al., "High Efficiency Video Coding (HEVC) Text Specification Draft 10", ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 Joint Collaborative Team On Video Coding (JCT-VC: Joint Collaborative Team on Video Coding), 12th Meeting: Geneva, Switzerland, January 14, 2013-January 23, 2013.
[0032] Various aspects of the new system, device, and method are described more fully below with reference to the accompanying drawings. However, this disclosure may be implemented in many different forms and should not be construed as being confined to any particular structure or function presented throughout this disclosure. Rather, these aspects are provided to ensure that the disclosure is meticulous and complete and that the scope of the disclosure is fully communicated to those skilled in the art. Based on the teachings of this specification, the scope of the present disclosure, whether implemented independently of other aspects of the invention or in combination with other aspects of the invention, will be disclosed herein. Those skilled in the art should appreciate that it covers any aspect of the system, equipment, and method of. For example, no matter how many aspects described herein are used, the device may be implemented or the method may be implemented. Moreover, the scope of the invention is such that it is practiced using other structures, functions, or structures and functions in addition to or in addition to the various aspects of the invention described herein. It shall cover the device or method. It should be understood that any aspect disclosed herein can be implemented by one or more elements of the claims.
Although specific embodiments are described herein, many variations and substitutions of these embodiments fall within the scope of the present disclosure. Although some of the benefits and benefits of preferred embodiments will be described, the scope of this disclosure is not limited to any particular benefit, use, or purpose. Rather, aspects of the present disclosure shall be widely applicable to a variety of wireless technologies, system configurations, networks, and transmission protocols, some of which are illustrated in the figures and in the following description of preferred embodiments. .. The embodiments and drawings for carrying out the invention are merely explanatory, but not limiting, to the present disclosure, and the scope of the present disclosure is defined by the appended claims and their equivalents.
[0034] The attached drawing shows an example. The elements indicated by reference numbers in the accompanying drawings correspond to the elements indicated by similar reference numbers in the following description. In the present disclosure, elements having names beginning with ordinal words (eg, "first", "second", "third", etc.) do not necessarily imply that they have a particular order. Is not always. Rather, such ordinal words are only used to refer to different elements of the same or similar type.
Video coding system [0035] FIG. 1 is a block diagram illustrating an exemplary video coding system 10 that may utilize the techniques according to the embodiments described herein. The term "video coder" as used and described herein collectively refers to both video encoders and video decoders. In the present disclosure, the term "video coding" or "coding" may collectively refer to video coding and video decoding.
[0036] As shown in FIG. 1, the video coding system 10 includes a source device 12 and a destination device 14. The source device 12 produces encoded video data. The destination device 14 may decode the encoded video data generated by the source device 12. Source device 12 and destination device 14 include desktop computers, notebook (eg laptops) computers, tablet computers, set-top boxes, so-called "smart" phones, so-called "smart" pads and other telephone handset, televisions, cameras. , Display devices, digital media players, video game consoles, in-car computers, and more. In some examples, the source device 12 and the destination device 14 may be equipped for wireless communication.
[0037] The destination device 14 may receive encoded video data from the source device 12 via channel 16. Channel 16 may include any type of medium or device capable of moving encoded video data from the source device 12 to the destination device 14. In one example, channel 16 may include a communication medium that allows the source device 12 to transmit encoded video data directly to the destination device 14 in real time. In this example, the source device 12 may modulate the encoded video data according to a communication standard such as a wireless communication protocol and may transmit the modulated video data to the destination device 14. The communication medium may include a wireless communication medium or a wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or other devices that allow communication from the source device 12 to the destination device 14.
[0038] In another example, channel 16 may correspond to a storage medium that stores the encoded video data generated by the source device 12. In this example, the destination device 14 may access the storage medium via disk access or card access. The storage medium is a variety of locally accessible data storage, such as Blu-ray® discs, DVDs, CD-ROMs, flash memory, or other suitable digital storage media for storing encoded video data. Can include media. In a further example, channel 16 may include a file server or another intermediate storage device that stores the encoded video produced by the source device 12. In this example, the destination device 14 may access the encoded video data stored on the file server or other intermediate storage device via streaming or download. The file server can be a type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Illustrative file servers include web servers (for example, for websites), FTP servers, network attached storage (NAS) devices, and local disk drives. The destination device 14 may access the encoded video data via any standard data connection, including an internet connection. Illustrative types of data connections include wireless channels (eg, Wi-Fi® connections, etc.), wired connections (eg, eg, Wi-Fi® connections, etc.) that are suitable for accessing encoded video data stored on a file server. DSL, cable modem, etc.), or a combination of both. Transmission of encoded video data from a file server can be streaming transmission, download transmission, or a combination of both.
[0039] The techniques of the present disclosure are not limited to wireless applications or configurations. The technique involves over-the-air television broadcasting, cable television transmission, satellite television transmission, eg streaming video transmission over the Internet (eg, dynamic adaptive streaming over HTTP (DASH), etc.), Applies to video coding that supports any of a variety of multimedia applications, such as digital video encoding for storage on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. obtain. In some examples, the video coding system 10 will support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony. Can be configured.
[0040] In the example of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. In some cases, the output interface 22 may include a modulator / demodulator (modem) and / or a transmitter. In the source device 12, the video source 18 is a video capture device, such as a video camera, a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and /. Alternatively, it may include a source such as a computer graphics system for generating video data, or a combination of such sources.
[0041] The video encoder 20 may be configured to encode captured video data, previously captured video data, or computer-generated video data. The encoded video data may be transmitted directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may also be stored on a storage medium or file server for later access by the destination device 14 for decryption and / or playback.
In the example of FIG. 1, the destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. In some cases, the input interface 28 may include a receiver and / or a modem. The input interface 28 of the destination device 14 receives the encoded video data via the channel 16. The encoded video data may include various syntax elements generated by the video encoder 20 that represent the video data. Syntax elements can describe the characteristics and / or processing of blocks and other coding units, such as group of pictures (GOP). Such syntax elements may be included with encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
[0043] The display device 32 may be integrated with or external to the destination device 14. In some examples, the destination device 14 may include an integrated display device and may be configured to interface with an external display device. In another example, the destination device 14 can be a display device. In general, the display device 32 displays the decoded video data to the user. The display device 32 may comprise any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0044] The video encoder 20 and video decoder 30 may operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard currently under development, and may comply with the HEVC test model (HM). Alternatively, the video encoder 20 and video decoder 30 are alternative to other proprietary or industry standards, such as MPEG-4, Part10, the ITU-T H.264 standard called Advanced Video Coding (AVC), or so. It can operate according to the extension of the standard. However, the techniques disclosed are not limited to any particular coding standard. Other examples of video compression standards are MPEG-2 and ITU-T H.263.
Although not shown in the example of Figure 1, the video encoder 20 and video decoder 30 can be integrated with the audio encoder and decoder, respectively, and include the appropriate MUX-DEMUX unit, or other hardware and software. Can handle both audio and video encoding in a common data stream or separate data streams. Where applicable, in some examples the MUX-DEMUX unit may comply with the ITU H.223 multiplexer protocol, or other protocols such as the user datagram protocol (UDP).
Again, FIG. 1 is merely an example, and the techniques of the present disclosure may include video coding settings (eg, video coding or) that do not necessarily include data communication between the encoding and decoding devices. Can be applied to video decoding). In other examples, data can be retrieved from local memory, streamed over a network, and so on. A coding device can encode the data and store it in memory, and / or a decoding device can retrieve the data from memory and decode it. In many examples, coding and decoding are done by devices that do not communicate with each other, but only encode the data into memory and / or retrieve and decode the data from memory.
[0047] The video encoder 20 and video decoder 30 are one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, and hardware, respectively. It can be implemented as any of a variety of suitable circuits, or any combination thereof. When the technique is partially implemented in software, the device may store software instructions in a suitable non-temporary computer-readable storage medium and use one or more processors to store the instructions in hardware. It can be carried out to perform the techniques of the present disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, both of which may be integrated as part of a composite encoder / decoder (codec) in their respective devices. The device including the video encoder 20 and / or the video decoder 30 may include wireless communication devices such as integrated circuits, microprocessors, and / or cellular phones.
[0048] As briefly mentioned above, the video encoder 20 encodes video data. Video data can include one or more pictures. Each of the pictures is a still image that forms part of the video. In some cases, pictures are sometimes referred to as video "frames." When the video encoder 20 encodes the video data, the video encoder 20 may generate a bitstream. A bitstream may contain a sequence of bits that form a coded representation of the video data. The bitstream can contain coded pictures and related data. A coded picture is a coded representation of a picture.
[0049] To generate a bitstream, the video encoder 20 may perform a coding operation on each picture in the video data. When the video encoder 20 performs a coding operation on a picture, the video encoder 20 may generate a series of coded pictures and related data. Relevant data may include a video parameter set (VPS), a sequence parameter set, a picture parameter set, an adaptive parameter set, and other syntax structures. A sequence parameter set (SPS) may contain parameters that are applicable to zero or more sequences of pictures. A picture parameter set (PPS) may contain parameters applicable to zero or more pictures. Adaptation parameter set (APS) set) may contain parameters applicable to zero or more pictures. The parameters in APS can be parameters that are more likely to change than the parameters in PPS.
[0050] To generate a coded picture, the video encoder 20 may divide the picture into blocks of equal size video. The video block can be a two-dimensional array of samples. Each of the video blocks is associated with a tree block. In some cases, the tree block is sometimes referred to as the largest coding unit (LCU). HEVC tree blocks can be broadly similar to previous standard macroblocks, such as H.264 / AVC. However, the tree block is not necessarily limited to a particular size and may include one or more coding units (CUs). The video encoder 20 can use quadtree partitioning to partition the video blocks of a tree block into the video blocks associated with the CU, hence the name "tree block".
[0051] In some examples, the video encoder 20 may divide the picture into multiple slices. Each slice can contain an integer number of CUs. In some cases, the slice comprises an integer number of tree blocks. In other cases, slice boundaries can be within tree blocks.
[0052] As part of performing a coding operation on a picture, the video encoder 20 may perform a coding operation on each slice of the picture. When the video encoder 20 performs a coding operation on a slice, the video encoder 20 may generate the coded data associated with the slice. The coded data associated with the slice is sometimes referred to as the "coded slice".
[0053] To generate a coded slice, the video encoder 20 may perform a coding operation on each tree block in the slice. When the video encoder 20 performs a coding operation on the tree block, the video encoder 20 may generate a coded tree block. A coded tree block may include data representing a coded version of the tree block.
[0054] When the video encoder 20 produces a coded slice, the video encoder 20 may perform coding operations on the tree blocks in the slice according to the raster scan order (eg, it may encode the tree blocks). .. For example, the video encoder 20 goes from left to right over the top row of the tree block being sliced, then from left to right over the next bottom row of the tree block, and so on, in order of video. The tree blocks of the slice may be encoded until the encoder 20 encodes each of the tree blocks in the slice.
As a result of encoding the tree blocks according to the raster scan order, the tree blocks above and to the left of a given tree block may be encoded, but the trees below and to the right of a given tree block. The block is not yet encoded. Thus, when encoding a given tree block, the video encoder 20 may be able to access the information generated by encoding the tree blocks above and to the left of the given tree block. However, when encoding a given tree block, the video encoder 20 may not be able to access the information generated by encoding the tree blocks below and to the right of the given tree block.
[0056] To generate a coded tree block, the video encoder 20 may recursively divide the video blocks of the tree block into quadtrees to gradually divide the video blocks into smaller video blocks. Each of the smaller video blocks can be associated with a different CU. For example, the video encoder 20 may divide a video block of a tree block into four equal sized subblocks and one or more of the subblocks into four equal sized subsubblocks, and so on. is there. A partitioned CU can be a CU in which its video block is partitioned into video blocks associated with other CUs. An unsegmented CU can be a CU whose video block is not partitioned into video blocks associated with other CUs.
[0057] One or more syntax elements in the bitstream may indicate the maximum number of times the video encoder 20 can segment a video block in a tree block. The CU video block can be square in shape. The size of a video block in a CU (eg, the size of a CU) can range from 8x8 pixels to the size of a video block in a tree block with up to 64x64 or more pixels (eg, the size of a tree block).
[0058] The video encoder 20 may perform coding operations on each CU in the tree block according to the z scan order (eg, each CU may be encoded). In other words, the video encoder 20 may encode the upper left CU, the upper right CU, the lower left CU, and then the lower right CU in that order. When the video encoder 20 performs a coding operation on the partitioned CU, the video encoder 20 may encode the CU associated with the subblock of the video block of the partitioned CU according to the z scan order. In other words, the video encoder 20 is associated with the CU associated with the upper left subblock, the CU associated with the upper right subblock, the CU associated with the lower left subblock, and then the lower right subblock. The CUs can be encoded in that order.
As a result of encoding the CUs in the tree block according to the z scan order, the top left, top right, left, and bottom left CUs of a given CU may be encoded. The CUs below and to the right of a given CU are not yet encoded. Thus, when encoding a given CU, the video encoder 20 may be able to access the information generated by encoding several CUs adjacent to the given CU. However, when encoding a given CU, the video encoder 20 may not be able to access the information generated by encoding other CUs adjacent to the given CU.
[0060] When the video encoder 20 encodes an unpartitioned CU, the video encoder 20 may generate one or more predictive units (PUs) for the CU. Each of the CU's PUs can be associated with different video blocks within the CU's video blocks. The video encoder 20 may generate predictive video blocks for each PU of the CU. The PU predictive video block can be a block of samples. The video encoder 20 may use intra-prediction or inter-prediction to generate a predictive video block for the PU.
[0061] When the video encoder 20 uses intra-prediction to generate a predictive video block of the PU, the video encoder 20 may generate a predictive video block of the PU based on the decoding sample of the picture associated with the PU. .. If the video encoder 20 uses intra-prediction to generate a predictive video block of the PU of the CU, then the CU is the intra-predicted CU. When the video encoder 20 uses inter-prediction to generate a PU predictive video block, the video encoder 20 uses the PU predictive video based on a decoded sample of one or more pictures other than the picture associated with the PU. Can generate blocks. If the video encoder 20 uses inter-prediction to generate a predictive video block of the CU's PU, then the CU is an inter-predicted CU.
[0062] Further, when the video encoder 20 uses inter-prediction to generate a predictive video block for the PU, the video encoder 20 may generate PU motion information. PU motion information can indicate one or more reference blocks of the PU. Each reference block in the PU can be a video block in the reference picture. The reference picture can be a picture other than the picture associated with the PU. In some cases, the PU reference block is sometimes referred to as the PU "reference sample". The video encoder 20 may generate a predictive video block for the PU based on the PU reference block.
[0063] After the video encoder 20 generates the predicted video block for one or more PUs of the CU, the video encoder 20 produces the residual data of the CU based on the predicted video block for the PUs of the CU. Can be generated. The CU residual data can indicate the difference between the sample in the predicted video block for the CU PU and the sample in the CU's original video block.
[0064] In addition, as part of performing the coding operation on the unpartitioned CU, the video encoder 20 performs a recursive quadrant on the residual data of the CU to perform the CU's residual data. The residual data can be divided into one or more blocks of residual data (eg, residual video blocks) associated with the CU's conversion unit (TU). Each TU of the CU can be associated with a different residual video block.
[0065] The videocoder 20 may apply one or more transformations to the residual video block associated with the TU to generate a transformation factor block associated with the TU (eg, a block of transformation coefficients). Conceptually, the transformation factor block can be a two-dimensional (2D) matrix of transformation coefficients.
[0066] After generating the conversion factor blocks, the video encoder 20 may perform a quantization process on the conversion factor blocks. Quantization generally refers to the process by which the conversion factors are quantized to achieve further compression in order to reduce the amount of data used to represent the conversion factors as much as possible. The quantization process can reduce the bit depth associated with some or all of the conversion factors. For example, during quantization, the n-bit conversion factor may be truncated to the m-bit conversion factor, where n is greater than m.
The video encoder 20 may associate each CU with a quantization parameter (QP) value. The QP value associated with the CU can determine how the video encoder 20 quantizes the transformation factor block associated with the CU. The video encoder 20 may adjust the degree of quantization applied to the conversion factor block associated with the CU by adjusting the QP value associated with the CU.
[0068] After the video encoder 20 has quantized the conversion factor block, the video encoder 20 may generate a set of syntax elements representing the conversion factor in the quantized conversion factor block. The video encoder 20 may apply entropy encoding operations such as Context Adaptive Binary Arithmetic Coding (CABAC) operations to some of these syntax elements. Other entropy coding techniques may be used, such as content adaptive variable length coding (CAVLC), probability interval partitioning entropy (PIPE) coding, or other binary arithmetic coding.
[0069] The bitstream produced by the video encoder 20 may include a series of Network Abstraction Layer (NAL) units. Each of the NAL units can be a syntax structure that contains an indication of the type of data in the NAL unit and the bytes that contain the data. For example, a NAL unit is a video parameter set, sequence parameter set, picture parameter set, coded slice, supplemental enhancement information (SEI), access unit delimiter, filler data, or data representing another type of data. May include. The data in the NAL unit can contain various syntactic structures.
[0070] The video decoder 30 may receive the bitstream generated by the video encoder 20. The bitstream may include a coded representation of the video data encoded by the video encoder 20. When the video decoder 30 receives the bitstream, the video decoder 30 may perform parsing operations on the bitstream. When the video decoder 30 performs a parsing operation, the video decoder 30 may extract syntax elements from the bitstream. The video decoder 30 may reconstruct a picture of video data based on syntax elements extracted from the bitstream. The process for reconstructing video data based on syntax elements can generally be the reverse of the process performed by the video encoder 20 to generate syntax elements.
After the video decoder 30 extracts the syntax elements associated with the CU, the video decoder 30 may generate predictive video blocks for the PU of the CU based on the syntax elements. In addition, the video decoder 30 can inverse quantize the conversion factor block associated with the TU of the CU. The video decoder 30 may perform an inverse transformation on the conversion factor block to reconstruct the residual video block associated with the TU of the CU. After generating the predicted video block and reconstructing the residual video block, the video decoder 30 may reconstruct the video block of the CU based on the predicted video block and the residual video block. In this way, the video decoder 30 can reconstruct the video block of the CU based on the syntax elements in the bitstream.
Video encoder [0072] FIG. 2 is a block diagram showing an example of a video encoder that can implement the technique according to the embodiment described in the present disclosure. The video encoder 20 may be configured to perform any or all of the techniques of the present disclosure. As an example, the prediction unit 100 may be configured to perform any or all of the techniques described in this disclosure. In another embodiment, the video encoder 20 includes a voluntary interlayer prediction unit 128 configured to perform any or all of the techniques described herein. In other embodiments, inter-layer prediction may be performed by prediction unit 100 (eg, inter-prediction unit 121 and / or intra-prediction unit 126), in which case inter-layer prediction unit 128 may be omitted. However, the aspects of the present disclosure are not so limited. In some examples, the techniques described in this disclosure may be shared among the various components of the video encoder 20. In some examples, additionally or instead, a processor (not shown) may be configured to perform any or all of the techniques described in this disclosure. As further described below with respect to FIG. 6, one or more components of the video encoder 20 may be configured to perform the method shown in FIG. For example, inter-prediction unit 121 (eg, via motion estimation unit 122 and / or motion compensation unit 124), intra-prediction unit 126, or inter-layer prediction unit 128 are shown in FIG. 6, together or separately. It can be configured to do the way it is.
[0073] In some cases, the video encoder 20 can be considered the same as the video encoder 400 (discussed below) shown in FIG. 4, but in each figure different aspects of the video encoder are highlighted. There is. In detail, the description of the video encoder 20 in FIG. 1 generally focuses on features related to block-based coding, while the description of the video encoder 400 in FIG. 4 is scalable video coding and increased bits. Particular focus is placed on features related to interlayer prediction of EL samples with depth. In some examples, the techniques described in the present disclosure may be shared between the various components of the video encoder 20 and the video encoder 400. In some examples, additionally or instead, a processor (not shown) may be configured to perform any or all of the techniques described in this disclosure.
[0074] For purposes of illustration, this disclosure describes the video encoder 20 in the context of HEVC coding. However, the techniques of the present disclosure may be applicable to other coding standards or methods.
[0075] The video encoder 20 may perform intracoding and intercoding of video blocks in a video slice. Intracoding relies on spatial prediction to reduce or eliminate the spatial redundancy of video within a given video frame or picture. Intercoding relies on temporal prediction to reduce or eliminate temporal redundancy of video in adjacent frames or pictures of a video sequence. Intra mode (I mode) may refer to any of several space-based coding modes. Intermodes such as unidirectional prediction (P mode) or bidirectional prediction (B mode) may refer to any of several time-based coding modes.
[0076] In the example of FIG. 2, the video encoder 20 includes a plurality of functional components. The functional components of the video encoder 20 are the prediction unit 100, the residual generation unit 102, the conversion unit 104, the quantization unit 106, the inverse quantization unit 108, the inverse conversion unit 110, and the reconstruction unit 112. , A filter unit 113, a decoding picture buffer 114, and an entropy coding unit 116. The prediction unit 100 includes an inter-prediction unit 121, a motion estimation unit 122, a motion compensation unit 124, an intra-prediction unit 126, and an inter-layer prediction unit 128. In another example, the video encoder 20 may include more, fewer, or different functional components. Further, the motion estimation unit 122 and the motion compensation unit 124 can be highly integrated, but in the example of FIG. 2, they are represented separately for illustration purposes.
[0077] The video encoder 20 may receive video data. The video encoder 20 may receive video data from various sources. For example, the video encoder 20 may receive video data from video source 18 (FIG. 1) or another source. Video data can represent a series of pictures. To encode the video data, the video encoder 20 may perform a coding operation on each of the pictures. As part of performing the coding operation on the picture, the video encoder 20 may perform the coding operation on each slice of the picture. As part of performing the coding operation on the slice, the video encoder 20 may perform the coding operation on the tree block in the slice.
[0078] As part of performing a coding operation on the tree block, the prediction unit 100 performs a quadtree division on the video block of the tree block and gradually divides the video block into smaller video blocks. Can be done. Each of the smaller video blocks can be associated with a different CU. For example, the prediction unit 100 may divide the video block of the tree block into four equal sized subblocks and one or more of the subblocks into four equal sized subsubblocks, and so on. is there.
The size of the video block associated with the CU can range from 8x8 samples to the size of a tree block with up to 64x64 or more samples. In the present disclosure, "NxN (NxN)" and "NxN (N by N)" are sample dimensions of a video block with respect to vertical and horizontal dimensions, such as 16x16 (16x16) samples or 16x16. (16 by 16) Can be used interchangeably to refer to samples. In general, a 16x16 video block has 16 samples vertically (y = 16) and 16 samples horizontally (x = 16). Similarly, an N × N block generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value.
[0080] Further, as part of performing a coding operation on the tree block, the prediction unit 100 may generate a hierarchical quadtree data structure for the tree block. For example, a tree block can correspond to a root node in a quadtree data structure. If the prediction unit 100 divides the video block of the tree block into four subblocks, the root node has four child nodes in the quadtree data structure. Each of the child nodes corresponds to the CU associated with one of the subblocks. If the prediction unit 100 divides one of the subblocks into four subsubblocks, the nodes corresponding to the CUs associated with the subblocks will each be in the CU associated with one of the subsubblocks. It can have four corresponding child nodes.
[0081] Each node in a quadtree data structure may contain syntax data (eg, syntax elements) for the corresponding tree block or CU. For example, a node in a quadtree may contain a split flag that indicates whether the video block of the CU corresponding to that node is divided into four subblocks (eg, divided). The syntax element for the CU can be defined recursively and can depend on whether the CU's video block is divided into subblocks. A CU whose video block is not partitioned can correspond to a leaf node in a quadtree data structure. The coded tree block may contain data based on the quadtree data structure for the corresponding tree block.
[0082] The video encoder 20 may perform a coding operation on each CU that is not divided into tree blocks. When the video encoder 20 performs a coding operation on an undivided CU, the video encoder 20 generates data representing a coded representation of the undivided CU.
[0083] As part of performing a coding operation on the CU, the prediction unit 100 may segment the video blocks of the CU within one or more PUs of the CU. The video encoder 20 and video decoder 30 may support a variety of PU sizes. Assuming that the size of a particular CU is 2Nx2N, the video encoder 20 and video decoder 30 have a PU size of 2Nx2N or NxN and 2Nx2N, 2NxN, Nx2N, Nx. It may support inter-prediction with N, 2N × nU, nL × 2N, nR × 2N, or similar symmetric PU sizes. The video encoder 20 and video decoder 30 may also support asymmetric divisions for 2N × nU, 2N × nD, nL × 2N, and nR × 2N PU sizes. In some examples, the prediction unit 100 makes geometric divisions, such as dividing the CU video blocks between the CU PUs, along boundaries that do not intersect the sides of the CU video blocks at right angles. obtain.
[0084] The inter-prediction unit 121 can perform inter-prediction for each PU of the CU. Inter-prediction can achieve time compression. In order to make an inter-prediction for the PU, the motion estimation unit 122 can generate the motion information of the PU. The motion compensation unit 124 may generate a predictive video block for the PU based on motion information and a decoded sample of a picture other than the picture associated with the CU (eg, a reference picture). In the present disclosure, the predictive video block generated by the motion compensation unit 124 may be referred to as an inter-predictive video block.
[0085] Slices can be I slices, P slices, or B slices. The motion estimation unit 122 and the motion compensation unit 124 can perform different operations on the PU of the CU depending on whether the PU is in the I slice, the P slice, or the B slice. In the I slice, all PUs are predicted intra. Therefore, when the PU is in the I slice, the motion estimation unit 122 and the motion compensation unit 124 do not make an inter-prediction for the PU.
[0086] When the PU is in a P-slice, the picture containing the PU is associated with a list of referenced pictures called "List 0". Each of the reference pictures in Listing 0 contains a sample that can be used for interprediction of other pictures. When the motion estimation unit 122 performs a motion estimation operation on the PU in the P slice, the motion estimation unit 122 can search the reference picture in the list 0 for the reference block for the PU. The PU reference block can be a set of samples that most closely corresponds to the sample in the PU video block, eg, a block of samples. The motion estimation unit 122 can use various metrics to determine how closely the set of samples in the reference picture corresponds to the samples in the PU video block. For example, the motion estimation unit 122 has a sum of absolute difference (SAD) and a sum of square (SSD). The difference), or other difference metric, can determine how closely the set of samples in the reference picture corresponds to the samples in the PU video block.
After identifying the reference block of the PU in the P-slice, the motion estimation unit 122 contains the reference block, which is the space between the PU and the reference block, indicating the reference picture in Listing 0. It is possible to generate a motion vector indicating the displacement. In various examples, the motion estimation unit 122 can generate motion vectors with different accuracy. For example, motion estimation unit 122 may generate motion vectors with 1/4 sample accuracy, 1/8 sample accuracy, or other fractional sample accuracy. For fractional sample precision, the reference block value can be interpolated from the sample value at an integer position in the reference picture. The motion estimation unit 122 can output a reference index and a motion vector as PU motion information. The motion compensation unit 124 may generate a predictive video block of the PU based on the reference block identified by the motion information of the PU.
[0088] If the PU is in a B slice, the picture containing the PU can be associated with two lists of reference pictures called "List 0" and "List 1". In some examples, a picture containing a B slice can be associated with a list combination, which is a combination of Listing 0 and Listing 1.
Further, when the PU is in the B slice, the motion estimation unit 122 may make a unidirectional or bidirectional prediction about the PU. When the motion estimation unit 122 makes a unidirectional prediction about the PU, the motion estimation unit 122 may search for a reference picture in Listing 0 or Listing 1 for a reference block for the PU. The motion estimation unit 122 may then generate a reference index indicating the reference picture in Listing 0 or Listing 1, including the reference block, and a motion vector indicating the spatial displacement between the PU and the reference block. The motion estimation unit 122 can output a reference index, a prediction direction indicator, and a motion vector as PU motion information. The predictive direction indicator may indicate whether the reference index points to a reference picture in Listing 0 or a reference picture in Listing 1. The motion compensation unit 124 may generate a predictive video block of the PU based on the reference block indicated by the motion information of the PU.
[0090] When the motion estimation unit 122 makes a bidirectional prediction about the PU, the motion estimation unit 122 can search the reference picture in Listing 0 for the reference block for the PU, and also for the PU. You can search for a reference picture in Listing 1 for another reference block. The motion estimation unit 122 may then generate a reference index indicating the reference picture in Listing 0 and Listing 1, including the reference block, and a motion vector indicating the spatial displacement between the reference block and the PU. The motion estimation unit 122 can output the PU reference index and the motion vector as the motion information of the PU. The motion compensation unit 124 may generate a predictive video block of the PU based on the reference block indicated by the motion information of the PU.
[0091] In some cases, motion estimation unit 122 does not output the full set of PU motion information to entropy coding unit 116. Instead, the motion estimation unit 122 may signal the motion information of the PU by referring to the motion information of another PU. For example, the motion estimation unit 122 may determine that the motion information of the PU is sufficiently similar to the motion information of the adjacent PU. In this example, the motion estimation unit 122 may indicate to the video decoder 30 that the PU has the same motion information as the adjacent PU in the syntax structure associated with the PU. In another example, motion estimation unit 122 may distinguish between adjacent PUs and motion vector decompression (MVD) in the syntax structure associated with the PU. The motion vector difference indicates the difference between the motion vector of the PU and the motion vector of the indicated adjacent PU. The video decoder 30 may use the motion vector of the indicated adjacent PU and the motion vector difference to determine the motion vector of the PU. By referencing the motion information of the first PU when signaling the motion information of the second PU, the video encoder 20 uses a smaller number of bits to signal the motion information of the second PU. Can be possible.
[0092] As part of performing the coding operation on the CU, the intra prediction unit 126 may perform the intra prediction on the PU of the CU. Intra-prediction can achieve spatial compression. When the intra-prediction unit 126 makes an intra-prediction to the PU, the intra-prediction unit 126 may generate PU prediction data based on decoded samples of other PUs in the same picture. The PU prediction data may include a prediction video block and various syntax elements. The intra prediction unit 126 can make intra predictions for PUs in I slices, P slices, and B slices.
[0093] To make an intra-prediction for a PU, the intra-prediction unit 126 may use multiple intra-prediction modes to generate multiple sets of PU prediction data. When the intra prediction unit 126 uses the intra prediction mode to generate a set of prediction data for the PU, the intra prediction unit 126 is from the video block of the adjacent PU in the direction and / or gradient associated with the intra prediction mode. Samples can be extended across PU video blocks. Adjacent PUs can be above, top right, top left, or left of the PU, assuming a left-to-right, top-to-bottom coding order for the PU, CU, and tree blocks. The intra prediction unit 126 may use various numbers of intra prediction modes, for example 33 directional intra prediction modes, depending on the size of the PU.
[0094] The prediction unit 100 may select the prediction data of the PU from the prediction data generated by the motion compensation unit 124 for the PU or the prediction data generated by the intra prediction unit 126 for the PU. .. In some examples, the prediction unit 100 selects the prediction data for the PU based on the rate / strain metrics of the set of prediction data.
[0095] If the prediction unit 100 selects the prediction data generated by the intra prediction unit 126, the prediction unit 100 uses the intra prediction mode used to generate the prediction data for the PU, eg, the selected intra. Predictive modes can be signaled. The prediction unit 100 may signal the selected intra prediction mode in various ways. For example, the selected intra-prediction mode may be the same as the intra-prediction mode of the adjacent PU. In other words, the intra-prediction mode of the adjacent PU can currently be the most probable mode for the PU. Therefore, the prediction unit 100 may generate a syntax element to indicate that the selected intra prediction mode is the same as the intra prediction mode of the adjacent PU.
[0096] As described above, the video encoder 20 may include an interlayer prediction unit 128. The inter-layer prediction unit 128 is configured to predict the current block (eg, the current block in the EL) using one or more different layers available in the SVC (eg, base layer or reference layer). Will be done. Such predictions are sometimes referred to as interlayer predictions. The inter-layer prediction unit 128 utilizes a prediction method to reduce inter-layer redundancy, thereby improving coding efficiency and reducing computational resource requirements. Some examples of inter-layer prediction are inter-layer intra-layer prediction, inter-layer motion prediction, and inter-layer residual prediction. Interlayer intra-layer prediction uses the reconstruction of collated blocks in the base layer to predict the current block in the enhancement layer. Inter-layer motion prediction uses motion information from the base layer to predict motion during the enhancement layer. Inter-layer residual prediction uses the residuals of the base layer to predict the residuals of the enhancement layer.
[0097] After the prediction unit 100 selects the prediction data for the PU of the CU, the residual generation unit 102 subtracts the prediction video block of the PU of the CU from the video block of the CU (eg, indicated by a minus sign). , CU residual data can be generated. The CU residual data may include 2D residual video blocks that correspond to different sample components of the sample in the CU video block. For example, the residual data may include a residual video block that corresponds to the difference between the luminance component of the sample in the predicted video block of the PU of the CU and the luminance component of the sample in the original video block of the CU. .. In addition, the CU residual data is a residual video block that corresponds to the difference between the sample chrominance component in the CU's PU predicted video block and the sample's chrominance component in the CU's original video block. Can include.
[0098] The prediction unit 100 may perform a quadtree division to divide the residual video blocks of the CU into subblocks. Each unsplit residual video block can be associated with a different TU in the CU. The size and location of the residual video blocks associated with the CU's TU may or may not be based on the size and location of the video blocks associated with the CU's PU. A quadtree structure known as a "residual quad tree" (RQT) may contain nodes associated with each of the residual video blocks. The TU of the CU can correspond to the leaf node of the RXT.
[0099] The conversion unit 104 may generate one or more conversion factor blocks for each TU of the CU by applying one or more conversions to the residual video blocks associated with the TU. Each of the transformation coefficient blocks can be a 2D matrix of transformation coefficients. The conversion unit 104 may apply various conversions to the residual video block associated with the TU. For example, the transform unit 104 may apply a discrete cosine transform (DCT), a directional transform, or a conceptually similar transform to the residual video block associated with the TU.
[00100] After the conversion unit 104 generates the conversion factor block associated with the TU, the quantization unit 106 can quantize the conversion factors in the conversion factor block. The quantization unit 106 can quantize the conversion factor block associated with the TU of the CU based on the QP value associated with the CU.
[00101] The video encoder 20 may associate a QP value with a CU in various ways. For example, the video encoder 20 may perform rate distortion analysis on the tree block associated with the CU. In rate distortion analysis, the video encoder 20 may generate multiple coded representations of the tree block by performing coding operations on the tree block multiple times. When the video encoder 20 produces different coded representations of the tree block, the video encoder 20 may associate different QP values with the CU. The video encoder 20 may signal that a given QP value is associated with the CU when the given QP value is associated with the CU in the coded representation of the tree block with the lowest bit rate and strain metrics.
[00102] The inverse quantization unit 108 and the inverse transformation unit 110 can respectively apply the inverse quantization and the inverse transformation to the transformation coefficient block to reconstruct the residual video block from the transformation coefficient block. Reconstruction unit 112 was reconstructed associated with the TU by adding the reconstructed residual video block to the corresponding sample from one or more prediction video blocks generated by prediction unit 100. Can generate video blocks. By reconstructing the video block for each TU of the CU in this way, the video encoder 20 can reconstruct the video block of the CU.
[00103] After the reconfiguration unit 112 reconfigures the video block of the CU, the filter unit 113 may perform a deblocking operation to reduce blocking artifacts in the video block associated with the CU. After performing one or more deblocking operations, the filter unit 113 may store the reconstructed video block of the CU in the decoding picture buffer 114. The motion estimation unit 122 and the motion compensation unit 124 may use the reference picture containing the reconstructed video block to make an inter-prediction for the PU of the subsequent picture . In addition, the intra-prediction unit 126 may use the reconstructed video block in the decrypted picture buffer 114 to make intra-prediction for other PUs in the same picture as the CU.
[00104] The entropy coding unit 116 may receive data from other functional components of the video encoder 20. For example, the entropy coding unit 116 may receive a transformation factor block from the quantization unit 106 and a syntax element from the prediction unit 100. When the entropy coding unit 116 receives the data, the entropy coding unit 116 may perform one or more entropy coding operations to generate the entropy coded data. For example, the video encoder 20 has context-adaptive variable-length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, and syntax-based context-adaptive binary arithmetic coding (SBAC: syntax-). based context-adaptive binary arithmetic A coding) operation, a probability interval segmentation entropy (PIPE) coding operation, or another type of entropy coding operation can be performed on the data. The entropy coding unit 116 may output a bit stream containing the entropy coded data.
[00105] As part of performing an entropy coding operation on the data, the entropy coding unit 116 may choose a context model. If the entropy coding unit 116 is performing a CABAC operation, the context model may give an estimate of the probability of a particular bin with a particular value. In the context of CABAC, the term "bin" is used to refer to the bits of the binarized version of the syntax element.
Video decoder [00106] FIG. 3 is a block diagram showing an example of a video decoder that can implement the technique according to the embodiments described in the present disclosure. The video decoder 30 may be configured to perform any or all of the techniques of the present disclosure. As an example, motion compensation unit 162 and / or intra-prediction unit 164 may be configured to perform any or all of the techniques described herein. In one embodiment, the video decoder 30 may optionally include an interlayer prediction unit 166 configured to perform any or all of the techniques described herein. In other embodiments, the inter-layer prediction can be made by the prediction unit 152 (eg, motion compensation unit 162 and / or intra-prediction unit 164), in which case the inter-layer prediction unit 166 can be omitted. However, the aspects of the present disclosure are not so limited. In some examples, the techniques described in the present disclosure may be shared among the various components of the video decoder 30. In some examples, additionally or instead, a processor (not shown) may be configured to perform any or all of the techniques described in this disclosure. As further described below with respect to FIG. 6, one or more components of the video decoder 30 may be configured to perform the method shown in FIG. For example, motion compensation unit 162, intra prediction unit 164, or interlayer prediction unit 166 may be configured to perform the method shown in FIG. 6, together or separately.
[00107] In the example of FIG. 3, the video decoder 30 includes a plurality of functional components. The functional components of the video decoder 30 include an entropy decoding unit 150, a prediction unit 152, an inverse quantization unit 154, an inverse transformation unit 156, a reconstruction unit 158, a filter unit 159, and a decoding picture buffer 160. Including. The prediction unit 152 includes a motion compensation unit 162, an intra prediction unit 164, and an interlayer prediction unit 166. In some examples, the video decoder 30 may perform a decoding path that is generally opposite to the coding path described for the video encoder 20 in FIG. In another example, the video decoder 30 may include more, fewer, or different functional components.
[00108] The video decoder 30 may receive a bitstream containing encoded video data. A bitstream can contain multiple syntax elements. When the video decoder 30 receives the bitstream, the entropy decoding unit 150 may perform a parsing operation on the bitstream. As a result of performing the parsing operation on the bitstream, the entropy decoding unit 150 can extract the syntax element from the bitstream. As part of performing the parsing operation, the entropy decoding unit 150 can entropy decode the entropy encoding syntax elements in the bitstream. Prediction unit 152, inverse quantization unit 154, inverse transformation unit 156, reconstruction unit 158, and filter unit 159 may perform reconstruction operations that generate decoded video data based on syntax elements extracted from the bitstream. ..
[00109] As described above, a bitstream may include a set of NAL units. The bitstream NAL unit may include a video parameter set NAL unit, a sequence parameter set NAL unit, a picture parameter set NAL unit, a SEI NAL unit, and the like. As part of performing parsing operations on the bitstream, the entropy decoding unit 150 has a sequence parameter set from the sequence parameter set NAL unit, a picture parameter set from the picture parameter set NAL unit, and a SEI NAL unit. It is possible to perform a parse operation that extracts SEI data from the data and decodes the entropy.
[00110] Further, the bitstream NAL unit may include a coded slice NAL unit. As part of performing the parse operation on the bitstream, the entropy decoding unit 150 may perform the parse operation of extracting the coded slice from the coded slice NAL unit and performing the entropy decoding. Each of the coded slices may contain a slice header and slice data. The slice header may contain syntax elements for the slice. The syntax element in the slice header may include a syntax element that identifies the picture parameter set associated with the picture containing the slice. The entropy decoding unit 150 can perform an entropy decoding operation such as a CABAC decoding operation on the syntax element in the coded slice header in order to restore the slice header.
[00111] As part of extracting slice data from the NAL unit of the coded slice, the entropy decoding unit 150 may perform a parsing operation to extract the syntax element from the coded CU in the slice data. The extracted syntax elements may include syntax elements associated with the transformation factor block. The entropy decoding unit 150 may then perform a CABAC decoding operation on some of the syntax elements.
[00112] After the entropy decoding unit 150 performs the parsing operation on the undivided CU, the video decoder 30 may perform the reconstructing operation on the undivided CU. In order to perform the reconstruction operation on the undivided CU, the video decoder 30 may perform the reconstruction operation on each TU of the CU. By performing a reconstruction operation on each TU of the CU, the video decoder 30 may reconstruct the residual video block associated with the CU.
[00113] As part of performing the reconstruction operation on the TU, the inverse quantization unit 154 inverse quantizes the transformation coefficient block associated with the TU, eg, inverse quantize (de-). quantize) possible. The dequantization unit 154 can dequantize the transformation factor blocks in a manner similar to the dequantization process proposed for HEVC or defined by the H.264 decoding standard. The dequantization unit 154 determines the degree of quantization and similarly by the video encoder 20 for the CU of the conversion factor block to determine the degree of dequantization to which the dequantization unit 154 should apply. The calculated quantization parameter QP can be used.
[00114] After the inverse quantization unit 154 dequantizes the transformation factor block, the inverse transformation unit 156 may generate a residual video block for the TU associated with the transformation factor block. The inverse transformation unit 156 may apply an inverse transformation to the transformation factor blocks to generate a residual video block for the TU. For example, the inverse transformation unit 156 transforms the transformation coefficient block into an inverse DCT, an inverse integer transformation, and an inverse Karhunen-Loeve transformation (KLT: Karhunen-Loeve). Transform), reverse rotation transformation, reverse transformation, or another inverse transformation may be applied. In some examples, the inverse transformation unit 156 may determine the inverse transformation to be applied to the transformation factor block based on the signaling from the video encoder 20. In such an example, the inverse transformation unit 156 may determine the inverse transformation based on the transformation signaled at the root node of the quadtree of the tree block associated with the transformation factor block. In another example, the inverse transformation unit 156 can infer the inverse transformation from one or more coding characteristics, such as block size, coding mode, and so on. In some examples, the inverse transformation unit 156 may apply cascading inverse transformation.
[00115] In some examples, motion compensation unit 162 may improve the predicted video block of the PU by performing interpolation based on an interpolation filter. An identifier for the interpolation filter to be used for motion compensation with subsample accuracy may be included in the syntax element. The motion compensation unit 162 may calculate the interpolated values for the sub-integer sample of the reference block using the same interpolation filter used by the video encoder 20 during the generation of the PU predictive video block. The motion compensation unit 162 may determine the interpolation filter used by the video encoder 20 according to the received syntax information and use the interpolation filter to generate a predictive video block.
[00116] If the PU is encoded using intra-prediction, intra-prediction unit 164 may make intra-prediction to generate a predictive video block for the PU. For example, the intra prediction unit 164 may determine the intra prediction mode for the PU based on the syntax elements in the bitstream. The bitstream may contain syntax elements that the intra-prediction unit 164 may use to determine the intra-prediction mode of the PU.
[00117] In some cases, the syntax factor may indicate that the intra-prediction unit 164 should use another PU's intra-prediction mode to determine the current PU's intra-prediction mode. For example, the intra-prediction mode of the current PU may be the same as the intra-prediction mode of the adjacent PU. In other words, the intra-prediction mode of the adjacent PU can currently be the most probable mode for the PU. Thus, in this example, the bitstream may contain small syntax elements that indicate that the intra-prediction mode of the PU is the same as the intra-prediction mode of the adjacent PU. The intra-prediction unit 164 can then use the intra-prediction mode to generate PU prediction data (eg, prediction samples) based on video blocks of spatially adjacent PUs.
[00118] Reconstruction unit 158, when applicable, uses the residual video block associated with the TU of the CU and the predicted video block of the PU of the CU, for example, either intra-predicted data or inter-predicted data. Can be used to reconstruct the video block of the CU. Thus, the video decoder 30 may generate a predictive video block and a residual video block based on the syntax elements in the bitstream, and generate a video block based on the predictive video block and the residual video block. obtain.
[00119] After the reconfiguration unit 158 reconstructs the video block of the CU, the filter unit 159 may perform a deblocking operation to reduce the blocking artifacts associated with the CU. After the filter unit 159 performs a deblocking operation to reduce the blocking artifacts associated with the CU, the video decoder 30 may store the CU's video block in the decrypted picture buffer 160. The decoded picture buffer 160 may provide a reference picture for subsequent motion compensation, intra-prediction, and presentation on a display device such as the display device 32 of FIG. For example, the video decoder 30 may perform an intra-prediction operation or an inter-prediction operation on the PU of another CU based on the video block in the decoding picture buffer 160.
Scalable Video Coding (SVC) and Bit Depth Scaling [00120] As explained above, scalable video coding (SVC) provides quality scalability (eg, signal-to-noise ratio (SNR) scalability, spatial scalability, time scalability, bit depth scalability, color gamut (color). gamut) can be used to provide scalability, or dynamic range scalability). The enhanced layer may include a sample with a higher bit depth than the corresponding base layer sample. For example, a sample in the enhancement layer can have a bit depth of 10 bits, while a corresponding sample in the base layer can have a bit depth of 8 bits. Each additional bit added to the bit depth of the sample doubles the number of discrete values that the sample can represent. Therefore, the number of discrete values that can be represented by a 10-bit sample is four times greater than the number of discrete values that can be represented by an 8-bit sample. Of course, the base layer sample can have a bit depth other than 8 bits, and the enhancement layer sample can have a bit depth other than 10 bits. For luminance samples, the additional bit depth in the enhancement layer allows coding for high dynamic range (HDR) video and supports increased contrast between the darkest and brightest parts of the video image. For chrominance samples, the additional bit depth in the enhancement layer allows coding of video with a wider range of colors.
[00121] Some implementations of the SVC may include prediction of samples or blocks in the enhancement layer based on the samples or blocks in the base layer. This type of prediction is sometimes referred to as inter-layer prediction, which can be utilized in SVC to reduce inter-layer redundancy. Some examples of inter-layer prediction may be inter-layer intra-layer prediction, inter-layer motion prediction, and inter-layer residual prediction. Interlayer intra-layer prediction uses the reconstruction of the corresponding block or sample in the base layer to predict the block or sample in the enhancement layer. Inter-layer motion prediction uses motion information in the base layer to predict motion information in the enhancement layer. Inter-layer residual prediction uses the residuals of the base layer to predict the residuals of the enhancement layer.
Interlayer prediction can be used in accordance with aspects of the present disclosure to predict higher bit depth samples in the enhancement layer using lower bit depth samples in the base layer. In some cases, the enhancement layer sample can be predicted from the base layer sample by simple operations such as constant multiplication or left bit shift. Left bit shift is equivalent to multiplication by 2, with the addition of one or more bits to the end of the base layer sample, thereby increasing the bit depth of the base layer sample. While this kind of simple operation may be sufficient in some cases, they may not give good results in other applications.
[00123] The usefulness of a simple operation for predicting an enhancement layer sample from a base layer sample depends on the relationship of the sample representation used by each layer. A simple operation is, for example, when the enhancement layer sample represents a different chromaticity component than the base layer sample, for example, the base layer sample is represented according to BT.709 and the enhancement layer sample is represented according to BT.2020. (Both BT.709 and BT.2020 are defined by the ITU-R, the International Telecommunication Union-Wireless Communications Sector), which can give inadequate predictions. A simple operation also means that the base layer sample represents a luminous value with a different gamma nonlinearity than the enhancement layer sample, or a sample in one layer represents a luminous value on a linear scale, but in another layer. Samples can give inadequate predictions when representing samples on a non-linear scale. The term non-linear scale as used herein has its general meaning, and is a scale that is partly linear and partly non-linear, a scale composed of several different linear components, And their equivalents.
[00124] In some embodiments, the enhancement layer sample can be predicted based on the base layer sample by using a lookup table that maps each possible base layer sample value to the corresponding enhancement layer sample value. ..
[00125] Predicted quality can depend not only on the range of possible chrominance and luminance values that can be represented by the samples in each layer, but also on the distribution of actual sample values in the particular video being coded. For example, the sample representation used by the enhancement layer may be linearly mapped to the sample representation used by the base layer, but the actual sample distribution may not be uniformly diffused over the full range of possible values. is there. Rather, the actual sample can be clustered into several parts of the scale. In this situation, better results can be obtained by skewing the prediction towards the part of the scale where the sample is clustered.
[00126] Embodiments according to aspects of the present disclosure provide advantages for inter-layer prediction in scalable video coding, including inter-layer prediction with heterogeneous sample representations and non-uniform distribution of sample values. Specific embodiments will be disclosed in more detail below with respect to the accompanying figures.
[00127] FIG. 4 is a block diagram illustrating an example of a scalable video encoder that may implement the technique according to aspects of the present disclosure. The video encoder 400 of FIG. 4 may correspond to the video encoder 20 of FIGS. 1 and 2. However, the encoder 400 diagram in FIG. 4 does not focus more generally on aspects related to block-based video coding, but specifically on aspects related to scalable video coding and interlayer prediction.
[00128] In the example of FIG. 4, the video encoder 400 is a scalable video encoder that includes a BL subsystem 420 and an EL subsystem 440. The BL subsystem 420 encodes the video data associated with the BL, and the EL subsystem 440 encodes the video data associated with the EL. The encoded video data generated by the BL subsystem 420 can be decoded alone to produce a reconstructed video with base level quality. The coded video data generated by the EL subsystem 440 may not be meaningfully decoded on its own, but it is with BL data to produce reconstructed video with enhanced quality. Can be decrypted in combination. In some embodiments, the video data associated with the BL is an older decoder, or a decoder that does not have sufficient computational resources to effectively decode and present a combined, higher quality video. Can be adapted to. As shown in FIG. 4, EL subsystem 440 encodes an EL that supports sample values with increased bit depth for BL encoded by BL subsystem 420. Samples with increased bit depth may allow, for example, the presentation of video with higher dynamic range or more diverse colors.
[00129] The BL subsystem 420 and the EL subsystem 440 may be implemented in hardware, in software, or in a combination of hardware and software. The BL subsystem 420 and the EL subsystem 440 are shown separately in FIG. 4 for conceptual purposes, but they may share several hardware components or software modules. For example, the entropy coding unit 428 in the BL subsystem 420 can be implemented in the same hardware component or software module as the entropy coding unit 448 in the EL subsystem 440.
[00130] The BL subsystem 420 includes a mapping unit 422, a residual calculation unit 424, an entropy coding unit 428, a reconstruction and storage unit 430, and a prediction unit 432. The EL subsystem 440 includes an inverse mapping unit 442, a residual calculation unit 444, an entropy coding unit 448, a reconstruction and storage unit 450, and a prediction unit 452. The units that are the various components of the BL subsystem 420 and the EL subsystem 440 are shown separately for conceptual purposes, but in some embodiments they are combined into a smaller number of units. Or can be subdivided into additional units. Most of the features described below are present in both BL subsystem 420 and EL subsystem 440. A detailed example covering common features shared by both subsystems has been described above for the video encoder 20 in FIG. In Figure 4, the description focuses on aspects of the video encoder 400 that allow the two subsystems to work together and produce scalable output.
[00131] During the coding process, the video encoder 400 receives the video data to be encoded. Video data received as input can be processed by both BL subsystem 420 and EL subsystem 440. In the BL subsystem 420, processing begins at mapping unit 422, where samples during video input are mapped from higher EL bit depths to lower BL bit depths. For example, the input to mapping unit 422 may include, for example, a sample representing an HDR video with a bit depth of 10, 12 or 14 bits. The output of mapping unit 422 may then include a sample representing an LDR video with a lower bit depth, such as 8 bits. The mapping unit 422 can calculate the values of the BL sample in various ways, such as by applying one or more arithmetic or mathematical functions to the EL sample. In some embodiments, the mapping unit 422 multiplicative the EL sample values in different ranges. BL samples can be calculated by applying a piecewise linear function that applies factor). In other embodiments, the mapping unit 422 may calculate the BL sample by applying a logarithmic function to the EL sample value. Moreover, mapping unit 422 can apply any functional inversion that can be used in inverse mapping unit 442, which is described in more detail below with respect to EL subsystem 440. In some embodiments, however, with the behavior applied by the mapping unit 422, except that the mapping unit 422 has the effect of reducing the bit depth and the inverse mapping unit 442 has the effect of increasing the bit depth. , There is no clear correspondence with the behavior applied by the inverse mapping unit 442. Mapping unit 422 may also apply a set of operations designed to approximate mathematical functions that may not be possible or feasible to apply accurately. In some embodiments, the mapping unit 422 may be configured to apply different arithmetic or mathematical functions to different EL samples, either separately from the values in each sample or based on criteria in addition to it. For example, the criteria can depend on the position of the sample relative to the block or frame, the values of other EL samples in the same slice, syntax information, or configuration parameters. Regardless of the specific operation applied by mapping unit 422, the EL sample of the input video slice is converted to a BL sample with a lower bit depth.
The BL sample generated by the mapping unit 422 can be used by the prediction unit 432 and the residual calculation unit 424. Prediction unit 432 can support various prediction modes, which are BL samples from those modes to determine which of several different modes produces the best prediction for a particular video slice. Can be compared with the predicted sample of. Prediction unit 432 can also compare different partitioning options, for example, by dividing the video frame into different combinations of maximum coding unit (LCU), coding unit (CU), and sub-CU. In some embodiments, various categories and predictability can be assessed using rate strain analysis. The classification and mode selection process applied by Prediction Unit 432 may match that used by Prediction Unit 100 in FIG. Examples of predictions that can be made by the prediction unit 432 have been described above in more detail with respect to the motion estimation unit 122, the motion compensation unit 124, and the intra prediction unit 126 of FIG.
[00133] In the residual calculation unit 424, the encoder 400 calculates the difference between the actual BL sample determined by the mapping unit 422 and the prediction sample generated from the video slice previously processed by the prediction unit 432. The difference between the actual BL sample and the corresponding predicted sample is sometimes referred to as the residual sample. Similarly, the difference between the actual block of the sample and the corresponding predicted block is sometimes referred to as the residual block. The residuals from the residual calculation unit 424 can be converted from the sample domain to an alternative domain such as the frequency domain. The resulting conversion factor can be quantized before being encoded by the entropy coding unit 428. The entropy coding unit 428 also encodes the syntax data from the prediction unit 432. This syntax data describes the divisions and predictions in which the quantized transformation coefficients were used to calculate the residuals derived from it. The output of the entropy coding unit 428 is a coded BL video that becomes part of the scalable video bitstream produced by the encoder 400. More detailed examples of conversion, quantization, and entropy coding are given above for the conversion unit 104, quantization unit 106, and entropy coding unit 116 of FIG.
[00134] In the reconstruction and storage unit 430, the transformation and quantization operations are inverted to reconstruct the residual values in the sample region. The reconstructed residual value can be combined with the prediction sample used to determine the original residual value prior to transformation and quantization. The combination of the reconstructed residuals and the corresponding prediction results in a reconstructed video slice. The reconstructed video slice may contain, for example, the strain introduced by the coding process during transformation and quantization. A more detailed example of the reconstruction process has been described above for the inverse quantization unit 108, the inverse transformation unit 110, and the reconstruction unit 112 of FIG.
[00135] Reconstruction and storage unit 430 may include memory for storing video data from the reconstructed video slice. The reconstructed video data stored in memory can be used as the basis for future rounds of prediction in BL subsystem 420 or EL subsystem 440. The encoder 400 takes into account the distortion introduced in the coding process and ensures that the predictions made by the encoder can be reproduced using the data available to the decoder. Make predictions based on the reconstructed data (rather than the original data generated by mapping unit 422). The prediction can be made, for example, by the prediction unit 432 or by the inverse mapping unit 442. For example, the reconstructed video data may include reference frames and the prediction unit 432 may use interframe prediction to predict subsequent frames. The reconstructed video data may also include reference blocks, and the prediction unit 432 may use intra-frame prediction to predict adjacent blocks. Encoder 400 may also use inverse mapping unit 442 to make inter-layer predictions, as described below for EL subsystem 440. A more detailed example of the storage process is given above for the decrypted picture buffer 114 of FIG. 2, and a more detailed example of the various prediction schemes is given in FIG. 2 for motion estimation unit 112, motion compensation unit 124, and intra. The prediction unit 126 has been described.
[00136] As previously described, the encoder 400 includes a BL subsystem 420 and an EL subsystem 440. The output of the BL subsystem 420 is sufficient on its own to produce a video output with base level quality. The output of EL subsystem 440, on the other hand, contains only the information needed to improve the quality of the rendered video from the base level quality associated with BL to the enhanced level quality associated with EL. .. Moreover, the output of the EL subsystem may not directly represent the difference between the BL video and the EL video. Instead, it can represent the difference between the actual EL video and some predicted version of the EL video derived from the BL video. Therefore, the output produced by the EL subsystem can be highly dependent on the method adopted for interlayer prediction between BL and EL. Better prediction methods result in predictions that are closer to the actual EL video, which allows the EL subsystem 140 to produce coding with increased spatial efficiency or higher visual quality.
[00137] General considerations related to the prediction of EL samples from BL samples have been described earlier in the Morphology section for practicing the present invention prior to the description in FIG. As described above, simple operations such as constant multiplication or left bit shift with a fixed number of bits can be used to predict an EL sample from a BL sample with a lower bit depth. These simple operations can be useful as they provide an easy way to increase the bit depth of the BL sample to match the expected bit depth of the EL sample. However, this kind of simple operation can give poor predictive performance in situations where the scale used for the EL sample is not directly proportional to the scale used for the BL sample. In such situations, better predictions can be obtained by applying different operations to BL samples that are on different parts of the BL scale. In other words, better results can be obtained by predicting the EL sample based on adaptive adjustments to the BL sample rather than fixed or constant adjustments.
[00138] As explained above, the inter-layer prediction performance depends not only on the respective scales used for the BL and EL samples, but also on how the individual samples are distributed with respect to those scales. Also depends. Since the samples may not be uniformly distributed, it may be beneficial to employ a prediction method that can be adapted to different sample distributions. Specifically, when adaptive adjustments are used to predict EL samples from BL samples, the specific adjustment parameters selected for a particular BL sample are only to the values of a particular sample on the BL scale. Instead, it may also favorably depend on the overall distribution of BL samples and the location of specific samples with respect to that distribution. In some applications, an exhaustive analysis of the complete sample distribution may not be computationally feasible, but heuristics are used to determine adaptive parameters that take into account at least some of the variances in the sample distribution. Can be used.
[00139] As shown in EL subsystem 440, the inverse mapping unit 442 can be used for interlayer prediction. More specifically, the inverse mapping unit 442 may make inter-layer predictions by applying the types of adaptive adjustments described above. For example, the inverse mapping unit 442 may multiply the reconstructed BL sample by a specific ratio to determine the predicted EL sample. The particular ratio is adaptive by the inverse mapping unit 442, for example, depending on the value of the BL sample, as well as one or more heuristics related to the overall distribution of the BL sample in the video slice to which the reconstructed sample belongs. Can be selected for.
The inverse mapping unit 442 may be configured to select a set of adaptive adjustment parameters that minimizes errors in the prediction EL sample. The error in the predicted sample can be measured, for example, by the mean of the signed differences between the predicted sample and the actual sample in the video slice, the mean of the absolute differences, or the mean of the squared differences. A computationally efficient approximation of any of these averages can also be used. In some embodiments, the inverse mapping unit 442 may use rate strain analysis to select adjustment parameters and minimize errors in the prediction sample. Once the inverse mapping unit 442 determines the adaptive adjustment parameters, they can be sent to the entropy coding unit 448 in the form of syntax data, which is used to make interlayer predictions while decoding the EL video data. For example, it can be used by the decoder 500 of FIG.
[00141] The EL subsystem 440 functions in a similar manner to the BL subsystem 420, except for the use of interlayer prediction. The inter-layer prediction function provided by the inverse mapping unit 442 is the same mode provided by the prediction unit 432 of the BL subsystem 420, as an additional prediction mode that complements the intra-picture and inter-picture prediction modes provided by the prediction unit 452. work. Prediction unit 452 may make the mode selection previously described for prediction unit 432 to select the optimal prediction mode for various video slices. Mode selection has also been described above in more detail with respect to the prediction unit 100 of FIG.
[00142] The remaining units of EL subsystem 440 function in the same manner as the corresponding units of BL subsystem 420, except that they operate for samples with larger bit depths. Therefore, the residual calculation unit 444 calculates the residual representing the difference between the predicted EL sample and the actual EL sample. The residuals can be transformed into alternative regions and the resulting coefficients can be quantized with the syntax data from the prediction unit 452 before being encoded by the entropy coding unit 448. In the reconstruction and storage unit 450, the quantized transformation factors can be dequantized, inversely transformed and combined with the predictions from the prediction unit 452 to form the reconstructed video slices. The reconstructed video slice can be used by the prediction unit 452 as the basis for additional rounds of prediction. The output from the EL subsystem 440 can be combined with the output from the BL subsystem 420 to form a coded wearable bitstream that is the output of the encoder 400.
[00143] FIG. 5 is a block diagram showing an example of a scalable video decoder that may implement the technique according to aspects of the present disclosure. The video decoder 500 of FIG. 5 may correspond to the video decoder 30 of FIGS. 1 and 3. However, the Decoder 500 diagram in FIG. 5 does not focus more generally on aspects related to block-based video coding, but specifically on aspects related to scalable video coding and interlayer prediction.
[00144] In the example of FIG. 5, the video decoder 500 is a scalable video decoder that includes a BL subsystem 520 and an EL subsystem 540. The video decoder 500 may perform a decoding process that is generally the opposite of the coding process performed by the video encoder 400 as described in FIG. The decoder 500 may receive as input a coded scalable bitstream with video that encodes both EL and BL. The BL subsystem 520 may decode the video data associated with the BL and the EL subsystem 540 may decode the video data associated with the EL. As shown in FIG. 5, the output of the decoder 500 may include a decoding BL bitstream and a decoding EL bitstream. In some embodiments, the decoder 500 provides output in only one of the BL or EL formats. For example, if the decoder 500 is a legacy decoder that does not support the higher bit depth video associated with the EL, it may only contain the BL subsystem 520, in which case the EL portion of the encoded scalable bitstream will be ignored and the BL Only the output will be given. Alternatively, the decoder 500 supports higher bit depth EL video, but uses the BL bitstream generated by the BL subsystem 520 internally for interlayer prediction only and outputs it in EL format only. obtain.
[00145] The BL subsystem 520 and EL subsystem 540 may be implemented in hardware, in software, or in a combination of hardware and software. The BL subsystem 520 and the EL subsystem 540 are shown separately in Figure 5 for conceptual purposes, but they may share several hardware components or software modules. For example, the entropy coding unit 524 in the BL subsystem 520 may be implemented in the same hardware component or software module as the entropy coding unit 546 in the EL subsystem 540.
[00146] The BL subsystem 520 includes a BL extraction unit 522, an entropy decoding unit 524, a prediction unit 526, and a reconstruction and storage unit 528. The BL extraction unit 522 receives encoded scalable video information including both EL video data and BL video data as input. The BL extraction unit 522 extracts a BL portion of the data that comprises a coded video sample with a certain bit depth corresponding to the base level video quality. The EL portion of the data, which provides the additional information needed to derive extended video samples with higher bit depth, may not be used within BL subsystem 520.
[00147] When the BL data is extracted from the scalable bitstream, it is entropy-decoded by the entropy decoding unit 524, resulting in the syntax data as well as the quantized conversion coefficients representing the residual video information, eg, in the frequency domain. In the prediction unit 526, the syntax data is used to generate a prediction video block or prediction video frame, for example, by intra-frame (spatial) prediction or inter-frame (motion) prediction. In the reconstruction and storage unit 528, the quantized transformation coefficients are inversely quantized and inversely transformed to produce residual information in the sample region. The residual information is added to the prediction generated by the prediction unit 526 to yield a reconstructed video frame or picture containing a video block consisting of the reconstructed BL video sample. The series of these reconstructed video frames constitutes the decoded BL video, which is the output of BL subsystem 520. The reconstructed video frames and blocks can then be used by the prediction unit 526 to make additional rounds of prediction. More detailed examples of the processes performed by (1) entropy decoding unit 524, (2) prediction unit 526, and (3) reconstruction and storage unit 528 are shown in FIG. 3, respectively, (1) entropy decoding unit 150, ( 2) motion compensation unit 162 and intra prediction unit 164, and (3) inverse quantization unit 154, inverse conversion unit 156, reconstruction unit 158, and decoding picture buffer 160 are given above.
[00148] As previously described, the decoder 500 includes a BL subsystem 520 and an EL subsystem 540. The EL subsystem 540 produces an extended decoded video by combining the extended information from the EL portion of the encoded scalable bitstream with the predictions generated from the decoded BL video generated by the BL subsystem 520. The EL subsystem 540 includes an inverse mapping unit 542, an EL extraction unit 544, an entropy decoding unit 546, a prediction unit 548, and a reconstruction and storage unit 550. The EL extraction unit 544 extracts EL data from the encoded scalable bitstream received as input to the EL subsystem 540. In the entropy decoding unit 546, the extracted EL data is entropy-decoded to produce quantized conversion coefficients representing the syntax data and, for example, the residual video information in the frequency domain. Syntax data can be used to generate predictive video frames or video blocks with predictive samples. The prediction can be generated by the prediction unit 548 according to the inter-frame prediction mode, the intra-frame prediction mode, and the like. The prediction unit 548 may also use the inter-layer prediction provided by the inverse mapping unit 542 instead of or in combination with the inter-frame and intra-frame prediction described above. The syntax data provided by the entropy decoding unit 546 may specify which prediction mode should be used for each part of the EL video sequence decoded by the EL subsystem 540.
[00149] The inverse mapping unit 542 receives adaptive adjustment parameters (eg, determination or extraction) from the syntax data given by the entropy coding unit 546, rather than selecting the parameters based on the optimization calculation. ), The inter-layer prediction is performed in the same manner as the inverse mapping unit 442 of FIG. The inverse mapping unit 542 does not perform such an optimization calculation because it may not have access to the original EL sample used to create the encoded EL video data. Conversely, the inverse mapping unit 442 of FIG. 4 creates syntax data and makes it available to the inverse mapping unit 542, or a similar interlayer prediction unit in another embodiment of the decoder 500. Perform optimization calculations.
[00150] In the reconstruction and storage unit 550, the prediction from the prediction unit 548 (possibly with interlayer prediction from the inverse mapping unit 542) is combined with the residuals in the sample area. The reconstruction and storage unit 550 determines the residuals by dequantizing and inverse transforming the quantized conversion factors from the entropy decoding unit 546. The combination of residuals and predictions yields a reconstructed EL video with decoded EL pictures or frames, which is the final output of the EL subsystem 540 and decoder 500 as shown in FIG. More detailed examples of certain features of the decoder 500 are given above for the video decoder 30 of FIGS. 1 and 3.
[00151] Next, referring to FIG. 6, a flowchart showing an exemplary method for determining the prediction of the EL sample from the BL sample is given. Method 600 of FIG. 6 is particularly suitable for predicting EL samples with higher bit depths than the corresponding BL samples. In some embodiments, the EL comprises a high dynamic range sample capable of representing a larger range of luminance values than the corresponding low dynamic range sample in BL. In another embodiment, the EL comprises a chrominance sample capable of representing a wider range of colors than the corresponding chrominance sample in BL. If Method 600 is implemented in a video coder (eg, encoder or decoder) that supports more than one EL and one BL, the steps in Method 600 are interleaved or interleaved with other interlayer prediction methods. Can be done at the same time. For example, if Method 600 is implemented in a video coder (eg, encoder or decoder) that supports both bit depth and spatial resolution scaling, any interlayer prediction associated with spatial scaling (such as upsampling). Steps can be performed before, during, or after Method 600.
[00152] For simplicity, the description of Method 600 focuses on BL and EL samples that represent intensity values or luminance values, sometimes referred to as intensities. However, those skilled in the art of video coding will appreciate that the techniques of the present disclosure performed in Method 600 can be similarly applied to samples measuring chrominance or other aspects of video pictures. In addition, the description of Method 600 may refer to the distribution of relative intensities between a small number of spatially adjacent samples, such as three samples arranged horizontally, vertically, or diagonally within a single video block. Mention the pattern. However, the techniques of the present disclosure include distributions with chrominance samples, distributions of three or more samples, distributions of samples that do not form a single line, distributions of samples in two or more blocks, random numbers from larger sets. Or to select and apply adaptive adjustments based on other types of sample distributions, such as distributions with a set of samples selected to statistically represent a larger set of samples, such as by pseudo-random selection. Can be used for.
[00153] The steps shown in FIG. 6 are an encoder (eg, the video encoder shown in FIG. 2 or 4), a decoder (eg, the video decoder shown in FIG. 3 or 5,), or any of them. It can be done by other components. For convenience, the steps are described as being performed by a coder, which can be an encoder, decoder or other component.
[00154] Method 600 starts at block 601. At block 605, the coder determines an intensity category and a pattern category for each sample in BL. In some embodiments, the range of possible intensities that can be represented by the BL sample can be divided into multiple bands. For example, if the BL sample has a bit depth of 8 bits, it may represent an intensity value in the range 0-255. This range can be divided into four bands of equal size, corresponding to ranges 0-63, 64-127, 128-191, and 192-255. If such a band is used, the intensity category associated with the sample may correspond to the intensity band in which the sample is located. In some embodiments, the defined band may not fill the entire range of possible sample values, so additional categories may be needed for samples that are outside the defined band. In a further embodiment, the band may not be uniformly sized. For example, the band is the center of the intensity spectrum to allow finer tuning of the midrange samples, where the overall sample distribution can be particularly advantageous for videos that contain most midrange samples. May be smaller near. The intensity category can be determined based on the luminance value of the BL sample, the chrominance value (s) of the BL sample, or the combination of the luminance and chrominance values of the BL sample.
[00155] The pattern category may be based, for example, on a categorized BL sample and a plurality of samples adjacent to the categorized sample. For example, an adjacent sample may include one sample on the left side of the categorized sample and one sample on the right side of the categorized sample. If the intensities of both the right and left samples are greater than the intensities of the categorized sample, the first category can be assigned and other combinations of relative intensities between the categorized sample and its neighbors. Can be assigned to other categories. In some embodiments, the pattern category may be determined on the basis of neighboring samples other than the left and right samples, such as upper or lower samples, or samples located diagonally to the categorized sample. .. As mentioned above, in some embodiments, three or more adjacent samples may be considered, and in other embodiments, the considered samples may not be adjacent. The pattern category is sometimes called the distribution category.
[00156] Method 600 proceeds to block 610 and applies a preliminary mapping to each of the BL samples categorized in block 302. In some embodiments, the preliminary mapping does not take into account the determined category. Preliminary mapping applies a mathematical function or set of computational operations to a BL sample to determine a preliminary prediction of the corresponding EL sample. In some embodiments, the preliminary mapping may have the effect of increasing the bit depth of the BL sample to the required bit depth of the EL sample. In some embodiments, the preliminary mapping may make coarser adjustments for block 306 than the adaptive adjustments described below. For example, preliminary mapping can involve exponentiation or multiplication, while adaptive adjustment can involve multiplication or addition. In some embodiments, the preliminary mapping utilizes a lookup table that maps each BL sample value (or set of values) to the corresponding EL sample value (or set of values). In some embodiments, the preliminary mapping can be skipped altogether and the BL sample itself can replace the preliminary prediction in a later step of Method 600. Additional examples of operations that may be used for preliminary mapping have been described above in the Modes section for practicing the present invention, along with relevant considerations for selection between such operations.
[00157] At block 615, the preliminary prediction from block 610 is improved by adjustment operations using adaptive adjustment parameters. One adjustment operation may be performed for each of the categories determined in block 605. For example, if one intensity category and one pattern category are determined for a particular BL sample, two adjustment operations can be applied to the preliminary predictions derived from that sample. Each adjustment operation uses the adaptive adjustment parameters associated with the corresponding category. In some embodiments, the adjustment operation is the same for each type of category. For example, the adjustment operation for both the intensity category and the pattern category can be multiplication, in which case the preliminary prediction will be multiplied by both the intensity and pattern adjustment parameters. Alternatively, different adjustment operations can be used for different types of categories, for example multiplication can be used to apply the adjustment parameters associated with the intensity category, and apply the adjustment parameters associated with the pattern category. Addition can be used to do so, and vice versa. As described above in the mode section for carrying out this invention, the adjustment parameters are multiplicative. It can include ratios or coefficients, additive or subtractive offsets, and so on. In addition, adjustment parameters for each category to minimize distortion or error between the predicted EL sample and the actual EL sample, as previously described for the inverse mapping units 442 and 542 in FIGS. 4 and 5. Can be selected (eg, can be determined when the EL video is encoded). In some embodiments, a single adjustment parameter is associated with that category for all samples in a particular video that fit into a single category. In other embodiments, however, different adjustment parameters can be associated with the same category for samples in different parts of the video. For example, a sample of the first block in a particular intensity band can be associated with adjustment parameter a, while a sample of a second block in the same intensity band can be associated with adjustment parameter b, where b is in a. Not equal. Different tuning parameters can be associated with a single category not only for different blocks, but also for different groups of blocks, different parts of blocks, different frames, and so on. To improve coding efficiency, tuning parameters for a particular region (eg, blocks, frames, etc.) can be predicted from regions that are temporally or spatially close together.
[00158] In block 620, when the adaptive adjustment parameters are applied to determine the improved prediction from the preliminary prediction, to the improved prediction to determine the reconstructed EL sample, which is the final product of Method 600. Add the residual values. Method 600 ends at block 625.
As described above, one or more components of the video encoder 20 of FIG. 2, the video encoder 30 of FIG. 3, the video encoder 400 of FIG. 4, or the video encoder 500 of FIG. 5 are each BL. Determining one or more categories for a sample, applying preliminary mappings applied to determine preliminary EL predictions, applying adaptive adjustments to each preliminary EL prediction, and remaining in improved EL predictions. It can be used to implement any of the techniques described in this disclosure, such as adding differences.
[00160] The information and signals disclosed herein can be represented using any of a wide variety of techniques and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description are voltages, currents, electromagnetic waves, magnetic or magnetic particles, light fields or optical particles, or any of them. It can be represented by a combination.
[00161] The various exemplary logic blocks, modules, circuits, and algorithm steps described with respect to the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination thereof. To articulate this compatibility of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above in general with respect to their functionality. Whether such functionality is implemented as hardware or software depends on specific application examples and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but decisions on such implementation should not be construed as causing a deviation from the scope of the invention.
[00162] The techniques described herein can be implemented in hardware, software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general purpose computers, wireless communication device handset, or integrated circuit devices with multiple uses, including applications in wireless communication device handset and other devices. The features described as modules or components can be implemented together in an integrated logical device or separately as a separate but interoperable logical device. When implemented in software, the technique, when implemented, is at least partially realized by a computer-readable data storage medium that contains program code containing instructions that perform one or more of the methods described above. obtain. Computer-readable data storage media may form part of a computer program product that may contain packaging material. Computer-readable media include random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), and electrically erasable programmable read-only memory (EEPROM (registration)). It may include memory or data storage media such as (trademark)), flash memory, magnetic or optical data storage media. The technique, in addition or as an alternative, carries or transmits program code in the form of instructions or data structures, such as propagated signals or radio waves, by a computer-readable communication medium that can be accessed, read, and / or executed by a computer. It can be realized at least partially.
[00163] Program code includes one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. It can be run by a processor that can include one or more processors. Such processors may be configured to perform any of the techniques described in this disclosure. The general purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. Processors can be implemented as a combination of computing devices, eg, a combination of DSP and microprocessor, multiple microprocessors, one or more microprocessors working with a DSP core, or any other such configuration. .. Thus, the term "processor" as used herein refers to any of the above structures, any combination of the above structures, or any other structure or device suitable for implementing the techniques described herein. .. Further, in some embodiments, the functionality described herein is provided within a dedicated software or hardware module configured for coding and decoding, or a composite video encoder / decoder (codec). Can be incorporated into. The technique can also be fully implemented in one or more circuits or logic elements.
[00164] The techniques of the present disclosure can be implemented in a wide variety of devices or devices, including wireless handsets, integrated circuits (ICs) or sets of ICs (eg, chipsets). Although various components, modules, or units have been described herein to emphasize the functional aspects of devices configured to perform the disclosed techniques, those components, modules, or units are not necessarily included. It does not necessarily require implementation by different hardware units. Rather, as described above, the various units can be combined or interacted with in the codec hardware unit, including the one or more processors described above, along with suitable software and / or firmware. Can be given by a set of hardware units.
[00165] Various embodiments of the present invention have been described. These and other embodiments fall within the scope of the following claims.<u style="single"> The inventions described in the claims at the time of filing the application of the present application are described below.</u><u style="single">[C1] A device configured to code video data, said device.</u><u style="single">A memory unit configured to store video data associated with a base layer and an enhancement layer, said base layer comprising a base layer sample having a first bit depth.</u><u style="single">The processor communicating with the memory unit and the processor</u><u style="single">Applying a function to the base layer sample to generate a preliminary prediction sample,</u><u style="single">Determining one or more categories associated with the base layer sample</u><u style="single">Determining the adjustment parameters for each of the one or more categories associated with the base layer sample.</u><u style="single">In order to determine an improved prediction sample from the preliminary prediction sample, an adjustment calculation is performed using each of the determined adjustment parameters, and a second in which the improved prediction sample is larger than the first bit depth. A device that is configured to have a bit depth of.</u><u style="single">[C2] The device according to C1, wherein the first bit depth is 8 bits and the second bit depth is one of 10, 12, and 14 bits.</u><u style="single">[C3] The device according to C1, wherein at least one of the one or more categories associated with the base layer sample is determined from one or more luminance or chrominance values of the base layer sample. ..</u><u style="single">[C4] At least one of the one or more categories associated with the base layer sample is one or more luminance or chrominance values of the base layer sample and at least one from the received video data. The device according to C1, which is determined from one or more luminance or chrominance values of one or more other samples.</u><u style="single">[C5] The base layer sample represents a luminance value on a scale of possible luminance values.</u><u style="single">The scale of possible luminance values is divided into multiple luminance bands.</u><u style="single">The luminance value represented by the base layer sample is within one of the luminance bands.</u><u style="single">One of the categories associated with the base layer sample is the apparatus according to C1, wherein the base layer sample corresponds to the luminance band within it.</u><u style="single">[C6] The base layer sample represents a chrominance value on a scale of possible chrominance values.</u><u style="single">The scale of possible chrominance values is divided into multiple chrominance bands.</u><u style="single">The chrominance value represented by the base layer sample is within one of the chrominance bands.</u><u style="single">One of the categories associated with the base layer sample is the device of C1, wherein the base layer sample corresponds to the chrominance band within it.</u><u style="single">[C7] One of the categories associated with the base layer sample is a plurality of comparisons between the base layer sample and other samples spatially adjacent to the base layer sample in the video data. The device according to C1, which depends on the results of.</u><u style="single">[C8] The device of C1, wherein the function comprises at least one logarithmic or exponential operation.</u><u style="single">[C9] The device of C1, wherein the function comprises a left bit shift, or multiplication by a number greater than or equal to 2.</u><u style="single">[C10] The device of C1, wherein the function comprises a look-up table that maps each possible base layer sample value to the corresponding enhancement layer sample value.</u><u style="single">[C11] The device of C1, wherein the preliminary prediction sample has a bit depth equal to the second bit depth.</u><u style="single">[C12] The device of C1, wherein the adjustment parameter has a base of ratio, coefficient, index, or logarithm.</u><u style="single">[C13] The device of C1, wherein the adjustment calculation comprises addition, subtraction, multiplication, division, exponentiation, or logarithm.</u><u style="single">[C14] The apparatus of C1, wherein the processor is further configured to add a residual value to the improved prediction sample to determine an enhancement layer sample.</u><u style="single">[C15] A method of coding video data, wherein the method is</u><u style="single">Receiving said video data with a base layer sample having a first bit depth,</u><u style="single">Applying a function to the base layer sample to generate a preliminary prediction sample,</u><u style="single">Determining one or more categories associated with the base layer sample</u><u style="single">Determining the adjustment parameters for each of the one or more categories associated with the base layer sample.</u><u style="single">In order to determine an improved prediction sample from the preliminary prediction sample, an adjustment calculation is performed using each of the determined adjustment parameters, and a second in which the improved prediction sample is larger than the first bit depth. A method of having a bit depth of.</u><u style="single">[C16] The method according to C15, wherein the first bit depth is 8 bits and the second bit depth is one of 10, 12, and 14 bits.</u><u style="single">[C17] The method of C15, wherein at least one of the one or more categories associated with the base layer sample is determined from one or more luminance or chrominance values of the base layer sample. ..</u><u style="single">[C18] At least one of the one or more categories associated with the base layer sample is one or more luminance or chrominance values of the base layer sample and at least one of the received video data. The method according to C15, which is determined from one or more luminance or chrominance values of one other sample.</u><u style="single">[C19] The base layer sample represents a scale 1 luminance value of possible luminance values.</u><u style="single">The scale of possible luminance values is divided into multiple luminance bands.</u><u style="single">The luminance value represented by the base layer sample is within one of the luminance bands.</u><u style="single">One of the categories associated with the base layer sample is the method of C15, wherein the base layer sample corresponds to the luminance band within it.</u><u style="single">[C20] The base layer sample represents a chrominance value on a scale of possible chrominance values.</u><u style="single">The scale of possible chrominance values is divided into multiple chrominance bands.</u><u style="single">The chrominance value represented by the base layer sample is within one of the chrominance bands.</u><u style="single">One of the categories associated with the base layer sample is the method of C15, wherein the base layer sample corresponds to the chrominance band within it.</u><u style="single">[C21] One of the categories associated with the base layer sample is a plurality of comparisons between the base layer sample and other samples spatially adjacent to the base layer sample in the video data. The method described in C15, which depends on the results of.</u><u style="single">[C22] The method of C15, wherein the function comprises at least one logarithmic or exponential operation.</u><u style="single">[C23] The method of C15, wherein the function comprises a left bit shift, or multiplication by a number greater than or equal to 2.</u><u style="single">[C24] The method of C15, wherein the function comprises a look-up table that maps each possible base layer sample value to the corresponding enhancement layer sample value.</u><u style="single">[C25] The method of C15, wherein the preliminary prediction sample has a bit depth equal to the second bit depth.</u><u style="single">[C26] The method of C15, wherein the adjustment parameter has a base of ratio, coefficient, exponent, or logarithm.</u><u style="single">[C27] The method of C15, wherein the adjustment calculation comprises addition, subtraction, multiplication, division, exponentiation, or logarithm.</u><u style="single">[C28] The method of C15, wherein the processor is further configured to add a residual value to the improved prediction sample to determine an enhancement layer sample.</u><u style="single">[C29] When executed</u><u style="single">Receiving video data with a base layer sample with a first bit depth,</u><u style="single">Applying a function to the base layer sample to generate a preliminary prediction sample,</u><u style="single">Determining one or more categories associated with the base layer sample</u><u style="single">Determining the adjustment parameters for each of the one or more categories associated with the base layer sample.</u><u style="single">In order to determine an improved predicted sample from the preliminary predicted sample, an adjustment calculation is performed using each of the determined adjustment parameters, and a second in which the improved predicted sample is larger than the first bit depth. A non-transitory computer-readable medium with a code that causes the device to have a bit depth of.</u><u style="single">[C30] A video coding device configured to code video data, said video coding device.</u><u style="single">A means for receiving said video data comprising a base layer sample having a first bit depth, and</u><u style="single">A means for applying a function to the base layer sample to generate a preliminary prediction sample, and</u><u style="single">A means for determining one or more categories associated with the base layer sample, and</u><u style="single">Means for determining the adjustment parameters corresponding to each of the one or more categories associated with the base layer sample, and</u><u style="single">A means for performing an adjustment calculation using each of the determined adjustment parameters to determine an improved prediction sample from the preliminary prediction sample, and the improved prediction sample is greater than the first bit depth. A video coding device, comprising, having a second bit depth.</u>
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
Every citation, both waysCites: the store holds 4 of 5
| Document | Relation | Office |
|---|---|---|
| US20090175338A1 | Cites | United States of America |
| US20110293003A1 | Cites | United States of America |
| JP2011509536A | Cites | Japan |
| JP2011501571A | Cites | Japan |
| Andrew Segall, et al.,”System for Bit-Depth Scalable Coding”,Joint Video Team (JVT) of ISO/IEC MPEG & ITU-T VCEG (ISO/IEC JTC1/SC29/WG11 and ITU-T SG16 Q.6) 23rd Meeting: San Jose, California, USA, 21.27 April, 2007,SW,ITU-T,2007年 4月27日,JVT-W113 | Non-patent | – |
| Chulkeun Kim et al.,”Description of scalable video coding technology proposal by LG Electronics and MediaTek (differential coding mode on)”,Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 11th Meeting: Shanghai, CN, 10-19 Oct., 2012,2012年10月19日,JCTVC-K0033 | Non-patent | – |
| Jill Boyce et al.,Description of low complexity scalable video coding technology proposal by Vidyo and Samsung,Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 11th Meeting: Shanghai, CN, 10-19 Oct., 2012,2012年10月19日,JCTVC-K0045,10-12頁 | Non-patent | – |
| Wei Pu et al.,High Frequency Pass Inter Layer Sample Adaptive Offset Filter[online],Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 12th Meeting: Geneva, CH, 14-23 Jan. 2013,2013年11月 6日,JCTVC-L0234 | Non-patent | – |
14 members in 7 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 201261746906 | United States of America | P | |
| 201261746906 | United States of America | P | |
| 61746906 | United States of America | – | |
| 14137031 | United States of America | – | |
| 201314137031 | United States of America | A | |
| 201314137031 | United States of America | A | |
| 2013077473 | United States of America | W | |
| 2013077473 | United States of America | W | |
| 14137031 | – | – | – |
| 61746906 | – | – | – |
| US201261746906P | – | – | – |
| US2013077473 | – | – | – |
| US201314137031 | – | – | – |
| WO2013US77473 | – | – | – |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| US2014185664A1 | United States of America | A1 | |
| WO2014105818A2 | World Intellectual Property Organization (WIPO) | A2 | |
| KR20150103065A | Republic of Korea | A | |
| WO2014105818A3 | World Intellectual Property Organization (WIPO) | A3 | |
| CN105025997A | China | A | |
| EP2939426A2 | European Patent Office (EPO) | A2 | |
| JP2016506684A | Japan | A | |
| US9532057B2 | United States of America | B2 | |
| KR101771336B1 | Republic of Korea | B1 | |
| JP6235040B2This record | Japan | B2 | |
| CN105025997B | China | B | |
| EP2939426B1 | European Patent Office (EPO) | B1 | |
| EP2939426C0 | European Patent Office (EPO) | C0 | |
| ES2987671T3 | Spain | T3 |
18 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on accelerated examinationJAPANESE INTERMEDIATE CODE: A971005A975 | A975 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Explanation of circumstances concerning accelerated examinationJAPANESE INTERMEDIATE CODE: A871A871 | A871 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6235040
- Publication, DOCDB
- 6235040
- Publication, EPODOC
- JP6235040B
- Application
- 2015550731
- Application, DOCDB
- 2015550731
- Application, EPODOC
- JP20150550731
Titles2
- Japanese
- ビット深度スケーラブルビデオコーディングのためのサンプル適応調整を使用するレイヤ間予測
- English
- Inter-Layer Prediction Using Sample Adaptation for Bit Depth Scalable Video Coding
Classification
- CPC, 7
- H04N19/187
- H04N19/103
- H04N19/50
- H04N19/196
- H04N19/136
- H04N19/182
- H04N19/33
- IPC, 3
- H04N19 36
- H04N19 117
- H04N19 14
