Adaptive coefficient scanning for video coding
Abstract
The present disclosure describes techniques for encoding video block header information. In particular, the techniques of the present disclosure include a unidirectional prediction mode and a multidirectional prediction mode that combines at least two unidirectional prediction modes for use in generating prediction blocks of video blocks of coding units. , Select one of multiple prediction modes. The encoding device encodes the prediction mode of the current video block based on the prediction mode of the previously encoded video block of one or more encoding units. Similarly, the decoding unit receives the encoded video data of the video block of the coding unit and of the video block based on the prediction mode of the video block decoded before one or more of the coding units. The encoded video data is decoded to identify one of a plurality of prediction modes for use in generating the prediction block.

Term
1.7 yearsto projected expiry
Projected expiry 12 June 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
104 claims: 13 independent, 91 dependent
- 1ビデオデータをエンコードする方法であって、 符号化単位のビデオブロックの予測ブロックの生成に使用するために、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを選択することと、 前記符号化単位の1つまたは複数の前にエンコードされたビデオブロックの予測モードに基づいて、前記現在のビデオブロックの前記予測モードをエンコードすることと、 を備える方法。
- 2前記符号化単位の前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードのエンコードに使用するために、複数の符号化コンテキストのうちの1つを選択することをさらに備え、エンコードすることが前記選択された符号化コンテキストに従ってエンコードすることを備える請求項1に記載の方法。
- 3前記符号化コンテキストのうちの1つを選択することが、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには、第1の符号化コンテキストを選択することと、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには、第2の符号化コンテキストを選択することと、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには、第3の符号化コンテキストを選択することと、 を備える請求項2に記載の方法。
- 4残差ブロックを形成するために、前記選択された予測モードを使用して生成された前記予測モードを、前記ビデオブロックから取り除くことと、 前記選択された予測モードに基づいて、前記残差ブロックに適用するための変換を選択することと、 残差変換係数を生成するために、前記選択された変換を前記残差ブロックに適用することと、 をさらに備える請求項1に記載の方法。
- 5前記残差ブロックに適用するための前記変換を選択することが、 前記選択された予測モードが限定された方向性を示すときには、前記残差ブロックに適用するために、離散的コサイン変換(DCT)および整数変換のうちの一方を選択することと、 前記選択された予測モードが方向性を示すときには、前記残差ブロックに適用するために、方向変換を選択することと、 を備える請求項4に記載の方法。
- 6前記DCTおよび前記整数変換のうちの一方を選択することが、前記選択された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードであるときには、前記残差ブロックに適用するために、前記DCTおよび前記整数変換のうちの一方を選択することを備える請求項5に記載の方法。
- 7複数の方向変換を格納することをさらに備え、前記複数の方向変換のそれぞれが、方向性を示す前記予測モードのうちの1つに対応し、前記方向変換を選択することが、前記選択された予測モードに対応する前記複数の方向変換のうちの前記1つを選択することを備える請求項5に記載の方法。
- 8前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納することをさらに備え、前記方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項4に記載の方法。
- 9前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納することをさらに備え、前記複数の方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、Nが前記ビデオブロックの大きさである請求項4に記載の方法。
- 10前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項1に記載の方法。
- 11前記符号化単位の1つまたは複数の前にエンコードされたビデオブロックの予測モードに基づいて、前記現在のビデオブロックの前記予測モードをエンコードすることが、 前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちの1つと同じであることを示すために、前記予測モードを表す第1のビットをエンコードすることと、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードが互いに同じではないとき、前記1つまたは複数の前にエンコードされたビデオブロックのうちのどれが前記ビデオブロックの前記予測モードと同じ予測モードを有するかを示すために、前記予測モードを表す少なくとも1つの追加ビットをエンコードすることと、 を備える請求項1に記載の方法。
- 12前記符号化単位の1つまたは複数の前にエンコードされたビデオブロックの予測モードに基づいて、前記現在のビデオブロックの前記予測モードをエンコードすることが、 前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちのいずれとも同じではないことを示すために、前記予測モードを表す第1のビットをエンコードすることと、 前記複数の予測モードから前記1つまたは複数の前にエンコードされたビデオブロックの少なくとも前記予測モードを削除することと、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの予測モードではない少なくとも1つの追加の予測モードを削除することと、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列することと、 前記現在のビデオブロックの前記予測モードに対応する前記予測モード識別子を識別する符号語をエンコードすることと、 を備える請求項1に記載の方法。
- 13ビデオデータをエンコードする装置であって、 符号化単位のビデオブロックの予測ブロックの生成に使用するために、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを選択する予測ユニットと、 前記符号化単位の1つまたは複数の前にエンコードされたビデオブロックの予測モードに基づいて、前記現在のビデオブロックの前記予測モードをエンコードするエントロピエンコーディングユニットと、 を備える装置。
- 14前記エントロピエンコーディングユニットが、前記符号化単位の前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードのエンコードに使用するために、複数の符号化コンテキストのうちの1つを選択し、前記選択された符号化コンテキストに従って前記予測モードをエンコードする請求項13に記載の装置。
- 15前記エントロピエンコーディングユニットが、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには第1の符号化コンテキストを選択し、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには第2の符号化コンテキストを選択し、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには第3の符号化コンテキストを選択する請求項14に記載の装置。
- 16前記選択された予測モードに基づいて、残差ブロックに適用するための変換を選択し、残差変換係数を生成するために、前記選択された変換を前記残差ブロックに適用する変換ユニットをさらに備え、前記エントロピエンコーディングユニットが前記残差変換係数をエンコードする請求項13に記載の装置。
- 17前記変換ユニットが、前記選択された予測モードが限定された方向性を示すときには、前記残差ブロックに適用するために、離散的コサイン変換(DCT)および整数変換のうちの一方を選択し、前記選択された予測モードが方向性を示すときには、前記残差ブロックに適用するために、方向変換を選択する請求項16に記載の装置。
- 18前記選択された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードであるときには、前記変換ユニットが、前記残差ブロックに適用するために、前記DCTおよび前記整数変換のうちの一方を選択する請求項17に記載の装置。
- 19複数の方向変換を格納するメモリをさらに備え、前記複数の方向変換のそれぞれが、方向性を示す前記予測モードのうちの1つに対応し、前記変換ユニットが、前記選択された予測モードに対応する前記複数の方向変換のうちの前記1つを選択する請求項17に記載の装置。
- 20前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納するメモリをさらに備え、前記方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項16に記載の装置。
- 21前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納するメモリをさらに備え、前記複数の方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、Nが前記ビデオブロックの大きさである請求項16に記載の装置。
- 22前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項13に記載の装置。
- 23前記エントロピエンコーディングユニットが、 前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちの1つと同じであることを示すために、前記予測モードを表す第1のビットをエンコードし、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードが互いに同じではないとき、前記1つまたは複数の前にエンコードされたビデオブロックのうちのどれが前記ビデオブロックの前記予測モードと同じ予測モードを有するかを示す少なくとも1つの追加ビットをエンコードする、 請求項13に記載の装置。
- 24前記予測ユニットが、 前記複数の予測モードから前記1つまたは複数の前にエンコードされたビデオブロックの少なくとも前記予測モードを削除し、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの予測モードではない少なくとも1つの追加の予測モードを削除し、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列し、 前記エントロピエンコーディングユニットが、 前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちのいずれとも同じではないことを示すために、前記予測モードを表す第1のビットをエンコードし、 前記現在のビデオブロックの前記予測モードに対応する前記予測モード識別子を識別する符号語をエンコードする、 請求項13に記載の装置。
- 25前記装置が無線通信装置を備える請求項13に記載の装置。
- 26前記装置が集積回路装置を備える請求項13に記載の装置。
- 27ビデオ符号化装置において実行されると、前記装置にビデオブロックを符号化させる命令を備えるコンピュータ可読媒体であって、前記命令が、前記装置に、 符号化単位のビデオブロックの予測ブロックの生成に使用するために、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを選択させ、 前記符号化単位の1つまたは複数の前にエンコードされたビデオブロックの予測モードに基づいて、前記現在のビデオブロックの前記予測モードをエンコードさせる、 コンピュータ可読媒体。
- 28前記装置に、前記符号化単位の前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードのエンコードに使用するために、複数の符号化コンテキストのうちの1つを選択させる命令をさらに備え、エンコードすることが前記選択された符号化コンテキストに従ってエンコードすることを備える請求項27に記載のコンピュータ可読媒体。
- 29前記命令が前記装置に、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには、第1の符号化コンテキストを選択させ、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには、第2の符号化コンテキストを選択させ、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには、第3の符号化コンテキストを選択させる、 請求項28に記載のコンピュータ可読媒体。
- 30前記装置に、 残差ブロックを形成するために、前記選択された予測モードを使用して生成された前記予測モードを、前記ビデオブロックから取り除かせ、 前記選択された予測モードに基づいて、前記残差ブロックに適用するための変換を選択させ、 残差変換係数を生成するために、前記選択された変換を前記残差ブロックに適用させる、 命令をさらに備える請求項27に記載のコンピュータ可読媒体。
- 31前記命令が前記装置に、 前記選択された予測モードが限定された方向性を示すときには、前記残差ブロックに適用するために、離散的コサイン変換(DCT)および整数変換のうちの一方を選択させ、 前記選択された予測モードが方向性を示すときには、前記残差ブロックに適用するために、方向変換を選択させる、 請求項30に記載のコンピュータ可読媒体。
- 32前記選択された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードであるときには、前記命令が前記装置に、前記残差ブロックに適用するために、前記DCTおよび前記整数変換のうちの一方を選択させる請求項31に記載のコンピュータ可読媒体。
- 33前記装置に複数の方向変換を格納させる命令をさらに備え、前記複数の方向変換のそれぞれが、方向性を示す前記予測モードのうちの1つに対応し、前記方向変換を選択することが、前記選択された予測モードに対応する前記複数の方向変換のうちの前記1つを選択することを備える請求項31に記載のコンピュータ可読媒体。
- 34前記装置に前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納させる命令をさらに備え、前記方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項30に記載のコンピュータ可読媒体。
- 35前記装置に前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納させる命令をさらに備え、前記複数の方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、Nが前記ビデオブロックの大きさである請求項30に記載のコンピュータ可読媒体。
- 36前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項27に記載のコンピュータ可読媒体。
- 37前記命令が前記装置に、 前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちの1つと同じであることを示すために、前記予測モードを表す第1のビットをエンコードさせ、 前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードが互いに同じではないとき、前記1つまたは複数の前にエンコードされたビデオブロックのうちのどれが前記ビデオブロックの前記予測モードと同じ予測モードを有するかを示すために、前記予測モードを表す少なくとも1つの追加ビットをエンコードさせる、 請求項27に記載のコンピュータ可読媒体。
- 38前記命令が前記装置に、 前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちのいずれとも同じではないことを示すために、前記予測モードを表す第1のビットをエンコードさせ、 前記複数の予測モードから前記1つまたは複数の前にエンコードされたビデオブロックの少なくとも前記予測モードを削除させ、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの予測モードではない少なくとも1つの追加の予測モードを削除させ、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列させ、 前記現在のビデオブロックの前記予測モードに対応する前記予測モード識別子を識別する符号語をエンコードさせる、 請求項27に記載のコンピュータ可読媒体。
- 39ビデオデータをエンコードする装置であって、 符号化単位のビデオブロックの予測ブロックの生成に使用するために、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを選択するための手段と、 前記符号化単位の1つまたは複数の前にエンコードされたビデオブロックの予測モードに基づいて、前記現在のビデオブロックの前記予測モードをエンコードするための手段と、 を備える装置。
- 40前記符号化単位の前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードのエンコードに使用するために、複数の符号化コンテキストのうちの1つを選択するための手段をさらに備え、エンコードすることが前記選択された符号化コンテキストに従ってエンコードすることを備える請求項39に記載の装置。
- 41前記選択手段が、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには第1の符号化コンテキストを選択し、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには第2の符号化コンテキストを選択し、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには第3の符号化コンテキストを選択する請求項40に記載の装置。
- 42残差ブロックを形成するために、前記選択された予測モードを使用して生成された前記予測ブロックを、前記ビデオブロックから取り除くための手段と、 前記選択された予測モードに基づいて、残差ブロックに適用するための変換を選択するための手段と、 残差変換係数を生成するために、前記選択された変換を前記残差ブロックに適用するための手段と、 をさらに備える請求項39に記載の装置。
- 43前記変換選択手段が、前記選択された予測モードが限定された方向性を示すときには、前記残差ブロックに適用するために、離散的コサイン変換(DCT)および整数変換のうちの一方を選択し、前記選択された予測モードが方向性を示すときには、前記残差ブロックに適用するために、方向変換を選択する請求項42に記載の装置。
- 44前記選択された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードであるときには、前記変換選択手段が、前記残差ブロックに適用するために、前記DCTおよび前記整数変換のうちの一方を選択する請求項43に記載の装置。
- 45複数の方向変換を格納するための手段をさらに備え、前記複数の方向変換のそれぞれが、方向性を示す前記予測モードのうちの1つに対応し、前記変換選択手段が、前記選択された予測モードに対応する前記複数の方向変換のうちの前記1つを選択する請求項43に記載の装置。
- 46前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納するための手段をさらに備え、前記方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項42に記載の装置。
- 47前記予測モードのうちの1つにそれぞれ対応する複数の方向変換を格納することをさらに備え、前記複数の方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、Nが前記ビデオブロックの大きさである請求項42に記載の装置。
- 48前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項39に記載の装置。
- 49前記エンコード手段が、前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちの1つと同じであることを示すために、前記予測モードを表す第1のビットをエンコードし、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードが互いに同じではないとき、前記1つまたは複数の前にエンコードされたビデオブロックのうちのどれが前記ビデオブロックの前記予測モードと同じ予測モードを有するかを示すために、前記予測モードを表す少なくとも1つの追加ビットをエンコードする請求項39に記載の装置。
- 50前記予測モード選択手段が、 前記複数の予測モードから前記1つまたは複数の前にエンコードされたビデオブロックの少なくとも前記予測モードを削除し、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの予測モードではない少なくとも1つの追加の予測モードを削除し、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列し、 前記エンコード手段が、 前記現在のブロックの前記予測モードが前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードのうちのいずれとも同じではないことを示すために、前記予測モードを表す第1のビットをエンコードし、 前記現在のビデオブロックの前記予測モードに対応する前記予測モード識別子を識別する符号語をエンコードする、 請求項39に記載の装置。
- 51ビデオデータを復号する方法であって、 符号化単位のビデオブロックのエンコードされたビデオデータを受信することと、 前記符号化単位の1つまたは複数の前に復号されたビデオブロックの予測モードに基づいて、前記ビデオブロックの予測ブロックの生成に使用するための、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを識別するために、前記エンコードされたビデオデータを復号することと、 を備える方法。
- 52前記符号化単位の前記1つまたは複数の前に復号されたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードの復号に使用するために、複数の符号化コンテキストのうちの1つを選択することをさらに備え、復号することが前記選択された符号化コンテキストに従って復号することを備える請求項51に記載の方法。
- 53前記複数の符号化コンテキストのうちの1つを選択することが、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには、第1の符号化コンテキストを選択することと、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには、第2の符号化コンテキストを選択することと、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには、第3の符号化コンテキストを選択することと、 を備える請求項52に記載の方法。
- 54前記識別された予測モードに基づいて、前記ビデオブロックの残差変換係数に適用するために、逆変換を選択することと、 残差データを生成するために、前記ビデオブロックの前記残差変換係数に前記選択された逆変換を適用することと、 をさらに備える請求項51に記載の方法。
- 55前記変換された残差係数に適用するための逆変換を選択することが、 前記識別された予測モードが限定された方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆離散的コサイン変換(DCT)および逆整数変換のうちの一方を選択することと、 前記識別された予測モードが方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆方向変換を選択することと、 を備える請求項54に記載の方法。
- 56前記逆DCTおよび前記逆整数変換のうちの一方を選択することが、前記識別された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードであるときには、前記ビデオブロックの前記残差変換係数に適用するために、前記逆DCTおよび逆整数変換のうちの一方を選択することを備える請求項55に記載の方法。
- 57方向性を示す前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納することをさらに備え、前記逆方向変換を選択することが、前記識別された予測モードに対応する前記複数の逆方向変換のうちの前記1つを選択することを備える請求項55に記載の方法。
- 58前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納することをさらに備え、前記複数の逆方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項54に記載の方法。
- 59前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納することをさらに備え、前記複数の逆方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項54に記載の方法。
- 60前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項51に記載の方法。
- 61前記符号化単位の1つまたは複数の前に復号されたビデオブロックの予測モードに基づいて、前記ビデオブロックの予測ブロックの生成に使用するための前記複数の予測モードのうちの1つを識別するために、前記エンコードされたビデオデータを復号することが、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのうちの1つとして識別することと、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じであるとき、前記1つまたは複数の前に復号されたビデオブロックのうちの任意のものの前記予測モードを選択することと、 を備える請求項51に記載の方法。
- 62前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じではないとき、前記予測モードを表す少なくとも1つの追加のエンコードされたビットに基づいて、前記1つまたは複数の前に復号されたビデオブロックのうちのどれが前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードと同じ予測モードを有するかを識別することと、 前記識別された前に復号されたビデオブロックの前記予測モードを選択することと、 をさらに備える請求項61に記載の方法。
- 63前記符号化単位の1つまたは複数の前に復号されたビデオブロックの予測モードに基づいて、前記ビデオブロックの予測ブロックの生成に使用するための前記複数の予測モードのうちの1つを識別するために、前記エンコードされたビデオデータを復号することが、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのいずれでもないと識別することと、 前記複数の予測モードから前記1つまたは複数の前に復号されたビデオブロックの少なくとも前記予測モードを削除することと、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードではない少なくとも1つの追加の予測モードを削除することと、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列することと、 前記現在のビデオブロックの予測ブロックの生成に使用するための前記予測モードに対応する前記予測モード識別子を識別するための符号語を復号することと、 を備える請求項51に記載の方法。
- 64ビデオデータを復号する装置であって、 符号化単位の1つまたは複数の前に復号されたビデオブロックの予測モードに基づいて、前記ビデオブロックの予測ブロックの生成に使用するための、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを識別するために、前記符号化単位のビデオブロックのエンコードされたビデオデータを復号するエントロピ復号ユニットと、 前記復号された予測モードを使用して、前記予測ブロックを生成する予測ユニットと、 を備える装置。
- 65前記エントロピ復号ユニットが、前記符号化単位の前記1つまたは複数の前に復号されたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードの復号に使用するために、複数の符号化コンテキストのうちの1つを選択し、復号することが前記選択された符号化コンテキストに従って復号することを備える請求項64に記載の装置。
- 66前記エントロピ復号ユニットが、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには、第1の符号化コンテキストを選択し、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには、第2の符号化コンテキストを選択し、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには、第3の符号化コンテキストを選択する、 請求項65に記載の装置。
- 67前記識別された予測モードに基づいて、前記ビデオブロックの残差変換係数に適用するために、逆変換を選択し、 残差データを生成するために、前記ビデオブロックの前記残差変換係数に前記選択された逆変換を適用する、 逆変換ユニットをさらに備える請求項64に記載の装置。
- 68前記逆変換ユニットが、 前記識別された予測モードが限定された方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆離散的コサイン変換(DCT)および逆整数変換のうちの一方を選択し、 前記識別された予測モードが方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆方向変換を選択する、 請求項67に記載の装置。
- 69前記識別された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードのいずれかであるとき、前記逆変換ユニットが、前記ビデオブロックの前記残差変換係数に適用するために、前記逆DCTおよび逆整数変換のうちの一方を選択する請求項68に記載の装置。
- 70前記逆変換ユニットが、方向性を示す前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納し、前記識別された予測モードに対応する前記複数の逆方向変換のうちの前記1つを選択する請求項68に記載の装置。
- 71前記逆変換ユニットが、前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納し、前記複数の逆方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項67に記載の装置。
- 72前記逆変換ユニットが、前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納し、前記複数の逆方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項67に記載の装置。
- 73前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項64に記載の装置。
- 74前記復号ユニットが、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのうちの1つとして識別し、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じであるとき、前記1つまたは複数の前に復号されたビデオブロックのうちの任意のものの前記予測モードを選択する、 請求項64に記載の装置。
- 75前記復号ユニットが、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じではないとき、前記予測モードを表す少なくとも1つの追加のエンコードされたビットに基づいて、前記1つまたは複数の前に復号されたビデオブロックのうちのどれが前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードと同じ予測モードを有するかを識別し、 前記識別された前に復号されたビデオブロックの前記予測モードを選択する、 請求項74に記載の装置。
- 76前記復号ユニットが、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのいずれでもないと識別し、 前記複数の予測モードから前記1つまたは複数の前に復号されたビデオブロックの少なくとも前記予測モードを削除し、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードではない少なくとも1つの追加の予測モードを削除し、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列し、 前記現在のビデオブロックの予測ブロックの生成に使用するための前記予測モードに対応する前記予測モード識別子を識別するための符号語を復号する、 請求項64に記載の装置。
- 77前記装置が無線通信装置を備える請求項64に記載の装置。
- 78前記装置が集積回路装置を備える請求項64に記載の装置。
- 79ビデオ符号化装置において実行されると、前記装置にビデオブロックを符号化させる命令を備えるコンピュータ可読媒体であって、前記命令が前記装置に、 符号化単位のビデオブロックのエンコードされたビデオデータを受信させ、 前記符号化単位の1つまたは複数の前に復号されたビデオブロックの予測モードに基づいて、前記ビデオブロックの予測ブロックの生成に使用するための、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを識別するために、前記エンコードされたビデオデータを復号させる、 コンピュータ可読媒体。
- 80前記装置に、前記符号化単位の前記1つまたは複数の前に復号されたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードの復号に使用するために、複数の符号化コンテキストのうちの1つを選択させる命令をさらに備え、復号することが前記選択された符号化コンテキストに従って復号することを備える請求項79に記載のコンピュータ可読媒体。
- 81前記命令が前記装置に、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには、第1の符号化コンテキストを選択させ、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには、第2の符号化コンテキストを選択させ、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには、第3の符号化コンテキストを選択させる、 請求項80に記載のコンピュータ可読媒体。
- 82前記装置に、 前記識別された予測モードに基づいて、前記ビデオブロックの残差変換係数に適用するために、逆変換を選択させ、 残差データを生成するために、前記ビデオブロックの前記残差変換係数に前記選択された逆変換を適用させる、 命令をさらに備える請求項79に記載のコンピュータ可読媒体。
- 83前記命令が前記装置に、 前記識別された予測モードが限定された方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆離散的コサイン変換(DCT)および逆整数変換のうちの一方を選択させ、 前記識別された予測モードが方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆方向変換を選択させる、 請求項82に記載のコンピュータ可読媒体。
- 84前記識別された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードであるとき、前記命令が前記装置に、前記ビデオブロックの前記残差ブロックに適用するために、前記逆DCTおよび逆整数変換のうちの1つを選択させる請求項83に記載のコンピュータ可読媒体。
- 85前記装置に方向性を示す前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納させる命令をさらに備え、前記逆方向変換を選択することが、前記識別された予測モードに対応する前記複数の逆方向変換のうちの前記1つを選択することを備える請求項83に記載のコンピュータ可読媒体。
- 86前記装置に前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納させる命令をさらに備え、前記複数の逆方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項82に記載のコンピュータ可読媒体。
- 87前記装置に前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納させる命令をさらに備え、前記複数の逆方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項82に記載のコンピュータ可読媒体。
- 88前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項79に記載のコンピュータ可読媒体。
- 89前記命令が前記装置に、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのうちの1つとして識別させ、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じであるとき、前記1つまたは複数の前に復号されたビデオブロックのうちの任意のものの前記予測モードを選択させる、 請求項79に記載のコンピュータ可読媒体。
- 90前記装置に、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じではないとき、前記予測モードを表す少なくとも1つの追加のエンコードされたビットに基づいて、前記1つまたは複数の前に復号されたビデオブロックのうちのどれが前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードと同じ予測モードを有するかを識別させ、 前記識別された前に復号されたビデオブロックの前記予測モードを選択させる、 命令をさらに備える請求項89に記載のコンピュータ可読媒体。
- 91前記命令が前記装置に、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのいずれでもないと識別させ、 前記複数の予測モードから前記1つまたは複数の前に復号されたビデオブロックの少なくとも前記予測モードを削除させ、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードではない少なくとも1つの追加の予測モードを削除させ、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列させ、 前記現在のビデオブロックの予測ブロックの生成に使用するための前記予測モードに対応する前記予測モード識別子を識別するための符号語を復号させる、 請求項79に記載のコンピュータ可読媒体。
- 92ビデオデータを復号する装置であって、 符号化単位のビデオブロックのエンコードされたビデオデータを受信するための手段と、 前記符号化単位の1つまたは複数の前に復号されたビデオブロックの予測モードに基づいて、前記ビデオブロックの予測ブロックの生成に使用するための、単方向の予測モードと少なくとも2つの単方向の予測モードを組み合わせた多方向の予測モードとを含む、複数の予測モードのうちの1つを識別するために、前記エンコードされたビデオデータを復号するための手段と、 を備える装置。
- 93前記復号手段が、前記符号化単位の前記1つまたは複数の前に復号されたビデオブロックの前記予測モードに基づいて、前記ビデオブロックの前記予測モードの復号に使用するために、複数の符号化コンテキストのうちの1つを選択し、復号することが前記選択された符号化コンテキストに従って復号することを備える請求項92に記載の装置。
- 94前記復号手段が、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向の予測モードであるときには、第1の符号化コンテキストを選択し、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて多方向の予測モードであるときには、第2の符号化コンテキストを選択し、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードがすべて単方向ではなく、すべて多方向でもないときには、第3の符号化コンテキストを選択する、 請求項93に記載の装置。
- 95残差変換係数を変換するための手段をさらに備え、前記変換手段が、 前記識別された予測モードに基づいて、前記ビデオブロックの残差変換係数に適用するために、逆変換を選択し、 残差データを生成するために、前記ビデオブロックの前記残差変換係数に前記選択された逆変換を適用する、 請求項92に記載の装置。
- 96前記変換手段が、 前記識別された予測モードが限定された方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆離散的コサイン変換(DCT)および逆整数変換のうちの一方を選択し、 前記識別された予測モードが方向性を示すときには、前記ビデオブロックの前記残差変換係数に適用するために、逆方向変換を選択する、 請求項95に記載の装置。
- 97前記識別された予測モードがDCの単方向の予測モード、または実質的に直交する方向を指す少なくとも2つの予測モードを組み合わせた多方向の予測モードのいずれかであるとき、前記変換手段が、前記ビデオブロックの前記残差変換係数に適用するために、前記逆DCTおよび逆整数変換のうちの一方を選択する請求項96に記載の装置。
- 98方向性を示す前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納するための手段をさらに備え、前記変換手段が前記識別された予測モードに対応する前記複数の逆方向変換のうちの前記1つを選択する請求項96に記載の装置。
- 99前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納するための手段をさらに備え、前記複数の逆方向変換のそれぞれが、サイズN×Nの列変換行列、およびサイズN×Nの行変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項95に記載の装置。
- 100前記予測モードのうちの1つにそれぞれ対応する複数の逆方向変換を格納するための手段をさらに備え、前記複数の逆方向変換のそれぞれが、サイズN 2 ×N 2 の変換行列を備え、N×Nが前記ビデオブロックの大きさである請求項95に記載の装置。
- 101前記複数の予測モードが単方向の予測モードと可能な双方向の予測モードのサブセットとを含み、前記双方向の予測モードのサブセットが前記単方向の予測モードのそれぞれを含む少なくとも1つの組合せを含む請求項92に記載の装置。
- 102前記復号手段が、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのうちの1つとして識別し、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じであるとき、前記1つまたは複数の前に復号されたビデオブロックのうちの任意のものの前記予測モードを選択する、 請求項60に記載の装置。
- 103前記復号手段が、 前記1つまたは複数の前に復号されたビデオブロックの前記予測モードが同じではないとき、前記予測モードを表す少なくとも1つの追加のエンコードされたビットに基づいて、前記1つまたは複数の前に復号されたビデオブロックのうちのどれが前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードと同じ予測モードを有するかを識別し、 前記識別された前に復号されたビデオブロックの前記予測モードを選択する、 請求項102に記載の装置。
- 104前記復号手段が、 前記予測モードを表す第1のエンコードされたビットに基づいて、前記ビデオブロックの前記予測ブロックの生成に使用するための前記予測モードを、前記1つまたは複数の前に復号されたビデオブロックの前記予測モードのいずれでもないと識別し、 前記複数の予測モードから前記1つまたは複数の前に復号されたビデオブロックの少なくとも前記予測モードを削除し、 前記符号化単位の前記前にエンコードされたビデオブロックのうちの2つ以上が同じ予測モードを有しているとき、前記1つまたは複数の前にエンコードされたビデオブロックの前記予測モードではない少なくとも1つの追加の予測モードを削除し、 前記複数の予測モードの前記残りの予測モード中の予測モード識別子を仮に再配列し、 前記現在のビデオブロックの予測ブロックの生成に使用するための前記予測モードに対応する前記予測モード識別子を識別するための符号語を復号する、 請求項92に記載の装置。
Independent claims104
132 paragraphs, as filed
This application is incorporated herein by reference in its entirety, U.S. Patent Provisional Application No. 60 / 944,470, filed June 15, 2007, and U.S. Patent Provisional, filed October 12, 2007. Claim the interests of Application No. 60 / 979,762.
The present disclosure relates to digital video coding, and more specifically to entropy coding of video blocks.
Digital video capabilities include digital television, digital direct broadcast systems, radio communication devices such as radio telephone handsets, radio broadcast systems, personal digital assistants (PDAs), laptops or desktop computers. It can be incorporated into a wide range of devices including digital cameras, digital recording devices, video game devices, video game consoles, and the like. Digital video equipment implements video compression technologies such as MPEG-2, MPEG-4, or H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC) to send and receive digital video more efficiently. To do. Video compression techniques make spatial and temporal predictions to reduce or eliminate the redundancy inherent in video sequences.
Video compression generally includes spatial and / or temporal predictions. In particular, intra-coding is used for spatial prediction to reduce or eliminate spatial redundancy between video blocks within a given coding unit, which may include video frames, slices of video frames, and so on. Dependent. Inter-coding, on the other hand, relies on temporal prediction to reduce or eliminate temporal redundancy between video blocks of successive coding units of the video sequence. For intra-coding, the video encoder makes spatial predictions to compress the data based on other data in the same coding unit. In the case of intercoding, the video encoder performs motion estimation and motion compensation to track the motion of matching video blocks of two or more adjacent coding units.
After spatial or temporal prediction, residual blocks are generated by removing the predicted video blocks generated during the prediction process from the original encoded video blocks. Therefore, the residual block indicates the difference between the predicted block and the current block being coded. The video encoder can apply a conversion process, a quantization process, and an entropy coding process to further reduce the bit rate associated with the communication of the residual blocks. The conversion technique can change a set of pixel values into a conversion factor that represents the energy of the pixel values in the frequency domain. Quantization is applied to transformation coefficients and generally involves the process of limiting the number of bits associated with any given coefficient. Prior to entropy encoding, the video encoder scans the quantized coefficient block into coefficients for a one-dimensional vector. The video encoding entropy encodes a vector of quantized conversion factors to further compress the residual data.
The video decoder can perform an inverse entropy coding operation to retrieve the coefficients. A reverse scan can also be performed on the decoder to form a 2D block from the coefficients of the received 1D vector. The video decoder then inverse quantizes and inverse transforms its coefficients in order to obtain the reconstructed residual block. The video decoder then decodes the predicted video block based on the predicted and motion information. The video decoder then generates the reconstructed video block and adds the predicted video block to the corresponding residual block to generate a sequence of decoded video information.
The present disclosure describes techniques for encoding video block header information. In particular, the techniques of the present disclosure include a unidirectional prediction mode and a multidirectional prediction mode that combines at least two unidirectional prediction modes for use in generating prediction blocks of video blocks of coding units. , Select one of multiple prediction modes. The video encoder may be configured to encode the prediction mode of the current video block based on the prediction mode of the previously encoded video block of one or more coding units. The video decoder can also be configured to perform the reciprocal decoding function of the encoding performed by the video encoder. Therefore, the video decoder uses a similar technique to decode the prediction mode for use in generating the prediction block of the video block.
The video encoder is used in some examples to encode a selected prediction mode based on, for example, the type of prediction mode of a previously encoded video block, such as unidirectional or multidirectional. , You can choose a different coding context. Moreover, the techniques of the present disclosure can further selectively apply the transformation to the residual information of the video block based on the selected prediction mode. In one example, the video encoder stores multiple directional transforms, each corresponding to one of the different predictive modes, and applies the corresponding directional transforms to the video block based on the selected predictive mode of the video block. can do. In another example, the video encoder stores at least one Discrete Cosine Transform (DCT) or Integer Transform, and multiple Orientation Transforms, and DCT or Integer when the selected prediction mode shows limited directionality. One of the transforms can be applied to the video block residual data when the transform is applied to the video block residual data and the selected prediction mode indicates directionality.
In one aspect, the method of encoding video data is to select one of a plurality of prediction modes for use in generating a prediction block of a video block of coding units, and one of the coding units. Alternatively, it comprises encoding the prediction mode of the current video block based on the prediction mode of a plurality of previously encoded video blocks. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes.
In another aspect, the device that encodes the video data is a prediction unit that selects one of a plurality of prediction modes for use in generating a prediction block of the video block of the coding unit, and a coding unit. It comprises an entropy encoding unit that encodes the prediction mode of the current video block based on the prediction mode of one or more previously encoded video blocks. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes.
In another aspect, the computer-readable medium comprises instructions that cause the device to encode video data when executed in a video encoding device, which commands the device to predict blocks of video blocks of coding units. The current video block prediction mode is based on the prediction mode of one or more previously encoded video blocks of the coding unit, with one of several prediction modes selected for use in the generation. To encode. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes.
In another aspect, the device encoding the video data is a means for selecting one of a plurality of prediction modes and a coding unit for use in generating a prediction block of the video block of the coding unit. It comprises means for encoding the prediction mode of the current video block, based on the prediction mode of one or more previously encoded video blocks. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes.
In another aspect, the method of decoding video data is to receive the encoded video data of the video block of the coding unit and the prediction mode of the video block decoded before one or more of the coding units. Based on, it comprises decoding the encoded video data to identify one of a plurality of prediction modes for use in generating the prediction block of the video block. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes.
In another aspect, a device for decoding video data is for use in generating a prediction block of a video block based on the prediction mode of one or more previously decoded video blocks of the coding unit. An entropy decoding unit that decodes the encoded video data of a video block of coding units is provided to identify one of a plurality of prediction modes. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes. The device also includes a prediction unit that uses the decoded prediction mode to generate a prediction block.
In another aspect, the computer-readable medium comprises an instruction to cause the device to encode a video block when executed in the video coding device. These instructions cause the device to receive the encoded video data of the video block of the coding unit and of the video block based on the prediction mode of the video block decoded before one or more of the coding units. The encoded video data is decoded to identify one of multiple prediction modes for use in generating the prediction block. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes.
In another aspect, the device for decoding the video data is a means for receiving the encoded video data of the video block of the coding unit and the video decoded before one or more of the coding units. A means for decoding encoded video data to identify one of a plurality of prediction modes to be used to generate a prediction block of a video block based on the prediction mode of the block. .. The prediction mode includes a unidirectional prediction mode and a multi-directional prediction mode that combines at least two unidirectional prediction modes.
The techniques described in this disclosure can be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the software may include microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or digital signal processors (DSPs), or other equivalent integrated or individual logic circuits. It can be run on a processor that can point to one or more processors. Software with instructions for performing these techniques can be first stored on a computer-readable medium, loaded by a processor, and executed.
Accordingly, the present disclosure also contemplates a computer-readable medium comprising instructions for causing a processor to perform any of the various techniques described in the present disclosure. In some cases, the computer-readable medium may form part of a computer program product that can be sold to the manufacturer and / or used in the device. Computer program products can include computer-readable media and, in some cases, packaging materials.
Details of one or more aspects of the present disclosure are given in the accompanying drawings and in the description below. Other features, objectives, and advantages of the techniques described in this disclosure will become apparent from the description, drawings, and claims.
<figref num="1">FIG. 1 is a block diagram showing a video encoding and decoding system that implements the coding techniques described in the present disclosure.</figref><figref num="2">FIG. 2 is a block diagram showing an example of the video encoder of FIG. 1 in more detail.</figref><figref num="3">FIG. 3 is a block diagram showing an example of the video decoder of FIG. 1 in more detail.</figref><figref num="4">FIG. 4 is a conceptual diagram showing a virtual example of adjusting the scanning order of the coefficients according to the present disclosure.</figref><figref num="5">FIG. 5 is a flow chart showing an operation example of an encoding device configured to adaptively adjust the scanning order of the coefficients.</figref><figref num="6">FIG. 6 is a flow chart showing an operation example of an encoding unit configured to encode the header information of the video block.</figref><figref num="7">FIG. 7 is a flow chart showing an example of selecting a coding context for coding.</figref><figref num="8">FIG. 8 is a flow chart showing an operation example of a decoding unit configured to decode the header information of the video block.</figref>
FIG. 1 is a block diagram showing a video encoding and decoding system 10 that implements the coding techniques described in the present disclosure. As shown in FIG. 1, the system 10 includes a source device 12 that transmits the encoded video data to the destination device 14 via the communication channel 16. The source device 12 produces encoded video data for transmission to the destination device 14. The source device 12 may include a video source 18, a video encoder 20, and a transmitter 22. The video source 18 of the source device 12 may include a video capture device such as a video camera, a video archive containing pre-captured video, or video supplied by a video content provider. Alternatively, the video source 18 can generate computer graphics-based data as source video, or a combination of live video and computer-generated video. In some cases, the source device 12 can be a so-called camera phone or videophone, in which case the video source 18 can be a video camera. In either case, the captured, pre-captured, or computer-generated video is sent from the source device 12 to the destination device 14 via the transmitter 22 and the communication channel 16 in the video encoder 20. Can be encoded by.
The video encoder 20 receives video data from the video source 18. The video data received from the video source 18 can be a series of video frames. The video encoder 20 divides this series of video frames into several coding units and processes these coding units in order to encode the series of video frames. The coding unit can be, for example, the entire frame or a part of the frame (ie, a slice). Therefore, in some examples, the frame can be split into slices. The video encoder 20 divides each coding unit into blocks of pixels (referred to herein as video blocks or blocks) to encode the video data, with respect to the video blocks within the individual coding units. Operate. Therefore, one coding unit (eg, one frame or slice) can contain multiple video blocks. In other words, one video sequence can contain multiple frames, one frame can contain multiple slices, and one slice can contain multiple video blocks.
The video block has a fixed size or a variable size and can vary in size depending on the specified coding standard. As an example, H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC) of the International Telecommunication Union Telecommunication Standardization Division (ITU-T) (hereinafter, "H.264 / MPEG-4 Part 10") AVC "standard) supports intra-prediction of various block sizes such as 16x16, 8x8, or 4x4 for the luma component, 8x8 for the chroma component, and 16x for the luma component. Supports inter-prediction of various block sizes, such as 16, 16x8, 8x16, 8x8, 8x4, 4x8, and 4x4, with corresponding scale sizes for chroma components. In H.264, for example, each 16x16 pixel video block, often referred to as a macroblock (MB), can be subdivided into smaller sized subblocks and predicted in the subblocks. In general, MB and various subblocks can be thought of as video blocks. Thus, an MB can be thought of as a video block, and if partitioned or subpartitioned, the MB itself can be thought of as defining a set of video blocks.
For each video block, the video encoder 20 selects the block type of the block. The block type can indicate the partition size of the block, in addition to whether the block is predicted using inter-prediction or intra-prediction. For example, the H.264 / MPEG-4 Part 10 AVC standard is Inter 16x16, Inter 16x8, Inter 8x16, Inter 8x8, Inter 8x4, Inter 4x8, Inter 4x4, Supports several inter and intra predictive block types, including intra 16x16, intra 8x8, and intra 4x4. As described in detail below, the video encoder 20 can select one of the block types for each video block.
The video encoder 20 also selects a prediction mode for each video block. For intra-encoded video blocks, the prediction mode can use one or more previously encoded video blocks to determine how to predict the current video block. In the H.264 / MPEG-4 Part 10 AVC standard, for example, the video encoder 20 has nine possible unidirectional prediction modes for each intra 4x4 block; ie vertical prediction mode, horizontal prediction mode, DC prediction mode. , Diagonal down / left prediction mode, Diagonal down / right prediction mode One of mode), vertical right prediction mode, horizontal lower prediction mode, vertical left prediction mode, and horizontal upper prediction mode can be selected. A similar prediction mode is used to predict each intra 8x8 block. In the intra 16x16 block, the video encoder 20 can select one of four possible unidirectional modes; ie, vertical prediction mode, horizontal prediction mode, DC prediction mode, and plane prediction mode. In some examples, the video encoder 20 selects a prediction mode from a set of prediction modes that includes not only a unidirectional prediction mode, but also one or more multi-directional prediction modes that define a combination of unidirectional modes. be able to. For example, one or more multi-directional prediction modes can be bi-directional prediction modes that combine two unidirectional prediction modes, as described in more detail below.
After selecting the prediction mode of the video block, the video encoder 20 uses the selected prediction mode to generate the predicted video block. The predicted video block is removed from the original video block to form a residual block. The residual block contains a set of pixel difference values that quantify the difference between the pixel values of the original video block and the pixel values of the generated predicted block. Residual blocks can be represented in a two-dimensional block format (eg, a two-dimensional matrix or array of pixel difference values).
After generating the residual block, the video encoder 20 can perform some other operations on the residual block before encoding the block. The video encoder 20 can apply transformations such as integer transform, DCT transform, direction transform, wavelet transform, etc. to the pixel value residual block to generate a block of transform coefficients. Therefore, the video encoder 20 converts the residual pixel value into a conversion factor (also called a residual conversion factor). Residual conversion coefficients may be referred to as conversion blocks or coefficient blocks. The transform or coefficient block can be a one-dimensional representation coefficient when a non-separable transform is applied, and a two-dimensional representation coefficient when a separation type transformation is applied. .. Non-separable conversions can include non-separable directional conversions. The discrete transform can include the discrete direction transform, the DCT transform, the integer transform, and the wavelet transform.
After the conversion, the video encoder 20 performs quantization to generate a quantized conversion factor (also called a quantization coefficient, or quantization residual coefficient). Again, the quantization coefficients can be expressed in one-dimensional vector format or two-dimensional block format. Quantization generally refers to the process by which coefficients are quantized in order to reduce the amount of data used to represent the coefficients as much as possible. The quantization process can reduce the bit depth associated with some or all of the coefficients. As used herein, the term "coefficient" can refer to a conversion factor, a quantization factor, or another type of factor. The techniques of the present disclosure may be applied to conversion and quantization conversion factors in addition to residual pixel values in some examples. However, for illustration purposes, the techniques of the present disclosure are described in the context of quantization conversion factors.
When a separable conversion is used and the coefficient blocks are represented in 2D block format, the video encoder 20 scans the coefficients from 2D format to 1D format. In other words, the video encoder 20 can scan the coefficients from a 2D block in order to serialize those coefficients into the coefficients of a 1D vector. According to one of the aspects of the disclosure, the video encoder 20 can adjust the scan order used to transform the coefficient blocks into one dimension based on the statistics collected. The statistic provides an indication of the likelihood that a given coefficient value at each position in the 2D block is zero or non-zero, for example, the count, probability or other associated with each of the coefficient positions in the 2D block. May have statistical metrics. In some examples, statistics can only be collected for a subset of the block's coefficient positions. For example, after a certain number of blocks, when the scan order is evaluated, the coefficient positions within the blocks that are determined to have a higher probability of having a non-zero coefficient are determined to have a lower probability of having a non-zero coefficient. The scan order can be changed so that it is scanned before the coefficient position in the block to be. Thus, the first scan sequence can be configured to more efficiently classify non-zero coefficients at the beginning of the one-dimensional coefficient vector and zero-valued coefficients at the end of the one-dimensional coefficient vector. And this is because at the beginning of the one-dimensional coefficient vector there are some shorter series of zeros between the non-zero coefficients, and at the end of the one-dimensional coefficient vector there is one longer series of zeros. , The number of bits consumed for entropy coding can be reduced.
After scanning the coefficients, the video encoder 20 receives any of a variety of entropy encoding methods, including context-adaptive variable-length encoding (CAVLC), context-adaptive binary arithmetic encoding (CABAC), and run-length encoding. Use the one to encode each of the video blocks in the encoding unit. The source device 12 transmits the encoded video data to the destination device 14 via the transmitter 22 and the channel 16. The communication channel 16 may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines, or any combination of wireless and wired media. The communication channel 16 can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication channel 16 generally represents any communication medium suitable for transmitting encoded video data from the source device 12 to the destination device 14, or a set of different communication media.
The destination device 14 may include a receiver 24, a video decoder 26, and a display device 28. The receiver 24 receives the video bitstream encoded from the source device 12 via channel 16. The video decoder 26 applies entropy decoding to decode the encoded video bitstream to obtain the header information and quantization residual coefficient of the encoded video block in the coding unit. As described above, the quantization residual coefficient encoded by the source device 12 is encoded as a one-dimensional vector. Therefore, the video decoder 26 scans the quantization residual coefficient of the encoded video block in order to convert the coefficient of the one-dimensional vector into the quantization residual coefficient of the two-dimensional block. Like the video encoder 20, the video decoder 26 collects statistics showing the likelihood that a given coefficient position in the video block is zero or non-zero, thereby the same method used in the encoding process. You can adjust the scan order with. Therefore, in order to change the serialized quantization conversion coefficient of the one-dimensional vector representation and back to the quantization conversion coefficient of the two-dimensional block, the reverse adaptive scan order is performed by the video decoder 26. Can be applied.
The video decoder 26 uses the decoded header information and the decoded residual information to reconstruct each block of coding units. In particular, the video decoder 26 can generate predictive video blocks for the current video block and combine the predictive blocks with the corresponding residual video blocks to reconstruct each of the video blocks. The destination device 14 can display the reconstructed video block to the user via the display device 28. The display device 28 can be any of a variety of display devices, such as brown tubes (CRTs), liquid crystal displays (LCDs), plasma displays, light emitting diode (LED) displays, organic LED displays, or other types of display devices. Can be prepared.
In some cases, the source device 12 and the destination device 14 may operate substantially symmetrically. For example, the source device 12 and the destination device 14 may each include a video encoding component and a decoding component. Thus, the system 10 may support unidirectional or bidirectional video transmission between device 12 and device 14, such as video streaming, video broadcasting, or video telephone. A device that includes a video encoding and decoding component can also form part of a common encoding, archiving and playback device, such as a digital video recorder (DVR).
The video encoder 20 and video decoder 26 are defined by the Moving Picture Experts Group (MPEG) MPEG-1, MPEG-2, and MPEG-4, ITU-T. H.263 Standard, Society of Motion Picture and Television Engineers (SMPTE) 421M Video Codec Standard (commonly referred to as "VC-1"), China Audio-Video Coding Standard Work Group (commonly referred to as "AVS"). And can operate according to any of the various video compression standards, including any other video coding standard defined by a standards body or developed by the organization as a proprietary standard. Although not shown in FIG. 1, in some embodiments, the video encoder 20 and the video decoder 26 may be integrated with the audio encoder and decoder, respectively, in a common or separate data stream. It may include a suitable MUX-DEMUX unit, or other hardware and software, to handle both audio and video encoders. In this way, the source device 12 and the destination device 14 can operate on multimedia data. If applicable, the MUX-DEMUX unit is ITU It may comply with the H.223 Multilayer Protocol, or other protocols such as User Datagram Protocol (UDP).
In some embodiments, in the case of video broadcasting, the techniques described in this disclosure are the Forward Link Only (FLO) Air Interface Specification, "Forward Link," published as Technical Standard TIA-1099 in July 2007. Can be applied to enhanced H.264 video encoding for delivering real-time video services in terrestrial mobile multimedia multicast (TM3) systems using the Only Air Interface Specification for Terrestrial Mobile Multimedia Multicast (FLO Specification) it can. That is, the communication channel 16 may include a radio information channel used to broadcast radio video information in accordance with FLO specifications and the like. The FLO specification includes examples that define the bitstream syntax, semantics, and decryption process suitable for the FLO radio interface.
Alternatively, the video is broadcast according to other standards such as DVB-H (Digital Video Broadcast Handheld), ISDB-T (Integrated services digital broadcast --terrestrial), or DMB (Digital Media Broadcast). be able to. Therefore, the source device 12 can be a mobile wireless terminal, a video streaming server, or a video broadcast server. However, the techniques described in this disclosure are not limited to any particular type of broadcast, multicast, or point-to-point system. In the case of broadcast, the source device 12 can broadcast the video data of several channels to a plurality of destination devices, each of which can be similar to the destination device 14 of FIG. Thus, although FIG. 1 shows a single destination device 14, for video broadcasting applications, the source device 12 will generally broadcast video content to many destination devices at the same time.
In another example, the transmitter 22, communication channel 16, and receiver 24 are any wired or wireless communication system, including Ethernet®, telephones (eg POTS), cables, power lines, and fiber optic systems, as well as / Or code division multiple access (CDMA or CDMA2000) communication system, frequency division multiple access (FDMA) system, orthogonal frequency division multiple access (OFDM) connection system, such as GSM (Global System for Mobile Communication), GPRS (General Purpose Packet Radio Service) ), EDGE (enhanced data GSM environment) and other time division multiple access (TDMA) systems, TETRA (terrestrial wireless infrastructure) mobile phone systems, wideband code division multiple access (WCDMA) systems, high-speed data rate 1xEV-DO (First generation Evolution) Data Only) or Ix EV-DO Gold Radio system with one or more of Multicast system, IEEE802.18 system, MediaFLO system, DMB system, DVB-H system, or another method for data communication between two or more devices. Can be configured for communication according to.
The video encoder 20 and video decoder 26 are each one or more macro processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and te. It can be implemented as cycle logic, software, hardware, firmware, or any combination thereof. Each of the video encoder 20 and the video decoder 26 can be included in one or more encoders or decoders, each of which is coupled in their mobile device, subscriber device, broadcast device, server, etc. Can be integrated as part of an encoder / decoder (CODEC). In addition, the source device 12 and destination device 14, respectively, are modulated, demodulated, suitable for transmitting and receiving encoded video, including radio frequency (RF) radio components and antennas sufficient to support radio communication, if applicable. It may include frequency conversion, filtering, and amplifier components. However, for simplicity of explanation, these components are summarized as transmitter 22 of source device 12 and receiver 24 of destination device 14 in FIG.
FIG. 2 is a block diagram showing an example of the video encoder 20 of FIG. 1 in more detail. The video encoder 20 performs intra-coding and inter-coding of blocks in the video frame. Intra-coding relies on spatial prediction to reduce or eliminate spatial redundancy in video data within a given video coding unit, such as frames and slices. For intra-encoding, the video encoder 20 forms a spatial prediction block based on one or more previously encoded blocks within the same coding unit as the encoded block. Intercoding relies on temporal prediction to reduce or eliminate temporal redundancy within adjacent frames of the video sequence. For intercoding, the video encoder 20 makes motion estimates to track the movement of well-matched video blocks between two or more adjacent frames.
In the example of FIG. 2, the video encoder 20 includes a block partition unit 30, a prediction unit 32, a frame storage device 34, a conversion unit 38, a quantization unit 40, a coefficient scan unit 41, an inverse quantization unit 42, and an inverse conversion unit 44. And includes the entropy encoding unit 46. The video encoder 20 also includes adders 48A and 48B (Adder 48). In-loop deblocking filters (not shown) can be applied to reconstructed video blocks to reduce or remove blocking artifacts. The various features described as units in Figure 2 are intended to highlight the different functional aspects of the devices shown, and these units need to be implemented by separate hardware or software components. It does not imply that there is. Rather, the functionality associated with one or more units can be integrated within common or separate hardware or software components.
The block partition unit 30 receives video information in the form of a series of video frames, for example, from the video source 18 (FIG. 1) (denoted as "VIDEO IN" in FIG. 2). The block partition unit 30 divides each video frame into coding units including a plurality of video blocks. As mentioned above, the coding unit can be an entire frame or a portion of a frame (eg, a slice of a frame). In one example, the block partition unit 30 can initially divide each of the coding units into a plurality of video blocks (ie, macroblocks) with a partition size of 16x16. The block partition unit 30 can further subdivide each of the 16x16 video blocks into smaller blocks such as 8x8 video blocks and 4x4 video blocks.
The video encoder 20 performs intra-coding or inter-coding on a block-by-block basis for each video block in the coding unit based on the block type of the block. In addition to the selected partition size of the block, the prediction unit 32 can indicate the block type of the video block that can indicate whether the block should be predicted using inter-prediction or intra-prediction. Assign to each. In the case of inter-prediction, the prediction unit 32 also determines the motion vector. For intra-prediction, the prediction unit 32 determines the prediction mode used to generate the prediction block.
The prediction unit 32 then generates a prediction block. The predictive block can be a predicted version of the current video block. The current video block refers to the video block currently being encoded. In the case of inter-prediction, for example, when a block is assigned an inter-block type, the prediction unit 32 can perform a temporal prediction of the inter-encoding of the current video block. Prediction unit 32 is used to identify blocks in adjacent frames that best match the current video block, for example, blocks in adjacent frames with the smallest MSE, SSD, SAD, or other different metrics. , Compare the current video block with blocks in one or more adjacent video frames. The prediction unit 32 selects the block identified in the adjacent frame as the prediction block.
For intra-prediction, that is, when a block is assigned an intra-block type, the prediction unit 32 is one or more previously encoded neighbor blocks within a common coding unit (eg, frame or slice). A prediction block can be generated based on. The prediction unit 32 performs a spatial prediction to generate a prediction block by performing interpolation, for example, using one or more previously encoded neighbor blocks in the current frame. Can be done. One or more neighbor blocks in the current frame may include, for example, any type of memory or datastore device for storing one or more previously encoded frames or blocks 34 Can be taken from.
The prediction unit 32 can perform interpolation according to one of a set of prediction modes. As mentioned above, the set of prediction modes can include unidirectional prediction modes and / or multidirectional prediction modes. Multidirectional prediction modes define a combination of unidirectional prediction modes. In one example, the set of prediction modes is a unidirectional prediction mode defined in the H.264 / MPEG-4 Part 10 AVC standard, and a bidirectional prediction mode that defines various combinations of two unidirectional prediction modes. May include.
For the intra 4x4 block type, for example, a set of prediction modes is a possible combination of nine unidirectional prediction modes defined in the H.264 / MPEG-4 Part 10 AVC standard, and unidirectional prediction modes. Can include a subset of. Thus, instead of supporting all 36 possible combinations of unidirectional predictive modes, the video encoder 20 can only support a portion of the possible combinations of unidirectional predictive modes. Doing so may not result in significant encoding degradation. An example of a set of intra-prediction modes, including a total of 18 intra-prediction modes, is shown below.
Mode 0: vertical Mode 1: Horizontal Mode 2: DC Mode 3: Diagonal lower left Mode 4: Diagonal lower right Mode 5: Vertical right Mode 6: Horizontal down Mode 7: Vertical left Mode 8: Horizontally above Mode 9: Vertical + Horizontal (Mode 0 + Mode 1) Mode 10: DC + vertical (mode 2 + mode 0) Mode 11: DC + horizontal (mode 2 + mode 1) Mode 12: Diagonal lower left + horizontal (mode 3 + mode 1) Mode 13: Diagonal lower right + vertical (mode 4 + mode 0) Mode 14: Vertical Right + Horizontal (Mode 5 + Mode 1) Mode 15: Horizontal Down + Vertical (Mode 6 + Mode 0) Mode 16: Vertical Left + Horizontal (Mode 7 + Mode 1) Mode 17: Horizontal Top + Vertical (Mode 8 + Mode 0)
In the example of the set shown above, modes 0 to 8 are unidirectional prediction modes, and modes 9 to 17 are bidirectional prediction modes. In particular, modes 0 to 8 are H.264 / MPEG-4 Part 10 Intra 4x4 prediction mode defined by the AVC standard. Modes 9-17 are a subset of the possible bidirectional prediction modes. A subset of the possible bidirectional prediction modes in the example shown includes at least one combination that incorporates each unidirectional prediction mode. Each bidirectional prediction mode is non-parallel, in addition to bidirectional prediction modes, including DC prediction modes (eg, modes 10 and 11), and in some examples, interpolation directions that are substantially orthogonal to each other. Combine the unidirectional prediction modes you have. In other words, a subset of bidirectional prediction modes generally includes bidirectional prediction modes that combine prediction modes from the "vertical" category with prediction modes from the "horizontal" category. These two-way prediction modes allow the intra-prediction process to combine predictive pixels that are available from a great distance, thus improving the prediction quality of more pixel positions within the current video block. it can.
The set of prediction modes described above are described for purposes of explanation. The set of prediction modes may include more or less prediction modes. For example, a set of prediction modes may or may not include more or less bidirectional prediction modes. In another example, the set of prediction modes may contain only a subset of unidirectional prediction modes. Further, the set of prediction modes may include, in addition to, or instead of, a bidirectional prediction mode, a multi-directional prediction mode that combines more than one unidirectional prediction mode. Further, as described above with reference to the intra 4x4 block type, the techniques of the present disclosure apply to other intrablock types (eg, intra8x8 block type or intra16x16 block type) or interblock types. can do.
To determine which one of the plurality of prediction modes to select for a particular block, the prediction unit 32 estimates the coding costs, such as the Lagrange cost, for each prediction mode in the set. You can choose the mode that predicts the lowest coding cost. In another example, the prediction unit 32 can estimate the coding cost of only a portion of a set of possible prediction modes. For example, the prediction unit 32 can select a portion of the set of prediction modes based on the prediction modes selected for one or more neighboring video blocks. The prediction unit 32 uses the selected prediction mode to generate a prediction block.
After generating the prediction block, the video encoder 20 generates a residual block by removing the prediction block generated by the prediction unit 32 from the current video block at the adder 48A. The residual block contains a set of pixel difference values that quantifies the difference between the pixel value of the current video block and the pixel value of the predicted block. Residual blocks can be represented in a two-dimensional block format (eg, a two-dimensional matrix or array of pixel values). In other words, the residual block is a two-dimensional representation of the pixel values.
The transformation unit 38 applies the transformation to the residual block to generate the residual transformation coefficients. The conversion unit 38 can apply, for example, DCT, integer transform, direction transform, wavelet transform, or a combination thereof. The transformation unit 38 can selectively apply the transformation to the residual block based on the prediction mode selected by the prediction unit 32 to generate the prediction block. In other words, the transformation applied to the residual information may depend on the prediction mode selected for the block by the prediction unit 32.
The transformation unit 38 can maintain a plurality of different transformations and selectively apply the transformations to the residual block based on the prediction mode of that block. Multiple different transforms may include DCT, integer transforms, direction transforms, wavelet transforms, or a combination thereof. In some examples, the conversion unit 38 maintains one DCT or integer conversion and multiple direction conversions and selectively applies the conversions based on the prediction mode selected for the current video block. be able to. For example, the transformation unit 38 applies a DCT or integer transformation to the residual block in predictive mode showing limited orientation, and residuals one of the orientation transformations in predictive mode showing significant directionality. Can be applied to blocks.
Using the example of the set of predictive modes described above, the conversion unit 38 can apply DCT or integer conversion to modes 2, 9, and 12-17. Since these modes are either DC predictions or a combination of two prediction modes that are nearly orthogonal, they can exhibit limited directionality. On the contrary, modes 1, 3-8, 10, and 11 are modes that can indicate direction, and therefore the conversion unit 38 uses these to achieve better energy compression of the residual video block. Different directional transformations can be applied to each of the modes. In other words, if a prediction mode with a stronger directionality is selected, the directionality can also be represented in the residual block of such prediction mode. Also, the residual blocks in different prediction modes show different directional characteristics. Therefore, compared to transformations such as DCT and integer transformations such as DCT, trained orientation transformations for each prediction mode provide better energy compression for residual blocks in a given prediction mode. Can be provided. On the other hand, for predictive modes that do not include strong directionality, transformations such as DCT or integer transformations such as DCT provide sufficient energy compression. Thus, the conversion unit 38 does not need to maintain a separate conversion for each of the possible predictive modes, thus lowering the conversion storage requirement. Moreover, the application of DCT and / or integer conversion is less complicated in terms of computational complexity.
In another example, the transformation unit 38 maintains different orientation transformations for each possible prediction mode and can apply the corresponding orientation transformations based on the block's selected prediction modes. In the example of the set of prediction modes described above, the transformation unit 38 can maintain 18 different orientation transformations, each corresponding to one of 18 possible intra 4x4 prediction modes. To do. In addition, the conversion unit 38 has 18 different orientation transformations in 18 possible intra 8x8 prediction modes, 4 different orientation transformations in 4 possible intra 16x16 prediction modes, and other partition sizes. The conversion of any other prediction mode can be maintained. Applying individual directional transformations based on the block's selected prediction mode increases the efficiency with which residual energy is captured, especially for blocks where a prediction mode that exhibits significant directionality is selected. .. The directional transformation can be, for example, a non-separable directional transformation derived from the non-separable Karunenlebe transformation (KLT), or a separate directional transformation. In some examples, the orientation changes may be pre-computed using a training set of data.
KLT is a linear transformation in which the basis functions are derived from signal statistics and can therefore be adaptive. KLT is designed to put as much energy as possible into as few coefficients as possible. KLTs are generally not separable, so transformation unit 38 performs full matrix multiplication, as described in detail below. The application of non-separable orientation transforms to 4x4 residual blocks is illustrated for illustrative purposes. Similar techniques are used for blocks of different sizes, for example 8x8 blocks and 16x16 blocks.
The 4x4 residual block X is represented in the form of a two-dimensional block with four rows and four columns of pixel values, i.e. a total of 16 pixel values. To apply the non-separable orientation transformation, the 4x4 residual blocks are rearranged into a one-dimensional vector x or length 16 of pixel values. The 4x4 residual block X is rearranged into a vector x by arranging the pixels in X in raster scan order. That is, the 4 × 4 residual block X is<maths num="1"><img file="JP2010530184A_D0001.tif" /></maths>
When written as, a residual block x of length 16 is<maths num="2"><img file="JP2010530184A_D0002.tif" /></maths>
It is written as.
The transformation coefficient vector y is obtained by performing matrix multiplication according to the following equation (1) . y = Tx (1) However, T is a transformation matrix of size 16 × 16 corresponding to the prediction mode selected for the block. The conversion coefficient vector y is also a one-dimensional vector having a coefficient of length 16.
The use of non-separable orientation conversions can inevitably involve increased computational costs and storage requirements. In general, for residual blocks of size N × N, non-separable orientation conversions are of size N.<sup>2</sup>× N<sup>2</sup>Requires the basis set of. That is, in the case of a 4x4 residual block, the non-separable directional conversion has a size of 16x16; in the case of an 8x8 residual block, the non-separable directional conversion is 64x64. Has a size of; in the case of a 16x16 residual block, the non-separable orientation transform has a size of 256x256. Since different non-separable directional transformations can be used for each set of prediction modes, the transformation unit 32 has 18 16x16 directional transformations in the case of 4x4 blocks, 8x8 blocks. In the case of, 18 64 × 64 transformations (in the case of the example of the prediction mode set described above) can be stored, and in some cases more if the prediction mode set is larger. This can result in the use of large memory resources to store the transformation matrix needed to run the transformation process. The computational cost of non-separable orientation conversion is also high. In general, applying a non-separable orientation transformation to an N × N block is N<sup>2</sup>× N<sup>2</sup>Multiplication of times and N<sup>2</sup>× (N<sup>2</sup>-1) Requires addition of times.
Instead of the non-separable directional conversion, the conversion unit 32 can maintain the separated directional conversion for each prediction mode. Separated directional transforms have lower storage and computational costs than non-separated directional transforms. For 4 × 4 residual block X, for example, the separable conversion is applied as shown by Eq. (2): Y = CXR (2) Where Y is the resulting transformation coefficient matrix, C is the column transformation matrix, and R is the row transformation matrix, all of which have a size equal to the size of the block (eg,). In this example, 4 × 4). Therefore, the resulting transformation coefficient matrix Y is also a two-dimensional matrix of size 4 × 4.
For each prediction mode, the transformation unit 32 can store two N × N transformation matrices (eg, matrix pairs C and R), where N × N corresponds to the block size (eg, for example). N = 4, 8, or 16). In the example of the set of 18 predictive modes for the 4x4 block described above, the transformation unit 32 is less than the 18 16x16 transformation matrices stored when the non-separable transformation is used. Stores 36 4x4 transformation matrices that only need storage space. In addition, the transform unit 32 is used to perform non-separable orientation transforms.<sup>2</sup>× N<sup>2</sup>Multiplication of times and N<sup>2</sup>× (N<sup>2</sup>Separate directional conversion can be performed using 2 × N × N × N multiplications and 2 × N × N × (N-1) additions, which are considerably less than -1) additions. it can. Table 1 compares storage and computational requirements between the use of 4x4 and 8x8 block size separated-to-non-separable orientation changes. Comparisons between separated and non-separable orientations for 16x16 blocks can be done in a similar manner. As shown in Table 1, using separated directional transforms reduces both computational complexity and storage requirements compared to non-separated directional transforms, which reduces for larger block sizes. , For example, the reduction in the case of 8x8 blocks is greater than the reduction in the case of 4x4 blocks.<tables num="1"><img file="JP2010530184A_D0003.tif" /></tables>
A separate transformation matrix for each prediction mode can be obtained using the prediction residuals from a set of training video series. Similar to the derivation of the non-separable KLT transformation, the Singular Value Decomposition (SVD) process first obtains the row and column transformation matrices, first in the row and then in the column, the predicted residuals of the training set. Can be applied to. Alternatively, a non-separable directional transformation matrix, i.e. a non-separable KLT transformation matrix, can first be trained using the predicted residuals from the training set; then the separated transformation matrix for each prediction mode , Can be obtained by further decomposing the non-separable transformation matrix into the separable transformation matrix.
Either way, the resulting transformation matrix usually has floating-point relative precision. Fixed-point precision numbers are used to approach coefficients in conversion matrices to allow the use of fixed-point operations in the conversion process and to reduce computational costs. The accuracy of the fixed-point approximation of the coefficients in the transformation matrix is determined by finding a balance between the computational complexity and maximum accuracy required during the transformation process using fixed-point arithmetic. In other words, a higher accuracy of the fixed-point approximation of the transformation matrix may result in a smaller error due to the use of the fixed-point approximation of the transformation matrix, which is desirable, but of the fixed-point approximation of the transformation matrix. Too high precision can also cause fixed-point arithmetic to overflow during the conversion process, which is not desirable.
After applying the transformation to the pixel value residual block, the quantization unit 40 quantizes the transformation coefficients to further reduce the bit rate. After quantization, the inverse quantization unit 42 and the inverse transform unit 44 apply the inverse quantization and inverse transform, respectively, to reconstruct the residual block (denoted as "RECON RESID BLOCK" in Figure 2). Can be done. The adder 48B adds the reconstructed residual block to the prediction block generated by the prediction unit 32 in order to generate a reconstructed video block for storage in the frame storage device 34. The reconstructed video block can be used by the prediction unit 32 to intra-code or inter-code the subsequent video block.
As mentioned above, when the integer conversion used in DCT, H.264 / AVC, and the separation type conversion including the separation type direction conversion are used, the resulting conversion coefficients are tabulated as a two-dimensional coefficient matrix. Will be done. Therefore, after quantization, the coefficient scan unit 41 scans the coefficients from a two-dimensional block format to a one-dimensional vector format, a process often referred to as coefficient scanning. In particular, the coefficient scanning unit 41 scans the coefficients according to the scanning order. According to one aspect of the present disclosure, the coefficient scan unit 41 can adaptively adjust the scan order used for coefficient scanning based on one or more coefficient statistics. In some examples, each prediction mode may have different coefficient statistics, so the coefficient scan unit 41 can adaptively adjust the scan order separately for each prediction mode.
The coefficient scan unit 41 can use the first scan sequence to scan the coefficients of the quantized residual block first. In one aspect, the first scan order can be the zigzag scan order commonly used in H.264 / MPEG-4 Part 10 AVC applications. Although the coefficient scan unit 41 is described as being scanned first using a zigzag scan order, the techniques of the present disclosure are not limited to any particular first scan sequence or technique. In addition, each of the predictive modes may have a different initial scan sequence, for example, a scan sequence specifically trained for that predictive mode. However, for the purposes of explanation, the zigzag scan order will be described. The zigzag scan order arranges the quantization coefficients in a one-dimensional vector so that the coefficients in the upper left corner of the two-dimensional block are compressed toward the beginning of the coefficient vector. The zigzag scan order can provide sufficient compactness for coefficient blocks with limited orientation.
If the residual block has some or significant orientation and is transformed using separate orientation transformation, the resulting 2D transformation coefficient block may still contain some amount of orientation. .. This is because using separated directional transforms benefits from lower computational complexity and storage requirements, but cannot provide directional in residual blocks to the same extent as using non-separated directional transforms. Because. As an example, after applying the orientation transformation to vertical prediction (mode 0 in the example above), the nonzero coefficients tend to exist along the horizontal direction. Therefore, the zigzag scan order may not result in all of the nonzero coefficients compressed towards the beginning of the coefficient vector. By adapting the coefficient scan order to direct the scan order horizontally instead of the fixed zigzag scan order, the non-zero coefficients of the coefficient block are one-dimensional coefficients than when scanned in zigzag scan order. It can be more compressed towards the beginning of the vector. This is because the one-dimensional coefficient vector then has some shorter series of zeros between the non-zero coefficients, and the one-dimensional coefficient vector has one longer series of zeros at the end. The number of bits consumed for coding can be reduced. The concept of adapting the scan order used to generate the one-dimensional coefficient vector also applies to other prediction modes. For example, the coefficient scan unit 41 can adaptively adjust the scan order separately for each prediction mode, as each prediction mode can have different directions in the coefficient blocks, and thus different coefficient statistics. .. Thus, the scan order can vary from prediction mode to prediction mode.
As mentioned above, the first scan order may not be the zigzag scan order, especially if the redirection is applied to the residual blocks. In such cases, the initial scan order can be pre-determined using one of the techniques described below. As an example, the initial scan order can be determined using a set of training video sequences. Non-zero coefficient statistics, such as those described below, are collected for each prediction mode and used to initialize the coefficient scan order. In particular, the position with the highest probability of non-zero coefficient is the first coefficient position of the first scan order, and the position with the next highest probability of non-zero coefficient is the second coefficient of the first scan order. The same applies to the position and the position with the lowest probability of non-zero, which is the last coefficient position of the first scan order. Alternatively, the initial scan order can be determined based on the magnitude of the eigenvalues of the isolated transformation matrix. For example, the eigenvalues can be sorted in descending order and the coefficients are scanned according to the corresponding order of the eigenvalues.
Even if the initial scan order is determined using one of the techniques described above, various types of video sources will be placed in different coefficient positions within the block to the quantized residual coefficient. Can be made to. For example, video sources with different resolutions, such as Common Intermediate Format (CIF), Quarter CIF (QCIF), and High Definition (eg 720p / i or 1080p / i) video sources, have different coefficient positions within the block to non-zero coefficients. Can be placed in. Therefore, even if the first scan order is selected based on the block's prediction mode, the coefficient scan unit 41 is still towards the beginning of the one-dimensional coefficient vector to improve the compactness of the non-zero coefficient. The scan order can be adapted.
To adapt the scan order, the coefficient scan unit 41 or other unit of the video encoder 20 can collect one or more coefficient statistics for one or more blocks. In other words, since the coefficient scan is performed block by block, the coefficient scan unit 41 can collect statistics indicating the number of times each position in the block has a non-zero coefficient. For example, the coefficient scan unit 41 may maintain a plurality of counters corresponding to coefficient positions in a two-dimensional block, and when a non-zero coefficient is placed at each position, the counter corresponding to that position may be incremented. it can. In this way, a high count value corresponds to a position in the block where the non-zero coefficient is very often present, and a low count value corresponds to a position in the block where the non-zero coefficient is rarely present. In some examples, the coefficient scan unit 41 can collect separate sets of coefficient statistics for each prediction mode.
As mentioned above, the coefficient scan unit 41 can adapt the scan order based on the statistics collected. The coefficient scan unit 41 is determined to have a higher likelihood of having a non-zero coefficient before a coefficient position that is determined to have a lower likelihood of having a non-zero coefficient based on the collected statistics. The scan order can be adapted to scan. For example, the coefficient scan unit 41 scans the coefficient positions of a two-dimensional block in descending order based on the count values when the count values represent the number of times each coefficient position has a non-zero value. Can be adapted. Alternatively, the counter tracks the number of times each position in the block was a zero-valued coefficient position and adapts the scan order to scan the coefficient positions in ascending order based on their counts. Can be done. In some examples, statistics can only be collected for a subset of the block's coefficient positions instead of all of the block's coefficient positions. In this case, the coefficient scan unit 41 can adapt only part of the scan order.
The coefficient scan unit 41 can adapt the scan order at fixed or non-fixed intervals. For example, the coefficient scan unit 41 can adapt the scan order at fixed intervals, such as block boundaries. In some examples, the coefficient scan unit 41 can adapt the scan order at 4x4 or 8x8 block boundaries, or at macroblock boundaries. In this way, the scan order can be applied block or macroblock by block. However, to reduce the complexity of the system, the coefficient scan unit 41 does not have to adapt the scan order too often, for example every n blocks or macroblocks. Alternatively, the coefficient scan unit 41 can adapt the scan order at non-fixed intervals. The coefficient scan unit 41 can adapt the scan order, for example, when one of the count values at a position in the block exceeds a threshold. After adapting the scan order, the coefficient scan unit 41 can use the adapted scan order to scan at least one subsequent quantized residual block of the subsequent video block. In some examples, the coefficient scan unit 41 uses an adapted scan order when at least one subsequent video block is in the coding unit of the first video block, and at least one subsequent video. Subsequent quantized residual blocks of the block can be scanned. The coefficient scan unit 41 can continue to scan subsequent video blocks until the scan order is re-adapted according to the collected statistics or the scan order is reinitialized. In this way, the coefficient scan unit 41 adapts the scan order so that it generates a one-dimensional coefficient vector in such a way that the quantization residual coefficient can be more efficiently encoded by the entropy encoding unit 46.
The coefficient scan unit 41 can normalize the collected statistics in some examples. Normalization of the collected statistics may be desirable when the coefficient count reaches the threshold. Within a block that has a count value that has reached a threshold, the coefficient position, referred to herein as coefficient position A, is most likely, for example, even when the coefficient position did not have a non-zero coefficient for a period of time. It can remain a coefficient position with a high count. The reason for this is that the coefficient count at position A is large and therefore within the block referred to herein as coefficient position B until other coefficient counts are obtained across multiple blocks (eg, tens or even hundreds of blocks). This is because the coefficient count at another position in the above may not exceed the coefficient count at position A, resulting in a change in scan order (ie, swapping) between coefficient positions A and B. .. Therefore, in order to allow the video encoder 20 to adapt more quickly to local coefficient statistics, the coefficient scan unit 41 can normalize the coefficient when one of the counts reaches the threshold. .. For example, the coefficient scan unit 41 reduces each of the count values by a predetermined multiple, such as reducing each of the count values by a factor of two, or redistributes the count values into the first set of count values. By setting, the coefficient can be normalized. The coefficient scan unit 41 can use other normalization methods. For example, the coefficient scan unit 41 can refresh the statistics after encoding a certain number of blocks.
The entropy encoding unit 46 receives block header information for a block in the form of one or more header syntax elements, in addition to a one-dimensional coefficient vector that represents the residual coefficient of the block. Header syntax elements can identify certain characteristics of the current video block, such as block type, predictive mode, Luma and chroma coded block pattern (CBP), block partition, and one or more motion vectors. .. These header syntax elements may be received from other components within the video encoder 20, for example from the prediction unit 32.
The entropy encoding unit 46 is an encoded bitstream (in Figure 2, "VIDEO". Encode the header and residual information of the current video block to generate "BITSTREAM"). The entropy encoding unit 46 encodes one or more of the respective syntactic elements of the block according to the techniques described in this disclosure. In particular, the entropy encoding unit 46 can encode the syntax elements of the current block based on the syntax elements of one or more previously encoded video blocks. Thus, the entropy encoding unit 46 may include one or more buffers for storing one or more previously encoded video block syntax elements. The entropy encoding unit 46 can analyze any number of neighboring blocks at any location to help encode the syntax elements of the current video block. For purposes of illustration, the entropy encoding unit 46 is placed immediately above the current block, the previously encoded block (ie, the upper neighbor block), and just to the left of the current block. It is described as encoding the prediction mode based on the block encoded in (that is, the neighboring block on the left side). However, similar techniques may be used to encode other header syntax elements such as block type, block partition, CBP, etc. Also, similar techniques may be used in which more neighbor blocks are involved in the encoding of the current video block, not just the upper and left neighbor blocks.
The operation of the entropy encoding unit 46 will be described with reference to the 18 sets of prediction modes described above and with reference to the following pseudo-code example.<maths num="3"><img file="JP2010530184A_D0004.tif" /></maths>
The entropy encoding unit 46 initializes the variables upMode, leftMode, and currMode so that they are equal to the prediction mode of the upper neighbor block, the prediction mode of the left neighbor block, and the prediction mode of the current block, respectively. As mentioned above, the predictive modes of the upper neighbor block, the left neighbor block, and the current block can be determined based on the Lagrange cost analysis. The entropy encoding unit 46 compares the current hook prediction mode (currMode) with the neighboring block prediction mode (upMode and leftMode). If the prediction mode of the current block is equal to the prediction mode of any of the neighboring blocks, the entropy encoding unit 46 encodes to "1". Therefore, the first bit encoded by the entropy encoding unit 46 to represent the prediction mode of the current block is either the prediction mode of the neighboring block whose current prediction mode is on the upper side or the prediction mode of the neighboring block on the left side. Indicates whether they are the same.
If the prediction mode of the current block is equal to the prediction mode of any of the neighboring blocks, that is, if the first encoded bit is "1", the entropy encoding unit 46 sets the prediction mode of the upper neighboring block. , Compare with the prediction mode of the neighboring block on the left. If the prediction mode of the upper neighbor block is the same as the prediction mode of the left neighbor block, the entropy encoding unit 46 no longer encodes the bits of the prediction mode. In this case, the predictive mode can be encoded using a single bit.
However, if the prediction mode of the upper neighbor block is not equal to the prediction mode of the left neighbor block, the entropy encoding unit 46 will specify which of the neighbor blocks has the same prediction mode as the current block. , Encode at least one additional bit that represents the predictive mode. For example, when the entropy encoding unit 46 analyzes the prediction mode of the upper and left neighbor blocks, the entropy encoding unit 46 will say "1" if the prediction mode of the current block is the same as the prediction mode of the upper neighbor block. If the prediction mode of the current block is the same as the prediction mode of the neighboring block on the left side, it is encoded to "0". Alternatively, the entropy encoding unit 46 encodes to "1" when the prediction mode of the current block is the same as the prediction mode of the neighboring block on the left side, and the prediction mode of the current block is the prediction mode of the neighboring block on the upper side. If it is the same as, encode it to "0". In each case, the second bit of the encoded prediction mode indicates which of the upper or left neighbor blocks has the same prediction mode as the prediction mode of the current block. In this way, the entropy encoding unit 46 uses only 1 bit, at most 2 bits, when the prediction mode of the current block is equal to the prediction mode of one of the neighboring blocks. Can be encoded. When the entropy encoding unit 46 analyzes more than one neighboring block, the entropy encoding unit 46 uses 1 to specify which of the previously encoded blocks has the same prediction mode as the current block. More than one additional bit can be encoded.
If the prediction mode of the current video block is not the same as the prediction mode of the upper neighbor block or the prediction mode of the left neighbor block, the entropy encoding unit 46 will either have the prediction mode of the current video block. Send "0" indicating that it is not the same as the prediction mode of the neighboring block. The entropy encoding unit 46 encodes a codeword that represents the prediction mode of the current block. Using a set of 18 prediction modes described above as an example, the entropy encoding unit 46 can encode the prediction mode of the current video block using 4-bit codewords. Although there are generally 18 possible prediction modes that require 5-bit codeword, the prediction modes for the upper and left neighbor blocks have already been compared to the prediction mode for the current block to predict the current block. Two of the above possible prediction modes have already been removed from the current set of blocks, namely the prediction modes of the upper and left neighbor blocks, as it has been determined that they are not equal to the mode. there is a possibility. However, when the upper neighbor block and the left neighbor block have the same prediction mode, instead of 16 prediction modes, 17 prediction modes are still possible, again to represent the 4-bit code. Requires 5-bit codewords, not words. In this case, during the prediction process, the prediction unit 32 is one of the remaining 17 encoding modes from the set so that the prediction mode of the current block can be represented using 4-bit codewords. Can be selectively deleted. In one example, the prediction unit 32 can delete the last prediction mode, eg, prediction mode 17 in this example. However, the prediction unit 32 can use any of the various other methods to select any of the set of prediction modes to be deleted. For example, the prediction unit 32 tracks and selects the probability that each prediction mode will be selected.
After deleting the selected prediction mode, the entropy encoding unit 46 adjusts the range of the remaining 16 prediction modes so that the prediction mode number is in the range [0,15]. In one example, the entropy encoding unit 46 starts by assigning 0 to the remaining predictive modes with the lowest mode numbers and finishes by assigning 15 to the remaining predictive modes with the highest predictive mode numbers. Temporarily renumber the modes from 0 to 15. For example, if the upper neighbor block's prediction mode is mode 12 and the left neighbor block's prediction mode is mode 14, the entropy encoding unit 46 will have prediction mode 13, prediction mode 15, prediction mode 16, and prediction mode. 17 can be renumbered as prediction mode 12, prediction mode 13, prediction mode 14, and prediction mode 15, respectively. The entropy encoding unit 46 then uses 4 bits to encode the predictive mode. In another example with a set of more or less predictive modes possible, the entropy encoding unit 46 can use a similar technique to encode more or less bit predictive modes. ..
The entropy encoding unit 46 can use CAVLC or CABAC to encode the prediction mode of the current video block. A strong correlation can exist between the prediction mode of the current block and the prediction mode of the upper and left neighbor blocks. In particular, when the prediction mode of the upper neighboring block and the prediction mode of the left neighboring block are both unidirectional prediction modes, the probability that the prediction mode of the current block is also one of the unidirectional prediction modes. Is high. Similarly, when both the upper neighbor block prediction mode and the left neighbor block prediction mode are bidirectional prediction modes, the current block prediction mode is also one of the two-way prediction modes. There is a high probability that In this way, changing the prediction mode category of the upper and left neighbor blocks (eg, unidirectional vs. bidirectional) changes the probability distribution of the prediction mode of the current block.
Thus, the entropy encoding unit 46, in some embodiments, has one or more previously encoded video blocks (eg, upper and left neighbor video blocks) in unidirectional or bidirectional prediction modes. Different encoding contexts can be selected depending on the presence. For CABAC, different coding contexts reflect different probabilities of pairs of predictive modes within a given context. The coding context referred to herein as the "first coding context" is taken as an example, corresponding to the case where both the upper and left neighbor coding blocks have a unidirectional prediction mode. Due to the correlation of neighbors, the first coding context can assign a higher probability to the unidirectional prediction mode than the bidirectional prediction mode. Therefore, when the first coding context is selected for CABAC encoding (ie, both the upper and left neighbor prediction modes are unidirectional), the current prediction mode is one of the two-way prediction modes. If the current prediction mode is one of the unidirectional prediction modes, less bits may be spent encoding the current prediction mode than in the case of. For CAVLC, different VLC coding tables can be defined for different coding contexts. For example, when the first coding context is selected (ie, both the upper and left neighbor blocks have a unidirectional prediction mode), a codeword shorter than the bidirectional prediction mode is in the unidirectional prediction mode. The VLC coded table assigned to can be used.
Thus, the entropy encoding unit 46 can select the first coding context when both the prediction mode of the upper video block and the prediction mode of the left video block are unidirectional prediction modes. The entropy encoding unit 46 can select different coding contexts when neither the prediction mode of the upper video block nor the prediction mode of the left video block is a unidirectional prediction mode. For example, the entropy encoding unit 46 can select a second coding context when both the upper neighbor video block prediction mode and the left neighbor video block prediction mode are bidirectional prediction modes. The second coding context models the probability distribution of the prediction mode of the current video block when the prediction modes of the upper and left neighbor blocks are both bidirectional. The probability distribution of the second coding context assigns a higher probability to the bidirectional prediction mode for CABAC coding than for the unidirectional prediction mode and shorter for the CAVLC encoding than for the unidirectional prediction mode. Codewords can be assigned to bidirectional prediction modes.
The entropy encoding unit 46 further has a third coding context when one of the neighboring blocks is a unidirectional prediction mode and the other of the neighboring blocks is a bidirectional prediction mode. Can be selected. The third coding context more evenly distributes the probabilities of the current prediction mode to the set of unidirectional and bidirectional prediction modes. Different encodings for use in encoding based on whether the prediction mode of one or more previously encoded video blocks (eg, upper and left video blocks) is unidirectional or bidirectional. Choosing a context allows for better compression of predictive mode information.
FIG. 3 is a block diagram showing an example of the video decoder 26 of FIG. 1 in more detail. The video decoder 26 can perform intra-decoding and inter-decoding of blocks within a coding unit, such as video frames and slices. In the example of FIG. 3, the video decoder 26 includes an entropy decoding unit 60, a prediction unit 62, a coefficient scan unit 63, an inverse quantization unit 64, an inverse conversion unit 66, and a frame storage device 68. The video decoder 26 also includes an adder 69 that combines the output of the inverse conversion unit 66 with the output of the prediction unit 62.
The entropy decoding unit 60 receives an encoded video bitstream (denoted as "VIDEO BITSTREAM" in FIG. 3) and receives residual information (eg, in the form of a quantization residual coefficient of a one-dimensional vector) and header information (eg, in the form of a one-dimensional vector quantization residual coefficient). Decrypts the encoded bitstream to get (in the form of one or more header syntax elements). The entropy decoding unit 60 performs the reverse decoding function of the encoding performed by the encoding module 46 of FIG. For illustrative purposes, it is described that the entropy decoding unit 60 decodes predictive mode syntax elements. These techniques can be extended to decrypt other syntax elements such as block types, block partitions, and CBP.
In particular, the entropy decoding unit 60 analyzes the first bit representing the prediction mode, and the prediction mode of the current block is a block that was decoded before it was analyzed, for example, an upper neighborhood block or a left neighbor block. Determine if it is equal to the prediction mode of any of the. The entropy decoding module 60 determines that the prediction mode of the current block is equal to the prediction mode of one of the neighboring blocks when the first bit is "1", and the first bit is "0". At some point, it can be determined that the prediction mode of the current block is not the same as any prediction mode of the neighboring blocks.
If the first bit is "1" and the prediction mode of the upper neighbor block is the same as the prediction mode of the left neighbor block, the entropy decoding unit 60 does not need to receive any more bits. The entropy decoding unit 60 selects any prediction mode of the neighboring block as the prediction mode of the current block. The entropy decoding unit 60 may include, for example, one or more buffers (or other memory) that store the previous prediction mode of one or more previously decoded blocks.
If the first bit is "1" and the prediction mode of the upper neighbor block is not the same as the prediction mode of the left neighbor block, the entropy decoding unit 60 receives the second bit representing the prediction mode. The entropy decoding unit 60 determines which of the neighboring blocks has the same prediction mode as the current block based on the second bit. For example, the entropy decoding unit 60 determines that the prediction mode of the current block is the same as the prediction mode of the upper neighboring block when the second bit is "1", and the second bit is "0". When, it can be determined that the prediction mode of the current block is the same as the prediction mode of the neighboring block on the left side. The entropy decoding unit 60 selects the prediction mode of the current neighboring block as the prediction mode of the current block.
However, when the first bit is "0", the entropy decoding unit 60 determines that the prediction mode of the current block is not the same as any prediction mode of the neighboring blocks. Therefore, the entropy decoding unit 60 can remove the prediction modes of the upper and left neighbor blocks from the set of possible prediction modes. A set of possible prediction modes may include one or more unidirectional prediction modes and / or one or more multi-directional prediction modes. An example of a set of prediction modes, including a total of 18 prediction modes, is described above in the description of FIG. If the upper and left neighbor blocks have the same prediction mode, the entropy decoding unit 60 can delete the prediction mode of the neighboring block and at least one other prediction mode. As an example, the entropy decoding unit 60 can delete the prediction mode of the maximum mode number (eg, mode 17 in the set of 18 prediction modes described above). However, the entropy decoding unit 60 should use any other method of various methods to delete the set of prediction modes, as long as the decoding unit 60 deletes the same prediction mode that was deleted by the prediction unit 32. You can choose any of them. For example, the entropy decoding unit 60 can remove the prediction mode that is least likely to be selected.
The entropy decoding unit 60 can adjust the prediction mode numbers of the remaining prediction modes so that the prediction mode numbers are in the range of 0 to 15. In one example, the entropy encoding unit 46 starts with the remaining predictive modes with the lowest mode numbers and ends with the remaining predictive modes with the highest predictive mode numbers, as described above with reference to FIG. The prediction mode can be tentatively renumbered from 0 to 15. The entropy decoding unit 60 decodes the remaining bits, for example, 4 bits in the described example, in order to obtain the prediction mode number of the remaining prediction mode corresponding to the prediction mode of the current block.
In some examples, the entropy decoding unit 60 can use CAVLC or CABAC to decode the predictive mode of the current video block. The entropy decoding unit 60 is configured because there can be a strong correlation between the prediction mode of the current block and one or more previously decoded blocks (eg, the prediction mode of the upper and left neighbor blocks). You can choose different encoding contexts for the block's predictive mode, based on the type of predictive mode for the previously decrypted video block. In other words, the entropy decoding unit 60 can select different coding contexts based on whether the prediction mode of the previously decoded block is unidirectional or bidirectional.
As an example, the entropy decoding unit 60 selects the first encoding context when both predictive modes of the previously decoded block are unidirectional predictive modes, and predicts both of the previously decoded blocks. When the mode is bidirectional predictive mode, a second coding context is selected and one of the previously decoded blocks is in unidirectional predictive mode and of the previously decoded block. A third coding context can be selected when one of the other prediction modes is a bidirectional prediction mode.
The prediction unit 62 uses at least a portion of the header information to generate a prediction block. For example, in the case of an intra-encoded block, the entropy decoding unit 60 may provide the prediction unit 62 with at least a portion of the header information (eg, the block type and prediction mode of this block) for the generation of the prediction block. The prediction unit 62 uses one or more neighboring blocks (or parts of neighboring blocks) within a common coding unit to generate prediction blocks, depending on the block type and prediction mode. As an example, the prediction unit 62 can use the prediction mode specified by the prediction mode syntax element to generate a prediction block of the partition size indicated by the block type syntax element. One or more neighboring blocks (or parts of neighboring blocks) within the current coding unit may be retrieved, for example, from frame storage 68.
In addition, the entropy decoding unit 60 decodes the encoded video data in order to acquire the residual information in the form of a one-dimensional coefficient vector. When separable transformations (eg DCT, H.264 / AVC integer transforms, separable orientation transforms) are used, the coefficient scan unit 63 scans the one-dimensional coefficient vector to generate a two-dimensional block. The coefficient scan unit 63 performs the reverse scanning function of the scan performed by the coefficient scan unit 41 of FIG. In particular, the coefficient scan unit 63 scans the coefficients according to the first scan order in order to convert the coefficients of the one-dimensional vector into a two-dimensional form. In other words, the coefficient scan unit 63 scans the one-dimensional vector to generate the quantization coefficient of the two-dimensional block.
The coefficient scan unit 63 adaptively adjusts the scan order used for coefficient scans based on one or more coefficient statistics to synchronize this scan order with the scan order used by the video encoder 20. be able to. To do so, the coefficient scan unit 63 can collect one or more coefficient statistics for one or more blocks and adapt the scan order based on the collected statistics. In other words, when the quantization coefficient of the 2D block is reconstructed, the coefficient scan unit 63 can collect statistics showing the number of times each position in the 2D block is a non-zero coefficient position. .. The coefficient scan unit 63 maintains a plurality of counters corresponding to the coefficient positions in the two-dimensional block, and when the non-zero coefficient is arranged at each position, the counter corresponding to the position can be incremented.
The coefficient scan unit 63 can adapt the scan order based on the statistics collected. The coefficient scan unit 63 now scans the higher likelihood position with the non-zero coefficient before the coefficient position determined to be less likely with the non-zero coefficient based on the collected statistics. , The scan order can be adapted. The coefficient scan unit 63 adapts the scan order at the same fixed or non-fixed intervals used by the video encoder 20. The coefficient scan unit 63 normalizes the collected statistics in the same way as described above for the video encoder 20.
As mentioned above, the coefficient scan unit 63 can collect separate coefficient statistics in some examples and can adaptively adjust the scan order separately for each prediction mode. For example, the coefficient scan unit 63 can do so because each of the prediction modes can have different coefficient statistics.
After generating the quantization residual coefficient of the two-dimensional block, the inverse quantization unit 64 dequantizes or de-quantizes the quantization residual coefficient. The inverse transformation unit 66 applies inverse transformations such as inverse DCT, inverse integer transformation, inverse transformation, etc. to the dequantized residual coefficient in order to generate a residual block of pixel values. The adder 69 sums the prediction blocks generated by the prediction unit 62 with the residual blocks from the inverse conversion unit 66 to form the reconstructed video block. In this way, the video decoder 26 reconstructs the frames of the video sequence block by block using the header information and the residual information.
Block-based video coding can sometimes result in visually perceptible blockiness at the block boundaries of the encoded video frame. In such cases, deblock filtering can smooth the block boundaries in order to reduce or eliminate visually perceptible block noise. Therefore, a deblocking filter (not shown) can also be applied to filter the decoded blocks to reduce or eliminate block noise. After optional deblocking filtering, the reconstructed block then provides a reference block for the spatial and temporal prediction of subsequent video blocks and displays (eg, display device 28 in FIG. 1). It is housed in frame storage 68, which produces decoded video for driving.
FIG. 4 is a conceptual diagram showing a virtual example of adaptive scanning according to the present disclosure. In this example, the coefficient positions are described in item 71 as c1 to c16. The actual coefficient values are shown in block 1 (72), block 2 (73), block 3 (74), and block 4 (75) as four consecutive blocks. The actual coefficient values in blocks 1-4 may represent a quantization residual coefficient, a conversion coefficient without quantization, or another type of coefficient. In another example, the position may represent the position of the pixel value of the residual block. Blocks 1-4 may include blocks associated with the same prediction mode. In the example shown in FIG. 4, blocks 1 to 4 are 4 × 4 blocks. However, as mentioned above, the techniques of the present disclosure can be extended to apply to blocks of any size. Further, although the coefficient scan unit 41 of the video encoder 20 will be described below, the coefficient scan unit 63 of the video decoder 26 can collect statistics and adapt the scan order in a similar manner.
First, the coefficient scan unit 41 may scan the coefficients of block 1 using a zigzag scan order. In this case, the coefficient scan unit 41 scans the coefficient position of block 1 in the order of c1, c2, c5, c9, c6, c3, c4, c7, c10, c13, c14, c11, c8, c12, c15, c16. To do. Therefore, after scanning the coefficient of block 1, the coefficient scan unit 41 outputs the one-dimensional coefficient vector v, in which case v = [9,4,6,1,1,0,0,0,0, 2,0,0,0,0,0,0]. In the example shown in Figure 4, the coefficient scan unit 41 first scans the coefficients of block 1 using a zigzag scan order, but this zigzag scan is the only possible starting point for adaptive scans. Absent. A horizontal scan, a vertical scan, or any other first scan sequence can be used as the first scan sequence. The use of zigzag scan results in a one-dimensional coefficient vector v with a series of four zeros between the two-dimensional coefficients.
Statistics 1 (76) represents the statistics for block 1. Statistics 1 (76) can be a count value for each coefficient position to track the number of times each coefficient position has a nonzero value. In the example of Figure 4, the coefficient statistics are all initialized to zero. However, other initialization methods may be used. For example, representative or average coefficient statistics for each prediction mode may be used to initialize the statistics for each prediction mode. After encoding block 1, statistic 1 (76) shows a value of 1 for any coefficient position in block 1 that is non-zero, and a value of 0 for any coefficient position in block 1 that has a value of zero. Have. Statistics 2 (77) represent the combined statistics of blocks 1 and 2. The coefficient scan module 41 increments the count in statistic 1 (76) when the coefficient position has a non-zero value in block 2 and keeps the count the same when the coefficient position has a zero value. Therefore, as shown in FIG. 4, the coefficient scan module 41 increments the statistics for coefficient positions c1, c2, c5, c9, and c13 to a value of 2, and the remaining statistics for the coefficient positions are statistics 1 (76). ) And keep it the same. Statistics 3 (78) represent the combined statistics of blocks 1-3, and statistic 4 (79) represents the combined statistics of blocks 1-4. As mentioned above, in some embodiments, the coefficient scan unit 41 can use multiple counters to collect block statistics.
The coefficient scan unit 41 can adapt the scan order based on the collected statistics. In the example shown, the coefficient scan unit 41 can be configured to adapt the scan order after the four video blocks based on statistic 4 (79). In this case, the coefficient scan unit 41 analyzes the collected statistics and adapts the scan order so that the coefficient positions are scanned in descending order by their corresponding count values. Therefore, the coefficient scan unit 41 scans blocks 1 to 4 according to the first scan order, and positions the next block, for example block 5 (not shown), at c1, c5, c9, c2, c13, c6. , C3, c4, c7, c10, c14, c11, c8, c12, c15, c16. The coefficient scan unit 41 then follows this new scan order until the scan order is re-adapted based on the statistics collected for the block, or, for example, it is reinitialized at the beginning of subsequent coding units. Continue scanning for blocks.
Adapting the scan order to change from the first scan order (eg, zigzag scan order) to the new scan order encourages the one-dimensional coefficient vector to start with a non-zero factor and end with a zero factor. .. In the example of Figure 4, the new scan sequence is faster than the coefficients in the horizontal dimension, reflecting the fact that for a given prediction mode, the coefficients in the vertical dimension are more likely to be non-zero than the coefficients in the horizontal dimension. Scan the coefficients in the vertical dimension. Blocks 1-4 may all have the same prediction mode, and past statistics may represent the prospect of future non-zero coefficient positions. Therefore, by using historical statistics to define the scan order, the techniques of the present disclosure have a non-zero coefficient near the beginning of the scanned 1D vector, and near the end of the scanned 1D vector. It encourages the classification of zero-valued coefficients and thus can eliminate or reduce the number of consecutive zeros between two non-zero coefficients. This can then improve the level of compression that can be achieved during entropy coding.
FIG. 5 is a flow chart showing a coding technique according to the present disclosure. The coding technique shown in FIG. 5 can be used for either encoding or decoding of video blocks. As shown in FIG. 5, the coefficient scan units 41, 63 scan the coefficients of a block according to the first scan sequence defined for the corresponding prediction mode of the current block (80). Seen from the video encoder 20, the scan converts the coefficients of the 2D block into 1D coefficient vectors. However, from the perspective of the video decoder 26, the scan would convert the 1D coefficient vector into a 2D coefficient block. As an example, the first scan order of the corresponding prediction mode can be a zigzag scan order. Zigzag scans are not the only possible first scan sequence. A horizontal scan, a vertical scan, or any other first scan sequence can be used as the first scan sequence.
Coefficient scan units 41, 63 collect statistics for one or more blocks (82). In particular, for each block scanned, the coefficient scan units 41, 63 can collect statistics that track the frequency with which each of the coefficient positions in the two-dimensional block is a non-zero coefficient, for example a counter. Coefficient scan units 41, 63 determine whether to evaluate the scan order (83). Coefficient scan units 41, 63 are fixed (eg, every block boundary, or after n block boundaries) or non-fixed intervals (eg, when one of the counts of positions within a block exceeds a threshold). ) Can be used to evaluate the scan order.
If the coefficient scan units 41, 63 decide not to evaluate the scan order, the coefficient scan units 41, 63 scan the next block according to the first scan order (80). For example, if coefficient scan units 41, 63 decide to evaluate the scan order after n blocks have been encoded / decoded, the coefficient scan units will adapt the scan order based on the statistics collected. Can be (84). For example, the coefficient scan units 41, 63 may scan the coefficient positions of a block in descending order based on those count values if the count values reflect the likelihood that a given position has a nonzero coefficient. The scan order can be adapted. After adapting the scan order, the coefficient scan units 41, 63 can, in some cases, determine whether any count value in the statistics exceeds the threshold (86). If one of the coefficient positions has a corresponding count value above the threshold, the coefficient scan units 41, 63 can normalize the collected statistics, eg, the coefficient count value (87). For example, the coefficient scan units 41, 63 reduce each count value by a predetermined multiple, eg, by reducing each count value by a half to reduce it in half, or by reducing the count value to the first count value. The coefficient count value can be normalized by resetting to the set of. Normalizing the coefficient counts can allow the video encoder 20 to adapt to local coefficient statistics more quickly.
After normalizing the collected statistics, or when no normalization occurs, the coefficient scan units 41, 63 scan subsequent blocks using the adapted scan order (88). The coefficient scan units 41, 63 scan at least one subsequent block using the adapted scan order when at least one subsequent block is within the coding unit of the previously scanned video block. can do. The coefficient scan units 41, 63 can continue to scan subsequent video blocks until the scan order is readjusted or reinitialized, for example, at the boundaries of the coding units. In this way, the coefficient scan units 41, 63 are the coefficients of the block determined to have a higher non-zero likelihood before the coefficient position of the block determined to have a lower non-zero likelihood. Adapt the scan order based on the statistics collected to scan the location. Therefore, the one-dimensional coefficient vectors are arranged so as to classify the non-zero coefficients near the beginning of the scanned one-dimensional vector and the zero-valued coefficients near the end of the scanned one-dimensional vector. This can then improve the level of compression that can be achieved during entropy coding.
In some examples, the coefficient scan units 41, 63 may adaptively adjust the scan order separately for each prediction mode, as each prediction mode may have different coefficient statistics. In other words, the coefficient scan units 41, 63 can maintain separate statistics for each prediction mode and adjust the scan order for each prediction mode to be different based on each statistic. Therefore, the flowchart example described above can be performed by the coefficient scan units 41, 63 for each prediction mode.
FIG. 6 is a flow chart showing an operation example of an encoding unit such as the entropy encoding unit 46 of the video encoder 20, which encodes the header information of the video block according to one of the techniques of the present disclosure. The entropy encoding unit 46 receives the block's header information in the form of one or more header syntax elements (90). Header syntax elements identify specific characteristics of the current video block, such as block type, predictive mode, Luma and / or chroma coded block pattern (CBP), block partition, and one or more motion vectors. Can be done. Figure 6 shows the encoding of the prediction mode of the current block. However, similar techniques may be used to encode other of the header syntax elements.
The entropy encoding unit 46 compares the prediction mode of the current block with the prediction mode of one or more previously encoded blocks (92). One or more previously encoded blocks may include, for example, one or more neighboring blocks. In the example of FIG. 6, two previously encoded blocks are analyzed, for example, the upper neighbor block and the left neighbor block. If the prediction mode of the current block is the same as the prediction mode of any of the previously encoded blocks, the entropy encoding unit 46 encodes the first bit to indicate that (94). As an example, the entropy encoding unit 46 may encode the first bit as "1" to indicate that the prediction mode of the current block is the same as the prediction mode of any of the previously encoded blocks. ..
The entropy encoding unit 46 compares the prediction mode of the upper neighbor block with the prediction mode of the left neighbor block (98). If the prediction mode of the upper neighbor block is the same as the prediction mode of the left neighbor block, the entropy encoding unit 46 does not encode any more prediction mode bits (100). In this case, the predictive mode can be encoded using a single bit.
However, if the prediction mode of the upper neighbor block is not equal to the prediction mode of the left neighbor block, the entropy encoding unit 46 indicates which of the neighbor blocks has the same prediction mode as the current block. Encodes a second bit that represents the prediction mode (102). For example, the entropy encoding unit 46 encodes to "1" when the prediction mode of the current block is the same as the prediction mode of the upper neighboring block, and the prediction mode of the current block is the prediction mode of the neighboring block on the left side. If it is the same as, encode it to "0". Therefore, the entropy encoding unit 46 encodes the prediction mode of the current block using only 1 bit, at most 2 bits, when the prediction mode of the current block is equal to the prediction mode of one of the neighboring blocks. be able to.
If the prediction mode of the current block is not the same as any prediction mode of the previously encoded block, the entropy encoding unit 46 encodes the first bit to indicate that (96). Continuing with the above example, the entropy encoding unit 46 sets the first bit to "0" to indicate that the prediction mode of the current block is not the same as any prediction mode of the previously encoded block. Can be encoded. The entropy encoding unit 46 can rearrange the possible set of predictive modes (104). The entropy encoding unit 46 can rearrange the possible prediction mode sets by removing the prediction modes or groups of neighboring blocks from the possible prediction mode sets. When the upper and left neighbor blocks have different prediction modes, the entropy encoding unit 46 can remove the two prediction modes from the set. When the upper and left neighbor blocks have the same prediction mode with each other, the entropy encoding unit 46 can remove one prediction mode from the set (ie, the prediction mode of the upper and left neighbor blocks). Moreover, in some examples, the entropy encoding unit 46 may selectively remove one or more additional encoding modes from its set. When the entropy encoding unit 46 removes one or more additional coding modes, the prediction unit 32 in FIG. 2 also adds them from the set of possible prediction modes so that the same additional encoding modes are not selected. Delete the encoding mode. After removing one or more prediction modes, the entropy encoding unit 46 adjusts the mode numbers for the remaining prediction modes in the set.
The entropy encoding unit 46 encodes a codeword that represents the prediction mode of the current block (106). The entropy encoding unit 46 can encode the prediction mode of the current video block using CAVLC, CABAC, or other entropy encoding methods. As described in more detail with reference to FIG. 7, the encoding unit 46, in some examples, is based on the prediction mode of one or more previously encoded blocks of the prediction mode of the current block. You can adaptively select the encoding context to use for encoding.
FIG. 7 is a flow chart showing a coding context selection according to one aspect of the present disclosure. As mentioned above, there can be a correlation between the type of prediction mode for the current block and the type of prediction mode for one or more previously encoded blocks, such as the upper and left neighbor blocks. For example, when the prediction modes of the upper and left neighboring blocks are both unidirectional prediction modes, the probability that the prediction mode of the current block is also unidirectional prediction mode is higher. Similarly, when the prediction modes of the upper and left neighbor blocks are both bidirectional prediction modes, it is more likely that the prediction mode of the current block is also bidirectional prediction mode.
Therefore, the entropy encoding unit 46 determines whether the prediction mode of the upper and left neighbor blocks is a unidirectional prediction mode (112), and the prediction modes of both the upper and left neighbor blocks are When in unidirectional predictive mode, a first coding context can be selected (114). The first coding context models the probability distribution of the prediction mode of the current video block when the prediction modes of the upper and left neighbor blocks are both unidirectional. The probability distribution of the first coding context can provide higher probabilities for the unidirectional prediction mode of the set than for the bidirectional prediction mode of the set. For example, in the case of CAVLC, the first coding context may use a coding table that associates shorter codewords with unidirectional prediction modes than with codewords associated with bidirectional prediction modes. it can.
When the respective prediction modes of the upper and left neighbor blocks are not unidirectional prediction modes, the entropy encoding unit 46 determines whether the respective prediction modes of the upper and left neighbor blocks are bidirectional prediction modes. Can be (116). The entropy encoding unit 46 can select a second encoding context when the prediction modes of the upper and left neighbor blocks are both bidirectional prediction modes (117). The second coding context models the probability distribution of the prediction mode of the current video block, based on the assumption that the current mode is more likely to be a bidirectional prediction mode than a unidirectional prediction mode. Again, in the case of CAVLC, for example, the second coding context uses a coding table that associates shorter codewords with bidirectional prediction modes than with codewords associated with unidirectional prediction modes. can do.
Entropy encoding when neither the upper and left neighbor block prediction modes are bidirectional prediction modes, that is, when the previously encoded block prediction mode is a combination of bidirectional and unidirectional prediction modes. Unit 46 can choose a third encoding context (118). The third coding context is generated under the assumption that the probabilities of the current prediction mode are more evenly distributed in the set of unidirectional and bidirectional prediction modes. For example, in the case of CAVLC, the third coding context can use a coding table that associates codewords of similar code length with bidirectional and unidirectional prediction modes.
The entropy encoding module 46 encodes the prediction mode of the current video block according to the selected encoding context (119). Choosing a different encoding context to use for encoding the prediction mode of the current video block based on the prediction mode of one or more previously encoded video blocks is a better compression of the prediction mode information. Can bring. The same coding context selection technique is performed by the decoding unit 60 so that the decoding unit 60 can accurately decode the predicted mode of the video block.
FIG. 8 is a flow chart showing an operation example of a decoding unit such as the entropy decoding unit 60 of the video decoder 26, which decodes the header information of the video block according to the technique of the present disclosure. The entropy decoding unit 60 decodes the encoded video bitstream to obtain the header information, for example in the form of one or more header syntax elements. A description of the entropy decoding unit 60 for decoding in predictive mode is given for purposes of illustration. These techniques can be extended to decrypt other header syntax elements such as block type, block partition, CBP, etc.
In particular, the entropy decoding unit 60 receives a first bit representing the prediction mode of the current block (120). Does the entropy decoding unit 60 indicate that the prediction mode of the current block is the same as the prediction mode of the previously decoded block, for example the upper or left neighbor block, with the first bit representing the prediction mode? Decide whether (122). For example, the entropy decoding module 60 states that when the first bit is "1", the prediction mode of the current block is the same as the prediction mode of one of the upper and left neighbor blocks, and the first bit. When is "0", it can be determined that the prediction mode of the current block is not the same as the prediction mode of the upper and left neighbor blocks.
When the entropy decoding unit 60 determines that the prediction mode of the current block is the same as the prediction mode of one of the upper and left neighbor blocks, the entropy decoding unit 60 sets the prediction mode of the upper neighbor block and the left side. Determines if the prediction mode of the neighboring blocks of is the same (124). When the prediction mode of the upper neighboring block and the prediction mode of the left neighboring block are the same, no more bits representing the prediction mode of the current video block are received, and the entropy decoding unit 60 is one of the neighboring blocks. Select the prediction mode for the current block as the prediction mode for the current block (126). When the prediction mode of the upper neighbor block and the prediction mode of the left neighbor block are different, one additional bit representing the prediction bit is received, and the entropy decoding unit 60 is sent to that next received bit representing the prediction mode. Based on this, the appropriate neighboring block prediction mode is selected as the prediction mode for the current block (128). For example, the entropy decoding unit 60 selects the prediction mode of the upper neighboring block as the prediction mode of the current block when the next received bit is "1", and the next received bit is "0". When is, the prediction mode of the neighboring block on the left side can be selected as the prediction mode of the current block.
When the entropy decoding unit 60 determines that the prediction mode of the current block is not the same as the prediction mode of either the upper or left neighbor block, that is, when the first bit representing the prediction mode is "0". , Entropy Decoding Unit 60 Entropy Decoding Unit 60 can delete one or more prediction modes in a set of possible prediction modes (130). The entropy decoding unit 60 can remove the prediction modes of the upper and left neighbor blocks from the set of possible prediction modes. If the upper and left neighbor blocks have the same prediction mode, the entropy decoding unit 60 can remove the prediction mode of the neighbor block and at least one other prediction mode, as described in detail above.
The entropy decoding unit 60 decodes the remaining bits, eg, 4 bits in the example described, in order to obtain the predicted mode number of the predicted mode of the current block (132). The entropy decoding unit 60 can adjust the predictive mode numbering of the remaining predictive modes in the reverse way of the predictive mode numbering adjustment process performed by the entropy encoding unit 46 (134). In one example, the entropy decoding unit 60 reinserts the decrypted predictive mode number (range 0-15) by reinserting the deleted predictive mode and reinserts the original predictive mode number (range 0-17). ) Can be. In some examples, the entropy decoding unit 60 is based on the prediction mode of one or more previously decrypted video blocks, for example, the prediction mode of the previously decoded block, as detailed above. Different coding contexts can be selected for the prediction mode of the block based on whether they are all unidirectional, both bidirectional, or one unidirectional and the other bidirectional. The entropy decoding unit 60 provides a prediction mode to the prediction unit 62 to generate a prediction block according to the prediction mode selected (136). As described with reference to FIG. 3, the predictive block is combined with the residual pixel value to generate a reconstructed block for presentation to the user.
The techniques described in this disclosure can be implemented in hardware, software, firmware, or any combination thereof. Any feature described as a unit or component can be implemented together or separately in an integrated logical unit, but separately as an interoperable logical unit. When implemented in software, such techniques, when implemented, can be at least partially realized by computer-readable media with instructions that perform one or more of the methods described above. The computer-readable medium may form part of a computer program product that may include packaging material. Computer-readable media include random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), and electrically erasable programmable read-only memory (EEPROM). , Flash memory, magnetic or optical data storage medium, etc. may be provided. These techniques, or instead, carry or convey code in the form of instructions or data structures, and are at least partially realized by computer-readable communication media that can be accessed, read, and / or executed by a computer. Can be done.
One code, including one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASCIs), field programmable logic arrays (FPGAs), or other equivalent integrated or individual logic circuits. Or it can be run by multiple processors. Thus, the term "processor" as used herein may refer to any of the above structures, or any other structure suitable for the practice of the techniques described herein. In addition, in some embodiments, the functionality described herein is a dedicated software or hardware unit built into a video encoder-decoder (CODEC) configured or combined for encoding and decoding. Can be provided inside. The various features described as units are intended to highlight the different functional aspects of the devices shown, and that these units must necessarily be implemented by separate hardware or software components. It does not imply. Rather, the functionality associated with one or more units can be integrated within common or separate hardware or software components.
Various embodiments of the present disclosure have been described. These and other embodiments are included within the scope of the following claims.
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2015115614A | Cited by | Japan | Search report |
| US10432956B2 | Cited by | United States of America | Applicant |
| JP2015027116A | Cited by | Japan | Search report |
| US10484700B2 | Cited by | United States of America | Applicant |
| JP2018121353A | Cited by | Japan | Search report |
| JP2015213367A | Cited by | Japan | Search report |
| JPWO2012081162A1 | Cited by | Japan | Search report |
| US10820001B2 | Cited by | United States of America | Applicant |
| JP2018113702A | Cited by | Japan | Search report |
| JP2012147268A | Cited by | Japan | Search report |
| TWI602428B | Cited by | Taiwan Province of China | Examiner |
| TWI628949B | Cited by | Taiwan Province of China | Examiner |
| US10820000B2 | Cited by | United States of America | Applicant |
| US11831896B2 | Cited by | United States of America | Applicant |
| US10827193B2 | Cited by | United States of America | Applicant |
| JP2021044840A | Cited by | Japan | Search report |
| US11831892B2 | Cited by | United States of America | Applicant |
| JP2017063450A | Cited by | Japan | Search report |
| US10075723B2 | Cited by | United States of America | Applicant |
| JP5711266B2 | Cited by | Japan | Examiner |
| US11350120B2 | Cited by | United States of America | Applicant |
| WO2012096095A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO2012081162A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| TWI650005B | Cited by | Taiwan Province of China | Examiner |
| TWI660625B | Cited by | Taiwan Province of China | Examiner |
| US10484699B2 | Cited by | United States of America | Applicant |
| AU2011354861B2 | Cited by | Australia | Search report |
| JP2017063450A | Cited by | Japan | Search report |
| US11831893B2 | Cited by | United States of America | Applicant |
| AU2015202844B2 | Cited by | Australia | Search report |
| JP2015130688A | Cited by | Japan | Search report |
| US10397597B2 | Cited by | United States of America | Applicant |
| JP2015130688A | Cited by | Japan | Search report |
| JP2016187228A | Cited by | Japan | Search report |
| US10178402B2 | Cited by | United States of America | Applicant |
| JP2002135126A | Cites | Japan | Examiner |
| JP2002232887A | Cites | Japan | Examiner |
| JP2007053561A | Cites | Japan | Examiner |
| JPH10271505A | Cites | Japan | Examiner |
| JPH1155678A | Cites | Japan | Examiner |
80 members in 14 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 60944470 | United States of America | – | |
| 94447007 | United States of America | P | |
| 60979762 | United States of America | – | |
| 97976207 | United States of America | P | |
| 12133227 | United States of America | – | |
| 13322708 | United States of America | A | |
| 2008066797 | United States of America | W |
Members80
| Document | Office | Kind | |
|---|---|---|---|
| US2008310504A1 | United States of America | A1 | |
| US2008310507A1 | United States of America | A1 | |
| US2008310512A1 | United States of America | A1 | |
| US2008310745A1 | United States of America | A1 | |
| CA2687253A1 | Canada | A1 | |
| CA2687260A1 | Canada | A1 | |
| CA2687263A1 | Canada | A1 | |
| CA2687725A1 | Canada | A1 | |
| WO2008157268A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008157269A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008157360A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008157431A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008157268A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008157431A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW200913723A | Taiwan Province of China | A | |
| TW200913727A | Taiwan Province of China | A | |
| WO2008157269A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008157360A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW200915880A | Taiwan Province of China | A | |
| TW200915885A | Taiwan Province of China | A | |
| KR20100021658A | Republic of Korea | A | |
| WO2008157431A8 | World Intellectual Property Organization (WIPO) | A8 | |
| KR20100029246A | Republic of Korea | A | |
| KR20100029837A | Republic of Korea | A | |
| KR20100029838A | Republic of Korea | A | |
| CN101682770A | China | A | |
| CN101682771A | China | A | |
| EP2165542A2 | European Patent Office (EPO) | A2 | |
| EP2165543A2 | European Patent Office (EPO) | A2 | |
| EP2168381A2 | European Patent Office (EPO) | A2 | |
| EP2172026A2 | European Patent Office (EPO) | A2 | |
| CN101743751A | China | A | |
| CN101803386A | China | A | |
| JP2010530183A | Japan | A | |
| JP2010530184AThis record | Japan | A | |
| JP2010530188A | Japan | A | |
| JP2010530190A | Japan | A | |
| RU2010101053A | Russian Federation | A | |
| RU2010101115A | Russian Federation | A | |
| RU2010101116A | Russian Federation | A | |
| RU2010101085A | Russian Federation | A | |
| RU2434360C2 | Russian Federation | C2 | |
| KR20110129493A | Republic of Korea | A | |
| KR101091479B1 | Republic of Korea | B1 | |
| KR101107867B1 | Republic of Korea | B1 | |
| CN101682771B | China | B | |
| RU2446615C2 | Russian Federation | C2 | |
| RU2447612C2 | Russian Federation | C2 | |
| KR101161065B1 | Republic of Korea | B1 | |
| CN101682770B | China | B | |
| EP2165542B1 | European Patent Office (EPO) | B1 | |
| JP5032657B2 | Japan | B2 | |
| RU2463729C2 | Russian Federation | C2 | |
| US2013044812A1 | United States of America | A1 | |
| KR101244229B1 | Republic of Korea | B1 | |
| US8428133B2 | United States of America | B2 | |
| CN101743751B | China | B | |
| TWI401959B | Taiwan Province of China | B | |
| US8488668B2 | United States of America | B2 | |
| JP5254324B2 | Japan | B2 | |
| JP2013153463A | Japan | A | |
| CA2687260C | Canada | C | |
| US8520732B2 | United States of America | B2 | |
| US8571104B2 | United States of America | B2 | |
| US8619853B2 | United States of America | B2 | |
| US2014112387A1 | United States of America | A1 | |
| JP5575940B2 | Japan | B2 | |
| EP2165543B1 | European Patent Office (EPO) | B1 | |
| BRPI0813275A2 | Brazil | A2 | |
| DK2165543T3 | Denmark | T3 | |
| PT2165543E | Portugal | E | |
| ES2530796T3 | Spain | T3 | |
| PL2165543T3 | Poland | T3 | |
| BRPI0813345A2 | Brazil | A2 | |
| BRPI0813351A2 | Brazil | A2 | |
| CA2687263C | Canada | C | |
| BRPI0813349A2 | Brazil | A2 | |
| US9578331B2 | United States of America | B2 | |
| BRPI0813351B1 | Brazil | B1 | |
| BRPI0813345B1 | Brazil | B1 |
27 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2010530184
- Application
- 2010512367
Titles2
- Japanese
- ビデオブロック予測モードの適応形符号化
- English
- Adaptive coding in video block prediction mode
Classification
- CPC, 21
- H04N19/625
- H04N19/197
- H04N19/61
- H04N19/103
- H04N19/11
- H04N19/12
- H04N19/122
- H04N19/129
- H04N19/13
- H04N19/147
- H04N19/157
- H04N19/176
- H04N19/18
- H04N19/19
- H04N19/196
- H04N19/42
- H04N19/46
- H04N19/463
- H04N19/48
- H04N19/593
- H04N19/70
- IPC, 3
- H04N7 32
- H04N7 30
- H04N19 593
Designated states137
- Regional, 73
- Botswana
- Ghana
- Gambia
- Kenya
- Lesotho
- Malawi
- Mozambique
- Namibia
- Sudan
- Sierra Leone
- Eswatini
- United Republic of Tanzania
- Uganda
- Zambia
- Zimbabwe
- Armenia
- Azerbaijan
- Belarus
- Kyrgyzstan
- Kazakhstan
- Republic of Moldova
- Russian Federation
- Tajikistan
- Turkmenistan
and 49 moreShow fewer
- Austria
- Belgium
- Bulgaria
- Switzerland
- Cyprus
- Czechia
- Germany
- Denmark
- Estonia
- Spain
- Finland
- France
- United Kingdom
- Greece
- Croatia
- Hungary
- Ireland
- Iceland
- Italy
- Lithuania
- Luxembourg
- Latvia
- Monaco
- Malta
- Netherlands (Kingdom of the)
- Norway
- Poland
- Portugal
- Romania
- Sweden
- Slovenia
- Slovakia
- Türkiye
- Burkina Faso
- Benin
- Central African Republic
- Congo
- Côte d’Ivoire
- Cameroon
- Gabon
- Guinea
- Equatorial Guinea
- Guinea-Bissau
- Mali
- Mauritania
- Niger
- Senegal
- Chad
- Togo
- National, 64
- United Arab Emirates
- Antigua and Barbuda
- Albania
- Angola
- Australia
- Bosnia and Herzegovina
- Barbados
- Bahrain
- Brazil
- Belize
- Canada
- China
- Colombia
- Costa Rica
- Cuba
- Dominica
- Dominican Republic
- Algeria
- Ecuador
- Egypt
- Grenada
- Georgia
- Guatemala
- Honduras
and 40 moreShow fewer
- Indonesia
- Israel
- India
- Japan
- Comoros
- Saint Kitts and Nevis
- Democratic People’s Republic of Korea
- Republic of Korea
- Lao People’s Democratic Republic
- Saint Lucia
- Sri Lanka
- Liberia
- Libya
- Morocco
- Montenegro
- Madagascar
- North Macedonia
- Mongolia
- Mexico
- Malaysia
- Nigeria
- Nicaragua
- New Zealand
- Oman
- Papua New Guinea
- Philippines
- Serbia
- Seychelles
- Singapore
- San Marino
- El Salvador
- Syrian Arab Republic
- Tunisia
- Trinidad and Tobago
- Ukraine
- United States of America
- Uzbekistan
- Saint Vincent and the Grenadines
- Viet Nam
- South Africa