Video data stream concept
Abstract
Problem to be solved.To provide a video data stream coding concept which makes it easier to identify a part of a data stream and enables low inter-terminal delay. Decoder search timing information, ROI information and tile identification information are transmitted in a video data stream at a level that allows easy access by network entities such as MANE or decoders. To reach this level, such types of information are transmitted within the video data stream via packets that are distributed among the packets of the access unit of the video data stream. Distributed packets are among the removable packet types, i.e., the ability of the decoder to completely recover the video content transmitted over the video data stream. To maintain. [Selection diagram] Fig. 11

Term
10.4 yearsto projected expiry
Projected expiry 6 February 2037, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
26 claims: 6 independent, 20 dependent
- 1ビデオ内容(16)の画像(18)のサブポーション(24)を単位にしてその中で符号化されるビデオ内容(16)を有し、各サブポーション(24)はビデオ・データ(22)のパケットのシーケンス(34)の1つ以上のペイロード・パケット(32)にそれぞれ符号化され、パケットのシーケンス(34)はアクセスユニット(30)のシーケンスに分割され、それにより各アクセスユニット(30)がビデオ内容(16)のそれぞれの画像(18)に関するペイロード・パケット(32)を集めるビデオ・データをその上に記憶されるデジタル記憶媒体であって、パケットのシーケンス(34)はその中にタイミング制御パケット(36)を分散し、それによりタイミング制御パケット(36)がアクセスユニット(30)を復号化ユニット(38)に再分割し、それにより少なくともいくつかのアクセスユニット(30)が複数の復号化ユニット(38)に再分割され、各タイミング制御パケット(36)が復号化ユニット(38)のためのデコーダバッファ検索時間の信号を送り、そのペイロード・パケット(32)がパケットのシーケンス(34)におけるそれぞれのタイミング制御パケット(38)に続き、復号化ユニット(38)のためのデコーダバッファ検索時間は、符号化ピクチャ・バッファからの検索時間および復号化ピクチャ・バッファからのすでに復号化された画像データの検索時間を含むように構成される、デジタル記憶媒体。
- 2サブポーション(24)はスライスであり、各ペイロード・パケット(32)は1つ以上のスライスを含む、請求項1に記載のデジタル記憶媒体。
- 3スライスは、独立して復号化可能なスライスと、WPP処理のために、スライス境界を越えたエントロピーおよび予測復号化を用いた復号化を可能にする従属性スライスとを含む、請求項2に記載のデジタル記憶媒体。
- 4パケットのシーケンス(34)の各パケットが異なるパケット・タイプであるペイロード・パケット(32)およびタイミング制御パケット(36)を有する複数のパケット・タイプから正確に1つのパケット・タイプに割り当てられ、パケットのシーケンス(34)における複数のパケット・タイプのパケットの発生は、各アクセスユニット(30)の範囲内でパケットにより従わされることになっているパケット・タイプの間の順序を規定するいくらかの限界に従属し、それにより、アクセスユニット境界は、限界が抵触するときを検出することにより限界を用いて検出可能であり、たとえいかなる除去可能なパケット・タイプがビデオ・データから除去されても、それにより、アクセスユニット境界は、パケットのシーケンスの範囲内の同じ位置に残り、ペイロード・パケット(32)は除去不可能なパケット・タイプであり、タイミング制御パケット(36)は除去可能なパケット・タイプである、請求項1ないし請求項3のいずれかに記載のデジタル記憶媒体。
- 5各パケットは、パケット・タイプを表示している構文要素部を含む、請求項1ないし請求項4のいずれかに記載のデジタル記憶媒体。
- 6各パケットは、パケット・タイプを表示している構文要素部によって含まれるパケット・タイプフィールドを含み、パケット・タイプフィールドの内容は、ペイロード・パケットおよびタイミング制御パケットとの間で異なり、さらに、タイミング制御パケットは、一方のタイミング制御パケットと他方の異なるタイプのSEIパケットとの間を区別しているSEIパケット・タイプフィールドを含む請求項5に記載のデジタル記憶媒体。
- 7各復号化ユニットは、サブピクチャを表し、符号化ピクチャ・バッファからの検索時間は、同じサブピクチャにおける先行するアクセスユニットの最近のタイミング制御パケットと関連し、復号化ユニットの符号化ピクチャ・バッファからの除去後、どれくらいのクロックが待つために刻んだかを特定する、請求項2に記載のデジタル記憶媒体。
- 8サブポーションはスライスであり、ビデオ内容はビデオ・データに符号化され、スライス(24)の中の符号化順序を用いて、スライス(24)を単位にして、予測およびエントロピー符号化を用いて、ビデオ内容の画像が空間的に再分割されるタイル(70)の内部に予測符号化の予測および/またはエントロピー符号化の予測を制限し、そこにおいて、スライス(24)のシーケンスはペイロード・パケット(32)にパケット化され、そこにおいて、パケットのシーケンス(34)がタイル識別パケットとして1つのアクセスユニットのペイロード・パケットとの間でそこに分散するタイミング制御パケット(36)を有し、パケットのシーケンス(34)のそれぞれのタイル識別パケット(72)に直ちに続く1つ以上のペイロード・パケット(32)にパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認する、請求項1ないし請求項7のいずれかに記載のデジタル記憶媒体。
- 9タイル識別パケットは、正確に直ちに続くペイロード・パケットにパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認する、請求項8に記載のデジタル記憶媒体。
- 10タイル識別パケットは、現在のアクセスユニット(30)の端部の前までパケットのシーケンス(34)のそれぞれのタイル識別パケットに直ちに続く1つ以上のペイロード・パケット(32)にパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認し、次のタイル識別パケットは、それぞれ、パケットのシーケンス(34)にある、請求項9に記載のデジタル記憶媒体。
- 11ビデオ・データは、スケーラブルビデオ符号化(SVC)データである、請求項1に記載のデジタル記憶媒体。
- 12ビデオ内容(16)の画像(18)のサブポーション(24)を単位にするビデオ内容(16)をビデオ・データ(22)に符号化するためのエンコーダであって、それぞれ各サブポーション(24)をビデオ・データ(22)のパケットのシーケンス(34)の1つ以上のペイロード・パケット(32)に符号化し、それにより、パケットのシーケンス(34)がアクセスユニット(30)のシーケンスに分割され、各アクセスユニット(30)がビデオ内容(16)のそれぞれの画像(18)に関するペイロード・パケット(32)を集め、エンコーダはパケットのシーケンス(34)にタイミング制御パケット(36)を分散し、それにより、タイミング制御パケット(36)はアクセスユニット(30)を復号化ユニット(38)に再分割し、それにより、少なくともいくつかのアクセスユニット(30)は複数の復号化ユニット(38)に再分割され、各タイミング制御パケット(36)は復号化ユニット(38)のためのデコーダバッファ検索時間の信号を送り、そのペイロード・パケット(32)はパケットのシーケンス(34)におけるそれぞれのタイミング制御パケット(36)に続き、復号化ユニット(38)のためのデコーダバッファ検索時間は、符号化ピクチャ・バッファからの検索時間および復号化ピクチャ・バッファからのすでに復号化された画像データの検索時間を含むように構成された、エンコーダ。
- 13ビデオ・データは、スケーラブルビデオ符号化(SVC)データである、請求項12に記載のエンコーダ。
- 14各復号化ユニットは、サブピクチャを表し、符号化ピクチャ・バッファからの検索時間は、同じサブピクチャにおける先行するアクセスユニットの最近のタイミング制御パケットと関連し、復号化ユニットの符号化ピクチャ・バッファからの除去後、どれくらいのクロックが待つために刻んだかを特定する、請求項12に記載のエンコーダ。
- 15サブポーションはスライスであり、ビデオ内容はビデオ・データに符号化され、スライス(24)の中の符号化順序を用いて、スライス(24)を単位にして、予測およびエントロピー符号化を用いて、ビデオ内容の画像が空間的に再分割されるタイル(70)の内部に予測符号化の予測および/またはエントロピー符号化の予測を制限し、そこにおいて、スライス(24)のシーケンスはペイロード・パケット(32)にパケット化され、そこにおいて、パケットのシーケンス(34)がタイル識別パケットとして1つのアクセスユニットのペイロード・パケットとの間でそこに分散するタイミング制御パケット(36)を有し、パケットのシーケンス(34)のそれぞれのタイル識別パケット(72)に直ちに続く1つ以上のペイロード・パケット(32)にパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認する、請求項12ないし請求項14のいずれかに記載のエンコーダ。
- 16タイル識別パケットは、正確に直ちに続くペイロード・パケットにパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認する、請求項15に記載のエンコーダ。
- 17タイル識別パケットは、現在のアクセスユニット(30)の端部の前までパケットのシーケンス(34)のそれぞれのタイル識別パケットに直ちに続く1つ以上のペイロード・パケット(32)にパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認し、次のタイル識別パケットは、それぞれ、パケットのシーケンス(34)にある、請求項15に記載のエンコーダ。
- 18ビデオ内容(16)の画像(18)のサブポーション(24)を単位にするビデオ内容(16)をビデオ・データ(22)に符号化するための方法であって、それぞれ各サブポーション(24)をビデオデータ(22)のパケットのシーケンス(34)の1つ以上のペイロード・パケット(32)に符号化し、それにより、パケットのシーケンス(34)がアクセスユニット(30)のシーケンスに分割され、各アクセスユニット(30)がビデオ内容(16)のそれぞれの画像(18)に関連するペイロード・パケット(32)を集め、この方法はパケットのシーケンス(34)にタイミング制御パケット(36)を分散し、それにより、タイミング制御パケット(36)はアクセスユニット(30)を復号化ユニット(38)に再分割し、それにより、少なくともいくつかのアクセスユニット(30)は複数の復号化ユニット(38)に再分割され、各タイミング制御パケット(36)は復号化ユニット(38)のためのデコーダバッファ検索時間の信号を送り、そのペイロード・パケット(32)はパケットのシーケンス(34)におけるそれぞれのタイミング制御パケット(36)に続き、復号化ユニット(38)のためのデコーダバッファ検索時間は、符号化ピクチャ・バッファからの検索時間および復号化ピクチャ・バッファからのすでに復号化された画像データの検索時間を含むように構成された、方法。
- 19ビデオ内容(16)の画像(18)のサブポーション(24)を単位にしてその中で符号化されるビデオ内容(16)を有するビデオ・データ(22)を復号化するためのデコーダであって、それぞれ各サブポーションをビデオ・データ(22)のパケットのシーケンス(34)の1つ以上のペイロード・パケット(32)に符号化し、パケットのシーケンス(34)がアクセスユニット(30)のシーケンスに分割され、それにより、各アクセスユニット(30)がビデオ内容(16)のそれぞれの画像(18)に関連するペイロード・パケット(32)を集め、デコーダはビデオ・データをバッファリングするための符号化ピクチャ・バッファおよびビデオ・データの復号化によりそこから得られるビデオ内容の再現をバッファリングするための復号化されたピクチャ・バッファを含み、パケットのシーケンスに分散されたタイミング制御パケット(36)を探し、アクセスユニット(30)をタイミング制御パケット(36)で復号化ユニット(38)に再分割し、それにより、少なくともいくつかのアクセスユニットが複数の復号化ユニットに再分割され、復号化ユニットを単位にするバッファを空にするように構成され、デコーダは、復号化ユニット(38)のための各タイミング制御パケット(36)によって信号の送られるデコーダバッファ検索時間を使用するように構成され、そのペイロード・パケット(32)はパケットのシーケンス(34)におけるそれぞれのタイミング制御パケット(36)に続き、それぞれの復号化ユニット(38)のためのデコーダバッファ検索時間は、符号化ピクチャ・バッファからの検索時間および復号化ピクチャ・バッファからのすでに復号化された画像データの検索時間を含む、デコーダ。
- 20ビデオ・データは、スケーラブルビデオ符号化(SVC)データである、請求項19に記載のデコーダ。
- 21各復号化ユニットは、サブピクチャを表し、符号化ピクチャ・バッファからの検索時間は、同じサブピクチャにおける先行するアクセスユニットの最近のタイミング制御パケットと関連し、復号化ユニットの符号化ピクチャ・バッファからの除去後、どれくらいのクロックが待つために刻んだかを特定する、請求項19に記載のデコーダ。
- 22サブポーションはスライスであり、ビデオ内容はビデオ・データに符号化され、スライス(24)の中の符号化順序を用いて、スライス(24)を単位にして、予測およびエントロピー符号化を用いて、ビデオ内容の画像が空間的に再分割されるタイル(70)の内部に予測符号化の予測および/またはエントロピー符号化の予測を制限し、そこにおいて、スライス(24)のシーケンスはペイロード・パケット(32)にパケット化され、そこにおいて、パケットのシーケンス(34)がタイル識別パケットとして1つのアクセスユニットのペイロード・パケットとの間でそこに分散するタイミング制御パケット(36)を有し、パケットのシーケンス(34)のそれぞれのタイル識別パケット(72)に直ちに続く1つ以上のペイロード・パケット(32)にパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認し、 デコーダは、タイル識別パケットに基づいて、タイル(70)を確認するように構成される、請求項19ないし請求項21のいずれかに記載のデコーダ。
- 23タイル識別パケットは、正確に直ちに続くペイロード・パケットにパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認する、請求項22に記載のデコーダ。
- 24タイル識別パケットは、現在のアクセスユニット(30)の端部の前までパケットのシーケンス(34)のそれぞれのタイル識別パケットに直ちに続く1つ以上のペイロード・パケット(32)にパケット化されたいくつかのスライス(24)によって覆われる1つ以上のタイル(70)を確認し、次のタイル識別パケットは、それぞれ、パケットのシーケンス(34)にある、請求項22に記載のデコーダ。
- 25ビデオ内容(16)の画像(18)のサブポーション(24)を単位にしてその中で符号化されるビデオ内容(16)を有するビデオ・データ(22)を復号化するための方法であって、それぞれ各サブポーションをビデオ・データ(22)のパケットのシーケンス(34)の1つ以上のペイロード・パケット(32)に符号化し、パケットのシーケンス(34)がアクセスユニット(30)のシーケンスに分割され、それにより、各アクセスユニット(30)がビデオ内容(16)のそれぞれの画像(18)に関連するペイロード・パケット(32)を集め、この方法は、ビデオ・データをバッファリングするための符号化ピクチャ・バッファおよびビデオ・データの復号化によりそこから得られるビデオ内容の再現をバッファリングするための復号化されたピクチャ・バッファを使用し、パケットのシーケンスに分散されたタイミング制御パケット(36)を探し、アクセスユニット(30)をタイミング制御パケット(36)で復号化ユニット(38)に再分割し、それにより、少なくともいくつかのアクセスユニットが複数の復号化ユニットに再分割され、復号化ユニットを単位にするバッファを空にし、デコーダは、復号化ユニット(38)のための各タイミング制御パケット(36)によって信号の送られるデコーダバッファ検索時間を使用するように構成され、そのペイロード・パケット(32)はパケットのシーケンス(34)におけるそれぞれのタイミング制御パケット(36)に続き、それぞれの復号化ユニット(38)のためのデコーダバッファ検索時間は、符号化ピクチャ・バッファからの検索時間および復号化ピクチャ・バッファからのすでに復号化された画像データの検索時間を含むように構成された、方法。
- 26コンピュータ上で動作するときに、請求項18または25に記載の方法を実行するためのプログラムコードを有する、コンピュータプログラム。
Independent claims26
290 paragraphs, as filed
0001The present invention relates specifically to a video data stream concept that is advantageous in connection with low latency applications.
0002HEVC [2] enables different means of High Level Syntax signaling to the application layer. This type of means is the NAL unit header, Parameter Sets and Supplemental Enhancement Information (SEI) Messages. The latter is not used in the decryption process. Other means of High Level Syntax signaling include transport protocols such as MPEG2 Transport Protocol [3] or Realtime Transport Protocol [4], and specifications specific to their payloads, such as H.264 / AVC [5]. Proposals for, extensible video coding (SVC) [6] or HEVC [7], etc. are of origin. Such transport protocols use a structure and mechanism similar to the High Level signaling of their respective application layer codec specifications, such as HEVC [2]. Level signaling can be introduced. One example of such signaling is the Payload Content Scalability Information (PACSI) NAL unit, as described in [6], which provides auxiliary information to the transport layer.
0003For parameter sets, HEVC includes a Video Parameter Set (VPS) that follows the most important stream information used by the application layer in one and center positions. In early methods, this information needed to be gathered from multiple Parameter Sets and NAL unit headers.
0004Prior to this application, along with the standard state of the Hypothetical Reference Decoder (HRD) Coded Picture Buffer (CPB) operation, and the definition of a decoding unit that represents the Dependent Slices subpictures and syntax as well as the Picture Parameter Set (PPS). All related syntaxes present in the Sequence Parameter Set (SPS) / Video Usability Information (VUI), Picture Timing SEI, and Buffering Period SEI were as follows:
0005To enable low-latency CPB operation at the sub-picture level, sub-picture CPB operation was proposed and incorporated into the HEVC draft standard 7JCTVC-I1003 [2]. Here, in particular, the decryption unit is defined in Section 3 of [2] as follows: Decryptor: An access unit or a subset of access units. If SubPicCpbFlag is equal to 0, then the decryption unit is an access unit. Otherwise, the decryption unit consists of one or more VCL NAL units in the access unit and associated non-VCL NAL units. For the first VCL NAL unit in the access unit, there is an associated non-VCL NAL unit, and the filler NAL unit, if any, immediately follows the first VCL NAL unit, followed by the first VCL NAL unit. Follows all non-VCL NAL units in the access unit that precedes it. For VCL NAL units that are not the first VCL NAL unit in the access unit, the associated non-VCL NAL unit is the filler data NAL unit that immediately follows the VCL NAL unit, if any.
0006"Timing of removal of decoding unit and decoding of decoding unit" is described in the standard set up to that time, and Annex C "Hypothential" Added to "reference decoder". Buffering along with HRD parameters in VUI to send sub-picture timing The period SEI and picture timing SEI messages have been extended to support decryption units as sub-picture units.
0007The Buffering period SEI message syntax of [2] is shown in Fig. 1.
0008When the NalHrdBpPresentFlag or VclHrdBpPresentFlag is equal to 1, the buffering period SEI message can be associated with any access unit in the bitstream, and the buffering period SEI message is with each RAP access device and the recovery point SEI message. Related to each access unit associated with.
0009In some applications, the frequent presence of buffering period SEI messages is preferred.
0010The buffering period is defined as a set of access units between two instances of a buffering SEI message in the decryption order.
0011The semantics were:
0012seq_parameter_set_id specifies a sequence parameter set that includes the sequence HRD attribute. The value of seq_parameter_set_id is equal to the value of seq_parameter_set_id in the image parameter set referenced by the primary coded image associated with the buffering period SEI message. The value of seq_parameter_set_id is in the range 0-31.
0013A rap_cpb_params_present_flag equal to 1 specifies the existence of the initial_alt_cpb_removal_delay [SchedSelIdx] and initial_alt_cpb_removal_delay_offset [SchedSelIdx] syntax elements. If it does not exist, the value of rap_cpb_params_present_flag is presumed to be equal to 0. If the associated image is neither a CRA image nor a BLA image, the value of rap_cpb_params_present_flag is equal to 0.
0014initial_cpb_removal_delay [SchedSelIdx] and initial_alt_cpb_removal_delay [SchedSelIdx] specify the initial CPB removal delay for ScheduleSelIdx-th-CPB. The syntax element has the bit length given by initial_cpb_removal_delay_length_minus1 + 1 and is in units of 90kHz clock. The value of the syntax element is not equal to 0 and does not exceed the CPB size time equivalent of 90000 * (CpbSize [SchedSelIdx] ÷ BitRate [SchedSelIdx]) at 90kHz clock time.
0015initial_cpb_removal_delay_offset [SchedSelIdx] and initial_alt_cpb_removal_delay_offset [SchedSelIdx] are used by the ScheduleSelIdx-th CPB to specify to the CPB the initial transport time of the encoded data unit. The syntax element has the bit length given by initial_cpb_removal_delay_length_minus1 + 1 and is in units of 90kHz clock. These syntax elements are not used by the decoder and are only needed for the Transport Scheduler (HSS).
0016On all encoded video sequences, the sum of initial_cpb_removal_delay [SchedSelIdx] and initial_cpb_removal_delay_offset [SchedSelIdx] is constant for each value of ScheduleSelIdx, and the sum of initial_alt_cpb_removal_delay [SchedSelIdx] and initial_alt_cp Is.
0017The image timing SEI message syntax of [2] is shown in Fig. 2.
0018The syntax of the image-timed SEI message relied on the contents of the sequence parameter set working for the coded image associated with the image-timed SEI message. However, unless there is a buffering period SEI message within the access unit prior to the image timing SEI message of the IDR or BLA access unit, the relevant parameter set (and the IDR or BLA image that is not the first image in the bitstream). Therefore, the determination that the coded image is an IDR image or a BLA image) does not occur until the first coded slice NAL unit of the coded image is decoded. Since the coded slice NAL unit of the coded image follows the image timing SEI message in NAL unit order, the RBSP contains the image timing SEI message until the decoder determines the parameters of the sequence parameters that work for the coded image. In some cases it may be necessary to store the image timing SEI message parsing.
0019The existence of image timing SEI messages in bitstreams is defined as follows. -If CpbDpbDelaysPresentFlag is equal to 1, one image timing SEI message is present in every access unit of the coded video sequence. -Otherwise, the image timing SEI message is not present in any access unit of the coded video sequence (CpbDpbDelaysPresentFlag equals 0).
0020The semantics were set out as follows:
0021cpb_removal_delay is the most recent buffering period in the preceding access unit before removing the access unit data associated with the image timing SEI message from the buffer How much clock after the removal of the access unit associated with the SEI message from the CPB. Specify whether to wait for a tick. This value is also used to calculate the earliest possible time for access unit data to appear in the CPB for HSS. The syntax element is a fixed length code whose bit length is given by cpb_removal_delay_length_minus1 + 1. cpb_removal_delay is a modulo remainder 2 (cpb_removal_delay_length_minus1 + 1) counter.
0022The value of cpb_removal_delay_length_minus1 that determines the length (in bits) of the syntax element cpb_removal_delay is the value of cpb_removal_delay_length_minus1 encoded in the sequence parameter set that works for the main encoded image associated with the image timing SEI message. However, however, cpb_removal_delay specifies the number of clock ticks associated with the removal time of the preceding access unit containing the buffering period SEI message, which may be an access unit of a different encoded video sequence.
0023dpb_output_delay is used to calculate the DPB output time for an image. It specifies how many clock ticks to wait before the decoded image is output from the DPB after the removal of the last decoder of the access device from the CPB.
0024The image is not removed from the DPB at its output time, when it is still marked as "used for short-term references" or "used for long-term references".
0025Only one dpb_output_delay is specified for the decoded image.
0026The length of the syntax element dpb_output_delay is given in bits by dpb_output_delay_length_minus1 + 1. When sps_max_dec_pic_buffering [max_temporal_layers_minus1] is equal to 0, dpb_output_delay will be equal to 0.
0027The output time obtained from the dpb_output_delay of some images output from the output timing tuning decoder precedes the output time obtained from the dpb_output_delay of all images in some next coded video sequence in the decoding sequence.
0028The image output order specified by the value of this syntax element is the same order specified by the value of PicOrderCntVal.
0029The output time obtained from dpb_output_delay is the same code for images that are not output by the "bumping" process because in the decoding order no_output_of_prior_pics_flag precedes an IDR or BLA that is equal to or is presumed to be equal to 1. It increases with increasing PicOrderCntVal values associated with all images in the video sequence.
0030num_decoding_units_minus1 plus1 specifies the number of decoding units in the access unit to which the image timing SEI message is associated. The value of num_decoding_units_minus1 is included in the range from 0 to PicWidthInCtbs * PicHeightInCtbs-1.
0031num_nalus_in_du_minus1 [i] plus1 specifies the number of NAL units in the i-th decoding unit of the access unit to which the image timing SEI message is associated. The value of num_nalus_in_du_minus1 [i] is in the range from 0 to PicWidthInCtbs * PicHeightInCtbs-1.
0032The first decoding unit of the access unit consists of a continuous NAL unit of the first num_nalus_in_du_minus1 [0] in the decoding order in the access unit. The i-th (i is greater than 0) decoding unit of the access unit is a continuous NAL of num_nalus_in_du_minus1 [i] + 1 immediately following the last NAL unit in the decoding unit before the access unit in the decoding order. Consists of units. Each decryption unit has at least one VCL NAL device. All non-VCL NAL units associated with a VCL NAL unit are included in the same decryption unit.
0033du_cpb_removal_delay [i] is the most recent buffering period in the previous access unit before removing the i-th decryption unit in the access unit associated with the image timing SEI message from the CPB in the access unit associated with the SEI message. Specifies how many subpicture clock ticks to wait after removing the first decoding unit from the CPB. This value is also used to calculate the earliest possible time for the decryption unit data to arrive at the CPB for HSS. The syntax element is a fixed length code whose bit length is given by cpb_removal_delay_length_minus1 + 1. du_cpb_removal_delay [i] is the remainder modulo the 2 (cpb_removal_delay_length_minus1 + 1) counter.
0034The value of cpb_removal_delay_length_minus1 that determines the length (in bits) of the syntax element du_cpb_removal_delay [i] is the value of cpb_removal_delay_length_minus1 encoded in the sequence parameter set that works for the encoded image associated with the image timing SEI message. However, du_cpb_removal_delay [i] identifies the number of sub-picture clock ticks in relation to the removal time of the first decryption unit in the preceding access unit containing the buffering period SEI message, which is a different encoded video. -It may be a sequence access unit.
0035Some information was included in the VUI syntax in [2]. The VUI parameter syntax of [2] is shown in Fig. 3. The HRD parameter syntax of [2] is shown in Fig. 4. The semantics were set out as follows:
0036A sub_pic_cpb_params_present_flag equal to 1 indicates that there is a sub-picture level CPB removal delay parameter and that the CPB can operate at the access unit level or sub-picture level. A sub_pic_cpb_params_present_flag equal to 0 indicates that there is no sub-picture level CPB removal delay parameter and the CPB operates at the access unit level. In the absence of sub_pic_cpb_params_present_flag, its value is presumed to be equal to 0.
0037num_units_in_sub_tick is the number of clock time units operating at the frequency time_scale Hz corresponding to one increase in the subpicture clock tick counter (called the subpicture clock tick). num_units_in_sub_tick is greater than 0. The sub-picture clock tick is the minimum interval of time that can be represented in the coded data when sub_pic_cpb_params_present_flag is equal to 1.
0038A tiles_fixed_structure_flag equal to 1 has the same values as the syntax elements num_tile_columns_minus1, num_tile_rows_minus1, uniform_spacing_flag, column_width [i] row_height [i] and loop_filter_across_tiles_enabled_flag, where each image parameter set functioning in the coded video sequence is shown. Show that. A tiles_fixed_structure_flag equal to 0 indicates whether tile syntax elements in different image parameter sets can or cannot have the same value. In the absence of the tiles_fixed_structure_flag syntax element, it is presumed to be equal to 0.
0039Signaling tiles_fixed_structure_flag equal to 1 guarantees to the decoder that each image in the encoded video sequence has the same number of tiles distributed in the same way that may help load distribution in the case of multithreaded decoding. Is.
0040The filler data in [2] was signaled using the filter data RBSP syntax shown in FIG.
0041The virtual reference decoder in [2] used to check bitstream and decoder compatibility was defined as follows:
0042The two types of bitstreams are subject to HRD compliance checking for this recommendation | international standards. The first type of bitstream, called a type I bitstream, is a NAL unit stream that contains only the VCL NAL unit and the filler data NAL unit for all access units of the bitstream. A second type of bitstream, called a type II bitstream, contains at least one of the following, in addition to the VCL NAL unit and filler data NAL unit for all access units in the bitstream: --Additional non-VCL NAL units other than filler data NAL units, --All leading_zero_8bits, zero_byte, start_code_prefix_one_3bytes and trailing_zero_8bits syntax elements that form a byte stream from the NAL unit stream.
0043Figure 6 shows the types of bitstream conformance points checked by the HRD in [2].
0044Two HRD parameter sets (NAL HRD and VCL HRD parameters) are used. The HRD parameter set is signaled by video usefulness information, which is part of the sequence parameter syntax structure.
0045All sequence and image parameter sets referred to the VCL NAL unit, as well as the corresponding buffering period and image timing SEI messages, are sent to the HRD in a timely manner, in the bitstream, or elsewhere. Communicated by means.
0046For the "existence" of non-VCL NAL units when those NAL units (or just some of them) are transmitted to the decoder (or to HRD) by other means not specified by this recommendation | international standards. Specifications are also met. To count the bits, only the appropriate bits that actually exist in the bitstream are counted.
0047As an example, synchronization of non-VCL NAL units that have NAL units that are transmitted in a way other than their presence in the bitstream and are present in the bitstream is achieved by showing two points in the bitstream, and they There was a non-VCL NAL unit in the bitstream between and had an encoder that was determined to carry it in the bitstream.
0048The representation of the contents of a non-VCL NAL unit uses the same syntax specified in this attachment when the contents of the non-VCL NAL unit are communicated for the application by some means other than being within the scope of the bitstream. Does not need to be done.
0049It should be noted that when the HRD information is contained within the bitstream, the match between the bitstream and this dependent request can only be verified based on the information contained within the bitstream. When the HRD information is not present in the bitstream and the HRD data is provided by some other means not specified in this Recommendation | International Standard, as is the case with all "standalone" Type I bitstreams. , Matches can only be confirmed.
0050The HRD includes a coded image buffer (CPB), an instantaneous decoding process, a decoded image buffer (DPB) and output cropping, as shown in FIG.
0051The CPB size (number of bits) is CpbSize [SchedSelIdx]. The DPB size (number of image storage buffers) for temporal layer X is sps_max_dec_pic_buffering [X] for each X in the range from 0 to sps_max_temporal_layers_minus1.
0052The variable SubPicCpbPreferredFlag is specified by external means, or 0 if not specified by external means.
0053The variable SubPicCpbFlag is pulled out as follows: SubPicCpbFlag = SubPicCpbPreferredFlag && sub_pic_cpb_params_present_flag
0054If SubPicCpbFlag is equal to 0, the CPB operates at the access unit level and each decryption unit is an access unit. Otherwise, the CPB operates at the subpicture level and each decryption unit is a subset of the access unit.
0055HRD works as follows. The data associated with the decryption unit that flows into the CPB according to the specified arrival schedule is distributed by the HSS. The data associated with each decryption unit is deleted and immediately decoded by a momentary decoding process during the CPB removal time. Each decrypted image is placed in the DPB. The decoded image is removed from the DPB at a later DPB output or when it is no longer needed for an inter-prediction reference.
0056The HRD is initialized to be specified by the buffer period SEI. The timing of removing the decoding unit from the CPB and the timing of outputting the decoded image from the DPB are specified in the image timing SEI message. All timing information for a particular decryption unit arrives before the decryption unit's CPB removal time.
0057HRD is used to check bitstream and decoder matches.
0058Matching is guaranteed under the assumption that all framerates and clocks used to generate the bitstream exactly match the values signaled in the bitstream, while in a real system each of these Can be signaled or varied from the specified value.
0059All calculations are done with real values and as a result rounding errors cannot be propagated. For example, the number of bits in the CPB just before or after the removal of the decryption unit is not necessarily an integer.
0060The variable tc is pulled out as follows and is called a clock tick: tc = num_units_in_tick time_scale The variable tc_sub is pulled out as follows and is called a sub-picture clock tick: tc_sub = num_units_in_sub_tick time_scale
0061The following is specified to represent the constraint: --Let access unit n be the nth access unit in the decoding sequence with the first access knit, which is access unit 0. --Suppose the image n is a coded image or a decoded image of the access unit n. --In the decoding order having the first decoding unit which is the decoding unit 0, the decoding unit m is set as the mth decoding unit.
0062In [2], the slice header syntax allows for so-called dependent slices.
0063FIG. 8 shows the slice header syntax of [2].
0064The slice header semantics were defined as follows:
0065Dependent_slice_flag equal to 1 is the value of each slice header syntax element not shown in the corresponding slice header in the previous slice containing a coded tree block whose coded tree block address is SliceCtbAddrRS. Indicates that it is presumed to be equal to the value. When not shown, the value of dependent_slice_flag is presumed to be equal to 0. The value of dependent_slice_flag is equal to 0 when SliceCtbAddrRS is equal to 0.
0066slice_address specifies the address of the slice granular resolution at which the slice starts. The length of the slice_address syntax element is (CEil (Log2 (PicWidthInCtbs * PicHeightInCtbs)) + Slice Granularity) bits.
0067The variable SliceCtbAddrRS defines the coded tree block in which the slice starts in the coded tree block raster scan order and is pulled out as follows. SliceCtbAddrRS = (slice_address >> Slice Granularity)
0068The variable SliceCbAddrZS is derived as follows, specifying the address of the first coded block of the slice with the minimum coded block particle size of the z-scan order. SliceCbAddrZS = slice_address << ((log2_diff_max_min_coding_block_size = Slice Granularity) << 1)
0069Slice decoding begins with the largest possible coding unit of the slice start coordinates.
0070first_slice_in_pic_flag indicates whether the slice is the first slice of the image. If first_slice_in_pic_flag is equal to 1, the variables SliceCbAddrZS and SliceCtbAddrRS are both set to 0 and decoding begins with the first coded tree block of the image.
0071pic_parameter_set_id defines the built-in image parameters. The value of pic_parameter_set_id contains the range 0-255.
0072num_entry_point_offsets specifies the number of entry_point_offset [i] syntax elements in the slice header. When tiles_or_entropy_coding_sync_idc is equal to 1, the value of num_entry_point_offsets contains the range from 0 to (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1). When tiles_or_entropy_coding_sync_idc is equal to 2, the value of num_entry_point_offsets contains the range from 0 to PicHeightInCtbs = 1. When not shown, the value of num_entry_point_offsets is estimated to be equal to 0.
0073offset_len_minus1 plus1 specifies the length of the entry_point_offset [i] syntax element in bits.
0074entry_point_offset [i] specifies the offset of the i-th entry point in bytes and is represented by the offset_len_minus1 plus 1 bit. The coded slice data after the slice header consists of a num_entry_point_offsets + 1 subset containing the range 0 to num_entry_point_offsets. Subset 0 consists of bytes containing 0 to entry_point_offset [0] -1 in the coded slice data k, where k contains the range 1 to num_entry_point_offset-1 and from the entry_point_offset [k-1] in the coded slice data. It consists of bytes containing entry_point_offset [k] + entry_point_offset [k-1] -1, and the last subset (with subset index equal to num_entry_point_offsets) consists of the remaining bytes of encoded slice data.
0075When tiles_or_entropy_coding_sync_idc is equal to 1 and num_entry_point_offsets is greater than 0, each subset contains exactly all the coding bits of one tile, and the number of subsets (ie, the value of num_entry_point_offsets + 1) is less than or equal to the number of tiles in the slice. is there. When tiles_or_entropy_coding_sync_idc is equal to 1, each slice must contain a subset of one tile or an integer value of the full tile (entry point case signaling is unnecessary).
0076When tiles_or_entropy_coding_sync_idc is equal to 2 and num_entry_point_offsets is greater than 0, each subset k with k containing the range 0 to num_entry_point_offsets-1 ensures that each subset k contains all the coding bits of one row of the coding tree block The last subset (with a subset index equal to num_entry_point_offsets) contains all the coded bits of the remaining coded blocks contained in the slice, and the remaining coded blocks are exactly one line of the coded tree block. It consists of one of a subset of one row of coded tree blocks, and the number of subsets (ie, the value of num_entry_point_offsets + 1) is equal to the number of rows in the sliced coded tree block, and the sliced coded tree. A subset of one row of blocks is also counted.
0077When tiles_or_entropy_coding_sync_idc is equal to 2, the slice contains many rows of coded tree blocks and a subset of the rows of coded tree blocks. For example, if the slice contains two and a half rows of coded tree blocks, the number of subsets (ie, the value of num_entry_point_offsets + 1) would be equal to 3.
0078Figure 9 shows the image parameter set RBSP syntax in [2] and the image parameter set RBSP semantics in [2] and is defined as follows:
0079Dependent_slice_enabled_flag equal to 1 specifies the existence of the syntax element dependent_slice_flag in the slice header for the coded image with reference to the image parameter set. Dependent_slice_enabled_flag equal to 0 refers to the image parameter set to identify the lack of the slice header syntax element dependent_slice_flag for the coded image. When tiles_or_entropy_coding_sync_idc is equal to 3, the value of dependent_slice_enabled_flag will be equal to 1.
0080Tiles_or_entropy_coding_sync_idc equal to 0 indicates that each image has only one tile with reference to the image parameter set, and the first coded tree in the row of the coded tree block in each image in the image parameter set. The values of cabac_independent_flag and dependent_slice_flag for encoded images with reference to the image parameter set, which have a specific synchronization process for context changes caused before decoding the block, are both not equal to 1. Let's go.
0081A slice is an entropy slice when both cabac_independent_flag and depedent_slice_flag are equal to 1 for the slice. ]
0082Tiles_or_entropy_coding_sync_idc equal to 1 indicates that each image has one or more tiles with reference to the image parameter set, and a row of coded tree blocks in each image with reference to the image parameter set. There will be no specific synchronization process for context changes that will be pulled out before decoding the first encoded tree box of, and the values of cabac_independent_flag and dependent_slice_flag for the encoded image with reference to the image parameter set will be. , Both would not be equal to 1.
0083A tiles_or_entropy_coding_sync_idc equal to 2 indicates that each image has only one tile with reference to the image parameter set, and the first row of the encoded tree block row in each image with reference to the image parameter set. A specific synchronization process for context changes is pulled out before decoding the coded tree block of the two coded tree blocks in the row of the coded tree box in each image with reference to the image parameter set. A specific storage process for context changes will be derived after decoding the image, and the values of cabac_independent_flag and dependent_slice_flag for the encoded image with reference to the image parameters will not both be 1.
0084Tiles_or_entropy_coding_sync_idc equal to 3 refers to the image parameter set to clarify that each image has only one tile, and refers to the image parameter set to be the first in the row of the encoded tree block in each image. There will be no specific synchronization process for context changes that will be pulled out before decoding the coded tree block of the image, and the values of cabac_independent_flag and dependent_slice_flag for the coded image with reference to the image parameter set are both. Will be equal to 1.
0085When dependent_slice_enabled_flag is equal to 0, tiles_or_entropy_coding_sync_idc will not be equal to 3.
0086It is a bitstream conformance requirement that the value of tiles_or_entropy_coding_sync_idc is common to all image parameter sets invoked in the coded video sequence.
0087For each slice with reference to the image parameter set, if tiles_or_entropy_coding_sync_idc is equal to 2 and the first coding block of the slice is not the first coding tree block in the row of the coding tree block, then the slice's The last coded block belongs to the same row of coded tree blocks as the first coded block of the slice.
0088num_tile_columns_minus1 plus 1 identifies the number of tile columns that divide the image.
0089num_tile_rows_minus1 plus 1 identifies the number of rows of tiles that divide the image. When num_tile_columns_minus1 is equal to 0, num_tile_rows_minus1 will not be equal to 0.
0090Uniform_spacing_flag equal to 1 defines its column boundaries, and similarly row boundaries are evenly distributed throughout the image. Uniform_spacing_flag equal to 0 defines its column boundaries, and similarly, row boundaries are not evenly distributed throughout the image, but the syntax elements column_width [i] and row_height [i] are used to explicitly signal.
0091column_width [i] specifies the width of the column of the i-th tile in the coded tree block as a unit.
0092row_height [i] specifies the row height of the i-th tile in the coded tree block.
0093The vector colWidth [i] defines and contains the width of the i-th tile column in units of CTB with column i extending from 0 to num_tile_columns_minus1.
0094The vector CtbAddrRStoTS [ctbAddrRS] is an index ctbAddrRS that extends from 0 to (picHeightInCtbs * picWidthInCtbs) -1 and defines and contains conversations from the CTB address in raster scan order to the CTB address in tile scan order. The vector CtbAddrTStoRS [ctbAddrTS] is an index ctbAddrTS extending from 0 to (picHeightInCtbs * picWidthInCtbs) -1 that defines and contains the conversation from the CTB address in the tile scan order to the CTB address in the raster scan order.
0095The vector TileId [ctbAddrTS] is a ctbAddrTS that extends from 0 to (picHeightInCtbs * picWidthInCtbs) -1 and defines and contains the CTB address-to-tile id conversation in the tile scan order.
0096The values of colWidth, CtbAddrRStoTS, CtbAddrTStoRS and TileId are extracted as inputs by calling CTB raster and tile scanning as specified in subclose 6.5.1, and the output is assigned to colWidth, CtbAddrRStoTS and TileId.
0097The value of ColumnWidthInLumaSamples [i], which specifies the width of the column of the i-th tile in luma samples, is set equal to colWidth [i] << Log2CtbSize.
0098X spanning 0 to picWidthInMinCbs-1 and y spanning 0 to picHeightInMinCbs-1 specify the minimum CBs-based location (x, y) to the minimum CB address in the z-scan sequence. The array MinCbAddrZS [x] [y] is a Z-scanning array as specified in sub-close 6.5.2 with Log2MinCbSize, Log2CtbSize, PicHeightInCtbs, PicWidthInCtbs, and vector CtbAddrRStoTS so that inputs and outputs are assigned to MinCbAddrZS. Obtained by calling the initialization process.
0099A loop_filter_across_tiles_enabled_flag equal to 1 indicates that the in-loop filtering operation will be performed across the tile boundaries. A loop_filter_across_tiles_enabled_flag equal to 0 indicates that the in-loop filtering operation will not be performed on the entire tile boundary. In-loop filtering operations include deblocking filters, sample adaptive offsets, and adaptive loop filter operations. If not, the value of loop_filter_across_tiles_enabled_flag is presumed to be equal to 1.
0100A cabac_independent_flag equal to 1 indicates that the CABAC decoding of the slice's coded block is independent of any aspect of the previously decoded slice. A cabac_independent_flag equal to 0 indicates that the CABAC decoding of the slice's coded block depends on the aspect of the previously decoded slice. If not, the value of cabac_independent_flag is presumed to be equal to 0.
0101The derivation process for the availability of a coded block with a minimal coded block address was described as follows: The inputs to this process are as follows: -z-Minimum coded block address for scan order minCbAddrZS -z-Current minimum coded block address for scan order currMinCBAddrZS
0102The output of this process is the availability of a coded block with the minimum coded block address cbAddrZS in the z-scan sequence cbAvailable.
0103Note that the meaning of validity is determined when this process is called.
0104Note that any coded block, regardless of its size, is associated with the minimum coded block address, which is the address of the coded block with the minimum coded block size in the z-scan order. I want to be.
0105-CbAvailable is set to FALSE if one or more of the following situations are true: -minCbAddrZS is less than 0 -minCbAddrZS is greater than currMinCBAddrZS -A coded block with a minimal coded block address minCbAddrZS belongs to a coded block with the current minimal coded block address currMinCBAddrZS and a code with the current minimal coded block address currMinCBAddrZS. A slice that differs from the dependent_slice_flag of the slice that contains the coded block is equal to 0. -The coded block with the minimum coded block address minCbAddrZS is contained in a different tile than the coded block with the current minimal coded block address currMinCBAddrZS. -Otherwise, cbAvailable is set to TRUE.
0106The CABAC parsing process for sliced data in [2] was as follows: This process is called when parsing a syntax element that has a descriptor ae (v).
0107The input to this process is a request for the value of the syntax element and the value of the previously parsed syntax element.
0108The output of this process is the value of the syntax element.
0109The initialization process of the CABAC parsing process is called when you start parsing the sliced data of a slice.
0110The minimum coded block address of the coded tree block containing the spatially adjacent block T (Figure 10a), ctbMinCbAddrT, is the location of the upper left luma sample of the current coded tree block (x0, It is pulled out using y0). x = x0 + 2 << Log2CtbSize-1 y = y0-1 ctbMinCbAddrT = MinCbAddrZS [x >> Log2MinCbSize] [y >> Log2MinCbSize]
0111The variable availableFlagT is obtained by calling the coded block validity derivation process with ctbMinCbAddrT as input.
0112The following steps apply when you start parsing the coded tree. 1. The arithmetic decoding engine is initialized as follows. -If CtbAddrRS is equal to slice_address, dependent_slice_flag is equal to 1 and entropy_coding_reset_flag is equal to 0, and the following applies: -The synchronization process of the CABAC parsing process is called by TableStateIdxDS and TableMPSValDS as inputs. -The decryption process for determining the binary before termination is called, followed by the initialization process for arithmetic decoding.-Otherwise , if tiles_or_entropy_cod ing_sync_idc is equal to 2 and CtbAddrRS% PicWidthInCtbs is equal to 0, then the following applies: -When availableFlagT is equal to 1, the synchronization process of the CABAC parsing process is called by TableStateIdxWPP and TableMPSValWPP as inputs. -The decryption process for determining the binary before termination is called, followed by the initialization process for the arithmetic decoding engine.
01132. When cabac_independent_flag is equal to 0 and dependent_slice_flag is equal to 1, or tiles_or_entropy_coding_sync_idc is equal to 2, the storage process is applied as follows. When -tiles_or_entropy_coding_sync_idc is equal to 2 and CtbAddrRS% PicWidthInCtbs is equal to 2, the storage process of the CABAC parsing process is called by TableStateIdxWPP and TableMPSValWPP as output. -When cabac_independent_flag is equal to 0, dependent_slice_flag is equal to 1, and end_of_slice_flag is equal to 1, the storage process of the CABAC parsing process is called as output by TableStateIdxDS and TableMPSValDS.
0114Parsing of syntactic elements proceeds as follows:
0115Binarization is obtained for each requested value of the syntax element.
0116The binarized and parsed sequence of bins for syntactic elements determines the decryption process flow.
0117For each binarization bin of the syntax element indexed by the variable binIdx, the context index ctxIdx is obtained.
0118For each ctxIdx, the arithmetic decoding process is started.
0119The resulting sequence (b0 ... bbinIdx) of the parsed bins is compared to the set of bin strings given by the binarization process after decoding each bin. When a sequence matches a bin in a given set, the corresponding value is assigned to the syntax element.
0120If the request for the value of the syntax element is equal to the decryption value of the syntax elements pcm-flag and pcm_flag, the decoding engine initializes any pcm_alignment_zero_bit after decoding num_subsequent_pcm and all pcm_sample_luma and pcm_sample_chroma data.
0121The following issues have arisen in the design frameworks described so far.
0122The timing of the decoding unit needs to be known before encoding and transmitting the data in low latency scenarios, the NAL unit is already sent by the encoder, while the encoder is the image, i.e. the other subpicture. The decoding unit is still encoded. This is because the NAL unit order in the access unit only allows SEI messages that precede the VCL (video-coded NAL unit) in the access unit, and in such low latency scenarios the encoder is in the decoding unit. This is because the non-VCL NAL unit must already be in communication, i.e. transmitted, when coding has begun. FIG. 10b illustrates the structure of the access unit described in [2]. [2] has not yet identified the end of the sequence or stream, so their presence in the access unit was tentative.
0123Also, the number of NAL units associated with the subpicture must be known in advance in low latency scenarios, which is that the image timing SEI contains and follows this information, and the encoder encodes the actual image. This is because it must be sent before it can be started. The application designer reluctantly inserts a filter data NAL unit for which there is potentially no filter data matching the NAL unit number, but it is sent per image timing SEI decoding unit and this information on the subpicture level. This is because it requires a means of transmitting. It holds for sub-picture timing that is currently determined in the presence of the access unit by the parameters given in the timing SEI message.
0124In addition, a further drawback of draft specification [2] involves a large number of sub-picture level signaling required for special use, such as ROI signaling or tile-dimensional signaling.
<p num="0125"><nplcit num="1"><text>Thomas Wiegand, Gary J. Sullivan, Gisle Bjontegaard, Ajay Luthra, "Overview of the H.264 / AVC Video Coding Standard", IEEE Trans. Circuits Syst. Video Technol., Vol. 13, N7, July 2003.</text></nplcit><nplcit num="2"><text>JCT-VC, "High-Efficiency Video Coding (HEVC) text specification Working Draft 7", JCTVC-I1003, May 2012.</text></nplcit><nplcit num="3"><text>ISO / IEC 13818-1: MPEG-2 Systems specification.</text></nplcit><nplcit num="4"><text>IETF RFC 3550 --Real-time Transport Protocol.</text></nplcit><nplcit num="5"><text>Y.-K. Wang et al., "RTP Payload Format for H.264 Video", IETF RFC6184, http://tools.ietf.org/html/</text></nplcit><nplcit num="6"><text>S. Wenger et al., "RTP Payload Format for Scalable Video Coding", IETF RFC6190, http://tools.ietf.org/html/rfc6190</text></nplcit><nplcit num="7"><text>T. Schierl et al., "RTP Payload Format for High Efficiency Video Coding", IETF internet draft, http://datatracker.ietf.org/doc/draft-schierl-payload-rtp-h265/</text></nplcit></p>
<p num="0126"> The issues outlined above are not specific to HEVC standards. Rather, this challenge arises in connection with other video codecs as well. More generally, FIG. 11 shows a video transmission in which a pair of encoders 10 and a decoder 12 are connected via a network 14 to transmit the video 16 from the encoder 10 to the decoder 12 with a short inter-terminal delay. Show the landscape. The issues already outlined above are: Encoder 10 substantially, but not necessarily, follows frame 18 reproduction sequence 20 and within each frame 18, for example, how many, such as raster scans with or without frame 18 tile-sectioning methods. The sequence of frame 18 of video 16 is encoded according to a particular decoding sequence transmitted through the frame area of frame 18 in that particular way. The decoding order is, for example, predictive and / or entropy coding, that is, the validity of information about the spatially and / or temporally adjacent parts of the available video 16 that serve as the basis for predictive or contextual selection. Controls the validity of the information for the coding technique used by the encoder 10. Even if encoder 10 may be able to use parallel processing to encode frame 18 of video 16, encoder 10 inevitably encodes a specific frame 18 like the current frame. It takes some time to do. FIG. 11 shows, for example, the moment when the encoder 10 has already finished coding part 18a of the current frame 18, while the other part 18b of the current frame 18 has not yet been coded. Since the encoder 10 has not yet encoded part 18b, the encoder 10 is currently in order to achieve the optimum conditions for how the available bit rates for encoding the current frame 18 are, for example, in terms of rate / distortion recognition. It is not possible to predict whether it should be spatially distributed through frame 18. Therefore, encoder 10 only has two choices. In any case, the network 14 is in the form of a coded picture buffer search time so that any transmission of the slice packet of the current coded frame 18 can be utilized before the end of its coding. The bit rate associated with each slice packet of this type should be informed. However, as mentioned above, the encoder 10 can vary the bit rate distributed over frame 18 by individually determining the decoder buffer search time for the sub-picture area according to the current version of HEVC. Nevertheless, the encoder 10 needs to send or send such information to the decoder 12 over the network 14 at the beginning of each access unit that collects all the data about the current frame 18. Encoder 10 prompts you to choose from the options outlined, one leading low delay but bad rate / distortion, and the other leading optimal rate / distortion but increased inter-terminal delay.</p><p num="0127"> Thus, so far, the encoder has allowed the start of transmission of a packet for part 18a of the current frame prior to encoding the remaining part 18b of the current frame, and the decoder sends from encoder 12 to decoder 14. No video codec has enabled low latency performance to take advantage of this intermediate transmission of packets for spare portion 18a by network 16 that follows the decryption buffer lookup timing transmitted within the video data stream. For example, applications that utilize this type of low latency as an example include industrial applications such as workpieces or production monitoring for automation or inspection purposes, etc. So far, the current frame allows intermediate network entities in network 16 to gather this kind of information from packets, a data stream that does not require deep inspection inside the slice syntax. There is also not enough solution to inform the relevant decryptor of the packet to the tiles that are configured and interested in the area (area of interest).</p><p num="0128"> Therefore, an object of the present invention is a video data stream code that is more efficient in allowing low terminal-to-terminal delay and / or makes it easier to identify parts of the data stream in a region of interest or in a particular tile. It is to provide the concept of conversion.</p>
<p num="0129"> This purpose is achieved by the subject matter of the independent claims of the claims.</p><p num="0130"> One idea on which this application is based is that decoder search timing information, ROI information and tile identification information are communicated within the video data stream at a level that allows easy access by network entities such as MANEs or decoders. This means that in order to reach this type of level, this type of information is transmitted within the video data stream via packets that are distributed among the packets of the access unit of the video data stream. It means that you have to. According to the embodiment, the distributed packets are among the removable packet types, i.e., the removal of these distributed packets is the video content carried entirely through the video data stream. Maintain the ability of the decoder to recover.</p><p num="0131"> According to aspects of the present application, achieving low inter-terminal delay is a decoder for the decoding unit formed by the payload packet following each timing control packet of the video data stream within the current access unit. It becomes more effective by using distributed packets to convey information during the buffer search time. This measurement allows the encoder to determine the decoder buffer search time on the fly while encoding the current frame, and is actually already encoded in the payload packet while encoding the current frame. , Spent on a portion of the current frame transmitted, or, on the other hand, to the current frame on the rest of the current frame that was preceded by a timing control packet and therefore not yet encoded. It will be possible to continuously determine the bitrates to which the distribution of the remaining available bitrates will be adapted. With this measurement, the available bitrates are effectively utilized, and the delay is nevertheless kept shorter because the encoder does not have to wait for the current frame to finish encoding completely.</p><p num="0132"> According to a further aspect of the application, the packets distributed in the payload packets of the access unit are used to convey information to the region of interest, thereby requiring inspection of intermediate payload packets, as described above. No network entity allows easy access to this information. In addition, the encoder does not need to determine the subpotion and the current frame subdivision for each payload packet in advance, and can flexibly determine which packet belongs to the ROI while encoding the current frame. I can still do it. Further, according to the embodiment in which the distributed packet is a mobile packet type, the ROI information is ignored by the recipient of the video data stream who is not interested in or cannot process the ROI information. Can be done.</p><p num="0133"> Similar ideas are utilized in this application according to other aspects in which distributed packets communicate which tile a particular packet belongs to within the scope of an access unit.</p><p num="0134"> An advantageous embodiment of the present invention is the subject of dependent claims. Preferred examples of the present application are described in more detail with reference to the following figures, with FIGS. 1-10b showing the current state of HEVC.</p>
0135<figref num="1">Figure 1 shows the buffer period SEI message syntax.</figref><figref num="2">Figure 2 shows the image timing SEI message syntax.</figref><figref num="3a">Figure 3a shows the VUI parameter syntax.</figref><figref num="3b">Figure 3b shows the VUI parameter syntax.</figref><figref num="4">Figure 4 shows the HRD parameter syntax.</figref><figref num="5">Figure 5 shows the filler data RBSP syntax.</figref><figref num="6">Figure 6 shows the structure of the byte stream and NAL unit stream for HRD match checking.</figref><figref num="7">Figure 7 shows the HRD buffer model.</figref><figref num="8">Figure 8 shows the slice header syntax.</figref><figref num="9">Figure 9 shows the image parameter set RBSP syntax.</figref><figref num="10a">FIG. 10a illustrates a spatially adjacent coding tree block T that is optionally used to initiate the coding tree block utilization guidance process associated with the current coding tree block.</figref><figref num="10b">Figure 10b shows the definition of the structure of the access unit.</figref><figref num="11">FIG. 11 schematically shows a pair of encoders and decoders connected over a network to illustrate the challenges that are occurring in video data stream transmission.</figref><figref num="12">FIG. 12 shows a schematic block diagram of an encoder according to an embodiment using a timing control packet.</figref><figref num="13">FIG. 13 shows a flow diagram illustrating an operating mode of the encoder of FIG. 12 according to an embodiment.</figref><figref num="14">FIG. 14 shows a block diagram of an embodiment of a decoder to illustrate its function in relation to the video data stream generated by the encoder according to FIG.</figref><figref num="15">FIG. 15 shows a schematic block diagram illustrating an encoder, network entity and video data stream according to a further embodiment using ROI packets.</figref><figref num="16">FIG. 16 shows a schematic block diagram illustrating encoders, network entities and video data streams according to a further embodiment using tile identification packets.</figref><figref num="17">FIG. 17 shows the structure of the access unit according to the embodiment. The dotted line reflects the case of any slice prefix NAL unit.</figref><figref num="18">Figure 18 shows the use of tiles in the area of interest signaling.</figref><figref num="19">Figure 19 shows the first simple syntax / version 1.</figref><figref num="20">Figure 20 shows an extended syntax / version 2 that includes tile_id signaling, decryption unit start identifier, slice prefix ID, and slice header data that differ from the SEI message concept.</figref><figref num="21">Figure 21 shows the NAL unit type code and NAL unit type class.</figref><figref num="22">Figure 22 shows the possible syntax for slice headers, and certain syntax elements present in slice headers according to the current version are moved to a lower hierarchical syntax element called slice_header_data.</figref><figref num="23a">Figure 23a shows a table in which all syntax elements apart from the slice header are signaled by the syntax element slice header data.</figref><figref num="23b">Figure 23b shows a table in which all syntax elements apart from the slice header are signaled by the syntax element slice header data.</figref><figref num="23c">Figure 23c shows a table in which all syntax elements apart from the slice header are signaled by the syntax element slice header data.</figref><figref num="24">Figure 24 shows the supplementary enhanced information message syntax.</figref><figref num="25a">Figure 25a shows the SEI payload syntax configured to introduce a new slice or subpicture SEI message type.</figref><figref num="25b">Figure 25b shows the SEI payload syntax configured to introduce a new slice or subpicture SEI message type.</figref><figref num="26">FIG. 26 shows an example for a subpicture buffer SEI message.</figref><figref num="27">FIG. 27 shows an example for a subpicture timing SEI message.</figref><figref num="28">FIG. 28 shows what the subpicture slice information SEI looks like.</figref><figref num="29">FIG. 29 shows an example for a subpicture tile information SEI message.</figref><figref num="30">Figure 30 shows an example syntax for a subpicture tile range information SEI message.</figref><figref num="31">Figure 31 shows the first variant of a syntactic example for a region of interest SEI message where each ROI is signaled in an individual SEI message.</figref><figref num="32">Figure 32 shows a second variant of the syntax example for region of interest SEI messages where all ROIs are signaled in a single SEI message.</figref><figref num="33">FIG. 33 shows a possible syntax for a timing control packet according to a further embodiment.</figref><figref num="34">FIG. 34 shows a possible syntax for a tile identification packet according to an embodiment.</figref><figref num="35">FIG. 35 shows possible subdivisions of an image according to different subdivision settings according to the embodiment.</figref><figref num="36">FIG. 36 shows possible subdivisions of an image according to different subdivision settings according to the embodiment.</figref><figref num="37">FIG. 37 shows possible subdivision of the image according to the different subdivision settings according to the embodiment.</figref><figref num="38">FIG. 38 shows a possible subdivision of an image according to a different subdivision setting according to an embodiment.</figref><figref num="39">FIG. 39 shows an example of a portion from a video data stream according to an embodiment using timing control packets distributed between access unit payload packets.</figref>
0136With respect to FIG. 12, an encoder 10 according to an embodiment of the present application and its operating mode is described. The encoder 10 is configured to encode the video content 16 into the video data stream 22. The encoder is configured to do this in units of frame / image 18 subpotions of video content 16, where the subpotions are, for example, slices 24 into which image 18 is divided, or eg tiles 26 or WPP substreams 28. It may be some other spatial part, such as, all of which require the encoder 10 to be able to support, for example, tiles or WPP parallel processing, or the subpotion is a slice. Is illustrated in FIG. 12 solely for the purposes illustrated, rather than suggesting the need for.
0137In encoding the video content 16 in units of subportion 24, the encoder 10 distributes the image 18 according to the raster scan order with subpotions indicating the continuous operation of such blocks along the decoding order. For example, the subpotion 24 that does not necessarily match the reproduction sequence 20 defined between the images 18 and interferes within each image 18 block, eg, interferes with the image 18 of the video 16 according to the image decoding order. The decoding order-or the coding order-defined within can be followed. In particular, the encoder 10 is assigned to the current part to be encoded to use, for example, the properties that indicate such adjacencies in predictive coding and / or entropy coding, such as to determine the predictive and / or entropy context. It is configured to follow this decoding order, which determines the usefulness of spatially and / or temporally adjacent parts: parts that were merely coming before the video (encoding / decoding) are available. .. Otherwise, the characteristics just mentioned will be set to default values, or some other alternative will be taken.
0138On the one hand, the encoder 10 does not need to continuously encode the subportions 24 in the decoding order. Rather, the encoder 10 is capable of performing parallel processing and performing more complex coding in real time to speed up the coding process. Similarly, the encoder 10 may or may not be configured to convey or send data encoding subportions according to the decoding order. For example, the encoder 10 encodes in some other order such that the coding of the subportions follows an order that is terminated by the encoder 10 which can deviate from the decoding order just mentioned, for example, by parallel processing. The converted data can be output / transmitted.
0139To make the encoded version of subportion 24 suitable for transmission over the network, encoder 10 encodes each subportion 24 into one or more payload packets of a sequence of packets in video stream 22. .. If the subportion 24 is a slice, the encoder 10 is configured to replace, for example, each slice data, eg, each coded data, with one payload packet, such as a NAL unit. This packetization can help make the video data stream 22 suitable for transmission over the network. Thus, the packet represents the smallest unit in which the video data stream 22 occurs, i.e., sent individually by the encoder 10 for transmission over the network to the recipient.
0140In addition, the propagation of payload and timing control packets distributed between them and other packets described below, such as rarely changing syntax elements or EOF (end of file) or AUE (access unit end) packets. There are other types of packets, such as fill data packets, images or sequence parameters.
0141The encoder performs encoding into the payload packet so that the sequence of packets is divided into a sequence of access units 30 and each access unit collects the payload packet 32 for one image 18 of the video content 16. That is, the sequence 34 of packets forming the video data stream 22 is subdivided into non-overlapping parts called access units 30, each associated with each one of image 18. The sequence of the access unit 30 follows the decoding order of the image 18 with which the access unit 30 is associated. In FIG. 12, for example, the access unit 30 centered on the portion of the illustrated data stream 22 contains one payload packet 32 per subportion 24 in which the image 18 is subdivided. That is, each payload packet 32 carries a corresponding subportion 24. The encoder 10 is distributed in sequence 34 of packet timing control packets 36, which causes the timing control packet to subdivide the access unit 30 into decoding units 38, thereby as shown in the center of FIG. At least some access units 30 are subdivided into two or more decoding units 38, and each timing control packet is of a decoding unit 38 whose payload packet 32 follows each timing control packet in the packet sequence 34. Send a signal to the decoder buffer search time for. In other words, the encoder 10 is preceded by each timing control packet 36, and the decoding unit 38, each timing control that sends a signal for each subsequence of the payload packet forming the decoder buffer search time. In packet 36, it precedes a subsequence of the payload packet 32 sequence in one access unit 30. FIG. 12 shows, for example, that every second packet 32 represents the first payload packet of decryption unit 38 of access unit 30. Illustrate the booth. As shown in FIG. 12, the amount of data or the bit rate spent per decoding unit 38 varies, and the decoder buffer search time of decoding unit 38 is the bit rate spent immediately before this decoding unit 38. The decoder buffer search time is within the decoding unit 38 in that it can immediately follow the decoder buffer search time signaled by the timing control packet 36 of the previous decoding unit 38 in addition to the time interval corresponding to. It can correlate with this bit rate variation.
0142That is, the encoder 10 can operate as shown in FIG. In particular, as mentioned above, the encoder 10 can encode the current subpotion 24 of the current image 18 in step 40. As already mentioned, the encoder 10 can be cycled sequentially by the decoding sequence subpotions 24 described above, as indicated by arrow 42, or the encoders 10 may have several "current subpotions" in parallel. Some parallel processing such as WPP and / or tile processing can be used to encode 24 at the same time. Whether or not parallel processing is used, the encoder 10 forms a decoding unit from one or several subportions encoded in step 40 and subsequent step 44, and the encoder 10 has a decoder buffer search time. Is set in this decoding unit and transmitted, and this decoding unit placed in front sends a signal of the decoding buffer search time set for the decoding unit in a time control packet. For example, the encoder 10, if any, subports encoded into all the additional intermediate packets in this decoding unit, the payload packets that form the current decoding unit, including the "prefix packet". Determine the decoder buffer search time in step 44 based on the bit rate spent encoding.
0143Then, in step 46, the encoder 10 can adapt the available bit rate based on the bit rate spent on the decoding device that was just being transmitted in step 44. For example, if the image content in the decoding unit just transmitted in step 44 is very complex in terms of compression ratio, the encoder 10 is faced, for example, in connection with the network transmitting the video data stream 22. Based on the current bandwidth situation, the bit rate available for the next decryption unit can be reduced to follow some externally set target bit rates that have been determined. Steps 40-46 are then repeated. By this measurement, the image 18 is encoded and transmitted, i.e., transmitted in units of decoding units preceded by the corresponding timing control packets.
0144In other words, the encoder 10 encodes the current subportion 24 of the current image 18 into the current payload packet 32 of the current decoding unit 28 while encoding the current image 18 of the video content 16.40 Then, at the first moment, in the data stream, the data stream the current decoding unit 38, which is preceded by the current timing control packet 36 at the decoder buffer search time signaled by the current timing control packet (36). By transmitting within 44 and returning to steps 46-40, the first temporal moment-the second temporal moment later than the first time visit step 44-the current image in the second time visit step 18 Encode 44 the further subportion 24 of.
0145Since the encoder can transmit the decoding unit prior to encoding the rest of the current image to which the decoding unit belongs, the encoder 10 can reduce the inter-terminal delay. On the one hand, the encoder 10 does not have to waste the available bit rate because it is possible for the encoder 10 to react to certain properties of its spatial distribution of current image content and complexity.
0146On the one hand, an intermediate network entity that is also important for the transmission of the video data stream from the encoder to the decoder receives the video data stream 22 to benefit from the decoding unit-like coding and transmission by the encoder 10. Timing control packets 36 can be used to ensure that some decoders receive the decryption unit in time. See, for example, FIG. 14, which shows a decoder for decoding the video data stream 22. The decoder 12 receives the video data stream 22 in the encoded image buffer CPB48 via the network in which the encoder 10 has transmitted the video data stream 22 to the decoder 12. In particular, since network 14 is considered capable of supporting low latency applications, network 10 sends the decoder buffer search time to send the sequence 34 of packets of the video data stream 22 to the encoded image buffer 48 of the decoder 12. The decoding unit is thereby present in the encoded image buffer 48 prior to the decoder buffer search time signaled by the timing control packet placed in front of each decoding unit. With this measurement, the decoder does not get stuck, i.e., without running out of payload packets available in the encoded image buffer 48, the decoder's encoded picture buffer in units of decoding units rather than full access units. Use the decoder buffer search time of the timing control packet to empty 48. FIG. 14 illustrates, for example, a processing unit 50 that is connected to the output of the coded image buffer 48 and whose input receives the video data stream 22. Like the encoder 10, the decoder 12 can perform parallel processing using, for example, tile parallel processing / decoding and / or WPP parallel processing / decoding.
0147As described in detail below, the decoder buffer search time is not necessarily related to the search time for the coded picture buffer 48 of the decoder 12. Rather, the timing control packet can additionally, or proceed with the search for already decoded image data in the corresponding decoded picture buffer of the decoder 12. FIG. 14 buffers, for example, the decoded version of the video content obtained by decoding the video data stream 22 by the processing unit 50, and is therefore stored in units of the decoded version of the decoding unit. , Output decoder A decoder 12 including a picture buffer is shown. The decoder's decoding picture buffer 22 is thus connected between the output of the decoder 12 and the output of the processing unit 50. By having the ability to set a search time to output a decoded version that decodes the unit from the decoded picture buffer 52, the encoder 10 is flexible, i.e., while encoding the current image. You are given the opportunity to control the reproduction, or between terminals, to get a delay in the reproduction of the video content on the decoding side, even with a particle size smaller than the image rate or frame rate. Obviously, over-dividing each image 18 into a large number of subportions 24 on the encoding side has a negative effect on the bit rate for transmitting the video data stream 22, but on the other hand, such Delays between terminals can be minimized by minimizing the time required to encode, transmit, decode, and output the decoding unit. On the one hand, increasing the size of the subpotion 24 increases the delay between terminals. Therefore, a compromise must be found. Using the decoder buffer search time mentioned to guide the output timing of the decrypted version of subportion 24 in units of decryption units is currently available to encoder 10 or some other unit on the coding side. Allows this compromise to be applied spatially beyond the content of the image To do. By this means, it is possible to control the inter-terminal delay in such a way that it spatially varies throughout the current image content.
0148When executing the above-described embodiment, it is possible to use a packet of a removable packet type as the timing control packet. Detachable packet type packets are not needed on the decryption side to recover the video content. Hereinafter, this type of packet is referred to as an SEI packet. In addition, there are also removable packet types of packets, i.e., other types of removable packets, such as redundant packets, when transmitted in a stream. Alternatively, the timing control packet may be a packet of a particular removable packet type, but in addition carries a particular SEI packet type field. For example, a timing control packet is an SEI packet with each SEI packet carrying one or several SEI messages, and only those SEI packets containing a particular type of SEI message carry the aforementioned timing control packet. Form.
0149Thus, the examples described so far with respect to FIGS. 12-14 are applied to the HEVC criteria according to further examples, thereby providing more effective HEVC at lower inter-terminal delays. Form a possible concept of. At this time, the above-mentioned packet is formed by the NAL unit, and the above-mentioned payload packet is a VCL NAL unit of the NAL unit stream having slices forming the above-mentioned subportion.
0150Prior to the description of the more detailed examples, it is consistent with the above outlined examples in that distributed packets are used to carry information indicating a video data stream in an efficient manner. A further embodiment is shown, but the classification of the information is different from the above-described embodiment in which the timing control packet is transmitted to the decoder buffer search timing information. Further in the embodiments described below, the type of information transferred over the distributed packets distributed to the payload packets belonging to the access unit relates to region of interest (ROI) information and / or tile identification information. Further, the examples described below can or cannot be combined with the examples described with respect to FIGS. 12-14.
0151FIG. 15 shows an encoder 10 that behaves similarly to that described above with respect to FIG. 12, except for the timing control packets and functional distribution described above with respect to FIG. 13, which are optional for the encoder 10 of FIG. Shown. However, the encoder 10 of FIG. 15 encodes the video content 16 into the video data stream 22 in units of the subportion 24 of the image 18 of the video content 16 as described above in connection with FIG. It is composed of. In encoding the video content 16, the encoder 10, along with the video data stream 22, is interested in transmitting information about the region of interest ROI 60 to the decoding side. The ROI 60 is the spatial subarea of the current image 18 for which the decoder must pay special attention, for example. The spatial position of the ROI 60 can be input to the encoder 10 from the outside as shown by the dotted line 62, such as the user's input, during the coding of the current image 18, or the encoder 10 or on the fly. Can be determined automatically by some other entity. In any case, Encoder 10 faces the following challenges: ROI The display of the 60 position is not a problem for the encoder 10 in principle. Therefore, the encoder 10 can easily indicate the position of the ROI 60 in the data stream 22. However, to show this easily accessible information, the encoder 10 in Figure 15 uses the distribution of ROI packets between the payload packets of the access unit, so that the encoder 10 has subportions 24 and / or subs. The portion 24 is packetized and freely and continuously selects the current image 18 segment for the number of payload packets spatially outer and spatially inside the ROI 60. With distributed ROI packets, any network entity can easily see the payload packets that belong to the ROI. On the one hand, if you use a removable packet type for these ROI packets, it can easily be ignored by any network entity.
0152FIG. 15 shows an example for distributing the ROI packet 64 among the payload packets 32 of the access unit 30. The ROI packet 64 indicates in the video data stream 22 where the encoded data associated with, i.e., to encode, the ROI 60 is contained. ROI packets 64 can indicate the location of ROI 60 in a wide variety of ways. For example, the pure presence / occurrence of ROI packet 64 continues in the sequential order of sequence 34, i.e., within one or more of the following payload packets 32 belonging to the preceding payload packet: The incorporation of coded data for 60 is shown. Alternatively, the syntax element inside the ROI packet 64 indicates whether one or more of the following payload packets 32 are related to the ROI 60, that is, at least partially coded. The variability also stems from the possible variations with respect to the "scope" of each ROI packet 64, that is, the number of payload packets placed before being preceded by one ROI packet 64. For example, ROI in one ROI packet Some indication of inclusion or non-incorporation of encoded data for 60 follows the sequential sequence of sequence 34 until the next ROI packet 64 occurs, or simply the immediate payload packet 32, ie the sequential sequence 34. It pertains to the payload packet 32 following each ROI packet 64 in the correct order. In FIG. 15, graph 66 is a sample of all payload packets 32 that cause the downstream of each ROI packet 64 until the next ROI packet 64 occurs, or whatever occurs early along the packet sequence 34. In connection with the end of the current access unit 30, the ROI packet 64 incorporates any encoded data with respect to the ROI, ie, the ROI 60, or is unrelated to the ROI, ie any encoded data associated with the ROI 60. Illustrate a case showing a deficiency. In particular, FIG. 15 illustrates the case where ROI packet 64 has a syntax element inside, which in any case the payload packet 32 following in the sequential order of packet sequence 34 has the ROI inside. Indicates whether or not it has coded data for 60. Such examples are also described below. However, another possibility is that, as just mentioned, each ROI packet 64 has a payload packet 32 that belongs to the "scope" of each ROI packet 64, which is the ROI 60 associated with the inside of the data, i.e. the data associated with the ROI 60. It is indicated by being present in the packet sequence 34 that it has. According to the embodiments described in more detail below, the ROI packet 64 even indicates the location of the portion of the ROI 60 encoded in the payload packet 32 that belongs to its "scope".
0153For example, any receiving video data stream 22 as understood using ROI packet 64 to treat the ROI-related part of packet sequence 34 with a higher priority than the rest of packet sequence 34. Network entity 68 can also take advantage of ROI-related pointing. Alternatively, network entity 68 can use ROI-related information, for example, to perform other tasks relating to the transmission of the video data stream 22. Network entity 68 is a MANE or decoder for decoding and playing video content 60, such as transmitted via video data streams 22, 28. In other words, network material 68 can use the results of ROI packet identification to determine the transmission work associated with the video data stream. The transmission operation can include a retransmission request for defective packets. Network entity 68 is configured to handle region of interest 70 with increased priority and assign higher priority to ROI packets 72 and their accompanying payload packets, that is, those preceded by them. It can be signaled to cover the area of interest as compared to ROI packets and their accompanying payload packets that are signaled so as not to cover the ROI. Before requesting the forwarding of some of the lower priority payload packets assigned to it, network entity 68 first requests the forwarding of the higher priority payload packets assigned to it. can do.
0154The embodiment of FIG. 15 can be easily combined with the embodiments previously described with respect to FIGS. 12-14. For example, the ROI packet 64 described above may be an SEI packet with a particular type of SEI message contained therein, i.e. an ROI SEI message. That is, the SEI packet may be, for example, a timing control packet and a ROI packet at the same time, that is, when each SEI packet consists of both timings, the information is controlled in the same manner as the ROI display information. Alternatively, the SEI packet may be one of the timing control packet and the ROI packet, rather than the other one, and may not be the ROI packet or the timing control packet.
0155According to the embodiment illustrated in FIG. 16, the distribution of packets between the payload packets of the access unit deals with the video data stream 22, which is the tile of the current image 18 with which the current access unit 30 is associated. Used as shown in a way that network entity 68 is easily accessible, each packet is covered by several subportions encoded in some of the payload packets 32 that serve as prefixes. In FIG. 16, for example, as a sample, it is shown here that the current image 18 is subdivided into four tiles 70, formed by the four quadrants of the current image 18. Also, for example, in a VPS or SPS packet distributed over a sequence of packets 34, the subdivision of the current image 18 to tile 70 is signaled, for example, in the video data stream of the unit consisting of the sequence of images. be able to. As described in more detail below, the current tile subdivision of image 18 may be a regular subdivision of image 18 in the columns and rows of tiles. The number of columns and the number of rows may vary as well as the width of the tile columns and the height of the rows. In particular, the tile columns / row widths and heights may be different for different rows and different columns. FIG. 16 additionally shows an example in which the subpotion 24 is a slice of image 18. Slice 24 subdivides image 18. As outlined in more detail below, 18 subdivisions of an image into slices 24 can include each slice 24 entirely within the range of one tile 70, or completely two or more. It is subject to the constraint that it can cover tile 70 of. FIG. 16 illustrates a case where image 18 is subdivided into five slices 24. These first four slices 24 in the decoding sequence described above cover the first two tiles 70, while the fifth slice completely covers the third and fourth tiles 70. In addition, FIG. 16 shows that each slice 24 has its own payload package. An example is illustrated in which the code is individually encoded in the device 32. Further, FIG. 16 illustrates, as a sample, a case where each payload packet 32 is preceded by a preceding tile identification packet 72. Each tile identification packet 72 then directs for its immediate follow-on payload packet 32 which of tiles 70 the subportion 24 encoded in this payload packet 32 covers. Therefore, the first two tile identification packets 72 within the access unit 30 for the current image 18 indicate the first tile, while the third and fourth tile identification packets 72 are the second tile 70 of image 18. The fifth tile identification packet 72 indicates the third and fourth tiles 70. With respect to the embodiment of FIG. 16, for example, the same variation is possible as described above with respect to FIG. That is, the "scope" of the tile identification packet 72 can only include, for example, the first immediately following payload packet 32 or the immediately following payload packet 32 until the next tile identification packet is generated.
0156For tiles, the encoder 10 can be configured to encode each tile 70 so that spatial prediction or context selection does not occur across the tile boundaries. The encoder 10 can encode tile 70 in parallel, for example. Similarly, any decoder, such as network entity 68, can decrypt tile 70 in parallel.
0157The network entity 68 may be a MANE or decoder or some other device between the encoder 10 and the decoder and is configured to use the information transmitted by the tile identification packet 72 to determine a particular transmission task. be able to. For example, network entity 68 can handle a particular tile of the current image 18 of video 16 with a higher priority, i.e. previously shown as using such tiles or safer FEC protection, etc. Each payload packet can be advanced. In other words, network entity 68 can use the results of the identification to determine the transmission work involved in the video data stream. The transmission operation can include a retransmission request for the received packet in a defective state-ie, surpassing any FEC protection of the video, if any. Network entities can, for example, handle different tiles 70 with different priorities. For this purpose, network entities can assign higher priority to identification packets 72 and their payload packets 72, i.e. prior to tile identification packets 72 and their payload packets associated with lower priority. Placed and thereby related to higher priority. Before requesting any forwarding of the lower priority payload packet assigned to it, network entity 68 first requests, for example, the forwarding of the higher priority payload packet assigned to it. be able to.
0158The examples described so far can be incorporated into the HEVC framework as described in the introductory part of this application, as described below.
0159In particular, SEI messages can be assigned to slices of the decryption unit in the subpicture CPB / HRD case. That is, buffering period and timing SEI messages can be assigned to NAL units that contain slices of decryption units. This can be achieved with a new NAL unit type, which is a non-VCL NAL unit that can precede one or more slice / VCL NAL units in a direct decoder. This new NAL unit can be called the slice prefix NAL unit. Figure 17 illustrates the structure of an access unit that omits any temporary NAL units for sequences and stream ends.
0160According to FIG. 17, the access unit 30 is configured as follows: In the sequential order of the packets in sequence 34 of the packets, the access unit 30 is a special type of packet, i.e. the access unit delimiter 80. You can start from the outbreak. Then, one or more SEI packets 82 of the SEI packet type for all access units can follow within the access unit 30. Both packet types 80 and 82 are optional. That is, this type of packet cannot occur within the access unit 30. Then the sequence of decryption unit 38 continues. Each decoding unit 38 contains, for example, timing control information, or according to the embodiment of FIG. 15 or 16, ROI information or tile information, or more generally, each subpicture SEI message 86. Prefix Arbitrarily starting with NAL unit 84. Then each payload packet or VCL The actual slice data 88 of the NAL unit continues as shown in 88. Thus, each decoding unit 38 contains a sequence of slice prefix NAL units 84 followed by their respective slice data NAL units 88. Bypass arrow 90 in FIG. 17, which avoids the slice prefix NAL unit, indicates that the slice prefix NAL unit 84 is absent in the case of the undecrypted unit subdivision of the current access unit 30.
0161As already mentioned above, all information signaled with the slice prefix and accompanying subpicture SEI messages are prefixed with the access unit or with a second prefix, depending on the flags transmitted in the NAL unit. It may be effective for all VCL NAL units up to the occurrence of the NAL unit, or for the following VCL-NAL units in the decoding order.
0162A slice VCL NAL unit for which information signaled with a slice prefix is valid is referred to below as a prefix slice. The prefix slice associated with the single slice that precedes it does not necessarily have to constitute a complete decryption unit, but it can be part of it. However, a single slice prefix cannot be valid for parallel decoding units (subpictures), and the beginning of the decoding unit is signaled with the slice prefix. If no means for sending a signal is given by the slice prefix syntax (such as the "simple syntax" / version shown below), the occurrence of the slice prefix NAL unit signals the beginning of the decoding unit. Only certain SEI messages (verified via the payload type in the syntax description below) can be sent at the subpicture level within the slice prefix NAL unit, while some SEI messages It can be sent in the slice prefix NAL unit at the subpicture level or as a regular SEI message at the access unit level.
0163As mentioned above with respect to FIG. 16, and / or tile ID SEI messages / tile ID signaling can be understood in high-level syntax. In HEVC's previous design, the slice header / slice data contained an identifier for the tile contained in each slice. For example, slice data semantics show:
0164tile_idx_minus_1 specifies the TileID of the raster scan order. The first tile in the image has a TileID of 0. The value of tile_idx_minus_1 is in the range 0 to (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1) -1.
0165If tiles_or_entropy_coding_sync_idc is equal to 1, this parameter is not found to be useful as this ID is easily derived from the slice address and from the slice dimensions that are signaled in the image parameter set.
0166Different tiles have different priorities for playback, even though the tile ID can be implicitly retrieved during the decryption process (these tiles typically have conversational speakers). Knowledge of this parameter in the application layer is important for different uses, as the video conferencing scenario (which forms the area of interest to include) has a higher priority than another tile. If multiple tiles of transmission network packets are lost, those network packets containing tiles representing the region of interest will be sent to a higher receiver terminal than for retransmission tiles without any priority order. It can be retransmitted with higher priority to maintain the quality of experience. Other use cases can assign tiles to different screens if the dimensions and their location are known, for example in a video conferencing scenario.
0167Because such application layers can handle tiles with a particular priority in transmission scenarios, tile_id can be as a subpicture or slice-specific SEI message, or before one or more NAL units in the tile. It can be provided in the special NAL unit of, or in the special header part of the NAL unit belonging to the tile.
0168As mentioned above with respect to FIG. 15, the area of interest SEI message can be further or optionally provided. Such SEI messages can allow signals in the region of interest (ROI), and in particular can be signaled in the region of interest (ROI), especially in the signaling of the ROI to which a particular tile_id / tile belongs. The message can allow the region of interest to be prioritized in addition to the region of interest ID.
0169FIG. 18 illustrates the use of region of interest signaling tiles.
0170In addition to what was described above, slice header signaling can be performed. The slice prefix NAL unit can also contain the following subordinate slices, ie slice headers for the slices preceded by each slice prefix. If the slice header is only provided in the slice prefix NAL unit, it belongs to the NAL unit type of the NAL unit that contains each dependent slice, or to the slice type in which the next slice data acts as a random access point. The flag in the slice prefix that signals whether or not pulls out the actual slice type.
0171In addition, the slice prefix NAL unit can have a slice or subpicture carry a specific SEI message to convey arbitrary information, such as subpicture timing or tile identifiers. The specific message transmission of any subpicture is not supported in the HEVC specification described in the introductory part of this application, but is important for a particular application.
0172Next, the possible syntax for implementing the above concept of slice prefixing is described. In particular, when using the HEVC situation outlined in the introductory part of the specification of the present application as a basis, it is explained which changes can be sufficient for the slice level.
0173In particular, in the following, two versions of the possible slice prefix syntax are given, one with the function of sending SEI messages only, and one part of the slice header for the next slice. It has an extended function to send a signal. The first simple syntax / version 1 is shown in Figure 19.
0174As a preliminary note, FIG. 19 thus shows a possible practice for performing any of the above embodiments with respect to FIGS. 11-16. The distributed packets shown therein can be interpreted as shown in FIG. 19, which are described in more detail by a particular embodiment.
0175Extended syntax / version 2 with tile_id signaling, decryption unit start identifier, slice prefix ID and slice header data away from the SEI message concept is given in the table in Figure 20.
0176Semantics can be defined as follows: A rap_flag with a value of 1 indicates that the access unit containing the slice prefix is a RAP image. A rap_flag with a value of 0 indicates that the access unit without the slice prefix is not a RAP image.
0177The decoding_unit_start_flag indicates the beginning of a decoding unit within the range of the access unit, so the next slice continues to the end of the access unit or to the beginning of another decoding unit belonging to the same decoding unit.
0178A single_slice_flag with a value of 0 is the information given within the prefix slice NAL unit and accompanying subpicture SEI message to the next access unit, the occurrence of another slice prefix or the beginning of another complete slice header. Indicates that this applies to all of the following VCL-NAL units. A single_slice_flag with the value 1 indicates that all the information shown in the slice prefix NAL unit and the accompanying subpicture SEI message is valid only for the next VCL-NAL unit in decoding the order.
0179tile_idc indicates the amount of tiles present in the following slices. A tile_idc equal to 0 indicates that the tile will not be used in the following slices. A tile_idc equal to 1 indicates that a single tile will be used in the following slices, so its tile identifier will be signaled. A tile_idc with a value of 2 indicates that multiple tiles will be used within the following slices, so the number of tiles and the first tile identifier will be signaled.
0180prefix_slice_header_data_present_flag indicates that the slice header data corresponding to subsequent slices in the decoding order will be signaled with the given slice prefix.
0181slice_header_data () is defined later in the text. It contains relevant slice header information and if dependent_slice_flag is set equal to 1, it is not covered by the slice header.
0182Note that decoupling the slice header and the actual slice data allows for a more flexible transmission of the header and slice data.
0183num_tiles_in_prefixed_slices_minus1 indicates the number of tiles used in the next decoding unit minus 1 without 1.
0184first_tile_id_in_prefixed_slices indicates the tile identifier of the first tile in the next decryption unit.
0185Due to the simple syntax / version 1 of the slice prefix, when not, the following syntax elements can be set to default values as follows: -decoding_unit_start is equal to 1, that is, the slice prefix always indicates the start of the decoding knit. -single_slice_flag is equal to 0, that is, the slice prefix applies to all slices in the decryption unit.
0186The slice prefix NAL unit has 24 NAL unit types, and the NAL unit overviews the table extended according to FIG.
0187That is, to briefly summarize Figures 19-21, the syntactic details found here are due to the confirmed distributed packets above, where the particular packet type is due to the NAL unit type 24 as a sample. Publish. Furthermore, in particular, the syntax example in Figure 20 is the "scope" of distributed packets, a switching mechanism controlled by each syntax element within the scope of these distributed packets themselves, where single_slice_flag controls this scope. That is, the above-mentioned options for switching between different options for defining this scope, respectively, are revealed. In addition, the above-described implementation of FIGS. 1-16, in that each distributed packet contains common slice header data for slice 24 contained in the packet belonging to the "scope" of the distributed packet. It was revealed that the example could be extended. That is, there can be a mechanism controlled by each flag within the range of these distributed packets that indicates whether common slice header data is included within the range of each distributed packet.
0188Of course, as specified in the current version of HEVC, the just-presented concept of which part of the slice header data is transferred to the slice header prefix requires a change in the slice header. The table in Figure 22 shows the possible syntax for such slice headers, shifting the definitive syntactic elements present in the slice headers to a lower hierarchy called slice_header_data () according to the current version. This syntax for slice headers and slice header data only applies to options that use the extended slice header prefix NAL unit concept.
0189In Figure 22, slice_header_data_present_flag is predicted from the value at which slice header data for the current slice is signaled in the access unit, that is, the last slice prefix NAL unit of the most recently occurring slice prefix NAL unit. Indicates that.
0190All syntax elements away from the slice header are signaled by the syntax element slice header data, as given in the table in Figure 23.
0191That is, the concepts of FIGS. 22 and 23 are transferred to the examples of FIGS. 12 to 16, and the distributed packets described therein are coded in the payload packets, which are the concepts to be incorporated into these distributed packets. Part of the slice header syntax (subportion) of the slice to be converted 24, that is, VCL Can be extended by NAL unit. The transfer may be optional. That is, each syntax element of a distributed packet can indicate whether such slice header syntax is included in each distributed packet. When included, each slice header data included in each distributed packet can be applied to all slices contained in a packet that belongs to the "scope" of each distributed packet. Whether the slice header data contained in the distributed packet is adopted by the slice encoded in any of the payload packets belonging to the scope of this distributed packet is signaled, for example, by slice_header_data_present_flag in Figure 22. Can be This measurement miniaturizes the slice headers of the slices encoded in the packets that belong to the "scope" of each distributed packet, and therefore, using the flags just mentioned in the slice, the figure above. Any decoder that receives slice headers and video data streams, such as the network entities found in 12-16, responds to the flag just mentioned in the slice header and is embedded in the distributed packet. The slice header data belongs to the scope of this distributed packet for each flag within the range of the slice that signals the movement of the slice header data to the slice prefix, that is, for each distributed packet. I try to copy it to the slice header of the slice that is encoded in the payload packet.
0192In addition, the SEI message syntax can be made as shown in FIG. 24, continuing with the syntax examples for implementing the examples of FIGS. 12-16. To introduce the slice or subpicture SEI message type, the SEI payload syntax can be configured as shown in the table in Figure 25. Only SEI messages with a payload Type in the range 180-184 can be sent exclusively at the subpicture level within the range of the slice prefix NAL unit. In addition, region of interest SEI messages with a payloadType equal to 140 can be sent either as a slice prefix NAL unit on the subpicture level or as a periodic SEI message on the access unit level.
0193That is, while moving the details shown in FIGS. 25 and 24 onto the embodiments described above with respect to FIGS. 12-16, the distributed packets shown in these examples of FIGS. 12-16 are slice prefix NAL units. Within the range of, for example, it can be achieved using a slice prefix NAL unit with a particular NAL unit type, eg 24, that includes a particular type of SEI message signaled by the payloadType at the beginning of each SEI message. it can. In the particular syntax examples currently described, payloadType = 180 and payloadType = 181 result in timing control packets according to the examples in Figures 11-14, while payloadType = 140 follows the ROI in Figure 15. The result is a packet, and payloadType = 182 results in a tile identification packet according to the embodiment of FIG. The particular syntax examples described herein can include a subset of one or just mentioned payloadType options. Beyond this, FIG. 25 reveals that any of the above-described embodiments of FIGS. 11-16 can be combined with each. In addition, FIG. 25 reveals that the examples of FIGS. 12-16 or combinations thereof can also be extended by distributed packets, which is subsequently explained by payloadType = 184. As already mentioned above, the extensions described below with respect to payloadType = 183 are common slice header data for slice headers of slices that are encoded by any distributed packet to any payload packet that belongs to its scope. Ends with the possibility that it could be incorporated.
0194The table below defines the SEI messages that can be used at the slice or subpicture level. Areas of SEI messages of interest that can be used at the subpicture and access unit level are also shown.
0195Whenever a slice prefix NAL unit of NAL unit type 24 has the SEI message type 180 contained therein, FIG. 26 shows, for example, an example for an occurring subpicture buffer SEI message. As such, the timing control packet is formed.
0196Semantics can be defined as follows: seq_parameter_set_id identifies the sequence parameter set that contains the sequence HRD attribute. The value of seq_parameter_set_id is equal to the value of seq_parameter_set_id of the image parameter referenced and set by the primary encoded image associated with the buffering period SEI message. The value of seq_parameter_set_id ranges from 0 to 31 and includes:
0197initial_cpb_removal_delay [SchedSelIdx] The initial_alt_cpb_removal_delay [SchedSelIdx] then identifies the first CPB removal delay for the Decoding unit (subpicture) ScheduleSelIdxth CPB. The syntax element has the bit length given by initial_cpb_removal_delay_length_minus1 + 1 and is in units of 90kHz clock. The value of the syntax element is not 0 and does not exceed the time equivalent of 90000 * (CpbSize [SchedSelIdx] ÷ BitRate [SchedSelIdx]), CPB size of the 90kHz clock unit.
0198On all encoded video sequences, the sum of initial_cpb_removal_delay [SchedSelIdx] and the initial_cpb_removal_delay_offset [SchedSelIdx] for each decoding unit (sub-picture) are constant for each value of ScheduleSelIdx, and initial_alt_cpb_removal_delay [SchedSelIdx] ] Is constant for each value of ScheduleSelIdx.
0199Figure 27 also shows an example for a subpicture timing SEI message, and the semantics can be described as follows:
0200Decoding units associated with the recent subpicture buffering period SEI message of the preceding access unit in the same decoding unit (subpicture), if any, or with the subpicture timing SEI message. How many clocks are ticked to wait after removing the decoding unit (subpicture) from the CPB associated with the recent buffering period SEI message in the preceding access unit before removing the data (subpicture) from the buffer. Identify the data. This value is also used to calculate the earliest possible time for decoding unit (subpicture) data to arrive at the CPB for HSS (hypothetical Stream Scheduler [2] 0). The syntax element is a fixed length code whose bit length is given by cpb_removal_delay_length_minus1 + 1. On the contrary, cpb_removal_delay is the remainder of the modulo2 (cpb_removal_delay_length_minus1 + 1) counter.
0201du_dpb_output_delay is used to calculate the DPB output time of the decoding unit (subpicture). It specifies how many clocks tick to wait after the image decoding unit (sub-picture) is removed from the CPB before it is output from the DPB.
0202Note that this allows sub-picture updates. In such a scenario, the unupdated decryption units remain immutable in the final decrypted image and they remain visible.
0203Summarizing FIGS. 26 and 27 and transferring the specific details contained therein to the examples of FIGS. 12-14, the decoder buffer search time for the decoding unit, ie, the additional decoding buffer search time. In connection with, it can be said that the signal can be discriminatively sent in the timing control packet associated with the coding method. That is, in order to obtain the decoder buffer search time for a particular decoding unit, the decoder receiving the video data stream determines the decoder search time obtained from the timing control packet placed in front of the particular decoding unit. Immediately add to the decoder search time of the preceding decoding unit, i.e., the procedure of this method such that the particular decoding unit precedes and the decoding unit follows. At the beginning of each coded video sequence of some image or its part, the timing control packet is further or instead fully coded than compared to the decoder buffer search time of the preceding decoding unit. Includes the coded decoder buffer search time.
0204FIG. 28 shows what the subpicture slice information SEI message looks like. Semantics can be defined as follows:
0205A slice_header_data_flag with a value of 1 indicates that slice header data is present in the SEI message. The slice header data given in the SEI is valid for the access device, the generation of slice data for other SEI messages, and all slices that inherit the decryption order to the end of the slice NAL device or slice prefix NAL device. ..
0206Figure 29 shows an example for a sub-picture tile information SEI message, the semantics can be defined as follows:
0207tile_priority indicates the priority of all tiles in the slice placed before taking over the decryption order. All tile_priority values are in the range 0-7, with 7 indicating the highest priority.
0208Multiple_tiles_in_prefixed_slices_flag with a value of 1 indicates that it has one or more tiles in the prefix slice to inherit the decryption order. A multiple_tiles_in_prefixed_slices_flag with a value of 0 indicates that the next prefix slice contains only one tile.
0209num_tiles_in_prefixed_slices_minus1 indicates the number of tiles in the prefix slice that inherits the decryption order.
0210first_tile_id_in_prefixed_slices indicates the tile_id of the first tile of the prefix slice that inherits the decryption order.
0211That is, the embodiment of FIG. 16 can be implemented using the syntax of FIG. 29 for implementing the tile identification packet referred to in FIG. As shown there, a particular flag, here multiple_tiles_in_prefixed_slices_flag, is encoded in either just one tile or a payload packet in which multiple tiles belong to the range of their distributed tile identification packets. It may be used to indicate within the distributed tile identification packet whether it is covered by the subportion of the current image 18 to be. If the flag signals to cover multiple tiles, then an additional syntax element is any sub of any payload packet that belongs to the range of each distributed packet, here as a sample, each distributed tile identification packet. Included in num_tiles_in_prefixed_slices_minus1 which indicates the number of tiles ending with a potion. Finally, a further syntax element, here as a sample, first_tile_id_in_prefixed_slices indicates the ID of the tile in the number of tiles indicated by the current distributed tile identification packet, which is the first according to the decryption order. .. Transferring the syntax of FIG. 29 to the embodiment of FIG. 16, the tile identification packet 72 preceded by the fifth payload packet 32 is set to, for example, multiple_tiles_in_prefixed_slices_flag, 1. It has all three just stated syntax elements with num_tiles_in_prefixed_slices_minus1, thereby indicating that the two tiles belong to the two tiles that belong to the current scope, and the first_tile_id_in_prefixed_slices that are set to 3, and the current tile. The progress of the decryption order tile belonging to the scope of identification packet 72 is the third tile (tile)
0212Figure 29 also reveals that tile identification packet 72 can probably also indicate tile_priority, the priority of tiles that belong to that scope. Similar to the ROI aspect, network entity 68 can use such priority information to control transmission operations, such as requesting the transfer of a particular payload packet.
0213Figure 30 shows a syntactic example for a subpicture tile range information SEI message, and the semantics can be defined as:
0214Multiple_tiles_in_prefixed_slices_flag with a value of 1 indicates that the prefix slice has one or more tiles to inherit the decryption order. Multiple_tiles_in_prefixed_slices_flag with a value of 0 indicates that the following prefix slices contain only one tile.
0215num_tiles_in_prefixed_slices_minus1 indicates the number of tiles in the prefix slice to inherit the decryption order.
0216tile_horz_start [i] indicates the horizontal start of the i-th tile of the pixel within the range of the image.
0217tile_width [i] indicates the width of the i-th tile of the pixel within the range of the image.
0218tile_vert_start [i] indicates the horizontal start of the i-th tile of the pixel within the range of the image.
0219tile_height [i] indicates the height of the i-th tile of the pixel within the range of the image.
0220Note that the tile range SEI message is used for display operations, such as assigning tiles to screens in multiple screen display scenarios.
0221FIG. 30 thus shows the tile identification of FIG. 16 in that the tiles belonging to the scope of each tile identification packet are indicated by their location within the range of the current image 18 rather than their tile ID. It is clarified that the implementation syntax example of FIG. 29 for packets may be varied. That is, it sends a signal for the tile ID of the first tile in the decoding sequence covered by each subportion encoded in one of the payload packets belonging to the scope of each distributed tile identification packet. Rather, for each tile that belongs to the current tile identification packet, its position is, for example, the top left corner position of each tile i, here as a sample tile_horz_start and tile_vert_start, and the width and height of tile i, here. It can be signaled by tile_width and tile_height as a sample.
0222A syntax example for the region of interest SEI message is shown in Figure 31. For more accuracy, FIG. 32 shows the first variant. In particular, region of interest SEI messages can be used, for example, at the access unit level or subpicture level to signal one or more regions of interest. According to the first variant of Figure 32, if multiple ROIs are in the current scope, each ROI is a signal for all ROIs in the scope of each ROI packet within the scope of one ROI SEI message. Rather than sending a signal, it is signaled once for each ROI SEI message.
0223According to FIG. 31, the region of interest SEI message signals each ROI individually. Semantics can be defined as follows:
0224roi_id indicates the identifier of the region of interest.
0225roi_priority indicates the priority of all tiles belonging to the region of interest or all slices in the prefix slice, depending on whether the SEI message is sent at the subpicture level or the access unit level to inherit the decryption order. .. The value of roi_priority is comprehensively in the range of 0 to 7, where 7 indicates the highest priority. In both cases, given the roi_priority of the roi info SEI message and the tile_priority of the subpicture tile info SEI message, the highest values of both apply to the priority of the individual tiles.
0226num_tiles_in_roi_minus1 indicates the number of tiles in the prefix slice that inherits the decoding order belonging to the region of interest.
0227roi_tile_id [i] indicates the tile_id of the i-th tile belonging to the region of interest for the prefix slice that inherits the decryption order.
0228That is, FIG. 31 shows that an ROI packet as shown in FIG. 15 can signal the ID of the region of interest referenced by each ROI packet and payload packet belonging to the scope. Optionally, the ROI priority index can be signaled with the ROI ID. However, both syntax elements are optional. Then the syntax element num_tiles_in_roi_minus1 can indicate the number of tiles in the scope of each ROI packet belonging to each ROI 60. Then roi_tile_id indicates the tile-ID of the i-th tile that belongs to ROI 60. For example, imagine that image 18 is subdivided into tiles 70 in the manner shown in FIG. 16, the ROI in FIG. It is formed by the first and third tiles in the decoding order that 60 corresponds to the image corresponding to the left half of image 18. The ROI packet can then be placed before the first payload packet 32 of access unit 30 in FIG. 16 and between the fourth and fifth payload packets 32 of this access unit 30. Further ROI packets follow. Then the first ROI packet has num_tile_in_roi_minus1 set to 0 and roi_tile_id [0] set to 0 (which pays attention to the first tile in the decryption order) and the fifth payload. The second ROI packet before packet 32 has roi_tile_id [0] set to 2 and num_tiles_in_roi_minus1 set to 0 (this causes the decryption order in the bottom quarter on the left side of image 18). Means the third tile of).
0229According to the second variant, the syntax of the area of interest SEI message can be as shown in Figure 32. Here, all ROIs in a single SEI message are signaled. In particular, the same syntax as for Figure 31 is used as described above, but with the syntax element, here the number signaled by num_rois_minus1 as a sample, for each of the many ROIs referenced by each ROI SEI message or ROI packet. Multiply the syntax elements for the ROI of. Optionally, a further syntax element, here as a sample, roi_presentation_on_separate_screen, can indicate for each ROI whether each ROI is suitable for being shown on a separate screen.
0230The semantics can be:
0231num_rois_minus1 indicates the number of regular slices that inherit the ROI or decryption order of the prefix slices.
0232roi_id [i] indicates the identifier of the i-th region of interest.
0233roi_priority [i] is all that belong to the i-th region of interest or all slices in the prefix slice, depending on whether the SEI message is sent at the subpicture level or the access unit level to inherit the decryption order. Indicates the tile priority. The value of roi_priority is comprehensively in the range of 0 to 7, where 7 indicates the highest priority. In both cases, given the roi_priority of the roi info SEI message and the tile_priority of the subpicture tile info SEI message, the highest values of both apply to the priority of the individual tiles.
0234num_tiles_in_roi_minus1 [i] indicates the number of tiles in the prefix slice that inherits the decoding order belonging to the i-th region of interest.
0235roi_tile_id [i] [n] indicates the tile_id of the nth tile belonging to the i-th region of interest in the prefix slice that inherits the decryption order.
0236roi_presentation_on_seperate_screen [i] indicates that the region of interest associated with the i-th roi_id is suitable for presentations on separate screens.
0237Thus, temporarily summarizing the various embodiments described so far, SEI messages can be applied as well as higher level syntax items beyond those contained in the NAL unit header per slice level. A higher level syntactic signaling strategy was presented. Therefore, we described the slice prefix NAL unit. The syntax and semantics of slice prefixes and slice_level / subpicture SEI messages have been described with use cases for low latency / subpicture CPB behavior, tile signaling and ROI signaling. The extended syntax was additionally presented in the signal portion of the slice header of the slice below the slice prefix.
0238For integrity, FIG. 33 shows a further embodiment for the syntax that can be used for timing control packets according to the examples in FIGS. 12-14. The semantics can be:
0239Du_spt_cpb_removal_delay_increment, in clock subtick units, between the last decryption unit and decryption unit information SEI message in the decryption order of the current access unit and the nominal CPB time of the associated decryption unit. To identify. As identified in Annex C, this value is also used to calculate the earliest possible time of arrival of decryption unit data in the CPB for HSS. The syntax element is represented by a fixed length code whose bit length is given by du_cpb_removal_delay_increment_length_minus1 + 1. Decryption Unit Information The value of du_spt_cpb_removal_delay_increment is equal to 0 when the decryption unit associated with the SEI message is the last decryption unit of the current access unit.
0240A dpb_output_du_delay_present_flag equal to 1 identifies the presence of the pic_spt_dpb_output_du_delay syntax element in the decryption unit information SEI message. A dpb_output_du_delay_present_flag equal to 0 identifies the lack of a pic_spt_dpb_output_du_delay syntax element in the decryption unit information SEI message.
0241When SubPicHrdFlag is equal to 1, pic_spt_dpb_output_du_delay is used to calculate the DPB output time of the image. It specifies how many subclock ticks to wait before the decoded image is output from the DPB after the removal of the last decoding unit of the access unit from the CPB. If not, the value of pic_spt_dpb_output_du_delay is presumed to be equal to pic_dpb_output_du_delay. The length of the syntax element pic_spt_dpb_output_du_delay is given in bits by dpb_output_delay_du_length_minus1 + 1.
0242It is a requirement for bitstream matching that all decryption unit information SEI messages associated with the same access unit apply to the same operating point and that dpb_output_du_delay_present_flag equal to 1 have the same value as pic_spt_dpb_output_du_delay. The output time derived from the pic_spt_dpb_output_du_delay of any image output from the output timing tuning decoder precedes the output time derived from the pic_spt_dpb_output_du_delay of all images of any next CVS in the decoding order.
0243The image output order determined by the value of this syntax element is the same order established by the value of PicOrderCntVal.
0244The output time from pic_spt_dpb_output_du_delay is all images in the same CVS for images that are not output by the "bumping" process because they precede IRAP images where no_output_of_prior_pics_flag is recognized to be equal to or equal to 1 in the decoding order On the other hand, it increases with the increase of the value of PicOrderCntVal. For any two images in CVS, the difference in output time between the two images when SubPicHrdFlag is equal to 1 is the same as the difference when SubPicHrdFlag is equal to 0.
0245In addition, FIG. 34 shows a further embodiment for sending a signal in the ROI region using ROI packets. According to Figure 34, the ROI packet syntax contains only one flag that indicates whether all subportions of the image encoded in some payload packets 32 belonging to that scope belong to the ROI. The "scope" extends to the occurrence of ROI packets or region_refresh_info SEI messages. If the flag is 1, the region is encoded in each next payload packet, if 0, the opposite applies, that is, each subpotion of image 18 does not belong to ROI 60.
0246Before re-discussing some of the above examples, in other words, before explaining some terms used above such as tiles, slices, WPP, substreams, subdivisioning, etc. It should be noted that High Level signaling can also be defined in transport specifications such as [3-7]. In other words, the packet and formation sequence 34 described above are forwarded packets, some of which have application layer subportions such as slices, and are included as whole or fragmented and packetized therein. Some is dispersed in the method between the latter and the objectives described above. In other words, the decentralized packet described above is defined as another type of SEI message that can be defined in the application layer video codec, but instead can be a special forwarding packet defined in the forwarding protocol. Not restricted to be defined.
0247In other words, according to one aspect of the present specification, the embodiment has video content encoded in a sub-packet (see encoded tree block or slice) of an image of video content. A video data stream is seen, and each subportion is encoded into one or more payload packets (see VCL NAL unit) of the packet sequence of the video data stream (NAL unit), respectively, and the packet sequence. Is divided into a sequence of access units, so that each access unit collects payload packets for each image of the video content, and the sequence of packets distributes timing control packets (slice prefixes) within it. The timing control packet subdivides the access unit into decoding units, thereby subdividing at least some access units into two or more decoding units, and each timing control packet is a decoder buffer for the decoding unit. Signals a payload packet followed by each timing control packet for search time and packet sequence.
0248As mentioned above, the region where the video content is encoded in the data stream in units of image subportions is a precursory encoding, such as an encoding mode (eg, intramode, intermode, subdivision information, etc.). , Predictive parameters (eg, motion vector, extrapolation direction, etc.) and / or residual data (eg, transformation coefficients, etc.) can be covered, and each of these syntax elements is, for example, a coded tree. Related to local parts of the image, such as blocks, predictive blocks and residual (eg, transform) blocks.
0249As mentioned above, each payload packet can contain one or more slices (completely, each). The slices may each be decodable or may exhibit interrelationships that prevent their independent decoding. For example, each entropy slice can be entropy-decoded, but predictions across slice boundaries can be prohibited. Dependent slices use WPP processing, that is, entropy and predictive coding across slice boundaries with the ability to code / decode dependent slices in parallel, depending on the individual dependent slices and dependent slices. It is possible to enable coding / decoding of alternating initiations of coding / decoding of so-called slices. Instructions over time that the payload packets of the access unit are placed within the range of each access device can be known to the decoder in advance.
0250The order in which the access unit payload packets are placed within the range of each access unit is known to the decoder in advance. For example, the coding / decoding order can be defined in subpotions such as, for example, the scanning order in the coding tree block of the above example.
0251For example, please refer to the figure below. The currently coded / decoded image 100 corresponds to a quarter of the image 110 as a sample in FIGS. 35 and 36 and can be divided into tiles indicated by reference numerals 112a-112d. That is, the entire image 110 can form one tile as in the case of FIG. 37 or can be divided into a plurality of tiles. Tiles can be limited to regular tiles that are placed only in columns and rows. Different examples are shown below.
0252As can be seen, the image 110 is further subdivided into a coding (tree) block (a small box in the figure called CTB above) 114, in which the coding sequence 116 is defined (in which the coding sequence 116 is defined. Here it is a raster scan order, but it can be different). Subdivision of the image of block 114 into tiles 112a-d can be restricted so that the tiles are relatively prime to block 114. In addition, blocks 114 and tiles 112a-d can be restricted to a regular array of columns and rows.
0253If there are tiles (ie, one or more), the raster in coding (decoding) order 116 scans the complete tile first, then-in raster scan tile order-in tile order to the next tile.
0254Since tiles are encoding / decoding that are independent of each other due to non-intersection of tile boundaries due to spatial prediction and context selection inferred from spatial adjacency, encoder 10 and decoder 12 have, for example, tile boundary intersections. With the exception of allowed in-loop or post-filtering, images subdivided into tile 112 (originally indicated by 70) can be encoded / decoded in parallel independently of each other.
0255Image 110 can be further subdivided into slices 118a-d, 180-originally indicated by reference numeral 24. A slice can completely contain only one part of a tile, one complete tile or multiple tiles. In this way, the division into slices can also subdivide the tiles as in the case of FIG. Each slice contains entirely one coding block 114 and consists of consecutive coding blocks 114 in coding sequence 116, so that the sequence is defined within the indexed slices 118a-d of the figure. Be done. The slice divisions shown in FIGS. 35 to 37 are selected for convenience of explanation. Tile boundaries signaled in the data stream. Image 110 can form a single tile, as illustrated in FIG.
0256The encoder 10 and decoder 12 can be configured to follow tile boundaries in that spatial predictions are not applied across tile boundaries. Context adaptation, or probabilistic adaptation of various entropy (arithmetic) contexts, can be continued throughout all slices. However, whenever the slices-along the coding sequence 116-cross the tile border (if inside the slice), as shown in FIG. 36 with respect to slices 118a, b, the slices then then each sub. It is subdivided into subsections (substreams or tiles) that have slices that contain a pointer (cpentry_point_offset) that marks the beginning of the section. In the decoder loop, the filter can cross tile boundaries. Such a filter can include one or more of a deblocking filter, a Sample Adaptive Offset (SAO) filter and an Adaptive loop filter (ALF). When activated, the latter can be applied through tile / slice boundaries.
0257Each arbitrary second and following subsection is byte-aligned and placed within a slice that has a pointer pointing to an offset from the beginning of one subsection at the beginning of the next subsection. Can have. Subsections are arranged within slices of scan sequence 116. FIG. 38 shows an example with slice 180c of FIG. 37 subdivided into subsections 119i as a sample.
0258Note that with respect to the figure, it does not mean that the slices forming the subparts of the tile must be the ends with the rows of tile 112a. See, for example, slice 118a in FIGS. 37 and 38.
0259The figure below shows a typical portion of the data stream for the access unit associated with image 110 in FIG. 38 above. Here, each payload packet 122a ~ d-previously indicated by reference numeral 32-applies to just one slice 118a as a sample. The two timing control packets 124a, b-previously shown with reference symbol 36-are shown as distributed to access unit 120 for convenience of explanation: 124a corresponds to packet sequence 126 (decoding / coding time axis). ) Precedes packet 122a, and 124b precedes packet 122c. Therefore, access unit 120 is divided into two decoding units 128a, b-previously indicated by reference numeral 38-the first of which is packet 122a, b (arbitrary filter data packet (respectively, first). It consists of the first and second packets 122a, b) and any access unit following the SEI packet (preceding the first packet 122a), the second of which is packets 118c, d (any). It consists of filter data packets (following packets 122c and d, respectively).
0260As mentioned above, each packet in a sequence of packets can be assigned to exactly one packet type from multiple packet types (nal_unit_type). Payload packets and timing control packets (and any filter data and SEI packets) are, for example, of different packet types. The instantiation of packets of a particular packet type in a sequence of packets may be subject to certain limits. These limits define the order within the packet types (see Figure 17) that are to be followed by packets within the scope of each access unit so that access device boundaries 130a, b are detectable. It can and remain in the same position in the sequence of packets, even if packets of any removable packet type are removed from the video data stream. For example, a payload packet is of a packet type that cannot be removed. However, timing control packets, filter data packets and SEI packets may be of the removable packet type as described above, i.e. they are non-VCL. It may be a NAL unit.
0261In the above embodiment, the timing control packet is clearly illustrated above by the syntax of slice_prefix_rbsp0.
0262This distribution of timing control packets allows the encoder to adjust the buffer scheduling on the decoder side in the process of encoding the individual images of the video content. For example, encoders are allowed to optimize buffer scheduling to minimize delays between terminals. In this regard, the encoder is allowed to take into account the individual distribution of the coding complex of the entire image area of the video content for the individual images of the video content. In particular, the encoder can continuously output a sequence of packets 122, 122a-d, 122a-d1-3 on a packet-by-packet basis (ie, it is output as soon as the current packet ends). Using the timing control packet, the encoder decodes if some of the subportions of the current image have already been encoded in each payload packet with the remaining subportions, but not yet encoded. It is possible to adjust the buffer scheduling from time to time on the conversion side.
0263Therefore, one of the packet sequences (NAL units) of the video data stream so that the packet sequence is divided into access unit sequences and each access unit collects the payload packets associated with each image of the video content. By encoding each subportion into each of the above payload packets (VCL NAL unit), the video data stream of the unit of the image subportion of the video content (see Encoded Tree Block, Tile or Slice) The encoder for encoding into the video content is configured to distribute the timing control packet (spice prefix) in the sequence of packets so that the timing control packet subdivides the access unit into decryption units. Allows at least some access units to be subdivided into multiple decoding units, each timing control packet signals a decoding buffer search time for the decoding unit, and its payload packet is each in the sequence of packets. Followed by the timing control packet of.
0264Any decoder receiving the video data stream just outlined is free to utilize the schedule information contained in the timing control packet. However, while the information is available to the decoder, the decoder following the codec level must be able to decode the data after the indicated timing. When used, the decoder supplies its decoder buffer and emptys its decoder buffer in units of decoding units. The "decoder buffer" can include a decoded picture buffer and / or a coded picture buffer as described above.
0265Therefore, so that each access unit collects payload packets associated with each image of the video content, it is coded there in units of subportions of the video content image (see Encoded Tree Blocks, Tile or Slice). A video data stream that has video content to be encoded, each subportion being encoded into one or more payload packets (see VCL NAL unit) of a sequence of packets (NAL unit) in the video data stream. The decoder for decoding is configured to look for timing control packets distributed in a sequence of packets in which the access unit is subdivided into decoding units with timing control packets, thereby at least some access units. Divided into multiple decoding units, each timing control packet derives a decoder buffer search time for the decoding unit, whose payload packet follows each timing control packet in the sequence of packets for the decoding unit. The decoding unit is searched from the buffer of the decoder scheduled at the time specified by the decoder buffer search time.
0266Searching for timing control packets can include a decoder inspecting the NAL unit headers and syntax elements it contains, i.e. nal_unit_type. If the value of the later flag is equal to some value, i.e. 124 in the example above, then the packet currently inspected is a timing control packet. That is, the timing control packet contains or conveys the information described above with respect to the pseudocode subpic_buffering as well as subpic_timing. That is, the timing control packet can specify whether to carry the first CPB removal delay for the decoder, or how much clock has elapsed after the removal of each decoder unit from the CPB.
0267To allow repeated transmission of timing control packets without unintentionally further dividing the access unit into decoding units, the flags within the timing control packet range are such that the current timing control packet is coded within the access unit. Clearly signal whether to subdivide into the conversion unit (compare decoding_unit_start_flag = 1 indicating the beginning of the decoding unit and decoding_unit_start_flag = 0 sending the signal in the opposite situation).
0268The mode of using the tile identification information related to the distributed decoding unit differs from the mode using the timing control packet related to the distributed decoding unit in that the tile identification packet is distributed to the data stream. The timing control packets described above can also be distributed across the data stream, or the decoder buffer search time is commonly propagated within the same packet with the tile identification information described below. Therefore, the details brought forward in the above section can be used to clarify the problem in the following description.
0269A further aspect of the specification that can be derived from the above embodiment reveals a video data stream having video content encoded therein and predicts in units of slices in which the image of the video content is spatially subdivided. And using entropy coding, the coding order using the predictive coding and / or coding order between slices that limits entropy coding inside the tile where the image of the video content is spatially subdivided. The sequence of slices in the sequence of slices of the encoded video data stream in the coding order packetizes the packets of the sequence of packets (NAL unit) in the video data stream of the coding order into the payload, and the sequence of packets is accessed. It is divided into a sequence of units, so that each access unit collects a payload packet in which slices for each image of the video content are packetized, and the sequence of packets is 1 after each tile identification packet in the sequence of packets. It has a tile identification packet that identifies a tile (potentially only one) that is immediately covered by a slice (potentially only one) packetized into one or more payload packets and distributes to it.
0270See, for example, that the previous figure shows a data stream. Packets 124a and 124b currently represent tile identification packets. By explicit signaling (compare single_slice_flag = 1), or by convention, the tile identification packet can only see the tiles covered by the packetized slices in the immediately following payload packet 122a. Alternatively, by explicit signaling or by convention, the tile identification packet 124a is one or more payloads after each tile identification packet 124a in the sequence of packets up to the earlier of the end 130b of the current access unit 120. It can be seen that the tiles covered by the slices packetized into the packet and each start the next decryption unit 128b. For example, see Figure 35: If each slice 118a-d1-3 is a separate packetization into its own packet 122a-d1-3, subdivision into the decryption unit will result in the packets being {122a1-3} and {122a1-3} and { Slices {118c1-3, 118d1-3 that are grouped into three decoding units according to 122b1-3} and packetized into packets {122c1-3, 122d1-3} of the third decoding unit. } Covers tiles 112c and 112d, for example, and the corresponding slice prefix indicates c and d, ie these tiles 112c and 112d, when referring to a complete decoding unit, for example.
0271Thus, the network entities described further below use this explicit signaling or convention to associate each tile identification packet with one or more payload packets immediately following the identification packet in the sequence of packets. Can be done. The way the identification can be signaled was illustrated above as a sample via the pseudo code subpic_tile_info. The associated payload packet was mentioned earlier as a "prefix slice". Of course, the embodiments can be modified. For example, the syntax element "tile_priority" can be removed. Moreover, the order within the syntax elements can be switched, and the coding of the descriptors and syntax element principles regarding possible bit lengths is merely illustrated.
0272Prediction and entropy code in units of slices that reveal the network entity that receives the video data stream (ie, the video data stream that has the video content encoded therein, and the image of the video content is spatially subdivided. Using encoding, the image of the video content is spatially subdivided inside the tile, using the coding order between the slices that limit predictive coding and / or entropy coding of the slices in the coding order. The sequence packets the packets of the sequence of packets (NAL units) in the video data stream in the coded order into a payload, and the sequence of packets is divided into a sequence of access units, so that each access unit has its own video content. A slice of the image collects packetized payload packets, and the sequence of packets has tile-identifying packets distributed within it), based on the tile-identifying packets, of each tile-identifying packet in the sequence of packets. It can later be configured to see the tiles covered by slice packetization in one or more payload packets. The network entity can use the identification result to determine the transmission operation. For example, network entities can handle different tiles with different priorities for playback. For example, in the case of packet loss, those payload packets for higher priority tiles are preferably retransmitted over the payload packets for lower priority tiles. That is, the network entity can first request the retransmission of lost payload packets for higher priority tiles. If only sufficient time is left (depending on the transfer rate), the network entity moves on requesting the retransmission of the lost payload packet for the lower priority tile. However, network entities are specific tiles
0273With respect to aspects of using distributed region of interest information, the following ROI packets combine the information content within the common packet as described above with respect to slice prefixes by the timing control packet and / or tile identification packet described above. It should be noted that they can coexist either by doing so or in the form of separate packets.
0274In embodiments using distributed ROI information as described above, in other words, the video data stream which smell has a video content to be encoded using the prediction and entropy coding Te, between slices Using the coding order, the image of the video content is spatially divided, limiting the prediction and / or entropy coding of the predicted coding inside the tile into which the image of the video content is divided, and of the coding order. The sequence of slices is packetized into a payload packet of the sequence of packets (NAL unit) of the video data stream in the coded order, and each access unit packets the slices for each image of the video content into it. -The packet sequence is divided into access unit sequences so as to collect packets, and each packet sequence has ROI packets that check the image tiles belonging to the image ROI and distribute them there.
0275For ROI packets, similar comments are valid as those previously given for tile identification packets: ROI packets are simply one or more for each ROI packet as it was mentioned above for "prefix slices". You can see the tiles of an image that belong to the ROI of the image only within those tiles covered by the slices contained in one or more payload packets associated with it to immediately precede the payload packet of.
0276ROI packets can allow you to see multiple ROIs per previously placed slice regarding checking the associated tiles for each of these ROIs (cpnum_rois_minus1). Then, for each ROI, a priority can be sent that allows the ROI to be ranked in terms of priority (cproi_priority [i]). To allow "tracking" of the ROI over time during the image sequence of the video, the ROIs shown in the ROI packets are related to each other across / across the image boundaries, i.e. over time (cproi_id [i]). As such, each ROI can be indexed by the ROI index.
0277The network entity that receives the video data stream (ie, the video data stream has video content encoded there using predictive and entropy coding, and the video content using the coding order between slices. The image of the video content is spatially divided, limiting the prediction of predictive coding inside the tile where the image of the video content is divided, and the sequence of slices in the coding order is the packet of the video data stream in the coding order. The sequence of packets into the sequence of access units is packetized into a payload packet of the sequence (NAL unit) of, and each access unit collects the packetized payload packets into which slices of each image of the video content. (Split) is configured to see the packet packetizing the slice covering the tile belonging to the ROI of the image, based on the tile identification packet.
0278The network entity can utilize the information transmitted by the ROI packet in a similar manner in the previous section as described above for the tile identification packet. For the current section as well as the previous section, simply signal the slice order of the slices of the image with respect to the position of the tiles in the image, either explicitly signaled in the data stream as described above, or customarily known to the encoder or decoder. By doing so, and by investigating the progress of the portion of the current image covered by these slices, simply by: Any network entity, such as a MANE or decoder, will currently inspect any tile of the payload packet. It should be noted that it is possible to see if it is covered by slices. Alternatively, for each slice (except for the first image in the scan order), the decoder can place each slice (its reconstruction) from this first coded block in the direction of the slice order on the image. As such, it can have a display / index (slice_address measured in units of coded tree blocks) of a first coded block (eg CTB) that references (same code). Therefore, if the index carried by a later tile identification packet differs from the previous one by a difference of one or more, the payload packet between those two tile identification packets should cover the tile with the tile index between them. As soon as the next tile identification packet is encountered according to, it becomes apparent to the network entity, so the tile information packet described above simply contains a slice of one or more payload packets that follow each tile identification packet. It can be sufficient if it only consists of the index of the first tile (first_tile_id_in_prefixed_slices). Slice order along this raster scan order in the coded block
0279The packetized distributed slice header aspect of delivering the signal described above, which is derived from the above embodiment, can be combined with any one of the above-mentioned aspects, which can be derived from the above-described embodiment, or any combination thereof. is there. The previously explicitly stated slice prefix unifies all these aspects, for example by version 2. The effect of the current aspect is to make slice header data more easily available to network entities so that they are transmitted in the outer self-sufficient packets relative to the preceding slice / payload packets. And the iterative transmission of slice header data is possible.
0280Thus, a further aspect of the specification is a packetized and distributed slice header signaling aspect, in other words, subportions of images of video content (see Encoded Tree Blocks or Slices). A video data stream with video content encoded in it in units can be considered to be exposed, with each subportion being one or more payloads of a packet sequence (NAL unit) of the video data stream. Encoded into packets (see VCL NAL unit), the sequence of packets is divided into a sequence of access units so that each access unit collects payload packets for each image of the video content, and the sequence of packets is there. , Distribute slice header packets (slice prefixes) that carry missing slice header data between one or more payload packets following each slice header packet in a sequence of packets. It was.
0281To be considered to publish a video data stream with video content encoded in a video data stream (ie, in units of image subportions of video content (see Encoded Tree Blocks or Slices)). Each subportion can be one or more payload packets (VCL) of a sequence of packets (NAL unit) in the video data stream. Encoded in (see NAL unit), the sequence of packets is divided into a sequence of access units so that each access unit collects payload packets for each image of the video content, and the sequence of packets is sliced headers there. The network entity that receives (distributed the packet) is along the payload data for slicing the packet derived from the slice header data of the slice header packet, and each slice header of the sequence of packets. Slice header that follows a packet but is followed by one or more payload packets Slice headers for one or more payload packets that apply a slice header that is pulled from the packet are skipped, read, and sliced. -It is configured to read the header.
0282Packets Here slice header packets are one or more payloads preceded by any network entity, such as MANE or decoder, at the beginning of a decryption unit or by each packet, as was true with respect to the aspects described above. -It can also have a function of instructing the start of packet execution. Therefore, a network entity that follows the current aspect sees this packet, eg, a payload packet that must be skipped, based on the syntax elements described above for single_slice_flag, which can be combined with decoding_unit_start_flag to read the slice header. In which the later flag allows the retransmission of a copy of a particular slice header packet within the decoding unit, as described above. This is useful, for example, because the slice headers of slices within the scope of one decryption unit can change along the sequence of slices, so the first slice header packet of the decryption unit is (1). Can have a decoding_unit_start_flag that is set (equal to), and slice header packets placed between them can have this flag that is not set, and any network entity can illegally start a new decryption unit. Prevents reading the occurrence of slice header packets.
0283Although some aspects have been described in the context of the device, it is clear that these aspects also represent a description of the corresponding method, and the block or device corresponds to a method step or feature of the method step. To do. Similarly, the aspects described in the context of the method steps represent a description of the corresponding block or member or feature of the corresponding device. The steps of some or all methods can be performed by a hardware device such as a microprocessor, a programmable computer or an electronic circuit (using). In some embodiments, one or more of the most important method steps can be performed by this type of device.
0284The video data stream of the invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
0285Depending on the particular implementation requirements, the embodiments of the present invention can be implemented in hardware or in software. The implementation shall be performed using a digital storage medium having an electronically readable control signal stored on it, such as a flexible disc, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM or flash memory. It can (or can) work with a programmable computer system so that each method can be performed. Therefore, the digital storage medium may be computer readable.
0286Some embodiments according to the invention are electronically readable controls that can work with a programmable computer system such that one of the methods described herein is performed. A data carrier having a signal includes a data carrier having a signal.
0287Usually, the embodiments of the present invention can be implemented as a computer program product having an implementation program code for the program code to execute one of the methods when the computer program product runs on a computer. .. The program code can be stored, for example, in a machine-readable readable carrier.
0288Other examples include computer programs described herein to perform one of the methods stored in a machine-readable readable carrier.
0289In other words, an embodiment of the method of the invention is therefore a computer program having program code for executing one of the methods described herein when the computer program runs on a computer.
0290Further examples of the methods of the invention are therefore recorded on it and consist of a computer program for performing one of the methods described herein (or digital storage). Medium or computer-readable medium). Data carriers, digital storage media or recording media are typically tangible and / or non-transitional.
0291A further embodiment of the method of the invention is therefore a sequence of data streams or signals representing a computer program for performing one of the methods described herein. A data stream or sequence of signals can be configured to be transferred over a data communication connection, such as the Internet.
0292Further embodiments include processing means configured to be applied to perform one of the methods herein, such as a computer or programmable logic device.
0293Further examples include a computer on which a computer program for performing one of the methods described herein is installed.
0294Further embodiments according to the invention are such that the receiver is transmitted (eg, electronically or optically) with a computer program to perform one of the methods described herein. Includes the equipment or system to be configured. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The device or system includes, for example, a file server for transferring a computer program to a recipient.
0295In some embodiments, programmable logic devices (eg, field programmable gate arrays) can be used to perform some or all of the functions of the methods described herein. In some embodiments, the field programmable gate array can work with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device.
0296The above examples are merely illustrated for the purposes of the present invention. It is understood that modifications and changes in arrangement and the details described herein will be apparent to those of ordinary skill in the art. Therefore, it is intended to be limited only by the imminent claims, not only by the specific details provided in the description and description of the examples herein. By the way.
0297reference [1] Thomas Wiegand, Gary J. Sullivan, Gisle Bjontegaard, Ajay Luthra, "Overview of the H.264 / AVC Video Coding Standard", IEEE Trans. Circuits Syst. Video Technol., Vol. 13, N7, July 2003. [2] JCT-VC, "High-Efficiency Video Coding (HEVC) text specification Working Draft 7", JCTVC-I1003, May 2012. [3] ISO / IEC 13818-1: MPEG-2 Systems specification. [4] IETF RFC 3550 --Real-time Transport Protocol. [5] Y.-K. Wang et al. , "RTP Payload Format for H.264 Video", IETF RFC6184, http://tools.ietf.org/html/ [6] S. Wenger et al., "RTP Payload Format for Scalable Video Coding", IETF RFC6190, http://tools.ietf.org/html/rfc6190 [6] T. Schierl et al., "RTP Payload Format for High Efficiency Video Coding", IETF internet draft, http://datatracker.ietf.org/doc/draft-schierl-payload-rtp-h265/
45 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45
Every citation, both ways
| Document | Relation | Office | Category | Cited during |
|---|---|---|---|---|
| US12020345B2 | Cited by | United States of America | – | Applicant |
| KR20200034504A | Cited by | Republic of Korea | – | Search report |
| JP2010232720A | Cites | Japan | A | Search report |
| JP2010232720A | Cites | Japan | A | Search report |
| US2010246662A1 | Cites | United States of America | A | Search report |
| US2010246662A1 | Cites | United States of America | A | Search report |
| JP2010516085A | Cites | Japan | A | Search report |
| JP2010516085A | Cites | Japan | A | Search report |
| JP2013132048A | Cites | Japan | A | Search report |
| JP2013132048A | Cites | Japan | A | Search report |
| WO2013151634A1 | Cites | World Intellectual Property Organization (WIPO) | A | Search report |
| WO2013151634A1 | Cites | World Intellectual Property Organization (WIPO) | A | Search report |
| WO2013151634A1 | Cites | World Intellectual Property Organization (WIPO) | A | Search report |
| 大久保榮監修, 「インプレス標準教科書シリーズ 改訂三版H.264/AVC教科書」, vol. 第1版, JPN6016006149, 1 January 2009 (2009-01-01), JP, pages 99 - 107, ISSN: 0003709151 | Non-patent | – | – | Search report |
| SCHIERL, T., ET.AL.: ""Dependent Slices"", JCTVC-I0229, JPN6016006150, 16 April 2012 (2012-04-16), pages 1 - 7, ISSN: 0003709152 | Non-patent | – | – | Search report |
| BENJAMIN BROSS ET AL.: "High efficiency video coding (HEVC) text specification draft 7", JOINT COLLABORATIVE TEAM ON VIDEO CODING (JCT-VC), vol. JCTVC-I1003_d4, JPN6017049655, 12 June 2012 (2012-06-12), pages 19 - 20, ISSN: 0003976881 | Non-patent | – | – | Search report |
| 山田悦久(外3名): "「高品質映像符号化技術の標準化動向」", 三菱電機技報, vol. 82, no. 12, JPN6016038021, 25 December 2008 (2008-12-25), JP, pages 7 - 10, ISSN: 0003709154 | Non-patent | – | – | Search report |
| KIMIHIKO KAZUI ET AL.: "AHG9: Improvement of HRD for sub-picture based operation", JOINT COLLABORATIVE TEAM ON VIDEO CODING (JCT-VC) OF ITU-T SG 16 WP 3 AND ISO/IEC JTC 1/SC 29/WG 11, vol. JCTVC-J0136, JPN6018026693, July 2012 (2012-07-01), pages 1 - 10, ISSN: 0003976882 | Non-patent | – | – | Search report |
| T. SCHIERL ET AL.: "Slice Prefix for sub-picture and slice level HLS signalling", JOINT COLLABORATIVE TEAM ON VIDEO CODING (JCT-VC) OF ITU-T SG 16 WP 3 AND ISO/IEC JTC 1/SC 29/WG 11, vol. JCTVC-J0255, JPN6017046320, July 2012 (2012-07-01), pages 1 - 12, ISSN: 0003976883 | Non-patent | – | – | Search report |
| R. SKUPIN, V. GEORGE AND T. SCHIERL: "Tile-based region-of-interest signalling with sub-picture SEI messages", JOINT COLLABORATIVE TEAM ON VIDEO CODING (JCT-VC) OF ITU-T SG 16 WP 3 AND ISO/IEC JTC 1/SC 29/WG 11, vol. JCTVC-K0218, JPN6017046321, October 2012 (2012-10-01), pages 1 - 3, ISSN: 0003976884 | Non-patent | – | – | Search report |
406 members in 29 offices
Members406
| Document | Office | Kind | |
|---|---|---|---|
| CA2870039A1 | Canada | A1 | |
| CA3056122A1 | Canada | A1 | |
| WO2013153226A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2013153227A2 | World Intellectual Property Organization (WIPO) | A2 | |
| TW201349878A | Taiwan Province of China | A | |
| WO2013153226A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2013153227A3 | World Intellectual Property Organization (WIPO) | A3 | |
| CA2877045A1 | Canada | A1 | |
| CA3095638A1 | Canada | A1 | |
| CA3214600A1 | Canada | A1 | |
| WO2014001573A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201408074A | Taiwan Province of China | A | |
| TW201409995A | Taiwan Province of China | A | |
| SG11201406493RA | Singapore | A | |
| MX2014012255A | Mexico | A | |
| AU2013246828A1 | Australia | A1 | |
| PH12014502303A1 | Philippines | A1 | |
| PH12014502303B1 | Philippines | B1 | |
| AU2013283173A1 | Australia | A1 | |
| US2015023409A1 | United States of America | A1 | |
| US2015023434A1 | United States of America | A1 | |
| SG11201408612TA | Singapore | A | |
| KR20150013521A | Republic of Korea | A | |
| PH12014502882A1 | Philippines | A1 | |
| PH12014502882B1 | Philippines | B1 | |
| IL236285A0 | Israel | A0 | |
| IL236285D0 | Israel | D0 | |
| KR20150020538A | Republic of Korea | A | |
| CL2014003507A1 | Chile | A1 | |
| EP2842313A2 | European Patent Office (EPO) | A2 | |
| EP2842318A2 | European Patent Office (EPO) | A2 | |
| KR20150029723A | Republic of Korea | A | |
| MX2014016063A | Mexico | A | |
| CL2014002739A1 | Chile | A1 | |
| EP2868103A1 | European Patent Office (EPO) | A1 | |
| CN104620584A | China | A | |
| CN104641647A | China | A | |
| CN104685893A | China | A | |
| JP2015516747A | Japan | A | |
| JP2015516748A | Japan | A | |
| US2015208095A1 | United States of America | A1 | |
| JP2015526006A | Japan | A | |
| HK1205839A | Hong Kong, China | A | |
| HK1205839A1 | Hong Kong, China | A1 | |
| ZA201407815B | South Africa | B | |
| ZA201500558B | South Africa | B | |
| TWI527466B | Taiwan Province of China | B | |
| AU2013283173B2 | Australia | B2 | |
| HK1210342A | Hong Kong, China | A | |
| HK1210342A1 | Hong Kong, China | A1 | |
| RU2014145559A | Russian Federation | A | |
| AU2016204304A1 | Australia | A1 | |
| TWI544803B | Taiwan Province of China | B | |
| RU2015102812A | Russian Federation | A | |
| AU2013246828B2 | Australia | B2 | |
| JP5993083B2 | Japan | B2 | |
| TW201633777A | Taiwan Province of China | A | |
| SG10201606616WA | Singapore | A | |
| KR101667341B1 | Republic of Korea | B1 | |
| EP2842313B1 | European Patent Office (EPO) | B1 | |
| TWI558182B | Taiwan Province of China | B | |
| RU2603531C2 | Russian Federation | C2 | |
| EP2868103B1 | European Patent Office (EPO) | B1 | |
| AU2016259446A1 | Australia | A1 | |
| KR101686088B1 | Republic of Korea | B1 | |
| MX344485B | Mexico | B | |
| KR20160145843A | Republic of Korea | A | |
| PT2842313T | Portugal | T | |
| EP2842318B1 | European Patent Office (EPO) | B1 | |
| DK2842313T3 | Denmark | T3 | |
| JP2017022724A | Japan | A | |
| TW201705765A | Taiwan Province of China | A | |
| PT2868103T | Portugal | T | |
| DK2868103T3 | Denmark | T3 | |
| TWI575940B | Taiwan Province of China | B | |
| ES2607438T3 | Spain | T3 | |
| PT2842318T | Portugal | T | |
| EP3151566A1 | European Patent Office (EPO) | A1 | |
| CL2016001115A1 | Chile | A1 | |
| DK2842318T3 | Denmark | T3 | |
| TWI584637B | Taiwan Province of China | B | |
| JP6133400B2 | Japan | B2 | |
| SG10201702988RA | Singapore | A | |
| EP3174295A1 | European Patent Office (EPO) | A1 | |
| TWI586179B | Taiwan Province of China | B | |
| ES2614910T3 | Spain | T3 | |
| HUE031183T2 | Hungary | T2 | |
| ES2620707T3 | Spain | T3 | |
| JP2017118564AThis record | Japan | A | |
| PL2842313T3 | Poland | T3 | |
| PL2842318T3 | Poland | T3 | |
| PL2868103T3 | Poland | T3 | |
| JP2017123659A | Japan | A | |
| HUE031264T2 | Hungary | T2 | |
| MX349567B | Mexico | B | |
| UA114909C2 | Ukraine | C2 | |
| BR112014025496A2 | Brazil | A2 | |
| UA115240C2 | Ukraine | C2 | |
| RU2635251C2 | Russian Federation | C2 | |
| TW201742452A | Taiwan Province of China | A |
22 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written request for registration of change of nameJAPANESE INTERMEDIATE CODE: R313533S533 | S533 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2017118564
- Application
- 19896
Titles2
- Japanese
- ビデオ・データストリーム・コンセプト
- English
- Video data stream concept
Classification
- CPC, 23
- H04N21/234327
- H04N19/46
- H04N19/167
- G06F15/173
- H04N21/4621
- H04N21/4728
- H04N19/70
- H04N19/174
- H04N19/423
- H04N19/436
- H04N19/188
- H04N19/67
- H04L47/10
- H04N19/68
- H04L12/56
- H04L12/66
- H04L47/31
- H04N19/55
- H04N19/91
- H04N19/30
- H04N21/23605
- H04N21/2383
- H04N21/64776
- IPC, 4
- H04N19 70
- H04N19 30
- H04L47 31
- H04L47 43