Method and apparatus for supporting avc in mp4
Abstract
Problem to be solved.To support a new capability provided by a new video coding standard and improve a storage method so as to remove restrictions in an existing storage method.
Solution.Subsample metadata that defines a subsample in each sample of multimedia data is created. In addition, a file related to this multimedia data is generated. This file contains subsample metadata and other information related to multimedia data. [Selection diagram] Fig. 2

Term
Projected expiry 4 January 2030.
- Priority
- Filed
- Published
- Today
- Projected expiry
74 claims: 30 independent, 44 dependent
- 1マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するステップと、 上記マルチメディアデータに関連し、上記サブサンプルグループメタデータを含むファイルを生成するステップとを有するデータ処理方法。
- 2上記複数のサブサンプルのそれぞれは、復号によりサンプルを部分的に再構築するデータが得られるサンプルのサブユニットである請求項1記載のデータ処理方法。
- 3上記サブサンプルグループメタデータを作成するステップは、 符号化されたマルチメディアデータを含むファイルを受け取るステップと、 上記マルチメディアデータ内の複数のサブサンプルの境界を特定する情報を抽出するステップと、 上記抽出された情報に基づいて、上記サブサンプルメタデータを定義するステップとを有する請求項1記載のデータ処理方法。
- 4上記サブサンプルグループメタデータを作成するステップは、 上記サブサンプルグループメタデータを予め定義された一組のデータ構造に組織化するステップを有することを特徴とする請求項1記載のデータ処理方法。
- 5上記サンプルグループメタデータを作成するステップは、 予め定義されたデータ構造の組内のデータの各繰り返しシーケンスをシーケンス出現への参照情報と出現回数とを表す情報に変換するステップと有することを特徴とする請求項4記載のデータ処理方法。
- 6上記予め定義されたデータ構造の組は、サブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造とを含むことを特徴とする請求項4記載のデータ処理方法。
- 7マルチメディアデータに関連するファイルを復号システムに送信するステップと、 復号システムにおいて、上記マルチメディアデータに関連するファイルを受信するステップと、 復号システムにおいて、上記マルチメディアデータに関連するファイルからサブサンプルグループメタデータを抽出するステップとを更に有し、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられることを特徴とする請求項1記載のデータ処理方法。
- 8マルチメディアデータ内の複数のサブサンプルのグループを定義するサブサンプルグループメタデータを含むマルチメディアデータに関連するファイルを受信するステップと、 上記ファイルからサブサンプルグループメタデータを抽出するステップとを有し、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられることを特徴とするデータ処理方法。
- 9上記複数のサブサンプルのそれぞれは、復号によりサンプルを部分的に再構築するデータが得られるサンプルのサブユニットであることを特徴とする請求項8記載のデータ処理方法。
- 10上記抽出されたサブサンプルメタデータを用いて、上記マルチメディアデータ内の複数のサブサンプルを特定するステップと、 上記複数のサブサンプルから選択されたサブサンプルを結合して、メディアデコーダに送信するためのパケットを生成するステップとを更に有する請求項8記載のデータ処理方法。
- 11上記抽出されたサブサンプルグループメタデータは、一組の予め定義されたデータ構造に組織化されることを特徴とする請求項8記載のデータ処理方法。
- 12上記予め定義されたデータ構造の組は、サブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造とを含むことを特徴とする請求項11記載のデータ処理方法。
- 13マルチメディアデータの各サンプル内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するステップと、 マルチメディアデータの複数の部分に関する1つ以上のパラメータセットを特定するパラメータセットメタデータを作成するステップと、 上記サブサンプルメタデータと、パラメータセットメタデータとを含む上記マルチメディアデータに関連するファイルを生成するステップとを有するデータ処理方法。
- 14上記サブサンプルグループメタデータを作成するステップは、 上記サブサンプルメタデータをサブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造と含む一組の予め定義されたデータ構造に組織化するステップを有することを特徴とする請求項13記載のデータ処理方法。
- 15上記マルチメディアデータの複数の部分のそれぞれは、上記マルチメディアデータ内のサンプル及びサブサンプルのうちのいずれかであることを特徴とする請求項13記載のデータ処理方法。
- 16上記パラメータセットメタデータを作成するステップは、 上記パラメータセットメタデータを上記1つ以上のパラメータセットに関する記述的な情報を含む第1のデータ構造と、上記1つ以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を定義する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造に組織化するステップを有することを特徴とする請求項13記載のデータ処理方法。
- 17マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータと、該マルチメディアデータに関する1つ以上のパラメータセットを特定するパラメータセットメタデータとを含むマルチメディアデータに関連するファイルを受け取るステップと、 上記ファイルから上記サブサンプルメタデータと上記パラメータセットメタデータとを抽出するステップとを有し、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられ、上記抽出されたパラメータセットメタデータは、後に上記1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定するために用いられることを特徴とするデータ処理方法。
- 18上記マルチメディアデータの複数の部分のそれぞれは、上記マルチメディアデータ内のサンプル及びサブサンプルのうちのいずれかであることを特徴とする請求項17記載のデータ処理方法。
- 19上記抽出されたパラメータセットメタデータは、上記1つ以上のパラメータセットに関する記述的な情報を含む第1のデータ構造と、上記1つ以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を定義する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造を含む予め定義されたデータ構造の組に組織化されることを特徴とする請求項17記載のデータ処理方法。
- 20上記抽出されたサブサンプルメタデータは、サブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造と含む一組の予め定義されたデータ構造に組織化されることを特徴とする請求項17記載のデータ処理方法。
- 21マルチメディアデータの各サンプル内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するステップと、 マルチメディアデータの複数の部分に関する1つ以上のパラメータセットを特定するパラメータセットメタデータを作成するステップと、 上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータを作成するステップと、 上記サブサンプルメタデータ、パラメータセットメタデータ、及びサンプルグループメタデータを含む上記マルチメディアデータに関連するファイルを生成するステップとを有するデータ処理方法。
- 22上記サブサンプルグループメタデータを作成するステップは、 上記サブサンプルメタデータをサブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造とに組織化するステップを有することを特徴とする請求項21記載のデータ処理方法。
- 23上記マルチメディアデータの複数の部分のそれぞれは、上記マルチメディアデータ内のサンプル及びサブサンプルのうちのいずれかである請求項21記載のデータ処理方法。
- 24上記パラメータセットメタデータを作成するステップは、 上記パラメータセットメタデータを上記1つ以上のパラメータセットに関する記述的な情報を含む第1のデータ構造と、上記1つ以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を定義する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造に組織化するステップとを有することを特徴とする請求項21記載のデータ処理方法。
- 25上記グループは、上記複数のサンプルの相互依存性に基づいていることを特徴とする請求項21記載のデータ処理方法。
- 26上記サンプルグループメタデータを作成するステップは、 上記サンプルグループメタデータを上記マルチメディアデータ内の複数のサンプルグループに関する記述的な情報を含む第1のデータ構造と、上記複数のサンプルグループ内のそれぞれのサンプルを特定する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造に組織化するステップを有することを特徴とする請求項21記載のデータ処理方法。
- 27マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータと、該マルチメディアデータに関する1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータとを含むマルチメディアデータに関連するファイルを受け取るステップと、 上記ファイルから上記サブサンプルメタデータと、パラメータセットメタデータと、サンプルグループメタデータとを抽出するステップとを有し、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられ、上記抽出されたパラメータセットメタデータは、後に上記1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定するために用いられ、上記抽出されたサンプルグループメタデータは、後に将来の処理において除外できるサンプルを特定するために用いられることを特徴とするデータ処理方法。
- 28上記マルチメディアデータの複数の部分のそれぞれは、上記マルチメディアデータ内のサンプル及びサブサンプルのうちのいずれかである請求項27記載のデータ処理方法。
- 29上記抽出されたパラメータセットメタデータは、上記1つ以上のパラメータセットに関する記述的な情報を含む第1のデータ構造と、上記1つ以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を定義する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造を含む予め定義されたデータ構造の組に組織化されることを特徴とする請求項27記載のデータ処理方法。
- 30上記抽出されたサブサンプルメタデータは、サブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造と含む一組の予め定義されたデータ構造に組織化されることを特徴とする請求項27記載のデータ処理方法。
- 31上記抽出されたサンプルグループメタデータは、上記マルチメディアデータ内の複数のサンプルグループに関する記述的な情報を含む第1のデータ構造と、上記複数のサンプルグループのそれぞれにおけるサンプルを特定する情報を含む第2のデータ構造とを含む予め定義されたデータ構造の組に組織化されることを特徴とする請求項27記載のデータ処理方法。
- 32マルチメディアデータの各サンプル内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するステップと、 マルチメディアデータの複数の部分に関する1つ以上のパラメータセットを特定するパラメータセットメタデータを作成するステップと、 上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータを作成するステップと、 上記マルチメディアデータに関連する複数の切換サンプルセットを定義する切換サンプルメタデータを作成するステップと、 上記サブサンプルメタデータ、マルチメディアデータ、サンプルグループメタデータ及び切換サンプルメタデータを含むマルチメディアデータに関連するファイルを生成するステップとを有するデータ処理方法。
- 33上記サブサンプルグループメタデータを作成するステップは、 上記サブサンプルメタデータをサブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造とを含む一組の予め定義されたデータ構造に組織化するステップを有することを特徴とする請求項32記載のデータ処理方法。
- 34上記マルチメディアデータの複数の部分のそれぞれは、上記マルチメディアデータ内のサンプル及びサブサンプルのうちのいずれかである請求項32記載のデータ処理方法。
- 35上記パラメータセットメタデータを作成するステップは、 上記パラメータセットメタデータを上記1つ以上のパラメータセットに関する記述的な情報を含む第1のデータ構造と、上記1つ以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を定義する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造に組織化するステップを有することを特徴とする請求項32記載のデータ処理方法。
- 36上記グループは、上記複数のサンプルの相互依存性に基づいていることを特徴とする請求項32記載のデータ処理方法。
- 37上記サンプルグループメタデータを作成するステップは、 上記サンプルグループメタデータを上記マルチメディアデータ内の複数のサンプルグループに関する記述的な情報を含む第1のデータ構造と、上記複数のサンプルグループ内のそれぞれのサンプルを特定する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造に組織化するステップを有することを特徴とする請求項32記載のデータ処理方法。
- 38上記複数の切換サンプルセットのそれぞれは、異なる参照サンプルを用いても同じ復号値に復号されるサンプルを含むことを特徴とする請求項32記載のデータ処理方法。
- 39上記切換サンプルメタデータを作成するステップは、 上記切換サンプルメタデータを、一組のネスト化されたテーブルを含むテーブルボックスとして表された予め定義されたデータ構造に組織化するステップを有することを特徴とする請求項32記載のデータ処理方法。
- 40マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータと、マルチメディアデータに関する1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータと、上記マルチメディアデータに関連する複数の切換サンプルセットを定義する切換サンプルメタデータとを含む、マルチメディアデータに関連するファイルを受け取るステップと、 上記ファイルから上記サブサンプルメタデータ、パラメータセットメタデータ、サンプルグループメタデータ及び切換サンプルメタデータを抽出するステップとを有し、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられ、上記抽出されたパラメータセットメタデータは、後に上記1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定するために用いられ、上記抽出されたサンプルグループメタデータは、後に将来の処理において除外できるサンプルを特定するために用いられ、上記抽出された切換サンプルメタデータは、後に特定のサンプルの置換を検出するために用いられることを特徴とするデータ処理方法。
- 41上記マルチメディアデータの複数の部分のそれぞれは、上記マルチメディアデータ内のサンプル及びサブサンプルのうちのいずれかであることを特徴とする請求項40記載のデータ処理方法。
- 42上記抽出されたパラメータセットメタデータは、上記1つ以上のパラメータセットに関する記述的な情報を含む第1のデータ構造と、上記1つ以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を定義する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造を含む予め定義されたデータ構造の組に組織化されることを特徴とする請求項40記載のデータ処理方法。
- 43上記抽出されたサブサンプルメタデータは、サブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、各サブサンプルを記述する情報を含む第3のデータ構造と含む一組の予め定義されたデータ構造に組織化されることを特徴とする請求項40記載のデータ処理方法。
- 44上記グループは、上記複数のサンプルの相互依存性に基づいていることを特徴とする請求項40記載のデータ処理方法。
- 45上記抽出されたサンプルグループメタデータは、上記マルチメディアデータ内の複数のサンプルグループに関する記述的な情報を含む第1のデータ構造及び上記複数のサンプルグループのそれぞれにおけるサンプルを特定する情報を含む第2のデータ構造とを含む一組の予め定義されたデータ構造に組織化されることを特徴とする請求項40記載のデータ処理方法。
- 46上記複数の切換サンプルセットのそれぞれは、異なる参照サンプルを用いても同じ復号値に復号されるサンプルを含むことを特徴とする請求項40記載のデータ処理方法。
- 47上記抽出された切換サンプルメタデータは、一組のネスト化されたテーブルを含むテーブルボックスとして表された予め定義されたデータ構造に組織化されることを特徴とする請求項40記載のデータ処理方法。
- 48データ処理システム上で実行されるアプリケーションプログラムがアクセスするためのデータを格納するメモリ装置において、 当該メモリ装置に格納され、アプリケーションプログラムに用いられるマルチメディアデータに関連するファイル内に設けられ、マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータを含むサンプルグループメタデータを含む複数のデータ構造を備えるメモリ装置。
- 49上記サブサンプルメタデータを含むファイルは、更に関連するマルチメディアデータを含むことを特徴とする請求項48記載のメモリ装置。
- 50上記サブサンプルメタデータを含むファイルは、関連するマルチメディアデータを含むファイルへの参照情報を含むことを特徴とする請求項48記載のメモリ装置。
- 51各サブサンプルを記述する情報を含む第3のデータ構造とを含むことを特徴とする上記複数のデータ構造は、サブサンプルサイズに関する情報を含む第1のデータ構造と、各サンプルにおけるサブサンプルの数に関する情報を含む第2のデータ構造と、請求項48記載のメモリ装置。
- 52データ処理システム上で実行されるアプリケーションプログラムがアクセスするためのデータを格納するメモリ装置において、 当該メモリ装置に格納され、アプリケーションプログラムに用いられるマルチメディアデータに関連するファイル内に設けられ、該マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータとを含むサンプルグループメタデータと、上記マルチメディアデータの複数の部分に関連する1以上のパラメータを定義するパラメータセットメタデータとを含む複数のデータ構造を備えるメモリ装置。
- 53データ処理システム上で実行されるアプリケーションプログラムがアクセスするためのデータを格納するメモリ装置において、 当該メモリ装置に格納され、アプリケーションプログラムに用いられるマルチメディアデータに関連するファイル内に設けられ、該マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータとを含むサンプルグループメタデータと、上記マルチメディアデータの複数の部分に関連する1以上のパラメータを定義するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータとを含む複数のデータ構造を備えるメモリ装置。
- 54データ処理システム上で実行されるアプリケーションプログラムがアクセスするためのデータを格納するメモリ装置において、 当該メモリ装置に格納され、アプリケーションプログラムに用いられるマルチメディアデータに関連するファイル内に設けられ、該マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータとを含むサンプルグループメタデータと、上記マルチメディアデータの複数の部分に関連する1以上のパラメータを定義するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータと、上記マルチメディアデータに関連する複数の切換サンプルセットを定義する切換サンプルメタデータとを含む複数のデータ構造を備えるメモリ装置。
- 55マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータをサブサンプルメタデータを作成するサブサンプルメタデータ作成器と、 上記マルチメディアデータに関連し、上記サブサンプルグループメタデータを含むファイルを生成するファイル生成器とを備えるデータ処理装置。
- 56上記複数のサブサンプルのそれぞれは、復号によりサンプルを部分的に再構築するデータが得られるサンプルのサブユニットである請求項55記載のデータ処理装置。
- 57上記メタデータ作成器は、符号化されたマルチメディアデータを含むファイルを受け取り、上記マルチメディアデータ内の複数のサブサンプルの境界を特定する情報を抽出し、該抽出された情報に基づいて、上記サブサンプルメタデータを定義することを特徴とする請求項55記載のデータ処理装置。
- 58マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータを含むマルチメディアデータに関連するファイルを受け取り、該ファイルからサブサンプルメタデータを抽出するメタデータ抽出器と、 上記抽出されたサブサンプルメタデータを用いて、複数のサブサンプルファイルのいずれかにアクセスするメディアデータストリームプロセッサとを備えるデータ処理装置。
- 59上記複数のサブサンプルのそれぞれは、復号によりサンプルを部分的に再構築するデータが得られるサンプルのサブユニットであることを特徴とする請求項58記載のデータ処理装置。
- 60上記メディアデータストリームプロセッサは、更に、上記抽出されたサブサンプルメタデータを用いて、マルチメディアファイル内の複数のサブサンプルを特定し、該複数のサブサンプルから選択されたサブサンプルを結合して、メディアデコーダに送信するためのパケットを生成することを特徴とすることを特徴とする請求項58記載のデータ処理装置。
- 61マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータをサブサンプルメタデータと、マルチメディアデータの複数の部分について、1つ以上のパラメータセットを特定するパラメータセットメタデータを作成するサブサンプルメタデータ作成器と、 上記マルチメディアデータに関連し、上記サブサンプルグループメタデータ及びパラメータセットメタデータを含むファイルを生成するファイル生成器とを備えるデータ処理装置。
- 62マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータと、マルチメディアデータの複数の部分について、1つ以上のパラメータセットを特定するパラメータセットメタデータとを含むマルチメディアデータに関連するファイルを受け取り、該ファイルから該サブサンプルメタデータ及びパラメータセットメタデータを抽出するメタデータ抽出器と、 上記抽出されたサブサンプルメタデータを用いて、複数のサブサンプルファイルのいずれかにアクセスし、上記抽出されたパラメータセットメタデータを用いて、1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定するメディアデータストリームプロセッサとを備えるデータ処理装置。
- 63マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータをサブサンプルメタデータと、マルチメディアデータの複数の部分について、1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータとを作成するサブサンプルメタデータ作成器と、 上記マルチメディアデータに関連し、上記サブサンプルグループメタデータと、パラメータセットメタデータと、サンプルグループメタデータとを含むファイルを生成するファイル生成器とを備えるデータ処理装置。
- 64マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータと、マルチメディアデータの複数の部分について、1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータとを含むマルチメディアデータに関連するファイルを受け取り、該ファイルから該サブサンプルメタデータと、パラメータセットメタデータと、サンプルグループメタデータとを抽出するメタデータ抽出器と、 上記抽出されたサブサンプルメタデータを用いて、複数のサブサンプルファイルのいずれかにアクセスし、上記抽出されたパラメータセットメタデータを用いて、1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定し、上記抽出されたサンプルグループメタデータを用いて、将来の処理において除外できるサンプルを特定するメディアデータストリームプロセッサとを備えるデータ処理装置。
- 65マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータをサブサンプルメタデータと、マルチメディアデータの複数の部分について、1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータと、上記マルチメディアデータに関連する複数の切換サンプルセットを定義する切換サンプルメタデータとを作成するサブサンプルメタデータ作成器と、 上記マルチメディアデータに関連し、上記サブサンプルグループメタデータと、パラメータセットメタデータと、サンプルグループメタデータと、切換サンプルメタデータとを含むファイルを生成するファイル生成器とを備えるデータ処理装置。
- 66マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータと、マルチメディアデータの複数の部分について、1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータと上記マルチメディアデータに関連する複数の切換サンプルセットを定義する切換サンプルメタデータとを含むマルチメディアデータに関連するファイルを受け取り、該ファイルから該サブサンプルメタデータと、パラメータセットメタデータと、サンプルグループメタデータと、切換サンプルメタデータとを抽出するメタデータ抽出器と、 上記抽出されたサブサンプルメタデータを用いて、複数のサブサンプルファイルのいずれかにアクセスし、上記抽出されたパラメータセットメタデータを用いて、1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定し、上記抽出されたサンプルグループメタデータを用いて、将来の処理において除外できるサンプルを特定し、上記抽出された切換サンプルメタデータを用いて、後に特定のサンプルの置換を検出するメディアデータストリームプロセッサとを備えるデータ処理装置。
- 67マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するサンプルメタデータ作成手段と、 上記マルチメディアデータに関連し、上記サブサンプルグループメタデータを含むファイルを生成するファイル生成手段とを備えるデータ処理装置。
- 68マルチメディアデータ内の複数のサブサンプルのグループを定義するサブサンプルグループメタデータを含むマルチメディアデータに関連するファイルを受信するファイル受信手段と、 上記ファイルからサブサンプルグループメタデータを抽出するメタデータ抽出手段とを備え、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられることを特徴とするデータ処理装置。
- 69マルチメディアデータの各サンプル内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するサブサンプルメタデータ作成手段と、 マルチメディアデータの複数の部分に関する1つ以上のパラメータセットを特定するパラメータセットメタデータを作成するパラメータセットメタデータ作成手段と、 上記サブサンプルメタデータと、パラメータセットメタデータとを含む上記マルチメディアデータに関連するファイルを生成するファイル生成手段とを備えるデータ処理装置。
- 70マルチメディアデータ内の複数のサンプルのそれぞれを定義するサブサンプルメタデータと、該マルチメディアデータに関する1つ以上のパラメータセットを特定するパラメータセットメタデータとを含むマルチメディアデータに関連するファイルをファイル受信手段と、 上記ファイルから上記サブサンプルメタデータと上記パラメータセットメタデータとを抽出するメタデータ抽出手段とを備え、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられ、上記抽出されたパラメータセットメタデータは、後に上記1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定するために用いられることを特徴とするデータ処理装置。
- 71マルチメディアデータの各サンプル内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するサブサンプルメタデータ作成手段と、 マルチメディアデータの複数の部分に関する1つ以上のパラメータセットを特定するパラメータセットメタデータをパラメータセットメタデータ作成手段と、 上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータを作成するサンプルメタデータ作成手段と、 上記サブサンプルメタデータ、パラメータセットメタデータ、及びサンプルグループメタデータを含む上記マルチメディアデータに関連するファイルを生成するファイル生成手段とを備えるデータ処理装置。
- 72マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータと、該マルチメディアデータに関する1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータとを含むマルチメディアデータに関連するファイルをファイル受信手段と、 上記ファイルから上記サブサンプルメタデータと、パラメータセットメタデータと、サンプルグループメタデータとを抽出するメタデータ抽出手段と、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられ、上記抽出されたパラメータセットメタデータは、後に上記1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定するために用いられ、上記抽出されたサンプルグループメタデータは、後に将来の処理において除外できるサンプルを特定するために用いられることを特徴とするデータ処理装置。
- 73マルチメディアデータの各サンプル内の複数のサンプルのそれぞれを定義するサブサンプルメタデータを作成するサブサンプルメタデータ作成手段と、 マルチメディアデータの複数の部分に関する1つ以上のパラメータセットを特定するパラメータセットメタデータを作成するパラメータセットメタデータ作成手段と、 上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータを作成するサンプルメタデータ作成手段と、 上記マルチメディアデータに関連する複数の切換サンプルセットを定義する切換サンプルメタデータを作成する切換サンプルメタデータ作成手段と、 上記サブサンプルメタデータ、マルチメディアデータ、サンプルグループメタデータ及び切換サンプルメタデータを含むマルチメディアデータに関連するファイルを生成するファイル生成手段とを備えるデータ処理装置。
- 74マルチメディアデータの各サンプル内の複数のサブサンプルのそれぞれを定義するサブサンプルメタデータと、マルチメディアデータに関する1つ以上のパラメータセットを特定するパラメータセットメタデータと、上記マルチメディアデータ内の複数のサンプルのグループを定義するサンプルグループメタデータと、上記マルチメディアデータに関連する複数の切換サンプルセットを定義する切換サンプルメタデータとを含む、マルチメディアデータに関連するファイルを受け取るファイル受信手段と、 上記ファイルから上記サブサンプルメタデータ、パラメータセットメタデータ、サンプルグループメタデータ及び切換サンプルメタデータを抽出するメタデータ抽出手段とを備え、 上記抽出されたサブサンプルグループメタデータは、後に複数のサンプルのいずれかにアクセスするために用いられ、上記抽出されたパラメータセットメタデータは、後に上記1以上のパラメータセットと上記マルチメディアデータの複数の部分との間の関係を判定するために用いられ、上記抽出されたサンプルグループメタデータは、後に将来の処理において除外できるサンプルを特定するために用いられ、上記抽出された切換サンプルメタデータは、後に特定のサンプルの置換を検出するために用いられることを特徴とするデータ処理装置。
Independent claims74
121 paragraphs, as filed
Related application This application was filed on February 25, 2002, US Patent Provisional Application No. 60 / 359,606, March 5, 2002, US Patent Provisional Application No. 60 / 361,773, March 8, 2002. Claims priority over the filed US Patent Provisional Application Nos. 60 / 363,643. These documents are incorporated herein by reference.
The present invention relates to storage and retrieval of audiovisual content in multimedia file formats, and more particularly to file formats compatible with ISO media file formats.
Copyright notice / permission Some of the disclosures in this specification include material subject to copyright protection. The copyright owner does not object to the fax copy in which this specification is found to be a patent application to the Patent and Trademark Office, but claims all other copyrights. The following indications apply to the software and data described below, as well as the accompanying drawings. : Copyright (c) 2003: All copyrights belong to Sony Electronics Inc. (Copyright (c) 2003, Sony Electronics, Inc., All Rights Reserved).
The demand for networks, multimedia, databases and other digital capacities has grown rapidly, and many multimedia coding and storage methods have been developed. One of the well-known file formats for encoding and storing audiovisual data is the Quicktime file format developed by Apple Competer. The International Organization for Standardization (ISO) Multimedia File Format ISO / IEC 14496-12 Information Technology Encoding of Audiovisual Objects-Part 12 (Multimedia file format, ISO / IEC 14496) -12, It was used as the basic technology of Information Technology-Coding of audio-visual objects-Part 12). The ISO media file format (also known as the ISO file format) is used as a template for two standard file formats: (1) Called MP4 developed by the Motion Picture Experts Group (MPEG) (ISO / IEC 14496-14, Information Technology, Audiovisual Object Coding, Part 14 (ISO / IEC 14496-14, Information Technology--Coding of) audio-visual objects--Part 14): MP4 file format) MPEG-4 file format. (2) File format for JPEG2000 (ISO / IEC15444-1) developed by Joint Photographic Expert Group (JPEG).
The ISO media file format consists of an object-oriented structure called a box (also called an atom or object). Two important top-level boxes contain media data or metadata. Most boxes describe a hierarchy of metadata that provides descriptive, structural, and temporal information about actual media data. This collection of boxes is contained in a box known as a movie box. The media data itself may be contained in the media data box or may exist outside. Each media data stream is called a track (also called an elementary stream or simply a stream).
The main metadata is the movie object. The movie box contains a track box that describes the media data expressed in time. There are various types of media data for tracks (eg, video data, audio data, format screen representations (BIFS), etc.). Each track is further divided into samples (also called access units or pictures).
The sample represents a unit of media data at a particular time. The sample metadata is contained in a set of sample boxes. Each trackbox contains a sample tablebox metadata box that contains a box that provides the time for each sample, the size expressed in bytes, the position of the media data (outside the file or inside the file), etc. There is. Samples are the smallest data components that can represent timing, location and other metadata information.
In recent years, the International Telecommunication Union (ITU) Video Group and Image Coding Expert Group (VCEG) of MPEG have joined the Joint Video Team (JVT), ITU Recommendation H.264 or MPEG4 Part 10, Advanced Video Coding / Collaboration has begun to develop a new image coding / decoding (codec) standard called the Decoding Standard (Advanced Video Codec: AVC) or the JVT Codec. These terms and abbreviations such as H.264, JVT, and AVC are used interchangeably here.
The JVT codec design distinguishes between two different conceptual layers: the image coding layer (VCL) and the network abstraction layer (NAL). The VCL includes parts of the codec related to motion compensation, coefficient conversion coding, entropy coding, and other coding. The output of the VCL is a slice, and each slice contains a set of macroblocks and associated header information. NAL extracts the VCL from the details of the transport layer used to carry the VCL data. The VCL defines a comprehensive, transmission-independent representation of information at higher levels than slices. NAL defines the interface between the video codec itself and the outside world. Internally, NAL uses NAL packets. The NAL packet contains a type field that indicates the type of payload and a bit set in the payload. The data in a single slice can be further subdivided into different pieces of data.
In many existing video coding formats, the encoded stream data contains various types of headers, including parameters that control the decoding process. For example, the MPEG-2 video standard places sequence headers, extended group of pictures (hereinafter referred to as GOP) and picture headers in front of the video data corresponding to those items. In JVT, the information needed to decrypt VCL data is grouped into parameter sets. Each parameter set is given an identifier that will later be used as reference information from the slice. The parameter set may be sent out of stream (out of band) instead of in in stream (in band).
Existing file formats do not have the ability to store parameter sets associated with encoded media data, and they are efficient so that they can be efficiently retrieved and transmitted. It also has no function for linking media data (ie, sample or subsample) to a parameter set.
In the ISO media file format, the smallest unit that can be accessed without using parsing media data is the sample, the entire AVC. In many coding formats, the sample can be further subdivided into smaller units called subsamples (also called sample fragments or access unit fragments). In AVC, subsamples correspond to slices. However, existing file formats do not support access to sample subparts. In a system that needs to flexibly generate packets for streaming based on the data stored in a file, the JVT media data for streaming cannot be flexibly packetized without access to the subsamples.
Existing storage formats have restrictions on switching between streams stored with different bandwidths as the network state changes when streaming media data. One of the main requirements in a typical streaming scenario is to scale the bit transmission rate of compressed data in response to changing network conditions. This is usually achieved by encoding multiple streams with different bandwidths and quality settings for typical network conditions and storing them in one or more files. The server can switch between these pre-encoded streams depending on network conditions. With existing file formats, stream switching is only possible if the sample can be reconstructed independently of the preceding sample. Such a sample is called an I-frame. Currently, stream switching is not supported for samples that are rebuilt depending on the previous sample (ie, P-frames or B-frames that are rebuilt with reference to multiple samples).
The AVC standard provides a tool called switching picture (called SI picture and SP picture) that provides efficient switching between streams, random access and error recovery, and other features. A switching picture is a special kind of picture whose value to be reconstructed is exactly equal to the picture being switched. As the switching picture, a reference picture different from the reference picture used for predicting the corresponding picture can be used, and as a result, encoding can be performed more efficiently using an I-frame. In order to make efficient use of the switching pictures stored in the file, it is necessary to know which set of pictures are equivalent and which pictures are used for prediction. Existing file formats do not provide this information, so this information needs to be extracted by parsing the coded (such processing is inefficient and time consuming).
Therefore, it is hoped that storage methods will be improved to accommodate the new capabilities provided by the new video coding standards and to remove the constraints of existing storage methods.
<p> Create parameter set metadata that identifies one or more parameter sets for multiple parts of multimedia data. In addition, a file related to this multimedia data is generated. This file contains parameter set metadata and other information related to multimedia data.</p><p> Create sample group metadata that defines groups of samples in multimedia data. The groups are based on sample interdependence. In addition, it will generate files related to multimedia data. This file contains sample group metadata and other information related to multimedia data.</p>
<figref num="1">It is a block diagram of one Example of a coding system.</figref><figref num="2">It is a block diagram of one Example of a decoding system.</figref><figref num="3">It is a block diagram of the computer environment suitable for the realization of this invention.</figref><figref num="4">It is a flowchart about the process for storing the subsample metadata in a coding system.</figref><figref num="5">It is a flowchart about the process for using the subsample metadata in a decoding system.</figref><figref num="6">It is a figure explaining the extended MP4 media stream model which has a subsample.</figref><figref num="7A">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7B">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7C">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7D">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7E">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7F">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7G">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7H">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7I">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7J">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="7K">It is a figure which shows an exemplary data structure for storing a subsample metadata.</figref><figref num="8">It is a flowchart about the process for storing the parameter set metadata in a coding system.</figref><figref num="9">It is a flowchart about the process for using the parameter set metadata in a decoding system.</figref><figref num="10A">FIG. 5 illustrates an exemplary data structure for storing parameter set metadata.</figref><figref num="10B">FIG. 5 illustrates an exemplary data structure for storing parameter set metadata.</figref><figref num="10C">FIG. 5 illustrates an exemplary data structure for storing parameter set metadata.</figref><figref num="10D">FIG. 5 illustrates an exemplary data structure for storing parameter set metadata.</figref><figref num="10E">FIG. 5 illustrates an exemplary data structure for storing parameter set metadata.</figref><figref num="11">It is a figure which shows an exemplary extended group of pictures (GOP).</figref><figref num="12">It is a flowchart about the process for storing sequence metadata in a coding system.</figref><figref num="13">It is a flowchart about the process for using sequence metadata in a decoding system.</figref><figref num="14A">FIG. 5 shows an exemplary data structure for storing sequence metadata.</figref><figref num="14B">FIG. 5 shows an exemplary data structure for storing sequence metadata.</figref><figref num="14C">FIG. 5 shows an exemplary data structure for storing sequence metadata.</figref><figref num="14D">FIG. 5 shows an exemplary data structure for storing sequence metadata.</figref><figref num="14E">FIG. 5 shows an exemplary data structure for storing sequence metadata.</figref><figref num="15A">It is a figure which shows the usage of the switching sample set for bit stream switching.</figref><figref num="15B">It is a figure which shows the usage of the switching sample set for bit stream switching.</figref><figref num="15C">It is a flowchart which shows an Example of the process for determining the point of switching between two bitstreams.</figref><figref num="16">It is a flowchart about the process for storing the switching sample metadata in a coding system.</figref><figref num="17">It is a flowchart about the process for using the switching sample metadata in a decoding system.</figref><figref num="18">It is a figure which shows the exemplary data structure for storing the switching sample metadata.</figref><figref num="19A">It is a figure which shows the usage of the switching sample set to realize a random access entry point.</figref><figref num="19B">It is a figure which shows the usage of the switching sample set to realize a random access entry point.</figref><figref num="19C">It is a flowchart about one Example of the process for determining a random access point of a sample.</figref><figref num="20A">It is a figure which shows the usage of the switching sample set to realize error recovery.</figref><figref num="20B">It is a figure which shows the usage of the switching sample set to realize error recovery.</figref><figref num="20C">It is a flowchart about one Example of the process for realizing error recovery at the time of sample transmission.</figref>
Hereinafter, examples of the present invention will be described in detail with reference to the accompanying drawings. In the accompanying drawings, similar elements are designated by similar reference numerals. The accompanying drawings exemplify specific embodiments that realize the present invention. These examples will be described in detail so that those skilled in the art can practice the present invention, but other examples are also possible and are logical and mechanical without departing from the scope of the present invention. Physical, electrical, functional and other changes can be made. Therefore, the following detailed description is not construed in a limited manner and the scope of the present invention is defined only by the appended claims.
Overview First, in order to give an overview of the operation of the present invention, FIG. 1 shows an embodiment of the coding system 100. The coding system 100 includes a media encoder 104, a metadata generator 106, and a file creator 108. The media encoder 104 may include, for example, video data (eg, a video object created from a natural source video scene and other external video objects), audio data (eg, natural source audio). Receives media data containing scenes) and audio objects created from other external audio objects), synthetic objects, or any combination of these. The media encoder 104 may include a plurality of separate encoders or sub-encoders for processing various types of media data. The media encoder 104 encodes the data on the media and passes it to the metadata generator 106. The metadata generator 106 generates metadata that provides information about the media data based on the media file format. The media file format may be based on the ISO media file format (or MPEG-4, JPEG2000, etc. derived from it), Quicktime, or any other media file format, with some additional data structures. It may be included. One embodiment defines an additional data structure for storing metadata related to subsamples within the media data. In another embodiment, an addition to store metadata that links a portion of the media data (eg, a sample or subsample) to a corresponding parameter set that contains decryption information previously stored in the media data. Define the target data structure. Yet another example defines an additional data structure for storing metadata related to various groups of samples in the metadata data, created based on the interdependence of the samples in the media data. In yet another embodiment, a switch sample associated with the media data. Define additional data structures for storing metadata related to set). A switching sample set refers to a set of samples that have the same decoding value but may depend on different samples. Yet another embodiment defines various combinations of additional data structures in the file format used. These additional data structures and their functions are described in detail below.
File Creator 108 stores the metadata in a file whose structure is defined by the media file format. Alternatively, the encoded media data may be contained in a partially or completely separate file and linked to the metadata by the reference information contained in the metadata file (eg, via a URL). .. The file created by the file creator 108 can be stored or transmitted via channel 110.
FIG. 2 shows an embodiment of the decoding system 200. The decoding system 200 includes a metadata extractor 204, a media data stream processor 206, a media decoder 210, a compositer 212, and a renderer 214. The decryption system 200 may be provided in the client device and used for local playback. Alternatively, the decryption system 200 may be used for streaming data and may include server and client devices that communicate with each other via a network (eg, the Internet) 208. The server device may include a metadata extractor 204 and a media data stream processor 206. The client device may include a media decoder 210, a synthesizer 212 and a renderer 214.
The metadata extractor 204 is responsible for extracting metadata from a file stored in database 216 or received over a network (eg, from a coding system 100). This file may or may not contain media data associated with the metadata to be extracted. The metadata extracted from the file contains one or more of the additional data structures mentioned above.
The extracted metadata is passed to the media data stream processor 206. The media data stream processor 206 also receives the associated encoded media data. The media data stream processor 206 uses the metadata to generate a media data stream and supplies it to the media decoder 210. In one embodiment, the media data stream processor 206 uses metadata related to the subsample (eg, for packetization) to locate the subsample in the media data. In another embodiment, the media data stream processor 206 uses the metadata associated with the parameter set to link a portion of the media data to the corresponding parameter set. In yet another embodiment, the media data stream processor 206 uses the metadata to define various groups of samples in the metadata to access the samples in a given group (eg, for scalability). In addition, depending on the transmission conditions, the transmission bit rate is reduced by excluding a group of samples that other samples do not depend on). In yet another embodiment, the media data stream processor 206 uses the metadata that defines the switching sample set to provide the same decoding value as the sample to be switched, which is independent of the sample on which the derived sample depends. Identify the switching samples that have (eg, allow switching to streams with different bitrates in P-frames or B-frames). (In still another embodiment, The created media data stream is fed to the media decoder 210 and decoded, either directly (eg, for local playback) or over network 208 (eg, for streaming data). The synthesizer 212 receives the output of the media decoder 210 and composes the scene, and the composited scene is rendered by the renderer 214 in the user display device.
The description with reference to FIG. 3 shown below is intended to provide an overview of computer hardware and other operating components suitable for the practice of the present invention, but this limits the applicable environment. is not it. FIG. 3 shows a computer system suitable for realizing the metadata generator 106 and / or the file creator 108 shown in FIG. 1 or the metadata extractor 204 and / or the media data stream processor 206 shown in FIG. ..
The computer system 340 includes a processor 350, a memory 355, and an input / output device (input / output), each connected to the system bus 365. It has capability) 360. The memory 355 is configured to store instructions that implement the processes described herein by being executed by the processor 350. The input / output device 360 includes various types of computer-readable media, including all types of storage devices accessible by the processor 350. It will be apparent to those skilled in the art that the term "computer-readable medium" also includes a carrier wave in which a digital signal is encoded. Computer system 340 is controlled by operating system software running in memory 355. The input / output device 360 and related media store the operating system software, instructions for processing according to the present invention, and an access unit. The metadata generator 106, file creator 108, metadata extractor 204, and media data stream processor 206 shown in FIGS. 1 and 2 may be independent elements connected to the processor 350, depending on the processor 350. It may be implemented as a computer-executable instruction to be executed. In one embodiment, the computer system 340 may be part of an Internet Service Provider (hereafter referred to as an ISP), or is connected to an ISP via an I / O device 360 to connect to the Internet. The access unit may be transmitted and received via. It is clear that the present invention is not limited to Internet access and Internet websites, but may be applied to directly connected computer systems and private networks.
It is clear that the computer system 340 is just one example of various possible computer systems with different architectures. A typical computer system often includes at least a processor, a memory, and a bus connecting the processor and the memory. It is clear to those skilled in the art that the present invention can also be realized by other computer system configurations including a multiprocessor system, a minicomputer, a mainframe computer and the like. Further, the present invention can also be realized in a distributed computer system environment in which tasks are executed by remote processing devices linked via a communication network.
Sub-Sample Accessibility 4 and 5 show specific examples of processing for storing and retrieving subsample metadata executed in the coding system 100 and the decoding system 200, respectively. This process may be performed by processing logic that includes hardware (eg, circuits, dedicated logic, etc.), software (software that runs on a general purpose computer system or dedicated machine), or a combination of both. By explaining the present invention using these flowcharts for the processing that can be realized by software, those skilled in the art can develop a program including instructions for executing this processing by a properly configured computer. (The computer's processor reads and executes instructions from a computer-readable medium that contains memory). Instructions that can be executed by a computer may be written as a computer programming language or implemented as firmware logic. Written in a programming language that conforms to commonly accepted standards, such instructions are interfaced to different operating systems and can be executed on different types of hardware platforms. Furthermore, the present invention describes the present invention without being based on any particular programming language. It is clear that various programming languages can be used to realize the processes of the present invention disclosed herein. Furthermore, in the art, software may be referred to in various ways as performing actions or producing results (eg, programs, procedures, processes, applications, modules, logic, etc.). These expressions are merely a simple expression that the computer's processor performs an operation or produces a result by executing software by the computer. Further, without departing from the scope of the present invention, the processing steps shown in FIGS. 4 and 5 may be omitted, other steps may be added, and the execution order of the processing steps described here may be changed. Change
FIG. 4 is a flowchart showing an embodiment of the process 400 for creating subsample metadata in the coding system 100. First, processing 400 is started by processing logic receiving a file containing encoded media data (processing block 402). The processing logic then extracts information that identifies the boundaries of the subsamples in the media data (processing block 404). Depending on the file format used, the smallest unit of data stream that can be given a temporal attribute is a sample (defined in ISO media file format or Quicktime), an access unit (defined in MPEG-4). Is called) or a picture (defined in JVT), etc. A subsample represents a contiguous portion of a data stream below the level of the sample. The definition of a subsample depends on the coding format, but comprehensively, a subsample can be decoded as a single entity or combination of subunits, allowing partial reconstruction of the sample. It is an important subunit of the sample. Subsamples are also called access unit fragments. Each subsample has little or no dependency on other subsamples within the same sample, as subsamples often represent a division of the sample's data stream. For example, in JVT, the subsample is a NAL packet. Similarly, in MPEG-4 video, the subsample is a video packet.
In one embodiment, the coding system 100 operates at the network abstraction layer defined in the JVT, as described above. The JVT media data stream is composed of a series of NAL packets, and each NAL packet (also called a NAL unit) includes a header part and a payload part. Certain NAL packets are used to store encoded VCL data for each slice or a single data partition of a slice. Further, the NAL packet may be an information packet containing a supplemental enhancement information (hereinafter referred to as SEI) message. The SEI message represents optional data used when decrypting the corresponding slice. In JVT, the subsample may be a complete NAL packet with both a header and a payload.
In processing block 406, the processing logic creates subsample metadata that defines the subsamples in the media data. In one embodiment, the subsample metadata is organized as a set of predefined data structures (eg, a set of boxes). A set of predefined data structures includes a data structure that contains information about the size of each subsample, a data structure that contains information about the total number of subsamples for each sample, and information that describes each subsample (eg, what is a sub). It can include data structures that include (is it defined as a sample), or other data structures that contain any data related to subsamples.
Next, in one embodiment, the processing logic determines if there is a data structure containing repeated sequences of data (decision box 408). If the result of this determination is positive, the processing logic translates each repeating sequence of data into information that represents a reference to a sequence occurrence and the number of occurrences of the repeating sequence (processing block). 410).
Next, in processing block 412, the processing logic uses a particular media file format (eg, JVT file format) to include subsample metadata in the file associated with the media data. Depending on the media file format, the subsample metadata may be stored with the sample metadata (for example, the sample table box containing the sample data structure can contain the subsample data structure) or stored separately from the sample metadata. You may.
FIG. 5 is a flowchart of a specific example of the process 500 for using the subsample metadata in the decoding system 200. First, processing 500 is started when the processing logic receives a file associated with the encoded media data (processing block 502). The file may be sourced from a database (local or external), a coding system 100 or any other device on the network. The file contains subsample metadata that defines the subsamples in the media data.
The processing logic then extracts the subsample metadata from the file (processing block 504). As mentioned above, the subsample metadata may be stored in a set of data structures (eg, a set of boxes).
Further, in processing block 506, the processing logic uses the extracted metadata to identify subsamples of encoded media data (stored in the same or different files) and feed them to the media decoder. Therefore, a plurality of subsamples are combined to generate a packet, which enables flexible packetization of media data for streaming (for example, error recovery, scalability, etc. can be supported).
Hereinafter, a specific example of the subsample metadata structure will be described based on the extended ISO media file format (called extended MP4). It will be apparent to those skilled in the art that other media file formats can be extended to incorporate similar data structures for storing subsample metadata.
Figure 6 shows an extended MP4 media stream model with subsamples. The presentation data (eg, a presentation containing synchronized audio and video) is represented as movie 602. Movie 602 includes a set of tracks 604. Each track 604 represents a media data stream. Each track 604 is divided into sample 606. Each sample 606 represents a unit of media data at a particular time. Sample 606 is further divided into subsample 608. In the JVT standard, subsample 608 can represent a unit such as a NAL packet or a slice of a single picture, one data portion of a slice containing multiple data portions, an in-band parameter set or an SEI information packet. Alternatively, the subsample 606 may represent any other structural element in the sample, such as code data representing a spatial or temporal domain in the media. In one embodiment, any portion of encoded media data based on some structural or semantic criterion can be treated as a subsample.
Specific examples of the data structure for storing the subsample metadata are shown in FIGS. 7A to 7L.
As shown in Figure 7A, the sample table box 700, which contains the sample metadata box defined by the ISO media file format, is a sub-sample size box 702, a sub-sample description association box. Expanded to include subsample access boxes such as association box 704, sub-sample to sample box 706, and sub-sample description box 708. In one embodiment, the sub-sample access box whether to use scan is optional.
As shown in FIG. 7B, sample 710 is divided into, for example, slices such as slice 712, data portions such as data partition 714, and regions of interest (ROI) such as ROI 716. be able to. Each of these examples represents various examples of dividing a sample into subsamples. Each subsample in a single sample can have different sizes from each other.
The subsample size box 718 consists of a version field that specifies the version of the subsample size box 718, a subsample size field that specifies the default subsample size, and a subsample count field that represents the number of subsamples in the track. Contains an entry size field that specifies the size of the subsample. If the subsample size field is set to 0, the subsamples stored in the subsample size table 720 have different sizes.
If the value of the subsample size field is not set to 0, this field indicates that the subsample size is constant and that the subsample size table 720 is empty. Table 720 may have a 32-bit fixed-length or variable-length field to represent the subsample size. If the field is variable length, the subsample table contains a field that represents the length of the subsample size field in bytes.
As shown in Figure 7C, the subsample-sample box 722 contains a version field that identifies the version of the subsample-sample box 722 and an entry count field that indicates the number of entries in table 723. Subsamples-Each entry in the sample table provides a first sample field that provides an index for the first sample in a sample run that shares the same number of subsamples per sample, and each sample in a sample run. Includes a per-sample subsample (sub-samples-per-sample) field that provides the number of subsamples in.
By using Table 723, we calculate how many samples there are in a run, multiply this number by the appropriate number of subsamples per sample, and sum the results of all the runs to get the track's results. The total number of subsamples can be calculated.
As shown in Figure 7D, the subsample description association box 724 has a version field that specifies the version of the subsample description association box 724 and a description type identifier that indicates the type of subsample to be described (eg, NAL packet, region of interest). And an entry count field that provides the number of entries in table 726. Each entry in Table 726 provides a subsample description type identifier field that indicates the subsample description ID and a first sub that provides an index to the first subsample in the run of subsamples that share the same subsample description ID. Includes sample fields.
The subsample description type identifier controls the use of the subsample description ID field. That is, depending on the type specified in the description type identifier, the subsample description ID field itself specifies the description ID that directly encodes the subsample description in the description ID itself, or the subsample description ID field is different. It can function as an index to a table (that is, a subsample description table described later). For example, if the description type identifier points to a JVT description, the subsample description ID field can contain code that specifies the characteristics of the JVT subsample. In this case, the subsample description ID field contains the lower 8 bits used as a bitmask to indicate the existence of a predefined data portion within the subsample, and to represent the NAL packet type or for future expansion. It may be a 32-bit field including the upper 24 bits of.
As shown in Figure 7E, the subsample description box 728 represents a version field that specifies the version of the subsample description box 728, an entry count field that indicates the number of entries in table 730, and information about the characteristics of the subsample. Includes a description type identifier field that indicates the description type of the subsample description field and a table that contains one or more subsample description entries 730. The subsample description type identifies the type to which descriptive information is relevant and corresponds to the same field in the subsample description association table 724. Each entry in Table 730 contains a subsample description entry that contains information about the characteristics of the associated subsample. The information and format of the description entry depends on the description type field. For example, and if the description type is a parameter set, each description entry contains the value of the parameter set.
Descriptive information can be associated with parameter set information, information related to ROI or any other information needed to characterize a subsample. For parameter sets, the subsample description association table 724 shows the parameter sets associated with each subsample. In such a case, the subsample description ID corresponds to the parameter set identifier. Similarly, subsamples can represent different regions of interest as follows: First, the subsample is defined as one or more encoded macroblocks, and then the subsample description association table is used to represent the division of the encoded microblock of the video frame or image into different regions. To do. For example, two subsample description IDs (eg, subsample description IDs 1 and 2) divide the coded macroblock in the frame into a foreground macroblock and a background macroblock that indicate the allocation to the foreground and background areas, respectively. Can be split.
Different types of subsamples are shown in Figure 7F. The subsamples are unsplit slice 732, slices split into multiple data 734, headers in slice 736. Central data part 738 of one slice, last data part 740 of one slice, SEI information packet It can represent 742 etc. Each of these subsample types may be associated with a specific value for the 8-bit mask 744 shown in Figure 7G. The 8-bit mask may constitute the least significant 8 bits of the 32-bit subsample description ID field, as described above. Figure 7H shows a subsample description association box 724 with a description type identifier equal to "jvtd". Table 726 has a 32-bit subsample description ID field that stores the values shown in Figure 7G.
The compression of data in the subsample description association table will be described with reference to FIGS. 7H to 7K.
As shown in Figure 7I, uncompressed table 726 contains sequence 750 with subsample description IDs that repeat sequence 748. In the compressed table 746, the iterative sequence 750 is compressed with reference information to sequence 748 and information representing the number of times this sequence appears.
In one embodiment shown in FIG. 7J, in the subsample description ID field, the most significant bit is used as the run 754 of the sequence flag, the next 23 bits are used as the occurrence index 756, and the bits lower than this are used as the occurrence length (occurrence). By using it as length) 758, sequence occurrence can be encoded. If run 754 of the sequence flag is set to 1, this indicates that the entry contains a repeating sequence occurrence. If run 754 of the sequence flag is set to 0, this entry indicates a subsample description ID. The appearance index 756 is an index in the subsample description association box 724 of the first appearance of the sequence, and the appearance length 758 indicates the length of the repeated sequence appearance.
In the other embodiment shown in FIG. 7K, the iterative sequence occurrence table 760 is used to represent the occurrence of the iterative sequence. The most significant bit of the subsample description ID field is as a run of sequence flag 762 indicating whether the entry is a subsample description ID, or in the iterative sequence occurrence table 760 that is part of the subsample description association box 724. Used as the sequence index 764 for the entry in. The repeating sequence occurrence table 760 contains a occurrence index field for identifying the index in the subsample description association box 724 of the first item in the repeating sequence, and a sequence length region for specifying the length of the repeating sequence. ..
Parameter set In certain media formats such as JVT, "header" information, including the critical control value needed to properly decode the media data, is separated from the rest of the encoded data / Detached and stored in the parameter set. Thereby, instead of mixing these control values in the stream with the encoded data, it is possible to indicate the required parameter set by the encoded data by using a mechanism such as a unique identifier. This technique allows the transmission of higher level coding parameters to be separated from the coded data. At the same time, the redundancy can be reduced by sharing a common set of control values as a parameter set.
To support efficient transmission of stored media streams using parameter sets, the sender or player is encoded to know when and where to transmit or access the parameter set. The data needs to be linked to the corresponding parameters at high speed. In one embodiment of the invention, this function is achieved by storing data that identifies the relationship between a parameter set of media data and the corresponding portion as parameter set metadata in a media file format.
8 and 9 show the processing for storing and retrieving the parameter set metadata performed by the coding system 100 and the decoding system 200, respectively. This process may be performed by processing logic that includes hardware (eg, circuits, dedicated logic, etc.), software (software that runs on a general purpose computer system or dedicated machine), or a combination of both.
FIG. 8 is a flowchart showing an embodiment of the process 800 for creating the parameter set metadata in the coding system 100. First, processing 800 is initiated when the processing logic receives a file associated with the encoded media data (processing block 802). This file contains a set of coding parameters that specify how part of the media data should be decrypted. Next, the processing logic examines the relationship between a set of coded parameters called a parameter set and the corresponding part of the media data (processing block 804), the parameter set, and the relationship between this parameter set and the media data part. Create parameter set metadata that defines the (processing block 806). The media data portion may be represented by a sample or subsample.
In one embodiment, the parameter set metadata is organized into a set of predefined data structures (eg, a set of boxes). A set of predefined data structures can include a data structure that contains descriptive information about the parameter set and a data structure that contains information that defines the relationship between the sample and the corresponding parameter set. In one embodiment, a predefined set of data structures includes a data structure that contains information that defines the relationship between the subsample and the corresponding parameter set. A data structure containing sub-sample to parameter set association information may and may take precedence over a data structure containing sample to parameter set association information. It does not have to be.
Next, in one embodiment, the processing logic determines if a parameter set data structure containing a repeating sequence of data exists (decision box 808). If the result of this determination is affirmative, the processing logic transforms each repeating sequence of data into reference information for the occurrence of the sequence and information representing the number of occurrences of the sequence (processing block 810).
Next, in processing block 812, the processing logic includes parameter set metadata in the file associated with the media data using a particular media file format (eg, JVT file format). Depending on the media file format, the parameter set metadata may be stored with the track metadata and / or sample metadata (eg, a data structure containing descriptive information about the parameter set may be included in the track box. A data structure containing association information may be included in the sample table box), track metadata and / or may be stored separately from the sample metadata.
FIG. 9 is a flowchart showing an embodiment of the process 900 for using the parameter set metadata in the decoding system 200. First, processing 900 is started when the processing logic receives a file associated with the encoded media data (processing block 902). The file may be received from a database (local or external database), coding system 100 or any other device in the network. The file contains a parameter set for the media data and parameter set metadata that defines the relationship between the parameter set of the media data and the corresponding part (eg, the corresponding sample or subsample).
The processing logic then extracts the parameter set metadata from the file (processing block 904). As mentioned above, parameter set metadata can be stored in a set of data structures (eg, a set of boxes).
Further, in processing block 906, processing logic uses the extracted metadata to determine which parameter set is relevant to a particular media data portion (eg, sample or subsample). This information can be used to control the transmission time of the media data portion and the corresponding parameter set. That is, the parameter set to be used to decode a particular sample or subsample must be transmitted before or with the packet containing the sample or subsample.
Thus, using parameter set metadata allows the parameter set to be transmitted separately over a more reliable channel, which can result in the loss of part of the media stream due to transmission errors or data loss. The sex can be reduced.
Hereinafter, the extended ISO media file format (called extended ISO) will be described using an exemplary parameter set metadata structure. However, it is clear that other file formats may be extended to include various data structures for storing parameter set metadata.
Illustrative data structures for storing parameter set metadata are shown in Figures 10A-10E.
As shown in Figure 10A, the track box 1002, which contains the track metadata box defined by the ISO file format, is extended to include the parameter set description box 1004. In addition, the sample table box 1006, which contains the sample metadata box defined by the ISO file format, is extended to include the sample-parameter set box 1008. In one embodiment, the sample table box 1006 may include a subsample-parameter set box, which may take precedence over the sample to the sample-parameter set box 1008, as described in detail later. it can.
In one embodiment, the parameter set metadata boxes 1004, 1008 are mandatory requirements. In other embodiments, only the parameter set description box 1004 is a mandatory requirement. In yet another embodiment, all parameter set metadata boxes are optional requirements.
As shown in Figure 10B, the parameter set description box 1010 is for a version field that specifies the version of the parameter set description box 1010, a parameter set description count field that indicates the number of entries in table 1012, and for the parameter set itself. Contains a parameter set entry field that contains the entry.
Parameter sets can be referenced from the sample level or subsample level. As shown in Figure 10C, the sample-parameter set box 1014 provides reference information from the sample level to the parameter set. The sample-parameter set box 1014 has a version field that specifies the version of the sample-parameter set box 1014, a default parameter set ID field that specifies the default parameter set ID, and an entry count field that indicates the number of entries in table 1016. Includes. Each entry in table 1016 contains a first sample field that provides an index for the first sample in a run of samples that share the same parameter set, and a parameter set index that specifies an index for the parameter set description box 1010. Includes. When the default parameter set ID is 0, the samples stored in table 1016 have different parameter sets. When the default parameter set ID is 1, only one constant parameter set is used for each sample.
In one embodiment, as described above in relation to the subsample description relation table, the data in table 1016 transforms each iterative sequence into reference information to the first sequence and information representing the number of times this sequence appears. It is compressed by doing.
The parameter set can be referenced from the subsample level by defining the relationship between the parameter set and the subsample. In one embodiment, the relationship between the parameter set and the subsamples is defined using the subsample description association box described above. Figure 10D shows a subsample description association box 1018 with a description type identifier that identifies the parameter set. (For example, the description type identifier is equal to "pars"). The subsample description ID in table 1020 indicates the index in the parameter set description box 1010 based on this description type identifier.
In one embodiment, if a subsample description association box 1018 containing a description type identifier that identifies the parameter set is present, it takes precedence over the sample-parameter set box 1014.
The parameter set may change between the time the parameter set is created and the time the parameter set is used to decode the corresponding portion of the media data. When such a change is made, the decoding system 200 receives a parameter update packet indicating a parameter set change. Parameter set metadata contains data that identifies the state of both pre-update and post-update parameter sets.
As shown in Figure 10E, the parameter set description box 1010 is at time t.<sub>0</sub>The entry for the initial parameter set 1022 created in, and the time t<sub>1</sub>Contains an entry for the update parameter set 1024 created in response to the parameter update packet 1026 received in. The subsample description association box 1018 associates these two parameter sets with the corresponding subsamples.
Sample group The samples in the track can be logically grouped (divided) into sequences (which may be discontinuous) that represent the high-level structure in the media data, but existing file formats are grouped in this way. It does not have a suitable mechanism for expressing and storing the converted data. Advanced encoding formats, such as JVT, organize samples within a single track into multiple groups based on their interdependence. These groups (also referred to here as sequences or sample groups) can be used to identify a set of samples that can be excluded, thereby providing temporal scalability when required by network conditions. Can be supported. By storing the metadata that defines the sample group in the file format, the media transmitter can easily and efficiently realize the above-mentioned features.
As a specific example of the sample group, for example, there is a set in which those samples can be decoded without depending on other samples due to the dependency between frames. In JVT, such a sample group is called an extended group of pictures (hereinafter referred to as an extended GOP). In the extended GOP, the sample can be divided into subsequences. Each subsequence contains a set of samples that are dependent on each other and can be processed as a unit. In addition, the extended GOP sample excludes the top layer sample without affecting the ability of the higher layer samples to be predicted from the lower layer samples only, thereby decoding other samples. It can be hierarchically structured into multiple layers so that it can be. The lowest layer that contains a sample that does not depend on the sample of any other layer is called the base layer. All layers other than the base layer are called enhancement layers.
Figure 11 shows a sample layer with two layers, a base layer 1102 and an enhancement layer 1104. And an exemplary extended GOP divided into two subsequences 1106, 1108. The two subsequences 1106 and 1108 can be discarded independently of each other.
12 and 13 show the processing for storing and retrieving sample group metadata performed by the coding system 100 and the decoding system 200, respectively. This process may be performed by processing logic that includes hardware (eg, circuits, dedicated logic, etc.), software (software that runs on a general purpose computer system or dedicated machine), or a combination of both.
FIG. 12 is a flowchart showing an embodiment of the process 1200 for creating sample group metadata in the coding system 100. First, processing 1200 is started when the processing logic receives a file associated with the encoded media data (processing block 1202). The samples in the track of media data have some kind of interdependence. For example, a track depends on two preceding samples, including an I-frame that does not depend on any other sample, a P-frame and I-frame that depend on a single preceding sample, and any combination of P-frame and B-frame. It may include a frame. Track samples can be logically combined into sample groups (eg, extended GOPs, layers, subsequences, etc.) based on their interdependence.
The processing logic then examines the media data to identify the sample groups within each track (processing block 1204), describes the sample groups, and defines which samples are contained in which sample group. Create metadata (processing block 1206). In one embodiment, the sample group metadata is organized into a set of predefined data structures (eg, a set of boxes). A set of predefined data structures can include a data structure that contains descriptive information about each sample group and a data structure that contains information that identifies the samples contained in each sample group.
Next, in one embodiment, the processing logic determines if a sample group data structure containing a repeating sequence of data exists (decision box 1208). If the result of this determination is positive, the processing logic transforms each repeating sequence of data into reference information for the occurrence of the sequence and information representing the number of occurrences of the sequence (processing block 1210).
Next, in processing block 1212, the processing logic includes sample group metadata within the file associated with the media data using a particular media file format (eg, JVT file format). Depending on the media file format, the sample group metadata may be stored with the sample metadata (the sample group data structure may be included in the sample table box) or may be stored separately from the sample metadata. ..
FIG. 13 is a flowchart showing an embodiment of the process 1300 for using the sample group metadata in the decoding system 200. First, processing 1300 is started when the processing logic receives a file associated with the encoded media data (processing block 1302). The file may be received from a database (local or external database), coding system 100 or any other device in the network. The file contains sample group metadata that defines the sample groups in the media data.
The processing logic then extracts sample group metadata from the file (processing block 1304). As mentioned above, sample group metadata can be stored in a set of data structures (eg, a set of boxes).
Further, in processing block 1306, the processing logic uses the extracted metadata to identify a set of samples that can be excluded without affecting their ability to decode other samples. In one embodiment, this information can be used to access samples in a particular sample group and determine which samples can be discarded in response to changes in network capacity. In another embodiment, the sample group metadata is used to filter the samples so that only a portion of the samples in the track is processed or rendered.
In this way, sample group metadata facilitates selective access to samples and scalability.
Hereinafter, a specific example of the sample group metadata structure will be described in relation to the extended ISO media file format (also referred to as extended MP4). However, it is clear that other file formats may be extended to include various data structures for storing sample group metadata.
Illustrative data structures for storing sample group metadata are shown in Figures 10A-10E.
As shown in Figure 14A, the sample table box 1400 containing the sample metadata box defined by MP4 is expanded to include the sample group box 1402 and the sample group description box 1404. In one embodiment, the sample group metadata boxes 1402 and 1404 are arbitrary elements.
As shown in FIG. 14B, the sample group box 1406 is used to detect a set of samples contained in a particular sample group. Multiple instances of the sample group box 1406 can accommodate different types of sample groups (eg, extended GOPs, subsequences, layers, parameter sets, etc.). The sample group box 1406 has a version field that identifies the version of the sample group box 1406, an entry count field that indicates the number of entries in table 1408, and a sample group identifier field that identifies the type of sample group, in the same sample group. It contains a first sample field that provides an index to the first sample in the run of the included sample, and a sample group description index that specifies the index to the sample group description box.
As shown in Figure 14C, the sample group description box 1410 provides information about the characteristics of the sample group. The sample group description box 1410 contains a version field that identifies the version of the sample group description box 1410, an entry count field that indicates the number of entries in table 1412, a sample group identifier field that specifies the type of sample group, and a sample group. Contains a sample group description field that provides a descriptor.
FIG. 14D illustrates the use of the sample group box 1416 for the layer sample group type. Samples 1 to 11 are divided into three layers based on the interdependence of the samples. At layer 0 (base layer), the samples (samples 1, 6, 11) are only interdependent and independent of the samples in any other layer. At layer 1, the samples (samples 2, 5, 7, 10) depend on the samples in the lower layers (ie, layer 0) and the samples in this layer 1. At Layer 2, the samples (Samples 3, 4, 8, 9) depend on the samples in the lower layers (ie, Layer 0 and Layer 1) and the samples in this Layer 2. Therefore, layer 2 samples can be excluded without affecting the ability to decode samples from lower layers 0 and 1.
The data in the sample group box 1416 shows the relationship between the sample and the layer as described above. As shown here, this data contains a repeating layer pattern 1414. As described above, the repeating layer pattern 1414 is compressed by converting each repeating layer pattern 1414 into reference information to the first layer pattern and information indicating the number of times this pattern appears.
FIG. 14E illustrates the use of sample group box 1418 for the subsequence (sseq) sample group type. Samples 1 to 11 are divided into four subsequences based on the interdependence of the samples. Each subsequence except subsequence 0 at layer 0 contains a sample that the other subsequences do not depend on. Therefore, the samples in the subsequence can be collectively excluded as needed.
The data in the sample group box 1418 shows the relationship between the sample and the subsequence. This data allows random access to the first sample of the corresponding subsequence.
Stream switching One of the main requirements in a typical streaming scenario is to scale the bit transmission rate of compressed data in response to changing network conditions. The simplest way to achieve this is to encode multiple streams with different bit rates and quality settings that correspond to typical network conditions. This allows the server to switch between these pre-encoded streams depending on network conditions.
The JVT standard provides a new kind of picture called a switching picture in which each of the two pictures can be reconstructed like the other picture without using the same frame for prediction. In particular, the JVT, like the I-frame, has two types of switching pictures: SI pictures, which are coded independently of any other picture, and SP pictures, which are coded with reference to other pictures. provide. By using the switching picture, it is possible to switch between streams having different bit transmission speeds and quality settings in response to changes in transmission conditions, and to realize trick modes such as error recovery, fast forward, and rewind.
Here, in order to make effective use of the JVT switching picture, when implementing stream switching, error recovery, trick mode and other features, the player has an alternative representation of which sample of stored media data. And we need to know what the dependencies of these samples are. Existing file formats do not provide such functionality.
In one embodiment of the invention, this problem is solved by defining a switching sample set. A switching sample set represents a set of samples that have the same decoding value but can use different reference samples. Reference samples are samples used to predict the values of other samples. Each element of the switching sample set is called a switching sample. FIG. 15A is a diagram illustrating bitstream switching using a switching sample set.
As shown in FIG. 15A, stream 1 and stream 2 represent two encodings of the same content with different quality and bit rate parameters. Sample S12 is an SP picture that does not appear in either stream and is used to realize switching from stream 1 to stream 2 (switching has directional properties). Samples S12 and S2 are included in the switching sample set. Both S1 and S12 are predicted from sample P12 of track 1, and S2 is predicted from sample P22 of track 2. Samples S12 and S2 use different reference samples, but their decoded values are the same. Therefore, switching from stream 1 to stream 2 (in sample S1 in stream 1 and sample S2 in stream 2) can be done via switching sample S12.
16 and 17 show the processing for storing and retrieving the switching sample metadata performed by the coding system 100 and the decoding system 200, respectively. This process may be performed by processing logic that includes hardware (eg, circuits, dedicated logic, etc.), software (software that runs on a general purpose computer system or dedicated machine), or a combination of both.
FIG. 16 is a flowchart showing an embodiment of the process 1600 for creating switching sample metadata in the coding system 100. First, processing 1600 is initiated when the processing logic receives a file associated with the encoded media data (processing block 1602). This file contains one or more alternative coded data for media data (eg, with different bandwidth and quality settings corresponding to typical network conditions). The alternative coded data includes one or more switching pictures. Such pictures may be included in an alternative media data stream or may be created as separate entities that implement special features such as error recovery or trick mode. Although the method for creating these tracks and switching pictures is not particularly limited in the present invention, it is clear to those skilled in the art that various methods can be used. For example, switching samples may be inserted periodically (eg, every second) between pairs of tracks containing alternative encoded data.
The processing logic then examines the file (processing block 1604) to create a switching sample set containing samples from which the same decoded value is derived, using different reference samples, and defines a switching sample set for the media data. , Create switching sample metadata that describes the samples in the switching sample set (processing block 1606). In one embodiment, the switching sample metadata is organized into a predefined data structure, such as a table box containing a set of nested tables.
Next, in one embodiment, the processing logic determines whether the switching sample metadata structure contains a repeating sequence of data (decision box 1608). If the result of this determination is positive, the processing logic transforms each repeating sequence of data into reference information for the occurrence of the sequence and information representing the number of occurrences of the sequence (processing block 1610).
Next, in processing block 1612, the processing logic includes switching sample metadata in the file associated with the media data using a particular media file format (eg, JVT file format). In one embodiment, the switching sample metadata may be stored on another track designated for stream switching. In other embodiments, the switching sample metadata is stored with the sample metadata (eg, the sequence data structure may be included in the sample table box).
FIG. 17 is a flowchart showing an embodiment of the process 1700 for using the switching sample metadata in the decoding system 200. First, processing 1700 is started when the processing logic receives a file associated with the encoded media data (processing block 1702). The file may be received from a database (local or external database), coding system 100 or any other device in the network. The file contains switching sample metadata that defines the switching sample set associated with the media data.
The processing logic then extracts the sample group metadata from the file (processing block 1704). As mentioned above, the switching sample metadata can be stored, for example, in a data structure such as a table box containing a set of nested tables.
Further, in processing block 1706, the processing logic uses the extracted metadata to detect a switching sample set containing a particular sample and select an alternative sample from the switching sample set. An alternative sample with the same decoding value as the initial sample may be used to switch between two differently encoded bitstreams as the network state changes, resulting in random access within the bitstream. It can provide an entry point and facilitate error recovery.
Hereinafter, a specific example of the switching sample metadata structure will be described in relation to the extended ISO media file format (also referred to as extended MP4). However, it is clear that other file formats may be extended to incorporate various data structures for storing switching sample metadata.
FIG. 18 shows an exemplary data structure for storing switching sample metadata. The exemplary data structure has the form of a switching sample table box containing a set of nested tables. Each entry in Table 1802 identifies one switching sample set. Each switching sample set may have objectively the same (or perceptually the same) reconstruction results, but may be predicted from different reference samples in the same or different tracks (streams) as the switching samples. It consists of a group of switching samples. Each entry in table 1802 is linked to the corresponding table 1804. Table 1804 identifies each switching sample included in the switching sample set. Each entry in Table 1804 is further linked to a corresponding Table 1806 that defines the location of the switching sample (ie its track and sample number), and the tracks are used by the switching sample and the reference sample used by the switching sample. It includes the total number of reference samples and each switching sample used by the switching samples.
As shown in FIG. 15A, in one embodiment, switching sample metadata can be used to switch between differently encoded versions of the same content. In MP4, each alternative coded data is stored as a separate MP4 track, and the "alternate group" in the track header means that the data is the coded data that substitutes for the specific content. Is shown.
FIG. 15B shows a table containing metadata that defines a switching sample set 1502 consisting of samples S2 and S12 shown in FIG. 15A.
FIG. 15C is a flowchart showing an embodiment of the process 1510 for determining a point for switching between two bitstreams. When switching from stream 1 to stream 2, process 1510 first searches for switching sample metadata, and then includes a switching sample with a reference track for stream 1 and a switching sample with a switching sample track for stream 2. Detects the switching sample set of (processing block 1512). Next, the detected switching sample set is evaluated and the switching sample set for which all reference samples of the switching sample having the reference track of stream 1 are available is selected (processing block 1514). For example, if the switching sample with the reference track of stream 1 is a P-frame, then one sample needs to be available before switching. In addition, the samples in the selected switching sample set are used to identify switching points (processing block 1516). That is, the highest reference sample of the switching sample having the reference track of stream 1 through the switching sample having the reference track of stream 1. The sample immediately after the reference sample) and immediately after the switching sample having the switching sample track of stream 2 is considered to be the switching point.
In another embodiment, switching sample metadata can be used to facilitate random access to entry points in a bitstream, as shown in FIGS. 19A-19C.
As shown in FIGS. 19A and 19B, switching sample 1902 includes samples S2, S12. S2 is a P frame predicted from P22 and used during normal stream playback. The S12 is used as a random access point (eg for splices). Once S12 is decoded, stream playback, including decoding P24, continues as if P24 had been decoded after S2.
FIG. 19C is a flow chart illustrating an embodiment of process 1910 for determining a random access point for a sample (eg, sample S on track T). Processing 1910 is initiated by searching the switching sample metadata and discovering all switching sample sets, including switching samples with a switching sample track T (processing block 1912). Next, the detected switching sample set is evaluated, and the switching sample having the switching sample track T selects the switching sample set which is the closest sample preceding the sample S in the decoding order (processing block 1514). Further, from the selected switching sample set, a switching sample (sample SS) other than the switching sample having the switching sample track T is selected as a random access point to the sample S (processing block 1916). During stream playback, the sample SS is decoded instead of the sample S (by decoding all the reference samples identified in the entry for the sample SS).
In yet another embodiment, as shown in FIGS. 20A-20C, switching sample metadata can be used to facilitate error recovery.
As shown in FIGS. 20A and 20B, switching sample 2002 includes samples S2, S12, S22. Sample S2 is predicted from sample P4. Sample S12 is predicted from sample S1. If an error occurs between P2 and P4 of the sample, the switching sample S12 can be decoded instead of the sample S2. Following this, streaming continues from sample P6 as usual. Also, if the error affects sample S1, similarly, switching sample S22 can be decoded instead of sample S2, and subsequent streaming continues from sample P6 as usual.
FIG. 20C is a flowchart showing an embodiment of the process 2012 for facilitating error recovery when transmitting a sample (for example, sample S). Processing 2012 is initiated by searching the switching sample metadata and finding all switching sample sets containing switching samples equal to sample S (processing block 2010). The detected switching sample set is then evaluated and the switching sample with the switching sample SS closest to sample S and known to have the correct reference sample value (via feedback or other source) is selected. (Processing block 2014). Then, the switching sample SS is transmitted instead of the sample S (processing block 2016).
Described storage and retrieval of audiovisual metadata. Although specific embodiments have been shown here, it will be apparent to those skilled in the art that any configuration that achieves the same objectives may be used in place of the particular embodiments shown herein. Therefore, the present application shall include all indications and variations of the present invention.
45 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45
Every citation, both waysCites: the store holds 0 of 1
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2015521825A | Cited by | Japan | Search report |
| JP2015521825A | Cited by | Japan | Search report |
| JPN6012023695; Hannuksela,M.: 'Enhanced Concept of GOP' Joint Video Team (JVT) of ISO/IEC MPEG & ITU-T VCEG(ISO/IEC JTC1/SC29/WG11 and ITU-T SG16 Q.6) JVT-B , 20020123 | Non-patent | – | Examiner |
| JPN7012002686; Lim,Y.-K.,Hannuksela,M.,Singer,D.: 'JVT Codec Transport Ad-Hoc Group Report' Joint Video Team (JVT) of ISO/IEC MPEG & ITU-T VCEG(ISO/IEC JTC1/SC29/WG11 and ITU-T SG16 Q.6) JVT- , 20020125 | Non-patent | – | Examiner |
66 members in 9 offices
Priority claims16
| Document | Office | Kind | Date |
|---|---|---|---|
| 35960602 | United States of America | P | |
| 60359606 | United States of America | – | |
| 36177302 | United States of America | P | |
| 60361773 | United States of America | – | |
| 36364302 | United States of America | P | |
| 60363643 | United States of America | – | |
| 10371464 | United States of America | – | |
| 37146403 | United States of America | A | |
| 2002359606 | – | – | – |
| 2002361773 | – | – | – |
| 2002363643 | – | – | – |
| 2003371464 | – | – | – |
| US20020359606P | – | – | – |
| US20020361773P | – | – | – |
| US20020363643P | – | – | – |
| US20030371464 | – | – | – |
Members66
| Document | Office | Kind | |
|---|---|---|---|
| US2003163477A1 | United States of America | A1 | |
| US2003163781A1 | United States of America | A1 | |
| WO03073767A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO03073768A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO03073769A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO03073770A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003213554A1 | Australia | A1 | |
| AU2003213555A1 | Australia | A1 | |
| AU2003219876A1 | Australia | A1 | |
| AU2003219877A1 | Australia | A1 | |
| WO03098475A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003237120A1 | Australia | A1 | |
| US2004006575A1 | United States of America | A1 | |
| US2004167925A1 | United States of America | A1 | |
| US2004199565A1 | United States of America | A1 | |
| KR20040088526A | Republic of Korea | A | |
| KR20040088541A | Republic of Korea | A | |
| GB0421361D0 | United Kingdom | D0 | |
| KR20040091659A | Republic of Korea | A | |
| KR20040091664A | Republic of Korea | A | |
| EP1481552A1 | European Patent Office (EPO) | A1 | |
| EP1481553A1 | European Patent Office (EPO) | A1 | |
| EP1481554A1 | European Patent Office (EPO) | A1 | |
| EP1481555A1 | European Patent Office (EPO) | A1 | |
| GB2402247A | United Kingdom | A | |
| GB2402248A | United Kingdom | A | |
| GB2402575A | United Kingdom | A | |
| GB2402590A | United Kingdom | A | |
| KR20040106414A | Republic of Korea | A | |
| GB2403835A | United Kingdom | A | |
| EP1500002A1 | European Patent Office (EPO) | A1 | |
| DE10392284T5 | Germany | T5 | |
| DE10392282T5 | Germany | T5 | |
| DE10392280T5 | Germany | T5 | |
| DE10392281T5 | Germany | T5 | |
| DE10392598T5 | Germany | T5 | |
| CN1650626A | China | A | |
| CN1650627A | China | A | |
| CN1650628A | China | A | |
| CN1653818A | China | A | |
| JP2005524128A | Japan | A | |
| JP2005525627A | Japan | A | |
| CN1666195A | China | A | |
| JP2005527885A | Japan | A | |
| GB2402248B | United Kingdom | B | |
| GB2402247B | United Kingdom | B | |
| GB2402575B | United Kingdom | B | |
| GB2403835B | United Kingdom | B | |
| GB2402590B | United Kingdom | B | |
| JP2006505024A | Japan | A | |
| JP2006507553A | Japan | A | |
| CN100379290C | China | C | |
| AU2003213555B2 | Australia | B2 | |
| AU2003219876B2 | Australia | B2 | |
| AU2003213554B2 | Australia | B2 | |
| AU2003219877B2 | Australia | B2 | |
| AU2003219876B8 | Australia | B8 | |
| CN100419748C | China | C | |
| AU2003237120B2 | Australia | B2 | |
| US7613727B2 | United States of America | B2 | |
| JP2010098763A | Japan | A | |
| JP2010104030A | Japan | A | |
| JP2010124479A | Japan | A | |
| JP2010141900AThis record | Japan | A | |
| CN1650628B | China | B | |
| CN1650627B | China | B |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 2010141900
- Publication, DOCDB
- 2010141900
- Publication, EPODOC
- JP2010141900
- Application
- 153
- Application, DOCDB
- 2010000153
- Application, EPODOC
- JP20100000153
Titles3
- Japanese
- MP4においてAVCをサポートするための方法及び装置
- English
- Methods and equipment to support AVC in MP4
- English
- METHOD AND APPARATUS FOR SUPPORTING AVC IN MP4
Classification
- CPC, 11
- H04N21/8451
- H04N5/91
- H04N21/23424
- H04N21/44016
- H04N21/4621
- H04N21/64792
- H04N21/84
- H04N21/85406
- H04N19/46
- H04N5/76
- H04N5/92
- IPC, 4
- H04N7 26
- H04N5 91
- G06F19 00
- G06F17 30