Audio splitting with codec-enforced frame sizes
21 claims: 6 independent, 15 dependent
- 1音声および映像を含むメディア・コンテンツをコンピューティング・システムにより受け取ることと、 前記映像を複数のビデオフレームに、所定のビデオ符号化フレーム・レートに従って前記コンピューティング・システムにより符号化することと、 前記音声を、所定の音声符号化コーデック適用フレーム・サイズを有する、複数の隙間のない音声フレームに、前記コンピューティング・システムにより符号化することと、 複数のコンテンツ・ファイルを前記コンピューティング・システムにより生成することであって、前記複数のコンテンツ・ファイルのそれぞれが、固定時間幅を有する符号化映像部分と符号化済音声部分と有し、前記符号化済映像部分は、1以上の複数の映像フレームを有し、前記符号化済音声部分は、1以上の複数の隙間のない音声フレームを有し、 前記複数のコンテンツ・ファイルのうちの1つまたは複数のコンテンツ・ファイルの前記符号化済音声部分の期間が、それぞれの1以上のコンテンツ・ファイルの前記符号化済映像部分の前記固定時間幅よりも大きいか、またはこれよりも小さ く、 前記1以上のコンテンツ・ファイルの現在のコンテンツ・ファイルの符号化済音声部分の時間長が互いに整数倍でない場合も、長さ調整のために0値の埋め込みを前記現在のコンテンツ・ファイルの前記符号化済音声部分に行わず、前記現在のコンテンツ・ファイルの前記符号化済音声部分を隙間のない音声フレームを用いて満たし、 次のコンテンツ・ファイルの前記音声に、固定時間幅から前記現在のコンテンツ・ファイルの音声のうちの符号化済部分の期間を引くことにより得られる、前記次のコンテンツ・ファイルの前記符号化済音声部分の再生開始位置を示すプレゼンテーション・オフセットを与えて再生し、前記音声が再生されず、前記現在のコンテンツ・ファイルと前記次のコンテンツ・ファイルのつなぎ目で、アーチファクトが発生しないようにする、 ことと、 を含む方法。
- 2前記1以上のコンテンツ・ファイルのそれぞれのコンテンツ・ファイルのそれぞれの符号化済音声部分の前記隙間のない音声フレームの最後が、0で埋められることがない、 請求項1に記載の方法。
- 3前記メディア・コンテンツを前記音声および前記映像にスプリットすることを更に含み、 前記映像を符号化することが、前記固定時間幅に従って映像コーデックを使用することにより前記映像を符号化することを含み、 前記音声を符号化することが、前記コーデック適用フレーム・サイズに従って音声コーデックを使用することにより前記音声を符号化することを含む、 請求項1に記載の方法。
- 4前記音声の符号化済フレームをバッファリングすることと、 前記複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの前記符号化済音声部分を満たすのに必要な符号化済フレーム数を求めることであって、前記フレーム数が、前記複数のファイルのうち前記現在のファイルの前記符号化済音声部分を満たすのに必要なサンプル数を、前記コーデック適用フレーム・サイズで割った数以上の最小の整数であり、割った数には端数があり、端数は、サンプル・オフセットとして保持されることと、 前記複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの前記符号化済音声部分を満たすための前記符号化済フレームがバッファリングされたかどうかを判定することと、 前記符号化済フレームが前記複数のコンテンツ・ファイルの現在のコンテンツ・ファイルの前記符号化済音声部分を満たすほどバッファリングされている場合、前記複数のコンテンツ・ファイルのうちの前記現在のコンテンツ・ファイルの前記符号化済音声部分を前記フレーム数のフレームで満たすことと、 前記符号化済フレームが前記複数のコンテンツ・ファイルの現在のコンテンツ・ファイルの前記符号化済音声部分を満たすほどバッファリングされていない場合、前記音声のフレームを更にバッファリングし、前記複数のコンテンツ・ファイルのうちの前記現在のコンテンツ・ファイルの前記符号化済音声部分を、前記フレーム数のフレームおよび前記追加のフレームで満たすことと、 を更に含む請求項1に記載の方法。
- 5前記複数のコンテンツ・ファイルの現在のコンテンツ・ファイルの前記符号化済音声部分を満たすだけの符号化済フレームがバッファリングされたかどうかを前記判定することが、 バッファリングされたフレームの数と、前記コーデック適用フレーム・サイズとを乗算することと、 前記複数のコンテンツ・ファイルのうちの前のコンテンツ・ファイルからの前記サンプル・オフセットがあればこれを、前記乗算の積に加えることと、 その和が、前記複数のコンテンツ・ファイルのうちの第1のコンテンツ・ファイルを満たすのに必要なサンプルの数以上であるかどうかを判定することと、 を含む 請求項4に記載の方法。
- 6前記複数のコンテンツ・ファイルのうちの次のコンテンツ・ファイルについてサンプル・オフセットがあればこれを求めることを更に含む、 請求項4に記載の方法。
- 7前記サンプル・オフセットを求めることが、 (前記符号化済フレームの数)×(前記コーデック適用フレーム・サイズ)-(前記複数のコンテンツ・ファイルのうちの第1のコンテンツ・ファイルを満たすのに必要なサンプルの数)+(前記複数のコンテンツ・ファイルのうちの前のコンテンツ・ファイルからの前記サンプル・オフセット)を計算することを含む、 請求項6に記載の方法。
- 8前記音声の符号化済フレームをバッファリングすることを更に含み、 前記複数のコンテンツ・ファイルを前記生成することが、 前記複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの前記符号化済音声部分を満たすのに必要なサンプルの数を算出することと、 前記複数のコンテンツ・ファイルのうちの前記現在のコンテンツ・ファイルの前記符号化済音声部分に必要なフレームの数を算出することと、 前記サンプルの数が前記コーデック適用フレーム・サイズで均等に割り切れない場合、前記フレーム数のフレームにフレームを追加することと、 前記複数のコンテンツ・ファイルのうちの前記現在のコンテンツ・ファイルの前記符号化済音声部分をフレームで満たすことと、 を含む、 請求項1に記載の方法。
- 9前記音声の符号化済フレームをバッファリングすることを更に含み、 前記複数のコンテンツ・ファイルを前記生成することが、 前記固定時間幅を映像フレームのサンプリング・レートに掛け、これに、前記複数のコンテンツ・ファイルのうちの直前のコンテンツ・ファイルからのサンプル・オフセットがあればこれを足すことにより、前記複数のコンテンツ・ファイルのうちの前記1つまたは複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの前記符号化済音声部分を満たすのに必要なサンプルの数を算出することと、 前記サンプルの数を、前記コーデック適用フレーム・サイズで割ることにより、前記現在のコンテンツ・ファイルの前記符号化済音声部分を満たすのに必要なフレームの数を算出することと、 前記除算の余りが0の場合、前記現在のコンテンツ・ファイルの前記符号化済音声部分を前記フレーム数のフレームで満たすことと、 前記除算の余りが0より大きい場合、前記フレーム数を1だけインクリメントし、前記インクリメントしたフレーム数のフレームで前記現在のコンテンツ・ファイルの前記符号化済音声部分を満たすことと を含む、 請求項 4 に記載の方法。
- 10前記複数のコンテンツ・ファイルを前記生成することが、 前記現在のコンテンツ・ファイルを満たすのに必要な前記サンプルの数に戻すために、前記フレームの数と、前記コーデック適用フレーム・サイズとを掛けることと、 前記サンプルの数を前記サンプリング・レートで割ることにより、前記現在のコンテンツ・ファイルの前記音声のうちの前記符号化済部分の前記期間を算出することと、 前記固定時間幅から前記期間を引くことにより、前記複数のコンテンツ・ファイルのうちの次のコンテンツ・ファイルについての、前記音声ファイルの再生開始位置を示すプレゼンテーション・オフセットを求めることと、 前記フレームの数に前記コーデック適用フレーム・サイズを掛け、ここから、前記複数のコンテンツ・ファイルのうちの前記第1のコンテンツ・ファイルを満たすのに必要な前記サンプルの数を引き、前記複数のコンテンツ・ファイルのうちの前記前のコンテンツ・ファイルからのサンプル・オフセットがあればこれを足すことにより、前記複数のコンテンツ・ファイルのうちの前記次のコンテンツ・ファイルについての前記サンプル・オフセットを、現在処理中のコンテンツ・ファイルのサンプル・オフセットとしてアップデートすることと を更に含む、 請求項9に記載の方法。
- 11前記受け取ることが、前記メディア・コンテンツを音声と映像の両方を含む複数の生の離散的なデータのストリーム(ストリームレット)として受け取ることを含み、 前記複数の生のストリームレットのそれぞれが、前記メディア・コンテンツのうちで前記固定時間幅を有する映像フレームを含む、 請求項1に記載の方法。
- 12前記メディア・コンテンツを前記受け取ることが、 前記複数の生のストリームレットのうちの第1の生のストリームレットおよび前記複数の生のストリームレットのうちの第2の生のストリームレットを受け取ることと、 前記第1の生のストリームレットの前記音声および前記映像をスプリットし、前記第2の生のストリームレットの前記音声および前記映像をスプリットすることと、 を含み、 前記映像を前記符号化することが、 前記第1の生のストリームレットの前記映像を符号化することであって、前記第1の生のストリームレットの前記映像が、前記複数のコンテンツ・ファイルのうちの第1のコンテンツ・ファイル内に格納されることと、 前記第2の生のストリームレットの前記映像を符号化することであって、前記第2の生のストリームレットの前記映像が、前記複数のコンテンツ・ファイルのうちの第2のコンテンツ・ファイル内に格納されることと、 を含み、 前記音声を前記符号化することが、 前記第1の生のストリームレットの前記音声を、第1の複数の音声フレームに符号化することと、 前記第1の複数の音声フレームをバッファリングすることと、 前記第1のコンテンツ・ファイルを満たすのに十分なフレームがバッファリングされたかどうかを判定することと、 前記第1のコンテンツ・ファイルを満たすのに十分なフレームがバッファリングされていない場合、前記第2の生のストリームレットの前記音声を、第2の複数の音声フレームに符号化し、前記第2の複数の音声フレームをバッファリングすることと、 前記第1のコンテンツ・ファイルを満たすのに十分なフレームがバッファリングされている場合、前記バッファリングされた音声フレームを前記第1のコンテンツ・ファイル内に格納することと、 を含む、 請求項11に記載の方法。
- 13前記固定時間幅が約2秒間であり、 前記音声が、1秒間に約48000サンプルでサンプリングされ、 前記コーデック適用フレーム・サイズが、1フレームにつき1024サンプルであり、 前記複数のコンテンツ・ファイルのうちの初めの3つのコンテンツ・ファイルの音声部分がそれぞれ、94の音声フレームを含み、 前記複数のコンテンツ・ファイルのうちの第4のコンテンツ・ファイルの音声部分が、93の音声フレームを含み、 前記第4のコンテンツ・ファイルの映像部分がそれぞれ、約60の映像フレームを含む、 請求項1に記載の方法。
- 14前記固定時間幅が約2秒間であり、 前記音声が、1秒間に約44100サンプルでサンプリングされ、 前記コーデック適用フレーム・サイズが、1フレームにつき1024サンプルであり、 前記複数のコンテンツ・ファイルのうちの第1のコンテンツ・ファイルの前記符号化済音声部分が、87の音声フレームを含み、前記複数のコンテンツ・ファイルのうちの第2のコンテンツ・ファイルの前記符号化済音声部分が、86の音声フレームを含む、 請求項1に記載の方法。
- 15前記コーデック適用フレーム・サイズが、1フレームにつき2048サンプルである、 請求項1に記載の方法。
- 16映像および音声を含むメディア・コンテンツを受け取る手段と、 前記映像をフレーム・レートに従って符号化する手段と、 前記音声を固定フレーム・サイズに従って符号化する手段と、 前記符号化済映像を複数の符号化済映像部分にセグメント化する手段であって、前記符号化済映像部分のそれぞれが、別々のコンテンツ・ファイル内に格納される手段と、 前記符号化済音声を複数の符号化済音声部分にスプリットする手段と、 を備えるコンピューティング・システムであって、 各符号化済音声部分が、複数の隙間のない音声フレームに含まれ、それぞれ個別のコンテンツ・ファイルに境界アーチファクトを導入することなく格納され、前記別々のコンテンツ・ファイルのうちの第1のコンテンツ・ファイルの前記符号化された音声が、前記第1のコンテンツ・ファイル内に格納された前記符号化済映像部分の期間よりも長い、または、これより短い期間を有 し 、 前記複数のコンテンツ・ファイルの現在のコンテンツ・ファイルの符号化済音声部分の時間長が互いに整数倍でない場合も、長さ調整のために0値の埋め込みを前記現在のコンテンツ・ファイルの前記符号化済音声部分に行わず、前記現在のコンテンツ・ファイルの前記符号化済音声部分を隙間のない音声フレームを用いて満たし、 次のコンテンツ・ファイルの前記音声に、固定時間幅から前記現在のコンテンツ・ファイルの音声のうちの符号化済部分の期間を引くことにより得られる、前記次のコンテンツ・ファイルの前記符号化済音声部分の再生開始位置を示すプレゼンテーション・オフセットを与えて再生し、前記音声が再生されず、前記現在のコンテンツ・ファイルと前記次のコンテンツ・ファイルのつなぎ目で、アーチファクトが発生しないようにする、 コンピューティング・システム。
- 17前記コンテンツ・ファイルのそれぞれについて、サンプル・オフセットがあればこれを記録する手段と、 前記コンテンツ・ファイルのそれぞれについて、プレゼンテーション・オフセットがあればこれを記録する手段と、 を更に備える請求項16に記載のコンピューティング・システム。
- 18音声および映像を含むメディア・コンテンツを受け取り、前記音声および前記映像をスプリットするためのスプリッタと、 前記スプリッタから前記映像を受け取るように結合され、前記映像をフレーム・レートに従って符号化するための映像エンコーダと、 前記スプリッタから前記音声を受け取るように結合され、前記音声を、コーデック適用フレーム・サイズに従って符号化するための音声エンコーダと、 複数のコンテンツ・ファイルを生成するための音声スプリッティング・マルチプレクサであって、前記複数のコンテンツ・ファイルのそれぞれが、前記映像のうちで固定時間幅を有する符号化済部分と、前記音声のうちで、前記コーデック適用フレーム・サイズを有する複数の隙間のない音声フレームを有する符号化済部分とを含み、前記複数のコンテンツ・ファイルのうちの1つまたは複数のコンテンツ・ファイルの前記音声の前記符号化済部分の期間が、前記固定時間幅よりも長い、または、これよりも短く、1以上の前記複数のコンテンツ・ファイルの音声の符号 化 された部分の隙間のない複数の音声フレームのうちの最後のものは、ゼロで埋められることがな く、 前記複数のコンテンツ・ファイルの現在のコンテンツ・ファイルの符号化済音声部分の時間長が互いに整数倍でない場合も、長さ調整のために0値の埋め込みを前記現在のコンテンツ・ファイルの前記符号化済音声部分に行わず、前記現在のコンテンツ・ファイルの前記符号化済音声部分を隙間のない音声フレームを用いて満たし、 次のコンテンツ・ファイルの前記音声に、固定時間幅から前記現在のコンテンツ・ファイルの音声のうちの符号化済部分の期間を引くことにより得られる、前記次のコンテンツ・ファイルの前記符号化済音声部分の再生開始位置を示すプレゼンテーション・オフセットを与えて再生し、前記音声が再生されず、前記現在のコンテンツ・ファイルと前記次のコンテンツ・ファイルのつなぎ目で、アーチファクトが発生しないようにする、 音声スプリッティング・マルチプレクサと、 を備えるコンピューティング・デバイス。
- 19前記コンピューティング・デバイスが、前記音声の符号化済フレームをバッファリングするための音声フレーム・バッファを更に備える、 請求項18に記載のコンピューティング・デバイス。
- 20コンピューティング・デバイスに実行されると前記コンピューティング・デバイスにある方法を行わせる命令を格納する非一時的コンピュータ可読記憶媒体であって、 前記方法が、 音声および映像を含むメディア・コンテンツを受け取ることと、 前記映像をフレーム・レートに従って符号化することと、 前記音声を、コーデック適用フレーム・サイズに従って符号化することと、 複数のコンテンツ・ファイルを生成することであって、前記複数のコンテンツ・ファイルのそれぞれが、前記映像のうちで固定時間幅を有する符号化済部分と、前記音声のうちで、前記コーデック適用フレーム・サイズを有する複数の隙間のない音声フレームを有する符号化済部分とを含み、前記複数のコンテンツ・ファイルのうちの1つまたは複数のコンテンツ・ファイルの前記音声の前記符号化済部分の期間が、前記固定時間幅よりも長い、または、これよりも短く、1以上の前記複数のコンテンツ・ファイルの音声の符号かされた部分の隙間のない複数の音声フレームのうちの最後のものは、ゼロで埋められることがな く、 前記複数のコンテンツ・ファイルの現在のコンテンツ・ファイルの符号化済音声部分の時間長が互いに整数倍でない場合も、長さ調整のために0値の埋め込みを前記現在のコンテンツ・ファイルの前記符号化済音声部分に行わず、前記現在のコンテンツ・ファイルの前記符号化済音声部分を隙間のない音声フレームを用いて満たし、 次のコンテンツ・ファイルの前記音声に、固定時間幅から前記現在のコンテンツ・ファイルの音声のうちの符号化済部分の期間を引くことにより得られる、前記次のコンテンツ・ファイルの前記符号化済音声部分の再生開始位置を示すプレゼンテーション・オフセットを与えて再生し、前記音声が再生されず、前記現在のコンテンツ・ファイルと前記次のコンテンツ・ファイルのつなぎ目で、アーチファクトが発生しないようにする、 ことと、 を含む、 コンピュータ可読記憶媒体。
- 21前記方法が、 前記音声の符号化済フレームをバッファリングすることと、 前記複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの音声の符号化された部分を満たすのに必要な符号化済フレーム数を求めることであって、前記フレーム数が、前記複数のファイルのうちの前記現在のファイルの音声の符号化された部分を満たすのに必要なサンプル数を、前記コーデック適用フレーム・サイズで割った数以上の最小の整数であることと、 前記複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの音声の符号化された部分を満たすだけの前記符号化済フレームがバッファリングされたかどうかを判定することと、 前記複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの音声の符号化された部分を満たすだけの前記符号化済フレームがバッファリングされている場合、前記複数のコンテンツ・ファイルのうちの前記現在のコンテンツ・ファイルの音声の前記符号化された部分を前記フレーム数のフレームで満たすことと、 前記複数のコンテンツ・ファイルのうちの現在のコンテンツ・ファイルの音声の符号化された部分を満たすだけの前記符号化済フレームがバッファリングされていない場合、前記音声のフレームをバッファリングし、前記複数のコンテンツ・ファイルのうちの前記現在のコンテンツ・ファイルの音声の前記符号化された部分を、前記バッファされたフレームで満たすことと、 を更に含む、 請求項20に記載のコンピュータ可読記憶媒体。
Independent claims21
83 paragraphs, as filed
0001Embodiments of the present invention relate to the field of distribution of media content over the Internet, more specifically, splitting the audio of media content into a plurality of separated content files without introducing boundary artifacts. Regarding what to do.
0002The Internet has become the primary method of distributing media content (eg, video and audio or audio) and other information to end users. It is now possible to download music, video, games and other media information to computers, mobile phones and virtually any networkable device. The percentage of people accessing the Internet for media content is growing rapidly. The quality of the viewer experience is a key barrier to the growth of online video viewing. The consumer's TV and movie viewing experience brings consumer expectations for online video.
0003The number of viewers streaming video on the web is growing rapidly, and the interest and demand for watching video on the Internet is increasing. Streaming data files or "streaming media" refers to the technology of delivering continuous media content at a rate sufficient to provide the media to the user at the originally expected playback speed without major interruptions. Unlike downloaded data in a media file, streamed data can be stored in memory until it is played, and then deleted after a specified amount of time.
0004Streaming media content over the Internet presents several challenges when compared to regular broadcasts over radio, satellite, or cable. One of the concerns that arises in the context of audio coding of media content is the introduction of boundary artifacts when segmenting video and audio into multiple fixed-time parts. In one of the conventional methods, the audio is segmented into a plurality of parts having a fixed time width that matches the fixed time width of the corresponding video, for example, 2 seconds. In this technique, the audio boundaries are always aligned with the video boundaries. In this conventional method, for example, by using AAC LC (Low Complexity Advanced Audio Coding), a new coding session of a voice codec is started in order to encode each voice part for each content file. By using a new coding session for each part of the voice, the voice codec interprets the start and end of the waveform as a transition from 0, thereby coding at the partial boundaries as shown in Figure 1. Pop noise during playback of part (pop Causes noise or click noise. Such pop noise or click noise is called a boundary artifact. The audio codec also encodes audio with a fixed time width according to the codec-enforced frame size. This also introduces boundary artifacts if the number of samples produced by the voice codec is not equally divisible by the codec applied frame size.
0005FIG. 1 is a diagram showing an exemplary speech waveform 100 for two parts of speech using a conventional technique. The audio waveform 100 shows the transition 102 from 0 between the first and second parts of the video. When the voice codec has a fixed frame size (hereinafter referred to as a codec application frame size in the present specification), in the voice codec, the number of samples of the portion is the number of samples per frame according to the codec application frame size. If not equally divisible, the final frame 104 must be filled with 0s. For example, using a sampling rate of 48kHz, 96000 samples are generated for a 2 second audio segment. This number of samples 96000 is the number of samples per frame (for example, in the case of AAC LC, the number of samples is 1024, HE AAC (High Efficiency) In the case of AAC), the number of samples is 93.75 when divided by 2048). Since 93.75 is not an integer, the voice codec fills the final frame 104 with 0s. In this example, the last 256 samples in the final frame are given the value 0. A value of 0 represents silent audio, but since the last frame is padded with 0, pop or click audio is seen at the partial boundary during playback of the encoded portion of the audio. The transition 102 from 0 and the 0 embedded in the final frame 104 introduce the boundary artifact. When boundary artifacts are introduced, the overall quality of the audio may be reduced, which affects the user experience during playback of media content.
0006Other conventional methods try to limit the number of boundary artifacts by using a longer period of speech to align with the frame boundaries. However, as the duration portion of the audio used becomes larger, it may be necessary to combine the audio and video separately. This can impede the streaming of media content with audio and video, and such impediments allow, for example, a shift to different quality levels during playback of media content. This is especially true when the same media content is encoded at different quality levels, as in the scene.
0007The present invention can be optimally understood by reference to the following description and the accompanying drawings used to illustrate embodiments of the present invention.<figref num="1">It is a figure which shows the exemplary speech waveform about two parts of speech using a certain conventional method.</figref><figref num="2">It is a schematic block diagram which shows one Embodiment of the computing environment which can use the encoder of this Embodiment.</figref><figref num="3A">FIG. 3 is a schematic block diagram showing another embodiment of a computing environment in which an encoding system including a plurality of hosts each using the encoder of FIG. 2 can be used.</figref><figref num="3B">It is a schematic block diagram which shows one Embodiment of the parallel coding of streamlet by one Embodiment.</figref><figref num="4">FIG. 5 is a flow chart of an embodiment of a method of encoding audio of media content according to a codec application frame size and splitting audio frames without gaps between content files having a fixed-time video portion of the media content.</figref><figref num="5A">FIG. 5 is a flow chart of an embodiment of generating a content file with a fixed-time video portion and a gapless audio frame having a codec application frame size.</figref><figref num="5B">FIG. 5 is a flow chart of an embodiment of generating a content file with a fixed-time video portion and a gapless audio frame having a codec application frame size.</figref><figref num="5C">FIG. 5 is a flow chart of an embodiment of generating a content file with a fixed-time video portion and a gapless audio frame having a codec application frame size.</figref><figref num="6A">It is a figure of an audio part, a video part, and a streamlet by one Embodiment of audio splitting.</figref><figref num="6B">It is a figure which shows one Embodiment of the voice waveform about four parts of voice using voice splitting.</figref><figref num="7">FIG. 5 is a diagram of a machine in an exemplary embodiment of a computer system for voice splitting according to an embodiment.</figref>
0008A method and device for splitting the audio of media content into separate content files without introducing boundary artifacts will be described. In one embodiment, the method performed by a computing system programmed to perform an operation is to receive media content, including audio and video, to encode the video according to a frame rate, and a codec. Coding the audio according to the applied frame size (ie, a fixed frame size) and the gap between the encoded portion of the video with a fixed time width and the audio with the codec applicable frame size. Includes generating a content file, each containing an encoded portion with no audio frames. In one embodiment, the end of the audio frame is not padded with zeros as is traditionally done.
0009Embodiments of the present invention provide improved techniques for streaming audio. Unlike traditional methods that use a new coded session for each audio portion of the media content, the embodiments described herein combine the media content into multiple subsections without introducing boundary artifacts. Allows segmentation. In the embodiments described herein, audio is segmented by using audio frames without gaps. When the audio is staged for playback, the audio is provided to the decoder as a single stream rather than a large number of small segments with boundary artifacts. In the embodiments described herein, the encoder has a codec frame size (eg, AAC-LC with 1024 samples, or HE). In the case of AAC, the number of samples is 2048), and it recognizes how many audio frames are created each time the codec is activated. The encoder stores as many audio frames as it can fit in an encoded streamlet (ie, a content file), which is the portion of the video based on a fixed time width. Has. Instead of padding the final audio frame with zeros, it encodes a tight frame for the next part of the audio and adds it to the current streamlet. This writes a small amount of audio to the current streamlet that would normally be written to the next streamlet. These next streamlets are then given a time offset for the audio stream to indicate a gap so that the audio can be provided to the decoder as a continuous stream when playing the audio. The same amount of time is deducted from the audio target period for this streamlet. If the end of the audio in the next streamlet above does not fall on the frame boundary, borrow the audio from the next streamlet again to fill the final frame. This process is repeated until the end of the stream of media content is reached. The gap inserted at the beginning of the streamlet where the audio is borrowed can be removed when the audio portion of the streamlet is staged before being decoded and played. When trying to get a random streamlet, silent audio can be played during the gap period to maintain audio / video synchronization.
0010According to the audio splitting embodiments described herein, a voice codec with a large codec applicable frame size (eg, AAC, AC3, etc.) is used to create boundary artifacts while maintaining the same fixed time width for the video. It will be possible to encode the audio of media content without introducing it.
0011In the following description, many details are mentioned. However, it will be apparent to those skilled in the art who will benefit from the disclosure of the present invention that the embodiments of the present invention can be practiced without such specific details. In some examples, well-known structures and devices are shown in block diagram format rather than in detail to avoid obscuring embodiments of the present invention.
0012Some parts of the detailed description below are provided in terms of symbolic representations and algorithms of operations on data bits in computer memory. Such algorithmic descriptions and representations are the means used by those skilled in the art of data processing techniques to most effectively convey the essence of their research to others. Here again, an algorithm is generally understood to be a coherent sequence of steps that yields the desired result. These steps require physical manipulation of physical quantities. Usually, but not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, or manipulated. Mainly because of its general usage, it may be convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc. I know.
0013However, keep in mind that all of these terms and similar terms should be associated with appropriate physical quantities and are only convenient labels that apply to these quantities. Unless otherwise stated, as is clear from the discussion below, "receiving," "encoding," "generating," and "split" throughout the description. "Splitting", "processing", "computing", "calculating", "determining", "displaying" Discussions using terms such as "manipulate data represented by physical (eg, electronic) quantities in computer system registers and memory, computer system memory or registers or other such information storage. It is understood to refer to the operation and processing of a computer system or similar electronic computing system that transforms into other data, also represented using physical quantities, within a device, information transmitting device or information displaying device. To.
0014Embodiments of the present invention also relate to devices for performing the operations herein. The device may be specifically constructed for the required purpose or may include a general purpose computer system specifically programmed into a computer program stored therein. These computer programs can be any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks, ROM (read-only memory), RAM (random access memory), EPROM, It can be stored in a computer-readable storage medium such as an EEPROM, a magnetic card or an optical card, or any kind of media suitable for storing electronic instructions, but the storage destination is not limited thereto.
0015As used herein, the term "encoded streamlet" refers to a single coded representation of a portion of media content. Each streamlet can be a separate content file that contains a portion of the media and can also be enclosed as a separate media object, which allows the streamlet to be cached individually, the media. It is possible for the player to request it independently and to play it independently. In the present specification, these individual files are also referred to as OSS files. In one embodiment, a streamlet is a static file that may be served by a non-specialized server rather than by a specialized media server. In one embodiment, the media content in the streamlet may have a play time (also referred to as a fixed time width) of a predetermined length. This predetermined length of time is, for example, about 0.1-8. It can be 0 seconds. Alternatively, other predetermined lengths can be used. The media content in the streamlet may have a unique time index with respect to the beginning of the media content contained in the stream. The file name may contain part of the time index. Alternatively, the streamlet can be split according to the size of the file instead of the time index. As used herein, the term "stream" may refer to a set of streamlets of media content encoded with the same video quality profile, eg, the same video bit rate of video. It is the part encoded by. A stream represents a copy of the original media content. Streamlets are separate files that are one of content servers, web servers, cache servers, proxy caches, or other devices on the network, such as those found on content delivery networks (CDNs). Or it can be stored on multiple. These separate files (eg, streamlets) can be requested by the client device from the web server using HTTP. By using a standard protocol such as HTTP, the network administrator can use RTSP (Real Time Streaming) for the firewall. It eliminates the need to configure network traffic to recognize and pass network traffic for specialized new protocols such as Protocol). Moreover, since the media player initiates the request, for example, the web server is only required to retrieve and supply the requested streamlet, not the entire stream. The media player can also retrieve streamlets from more than one web server. These web servers may not be accompanied by specialized server-side intelligence to retrieve the requested portion. In other embodiments, the streamlet is stored as a separate file on the cache server of the network infrastructure operator (eg, ISP) or other component of the CDN. Although some of the embodiments describe the use of streamlets, the embodiments described herein are not limited to use within computing systems that use streamlets, but live over the Internet. It can also be run within other systems that use other techniques for delivering media content. For example, in other embodiments, the media content is stored in a single, multi-part file that can be requested by using an HTTP range request and can be cached within the CDN.
0016There are two common types of media streaming, that is, push-based streaming and pull-based streaming. Push technology describes a method of internet-based communication in which a server, such as the issuer's content server, initiates a request for a given transaction. In contrast, pull technology describes a method of Internet-based communication in which a request to send information is initiated by a client device and then replied by a server. HTTP requests (eg HTTP GET requests) are a type of request in pull technology. In contrast, in push-based technology, a dedicated server typically uses a dedicated protocol, such as RTSP, to send data to client devices. Alternatively, some push-based technologies may use HTTP to deliver media content. Pull-based technology may use a CDN to deliver media to multiple client devices.
0017Although the various embodiments described herein are intended for pull-based models, it should be noted that these embodiments may be implemented within other configurations, such as push-based configurations. I want to. In the push-based configuration, an embodiment of voice splitting by the encoder can be executed in the same manner as the pull-based configuration described in connection with FIG. 2, and the encoded content file can be executed by a media server or the like. This media content can be delivered to client devices for playback by storing it on a content server and using push-based technology. These embodiments can be used to bring about different quality levels of media content, and these embodiments switch between different quality levels, commonly referred to as adaptive streaming. Also note that this is possible. In the push-based model, the media server determines which content files to send to the client device, and in the pull-based model, the client device content which (one or more) content files. -Determining whether to request from the server can be one of the differences.
0018FIG. 2 is a schematic block diagram showing an embodiment of the computing environment 200 in which the encoder 220 of the present embodiment can be used. The computing environment 200 includes a source 205, an encoder 220, an origin content server 210 (also referred to as a media server or origin server) of the content distribution network 240, and a plurality of media running on the client device 204, respectively. -Including player 200. The content server 210, the encoder 220, and the client device 204 can be coupled by a data communication network. Data communication networks may include the Internet. Alternatively, the content server 210, encoder 220, client device 204 can be connected to a common LAN (local area network), PAN (personal area network), CAN (campus area network), MAN (metropolitan area). -It can also be placed on a network), WAN (wide area network), wireless LAN, cellular network, virtual LAN, etc. The client device 204 is an entertainment configured to communicate over a client workstation, server, computer, portable electronic device, or network such as a set-top box, digital receiver, or digital television. -Can be a system or other electronic device. For example, portable electronic devices may include, but are not limited to, mobile phones, portable game systems, portable computing devices, and the like. Client device 204 can access the Internet through a firewall, router, or other packet-switched device.
0019In the illustrated embodiment, the source 205 may be a publisher server or a publisher content repository. Source 205 may also be the creator or distributor of media content. For example, if the media content to be streamed is a broadcast of a television program, the source 205 is a television server or a cable network channel such as an ABC® or MTV® channel. Sometimes. The publisher can transfer media content to the encoder 220 over the Internet, which receives and processes the media content and the content file (s) of this media content. Can be configured to be stored within the Origin Content Server 210. In one embodiment, the content server 210 delivers the media content to the client device 204, which client device 204 is configured to play the content on a media player running on it. To. The content server 210 delivers the media content to the client device 204 by streaming it. As will be described in more detail later, in a further embodiment, the client / server 204 is configured to receive various parts of the media content from multiple locations simultaneously or simultaneously.
0020The media content stored on the content server 210 can be replicated to other web servers or to the CDN240 proxy cache server. Replication occurs either by being systematically pushed from content server 210 or by a web server, cache server, or proxy server outside content server 210 requesting content for client device 204. I have something to do. For example, client device 204 can request and receive content from one of a plurality of web servers, edge caches, or proxy cache servers. In the illustrated embodiment, the web server, proxy cache, edge cache, and content server 210 are organized in a hierarchy of CDN240 to deliver media content to the client device 204. A CDN is a system of multiple computers networked over the Internet that work transparently to deliver content, such as one or more origin content servers, web servers, and caches. May include servers, edge servers, etc. A CDN is typically configured in a hierarchy where, for example, a client device requests data from the edge cache, and if the edge cache does not contain the requested data, then the request is then the parent cache. It will be sent to the origin content server as well. A CDN may also include an interconnected computer network or node for delivering media content. Some examples of CDNs include Akamai Technologies, Level 3 Communications, or Limelight. There is a CDN developed by Networks. Alternatively, other types of CDNs can be used. In another embodiment, the origin content server 210 delivers media content to client device 204 by using other configurations that would be understood by those skilled in the art who would benefit from the disclosure. be able to.
0021In one embodiment, the publisher stores the media content within the original content file to be distributed from the source 205. Content files may include video and / or audio equivalent data such as television broadcasts, sporting events, movies, music, concerts, and the like. The original content file may contain uncompressed video and audio, or uncompressed video or audio. Alternatively, the content file may contain compressed content (eg, video and / or audio) using a standard encoding scheme or a proprietary encoding scheme. The original content file from Source 205 may be in digital format and may contain high bit rate media content, such as about 5 Mbps or higher.
0022In the illustrated embodiment, the encoder 220 receives the original content file, the signal from the direct supply of the live event broadcast, the stream of the live television event broadcast, etc., so that the original media from the source 205. -Receive content 231. Encoder 220 can be implemented on one or more machines including one or more server computers, gateways or other computing devices. In one embodiment, the encoder 220 receives the original media content 231 as one or more content files from the publishing system (not shown) (eg, the publisher's server or the publisher's content repository). .. Alternatively, the encoder 220 receives the original media content 231 at the time of capture. For example, the encoder 220 is captured. Direct supply of live television broadcasts such as broadcast) can be received in stream or signal format. Original media content 231 should be captured with a capture card configured for TV capture and / or video capture, such as the DRC-2600 capture card available from Digital Rapids, Ontario, Canada. Can be done. Alternatively, any capture card capable of capturing audio and video can be used in the present invention. The capture card can be on the same server as the encoder or on a different server. The original media content 231 can be captured at a point in time, such as a broadcast, such as a simultaneous broadcast over radio waves, cables, and / or satellites, or according to a live event schedule. It may be a scheduled pre-recorded broadcast. Encoder 220 is a DivX® codec, Windows Media Video9 (registered trademark) series codec, Sorenson Video (registered trademark) 3 video codec, On2 Technologies (registered trademark) TrueMotion VP7 codec, MPEG-4 video codec, H.263 video codec, RealVideo10 codec, OGG Vorbis, MP3 Etc. can be used. Alternatively, a customized encoding scheme can be used.
0023In another embodiment, the encoder 220 is the original video content 231 as a fixed time width portion of the video and audio (referred to herein as "a portion of the media content"), such as a 2-second chunk. To receive. This 2-second chunk is raw audio and raw may include video). Alternatively, this 2-second chunk may be encoded audio and live video. In such a case, the encoder 220 expands the media content. In another embodiment, the encoder 220 receives the original media content 221 as multiple raw streamlets, each of which includes a fixed time portion of the media content (eg, live audio and video). Includes multiple 2-second raw streamlets). As used herein, the term "raw streamlet" refers to an uncompressed streamlet or a lightly compressed streamlet to substantially reduce its size without significant loss of quality. Point to. Lightly compressed raw streamlets can be transmitted more rapidly. In another embodiment, the encoder 220 receives the original media content 231 as a stream or signal and segments it into multiple fixed time portions of the media content, such as a raw streamlet.
0024In the illustrated embodiment, the encoder 220 includes a splitter 222, a fixed frame audio encoder 224, an audio frame buffer 225, a fixed time video encoder 226, a video frame buffer 227, and an audio splitting multiplexer 228. The splitter 222 receives the original media content 231 as, for example, a continuous stream of audio and video, which is split into live audio 233 and live video 235. In one embodiment, the fixed frame audio encoder 224 is an audio codec. In one embodiment, the splitter 222 splits a continuous stream of audio and video into multiple 2-second chunks of audio and video. Codec (compressor-decompressor or coder-decoderAlso referred to as) is a device or computer program capable of encoding and / or decoding a digital data stream or digital data signal. In one embodiment, the fixed-frame voice codec 224 is software that is executed on one or more computing devices of the encoder 220 to encode the live voice 233. Alternatively, the fixed-frame voice codec 224 may be a hardware logic element used to encode the live voice 233. Specifically, the fixed-frame audio encoder 224 receives live audio 233 and encodes this audio according to the codec application frame size, for example, 1024 samples for AAC-LC or 2048 samples for HE AAC. The fixed-frame audio encoder 224 outputs the encoded audio frame 237 to the audio frame buffer 225. Similarly, the fixed-time video encoder 226 receives the live video 235 from the splitter 220, but encodes this video according to a fixed time width, for example 60 frames (30 fps (frames per second)) every 2 seconds. The fixed-time video encoder 226 outputs the encoded video frame 239 to the video frame buffer 227. In one embodiment, the fixed time video codec 226 is software that is executed by one or more computing devices of the encoder 220 to encode the live video 235. Alternatively, the fixed time video codec 226 may be a hardware logic element used to encode the live video 235.
0025The audio splitting multiplexer 228 uses the encoded audio frame 237 and the encoded video frame 239 to generate an encoded media content file 232 (referred to herein as a "QSS file"). As explained above, traditional encoders generate content files with a portion of video and a portion of audio, each with a fixed time width, which is used in audio codecs. Since the number of samples in that part cannot be evenly divided by the number of samples for each frame according to the codec application frame size, the final frame of the audio is filled with 0. Unlike traditional encoders that fill the final frame, the audio splitting multiplexer 228 uses tight audio frames to provide a fixed-time video portion and an audio portion with a tight audio frame with a codec-applied frame size. Generate a content file with and. The audio splitting multiplexer 228 uses tight audio frames to fill the content file 232, so instead of filling the last few samples of the frame with zeros as is traditionally done, the current content Encode the next part of the audio to add a tight frame to file 232.
0026In one embodiment, the audio splitting multiplexer 228 records a sample offset that represents the amount of sample used from the next part in order to determine the number of frames used for the next content file. The audio splitting multiplexer 228 also records a presentation offset that indicates a gap in audio reproduction. Since the sample that would normally be played as part of the next content file becomes part of the current content file, the presentation offset of the next content file indicates a gap in audio playback, thereby The audio portion of the current and next content files is provided to the decoder as a continuous stream. In essence, during audio playback, if the audio portion of the content file is staged before decoding and playback, the gap inserted at the beginning of the content file can be removed. The presentation offset allows the audio to be provided to the decoder as a continuous stream rather than many small segments with boundary artifacts. In one embodiment, when searching for any part of the video, silent audio can be played for the duration of the gap in order to maintain audio / video synchronization.
0027In one embodiment, the audio splitting multiplexer 228 has a first video portion (eg, 60 frames) with a fixed time width (eg, 2 seconds) and some buffered, tight audio frames. The first content file is generated by filling with the audio part of 1. The duration of this buffered audio frame is longer than the fixed time width.
0028In one embodiment, the audio splitting multiplexer 228 generates the content file 232 by determining the number of encoded audio frames 237 needed to fill the current content file. In one embodiment, the number of frames is the smallest integer greater than or equal to the number of samples required to fill the current content file divided by the codec-applied frame size (eg, samples per frame). .. In one embodiment, this number is calculated by using a ceiling function that maps a real number to the next largest integer, for example "ceiling (x) = [x] is the smallest integer greater than or equal to x". can do. An example of the ceiling function is shown in the following equation (1). ceil ((Sample per Streamlet-Offset Sample) / Sample Per Frame) (1) Alternatively, other expressions can be used.
0029The voice splitting multiplexer 228 determines if there are enough encoded voice frames 237 in the voice frame buffer 225 to fill the current content file. If enough encoded frames are buffered, the audio splitting multiplexer 228 fills the current content file with the desired number of frames. If enough coded frames are not buffered, the voice splitting multiplexer 228 waits for enough coded frames to be stored in buffer 225 and finds the number of codes stored in buffer 225. Fill the current content file with the converted frame. In one embodiment, the voice splitting multiplexer 228 1) multiplies the number of buffered frames by the sample per frame, and 2) multiplies the sample offset from the previous content file, if any. In addition to 3), it is determined whether sufficient encoded frames are buffered by determining whether this sum is greater than or equal to the number of samples required to fill the current content file. An example of this operation is shown in the following equation (2). Number of buffered frames * Samples per frame + Offset sample> = Samples per streamlet (2)
0030The audio splitting multiplexer 228 seeks sample offsets for the following content files, if any. In one embodiment, the voice splitting multiplexer 228 multiplies the codec applied frame size (ie, the sample per frame) by the number of encoded frames, from which the sample required to fill the current content file. The sample offset is calculated by subtracting the number and adding the sample offset from the previous content file, if any. An example of this operation is shown in the following equations (3) and (4). Offset sample = Frame to send * Sample per frame-Sample per streamlet-Offset sample (3) Frame to send = ceil ((Sample per streamlet-Offset sample) / Sample per frame) (4)
0031In another embodiment, the audio splitting multiplexer 228 generates the content file 221 by calculating the number of samples (eg, 96000) required to fill the current content file. The audio splitting multiplexer 228 calculates the number of frames required for the current content file (eg 93 frames at a 48K sampling rate per 2 seconds portion) and the number of samples is evenly divisible by the samples per frame. If not, add the number of frames (for example, 94 frames in total). This effectively rounds up the number of frames to the next largest integer. The voice splitting multiplexer 228 fills the current content with this rounded up number of frames.
0032In one embodiment, the audio splitting multiplexer 228 multiplies the sampling rate (eg, 48K) by a period of fixed time width (eg, 2 seconds) to fill the current content file with the sample required. Generate content file 221 by calculating a number (eg 96000). The audio splitting multiplexer 228 calculates the number of frames required for the current content file by dividing the number of samples by the codec-applied frame size (eg, 1024 samples per frame). If the remainder of this division is 0, the audio splitting multiplexer 228 fills the current content file with that number of frames. However, if the remainder of this division is greater than 0, the voice splitting multiplexer 228 increments the number of frames by 1 and fills the current content file with this incremented number of frames.
0033In a further embodiment, the voice splitting multiplexer 228 multiplies the codec-applied frame size by the number of frames back to the number of samples needed to fill the current content file, and divides the number of samples by the sampling rate. The content file 221 is generated by calculating the audio period of the current content file (eg, streamlet period = sample / sampling rate per streamlet). The audio splitting multiplexer 228 obtains the presentation offset for the next content file by subtracting this period from the fixed time width. The voice splitting multiplexer 228 multiplies the number of frames by the codec-applied frame size, from which the number of samples used to fill the current content file is subtracted, and the sample offset from the previous content file is Add this, if any (eg, equation (3)) to update the sample offset for the next content file.
0034Seeing FIG. 2 again, in one embodiment, when the splitter 222 receives the original media content 231 as a raw streamlet, the splitter 222 receives the first and second raw streamlets, the first and second. Split the audio and video of the 2 raw streamlets. The fixed-time video encoder 226 encodes the video of the first and second raw streamlets, and the audio splitting multiplexer 228 stores the encoded video of the first raw streamlet in the first content file. Then, the encoded video of the second raw streamlet is stored in the second content file. The fixed-frame audio encoder 224 encodes the audio of the first raw streamlet into a first set of audio frames, and stores this first set in the audio frame buffer 225. The audio splitting multiplexer 228 determines if there are enough buffered frames to fill the first content file. If not sufficient, the fixed frame audio encoder 224 encodes the audio of the second raw streamlet into a second set of audio frames and stores this second set in the audio frame buffer 225. If there are enough buffered frames to fill the first content file (and in some cases another tight frame is stored in buffer 225), the audio splitting multiplexer 228 will Store the buffered audio frame in the first content file. Encoder 220 continues this process until the media content is finished.
0035Also, since the audio splitting multiplexer 228 uses tight audio frames, the audio frames in one content file 232 do not necessarily have to align with the video partial boundaries, as shown in FIGS. 6A and 6B. For example, the duration of the audio part of content file 232 is 2.0053 seconds, and the fixed time width of the video part of content file 232 is 2. It can be 00 seconds. In this example, the codec applied frame size is 1024 samples per frame, the audio sampling rate is 48K, and 96256 samples of 94 frames are stored in the audio portion stored in the content file 232. .. Due to the extra 53ms (ms) present in content file 232, a sample with a duration of 53ms that should be in the next content file when using the fixed time width voice coding scheme is currently available. As used by the content file 232 of, the audio splitting multiplexer 228 gives the next content file a presentation offset of 53 ms. The audio splitting multiplexer 228 also records a sample offset to determine the number of audio frames required to fill the next content file. In one embodiment, the audio splitting multiplexer 228 is one of the encoded video portions having a fixed time width and fills each content file (eg, 60 video if the frame rate is 30 fps). 2 seconds per frame). The audio splitting multiplexer 228 fills some of the content files with some buffered audio frames, but the duration of these audio frames is the audio frame at the video partial boundary determined by the audio splitting multiplexer 228. May be greater than the fixed time width, less than the fixed time width, or equal to the fixed time width, depending on whether they match.
0036Referring to FIG. 6A, in one embodiment, the audio splitting multiplexer 228 has a first video portion 611 having approximately 60 video frames for a period equal to a fixed time width of 2 seconds, each with 1024 per frame. A first streamlet (ie, content file) 601 is generated by filling with a first audio portion 621 having samples and having 94 audio frames totaling 96256 samples. The duration of the first audio part 621 is about 2.0053 seconds. The audio splitting multiplexer 228 sets the presentation offset of the first audio portion 631 of the first streamlet 603 to 0 because the audio boundary 652 and the video boundary 654 of the first streamlet 601 are play-aligned. Is determined.
0037The audio splitting multiplexer 228 produces a second streamlet 602 by filling with a second video portion 612 (60 frames, 2 seconds) and a second audio portion 622 having 94 audio frames. The duration of the second audio part 622 is about 2.0053 seconds. Since the duration of the first audio portion 621 of the first streamlet 601 is approximately 2.0053 seconds, the audio splitting multiplexer 228 sets the presentation offset of the second audio portion 632 of the second streamlet 602 to approximately 5.3 ms. Judge as (ms). This presentation offset indicates the audio gap between the first streamlet 601 and the second streamlet 602. As shown in FIG. 6B, the audio boundary 652 and the video boundary 654 of the second streamlet 602 are not consistent for reproduction. This presentation offset can be used to allow the audio portion of the first and second streamlets 601 and 602 to be staged to be provided to the decoder as a continuous stream.
0038The audio splitting multiplexer 228 produces a third streamlet 603 by filling with a third video portion 613 (60 frames, 2 seconds) and a third audio portion 623 having 94 audio frames. The duration of the third audio part 623 is about 2.0053 seconds. Since the duration of the second audio portion 622 of the second streamlet 602 is approximately 2.0053 seconds, the audio splitting multiplexer 228 sets the presentation offset of the third audio portion 633 of the third streamlet 603 to approximately 10.66 ms. judge. This presentation offset indicates the audio gap between the second streamlet 602 and the third streamlet 603. As shown in FIG. 6B, the audio boundary 652 and the video boundary 654 of the third streamlet 603 are not consistent for reproduction. This presentation offset can be used to allow the audio portion of the second and third streamlets 602 and 603 to be staged to be provided to the decoder as a continuous stream.
0039The audio splitting multiplexer 228 produces a fourth streamlet 604 by filling a fourth video portion 614 (60 frames, 2 seconds) with a fourth audio portion 624 having 93 audio frames. The duration of the fourth audio part 624 is about 1.984 seconds. Since the duration of the third audio portion 623 of the third streamlet 603 is approximately 2.0053 seconds, the audio splitting multiplexer 228 determines that the presentation offset of the fourth audio portion 634 of the fourth streamlet 604 is approximately 16 ms. To do. This presentation offset indicates the audio gap between the third streamlet 603 and the fourth streamlet 604. As shown in FIG. 6B, the audio boundary 652 and the video boundary 654 of the fourth streamlet 603 are not consistent for reproduction. This presentation offset can be used to allow the audio portion of the third and fourth streamlets 603 and 604 to be staged to be provided to the decoder as a continuous stream. However, after the fourth streamlet 604, the audio boundary 652 and the video boundary 654 are aligned, which means that the fifth streamlet (not shown) has a presentation offset of zero. Note that in the embodiments of FIGS. 6A and 6B, the sampling rate is 48 kHz, the fixed time width is 2 seconds, and the codec applied frame size is 1024 samples per frame.
0040In the above embodiment, the audio portion of the three first streamlets 601 to 603 has 94 audio frames and the audio portion of the fourth streamlet 604 has 93 audio frames. In this embodiment, when the video is encoded at 30 fps, each of the video portions of the fourth content files 601 to 604 has about 60 video frames. This pattern is repeated until the end of the media content is reached. In this embodiment, after each fourth content file, the presentation offset and sample offset are zero, which means that after each fourth content file, the audio boundary 652 and the video boundary 654 are Note that this means consistency.
0041As can be seen from Figure 6B, after 8 seconds of media content, the video and audio boundaries align. Therefore, one of the other techniques to reduce the frequency of boundary artifacts and match the AAC frame size should be to use 8 seconds for a fixed time width. However, these methods have the following disadvantages. 1) This method requires images with large chunk sizes such as 8, 16 and 32 seconds. 2) This approach is constrained to a specific frame size, ie 1024 samples per frame. When changing the frame size to, for example, 2048, this technique requires switching to an audio codec with a different frame size and also changing the chunk period of the video. 3) This technique requires the audio sample rate to always be 48kHz. Other common sample rates, such as 44.1kHz, require different and possibly much larger chunk sizes. Alternatively, the source audio must be upsampled to 48kHz. However, up-sampling can introduce artifacts and reduce the efficiency of voice codecs. However, in the embodiments described herein, codecs are encoded by using a voice codec with a large frame size (AAC, AC3, etc.) while maintaining the same chunk duration, without introducing chunk boundary artifacts. Can be done.
0042Alternatively, other sampling rates (eg, 44.1kHz), fixed time widths (eg, 0.1-5.0 seconds), video frame rates (eg, 24fps, 30fps, etc.), and / or codec-applied frame sizes (eg, .. 2048) can also be used. Different source footage uses different frame rates. Most radio signals in the United States are at 30 fps (actually 29.97). Some HD signals are 60fps (59. 94). Some of the file-based content is 24fps. In one embodiment, the encoder 220 does not increase the frame rate of the video because it would need to generate additional frames. However, generating additional frames does not bring much benefit with this additional load. So, for example, if the original media content has a frame rate of 24 fps, the encoder 220 will use a frame rate of 24 fps instead of upsample to 30 fps. However, in some embodiments, the encoder 220 may downsample the frame rate. For example, if the original media content has a frame rate of 60 fps, the encoder 220 may downsample to 30 fps. Such downsampling can occur because using 60fps doubles the amount of data that needs to be encoded at the target bit rate, which can compromise quality. In one embodiment, the encoder 220 uses this frame rate for most of the quality profiles as it seeks the frame rate after downsampling or received (typically 30 fps or 24 fps). Lower frame rates may be used for some quality profiles, such as the lowest quality profile. However, in other embodiments, the encoder 220 has different frames for different quality profiles, for example to target other resource-constrained devices such as mobile phones and inferior computing power. You can use the rate. In these cases, it may be beneficial to have more profiles with lower frame rates.
0043Note that if other values are used for these parameters, the audio boundary 652 and the video boundary 654 may differ from the embodiments shown in FIG. 6B. For example, if you use 44.1kHz as the sampling rate, 1024 as the codec application frame size, and 2 seconds as the fixed time width, the audio part of the first content file will have 87 audio frames and the second. The 7th content file will have 86 audio frames. This pattern is repeated until there is insufficient video left in the media content. In this embodiment, after every 128 content files, the presentation offset and sample offset are 0, which means that the audio boundary 652 and the video boundary 654 are as shown in approximately Table 1-1. Note that this means matching after the 128th content file.<tables num="1"><img id="000002" he="146" wi="115" file="JP5728736B2_D0001.tif" img-format="tif" img-content="drawing" /></tables>Note that the sample offsets in the table above are shown in sample units, not seconds or milliseconds, for simplicity of explanation. When converting a sample offset to a presentation offset, you can divide the sample offset by 44100 to get the presentation offset in seconds and then multiply this by 1000 to get the presentation offset in milliseconds. it can. In one embodiment, the presentation offset in milliseconds may be stored in the streamlet header. Alternatively, the presentation offset or sample offset may be stored in the streamlet header in other units.
0044In another embodiment, the audio splitting multiplexer 228 generates a plurality of encoded content files 232 by filling each with a coded video frame 239 having a fixed time width (eg, a fixed time width portion). , These content files 232 are filled with some tight audio frames 237, but the duration of the audio frame 237 is fixed so as to accommodate the tight audio frames used in the content file 232. It is shorter or longer than the width. For example, the first content file can be filled with a portion of the video having a fixed time width such as 2 seconds and an audio portion having a plurality of audio frames without gaps having a period longer than the fixed time width. .. Over time, the sample offset increases and the number of audio frames that can be used decreases, in which case the duration of the audio frames may be shorter than the fixed time width. Occasionally, the audio boundaries of audio can match the video boundaries of video.
0045In another embodiment, the audio splitting multiplexer 228 has a first portion having a video frame of a first portion of video, an audio frame from a first portion of audio, and an audio frame from a second portion. Generates the encoded content file 232 by generating the content file. The audio splitting multiplexer 228 produces a second content file with a video frame for the second part of the video. In the case of audio, the audio splitting multiplexer 228 determines whether the audio boundary is on the video boundary. If the audio boundary is on the video boundary, the audio splitting multiplexer 228 fills the second content file with the remaining audio frames in the second part. However, if the audio boundary does not rest on the video boundary, the audio splitting multiplexer 228 encodes the audio frame of the third part of the media content, from the remaining audio frame of the second part and from the third part. Fill the second content file with an audio frame. This process is repeated until the end of the content file is reached.
0046Seeing FIG. 2 again, when the encoder 220 encodes the original media content 231, it sends the encoded media content file 232 to the origin content server 210, where the origin content server 210 is networked. The encoded media content 232 is delivered to the media player 200 via the connection 241. When the media player 200 receives a content file having a fixed time width of video and a variable time width of audio, the media player 200 uses the presentation offset of these content files to provide the audio to the decoder as a continuous stream. Stages to remove or reduce pop or click noise that results in boundary artifacts. In essence, during audio playback, the media player 200 removes the gap inserted at the beginning of the content file when the audio portion of the content file is staged prior to decoding and playback. In another embodiment, if the audio splitting described herein is not performed and the final frame is padded with 0s, the media player 200 removes the padded sample of the final frame before sending the audio to the decoder. Can be configured as follows. However, this technique may not be practical in certain situations, for example when the media player is provided by a third party or when access to the decrypted audio frame data is restricted. is there.
0047Although one line is shown for each media player 200, it should be noted that each line 241 may represent multiple network connections to the CDN 240. In one embodiment, each media player 200 may establish multiple TCP (Transmission Control Protocol) connections to the CDN240. In other embodiments, the media content is stored on a plurality of CDNs, such as the origin server associated with each of the plurality of CDNs. The CDN240 can be used to improve performance, scalability, and cost effectiveness for end users (eg, viewers) by reducing bandwidth costs and increasing the overall availability of content. .. CDNs can be implemented in a variety of ways, and details of their operations should be understood by those skilled in the art. Therefore, it does not include further details about those operations. In other embodiments, other delivery techniques, such as peer-to-peer networks, can be used to deliver media content from the origin server to the media player.
0048In the embodiment described above, the content file 232 represents one of the copies of the original media content stream 231. However, in other embodiments, each portion of the original media content 231 can be encoded into a plurality of coded representations of the same portion of the content. These multiple coded representations can be encoded according to various quality profiles, requested independently of the client device 204, and stored as separate files that can be played independently by the client device 204. it can. Each of these files can be stored within one or more content servers 210, on the CDN240's web server, proxy cache, and edge cache, requested separately, and delivered to client device 204. be able to. In one embodiment, the encoder 220 simultaneously encodes the original content media 231 at several different quality levels, for example 10 or 13. Each quality level is referred to as a quality profile or profile. For example, if the media content has a duration of 1 hour and is segmented into multiple QSS files with a duration of 2 seconds, there will be 1800 QSS files for each coded representation of the media content. If the media content is encoded according to 10 different quality profiles, there are 18000 QSS files for this media content. These quality profiles may indicate how to encode the stream, such as the width and height of the image (ie, the image size), the video bit rate (ie, the speed at which the video is encoded). , Audio bit rate, audio sample rate (ie, the speed at which audio is sampled when capturing), number of audio tracks (eg, mono, stereo, etc.), frame rate (eg, fps), staging size, etc. Pa The parameter may be specified. For example, a plurality of media players 200 may individually request the same media content 232 with different quality levels. For example, each media player 200 can request the same portion of media content 232 (eg, the same time index) at different quality levels. For example, one media player requires a streamlet with HD quality video because its computing device has sufficient computing power and sufficient network bandwidth, and another media player, for example, It may require lower quality streamlets because its computing device cannot have sufficient network bandwidth. In one embodiment, the Media Player 200 has different copies of the media content (eg, different qualities, as described in U.S. Patent Application Publication No. 2005/0262257, filing date April 28, 2005. Shift the quality level at the partial boundaries by requesting each part from the streamlet). Alternatively, the media player 200 may request those parts by using other techniques that should be understood by those skilled in the art who will benefit from the disclosure. It may require lower quality streamlets because the ing device cannot have sufficient network bandwidth. In one embodiment, the Media Player 200 has different copies of the media content (eg, different qualities, as described in U.S. Patent Application Publication No. 2005/0262257, filing date April 28, 2005. Shift the quality level at the partial boundaries by requesting each part from the streamlet). Alternatively, the media player 200 may request those parts by using other techniques that should be understood by those skilled in the art who will benefit from the disclosure. It may require lower quality streamlets because the ing device cannot have sufficient network bandwidth. In one embodiment, the Media Player 200 has different copies of the media content (eg, different qualities, as described in U.S. Patent Application Publication No. 2005/0262257, filing date April 28, 2005. Shift the quality level at the partial boundaries by requesting each part from the streamlet). Alternatively, the media player 200 may request those parts by using other techniques that should be understood by those skilled in the art who will benefit from the disclosure.
0049The encoder 220 can also identify which quality profile is available for a particular piece of media content, and how much media content is available for delivery, for example by using a QMX file. Can be identified. The QMX file shows the current duration of the media content represented by the available QSS files. The QMX file can act as a table of contents for the media content, showing which QSS files are available for distribution and where the QSS files can be retrieved. The QMX file can be sent to the media player 200 via, for example, the CDN240. Alternatively, the media player 200 may request a quality profile available for specific media content. In other embodiments, the scaling capabilities of the CDN can be used to scale this configuration to deliver HTTP traffic to multiple media players 200. For example, from a data center that stores encoded media content from the origin content server 210 to serve multiple media players requesting encoded media content from this data center. May have clusters of Alternatively, other configurations may be used that will be understood by those skilled in the art who will benefit from the disclosure.
0050In one pondered embodiment, the media player 200 requests each piece of media content by requesting individual streamlet files (eg, QSS files). Media player 200 requests a QSS file according to a metadata descriptor file (eg, OMX file). The media player 200 fetches a QMX file in response to the user selecting media content for serving, for example, and also reads the QMX file and uses the current period of media. Decide when to start playing the content and where to request the QSS file. The QMX file has a QMX timestamp, such as a UTC (Coordinated Universal Time) indicator that indicates when the encoding process will start (eg, the start time of the media content), and how much media content is available for distribution. Indicates the current period (current) duration) and is included. For example, a QMX timestamp can indicate that the encoding process starts at 6 pm (MDT) and that 4500 QSS files of media content are available for delivery. The media player 200 determines that the content period (live playback) is about 15 minutes, and starts requesting a QSS file corresponding to the playback of the program at a point 15 minutes after the program starts or slightly earlier than the program start. You can decide that. In one embodiment, the media player 200 may obtain a point in the media content at which the media player 200 should start playing the content by fetching the corresponding streamlet at the offset in the media content. it can. Each time the encoder stores another set of QSS files on the content server (for example, a set of 10 QSS files representing the next 2 seconds of media content with 10 different quality levels), the QMX file Has been updated to allow the Media Player 200 to fetch the QMX file to indicate over the internet that another 2 seconds are available for distribution. Media player 200 can periodically check for updated QMX files. Alternatively, a QMX file and any updates can be sent to the media player 200 to indicate when the media content is available for distribution over the Internet.
0051It should be noted that the Origin Content Server 210 is illustrated as being inside the CDN240, but it is outside the CDN240 and can also be associated with the CDN240. For example, a CDN240 where one entity can own and operate a content server that stores a streamlet, but its device may be owned and operated by one or more separate entities, is a stream. Deliver the let.
0052When the media content is processed by the media player 200 (running on an electronic device (ie, the client device)), the media player 200 gives the media player 200 a visual and / or audio representation of the event. Please note that it is the data that can be provided to the viewers of. The media player 200 may be one of the software that plays media content (for example, displays video and plays audio), in a standalone software application, web browser plug-in, or the like. May be, or browser plug-in and supporting web page It may be a combination with logic). For example, the event may be a television broadcast such as a sporting event, a live or recorded or recorded performance, live coverage or recorded or recorded coverage. A live event or scheduled television event in this context refers to media content that is scheduled to be played at a particular point in time according to a schedule. Live events may have pre-recorded content mixed with live media content, such as slow-motion clips of important events in it (eg replays), but such content is raw. Played between TV broadcasts. It should be noted that the embodiments described herein can also be used for video on demand (VOD) streaming.
0053FIG. 3A is a schematic block diagram showing another embodiment of the computing environment 300 in which the encoding system 320 including the plurality of hosts 314 each using the encoder 220 can be used. In one embodiment, the coding system 320 includes a master module 322 and a plurality of host computing modules (hereinafter hosts) 314. As described in connection with FIG. 2, each host 314 uses an encoder 220. Host 314 can be implemented on one or more personal computers, servers, and the like. In a further embodiment, the host 314 may be dedicated hardware, such as a card that plugs into a single computer.
0054In one embodiment, the master module (hereinafter "master") 322 is configured to receive raw streamlet 312 from streamlet generation system 301, which streamlet generation system 301 produces media content. It includes a receiving module 302 that receives from issuer 310 and a streamlet module 303 that segments media content into raw streamlets 312. Master module 322 stages the raw streamlet 312 for processing. In other embodiments, the master 322 can receive coded and / or compressed source streamlets and expand each source streamlet to produce a raw streamlet. As used herein, the term "raw streamlet" refers to uncompressed streamlet 312, or lightly compressed streamlet 312 to substantially reduce size without significant loss of quality. Point to that. A lightly compressed raw streamlet can be sent to more hosts more rapidly. Each host 314 is attached to a master 322 and is configured to receive a raw streamlet to be encoded from the master 322. In one example, host 314 produces multiple streamlets with the same time index, the same fixed time width, and different bit rates. In one embodiment, each host 314 is configured to generate a set of encoded streamlets 306 from a raw streamlet 312 sent from the master 322, where the coded streamlets of set 306 are media. Represents the same piece of content for each supported bit rate (ie, each streamlet is encoded according to one of the available quality profiles). Alternatively, to reduce the time it takes to encode, each host 314 is a single encoded stream at one of the supported bit rates.
0055When the coding is complete, host 314 returns set 306 to master 322, which allows the coding system 320 to store set 306 in the streamlet database 308. Master 322 is further configured to assign encoding jobs to host 314. In one embodiment, each host 314 is configured to present a coded job completion bid (hereinafter "bid") to the master 322. Master 322 allocates encoding jobs according to the bid from host 314. Each host 314 generates bids in response to multiple compute variables, which include the current encoded job completion percentage, average job completion time, processor speed, physical memory capacity, and so on. It may, but is not limited to these.
0056For example, host 314 can present a bid that indicates that host 314 could complete the encoding job in 15 seconds, based on past performance history. The master 322 is configured to select the best bid from a plurality of bids and then present the encoding job to the host 314 with the best bid. Therefore, the coding system 320 described does not require each host 314 to have the same hardware, and beneficially takes advantage of the available computing power of the host 314. Alternatively, the master 322 selects host 314 in order from earliest, or selects host 314 based on other algorithms that are considered suitable for a particular coding job.
0057The time it takes to encode a single streamlet depends on the computing power of host 314 and the encoding requirements for the content file of the original media content. Examples of coding requirements may include, but are not limited to, two-pass or multi-pass coding, multiple streams at various bit rates. One of the advantages of the present invention is the ability to perform 2-pass encoding on live content files. Typically, in order to perform two-pass encoding, prior art systems must wait for the content file to complete before encoding. However, streamlets can be encoded as many times as needed. Streamlets are sealed media objects with a short duration (eg, 2 seconds), so once the first streamlet is captured, multi-pass encoding can be initiated for live events.
0058In one embodiment, the encoder 220 segments the original content file into multiple source streamlets, eg, two-pass codes for multiple copies (eg, streams) without waiting for the TV show to end. Perform each of the corresponding raw streamlets 312. Thus, shortly after the streamlet generation system 301 begins capturing the original content file, the web server 316 can stream the streamlet over the Internet. The delay between the live broadcast sent by publisher 310 and the availability of content depends on the computing power of host 314.
0059FIG. 3B is a schematic block diagram showing an embodiment of parallel coding of streamlets 312 according to one embodiment. In one example, the streamlet generation system 301 starts capturing the original content file, generates a first streamlet 312a, and passes it to the encoding system 320. The coding system 320 includes a plurality of streamlets 304a (304a).<sub>1</sub>, 304a<sub>2</sub>, 304a<sub>3</sub>Etc. may take, for example, 10 seconds to generate the first set 306a consisting of streamlets 304 with different bit rates). To visually show the time width required to process the raw or lightly coded streamlet 312, as described above in connection with the coding system 320, in FIG. 3B, the coding is performed. The process is generally shown as block 308. The coding system 320 can process two or more streamlets 312 at the same time, and when the streamlets arrive from the streamlet generation module 301, the processing of the streamlets is started.
0060During the 10 seconds required to encode the first streamlet 312a, the streamlet module 404 has five additional 2 seconds of streamlet to encode, streamlet 312b, 312c, 312d, 312e, 312f. Is generated, and the master 322 creates the corresponding raw streamlet and stages it. When the first set 306a becomes available, the next set 306b becomes available two seconds later, and this flow is repeated thereafter. In this way, the original content files are encoded at various quality levels, streamed over the Internet, and appear live. The 10 second delay given herein is for example only. Multiple hosts 314 can be added to the coding system 320 to increase the processing power of the coding system 320. By adding a high CPU performance system, or by adding multiple low performance systems, the delay can be shortened to a level that is almost imperceptible.
0061Any particular encoding scheme applied to a streamlet can take longer than the time width of the streamlet itself to be complete. For example, for a 2-second streamlet, it can take 5 seconds to complete a very high quality encoding. Alternatively, the processing time required for each streamlet may be less than the time width of the streamlet. However, offset parallel coding of consecutive streamlets is coded at regular intervals by the coding system 320 (these streamlets match the intervals presented to the coding system 320, eg 2 seconds). Therefore, the output timing of the coding system 320 does not lag behind the real-time presentation speed of the uncoded streamlet 312.
0062With reference to FIG. 3A here, the master 322 and the host 314 can be located within a single local area network, as shown, in other words, the host 314 is physically close to the master 322. Can be made to. Alternatively, host 314 can also receive encoding jobs from master 322 via the Internet or other communication networks. For example, consider a live sporting event in a remote location where it is difficult to set up multiple hosts. In this example, the master does not encode or lightly encode the streamlet before issuing it online. Therefore, host 314 retrieves those streamlets and encodes them into a plurality of bit rate sets 306 as described above.
0063In addition, host 314 can be dynamically added to or removed from the coding system 320 without restarting the coding job and / or interrupting the issuance of the streamlet. .. If one host 314 crashes or fails, the encoding work is simply reassigned to another host.
0064In one embodiment, the coding system 320 can also be configured to produce a streamlet specific to a particular playback platform. For example, for a single raw streamlet, a single host 314 streamlets for different quality levels for personal computer playback, streams for playback on multiple mobile phones with different unique codecs. You can create a let, a small streamlet dedicated to video for use during playback with only a thumbnail view of the stream (such as a programming guide), and a very high quality streamlet for use in archiving. ..
0065In the illustrated embodiment, the computing environment 340 includes a content management system (CMS) 300. The CMS340 manages the encoded media content 220, for example by using the streamlet database 308, and the publisher creates and modifies a timeline (referred to herein as a virtual timeline (QVT)). It is a publishing system that makes it possible to plan the playback of media content 232. QVT is metadata that defines a playlist for viewers and can indicate when the media player 200 should play media content. For example, the timeline follows a schedule, specifying the start time for media content 232 and the current time period for media content 232 (eg, the amount of media content that can be used for distribution). It can be possible to play media events. In the example above, the encoder 220 updates the CMS240 with information about the stream (eg, a copy of media content 232), and a particular part of the stream (eg, streamlet) is the origin associated with the CDN240. Indicates that it was sent to the content server 210. In this embodiment, the CMS 340 receives information from the encoder 220, for example one of the following: Which quality for a particular portion of Availability Information / Media Content 232 indicates that the encryption key / encoder 220 pair sent a portion of the encoded Media Content 232 to the Origin Content Server 210. Information indicating whether a level is available / For example, content broadcast date, title, actress, actor, start index, end index, rights-owning issuer data, encryption level, content duration, episode or program name, publication Metadata including people / available menus, thumbnails, sidebars, ads, fast forward, rewind, pause, play, etc. A bit rate value that includes tools / frame size, audio channel information, codecs, sample rates, and frame parser information that can be used in the navigation environment. Alternatively, the encoder 220 may send more or less information than the information mentioned above.
0066In the illustrated embodiment, the computing environment 300 includes a digital rights management server (DRM) 350 that provides the system with digital rights management capabilities. The DRM350 is further configured to provide an encryption key to this end user by authenticating the end user. In one embodiment, the DRM server 350 is configured to authenticate the user based on the login proof. Those skilled in the art will appreciate a variety of different ways in which the DRAM server 350 can authenticate end users, such as encrypted cookies, user profiles, geographic location, etc. Sources, websites, etc. are included, but not limited to these.
0067In other embodiments, the computing environment 300 may include other devices such as directory servers, management servers, messaging servers, statistics servers, and devices for network infrastructure operators (eg, ISPs).
0068FIG. 4 shows an embodiment of method 400 in which the audio of the media content is encoded according to the codec application frame size and the audio frame without gaps is split between the content files having the fixed time video portion of the media content. It is a flow chart of. Method 400 may include hardware (circuits, dedicated logic circuits, etc.), software (such as running on a general purpose computer system or dedicated machine), or firmware (eg, embedded software), or It is performed by processing logic that may contain any combination of these elements. In one embodiment, method 400 is performed by encoder 220 in FIGS. 2 and 3A. In other embodiments, some of the operations of these methods may be performed by the fixed frame audio encoder 224 and audio splitting multiplexer 228 of FIG.
0069In Figure 4, the processing logic is started by initializing the sample offset to 0 (block 402) and receives the raw audio portion of the media content (block 404). The processing logic encodes the raw portion of the audio by using a fixed-frame audio codec (block 406) and buffers the encoded audio frame output by the audio codec (block 408). The processing logic determines if there are enough audio frames to fill the streamlet (block 410). In this embodiment, each streamlet also includes a fixed duration video frame, as described herein. If there are not enough audio frames to fill the streamlet, the processing logic returns to block 404, receives the next raw part of the audio, encodes this raw part of the audio, and buffers them in block 408. .. If the processing logic determines in block 410 that there are enough audio frames to fill the streamlet, it sends the audio frames to the audio splitting multiplexer and removes the transmit frames from the buffer (block 412). The processing logic updates the sample offset (block 414) to determine if the media content is finished (block 416). If the media content is not finished at block 416, the processing logic returns to block 404 to receive the other raw part of the audio. Otherwise, the method ends.
0070As described above for FIG. 2, the processing logic can be configured to perform various operations on the components of the encoder 220. For example, method 400 may be performed by a fixed frame audio encoder 224, which receives live audio 233 from the splitter 222, encodes the audio frame, and audios the encoded audio frame 237. Store in framebuffer 225. In this embodiment, the operations of blocks 402 to 408 can be performed by the fixed frame voice encoder 224, and the operations of blocks 410 to 416 can be performed by the voice splitting multiplexer 228. Alternatively, these operations can be performed by other combinations of encoder 220 components.
00715A-5C are flow diagrams of an embodiment of generating a content file with a fixed-time video portion and a gapless audio frame having a codec application frame size. Methods 500, 550, 570 can include hardware (circuits, dedicated logic circuits, etc.), software (such as running on a general purpose computer system or dedicated machine), or firmware (eg, embedded software). It is executed by processing logic that may or may contain any combination of these elements. In one embodiment, methods 500, 550, 570 are performed by encoder 220 in FIGS. 2 and 3A. In another embodiment, method 500 is performed on the fixed frame audio encoder 224, method 550 is performed on the fixed time video encoder 226, and method 570 is performed on the audio splitting multiplexer 228. Alternatively, the operations of methods 500, 550 and 570 can be performed by other combinations of encoder 220 components.
0072In Figure 5A, the processing logic of Method 500 is initiated by receiving the raw part of the audio (block 502). The processing logic encodes the raw part of the audio according to the codec application frame size (block 504) and buffers the encoded audio frame (block 506). The processing logic determines if the media content is finished (block 508). If the media content is not finished in block 508, the processing logic returns to block 502 to receive the other raw part of the audio. Otherwise, the method ends.
0073In Figure 5B, the processing logic of method 550 is initiated by receiving the raw portion of the video (block 552). The processing logic encodes the raw part of the video according to the frame rate (block 554) and buffers the encoded video frame (block 556). The processing logic determines if the media content is finished (block 558). If the media content is not finished at block 558, the processing logic returns to block 552 to receive the other raw part of the video. Otherwise, the method ends.
0074In Figure 5C, the processing logic of method 570 is initiated by receiving the encoded audio frame from the buffer (block 572) and the video frame from the buffer (block 574). The processing logic creates a streamlet (block 576) and sends it to the origin content server (block 578). The processing logic determines if the media content is finished (block 580). If the media content is not finished in block 580, the processing logic returns to block 572. Otherwise, the method ends.
0075In one embodiment, the processing logic determines in block 576 the number of video frames required to fill the streamlet and the number of audio frames required to fill the streamlet. In one embodiment, the number of video frames per streamlet is approximately fixed according to a fixed time width. For example, if the frame rate is 30fps, there are 60 frames in the 2-second streamlet. However, in reality, the video is not always exactly 30fps, but rather 29. Note that it is 97fps. So some 2-second streamlets may have 59 frames, some may have 60 frames, and some may even have 61 frames. Each frame in the streamlet has an offer time for the start of the streamlet. Thus, if a streamlet represents 30-32 seconds, the first frame within that streamlet may have a serving time of 6ms instead of 0ms. This frame should appear 30006ms from the start of the stream. For live performances, the encoder may drop frames to catch up if computational resources are limited and the encoder cannot keep up with the live flow. Therefore, some streamlets may have gaps in the video, which may be another cause of variations in the number of frames per streamlet. Alternatively, other frame rates other than 30fps, such as 24fps, can be used. The number of audio frames per streamlet is not fixed. The number of audio frames can be determined by the operation described above for the audio splitting multiplexer 228. The processing logic determines if there are enough frames in the buffer to fill the current streamlet. If there are not enough audio frames, the processing logic receives and encodes the next part of the audio, eg, one open frame of the audio, from the next part, as described herein. In some cases, the duration of the audio frame in the streamlet may be greater than the fixed time width, and in other cases, the duration of the audio frame may be less than the fixed time width.
0076FIG. 7 is a diagram of a machine of an exemplary embodiment of a computer system 700 for voice splitting according to one embodiment. Within the computer system 700, a set of instructions can be executed that causes the machine to perform either one or more of the speech splitting methodologies discussed herein. In an alternative embodiment, a machine can be connected (eg, networked) to another machine within a LAN, intranet, extranet, or Internet. The machine can operate within the capabilities of the server or client machine in a client-server network environment, or as a peer machine in a peer-to-peer (distributed) network environment. This machine is a PC, tablet PC, STB, PDA, mobile phone, web appliance, server, network router, switch or bridge, or a (continuous or other) command that specifies the action the machine should take. It can be any machine capable of running the set of. Further, although only a single machine is illustrated, one or more of the methodologies discussed herein, such as methods 400, 500, 550, 570, described herein for voice splitting operations. The term "machine" shall also include any set of machines that execute a set (or sets) of instructions individually or jointly to do any of the above. In one embodiment, the computer system 700 represents various components that can be implemented in the encoder 220 or coding system 320 as described above. Alternatively, the encoder 220 or coding system 320 may include more or fewer components than those illustrated in the computer system 700.
0077This exemplary computer system 700 includes a processing device 702, main memory 704 (eg, ROM (read-only memory), flash memory, DRAM (synchronous DRAM), DRAM (RDRAM), and other DRAM (dynamic random). (Access memory), etc.), static memory 706 (eg, flash memory, SRAM (static random access memory), etc.), data storage device 716, each of which communicates with each other via bus 730. It is carried out.
0078The processing device 702 represents one or more general purpose processing devices such as a microprocessor and a central processing unit. More specifically, the processing device 702 is a CISC (Compound Instruction Set Computing) microprocessor, RISC (Reduced Instruction Set Computing) microprocessor, VLIW (Ultra-Long Instruction) microprocessor, or other instruction. It may be a processor that implements a set of instructions or a processor that implements a combination of instruction sets. The processing device 702 may be one or more special purpose processing devices such as ASICs (application specific integrated circuits), FPGAs (field programmable gate arrays), DSPs (digital signal processors), network processors, etc. Good. The processing device 702 is configured to perform processing logic (eg, voice splitting 726) for performing the operations and steps discussed herein.
0079The computer system 700 may further include a network interface device 722. The computer system 700 includes a video display unit 710 (eg, liquid crystal display (LCD) or cathode ray tube (CRT)), alpha numeric input device 712 (eg keyboard), cursor control device 714 (eg mouse), signal generation device. It may also include a 720 (eg speaker).
0080The data storage device 716 provides a computer-readable storage medium 724 that stores one or more sets of instructions (eg, voice splitting 726) that implement one or more of the methodologies or functions described herein. May include. The voice splitting 726 may be present entirely or at least partially within the main memory 704 and / or the processing device 702 during execution by computer system 700, but these main memory 704 and processing Device 702 also constitutes a computer-readable storage medium. In addition, the network interface device 722 allows voice splitting 726 to be sent or received over the network.
0081In one exemplary embodiment, the computer-readable storage medium 724 is illustrated as a single medium, but the term "computer-readable medium" refers to a single or multiple medium (s) that store one or more sets of instructions. For example, centralized or distributed databases, and / or associated caches and servers) should also be considered. The term "computer-readable storage medium" also includes any medium that can store a set of instructions for execution by a machine and cause the machine to execute one or more of the methodologies of the present embodiment. It shall be. Thus, the term "computer-readable storage medium" includes, but is not limited to, solid-state memory, optical media, magnetic media, or other types of media for storing instructions. .. The term "computer-readable transmission medium" includes any medium capable of transmitting a set of instructions for execution by a machine to cause the machine to perform one or more of the methodologies of this embodiment. It shall be.
0082The voice splitting module 732, its respective components, and other features described herein (eg, in relation to Figures 2 and 3A) can be implemented as separate hardware components, or ASICS, FPGA. Can be integrated within the functionality of hardware components such as DSPs, DSPs, and similar devices. In addition, the voice splitting module 732 can also be implemented as firmware or functional circuits within a hardware device. In addition, the voice splitting module 732 can also be implemented within any combination of hardware devices and software components.
0083For purposes of explanation, the above description has been provided in connection with a particular embodiment. However, the above exemplary discussion is not intended to be exhaustive and is not intended to limit the invention to the exact forms disclosed. In view of the above teaching contents, many modified forms and modified forms can be considered. The above embodiments have been selected and described in order to best explain the principles of the present invention and examples of practical applications thereof, whereby those skilled in the art will appreciate the present invention and various embodiments with various modifications. , It will be possible to use it in a form suitable for the specific intended use.
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2001128119A | Cites | Japan |
| JP2008072718A | Cites | Japan |
| WO2006006714A1 | Cites | World Intellectual Property Organization (WIPO) |
32 members in 12 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 12643700 | United States of America | – | |
| 64370009 | United States of America | A | |
| 2010061658 | United States of America | W |
Members32
| Document | Office | Kind | |
|---|---|---|---|
| US2011150099A1 | United States of America | A1 | |
| CA2784779A1 | Canada | A1 | |
| WO2011084823A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2010339666A1 | Australia | A1 | |
| SG181840A1 | Singapore | A1 | |
| KR20120101710A | Republic of Korea | A | |
| CN102713883A | China | A | |
| EP2517121A1 | European Patent Office (EPO) | A1 | |
| MX2012007243A | Mexico | A | |
| JP2013515401A | Japan | A | |
| AU2010339666B2 | Australia | B2 | |
| KR20140097580A | Republic of Korea | A | |
| IL220482A | Israel | A | |
| KR101484900B1 | Republic of Korea | B1 | |
| JP5728736B2This record | Japan | B2 | |
| CA2784779C | Canada | C | |
| BR112012014872A2 | Brazil | A2 | |
| US9338523B2 | United States of America | B2 | |
| US2016240205A1 | United States of America | A1 | |
| CN102713883B | China | B | |
| EP2517121A4 | European Patent Office (EPO) | A4 | |
| CN106210768A | China | A | |
| US9601126B2 | United States of America | B2 | |
| US2017155910A1 | United States of America | A1 | |
| US9961349B2 | United States of America | B2 | |
| US2018234682A1 | United States of America | A1 | |
| US10230958B2 | United States of America | B2 | |
| CN106210768B | China | B | |
| US2019182488A1 | United States of America | A1 | |
| EP2517121B1 | European Patent Office (EPO) | B1 | |
| US10547850B2 | United States of America | B2 | |
| BR112012014872B1 | Brazil | B1 |
28 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Written request for registration of change of nameJAPANESE INTERMEDIATE CODE: R313533S533 | S533 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 5728736
- Application
- 2012544958
Titles2
- Japanese
- コーデック適用フレーム・サイズでの音声スプリッティング
- English
- Audio splitting at codec-applied frame size
Classification
- CPC, 15
- G10L19/167
- H04N21/23406
- H04N21/434
- H04N19/147
- H04N21/2343
- H04N21/2662
- H04N21/8456
- H04N21/439
- H04N21/236
- H04N7/52
- H04L65/65
- G06F3/165
- G10L21/055
- H04N19/15
- H04N19/172
- IPC, 2
- H04N21 2368
- H04N21 439
