JP4878437B2

Systems and methods for generating audio thumbnails

Abstract

This record has no abstract on file.

Term

Term ended

Expired 23 February 2025, 1.6 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

37 claims: 5 independent, 32 dependent

  1. 1
    A system for summarizing audio information, an analyzer that converts audio into frames and a fingerprinting component that converts said frames into fingerprints, where each fingerprint is partially based on multiple frames. A similarity detector that calculates the similarity between a component and a fingerprint, and the similarity detector is clustered.functionThe clusteringfunctionIs the initial threshold for similarityAll fingerprints that fitA heuristic module that generates thumbnails of audio files from a similarity detector and a set of clusters that have at least two gaps between fingerprints, which produces one or more sets of fingerprint clusters based on. The gap is characterized by having a heuristic module, which is the time interval between two adjacent fingerprints that exceed a predetermined threshold when the fingerprints in the cluster set are arranged in a sequential time order. System. オーディオ情報を要約するためのシステムであって、 オーディオをフレームに変換するアナライザと、 前記フレームをフィンガープリントに変換するフィンガープリンティングコンポーネントであって、各フィンガープリントが複数のフレームに部分的に基づくフィンガープリンティングコンポーネントと、 フィンガープリント間の類似性を計算する類似性ディテクタであって、前記類似性ディテクタは、クラスタリング機能を備え、前記クラスタリング機能は、類似性を示す初期のしきい値にかなうすべてのフィンガープリントに基づいてフィンガープリントのクラスタの1つまたは複数の集合を生成する、類似性ディテクタと、 フィンガープリント間の少なくとも2つのギャップを有するクラスタの集合からオーディオファイルのサムネイルを生成するヒューリスティックモジュールであって、ギャップは、クラスタの集合内のフィンガープリントが順次的な時間順序で配置されるとき所定のしきい値を超える2つの隣接するフィンガープリント間の時間間隔である、ヒューリスティックモジュールと を備えたことを特徴とするシステム。
  2. 8
    7. A claim 7, wherein for each frame, the average normalized energy E is calculated by dividing the average energy per frequency component within that frame by the average of that amount over the frames in the audio file. system. 各フレームについて、そのフレーム内の周波数成分あたりの平均エネルギーをオーディオファイル中のフレームにわたるその量の平均で割ることによって平均の正規化したエネルギーEを計算することを特徴とする請求項7に記載のシステム。
  3. 17
    A means for converting an audio file into frames, a means for fingerprinting the audio file and generating a fingerprint based partially on multiple frames, and a predefined similarity threshold.All fingerprints that fitA means for generating one or more sets of fingerprint clusters based on, and a means for generating audio thumbnails by selecting a set of clusters that have at least two gaps between fingerprints. The gap is characterized by having a time interval between two adjacent fingerprints that exceed a predetermined threshold when the fingerprints within the set of clusters are placed in sequential time order. Automatic thumbnail generator. オーディオファイルをフレームに変換するための手段と、 前記オーディオファイルをフィンガープリンティングし、複数のフレームに部分的に基づいてフィンガープリントを生成するための手段と、 予め定義された類似性しきい値にかなうすべてのフィンガープリントに基づいてフィンガープリントのクラスタの1つまたは複数の集合を生成する手段と、 フィンガープリント間の少なくとも2つのギャップを有するクラスタの集合を選択することによってオーディオサムネイルを生成するための手段であって、ギャップは、クラスタの集合内のフィンガープリントが順次的な時間順序で配置されるとき所定のしきい値を超える2つの隣接するフィンガープリント間の時間間隔であることと を備えたことを特徴とする自動サムネイルジェネレータ。
  4. 18
    A way to generate audio thumbnails, which is to generate multiple audio fingerprints, where each audio fingerprint is partially based on multiple audio frames and the similarity threshold.All fingerprints that fitTo generate one or more sets of fingerprint clusters based on and to create thumbnails based on a set of clusters that have at least two gaps between fingerprints. A method characterized by having a time interval between two adjacent fingerprints that exceed a predetermined threshold when the fingerprints in the set are arranged in a sequential time order. オーディオサムネイルを生成する方法であって、 複数のオーディオフィンガープリントを生成することであって、各オーディオフィンガープリントが複数のオーディオフレームに部分的に基づくことと、 類似性しきい値にかなうすべてのフィンガープリントに基づいてフィンガープリントのクラスタの1つまたは複数の集合を生成することと、 フィンガープリント間の少なくとも2つのギャップを有するクラスタの集合に基づいてサムネイルを作成することであって、ギャップは、クラスタの集合内のフィンガープリントが順次的な時間順序で配置されるとき所定のしきい値を超える2つの隣接するフィンガープリント間の時間間隔であることと を備えることを特徴とする方法。
  5. 29
    It is characterized in that the average spectral flatness and the parameter D are combined into a single parameter associated with each cluster set, whereby the set having an external value of the parameter is selected to be the best set. 28. 前記平均のスペクトル平坦性およびパラメータDを組み合わせて各クラスタ集合に関連付けられた単一のパラメータとし、それによって前記パラメータの外部値を有する集合を前記最良の集合とするように選択することを特徴とする請求項28に記載の方法。