US9195649B2

Audio processing techniques for semantic audio recognition and report generation

Summary by NHIP

Audio Template Semantic Recognition

The method forms audio templates by extracting distinct temporal, spectral, harmonic, and rhythmic features to determine specific feature ranges. Stored ranges compare against subsequent audio to generate tags identifying genre, instrumentation, style, acoustical dynamics, and emotive descriptors.

Claim Score by NHIP

Read claim 20, the broadest

Abstract

System, apparatus and method for determining semantic information from audio, where incoming audio is sampled and processed to extract audio features, including temporal, spectral, harmonic and rhythmic features. The extracted audio features are compared to stored audio templates that include ranges and/or values for certain features and are tagged for specific ranges and/or values. Extracted audio features that are most similar to one or more templates from the comparison are identified according to the tagged information. The tags are used to determine the semantic audio data that includes genre, instrumentation, style, acoustical dynamics, and emotive descriptor for the audio signal.

US9195649B2, drawing sheet 1
Sheet 1 of 19

Term

6.4 yearsleft in the term

Expires 26 February 2033, including 67 days of term adjustment.

  1. Priority and filed
  2. Granted
  3. Today
  4. Expires

38 claims: 6 independent, 32 dependent

  1. 1
    A method for forming an audio template for determining semantic audio information, comprising:extracting a first audio feature from audio, the first audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature;extracting a second audio feature from the audio, the second audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature, wherein the second audio feature is different from the first audio feature;determining a first range for the first audio feature and a second range for the second audio feature;and storing the first and second ranges to compare against other audio features from subsequent audio to generate tags signifying semantic audio information for the subsequent audio.
  2. 8
    A processor-based method for determining semantic audio information for audio, comprising:extracting a first audio feature from the audio, the first audio feature including at least one of a rhythmic structure, a beat period, a rhythmic fluctuation, or an average tempo;extracting a second audio feature from the audio, the second audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature, wherein the second audio feature is different from the first audio feature;comparing the first and second audio features to a plurality of stored audio feature ranges having tags associated therewith;and determining the stored audio feature ranges having the closest matches to the first and second audio features, the tags associated with the audio feature ranges having the closest matches to be used to determine the semantic audio information for the audio.
  3. 14
    An apparatus to form an audio template for determining semantic audio information, comprising:a processor to: extract a first audio feature from audio, the first audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature;extract a second audio feature from the audio, the second audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature, and the second audio feature is different from the first audio feature;and determine a first range for the first audio feature and a second range for the second audio feature;and a storage to store the first and second ranges to compare against other audio features from subsequent audio to generate semantic audio information for the subsequent audio.
  4. 20
    Broadest claimClaim Score 52, average(NHIP)An article of manufacture comprising instructions that, when executed, cause a computing device to at least:extract a first audio feature from audio, the first audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature;extract a second audio feature from the audio, the second audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature, wherein the second audio feature is different from the first audio feature;determine a first range for the first audio feature and a second range for the second audio feature;and store the first range and the second range to compare against other audio features from subsequent audio to generate tags signifying semantic audio information for the subsequent audio.
  5. 27
    An apparatus to determine semantic audio information from audio, comprising:a processor to: extract a first audio feature from the audio, the first audio feature including at least one of a rhythmic structure, a beat period, a rhythmic fluctuation, or an average tempo;extract a second audio feature from the audio, the second audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature, wherein the second audio feature is different from the first audio feature;compare the first and second audio features to a plurality of stored audio feature ranges having tags associated therewith;and determine the stored audio feature ranges matching the first and second audio features, the tags associated with the matching audio feature ranges to be used to determine the semantic audio information for the audio.
  6. 33
    An article of manufacture comprising instructions that, when executed, cause a computing device to at least:extract a first audio feature from audio, the first audio feature including at least one of a rhythmic structure, a beat period, a rhythmic fluctuation, or an average tempo;extract a second audio feature from the audio, the second audio feature including at least one of a temporal feature, a spectral feature, a harmonic feature, or a rhythmic feature, wherein the second audio feature is different from the first audio feature;compare the first and second audio features to a plurality of stored audio feature ranges having tags associated therewith;and determine the stored audio feature ranges matching the first and second audio features, the tags associated with the matching audio feature ranges to be used to determine semantic audio information for the audio.