EP1566748A1

Method and apparatus for detecting anchorperson shot

Abstract

A method and an apparatus for detecting an anchorperson shot are provided. The method includes separating a moving image into audio signals and video signals; deciding boundaries between shots using the video signals; and extracting shots having a length larger than a first threshold value and a silent section having a length larger than a second threshold value from the audio signals using the boundaries and deciding the extracted shots as anchorperson speech shots.

EP1566748A1, drawing sheet 1
Sheet 1 of 27

Term

Term ended

Projected expiry passed 21 December 2024, 1.8 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

41 claims: 15 independent, 26 dependent

  1. 1
    A method of detecting an anchorperson shot, the method comprising:separating a moving image into audio signals and video signals;deciding boundaries between shots using the video signals;and extracting shots having a length larger than a first threshold value and a silent section having a length larger than a second threshold value from the audio signals using the boundaries and deciding the extracted shots as anchorperson speech shots.
  2. 4
    The method of any preceding claim, further comprising down-sampling the audio signals, and the shots having the length larger than the first threshold value and the silent section having the length larger than the second threshold value are extracted from the down-sampled audio signals using the boundaries and are decided as the anchorperson speech shots.
  3. 5
    The method of any preceding claim, wherein the deciding of the anchorperson speech shots comprises:obtaining the length of each of the shots using the boundaries between the shots;selecting the shots having a length larger than the first threshold value from the shots;obtaining a length of the silent section of each of the selected shots;and extracting shots having the silent section having a length larger than the second threshold value from the selected shots.
  4. 9
    The method of any one of claims 6 to 8, wherein in the counting of the number of the frames, a last frame of each of the selected shots is not counted.
  5. 10
    The method of any one of claims 6 to 9, wherein in the counting of the number of the frames is stopped when the frames having an energy larger than the silent threshold value exist continuously.
  6. 11
    The method of any preceding claim, wherein the deciding of the anchorperson speech shots further comprises selecting only shots of a predetermined percentage having a relatively large length from the extracted shots and deciding the selected shots as the anchorperson speech shots.
  7. 12
    The method of any preceding claim, further comprising:separating anchorpersons' speech shots that contain anchorpersons' voices, from the anchorperson speech shots;grouping anchorperson's speech shots excluding the anchorpersons' speech shots from the anchorperson speech shots, grouping the anchorpersons' speech shots, and deciding the grouped results as similar groups;and obtaining a representative value of each of the similar groups as an anchorperson speech model.
  8. 18
    The method of any one of claims 13 to 17, wherein the detecting of the anchorpersons' speech shots comprises:obtaining average values of the MFCCs according to each coefficient of the frame of each window while moving a window having a predetermined length at predetermined time intervals with respect to each of the anchorperson speech shots from which the silent frame and the consonant frame are removed;obtaining a difference between the average values of the MFCCs between adjacent windows;and deciding the anchorperson speech shots as anchorpersons' speech shots having the difference larger than a third threshold value with respect to each of the anchorperson speech shots from which the silent frame and the consonant frame are removed.
  9. 19
    The method of any one of claims 13 to 18, wherein in the detecting of the anchorpersons' speech shots, the MFCCs according to each coefficient and power spectral densities (PSDs) in a predetermined frequency bandwidth are obtained in each of the frames included in each of the anchorperson speech shots from which the silent frame and the consonant frame are removed, and the anchorpersons' speech shots are detected using the MFCCs according to each coefficient and the PSDs.
  10. 22
    The method of any one of claims 12 to 21, wherein the grouping of the anchorperson's speech shots and deciding the similar groups comprises:obtaining average values of the MFCCs in each of the anchorperson's speech shots;when a MFCC distance calculated using the average values of the MFCCs according to each coefficient of two anchorpersons' speech shots is the closest among the anchorperson speech shots and smaller than a fifth threshold value, deciding the two anchorpersons' speech shots as similar candidate shots;obtaining a difference between average decibel values of PSDs in a predetermined frequency bandwidth of the similar candidate shots;when the difference between the average decibel values is smaller than a sixth threshold value, grouping the similar candidate shots and deciding the grouped similar candidate shots as the similar groups;and determining whether all of the anchorperson's speech shots are grouped, and wherein if it is determined that all of the anchorperson's speech shots are not grouped, deciding the similar candidate shots with respect to other two anchorperson's speech shots, obtaining the difference, and deciding the similar groups are performed.
  11. 24
    The method of any one of claims 12 to 23, wherein the representative value is the average value of MFCCs according to each coefficient of shots that belong to the similar groups and the average decibel value of PSDs in the predetermined frequency bandwidth of the shots that belong to the similar groups.
  12. 25
    The method of any one of claims 12 to 24, further comprising generating a separate speech model using information about initial frames among frames included in each of the similar groups.
  13. 26
    The method of any one of claims 12 to 25, further comprising generating an anchorperson image model.
  14. 33
    An apparatus for detecting an anchorperson shot, the apparatus comprising:a signal separating unit separating a moving image into audio signals and video signals;a boundary deciding unit deciding boundaries between shots using the video signals;and an anchorperson speech shot extracting unit extracting shots having a length larger than a first threshold value and a silent section having a length larger than a second threshold value from the audio signals using the boundaries and outputting the extracted shots as anchorperson speech shots.
  15. 39
    The apparatus of any one of claims 36 to 38, further comprising an image model generating unit generating an anchorperson speech model.