Nova Patents
EP1659518A2

Automatic face extraction

Abstract

Faces of speakers in a meeting or conference are automatically detected and facial images corresponding to each speaker are stored in a faces database. A timeline is created to graphically identify when each speaker is speaking during playback of a recording of the meeting. Instead of generically identifying each speaker in the timeline, a facial image is shown to identify each speaker associated with the timeline.

EP1659518A2, drawing sheet 1
Sheet 1 of 7

Term

Term ended

Projected expiry passed 18 October 2025, 0.9 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

32 claims: 5 independent, 27 dependent

  1. 1
    A method, comprising:detecting one or more facial images in a video sample;detecting one or more speakers in an audio sample that corresponds to the video sample;storing a speaker timeline that identifies a speaker by a speaker identifier and a speaker location at each time along a the speaker timeline;storing at least one facial image for each detected speaker in a faces database;and associating a speaker timeline and a facial image with each detected speaker.
  2. 10
    A method, comprising:displaying an audio/visual (A/V) sample having one or more speakers included therein;displaying a speaker timeline corresponding to each speaker, the speaker timeline indicating at what points along a temporal continuum the speaker corresponding to the speaker timeline is speaking;associating a speaker facial image with each speaker timeline, the speaker facial image corresponding to the speaker associated with the speaker timeline;and displaying the facial image with the corresponding speaker timeline.
  3. 13
    One or more computer-readable media containing executable instructions that, when executed, implement the following method:identifying each speaker in an A/V sample by a speaker identifier;identifying location for each speaker in the A/V sample;extracting at least one facial image for each speaker identified in the A/V sample;creating a speaker timeline for each speaker identified in the A/V sample, each speaker timeline indicating a time, a speaker identifier and a speaker location;and associating the facial image for a speaker with a speaker timeline that corresponds to the same speaker.
  4. 23
    One or more computer-readable media, comprising:a speaker timeline database that includes a speaker timeline for each speaker in an A/V sample, each speaker timeline identifying a speaker and a speaker location for multiple times along a time continuum;and a faces database that includes at least one facial image for each speaker identified in a speaker timeline and a speaker identifier that links each facial image with the appropriate speaker timeline in the speaker timeline database.
  5. 25
    A system, comprising:an A/V sample;means for identifying each speaker appearing in the A/V sample;means for identifying a facial image for each speaker identified in the A/V sample;means for creating a speaker timeline for each speaker identified in the A/V sample;and means for associating a facial image with an appropriate speaker timeline.