Automatic face extraction
Abstract
A method implemented by computer, comprising: detecting one or more speakers in an audio sample that corresponds to a video sample; storing a speaker's timeline, which identifies a speaker, by a speaker identifier and a speaker's location at all times along a speaker's timeline; detect, in the video sample, one or more facial images for each speaker detected; storing at least one facial image detected for each detected speaker in a face database; and associate a speaker timeline and a facial image with each speaker detected.

Term
Term ended
Projected expiry passed 18 October 2025, 0.9 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
30 claims: 5 independent, 25 dependent
- 1ES 2 645 313 T3 REIVINDICACIONES 1. Un procedimiento implementado por ordenador, que comprende:detectar uno o más hablantes en una muestra de audio que corresponde a una muestra de vídeo;almacenar una línea de tiempo de hablante, que identifica un hablante, por un identificador de hablante y una ubicación de hablante en todo momento a lo largo de una línea de tiempo de hablante;detectar, en la muestra de vídeo, una o más imágenes faciales para cada hablante detectado;almacenar en una base de datos de rostros al menos una imagen facial detectada para cada hablante detectado;y asociar una línea de tiempo de hablante y una imagen facial con cada hablante detectado.
- 2El procedimiento implementado por ordenador de acuerdo con la reivindicación 1, en el que la detección de una o más imágenes faciales comprende adicionalmente el uso de un seguimiento de rostros para detectar una o más imágenes faciales.
- 3El procedimiento implementado por ordenador de acuerdo con la reivindicación 1, en el que la detección de uno o más hablantes comprende adicionalmente el uso de una localización de fuentes de sonido para detectar uno o más hablantes.
- 4El procedimiento implementado por ordenador de acuerdo con la reivindicación 1, que comprende adicionalmente:identificar más de una imagen facial para cada hablante;y seleccionar la mejor imagen facial para almacenar en la base de datos de rostros.
- 5El procedimiento implementado por ordenador de acuerdo con la reivindicación 4, en el que la selección comprende adicionalmente seleccionar como mejor imagen facial una imagen facial que incluya una vista facial más frontal.
- 6El procedimiento implementado por ordenador de acuerdo con la reivindicación 4, en el que la selección comprende adicionalmente seleccionar como mejor imagen facial una imagen facial que presente el menor movimiento.
- 7Procedimiento implementado por ordenador de acuerdo con la reivindicación 4, en el que la selección comprende adicionalmente seleccionar como mejor imagen facial una imagen facial que presente la máxima simetría.
- 8El procedimiento implementado por ordenador de acuerdo con la reivindicación 1, en el que la ubicación del hablante es señalada por una caja delimitadora de hablante identificada por coordenadas de muestra de vídeo.
- 9El procedimiento implementado por ordenador de acuerdo con la reivindicación 1, en el que la ubicación del hablante es señalada por ángulos faciales del hablante identificados por azimut y elevación en la muestra de vídeo.
- 10Un procedimiento implementado por ordenador, que comprende:mostrar visualmente una muestra audio/visual (A/V) que tenga uno o más hablantes incluidos en la misma;mostrar una línea de tiempo de hablante correspondiente a cada hablante, indicando la línea de tiempo de hablante en qué puntos a lo largo de un continuo temporal y desde qué ubicación está hablando el hablante correspondiente a la línea de tiempo de hablante;asociar una imagen facial de hablante con cada línea de tiempo de hablante, correspondiendo la imagen facial de hablante al hablante asociado con la línea de tiempo de hablante;y mostrar la imagen facial con la correspondiente línea de tiempo de hablante.
- 11El procedimiento implementado por ordenador de acuerdo con la reivindicación 10, que comprende adicionalmente recuperar la imagen facial de hablante de una base de datos de rostros que asocia cada identificador de hablante con al menos una imagen facial de un hablante que corresponda al identificador de hablante.
- 12Uno o más medios legibles por ordenador que contienen instrucciones ejecutables que, cuando se ejecutan, implementan el siguiente procedimiento:identificar cada hablante de una muestra A/V por un identificador de hablante;identificar una ubicación para cada hablante de la muestra A/V;extraer al menos una imagen facial para cada hablante identificado en la muestra A/V;crear una línea de tiempo de hablante para cada hablante identificado en la muestra A/V, indicando cada línea de tiempo de hablante un tiempo, un identificador de hablante y una ubicación de hablante;y asociar la imagen facial para un hablante con una línea de tiempo de hablante que corresponda al mismo hablante.
- 13Uno o más medios legibles por ordenador de acuerdo con la reivindicación 12, que comprenden adicionalmente identificar cada hablante utilizando localización de fuentes de sonido. ES 2 645 313 T3
- 14Uno o más medios legibles por ordenador de acuerdo con la reivindicación 12, que comprenden adlclonalmente Identificar cada ubicación de hablante utilizando un seguidor de rostros.
- 15Uno o más medios legibles por ordenador de acuerdo con la reivindicación 12, en el que la ubicación de los hablantes es identificada por una caja delimitadora de hablante en la muestra A/V.
- 16Uno o más medios legibles por ordenador de acuerdo con la reivindicación 12, que comprenden adicionalmente almacenar las líneas de tiempo de hablante y las imágenes faciales y enlazar cada línea de tiempo de hablante con la imagen facial apropiada.
- 17Uno o más medios legibles por ordenador de acuerdo con la reivindicación 12, que comprenden adicionalmente extraer más de una imagen facial para cada hablante.
- 18Uno o más medios legibles por ordenador de acuerdo con la reivindicación 17, que comprenden adicionalmente seleccionar la mejor imagen facial para asociar con la línea de tiempo de hablante.
- 19Uno o más medios legibles por ordenador de acuerdo con la reivindicación 18, en el que la selección de una mejor imagen facial comprende adicionalmente seleccionar una imagen facial que tenga una imagen facial frontal máxima.
- 20Uno o más medios legibles por ordenador de acuerdo con la reivindicación 18, en el que la selección de una mejor imagen facial comprende adicionalmente seleccionar una imagen facial que presente el menor movimiento.
- 21Uno o más medios legibles por ordenador de acuerdo con la reivindicación 18, en el que la selección de una mejor imagen facial comprende adicionalmente seleccionar una imagen facial que presente la máxima simetría facial.
- 22Uno o más medios legibles por ordenador, que comprenden:una base de datos de líneas de tiempo de hablantes que incluye una línea de tiempo de hablante para cada hablante de una muestra A/V, identificando cada línea de tiempo de hablante un hablante y una ubicación de hablante en múltiples ocasiones a lo largo de un continuo de tiempo;y una base de datos de rostros que incluye al menos una imagen facial para cada hablante identificado en una línea de tiempo de hablante y un identificador de hablante que vincula cada imagen facial con la apropiada línea de tiempo de hablante de la base de datos de líneas de tiempo de hablantes.
- 23Uno o más medios legibles por ordenador de acuerdo con la reivindicación 22, en el que cada línea de tiempo de hablante de la base de datos de líneas de tiempo de hablantes incluye el identificador de hablante apropiado para enlazar la base de datos de líneas de tiempo de hablantes con la base de datos de rostros.
- 24Un sistema, que comprende:una muestra A/V;medios para identificar cada hablante que aparece en la muestra A/V;medios para identificar una imagen facial para cada hablante identificado en la muestra A/V;medios para crear una línea de tiempo de hablante para cada hablante identificado en la muestra de A/V;y medios para asociar una imagen facial con una línea de tiempo de hablante apropiada;en el que los medios para identificar cada hablante comprenden además un localizador de fuentes de sonido.
- 25El sistema de acuerdo con la reivindicación 24, en el que los medios para identificar una imagen facial comprenden además un seguidor de rostros.
- 26El sistema de acuerdo con la reivindicación 24, en el que una línea de tiempo de hablante identifica un hablante asociado con la línea de tiempo de hablante mediante un identificador de hablante y una ubicación de hablante para cada una de múltiples ocasiones a lo largo de un continuo de tiempo.
- 27El sistema de acuerdo con la reivindicación 26, en el que la asociación de una imagen facial con una línea de tiempo de hablante apropiada comprende adicionalmente asociar cada imagen facial con el identificador de hablante.
- 28El sistema de acuerdo con la reivindicación 24, que comprende adicionalmente almacenar las líneas de tiempo de hablante y las imágenes faciales.
- 29El sistema de acuerdo con la reivindicación 28, en el que las líneas de tiempo de hablante y las imágenes faciales se almacenan por separado.
- 30El sistema de acuerdo con la reivindicación 24, en el que la muestra A/V comprende además una reunión grabada.
Independent claims30
87 paragraphs in 1 section, as filed
ES 2 645 313 T3
DESCRIPTION
Automatic face extraction
Field of the invention
The following description generally refers to the processing of video images. More particularly, the following description refers to providing an indexed timeline for video playback.
Background
Playback of recorded video in scenarios that include more than one speaker, such as playback of a recorded meeting, is often displayed simultaneously with an indexed timeline. Using the timeline, a user can quickly move to a particular moment in the meeting by manipulating one or more controls on the timeline. When the video includes more than one speaker, multiple timelines can be used, with one timeline being associated with a particular speaker. Each timeline indicates when a corresponding speaker is speaking. In this way, a user can navigate to portions of the meeting in which a particular speaker is speaking.
Such multiple timelines can be generically labeled to identify each speaker, such as Speaker 1, Speaker 2, and so on. Current techniques for automatically tagging timelines with specific speaker names are inaccurate and may also require a database of users with their corresponding fingerprints and facial prints, which could pose security and privacy concerns.
From GB 2 395 852 a media handling system is known comprising means for detecting human faces.
From WO 99/607 88 a system is known, such as a video conference system, comprising an audio source locator.
Brief description of the drawings
The foregoing aspects and many of the associated advantages of this invention will be more readily appreciated as they are better understood by reference to the following detailed description, taken in conjunction with the accompanying drawings, in which:
Figure 1 is a block diagram depicting an exemplary general purpose camera and computing device.
Figure 2 is a block diagram depicting an exemplary panoramic camera and client device.
Figure 3 is a representation of an exemplary playback screen with a panoramic image and a facial image timeline.
Figure 4 is an exemplary playback screen with a panoramic image and a facial image timeline.
Figure 5 is an exemplary flow chart of a methodological implementation for creating a timeline with facial images.
Figure 6 is an exemplary flow chart depicting a methodological implementation for creating a database of faces.
Detailed description
The following description refers to various implementations and embodiments for automatically detecting the face of each speaker in a multi-speaker environment and associating one or more images of the face of a speaker with a portion of a timeline that corresponds to the speaker. This type of specific tagging has the advantage over generic tagging that a viewer can more easily determine which portion of a timeline corresponds to a particular of multiple speakers.
In the following discussion, an example of a panoramic camera is described in which the panoramic camera is used to record a meeting that has more than one participant and / or speaker. Although a panoramic camera that includes multiple cameras is described, the following description also refers to individual cameras and multi-camera devices having two or more cameras.
A panoramic image is fed into a face tracker (FT) that detects and tracks the faces in the meeting. An array of microphones is fed into a sound source locator (SSL) that detects speaker locations based on sound. The outputs of the face tracker and the sound source locator are fed into a virtual filmmaker to detect the locations of the speakers.
ES 2 645 313 T3
Speakers are post-processed with a speaker grouping module that groups speakers temporally and spatially to better delineate an aggregated timeline that includes two or more individual timelines. The (aggregated) timeline is stored in a timeline database. A face database is created to store one or more images for each speaker, of which at least one of each face will be used in a timeline associated with a speaker.
The concepts presented and claimed herein are described in greater detail below with respect to one or more appropriate operating environments. Some of the items described below are also described in US Patent Application No. 10 / 177,315, entitled System and Procedure for Distributed Meetings, filed 06/21/2002.
Exemplary Operating Environment
Fig. 1 is a block diagram depicting a general purpose camera and computing device. The computing system environment 100 is only one example of a suitable computing environment and is not intended to suggest any limitation as to the scope of use or functionality of the claimed object. Nor should the computing environment 100 be construed to have any dependencies or requirements related to any one or a combination of the components illustrated in the exemplary operating environment 100.
The techniques and objects described may operate with numerous other environments or configurations of general-purpose or special-purpose computer systems. Examples of well known computer systems, environments, and / or configurations that may be suitable for use include, but are not limited to, personal computers, server computers, handheld or portable devices, multiprocessor systems, microprocessor-based systems, television set-top boxes. , programmable consumer electronics, network PCs, minicomputers, mainframes, distributed computing environments that include any of the above systems or devices, and the like.
The following description can be formulated in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and so on. that perform particular tasks or implement particular types of abstract data. The described implementations can also be implemented in distributed computing environments where tasks are performed by remote processing devices that are linked through a communication network. In a distributed computing environment, the program modules can be located on both local and remote computing storage media, including memory storage devices.
Referring to FIG. 1, an exemplary system for implementing the invention includes a general purpose computing device in the form of a computer 110. The components of the computer 110 may include, but are not limited to, a processing unit 120, a memory 130 system and a system bus 121 that couples various system components, including system memory, to processing unit 120. The system bus 121 can be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. By way of example, and without limitation, such architectures include ISA (Industry Standard Architecture) bus, MCA (Micro Channel Architecture) bus, Enhanced ISA (EISA) bus, VESA (Electronic Video Standards Association) local bus and the Peripheral Component Interconnect (PCI) bus also known as the Mezzanine bus.
Computer 110 typically includes a variety of computer-readable media. The computer-readable medium can be any available medium that the computer 110 can access and includes both volatile and non-volatile media, removable and non-removable media. By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile discs (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, storage on magnetic disk or other magnetic storage devices or any other medium that can be used to store the desired information and that can be accessed by the computer 110. Communication media typically incorporate computer-readable instructions, data structures, program modules, or other data into a modulated data signal, such as a carrier wave or other transport mechanism, and include any information delivery medium. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a way that they encode information in the signal. By way of example, and without limitation, communication media includes wired media, such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above should also fall within the scope of computer-readable media.
ES 2 645 313 T3
System memory 130 includes Computer storage media in the form of volatile and / or non-volatile memory, such as read-only memory (ROM) 131 and random access memory (RAM) 132. A basic input / output system (BIOS) 133, which contains the basic routines that help transfer information between items within computer 110, such as during startup, is typically stored in ROM 131. RAM 132 typically contains data and / or program modules that processing unit 120 operates with and / or can immediately access. By way of example, and not limitation, FIG. 1 illustrates an operating system 134, application programs 135, other program modules 136, and program data 137.
Computer 110 may also include other volatile or non-volatile, removable or non-removable storage media. By way of example only, Fig. 1 illustrates a hard disk drive 141 that reads or writes to non-volatile, non-removable magnetic media, a magnetic disk drive 151 that reads or writes to a removable, non-volatile magnetic disk 152, and an optical disk drive 155 that reads or writes to a removable, non-volatile optical disk 156, such as a CD-ROM or other optical media. Other volatile or non-volatile removable or non-removable computer storage media that may be used in the exemplary operating environment include, but are not limited to, magnetic tape cassettes, flash memory cards, digital versatile disks, digital video tape, memory RAM. solid state, solid state ROM, and the like. Hard disk drive 141 is typically connected to system bus 121 through a removable memory interface, such as interface 140, and magnetic disk drive 151 and optical disk drive 155 are typically connected to bus 121 of the system. system using a removable memory interface, such as interface 150.
The drives and their corresponding computer storage media discussed above, and illustrated in FIG. 1, provide storage of computer-readable instructions, data structures, program modules, and other data for the computer 110. In FIG. 1, by For example, hard disk drive 141 is illustrated as storing operating system 144, application programs 145, other program modules 146, and program data 147. Note that these components may be the same as or different from the operating system 134, the application programs 135, the other program modules 136, and the program data 137. The operating system 144, the application programs 145, the other program modules 146, and the program data 147 are given different numbers herein to illustrate that they are at least different copies. A user can enter commands and information into the computer 110 through input devices such as a keyboard 162 and a pointing device 161, commonly referred to as a mouse, trackball, or touch pad. Other input devices (not shown) can include a microphone, a joystick, a game pad, a satellite dish, a scanner, or the like. These and other input devices are often connected to the processing unit 120 through a user input interface 160 that is coupled to the system bus 121, but may be connected by other interface and bus structures, such as a parallel port, game port, or universal serial bus (USB). A monitor 191 or other type of display device is also connected to the system bus 121 through an interface, such as a video interface 190. In addition to the monitor, the computers may also include other peripheral output devices such as speakers 197 and a printer 196, which may be connected through a peripheral output interface 195. Of particular importance to the present invention, a camera 163 (such as a still or video, digital or electronic camera, or a film or photographic scanner) capable of capturing a sequence may also be included as an input device to the personal computer 110. Imager 164. Additionally, although only one camera is depicted, multiple cameras could be included as input devices to the personal computer 110. Images 164 from one or more cameras are input to computer 110 through an appropriate camera interface 165. This interface 165 is connected to the bus 121 of the system, thus allowing the images to be sent and stored in the RAM 132, or in one of the other data storage devices associated with the computer 110. However, it is noted that image data can be input to computer 110 from any of the aforementioned computer-readable media, without the need to use camera 163.
Computer 110 may operate in a network environment that uses logical connections to one or more remote computers, such as remote computer 180. Remote computer 180 may be a personal computer, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the items described above in relation to computer 110, although only one storage memory device 181 is illustrated in Fig. 1. The logical connections depicted in FIG. 1 include a local area network (LAN) 171 and a wide area network (WAN) 173, but may also include other networks. Such network environments are common in offices, corporate computer networks, intranets, and the Internet.
When used in a LAN environment, the computer 110 is connected to the LAN 171 through a network interface or adapter 170. When used in a WAN network environment, computer 110 typically includes a modem 172 or other means for establishing communications over the WAN 173, such as the Internet. Modem 172, which may be internal or external, may be connected to system bus 121 through user input interface 160, or other appropriate mechanism. In a network environment, the program modules depicted relative to the computer 110, or parts thereof, may be stored in the remote memory storage device. By way of example, and not limitation, Fig. 1 illustrates the remote application programs 185 residing in memory device 181. It will be appreciated that the network connections shown are exemplary and other means may be used to establish a network link. communications between computers.
ES 2 645 313 T3
Exemplary panoramic camera and client device
FIG. 2 is a block diagram depicting an exemplary panoramic camera apparatus 200 and an exemplary client device 222. Although shown in a particular configuration, it is noted that panoramic camera apparatus 200 can be any apparatus that includes a panoramic camera or its functional equivalent. More or fewer components than those shown included with the panoramic camera apparatus 200 can be included in a practical application that incorporates one or more of the techniques described herein.
Panoramic camera apparatus 200 includes processor 202 and memory 204. Panoramic camera apparatus 200 creates a panoramic image by stitching together several individual images produced by multiple cameras 206 (designated 206_1 through 206_n). The panoramic image can be a complete 360 ° panoramic image or it can be just a portion of it. It is noted that although a panoramic camera apparatus 200 is shown and described in this document, the disclosed techniques can also be used with a single camera.
Panoramic camera apparatus 200 also includes a microphone array 208. As will be described in greater detail below, the microphone array is configured so that the direction of the sound can be localized. In other words, the analysis of the sound input into the microphone array gives a direction from which a detected sound is produced. A speaker 210 may also be included in the panoramic camera apparatus 200 to enable a hands-free function or to broadcast notification signals and the like to users.
Memory 204 stores various camera settings 212 such as calibration data, exposure settings, junction tables, etc. An operating system 214 that controls the functions of the camera is also stored in memory 204 along with one or more camera software applications 216.
Panoramic camera apparatus 200 also includes an input / output (I / O) module 218 for transmitting data and receiving data to and from panoramic camera apparatus 200, and other additional hardware 220 that may be necessary for functionality. Of camera.
The panoramic camera apparatus 200 communicates with at least one client device 222, which includes a processor 224, a memory 226, a mass storage device 242 (such as a hard disk drive), and other hardware 230 that may be necessary to execute the functionality attributed to the client device 222 listed below.
Memory 226 stores a face tracking (FT) module 230 and a sound source locating (SSL) module 232. The face tracking module 230 and the sound source locating module 232 are used in conjunction with a virtual filmmaker 234 to detect a person in a camera scene and determine if and when the person is speaking. Any of a number of conventional sound source localization procedures can be used. Various face tracking procedures (or human detection and tracking systems) can be used, including that described in the main application, as described herein.
Memory 226 also stores a speaker grouping module 236 that is configured to determine a primary speaker when two or more people are speaking and focus a particular timeline portion on the primary speaker. In most meeting situations, there are cases where more than one person speaks at the same time. Typically, a main speaker is speaking when another person interrupts the speaker for a short period of time or speaks at the same time as the speaker. The speaker grouping module 236 is configured to group the speakers temporally and spatially to clear the timeline.
The virtual filmmaker 234 creates a timeline 238. The timeline 238 is stored in a timeline database 244 on the mass storage device 242. Timeline database 238 includes a plurality of fields including, but not necessarily limited to, time, speaker number, and bounding box within a camera image (x, y, width, height). The timeline database 238 may also include one or more angles of the speaker's face (azimuth and elevation).
A face extractor module 240 is also stored in memory 226 and is configured to extract an image of a speaker's face from a face bounding box (identified by face tracker 230) from a camera image. The face extractor module 240 stores the extracted facial images in a face database 246 on the mass storage device 242.
In at least one implementation, multiple face images can be stored for one or more speakers. Parameters can be specified to determine which facial image is used at which specific times. Or, a user can manually select a particular facial image from the multiple facial images.
In at least one alternative implementation, only a single facial image is stored for each speaker. The stored facial image may be a single image extracted by the face extractor module 240, but the
ES 2 645 313 T3 module 240 face extractor can also be configured to select the best Image of a speaker.
Selection of the best image of a speaker can be achieved by identifying frontal facial angles (assuming that an image with a frontal facial image is a better representation than an alternate image), identifying a facial image that exhibits minimal movement or Identifying a facial image that maximizes facial symmetry.
The recorded meeting 248 is also stored in the mass storage device 242 so that it can be retrieved and played back later.
The elements and functionality shown and described with respect to Fig. 2 will be described more fully below, with respect to subsequent figures.
Sample Playback Screen
FIG. 3 is a line drawing representation of a playback screen 300 including a panoramic image 302 and a facial image timeline 304. Panoramic image 302 is shown with a first meeting participant 303 and a second meeting participant 305. The playback screen 300 is also displayed with a title bar 306 and a single image 308. Individual image 308 is an optional feature that is focused on by a particular individual, typically a primary speaker. In Fig. 3, individual image 308 shows a facial image of the first meeting participant 303.
The exemplary playback screen 300 also includes a controls section 310 that contains the controls normally found on a media player, such as a play button, a fast forward button, a rewind button, and the like. An information area 312 is included in the playback screen 300, where information regarding the object of the playback screen 300 can be displayed. For example, a meeting title, a meeting room number, a list of meeting attendees, and the like can be displayed in the information area 312.
The facial image timeline 304 includes a first sub-timeline 314 that corresponds to the first meeting participant 303 and a second sub-timeline 316 that corresponds to the second meeting participant. Each sub-timeline 314, 316 indicates sections along a time continuum in which the corresponding meeting participant is speaking. A user can go directly to any point on a part-time line 314, 316 to immediately access a part of the meeting in which a particular meeting participant is speaking.
A first facial image 318 of the first meeting participant 303 appears adjacent to the first sub-timeline 314 to indicate that the first sub-timeline 314 is associated with the first meeting participant 318. A facial image 320 of the second meeting participant 305 appears adjacent to the second sub-timeline 316 to indicate that the second sub-timeline 316 is associated with the second meeting participant 305.
Fig. 4 shows an exemplary playback screen 400 that includes elements similar to the exemplary playback screen 300 shown and described in Fig. 3. Items and reference numerals shown and described with respect to Fig. 3 will be used with reference to the exemplary playback screen 400 of FIG. 4.
The exemplary playback screen 400 includes a panoramic image 302 and a facial image timeline 304. Panoramic image 302 shows a first meeting participant 303 and a second meeting participant 305. A title bar 306 runs across the top of the playback screen 400 and a single image 408 shows the second meeting participant 303.
The exemplary playback screen 400 also includes a whiteboard speaker image 402 showing a meeting participant (in this case, the second meeting participant 305) positioned in front of a whiteboard. The blackboard talker image 402 is not included in the playback screen 300 of FIG. 3 and is used here to show how other images can be included in any particular playback screen 300, 400.
A controls section 310 includes multimedia controls and an information area 312 presents information regarding the meeting displayed on the playback screen 400.
The facial image timeline 304 includes a first sub-timeline 314, a second sub-timeline 316, and a third sub-timeline 404. It is noted that, although only two sub-timelines are shown In Fig. 3, a timeline can contain any manageable number of sub timelines. In Fig. 4, for example, there are three sub-timelines.
It is noted that, although there are only two meeting participants in this example, there are three sub-timelines. This is because a single speaker can be associated with more than one single sub-timeline. In the present example, the second sub-timeline 316 is associated with the second meeting participant 305 while the second meeting participant 305 is on the board, and the third sub-timeline 404 is
ES 2 645 313 T3 associated with the second meeting participant 305 while the second meeting participant 305 is located at a different location on the whiteboard.
This situation can occur when a meeting participant occupies more than one location during a meeting. The virtual filmmaker 234 has detected speakers in three locations in this case. He does not necessarily know that only two speakers are present in those places. This feature helps a user in cases where the user is primarily interested in a speaker when the speaker is in a certain position. For example, a user may want to play back only the part of a recorded meeting in which a speaker is at the whiteboard.
The exemplary playback screen 400 also includes a first facial image 318 of the first meeting participant 303 and a second facial image 320 of the second meeting participant 305. In addition, a third facial image 406 is included that is associated with the third sub-timeline 404. The third facial image 406 corresponds to a second location of the second meeting participant 305.
The techniques used to display the exemplary replay screens 300, 400 will be described in greater detail below, with respect to the other figures.
Exemplary Methodological Implementation: Creation of the Facial Image Timeline
FIG. 5 is an exemplary flow chart 500 of a methodological implementation for creating a timeline with facial images. In the following description of the exemplary flow chart 500, continuous reference is made to the elements and reference numerals shown in the preceding figures.
At block 502, panoramic camera apparatus 200 samples one or more video images to create a panoramic image. The panoramic image is input to face tracker 230 (block 504) which detects and tracks faces in the image. Almost simultaneously, at block 506, the microphone array 208 samples the sound corresponding to the panoramic image and feeds the sound into the sound source locator 232, which detects speaker locations based on the sound sampled at block 508.
Virtual filmmaker 234 processes data from face tracker 230 and sound source locator 232 to create timeline 238, at block 510. At block 512, speaker grouping module 236 groups speakers temporally and spatially to consolidate and clarify portions of timeline 238 as described above.
The timeline is stored in the timeline database 244 with the following fields: time, speaker number, bounding box of the speaker in the image (x, y, width, height), facial angles of the speaker (azimuth , elevation) etc.
Using the panoramic image and face identification coordinates (ie, face bounding boxes) derived by face tracker 230, face extractor 240 extracts a facial image of the speakers at block 514. The extracted facial images they are stored in face database 246 and associated with a speaker number.
As noted above, face extractor 240 may be configured to extract more than one image for each speaker and use what face extractor 240 determines to be the best image on timeline 238.
An exemplary methodological implementation of selecting a better facial image and creating the face database 246 is shown and described below, with respect to FIG. 6.
Exemplary Methodological Implementation: Creation of a Database of Faces
FIG. 6 is an exemplary flow chart 600 depicting a methodological implementation for creating a database of faces. In the following description of Fig. 6 continued reference is made to elements and reference numerals shown in one or more previous figures.
At block 602, face extractor 240 extracts a facial image from the panoramic image as described above. If a facial image for the speaker is not already stored in the face database 246 (branch No, block 604), then the facial image is stored in the face database 246, at block 610. It is observed that determining if the facial image is stored does not necessarily depend on the person appearing in the facial image already having an image stored in their likeness, but rather on the identified speaker having an already stored image corresponding to the speaker. Therefore, if a speaker located in a first position has a stored facial image and then the speaker is detected in a second location, the facial image of the speaker in the second location will not be compared with the stored facial image of the speaker in the first location. position to determine if the speaker already has a stored facial image.
If a facial image of the speaker is already stored in the face database 246 - hereinafter, the stored facial image - (Yes branch, block 604), then in block 606 the facial image is compared with the
ES 2 645 313 T3 facial image stored. If the face extractor 240 determines that the facial image is better or more acceptable than the stored facial image (branch Yes, block 608), then the facial image is stored in the face database 246, thus overwriting the facial image. previously stored.
If the facial image is no better than the stored facial image (branch No, block 608), then the facial image 5 is discarded and the stored facial image is retained.
The criteria for determining which facial image is a better facial image can be numerous and varied. For example, the face extractor 234 may be configured to determine that a best facial image is one that is captured by a speaker in a position where the speaker's face is in a more frontal position. Or, if a first facial image shows signs of movement and a second facial image does not, then the face extractor 246 may determine that the second facial image is the best facial image. Or, face extractor 246 can be configured to determine which of multiple images of a speaker exhibits maximum symmetry and to use that facial image on the timeline. Other criteria not listed in this document can also be used to determine the most appropriate facial image to use with the timeline.
If there is another speaker (Yes branch, block 612), then the process returns to block 602 and repeats for each unique speaker. Again, single speaker, as used in this context, does not necessarily mean a single person, since a person who appears speaking in different locations can be interpreted as different speakers. The process ends when there are no more unique speakers to identify (branch No, block 612).
Conclution
Although one or more exemplary implementations have been illustrated and described, it will be appreciated that various changes may be made therein without departing from the scope of the claims appended hereto.
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
59 members in 11 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 978172 | United States of America | – | |
| 97817204 | United States of America | A |
Members59
| Document | Office | Kind | |
|---|---|---|---|
| US2003234866A1 | United States of America | A1 | |
| EP1377026A2 | European Patent Office (EPO) | A2 | |
| JP2004054928A | Japan | A | |
| EP1377026A3 | European Patent Office (EPO) | A3 | |
| US2004263611A1 | United States of America | A1 | |
| US2005046703A1 | United States of America | A1 | |
| US2005117015A1 | United States of America | A1 | |
| US2005117034A1 | United States of America | A1 | |
| US2005151837A1 | United States of America | A1 | |
| US2005285943A1 | United States of America | A1 | |
| CA2513760A1 | Canada | A1 | |
| CN1727991A | China | A | |
| MXPA05008090A | Mexico | A | |
| US2006023074A1 | United States of America | A1 | |
| US2006023075A1 | United States of America | A1 | |
| EP1624687A2 | European Patent Office (EPO) | A2 | |
| JP2006039564A | Japan | A | |
| AU2005203094A1 | Australia | A1 | |
| BRPI0503143A | Brazil | A | |
| CA2521670A1 | Canada | A1 | |
| MXPA05010595A | Mexico | A | |
| MXPA05010595A | Mexico | A | |
| AU2005220252A1 | Australia | A1 | |
| JP2006129480A | Japan | A | |
| KR20060049992A | Republic of Korea | A | |
| KR20060051672A | Republic of Korea | A | |
| EP1659518A2 | European Patent Office (EPO) | A2 | |
| CN1783998A | China | A | |
| BRPI0504224A | Brazil | A | |
| BRPI0504224A | Brazil | A | |
| EP1677534A1 | European Patent Office (EPO) | A1 | |
| KR20060079079A | Republic of Korea | A | |
| JP2006191535A | Japan | A | |
| CN1837952A | China | A | |
| US2006268131A1 | United States of America | A1 | |
| RU2005123981A | Russian Federation | A | |
| RU2005133403A | Russian Federation | A | |
| US7259784B2 | United States of America | B2 | |
| US7298392B2 | United States of America | B2 | |
| US7495694B2 | United States of America | B2 | |
| CN100524017C | China | C | |
| US7593042B2 | United States of America | B2 | |
| US7598975B2 | United States of America | B2 | |
| US7602412B2 | United States of America | B2 | |
| EP1659518A3 | European Patent Office (EPO) | A3 | |
| JP4421843B2 | Japan | B2 | |
| CN1783998B | China | B | |
| US7782357B2 | United States of America | B2 | |
| RU2398277C2 | Russian Federation | C2 | |
| CN1837952B | China | B | |
| US7936374B2 | United States of America | B2 | |
| JP4890005B2 | Japan | B2 | |
| JP5027400B2 | Japan | B2 | |
| KR101201107B1 | Republic of Korea | B1 | |
| KR101238586B1 | Republic of Korea | B1 | |
| CA2521670C | Canada | C | |
| EP1677534B1 | European Patent Office (EPO) | B1 | |
| EP1659518B1 | European Patent Office (EPO) | B1 | |
| ES2645313T3This record | Spain | T3 |
Numbers
- Publication
- 2645313
- Application
- 5109683
Titles2
- Spanish
- Extracción automática de rostros
- English
- Automatic face extraction
Classification
- CPC, 13
- H04N1/3876
- G06V40/173
- H04N5/93
- H04N1/6027
- H04N7/147
- H04N7/15
- H04N7/155
- H04N17/002
- H04N23/698
- H04N23/90
- H04N23/88
- G06T7/20
- G06T7/40
- IPC, 15
- G06K9 00
- H04N7 14
- G03B37 00
- G06N99 00
- G06T1 00
- G06T3 00
- G06T5 40
- G06T7 00
- H04N1 387
- H04N1 60
- H04N5 225
- H04N5 265
- H04N5 76
- H04N9 73
- H04N17 00