EP2584530A2

Method and device for identifying and extracting images of multiple users, and for recognizing user gestures

Abstract

The invention relates to a method for identifying and extracting images of one or more users in an interactive environment comprising the steps of : - obtaining a depth map (7) of a scene in the form of an array of depth values, and an image (8) of said scene in the form of a corresponding array of pixel values, said depth map (7) and said image (8) being registered; - applying a coordinate transformation to said depth map (7) and said image (8) for obtaining a corresponding array (15) containing the 3D positions in a real-world coordinates system and pixel values points; - grouping said points according to their relative positions, by using a clustering process (18) so that each group contains points that are in the same region of space and correspond to a user location (19); - defining individual volumes of interest (20) each corresponding to one of said user locations (19); - selecting, from said array (15) containing the 3D positions and pixel values, the points located in said volumes of interest for obtaining segmentation masks (35) for each user; - applying said segmentation masks (35) to said image (8) for extracting images of said users. The invention also relates to a method for recognizing gestures of said users

EP2584530A2, drawing sheet 1
Sheet 1 of 38

Term

Term ended

Projected expiry passed 3 August 2026, 0.1 years ago.

  1. Priority and filed
  2. Published
  3. Projected expiry
  4. Today

15 claims: 6 independent, 9 dependent

  1. 1
    A method for identifying and extracting images of one or more users in an interactive environment comprising the steps of :- obtaining a depth map (7) of a scene in the form of an array of depth values, and an image (8) of said scene in the form of a corresponding array of pixel values, said depth map (7) and said image (8) being registered;- applying a coordinate transformation to said depth map (7) and said image (8) for obtaining a corresponding array (15) containing the 3D positions in a real-world coordinates system and pixel values points;- grouping said points according to their relative positions, by using a clustering process (18) so that each group contains points that are in the same region of space and correspond to a user location (19);- defining individual volumes of interest (20) each corresponding to one of said user locations (19);- selecting, from said array (15) containing the 3D positions and pixel values, the points located in said volumes of interest for obtaining segmentation masks (35) for each user;- applying said segmentation masks (35) to said image (8) for extracting images of said users;characterized in that the method further comprises a background segmentation process (32) for separating foreground objects from background objects, said background segmentation process (32) comprising the steps of: - obtaining successive images (8) of said scene;- detecting motion of objects in said scene by subtracting two successive images (8) and performing a thresholding;- computing a colour mean array (25) containing for each pixel a mean value µ and a colour variance array (26) containing for each pixel a variance var from corresponding pixel of said successive images (8) of said scene;- determining pixels of said images that change between successive images;- masking the pixels that did not change for updating said colour mean array (25) and colour variance array (26);- determining a background model (24) from said colour mean array (25) and colour variance array (26);- deciding that a pixel is part of the background if its pixel value x meets the condition x - μ ⁢ 2 β * var where β is a variable threshold;- determining a background-based segmentation mask (30) by selecting all pixels in image (8) that do not meet above criteria.
  2. 6
    The method according to any of preceding claims, for recognizing a gesture of one of said users, comprising the steps of:- dividing the individual volume of interest (20) of said user in small elementary volumes of pre-determined size;- allocating each pixel of the mask (35) of said user to the elementary volume wherein it it located;- counting the number of pixels within each elementary volume, and selecting (36) the elementary volumes containing at least one pixel;- determining the centres of said selected elementary volumes (36), said centres being either the geometric centre or centre of gravity of the pixels within said elementary volumes;- performing a graph construction step (39), wherein each centre of said selected elementary volumes is a node, edges linking two nodes n, m are given a weight w(n,m) equal to the Euclidian distance between the centres of said nodes;- performing a centre estimation step (40), wherein an elementary volume is determined that is closest to the centre of gravity of all pixels is named "source cube" and the corresponding node in the graph is named the "source node";- computing (41) a curvilinear distance D(v) from the source node to all nodes v that are connected to it;- performing an extremity detection step (42) by checking for all nodes u, whether D (u) is greater or equal to D (v) for all nodes v connected to u, and if it is the case, marking node u as an extremity node;- performing a body part labelling step (43) for obtaining the locations (44) of body parts comprising head, hands, feet, by allocating body parts to extremity nodes using rules;- determining (45) a sequence of body parts locations, by performing the above steps on successive depth maps (7) and images (8) of the scene;- performing a tracking (46) of said body parts;- performing gesture recognition step (48) by using a hidden Markov model method.
  3. 8
    A device for identifying and extracting images of multiple users in an interactive environment scene comprising:- a video camera (4) for capturing an image from the scene;- a depth perception devices (3) for providing depth information about said scene;- at least one computer processor (2) for processing said depth information and said image information;characterized in that it comprises : - means (2) for using individual volumes of interest (20) from said scene for each user (6, 6');- means (2) for obtaining adaptive volumes of interest for each user;- means for dividing the individual volume of interest (20) of said user in small elementary volumes of pre-determined size;- means for allocating each pixel of the mask (35) of said user to the elementary volume wherein it is located;- means for counting the number of pixels within each elementary volume, and selecting (36) the elementary volumes containing at least one pixel;- means for determining the centres of said selected elementary volumes (36), said centres being either the geometric centre or centre of gravity of the pixels within said elementary volumes;- means for performing a graph construction step (39), wherein each centre of said selected elementary volumes is a node, edges linking two nodes n, m are given a weight w(n,m) equal to the Euclidian distance between the centres of said nodes;- means for performing a centre estimation step (40), wherein an elementary volume is determined that is closest to the centre of gravity of all pixels is named "source cube" and the corresponding node in the graph is named the "source node";- means for computing (41) a curvilinear distance D(v) from the source node to all nodes v that are connected to it;- means for performing an extremity detection step (42) by checking for all nodes u, whether D (u) is greater or equal to D (v) for all nodes v connected to u, and if it is the case, marking node u as an extremity node;- means for performing a body part labelling step (43) for obtaining the locations (44) of body parts comprising head, hands, feet, by allocating body parts to extremity nodes using rules;- means for determining (45) a sequence of body parts locations, by performing the above steps on successive depth maps (7) and images (8) of the scene;- means for performing a tracking (46) of said body parts;- means for performing gesture recognition step (48) by using a hidden Markov model method.
  4. 9
    Use of a method and/or of a device for identifying and segmenting one or more users in an interactive environment for providing an interactive virtual reality game wherein a plurality of users are separated from each other and from a background.
  5. 12
    Use according to any of claims 9 to 11 wherein one determines the location of a user and one uses said location for directing one or more light beams.
  6. 14
    Use according to any of claims 9 to 13 wherein users are directed to point at each other or to give each other virtual objects.