Visual target tracking
Abstract
A method of tracking a target includes receiving an observed depth image of the target from a source and obtaining a posed model of the target. The model is rasterized into a synthesized depth image, and the pose of the model is adjusted based, at least in part, on differences between the observed depth image and the synthesized depth image.

Term
3.3 yearsleft in the term
Expires 12 January 2030.
- Priority
- Filed
- Granted
- Today
- Expires
15 claims: 2 independent, 13 dependent
- 11/2 1/2 CLAIMS REIVINDICAÇÕES 1. Method for tracking a target, CHARACTERIZED by comprising:1. Método para rastrear um alvo, CARACTERIZADO por compreender: receber (102) uma imagem com profundidade observada (140) do alvo (18) a partir de uma fonte (20);receiving (102) an observed depth image (140) of the target (18) from a source (20);obtaining (112) a model (70) of the target, the model having a pose;obter (112) um modelo (70) do alvo, o modelo tendo uma pose;rasterizar (116) o modelo em uma imagem com profundidade sintetizada (150);rasterize (116) the model into an image with synthesized depth (150);adjust (142) the pose of the model based, at least in part, on the differences between the image with observed depth and the image with synthesized depth. ajustar (142) a pose do modelo baseando-se, ao menos em parte, nas diferenças entre a imagem com profundidade observada e a imagem com profundidade sintetizada.
- 15Computing system, CHARACTERIZED by comprising:15. Sistema de computação, CARACTERIZADO por compreender: a source (48) configured to capture depth information;uma fonte (48) configurada para capturar informações de profundidade;a logic subsystem (42) operatively connected to the source;and a data storage subsystem (44) that stores instructions executable by the logic subsystem for: um subsistema de lógica (42) operativamente conectado à fonte;e um subsistema de armazenamento de dados (44) que armazena instruções executáveis pelo subsistema de lógica para: receber (102) uma imagem com profundidade observada de um alvo a partir da fonte;receiving (102) an observed depth image of a target from the source;obtaining (112) a model of the target, the model having a pose;obter (112) um modelo do alvo, o modelo tendo uma pose;rasterizar (116) o modelo em uma imagem com profundidade sintetizada;rasterize (116) the model into an image with synthesized depth;adjust (142) the pose of the model based, at least in part, on the differences between the image with observed depth and the image with synthesized depth. ajustar (142) a pose do modelo baseando-se, ao menos em parte, nas diferenças entre a imagem com profundidade observada e a imagem com profundidade sintetizada.
Independent claims2
199 paragraphs in 4 sections, as filed
1/32 “VISUAL TARGET TRACKING”
BACKGROUND OF THE INVENTION
Many computer games and other computer vision applications use complicated controls to allow users to manipulate game characters or other aspects of an application. Such controls can be difficult to learn, thus creating a barrier to getting started in many games or applications. Furthermore, such controls can be quite different from the actual actions in the game or the actions of the applications for which they are used. For example, a game controller that causes a game character to swing a baseball bat may bear no resemblance to the actual swinging motion of a baseball bat.
SUMMARY
The intention of this summary is to present, in a simplified way, a selection of concepts that are described in detail below in the Detailed Description. This Summary is not intended to identify crucial or essential aspects of the claimed matter and should not be used to limit the scope of the claimed matter. Furthermore, the subject matter claimed is not limited to implementations that resolve some or all of the drawbacks mentioned elsewhere in this disclosure.
Various embodiments related to visual target tracking are discussed in the present invention. A disclosed embodiment includes tracking a target by receiving an image with observed depth of the target from a source and obtaining a pose model of the target. The pose model is rasterized into an image with synthesized depth. The model pose is then adjusted based, at least in part, on the differences between the observed depth image and the synthesized depth image.
BRIEF DESCRIPTION OF THE DRAWINGS
FIG. 1A shows an embodiment of an exemplary target recognition, analysis and tracking system that tracks a player playing a boxing game.
FIG. 1B shows the player of FIG. 1A throwing a punch, which is tracked and interpreted as a game control that causes the player's avatar to throw a game in the game's 30 space.
FIG. 2 schematically illustrates a computing system in accordance with one embodiment of the present disclosure.
FIG. 3 shows an exemplary body model used to represent a human target.
FIG. 4 illustrates a substantially front view of an exemplary skeletal model used to represent a human target.
FIG. 5 illustrates an asymmetrical view of an exemplary skeletal model
2/32 used to represent a human target.
FIG. 6 shows an exemplary mesh model used to represent a human target.
FIG. 7 shows a flowchart of an exemplary method of visually tracking a target.
FIG. 8 shows an exemplary observed depth image.
FIG. 9 shows an exemplary synthesized depth image.
FIG. 10 schematically illustrates some of the pixels that make up an image with synthesized depth.
FIG. 11A schematically illustrates the application of a force to a force receiving location of a model.
FIG. 11B schematically illustrates a result of applying force to the force receiving location of the FIG. 11 a.
FIG. 12A shows a player avatar rendered from the model of FIG. 11 a.
FIG. 12B shows a player avatar rendered from the model of FIG. 11B.
DETAILED DESCRIPTION
The present revelation is concerned with the recognition, analysis and tracking of a target. In particular, the use of a depth-sensing camera or other source to obtain depth information for one or more targets is disclosed. Such depth information can then be used to effectively and accurately model and track the one or more targets, as described in detail below. The target recognition, analysis, and tracking described here provide a robust platform where one or more targets can be consistently tracked at a relatively fast frame rate, even when the target(s) move into poses considered difficult to analyze. using other approaches (eg when two or more targets partially overlap and/or block each other; when one part of a target blocks another part of the same target, when a target changes its topographical appearance (eg, a person touching its head), etc.).
FIG. 1A shows a non-limiting example of a target recognition, analysis and tracking system 10. In particular, FIG. 1A shows a computer game system 12 that can be used to run a variety of different games, play one or more different types of media, and/or control or manipulate applications other than games. FIG. 1A also shows a display medium 14 in the form of a high definition television, or HDTV 16, which can be used to present visual graphics of the game to players, such as player 18. Furthermore, FIG. 1A shows a device for
3/32 captures in the form of a depth-sensing camera 20, which can be used to visually monitor one or more players, such as player 18. The example illustrated in FIG. 1A is not limiting. (As described below with reference to FIG. 2, it is possible to use a variety of different types of target recognition, analysis and tracking systems without departing from the scope of the present disclosure.
A target recognition, analysis and tracking system can be used to recognize, analyze and/or track one or more targets, such as player 18. FIG. 1A shows a scenario where player 18 is tracked using depth sensing camera 20 so that player 18's movements can be interpreted by game system 12 as controls that can be used to affect the game being played by game system 12. game 12. In other words, player 18 can use his moves to control the game. Player 18's movements can be interpreted pretty much like any kind of game controller.
The example scenario illustrated in FIG. 1A shows player 18 playing a boxing game being played by game system 12. The game system uses HDTV 16 to visually present a boxing opponent 22 to player 18. In addition, the game system uses HDTV 16 to visually present an avatar of player 25 that player 18 controls with his movements. As shown in FIG. 1B, player 18 may punch the physical space as an instruction to the player avatar 24 to punch the game space. The game system 12 and the depth-sensing camera 20 can be used to recognize and analyze player 18's punch in physical space, so that the punch can be interpreted as a game control that makes the player's avatar 24 punch the game space. For example, FIG. 1B shows HDTV 16 visually displaying the avatar of player 24 throwing a punch that strikes boxing opponent 22 in response to the movement of player 18, who punches in physical space. Other player 18 moves can be interpreted as other controls, such as controls to swing, lean to one side, move quickly, block, jab, or deliver a variety of punches with different strengths. Furthermore, some moves can be interpreted in controls that serve other purposes than controlling the player's avatar 24. For example, the player can use moves to end, pause or save a game, choose a level, view high scores, communicate with a friend, etc.
In some embodiments, a target may include a human and an object. In such embodiments, for example, a player of an electronic game may be holding an object, such that the movements of the player and the object are used to adjust and/or control parameters of the electronic game. For example, the movement of a player holding a racket can be tracked and used to control a racket displayed on the screen in an electronic sports game. In another example, the movement of a player holding an object can be tracked and used to control a weapon displayed on the screen in an electronic combat game.
Target recognition, analysis and tracking systems can be used to interpret target movements as operating system and/or application controls that are outside the scope of the game. Virtually any controllable aspect of an operating system and/or application, such as the boxing game illustrated in FIGs. 1A and 1B, can be controlled by the movements of a target, such as player 18. The boxing scenario illustrated is presented as an example, but is not intended to be limiting in any way. Rather, the illustrated scenario is intended to demonstrate a general concept, which can be applied to a variety of different applications without departing from the scope of the present disclosure.
The methods and processes described in the present invention can be linked to a variety of different computer system types. FIGs. 1A and 1B show a non-limiting example in the form of gaming system 12, HDTV 16 and depth-sensing camera 20. As another more general example, FIG. 2 schematically illustrates a computing system 40 that can perform one or more of the target recognition, tracking, and analysis methods and processes described in the present invention. Computer system 40 can take a variety of different forms, including, but not limited to, game consoles, personal computer game systems, military tracking and/or targeting systems, and character acquisition systems that provide green screen” or motion capture, among others.
Computing system 40 may include a logic subsystem 42, a data storage subsystem 44, a display subsystem 46, and/or a capture device 48. The computing system may optionally include components not illustrated in FIG. 2 and/or some components shown in FIG. 2 may be peripheral components that are not integrated into the computing system.
Logical subsystem 42 may include one or more physical devices configured to execute one or more instructions. For example, the logical subsystem can be configured to execute one or more instructions that are part of one or more programs, routines, objects, components, data structures, or other logical constructs. Such instructions can be implemented to perform a task, implement a data type, transform the state of one or more devices, or else arrive at the desired result. The logical subsystem may include one or more processors that are configured to execute software instructions. In addition, or alternatively, the logic subsystem may include one or more logic engines in hardware or firmware configured to execute hardware or firmware instructions. The logical subsystem may optionally include individual components that are distributed across two or more devices, which may be
5/32 remotely located in some embodiments.
Data storage subsystem 44 may include one or more physical devices configured to store data and/or instructions executable by the logical subsystem to implement the methods and processes described herein. When such methods and processes are implemented, the state of the data storage subsystem 44 can be transformed (eg, to store different data). Data storage subsystem 44 may include removable media and/or embedded devices. Data storage subsystem 44 may include optical memory devices, semiconductor memory devices (eg, RAM, EEPROM, flash, etc.) and/or magnetic memory devices, among others. Data storage subsystem 44 may include devices having one or more of the following characteristics: volatile, non-volatile, dynamic, static, read/write, read-only, random access, sequential access, location-addressable, file-addressable, and content addressable. In some embodiments, logic subsystem 42 and data storage subsystem 44 may be integrated into one or more common devices, such as an application-specific integrated circuit or a system on a chip.
FIG. 2 also shows an aspect of the data storage subsystem in the form of removable computer readable media 50 that can be used to store and/or transfer data and/or executable instructions to implement the methods and processes described herein. Display subsystem 46 may be used to present a visual representation of data stored by data storage subsystem 44. As the methods and processes described herein change the data stored by the data storage subsystem, and therefore transform the state of the data storage subsystem, the state of the display subsystem 46 can be similarly transformed to visually represent changes in the data. underlying. As a non-limiting example, the target recognition, tracking, and analysis described herein may be reflected through the display subsystem 46 in the form of a game character that changes pose in game space in response to a player's movements in space. physicist. Display subsystem 46 may include one or more display devices utilizing virtually any type of technology. Such display devices may be combined with logic subsystem 42 and/or data storage subsystem 44 in a shared compartment, or such display devices may be peripheral display devices, as shown in FIGs. IA and 1B.
Computing system 40 additionally includes a capture device 48 configured to acquire depth images of one or more targets. The capture device 48 can be configured to capture video with depth information using any suitable technique (e.g. time of flight, structured light, image
6/32 stereo, etc.). As such, the capture device 48 may include a depth-sensing camera, a video camera, stereo cameras, and/or other suitable capture devices.
For example, in time-of-flight analysis, the capture device 48 may emit infrared light to the target and may then use sensors to detect backscattered light from the surface of the target. In some cases, pulsed infrared light can be used, where the time between an outgoing light pulse and a corresponding incoming light pulse can be measured and used to determine a physical distance from the capture device to a particular location in the target. In some cases, the phase of the outgoing light wave can be compared with the phase of the incoming light wave to determine a phase shift, and the phase shift can be used to determine the physical distance from the capture device to a specific location on the target.
In another example, time-of-flight analysis can be used to indirectly determine the physical distance of the capture device to a specific location on the target by analyzing the intensity of the reflected light beam over time, using a technique, such as obturated light pulse image.
In another example, structured light analysis may be used by capture device 48 to capture depth information. In such an analysis, a patterned light (that is, light displayed in a known pattern, such as a grid pattern or band pattern) can be projected onto the target. Upon reaching the surface of the target, the pattern can become deformed in response, and this deformation of the pattern can be studied to determine a physical distance from the capture device to a specific location on the target.
In another example, the capture device may include two or more physically separate cameras that view a target from different angles to obtain stereo visual data. In such cases, stereo visual data can be resolved to generate an image with depth.
In other embodiments, capture device 48 may use other technologies to measure and/or calculate depth values. In addition, the capture device 48 can organize the calculated depth information into "Z layers", that is, layers perpendicular to a Z axis extending from the depth sensing camera along its line of sight to the observer.
In some embodiments, two or more different cameras can be incorporated into an integrated capture device. For example, a depth camera and a video camera (eg RGB video camera) can be built into a common capture device. In some embodiments, two or more separate capture devices may be used cooperatively. For example, a camera with
7/32 depth sensor and a separate video camera can be used. When a video camera is used, it can be used to provide target tracking data, confirmation data for target tracking error correction, image capture, facial recognition, high precision finger tracking (or other small details ), light detection and/or other functions.
It should be understood that at least some target tracking and analysis operations can be performed by a logic machine of one or more capture devices. A capture device may include one or more integrated processing units configured to perform one or more target analysis and/or tracking functions. A capture device may include firmware to facilitate the upgrade of such built-in processing logic.
Computing system 40 may optionally include one or more input devices, such as controller 52 and controller 54. The input devices may be used to control the operation of the computing system. In the context of a game, input devices such as controller 52 and/or controller 54 can be used to control uncontrolled aspects of a game through the described target recognition, tracking and analysis methods and procedures. in the present invention. In some embodiments, input devices, such as controller 52 and/or controller 54, may include one or more accelerometers, gyroscopes, infrared target/sensor systems, etc. controllers in physical space. In some embodiments, the computing system may optionally include and/or utilize input gloves, keyboards, mice, trackpads, trackballs, touch screens, buttons, switches, disks, and/or other input devices. As will be appreciated, target recognition, tracking and analysis can be used to control or enhance aspects of a game, or other application, conventionally controlled by an input device, such as a game controller. In some embodiments, the target tracking described in the present invention may be used as a complete replacement for other forms of user input, whereas in other embodiments, such target tracking may be used to complement one or more other forms of user input. of user.
Computing system 40 may be configured to perform the target tracking methods described in the present invention. However, it should be understood that the computing system 40 is presented as a non-restrictive example of a device capable of performing such target tracking. Other devices are within the scope of this disclosure.
Computing system 40, or other suitable device, can be configured to represent each target with a template. As described in more detail below, the
8/32 information derived from such a model can be compared with information obtained from a capture device, such as a depth-sensing camera, so that the fundamental proportions or shape of the model, as well as its current pose, can be adjusted to more accurately represent the modeled target. The model may be represented by one or more polygonal meshes, by a set of mathematical primitives, and/or by other suitable machine representations of the model target.
FIG. 3 shows a non-limiting visual representation of an exemplary body model 70. The body model 70 is a machine representation of a modeled target (e.g., the player 18 of FIGs. 1A and 1B). The body model can include one or 10 more data structures that include a set of variables that collectively define the target modeled in the language of a game or other application/operating system.
A model of a target can be configured in a variety of ways without departing from the scope of the present disclosure. In some examples, a model may include one or more data structures that represent a target as a three-dimensional model comprising rigid and/or deformable shapes, or body parts. Each body part can be characterized as a mathematical primitive, examples of which include, but are not limited to, spheres, anisotropically sized spheres, cylinders, anisotropic cylinders, smooth cylinders, boxes, oblique boxes, prisms, among others.
For example, the body model 70 of FIG. 3 includes body parts bpl through bp14, each of which represents a different part of the modeled target. Each body part is a three-dimensional shape. For example, bp3 is a rectangular prism that represents the left hand of a modeled target, and bp5 is an octagonal prism that represents the left arm of the modeled target. Body model 70 is exemplary, as a body model may contain any number of body parts, each of which may be any machine-understandable representation of the corresponding part of the modeled target.
A model including two or more body parts may also include one or more joints. Each joint can allow one or more body parts to move relative to one or more other body parts. For example, a model representing a human target may include a plurality of rigid and/or deformable body parts, where some body parts may represent a corresponding anatomical body part of the human target. In addition, each body part in the model may comprise one or more structural members (ie bones'j, with joints located at the intersection of adjacent bones. It should be understood that some bones may correspond to anatomical bones in a human target and /or some bones may not have corresponding anatomical bones in the human target.
Bones and joints can collectively constitute a skeletal model,
9/32 which can be a constituent element of the model. The skeletal model can include one or more skeletal members for each body part and a joint between adjacent skeletal members. Exemplary skeletal model 80 and exemplary skeletal model 82 are illustrated in FIGs. 4 and 5, respectively. FIG. 4 shows a skeletal model 80 seen from the front, with joints j1 through j33. FIG. 5 shows a skeletal model 82 seen from an asymmetrical view, also with joints j1 to j33. Skeleton model 83 additionally includes rotation joints j34 through j47, where each rotation joint can be used to track axial angles of rotation. For example, an axial rotation angle can be used to define a rotational orientation of a body limb relative to its parent limb and/or the bust. For example, if a skeletal model is illustrating an axial rotation of an arm, the j40 rotation joint can be used to indicate the direction the associated wrist is pointing (eg, palm facing up). Therefore, while joints can receive forces and adjust the skeletal model as described below, rotation joints can instead be constructed and used to track axial rotation angles. More generally, by analyzing the orientation of a member in relation to its parent member and/or the bust, it is possible to determine an axial rotation angle. For example, if looking at the lower leg, the orientation of the lower leg in relation to the upper leg and hips can be examined to determine an axial rotation angle.
As described above, some models may include a skeleton and/or body parts that serve as a machine representation of a modeled target. In some embodiments, a model may alternatively or additionally include a wire skeleton mesh, which may include hierarchies of rigid polygonal meshes, one or more deformable meshes, or any combination of the two. As a non-limiting example, FIG. 6 shows a model 90 including a plurality of triangles (eg, triangle 92) arranged in a mesh that defines the shape of the body model. Such a mesh may include bending limits at each polygonal edge. When such a mesh is used, the number of triangles, and/or other polygons, which collectively constitute the mesh, can be selected to achieve a desired balance between quality and computational cost. More triangles can provide higher quality and/or more accurate models, while fewer triangles may require less computational resources. A body model including a polygonal mesh does not need to include a skeleton, although it may in some embodiments. The body part models described above and polygonal meshes are exemplary non-limiting types of models that can be used as machine representations of a modeled target. Other models are also within the scope of this revelation. For example, some models may include patches, B-splines
10/32 non-uniform rationals, subdivision surfaces, or other higher-order surfaces. A model may also include surface textures and/or other information to more accurately represent the clothing, hair, and/or other aspects of a modeled target. A model can optionally include information pertaining to the current pose, one or more previous poses, and/or the physical properties of the model. It must be understood that any model capable of striking a pose and then being rasterized (or otherwise rendered or expressed by) a depth-synthesized image is compatible with the target recognition, analysis and tracking described in the present invention.
As mentioned above, a model serves as a representation of a target, such as player 18 in FIGs. 1A and 1B. As the target moves in physical space, information from a capture device, such as a depth-sensing camera 20 in FIGs. 1A and IB, can be used to adjust the pose and/or fundamental size/shape of the model so that it more accurately represents the target. In particular, one or more forces can be applied to one or more force receiving aspects of the model to adjust the model to a pose that most closely matches the pose of the target in physical space. Depending on the type of model being used, force can be applied to a joint, a centroid of a body part, a vertex of a triangle, or any other suitable force-receiving aspect of the model. Also, in some embodiments, two or more different calculations may be used when determining the direction and/or magnitude of the force. As described in more detail below, the differences between an observed image of the target, as retrieved by a capture device, and a rasterized (i.e., synthesized) image of the model can be used to determine the forces that are applied to the model in order to to adjust the body to a different pose.
FIG. 7 shows a flowchart of an exemplary method 100 of tracking a target using a model (e.g., the body model 70 of FIG. 3). In some embodiments, the target may be a human, and the human may be one of two or more targets being tracked. As such, in some embodiments, method 100 may be performed by a computing system (e.g., gaming system 12 illustrated in FIG. 1 and/or computing system 40 illustrated in FIG. 2) to track one or more players interacting with an electronic game being played on the computing system. As presented above, player tracking allows the physical movements of these players to act as a real-time user interface that adjusts and/or controls electronic game parameters. For example, a player's tracked movements can be used to move a character or avatar around the screen in a role-playing game. In another example, a player's tracked movements can be used to control an on-screen vehicle in an electronic racing game. In yet another example, a player's tracked movements can be used to control the building or
11/32 organization of objects in a virtual environment.
At 102, method 100 includes receiving an observed depth image of the target from a source. In some embodiments, the source may be a depth-sensing camera configured to obtain detailed information about the target through a suitable technique, such as time-of-flight analysis, structured light analysis, stereo vision analysis, or other techniques. suitable. The observed depth image may include a plurality of observed pixels, where each observed pixel has an observed depth value. The observed depth value includes target depth information as seen from the source. FIG. 8 shows a visual representation of an exemplary observed depth image 140. As illustrated, the observed depth image 140 captures an exemplary observed pose of a person (e.g., player 18) standing with arms raised.
As illustrated at 104 of FIG. 7 , upon receiving the observed depth image, method 100 may optionally include downscaling the observed depth image to a lower processing resolution. Downsizing to a lower processing resolution can allow the image with observed depth to be used more easily and/or processed faster with less computational overhead.
As illustrated at 106, upon receiving the observed depth image, method 100 may optionally include removing background elements (other than the player) from the observed depth image. Removal of such background elements may include separating various image regions with observed depth into background regions and regions occupied by the target image. Background regions can be removed from the image or labeled so that they can be ignored during one or more subsequent processing steps. Virtually any background removal technique can be used, and information from the tracking (and previous frame) can optionally be used to assist and improve the quality of background removal.
As illustrated at 108, upon receiving the observed depth image, method 100 may optionally include removing and/or smoothing one or more high variance and/or noisy depth values from the observed depth image. Such high variance and/or noisy depth values in the observed depth image can come from a number of different sources, such as random and/or systematic errors that occur during the image capture process, defects and/or aberrations resulting from the capture device, etc. Since such high variance and/or noisy depth values may be artifacts of the image capture process, including these values in any future image analysis may skew the results and/or cause calculations to be slow. Thus, removing such values may provide12/32 better data integrity for future calculations.
Other depth values can also be filtered. For example, the accuracy of the enlargement operations described below with reference to step 118 can be improved by selectively removing pixels that satisfy one or more removal criteria. For example, if a depth value is midway between the hand and the bust that the hand is covering, removing this pixel can prevent magnification operations from blending one body part with another during subsequent processing steps.
As illustrated at 110, method 100 may optionally include filling and/or rebuilding parts whose depth information is missing or removed. Such filling can be performed by averaging the nearest neighbors, by filtering and/or by any other suitable method.
As illustrated at 112 of FIG. 7 , method 100 may include obtaining a model (e.g., body model 70 of FIG. 3). As described above, the model may include one or more polygonal meshes, one or more mathematical primitives, one or more higher-order surfaces, and/or other features used to provide a machine representation of the target. Furthermore, the model can exist as an instance of one or more existing data structures in a computing system.
In some embodiments of method 100, the model may be a pose model obtained from a previous step. For example, if method 100 is performed continuously, a model pose resulting from a previous iteration of method 100, corresponding to a step at an earlier time, can be obtained.
In some embodiments, a pose can be determined by one or more algorithms, which can analyze an image in depth and identify, at a general level, where the target(s) of interest (e.g., human(s)) are located and/or the pose of such target(s). Algorithms can be used to select a pose during an initial iteration or whenever it is believed that the algorithm can select a more accurate pose than the pose calculated during a step taken at an earlier time.
In some embodiments, the model can be obtained from a database and/or other program. For example, a model may not be available during a first iteration of method 100, in which case the model may be obtained from a database including one or more models. In such a model, one can choose from the database using a search algorithm designed to select a model exhibiting a pose similar to that of the target. Even if a previously performed one-step model is available, a model from a database can be used. For example, a model from a database can be used after a certain number of frames, if the target has changed poses by more than a predetermined threshold, and/or accordingly.
13/32 with other criteria.
In other embodiments, the model, or parts thereof, may be synthesized. For example, if the core of the target's body (bust, midsection, and hips) is represented by a deformable polygonal model, that model can be originally constructed using the contents of an image with observed depth, where the contour of the target in the image (i.e. the silhouette) can be used to shape the mesh in X and Y dimensions. Furthermore, in such an approach, the observed depth value(s) in the image area with observed depth can be used to shape the mesh in the XY direction as well as the Z direction. , of the model to more favorably represent the target's body shape.
Method 100 may additionally include depicting any clothing that appears on the target using a suitable approach. Such a suitable approach might include adding auxiliary geometry to the model in the form of polygonal meshes or primitives, and optionally adjusting auxiliary geometry based on poses to reflect gravity, simulating clothing, etc. Such an approach can facilitate molding models into more realistic representations of targets.
As illustrated at 114, method 100 may optionally comprise applying a momentum algorithm to the model. Since the timing of the various parts of a target can predict changes in a sequence of images, such an algorithm can help to obtain the pose of the model. The momentum algorithm can use a trajectory of each of the articulations or vertices of a model along a fixed number of a plurality of previous frames to assist in obtaining the model.
In some embodiments, the knowledge that different parts of a target can move a limited distance in a time interval (e.g. 1/30 or 1/60<sup>2</sup> of one second) can be used as a constraint on obtaining a model. Such a constraint can be used to rule out certain poses when a previous frame is known.
At 116 of FIG. 7 , method 100 can also include rasterizing the model into an image with synthesized depth. Rasterization allows the model described by mathematical primitives, polygonal meshes, or other objects to be converted into an image with synthesized depth described by a plurality of pixels.
Rasterization can be performed using one or more different techniques and/or algorithms. For example, model rasterization may include projecting a representation of the model onto a two-dimensional plane. In the case of a model including a plurality of body part shapes (e.g., body model 70 of FIG. 3), rasterization may include projecting and rasterizing the set of body part shapes in a two-dimensional plane. For each pixel in the two-dimensional plane on which the model is projected, several different types of information can be stored.
FIG. 9 shows a visual representation 150 of an exemplary synthesized depth image corresponding to the body model 70 of FIG. 3. FIG. 10 shows a pixel array 160 of a portion of the same image with synthesized depth. As indicated at 170, each synthesized pixel in the synthesized depth image can include a synthesized depth value. The synthesized depth value for a given synthesized pixel can be the depth value of the corresponding part of the model that is represented by this synthesized pixel, as determined during rasterization. In other words, if a part of a forearm body part (e.g., forearm body part bp4 of FIG. 3) is projected onto a two-dimensional plane, a corresponding synthesized pixel (e.g., synthesized pixel 162 of FIG. 10) can receive a synthesized depth value (e.g., the synthesized depth value 164 of FIG. 10) equal to the depth value of that forearm body part. In the illustrated example, the synthesized pixel 162 has a synthesized depth value of 382 cm. Similarly, if an adjacent hand body part (e.g., the bp3 hand body part of FIG. 3) is projected onto a two-dimensional plane, a corresponding synthesized pixel (e.g., synthesized pixel 166 of Fig. 10) may receive a synthesized depth value (e.g., synthesized depth value 168 of Fig. 10) equal to the depth value of that part of the hand body part. In the illustrated example, the synthesized pixel 166 has a synthesized depth value of 383 cm. It should be understood that the previous segment is presented as an example. The synthesized depth values can be saved in any unit of measure or as a dimensionless number.
As indicated at 170, each synthesized pixel in the synthesized depth image may include an original body part index determined during rasterization. Such an original body part index can indicate which of the model's body parts this pixel corresponds to. In the illustrated example of FIG. 10, the synthesized pixel 162 has an original body part index of bp4, and the synthesized pixel 166 has an original body part index of bp3. In some embodiments, the original body part index of a synthesized pixel may be null if the synthesized pixel does not correspond to a target body part (e.g., if the synthesized pixel is a background pixel). In some embodiments, synthesized pixels that do not correspond to a body part may be assigned a different index type.
As indicated at 170, each synthesized pixel in the synthesized depth image may include an original player index determined during rasterization, the original player index corresponding to the target. For example, if there are two targets, the synthesized pixels corresponding to the first target will have a first player index and the synthesized pi15/32 xels corresponding to the second target will have a second player index. In the illustrated example, the pixel array 160 corresponds to only one target, therefore, the synthesized pixel 162 has an original player index of P1 and the synthesized pixel 166 has a original player index of P1. Other types of indexing systems may be used without departing from the scope of the present disclosure.
As indicated at 170, each synthesized pixel in the depth synthesized image can include a pixel address. The pixel address can define the position of a pixel in relation to other pixels. In the illustrated example, the synthesized pixels 162 have a pixel address of [5,7], and the synthesized pixel 166 has an address of [4,8]. scope of this revelation.
As indicated at 170, each synthesized pixel may optionally include other types of information, some of which may be obtained after rasterization. For example, each synthesized pixel can include an updated body part index, which can be determined as part of an association operation performed during rasterization, as described below. Each synthesized pixel can include an updated player index, which can be determined as part of an association operation performed during rasterization. Each synthesized pixel can include an updated body part index, which can be obtained as part of a zoom/fit operation, as described below. Each synthesized pixel can include an updated player index, which can be obtained as part of a zoom/adjust operation as described above.
The exemplary types of pixel information provided above are not limiting. Several different types of information can be stored as part of each pixel. Such information may be stored as part of a common data structure, or the different types of information may be stored in different data structures that may be mapped to particular pixel locations (eg via a pixel address). As an example, player indices and/or body part indices obtained as part of an association operation during rasterization can be stored in a raster map and/or association map, whereas player indices and/or or the body part indices obtained as part of a zooming/fitting operation after rasterization can be stored in a zoom map as described below. Non-limiting examples of other types of pixel information that can be assigned to each pixel include, but are not limited to, joint indices, bone indices, vertex indices, triangle indices, centroid indices, among others.
At 118, the method 100 of FIG. 7 may optionally include association and/or amplification of body part indices and/or player indices. In other words, the ima
16/32 gem with synthesized depth can be increased so that the body part index and/or player index of a few pixels are altered in an attempt to more accurately match the modeled target.
By performing the rasterizations described above, one or more Z-Buffers and/or body part/player index maps can be constructed. As a non-limiting example, a first version of a buffer! map can be constructed by performing a Z test where a surface closest to the viewer (e.g. depth sensing camera) is selected and a body part index and/or player index associated with that surface is recorded in the corresponding pixel. This map can be called a raster map or an original synthesized depth map. A second version of such a buffer! map can be constructed by performing a Z test in which a surface that is closest to an observed depth value in that pixel is selected and a body part index and/or player index associated with that surface is recorded in the corresponding pixel . This can be called an association map. Such tests can be constrained to reject a distance Z between a synthesized depth value and an observed depth value that is beyond a predetermined threshold. In some embodiments, two or more Z-buffers and/or two or more body part/player index maps may be maintained, thus allowing two or more of the tests described above to be performed.
A third version of a buffer/map can be constructed by enlarging and/or correcting a body part/player index map. This can be called a magnification map. Starting with a copy of the association map described above, the values can be increased relative to any “unknown” values within a predetermined Z distance, so that a space being occupied by the target but not yet occupied by the body model can be filled with appropriate body part/player indices. Such an approach may also include exceeding a known value if a more favorable correlation is identified.
The zoom map can start with a pass through the synthesized pixels of the association map to detect pixels containing adjacent pixels with a different body part/player index. These can be considered edge pixels, that is, boundaries along which values can optionally be propagated. As introduced above, scaling up pixel values can include scaling for both unknown and known pixels. For unknown pixels, the body part/player index value, for example in a scenario, may have been zero before, but now may have a non-zero adjacent pixel. In such a case, the four direct adjacent pixels can be examined, and the adjacent pixel with an observed depth value most similar to that of the pixel of interest can be selected and assigned to the pixel of interest.
17/32 interest. In the case of known pixels, it may be possible that a pixel with a non-zero body part/player index value could be overtaken, if one of its adjacent pixels has a depth value recorded during rasterization that most closely matches the value. depth of the pixel of interest than the depth value synthesized for that pixel.
Also, for efficiency, updating a body part/player index value of a synthesized pixel may include adding its four adjacent pixels to a row of pixels to be revisited in a later pass. As such, values can continue to propagate across borders without making an entire pass through every pixel. As another optimization, different blocks of NxN pixels (eg blocks of 16x16 pixels) occupied by a target of interest can be tracked so that other blocks that are not occupied by a target of interest can be ignored. Such optimization can be applied at any time during target analysis after rasterization in various ways.
It should be noted, however, that enlargement operations can take a variety of different forms. For example, several fills can first be performed to identify regions of similar values, and then it can be decided which regions belong to which body parts. Also, the number of pixels that any body part/player index object (e.g., forearm body part bp4 from FIG. 3) can increase can be limited based on how many pixels such an object is expected to occupy (e.g. given its shape, distance and angle) versus how many pixels in the association map have been assigned to that body part/player index. Also, the approaches mentioned above may include adding advantages or disadvantages, for certain poses, to condition the magnification for certain body parts so that the magnification can be correct.
A progressive association adjustment can be made to the association map if it is determined that one distribution of pixels from one body part is clustered at one depth, and another distribution from pixels from the same body part is clustered at another depth, so that there is a gap between these two distributions. For example, an arm waving in front of a bust, and close to the bust, can be spilled over the bust. Such a case may produce a group of bust pixels with a body part index indicating that they are arm pixels, when in fact they should be bust pixels. By analyzing the distribution of depth values synthesized at the bottom of the arm, it can be determined that some of the arm pixels can be clustered at one depth, and the rest can be clustered at another depth. The gap between these two groups of depth values indicates a jump between arm pixels and bust pixels. Thus, in response to identifying such a gap, the spill may then be remedied by assigning body part indices to the spill pixels. As another example, a progressive association adjustment can be useful in the case of an arm over a background object. In this case, a histogram can be used to identify a gap in the observed depth of the pixels of interest (ie, the pixels thought to belong to the arm). Based on such a gap, one or more groups of pixels can be identified as properly belonging to an arm and/or other group(s) can be rejected as background pixels. The histogram can be based on a variety of metrics, such as absolute depth; the depth error (synthesized depth - observed depth), etc. Progressive association adjustment can be performed in-line during rasterization, before any enlargement operations.
At 120, the method 100 of FIG. 7 may optionally include creating a heightmap from the observed depth image, the synthesized depth image, and the body part/player index maps in the three processing stages described above. The gradient of such a heightmap and/or a blurred version of such a heightmap can be used when determining the directions of adjustments that should be made to the model, as described below. However, the heightmap is just an optimization; alternatively or additionally, a search in all directions can be performed to identify the closest joints where adjustments can be applied and/or the direction in which such adjustments should be made. When using a heightmap, it can be created before, after, or in parallel with the pixel class determinations described below. When used, the heightmap is designed to set the player's actual body at a low elevation and the background elements at a high elevation. A watershed-type technique can then be used to track a slope on the height map, to find the player's closest point from the bottom, or vice versa (i.e., look for a slope on the map). height to find the closest background pixel for a given player pixel).
The depth synthesized image and the observed depth image cannot be identical, and thus the synthesized depth image may use adjustments and/or modifications so that it more accurately corresponds to an observed depth image and can therefore represent the target more accurately. It should be understood that adjustments can be made to the depth synthesized image by first making adjustments to the model (for example, changing the model's pose), and then synthesizing the adjusted model into a new version of the depth synthesized image.
Several different approaches can be employed to modify an image with synthesized depth. In one approach, two or more different models can be taken and rasterized to produce two or more images with depth.
19/32 synthesized. Each synthesized depth image can then be compared to the depth image observed by a predetermined set of comparison metrics. The image with synthesized depth demonstrating a closer correspondence with the image with observed depth can be selected, and this process can be repeated, optionally, in order to improve the model. When used, this process can be particularly useful for refining the body model to match the player's body type and/or dimensions.
In another approach, the two or more depth synthesized images can be combined by interpellation or extrapolation to produce an image with combined synthesized depth. In yet another approach, two or more depth synthesized images can be blended in such a way that the blending techniques and parameters vary across the depth synthesized image. For example, if a first depth synthesized image favorably matches the depth image observed in one region, and a second depth synthesized image is favorably correlated with a second region, the pose selected in the blended depth synthesized image could be a mixture similar to the pose used to create the first depth synthesized image in the first region and the pose used to create the second depth synthesized image in the second region.
In yet another approach, and as indicated at 122 in FIG. 7, the synthesized depth image can be compared to the observed depth image. Each synthesized pixel of the synthesized image can be classified based on the results of the comparison. Such classification can be designated as determining the pixel case for each pixel. The model used to create the synthesized depth image (e.g., body model 70 in FIG. 3) can be systematically adjusted according to certain pixel cases. In particular, a force vector (quantity and direction) can be calculated on each pixel based on the given pixel case and, depending on the model type, the calculated force vector can be applied to a closer joint, to a centroid. from a body part, to a point on a body part, to a vertex of a triangle, or to another predetermined force reception location of the model used to generate the image with synthesized depth. In some embodiments, the force assigned to a given pixel may be distributed between two or more force receiving locations in the model.
One or more pixel cases can be selected for each synthesized pixel based on one or more factors, which include, but are not limited to: the difference between an observed depth value and a synthesized depth value for that synthesized pixel; the difference between the original Body Part Index, the Body Part Index
20/32 (association) and/or the body part I index (magnification) for that synthesized pixel; and/or the difference between the original player index, the player (association) index, and/or the player (magnification) index for that synthesized pixel.
As indicated at 124 of FIG. 7 , determining a pixel case may include selecting a z-refinement pixel case. The z-refinement pixel case can be selected when the observed depth value of an observed pixel (or in a region of observed pixels) of the image with observed depth does not match the synthesized depth value(s). ) in the image with synthesized depth, but it is close enough to likely belong to the same object in both images, and the body part indices match (or, in some cases, correspond to body parts or neighboring regions). A z-refinement pixel case can be selected for a synthesized pixel if a difference between an observed depth value and a synthesized depth value for the synthesized pixel is within a predetermined range and, optionally, if the index of body part (magnification) of the synthesized pixel corresponds to a body part that has not been designed to receive the forces of magnetism. The z-refinement pixel case corresponds to a calculated force vector that can exert a force on the model to move the model to the correct position. The calculated force vector can be applied along the Z axis perpendicular to the image plane, along a vector normal to an aspect of the model (e.g. the face of the corresponding body part), and/or along a vector normal to the nearby pixels observed. The magnitude of the force vector is based on the difference in observed and synthesized depth values, with larger differences corresponding to larger forces. The force receiving site to which force is applied can be selected to be the closest qualifying force receiving site to the pixel of interest (e.g. nearest bust joint), or the force can be distributed between a weighted mixing of the nearest power receiving locations. The nearest power reception location can be chosen; however, in some cases, applying bias can be helpful. For example, if a pixel is located midway down the lower leg, and it has been established that the hip joint is less mobile (or agile) than the knee, it may be useful to induce joint forces for pixels in the lower leg. middle of the leg to act on the knee instead of the hip.
Determining which force reception location is closest to the pixel of interest can be found by a brute force search, with or without the biases mentioned above. To speed up the search, the set of force reception locations searched can be limited only to those on or near the body part that is associated with the body part index of this pixel. BSP (Binary Space Partitioning) trees can also be created, every time the pose is changed, to help
21/32 speed up these searches. Each body region, or each body part that corresponds to a body part index, can be given its own BSP tree. In this case, biases can be applied differently for each body part, which additionally allows for the correct selection of appropriate force reception sites.
As indicated at 126 of FIG. 7 , determining a pixel case may include selecting a magnetism pixel case. The magnetism pixel case can be used when the synthesized pixel being examined, on the (magnification) map corresponds to a predetermined subset of the body parts (e.g. the arms, or bp3, bp4, bp5, bp7, bp8 and bp9 of the Fig. 3). Although the arms are presented as an example, other parts of the body such as the legs or the whole body can optionally be associated with the magnetism pixel case in some situations. Likewise, in some scenarios, arms cannot be associated with the magnetism pixel case.
The pixels marked for the magnetism case can be grouped into regions, each region being associated with a specific body part (such as, in this example, the upper left arm, the left forearm, the left hand, and so on). It is possible to determine to which region a pixel belongs from its body part index, or by performing a more precise test (to reduce the error potentially introduced in the enlargement operation), comparing the pixel position with several points in or on the body model (but not restricted to the body part indicated by the pixel body part index). For example, for a pixel somewhere on the left arm, various metrics can be used to determine which bone segment (shoulder-elbow, elbow-wrist, or wrist-tip of hand) the pixel is most likely to belong. Each of these bone segments can be considered a “region”.
For each of these magnetism regions, the centroids of pixels belonging to the region can be calculated. These centroids can either be orthodox (all contributing pixels are weighted equally) or biased, with some pixels carrying more weight than others. For example, for the upper arm, three centroids can be tracked: 1) an unbiased centroid, 2) a near centroid, whose contributing pixels are weighted more heavily when they are closer to the shoulder, and 3) a centroid far, whose contributing pixels are weighted higher when closer to the elbow. These coefficients can be linear (eg 2X) or non-linear (eg X<sup>2</sup>) or follow some curve.
Once these centroids are calculated, a variety of options are available (and can be dynamically chosen) to calculate the position and orientation of the body part of interest, even if some are partially obscured. For example, when trying to determine the new position for the elbow, if the centroid in that area is sufficiently visible (if the sum of the weights of the contributing pixels exceeds one
22/32 predetermined limit), then the centroid itself marks the elbow (estimate #1). However, if the elbow area is not visible (perhaps because it is covered by some other object or body part), the location of the elbow can often still be determined, as described in the following non-limiting example. If the distant centroid of the upper arm is visible, then a projection can be made from the shoulder, through this centroid, for the length of the upper arm, to obtain a very likely position for the elbow (estimate #2). If the centroid near the lower arm is visible, then a projection can be made from the wrist, through that centroid, down the length of the lower arm, to obtain a very likely position for the elbow (estimate #3).
Selection of one of three possible estimates can be made, or a combination of the three possible estimates can be made, giving priority (or greater weight) to the estimates that have the highest visibility, confidence, pixel count, or any number of other metrics. . Finally, in this example, a single force vector can be applied to the model at the elbow location; however, it can be more heavily weighted (when accumulated with pixel force vectors resulting from other pixel cases, but acting on this same force reception location), to represent the fact that many pixels were used to construct it. When applied, the computerized force vector can move the model so that the corresponding model more favorably matches the target illustrated in the observed image. One advantage of the magnetism pixel case is its ability to work well with extremely agile body parts such as the arms.
In some embodiments, a model with no defined joints or body parts can be fitted using just the magnetism pixel case.
As indicated at 128 and 130 of FIG. 7, determining a pixel case may include selecting a pull pixel case and/or a push pixel case. These pixel cases can be invoked in the silhouette, where the synthesized and observed depth values can be severely divergent at the same pixel address. Note that the pull pixel case and the push pixel case can also be used when the original player index does not match the player index (magnification). The push versus pull determination is as follows. If the synthesized depth image contains a depth value that is greater than (farther than) the depth value in the observed depth image at that same pixel address, then the model can be pulled towards the true silhouette seen in the image. enlarged image. On the other hand, if the original synthesized image contains a depth value that is less than (closer than) the depth value in the image with observed depth, then the model may be pushed out of space.
23/32 that the player no longer occupies (and towards the actual silhouette in the enlarged image). In both cases, for each of these pixels or pixel regions, a two- or three-dimensional calculated force vector can be exerted on the model to correct the silhouette divergence, or by pushing or pulling parts of the body model to a position that corresponds more precisely to the position of the target in the image with observed depth. The direction of such pushing and/or pulling motion is often predominantly in the XY plane, although a Z component may be added to the force in some scenarios.
In order to produce the appropriate force vector for a push or pull case, the point closest to both the player's silhouette in the image with synthesized depth (for a pull case) and the player's silhouette in the image with observed depth (from a push case) can be found first. This point can be found, for each source pixel (or for each group of source pixels), by performing an exhaustive 2D brute force search for the closest point (in the desired silhouette) that meets the following criteria. In the case of a pull pixel, the closest pixel with a player index on the original map (at the search position) that matches the player index on the augmented map (at the pixel or region of origin) is found. In the case of the push pixel, the closest pixel with an increased map player index (at the search position) that matches the original map player index (at the source pixel or region) is found.
However, a brute force search can be very expensive from a computational point of view, and optimizations can be used to reduce the computational cost. An illustrative non-limiting optimization to find this point more efficiently is to follow the heightmap gradient described above, or a blurred version of it, and just examine pixels in a straight line, in the direction of the gradient. In this height map, the height values are low when the player index is the same in both the original and augmented player index maps, and the height values are high when the player index (on both maps) is zero. . The gradient can be defined as the vector, at any given pixel, pointing downwards in this heightmap. Both push and pull pixels can then seek along that gradient (downhill) until they reach their respective stop condition, as described above. Other basic optimizations for this seek operation include ignoring pixels, using division by range, or using a slope-based approach; re-sampling the gradient, at intervals, as the search progresses: as well as checking for better/closer correlations (not directly along the gradient) once stopping criteria are satisfied.
No matter which technique is used to find the closest point to the silhouette of interest, the distance traveled (the distance between the source pixel and the
24/32 silhouette), D1, can be used to calculate the magnitude (length), D2 , of the force vector that will push or pull the model. In some embodiments, D2 may or may not be linearly related to D1 (e.g., D2 = 2 * D1 or D2 = D1<sup>2</sup>). As a non-limiting example, one can use the following formula: D2 = (D1 - 0.5 pixels) * 2. For example, if there is a gap of 5 pixels between the silhouette in the two images with depth, each pixel in that gap can perform a small search and produce a force vector. Pixels close to the actual silhouette can search for only 1 pixel to reach the silhouette, so the force magnitude in these pixels is (1 - 0.5) * 2 = 1. Pixels far from the real silhouette can search for 5 pixels, so the force magnitude is (5 - 0.5) * 2 = 9. In general, moving from the pixels closest to the real silhouette to the farthest, the search distances will be D1 = {1, 2, 3, 4, 5} and the force quantities produced will be: D2 = {1, 3, 5, 7, 9). The mean of D2 in this case is 5, as desired - the mean magnitudes of the resulting force vectors are equivalent to the distance between the silhouettes (near each force receiving location), which is the distance the model can be moved. to place the model in the proper place.
The final force vector, for each source pixel, can then be constructed with a direction and a magnitude (ie, length). For pull pixels, the direction is determined by the vector from the silhouette pixel to the source pixel; for push pixels, it's the opposite vector. The length of this force vector is D2. At each pixel, then, the force can be applied to a better-qualified (e.g., closest) force reception location (or distributed across several), and these forces can be averaged, at each reception location. force, to produce the appropriate localized movements of the body model.
As indicated at 132 and 134 of FIG. 7 , determining a pixel case may include selecting a pull and/or push pixel case of auto-occlusion. Whereas in the push and pull pixel cases mentioned above, a body part may be moving in the foreground relative to a background or other target, the auto-occlusion push and pull pixel cases consider the scenarios in that the body part is in front of another body part of the same target (eg one leg in front of the other, the arm in front of the bust, etc). These cases can be identified when the pixel's player (association) index matches its corresponding player (magnification) index, but when the body part (association) index does not match its body part (magnification) index. ) corresponding. In such cases, the search direction (to find the silhouette) can be derived in several ways. As non-limiting examples, a brute force 2D search can be performed, a second set of occlusion height maps can be adapted for this case, so that a gradient can guide a 1D search; or the direction can be set in
25/32 towards the nearest point of the closest skeletal member. The details for these two cases are otherwise similar to the standard push and pull cases.
Push, pull, auto-occlude push, and/or auto-occlude pull pixel cases can be selected for a synthesized pixel if the body part (magnification) index of that synthesized pixel corresponds to a body part that was not designed to receive forces of magnetism. It must be understood that in some scenarios, a single pixel may be responsible for one or more pixel cases. As a non-limiting example, a pixel can be responsible for both an auto-occlusion pixel force and a z-refnance pixel force, where the auto-occlusion push pixel force is applied to a receiving location of force on the occluding body part and the z-refinement pixel force is applied to a force receiving site on the body part being occluded.
As indicated at 136 of FIG. 7 , determining a pixel case may include selecting no pixel case for a synthesized pixel. Often, a force vector will not need to be calculated for all synthesized pixels in the synthesized depth image. For example, synthesized pixels that are furthest from the body model illustrated in the depth synthesized image, and observed pixels that are furthest from the target illustrated in the observed depth image (i.e., background pixels) may not influence any location. reception or body part. A pixel case need not be determined for such pixels, although it may be in some scenarios. As another example, a difference between an observed depth value and a synthesized depth value for that synthesized pixel may be below a predetermined threshold value (eg, the model already matches the observed image). As such, a pixel case need not be determined for such pixels, although it may be in some scenarios.
The table presented below details an illustrative relationship between the pixel cases described above and the joints illustrated in the skeleton model 82 of FIG. 5. Pixel cases 1 to 7 are abbreviated in the table as follows: 1-Pull (regular), 2-Pull (occlusion), 3-Push (regular), 4-Push (occlusion), 5-Refinement-Z , 6-Magnetic Pull, and 7Occlusion (no action). A yes record in the Receives Forces column?” indicates that the joint of this line can receive forces from a force vector. An X-register in a pixel case column indicates that the joint of that line can receive a force from a force vector corresponding to the pixel case of that column. It should be understood that the following table is presented as an example. It should not be considered limiting. Other relationships between models and pixel cases can be established without departing from the scope of this disclosure.
26/32
<td colspan="2"></td><td colspan="7">Pixel cases</td>
<td>Articulation</td><td>Do you receive Forces?</td><td> 1</td><td> 2</td><td> 3</td><td> 4</td><td> 5</td><td> 6</td><td> 7</td>
<td>ji</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j2</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>J3</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j4</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j5</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j6</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j7</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j®</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j9</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j10</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j11</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j12</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td></td><td>cases of</td><td colspan="2">s Pixel</td>
<td>Does Articulation Receive Forces?</td><td> 1 2</td><td> 3 4</td><td> 5 6 7</td>
<td>j13 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j14 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j15 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j16 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j17 Yes</td><td></td><td></td><td>XX</td>
<td>j18 Yes</td><td></td><td></td><td>XX</td>
<td>j19 Yes</td><td></td><td></td><td>XX</td>
<td>j20 Yes</td><td></td><td></td><td>XX</td>
<td>j21 Yes</td><td></td><td></td><td>XX</td>
<td>j22 Yes</td><td></td><td></td><td>XX</td>
<td>j23 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j24 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j25 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j26 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>j27 Yes</td><td>XX</td><td>XX</td><td>XX</td>
<td>]28 Yes</td><td>XX</td><td>XX</td><td>XX</td>
27/32
<td>j29</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j30</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j31</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j32</td><td>Yea</td><td>X</td><td>x</td><td>X</td><td>X</td><td>x</td><td></td><td>x</td>
<td>j33</td><td>Yea</td><td>X</td><td>X</td><td>X</td><td>X</td><td>X</td><td></td><td>X</td>
<td>j34</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j35</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j36</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j37</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j38</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j39</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j40</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j41</td><td>No</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>j42</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j43</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j44</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j45</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j46</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
<td>j47</td><td>No</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td><td> -</td>
At 140, the method 100 of FIG. 7 includes, for each synthesized pixel for which a pixel case has been determined, calculating a force vector based on the selected pixel case for that synthesized pixel. As described above, each pixel case corresponds to a different algorithm and/or methodology for selecting the magnitude, direction and/or force reception location of a force vector. Force vectors can be calculated and/or accumulated in any coordinate space, such as world space, screen space (pre-division Z), projection space (post-division Z), model space, and so on. on.
At 142, method 100 includes mapping each calculated force vector to one or more model force receiving locations. Mapping may include mapping a calculated force vector to a best-matched force reception location.” The selection of a best-matched strength reception location is dependent on the pixel case selected for the corresponding pixel. The best matched force receiving site may be the nearest joint, vertex or centroid, for example. In some embodiments, moments (i.e. rotational forces) can be applied to a model.
In general, translations can result from forces with similar directions acting on force-receiving locations on a model, and rotations can result from forces from different directions acting on force-receiving locations on a model. For deformable objects, some of the force vector components can be used to deform the model within its deformation limits, and the remaining force vector components can be used to translate and/or rotate the model.
In some embodiments, force vectors can be mapped to the rigid or deformable object, sub-object, and/or polygon set of an object that best matches. Therefore, some of the force vectors can be used to deform the model, and the other components of the force vectors can be used to perform the rigid translation of the model. Such a technique can result in a fragmented model (eg, an arm can be separated from the body). As discussed in more detail below, a rectification step can then be used to transform translations into rotations and/or apply constraints in order to connect the body parts again along a low-energy trajectory.
FIGs. 11A and 1IB show a very simplified example of applying force vectors to a model - in the illustrated example, a skeleton model 180. For simplicity, only two force vectors are shown in the illustrated example. Each of these force vectors can be the result of the sum of two or more different force vectors resulting from pixel case determinations and force vector calculations of two or more different pixels. Often, a model will be fitted by many different force vectors, each of which is the sum of many different force vectors resulting from pixel case determinations and force vector calculations from many different pixels.
FIG. 11A shows a skeletal model 180, where force vector 182 is to be applied to joint j18 (i.e., an elbow) and force vector 184 is to be applied to joint j20 (i.e., a wrist), for the purpose of to straighten an arm of the Skeleton Model 180 to more accurately match the observed depth image. FIG. 11B shows skeletal model 180 after forces are applied. FIG. 11B illustrates how applied forces adjust the model's pose. As shown in FIG. 11B, the lengths of the skeleton members can be preserved. As further illustrated, the position of joint j2 remains at the shoulder of the skeletal model, as expected in the case of a human straightening his arm. In other words, the skeletal model remains intact after forces have been applied. Maintaining the integrity of the skeletal model during the application of forces results from the application of one or more constraints, as discussed in more detail below. A variety of different constraints can be applied to maintain the integrity of different types of possible models.
29/32
At 144, method 100 of FIG. 7 optionally includes rectifying the model to a pose that satisfies one or more constraints. As described above, after collecting and mapping the calculated force vectors to the model's force receiving locations, the calculated force vectors can then be applied to the model. If performed unconstrained, this can break the model, stretching out of proportion and/or moving body parts to invalid configurations for the target's actual body. Iterations of various functions can then be used to relax the new model position into a “close” legal configuration. During each iteration of model rectification, constraints can be smoothly and/or gradually applied to the pose in order to limit the set of poses to those that can be physically expressed by one or more real bodies of one or more targets. In other embodiments, such a rectification step can be done non-iteratively.
In some embodiments, constraints may include one or more of: skeletal member length constraints, hinge angle constraints, polygon edge angle constraints, and crash tests, as described below.
As an example where a skeletal model is used, skeletal member (ie bone) length constraints can be applied. Force vectors that can be detected (i.e. force vectors at locations where joints and/or body parts are visible and not hidden) can be propagated along a network of skeletal limbs of the skeletal model. By applying skeletal member length constraints, the propagated forces can accommodate once all skeletal members have acceptable lengths. In some embodiments, one or more of the skeleton member lengths may be variable within a predetermined range. For example, the length of the skeletal limbs that make up the sides of the bust can be varied to simulate a deformable center section. As another example, the length of the skeletal limbs that make up the upper arm can be varied to simulate a complex shoulder fit.
A skeletal model can additionally or alternatively be limited by calculating a target-based length of each skeletal member, such that these lengths can be used as constraints during grinding. For example, desired bone lengths are known from the body model; and the difference between current bone lengths (ie, distances between new joint positions) and desired bone lengths can be evaluated. The model can be adjusted to reduce any errors between desired lengths and actual lengths. Priority can be given to certain joints and/or bones that are considered more important, as well as to joints or parts of the body that, at a given time, are more visible than others. In addition, large-scale changes
30/32 may be given higher priority than low magnitude changes.
Joint visibility and/or reliability can be tracked separately in X, Y, Z dimensions to allow for more accurate application of bone length constraints. For example, if a bone connects the chest to the left shoulder, and the Z position of the chest joint is high confidence (that is, many z-refinement pixels correspond to the joint) and the Y position of the shoulder is high confidence ( many push/pull pixels correspond to the joint), so any error in bone length can be corrected while partially or fully limiting shoulder movement in the Y direction or the chest movement in the Z direction.
In some embodiments, hinge positions before grinding can be compared to hinge positions after grinding. If it is determined that a consistent set of adjustments is being made to the skeletal model at each frame, method 100 can use this information to perform progressive refinement on the skeleton and/or body model. For example, by comparing the joint positions before and after the grinding, it can be determined that, in each frame, the shoulders are being pushed farther during the grinding. Such a consistent fit suggests that the shoulders of the skeleton model are smaller than that of the target being depicted, and consequently the shoulder width is being adjusted every frame during the rectification to correct for this. In that case, progressive refinement, such as increasing the shoulder width of the skeletal model, can be performed to correct the skeletal and/or body model to better match the target.
With regard to joint angle restrictions, certain limbs and body parts may be limited in their range of motion relative to an adjacent body part. Also, this range of motion can change according to the orientation of adjacent body parts. Thus, the application of articulation angle restrictions may allow body limb segments to be restricted to possible configurations, given the orientation of the parent limbs and/or parent body parts. For example, the lower leg can be configured to bend backwards (at the knee) but not forward. If illegal angles are detected, the violating body part(s) and/or their parents (or, in the case of a mesh model, the violating triangles and their neighbors) are adjusted to keep the pose within a range of predetermined possibilities, thus helping to avoid the case where the model flexes into a pose that is considered unacceptable. In certain cases of extreme angle violations, the pose may be recognized as backwards, i.e. what is being tracked as the chest is actually the player's back, the left hand is actually the right hand, and so on. against. When such an impossible angle is clearly visible (and sufficiently noticeable), it can be interpreted to mean that the pose has been mapped backwards onto the player's body, and the pose can be inverted to accurately model the target.
Crash tests can be applied to prevent the model from interpenetrating. For example, crash tests can prevent any part of the forearms/hands from penetrating the bust, or prevent the forearms/hands from penetrating each other. In other examples, crash tests can prevent one leg from penetrating the other leg. In some embodiments, crash tests can be applied to models of two or more players to prevent similar situations from occurring between models. In some embodiments, the crash tests can be applied to a body model and/or a skeletal model. In some embodiments, crash tests can be applied to certain polygons of a mesh model.
Crash tests can be applied in any suitable way. One approach examines collisions of one volumetric line segment versus another, where a volumetric line segment can be a line segment with a radius that extends in 3-D. An example of such a crash test might be examining one forearm versus another forearm. In some embodiments, the volumetric line segment may have a different radius at each end of the segment.
Another approach examines collisions of a volumetric line segment versus a polygonal object in a given pose. An example of such a crash test might be examining a forearm versus a torso. In some embodiments, the posed polygonal object may be a deformed polygonal object.
In some embodiments, the knowledge that different parts of a target can move a limited distance in a time interval (eg, 1/30® or 1/60® of a second) can be used as a constraint. Such a constraint can be used to rule out certain poses resulting from the application of forces to the pixel receiving locations of the model.
As indicated in 145, after the model has been fitted and optionally given constraints, the process can restart to begin a new rasterization of the model to a new image with synthesized depth, which can then be compared to the image with observed depth so that other adjustments can be made to the model. In this way, the model can be progressively adjusted to more accurately represent the modeled target. Virtually any number of iterations can be completed for each frame. More iterations can achieve more accurate results, but more iterations can also require more computational overhead. Two or three iterations per frame is believed to be appropriate in many scenarios, although one iteration may be sufficient in some embodiments.
At 146, the method 100 of FIG. 7 optionally includes changing the visual appearance of an on-screen character (e.g., player avatar 190 of FIG. 12A) in response to
32/32 changes to the model, such as the changes shown in FIG. 11B. For example, a user playing an electronic game on a game console (eg, the game system of FIGs. 1A and 1B) may be tracked by the game console as described herein. In particular, a body model (e.g., body model 70 of FIG. 3) including a skeletal model (e.g., skeletal model 180 of FIG. 11A) can be used to model the target player, and the body model can be used to render a player avatar on screen. As the player straightens an arm, the game console may track that movement, and then, in response to the tracked movement, adjust the model 180, as illustrated in FIG. 11B. The game console may also apply one or more restrictions as described above. By making such adjustments and applying such restrictions, the game console may display the adjusted player avatar 192, as illustrated in FIG. 12B. This is also illustrated by way of example in FIG. 1A, in which player avatar 24 is illustrated punching boxing opponent 22 in response to player 18 throwing a punch in real space.
As discussed above, visual target recognition can be performed for purposes other than altering the visual appearance of an on-screen character or avatar. As such, the visual appearance of a character or avatar on screen does not need to be changed in every embodiment. As discussed above, target tracking can be used for virtually any number of different purposes, many of which do not result in changing a character on screen. Target tracking and/or model pose, as adjusted, can be used as a parameter to affect virtually any element of an application, such as a game.
As indicated at 147, the process described above can be repeated for subsequent frames.
It is to be understood that the configurations and/or approaches described herein are exemplary in nature, and that such specific embodiments or examples are not to be considered in a limiting sense, as numerous variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. As such, several illustrated steps can be performed in the illustrated sequence, in other sequences, in parallel, or, in some cases, omitted. Likewise, the order of the processes described above can be changed.
The subject matter of the present disclosure includes all new and non-obvious combinations and subcombinations of the various processes, systems and configurations, and other features, functions, steps and/or properties disclosed in the present invention, as well as any and all equivalents thereof.
Contents4
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
62 members in 11 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 12363604 | United States of America | – | |
| 36360409 | United States of America | A | |
| 2010020791 | United States of America | W | |
| 12363604 | – | – | – |
| PCTUS2010020791 | – | – | – |
| US20090363604 | – | – | – |
| WO2010US20791 | – | – | – |
Members62
| Document | Office | Kind | |
|---|---|---|---|
| CA2748557A1 | Canada | A1 | |
| US2010195869A1 | United States of America | A1 | |
| US2010197391A1 | United States of America | A1 | |
| US2010197392A1 | United States of America | A1 | |
| US2010197393A1 | United States of America | A1 | |
| US2010197395A1 | United States of America | A1 | |
| US2010197399A1 | United States of America | A1 | |
| US2010197400A1 | United States of America | A1 | |
| WO2010088032A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2010088032A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2011071696A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2011071801A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2011071804A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2011071808A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2011071811A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2011071815A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2011071696A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2011071811A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2011071801A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20110117114A | Republic of Korea | A | |
| WO2011071808A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2011071815A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2011071804A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP2391988A2 | European Patent Office (EPO) | A2 | |
| CN102301313A | China | A | |
| US2012077591A1 | United States of America | A1 | |
| TW201215435A | Taiwan Province of China | A | |
| JP2012516504A | Japan | A | |
| CN102639198A | China | A | |
| CN102640186A | China | A | |
| CN102640187A | China | A | |
| CN102648032A | China | A | |
| CN102648484A | China | A | |
| CN102665837A | China | A | |
| US8267781B2 | United States of America | B2 | |
| RU2011132029A | Russian Federation | A | |
| HK1172134A1 | Hong Kong, China | A1 | |
| HK1173691A1 | Hong Kong, China | A1 | |
| JP5227463B2 | Japan | B2 | |
| CN102648032B | China | B | |
| US8565476B2 | United States of America | B2 | |
| US8565477B2 | United States of America | B2 | |
| CN102665837B | China | B | |
| CN102301313B | China | B | |
| US8577084B2 | United States of America | B2 | |
| US8577085B2 | United States of America | B2 | |
| US8588465B2 | United States of America | B2 | |
| US2014051515A1 | United States of America | A1 | |
| US8682028B2 | United States of America | B2 | |
| CN102648484B | China | B | |
| CN102639198B | China | B | |
| RU2530334C2 | Russian Federation | C2 | |
| TWI469812B | Taiwan Province of China | B | |
| CN102640186B | China | B | |
| US9039528B2 | United States of America | B2 | |
| CN102640187B | China | B | |
| BRPI1006111A2This record | Brazil | A2 | |
| KR101619562B1 | Republic of Korea | B1 | |
| CA2748557C | Canada | C | |
| US9842405B2 | United States of America | B2 | |
| EP2391988A4 | European Patent Office (EPO) | A4 | |
| EP2391988B1 | European Patent Office (EPO) | B1 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Objections, documents and/or translations needed after an examination request according art. 34 industrial property lawB06F | B06F | |
| Requested transfer of rights approvedB25A | B25A |
Numbers
- Publication
- PI1006111
- Publication, DOCDB
- PI1006111
- Publication, EPODOC
- BRPI1006111
- Application
- 6111
- Application, DOCDB
- PI1006111
- Application, EPODOC
- BR2010PI06111
Titles2
- Portuguese
- RASTREAMENTO DE ALVO VISUAL
- English
- visual target tracking
Classification
- CPC, 8
- G06T7/251
- A63F2300/1093
- A63F2300/6607
- A63F13/213
- A63F13/428
- G06T2207/10028
- G06T2207/30196
- G06T2207/10021
- IPC, 2
- G06T7 20
- G06T15 00