Speaker detection and tracking using audiovisual data
Summary by NHIP
Audiovisual object tracking
The system tracks objects by combining audio and video probabilistic generative models. It uses specific equations defining parameters like π, η, λ, and ν for audio signals alongside equations with π, μ, φ, and ψ for video inputs.
Claim Score by NHIP
Abstract
A system and method facilitating object tracking is provided. The system includes an audio model that receives at least two audio input signals and a video model that receives a video input. The audio model and the video model employ probabilistic generative models which are combined to facilitate object tracking. Expectation maximization can be employed to modify trainable parameters of the audio model and the video model.

Term
Term ended
Expired 17 September 2023, 3 years ago.
- Priority and filed
- Granted
- Expired
- Today
24 claims: 7 independent, 17 dependent
- 1An object tracker system, comprising:an audio model that models an original audio signal of an object, a time delay between at least two audio input signals and a variability component of the original audio signal, the audio model employing a probabilistic generative model, and employing, at least in part, the following equations: p ( r )=π r , p ( a|r )= N ( a| 0,η r ), p ( x 1 |a )= N ( x 1 |λ 1 a,ν 1 ), p ( x 2 |a ,τ)= N ( x 2 |λ 2 L τ a,ν 2 ), where r is variability component of the original audio signal, π is a prior probability parameter of r, a is the original audio signal of the object, x 1 is a first audio input signal, x 2 is a second audio input signal, τ is the time delay between x 1 and x 2 , λ 1 is an attenuation parameter associated with x 1 , λ 2 is an attenuation parameter associated with x 2 , η r is a precision matrix parameter associated with r, ν 1 is a precision matrix parameter associated with additive noise of x 1 , ν 2 is a precision matrix parameter associated with additive noise of x 2 , L r denotes a temporal shift operator: a video model that models a location of the object, an original image of the object and a variability component of the original image, the video model employing a probabilistic generative model, the video model receiving a video input;and, an audio video tracker that models the location of the object based, at least in part, upon the audio model and the video model, the audio video tracker providing an output associated with the location of the object.
- 17Broadest claimClaim Score 30, narrow(NHIP)A method for object tracking, comprising:updating a posterior distribution over unobserved variables of an audio model and a video model;updating trainable parameters of the audio model and the video model;employing, at least in part, the following equations in the audio model: p ( r )=π r , p ( a|r )= N ( a| 0,η r ), p ( x 1 |a )= N ( x 1 |λ 1 a,ν 1 ), p ( x 2 |a ,τ)= N ( x 2 |λ 2 L τ a,ν 2 ), where r is variability component of the original audio signal, π a prior probability parameter of r, a is the original audio signal of the object, x 1 is a first audio input signal, x 2 is a second audio input signal, τ is the time delay between x 1 and x 2 , λ 1 is an attenuation parameter associated with x 1 , λ 2 is an attenuation parameter associated with x 2 , η r is a precision matrix parameter associated with r, ν 1 is a precision matrix parameter associated with additive noise of x 1 , ν 2 is a precision matrix parameter associated with additive noise of x 2 , L r denotes a temporal shift operator;and, providing an output associated with a location of an object.
- 20A data packet transmitted between two or more computer components that facilitates object tracking, the data packet comprising:a first data field comprising information associated with a horizontal location of an object;and, a second data field comprising information associated with a vertical location of the object, the horizontal location and the vertical location being based, at least in part, upon an object tracker system receiving at least two audio signal inputs and a video input signal;wherein the object tracker system comprising at least an audio model employing, at least in part, the following equations: p ( r )=π r , p ( a|r )= N ( a| 0,η r ), p ( x 1 |a )= N ( x 1 |λ 1 a,ν 1 ), p ( x 2 |a ,τ)= N ( x 2 |λ 2 L τ a,ν 2 ), where r is variability component of the original audio signal, π r is a prior probability parameter of r, a is the original audio signal of the object, x 1 is a first audio input signal, x 2 is a second audio input signal. π is the time delay between x 1 and x 2 , λ 1 is an attenuation parameter associated with x 1 , λ 2 is an attenuation parameter associated with x 2 , η r is a precision matrix parameter associated with r, ν 1 is a precision matrix parameter associated with additive noise of x 1 , ν 2 is a precision matrix parameter associated with additive noise of x 2 , L r denotes a temporal shift operator.
- 21A computer readable medium storing computer executable components of an object tracker system, comprising:an audio model component that models an original audio signal of an object, a time delay between at least two audio input signals and a variability component of the original audio signal, the audio model employing a probabilistic generative model;a video model component that models a location of the object, an original image of the object and a variability component of the original image, the video model employing a probabilistic generative model, the video model receiving a video input;and employing, at least in part, the following equations: p ( s )=π s , p (ν| s )= N (ν|μ s ,φ s ), p ( y|ν,l )= N ( y|G l ν,ψ), where π s is a prior probability parameter of s, y is the video input signal, l is the location of the object, ν is the original image of the object, μ s is a mean parameter associated with s, φ s is a precision matrix parameter associated with s, ψ is a precision matrix parameter associated with additive noise of y, G l denotes a shift operator;and, an audio video tracker component that models the location of the object based, at least in part, upon the audio model and the video model, the audio video tracker providing an output associated with the location of the object.
- 22An means for modeling audio that models an original audio signal of an object, a time delay between at least two audio input signals and a variability component of the original audio signal, the means for modeling audio employing a probabilistic generative model, and employing, at least in part, the following equations in the audio model:p ( r )=π r , p ( a|r )= N ( a| 0,η r ), p ( x 1 |a )= N ( x 1 |λ 1 a,ν 1 ), p ( x 2 |a ,τ)= N ( x 2 |λ 2 L τ a,ν 2 ), where r is variability component of the original audio signal, π a prior probability parameter of r, a is the original audio signal of the object, x 1 is a first audio input signal, x 2 is a second audio input signal, τ is the time delay between x 1 and x 2 , λ 1 is an attenuation parameter associated with x 1 , λ 2 is an attenuation parameter associated with x 2 , η r is a precision matrix parameter associated with r, ν 1 is a precision matrix parameter associated with additive noise of x 1 , ν 2 is a precision matrix parameter associated with additive noise of x 2 , L r denotes a temporal shift operator;means for modeling video that models a location of the object, an original image of the object and a variability component of the original image, the means for modeling video employing a probabilistic generative model;and, means for tracking the location of the object based, at least in part, upon the means for modeling audio and the means for model video, the means for tracking the location of the object providing an output associated with the location of the object.
- 23An object tracker system, comprising:an audio model that models an original audio signal of an object, a time delay between at least two audio input signals and a variability component of the original audio signal, the audio model employing a probabilistic generative model;a video model that models a location of the object, an original image of the object, a variability component of the original image and a background image, the video model employing a probabilistic generative model, the video model receiving a video input, and employing, at least in part, the following equations: p ( s )=π s , p (ν| s )= N (ν|μ s ,φ s ), p ( y|ν,l )= N ( y|G l ν,ψ), where π s is a prior probability parameter of s, y is the video input signal, l is the location of the object, ν is the original image of the object, μ s is a mean parameter associated with s, φ s is a precision matrix parameter associated with s, ψ is a precision matrix parameter associated with additive noise of ν, G l denotes a shift operator. an audio video tracker that models the location of the object based, at least in part, upon the audio model and the video model, the audio video tracker providing an output associated with the location of the object.
- 24An object tracker system, comprising:an audio model that models an original audio signal of an object, a time delay between at least two audio input signals, a variability component of the original audio signal and a previous original audio signal of the object, the audio model employing a probabilistic generative model, and employing, at least in part, the following equations: p ( r )=π r , p ( a|r )= N ( a| 0,η r ), p ( x 1 |a )= N ( x 1 |λ 1 a,ν 1 ), p ( x 2 |a ,τ)= N ( x 2 |λ 2 L τ a,ν 2 ), where r is variability component of the original audio signal, π a prior probability parameter of r, a is the original audio signal of the object, x 1 is a first audio input signal, x 2 is a second audio input signal, τ is the time delay between x 1 and x 2 , λ 1 is an attenuation parameter associated with x 1 , λ 2 is an attenuation parameter associated with x 2 , η r is a precision matrix parameter associated with r, ν 1 is a precision matrix parameter associated with additive noise of x 1 , ν 2 is a precision matrix parameter associated with additive noise of x 2 , L r denotes a temporal shift operator: a video model that models a location of the object, an original image of the object and a variability component of the original image, the video model employing a probabilistic generative model, the video model receiving a video input;and, an audio video tracker that models the location of the object based, at least in part, upon the audio model, the video model and a previous location of the object, the audio video tracker providing an output associated with the location of the object.
Independent claims7
80 paragraphs in 5 sections, as filed
TECHNICAL FIELD
0001The present invention relates generally to object (speaker) detection and tracking, and, more particularly to a system and method for object (speaker) detection and tracking using audiovisual data.
BACKGROUND OF THE INVENTION
0002Video conferencing has become increasingly effective in order to facilitate discussion among physically remote participants. A video input device, such as a camera, generally provides the video input signal portion of a video conference. Many conventional systems employ an operator to manually operate (e.g., move) the video input device.
0003Other systems employ a tracking system to facilitate tracking of speakers. However, in many conventional system(s) that process digital media, audio and video data are generally treated separately. Such systems usually have subsystems that are specialized for the different modalities and are optimized for each modality separately. Combining the two modalities is performed at a higher level. This process generally requires scenario dependent treatment, including precise and often manual calibration. A tracker using only video data may mistake the background for the object or lose the object altogether due to occlusion. Further, a tracker using only audio data can lose the object as it stops emitting sound or is masked by background noise.
SUMMARY OF THE INVENTION
0004The following presents a simplified summary of the invention in order to provide a basic understanding of some aspects of the invention. This summary is not an extensive overview of the invention. It is not intended to identify key/critical elements of the invention or to delineate the scope of the invention. Its sole purpose is to present some concepts of the invention in a simplified form as a prelude to the more detailed description that is presented later.
0005The present invention provides for an object tracker system having an audio model, a video model and an audio video tracker. For example, the system can be used to track a human speaker.
0006The object tracker system employs modeling and processing multimedia data to facilitate object tracking based on graphical models that combine audio and video variables. The object tracker system can utilize an algorithm for tracking a moving object, for example, in a cluttered, noisy scene using at least two audio input signals (e.g., from microphones) and a video input signal (e.g., from a camera). The object tracker system utilizes unobserved (hidden) variables to describe the observed data in terms of the process that generates them. The object tracker system is therefore able to capture and exploit the statistical structure of the audio and video data separately as well as their mutual dependencies. Parameters of the system can be learned from data, for example, via an expectation maximization (EM) algorithm, and automatic calibration is performed as part of this procedure. Tracking can be done by Bayesian inference of the object location from the observed data.
0007The object tracker system uses probabilistic generative models (also termed graphical models) to describe the observed data (e.g., audio input signals and video input signal).
0008The object's original audio signal and a time delay between observed audio signals are unobserved variables in the audio model. Similarly, a video input signal is generated by the object's original image, which is shifted as the object's spatial location changes. Thus, the object's original image and location are also unobserved variables in the video model. The presence of unobserved variables is typical of probabilistic generative models and constitutes one source of their power and flexibility. The time delay between the audio signals can be reflective of the object's position.
0009The object tracker system combines the audio model and the video model in a principled manner using a single probabilistic model. Probabilistic generative models have several important advantages. First, since they explicitly model the actual sources of variability in the problem, such as object appearance and background noise, the resulting algorithm turns out to be quite robust. Second, using a probabilistic framework leads to a solution by an estimation algorithm that is Bayes-optimal. Third, parameter estimation and object tracking can both be performed efficiently using the expectation maximization (EM) algorithm.
0010Within the probabilistic modeling framework, the problem of calibration becomes the problem of estimating the parametric dependence of the time delay on the object position. The object tracker system estimates these parameters as part of the EM algorithm and thus no special treatment is required. Hence, the object tracker system does not need prior calibration and/or manual initialization (e.g., defining the templates or the contours of the object to be tracked, knowledge of microphone based line, camera focal length and/or various threshold(s) used in visual feature extraction) as with conventional systems.
0011The audio model receives at least two audio signal inputs and models a speech signal of an object, a time delay between the audio input signals and a variability component of the original, audio signal. The video model models a location of the object, an original image of the object and a variability component of the original image. The audio video tracker links the time delay between the audio signals of the audio model to the spatial location of the object's image of the video model. Further, the audio video tracker provides an output associated with the location of the object.
0012The audio video tracker thus fuses the audio model and the video model into a single probabilistic graphical model. The audio video tracker can exploit the fact that the relative time delay between the microphone signals is related to the object position. The parameters of the object tracker system can be trained using, for example, variational method(s). In one implementation, an expectation maximization (EM) algorithm is employed. The E-step of an iteration updates the posterior distribution over the unobserved variables conditioned on the data. The M-step of an iteration updates parameter estimates.
0013Tracking of the object tracker system is performed as part of the E-step. The object tracker system can provide an output associated with the most likely object position. Further, the object tracker system can be employed, for example, in a video conferencing system and/or a multi-media processing system.
0014To the accomplishment of the foregoing and related ends, certain illustrative aspects of the invention are described herein in connection with the following description and the annexed drawings. These aspects are indicative, however, of but a few of the various ways in which the principles of the invention may be employed and the present invention is intended to include all such aspects and their equivalents. Other advantages and novel features of the invention may become apparent from the following detailed description of the invention when considered in conjunction with the drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
<figref idref="DRAWINGS">FIG. 1</figref> is a block diagram of an object tracker system in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram of an object tracker system in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 3</figref> is a diagram illustrating horizontal and vertical position of an object in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 4</figref> is a graphical representation of an audio model in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 5</figref> is a graphical representation of a video model in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 6</figref> is a graphical representation of an audio video tracker in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 7</figref> is a block diagram of an object tracker system in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 8</figref> is a flow chart illustrating a method for object tracking in accordance with an aspect of the present invention.
<figref idref="DRAWINGS">FIG. 9</figref> is a flow chart further illustrating the method of FIG. <b>8</b>.
<figref idref="DRAWINGS">FIG. 10</figref> illustrates an example operating environment in which the present invention may function.
DETAILED DESCRIPTION OF THE INVENTION
0025The present invention is now described with reference to the drawings, wherein like reference numerals are used to refer to like elements throughout. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It may be evident, however, that the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to facilitate describing the present invention.
0026As used in this application, the term “computer component” is intended to refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution. For example, a computer component may be, but is not limited to being, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and/or a computer. By way of illustration, both an application running on a server and the server can be a computer component. One or more computer components may reside within a process and/or thread of execution and a component may be localized on one computer and/or distributed between two or more computers.
0027Referring to <figref idref="DRAWINGS">FIG. 1</figref>, an object tracker system <b>100</b> in accordance with an aspect of the present invention is illustrated. The system <b>100</b> includes an audio model <b>110</b>, a video model <b>120</b> and an audio video tracker <b>130</b>. For example, the object can be a human speaker.
0028The object tracker system <b>100</b> employs modeling and processing multimedia data to facilitate object tracking based on graphical models that combine audio and video variables. The object tracker system <b>100</b> can utilize an algorithm for tracking a moving object, for example, in a cluttered, noisy scene using at least two audio input signals (e.g., from microphones) and a video input signal (e.g., from a camera). The object tracker system <b>100</b> utilizes unobserved variables to describe the observed data in terms of the process that generates them. The object tracker system <b>100</b> is therefore able to capture and exploit the statistical structure of the audio and video data separately as well as their mutual dependencies. Parameters of the system <b>100</b> can be learned from data, for example, via an expectation maximization (EM) algorithm, and automatic calibration is performed as part of this procedure. Tracking can be done by Bayesian inference of the object location from the observed data.
0029The object tracker system <b>100</b> uses probabilistic generative models (also termed graphical models) to describe the observed data (e.g., audio input signals and video input signal). The models are termed generative, since they describe the observed data in terms of the process that generated them, using additional variables that are not observable. The models are termed probabilistic, because rather than describing signals, they describe probability distributions over signals. These two properties combine to create flexible and powerful models. The models are also termed graphical since they have a useful graphical representation.
0030Turning briefly to <figref idref="DRAWINGS">FIG. 2</figref>, an object tracker system <b>200</b> in accordance with an aspect of the present invention is illustrated. The system <b>200</b> includes an object tracker system <b>100</b>, a first audio input device <b>210</b><sub>1</sub>, a second audio input device <b>210</b><sub>2 </sub>and a video input device <b>220</b>. For example, the first audio input device <b>210</b><sub>1 </sub>and the second audio input device <b>210</b><sub>2 </sub>can be microphones and the video input device <b>220</b> can be a camera.
0031The observed audio signals are generated by, for example, an object's original audio signal, which arrives at the second audio input device <b>210</b><sub>2 </sub>with a time delay relative to the first audio input device <b>210</b><sub>1</sub>. The object's original signal and the time delay are unobserved variables in the audio model <b>110</b>. Similarly, a video input signal is generated by the object's original image, which is shifted as the object's spatial location changes. Thus, the object's original image and location are also unobserved variables in the video model <b>120</b>. The presence of unobserved (hidden) variables is typical of probabilistic generative models and constitutes one source of their power and flexibility.
0032The time delay between the audio signals captured by the first audio input device <b>210</b><sub>1 </sub>and the second audio input device <b>220</b><sub>2 </sub>is reflective of the object's horizontal position l<sub>x</sub>.
0033Referring briefly to <figref idref="DRAWINGS">FIG. 3</figref>, a diagram illustrating horizontal and vertical position of an object in accordance with an aspect of the present invention is illustrated. The object has a horizontal position l<sub>x </sub>and a vertical position l<sub>y</sub>, collectively referred to as location l.
0034Turning back to <figref idref="DRAWINGS">FIG. 1</figref>, the object tracker system <b>100</b> combines the audio model <b>110</b> and the video model <b>120</b> in a principled manner using a single probabilistic model. Probabilistic generative models have several important advantages. First, since they explicitly model the actual sources of variability in the problem, such as object appearance and background noise, the resulting algorithm turns out to be quite robust. Second, using a probabilistic framework leads to a solution by an estimation algorithm that is Bayes-optimal. Third, parameter estimation and object tracking can both be performed efficiently using the expectation maximization (EM) algorithm.
0035Within the probabilistic modeling framework, the problem of calibration becomes the problem of estimating the parametric dependence of the time delay on the object position. The object tracker system <b>100</b> estimates these parameters as part of the EM algorithm and thus no special treatment is required. Hence, the object tracker system <b>100</b> does not need prior calibration and/or manual initialization (e.g., defining the templates or the contours of the object to be tracked, knowledge of microphone based line, camera focal length and/or various threshold(s) used in visual feature extraction) as with conventional systems.
0036The audio model <b>110</b> receives a first audio input signal through an Nth audio input signal, N being an integer greater than or equal to two, hereinafter referred to collectively as the audio input signals. The audio input signals are represented by a sound pressure waveform at each microphone for each frame. For purposes of discussion, two audio input signals will be employed; however, it is to be appreciated that any suitable quantity of audio input signals suitable for carrying out the present invention can be employed and are intended to fall within the scope of the hereto appended claims.
0037The audio model <b>110</b> models a speech signal of an object, a time delay between the audio input signals and a variability component of the original audio signal. For example, the audio model <b>110</b> can employ a hidden Markov model. The audio model <b>110</b> models audio input signals x<sub>1</sub>, x<sub>2 </sub>as follows. First, each audio input signal is chopped into equal length segments termed frames. For example, the frame length can be determined by the frame rate of the video. Hence, 30 video frames per second translates into 1/30 second long audio frames. Each audio frame is a vector with entries x<sub>1n</sub>, x<sub>2n </sub>corresponding to the audio input signal values at time point n. The audio model <b>110</b> can be trained online (e.g., from available audio data) and/or offline using pre-collected data. For example, the audio model <b>110</b> can be trained offline using clean speech data from another source (e.g., not contemporaneous with object tracking).
0038x<sub>1</sub>, x<sub>2 </sub>are described in terms of an original audio signal a. It is assumed that a is attenuated by a factor λ<sub>1 </sub>on its way to audio input device i=1, 2, and that it is received at the second audio input device with a delay of τ time points relative to the first audio input device: <br /><i>x</i><sub>1n</sub>=λ<sub>1</sub><i>a</i><sub>n</sub>,<br /><i>x</i><sub>2n</sub>=λ<sub>2</sub><i>a</i><sub>n−τ</sub> (1)<br /> It can further be assumed that a is contaminated by additive sensor noise with precision matrices ν<sub>1</sub>, v<sub>2</sub>. To account for the variability of that signal, it is described by a mixture model. Denoting the component label by r, each component has mean zero, a precision matrix η<sub>r</sub>, and a prior probability π<sub>r</sub>. Viewing it in the frequency domain, the precision matrix corresponds to the inverse of the spectral template for each component, thus: <br /><i>p</i>(<i>r</i>)=π<sub>r</sub>,<br /><i>p</i>(<i>a|r</i>)=<i>N</i>(<i>a|</i>0,η<sub>r</sub>),<br /><i>p</i>(<i>x</i><sub>1</sub><i>|a</i>)=<i>N</i>(<i>x</i><sub>1</sub>|λ<sub>1</sub><i>a,ν</i><sub>1</sub>),<br /><i>p</i>(<i>x</i><sub>2</sub><i>|a</i>,τ)=<i>N</i>(<i>x</i><sub>2</sub>|λ<sub>2</sub><i>L</i><sub>τ</sub><i>a,ν</i><sub>2</sub>), (2)<br /> where L<sub>r </sub>denotes the temporal shift operator (e.g., (L<sub>r</sub>a)<sub>n</sub>=a<sub>n−r</sub>). The prior probability for a delay τ is assumed flat, p(τ)=const. N(x|μ,ν) denotes a Gaussian distribution over the random vector x with mean μ and precision matrix (defined as the inverse covariance matrix) ν: <maths id="MATH-US-00001" num="00001"><math overflow="scroll"><mtable><mtr><mtd><mrow><mrow><mi>𝒩</mi><mo></mo><mrow><mo>(</mo><mrow><mrow><mi>x</mi><mo>|</mo><mi>μ</mi></mrow><mo>,</mo><mi>v</mi></mrow><mo>)</mo></mrow></mrow><mo>∝</mo><mrow><mi>exp</mi><mo></mo><mrow><mo>[</mo><mrow><mrow><mo>-</mo><mfrac><mn>1</mn><mn>2</mn></mfrac></mrow><mo></mo><msup><mrow><mo>(</mo><mrow><mi>x</mi><mo>-</mo><mi>μ</mi></mrow><mo>)</mo></mrow><mi>T</mi></msup><mo></mo><mrow><mi>v</mi><mo></mo><mrow><mo>(</mo><mrow><mi>x</mi><mo>-</mo><mi>μ</mi></mrow><mo>)</mo></mrow></mrow></mrow><mo>]</mo></mrow></mrow></mrow></mtd><mtd><mrow><mo>(</mo><mn>3</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths>
0039Referring briefly to <figref idref="DRAWINGS">FIG. 4</figref>, a graphical representation <b>400</b> of the audio model <b>110</b> in accordance with an aspect of the present invention is illustrated. The graphical representation <b>400</b> includes nodes and edges. The nodes include observed variables (illustrated as shaded circles), unobserved variables (illustrated as unshaded circles) and model parameters (illustrated as square boxes). The edges (directed arrow) correspond to a probabilistic conditional dependence of the nose at the arrow's head on the node at tail.
0040A probabilistic graphical model has a generative interpretation: according to the graphical representation <b>400</b>, the process of generating the observed microphone signals starts with picking a spectral component r with probability p(r), followed by drawing a signal a from the Gaussian p(a|r). Separately, a time delay τ is also picked. The signals x<sub>1</sub>, x<sub>2 </sub>are then drawn from the undelayed Gaussian p(x<sub>1</sub>|a) and the delayed Gaussian p(x<sub>2</sub>|a, τ), respectively.
0041Turning back to <figref idref="DRAWINGS">FIG. 1</figref>, the video model <b>120</b> models a location of the object (l<sub>x</sub>, l<sub>y</sub>), an original image of the object (ν) and a variability component of the original image (s). For example the video model <b>120</b> can employ a hidden Markov model. The video model <b>120</b> can be trained online (e.g., from available video data) and/or offline using pre-collected data.
0042The video input can be represented by a vector of pixel intensities for each frame. An observed frame is denoted by y, which is a vector with entries y<sub>n </sub>corresponding to the intensity of pixel n. This vector is described in terms of an original image ν that has been shifted by l=(l<sub>x</sub>, l<sub>y</sub>) pixels in the x and y directions, respectively, <br /><i>y</i><sub>n</sub>=ν<sub>n−1</sub>, (4)<br /> and has been further contaminated by additive noise with precision matrix ψ. To account for the variability in the original image, ν is modeled by a mixture model. Denoting its component label by s, each component is a Gaussian with mean μ<sub>s </sub>and precision matrix φ<sub>s</sub>, and has a prior probability π<sub>s</sub>. The means serve as image templates. Hence: <br /><i>p</i>(<i>s</i>)=π<sub>s</sub>,<br /><i>p</i>(ν|<i>s</i>)=<i>N</i>(ν|μ<sub>s</sub>,φ<sub>s</sub>),<br /><i>p</i>(<i>y|ν,l</i>)=<i>N</i>(<i>y|G</i><sub>l</sub>ν,ψ), (5)<br /> where G<sub>l </sub>denotes the shift operator (e.g., (G<sub>l</sub>ν)<sub>n</sub>=ν<sub>n−l</sub>). The prior probability for a shift l is assumed flat, p(l)=const.
0043Referring briefly to <figref idref="DRAWINGS">FIG. 5</figref>, a graphical representation <b>500</b> of the video model <b>120</b> in accordance with an aspect of the present invention is illustrated. The graphical representation <b>500</b> includes nodes and edges. The nodes include observed variables (illustrated as shaded circles), unobserved variables (illustrated as unshaded circles) and model parameters (illustrated as square boxes).
0044The graphical representation <b>500</b> has a generative interpretation. The process of generating the observed image starts with picking an appearance component s from the distribution p(s)=π<sub>s</sub>, followed by drawing an image ν from the Gaussian p(ν|s). The image is represented as a vector of pixel intensities, where the elements of the diagonal precision matrix define the level of confidence in those intensities. Separately, a discrete shift l is picked. The image y is then drawn from the shifted Gaussian p(y|ν,l).
0045Notice the symmetry between the audio model <b>110</b> and video model <b>120</b>. In each model, the original signal is hidden and described by a mixture model. In the video model <b>120</b> the templates describe the image, and in the audio model <b>110</b> the templates describe the spectrum. In each model, the data are obtained by shifting the original signal, where in the video model <b>120</b> the shift is spatial and in the audio model <b>110</b> the shift is temporal. Finally, in each model the shifted signal is corrupted by additive noise.
0046Referring back to <figref idref="DRAWINGS">FIG. 1</figref>, the audio video tracker <b>130</b> links the time delay τ between the audio signals of the audio model <b>110</b> to the spatial location of the object's image of the video model <b>120</b>. Further, the audio video tracker <b>130</b> provides an output associated with the location of the object.
0047The audio video tracker <b>130</b> thus fuses the audio model <b>110</b> and the video model <b>120</b> into a single probabilistic graphical model. The audio video tracker <b>130</b> can exploit the fact that the relative time delay τ between the microphone signals is related to the object position l. In particular, as the distance of the object from the sensor setup becomes much larger than the distance between the microphones, τ becomes linear in l. Therefore, a linear mapping can be used to approximate this dependence and the approximation error can be modeled by a zero mean Gaussian with precision ν<sub>τ</sub>, <br /><i>p</i>(τ|<i>l</i>)=<i>N</i>(τ|<i>al</i><sub>x</sub><i>+al</i><sub>y</sub>+β,ν<sub>τ</sub>). (6)
0048Note that the mapping involves the horizontal position, as the vertical movement has a significantly smaller affect on the signal delay due to the horizontal alignment of the microphones (e.g., α′≈0). The link formed by Eq. (6) fuses the audio model <b>110</b> and the video model <b>120</b> into a single model.
0049Referring to <figref idref="DRAWINGS">FIG. 6</figref>, a graphical representation <b>600</b> of the audio video tracker <b>130</b> in accordance with an aspect of the present invention is illustrated. The graphical representation <b>600</b> includes the observed variables, unobserved variables and model parameters of the audio model <b>110</b> and the video model <b>120</b>. The graphical representation <b>600</b> includes a probabilistic conditional dependence of the unobserved time delay parameter τ of the audio model <b>110</b> upon the unobserved object positional variables l<sub>x </sub>and l<sub>y </sub>of the video model <b>120</b>.
0050The parameters of the object tracker system <b>100</b> can be trained using, for example, variational method(s). In one implementation, an expectation maximization (EM) algorithm is employed.
0051Generally, an iteration in the EM algorithm consists of an expectation step (or E-step) and a maximization step (or M-step). For each iteration, the algorithm gradually improves the parameterization until convergence. The EM algorithm can be performed as many EM iterations as necessary (e.g., to substantial convergence).
0052With regard to the object tracker system <b>100</b>, the E-step of an iteration updates the posterior distribution over the unobserved variables conditioned on the data. The M-step of an iteration updates parameter estimates.
0053The joint distribution over the variables of the audio video tracker <b>130</b> including the observed variables x<sub>1</sub>, x<sub>2</sub>, y, the unobserved variables a, τ, r, ν, l, s factorizes as: <br /><i>p</i>(<i>x</i><sub>1</sub><i>,x</i><sub>2</sub><i>,y,a,τ,r,ν,l,s</i>|θ)=<i>p</i>(<i>x</i><sub>1</sub><i>|a</i>)<i>p</i>(<i>x</i><sub>2</sub><i>|a</i>,τ)<i>p</i>(<i>a|r</i>)·<i>p</i>(<i>r</i>)<i>p</i>(<i>y|ν,l</i>)<i>p</i>(ν|<i>s</i>)<i>p</i>(<i>s</i>)<i>p</i>(τ|<i>l</i>)<i>p</i>(<i>l</i>). (7)
0054This is the product of the joint distributions defined by the audio model <b>110</b>, and the video model <b>120</b> as linked by the audio video tracker <b>130</b>. The parameters of the audio video tracker are: <br />θ={λ<sub>1</sub>,ν<sub>1</sub>,λ<sub>2</sub>,ν<sub>2</sub>,n<sub>r</sub>,π<sub>r</sub>,ψ,μ<sub>s</sub>,φ<sub>s</sub>,π<sub>s</sub>,α,α′,β,ν<sub>τ</sub>}. (8)
0055Ultimately, tracking of the object based on the data is desired, that is, obtaining a position estimate {circumflex over (l)} at each frame. In the framework of probabilistic modeling, more than a single value of l is generally computed. Thus, the full posterior distribution over l given the data, p(l|x<sub>1</sub>, x<sub>2</sub>, y), for each frame, is computed. This distribution provides the most likely position value via <maths id="MATH-US-00002" num="00002"><math overflow="scroll"><mtable><mtr><mtd><mrow><mrow><mover><mi>l</mi><mo>^</mo></mover><mo>=</mo><mrow><mi>arg</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><munder><mi>max</mi><mi>l</mi></munder><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mrow><mrow><mi>l</mi><mo>|</mo><msub><mi>x</mi><mn>1</mn></msub></mrow><mo>,</mo><msub><mi>x</mi><mn>2</mn></msub><mo>,</mo><mi>y</mi></mrow><mo>)</mo></mrow></mrow></mrow></mrow></mrow><mo>,</mo></mrow></mtd><mtd><mrow><mo>(</mo><mn>9</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> as well as a measure of how confident the model is of that value. It can also handle situations where the position is ambiguous (e.g., by exhibiting more than one mode). An example is when the object (e.g., speaker) is occluded by either of two objects. However, in one example, the position posterior is always unimodal.
0056For the E-step, generally, the posterior distribution over the unobserved variables is computed from the model distribution by Bayes' Rule, <maths id="MATH-US-00003" num="00003"><math overflow="scroll"><mtable><mtr><mtd><mrow><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mrow><mi>a</mi><mo>,</mo><mi>τ</mi><mo>,</mo><mi>r</mi><mo>,</mo><mi>v</mi><mo>,</mo><mi>l</mi><mo>,</mo><mrow><mi>s</mi><mo>|</mo><msub><mi>x</mi><mn>1</mn></msub></mrow><mo>,</mo><msub><mi>x</mi><mn>2</mn></msub><mo>,</mo><mi>y</mi><mo>,</mo><mi>θ</mi></mrow><mo>)</mo></mrow></mrow><mo>=</mo><mfrac><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mrow><msub><mi>x</mi><mn>1</mn></msub><mo>,</mo><msub><mi>x</mi><mn>2</mn></msub><mo>,</mo><mi>y</mi><mo>,</mo><mi>a</mi><mo>,</mo><mi>τ</mi><mo>,</mo><mi>r</mi><mo>,</mo><mi>v</mi><mo>,</mo><mi>l</mi><mo>,</mo><mrow><mi>s</mi><mo>|</mo><mi>θ</mi></mrow></mrow><mo>)</mo></mrow></mrow><mrow><mi>p</mi><mo></mo><mrow><mo>(</mo><mrow><msub><mi>x</mi><mn>1</mn></msub><mo>,</mo><msub><mi>x</mi><mn>2</mn></msub><mo>,</mo><mrow><mi>y</mi><mo>|</mo><mi>θ</mi></mrow></mrow><mo>)</mo></mrow></mrow></mfrac></mrow></mtd><mtd><mrow><mo>(</mo><mn>10</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> where p(x<sub>1</sub>, x<sub>2</sub>, y|θ) is obtained from the model distribution by marginalizing over the unobserved variables. It can be shown that the posterior distribution of the audio video tracker <b>130</b> has a factorized form, as does the distribution of the audio video tracker <b>130</b> (Eq. (7)). To describe it, a notation that uses q to denote a posterior distribution conditioned on the data can be used. Hence, <br /><i>p</i>(<i>a,τ,r,ν,l,s|x</i><sub>1</sub><i>,x</i><sub>2</sub><i>,y</i>,θ)=<i>q</i>(<i>a|τ,r</i>)<i>q</i>(ν|<i>l,s</i>)<i>q</i>(τ|<i>l</i>)<i>q</i>(<i>l,r,s</i>). (11)<br /> This factorized form follows from the audio video tracker <b>130</b>. The q notation omits the data, as well as the parameters. Therefore, q(a|τ, r)=p(a|τ, r, x<sub>1</sub>, x<sub>2</sub>, y, θ), and so on.
0057The functional forms of the posterior components q also follow from the model distribution. As the model is constructed from Gaussian components tied together by discrete variables, it can be shown that the audio posterior q(a|τ, r) and the video posterior q(ν|l, s) are both Gaussian, <maths id="MATH-US-00004" num="00004"><math overflow="scroll"><mtable><mtr><mtd><mtable><mtr><mtd><mrow><mrow><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mrow><mi>a</mi><mo>|</mo><mi>τ</mi></mrow><mo>)</mo></mrow></mrow><mo>=</mo><mrow><mi>𝒩</mi><mo></mo><mrow><mo>(</mo><mrow><mrow><mi>a</mi><mo>|</mo><msubsup><mi>μ</mi><mrow><mi>τ</mi><mo>,</mo><mi>r</mi></mrow><mi>a</mi></msubsup></mrow><mo>,</mo><msubsup><mi>v</mi><mi>r</mi><mi>a</mi></msubsup></mrow><mo>)</mo></mrow></mrow></mrow><mo>,</mo></mrow></mtd></mtr><mtr><mtd><mrow><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mrow><mrow><mi>v</mi><mo>|</mo><mi>l</mi></mrow><mo>,</mo><mi>s</mi></mrow><mo>)</mo></mrow></mrow><mo>=</mo><mrow><mi>𝒩</mi><mo></mo><mrow><mo>(</mo><mrow><mrow><mi>v</mi><mo>|</mo><msubsup><mi>μ</mi><mrow><mi>l</mi><mo>,</mo><mi>s</mi></mrow><mi>v</mi></msubsup></mrow><mo>,</mo><msubsup><mi>v</mi><mi>s</mi><mi>v</mi></msubsup></mrow><mo>)</mo></mrow></mrow></mrow></mtd></mtr></mtable></mtd><mtd><mrow><mo>(</mo><mn>12</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> The means μ<sub>τ,r</sub><sup>a</sup>, μ<sub>l,s</sub><sup>ν</sup> and precisions ν<sub>r</sub><sup>a</sup>, ν<sub>s</sub><sup>ν</sup> are straightforward to compute; note that the precisions do not depend on the shift variables τ, l. One particularly simple way to obtain them is to consider Eq. (11) and observe that its logarithms satisfies: <br />log <i>p</i>(<i>a,τ,r,ν,l,s|x</i><sub>1</sub><i>,x</i><sub>2</sub><i>,y</i>,θ)=log(<i>x</i><sub>1</sub><i>,x</i><sub>2</sub><i>,y,a,τ,r,ν,l,s</i>|θ)+const. (13)<br /> where the constant is independent of the hidden variables. Due to the nature of the model, this logarithm is quadratic in a and ν. To find the mean of the posterior over ν, the gradient of the log probability with respect to is set to zero. The precision is then given by the negative Hessian, leading to: <maths id="MATH-US-00005" num="00005"><math overflow="scroll"><mtable><mtr><mtd><mtable><mtr><mtd><mrow><mrow><msubsup><mi>μ</mi><mrow><mi>l</mi><mo>,</mo><mi>s</mi></mrow><mi>v</mi></msubsup><mo>=</mo><mrow><msup><mrow><mo>(</mo><msubsup><mi>v</mi><mi>s</mi><mi>v</mi></msubsup><mo>)</mo></mrow><mrow><mo>-</mo><mn>1</mn></mrow></msup><mo></mo><mrow><mo>(</mo><mrow><mrow><msub><mi>ϕ</mi><mi>s</mi></msub><mo></mo><msub><mi>μ</mi><mi>s</mi></msub></mrow><mo>+</mo><mrow><msubsup><mi>G</mi><mi>l</mi><mi>T</mi></msubsup><mo></mo><mi>ψ</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><mi>y</mi></mrow></mrow><mo>)</mo></mrow></mrow></mrow><mo>,</mo></mrow></mtd></mtr><mtr><mtd><mrow><msubsup><mi>v</mi><mi>s</mi><mi>v</mi></msubsup><mo>=</mo><mrow><msub><mi>ϕ</mi><mi>s</mi></msub><mo>+</mo><mrow><mi>ψ</mi><mo>.</mo></mrow></mrow></mrow></mtd></mtr></mtable></mtd><mtd><mrow><mo>(</mo><mn>14</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> Equations for the mean and precision of the posterior over a are obtained in a similar fashion. Another component of the posterior is the conditional probability table q(τ|l)=p (τ|l, x<sub>1</sub>, x<sub>2</sub>, y, θ), which turns out to be: <br />q(τ|<i>l</i>)∝<i>p</i>(τ|<i>l)exp(λ</i><sub>1</sub>λ<sub>2</sub>ν<sub>1</sub>ν<sub>2</sub>(ν<sub>r</sub><sup>a</sup>)<sup>−1</sup>c<sub>τ</sub>), (15)<br /> where <maths id="MATH-US-00006" num="00006"><math overflow="scroll"><mtable><mtr><mtd><mrow><msub><mi>c</mi><mi>τ</mi></msub><mo>=</mo><mrow><munder><mo>∑</mo><mi>n</mi></munder><mo></mo><mrow><msub><mi>x</mi><mrow><mn>1</mn><mo></mo><mi>n</mi></mrow></msub><mo></mo><msub><mi>x</mi><mrow><mn>2</mn><mo>,</mo><mrow><mi>n</mi><mo>+</mo><mi>τ</mi></mrow></mrow></msub></mrow></mrow></mrow></mtd><mtd><mrow><mo>(</mo><mn>16</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> is the cross-correlation between the audio input signals (e.g., microphone signal) x<sub>1 </sub>and x<sub>2</sub>. Finally, the last component of the posterior is the probability table q(l, r, s), whose form is omitted for brevity.
0058The calculation of q(τ|l) involves a minor but somewhat subtle point. The delay τ has generally been regarded as a discrete variable since the object tracker system <b>100</b> has been described in discrete time. In particular, q(τ|l) is a discrete probability table. However, for reasons of mathematical convenience, the model distribution p(τ|l) (Eq. (6)) treats τ as continuous. Accordingly, the posterior q(τ|l) computed by the algorithm of this implementation, is, strictly speaking, an approximation, as the true posterior in this model also treats τ as continuous. It turns out that this approximation is of the variational type. The rigorous derivation proceeds as follows. First, the form of the approximate posterior as a sum of delta functions is noted to be: <maths id="MATH-US-00007" num="00007"><math overflow="scroll"><mtable><mtr><mtd><mrow><mrow><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mrow><mi>τ</mi><mo>|</mo><mi>l</mi></mrow><mo>)</mo></mrow></mrow><mo>=</mo><mrow><munder><mo>∑</mo><mi>n</mi></munder><mo></mo><mrow><mrow><msub><mi>q</mi><mi>n</mi></msub><mo></mo><mrow><mo>(</mo><mi>l</mi><mo>)</mo></mrow></mrow><mo></mo><mrow><mi>δ</mi><mo></mo><mrow><mo>(</mo><mrow><mi>τ</mi><mo>-</mo><msub><mi>τ</mi><mi>n</mi></msub></mrow><mo>)</mo></mrow></mrow></mrow></mrow></mrow><mo>,</mo></mrow></mtd><mtd><mrow><mo>(</mo><mn>17</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> where the τ<sub>n </sub>are spaced one time point apart. The coefficients q<sub>n </sub>are nonnegative and sum up to one, and their dependence on l is initially unspecified. Next, the q<sub>n</sub>(l) is computed by minimizing the Kullback Leibler (KL) distance between the approximate posterior and the true posterior. This produces the optimal approximate posterior out of substantially all possible posteriors which satisfy the restriction of Eq. (17). For ease of notation, q(τ|l) will be utilized rather than q<sub>n</sub>(l).
0059Next, the M-step performs updates of the model parameters θ(Eq. (8)). The update rules are derived by considering the objective function: <br /><i>F</i>(θ)=<log <i>p</i>(<i>x</i><sub>1</sub><i>,x</i><sub>2</sub><i>,y,a,τ,r,ν,l,s</i>|θ)>, (18)<br /> known as the averaged complete data likelihood. The notation <.> will be used to denote averaging with respect to the posterior (Eq. (11)) over the hidden variables that do not appear on the left hand side and, in addition, averaging over all frames. Hence, F is essentially the log probability of the object tracker system <b>100</b> for each frame, where values for the hidden variables are filled in by the posterior distribution for that frame, followed by summing over frames. Each parameter update rule is obtained by setting the derivative of F with respect to that parameter to zero. For the video model <b>120</b> parameters μ<sub>s</sub>, φ<sub>s</sub>, π<sub>s</sub>, thus: <maths id="MATH-US-00008" num="00008"><math overflow="scroll"><mtable><mtr><mtd><mtable><mtr><mtd><mtable><mtr><mtd><mrow><mrow><msub><mi>μ</mi><mi>s</mi></msub><mo>=</mo><mfrac><mrow><mo><</mo><mrow><munder><mo>∑</mo><mi>l</mi></munder><mo></mo><mrow><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mrow><mi>l</mi><mo>,</mo><mi>s</mi></mrow><mo>)</mo></mrow></mrow><mo></mo><msubsup><mi>μ</mi><mrow><mi>l</mi><mo>,</mo><mi>s</mi></mrow><mi>v</mi></msubsup></mrow></mrow><mo>></mo></mrow><mrow><mo><</mo><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow></mrow><mo>></mo></mrow></mfrac></mrow><mo>,</mo></mrow></mtd></mtr><mtr><mtd><mrow><msubsup><mi>ϕ</mi><mi>s</mi><mrow><mo>-</mo><mn>1</mn></mrow></msubsup><mo>=</mo><mfrac><mrow><mo><</mo><mrow><mrow><munder><mo>∑</mo><mi>l</mi></munder><mo></mo><mrow><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mrow><mi>l</mi><mo>,</mo><mi>s</mi></mrow><mo>)</mo></mrow></mrow><mo></mo><msup><mrow><mo>(</mo><mrow><msubsup><mi>μ</mi><mrow><mi>l</mi><mo>,</mo><mi>s</mi></mrow><mi>v</mi></msubsup><mo>-</mo><msub><mi>μ</mi><mi>s</mi></msub></mrow><mo>)</mo></mrow><mn>2</mn></msup></mrow></mrow><mo>+</mo><mrow><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow></mrow><mo></mo><msup><mrow><mo>(</mo><msubsup><mi>v</mi><mrow><mi>l</mi><mo>,</mo><mi>s</mi></mrow><mi>v</mi></msubsup><mo>)</mo></mrow><mrow><mo>-</mo><mn>1</mn></mrow></msup></mrow></mrow><mo>></mo></mrow><mrow><mo><</mo><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow></mrow><mo>></mo></mrow></mfrac></mrow></mtd></mtr></mtable></mtd></mtr><mtr><mtd><mrow><mrow><msup><mi>π</mi><mi>s</mi></msup><mo>=</mo><mrow><mo><</mo><mrow><mi>q</mi><mo></mo><mrow><mo>(</mo><mi>s</mi><mo>)</mo></mrow></mrow><mo>></mo></mrow></mrow><mo>,</mo></mrow></mtd></mtr></mtable></mtd><mtd><mrow><mo>(</mo><mn>19</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> where the q's are computed by appropriate marginalizations over q(l, r, s) from the E-step. Notice that here, the notation <.> only average over frames. Update rules for the audio model <b>110</b> parameters η<sub>r</sub>, π<sub>r </sub>are obtained in a similar fashion.
0060For the audio video link parameters α, β, for simplicity, it has been assumed that α′=0, <maths id="MATH-US-00009" num="00009"><math overflow="scroll"><mtable><mtr><mtd><mtable><mtr><mtd><mtable><mtr><mtd><mrow><mi>α</mi><mo>=</mo><mfrac><mrow><mo><</mo><mrow><msub><mi>l</mi><mi>x</mi></msub><mo></mo><mi>τ</mi></mrow><mo>></mo><mrow><mo>-</mo><mrow><mo><</mo><mi>τ</mi><mo>></mo><mo><</mo><msub><mi>l</mi><mi>x</mi></msub><mo>></mo></mrow></mrow></mrow><mrow><mo><</mo><msubsup><mi>l</mi><mi>x</mi><mn>2</mn></msubsup><mo>></mo><mrow><mo>-</mo><mrow><mo><</mo><msub><mi>l</mi><mi>x</mi></msub><mo></mo><msup><mo>></mo><mn>2</mn></msup></mrow></mrow></mrow></mfrac></mrow></mtd></mtr><mtr><mtd><mrow><mi>β</mi><mo>=</mo><mrow><mo><</mo><mi>τ</mi><mo>></mo><mrow><mo>-</mo><mi>α</mi></mrow><mo><</mo><msub><mi>l</mi><mi>x</mi></msub><mo>></mo></mrow></mrow></mtd></mtr></mtable></mtd></mtr><mtr><mtd><mrow><mrow><msubsup><mi>v</mi><mi>τ</mi><mrow><mo>-</mo><mn>1</mn></mrow></msubsup><mo>=</mo><mrow><mo><</mo><msup><mi>τ</mi><mn>2</mn></msup><mo>></mo><mrow><mo>+</mo><msup><mi>α</mi><mn>2</mn></msup></mrow><mo><</mo><msubsup><mi>l</mi><mi>x</mi><mn>2</mn></msubsup><mo>></mo><mrow><mrow><mo>+</mo><msup><mi>β</mi><mn>2</mn></msup></mrow><mo>+</mo><mrow><mn>2</mn><mo></mo><mi>αβ</mi></mrow></mrow><mo><</mo><msub><mi>l</mi><mi>x</mi></msub><mo>></mo><mrow><mrow><mo>-</mo><mn>2</mn></mrow><mo></mo><mi>a</mi></mrow><mo><</mo><mrow><mi>τ</mi><mo></mo><mstyle><mtext> </mtext></mstyle><mo></mo><msub><mi>l</mi><mi>x</mi></msub></mrow><mo>></mo><mrow><mrow><mo>-</mo><mn>2</mn></mrow><mo></mo><mi>β</mi></mrow><mo><</mo><mi>τ</mi><mo>></mo></mrow></mrow><mo>,</mo></mrow></mtd></mtr></mtable></mtd><mtd><mrow><mo>(</mo><mn>20</mn><mo>)</mo></mrow></mtd></mtr></mtable></math></maths><br /> where in addition to averaging over frames, <.> here implies averaging for each frame with respect to q(τ, l) for that frame, which is obtained by marginalizing q(τ|l) q(l, r, s) over r, s.
0061It is be to appreciated that according to Eq. (19), computing the mean (μ<sub>s</sub>)<sub>n </sub>for each pixel n requires summing over substantially all possible spatial shifts l. Since the number of possible shifts equals the number of pixels, this seems to imply that the complexity of the algorithm of this implementation is quadratic in the number of pixels N. If that were the case, a standard N=120×160 pixel array would render the computation practically intractable. However, a more careful examination of Eq. (19), in combination with Eq. (14), shows that it can be written in the form of an inverse Fast Fourier Transform (FFT). Consequently, the actual complexity is not O(N<sup>2</sup>) but rather O (N log N). This result, which extends to the corresponding quantities in the audio model <b>110</b>, significantly increases the efficiency of the EM algorithm of this implementation.
0062Tracking of the object tracker system <b>100</b> is performed as part of the E-step using Eq. (9), where p(l|x<sub>1</sub>, x<sub>2</sub>, y) is computed from q(τ, l) above by marginalization. For each frame, the mode of this posterior distribution represents the most likely object position, and the width of the mode a degree of uncertainty in this inference. The object tracker system <b>100</b> can provide an output associated with the most likely object position. Further, the object tracker system <b>100</b> can be employed, for example, in a video conferencing system and/or a multi-media processing system.
0063Those skilled in the art will recognize that other variants of the system can be employed in accordance with the present invention and all such variants are intended to fall within the scope of the appended claims. For example, one variant of the system is based on incorporating temporal correlation(s) between the audio-video data, for instance, the location l at time n depends on the location at time n−1, and the speech component r at time n depends on the speech component at time n−1. Another variant is based on the system including a variable for the background image (e.g., the background against which the object is moving), modeling the mean background and its variability.
0064While <figref idref="DRAWINGS">FIG. 1</figref> is a block diagram illustrating components for the system <b>100</b>, it is to be appreciated that the audio model <b>110</b>, the video model <b>120</b> and/or the audio video tracker <b>130</b> can be implemented as one or more computer components, as that term is defined herein. Thus, it is to be appreciated that computer executable components operable to implement the audio model <b>110</b>, the video model <b>120</b> and/or the audio video tracker <b>130</b> can be stored on computer readable media including, but not limited to, an ASIC (application specific integrated circuit), CD (compact disc), DVD (digital video disk), ROM (read only memory), floppy disk, hard disk, EEPROM (electrically erasable programmable read only memory) and memory stick in accordance with the present invention.
0065Turning to <figref idref="DRAWINGS">FIG. 7</figref>, an object tracker system <b>700</b> includes an audio model <b>110</b>, a video model <b>120</b> and an audio video tracker <b>130</b>. The system <b>700</b> can further include a first audio input device <b>140</b><sub>l </sub>through an Mth audio input device <b>140</b><sub>M</sub>, M being an integer greater than or equal to two. The first audio input device <b>140</b><sub>l </sub>through the Mth audio input device <b>140</b><sub>M </sub>can be referred to collectively as the audio input devices <b>140</b>. Additionally and/or alternatively, the system <b>700</b> can further includes a video input device <b>150</b>.
0066The audio input devices <b>140</b> can be include, for example, a microphone, a telephone and/or a speaker phone. The video input device <b>150</b> can include a camera.
0067In view of the exemplary systems shown and described above, methodologies that may be implemented in accordance with the present invention will be better appreciated with reference to the flow chart of <figref idref="DRAWINGS">FIGS. 8 and 9</figref>. While, for purposes of simplicity of explanation, the methodologies are shown and described as a series of blocks, it is to be understood and appreciated that the present invention is not limited by the order of the blocks, as some blocks may, in accordance with the present invention, occur in different orders and/or concurrently with other blocks from that shown and described herein. Moreover, not all illustrated blocks may be required to implement the methodologies in accordance with the present invention.
0068The invention may be described in the general context of computer-executable instructions, such as program modules, executed by one or more components. Generally, program modules include routines, programs, objects, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically the functionality of the program modules may be combined or distributed as desired in various embodiments.
0069Turning to <figref idref="DRAWINGS">FIGS. 8 and 9</figref>, a method <b>800</b> for object tracking in accordance with an aspect of the present invention is illustrated. At <b>810</b>, an audio model having trainable parameters, and, an original audio signal of an object, a time delay between at least two audio input signals and a variability component of the original audio signal as unobserved variables is provided. At <b>820</b>, a video model having trainable parameters, and, a location of the object, an original image of the object and a variability component of the original image as unobserved variables is provided.
0070At <b>830</b>, at least two audio input signals are received. At <b>840</b>, a video input signal is received. At <b>850</b>, a posterior distribution over the unobserved variables of the audio model and the video model is updated (e.g., based on Eqs. (10)-(17)). At <b>860</b>, trainable parameters of the audio model and the video model are updated (e.g., based on Eqs. (18)-(20)). At <b>870</b>, an output associated with the location of the object is provided. Thereafter, processing continues at <b>830</b>.
0071In order to provide additional context for various aspects of the present invention, FIG. <b>10</b> and the following discussion are intended to provide a brief, general description of a suitable operating environment <b>1010</b> in which various aspects of the present invention may be implemented. While the invention is described in the general context of computer-executable instructions, such as program modules, executed by one or more computers or other devices, those skilled in the art will recognize that the invention can also be implemented in combination with other program modules and/or as a combination of hardware and software. Generally, however, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular data types. The operating environment <b>1010</b> is only one example of a suitable operating environment and is not intended to suggest any limitation as to the scope of use or functionality of the invention. Other well known computer systems, environments, and/or configurations that may be suitable for use with the invention include but are not limited to, personal computers, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include the above systems or devices, and the like.
0072With reference to <figref idref="DRAWINGS">FIG. 10</figref>, an exemplary environment <b>1010</b> for implementing various aspects of the invention includes a computer <b>1012</b>. The computer <b>1012</b> includes a processing unit <b>1014</b>, a system memory <b>1016</b>, and a system bus <b>1018</b>. The system bus <b>1018</b> couples system components including, but not limited to, the system memory <b>1016</b> to the processing unit <b>1014</b>. The processing unit <b>1014</b> can be any of various available processors. Dual microprocessors and other multiprocessor architectures also can be employed as the processing unit <b>1014</b>.
0073The system bus <b>1018</b> can be any of several types of bus structure(s) including the memory bus or memory controller, a peripheral bus or external bus, and/or a local bus using any variety of available bus architectures including, but not limited to, 10-bit bus, Industrial Standard Architecture (ISA), Micro-Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics (IDE), VESA Local Bus (VLB), Peripheral Component Interconnect (PCI), Universal Serial Bus (USB), Advanced Graphics Port (AGP), Personal Computer Memory Card International Association bus (PCMCIA), and Small Computer Systems Interface (SCSI).
0074The system memory <b>1016</b> includes volatile memory <b>1020</b> and nonvolatile memory <b>1022</b>. The basic input/output system (BIOS), containing the basic routines to transfer information between elements within the computer <b>1012</b>, such as during start-up, is stored in nonvolatile memory <b>1022</b>. By way of illustration, and not limitation, nonvolatile memory <b>1022</b> can include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory <b>1020</b> includes random access memory (RAM), which acts as external cache memory. By way of illustration and not limitation, RAM is available in many forms such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM).
0075Computer <b>1012</b> also includes removable/nonremovable, volatile/nonvolatile computer storage media. <figref idref="DRAWINGS">FIG. 10</figref> illustrates, for example a disk storage <b>1024</b>. Disk storage <b>1024</b> includes, but is not limited to, devices like a magnetic disk drive, floppy disk drive, tape drive, Jaz drive, Zip drive, LS-100 drive, flash memory card, or memory stick. In addition, disk storage <b>1024</b> can include storage media separately or in combination with other storage media including, but not limited to, an optical disk drive such as a compact disk ROM device (CD-ROM), CD recordable drive (CD-R Drive), CD rewritable drive (CD-RW Drive) or a digital versatile disk ROM drive (DVD-ROM). To facilitate connection of the disk storage devices <b>1024</b> to the system bus <b>1018</b>, a removable or non-removable interface is typically used such as interface <b>1026</b>.
0076It is to be appreciated that <figref idref="DRAWINGS">FIG. 10</figref> describes software that acts as an intermediary between users and the basic computer resources described in suitable operating environment <b>1010</b>. Such software includes an operating system <b>1028</b>. Operating system <b>1028</b>, which can be stored on disk storage <b>1024</b>, acts to control and allocate resources of the computer system <b>1012</b>. System applications <b>1030</b> take advantage of the management of resources by operating system <b>1028</b> through program modules <b>1032</b> and program data <b>1034</b> stored either in system memory <b>1016</b> or on disk storage <b>1024</b>. It is to be appreciated that the present invention can be implemented with various operating systems or combinations of operating systems.
0077A user enters commands or information into the computer <b>1012</b> through input device(s) <b>1036</b>. Input devices <b>1036</b> include, but are not limited to, a pointing device such as a mouse, trackball, stylus, touch pad, keyboard, microphone, joystick, game pad, satellite dish, scanner, TV tuner card, digital camera, digital video camera, web camera, and the like. These and other input devices connect to the processing unit <b>1014</b> through the system bus <b>1018</b> via interface port(s) <b>1038</b>. Interface port(s) <b>1038</b> include, for example, a serial port, a parallel port, a game port, and a universal serial bus (USB). Output device(s) <b>1040</b> use some of the same type of ports as input device(s) <b>1036</b>. Thus, for example, a USB port may be used to provide input to computer <b>1012</b>, and to output information from computer <b>1012</b> to an output device <b>1040</b>. Output adapter <b>1042</b> is provided to illustrate that there are some output devices <b>1040</b> like monitors, speakers, and printers among other output devices <b>1040</b> that require special adapters. The output adapters <b>1042</b> include, by way of illustration and not limitation, video and sound cards that provide a means of connection between the output device <b>1040</b> and the system bus <b>1018</b>. It should be noted that other devices and/or systems of devices provide both input and output capabilities such as remote computer(s) <b>1044</b>.
0078Computer <b>1012</b> can operate in a networked environment using logical connections to one or more remote computers, such as remote computer(s) <b>1044</b>. The remote computer(s) <b>1044</b> can be a personal computer, a server, a router, a network PC, a workstation, a microprocessor based appliance, a peer device or other common network node and the like, and typically includes many or all of the elements described relative to computer <b>1012</b>. For purposes of brevity, only a memory storage device <b>1046</b> is illustrated with remote computer(s) <b>1044</b>. Remote computer(s) <b>1044</b> is logically connected to computer <b>1012</b> through a network interface <b>1048</b> and then physically connected via communication connection <b>1050</b>. Network interface <b>1048</b> encompasses communication networks such as local-area networks (LAN) and wide-area networks (WAN). LAN technologies include Fiber Distributed Data Interface (FDDI), Copper Distributed Data Interface (CDDI), Ethernet/IEEE 1002.3, Token Ring/IEEE 1002.5 and the like. WAN technologies include, but are not limited to, point-to-point links, circuit switching networks like Integrated Services Digital Networks (ISDN) and variations thereon, packet switching networks, and Digital Subscriber Lines (DSL).
0079Communication connection(s) <b>1050</b> refers to the hardware/software employed to connect the network interface <b>1048</b> to the bus <b>1018</b>. While communication connection <b>1050</b> is shown for illustrative clarity inside computer <b>1012</b>, it can also be external to computer <b>1012</b>. The hardware/software necessary for connection to the network interface <b>1048</b> includes, for exemplary purposes only, internal and external technologies such as, modems including regular telephone grade modems, cable modems and DSL modems, ISDN adapters, and Ethernet cards.
0080What has been described above includes examples of the present invention. It is, of course, not possible to describe every conceivable combination of components or methodologies for purposes of describing the present invention, but one of ordinary skill in the art may recognize that many further combinations and permutations of the present invention are possible. Accordingly, the present invention is intended to embrace all such alterations, modifications and variations that fall within the spirit and scope of the appended claims. Furthermore, to the extent that the term “includes” is used in either the detailed description or the claims, such term is intended to be inclusive in a manner similar to the term “comprising” as “comprising” is interpreted when employed as a transitional word in a claim.
Contents5
20 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US7475014B2 | Cited by | United States of America | Search report |
| US8238562B2 | Cited by | United States of America | Applicant |
| US2006153408A1 | Cited by | United States of America | Pre-grant |
| US7805313B2 | Cited by | United States of America | Applicant |
| US2003236583A1 | Cited by | United States of America | Pre-grant |
| US2009319281A1 | Cited by | United States of America | Pre-grant |
| US2019141445A1 | Cited by | United States of America | Search report |
| US8024189B2 | Cited by | United States of America | Applicant |
| US2006083385A1 | Cited by | United States of America | Pre-grant |
| US10489917B2 | Cited by | United States of America | Applicant |
| US8204261B2 | Cited by | United States of America | Applicant |
| US2019141445A1 | Cited by | United States of America | Search report |
| US7693721B2 | Cited by | United States of America | Applicant |
| US2009319282A1 | Cited by | United States of America | Pre-grant |
| US2006125953A1 | Cited by | United States of America | Pre-grant |
| US2011164756A1 | Cited by | United States of America | Pre-grant |
| US2009002489A1 | Cited by | United States of America | Pre-grant |
| US9741129B2 | Cited by | United States of America | Search report |
| US2011007158A1 | Cited by | United States of America | Pre-grant |
| US2003035553A1 | Cited by | United States of America | Pre-grant |
| US2005058304A1 | Cited by | United States of America | Pre-grant |
| US2005195981A1 | Cited by | United States of America | Pre-grant |
| US2008294375A1 | Cited by | United States of America | Pre-grant |
| US2014257968A1 | Cited by | United States of America | Pre-grant |
| US9538156B2 | Cited by | United States of America | Applicant |
| US2004105004A1 | Cited by | United States of America | Pre-grant |
| US2007003069A1 | Cited by | United States of America | Pre-grant |
| US10951859B2 | Cited by | United States of America | Applicant |
| US2006115100A1 | Cited by | United States of America | Pre-grant |
| US7941320B2 | Cited by | United States of America | Applicant |
| US2007291104A1 | Cited by | United States of America | Pre-grant |
| US7583805B2 | Cited by | United States of America | Applicant |
| US2005140779A1 | Cited by | United States of America | Pre-grant |
| US8200500B2 | Cited by | United States of America | Applicant |
| US7720230B2 | Cited by | United States of America | Applicant |
| US2007033045A1 | Cited by | United States of America | Pre-grant |
| US7403217B2 | Cited by | United States of America | Applicant |
| US7903824B2 | Cited by | United States of America | Applicant |
| US8510110B2 | Cited by | United States of America | Applicant |
| US8340306B2 | Cited by | United States of America | Applicant |
| US7400357B2 | Cited by | United States of America | Search report |
| US7761304B2 | Cited by | United States of America | Applicant |
| US7644003B2 | Cited by | United States of America | Applicant |
| US2009150161A1 | Cited by | United States of America | Pre-grant |
| US10887690B2 | Cited by | United States of America | Search report |
| WO2012103649A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US2006085200A1 | Cited by | United States of America | Pre-grant |
| US2008130904A1 | Cited by | United States of America | Pre-grant |
| US7349008B2 | Cited by | United States of America | Search report |
| US2005180579A1 | Cited by | United States of America | Pre-grant |
| US7292901B2 | Cited by | United States of America | Applicant |
| US7787631B2 | Cited by | United States of America | Applicant |
| US2002093591A1 | Cites | United States of America | Search report |
| US2002101505A1 | Cites | United States of America | Search report |
| US2003167148A1 | Cites | United States of America | Search report |
| US6005610A | Cites | United States of America | Search report |
| US6014167A | Cites | United States of America | Applicant |
| US6275258B1 | Cites | United States of America | Applicant |
| US6766035B1 | Cites | United States of America | Search report |
| US6801656B1 | Cites | United States of America | Search report |
8 members in 2 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 18357502 | United States of America | A | |
| US20020183575 | – | – | – |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| US2004001143A1 | United States of America | A1 | |
| EP1377057A2 | European Patent Office (EPO) | A2 | |
| EP1377057A3 | European Patent Office (EPO) | A3 | |
| US2005171971A1 | United States of America | A1 | |
| US6940540B2This record | United States of America | B2 | |
| US7692685B2 | United States of America | B2 | |
| US2010194881A1 | United States of America | A1 | |
| US8842177B2 | United States of America | B2 |
31 transactions on the USPTO file
Allowed after 1 non-final rejection.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | |
|---|---|
| Change in Power of Attorney (May Include Associate POA) | |
| Correspondence Address Change | |
| Post Issue Communication - Certificate of Correction | |
| Email Notification | |
| Mail Miscellaneous Communication to Applicant | |
| Miscellaneous Communication to Applicant - No Action Count | |
| Recordation of Patent Grant Mailed | |
| Patent Issue Date Used in PTA CalculationAllowed | |
| Issue Notification MailedAllowed | |
| Receipt into Pubs | |
| Dispatch to FDC | |
| Application Is Considered Ready for Issue | |
| Receipt into Pubs | |
| Issue Fee Payment Verified | |
| Issue Fee Payment Received | |
| Workflow - File Sent to Contractor | |
| Mail Notice of AllowanceAllowed | |
| Notice of Allowance Data Verification CompletedAllowed | |
| Date Forwarded to Examiner | |
| Response after Non-Final Action | |
| Mail Non-Final RejectionNon-final rejection | |
| Non-Final RejectionNon-final rejection | |
| IFW TSS Processing by Tech Center Complete | |
| Case Docketed to Examiner in GAU | |
| Reference capture on IDS | |
| Information Disclosure Statement (IDS) Filed | |
| Information Disclosure Statement (IDS) Filed | |
| Application Dispatched from OIPE | |
| Application Is Now Complete | |
| IFW Scan & PACR Auto Security Review | |
| Initial Exam Team nn |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Fee paymentFPAY | FPAY | |
| AssignmentAS | AS | |
| Fee paymentFPAY | FPAY | |
| Fee paymentFPAY | FPAY | |
| Certificate of correctionCC | CC | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS |
Numbers
- Publication
- 06940540
- Publication, DOCDB
- 6940540
- Publication, EPODOC
- US6940540
- Application
- 10183575
- Application, DOCDB
- 18357502
- Application, EPODOC
- US20020183575
Titles
- English
- Speaker detection and tracking using audiovisual data
Patent term adjustment
- A delay
- +447 daysthe office missed an examination deadline
- Net adjustment
- 447 days
Classification
- CPC, 5
- H04N7/15
- G06V10/24
- G06V10/811
- G06F2218/22
- G06F18/256
- IPC, 5
- G06F15 00
- G06F17 00
- G06V10 24
- H04N5 225
- H04N7 15
- USPC, 3
- 348169000
- 348E07083
- 702181000