Image stabilizing apparatus, image pick-up apparatus and image stabilizing method
Abstract
An image stabilizing apparatus includes a motion vector calculating part (104) that calculates a motion vector between a plurality of images including a displacement caused by a motion of an image-pickup apparatus, a shake-correction parameter calculating part (106) that receives the motion vector as input to calculate a shake correction amount, and an image transforming part (107) that performs geometric transformation of the image in accordance with the shake correction amount. The shake-correction parameter calculating part performs variation amount calculation (S502), variation amount correction (S503) and correction amount calculation (SS04) based on the motion information between the plurality of images. The image stabilizing apparatus preserves a motion in video from an intended camera work and allows image stabilization for an unintended shake.

Term
1.5 yearsto projected expiry
Projected expiry 31 March 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
15 claims: 4 independent, 11 dependent
- 1An image stabilizing apparatus comprising:a motion vector calculating part (104) that calculates a motion vector between a plurality of images including a displacement caused by a motion of an image-pickup apparatus;a shake-correction parameter calculating part (106) that receives the motion vector as input to calculate a shake correction amount;and an image transforming part (107) that performs geometric transformation of the image in accordance with the shake correction amount, characterized in that the shake-correction parameter calculating part includes: a variation amount calculating part (S502) that calculates an image variation amount between the plurality of images based on the motion vector;a variation amount correcting part (S503) that calculates, based on the image variation amount, motion information in which a component distorting the image is excluded;and a correction amount calculating part (S504) that calculates the shake correction amount based on the motion information between the plurality of images.
- 10An image stabilizing apparatus comprising:a motion vector calculating part (104) that calculates a motion vector that represents displacement of corresponding feature portions between a plurality of images, the displacement being caused by a motion of an image-pickup apparatus;a shake-correction parameter calculating part (106) that receives the motion vector as input to calculate a shake correction amount;and an image transforming part (107) that performs geometric transformation of the image in accordance with the shake correction amount, characterized in that the shake-correction parameter calculating part includes: a variation amount calculating part (S502) that calculates an image variation amount between the plurality of images based on the motion vector;a variation amount correcting part (S503) that calculates motion information based on the image variation amount between the plurality of images;and a correction amount calculating part (S504) that calculates the shake correction amount based on the motion information.
- 14An image-stabilizing method comprising:a motion vector calculating step (S301) of calculating a motion vector between a plurality of images including a displacement caused by a motion of an image-pickup apparatus;a shake-correction parameter calculating step (S302) of receiving the motion vector as input to calculate a shake correction amount;and an image transforming step (S303) of performing geometric transformation of the image in accordance with the shake correction amount, characterized in that the shake-correction parameter calculating step includes: a variation amount calculating step (S502) of calculating an image variation amount between the plurality of images based on the motion vector;a variation amount correcting step (S503) of calculating, based on the image variation amount, motion information in which a component distorting the image is excluded;and a correction amount calculating step (S504) of calculating the shake correction amount based on the motion information between the plurality of images.
- 15An image-stabilizing method comprising:a motion vector calculating step (S301) of calculating a motion vector that represents displacement of corresponding feature portions between a plurality of images, the displacement being caused by a motion of an image-pickup apparatus;a shake-correction parameter calculating step (S302) of receiving the motion vector as input to calculate a shake correction amount;and an image transforming step (S303) of performing geometric transformation of the image in accordance with the shake correction amount, characterized in that the shake-correction parameter calculating step includes: a variation amount calculating step (S502) of calculating an image variation amount between the plurality of images based on the motion vector;a variation amount correcting step (S503) of calculating motion information based on the image variation amount between the plurality of images;and a correction amount calculating step (S504) of calculating the shake correction amount based on the motion information.
Independent claims6
324 paragraphs in 4 sections, as filed
BACKGROUND OF THE INVENTION
0001The present invention relates to an image stabilizing apparatus and an image stabilizing method for performing image-stabilization processing in moving images. The present invention also relates to an image-pickup apparatus on which the image stabilizing apparatus is mounted.
0002Image stabilizing techniques involving image processing for reducing shakes of video (moving image) due to camera shakes are widely used in image-pickup apparatuses for moving images such as video cameras. Especially when an image-pickup optical system of a long focal length is used to pick up images, a slight camera shake leads to a violent shake of video and thus an image stabilizing function is essential for the camera. Even for an image-pickup optical system of a short focal length, effective operation of the image stabilizing function is desirable when a user attempts to pick up an object image while the user is moving.
0003When picking up images while the use is moving, an advanced image stabilizing function is necessary in which an unintended camera shake is discriminated from an intended camera work and only the image shake due to the camera shake is suppressed. Already proposed image stabilizing techniques for supporting such movement include use of inertial motion filtering (see, <nplcit id="ncit0001" npl-type="s"><text>Z. Zhu, et al. "Camera stabilization based on 2.5D motion estimation and inertial filtering," ICIV, 1998</text></nplcit>), and use of low-order model fitting (see, <nplcit id="ncit0002" npl-type="s"><text>A. Ltvin, J. Konrad, W.C. Karl, "Probabilistic video stabilization using Kalman filtering and mosaicing," Proceedings of SPIE. Jan. 2003, p.p. 20-24</text></nplcit>).
0004In these image stabilizing techniques, an approximation model such as a translation model and a Helmert model (similarity model) is used in motion estimation from images (estimation of a global motion or a camera work). A motion estimate value is given as one-dimensional time-series data set corresponding to camera work components such as horizontal translation, vertical translation, in-plane rotation, scaling, and shear. Thus, a filtering mechanism such as an inertial filter and a Kalman filter, which receives as input the one-dimensional time-series data set for use in signal processing can be used without any change.
0005Since motions in the image are in one-to-one correspondence with camera works, intended image stabilization is realized simply by causing the result of the filtering of the abovementioned motion amount determined from the image or the difference between the original motion amount and the filtering result to act on the image.
0006An image-pickup apparatus of a short focal length may be mounted on a walking robot, a helicopter, or a wearable camera which can violently shake. As the focal length is further reduced, motions appearing in a picked-up image are changed.
0007Specifically, the degree of the camera work allowable for a motion in the picked-up image is inversely proportional to the focal length, so that a motion referred to as "foreshortening" occurs which is not seen in video at an immediate focal length, thereby making it impossible to achieve image stabilization by the estimation of an image variation amount in the conventional approximation model and the image-stabilization processing. To address this, a proposal has been made in which the estimation of an image variation amount is performed with a projective model instead of the abovementioned approximation model and geometric correction in the image-stabilization processing is performed with projective transformation.
0008The abovementioned uses are based on the premise that the motion estimation is performed from the image and the image-stabilization processing is performed from a combination of image geometric transformation. In this case, it is necessary to accurately detect an image variation in response to a large and complicated camera work and to correct a large motion based on the movement of the user. However, it is difficult for only a motion sensor often used in the conventional image-stabilization processing to sense a multi-axis variation with high accuracy and at low cost. In addition, optical image-stabilization processing cannot correct violent shakes.
0009When image-stabilization processing of video which is picked up by a moving user is performed with the projective model, the following problems arise.
0010One of the problems is that the filtering method based on the conventional signal processing technique does not appropriately function as it is in the discrimination between an intended camera work and an unintended shake (motion). This is because a projective homography representing the image variation amount is a multi-dimensional amount represented by a matrix of 3 x 3.
0011One component of the projective homography is affected by a plurality of camera works. Thus, especially when a large forward camera work occurs, appropriate image stabilization cannot be achieved even when the camera work corresponds to a linear motion at a constant speed. This is because variation of each term component in terms of the homography is the linear sum of a non-linear image variation by the forward camera work and a linear image variation by a camera work such as translation and rotation that are perpendicular to an optical axis. As a result, even when filtering premised on a linear change is applied to each term of the projective homography, appropriate image-stabilization effects cannot be provided.
0012Second, one of the problems results from the extension of the estimation of the image variation amount and the image-stabilization processing to the projective model. The extension to the projective model allows detection of the image variation due to a large rotational camera work. Conversely, if the projective homography determined from motion vectors between frame images constituting video is inversely transformed directly or through the motion determination and then is used as a shake correction amount, appropriate image stabilization cannot- be achieved. The image stabilizing method is widely used in image stabilization with the approximation model.
0013However, in the projective model, the influence of a translation camera work upon the image variation amount is relevant to the orientation of a reference plane associated with spatial distribution of motion vector extraction points in calculating a new projective homography. This causes the problem.
0014The relationship between a projective homography representing an image variation amount between frame images, a camera work, and a reference plane is expressed as follows: <maths id="math0001"><math display="block"><mi>H</mi><mo>=</mo><mi>R</mi><mo>+</mo><mfrac><mn>1</mn><mi>d</mi></mfrac><mo></mo><mover><mi>t</mi><mo>→</mo></mover><mo></mo><msup><mover><mi>n</mi><mo>→</mo></mover><mi>T</mi></msup></math><img file="EP1978731A2_D0001.tif" /></maths>
0015where <i>H</i> represents the projective homography, <i>R</i> and <o ostyle="rightarrow"><i>t</i></o> represent rotation and translation of the camera, respectively, and d and n represent the distance betweeen the reference plane determined by the spatial positions of corresponding points and one camera, and the orientation of the normal to the reference plane, respectively.
0016Since the reference plane provided by the spatial positions of corresponding points for which a motion vector is extracted is often different from the position of a plane in space for which an observer wishes image stabilization, a problem arises. As seen from the abovementioned expression, the problem occurs only when the translation camera work is performed. For example, the problem involves distortion of the image in which an image plane is inclined in an advancing scene or an image plane is collapsed in a panning scene.
0017It is possible to adopt a compromise in which image stabilization is performed by using only triaxial rotation information of a camera work determined from a motion vector between frame images as proposed in <nplcit id="ncit0003" npl-type="s"><text>Michal Irani, et al. "Recovery of Ego-Motion Using Image Stabilization," CVPR ('94), Seattle, Jun 1994</text></nplcit>. However, a camera work of a translation component from an up-and-down motion caused by a walking shake is not ignorable as a motion for which image stabilization should be performed.
SUMMARY OF THE INVENTION
0018The present invention in its first aspect provides an image stabilizing apparatus as specified in claims 1 to 9.
0019The present invention in its second aspect provides an image stabilizing apparatus as specified in claims 10 to 12.
0020The present invention in its third aspect provides an image-pickup apparatus as specified in claim 13.
0021The present invention in its fourth aspect provides an image-pickup apparatus as specified in claim 14.
0022The present invention in its fifth aspect provides an image stabilizing method as specified in claim 15.
0023Other aspects of the present invention will be apparent from the embodiments described below with reference to the drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
0024<figref idref="f0001">Fig. 1</figref> is a diagram showing the configuration of a video camera which is Embodiments 1 to 4, 6, and 7 of the present invention.
0025<figref idref="f0002">Fig. 2</figref> is a schematic diagram for explaining the mechanism of a work memory in the embodiments.
0026<figref idref="f0003">Fig. 3</figref> is a flow chart showing an image stabilization processing procedure in Embodiment 1.
0027<figref idref="f0004">Fig. 4</figref> is a diagram for explaining block matching performed in Embodiment 1.
0028<figref idref="f0005">Fig. 5</figref> is a flow chart showing a shake correction parameter calculating procedure in Embodiment 1.
0029<figref idref="f0006">Fig. 6</figref> is a flow chart showing a correction amount calculating procedure in Embodiment 1.
0030<figref idref="f0007">Fig. 7</figref> is a schematic diagram showing a digital filtering procedure in the correction amount calculating procedure in Embodiment 1.
0031<figref idref="f0008">Fig. 8</figref> is a schematic diagram showing a procedure for handling a projective homography which is a multi-dimensional amount in the digital filtering of Embodiment 1.
0032<figref idref="f0009">Fig. 9</figref> is a flow chart showing a correction amount calculating procedure in Embodiment 2.
0033<figref idref="f0010">Fig. 10</figref> is a flow chart showing a correction amount calculating procedure in Embodiment 3.
0034<figref idref="f0011">Fig. 11</figref> is a flow chart showing a shake correction parameter calculating procedure in Embodiment 3.
0035<figref idref="f0012">Fig. 12</figref> is a flow chart showing a correction amount calculating procedure in Embodiment 4.
0036<figref idref="f0013">Fig. 13</figref> is a diagram showing the configuration of an image stabilizing apparatus which is Embodiment 5.
0037<figref idref="f0014">Fig. 14</figref> is a flow chart showing a correction amount calculating procedure in Embodiment 6.
0038<figref idref="f0015">Fig. 15</figref> is a diagram showing an exemplary display menu in Embodiment 6.
0039<figref idref="f0015">Fig. 16</figref> is a diagram showing another exemplary display menu in Embodiment 6.
0040<figref idref="f0016">Fig. 17</figref> is a flow chart showing a shake correction parameter calculating procedure in Embodiment 7.
0041<figref idref="f0017">Fig. 18</figref> is a flow chart showing a digital filtering procedure in the shake correction parameter calculating procedure in Embodiment 1.
0042<figref idref="f0018">Fig. 19</figref> is a flow chart showing the digital filtering procedure in Embodiment 1 in detail.
DESCRIPTION OF THE EMBODIMENTS
0043Exemplary embodiments of the present invention will be described below with reference to the accompanied drawings.
[Embodiment 1]
0044<figref idref="f0001">Fig. 1</figref> shows the configuration of a video camera (image-pickup apparatus) serving as a video input apparatus on which an image stabilizing apparatus which is Embodiment 1 of the present invention is mounted. In <figref idref="f0001">Fig. 1</figref>, reference numeral 101 shows a lens optical system serving as an image-pickup optical system and reference numeral 102 shows an image-pickup element such as a CCD sensor and a CMOS sensor.
0045Reference numeral 103 shows a preprocessing part, 104 a motion vector detecting part (motion vector calculating part), and 105 a work memory. Reference numeral 106 shows a shake-correction parameter calculating part, 107 a geometric transformation processing part (image transforming part), 108 an encoding/decoding part, and 109 a work memory. Reference numeral 110 shows a system controlling part, 111 a zoom adjusting part, 112 a non-volatile memory part, 113 a recording part, 114 a displaying part, 115 an operation signal inputting part, and 116 an external I/F.
0046The preprocessing part 103, the motion vector detecting part 104, the shake-correction parameter calculating part 106, the geometric transformation processing part 107, and the encoding/decoding part 108 constitute a video signal processing section. The video signal processing part forms a main part of the image stabilizing apparatus.
0047The lens optical system 101 includes a plurality of lenses and forms an optical image of an object (object image).
0048The image-pickup element 102 photoelectrically converts the optical image formed on a light-receiving surface by the lens optical system 101 into an image-pickup signal.
0049The preprocessing part 103 performs video processing on the image-pickup signal output from the image-pickup element 102 and outputs a video signal after the processing. The image-pickup element 102 and the preprocessing part 103 constitute an image-pickup system which photoelectrically converts the object image into the image (video). The video processing performed by the preprocessing part 103 includes autogain control, luminance/color difference separation, sharpening, white balance adjustment, black level adjustment, and colorimetric system transformation.
0050The motion vector detecting part 104 receives, as input, video frames (frame images) such as successive luminance frames, successive luminance and color difference frames, or successive RGB frames transformed from luminance and color difference frames provided by the preprocessing part 103. It may receive, as input, differential processing frames processed for motion vector detection or binary code frames.
0051The motion vector detecting part 104 detects motion vectors (motion information) between the successive frames input thereto. Specifically, it calculates the motion vector between a present frame image input from the preprocessing part 103, that is, a current frame, and a previous frame image input previously and accumulated in the work memory 105, that is, a past frame. The past frame is a frame subsequent to the current frame or a much older frame.
0052The work memory 105 is, for example, a socalled FIFO (first in, first out) memory. A delay amount of output is controlled on the basis of the number of memory blocks of the work memory 105.
0053<figref idref="f0002">Fig. 2</figref> schematically shows a FIFO memory formed of two blocks. The FIFO memory is constituted with a list arrangement in terms of program, and insertion (push) operation and extraction (pop) operation are performed simultaneously. When the current frame is pushed as an n-th frame, the pop operation is simultaneously performed in which a previously pushed n-2th frame overflows and is output from the memory. If the number of memory block is three, an n-3th frame is output. In this manner, the number of memory blocks controls the delay relationship between the pushed frame and the popped frame.
0054The shake-correction parameter calculating part 106 receives, as input, the motion vector output from the motion vector detecting part 104 and camera calibration information such as in-camera parameters and a distortion coefficient provided by the system controlling part 110, later described, to calculate a shake correction amount. Although described later in detail, the shake-correction parameter calculating part 106 serves as a variation amount calculating part which calculates an image variation amount between a plurality of images (frame images) based on the motion vector and a variation amount correcting part which calculates, based on the image variation amount, motion information in which a component distorting the image is excluded. It also serves as a correction amount calculating part which calculates the shake correction amount based on the motion information between the frame images.
0055The in-camera parameters include a focal length, a pixel size of the image-pickup element 102, an image offset, and a shear amount.
0056The focal length is a focal length of the lens optical system 101 and is associated with a zoom state of the lens optical system 101 in picking up the frame images.
0057The pixel size of the image-pickup element 102 is a size of each pixel in horizontal and vertical directions.
0058The offset is provided to set the center of the image crossed by the optical axis of the lens optical system 101 on an image plane as the origin of image coordinates in contrast to a typical case where the upper-left point of the image is regarded as the origin.
0059The shear represents distortion of a pixel resulting from the shape of the pixel or the fact that the optical axis is not orthogonal to the image plane. The distortion coefficient represents a distortion amount due to aberration of the lens optical system 101.
0060The shake-correction parameter calculating part 106 outputs shake correction parameters including the calculated shake correction amount, the in-camera parameters, and the distortion coefficient.
0061The geometric transformation processing part 107 receives, as input, the shake correction parameters calculated by the shake-correction parameter calculating part 106 and the associated video frames to perform geometric transformation processing of the video frames. The shake correction parameters may be subjected to filtering processing or the like before the processing in this part 107, so that they may be delayed relative to the associated video frames. In this case, the video frames are once passed through the work memory 109 to match the video frames with the shake correction parameters. The work memory is a FIFO memory similar to the work memory 105.
0062The encoding/decoding part 108 encodes the video frame signal successively output from the geometric transformation processing part 107 in a video format such as NTSC and MPEG4. To reproduce a recorded and encoded video signal, the encoding/decoding part 108 decodes the video signal read out from the recording part 113 and displays it on the displaying part 114.
0063The system controlling part 110 transmits the video signal encoded in the abovementioned format and output from the encoding/decoding part 108 to the recording part 113 for recording. The system controlling part 110 also controls parameters of the processing blocks such as the motion vector detecting part 104, the shake-correction parameter calculating part 106, the geometric transformation processing part 107, and the encoding/decoding part 108. Initial values of the parameters are read out from the non-volatile memory part 112. The various parameters are displayed on the displaying part 114 and the values of the parameters can be changed with the operation signal inputting part 115 or a GUI.
0064The system controlling part 110 holds control parameters such as the number of the motion vectors, a search range of the motion vectors, and a template size for the motion vector detecting part 104. The system controlling part 110 provides the geometric transformation processing part 107 with control parameters such as the shake correction parameters calculated by the shake-correction parameter calculating part 106, and the in-camera parameters and distortion coefficient used in the calculation. The system controlling part 110 provides the encoding/decoding part 108 with control parameters such as an encoding format and a compression rate.
0065The system controlling part 110 performs control of the work memories 105 and 109 to control the delay amount of output. The system controlling part 110 also controls the zoom adjusting part 111 which performs zoom operation of the lens optical system 101. Specifically, the system controlling part 110 reads a zoom value representing a zoom state with an encoder in the zoom adjusting part 111. The system controlling part 110 uses a lookup table or a transforming expression showing the relationship between the zoom value and the focal length stored in the non-volatile memory part 112 to calculate and hold the focal length of the lens optical system 101 in an arbitrary zoom state.
0066The distortion coefficient varies depending on the focal length. Thus, the system controlling part 110 calculates the distortion coefficient corresponding to a focal length. Specifically, it uses a lookup table or a transforming expression showing the relationship between the focal length and the distortion coefficient stored in the non-volatile memory part 112 to calculate and hold the distortion coefficient in an arbitrary focal length. In addition, it reads and holds the in-camera parameters other than the focal length from the non-volatile memory part 112.
0067The in-camera parameters other than the focal length f include pixel sizes k<sub>u</sub>, k<sub>v</sub> in horizontal and vertical directions, a shear amount φ, and offset amounts u<sub>0</sub>, v<sub>0</sub> in horizontal and vertical directions. The in-camera parameters are provided from camera design specifications or camera calibration. The system controlling part 110 transmits the in-camera parameters and the distortion coefficient to the shake-correction parameter calculating part 106.
0068The non-volatile memory part 112 stores the initial values of the control parameters necessary to system control for the motion vector detecting part 103, the shake-correction parameter calculating part 106, the encoding/decoding part 108, the preprocessing part 103, the image-pickup element 102 and the like. The control parameters include the in-camera parameters, the lookup table or the transforming expression showing the relationship between the zoom position (zoom value) and the focal length, and the lookup table or the transforming expression showing the relationship between the focal length and the distortion coefficient. The control parameters are read out by the system controlling part 110.
0069The recording part 113 performs writing (recording) and reading (reproduction) of the video signal encoded by the encoding/decoding part 108 to and from a recording medium on which the video signal can be recorded such as a semiconductor memory, a magnetic tape, and an optical disk.
0070The displaying part 114 is formed of a display element such as an LCD, an LED, and an EL. The displaying part 114 performs, for example, parameter setting display, alarm display, display of picked-up video data, and display of recorded video data read by the recording part 113. In reproducing the recorded video data, the encoded video signal is read from the recording part 113 and the read signal is transmitted the read signal to the encoding/decoding part 108 via the system controlling part 110. The recorded video data after it is decoded is displayed on the displaying part 114.
0071The operation signal inputting part 115 includes setting buttons for performing selection of functions and various settings in the camera from the outside and a button for instructing start and end of an image pick-up operation. The operation signal inputting part 115 may be integrated with the displaying part 114 by using a touch panel display.
0072The external I/F 116 receives an input signal from the outside instead of an operation signal input from the operation signal inputting part 115 or outputs the encoded video signal to an external device. The external I/F 116 is realized with an I/F protocol such as USB, IEEE1394, and wireless LAN. It can receive from the outside a video signal including information necessary for image stabilization such as the focal length or the zoom state in image-pickup operation, the in-camera parameters, and the distortion coefficient to allow image stabilization processing for recorded video.
0073<figref idref="f0003">Fig. 3</figref> shows an image stabilization processing procedure in Embodiment 1. The image-stabilization processing includes a motion vector calculating step, a shake-correction parameter calculating step, and a geometric transformation processing step, and is repeated for the input video frames. The motion vector calculating step, the shake-correction parameter calculating step, and the geometric transformation processing step are performed by the motion vector detecting part 104, the shake-correction parameter calculating part 106, and the geometric transformation processing part 107, respectively. The processing in each part is controlled with a computer program (image stabilizing program) stored in the system controlling part 110.
0074S (step) 301 is the motion vector calculating step. At this step, the motion vectors are calculated between the current frame directly input from the preprocessing part 103 and the past frame input from the work memory 105. In the calculation of the motion vectors, template matching or matching by a gradient method is performed.
0075<figref idref="f0004">Fig. 4</figref> shows an example of block matching which is a type of the template matching. A video frame (frame image) 401 on the left is used as a reference image, while a video frame (frame image) 402 on the right is used as a search image. For example, the previously input past frame is used as the reference image and the current frame input after that is used as the search image to detect the motion vectors. A template 403 is defined in the left image 401 as a partial area of a predetermined size including points arranged in a grid pattern in which an attention point 404 is located at the center. An arbitrary search area 407 is set in the right image 402. While the search area 407 is gradually shifted, the position best matching the template 403 is searched for.
0076Specifically, similarity is calculated between an area 406 including the attention pixel 405 as a reference in the right image 402 and the template 403 in the left image 401. SSD (Sum of Square Difference), SAD (Sum of Absolute Difference), the result of correlation calculation such as normalized cross-correlation can be used as the index of the similarity. When the luminance significantly varies between frames as in video taken from a real scene, the normalized cross-correlation is mainly used. The following is the expression for calculating the similarity score in the normalized cross-correlation: <maths id="math0002"><math display="block"><mi>R</mi><mfenced><mi>x</mi><mo></mo><mi>y</mi><mo></mo><mi mathvariant="italic">xʹ</mi><mo></mo><mi mathvariant="italic">yʹ</mi></mfenced><mo>=</mo><mfrac><mrow><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mo>-</mo><msub><mi>M</mi><mi>T</mi></msub></mrow><msub><mi>M</mi><mi>T</mi></msub></munderover></mstyle><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mo>-</mo><msub><mi>N</mi><mi>T</mi></msub></mrow><msub><mi>N</mi><mi>T</mi></msub></munderover></mstyle><mfenced open="{" close="}"><msub><mi>I</mi><mfenced><mi>x</mi><mo></mo><mi>y</mi></mfenced></msub><mfenced><mi>i</mi><mo></mo><mi>j</mi></mfenced><mo>-</mo><mover><mi>I</mi><mo>‾</mo></mover></mfenced><mfenced open="{" close="}"><msub><mi mathvariant="italic">Iʹ</mi><mfenced><mi mathvariant="italic">xʹ</mi><mo></mo><mi mathvariant="italic">yʹ</mi></mfenced></msub><mfenced><mi>i</mi><mo></mo><mi>j</mi></mfenced><mo>-</mo><mover><mi>I</mi><mo>‾</mo></mover><mo></mo><mi mathvariant="italic">ʹ</mi></mfenced></mrow><mrow><msqrt><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mo>-</mo><msub><mi>M</mi><mi>T</mi></msub></mrow><msub><mi>M</mi><mi>T</mi></msub></munderover></mstyle><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mo>-</mo><msub><mi>N</mi><mi>T</mi></msub></mrow><msub><mi>N</mi><mi>T</mi></msub></munderover></mstyle><msup><mfenced open="{" close="}"><msub><mi>I</mi><mfenced><mi>x</mi><mo></mo><mi>y</mi></mfenced></msub><mfenced><mi>i</mi><mo></mo><mi>j</mi></mfenced><mo>-</mo><mover><mi>I</mi><mo>‾</mo></mover></mfenced><mn>2</mn></msup></msqrt><mo></mo><msqrt><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mo>-</mo><msub><mi>M</mi><mi>T</mi></msub></mrow><msub><mi>M</mi><mi>T</mi></msub></munderover></mstyle><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mo>-</mo><msub><mi>N</mi><mi>T</mi></msub></mrow><msub><mi>N</mi><mi>T</mi></msub></munderover></mstyle><msup><mfenced open="{" close="}"><msub><mi mathvariant="italic">Iʹ</mi><mfenced><mi mathvariant="italic">xʹ</mi><mo></mo><mi mathvariant="italic">yʹ</mi></mfenced></msub><mfenced><mi>i</mi><mo></mo><mi>j</mi></mfenced><mo>-</mo><mover><mi>I</mi><mo>‾</mo></mover><mo></mo><mi mathvariant="italic">ʹ</mi></mfenced><mn>2</mn></msup></msqrt></mrow></mfrac></math><img file="EP1978731A2_D0002.tif" /></maths>
0077where <maths id="math0003"><math display="block"><mover><mi>I</mi><mo>‾</mo></mover><mo>=</mo><mfrac><mn>1</mn><mrow><msub><mi>M</mi><mi>T</mi></msub><mo></mo><msub><mi>N</mi><mi>T</mi></msub></mrow></mfrac><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mo>-</mo><msub><mi>M</mi><mi>T</mi></msub></mrow><msub><mi>M</mi><mi>T</mi></msub></munderover></mstyle><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mo>-</mo><msub><mi>N</mi><mi>T</mi></msub></mrow><msub><mi>N</mi><mi>T</mi></msub></munderover></mstyle><msub><mi>I</mi><mfenced><mi>x</mi><mo></mo><mi>y</mi></mfenced></msub><mfenced><mi>i</mi><mo></mo><mi>j</mi></mfenced><mo>,</mo><mspace width="1em" /><mover><mi>I</mi><mo>‾</mo></mover><mo></mo><mi mathvariant="italic">ʹ</mi><mo>=</mo><mfrac><mn>1</mn><mrow><msub><mi>M</mi><mi>T</mi></msub><mo></mo><msub><mi>N</mi><mi>T</mi></msub></mrow></mfrac><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mo>-</mo><msub><mi>M</mi><mi>T</mi></msub></mrow><msub><mi>M</mi><mi>T</mi></msub></munderover></mstyle><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mo>-</mo><msub><mi>N</mi><mi>T</mi></msub></mrow><msub><mi>N</mi><mi>T</mi></msub></munderover></mstyle><msub><mi mathvariant="italic">Iʹ</mi><mfenced><mi mathvariant="italic">xʹ</mi><mo></mo><mi mathvariant="italic">yʹ</mi></mfenced></msub><mfenced><mi>i</mi><mo></mo><mi>j</mi></mfenced><mo>,</mo></math><img file="EP1978731A2_D0003.tif" /></maths> and (x, y) and (x', y') represent the positions of the templates in the reference image I and the search image <i>I</i>', respectively. <i>I</i><sub>(x,y)</sub> (i, j) and <i>I</i>'<sub>(x',y')</sub>(i, j) represent partial images.
0078After the calculation of the similarity in all of the search areas, the position with the highest similarity is regarded as the corresponding position to calculate the motion vectors. If no occlusion is present, as many motion vectors are calculated as the number of the attention points 404 set on the reference image (left image 401). Each of the motion vectors is represented as follows by a vector starting from the position of the attention point 404 in the reference image and ending at the position of the corresponding point in the search image (right image 402): <maths id="math0004"><math display="block"><msup><mfenced><mi>x</mi><mo></mo><mi>y</mi><mo></mo><mi mathvariant="italic">xʹ</mi><mo></mo><mi mathvariant="italic">yʹ</mi></mfenced><mi>i</mi></msup><mo>,</mo><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>…</mo><mo>,</mo><mi>m</mi></math><img file="EP1978731A2_D0004.tif" /></maths> (m represents the number of the motion vectors)
0079The example of the block matching in which the attention points are fixedly arranged in the grid pattern has been shown. Alternatively, it is possible to extract a feature point with which a motion vector is readily calculated on the reference image and to define the position of the feature point as the attention point.
0080The extraction of the attention point is typically performed by using an image processing filter such as Harris operator (<nplcit id="ncit0004" npl-type="s"><text>C. Harris and M. Stephens, "A combined corner and edge detector". Fourth Alvey Vision Conference, pp. 147-151, 1988</text></nplcit>).
0081The Harris operator first determines the size of a window W and calculates differential images (I<sub>dx</sub>, I<sub>dy</sub>) in horizontal and vertical directions. The calculation may be performed with a Sobl filter for calculation of the differential images. For example, by using a filter defined as h = 1,√2, 1]/[2+√2], (I<sub>dx</sub>, I<sub>dy</sub>) is provided by applying 3 x 3 filter h<sub>x</sub> arranged in a vertical direction and 3 x 3 filter h<sub>y</sub> arranged in a horizontal direction to the image.
0082For all of coordinates (x, y) in the images, the following matrix G is calculated using the window <maths id="math0005"><math display="block"><mtable><mtr><mtd><mi mathvariant="normal">W</mi><mo>:</mo></mtd><mtd><mi>G</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mstyle displaystyle="false"><mstyle displaystyle="true"><munder><mo>∑</mo><mi>W</mi></munder></mstyle><msubsup><mi>I</mi><mi>x</mi><mn>2</mn></msubsup></mstyle></mtd><mtd><mstyle displaystyle="false"><mstyle displaystyle="true"><munder><mo>∑</mo><mi>W</mi></munder></mstyle><msub><mi>I</mi><mi>x</mi></msub><mo></mo><msub><mi>I</mi><mi>x</mi></msub></mstyle></mtd></mtr><mtr><mtd><mstyle displaystyle="false"><mstyle displaystyle="true"><munder><mo>∑</mo><mi>W</mi></munder></mstyle><msub><mi>I</mi><mi>x</mi></msub><mo></mo><msub><mi>I</mi><mi>x</mi></msub></mstyle></mtd><mtd><mstyle displaystyle="false"><mstyle displaystyle="true"><munder><mo>∑</mo><mi>W</mi></munder></mstyle><msubsup><mi>I</mi><mi>y</mi><mn>2</mn></msubsup></mstyle></mtd></mtr></mtable></mfenced></mtd></mtr></mtable><mn>.</mn></math><img file="EP1978731A2_D0005.tif" /></maths>
0083In addition, feature points are extracted in order from coordinates (x, y) having a larger minimum singular value of the matrix G. In this case, it is preferable to prevent dense distribution of the singular points. Thus, it is possible to make a rule not to extract a feature point in an area close to the window W including the coordinates (x, y) at which a feature point is already extracted.
0084S302 is the shake-correction parameter calculating step. At this step, the motion vectors between the current frame and the past frame are received as input, and the shake correction parameter for the target frame is output. When no delay is produced in the calculation, the shake correction parameter is for the current frame as the target frame. When delay is produced, a frame of calculation target of the shake correction parameter is a frame which is older corresponding to the delay amount.
0085At the shake-correction parameter calculating step S302, the shake correction parameter is calculated at subdivided steps as shown in <figref idref="f0005">Fig. 5</figref>. The shake correction parameter is formed of a geometric transforming matrix which represents the in-camera parameters, the distortion coefficient, and the shake correction amount.
0086First, at a normalization step of S501, the values of the motion vectors in a pixel coordinate system of the input frame are transformed into the values of motion vectors in a normalized image coordinate system. Coordinates (x, y) represent the pixel coordinates on the input frame, coordinates (u<sub>o</sub>, v<sub>o</sub>) represent the normalized image coordinates including distortion, and coordinates (u, v) represent the normalized image coordinates where the distortion was excluded. In this case, the motion vectors are first transformed into the normalized image coordinates with the in-camera parameters. In the following expression, inv() represents an inverse matrix of the matrix in the parentheses: <maths id="math0006"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>u</mi><mi>d</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>d</mi></msub></mtd></mtr><mtr><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>=</mo><mi mathvariant="italic">inv</mi><mfenced><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>f</mi><mi mathvariant="italic">c_new</mi></msub><mo></mo><msub><mi>k</mi><mi>u</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><msub><mi>u</mi><mn>0</mn></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><msub><mi>f</mi><mi mathvariant="italic">c_new</mi></msub><mo></mo><msub><mi>k</mi><mi>v</mi></msub></mtd><mtd><msub><mi>v</mi><mn>0</mn></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>x</mi></mtd></mtr><mtr><mtd><mi>y</mi></mtd></mtr><mtr><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0006.tif" /></maths> Then, the distortion is removed with the distortion coefficient as follows: <maths id="math0007"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>u</mi><mi>d</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>d</mi></msub></mtd></mtr></mtable></mfenced><mo>→</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>u</mi></mtd></mtr><mtr><mtd><mi>v</mi></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0007.tif" /></maths>
0087The calculation represented by "→" is performed by the following processing:
0088The distortion removal is performed by using the following expressions which represent the relationship of radial distortion: <maths id="math0008"><math display="block"><mi>K</mi><mo>=</mo><mn>1</mn><mo>+</mo><msub><mi>k</mi><mn>1</mn></msub><mo></mo><mi>r</mi><mo>+</mo><msub><mi>k</mi><mn>2</mn></msub><mo></mo><msup><mi>r</mi><mn>2</mn></msup><mo>+</mo><mo>⋯</mo><mo>,</mo><msup><mi>r</mi><mn>2</mn></msup><mo>=</mo><msubsup><mi>u</mi><mi>d</mi><mn>2</mn></msubsup><mo>+</mo><msubsup><mi>v</mi><mi>d</mi><mn>2</mn></msubsup></math><img file="EP1978731A2_D0008.tif" /></maths><maths id="math0009"><math display="block"><mi>u</mi><mo>=</mo><msub><mi>u</mi><mi>d</mi></msub><mo>/</mo><mi>K</mi><mo>,</mo><mi>v</mi><mo>=</mo><msub><mi>v</mi><mi>d</mi></msub><mo>/</mo><mi>K</mi></math><img file="EP1978731A2_D0009.tif" /></maths> where k<sub>1</sub>, k<sub>2</sub>, and k<sub>3</sub> represent distortion coefficients in first, second, and third order radial directions, respectively. These distortions are caused by the aberration of the lens optical system 101.
0089The distortion varies with the focal length of the lens optical system 101. Thus, the relationship between the distortion and the focal length is previously provided through calculation from designed values or measurement with variation of the focal length. The relationship is stored on the non-volatile memory part 112 as the lookup table associated with the focal length or the transformation expression relating to the focal length.
0090The system controlling part 110 calculates the focal length based on the zoom state of the lens optical system 110 sent from the zoom adjusting part 111, obtains the corresponding distortion coefficient from the calculation with the calculation expression or with reference to the lookup table, and provides the obtained distortion coefficient to each processing part.
0091Only the distortion in the radial directions is removed in Embodiment 1. If another distortion is serious such as distortion in a moving radius direction, additional distortion removal processing may be performed.
0092At an image variation amount calculating step of S502, the motion vectors between the frames transformed into the normalized image coordinate system are used as input to calculate the image variation amount between the frames. The projective homography is used as the index of the image variation amount. The following linear expression for the projective homography can be provided by setting the normalized image coordinates in the past frame to (u<sub>1</sub>, v<sub>1</sub>), the normalized image coordinates in the current frame to (u<sub>1</sub>',v<sub>1</sub>'), and i = 1, A, m (m represents the number of the motion vectors): <maths id="math0010"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>i</mi></msub></mtd><mtd><mo>-</mo><msub><mi>v</mi><mi>i</mi></msub></mtd><mtd><mo>-</mo><mn>1</mn></mtd><mtd><msub><mi>v</mi><mi>i</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>u</mi><mi>i</mi></msub></mtd><mtd><msub><mi>v</mi><mi>i</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>v</mi><mi>i</mi></msub></mtd><mtd><msub><mi>v</mi><mi>i</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><msub><mi>u</mi><mi>i</mi></msub></mtd><mtd><msub><mi>v</mi><mi>i</mi></msub></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>i</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>u</mi><mi>i</mi></msub></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>i</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>v</mi><mi>i</mi></msub></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>i</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>m</mi></msub></mtd><mtd><mo>-</mo><msub><mi>v</mi><mi>m</mi></msub></mtd><mtd><mo>-</mo><mn>1</mn></mtd><mtd><msub><mi>v</mi><mi>m</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>u</mi><mi>m</mi></msub></mtd><mtd><msub><mi>v</mi><mi>m</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>v</mi><mi>m</mi></msub></mtd><mtd><msub><mi>v</mi><mi>m</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><msub><mi>u</mi><mi>m</mi></msub></mtd><mtd><msub><mi>v</mi><mi>m</mi></msub></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>m</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>u</mi><mi>m</mi></msub></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>m</mi></msub><mo></mo><mi>ʹ</mi><mo></mo><msub><mi>v</mi><mi>m</mi></msub></mtd><mtd><mo>-</mo><msub><mi>u</mi><mi>m</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>h</mi><mn>11</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>12</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>13</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>21</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>22</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>23</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>31</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>32</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>33</mn></msub></mtd></mtr></mtable></mfenced><mo>=</mo><mn>0.</mn></math><img file="EP1978731A2_D0010.tif" /></maths>
0093The linear expression is overdetermined if the number m of the corresponding points is equal to or larger than eight. The expression can be solved as a linear least square expression to provide the following: <maths id="math0011"><math display="block"><mi>h</mi><mo>=</mo><mfenced open="{" close="}"><msub><mi>h</mi><mn>11</mn></msub><mo>⋯</mo><msub><mi>h</mi><mn>33</mn></msub></mfenced><mn>.</mn></math><img file="EP1978731A2_D0011.tif" /></maths>
0094This is shaped into a matrix of 3 x 3 to provide the projective homography represented as follows, that is, the image variation amount: <maths id="math0012"><math display="block"><mi>H</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>h</mi><mn>11</mn></msub></mtd><mtd><msub><mi>h</mi><mn>12</mn></msub></mtd><mtd><msub><mi>h</mi><mn>13</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>21</mn></msub></mtd><mtd><msub><mi>h</mi><mn>22</mn></msub></mtd><mtd><msub><mi>h</mi><mn>23</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>31</mn></msub></mtd><mtd><msub><mi>h</mi><mn>32</mn></msub></mtd><mtd><msub><mi>h</mi><mn>33</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>41</mn></msub></mtd><mtd><msub><mi>h</mi><mn>42</mn></msub></mtd><mtd><msub><mi>h</mi><mn>43</mn></msub></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0012.tif" /></maths>
0095At an appearance variation component removal step of S503, an appearance variation component is removed from the image variation amount determined between the frames. The image variation amount, that is, the projective homography is the index representing the motion between the frames and is formed of information on camera works including the rotation and translation (information on the rotation and translation) and information on scene arrangement including the depth position and direction of the reference plane in space.
0096The relationship between the projective homography and the camera works and scene arrangement is represented by the following expression: <maths id="math0013"><math display="block"><mi>H</mi><mo>=</mo><mi mathvariant="italic">λ</mi><mo></mo><mfenced><mi>R</mi><mo>+</mo><mfrac><mn>1</mn><mi>d</mi></mfrac><mo></mo><mover><mi>t</mi><mo>→</mo></mover><mo></mo><msup><mover><mi>n</mi><mo>→</mo></mover><mi>T</mi></msup></mfenced></math><img file="EP1978731A2_D0013.tif" /></maths> where <i>R</i> represents the rotation of the camera, <o ostyle="rightarrow"><i>t</i></o> the translation of the camera, d the distance to the reference plane, <o ostyle="rightarrow"><i>n</i></o> the normal to the reference plane in the direction away from the camera, and λ an arbitrary constant.
0097In calculation from two images, the product of the distance d to a spatial plane and a norm of a translation camera work expressed as below cannot be resoluved: <maths id="math0014"><math display="block"><mi mathvariant="italic">norm</mi><mfenced><mover><mi>t</mi><mo>→</mo></mover></mfenced><mn>.</mn></math><img file="EP1978731A2_D0014.tif" /></maths>
0098The norm refers to the amount representing the size of the vector. In this case, <o ostyle="rightarrow"><i>t</i></o> is handled as <i>norm</i> (<o ostyle="rightarrow"><i>t</i></o>) = 1 that is a unit direction vector representing the translation direction. d is handled as the product of the distance to the spatial plane and the translation amount.
0099The appearance variation component is defined as a difference between a projective homography for a reference plane oriented in an arbitrary direction and a projective homography for a reference plane oriented perpendicularly to the optical axis.
0100Specifically, an image variation amount is produced by translation with respect to the reference plane oriented in the arbitrary direction, and an image variation amount is calculated which is produced by translation of the same amount with respect to the reference plane present at the same depth position and perpendicular to the optical axis. The difference between them is defined as the appearance change amount.
0101In other words, the appearance variation component shows shear in geometric transformation between images. The component corresponds to a component which distorts an image when image stabilization is performed by using the inversely transformed image variation amount as the correction amount. The appearance variation component should be removed in determining motion information appropriate for image stabilization.
0102From a different viewpoint, the removal of the appearance variation component corresponds to the turning of the direction of the reference plane of the homography arbitrarily determined on the basis of the distribution of the corresponding points in space to a direction in parallel with the image plane before the displacement due to the motion. It corresponds to the turning of the direction of the normal to the reference plane to a direction in parallel with the optical axis of the image plane before the displacement.
0103To remove the appearance variation component, the projective homography is decomposed into camera work rotation <i>R</i> , a direction <o ostyle="rightarrow"><i>n</i></o> of a plane approximate to an object in a scene, and a product of translation <o ostyle="rightarrow"><i>t</i></o> and the reciprocal of depth d as follows: <maths id="math0015"><math display="block"><mover><mi>t</mi><mo>→</mo></mover><mo>/</mo><mi>d</mi><mn>.</mn></math><img file="EP1978731A2_D0015.tif" /></maths>
0104Then, the plane direction <o ostyle="rightarrow"><i>n</i></o> is replaced with the direction perpendicular to the camera optical axis, thereby calculating the homography in which the appearance variation component is excluded.
0105First, in resolution of the projective homography, two possible solutions are calculated with the following procedure. The decomposition of the projective homography into the two solutions is performed by using eigenvalue resolution or singular value resolution to find an invariant. Although various manners of solution may be used, the following description will be made with reference to the approach used in B, Triggs, "Autocalibration from Planar Scene" ECCV98.
0106First, assuming that the sign of <i>H</i> is selected to satisfy <o ostyle="rightarrow"><i>x</i></o><sub>2</sub><i><sup>T</sup>H<o ostyle="rightarrow">x</o></i><sub>1</sub> > 0 at all of the corresponding points <o ostyle="rightarrow">x</o><sub>1</sub>, <o ostyle="rightarrow">x</o><sub>2</sub> on a plane.
0107The singular value resolution of <i>H</i> is given as <i>H = USV<sup>T</sup></i> , where <i>U</i> and <i>V</i> represent 3 x 3 rotation matrixes.
0108Further, <i>S = diag</i>(<i>σ</i><sub>1</sub>,<i>σ</i><sub>2</sub>,<i>σ</i><sub>3</sub>) represents a positive descending diagonal element (σ1 ≥ σ2 ≥ σ3 ≥ 0) and is set to the singular value of <i>H.</i> Column elements of <i>U</i> and <i>V</i> that are associated orthogonal matrixes are represented as u1, u2, u3 and ν1, ν2, ν3.
0109The reference system of a first camera is employed and a three-dimensional plane is represented by: <maths id="math0016"><math display="block"><msup><mover><mi>n</mi><mo>→</mo></mover><mi>T</mi></msup><mo></mo><mover><mi>x</mi><mo>→</mo></mover><mo>=</mo><mi>d</mi><mo>=</mo><mn>1</mn><mo>/</mo><mi mathvariant="italic">ζ</mi></math><img file="EP1978731A2_D0016.tif" /></maths> where <img file="EP1978731A2_D0017.tif" /> represents the outward normal (direction away from the camera), and ζ (= 1/d≥0) represents the reciprocal of the distance to the plane. In the reference system, the first camera has a 3 x 4 projection matrix P1 = [I3x3|<o ostyle="rightarrow">0</o>].
0110In a second camera, P2 = <i>R</i> [I3x3|t] = [<i>R</i> |t']. where t'= - <i>R</i> t, and t and t' represent translation between the cameras (translation from the optical axis center of the first camera to the optical axis center of the second camera), and <i>R</i> represents rotation between the cameras.
0111The homography representing the image variation from an image 1 of the first camera to an image 2 of the second camera is <i>H</i>=<i>RH</i><sub>1</sub>, and <i>H</i><sub>1</sub> = <i>I</i><sub>3×3</sub> - ζ<i><o ostyle="rightarrow">tn</o><sup>T</sup></i> holds. For a three-dimensional point <o ostyle="rightarrow"><i>x</i></o> on the plane, <i>H<o ostyle="rightarrow">x</o></i> = <i>R</i>(<o ostyle="rightarrow"><i>x</i></o> - ζ<i><o ostyle="rightarrow">tn</o><sup>T</sup><o ostyle="rightarrow">x</o></i>) = <i>R</i>(<o ostyle="rightarrow"><i>x</i></o> - <o ostyle="rightarrow"><i>t</i></o>) ≈ <i>P</i><sub>2</sub><o ostyle="rightarrow"><i>x</i></o> holds because ζ<i><o ostyle="rightarrow">n</o><sup>T</sup><o ostyle="rightarrow">x</o></i> = 1 is given. When <o ostyle="rightarrow"><i>x</i></o> is handled as an arbitrary point in the image 1, the difference is only the whole scale factor.
0112Only the product ζ<i><o ostyle="rightarrow">tn</o><sup>T</sup></i> is restorable, so that normalization is performed with ||t||=||n||=1, that is, the plane distance 1/ξ is measured in a unit base length ||t||. A depth positive constraint test, later described, is performed to determine the possible sign.
0113<i>H</i>=<i>USV<sup>T</sup></i> and <i>H</i><sub>1</sub>=<i>U</i><sub>1</sub><i>SV<sup>T</sup></i> in the singular value resolution are identical for the element of <i>R</i>, that is, <i>U</i>=<i>RU</i><sub>1</sub>. In <i>H</i><sub>1</sub>, the outer vector<o ostyle="rightarrow"><i>t</i></o> × <o ostyle="rightarrow"><i>n</i></o> is invariant. If the singular value is obvious,<o ostyle="rightarrow"><i>t</i></o> × <o ostyle="rightarrow"><i>n</i></o> should correspond to a singular vector. It is apparent that this is always the second singular vector v2. Thus, correction normalization of <i>H</i> is performed as <i>H</i>→<i>H</i> /σ2, that is, (σ1, σ2, σ3) →(σ1/σ2, 1, σ3/σ2,). In the following, it is assumed that normalization with σ2 is already performed.
0114In the image frame 1, it is given that <o ostyle="rightarrow"><i>t</i></o> × <o ostyle="rightarrow"><i>n</i></o> corresponds to v2, a partial space {<o ostyle="rightarrow"><i>t</i></o> × <o ostyle="rightarrow"><i>n</i></o>} should be occupied by {v1 , v3} . That is, <o ostyle="rightarrow"><i>n</i></o> = β<o ostyle="rightarrow">ν</o><sub>1</sub> - α<o ostyle="rightarrow">ν</o><sub>3</sub> and <o ostyle="rightarrow"><i>n</i></o> × (<o ostyle="rightarrow"><i>t</i></o> × <o ostyle="rightarrow"><i>n</i></o>) ≈ α<o ostyle="rightarrow">ν</o><sub>1</sub> + β<o ostyle="rightarrow">ν</o><sub>3</sub> hold for arbitrary parameters α,β (where α2 + β2 = 1). An arbitrary direction (especially <o ostyle="rightarrow"><i>n</i></o> × (<o ostyle="rightarrow"><i>t</i></o> × <o ostyle="rightarrow"><i>n</i></o>) orthogonal to <o ostyle="rightarrow"><i>n</i></o> has a norm which is invariant with H<sub>1</sub>.
0115In this case, <i>(α</i>σ<sub>1</sub>)<sup>2</sup>+(βσ<sub>3</sub>)<sup>2</sup>=α<sup>2</sup>+β<sup>2</sup> or <maths id="math0017"><math display="inline"><mfenced><mi>α</mi><mo></mo><mi>β</mi></mfenced><mo>=</mo><mfenced><mo>±</mo><msqrt><mn>1</mn><mo>-</mo><msubsup><mi>σ</mi><mn>3</mn><mn>2</mn></msubsup></msqrt><mo>±</mo><msqrt><msubsup><mi>σ</mi><mn>1</mn><mn>2</mn></msubsup><mo>-</mo><mn>1</mn></msqrt></mfenced></math><img file="EP1978731A2_D0018.tif" /></maths> holds.
0116If <o ostyle="rightarrow"><i>t</i></o> × <o ostyle="rightarrow"><i>n</i></o> corresponds to the abovementioned ν1 or v3, no solution is found. Thus, it can correspond to only v2.
0117Strictly, the same argument on the left-hand side shows <i>R<o ostyle="rightarrow">t</o></i> ≈ -(β<i>u</i><sub>1</sub> + α<i>u</i><sub>3</sub>). If <o ostyle="rightarrow"><i>t</i></o> satisfies an eigenvector 1 - ζ<i><o ostyle="rightarrow">nt</o><sup>T</sup></i> which is an eigenvalue of <i>H</i><sub>1</sub>, <i>H<o ostyle="rightarrow">t</o></i>= (1 - ζ<i><o ostyle="rightarrow">n</o><sup>T</sup><o ostyle="rightarrow">t</o></i>)<i>R<o ostyle="rightarrow">t</o></i> is given. Thus, <i>t</i> ≈ <i>H</i><sup>-1</sup>(<i>R<o ostyle="rightarrow">t</o></i>) ≈ β/σ<sub>1</sub><o ostyle="rightarrow">ν</o><sub>1</sub> + α/σ<sub>3</sub><o ostyle="rightarrow">ν</o><sub>3</sub> holds. After simplification, ξ = σ1 - σ3 holds.
0118The columns (<o ostyle="rightarrow"><i>u</i></o><sub>1</sub>, <o ostyle="rightarrow"><i>u</i></o><sub>2</sub>, <o ostyle="rightarrow"><i>u</i></o><sub>3</sub>) of <i>U</i><sub>1</sub> that is the left-hand side of the singular value resolution of <i>H</i><sub>1</sub> is restorable with the notation of <o ostyle="rightarrow"><i>u</i></o><sub>2</sub> = <o ostyle="rightarrow">ν</o><sub>2</sub>, and <o ostyle="rightarrow"><i>t</i></o> needs to be an eigenvector of <i>H</i><sub>1</sub>.
0119In this case, <o ostyle="rightarrow"><i>u</i></o><sub>1</sub> = γ<o ostyle="rightarrow">ν</o><sub>1</sub> + δ<o ostyle="rightarrow">ν</o><sub>3</sub> and <o ostyle="rightarrow"><i>u</i></o><sub>3</sub> = δ<o ostyle="rightarrow">ν</o><sub>1</sub> - γ<o ostyle="rightarrow">ν</o><sub>3</sub> hold. After simplification, (γ,δ)≈(1+σ<sub>1</sub>σ<sub>3</sub>,±αβ) holds. Thus, <maths id="math0018"><math display="block"><mi>R</mi><mo>=</mo><mi>U</mi><mo></mo><msubsup><mi>U</mi><mn>1</mn><mi>T</mi></msubsup><mo>=</mo><mi>U</mi><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi mathvariant="italic">γ</mi></mtd><mtd><mn>0</mn></mtd><mtd><mi mathvariant="italic">δ</mi></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mo>-</mo><mi mathvariant="italic">δ</mi></mtd><mtd><mn>0</mn></mtd><mtd><mi mathvariant="italic">γ</mi></mtd></mtr></mtable></mfenced><mo></mo><msup><mi>V</mi><mi>T</mi></msup></math><img file="EP1978731A2_D0019.tif" /></maths> is assumed and finally the rotation <i>R</i> is obtained.
0120Next, a series of specific processing is shown for calculating the two possible solutions for resolving the image variation amount into the camera work <i>R</i> including the rotation and translation, and the scene arrangement including <o ostyle="rightarrow"><i>t</i></o> (direction vector), the depth position d and direction <o ostyle="rightarrow"><i>n</i></o> of the reference plane in space. <maths id="math0019"><math display="block"><mfenced open="[" close="]"><mi>U</mi><mo></mo><mi>S</mi><mo></mo><mi>V</mi></mfenced><mo>=</mo><mi mathvariant="italic">svd</mi><mfenced><mi>H</mi></mfenced></math><img file="EP1978731A2_D0020.tif" /></maths><maths id="math0020"><math display="block"><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub><mo>=</mo><msub><mi>σ</mi><mn>1</mn></msub><mo>/</mo><msub><mi>σ</mi><mn>2</mn></msub><mo>,</mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub><mo>=</mo><msub><mi>σ</mi><mn>3</mn></msub><mo>/</mo><msub><mi>σ</mi><mn>2</mn></msub></math><img file="EP1978731A2_D0021.tif" /></maths> where <maths id="math0021"><math display="block"><mi>S</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>σ</mi><mn>1</mn></msub></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><msub><mi>σ</mi><mn>2</mn></msub></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><msub><mi>σ</mi><mn>3</mn></msub></mtd></mtr></mtable></mfenced><mo>,</mo><msub><mi>σ</mi><mn>1</mn></msub><mo>≥</mo><msub><mi>σ</mi><mn>2</mn></msub><mo>≥</mo><msub><mi>σ</mi><mn>3</mn></msub><mo>≥</mo><mn>0</mn></math><img file="EP1978731A2_D0022.tif" /></maths><maths id="math0022"><math display="block"><mi mathvariant="italic">ζ</mi><mo>=</mo><mfenced><mn>1</mn><mo>/</mo><mi>d</mi></mfenced><mo>=</mo><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub><mo>-</mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub></math><img file="EP1978731A2_D0023.tif" /></maths><maths id="math0023"><math display="block"><msub><mi>a</mi><mn>1</mn></msub><mo>=</mo><msqrt><mn>1</mn><mo>-</mo><msubsup><mi mathvariant="italic">σʹ</mi><mn>3</mn><mn>2</mn></msubsup></msqrt><mo>,</mo><msub><mi>b</mi><mn>1</mn></msub><mo>=</mo><msqrt><msubsup><mi mathvariant="italic">σʹ</mi><mn>1</mn><mn>2</mn></msubsup><mo>-</mo><mn>1</mn></msqrt></math><img file="EP1978731A2_D0024.tif" /></maths><maths id="math0024"><math display="block"><mi>a</mi><mo>=</mo><msub><mi>a</mi><mn>1</mn></msub><mo>/</mo><msqrt><msubsup><mi>a</mi><mn>1</mn><mn>2</mn></msubsup><mo>+</mo><msubsup><mi>b</mi><mn>1</mn><mn>2</mn></msubsup></msqrt><mo>,</mo><mi>b</mi><mo>=</mo><msub><mi>b</mi><mn>1</mn></msub><mo>/</mo><msqrt><msubsup><mi>a</mi><mn>1</mn><mn>2</mn></msubsup><mo>+</mo><msubsup><mi>b</mi><mn>1</mn><mn>2</mn></msubsup></msqrt></math><img file="EP1978731A2_D0025.tif" /></maths><maths id="math0025"><math display="block"><mi>c</mi><mo>=</mo><mfenced><mn>1</mn><mo>+</mo><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub><mo></mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub></mfenced><mo>/</mo><msqrt><msup><mfenced><mn>1</mn><mo>+</mo><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub><mo></mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub></mfenced><mn>2</mn></msup><mo>+</mo><msup><mfenced><msub><mi>a</mi><mn>1</mn></msub><mo></mo><msub><mi>b</mi><mn>1</mn></msub></mfenced><mn>2</mn></msup></msqrt></math><img file="EP1978731A2_D0026.tif" /></maths><maths id="math0026"><math display="block"><mi>d</mi><mo>=</mo><mfenced><msub><mi>a</mi><mn>1</mn></msub><mo></mo><msub><mi>b</mi><mn>1</mn></msub></mfenced><mo>/</mo><msqrt><msup><mfenced><mn>1</mn><mo>+</mo><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub><mo></mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub></mfenced><mn>2</mn></msup><mo>+</mo><msup><mfenced><msub><mi>a</mi><mn>1</mn></msub><mo></mo><msub><mi>b</mi><mn>1</mn></msub></mfenced><mn>2</mn></msup></msqrt></math><img file="EP1978731A2_D0027.tif" /></maths><maths id="math0027"><math display="block"><mi>e</mi><mo>=</mo><mfenced><mo>-</mo><mi>b</mi><mo>/</mo><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub></mfenced><mo>/</mo><msqrt><msup><mfenced><mo>-</mo><mi>b</mi><mo>/</mo><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub></mfenced><mn>2</mn></msup><mo>+</mo><msup><mfenced><mo>-</mo><mi>a</mi><mo>/</mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub></mfenced><mn>2</mn></msup></msqrt></math><img file="EP1978731A2_D0028.tif" /></maths><maths id="math0028"><math display="block"><mi>f</mi><mo>=</mo><mfenced><mo>-</mo><mi>a</mi><mo>/</mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub></mfenced><mo>/</mo><msqrt><msup><mfenced><mo>-</mo><mi>b</mi><mo>/</mo><msub><mi mathvariant="italic">σʹ</mi><mn>1</mn></msub></mfenced><mn>2</mn></msup><mo>+</mo><msup><mfenced><mo>-</mo><mi>a</mi><mo>/</mo><msub><mi mathvariant="italic">σʹ</mi><mn>3</mn></msub></mfenced><mn>2</mn></msup></msqrt></math><img file="EP1978731A2_D0029.tif" /></maths><maths id="math0029"><math display="block"><msub><mover><mi>v</mi><mo>→</mo></mover><mn>1</mn></msub><mo>=</mo><mi>V</mi><mfenced><mo>:</mo><mn>1</mn></mfenced><mo>,</mo><msub><mover><mi>v</mi><mo>→</mo></mover><mn>3</mn></msub><mo>=</mo><mi>V</mi><mfenced><mo>:</mo><mn>3</mn></mfenced></math><img file="EP1978731A2_D0030.tif" /></maths><maths id="math0030"><math display="block"><msub><mover><mi>u</mi><mo>→</mo></mover><mn>1</mn></msub><mo>=</mo><mi>U</mi><mfenced><mo>:</mo><mn>1</mn></mfenced><mo>,</mo><msub><mover><mi>u</mi><mo>→</mo></mover><mn>3</mn></msub><mo>=</mo><mi>U</mi><mfenced><mo>:</mo><mn>3</mn></mfenced><mn>.</mn></math><img file="EP1978731A2_D0031.tif" /></maths>
0121The above can be used to determine the two possible solutions expressed by: <maths id="math0031"><math display="block"><mfenced open="{" close="}"><msub><mi>R</mi><mn>1</mn></msub><mo></mo><msub><mover><mi>t</mi><mo>→</mo></mover><mn>1</mn></msub><mo></mo><msub><mover><mi>n</mi><mo>→</mo></mover><mn>1</mn></msub></mfenced><mo>,</mo><mfenced open="{" close="}"><msub><mi>R</mi><mn>2</mn></msub><mo></mo><msub><mover><mi>t</mi><mo>→</mo></mover><mn>2</mn></msub><mo></mo><msub><mover><mi>n</mi><mo>→</mo></mover><mn>2</mn></msub></mfenced></math><img file="EP1978731A2_D0032.tif" /></maths> where <o ostyle="rightarrow"><i>n</i></o><sub>1</sub> = <i>b</i><o ostyle="rightarrow">ν</o><sub>1</sub> - α<o ostyle="rightarrow">ν</o><sub>3</sub>, <o ostyle="rightarrow"><i>n</i></o><sub>2</sub> = <i>b</i><o ostyle="rightarrow">ν</o><sub>1</sub> + α<o ostyle="rightarrow">ν</o><sub>3</sub><maths id="math0032"><math display="block"><msub><mi>R</mi><mn>1</mn></msub><mo>=</mo><mi>U</mi><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>c</mi></mtd><mtd><mn>0</mn></mtd><mtd><mi>d</mi></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mo>-</mo><mi>d</mi></mtd><mtd><mn>0</mn></mtd><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo></mo><msup><mi>V</mi><mi>T</mi></msup><mo>,</mo><msub><mi>R</mi><mn>2</mn></msub><mo>=</mo><mi>U</mi><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>c</mi></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><mi>d</mi></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mi>d</mi></mtd><mtd><mn>0</mn></mtd><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo></mo><msup><mi>V</mi><mi>T</mi></msup></math><img file="EP1978731A2_D0033.tif" /></maths><o ostyle="rightarrow"><i>t</i></o><sub>1</sub> = (<i>b<o ostyle="rightarrow">u</o></i><sub>1</sub> + α<o ostyle="rightarrow"><i>u</i></o><sub>3</sub>), <o ostyle="rightarrow"><i>t</i></o>2 = -(<i>b<o ostyle="rightarrow">u</o></i>1 - α<o ostyle="rightarrow"><i>u</i></o><sub>3</sub>) (corresponding to <i>P</i><sub>2</sub>=[<i>R</i>|<i>t</i>],
0122A promise (depth positive constraint) that the direction vector <o ostyle="rightarrow"><i>n</i></o> is outward is introduced to the two possible solutions.
0123The two possible solutions are calculated by achieving consistency with the sign of <i>if</i> (<o ostyle="rightarrow"><i>n</i></o><sub>1</sub>(3) <0)<o ostyle="rightarrow"><i>t</i></o><sub>1</sub> = -<o ostyle="rightarrow"><i>t</i></o><sub>1</sub>,<o ostyle="rightarrow"><i>n</i></o><sub>1</sub> = -<o ostyle="rightarrow"><i>n</i></o><sub>1</sub> and <i>if</i> (<o ostyle="rightarrow"><i>n</i></o><sub>2</sub>(3) <0)<o ostyle="rightarrow"><i>t</i></o><sub>2</sub> = -<o ostyle="rightarrow"><i>t</i></o><sub>2</sub>,<o ostyle="rightarrow"><i>n</i></o><sub>2</sub> = -<o ostyle="rightarrow"><i>n</i></o><sub>2</sub>. Then, Epipolar error check is performed to extract one solution with less error.
0124The Epipolar error check is performed as follows. For a set of two solutions {<i>R</i><sub>1</sub>,<o ostyle="rightarrow"><i>t</i></o><sub>1</sub>/<i>d</i>,<o ostyle="rightarrow"><i>n</i></o><sub>1</sub>}.. {<i>R</i><sub>2</sub>,<o ostyle="rightarrow"><i>t</i></o><sub>2</sub>/<i>d</i>,<o ostyle="rightarrow"><i>n</i></o><sub>2</sub>} for attitude change and scene information obtained by resolving the homography calculated using the corresponding points <o ostyle="rightarrow"><i>x</i></o><sub>1</sub>, <o ostyle="rightarrow"><i>x</i></o><sub>2</sub>, Epipolar errors are calculated using the corresponding points.
0125The Epipolar error is represented by: <maths id="math0033"><math display="block"><msub><mi>e</mi><mi>i</mi></msub><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mi>j</mi><mi>n</mi></munderover></mstyle><mfenced><msup><msubsup><mover><mi>x</mi><mo>→</mo></mover><mn>2</mn><mi>j</mi></msubsup><mi>T</mi></msup><mfenced><msub><mfenced open="[" close="]"><msub><mover><mi>t</mi><mo>→</mo></mover><mi>i</mi></msub></mfenced><mi>x</mi></msub><mo></mo><msub><mi>R</mi><mi>i</mi></msub></mfenced><mo></mo><msubsup><mover><mi>x</mi><mo>→</mo></mover><mn>1</mn><mi>j</mi></msubsup></mfenced><mo>,</mo><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>2</mn><mo>,</mo><mi>j</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>2</mn><mo>,</mo><mo>…</mo><mo>,</mo><mi>n</mi></math><img file="EP1978731A2_D0034.tif" /></maths>
0126where n represents the number of the corresponding points. The solution with less error is selected as a true solution. Then, the only one solution of {<i>R, <o ostyle="rightarrow">t</o></i> , <i><o ostyle="rightarrow">n</o> }</i> is determined.
0127The reference plane normal <o ostyle="rightarrow"><i>n</i></o> in {<i>R</i>, <o ostyle="rightarrow"><i>t</i></o>/<i>d</i>, <o ostyle="rightarrow"><i>n</i></o>} obtained by resolving the image variation amount is replaced with <o ostyle="rightarrow"><i>e</i></o><sub>3</sub> = [0,0,1]<i><sup>T</sup></i> representing the normal perpendicular to the optical axis to recalculate the image variation amount as follows: <maths id="math0034"><math display="block"><mi>H</mi><mo>=</mo><mi>R</mi><mo>+</mo><mfrac><mn>1</mn><mi>d</mi></mfrac><mo></mo><msub><mover><mi>e</mi><mo>→</mo></mover><mn>3</mn></msub><mo></mo><msup><mover><mi>t</mi><mo>→</mo></mover><mi>T</mi></msup><mn>.</mn></math><img file="EP1978731A2_D0035.tif" /></maths>
0128In this manner, the image variation amount in which the appearance variation amount is excluded is calculated.
0129The recalculation of the image variation amount in which the appearance variation component is excluded may be performed by using the rotation <i>R</i> provided from the resolution of the image variation amount and the corresponding points <o ostyle="rightarrow"><i>x</i></o><sub>1</sub> , <o ostyle="rightarrow"><i>x</i></o><sub>2</sub> in the normalized image coordinate system, not by changing the reference plane normal <o ostyle="rightarrow"><i>n</i></o>.
0130First, <o ostyle="rightarrow"><i>x</i></o><sub>2</sub><sup>'</sup> = <i>R<sup>T</sup><o ostyle="rightarrow">x</o></i><sub>2</sub> is calculated. Then, calculation is performed in a least-square manner as follows to determine scaling and translation (vertical and horizontal) components which represent the influence of translation produced when the reference plane is perpendicular to the optical axis of the first camera upon the image variation amount between the corresponding points <o ostyle="rightarrow"><i>x</i></o><sub>1</sub>, <o ostyle="rightarrow"><i>x</i></o><sub>2</sub><sup>'</sup> in which the influence of the rotation <i>R</i> of the camera work was excluded: <maths id="math0035"><math display="block"><mfenced open="[" close="]"><mi>s</mi><mo></mo><msub><mi>t</mi><mi>x</mi></msub><mo></mo><msub><mi>t</mi><mi>y</mi></msub></mfenced><mo>=</mo><mi mathvariant="italic">est</mi><mfenced><msub><mover><mi>x</mi><mo>→</mo></mover><mn>1</mn></msub><mo>,</mo><msub><mover><mi>x</mi><mo>→</mo></mover><mn>2</mn></msub><mo></mo><mi>ʹ</mi></mfenced></math><img file="EP1978731A2_D0036.tif" /></maths> where est() represents processing for calculating the displacement components of scaling and translation (vertical and horizontal) between the corresponding points in the parentheses in the least-square manner.
0131Then, <maths id="math0036"><math display="block"><mi>H</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd><mtd><msub><mi>t</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd><mtd><msub><mi>t</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>s</mi></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mi>s</mi></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>⋅</mo><mi>R</mi></math><img file="EP1978731A2_D0037.tif" /></maths> is calculated. As a result, it is thus possible to stably determine in another approach the contribution of the translation camera work to the homography obtained by providing a plane. This allows calculation of the image variation amount in which the appearance variation component was excluded, similarly to the case where the reference plane normal <o ostyle="rightarrow"><i>n</i></o> is changed.
0132At a correction amount calculating step of S504, the image variation amount between the frames in which the appearance variation amount was excluded is used as input. The shake correction amount is calculated for a certain target frame by using a series including the series represented by <i>H<sup>n-k+1</sup>,H<sup>n-k+2</sup>,..., H<sup>n</sup></i> which is calculated between the past frames for the variation amount, where n represents the current frame number and k represents the number of constituent frames included in the series.
0133Next, shake correction is performed such that a motion component at high frequency is regarded as a component to be subjected to image stabilization and is removed from a video sequence. A motion component at low frequency is regarded as an intended motion component and is saved in the video.
0134Specifically, these signal components are separated through filtering. The filtering is realized by digital filtering. The number of the constituent frames of the input series corresponds to the number of taps of the digital filter.
0135<figref idref="f0007">Fig. 7</figref> is a schematic diagram for explaining the processing of calculation of the shake correction amount through the digital filtering. The digital filter is an FIR filter having five taps, by way of example. Calculation of the shake correction amount for one frame requires image variation amounts among five frames.
0136<figref idref="f0017">Fig. 18</figref> shows the procedure of the correction amount calculation.
0137First, at an accumulated variation amount calculating step of S1801, accumulated variation amounts represented by <maths id="math0037"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>1</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0038.tif" /></maths><maths id="math0038"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>2</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0039.tif" /></maths> ... , <maths id="math0039"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>5</mn></msubsup></math><img file="EP1978731A2_D0040.tif" /></maths> that are based on the top of the input series are calculated from the image variation amounts calculated between the current frame and the past frame and between the past frames at different points of time (for example, the image variation amount between the current frame and a first past frame, the image variation amount between the first past frame and a second past frame, and the image variation amount between the second past frame and a third past frame), where <maths id="math0040"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>i</mi></msubsup><mo>=</mo><msup><mi>H</mi><mrow><mi>n</mi><mo>-</mo><mi>k</mi><mo>+</mo><mi>i</mi></mrow></msup><mo>⋯</mo><msup><mi>H</mi><mrow><mi>n</mi><mo>-</mo><mi>k</mi><mo>+</mo><mn>2</mn></mrow></msup><mo></mo><msup><mi>H</mi><mrow><mi>n</mi><mo>-</mo><mi>k</mi><mo>+</mo><mn>1</mn></mrow></msup><mspace width="1em" /><mfenced><mi>i</mi><mo>≤</mo><mi>k</mi></mfenced><mn>.</mn></math><img file="EP1978731A2_D0041.tif" /></maths>
0138Thus, an example in this case is given as follows: <maths id="math0041"><math display="block"><msubsup><mi mathvariant="italic">H</mi><mi mathvariant="italic">acc</mi><mn mathvariant="italic">3</mn></msubsup><mo mathvariant="italic">=</mo><msup><mi mathvariant="italic">H</mi><mrow><mi mathvariant="italic">n</mi><mo mathvariant="italic">-</mo><mn mathvariant="italic">2</mn></mrow></msup><mo></mo><msup><mi mathvariant="italic">H</mi><mrow><mi mathvariant="italic">n</mi><mo mathvariant="italic">-</mo><mn mathvariant="italic">3</mn></mrow></msup><mo></mo><msup><mi mathvariant="italic">H</mi><mrow><mi mathvariant="italic">n</mi><mo mathvariant="italic">-</mo><mn mathvariant="italic">4</mn></mrow></msup></math><img file="EP1978731A2_D0042.tif" /></maths>
0139At a homography filtering step of S1802, filtering is performed on the series of the accumulated variation amount homography. To design the digital filter and determine the coefficient thereof, a Fourier series method and a window function method are used in combination. Characteristics including a transition area and the number of taps are determined to calculate the coefficient of the digital filter.
0140The accumulated variation amount series <maths id="math0042"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>1</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0043.tif" /></maths><maths id="math0043"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>2</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0044.tif" /></maths><i>... ,</i><maths id="math0044"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>5</mn></msubsup></math><img file="EP1978731A2_D0045.tif" /></maths> according to the number of taps of the digital filter (TAP = 5) is input and the digital filtering is performed. As a result, the filtering result <maths id="math0045"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc_filter</mi><mn>3</mn></msubsup></math><img file="EP1978731A2_D0046.tif" /></maths> influenced by the delay is output. When the digital filter is formed of the FIR filter, the delay amount is proportional to the number of taps.
0141Specifically, the delay amount is represented by (TAP)/2.
0142Accordingly, for the digital filter including five taps, the delay for two frames is produced. Therefore, when the image variation amount (<i>H</i><sup><i>n</i>-4</sup>...,<i>H<sup>n</sup></i>) from the current frame to the frame four frames before the current frame is used to calculate the accumulated variation amount ( <maths id="math0046"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>1</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0047.tif" /></maths><maths id="math0047"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>2</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0048.tif" /></maths> ... , <maths id="math0048"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>5</mn></msubsup></math><img file="EP1978731A2_D0049.tif" /></maths>) to perform the digital filtering, the result of the filtering corresponds to an accumulated variation amount <maths id="math0049"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>3</mn></msubsup></math><img file="EP1978731A2_D0050.tif" /></maths> for the frame two frames before the current frame.
0143At a correction amount calculating step of S1803, the shake correction amount is calculated by using the image variation amount <maths id="math0050"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc_filter</mi><mi>i</mi></msubsup></math><img file="EP1978731A2_D0051.tif" /></maths> restored from the filtering result and the accumulated variation amount <maths id="math0051"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>i</mi></msubsup></math><img file="EP1978731A2_D0052.tif" /></maths> of the image variation amount to the target frame, the accumulated variation amount corresponding to the target frame as a result of the delay.
0144When the digital filter is formed of a low-pass filter, <maths id="math0052"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">stb</mi><mrow><mi>n</mi><mo>-</mo><mfenced><mi>k</mi><mo>-</mo><mn>1</mn></mfenced><mo>/</mo><mn>2</mn></mrow></msubsup><mo>=</mo><msubsup><mi>H</mi><mi mathvariant="italic">acc_filter</mi><mrow><mi>k</mi><mo>-</mo><mfenced><mi>k</mi><mo>-</mo><mn>1</mn></mfenced><mo>/</mo><mn>2</mn></mrow></msubsup><mo></mo><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mrow><mi>k</mi><mo>-</mo><mfenced><mi>k</mi><mo>-</mo><mn>1</mn></mfenced><mo>/</mo><mn>2</mn></mrow></msubsup></math><img file="EP1978731A2_D0053.tif" /></maths> is calculated to determine the shake correction amount for the target frame, where k represents the number of taps of the digital filter. In this example with five taps, <maths id="math0053"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">stb</mi><mrow><mi>n</mi><mo>-</mo><mn>2</mn></mrow></msubsup><mo>=</mo><msubsup><mi>H</mi><mi mathvariant="italic">acc_filter</mi><mn>3</mn></msubsup><mo></mo><msup><mfenced><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>3</mn></msubsup></mfenced><mrow><mo>-</mo><mn>1</mn></mrow></msup></math><img file="EP1978731A2_D0054.tif" /></maths> is used to calculate the shake correction amount for the frame two frames before the current frame. In this example, if an n+1 frame is set as the current frame, an n-1 frame is subjected to image-stabilization processing.
0145With the abovementioned procedure, the shake correction amount is calculated for the corresponding frame. However, the digital filtering is typically based on the premise that the input signal is a one-dimensional signal having only the time axis.
0146Therefore, it is necessary to perform transformation (component resolution) of the homography series which is the multi-dimensional amount into a plurality of one-dimensional amount series, for example, sets of series <maths id="math0054"><math display="inline"><msubsup><mi>a</mi><mn>1</mn><mn>1</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0055.tif" /></maths><maths id="math0055"><math display="inline"><msubsup><mi>a</mi><mn>1</mn><mn>2</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0056.tif" /></maths> ... , <maths id="math0056"><math display="inline"><msubsup><mi>a</mi><mn>1</mn><mi>i</mi></msubsup></math><img file="EP1978731A2_D0057.tif" /></maths> and <maths id="math0057"><math display="inline"><msubsup><mi>a</mi><mn>2</mn><mn>1</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0058.tif" /></maths><maths id="math0058"><math display="inline"><msubsup><mi>a</mi><mn>2</mn><mn>2</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0059.tif" /></maths> ... , <maths id="math0059"><math display="inline"><msubsup><mi>a</mi><mn>2</mn><mi>i</mi></msubsup></math><img file="EP1978731A2_D0060.tif" /></maths> before the filtering step.
0147In Embodiment 1, the projective homography <maths id="math0060"><math display="block"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>i</mi></msubsup><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>h</mi><mn>1</mn></msub></mtd><mtd><msub><mi>h</mi><mn>2</mn></msub></mtd><mtd><msub><mi>h</mi><mn>3</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>4</mn></msub></mtd><mtd><msub><mi>h</mi><mn>5</mn></msub></mtd><mtd><msub><mi>h</mi><mn>6</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>7</mn></msub></mtd><mtd><msub><mi>h</mi><mn>8</mn></msub></mtd><mtd><msub><mi>h</mi><mn>9</mn></msub></mtd></mtr></mtable></mfenced><mo></mo><mfenced><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>…</mo><mo>,</mo><mi>k</mi></mfenced></math><img file="EP1978731A2_D0061.tif" /></maths>
0148which is the image variation amount between the frames is transformed into a set of one-dimensional amount series including resolved components similar to the camera works. Then, the digital filtering is performed. Thereafter, the set of one-dimensional amount series after the filtering is inversely transformed to provide the projective homography after the filtering represented as follows: <maths id="math0061"><math display="block"><msubsup><mi>H</mi><mi mathvariant="italic">acc_filter</mi><mrow><mi>k</mi><mo>-</mo><mfenced><mi>k</mi><mo>-</mo><mn>1</mn></mfenced><mo>/</mo><mn>2</mn></mrow></msubsup><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>1</mn></msub></mtd><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>2</mn></msub></mtd><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>3</mn></msub></mtd></mtr><mtr><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>4</mn></msub></mtd><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>5</mn></msub></mtd><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>6</mn></msub></mtd></mtr><mtr><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>7</mn></msub></mtd><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>8</mn></msub></mtd><mtd><msub><mi mathvariant="italic">hʹ</mi><mn>9</mn></msub></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0062.tif" /></maths>
0149<figref idref="f0008">Fig. 8</figref> is a schematic diagram for explaining the internal processing of the filtering processing of <figref idref="f0007">Fig. 7</figref> and the homography filtering step S1802 in the processing procedure of <figref idref="f0017">Fig. 18</figref> in more detail.
0150The projective homography represented as the multi-dimensional amount is transformed into a one-dimensional amount series. The one-dimensional amount series (time series) is then subjected to digital filtering. Thereafter, restoration of the filtering result of the one-dimensional amount series is performed to provide the filtered homography which is a multi-dimensional amount. The processing procedure is shown in <figref idref="f0018">Fig. 19</figref>.
0151At a component transforming step of S1901, first, each component of <maths id="math0062"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>i</mi></msubsup></math><img file="EP1978731A2_D0063.tif" /></maths> is divided by <i>h</i><sub>9</sub> to perform normalization such that <i>h</i><sub>9</sub> =1 holds for each accumulated variation amount homography represented by <maths id="math0063"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>i</mi></msubsup><mo>=</mo><mfenced open="{" close="}"><msub><mi>h</mi><mn>1</mn></msub><mo>…</mo><msub><mi>h</mi><mn>9</mn></msub></mfenced><mn>.</mn></math><img file="EP1978731A2_D0064.tif" /></maths> Then, the homography is resolved into seven components including translation (horizontal and vertical), scaling, rotation, shear, foreshortening (horizontal and vertical) which are motions in the image with the following expression: <maths id="math0064"><math display="block"><mi>H</mi><mo>=</mo><msub><mi>H</mi><mi>S</mi></msub><mo></mo><msub><mi>H</mi><mi>A</mi></msub><mo></mo><msub><mi>H</mi><mi>P</mi></msub><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi mathvariant="italic">sR</mi></mtd><mtd><mover><mi>t</mi><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>K</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>I</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mi>v</mi><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>A</mi></mtd><mtd><mover><mi>t</mi><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mi>v</mi><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0065.tif" /></maths> where <maths id="math0065"><math display="block"><mi>A</mi><mo>=</mo><mi mathvariant="italic">RK</mi><mo>+</mo><mover><mi>t</mi><mo>→</mo></mover><mo></mo><msup><mover><mi>v</mi><mo>→</mo></mover><mi>T</mi></msup><mn>.</mn></math><img file="EP1978731A2_D0066.tif" /></maths>
0152Then, <i>RK</i> = <i>A</i> - <i><o ostyle="rightarrow">tν</o><sup>T</sup></i> is calculated, and <i>R</i> and <i>K</i> are resolved by qr resolution using the property of <i><sup>K</sup></i> that is an <u style="single">upper triangular matrix.</u>
0153This achieves the resolution into eight parameters including horizontal translation tx, vertical translation ty, scaling s, rotation (in-plane rotation) θ, anisotropic magnification α of shear, direction angle φ of shear, horizontal foreshortening νx, and vertical foreshortening νy. <o ostyle="rightarrow"><i>t</i></o>, <o ostyle="rightarrow"><i>ν</i></o> , <i>R</i> and <i>K</i> are expressed as follows: <maths id="math0066"><math display="block"><mover><mi>t</mi><mo>→</mo></mover><mo>=</mo><msup><mfenced open="[" close="]"><msub><mi>t</mi><mi>x</mi></msub><mo></mo><msub><mi>t</mi><mi>y</mi></msub></mfenced><mi>T</mi></msup></math><img file="EP1978731A2_D0067.tif" /></maths><maths id="math0067"><math display="block"><mover><mi>v</mi><mo>→</mo></mover><mo>=</mo><msup><mfenced open="[" close="]"><msub><mi>v</mi><mi>x</mi></msub><mo></mo><msub><mi>v</mi><mi>y</mi></msub></mfenced><mi>T</mi></msup></math><img file="EP1978731A2_D0068.tif" /></maths><maths id="math0068"><math display="block"><mi>R</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>cos</mi><mi>θ</mi></mtd><mtd><mo>-</mo><mi>sin</mi><mi>θ</mi></mtd></mtr><mtr><mtd><mi>sin</mi><mi>θ</mi></mtd><mtd><mi>cos</mi><mi>θ</mi></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0069.tif" /></maths><maths id="math0069"><math display="block"><mi>K</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>α</mi></mtd><mtd><mi>tan</mi><mi>ϕ</mi></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0070.tif" /></maths>
0154The accumulated variation amount series <maths id="math0070"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>1</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0071.tif" /></maths><maths id="math0071"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mn>2</mn></msubsup><mo>,</mo></math><img file="EP1978731A2_D0072.tif" /></maths> ... , <maths id="math0072"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>k</mi></msubsup></math><img file="EP1978731A2_D0073.tif" /></maths> is resolved into components similar to the camera works to provide a set of one-dimensional time-series amounts represented by: <maths id="math0073"><math display="block"><mtable><mtr><mtd><msup><mfenced open="[" close="]"><msub><mi>t</mi><mi>x</mi></msub><mo></mo><msub><mi>t</mi><mi>y</mi></msub><mo></mo><mi>s</mi><mo></mo><mi>θ</mi><mo></mo><mi>α</mi><mo></mo><mi>ϕ</mi><mo></mo><msub><mi>v</mi><mi>x</mi></msub><mo></mo><msub><mi>v</mi><mi>y</mi></msub></mfenced><mi>i</mi></msup></mtd><mtd><mfenced><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>⋯</mo><mo>,</mo><mi>k</mi></mfenced></mtd></mtr></mtable><mo>,</mo></math><img file="EP1978731A2_D0074.tif" /></maths> which is then used as input to perform digital filtering for each component.
0155Since the appearance variation component has been excluded and therefore the shear component should be always α = 1 and φ = 0, the filtering may not be performed. In other words, [<i>t<sub>x</sub></i>,<i>t<sub>y</sub></i>,<i>s</i>,<i>θ</i>,<i>v<sub>x</sub></i>,<i>v<sub>y</sub></i>,]<i><sup>i</sup></i> (<i>i</i>=1,···, <i>k</i>,) may be used as a set of input time-series signals. The digital filtering is performed for each component.
0156The digital filtering processing for each component will hereinafter be described. The processing corresponds to the processing in the parentheses of <figref idref="f0008">Fig. 8</figref>. For tx component, by way of example, the digital filtering is applied to the one-dimensional time-series signal having terms corresponding to the number of taps of <maths id="math0074"><math display="inline"><mo>⌊</mo><msubsup><mi>t</mi><mi>x</mi><mn>1</mn></msubsup><mo>,</mo><msubsup><mi>t</mi><mi>x</mi><mn>2</mn></msubsup><mo>,</mo><mo>⋯</mo><mo>,</mo><msubsup><mi>t</mi><mi>x</mi><mi>k</mi></msubsup><mo>⌋</mo><mn>.</mn></math><img file="EP1978731A2_D0075.tif" /></maths>
0157Since the time-series signal is regarded as being steady, offset <o ostyle="rightarrow"><i>t</i></o> is subtracted. The offset <o ostyle="rightarrow"><i>t</i></o> represents the average value of <maths id="math0075"><math display="inline"><mo>⌊</mo><msubsup><mi>t</mi><mi>x</mi><mn>1</mn></msubsup><mo>,</mo><msubsup><mi>t</mi><mi>x</mi><mn>2</mn></msubsup><mo>,</mo><mo>⋯</mo><mo>,</mo><msubsup><mi>t</mi><mi>x</mi><mi>k</mi></msubsup><mo>⌋</mo></math><img file="EP1978731A2_D0076.tif" /></maths> that is the one-dimensional signal series. The digital filtering is applied to the time-series signal from which the offset has been subtracted.
0158At a linearization step of S1902, to linearize the variation of the scaling term (scaling component) s, logarithmic transformation is performed on the scaling term with s' = loges and then the digital filtering is applied.
0159For the foreshortening terms νx and νy, the resulting values include the influence of the scaling term due to the calculation order in the transformation into the components similar to the camera works. To remove the influence, the digital filtering is performed as appearance on the image. Thus, the foreshortening terms (components) νx and νy are multiplied by the scaling component s to be transformed to the appearance on the image (νx' = sνx, νy' = sνy), and then the filtering on the appearance on the image as the product is performed.
0160At a time-series filtering step of S1903, the filtering is performed on the dimensional signal series for each component.
0161At a non-linear restoration step of S1904, only the final time-series term (output value for the current signal) is first extracted from the time-series signal of the filtering result.
0162Here, [<i>t<sub>x</sub></i>,<i>t<sub>y</sub></i>,<i>s</i>,<i>θ</i>,<i>v<sub>x</sub></i>,<i>v<sub>y</sub></i>,]<i><sup>out</sup></i> is used as the output set. Then, <maths id="math0076"><math display="inline"><mo>⌊</mo><msub><mover><mi>t</mi><mo>‾</mo></mover><mi>x</mi></msub><mo>,</mo><msub><mover><mi>t</mi><mo>‾</mo></mover><mi>y</mi></msub><mo>,</mo><mover><mi>s</mi><mo>‾</mo></mover><mo></mo><mi>ʹ</mi><mo>,</mo><mi>θ</mi><mo>,</mo><msub><mrow><mover><mi>v</mi><mo>‾</mo></mover><mo></mo><mi>ʹ</mi></mrow><mi>x</mi></msub><mo>,</mo><msub><mrow><mover><mi>v</mi><mo>‾</mo></mover><mo></mo><mi>ʹ</mi></mrow><mi>y</mi></msub><mo>⌋</mo></math><img file="EP1978731A2_D0077.tif" /></maths> which represents each offset term subtracted before the filtering is added thereto for restoration. In addition, the scaling term and the foreshortening term are restored. Specifically, the calculation (exponential transformation) of <i>s</i>'=<i>e<sup>5</sup></i> , <i>ν<sub>x</sub>'=ν<sub>x</sub></i>/<i>s'</i> and <i>ν<sub>y</sub></i>'=<i>ν<sub>y</sub></i>/<i>s</i>' are performed.
0163However, the filtering result [<i>t<sub>x</sub>,t<sub>y</sub>,S'θ,α,φ,ν'<sub>x</sub>,ν'<sub>y</sub></i>] <i><sup>out</sup></i> determined in this manner is affected by the delay, so that this is not the motion component between the final frames. For the FIR filter which has k taps, the delay of (k-1)/2 occurs. In other words, the result corresponds to the filtering result of motions between frames (k-1)/2 before the current frame.
0164At a component inverse-transformation step of S1905, the homography form is restored from the one-directional amount set of the filtering result [<i>t<sub>x</sub>,t<sub>y</sub>,s',θ,α,φ,ν'<sub>x</sub>,ν'<sub>y</sub></i>]<i><sup>out</sup></i> with the following expression: <maths id="math0077"><math display="block"><msub><mi>H</mi><mi mathvariant="italic">fil</mi></msub><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msup><mi mathvariant="italic">sʹ</mi><mi mathvariant="italic">out</mi></msup><mo></mo><mi mathvariant="italic">Rʹ</mi></mtd><mtd><mover><mi>t</mi><mo>→</mo></mover><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi mathvariant="italic">Kʹ</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>I</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mrow><mover><mi>v</mi><mo>→</mo></mover><mo></mo><mi>ʹ</mi></mrow><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0078.tif" /></maths>
0165where <maths id="math0078"><math display="inline"><mover><mi>t</mi><mo>→</mo></mover><mo></mo><mi>ʹ</mi><mo>=</mo><msup><mfenced open="[" close="]"><msubsup><mi>t</mi><mi>x</mi><mi mathvariant="italic">out</mi></msubsup><mo></mo><msubsup><mi>t</mi><mi>y</mi><mi mathvariant="italic">out</mi></msubsup></mfenced><mi>T</mi></msup><mo>,</mo></math><img file="EP1978731A2_D0079.tif" /></maths><maths id="math0079"><math display="inline"><mover><mi>v</mi><mo>→</mo></mover><mo>=</mo><msup><mfenced open="[" close="]"><msubsup><mi mathvariant="italic">vʹ</mi><mi>x</mi><mi mathvariant="italic">out</mi></msubsup><mo></mo><msubsup><mi mathvariant="italic">vʹ</mi><mi>y</mi><mi mathvariant="italic">out</mi></msubsup></mfenced><mi>T</mi></msup><mo>,</mo></math><img file="EP1978731A2_D0080.tif" /></maths><maths id="math0080"><math display="block"><mi>R</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>cos</mi><msup><mi>θ</mi><mi mathvariant="italic">out</mi></msup></mtd><mtd><mo>-</mo><mi>sin</mi><msup><mi>θ</mi><mi mathvariant="italic">out</mi></msup></mtd></mtr><mtr><mtd><mi>sin</mi><msup><mi>θ</mi><mi mathvariant="italic">out</mi></msup></mtd><mtd><mi>cos</mi><msup><mi>θ</mi><mi mathvariant="italic">out</mi></msup></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0081.tif" /></maths> and <maths id="math0081"><math display="block"><mi>K</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msup><mtable><mtr><mtd><mi>α</mi></mtd></mtr></mtable><mi mathvariant="italic">out</mi></msup></mtd><mtd><mi>tan</mi><msup><mi>ϕ</mi><mi mathvariant="italic">out</mi></msup></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0082.tif" /></maths>
0166Then, as described above, the shake correction amount is calculated using the image variation amount restored from the filtering result <maths id="math0082"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc_filter</mi><mi>i</mi></msubsup></math><img file="EP1978731A2_D0083.tif" /></maths> and the accumulated variation amount <maths id="math0083"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>i</mi></msubsup></math><img file="EP1978731A2_D0084.tif" /></maths> of the image variation amount to the target frame, the accumulated variation amount corresponding to the target frame as a result of the delay.
0167S303 is a geometric transformation step. This step uses, as input, the shake correction parameters including the shake correction amount calculated in the shake-correction parameter calculating part 106, and the in-camera parameters and the distortion coefficient sent from the system controlling part 110. The target frame input from the work memory 105 is also used as input.
0168The geometric transformation processing part 107 applies the shake correction parameters on the target frame to perform geometric transformation, thereby providing a frame after the image stabilization processing. The geometric transformation is realized with backward mapping, for example.
0169The shake correction parameters include the in-camera parameters in image-pickup operation, the distortion coefficient, and the shake correction amount, as well as the in-camera parameters and distortion coefficient after the geometric transformation. Typically, the in-camera parameters after the geometric transformation are set to be equal to those before the geometric transformation except for the focal length. In contrast, the focal length is set to be longer than that before the geometric transformation in order to ensure redundant pixels for image stabilization, and the angle of view is determined so as not to produce a loss in video sequence. The video after the geometric transformation is typically output without any distortion.
0170The shake correction amount is represented by a 3 x 3 geometric transformation matrix for transforming image homogeneous coordinates in the normalized image coordinate system, for example.
0171<figref idref="f0006">Fig. 6</figref> shows a processing procedure for calculating the pixel coordinate position before the geometric transformation corresponding to the pixel coordinate position after the geometric transformation in the backward mapping. The pixel coordinate position before the geometric transformation corresponding to the pixel coordinate position after the geometric transformation is calculated, and the pixel value of the pixel coordinate position before the geometric transformation is calculated with interpolation. This procedure is performed for all of the pixel positions after the geometric transformation to provide the frames after image-stabilization processing.
0172At a normalization step of S601, the pixel coordinates (<i>x'y'</i>) of the frame after the image stabilization processing are transformed into coordinate values in the normalized image coordinate system. In other words, the image coordinates are determined on the camera coordinate system of focal length f = 1 in which the influence of the in-camera parameters is excluded. The in-camera parameters including the focal length <i>f<sub>c_new</sub></i> after the image stabilization processing are used to perform the transformation into the coordinate values in the normalized image coordinate system by the following expression: <maths id="math0084"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>u</mi><mi>d</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>d</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>=</mo><mi mathvariant="italic">inv</mi><mfenced><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>f</mi><mi mathvariant="italic">c_new</mi></msub><mo></mo><msub><mi>k</mi><mi>u</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><msub><mi>u</mi><mn>0</mn></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><msub><mi>f</mi><mi mathvariant="italic">c_new</mi></msub><mo></mo><msub><mi>k</mi><mi>v</mi></msub></mtd><mtd><msub><mi>v</mi><mn>0</mn></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>x</mi><mi>i</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><msub><mi>y</mi><mi>j</mi></msub><mo></mo><mi>ʹ</mi></mtd></mtr><mtr><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0085.tif" /></maths> where inv() represents the inverse matrix of the matrix in the parentheses.
0173A distortion removing step of S602 is provided for removing distortion added to the image after the geometric transformation. Generally, the image (video) after the geometric transformation includes no distortion, so that this step is omitted if the output video includes no distortion. Thus, (<i>u<sub>d</sub>',ν<sub>d</sub></i>')<i>→</i>(<i>u',v'</i>) holds.
0174In contrast, if the output video includes distortion, non-distortion coordinates (u', v') on the normalized image coordinates are calculated from distortion coordinates (ud', vd'). Specifically, the non-distortion coordinates (u', v') are determined with the procedure represented by the following expressions: <maths id="math0085"><math display="block"><mtable columnalign="left"><mtr><mtd><msup><mi>r</mi><mn>2</mn></msup></mtd><mtd><mo>=</mo><msup><mrow><msub><mi>u</mi><mi>d</mi></msub><mo></mo><mi>ʹ</mi></mrow><mn>2</mn></msup><mo>+</mo><msup><mrow><msub><mi>v</mi><mi>d</mi></msub><mo></mo><mi>ʹ</mi></mrow><mn>2</mn></msup></mtd></mtr><mtr><mtd><mi>K</mi></mtd><mtd><mo>=</mo><mn>1</mn><mo>+</mo><msub><mi>k</mi><mn>1</mn></msub><mo></mo><mi>r</mi><mo>+</mo><msub><mi>k</mi><mn>2</mn></msub><mo></mo><msup><mi>r</mi><mn>2</mn></msup><mo>+</mo><msub><mi>k</mi><mn>3</mn></msub><mo></mo><msup><mi>r</mi><mn>3</mn></msup><mo>+</mo><mo>⋯</mo></mtd></mtr><mtr><mtd><mi mathvariant="italic">uʹ</mi></mtd><mtd><mtable><mtr><mtd><mtable><mtr><mtd><mo>=</mo><msub><mi>u</mi><mi>d</mi></msub><mo></mo><mi>ʹ</mi><mo>/</mo><mi>K</mi></mtd></mtr></mtable></mtd><mtd><mi mathvariant="italic">vʹ</mi><mo>=</mo><msub><mi>v</mi><mi>d</mi></msub><mo></mo><mi>ʹ</mi><mo>/</mo><mi>K</mi></mtd></mtr></mtable><mn>.</mn></mtd></mtr></mtable></math><img file="EP1978731A2_D0086.tif" /></maths>
0175At a geometric transformation step of S603, the inverse transformation of shake correction is performed on the normalized image coordinates. If a geometric transformation matrix for representing the shake correction amount is a 3 x 3 matrix <i>H</i>, the inverse matrix inv(<i>H</i>) is applied to normalized coordinate points (u', ν') for the backward matching. Specifically, the normalized image coordinates (u, v) before the shake correction are calculated by the following expression: <maths id="math0086"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>u</mi><mi>p</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>p</mi></msub></mtd></mtr><mtr><mtd><mi>m</mi></mtd></mtr></mtable></mfenced><mo>=</mo><mi mathvariant="italic">inv</mi><mfenced><mi>H</mi></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi mathvariant="italic">uʹ</mi></mtd></mtr><mtr><mtd><mi mathvariant="italic">vʹ</mi></mtd></mtr><mtr><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0087.tif" /></maths> where u=u<i><sub>p</sub></i>/<i>m</i> and ν=ν<i><sub>p</sub></i>/<i>m</i>.
0176At a distortion adding step of S604, the distortion before the geometric transformation is added to the normalized image coordinate values. The following expression is used to add displacement from the distortion in the radial directions: <maths id="math0087"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>u</mi><mi>d</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>d</mi></msub></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced><mn>1</mn><mo>+</mo><msub><mi>k</mi><mn>1</mn></msub><mo></mo><mi>r</mi><mo>+</mo><msub><mi>k</mi><mn>2</mn></msub><mo></mo><msup><mi>r</mi><mn>2</mn></msup><mo>+</mo><msub><mi>k</mi><mn>3</mn></msub><mo></mo><msup><mi>r</mi><mn>3</mn></msup><mo>+</mo><mo>⋯</mo></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>u</mi></mtd></mtr><mtr><mtd><mi>v</mi></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0088.tif" /></maths> where <i>r</i><sup>2</sup><i>= u</i><sup>2</sup>+v<sup>2</sup>, and k1, k2, and k3 represent radial distortion coefficients of first, second, and third orders, respectively.
0177S605 is a normalization restoration step. At this step, the in-camera parameters are applied to the normalized image coordinates (ud, vd) before the shake correction having the distortion by the following expression to provide pixel coordinates on the input frame: <maths id="math0088"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><mi>x</mi></mtd></mtr><mtr><mtd><mi>y</mi></mtd></mtr><mtr><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>f</mi><mi>u</mi></msub><mo></mo><msub><mi>k</mi><mi>u</mi></msub></mtd><mtd><msub><mi>f</mi><mi>u</mi></msub><mo></mo><msub><mi>k</mi><mi>u</mi></msub><mo></mo><mi>cot</mi><mi>ϕ</mi></mtd><mtd><msub><mi>u</mi><mn>0</mn></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><msub><mi>f</mi><mi>v</mi></msub><mo></mo><msub><mi>k</mi><mi>v</mi></msub><mo></mo><mi>sin</mi><mi>ϕ</mi></mtd><mtd><msub><mi>v</mi><mn>0</mn></msub></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>u</mi><mi>d</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>d</mi></msub></mtd></mtr><mtr><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0089.tif" /></maths>
0178The pixel values of the pixel coordinates are sampled with interpolation such as bi-cubic interpolation to provide pixel values of each pixel of the frame after the image-stabilization processing. The backward mapping is performed on all of the frames after the image-stabilization processing to complete the geometric transformation processing.
0179With the abovementioned processing steps, the image-stabilization processing is performed on each frame of the video signal. The video stream after the image-stabilization processing is encoded in a video format such as NTSC and MPEG4 in the encoding/coding part 108.
0180Finally, the encoded video stream is recoded on the recording medium in the recording part 113.
0181The processing in Embodiment 1 allows preservation of a motion due to an intended camera work included in video picked up by a video camera having an image-pickup optical system of a very short focal length and effective suppression of an image shake due to an unintended camera shake.
0182The abovementioned time-series filtering in Embodiment 1 is performed by using the digital filtering. However, another filtering method may be used to separate image variation amounts produced from an intended camera work and an unintended camera shake.
0183While Embodiment 1 has been described in conjunction with the image stabilizing apparatus mounted on the video camera, the present invention can be realized as an image stabilizing apparatus which functions alone without having an image-pickup optical system or an image-pickup element. For example, a computer program for realizing the abovementioned image-stabilization processing function is installed on a personal computer to allow the personal computer to be used as the image stabilizing apparatus. In this case, video information taken by a video camera is input through a cable, a wireless LAN or the like to the personal computer which performs the image-stabilization processing.
[Embodiment 2]
0184An image stabilizing apparatus which is Embodiment 2 and a video camera including the apparatus will hereinafter be described. The configurations of the image stabilizing apparatus and the video camera are identical to those in Embodiment 1. Basic portions of the image-stabilization processing procedure are identical to the processing procedure in Embodiment 1 described with <figref idref="f0003">Fig. 3</figref>. <figref idref="f0009">Fig. 9</figref> shows a shake-correction parameter calculating step which represents a difference between the processing procedure of Embodiment 2 and that in Embodiment 1.
0185The processing at a normalization step of S901 is similar to the processing at the normalization step S501 of <figref idref="f0005">Fig. 5</figref>.
0186At an attitude amount calculating step of S902, the motion vectors between frames transformed into a normalized image coordinate system are used as input to calculate an image variation amount between the frames. An attitude variation amount of the camera determined between images is used as the index of the image variation amount.
0187In the following, two possible solutions for the attitude variation amount: <maths id="math0089"><math display="block"><mfenced open="{" close="}"><msub><mi>R</mi><mrow><mi>x</mi><mo></mo><mn>1</mn></mrow></msub><mo></mo><msub><mi>R</mi><mrow><mi>y</mi><mo></mo><mn>1</mn></mrow></msub><mo></mo><msub><mi>R</mi><mrow><mi>z</mi><mo></mo><mn>1</mn></mrow></msub><mo></mo><msub><mi>t</mi><mrow><mi>x</mi><mo></mo><mn>1</mn></mrow></msub><mo></mo><msub><mi>t</mi><mrow><mi>y</mi><mo></mo><mn>1</mn></mrow></msub><mo></mo><msub><mi>t</mi><mrow><mi>z</mi><mo></mo><mn>1</mn></mrow></msub><mo></mo><msub><mover><mi>n</mi><mo>→</mo></mover><mn>1</mn></msub></mfenced><mo>,</mo><mfenced open="{" close="}"><msub><mi>R</mi><mrow><mi>x</mi><mo></mo><mn>2</mn></mrow></msub><mo></mo><msub><mi>R</mi><mrow><mi>y</mi><mo></mo><mn>2</mn></mrow></msub><mo></mo><msub><mi>R</mi><mrow><mi>z</mi><mo></mo><mn>2</mn></mrow></msub><mo></mo><msub><mi>t</mi><mrow><mi>x</mi><mo></mo><mn>2</mn></mrow></msub><mo></mo><msub><mi>t</mi><mrow><mi>y</mi><mo></mo><mn>2</mn></mrow></msub><mo></mo><msub><mi>t</mi><mrow><mi>z</mi><mo></mo><mn>2</mn></mrow></msub><mo></mo><msub><mover><mi>n</mi><mo>→</mo></mover><mn>2</mn></msub></mfenced></math><img file="EP1978731A2_D0090.tif" /></maths> are calculated from the motion vectors between frames transformed into the normalized image coordinate system.
0188The two possible solutions for the attitude variation amount are calculated by using, for example, the method described in "<nplcit id="ncit0005" npl-type="b"><text>Understanding Images - Mathematics of Three-Dimension Recognition," Kenichi Kanatani, Morikita Publishing Co., Ltd.</text></nplcit> In this case, corresponding points between frames are needed to be optical flow. That is, it is necessary that the camera attitude variation between frames represented by {<i>R<sub>x</sub></i>,<i>R<sub>y</sub></i>,<i>R<sub>z</sub></i>,<i>t<sub>x</sub></i>,<i>t<sub>y</sub></i>,<i>t<sub>z</sub></i>} is extremely small. In other words, it is necessary that the frame rate of video is sufficiently high as compared with camera works and ccs(<i>R</i><sub>i</sub>) ≅ 0, sin(<i>R<sub>i</sub></i>) ≅ <i>R<sub>i</sub></i> are satisfied.
0189In the following expressions, a minute rotation {<i>R<sub>x</sub>,R<sub>y</sub>,R<sub>z</sub></i>} is represented as {ω<sub>1</sub>,ω<sub>2</sub>,ω<sub>3</sub>}. <maths id="math0090"><math display="block"><mtable columnalign="left"><mtr><mtd><mi>W</mi></mtd><mtd><mo>=</mo><mfenced><mtable><mtr><mtd><mfenced><mn>2</mn><mo></mo><mi>A</mi><mo>-</mo><mi>D</mi></mfenced><mo>/</mo><mn>3</mn></mtd><mtd><mi>C</mi></mtd><mtd><mo>-</mo><mi>E</mi></mtd></mtr><mtr><mtd><mi>B</mi></mtd><mtd><mfenced><mo>-</mo><mi>A</mi><mo>+</mo><mn>2</mn><mo></mo><mi>D</mi></mfenced><mo>/</mo><mn>3</mn></mtd><mtd><mo>-</mo><mi>F</mi></mtd></mtr><mtr><mtd><mi>U</mi></mtd><mtd><mi>V</mi></mtd><mtd><mo>-</mo><mfenced><mi>A</mi><mo>+</mo><mi>D</mi></mfenced><mo>/</mo><mn>3</mn></mtd></mtr></mtable></mfenced></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mo>=</mo><mo>-</mo><mfrac><mrow><mi mathvariant="italic">pa</mi><mo>+</mo><mi mathvariant="italic">qb</mi><mo>-</mo><mi>c</mi></mrow><mrow><mn>3</mn><mo></mo><mi>r</mi></mrow></mfrac><mfenced><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfrac><mn>1</mn><mi>r</mi></mfrac><mo></mo><mfenced><mtable><mtr><mtd><mi>p</mi></mtd></mtr><mtr><mtd><mi>q</mi></mtd></mtr><mtr><mtd><mo>-</mo><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced><mtable><mtr><mtd><mi>a</mi></mtd><mtd><mi>b</mi></mtd><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo>+</mo><mfenced><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>2</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>1</mn></msub></mtd><mtd><mn>0</mn></mtd></mtr></mtable></mfenced></mtd></mtr></mtable></math><img file="EP1978731A2_D0091.tif" /></maths>
0190A symmetric portion <i>W</i><sub>s</sub> of <i>W</i> and an asymmetric portion <i>W</i><sub>a</sub> thereof are defined as follows: <maths id="math0091"><math display="block"><msub><mi>W</mi><mi>s</mi></msub><mo>=</mo><mfrac><mn>1</mn><mn>2</mn></mfrac><mo></mo><mfenced><mi>W</mi><mo>+</mo><msup><mi>W</mi><mi>T</mi></msup></mfenced></math><img file="EP1978731A2_D0092.tif" /></maths><maths id="math0092"><math display="block"><msub><mi>W</mi><mi>a</mi></msub><mo>=</mo><mfrac><mn>1</mn><mn>2</mn></mfrac><mo></mo><mfenced><mi>W</mi><mo>-</mo><msup><mi>W</mi><mi>T</mi></msup></mfenced><mn>.</mn></math><img file="EP1978731A2_D0093.tif" /></maths>
0191These have the following meanings, respectively: <maths id="math0093"><math display="block"><mtable columnalign="left"><mtr><mtd><msub><mi>W</mi><mi>s</mi></msub></mtd><mtd><mo>=</mo><mfenced><mtable><mtr><mtd><mfenced><mn>2</mn><mo></mo><mi>A</mi><mo>-</mo><mi>D</mi></mfenced><mo>/</mo><mn>3</mn></mtd><mtd><mfenced><mi>B</mi><mo>+</mo><mi>C</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mfenced><mi>U</mi><mo>-</mo><mi>E</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mfenced><mi>B</mi><mo>+</mo><mi>C</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mfenced><mo>-</mo><mi>A</mi><mo>+</mo><mn>2</mn><mo></mo><mi>D</mi></mfenced><mo>/</mo><mn>3</mn></mtd><mtd><mfenced><mi>V</mi><mo>-</mo><mi>F</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mfenced><mi>U</mi><mo>-</mo><mi>E</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mfenced><mi>V</mi><mo>-</mo><mi>F</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mo>-</mo><mfenced><mi>A</mi><mo>+</mo><mi>D</mi></mfenced><mo>/</mo><mn>3</mn></mtd></mtr></mtable></mfenced></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mo>=</mo><mo>-</mo><mfrac><mrow><mi mathvariant="italic">pa</mi><mo>+</mo><mi mathvariant="italic">qb</mi><mo>-</mo><mi>c</mi></mrow><mrow><mn>3</mn><mo></mo><mi>r</mi></mrow></mfrac><mrow><mo>(</mo><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable><mo>)</mo><mo>+</mo></mrow><mfrac><mn>1</mn><mrow><mn>2</mn><mo></mo><mi>r</mi></mrow></mfrac><mo></mo><mfenced open="[" close="]"><mfenced><mtable><mtr><mtd><mi>p</mi></mtd></mtr><mtr><mtd><mi>q</mi></mtd></mtr><mtr><mtd><mo>-</mo><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced><mtable><mtr><mtd><mi>a</mi></mtd><mtd><mi>b</mi></mtd><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo>+</mo><mfenced><mtable><mtr><mtd><mi>a</mi></mtd></mtr><mtr><mtd><mi>b</mi></mtd></mtr><mtr><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo></mo><mfenced><mtable><mtr><mtd><mi>p</mi></mtd><mtd><mi>q</mi></mtd><mtd><mo>-</mo><mn>1</mn></mtd></mtr></mtable></mfenced></mfenced></mtd></mtr></mtable></math><img file="EP1978731A2_D0094.tif" /></maths><maths id="math0094"><math display="block"><mtable columnalign="left"><mtr><mtd><msub><mi>W</mi><mi>a</mi></msub></mtd><mtd><mo>=</mo><mfenced><mtable><mtr><mtd><mn>0</mn></mtd><mtd><mo>-</mo><mfenced><mi>B</mi><mo>-</mo><mi>C</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mo>-</mo><mfenced><mi>U</mi><mo>+</mo><mi>E</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mfenced><mi>B</mi><mo>-</mo><mi>C</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><mfenced><mi>V</mi><mo>+</mo><mi>F</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mfenced><mi>U</mi><mo>+</mo><mi>E</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mfenced><mi>V</mi><mo>+</mo><mi>F</mi></mfenced><mo>/</mo><mn>2</mn></mtd><mtd><mn>0</mn></mtd></mtr></mtable></mfenced></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mo>=</mo><mfrac><mn>1</mn><mrow><mn>2</mn><mo></mo><mi>r</mi></mrow></mfrac><mo></mo><mfenced open="[" close="]"><mfenced><mtable><mtr><mtd><mi>p</mi></mtd></mtr><mtr><mtd><mi>q</mi></mtd></mtr><mtr><mtd><mo>-</mo><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced><mtable><mtr><mtd><mi>a</mi></mtd><mtd><mi>b</mi></mtd><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo>+</mo><mfenced><mtable><mtr><mtd><mi>a</mi></mtd></mtr><mtr><mtd><mi>b</mi></mtd></mtr><mtr><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo></mo><mfenced><mtable><mtr><mtd><mi>p</mi></mtd><mtd><mi>q</mi></mtd><mtd><mo>-</mo><mn>1</mn></mtd></mtr></mtable></mfenced></mfenced><mo>+</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>2</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>1</mn></msub></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></mtd></mtr></mtable></math><img file="EP1978731A2_D0095.tif" /></maths>
0192These are used to provide solutions.
0193The eigenvalues of Ws are set to σ1 ≥ σ2 ≥ σ3 ≥ 0, and the corresponding eigenvectors {<o ostyle="rightarrow"><i>u</i></o><sub>1</sub>, <o ostyle="rightarrow"><i>u</i></o><sub>2</sub>, <o ostyle="rightarrow"><i>u</i></o><sub>3</sub>} are set to unit vectors orthogonal to each other.
0194If σ1 = σ2 = σ3 = 0, that is, <i>W<sub>s</sub></i> = <o ostyle="rightarrow"><i>0</i></o>holds, then motion parameters are given as follows: <maths id="math0095"><math display="block"><mfenced><mtable><mtr><mtd><mi>a</mi></mtd></mtr><mtr><mtd><mi>b</mi></mtd></mtr><mtr><mtd><mi>c</mi></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced><mtable><mtr><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0096.tif" /></maths><maths id="math0096"><math display="block"><mfenced><mtable><mtr><mtd><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>3</mn></msub></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced><mtable><mtr><mtd><mfenced><mi>V</mi><mo>+</mo><mi>F</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mo>-</mo><mfenced><mi>U</mi><mo>+</mo><mi>E</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mfenced><mi>B</mi><mo>-</mo><mi>C</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0097.tif" /></maths> and plane parameters {p, q, r} are indefinite.
0195If not, the two possible solutions are determined as follows.
0196First, the <u style="single">gradient</u> {p, q} of a plane serving as the reference for calculating the image variation amount is determined with the following expressions: <maths id="math0097"><math display="block"><mi mathvariant="italic">p</mi><mo mathvariant="italic">=</mo><mfrac><mi mathvariant="italic">pʹ</mi><mi mathvariant="italic">lʹ</mi></mfrac><mo mathvariant="italic">,</mo><mi mathvariant="italic">q</mi><mo mathvariant="italic">=</mo><mfrac><mi mathvariant="italic">qʹ</mi><mi mathvariant="italic">lʹ</mi></mfrac></math><img file="EP1978731A2_D0098.tif" /></maths><maths id="math0098"><math display="block"><mfenced><mtable><mtr><mtd><mi mathvariant="italic">pʹ</mi></mtd></mtr><mtr><mtd><mi mathvariant="italic">qʹ</mi></mtd></mtr><mtr><mtd><mi mathvariant="italic">rʹ</mi></mtd></mtr></mtable></mfenced><mo>=</mo><mo>±</mo><msqrt><msub><mi>σ</mi><mn>1</mn></msub><mo>-</mo><msub><mi>σ</mi><mn>2</mn></msub></msqrt><mo></mo><msub><mover><mi>u</mi><mo>→</mo></mover><mn>1</mn></msub><mo>-</mo><msqrt><msub><mi>σ</mi><mn>2</mn></msub><mo>-</mo><msub><mi>σ</mi><mn>3</mn></msub></msqrt><mo></mo><msub><mover><mi>u</mi><mo>→</mo></mover><mn>3</mn></msub><mn>.</mn></math><img file="EP1978731A2_D0099.tif" /></maths>
0197Next, a ratio of the translation speed (a, b, c) to the distance r is determined as follows: <maths id="math0099"><math display="block"><mfenced><mtable><mtr><mtd><mi>a</mi><mo>/</mo><mi>r</mi></mtd></mtr><mtr><mtd><mi>b</mi><mo>/</mo><mi>r</mi></mtd></mtr><mtr><mtd><mi>c</mi><mo>/</mo><mi>r</mi></mtd></mtr></mtable></mfenced><mo>=</mo><mo>-</mo><mi mathvariant="italic">lʹ</mi><mo></mo><mfenced><mo>±</mo><msqrt><msub><mi>σ</mi><mn>1</mn></msub><mo>-</mo><msub><mi>σ</mi><mn>2</mn></msub></msqrt><mo></mo><msub><mover><mi>u</mi><mo>→</mo></mover><mn>1</mn></msub><mo>-</mo><msqrt><msub><mi>σ</mi><mn>2</mn></msub><mo>-</mo><msub><mi>σ</mi><mn>3</mn></msub></msqrt><mo></mo><msub><mover><mi>u</mi><mo>→</mo></mover><mn>3</mn></msub></mfenced><mn>.</mn></math><img file="EP1978731A2_D0100.tif" /></maths>
0198Finally, the rotation speed (ω1, ω2, ω3) is calculated as follows: <maths id="math0100"><math display="block"><mfenced><mtable><mtr><mtd><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>3</mn></msub></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced><mtable><mtr><mtd><mfenced><mi>V</mi><mo>+</mo><mi>F</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mo>-</mo><mfenced><mi>U</mi><mo>+</mo><mi>E</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr><mtr><mtd><mfenced><mi>B</mi><mo>-</mo><mi>C</mi></mfenced><mo>/</mo><mn>2</mn></mtd></mtr></mtable></mfenced><mo>+</mo><mfrac><mn>1</mn><mn>2</mn></mfrac><mo></mo><mfenced><mtable><mtr><mtd><mi>p</mi></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mi>q</mi></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced><mtable><mtr><mtd><mi>a</mi><mo>/</mo><mi>r</mi></mtd></mtr><mtr><mtd><mi>b</mi><mo>/</mo><mi>r</mi></mtd></mtr><mtr><mtd><mi>c</mi><mo>/</mo><mi>r</mi></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0101.tif" /></maths> In addition, the following expressions hold: <maths id="math0101"><math display="block"><mover><mi>n</mi><mo>→</mo></mover><mo>=</mo><mfenced open="{" close="}"><mi>p</mi><mo>/</mo><msqrt><msup><mi>p</mi><mn>2</mn></msup><mo>+</mo><msup><mi>q</mi><mn>2</mn></msup><mo>+</mo><mn>1</mn></msqrt><mo>,</mo><mi>q</mi><mo>/</mo><msqrt><msup><mi>p</mi><mn>2</mn></msup><mo>+</mo><msup><mi>q</mi><mn>2</mn></msup><mo>+</mo><mn>1</mn></msqrt><mo>,</mo><mo>-</mo><mn>1</mn><mo>/</mo><msqrt><msup><mi>p</mi><mn>2</mn></msup><mo>+</mo><msup><mi>q</mi><mn>2</mn></msup><mo>+</mo><mn>1</mn></msqrt></mfenced></math><img file="EP1978731A2_D0102.tif" /></maths><maths id="math0102"><math display="block"><mi>d</mi><mo>=</mo><mo>-</mo><mi>r</mi><mo>/</mo><msqrt><msup><mi>p</mi><mn>2</mn></msup><mo>+</mo><msup><mi>q</mi><mn>2</mn></msup><mo>+</mo><mn>1</mn></msqrt><mn>.</mn></math><img file="EP1978731A2_D0103.tif" /></maths>
0199With the abovementioned processing, the two possible solutions for the attitude variation amount {<i>R</i><sub><i>x</i>1</sub>, <i>R</i><sub><i>y</i>1</sub>, <i>R</i><sub><i>z</i>1</sub>, <i>t</i><sub><i>x</i>1</sub>, <i>t</i><sub><i>y</i>1</sub>, <i>t</i><sub><i>z</i>1</sub>, <o ostyle="rightarrow"><i>n</i></o><sub>1</sub>}, {<i>R</i><sub><i>x</i>2</sub>, <i>R</i><sub><i>y</i>2</sub>, <i>R</i><sub><i>z</i>2</sub>, <i>t</i><sub><i>x</i>2</sub>, <i>t</i><sub><i>y</i>2</sub>, <i>t</i><sub><i>z</i>2</sub>, <o ostyle="rightarrow"><i>n</i></o><sub>2</sub>} are calculated. The shift of the coordinate system with (a, b, c) and (ω1, ω2, ω3) is represented as follows: <maths id="math0103"><math display="block"><mover><mi>X</mi><mo>˙</mo></mover><mo>=</mo><mfenced><mtable><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>3</mn></msub></mtd></mtr></mtable></mfenced><mo>×</mo><mi>X</mi><mo>-</mo><mfenced><mtable><mtr><mtd><mi>a</mi></mtd></mtr><mtr><mtd><mi>b</mi></mtd></mtr><mtr><mtd><mi>c</mi></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0104.tif" /></maths> where x represents multiplication of elements.
0200Thus, the rotation matrix is approximated as follows: <maths id="math0104"><math display="block"><mi>R</mi><mo>≅</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>2</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>1</mn></msub></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0105.tif" /></maths>
0201Based on this relationship, for the two possible solutions {<i>R</i><sub>1</sub>,<o ostyle="rightarrow">t</o><sub>1</sub>/<i>d</i>,<o ostyle="rightarrow"><i>n</i></o><sub>1</sub>}, {<i>R</i><sub>2</sub>,<o ostyle="rightarrow">t</o><sub>2</sub>/<i>d</i>,<o ostyle="rightarrow"><i>n</i></o><sub>2</sub>} for the attitude variation and scene information provided by resolution of the homography determined from the corresponding points <o ostyle="rightarrow"><i>x</i></o><sub>1</sub>, <o ostyle="rightarrow"><i>x</i></o><sub>2</sub>, <img file="EP1978731A2_D0106.tif" /> , the Epipolar error represented by: <maths id="math0105"><math display="block"><msub><mi>e</mi><mi>i</mi></msub><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mi>j</mi><mi>n</mi></munderover></mstyle><mfenced><msup><msubsup><mover><mi>x</mi><mo>→</mo></mover><mn>2</mn><mi>j</mi></msubsup><mi>T</mi></msup><mfenced><msub><mfenced open="[" close="]"><msub><mover><mi>t</mi><mo>→</mo></mover><mi>i</mi></msub></mfenced><mi>x</mi></msub><mo></mo><msub><mi>R</mi><mi>i</mi></msub></mfenced><mo></mo><msubsup><mover><mi>x</mi><mo>→</mo></mover><mn>1</mn><mi>j</mi></msubsup></mfenced><mo>,</mo><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>2</mn><mo>,</mo><mi>j</mi><mo>=</mo><mn>1</mn><mo>,</mo><mn>2</mn><mo>,</mo><mo>…</mo><mo>,</mo><mi>n</mi></math><img file="EP1978731A2_D0107.tif" /></maths> is calculated with the corresponding points. A set with less error is selected as a true set represented by {<i>R<sub>x</sub>,R<sub>y</sub>,R<sub>z</sub>,t<sub>x</sub>,t<sub>y</sub>,t<sub>z</sub></i>}
0202At a homography calculating step of S903, the homography is calculated as follows from the attitude variation amount {<i>R<sub>x</sub>,R<sub>y</sub>,R<sub>z</sub>,t<sub>x</sub>,t<sub>y</sub>,t<sub>z</sub></i>} which is the image variation amount determined between the frames: <maths id="math0106"><math display="block"><mi>H</mi><mo>=</mo><mfenced><mi>R</mi><mo>+</mo><mfrac><mn>1</mn><mi>d</mi></mfrac><mo></mo><mover><mi>t</mi><mo>→</mo></mover><mo></mo><msup><mover><mi>n</mi><mo>→</mo></mover><mi>T</mi></msup></mfenced></math><img file="EP1978731A2_D0108.tif" /></maths> where <maths id="math0107"><math display="block"><mi>R</mi><mo>≅</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>2</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>1</mn></msub></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0109.tif" /></maths>
0203The processing at a correction amount calculating step of S904 is similar to the processing at the correction amount calculating step S504 in <figref idref="f0005">Fig. 5</figref>.
0204In Embodiment 2, the processing step with the abovementioned changes made to Embodiment 1 is performed to apply image-stabilization processing to each frame of the video signal. Thus, Embodiment 2 has the advantage of providing the same effects as those in the method of Embodiment 1 through simple processing when slight motion changes occur between the frames.
0205The abovementioned time-series filtering in Embodiment 2 is performed by using the digital filtering. However, another filtering method may be used to separate image variation amounts produced due to an intended camera work and an unintended camera shake.
[Embodiment 3]
0206An image stabilizing apparatus which is Embodiment 3 and a video camera including the apparatus will hereinafter be described. The configurations of the image stabilizing apparatus and the video camera are identical to those in Embodiment 1. Basic portions of the image-stabilization processing procedure are identical to the processing procedure in Embodiment 1 described with <figref idref="f0003">Fig. 3</figref>. <figref idref="f0010">Fig. 10</figref> shows a shake-correction parameter calculating step which represents a difference between the processing procedure of Embodiment 3 and that in Embodiment 1.
0207The processing at a normalization step of S1001 is similar to the processing at the normalization step S501 of <figref idref="f0005">Fig. 5</figref>.
0208At a fundamental matrix calculating step of S1002, the motion vectors between the frames transformed into the normalized image coordinate system are used as input to calculate an image variation amount between the frames. A fundamental matrix <i>E</i> determined between images is used as the index of the image variation amount. The fundamental matrix <i>E</i> is calculated by using information on corresponding points <o ostyle="rightarrow"><i>x</i></o><sub>1</sub>, <o ostyle="rightarrow"><i>x</i></o><sub>2</sub> between the frames.
0209Specifically, assuming that <o ostyle="rightarrow"><i>x</i></o><sub>1</sub> = [<i>x,y</i>,1]<i><sup>T</sup></i> and <o ostyle="rightarrow"><i>x</i></o><sub>2</sub> = [<i>x',y'</i>,1]<i><sup>T</sup></i>, a linear equation: <maths id="math0108"><math display="block"><mi>A</mi><mo></mo><mover><mi>e</mi><mo>→</mo></mover><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msup><mi mathvariant="italic">xʹ</mi><mn>1</mn></msup><mo></mo><msup><mi>x</mi><mn>1</mn></msup></mtd><mtd><msup><mi mathvariant="italic">xʹ</mi><mn>1</mn></msup><mo></mo><msup><mi>y</mi><mn>1</mn></msup></mtd><mtd><msup><mi mathvariant="italic">xʹ</mi><mn>1</mn></msup></mtd><mtd><msup><mi mathvariant="italic">xʹ</mi><mn>1</mn></msup><mo></mo><msup><mi>x</mi><mn>1</mn></msup></mtd><mtd><msup><mi mathvariant="italic">yʹ</mi><mn>1</mn></msup><mo></mo><msup><mi>y</mi><mn>1</mn></msup></mtd><mtd><msup><mi mathvariant="italic">yʹ</mi><mn>1</mn></msup></mtd><mtd><msup><mi>x</mi><mn>1</mn></msup></mtd><mtd><msup><mi>y</mi><mn>1</mn></msup></mtd><mtd><mn>1</mn></mtd></mtr><mtr><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd><mtd><mo>⋮</mo></mtd></mtr><mtr><mtd><msup><mi mathvariant="italic">xʹ</mi><mi>n</mi></msup><mo></mo><msup><mi>x</mi><mi>n</mi></msup></mtd><mtd><msup><mi mathvariant="italic">xʹ</mi><mi>n</mi></msup><mo></mo><msup><mi>y</mi><mi>n</mi></msup></mtd><mtd><msup><mi mathvariant="italic">xʹ</mi><mi>n</mi></msup></mtd><mtd><msup><mi mathvariant="italic">xʹ</mi><mi>n</mi></msup><mo></mo><msup><mi>x</mi><mi>n</mi></msup></mtd><mtd><msup><mi mathvariant="italic">yʹ</mi><mi>n</mi></msup><mo></mo><msup><mi>y</mi><mi>n</mi></msup></mtd><mtd><msup><mi mathvariant="italic">yʹ</mi><mi>n</mi></msup></mtd><mtd><msup><mi>x</mi><mi>n</mi></msup></mtd><mtd><msup><mi>y</mi><mi>n</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mover><mi>e</mi><mo>→</mo></mover><mo>=</mo><mover><mn>0</mn><mo>→</mo></mover></math><img file="EP1978731A2_D0110.tif" /></maths> is formed. It is overdetermined if the number n of the corresponding points is equal to or larger than eight, and the vector form of the fundamental matrix represented by (9 × 1)<o ostyle="rightarrow"><i>e</i></o> can be obtained in the least-square manner. The vector form <o ostyle="rightarrow"><i>e</i></o> is shaped into a 3 x 3 matrix form to provide the fundamental matrix <i>E</i>.
0210At a homography calculating step of S1003, the projective homography is calculated from the fundamental matrix which represents the image variation amount determined between the frames and the corresponding points in the normalized coordinates used in the calculation of the fundamental matrix. The fundamental matrix <i>E</i> between the frames is formed with camera work rotation <i>R</i> and translation <o ostyle="rightarrow"><i>t</i></o> between the frames as follows: <maths id="math0109"><math display="block"><mi>E</mi><mo>=</mo><msub><mfenced open="[" close="]"><mi>T</mi></mfenced><mi>x</mi></msub><mo></mo><mi>R</mi></math><img file="EP1978731A2_D0111.tif" /></maths> where └<o ostyle="rightarrow"><i>t</i></o>┘<sub>x</sub> represents a torsional symmetric vector of the translation vector <o ostyle="rightarrow"><i>t</i></o> : <maths id="math0110"><math display="block"><msub><mrow><mo>⌊</mo><mover><mi>t</mi><mo>→</mo></mover><mo>⌋</mo></mrow><mi>x</mi></msub><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>0</mn></mtd><mtd><mo>-</mo><msub><mi>t</mi><mi>Z</mi></msub></mtd><mtd><msub><mi>t</mi><mi>Y</mi></msub></mtd></mtr><mtr><mtd><msub><mi>t</mi><mi>Z</mi></msub></mtd><mtd><mn>0</mn></mtd><mtd><mo>-</mo><msub><mi>t</mi><mi>X</mi></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>t</mi><mi>Y</mi></msub></mtd><mtd><msub><mi>t</mi><mi>X</mi></msub></mtd><mtd><mn>0</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0112.tif" /></maths>
0211The fundamental matrix is resolved by using a singular value resolution <i>USV<sup>T</sup> = SVD</i>(<i>X</i>)<i>. S</i> represents a matrix having a singular value in diagonal elements, and <i>U</i> and <i>V</i> represent matrices formed of a singular vector corresponding to the singular value. Further, the following expressions hold: <maths id="math0111"><math display="block"><msub><mi>R</mi><mn>1</mn></msub><mo>=</mo><mi>U</mi><mo>*</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mo>-</mo><mn>1</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>*</mo><msup><mi>V</mi><mi>T</mi></msup><mo>,</mo><msub><mi>R</mi><mn>2</mn></msub><mo>=</mo><mi>U</mi><mo>*</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>0</mn></mtd><mtd><mo>-</mo><mn>1</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>1</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>*</mo><msup><mi>V</mi><mi>T</mi></msup></math><img file="EP1978731A2_D0113.tif" /></maths><maths id="math0112"><math display="block"><mover><mi>t</mi><mo>→</mo></mover><mo>=</mo><msub><mover><mi>v</mi><mo>→</mo></mover><mn>3</mn></msub></math><img file="EP1978731A2_D0114.tif" /></maths> where <o ostyle="rightarrow">ν</o><sub>3</sub> is the singular vector of the third column of <i>V.</i>
0212The fundamental matrix <i>E</i> has indetermination for scale. Thus, a redundant scale component of <i>E</i> is ignored when |<o ostyle="rightarrow"><i>t</i></o>| = 1 and |<i>R</i>| = 1 are set. As a result, <o ostyle="rightarrow"><i>t</i></o> represents a direction vector indicating the direction of the translation. Since the indetermination of the sign occurs, four solutions are possible for the translation and the rotation. In this case, one solution is selected by adding <i>R</i> ≥ 0 and the depth positive constraint as follows: <maths id="math0113"><math display="block"><mi mathvariant="italic">nsign</mi><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover></mstyle><mi mathvariant="italic">sign</mi><mfenced><msup><msubsup><mi>x</mi><mn>2</mn><mi>i</mi></msubsup><mi>T</mi></msup><mfenced><msub><mfenced open="[" close="]"><mi>T</mi></mfenced><mi>x</mi></msub><mo></mo><mi>R</mi></mfenced><mo></mo><msubsup><mi>x</mi><mn>1</mn><mi>i</mi></msubsup></mfenced></math><img file="EP1978731A2_D0115.tif" /></maths> where sign( ) represents a function for calculating +1 if the numerical value is the parentheses is positive, -1 if the value is negative, or -1 if the value is zero.
0213The sign of the direction vector <o ostyle="rightarrow"><i>t</i></o> is not changed if nsign is positive or it is inversed if nsign is negative.
0214Next, the homography is calculated. If the only one solution set of the attitude variation represented by {<i>R</i>, <o ostyle="rightarrow"><i>t</i></o>} is determined from the fundamental matrix, the information is used to calculate the homography excluding an appearance variation component (image variation due to the shear and foreshortening caused by the translation). To exclude the appearance variation component, <o ostyle="rightarrow"><i>n</i></o> = <o ostyle="rightarrow"><i>e</i></o><sub>3</sub> = [0,0,1]<i><sup>T</sup></i> is defined.
0215Then, an insufficient element d is determined from the relationship between the projective homography, and camera work {<i>R</i>, <o ostyle="rightarrow"><i>t</i></o>} and scene information {<i>d</i>, <o ostyle="rightarrow"><i>n</i></o>}. Specifically, the product d of the reference plane distance in the depth direction and the translation magnitude is determined so as to minimize sum represented by the following expression: <maths id="math0114"><math display="block"><mi mathvariant="italic">sum</mi><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover></mstyle><msup><mfenced><msup><msubsup><mi>x</mi><mn>2</mn><mi>i</mi></msubsup><mi>T</mi></msup><mo></mo><mfenced><mi>R</mi><mo>+</mo><mfrac><mn>1</mn><mi>d</mi></mfrac><mo></mo><mover><mi>n</mi><mo>→</mo></mover><mo></mo><msup><mover><mi>t</mi><mo>→</mo></mover><mi>T</mi></msup></mfenced><mo></mo><msubsup><mi>x</mi><mn>1</mn><mi>i</mi></msubsup></mfenced><mn>2</mn></msup><mn>.</mn></math><img file="EP1978731A2_D0116.tif" /></maths>
0216As a result, the projective homography <i>H</i> in which the appearance variation component was excluded and which depends on the distribution of the corresponding points is determined from the following relational expression, the fundamental matrix, and the corresponding points used in the calculation thereof: <maths id="math0115"><math display="block"><mi>H</mi><mo>=</mo><mi>R</mi><mo>+</mo><mfrac><mn>1</mn><mi>d</mi></mfrac><mo></mo><msub><mover><mi>e</mi><mo>→</mo></mover><mn>3</mn></msub><mo></mo><msup><mover><mi>t</mi><mo>→</mo></mover><mi>T</mi></msup><mn>.</mn></math><img file="EP1978731A2_D0117.tif" /></maths>
0217The processing at a correction amount calculating step of S1004 is similar to the processing at the correction amount calculating step of S504 in <figref idref="f0005">Fig. 5</figref>.
0218In Embodiment 3, the processing step with the abovementioned changes made to Embodiment 1 is performed to apply image-stabilization processing to each frame of the video signal.
0219The shake-correction parameter calculating method in Embodiment 3 has the property of instability if the corresponding points in a scene input at the shake-correction parameter calculating step are distributed on a single plane. In contrast, the shake-correction parameter calculating method in Embodiment 1 has the property of instability if the corresponding points in a scene are uniformly distributed in a wide depth range. In other words, Embodiments 3 and 1 have the complementary properties.
0220The properties may be utilized. For example, <figref idref="f0011">Fig. 11</figref> shows a procedure of shake-correction parameter calculation which involves investigation of planarity of normalized corresponding point distribution (that is, planarity of spatial distribution of motion vector points) before calculation of an image variation amount, and involves switching of processing in response to the result of the investigation.
0221The processing at a normalization step of S1101 is similar to the processing at the normalization step of S501 in <figref idref="f0005">Fig. 5</figref> of Embodiment 1.
0222At a planarity calculating step of S1102, the projective homography is calculated from the corresponding points to investigate fitting of the reference plane in space determined by the projective homography and the distribution of the corresponding points in space with planarity. The corresponding points between the frames are defined as <o ostyle="rightarrow"><i>x</i></o><sub>1</sub>, <o ostyle="rightarrow"><i>x</i></o><sub>2</sub> and planarity P is calculated with the following expression: <maths id="math0116"><math display="block"><mi>P</mi><mo>=</mo><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover></mstyle><msup><mfenced><msup><msubsup><mover><mi>x</mi><mo>→</mo></mover><mn>2</mn><mi>i</mi></msubsup><mi>T</mi></msup><mo></mo><mi>H</mi><mo></mo><msubsup><mover><mi>x</mi><mo>→</mo></mover><mn>1</mn><mi>i</mi></msubsup></mfenced><mn>2</mn></msup><mo>,</mo><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>…</mo><mo>,</mo><mi>n</mi></math><img file="EP1978731A2_D0118.tif" /></maths> where <i>H</i> represents the projective homography determined from the corresponding points and n represents the number of the corresponding points.
0223Then, a threshold value th is set. If P is equal to or smaller than th, it is determined that the corresponding points are distribution in a planar manner. If P is larger than th, it is determined that the corresponding pointes are distributed in a nonplanar manner.
0224When the corresponding point distribution is close to planar distribution in space, the control proceeds to the projective homography calculating step of S1104 in Embodiment 1. Then, the shake correction parameters are calculated by using the projective homography as the index of the image variation amount. In contrast, if the corresponding point distribution is not close to planar distribution, the control proceeds to a fundamental matrix calculating step of S1114 in the shake-correction parameter calculating method of this embodiment. Then, the shake correction parameters are calculated by using the fundamental matrix as the index of the image variation amount.
0225Thereafter, in S1105, S1115, and S1116, the shake correction parameters are calculated with the shake-correction parameter calculating step of Embodiment 1 or this embodiment. If the absolute amount of the motion is small, the method based on the fundamental matrix in Embodiment 3 readily causes error. For this reason, the index of the absolute amount of motion may be used additionally in the determination of the planarity. As the index of the absolute amount of motion, the mean square of displacement of the corresponding points between the frames may be used, for example.
0226According to Embodiment 3, it is possible to preserve a motion due to an intended camera work included in video picked up by a video camera having an image-pickup optical system of a very short focal length and to effectively suppress an image shake due to an unintended camera shake. Especially for video in a deep scene and a widely moving scene, stable image-stabilization processing result can be achieved.
0227The abovementioned time-series filtering in Embodiment 3 is performed by using the digital filtering. However, another filtering method may be used to separate the image variation amounts produced due to an intended camera work and an unintended camera shake.
[Embodiment4]
0228An image stabilizing apparatus which is Embodiment 4 and a video camera including the apparatus will hereinafter be described. The configurations of the image stabilizing apparatus and the video camera are identical to those in Embodiment 1. Basic portions of the image-stabilization processing procedure are identical to the processing procedure in Embodiment 1 described with <figref idref="f0003">Figs. 3</figref> and <figref idref="f0005">5</figref>. <figref idref="f0012">Fig. 12</figref> shows a shake-correction amount calculating step which represents a difference between the processing procedure of Embodiment 4 and that in Embodiment 1.
0229In Embodiment 4, filtering in calculating a correction amount is performed through low-order model fitting. First, the image variation amount between the input frames is handled as an observation value including noise of an unintended camera work mixed into the result of an intended camera work. The intended camera work formed of motion at low frequency is modeled with a low-order model. Then, the observation value is filtered through the fitting. The low-order model is sequentially updated with a Kalman filter.
0230In Embodiment 4, the image variation amount between the frames is used as input. The projective homography in which the appearance variation component was excluded is used as the image variation amount.
0231At a homography transforming step of S1201, the input projective homography is transformed and resolved into component representation similar to the camera works. The filtering using the low-order model fitting allows input of a multi-dimensional amount series. However, if non-linearity is present in the correspondence between the camera works and the input multi-dimensional amount, the filtering with the low-order model fitting and the model update with the Kalman filter are not performed successfully. To prevent this problem, the step involves processing of resolution into linearly changing terms and non-linearly changing terms for the camera works.
0232The input projective homography between the frames is represented by the following expression if the scale component is normalized with h9 = 1: <maths id="math0117"><math display="block"><mi>H</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>h</mi><mn>1</mn></msub></mtd><mtd><msub><mi>h</mi><mn>2</mn></msub></mtd><mtd><msub><mi>h</mi><mn>3</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>4</mn></msub></mtd><mtd><msub><mi>h</mi><mn>5</mn></msub></mtd><mtd><msub><mi>h</mi><mn>6</mn></msub></mtd></mtr><mtr><mtd><msub><mi>h</mi><mn>7</mn></msub></mtd><mtd><msub><mi>h</mi><mn>8</mn></msub></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0119.tif" /></maths> When each term of the projective homography is compared with the camera work (triaxial rotation and triaxial translation), camera works (advancing, in-plane rotation, panning, and tilting) influencing the 2 x 2 terms on the upper left are mixed. As a result, if these camera works occur simultaneously, non-linear influence causes the possibility that the filtering and the model update are not performed successfully. Thus, the transformation of the input projective homography into image variation components (horizontal translation, vertical translation, scaling, in-plane rotation, shear, horizontal foreshortening, and vertical foreshortening), which are similar to the camera works, is performed by the following expression: <maths id="math0118"><math display="block"><mi>H</mi><mo>=</mo><msub><mi>H</mi><mi>S</mi></msub><mo></mo><msub><mi>H</mi><mi>A</mi></msub><mo></mo><msub><mi>H</mi><mi>P</mi></msub><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi mathvariant="italic">sR</mi></mtd><mtd><mover><mi>t</mi><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>K</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>I</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mi>v</mi><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>A</mi></mtd><mtd><mover><mi>t</mi><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mi>v</mi><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0120.tif" /></maths> where <i>A</i> = <i>RK</i> + <i><o ostyle="rightarrow">tν</o><sup>T</sup></i>
0233Then, <i>RK</i> = <i>A</i> - <i><o ostyle="rightarrow">tν</o><sup>T</sup></i> is calculated, and <i>R</i> and <i>K</i> are resolved by qr resolution using the property of <i>K</i> that is an <u style="single">upper triangular matrix.</u>
0234This achieves the resolution of the input projective homography into eight parameters including horizontal translation tx, vertical translation ty, scaling s, rotation θ, anisotropic magnification α of shear, direction angle φ of shear, horizontal foreshortening vx, and vertical foreshortening νy. <o ostyle="rightarrow"><i>t</i></o>, <o ostyle="rightarrow"><i>ν</i></o>, <i>R</i> and <i>K</i> are expressed as follows: <maths id="math0119"><math display="block"><mover><mi>t</mi><mo>→</mo></mover><mo>=</mo><msup><mfenced open="[" close="]"><msub><mi>t</mi><mi>x</mi></msub><mo></mo><msub><mi>t</mi><mi>y</mi></msub></mfenced><mi>T</mi></msup></math><img file="EP1978731A2_D0121.tif" /></maths><maths id="math0120"><math display="block"><mover><mi>v</mi><mo>→</mo></mover><mo>=</mo><msup><mfenced open="[" close="]"><msub><mi>v</mi><mi>x</mi></msub><mo></mo><msub><mi>v</mi><mi>y</mi></msub></mfenced><mi>T</mi></msup></math><img file="EP1978731A2_D0122.tif" /></maths><maths id="math0121"><math display="block"><mi>R</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>cos</mi><mi>θ</mi></mtd><mtd><mo>-</mo><mi>sin</mi><mi>θ</mi></mtd></mtr><mtr><mtd><mi>sin</mi><mi>θ</mi></mtd><mtd><mi>cos</mi><mi>θ</mi></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0123.tif" /></maths><maths id="math0122"><math display="block"><mi>K</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>α</mi></mtd><mtd><mi>tan</mi><mi>ϕ</mi></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0124.tif" /></maths>
0235In this manner, the input projective homography is transformed into a set of one-dimensional time-series amounts representing the image variation amount between the frames as follows: <maths id="math0123"><math display="block"><mtable><mtr><mtd><msup><mfenced open="[" close="]"><msub><mi>t</mi><mi>x</mi></msub><mo></mo><msub><mi>t</mi><mi>y</mi></msub><mo></mo><mi>s</mi><mo></mo><mi>θ</mi><mo></mo><mi>α</mi><mo></mo><mi>ϕ</mi><mo></mo><msub><mi>v</mi><mi>x</mi></msub><mo></mo><msub><mi>v</mi><mi>y</mi></msub></mfenced><mi>i</mi></msup><mo>,</mo></mtd><mtd><mi>i</mi><mo>=</mo><mn>1</mn><mo>,</mo><mo>⋯</mo><mo>,</mo><mi>k</mi></mtd></mtr></mtable><mn>.</mn></math><img file="EP1978731A2_D0125.tif" /></maths>
0236A linearization step of S1202 involves removing the remaining non-linearity and the influence of the scaling component upon the horizontal and vertical foreshortening components after the transformation into the image variation components similar to the camera works. Because of the order of the abovementioned resolution calculation, the foreshortening component is influenced by the scaling component if the parameters are seen from the viewpoint of appearance on the image. When a camera work relating to scaling and a camera work relating to foreshortening occur simultaneously, a non-linear parameter change may be caused. Thus, calculations of νx' = svx and νy' = svy are performed to remove the influence of the scaling from the foreshortening component.
0237S1203 is a step for performing the filtering with the low-order model fitting. In Embodiment 4, a constant-velocity model is used as a variation model of the image variation amount. Thus, a constant-velocity variation of the component of the image variation amount is included in the intended motion. As a result, a larger variation than the constant-velocity variation can be discriminated as an unintended image variation amount.
0238First, the Kalman filter is used to efficiently perform the sequential state model update and filtering. To use the Kalman filter, the following state space model representing time series is built as follows: <maths id="math0124"><math display="block"><msub><mi>x</mi><mrow><mi>n</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>=</mo><msub><mi>F</mi><mi>n</mi></msub><mo></mo><msub><mi>x</mi><mi>n</mi></msub><mo>+</mo><msub><mi>G</mi><mi>n</mi></msub><mo></mo><msub><mi>v</mi><mi>n</mi></msub><mspace width="1em" /><mfenced><mi>system model</mi></mfenced></math><img file="EP1978731A2_D0126.tif" /></maths><maths id="math0125"><math display="block"><msub><mi>y</mi><mi>n</mi></msub><mo>=</mo><msub><mi>H</mi><mi>n</mi></msub><mo></mo><msub><mi>x</mi><mi>n</mi></msub><mo>+</mo><msub><mi>w</mi><mi>n</mi></msub><mspace width="1em" /><mfenced><mi>observation model</mi></mfenced></math><img file="EP1978731A2_D0127.tif" /></maths> where <i>x<sub>n</sub></i> represents a vector of k dimension which cannot be directly observed and is called a 'state', vn represents system noise that is m-dimensional normalized white noise according to an average vector of 0 and a variance-covariance matrix <i>Q<sub>n</sub></i>, <i>W<sub>n</sub></i> represents observation noise that is one-dimensional normalized white noise according to the average vector of 0 and a variance-covariance matrix <i>R<sub>n</sub>,</i> and <i>F<sub>n</sub></i>,<i>G<sub>n</sub></i>,<i>H<sub>n</sub></i> respectively represent matrixes of k x k, k x m, and l x k.
0239A system model of the constant-velocity model is defined with a state variable x and a velocity variable Δx (<i>x:t<sub>x</sub>,t<sub>y</sub>,s,θ,α,φ,ν<sub>x</sub>,ν<sub>y</sub></i>). The velocity variable is an inside parameter which is not exposed.
0240A velocity variation element is handled as white Gaussian noise N(0,σ) which represents a white Gaussian noise with average zero and variance σ.
0241First, a system model for one component is represented as follows: <maths id="math0126"><math display="block"><msup><mfenced open="[" close="]"><mtable><mtr><mtd><mi>x</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>x</mi></mtd></mtr></mtable></mfenced><mrow><mi>n</mi><mo>+</mo><mn>1</mn></mrow></msup><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mn>1</mn></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><msup><mfenced open="[" close="]"><mtable><mtr><mtd><mi>x</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>x</mi></mtd></mtr></mtable></mfenced><mi>n</mi></msup><mo>+</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>0</mn></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><mi>σ</mi></mfenced></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0128.tif" /></maths>
0242Thus, a state space system model for all of the input image variation amount components is given as follows: <maths id="math0127"><math display="block"><msup><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>t</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mi>t</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mi>s</mi></mtd></mtr><mtr><mtd><mi>θ</mi></mtd></mtr><mtr><mtd><mi>α</mi></mtd></mtr><mtr><mtd><mi>ϕ</mi></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>t</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>t</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>s</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>θ</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>α</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>ϕ</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>v</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>v</mi><mi>y</mi></msub></mtd></mtr></mtable></mfenced><mrow><mi>t</mi><mo>+</mo><mn>1</mn></mrow></msup><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mspace width="1em" /></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><msup><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>t</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mi>t</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mi>s</mi></mtd></mtr><mtr><mtd><mi>θ</mi></mtd></mtr><mtr><mtd><mi>α</mi></mtd></mtr><mtr><mtd><mi>ϕ</mi></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>t</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>t</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>s</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>θ</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>α</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><mi>ϕ</mi></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>v</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><mi mathvariant="normal">Δ</mi><mo></mo><msub><mi>v</mi><mi>y</mi></msub></mtd></mtr></mtable></mfenced><mi>t</mi></msup><mo>+</mo><mfenced><mtable><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mspace width="1em" /></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><msub><mi>t</mi><mi>x</mi></msub></msub></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><msub><mi>t</mi><mi>y</mi></msub></msub></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><mi>s</mi></msub></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><mi>θ</mi></msub></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><mi>α</mi></msub></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><mi>ϕ</mi></msub></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><msub><mi>v</mi><mi>x</mi></msub></msub></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msub><mi>σ</mi><msub><mi>v</mi><mi>y</mi></msub></msub></mfenced></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0129.tif" /></maths>
0243An observation model for each parameter is represented as follows: <maths id="math0128"><math display="block"><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mover><mi>t</mi><mo>˜</mo></mover><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mover><mi>t</mi><mo>˜</mo></mover><mi>y</mi></msub></mtd></mtr><mtr><mtd><mover><mtable><mtr><mtd><mi>s</mi></mtd></mtr></mtable><mo>˜</mo></mover></mtd></mtr><mtr><mtd><mover><mtable><mtr><mtd><mi>θ</mi></mtd></mtr></mtable><mo>˜</mo></mover></mtd></mtr><mtr><mtd><mover><mtable><mtr><mtd><mi>α</mi></mtd></mtr></mtable><mo>˜</mo></mover></mtd></mtr><mtr><mtd><mover><mtable><mtr><mtd><mi>ϕ</mi></mtd></mtr></mtable><mo>˜</mo></mover></mtd></mtr><mtr><mtd><msub><mover><mi>v</mi><mo>˜</mo></mover><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mover><mi>v</mi><mo>˜</mo></mover><mi>y</mi></msub></mtd></mtr></mtable></mfenced><mo>=</mo><mfenced><mtable><mtr><mtd><msub><mi>t</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mi>t</mi><mi>y</mi></msub></mtd></mtr><mtr><mtd><mi>s</mi></mtd></mtr><mtr><mtd><mi>θ</mi></mtd></mtr><mtr><mtd><mi>α</mi></mtd></mtr><mtr><mtd><mi>ϕ</mi></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>x</mi></msub></mtd></mtr><mtr><mtd><msub><mi>v</mi><mi>y</mi></msub></mtd></mtr></mtable></mfenced><mo>+</mo><mfenced><mtable><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><msub><mi>t</mi><mi>x</mi></msub><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><msub><mi>t</mi><mi>y</mi></msub><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><mi>s</mi><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><mi>θ</mi><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><mi>α</mi><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><mi>ϕ</mi><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><msub><mi>v</mi><mi>x</mi></msub><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr><mtr><mtd><mi>N</mi><mfenced><mn>0</mn><mo></mo><msubsup><mi>σ</mi><msub><mi>v</mi><mi>y</mi></msub><mi mathvariant="italic">obs</mi></msubsup></mfenced></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0130.tif" /></maths> where <img file="EP1978731A2_D0131.tif" /> represents an observation value, and N(0,oxobs) represents white Gaussian observation noise for the x component. The white Gaussian observation noise component represents an unintended motion. The variance of the observation noise and the variance of the system noise are adjusted to allow adjustment of smoothness of camera motions.
0244The abovementioned system model and observation model are represented in the matrix form of the state space model as follows: <maths id="math0129"><math display="block"><mi>F</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>I</mi><mrow><mn>8</mn><mo>×</mo><mn>0</mn></mrow></msub></mtd><mtd><msub><mi>I</mi><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub></mtd></mtr><mtr><mtd><msub><mn>0</mn><mrow><mn>8</mn><mo>×</mo><mn>0</mn></mrow></msub></mtd><mtd><msub><mi>I</mi><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0132.tif" /></maths><maths id="math0130"><math display="block"><mi>G</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mn>0</mn><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub></mtd></mtr><mtr><mtd><msub><mi>I</mi><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0133.tif" /></maths><maths id="math0131"><math display="block"><mi>H</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><msub><mi>I</mi><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub></mtd><mtd><msub><mn>0</mn><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0134.tif" /></maths> In addition, the following can be assumed: <maths id="math0132"><math display="block"><mi>Q</mi><mo>=</mo><msup><mi>σ</mi><mi mathvariant="italic">sys</mi></msup><mo></mo><msub><mi>I</mi><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub></math><img file="EP1978731A2_D0135.tif" /></maths><maths id="math0133"><math display="block"><mi>R</mi><mo>=</mo><msup><mi>σ</mi><mi mathvariant="italic">obs</mi></msup><mo></mo><msub><mi>I</mi><mrow><mn>8</mn><mo>×</mo><mn>8</mn></mrow></msub><mn>.</mn></math><img file="EP1978731A2_D0136.tif" /></maths>
0245Thus, the sequential update of the model (<i>x</i>(<i>t</i> +1|<i>t</i>) ← <i>x</i>(<i>t</i>|<i>t</i>) ) is performed with <i>x = Fx</i> and <i>P = FPF<sup>T</sup></i> + <i>GQG<sup>T</sup></i>
0246The filtering results (<i>x</i>(<i>t</i>|<i>t</i>) ← <i>x</i>(<i>t</i>|<i>t</i> - 1), <i>P</i>(<i>t</i>|<i>t</i>) ← <i>P</i>(<i>t</i>|<i>t</i> -1 )) are obtained from the following expressions: <maths id="math0134"><math display="block"><mi>K</mi><mo>=</mo><mi>P</mi><mo></mo><msup><mi>H</mi><mi>T</mi></msup><mo>/</mo><mfenced><mi mathvariant="italic">HP</mi><mo></mo><msup><mi>H</mi><mi>T</mi></msup><mo>+</mo><mi>R</mi></mfenced></math><img file="EP1978731A2_D0137.tif" /></maths><maths id="math0135"><math display="block"><msub><mi>x</mi><mi mathvariant="italic">fil</mi></msub><mo>=</mo><mi>x</mi><mo>+</mo><mi>K</mi><mo></mo><mfenced><msub><mi>y</mi><mi mathvariant="italic">obs</mi></msub><mo>-</mo><mi mathvariant="italic">Hx</mi></mfenced></math><img file="EP1978731A2_D0138.tif" /></maths><maths id="math0136"><math display="block"><msub><mi>y</mi><mi mathvariant="italic">fil</mi></msub><mo>=</mo><mi>H</mi><mo></mo><msub><mi>x</mi><mi mathvariant="italic">fil</mi></msub><mn>.</mn></math><img file="EP1978731A2_D0139.tif" /></maths>
0247That is, the filtering value represented by <i>y<sub>fil</sub> =</i> [<i>t<sub>xy</sub>,t<sub>y</sub>,s,θ,α,φ,v'<sub>x</sub>,v'<sub>y</sub></i>]<i><sup>T</sup></i> is provided as an intended motion component from the current frame. The difference between the current frame and the predicted value can be provided as a value to be a correction amount.
0248At a non-linearity restoration step of S1204, the image variation amount component deformed to provide the linear change for the camera work is transformed to provide the original non-linear component. In Embodiment 4, the foreshortening term is restored with the following: <maths id="math0137"><math display="block"><mi mathvariant="normal">νx</mi><mo mathvariant="normal">=</mo><mi mathvariant="normal">νxʹ</mi><mo mathvariant="normal">/</mo><mi mathvariant="normal">s</mi></math><img file="EP1978731A2_D0140.tif" /></maths><maths id="math0138"><math display="block"><mi mathvariant="normal">νy</mi><mo mathvariant="normal">=</mo><mi mathvariant="normal">νyʹ</mi><mo mathvariant="normal">/</mo><mi mathvariant="normal">s</mi><mn>.</mn></math><img file="EP1978731A2_D0141.tif" /></maths>
0249A homography restoration step of S1205 is a step for restoring the image variation amount, which has been transformed into the image variation amount components similar to the camera works, to the representation of the homography. The one-dimensional amount set of the filtering result represented by [<i>t<sub>x</sub>,t<sub>x</sub>,s,θ,α,φ,ν<sub>x</sub>,ν<sub>y</sub></i>]<i><sub>filter</sub></i> is transformed with the following expression: <maths id="math0139"><math display="block"><msub><mi>H</mi><mi mathvariant="italic">filter</mi></msub><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi mathvariant="italic">sR</mi></mtd><mtd><mover><mi>t</mi><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>K</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mn>0</mn><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mo></mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>I</mi></mtd><mtd><mover><mn>0</mn><mo>→</mo></mover></mtd></mtr><mtr><mtd><msup><mover><mi>v</mi><mo>→</mo></mover><mi>T</mi></msup></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced></math><img file="EP1978731A2_D0142.tif" /></maths> where <maths id="math0140"><math display="inline"><mover><mi>t</mi><mo>→</mo></mover><mo>=</mo><msup><mfenced open="[" close="]"><msub><mi>t</mi><mi>x</mi></msub><mo></mo><msub><mi>t</mi><mi>y</mi></msub></mfenced><mi>T</mi></msup><mo>,</mo></math><img file="EP1978731A2_D0143.tif" /></maths><maths id="math0141"><math display="inline"><mover><mi>v</mi><mo>→</mo></mover><mo>=</mo><msup><mfenced open="[" close="]"><msub><mi>v</mi><mi>x</mi></msub><mo></mo><msub><mi>v</mi><mi>y</mi></msub></mfenced><mi>T</mi></msup><mo>,</mo></math><img file="EP1978731A2_D0144.tif" /></maths><maths id="math0142"><math display="inline"><mi>R</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>cos</mi><mi>θ</mi></mtd><mtd><mo>-</mo><mi>sin</mi><mi>θ</mi></mtd></mtr><mtr><mtd><mi>sin</mi><mi>θ</mi></mtd><mtd><mi>cos</mi><mi>θ</mi></mtd></mtr></mtable></mfenced><mo>,</mo></math><img file="EP1978731A2_D0145.tif" /></maths> and <maths id="math0143"><math display="block"><mi>K</mi><mo>=</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mi>α</mi></mtd><mtd><mi>tan</mi><mi>ϕ</mi></mtd></mtr><mtr><mtd><mn>0</mn></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0146.tif" /></maths>
0250A correction amount calculating step of S1206 is a step for calculating a projective matrix serving as an shake correction amount. A difference between an image variation amount <i>H<sub>filter</sub></i> between the frames restored from the filtering result and the image variation amount <i>H</i> between the target frame and the past frame is calculated as the shake correction amount.
0251When delay is not considered, an image motion <i>H<sup>i</sup></i> between the target frame and the past frame input to calculate <i>H<sup>i</sup><sub>filter</sub></i> is used as the target frame. If the shake correction amount is represented by <i>H<sub>stb</sub></i>, it is calculated with: <maths id="math0144"><math display="block"><msubsup><mi>H</mi><mi mathvariant="italic">stb</mi><mi>i</mi></msubsup><mo>=</mo><msubsup><mi>H</mi><mi mathvariant="italic">filter</mi><mi>i</mi></msubsup><mo></mo><msup><mfenced><msup><mi>H</mi><mi>i</mi></msup></mfenced><mrow><mo>-</mo><mn>1</mn></mrow></msup></math><img file="EP1978731A2_D0147.tif" /></maths> where i represents the frame number between the current frame and the past frame.
0252If delay is present, the shake correction amount between the frames is calculated with: <maths id="math0145"><math display="block"><msubsup><mi>H</mi><mi mathvariant="italic">stb</mi><mi mathvariant="italic">iʹ</mi></msubsup><mo>=</mo><msubsup><mi>H</mi><mi mathvariant="italic">filter</mi><mrow><mi>i</mi><mo>-</mo><mi mathvariant="italic">delay</mi></mrow></msubsup><mo></mo><msup><mfenced><msup><mi>H</mi><mi mathvariant="italic">iʹ</mi></msup></mfenced><mrow><mo>-</mo><mn>1</mn></mrow></msup></math><img file="EP1978731A2_D0148.tif" /></maths> where i represents the frame number between the current frame and the past frame and <i>delay</i> represents the delay amount when the relationship i' = i- <i>delay</i> is satisfied. When no delay is present, <i>delay</i> = 0 holds in the above expression.
0253A model update step of S1207 is a step for updating the state space model. Typically, the filtering is performed on the accumulated variation amount series from an arbitrary reference frame to extract a low-frequency variation amount component or a high-frequency variation amount component. The reference frame is typically an initial frame.
0254However, if image stabilization is performed during movement of a user with the initial frame used as the reference, the accumulated variation amount component of scaling which non-linearly changes for a camera work readily becomes an extremely small (large) value if an advancing (backing) movement is included. As a result, a minute variation component cannot be filtered.
0255To prevent this problem, the start frame of the current frame and the past frame is used as the reference frame of the accumulated variation amount. In other words, update is performed such that the reference frame of the state variable of the state space model is shifted.
0256Specifically, the reference frame of the state variable is shifted by one frame with the image variation amount between the current frame and the past frame used for updating the state space model. First, the state variable represented by {<i>t<sub>x</sub>,t<sub>y</sub>,s,θ,α,φ,v'<sub>x</sub>,v'<sub>y</sub></i>} is restored to the projective homography <i>H<sub>state</sub>.</i> Then, the processing of canceling the change of the image variation amount between the current frame and the past frame is performed with the following expression: <maths id="math0146"><math display="block"><msub><mi>H</mi><mi mathvariant="italic">state</mi></msub><mo>=</mo><msub><mi>H</mi><mi mathvariant="italic">state</mi></msub><mo></mo><msup><mi>H</mi><mrow><mo>-</mo><mn>1</mn></mrow></msup><mn>.</mn></math><img file="EP1978731A2_D0149.tif" /></maths>
0257The projective homography <i>H<sub>state</sub></i> is again resolved to the state variable term represented by {<i>t<sub>x</sub>,t<sub>y</sub>,s,θ,α,φ,ν'<sub>x</sub>,ν'<sub>y</sub></i>}. As a result, even when the image variation amount between the current frame and the past frame is input in the filtering, the result is given as if the accumulated variation amount was input to perform the filtering.
0258The filtering and the correction amount calculation with the low-order model fitting as described above are repeated to calculate the shake correction amount for the frame.
0259According to Embodiment 4, it is possible to preserve a motion due to an intended camera work included in video picked up by a video camera having an image-pickup optical system of a very short focal length and to effectively suppress an image shake due to an unintended camera shake. With the processing of Embodiment 4, the image stabilizing degree can be adjusted seamlessly from a full image stabilization state to a non-image stabilization state only by adjusting the Kalman filter coefficient.
0260While Embodiment 4 has been described of the case where the image variation amount calculating step similar to that in Embodiment 1 is performed, the image variation amount calculating step described in Embodiments 2 and 3 may be used.
[Embodiment 5]
0261<figref idref="f0013">Fig. 13</figref> shows the configuration of an image stabilizing apparatus which is Embodiment 5 of the present invention. The image stabilizing apparatus is not mounted on a video camera but is formed to function alone and is realized by a personal computer or the like. For example, a computer program for executing each processing described below is installed on the personal computer, so that the personal computer can be used as the image stabilizing apparatus.
0262In <figref idref="f0013">Fig. 13</figref>, reference numeral 1301 shows a read-out part, 1302 a decoding part, 1303 a preprocessing part, 1304 a motion vector detecting part, and 1305 a work memory. Reference numeral 1306 shows a shake-correction parameter calculating part, 1307 a geometric transformation processing part, 1308 an encoding/decoding part, and 1309 a work memory. Reference numeral 1310 a system controlling part, 133 a non-volatile memory part, 1312 a recording part, 1313 a displaying part, 1314 an operation signal inputting part, and 1315 an external I/F.
0263The preprocessing part 1303, the motion vector detecting part 1304, the shake-correction parameter calculating part 1306, the geometric transformation processing part 1307, and the encoding/decoding part 1308 constitute a video signal processing part.
0264The read-out part 1301 is formed of a mechanism for reading a video signal and image-pickup information including inside parameters of a camera used in an image-pickup operation from a recording medium such as a semiconductor memory, a magnetic tape, and an optic disk. The video signal is sent to the decoding part 1302, while the image-pickup information is sent to the system controlling part 1310.
0265The decoding part 1302 decodes the video signal read by the read-out part 1301 if it is an encoded signal.
0266The preprocessing part 1303 performs video processing for detecting a motion vector on the image-pickup signal output from the decoding part 1302. The video processing performed by the preprocessing part 130 includes, for example, gain adjustment, gamma adjustment, luminance /color difference separation, sharpening, white balance adjustment, black level adjustment, colorimetric system transformation, and coding.
0267The motion vector detecting part 1204 receives, as input, video frames such as successive luminance frames, luminance and color difference frames, or RGB frames transformed from the luminance and color difference frames provided by the preprocessing part 1303. It may receive, as input, differential processing frames processed for motion vector detection or binary code frames.
0268The motion vector detecting part 1304 detects the motion vector between successive frames input thereto. Specifically, it calculates motion vectors between a current frame input from the preprocessing part 1303 and a past frame input previously and accumulated in the work memory 1305. The past frame is a frame subsequent to the current frame or a much order frame.
0269The work memory 1305 is formed of a FIFO memory, for example.
0270The shake-correction parameter calculating part 1306 receives, as input, the motion vector output from the motion vector detecting part 1304 and camera calibration information such as in-camera parameters and a distortion coefficient provided by the system controlling part 1310 to calculate a shake correction amount.
0271The in-camera parameters include a focal length, a pixel size, and an offset, and a shear amount similarly to Embodiment 1. The distortion coefficient represents a distortion amount due to aberration of a lens optical system similarly to Embodiment 1.
0272The shake-correction parameter calculating part 1306 outputs shake correction parameters including the calculated shake correction amount, the in-camera parameters, and the distortion coefficient.
0273The geometric transformation processing part 1307 receives, as input, the shake correction parameters calculated by the shake-correction parameter calculating part 1306 and the corresponding video frames to perform geometric transformation processing of the video frames. As described in Embodiment 1, the shake correction parameters may be subjected to filtering processing or the like before the processing to this part 1307, so that they may be delayed relative to the corresponding video frames. In this case, the video frames are once passed through the work memory 1309 to match the video frames with the shake correction parameters. The work memory 1309 is a FIFO memory similar to the work memory 1305.
0274The encoding/decoding part 1308 encodes the video frame signal successively output from the geometric transformation processing part 1307 in a video format such as NTSC and MPEG4. To reproduce a recorded and encoded video signal, the encoding/decoding part 1308 decodes the video signal read out from the recording part 1312 and displays it on the displaying part 1313.
0275The system controlling part 1310 transmits a read-out instruction to the read-out part 130 to start reading of the video signal and the image-pickup information. The system controlling part 1310 takes the image-pickup information and control parameters in image-pickup operation. The control parameters include the in-camera parameters, a lookup table or a transforming expression showing the relationship between a zoom state and a focal length, and a lookup table or a transforming expression showing the relationship between a focal length and a distortion coefficient.
0276The focal length information or zoom state information for the video signal has a format of a time-series signal or for recording the state at the time of change and takes a form in which the focal length of all of the frames is restorable.
0277The system controlling part 1310 transmits the video information decoded by the decoding part 1302 to the preprocessing part 1303. It also sends the video signal encoded in the abovementioned video format and output from the encoding/decoding part 1308 to the recording part 1312 for recording.
0278The system controlling part 1310 also controls parameters for the processing blocks such as the motion vector detecting part 1304, the shake-correction parameter calculating part 1306, the geometric transformation processing part 1307, and the encoding/decoding part 1308. Initial values of the parameters are read out from the non-volatile memory part 1311. The various parameters are displayed on the displaying part 1313 and the values of the parameters can be changed with the operation signal inputting part 1314 or a GUI.
0279The system controlling part 1310 holds control parameters such as the number of the motion vectors, a search range of the motion vectors, and a template size for the motion vector detecting part 1304. The system controlling part 1310 provides the geometric transformation processing part 1307 with control parameters such as the shake correction parameters calculated by the shake-correction parameter calculating part 1306, and the inside parameters and distortion coefficient used in the calculation. The system controlling part 1310 provides the encoding/decoding part 1308 with control parameters such as the encoding method and the compression rate.
0280The system controlling part 1310 performs control of the work memories 1305 and 1309 to control the delay amount of output.
0281The system controlling part 1310 matches the video sequence with the image-pickup information. Specifically, it reads a zoom value representing a zoom state and uses the lookup table or the transforming expression showing the relationship between the zoom value and the focal length provided as the image-pickup information to acquire a focal length of the optical system in an arbitrary zoom state.
0282As described in Embodiment 1, the distortion coefficient varies depending on the focal length. Thus, the system controlling part 1310 also calculates the distortion coefficient corresponding to the focal length. It uses the lookup table or the transforming expression showing the relationship between the focal length and the distortion coefficient provided from the image-pickup information to calculate the distortion coefficient at an arbitrary focal length. In addition, the system controlling part 1310 takes and holds the in-camera parameters other than the focal length from the image-pickup information.
0283The inside parameters other than the focal length f include pixel sizes ku, kv in horizontal and vertical directions, a shear amount φ, and offset amounts u0 v0 in horizontal and vertical directions. The inside parameters are provided from camera design specifications or camera calibration. The system controlling part 1310 transmits the inside parameters and the distortion coefficient to the shake-correction parameter calculating part 1306.
0284The non-volatile memory part 1311 stores the initial values of the control parameters necessary to system control for the motion vector detecting part 1304, the shake-correction calculating part 1306, the encoding/decoding part 1308, the preprocessing part 1303 and the like. The control parameters are read out by the system controlling part 1310.
0285The recording part 1312 performs writing (recording) and reading (reproduction) of the video signal encoded by the encoding/decoding part 1308 to and from a recording medium on which the video signal can be recorded such as a semiconductor memory, a magnetic tape, and an optical disk.
0286The displaying part 1313 is formed of a display element such as an LCD, an LED, and an EL. The displaying part 1313 performs, for example, parameter setting display, alarm display, display of picked-up video data, and display of recorded video data read by the recording part 1312. In reproducing the recorded video data, the displaying part 1313 reads the encoded video signal from the recording part 1312 and transmits the read signal to the encoding/decoding part 1308 via the system controlling part 1310. The recorded video data after it is decoded is displayed on the displaying part 1313.
0287The operation signal inputting part 1314 includes setting buttons for performing selection of functions of the image stabilizing apparatus and various settings from the outside and a button for directing start and end of image-stabilization processing. The operation signal inputting part 1314 may be integrated with the displaying part 1313 by using a touch panel display method.
0288The external I/F 1315 receives an input signal from the outside instead of an operation signal input from the operation signal inputting part 1314 or outputs the encoded video signal to an external device. The external I/F 1315 is realized with an I/F protocol such as USB, IEEE1394, and wireless LAN. It can receive from the outside a video signal including information necessary for image stabilization such as the focal length or the zoom state in image-pickup operation, the in-camera parameters, and the distortion coefficient to allow image-stabilization processing of recorded video.
0289The procedure of image-stabilization processing in Embodiment 5 is identical to that in Embodiment 1. However, in Embodiment 5, the video information and the image-pickup information read out by the recording part 1312 are used to perform image-stabilization processing on the recorded video information. Then, the video stream after the image-stabilization processing is encoded in the video format such as NTSC and MPEG4 by the encoding/decoding part 1308. The encoded video stream is again recorded on the recording medium by the recording part 1312.
0290It is thus possible to preserve a motion due to an intended camera work included in video picked up previously by a video camera having an image-pickup optical system of a very short focal length and to effectively suppress an image shake due to an unintended camera shake.
0291While the image variation amount calculating step identical to that in Embodiment 1 is performed in Embodiment 5, the image variation amount calculating step described in Embodiments 2 and 3 may be used. While the time-series filtering identical to that in Embodiment 1 is performed, another filtering method may be used. For example, the filtering may be performed with the low-model fitting described in Embodiment 4.
0292The image-stabilization processing in Embodiment 5 can be performed not only on the video information recorded on the recording medium but also on video information recorded and saved across a network connected through the external I/F.
[Embodiment 6]
0293An image stabilizing apparatus which is Embodiment 6 and a video camera including the apparatus will hereinafter be described. Since the configurations of the image stabilizing apparatus and the video camera are identical to those in Embodiment 1, components identical to those in Embodiment 1 are designated with the same reference numerals as those in Embodiment 1. Basic portions of the image-stabilization processing procedure are identical to the processing procedure in Embodiment 1 described with <figref idref="f0003">Figs. 3</figref> and <figref idref="f0005">5</figref>.
0294<figref idref="f0014">Fig. 14</figref> shows a shake-correction amount calculating step which represents a difference between the processing procedure of Embodiment 6 and that in Embodiment 1.
0295Processing at a homography transforming step of S1401 is identical to the processing at the homography transforming step of Embodiment 1. The processing at a linearization step of S1402 is also identical to the processing at the linearization step of Embodiment 1. The processing at a time-series filtering step of S1404 is also identical to the filtering processing step of Embodiment 1.
0296At an empirical filtering of S1404, filtering with empirical knowledge of camera works in video is performed on a set of image variation amount components output from the time-series filtering step of S1403 and represented by [<i>t<sub>x</sub>,t<sub>x</sub>,s',θ,α,φ,ν'<sub>x</sub>,ν'<sub>y</sub></i>]<i><sup>out</sup>.</i> The empirical knowledge is input, for example, by presenting a menu on the displaying part 114 and operating the operation signal inputting part 115.
0297<figref idref="f0015">Fig. 15</figref> shows an example of the image-stabilization menu presented on the displaying part 114. As the image-stabilization menu, the state of a camera work in video is selected with a button and is confirmed with an OK button. The image-stabilization menu may be formed as shown in <figref idref="f0015">Fig. 16</figref> in which an image variation amount can be directly determined. The result is sent to the system controlling part 110. The initial value (image-pickup mode) is recorded on the non-volatile memory part 112. For example, the initial value is set to normal image-pickup.
0298For example, when empirical knowledge of wishing a stabilized image including only a restored forward camera work is selected as the image-pickup mode, intended image variation amounts other than scaling are assumed to be absent.
0299Specifically, the filtering represented as: <maths id="math0147"><math display="block"><msub><mi>t</mi><mi>x</mi></msub><mo>=</mo><mn>0</mn></math><img file="EP1978731A2_D0150.tif" /></maths><maths id="math0148"><math display="block"><msub><mi>t</mi><mi>y</mi></msub><mo>=</mo><mn>0</mn></math><img file="EP1978731A2_D0151.tif" /></maths><maths id="math0149"><math display="block"><mi>s</mi><mo>=</mo><mi>s</mi></math><img file="EP1978731A2_D0152.tif" /></maths><maths id="math0150"><math display="block"><mi>θ</mi><mo>=</mo><mn>0</mn></math><img file="EP1978731A2_D0153.tif" /></maths><maths id="math0151"><math display="block"><mi>α</mi><mo>=</mo><mn>0</mn></math><img file="EP1978731A2_D0154.tif" /></maths><maths id="math0152"><math display="block"><mi>ϕ</mi><mo>=</mo><mn>0</mn></math><img file="EP1978731A2_D0155.tif" /></maths><maths id="math0153"><math display="block"><msub><mi>v</mi><mi>x</mi></msub><mo>=</mo><mn>0</mn></math><img file="EP1978731A2_D0156.tif" /></maths><maths id="math0154"><math display="block"><msub><mi>v</mi><mi>y</mi></msub><mo>=</mo><mn>0</mn></math><img file="EP1978731A2_D0157.tif" /></maths>
0300is performed. In other words, the filtering is performed on the set of the image variation amount components after the time-series filtering such that only the image motion of scaling remains in the video.
0301Processing at a non-linear restoration step of S1405 is identical to the processing at the non-linear restoring step of Embodiment 1. Processing at a homography restoring step of S1406 is also identical to the processing at the homography restoring step of Embodiment 1. Processing at a correction amount calculating step of S1405 is also identical to the processing at the correction amount calculating step of Embodiment 1.
0302According to Embodiment 6, it is thus possible to preserve a motion due to an intended camera work included in video picked up by a video camera having an image-pickup optical system of a very short focal length and to effectively suppress an image shake due to an unintended camera shake. In addition, it is possible to provide video resulting from image-stabilization processing performed only on a particular image motion (shake).
0303While Embodiment 6 has been described in conjunction with the empirical filtering performed on the result of the time-series filtering, the empirical filtering may be performed before the time-series filtering.
0304The time-series filtering may not be performed but only the empirical filtering may be performed.
0305The empirical filtering coefficient may be a continuous value from zero to one, not a binary value of one or zero.
0306While the image variation amount calculating step identical to that in Embodiment 1 is performed in Embodiment 6, the image variation amount calculating step described in Embodiments 2 and 3 may be used. While the time-series filtering identical to that in Embodiment 1 is performed, another filtering method may be used. For example, the filtering may be performed with the low-model fitting described in Embodiment 4.
[Embodiment 7]
0307An image stabilizing apparatus which is Embodiment 7 and a video camera including the apparatus will hereinafter be described. Since the configurations of the image stabilizing apparatus and the video camera are identical to those in Embodiment 1. Basic portions of the image-stabilization processing procedure are identical to the processing procedure in Embodiment 2 described with <figref idref="f0009">Fig. 9</figref>. <figref idref="f0016">Fig. 17</figref> shows a shake-correction parameter calculating step which represents a difference between the processing procedure of Embodiment 7 and that in Embodiment 2.
0308Processing at a normalization step of S1701 is identical to the processing at the normalization step of S1901 described with <figref idref="f0009">Fig. 9</figref> in Embodiment 2.
0309At an attitude amount calculating step of S1702, motion vectors between frames transformed into a normalized image coordinate system are used as input to calculate an image variation amount between the frames. An attitude variation amount of a camera determined between the frames is calculated as the index of the image variation amount. In other words, the same processing as that at the attitude amount calculating step described in Embodiment 2 is performed.
0310At a filtering step of S1703, time-series data of the camera attitude variation between the frames represented by {<i>R<sub>x</sub>,R<sub>y</sub>,R<sub>z</sub>,t<sub>x</sub>,t<sub>y</sub>,t<sub>z</sub></i>}is formed and filtering is performed. The filtering is performed with digital filtering identical to that in Embodiment 1.
0311At a homography calculating step of S1704, a homography is calculated from the attitude variation amount which is the image variation amount determined between the frames and the filtering result.
0312Specifically, the attitude variation amount as the image variation amount determined between the frames represented by {R<sub>x</sub>,R<sub>y</sub>,R<sub>z</sub>,t<sub>x</sub>,t<sub>y</sub>,t<sub>z</sub>} is used to calculate the homography <i>H</i> that is represented by: <maths id="math0155"><math display="block"><mi>H</mi><mo>=</mo><mfenced><mi>R</mi><mo>+</mo><mfrac><mn>1</mn><mi>d</mi></mfrac><mo></mo><mover><mi>t</mi><mo>→</mo></mover><mo></mo><msup><mover><mi>n</mi><mo>→</mo></mover><mi>T</mi></msup></mfenced></math><img file="EP1978731A2_D0158.tif" /></maths> where <maths id="math0156"><math display="block"><mi>R</mi><mo>≅</mo><mfenced open="[" close="]"><mtable><mtr><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>2</mn></msub></mtd></mtr><mtr><mtd><msub><mi>ω</mi><mn>3</mn></msub></mtd><mtd><mn>1</mn></mtd><mtd><mo>-</mo><msub><mi>ω</mi><mn>1</mn></msub></mtd></mtr><mtr><mtd><mo>-</mo><msub><mi>ω</mi><mn>2</mn></msub></mtd><mtd><msub><mi>ω</mi><mn>1</mn></msub></mtd><mtd><mn>1</mn></mtd></mtr></mtable></mfenced><mn>.</mn></math><img file="EP1978731A2_D0159.tif" /></maths>
0313In this manner, the homography representing the image variation amount before and after the filtering is performed.
0314Processing at a correction amount calculating step of S1705 is identical to the processing at the correction amount calculating step of Embodiment 1. A shake correction amount is calculated by using the image variation amount restored from the filtering result <maths id="math0157"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc_filter</mi><mi>i</mi></msubsup></math><img file="EP1978731A2_D0160.tif" /></maths> and the accumulated variation amount <maths id="math0158"><math display="inline"><msubsup><mi>H</mi><mi mathvariant="italic">acc</mi><mi>i</mi></msubsup></math><img file="EP1978731A2_D0161.tif" /></maths> of the image variation amount to the target frame, the accumulated variation amount corresponding to the target frame as a result of the delay.
0315According to Embodiment 7, it is thus possible to preserve a motion due to an intended camera work included in video picked up by a video camera having an image-pickup optical system of a very short focal length and to effectively suppress an image shake due to an unintended camera shake. In addition, Embodiment 7 can realize the image stabilization with simpler processing than those of the methods of other embodiments.
0316While the time-series filtering identical to that of Embodiment 1 is performed in Embodiment 7, another filtering method may be used. For example, the filtering may be performed on the one-dimensional signal time-series with the low-model fitting described in Embodiment 4.
0317While the present invention has been described with reference to exemplary embodiments, it is to be understood that the invention is not limited to the disclosed exemplary embodiments. The scope of the following claims is to be accorded the broadest interpretation so as to encompass all modifications, equivalent structures and functions.
Contents4
180 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111 Sheet 112 Sheet 113 Sheet 114 Sheet 115 Sheet 116 Sheet 117 Sheet 118 Sheet 119 Sheet 120 Sheet 121 Sheet 122 Sheet 123 Sheet 124 Sheet 125 Sheet 126 Sheet 127 Sheet 128 Sheet 129 Sheet 130 Sheet 131 Sheet 132 Sheet 133 Sheet 134 Sheet 135 Sheet 136 Sheet 137 Sheet 138 Sheet 139 Sheet 140 Sheet 141 Sheet 142 Sheet 143 Sheet 144 Sheet 145 Sheet 146 Sheet 147 Sheet 148 Sheet 149 Sheet 150 Sheet 151 Sheet 152 Sheet 153 Sheet 154 Sheet 155 Sheet 156 Sheet 157 Sheet 158 Sheet 159 Sheet 160 Sheet 161 Sheet 162 Sheet 163 Sheet 164 Sheet 165 Sheet 166 Sheet 167 Sheet 168 Sheet 169 Sheet 170 Sheet 171 Sheet 172 Sheet 173 Sheet 174 Sheet 175 Sheet 176 Sheet 177 Sheet 178 Sheet 179 Sheet 180
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR20200137039A | Cited by | Republic of Korea | Search report |
| US9324160B2 | Cited by | United States of America | Applicant |
| US9001222B2 | Cited by | United States of America | Applicant |
| WO2013117959A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO2013056202A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| EP2640057A1 | Cited by | European Patent Office (EPO) | Search report |
| US9635261B2 | Cited by | United States of America | Applicant |
| EP2247097A3 | Cited by | European Patent Office (EPO) | Search report |
| US9888180B2 | Cited by | United States of America | Applicant |
| EP2640059A4 | Cited by | European Patent Office (EPO) | Search report |
| EP3800878A1 | Cited by | European Patent Office (EPO) | Search report |
| US9762799B2 | Cited by | United States of America | Applicant |
| US9305361B2 | Cited by | United States of America | Applicant |
| EP2640059A1 | Cited by | European Patent Office (EPO) | Search report |
| US9635256B2 | Cited by | United States of America | Applicant |
| EP2640057A4 | Cited by | European Patent Office (EPO) | Search report |
| US10412305B2 | Cited by | United States of America | Applicant |
| EP2247097A2 | Cited by | European Patent Office (EPO) | Search report |
| EP2974274A4 | Cited by | European Patent Office (EPO) | Search report |
| KR20150132846A | Cited by | Republic of Korea | Search report |
| US8798387B2 | Cited by | United States of America | Applicant |
| WO2013117959A1 | Cited by | World Intellectual Property Organization (WIPO) | Search report |
| GB2429867A | Cites | United Kingdom | Search report |
| Z. ZHU ET AL.: "Camera stabilization based on 2.5D motion estimation and inertial filtering", ICIV, 1998 | Non-patent | – | Applicant |
| A. LTVIN; J. KONRAD; W.C. KARL: "Probabilistic video stabilization using Kalman filtering and mosaicing", PROCEEDINGS OF SPIE., January 2003 (2003-01-01), pages 20 - 24 | Non-patent | – | Applicant |
| MICHAL IRANI ET AL.: "Recovery of Ego-Motion Using Image Stabilization", CVPR ('94, June 1994 (1994-06-01) | Non-patent | – | Applicant |
| C. HARRIS; M. STEPHENS: "A combined corner and edge detector", FOURTH ALVEY VISION CONFERENCE, 1988, pages 147 - 151 | Non-patent | – | Applicant |
12 members in 5 offices; this record represents the family
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 2007101162 | Japan | A | |
| 2007101162 | Japan | A | |
| 2007101162 | Japan | – | |
| 2007101162 | – | – | – |
| JP20070101162 | – | – | – |
Members12
| Document | Office | Kind | |
|---|---|---|---|
| EP1978731A2This record | European Patent Office (EPO) | A2 | |
| US2008246848A1 | United States of America | A1 | |
| JP2008259076A | Japan | A | |
| EP1978731A3 | European Patent Office (EPO) | A3 | |
| EP1978731B1 | European Patent Office (EPO) | B1 | |
| AT476827T | Austria | T | |
| ATE476827T1 | Austria | T1 | |
| DE602008002001D1 | Germany | D1 | |
| US7929043B2 | United States of America | B2 | |
| US2011211081A1 | United States of America | A1 | |
| JP4958610B2 | Japan | B2 | |
| US8508651B2 | United States of America | B2 |
67 legal events, as 7 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Application deemed withdrawn, or ip right lapsed, due to non-payment of renewal feeWithdrawnR119 | R119 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Amendment of ipc main classPREVIOUS MAIN CLASS: H04N0005232000R079 | R079 | DE | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Gb: european patent ceased through non-payment of renewal feeCeasedGBPC | GBPC | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent ceasedCeasedPL | PL | CH | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent lapsedLapsedMM4A | MM4A | IE | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Notification of lapseLapsedST | ST | FR | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| No opposition filed against granted patent, or epo opposition proceedings concluded without decisionGrantedR097 | R097 | DE | |
| No opposition filedOpposition26N | 26N | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lt: invalidation of european patent or patent extensionLTIE | LTIE | EP | |
| Discontinued in the netherlands as no translation has been filedVDEP | VDEP | NL | |
| Corresponds to:REF | REF | EP | |
| European patents granted designating irelandGrantedFG4D | FG4D | IE | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedFG4D | FG4D | GB | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| Designation fees paidAKX | AKX | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAX | AX | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAX | AX | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 1978731
- Publication, DOCDB
- 1978731
- Publication, EPODOC
- EP1978731
- Application
- 8006491
- Application, DOCDB
- 08006491
- Application, EPODOC
- EP20080006491
Titles3
- German
- Bildstabilisierungsvorrichtung, Bildaufnahmevorrichtung und Bildstabilisierungsverfahren
- English
- Image stabilizing apparatus, image pick-up apparatus and image stabilizing method
- French
- Appareil de stabilisation d'image, appareil de capture d'image et procédé de stabilisation d'image
Classification
- CPC, 6
- G06T7/20
- H04N23/68
- G06T2207/10016
- H04N5/145
- H04N23/6811
- H04N23/682
- IPC, 2
- H04N5 232
- H04N23 40
Designated states2
- Contracting states, 1
- Türkiye
- Extension states, 1
- Serbia