User interface device and user interface program
Abstract
Problem to be solved.To provide a stable user interface for a computer that eliminates the need for a user to move his or her hand by using facial gestures as a switch.
Solution.In the user interface, the position of an eye is extracted within a target image area (S402); the average value of brightnesses (average density of pixels) within a predetermined area around the eye is calculated (S406); based on comparison between the average value and information about the image including the eye previously registered in a storage device, the conditions of the eye as to whether it is open or closed and its line of sight, are detected (S418-S428). A system implements previously associated processes out of a plurality of predetermined processes according to the detected conditions of the eye.
Copyright (C)2006,JPO&NCIPI

Term
No projected expiry on record.
- Priority and filed
- Published
- Today
6 claims: 2 independent, 4 dependent
- 1A photographing means for acquiring digital data of the value of each pixel in the target image area including the user's face area, an eye detecting means for extracting the position of the eyes in the target image area, and the detected eyes. An eye condition determining means for determining the eye condition including the opening / closing of the eye and the eye direction based on the comparison with the information about the image including the eye registered in the storage device in advance based on the position of the eye, and at least detected. A processing means for selecting and executing a process associated in advance from a plurality of predetermined processes according to the state of the eyes, and an output means for outputting the result corresponding to the executed process. A user interface device provided. ユーザの顔領域を含む対象画像領域内の各画素の値のデジタルデータを獲得する撮影手段と、 前記対象となる画像領域内において、目の位置を抽出する目検出手段と、 前記検出された目の位置に基づき、予め記憶装置に登録された前記目を含む画像についての情報との比較に基づいて、目の開閉および視線方向を含む目の状態を判定する目状態判定手段と、 少なくとも検出された前記目の状態に応じて、複数の所定の処理のうち、予め対応づけられた処理を選択して実行する処理手段と、 実行された前記処理に対応する結果を、出力する出力手段とを備える、ユーザインタフェース装置。
- 5A user interface program for causing a computer to execute user interface processing, in which a step of acquiring digital data of the value of each pixel in a target image area including a user's face area and a step of acquiring digital data of the value of each pixel in the target image area and the target image area. Based on the step of extracting the eye position and the comparison with the information about the image including the eye registered in the storage device in advance based on the detected eye position, the eye including the opening / closing of the eye and the line-of-sight direction. Corresponds to the step of determining the state of the above, the step of selecting and executing the pre-associated process from among a plurality of predetermined processes at least according to the detected state of the eyes, and the executed process. A user interface program that causes a computer to perform steps to output the results to an output device. コンピュータにユーザインタフェース処理を実行させるためのユーザインタフェースプログラムであって、 ユーザの顔領域を含む対象画像領域内の各画素の値のデジタルデータを獲得するステップと、 前記対象となる画像領域内において、目の位置を抽出するステップと、 前記検出された目の位置に基づき、予め記憶装置に登録された前記目を含む画像についての情報との比較に基づいて、目の開閉および視線方向を含む目の状態を判定するステップと、 少なくとも検出された前記目の状態に応じて、複数の所定の処理のうち、予め対応づけられた処理を選択して実行するステップと、 実行された前記処理に対応する結果を、出力装置に出力するステップとをコンピュータに実行させる、ユーザインタフェースプログラム。
Independent claims2
132 paragraphs, as filed
The present invention relates to a user interface device and a user interface program that use a user's image taken by a camera or the like as a computer user interface.
In the user interface of a computer, as an input device, a device operated by a human hand such as a keyboard or a mouse is generally used.
On the other hand, using so-called voice recognition technology, software for directly inputting characters etc. to a computer by human utterance has also been developed, and by using the voice input function of a general-purpose computer, practical level ones are sold. Has been done.
However, for example, for a person who cannot move the muscles below the neck due to a spinal cord injury due to an accident and has difficulty in speaking, the conventional input device as described above cannot provide a sufficient input operation.
Therefore, as an input interface to a computer, and further, for communication with another person using such a computer, a device capable of detecting a user's line of sight and spelling characters has been proposed (for example, a patent). See Reference 1).
However, the operation using only the line of sight is not suitable for the pointing operation as described in Non-Patent Document 1, for example, and there is a problem that it is difficult to distinguish between pointing and mere movement of the line of sight.
Further, in order to analyze the line of sight, not only is it necessary to magnify the eye, but also a device other than the camera such as irradiating infrared light is required. There is also a type of device that attaches a device that detects the line of sight to the head, but it is difficult to use because people with disabilities cannot put it on and take it off by themselves.
On the other hand, technology for detecting a person from captured images is being actively researched as a technology indispensable for the development of fields such as human-computer interaction, gesture recognition, and security.
To detect a person, a method of detecting a face is effective. The face has important information such as facial expressions, and if the face can be detected, it becomes easy to estimate and search the positions of the limbs.
So far, many reports have been made on face detection systems using skin color information (see, for example, Patent Document 2, Non-Patent Document 2 to Non-Patent Document 3).
However, in these methods, the skin color region is extracted from the image and the face candidate region is obtained. Since the face candidate area can be limited, the processing range is limited and the amount of calculation can be significantly reduced, so that a high-speed system can be constructed. However, the method using color information is vulnerable to fluctuations in the lighting environment, and stable performance cannot be expected when operating in a general environment. In addition, it is necessary to extract the skin color area and perform pretreatment to limit the area, and a face with bangs extending to the eyebrows may not be detected because the above pattern does not appear. There was such a problem.<nplcit num="1"><text>Takehiko Ohno, "Interface Using Gaze", Information Processing, July 2003, Vol. 44, No. 7, pp.726-732</text></nplcit><nplcit num="2"><text>Shinjiro Kawato, Shinji Tetsuya, "Real-time detection of the eyebrows using a ring frequency filter" Shinjiron (D-II), vol.J84-D-II, no12, pp.2577-2584, Dec.2001.</text></nplcit><nplcit num="3"><text>Shinjiro Kawato, Shinji Tetsuya, "Real-time detection and tracking of eyes", Shingaku Giho, PRMU2000-63, pp.15-22, September 2000.</text></nplcit><patcit num="1"><text>Japanese Patent Application Laid-Open No. 2000-020196</text></patcit><patcit num="2"><text>Japanese Patent Application Laid-Open No. 2001-52176</text></patcit>
<p> An object of the present invention is to provide a user interface device and a user interface program that can be used as a stable interface of a computer by using facial gestures as switches, eliminating the need for user hand movements.</p>
<p> According to a certain aspect of the present invention, the user interface device is a photographing means for acquiring digital data of the value of each pixel in the target image area including the user's face area, and the eye in the target image area. Eye condition including eye opening / closing and line-of-sight direction based on comparison between the eye detecting means for extracting the position and the information about the image including the eye registered in the storage device in advance based on the detected eye position. Corresponds to the eye condition determining means for determining, a processing means for selecting and executing a predetermined process associated with a plurality of predetermined processes at least according to the detected eye state, and the executed process. It is provided with an output means for outputting the result to be output.</p><p> Preferably, the position of the user's mouth is detected and the mouth including the open and closed states is compared with the information about the image including the mouth registered in the storage device in advance with respect to the user's mouth. Further, the mouth state determining means for determining the state is further provided, and the processing means selects and executes a predetermined process associated with the plurality of predetermined processes according to the combination of the eye state and the mouth state.</p><p> Preferably, the information about the image containing the eyes registered in the storage device in advance includes the average value of the brightness of the pixels in a predetermined range including the eyes normalized by the eye spacing.</p><p> Preferably, the information about the image including the mouth pre-registered in the storage device includes the average value of the brightness of the pixels in a predetermined range including the mouth normalized by the eye spacing.</p><p> According to another aspect of the present invention, a user interface program for causing a computer to perform user interface processing, which is a step of acquiring digital data of the value of each pixel in a target image area including a user's face area. Opening and closing the eyes based on the step of extracting the eye position in the target image area and the comparison with the information about the image including the eyes registered in the storage device in advance based on the detected eye position. And a step of determining the eye condition including the line-of-sight direction, and a step of selecting and executing a pre-associated process from a plurality of predetermined processes according to at least the detected eye condition. Have the computer execute the step of outputting the result corresponding to the processing to the output device.</p><p> Preferably, the mouth including the open and closed states is based on a comparison of the step of detecting the position of the user's mouth with information about the image including the mouth previously registered in the storage device with respect to the user's mouth. The step for determining the state of the above is further provided, and the step to be executed selects and executes a predetermined process associated with the plurality of predetermined processes according to the combination of the eye state and the mouth state.</p>
[Embodiment 1]
[Hardware configuration]
Hereinafter, the user interface device according to the embodiment of the present invention will be described. This user interface device is realized by software executed on a computer such as a personal computer or a workstation, extracts a person's face from a target image, and further extracts eyes and mouth from the image of the person's face. The purpose is to detect the position of the computer, determine the state of the eyes and mouth, and realize the function as an input device of a computer by combining these states.
FIG. 1 is a block diagram showing a configuration of a system 100 in which the user interface device of the present invention operates.
The system 100 includes a computer body 40 equipped with a CD-ROM (Compact Disc Read-Only Memory) drive 50 and an FD (Flexible Disk) drive 52, a display 42 as a display device connected to the computer body 40, and a computer as well. It includes a keyboard 46 and a mouse 48 as input devices connected to the main body 40, and a camera 30 connected to the computer main body 40 for capturing images. In the apparatus of this embodiment, a camera including a solid-state image sensor such as a CCD (solid-state image sensor) is used as the camera 30, and the positions and eyes of the eyes and mouth of a person who operates the system 100 in front of the camera 30. And the process of detecting the condition of the mouth shall be performed.
That is, the camera 30 prepares digital data of the value of each pixel in the target image region, which is an image including the human face region.
As shown in FIG. 1, the computer main body 40 constituting this system 100 includes a CPU (Central Processing Unit) 56 connected to the bus 66 and a ROM (Central Processing Unit) 56, respectively, in addition to the CD-ROM drive 50 and the FD drive 52. It includes a Read Only Memory) 58, a RAM (Random Access Memory) 60, a hard disk 54, and an image capture device 68 for capturing images from the camera 30. The CD-ROM 62 is installed in the CD-ROM drive 50. The FD drive 52 is equipped with the FD64.
As already mentioned, the main part of this user interface device is realized by computer hardware and software executed by CPU 56. Generally, such software is stored in a storage medium such as a CD-ROM 62 or FD64 and distributed, and is read from the storage medium by a CD-ROM drive 50 or an FD drive 52 or the like and temporarily stored in a hard disk 54. Alternatively, if the device is connected to the network, it is temporarily copied from the server on the network to the hard disk 54. Then, it is further read from the hard disk 54 to the RAM 60 and executed by the CPU 56. If it is connected to the network, it may be directly loaded into the RAM 60 and executed without being stored in the hard disk 54.
The computer hardware itself and its operating principle shown in Fig. 1 are general. Therefore, the most essential part of the present invention is software stored in a storage medium such as an FD64 or a hard disk 54. In addition to this, the recording medium may be a memory card or a DVD (Digital Versatile Disc) -ROM. In this case, a reading drive device corresponding to such a medium is provided in the main body 40.
As a recent general tendency, it is common to prepare various program modules as a part of the operating system of a computer, and to call these modules in a predetermined array when necessary to proceed with processing. is there. In such a case, the software itself for realizing the user interface device does not include such a module, and the user interface device is realized only in cooperation with the operating system on the computer. However, as long as a general platform is used, it is not necessary to distribute software containing such modules, and the software itself without those modules and the recording medium on which the software is recorded (and the software is distributed on the network). Case data signals) can be considered to constitute an embodiment.
[User interface processing]
Next, the operation of the system 100 as a user interface device will be described.
A method of detecting a human face from an image and detecting the position of an eye from the inside of the face is disclosed in Non-Patent Documents 1 and 2 described above. It is also disclosed in Japanese Patent Application No. 2002-338175 and Japanese Patent Application No. 2003-391148 of the patent application by the inventor of the present application.
Therefore, in the following description, it is assumed that the detection of the eye position from the image is performed. After that, the method of detecting the position of the eyes disclosed in Japanese Patent Application No. 2002-338175 will be described.
(Processing at system startup) System 100 uses facial gestures as switches and as a computer interface.
Therefore, the detected condition of the eyes and mouth, for example, the condition of the eyes, is "looking at the front", "looking at the top", "looking at the bottom", "looking at the right", It is a digital switch with 6 states of "looking to the left" and "closed eyes", and the state of b is "closed", "open", and "protruding tongue". It is a digital switch with 3 states. When combined, 18 states can be expressed. Of course, a smaller number of combinations may be used without using all the cases described above. For example, it is possible to configure the user interface only by the state of the eyes. On the contrary, for example, a state such as "open the mouth wide" may be added to the state of the mouth to express more combinations.
For this purpose, the system 100 performs a process of registering such a mouth state and an eye state for each user in advance at the time of system startup. It should be noted that once registered, basically, for the same user, the state of the mouth and the state of the eyes can be determined from the second time onward based on the information registered in the hard disk 54, for example.
FIG. 2 is a flowchart for explaining such a registration process at the time of system startup.
First, when the system 100 starts up, the message "Close your mouth and look at the center of the monitor screen" is displayed on the display 42 (step S102).
According to the procedure described later, the eyes and the corners of the mouth are detected, and the detection position is displayed on the image on the display 42 (step S104).
Then, "The system is detecting the eyes and mouth. If you like, please stick out your tongue to signal." Is displayed on the display 42 (step S106).
The system 100 detects that the average density of the mouth region has changed brightly above a predetermined threshold value, and sets the average density before and after the change as a state in which the mouth is closed and the tongue is sticking out, for example, on the hard disk 54. Remember. At the same time, the reference pattern between the eyes of the user (the reference pattern of the rectangular area including the space between both eyes in the front view) and the reference eye position (the position of the eyes in the reference pattern between the eyes (the position of the black eye)). Relative data (corresponding to front view) and the average density of the eye area when the eyes are opened are calculated and stored in the hard disk 54 (S108).
System 100 then displays, "The tongue has been detected. Next, open your mouth slightly, close your eyes, and slowly count from 1 to 3" (S110).
It is detected that the average concentration of the mouth region has changed to be darker than that of the closed mouth state, and the average concentration after the change is stored as the open mouth state. At the same time, the reference eye position is calculated from the already acquired inter-eye reference pattern and the reference eye position relative data, and the average density of the eye area is stored in the hard disk 54 as the average density of the eye area when the eyes are closed (step S112). Sequentially instruct "Look to the right", "Look to the left", "Look up", "Look down", and the reference eye position of the eye position corresponding to the viewing direction. The amount of deviation from is detected and stored in the hard disk 54 (step S114).
Then, when the registration process is completed, the system 100 displays Initialization process completed on the display 42 (step S116).
After that, the system 100 determines the eye state as 3-bit data and the mouth state as 2-bit data each time a new image is input, and supports predetermined processing according to the combination of these. Output the processing result (step S118).
For example, when System 100 is used as a character editing system, the cursor is moved by the change in the state of the eyes with the tongue out, and the change in the opening of the mouth is regarded as a click action, and the characters are edited. Processing can be performed. No action is performed when the state of the eyes changes with the mouth closed. By doing so, it is possible to distinguish between a change of state of the eye intended for input and a change of state of the eye due to mere movement of the line of sight that is not intended for input.
(Mouth position detection flow) Next, the mouth position detection flow will be described.
As described above, the flow of mouth detection will be described assuming that the positions of both eyes are detected in advance.
It is assumed that the image taken by the camera 30 is rotationally corrected by the system 100 so that both eyes are aligned horizontally based on the detected eye positions.
At this time, if the distance between the eyes is Le, the mouth is below the eyes, for example, in the range of 0.7Le to 1.4Le, and the width of the mouth is almost equal to Le in the normal state with the mouth closed, and the mouth is opened. , When the mouth is sharpened, the width of the mouth becomes smaller than Le.
FIG. 3 is a flowchart for explaining the detection of the mouth position.
First, for each scanning line (y coordinate: vertical coordinate) of the image, the average density of pixels between the x coordinate (horizontal coordinate) of the left eye and the x coordinate of the right eye is calculated and plotted (step S202). ..
FIG. 4 is a diagram in which the average density of pixels between the x-coordinate of the right eye and the x-coordinate of the left eye is plotted for each scanning line (horizontal direction) in the lower half of the face.
As shown in FIG. 4, when the average concentration of the width Le is seen, the position of the mouth is the darkest, so this can be specified as the y-coordinate of the mouth. That is, the y-coordinate with the lowest average density (dark) is found between 0.7Le and 1.4Le under the eyes and used as the y-coordinate of the mouth.
FIG. 5 is a conceptual diagram showing a left corner template and a right corner template.
On the y-coordinate of the mouth determined as described above, the x-coordinate that best matches the left mouth corner template and the right mouth corner template as shown in FIG. 5 is searched for and used as the left and right edges of the mouth.
That is, first, on the y-coordinate of the mouth, the x-coordinate that best matches the left mouth angle template is set as the left end of the mouth (step S206).
Next, on the y-coordinate of the mouth, the x-coordinate that best matches the right mouth angle template is set as the right edge of the mouth (step S208).
In this case, since the degree of match does not matter, both ends of the mouth can be detected regardless of the opening and closing of the mouth.
(Mouth condition detection flow) Next, the mouth condition detection flow will be described.
FIG. 6 is a conceptual diagram showing the detected shape of the mouth.
In FIG. 6, as described below, let M be the distance between the detected left and right corners of the mouth.
FIG. 7 is a flowchart for explaining a process for detecting the state of the mouth.
With reference to FIG. 7, first, the distance between the detected left and right corners of the mouth is M (step S302), and the midpoint of the left and right corners of the mouth is the center of the mouth (step S304).
Next, in order to detect the state of the mouth, the average density (brightness) in the rectangle of, for example, 0.7M × 0.2Le shown in FIG. 6 is calculated (step S306).
This average concentration is darker when the mouth is open than when the mouth is closed, and brighter when the tongue is out than when the mouth is closed. Therefore, as described above, when the system is started, the system 100 is instructed in the three states, and the average density of the three states is stored in the hard disk 54.
Next, based on the calculated average concentration, the value is compared to which of the above three states is close (step S308), and the current mouth state is determined (step S310).
When the system is operating, each time the corner of the mouth is detected, this average concentration is calculated to determine which of the three states the value is close to, and the state of the mouth is determined.
(Eye condition detection flow) Next, the eye condition detection flow will be described.
FIG. 8 is a flowchart for explaining a process for detecting the state of the eyes.
With reference to FIG. 8, when the system is operating, the eye-to-eye position is first detected by template matching of the eye-to-eye reference pattern for the input face image (step S402), and the relative position of the eyes from the eye-to-eye when viewed from the front. Based on the data, the left and right reference eye positions are calculated (step S404).
Next, the average density of a rectangular region centered on the calculated reference eye position, for example, 0.3Le × 0.2Le, is calculated (step S406). It is determined whether the calculated average concentration is close to the value at the time of closing the eye or the value at the time of opening the eye, and whether it is in the closed state or the open state (step S408).
If the eyes are closed (step S410), it is determined that the eyes are closed and the result is returned to the main routine of the user interface (step S428).
On the other hand, if the eye is open (step S410), the deviation between the detected eye position and the reference eye position is calculated (step S412), and if the deviation is less than or equal to the threshold value (step S414), it is determined to be in the front view state. (Step S418).
Although not particularly limited, the amount of misalignment of the eyes is determined by the template image of the eye part in the front view and the eye part when the eyes are directed in the specified direction, as explained in FIG. It is possible to acquire the template image of the above in advance and store it in the hard disk 54, and identify it by comparing with these.
If the deviation is greater than or equal to the threshold value (step S414), compare which case of left / right / up / down vision is closest (step S416), and depending on the closest case, look up (look up) with the eyes open. State) (step S420), lower view with eyes open (looking down) (step S422), left view with eyes open (looking left) (step S424), right view with eyes open (right) It is judged as (as seen) (step S426), and the result is returned to the main routine of the user interface.
The 6 states of the eyes and the 3 states of the mouth can be distinguished from the obtained image by an inexpensive image input device such as a USB (Universal Serial Bus) camera, and infrared irradiation is not required, so if you use a zoom lens, you can distinguish between the user and the user. The distance can be set freely, and it can be operated without wearing it.
As another example of such processing based on the user interface, for example, the cursor at the cursor position is selected by operating the cursor on the screen with the direction of the eyes and closing the eyes. However, when the tongue is sticking out, it is assumed that the eyes are used for the purpose of operation, and if the tongue is not sticking out, it is judged that the eyes are moving only for the purpose of seeing.
In the above processing, detailed information such as which position the user's line of sight is directed on the display 42 is not required, so it is easy to distinguish between the computer and the user by distinguishing between the eye condition and the mouth condition. An interface between them can be realized.
FIG. 9 is a diagram showing an example of an image in which the combination of opening and closing of the mouth and the direction of the line of sight is changed.
FIG. 9 shows 8 of the 18 images described above.
Next, the method described in Japanese Patent Application No. 2002-338175 of the above-mentioned patent application will be described below as an example of a method of detecting a human face from an image and detecting the position of an eye from the inside of the face. ..
[Basic principle of face image extraction]
First, to summarize the outline of the procedure for detecting the position of the eyes, when processing a video image in which a face is continuously photographed, the screen is scanned with a rectangular filter having a width of the face and a height of about half of the width of the face. For example, the rectangle is divided into 6 parts of 3 × 2, and when the average brightness of each divided area is calculated and the relative lightness-dark relationship satisfies the condition, the center of the rectangle is set as the glabellar candidate.
When a continuous pixel becomes a glabellar candidate, only the center candidate of the frame surrounding it is left as the glabellar candidate. By comparing the remaining glabellar candidates with the standard pattern (the above-mentioned standard pattern between the eyes) and performing template matching, etc., the fake glabellar candidates are discarded from the glabellar candidates obtained by the above procedure, and the true glabellar candidates are discarded. Extract the glabellar.
Hereinafter, the procedure for face detection of the present invention will be described in more detail.
(6-Divided Rectangle Filter) FIG. 10 is a diagram showing the above-mentioned 3 × 2 6-divided rectangular filter (hereinafter referred to as 6-divided rectangular filter).
The 6-segment rectangular filter is a filter that finds the position between the eyebrows of the face by extracting facial features such as 1) the nose muscles are brighter than the binocular region and 2) the eye region is darker than the cheeks. A rectangular frame with horizontal i pixels and vertical j pixels (i, j: natural numbers) is provided around the point (x, y).
As shown in FIG. 10, this rectangular frame is divided into three equal parts horizontally and two equal parts vertically, and is divided into six blocks S1 to S6.
FIG. 11 is a conceptual diagram showing a case where such a 6-segment rectangular filter is applied to a face image. FIG. 11 (a) shows the shape of the 6-split rectangular filter, and FIG. 11 (b) shows the state in which the 6-split rectangular filter is applied to the binocular region and the cheek portion of the face image.
Considering that the nasal muscle is usually narrower than the eye area, it is more desirable that the width w2 of blocks S2 and S5 is narrower than the width w1 of blocks S1, S3, S4 and S6. Preferably, the width w2 can be half the width w1. FIG. 12 is a conceptual diagram showing the configuration of a 6-segment rectangular filter in such a case.
In the following description, it is assumed that a 6-segment rectangular filter as shown in FIG. 12 is used.
Further, the vertical width h1 of the blocks S1, S2 and S3 and the vertical width h2 of the blocks S4, S5 and S6 do not necessarily have to be the same. However, in the following description, the vertical width h1 and the vertical width h2 will be described as being equal to each other.
In the 6-segment rectangular filter shown in FIG. 12, for each block Si (1 i 6), the average value bar Si of the brightness of the pixels (Si is superscripted with -) is obtained.
Assuming that one eye and eyebrow exist in block S1 and another eye and eyebrow exist in block S3, the following relational expression (1) holds.
<maths num="1"><img file="JP2005293061A_D0001.tif" /></maths>
FIG. 13 is a conceptual diagram showing an image to be scanned by such a 6-segment rectangular filter.
As shown in FIG. 13, the target image for detecting the face image is composed of M × N pixels of M pixels in the horizontal direction and N pixels in the vertical direction. In principle, the work of checking the validity of the above relational expression (1) by applying the above 6-division rectangular filter while sequentially shifting the pixel (0,0) in the upper left corner by one pixel in the horizontal direction and the vertical direction. Will be done. However, it is inefficient to obtain the average value of the brightness in each block every time the 6-division rectangular filter is shifted in this way.
Therefore, in the present invention, regarding the process of obtaining the sum of the pixels in the rectangular frame, a known document (P. Viola and M. Jones, Rapid Object Detection using a Boosted Cascade of Simple Features, Proc. Of IEEE Conf. CVPR. , 1, pp.511-518, 2001), adopts the method of speeding up the calculation using the integral image.
From the image i (x, y), the "integral image" is defined by the following equation (2).
<maths num="2"><img file="JP2005293061A_D0002.tif" /></maths>
The integral image can be obtained by repeating the following.
<maths num="3"><img file="JP2005293061A_D0003.tif" /></maths>
s (x, y) represents the sum of the pixels in the row. However, s (x, -1) = 0, ii (-1, y) = 0. The important point is that the integral image can be obtained by scanning the entire image once.
By using the integral image, the sum of the brightness values of the pixels in the rectangular region can be easily obtained. FIG. 14 is a diagram showing a rectangular region for which the sum is calculated by using such an integral image.
Using an integral image, the total luminance Sr of the pixels in the frame of the rectangle D shown in FIG. 14 can be obtained by calculating the values of four points as follows.
<maths num="4"><img file="JP2005293061A_D0004.tif" /></maths>
In this way, by using the integral image, the sum of the luminance values of the pixels in the rectangular region, and by extension, the average of the luminance values of the pixels can be obtained at high speed, so that the 6-division rectangular filter is processed at high speed. It is possible.
(Glabellar Candidate Point Extraction Process) The process of extracting the glabellar candidate points using the above-mentioned 6-division rectangular filter will be described below.
FIG. 15 is a flowchart for explaining a process of extracting candidate points between the eyebrows.
With reference to FIG. 15, first, as the initialization process, the values of the variables m and n are set to m = 0 and n = 0 (step S1000).
Then, the upper left corner of the 6-segment filter is aligned with the (m, n) pixels of the image (step S1020). Further, the average density bar Si of the pixels in the region of block Si is calculated (step S1040).
Next, it is tested whether the magnitude of the value of the average density bar Si satisfies the glabellar candidate condition according to the equation (1) (step S1060).
If the test condition is satisfied (step S1080), a glabellar candidate mark is added to the pixel at the position (m + i / 2, n + j / 2) corresponding to the center point of the filter (step S1100). On the other hand, if the test conditions are not satisfied (step S1080), the process proceeds to step S1120.
In step S1120, the value of the variable m is incremented by 1. Next, it is determined whether the value of the variable m is within the range in which the filter can move in the horizontal direction in the target image (step S1140). When the filter is within the movable range, the process returns to step S1020. On the other hand, when the filter is at the limit of lateral movement, the value of the variable m is reset to 0 and the value of the variable n is incremented by 1 (step S1160).
Next, it is determined whether the value of the variable n is within the range in which the filter can move in the vertical direction in the target image (step S1180). When the filter is within the movable range, the process returns to step S1020. On the other hand, when the filter is at the limit of being able to move in the vertical direction, the eyebrows candidate mark is attached to check the pixel connectivity, and the pixel in the center of the outer frame of the connecting element is set as the eyebrows candidate point for each connecting element ( Step S1200). Here, the center pixel is not particularly limited, but may be, for example, the position of the center of gravity of each connecting element.
(Extraction of eye candidate points and extraction of true glabellar candidate points) The glabellar candidate points extracted as described above include false glabellar candidate points in addition to the true glabellar candidate points. Therefore, the true glabellar candidate points are extracted by the procedure described below.
First, the candidate points for the eye positions are extracted based on the information on the candidate points between the eyebrows.
For that purpose, a plurality of eye images are extracted from the face image database, and the average image is obtained. FIG. 16 is a diagram showing the template of the right eye thus obtained. For the left eye template, this right eye template may be inverted horizontally.
Using this right-eye template and left-eye template, if template matching processing is performed in the areas of blocks S1 and S3 of the 6-segment rectangular filter centered on the glabellar candidate point shown in FIG. 10, each candidate for the right eye and the left eye can be obtained. Points can be extracted.
FIG. 17 is a flowchart for explaining a process of extracting true eyebrows candidate points after extracting eye candidate points.
With reference to FIG. 17, first, in each region of the block S1 and S3 of the glabellar candidate extraction filter, the point that best matches the eye template is searched for and used as the candidate points for the left and right eyes (step S2000).
Next, the position of the candidate point between the eyebrows is corrected to the midpoint of the candidate points of the left and right eyes (step S2020). Then, the input image is rotated so that the candidate points of the left and right eyes are arranged horizontally around the position of the candidate point between the corrected eyebrows (step S2040).
The similarity between the pattern centered on the corrected glabellar candidate point after rotation and the glabellar template formed in advance by the procedure described later is calculated (step S2060).
It is determined whether the similarity is equal to or higher than a predetermined threshold value (step S2080), and if it is equal to or higher than the threshold value, it is set as a true glabellar candidate point (step S2100). On the other hand, if it is less than the threshold value, it is set as a false glabellar candidate point (step S2120).
Such processing is performed for all the candidate points between the eyebrows.
In the following, the above-mentioned "standard pattern between eyes" will be referred to as "glabella template".
In the present application, the glabellar template is set for each user as described above.
Next, the template matching process of step S2060 in FIG. 17 will be described in more detail.
FIG. 18 is a flowchart for explaining the template matching procedure of step S2060.
With reference to FIG. 18, first, the glabellar candidate points are extracted (step S4000), and if necessary, rotation is performed around the glabellar candidate points to perform scale correction (step S4020).
Next, an image of the same size as the template is cut out centering on the candidate point between the eyebrows (step S4040). The correlation value between the cut out glabellar candidate pattern and the glabellar template is calculated and used as the similarity (step S4060).
To calculate the similarity, the density of the cut-out candidate pattern between the eyebrows is normalized (mean zero, dispersion 1.0), the square of the difference from the corresponding pixel of the template is calculated for each pixel, and the sum is calculated. It may be that. That is, in this case, since the value of the sum can be regarded as the degree of dissimilarity, the degree of similarity may be evaluated by the reciprocal of this.
The position of the eyes can be detected by the above procedure.
It should be considered that the embodiments disclosed this time are exemplary in all respects and not restrictive. The scope of the present invention is shown by the scope of claims rather than the above description, and it is intended to include all modifications within the meaning and scope equivalent to the scope of claims.
<figref num="1">It is a block diagram which shows the hardware structure of the system 100 which concerns on this invention.</figref><figref num="2">It is a flowchart for demonstrating the registration process at the time of system startup.</figref><figref num="3">It is a flowchart for demonstrating the detection of the mouth position.</figref><figref num="4">It is the figure which plotted the average density of the pixel from the x-coordinate of the right eye to the x-coordinate of the left eye for each scanning line (horizontal direction) about the lower half of a face.</figref><figref num="5">It is a conceptual diagram which shows the left mouth corner template and the right mouth corner template.</figref><figref num="6">It is a conceptual diagram which shows the shape of the detected mouth.</figref><figref num="7">It is a flowchart for demonstrating the process for detecting the state of a mouth.</figref><figref num="8">It is a flowchart for demonstrating the process for detecting the state of an eye.</figref><figref num="9">It is a figure which shows an example of the image which changed the combination of opening and closing of a mouth, direction of a line of sight.</figref><figref num="10">It is a figure which shows the 6-division rectangular filter.</figref><figref num="11">It is a conceptual diagram which shows the case where a 6-division rectangular filter is applied to a face image.</figref><figref num="12">It is a conceptual diagram which shows the other structure of a 6-division rectangular filter.</figref><figref num="13">It is a conceptual diagram which shows the image which is the object to scan the division rectangle filter.</figref><figref num="14">It is a figure which shows the rectangular area for which the sum is calculated using an integral image.</figref><figref num="15">It is a flowchart for demonstrating the process of extracting the candidate point between the eyebrows.</figref><figref num="16">It is a figure which shows the template of the right eye.</figref><figref num="17">It is a flowchart for demonstrating the process of extracting the true eyebrows candidate point after extracting the eye candidate point.</figref><figref num="18">It is a flowchart for demonstrating the procedure of template matching of step S2060.</figref>
Code description
20 face position extractor, 30 cameras, 40 computer body, 42 monitors.
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO2021260830A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2017068576A | Cited by | Japan | Search report |
| US9311528B2 | Cited by | United States of America | Applicant |
| JP2008299627A | Cited by | Japan | Examiner |
| WO2011118224A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US8413075B2 | Cited by | United States of America | Applicant |
| US8810624B2 | Cited by | United States of America | Applicant |
| JP2012190126A | Cited by | Japan | Search report |
| US8970614B2 | Cited by | United States of America | Applicant |
| US7840912B2 | Cited by | United States of America | Applicant |
| US8432367B2 | Cited by | United States of America | Applicant |
| KR100947990B1 | Cited by | Republic of Korea | Examiner |
| JP2007249595A | Cited by | Japan | Examiner |
| JP2010134057A | Cited by | Japan | Search report |
| WO2008146934A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO2021260831A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US9021347B2 | Cited by | United States of America | Applicant |
| US9014483B2 | Cited by | United States of America | Applicant |
| WO2021260829A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO2008085783A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2015014938A | Cited by | Japan | Examiner |
| WO2010064361A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| CN106557735A | Cited by | China | Search report |
| JP5689871B2 | Cited by | Japan | Search report |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004105250 | Japan | A | |
| JP20040105250 | – | – | – |
18 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentA521 | A521 | |
| Notification of reasons for refusalA131 | A131 | |
| Written amendmentA521 | A521 | |
| Notification of reasons for refusalA131 | A131 | |
| Report on retrievalA977 | A977 | |
| Written request for application examinationA621 | A621 |
Numbers
- Publication
- 2005293061
- Publication, DOCDB
- 2005293061
- Publication, EPODOC
- JP2005293061
- Application
- 105250
- Application, DOCDB
- 2004105250
- Application, EPODOC
- JP20040105250
Titles2
- Japanese
- ユーザインタフェース装置およびユーザインタフェースプログラム
- English
- User interface device and user interface program
Classification
- IPC, 2
- G06T1 00
- G06F3 033