Face image detection method, face image detection system, and face image detection program
Abstract
Provides a new face image detection method, face image detection system, and face image detection program, which can detect the presence of a human face image at high speed and with good accuracy in an image that has not yet been determined whether it contains a human face Highly likely areas. Divide the detection target area into plural blocks and perform dimensional compression to calculate the feature vector constituted by each representative value of each block, and use the feature vector to identify the presence or absence of the previous detection target area by the recognizer There are facial images. In other words, the recognition is performed after dimensional compression is performed to the extent that the original feature quantity of the image is not damaged. In this way, because the pixel feature amount used in recognition is greatly reduced from the number of pixels in the detection target area to the number of blocks, the amount of calculation can be greatly reduced to achieve high-speed facial image detection.
Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
14 claims: 14 independent, 0 dependent
- 1一種臉部影像偵測方法,係屬於在未判明是否含有臉部影像之偵測對象影像中,偵測是否有臉部影像存在之方法,其特徵為,將前記偵測對象影像內的所定領域當作偵測對象領域予以選擇,除了算出所選擇之偵測對象領域內的邊緣(edge)強度,還根據所算出之邊緣強度而將該當偵測對象領域內分割成複數區塊後,算出以每一區塊之代表值所構成之特徵向量,然後,將這些特徵向量輸入識別器以偵測前記偵測對象領域內是否有臉部影像存在。
- 2如申請專利範圍第1項所記載之臉部影像偵測方法,其中,前記區塊的大小,係根據自我相關係數而決定。
- 3如申請專利範圍第1項或第2項所記載之臉部影像偵測方法,其中,取代前記邊緣強度,改以求出邊緣強度和前記偵測對象領域的亮度值,根據該亮度值而算出以每一區塊之代表值所構成之特徵向量。
- 4如申請專利範圍第1項或第2項所記載之臉部影像偵測方法,其中,前記每一區塊之代表值,是採用構成前記每一區塊之像素的像素特徵量之分散值或平均值。
- 5如申請專利範圍第1項或第2項所記載之臉部影像偵測方法,其中,前記識別器,是採用預先學習了複數學習用樣本臉部影像和樣本非臉部影像的支撐向量機(Support Vector Machine)。
- 6如申請專利範圍第5項所記載之臉部影像偵測方法,其中,前記支撐向量機的識別函數,是使用非線性的基核函數(kernel function)。
- 7如申請專利範圍第1項或第2項所記載之臉部影像偵測方法,其中,前記識別器,是採用預先學習了複數學習用樣本臉部影像和樣本非臉部影像的類神經網路。
- 8如申請專利範圍第1項或第2項所記載之臉部影像偵測方法,其中,前記偵測對象影像內的邊緣強度,係使用各像素中的索貝爾運算子(Sobel operator)來予以算出。
- 9一種臉部影像偵測系統,係屬於在未判明是否含有臉部影像之偵測對象影像中,偵測是否有臉部影像存在之系統,其特徵為,具備:影像讀取部,將前記偵測對象影像及該當偵測對象影像內的所定領域當作偵測對象領域而予以讀取;及特徵向量算出部,將前記影像讀取部所讀取到的偵測對象領域內再次分割成複數區塊而將該每一區塊的代表值所構成之特徵向量予以算出;及識別部,根據前記特徵向量算出部所得之每一區塊之代表值所構成之特徵向量,識別前記偵測對象領域內是否有臉部影像存在。
- 10如申請專利範圍第9項所記載之臉部影像偵測系統,其中前記特徵向量算出部,係由以下各部所構成:亮度算出部,將前記影像讀取部所讀取到的偵測對象領域內之各像素的亮度值予以算出;及邊緣算出部,算出前記偵測對象領域內之邊緣強度;及平均.分散值算出部,將前記亮度算出部所得之亮度值或前記邊緣算出部所得之邊緣強度或者兩者之值的平均值或分散值予以算出。
- 11如申請專利範圍第9項或第10項所記載之臉部影像偵測系統,其中前記識別部,是由預先學習了複數學習用樣本臉部影像和樣本非臉部影像的支撐向量機(Support Vector Machine)所成。
- 12一種記錄有臉部影像偵測程式之電腦可讀取之媒體,該程式係屬於在未判明是否含有臉部影像之偵測對象影像中,偵測是否有臉部影像存在之程式,其特徵為,可令電腦發揮以下的機能:影像讀取部,將前記偵測對象影像及該當偵測對象影像內的所定領域當作偵測對象領域而予以讀取;及特徵向量算出部,將前記影像讀取部所讀取到的偵測對象領域內再次分割成複數區塊而將該每一區塊的代表值所構成之特徵向量予以算出;及識別部,根據前記特徵向量算出部所得之每一區塊之代表值所構成之特徵向量,識別前記偵測對象領域內是否有臉部影像存在。
- 13如申請專利範圍第12項所記載之記錄有臉部影像偵測程式之電腦可讀取之媒體,其中,前記特徵向量算出部,係由以下各部所構成:亮度算出部,將前記影像讀取部所讀取到的偵測對象領域內之各像素的亮度值予以算出;及邊緣算出部,算出前記偵測對象領域內之邊緣強度;及平均.分散值算出部,將前記亮度算出部所得之亮度值或前記邊緣算出部所得之邊緣強度或者兩者之值的平均值或分散值予以算出。
- 14如申請專利範圍第12項或第13項所記載之記錄有臉部影像偵測程式之電腦可讀取之媒體,其中,前記識別部,是由預先學習了複數學習用樣本臉部影像和樣本非臉部影像的支撐向量機(Support Vector Machine)所成。
Independent claims14
61 paragraphs, as filed
Face image detection method and face image detection system, and computer readable media recorded with face image detection program
The present invention relates to pattern recognition or object recognition technology, in particular, it is used to detect at high speed whether there is a human face in an image that has not yet been determined whether it contains a face image detection method and a face image detection method Image detection system and facial image detection program.
Although the recognition accuracy of text or voice has been greatly improved with recent pattern recognition technology or the high performance of computer and other information processing devices, images of people, objects, scenery, etc. are reflected, for example, by borrowing In the pattern recognition of an image captured by a digital camera, etc., it is still a very difficult task to accurately and quickly recognize whether a human face is reflected in the image.
At the same time, however, it is necessary for the computer to automatically and correctly identify whether there is a human face in such images, and even who the person is. This is in the improvement of biometric technology or security, the rapidization of criminal investigations, and image data. The treatment. The realization of high-speed search operations is a very important issue, and there have been many proposals for such issues.
For example, in the following patent document 1 etc., for a certain input image, firstly, it is determined whether there is a human skin color area, the mosaic size is automatically determined for the human skin color area, the candidate area is mosaicked, and the distance to the face dictionary is calculated. It is determined whether there is a human face, and the human face is cut out, so as to reduce the false extraction caused by the influence of the background and the like, and more efficiently find the human face from the image.
[Patent Document 1] Japanese Patent Laid-Open No. 9-50528
<p>However, at the same time, in the prior art, although human faces are detected from images based on the "skin color", the color range of the "skin color" will be different due to the effects of lighting, etc., and there are often faces. The detection of some images is omitted or the screening cannot be carried out efficiently due to the background.</p><p>Therefore, the present invention is proposed in order to effectively solve these problems, and its purpose is to provide a new face image detection method, face image detection system and face image detection program, which can be used when it has not yet been determined whether there is an image of a human face. Medium, high-speed and high-precision detection of areas with high possibility of human face images.</p>
<heading level="1">[Invention 1]</heading><p>In order to solve the above problem, the face image detection method of Invention 1 is a method of detecting whether there is a face image in the detection target image that has not been determined whether it contains a face image. The predetermined area in the image of the object to be measured is selected as the area of the object to be detected. In addition to calculating the edge strength in the selected area of the object to be detected, the area of the object to be detected is divided into segments based on the calculated edge strength. After the complex number of blocks, the feature vector composed of the representative value of each block is calculated, and then these feature vectors are input to the recognizer to detect whether there is a face image in the pre-detection target area.</p><p>That is, as a technique for extracting a facial image from an image that has not yet been known whether it contains a facial image, or where there is no knowledge about the position it contains, in addition to the aforementioned method of using skin color, it is also based on brightness. It is a method of detecting the unique feature vector of the calculated facial image.</p><p>However, at the same time, in the method of using the usual feature vector, for example, even when only a 24×24 pixel facial image is detected, a huge number of feature vectors of 576 (24×24) dimensions must be used ( There are 576 elements of the vector), so high-speed facial image detection is not possible.</p><p>Therefore, as described above, the present invention divides the current detection target area into plural blocks, calculates the feature vector constituted by each representative value of each block, and uses the feature vector by the recognizer. Recognize whether there is a face image in the detection target area. In other words, the feature quantity of the image is dimensionally compressed to the extent that the feature of the face image is not compromised, and then the recognition is performed.</p><p>In this way, the image feature amount used in the recognition is greatly reduced from the number of pixels in the detection target area to the number of blocks, so that the amount of calculation can be drastically reduced to achieve face image detection. Furthermore, because of the use of edges, facial images with strong lighting fluctuations can also be detected.</p><heading level="1">[Invention 2]</heading><p>The face image detection method of the invention 2 is the face image detection method of the invention 1. The size of the pre-marked block is determined according to the self-correlation coefficient.</p><p>That is, as will be described in detail later, the self-correlation coefficient is used. According to the coefficient, the dimensional compression caused by blockization can be performed to the extent that the original features of the face image are not greatly damaged, so that the higher speed can be achieved. And implement facial image detection with high accuracy.</p><heading level="1">[Invention 3]</heading><p>The face image detection method of Invention 3 is the face image detection method described in Invention 1 or 2, instead of the edge intensity of the previous note, the edge intensity and the brightness value of the detection target area of the previous note are obtained according to the brightness Value and calculate the feature vector constituted by the representative value of each block.</p><p>In this way, when there is a facial image in the detection target area, the facial image can be recognized with high accuracy and high speed.</p><heading level="1">[Invention 4]</heading><p>The face image detection method of invention 4 is the face image detection method described in any one of inventions 1 to 3. The representative value of each block in the preceding note is the pixel that constitutes the pixel of each block in the preceding note The dispersion value or average value of the characteristic quantity.</p><p>In this way, it is possible to reliably calculate the prescriptive feature vector required for input to the recognition unit.</p><heading level="1">[Invention 5]</heading><p>The facial image detection method of invention 5 is the pre-notation recognizer in the facial image detection method described in any one of inventions 1 to 4, which uses sample facial images and sample non-face images for plural learning in advance. Support Vector Machine (Support Vector Machine).</p><p>That is, in the present invention, a support vector machine (Support Vector Machine) is used as the recognition part of the generated feature vector, so that it can recognize whether there is a human face in the selected detection target area with high speed and high accuracy. The image exists.</p><p>The so-called "Support Vector Machine (hereinafter referred to as SVM)" used in the present invention is described in detail later. It was developed by AT&T's V.Vapnik in the framework of statistical learning theory in 1995. The proposed learning machine that uses an index called "margin" to linearly separate all input data of two classes can find the best hyperplane. The ability of pattern recognition is Recognized as one of the best learning models. Also, as described later, even when linear separation is not possible, a technique called "kernel trick" can be used to achieve high recognition capabilities.</p><heading level="1">[Invention 6]</heading><p>The face image detection method of Invention 6 is the face image detection method of Invention 5, in which the recognition function of the aforementioned support vector machine uses a non-linear kernel function.</p><p>That is, although the basic structure of the support vector machine is a linear valve element, in principle, it cannot be applied to non-linearly separable data, that is, high-dimensional image feature vectors.</p><p>On the other hand, as a method of making non-linear classification possible by the support vector machine, high-dimensionalization can be exemplified. It is a method of mapping the original input data into a high-dimensional feature space by non-linear mapping, and performing linear separation in the feature space. As a result, it will be non-linear in the original input space. The result of linear recognition.</p><p>However, in order to obtain the non-linear mapping, a huge calculation is required. Therefore, the calculation of the non-linear mapping can be replaced with the calculation of a recognition function called a "kernel function" in fact. This is called the "kernel trick", by which the non-linear mapping can be avoided directly to overcome computational difficulties.</p><p>Therefore, if the identification function of the support vector machine used in the present invention adopts the non-linear "base kernel function", the high-dimensional image feature vectors that are originally non-linearly separable data can also be easily separated.</p><heading level="1">[Invention 7]</heading><p>The face image detection method of Invention 7 is the face image detection method described in any one of Inventions 1 to 4. The pre-notation recognizer uses sample facial images for plural learning and sample non-faces that have been learned in advance. Image-like neural network.</p><p>This "neural network" is a computer model that imitates the brain neural circuit network of biology, especially the PDP (Parallel Distributed Processing) model, which is a multi-layer neural network, which makes it possible to learn patterns that are not linearly separable. The representative of the classification technique of pattern recognition technology. However, generally speaking, when high-order feature quantities are used, the recognition ability on the quasi-neural network will gradually decrease. In the present invention, since the dimension of the image feature quantity is compressed, this kind of problem does not occur.</p><p>Therefore, even if the pre-note SVM is changed to use this type of neural network as the pre-note recognizer, high-speed and high-precision recognition can be implemented.</p><heading level="1">[Invention 8]</heading><p>The face image detection method of Invention 8 is the face image detection method described in any one of Inventions 1 to 7. The edge intensity in the image of the detection target is detected by using the Sobel operator in each pixel ( Sobel operator) to calculate it.</p><p>That is, the "Sobel operator" is a differential edge detection operator used to detect locations with sharp changes in shades such as edges or lines in an image.</p><p>Therefore, by using this "Sobel operator" to generate the edge intensity or edge dispersion value in each pixel, the image feature vector can be generated.</p><p>In addition, the shape of the "Sobel operator" is as shown in Figure 9 (a: horizontal edge) and (b: vertical edge). The result generated by each operator is squared and then taken. The square root can be used to find the edge strength.</p><heading level="1">[Invention 9]</heading><p>The face image detection system of the invention 9 belongs to a system that detects whether there is a face image in the detection target image that has not been determined whether it contains a face image. It is characterized by having: an image reading unit, The predetermined area in the previous detection target image and the current detection target image is read as the detection target area; and the feature vector calculation unit divides the detection target area read by the previous image reading unit again Into a complex number of blocks and calculate the feature vector constituted by the representative value of each block; and the identification part, according to the feature vector constituted by the representative value of each block obtained by the feature vector calculation part of the previous note, identify the pre-recorded detection Check whether there are facial images in the subject area.</p><p>As a result, as in Invention 1, the image feature value used in the recognition of the recognition unit is greatly reduced from the number of pixels in the detection target area to the number of blocks, so that the amount of calculation can be drastically reduced to achieve a facial image. Detection.</p><heading level="1">[Invention 10]</heading><p>The facial image detection system of Invention 10 is the facial image detection system described in Invention 9. The prescriptive feature vector calculation unit is composed of the following parts: the brightness calculation unit reads the prescriptive image reading unit The brightness value of each pixel in the detected object area is calculated; and the edge calculation part calculates the edge intensity in the detected object area before recording; and the average. The dispersion value calculation unit calculates the brightness value obtained by the preceding brightness calculation unit or the edge intensity obtained by the preceding edge calculation unit, or the average value or the dispersion value of the two values.</p><p>Thereby, as in Invention 4, it is possible to reliably calculate the prescriptive feature vector required for input to the recognition unit.</p><heading level="1">[Invention 11]</heading><p>The face image detection system of Invention 11 is the face image detection system described in Invention 9 or 10. The pre-recognition part is supported by sample facial images and sample non-face images for plural learning used in advance. Support Vector Machine (Support Vector Machine).</p><p>As a result, as in Invention 5, it is possible to quickly and accurately identify whether there is a human face image in the selected detection target area.</p><heading level="1">[Invention 12]</heading><p>The facial image detection program of Invention 12 is a program that detects whether there is a facial image in a detection target image that has not been determined whether it contains a facial image. Its feature is that it can make the computer perform the following functions: The image reading part reads the pre-detection target image and the predetermined area in the current detection target image as the detection target field; and the feature vector calculation part reads the detection read by the pre-image reading part The measurement target area is divided into plural blocks again, and the feature vector composed of the representative value of each block is calculated; and the recognition part is composed of the representative value of each block obtained by the aforementioned feature vector calculation part The feature vector is used to identify whether there is a face image in the field of the detection object.</p><p>In this way, in addition to obtaining the same effects as in Invention 1, these functions can also be realized one by one on software with general-purpose computer systems such as personal computers. Therefore, it can be more economical and more economical than realizing a dedicated device. Easily implement it. Moreover, each function can be improved easily by simply rewriting the program.</p><heading level="1">[Invention 13]</heading><p>The facial image detection program of Invention 13 is the facial image detection program described in Invention 12. The prescriptive feature vector calculation unit is composed of the following parts: the brightness calculation unit reads the prescriptive image reading unit The brightness value of each pixel in the detected object area is calculated; and the edge calculation unit calculates the edge intensity in the detected object area beforehand; and the average. The dispersion value calculation unit calculates the brightness value obtained by the preceding brightness calculation unit or the edge intensity obtained by the preceding edge calculation unit, or the average value or the dispersion value of the two values.</p><p>Thereby, as in Invention 4, it is possible to reliably calculate the prescriptive feature vector required for input to the recognition unit. In addition, as in Invention 12, these functions can be realized one by one on software using a general-purpose computer system such as a personal computer, so that it can be realized more economically and easily.</p><heading level="1">[Invention 14]</heading><p>The facial image detection program of Invention 14 is the facial image detection program described in Invention 12 or 13. The pre-recognition part is supported by sample facial images and sample non-face images for plural learning used in advance. Support Vector Machine (Support Vector Machine).</p><p>As a result, as in Invention 5, it is possible to quickly and accurately recognize whether there is a human face image in the selected detection target area. Also, as in Invention 12, a general-purpose computer such as a personal computer can be used. The system realizes these functions one by one on the software, so it can be realized more economically and easily.</p>
Hereinafter, the best mode for implementing the present invention will be described with reference to the drawings.
FIG. 1 is a diagram of an embodiment of the facial image detection system 100 according to the present invention. As shown in the figure, the face image detection system 100 is mainly composed of the following parts: an image reading part 10 for reading a sample image for learning and an image of a detection target, and generating an image by the image reading part. 10 The feature vector calculation unit 20 of the feature vector of the read image, the recognition unit 30 that recognizes from the feature vector generated by the feature vector calculation unit 20 whether the detection target image is a face image candidate area or not, that is, SVM (Support Vector Machine).
The image reading unit 10, specifically, a CCD (Charge Coupled Device) camera or a vidicon camera, an image scanner, a roller scanner, etc., such as a coefficient still camera or a digital camera, etc. It also provides the following functions: A/D conversion is performed on a predetermined area in the detected detection target image, and multiple facial images and non-face images as sample images for learning, and the digital image data are sequentially converted It is sent to the feature vector calculation unit 20.
The feature vector calculation unit 20 is composed of the following units: a brightness calculation unit 22 that calculates the brightness (Y) in the image, an edge calculation unit 24 that calculates the edge strength in the image, and an edge generated by the edge calculation unit 24 The intensity or the average value of the brightness or the edge intensity dispersion value generated by the previously noted brightness calculation unit 22. The dispersion value calculation unit 26; and provides the following functions: from the average. Among the pixel values sampled by the dispersion value calculation unit 26, a sample image and an image feature vector of each search target image are generated, and they are sequentially sent to the SVM 30.
SVM 30 provides the following functions: in addition to the image feature vectors of complex facial images and non-face images generated by the learning feature vector calculation unit 20 as learning samples, it also recognizes the feature vector calculation unit 20 based on the learning results Whether the predetermined area in the generated detection target image is a face image candidate area.
The SVM30 is a learning machine that uses an index called "margin" as described above to find the optimal hyperplane that is most suitable for linear separation of all input data, even when linear separation is not possible. The technique called "kernel trick" can also be used for downloading, and it can exert high recognition ability.
Then, the SVM30 used in this embodiment is divided into: 1. the step of learning, and 2. the step of identifying.
First, 1. The step of learning is to read most facial images and non-face images used as sample images for learning by the image reading unit 10 as shown in FIG. 1, and then generate them by the feature vector calculation unit 20 The feature vector of each face image is learned as an image feature vector.
After that, in the step of recognition, the selected areas in the detection target image are sequentially read, and the feature vector calculation unit 20 generates the feature vector of the image, and treats it as the feature vector. Input, and detect the high possibility area of the facial image based on the input image feature vector having any proper area for the recognition hyperplane.
Here, although the size of the face image and non-face image for the sample used for learning will be described later, for example, 24 pixel×24 pixel (pixel) is divided into a predetermined number of regions. It is performed to detect the area of the size of the block of the object area.
Furthermore, if the SVM is explained in detail based on the "Statistics of Pattern Recognition and Learning Type" (Iwanami Shoten, Hideki Aso, Hiroharu Tsuda, Masaki Murata) pp.107~118, the problem of recognition is When it is non-linear, a non-linear basic kernel function can be used in the SVM, and the recognition function at this time is shown in the following equation 1.
That is, when the value of Equation 1 is "0", it becomes the recognition hyperplane, and for cases other than "0", the distance between the recognition hyperplane and the recognition hyperplane calculated based on the given image feature vector is taken. Moreover, if the result of Equation 1 is non-negative, it is a face image, and if it is negative, it is a non-face image.
<maths><img file="TW200529093A_D0001.tif" /></maths>
x is the feature vector, x<sub>i</sub>System support vector; is a value generated by the feature vector calculation unit 20. The K-based base kernel function, and in this embodiment, the function of the following formula 2 is used.
[Number 2]K(x,x<sub>i</sub>)=(a*x*x<sub>i</sub>+b)<sup>T</sup>Let a=1, b=0, T=2
In addition, the feature vector calculation unit 20, SVM 30, and image reading unit 10 constituting the facial image detection system 100 are actually hardware formed by CPU or RAM, etc., and a dedicated computer program ( It can be realized by a personal computer (PC) and other computer systems formed by software).
That is, the computer system used to implement the facial image detection system 100 is, for example, as shown in FIG. RAM (Random Access Memory) 41 used in the main storage, read-only memory device that is ROM (Read Only Memory) 42, hard disk drive device (HDD) or semiconductor memory and other auxiliary memory devices ( Secondary Storage) 43, and display (LCD (liquid crystal display) or CRT (cathode picture tube)) and other output devices 44, image scanner or keyboard, mouse, CCD (Charge Coupled Device) or CMOS (Complementary Metal Oxide) The input device 45 formed by camera sensors such as Semiconductor) and the input/output interface (I/F) 46 of these devices are connected by PCI (Peripheral Component Interconnect) bus or ISA (Industrial Standard Architecture) bus, etc., which are composed of processor bus, memory bus, system bus, input/output bus, and other internal exchange bus 47 which are connected together.
Then, for example, install CD-ROM or DVD-ROM, floppy disk (FD) and other storage media, or various control programs or data supplied through the communication network (LAN, WAN, Internet, etc.) N to the auxiliary memory Device 43, etc., and load the program or data into the main memory device 41 as required, and follow the program loaded into the main memory device 41 to drive various resources by the CPU 40 to perform predetermined control and calculation processing, and process it The result (processed data) is output to the output device 44 through the bus 47 and displayed. At the same time, the data is appropriately memorized and saved (updated) in the database formed by the auxiliary memory device 43 as needed.
Next, an example of a facial image detection method using the facial image detection system 100 with such a configuration will be described.
Figure 3 is a flowchart of an example of a face image detection method that is actually used as a detection target image. However, before the actual use of the detection target image to perform recognition, it must first go through the above-mentioned recognition method. The obtained SVM30 is used as the step of learning the facial image and non-face image of the sample image.
The learning step is the same as before, generating a feature vector of each facial image and non-face image as a sample image for learning, and inputting whether the feature vector is a facial image or a non-face image at the same time. In addition, the learning image used for learning here is ideally an image that has been processed in the same field as the actual detection target image. That is, as will be described in detail later, the image area as the detection target of the present invention is dimensionally compressed. Therefore, by using images compressed to the same dimensionality in advance, recognition can be performed at a higher speed and with high accuracy.
Then, if learning the feature vector of the sample image is performed on the SVM 30 in this way, as shown in step S101 of FIG. 3, firstly determine (select) the area of the detection target image as the detection target. In addition, the method for determining the detection target field is not particularly limited. The field obtained by other facial image recognition units can be used directly, or the user of this system can arbitrarily specify in the detection target image. Domain, in principle, the detection target image, of course, does not know where the face image is contained, and it is almost impossible to know whether it contains the face image. Therefore, for example, take the upper left corner of the detection target image as the starting point. Starting from a certain field, one by one horizontally and vertically shifting a certain pixel to scan all the fields one by one to choose this field is ideal. In addition, the size of the field does not have to be fixed, and it can also be selected while changing the size appropriately.
After that, if the initial area to be the detection target of the face image is selected in this way, as shown in FIG. 3, move to the next step S103 to normalize the size of the initial detection object area (resize, change) Size) into a predetermined size, for example, 24×24 pixels. That is, in principle, of course, it is not known whether the image to be detected contains a face image, and even its size is unknown. Therefore, the number of pixels will vary with the size of the face image in the selected area. There is a big difference. In short, the selected field is first normalized (resize) to the size of the benchmark (24×24 pixels).
Next, if the selected area has been normalized in this way, it proceeds to the next step S105 to obtain the edge intensity of the normalized area for each pixel, and then divide the area into plural blocks to calculate the area within each block. The average or dispersion value of the edge.
Fig. 4 is a graph (image) of the change in edge intensity after normalization in this way, and the calculated edge intensity is displayed in a format of 24×24 pixels. In addition, Fig. 5 is the area that is re-blocked into 6×8, and the average value of the edges in each block is displayed as the representative value of each block. Then, Fig. 6 is the same. The area is re-blocked into 6×8, and the dispersion value of the edge in each block is displayed as the representative value of each block. In addition, the edges at both ends of the upper part of the picture are the "two eyes" of the face, the edges of the middle part of the picture are the "nose", and the edges of the lower part of the middle part of the picture are the "lips" of the face. share". It can be seen that even after the dimensional compression caused by the present invention, the features of the face image will still remain directly.
Here, as the number of partitions in the field, it is important to block the feature quantity of the image based on the self-correlation coefficient to the extent that it does not significantly damage its feature quantity. If the number of blocks is too large, the number of calculated image feature vectors will also increase, which increases the processing load, and cannot achieve high-speed detection. That is, if the autocorrelation coefficient is above the threshold, it can be thought of as the value of the image feature quantity in the block, or the variation pattern is converged within a certain range.
The calculation method of this self-correlation coefficient can be easily obtained by the following formula 3 and formula 4. Equation 3 is used to calculate the self-correlation coefficient in the horizontal (width) direction (H) for the detected object image, and Equation 4 is used to calculate the self-correlation coefficient in the vertical (high) direction (V) of the detected object image. The formula of the correlation coefficient.
<maths><img file="TW200529093A_D0002.tif" /></maths>r: correlation coefficient e: brightness or edge strength Width: number of pixels in the horizontal direction i: pixel position in the horizontal direction j: pixel position in the vertical direction dx: distance between pixels<maths><img file="TW200529093A_D0003.tif" /></maths>v: correlation coefficient e: brightness or edge strength height: number of pixels in the horizontal direction i: pixel position in the horizontal direction j: pixel position in the vertical direction dy: distance between pixels
Then, FIGS. 7 and 8 are examples of the correlation coefficients in the horizontal direction (H) and vertical direction (V) of the image obtained using Equation 3 and Equation 4 above.
As shown in Figure 7, relative to the image used as the reference, the stagger of one of the images is "0" in the horizontal direction, that is, when the two images are completely overlapped, the correlation between the two images is the largest "1.0" "; But if one of the images is offset by "1" pixels in the horizontal direction relative to the image used as the reference, the correlation between the two images will become about "0.9", and if the image is offset by "2" For pixels, the correlation between the two images will become approximately "0.75". In this way, the correlation between the two images will gradually decrease as the amount of shift (number of pixels) relative to the horizontal direction increases.
Also, as shown in Figure 8, with respect to the image used as the reference, the stagger of one of the images is "0" in the vertical direction, that is, the correlation between the two images is the largest when the two images are completely overlapped. "1.0"; but if one of the images is staggered by "1" in the vertical direction relative to the image used as the reference, the correlation between the two images will become about "0.8", and if the image is staggered " 2" pixels, the correlation between the two images will become about "0.65". In this way, the correlation between the two images will gradually decrease as the amount of shift (number of pixels) relative to the vertical direction increases.
As a result, when the amount of misalignment is relatively small, that is, within a certain number of pixels, the image feature amounts between the two images are not much different, and it can be thought that they are almost the same.
It can be assumed that the value of the image feature value or the variation pattern is within a certain range (below the threshold). Although it varies with the detection speed or the reliability of the detection, in this embodiment, it is assumed to be As shown by the arrows in the figure, it is up to "4" pixels in the horizontal direction and up to "3" pixels in the vertical direction. That is, as long as it is an image with a shift amount within this range, the change in the image feature amount is small, and the operation can be performed as if the variation range is within a certain range. As a result, in this embodiment, it is possible to perform dimensional compression to 1/12 (6×8=48 dimensions/24×24=576 dimensions) without greatly impairing the characteristics of the original selection area.
The present invention is based on the fact that this image feature quantity has a certain range. The self-correlation coefficient is not reduced to a certain value range as a block to operate, and the representative value in the block is used. It is composed of image feature vectors.
Then, if the area as the detection target is dimensionally compressed in this way, after calculating the image feature vector constituted by the representative values of each block, the obtained image feature vector is input to the recognizer (SVM) 30 to determine the appropriate Whether there is a face image in the area (step S109).
After that, the determination result can be displayed to the user every time the determination is completed, or together with other determination results, and then proceed to the next step S110 until the determination processing is completed in all areas and the processing ends.
That is, the example of FIG. 4 to FIG. 6, such that each block based autocorrelation coefficient of not less than a certain value or less, by the crossbar points respectively adjacent pixels 12 (3 × 4) formed by the 12 The average value (Figure 5) and dispersion value (Figure 6) of the image feature quantity (edge intensity) of each pixel are calculated as the representative value of each block, and the image feature vector obtained from the representative value is input to the recognition The device (SVM) 30 performs determination processing.
In this way, the present invention does not directly use the feature quantity of all pixels in the detection target area, but first performs sub-dimension compression to the extent that the original feature quantity of the image is not damaged, and then recognizes it, so the calculation can be greatly reduced. It can identify whether there are facial images in the selected area with high speed and high accuracy.
In addition, in this embodiment, although the image feature value based on the edge intensity is used, depending on the type of image, the brightness value of the pixel may be used to perform dimensional compression more efficiently than the edge intensity. Therefore, in this case It can be the image feature value that can be used solely by the brightness value, or combined with the edge intensity.
Furthermore, in the present invention, the "human face", which is extremely useful in the future as the detection target image, is the object, but it is not the "human face", "human body shape" or "animal face and posture", "cars, etc." Any other objects such as "vehicles", "buildings", "plants", and "terrain" are applicable.
In addition, FIG. 9 shows the "Sobel operator" which is one of the differential edge detection operators that can be used in the present invention. The operator (filter) shown in Fig. 9(a) surrounds the 8 pixel values of the pixel of interest, and adjusts each of the 3 pixel values in the left and right columns to emphasize the horizontal Edge; Figure 9(b) shows the operator, which surrounds the 8 pixel values of the pixel of interest, and adjusts each of the 3 pixel values at the upper and lower positions to emphasize the edge in the vertical direction; Thus, the vertical and horizontal edges are detected.
Then, after the square sum of the result generated by such an operator, the edge intensity can be obtained by taking its square root. By generating the edge intensity or edge dispersion value in each pixel, it can be accurately The image feature vector is detected. In addition, as mentioned above, the "Sobel operator" can also be replaced with other differential edge detection operators such as "Roberts" or "Prewitt", template-type edge detection operators, etc., to apply.
Moreover, it is also possible to replace the SVM and use a neural network as the pre-mark recognizer 30, which can also implement high-speed and high-precision recognition.
<p>10Image Reading Unit</p><p>20Eigenvector calculation section</p><p>22Brightness calculation unit</p><p>24Edge calculation section</p><p>26Average. Dispersion calculation unit</p><p>30SVM (Support Vector Machine)</p><p>100Face image detection system</p><p>40CPU</p><p>41RAM</p><p>42ROM</p><p>43Auxiliary Memory Device</p><p>44Output device</p><p>45Input device</p><p>46Input/output interface (I/F)</p><p>47Bus</p>
[Figure 1] A block diagram of an implementation form of a facial image detection system.
[Figure 2] The hardware configuration diagram of the facial image detection system.
[Figure 3] A flowchart of an implementation form of a face image detection method.
[Figure 4] Graphical representation of the change in edge intensity.
[Figure 5] Graphical representation of the average value of edge intensity.
[Figure 6] Graphical representation of the dispersion value of edge strength.
[Figure 7] A diagram showing the relationship between the amount of shift in the horizontal direction of the image and the correlation coefficient.
[Figure 8] A diagram showing the relationship between the amount of shift in the vertical direction of the image and the correlation coefficient.
[Figure 9] An illustration of the shape of the Sobel filter.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8463049B2 | Cited by | United States of America | Applicant |
| TWI452540B | Cited by | Taiwan Province of China | Examiner |
| TWI407800B | Cited by | Taiwan Province of China | Examiner |
| US9058744B2 | Cited by | United States of America | Applicant |
5 members in 4 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2003434177 | Japan | – | |
| 2003434177 | Japan | A |
Members5
| Document | Office | Kind | |
|---|---|---|---|
| US2005139782A1 | United States of America | A1 | |
| JP2005190400A | Japan | A | |
| WO2005064540A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200529093AThis record | Taiwan Province of China | A | |
| TWI254891B | Taiwan Province of China | B |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Annulment or lapse of patent due to non-payment of feesLapsedMM4A | MM4A |
Numbers
- Publication
- 200529093
- Application
- 93140626
Titles4
- Chinese
- 臉部影像偵測方法及臉部影像偵測系統以及記錄有臉部影像偵測程式之電腦可讀取之媒體
- English
- Face image detection method and face image detection system, and computer readable media recorded with face image detection program
- Unlabeled
- 臉部影像偵測方法及臉部影像偵測系統以及記錄有臉部影像偵測程式之電腦可讀取之媒體
- Unlabeled
- Face image detection method and face image detection system, and computer readable media recorded with face image detection program
Classification
- CPC, 1
- G06V40/161
- IPC, 4
- G06T1 00
- G01J1 58
- G06K9 00
- G06T7 00