A database search system and method for resemblance comparison of image data
Abstract
The present invention proposes a system control method for searching based on the similarity of image information content. It uses eigenvalue characteristics to find obvious common features, and generates new feature vectors and feature weights as the basis for retrieval. It can also be directly included in the new feature group, which can effectively improve the accuracy of the search without losing the stability of the search, and this method has the characteristics of aggregation, which can be applied to most search systems.
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
20 claims: 20 independent, 0 dependent
- 1一種自多媒體資料庫中搜尋至少一種或更多相似影像資訊內容之的檢索方法,包含下列步驟:建構至少一種詢例影像;擷取顯著且共同的特徵自該種詢例影像中;產生至少一調整特徵權數自該根據顯著且共同之的特徵中;產生至少一種新的檢索特徵向量和特徵權值自該調整特徵權數中;以及產生至少一種新的詢例影像自該新的特徵向量中。
- 2如專利申請範圍第1項所述之影像資訊內容檢索方法,其中該擷取顯著且共同特徵之向量的步驟,包含:擷取每一詢例影像的特徵值向量;以及衡量並儲存該特徵值的分佈特性和特徵值在詢例間的均值和標準差。
- 3如專利申請範圍第2項所述之影像資訊內容檢索方法,其中該特徵在詢例間之均值和標準差值可分別表示為:
- 4如專利申請範圍第2項所述之影像資訊內容檢索方法,其中該特徵值之分佈特性可表示為:
- 5如專利申請範圍第2項所述之影像資訊內容檢索方法,其中該特徵在詢例間是否具明顯的共同性,可以下方式衡量:
- 6如專利申請範圍第1項所述之影像資訊內容檢索方法,其中該調整特徵權值可以下方式衡量;
- 7一種自多媒體資料庫中搜尋至少一種或更多相似影像資訊內容之檢索方法,包含下列步驟:由使用者輸入一或多張詢例影像;擷取詢例影像之的特徵向量,自該使用者輸入一或多張詢例初始影像中;擷取顯著之共同特徵自該的特徵向量中;產生調整之特徵權值自該顯著且共同之的特徵中;產生至少一組新的特徵向量自該和調整之特徵權值中;以及產生至少一組詢例影像自該新的特徵向量中。
- 8如專利申請範圍第7項所述之影像資訊內容檢索方法,其中該擷取特徵向量係儲存在一種或更多之特徵資料庫中。
- 9如專利申請範圍第7項所述之影像資訊內容檢索方法,其中該擷取顯著且共同的特徵,之包含下列步驟:擷取每一詢例影像之的特徵向量;以及衡量並儲存每一特徵分佈之特性,特徵值之的均值和標準差自該特徵向量中。
- 10如專利申請範圍第9項所述之影像資訊內容檢索方法,其中該特徵值的均值和標準差可分別衡量如下:
- 11如專利申請範圍第9項所述之影像資訊內容檢索方法,其中該使用的分佈特性可以如下方式衡量:
- 12如專利申請範圍第9項所述之影像資訊內容檢索方法,其中該顯著的共同特徵之選取可以如下方式衡量:
- 13如專利申請範圍第7項所述之影像資訊內容檢索方法,其中該調整特徵權數的方式,係依據:
- 14一種自多媒體資料庫中搜尋至少一種或更多相似影像資訊內容之檢索方法,包含下列步驟:選取一或多張詢例影像自該從多媒體資料庫中;取得相關的特徵向量自該初始之複數個詢例影像中;選取顯著的共通特徵自該從詢例樣本之的特徵向量中;產生至少一組新的調整特徵權值自該顯著之共同特徵中;產生至少一組新的特徵向量值自該新的調整特徵權值中;以及產生至少一組新的詢例影像自該新的特徵向量中。
- 15如專利申請範圍第14項所述之影像資訊內容檢索方法,其中該選取顯著之共同特徵向量,包含下列步驟:擷取每一詢例樣本的特徵向量;以及衡量並儲存特徵值在詢例樣本間的分佈特性,及特徵值在詢例間的均值和標準差。
- 16如專利申請範圍第15項所述之影像資訊內容檢索方法,其中該均值和標準差可以如下方式衡量;
- 17如專利申請範圍第15項所述之影像資訊內容檢索方法,其中該分佈特性可以如下方式衡量;
- 18如專利申請範圍第15項所述之影像資訊內容檢索方法,其中該,選取顯著之共同特徵可以如下方式衡量:
- 19如專利申請範圍第15項所述之影像資訊內容檢索方法,其中該調整特徵權數的方法可以如下方式衡量:
- 20一種自多媒體資料庫中搜尋至少一種或更多相似影像資訊內容之檢索方法,包含:至少一個多媒體資料庫,以儲存詢例樣本之影像;一或多個特徵擷取單元,以擷取顯著且共同的特徵向量自該詢例影像中;一或多個特徵資料庫,以儲存該顯著共同之特徵向量;以及一或多個匯歸特徵檢索單元,以產生調整權值自該顯著之共同特徵向量,產生新的特徵向量自該和特徵權值,並以此產生新的詢例影像自該新的特徵特徵向量中。
Independent claims20
59 paragraphs, as filed
Database retrieval system and method for image information similarity comparison
<p>10. . . Multimedia Database</p><p>12. . . Feature extraction unit</p><p>14. . . Feature extraction unit</p><p>16. . . Feature database</p><p>18. . . Feature database</p><p>20. . . Online editing interface</p><p>twenty two. . . Universal Search Interface Unit</p><p>twenty four. . . General search interface sub-unit</p><p>26. . . General search interface sub-unit</p><p>42. . . Retrieve the query image</p><p>43. . . Positive correlation image</p><p>44. . . Similar to the sample of the inquiry image</p><p>45. . . Similar to the sample of the inquiry image</p><p>46~48. . . Similar to the sample of the inquiry image</p><p>49. . . Similar to the sample of the inquiry image</p><p>50~51. . . Similar to the sample of the inquiry image</p>
FIG. 1 is a schematic diagram of a database system method for searching based on the similarity of image information content in a preferred embodiment of the present invention.
Fig. 2 is a schematic diagram of the general retrieval method of the present invention related to the retrieving feature.
FIG. 3 is a schematic diagram illustrating whether a certain feature is significantly similar between the sample inquiries in a preferred embodiment of the present invention.
Figures 4 and 5 show the search results used in the image database in the preferred embodiment of the present invention.
Field of invention
The present invention is suitable for a retrieval system and method that uses the similarity of the image information content as a feature comparison. More specifically, it can be used in retrieval systems and methods that use multiple queries or relevance feedback.
Background of the invention
Multimedia compression methods, such as JPEG, MPEC-1, MPEG-2, MPEG-4, etc., have been formulated as international standards, and then mass-produced and circulated with the popularization of the Internet. So began to formulate MPEG-7 as a standard description method for multimedia retrieval. In MPEC-7, the description of the shape in the two-dimensional image is divided into two description methods: the outer contour of the shape and the shape block. Regarding the description of the outer contour of the shape, transforming the coordinate points of the contour into Fourier coefficients can provide feature coefficients that have nothing to do with the shape size, displacement and rotation angle as reference values for retrieval. However, using Fourier coefficients to describe the characteristics of a shape is susceptible to interference by noise, and is only suitable for describing the outer contour of a shape. The description method of the comparison summary feature should show the characteristics of the shape block on the feature value. Regarding such descriptions, zernike and pseudo-zernike moments can provide more effective block descriptions as parameters for similarity retrieval. Other features that describe the shape, such as the directionality of the edge points, the constant-width ratio of the shape, and the content complexity, are captured according to the application and content. Generally, statistical characteristics are used as the description and comparison of such characteristics.
The shape description index can also be extracted from a database that has been classified by human vision. It is usually used to describe the probability distribution characteristics of indicators to find visually obvious features as an important reference for retrieval. In addition, there are also multiple-resolution description indicators, such as the mean and standard deviation of wavelet sub-bands, as reference values for describing textures and shapes. Similar feature description indicators can be retrieved according to different application requirements, but if the features retrieved according to a specific application are used in another application, good retrieval results may not necessarily be obtained. In short, the previous technology can only effectively deal with the retrieval of multiple feature groups. The general retrieval method of one integrated feature is developed to handle the retrieval of multiple feature groups. The present invention is to develop a universal retrieval method of one integrated feature to overcome the inability of the previous technology. The shortcomings of significant common features are effectively selected among multiple feature groups.
Summary of the invention
The main purpose of the present invention is to provide a database system and method based on the similarity of media information content as a search basis, which leads the search by finding obvious common features among the query images in different types of feature combinations The similarity is estimated to meet the subjective definition of similarity.
Another object of the present invention is to provide a method for updating the search feature vector in a multi-routine query system.
Another object of the present invention is to provide a database system that uses a combination of multiple feature descriptions and a method for retrieving aggregated features, which can appropriately find effective feature vectors among multiple feature combinations to reflect users perceptions of similarity. The setting.
Another object of the present invention is to provide an effective retrieval method for a database system that uses content similarity to retrieve. By finding common and obvious feature vectors among routines, an effective retrieval method that meets user-defined similarity can be performed. Retrieval, the result of a general search is a sample in the database that is more similar to the routine image content. The search result is relatively similar rather than absolutely similar to the routine image. Therefore, a more reasonable approach is to calculate the probability distribution function of the feature value to find the statistically similar and obvious feature vectors among the conventional sample images, so as to lead the new retrieval process to effectively meet the users' perception of similarity. After the feature combination of is defined, this control method can effectively use the feature value to improve the retrieval accuracy and is not interfered by the inclusion of multiple feature combinations.
In a database system that uses content similarity for retrieval, different users have different definitions of similarity, and there is no set of universal feature description groups that can meet all application requirements. Usually the system provides an interface with multiple queries. , Or the user will return relevant information so that the system can recognize the users definition of similarity and search for appropriate samples. In this multi-inquiry interface environment, it is the simplest method to take the feature mean value of the routine sample as the new retrieval feature vector. Although this feature value averaging method is simple and has certain effects, it cannot effectively reflect the user's definition of similarity in the search results. Therefore, we propose to use similar and obvious feature values as the main measurement value for further search. This method can not only effectively cover new feature groups, but also more accurately meet the user's needs for similarity retrieval.
The key points and benefits of the present invention will be described in the following and shown in the drawings more clearly.
Detailed description of the invention
The following related drawings will be used to illustrate the preferred embodiments of the present invention in detail. Fig. 1 is a retrieval system according to a preferred embodiment of the present invention. In the first step, the image of the database is first extracted by the feature extraction units 12 and 14 to store the feature data in the image. In libraries 16 and 18. The second step is similarity retrieval. The user can customize the initial inquiry image via the online editing interface 20, or select one or several images in the database sample as the inquiry example. If the new image 20 is edited online, the image feature must be extracted through the feature extraction units 12 and 14 before the database retrieval, and then the feature database 16 and 18 are compared for similarity. If a sample is selected from the database as an inquiry example, the feature vector corresponding to the image sample is taken out directly from the feature database and used as the standard vector for retrieval. This standard feature vector is sent to the universal retrieval interface unit 22 (UQM) for collecting features. Among them, the general retrieval subunits 24 and 26, as shown in FIG. 2, use the feature vector of the query sample to generate a new retrieval vector. Then send it to the feature database for a new comparison.
Because there are usually multiple sets of characteristic parameters for the search unit to use in the search of the database. Or users will add new feature groups according to specific needs. The general retrieval system of the integrated features proposed in the present invention is to find the obvious common features among the feature groups as the main feature for further retrieval. Its main feature is that the new feature group can be directly embedded in the retrieval system for use by the system to improve the accuracy of retrieval, and it is not interfered by irrelevant feature groups to affect retrieval accuracy, and it can take into account both accuracy and stability.
Suppose there are a total of n sets of features, then the features of each image can be expressed as <img file="TW511010B_D0001.tif" /><sub>s</sub> =( <img file="TW511010B_D0002.tif" /><sub>1</sub> , <img file="TW511010B_D0003.tif" /><sub>2</sub> ,..., <img file="TW511010B_D0004.tif" /><sub>n</sub> ), and each element can be expressed as <img file="TW511010B_D0005.tif" /><img file="TW511010B_D0006.tif" /> =(f <sub>i1</sub> ,f <sub>i2</sub> ,..., <img file="TW511010B_D0007.tif" /> ). Assumption = <img file="TW511010B_D0008.tif" /><sub>1</sub> ={iji=1,...,n,j=l....,N <sub>i</sub> }, the characteristic value can be expressed as <img file="TW511010B_D0009.tif" /><sub>s</sub> ={f <sub>k</sub><img file="TW511010B_D0010.tif" /><img file="TW511010B_D0011.tif" /><sub>1</sub> }. If the i-th feature f <sub>i</sub> The probability step function is p(f <sub>i</sub> ), then the probability function is
<maths><img file="TW511010B_D0012.tif" /></maths>
In the search interface that provides multiple enquiries or similarity feedback, search engines
In the user's query example, select the common feature as the search condition. We achieve this goal by adjusting the weight of the new feature vector. Assuming that there are q enquiries, the relative characteristic value is <img file="TW511010B_D0013.tif" /><sub>s</sub> ,s=1,...,q. The mean and standard deviation between each eigenvalue of these eigenvectors are
<maths><img file="TW511010B_D0014.tif" /></maths>
In the case of positive and negative proofs, in order to evaluate the statistical similarity, we define the feature convergence value F in the probability space <sub>C</sub> (.) and the divergence value F <sub>D</sub> (.)as follows:
<maths><img file="TW511010B_D0015.tif" /></maths>
When there is no negative proof, then F <sub>D</sub> (.)=1 The way to adjust the weight of the i-th feature is as follows:
<maths><img file="TW511010B_D0016.tif" /></maths>
If the distance between the j-th eigenvalues in the statistical space of each query is similar, the corresponding F <sub>C</sub> (.) The value must be smaller, so the weight is larger, and the importance of this feature in the new search is higher. When a new feature group is generated <sub>,</sub> The search is still dominated by more similar features in the statistical space. Therefore, integrating new feature groups or eliminating improper feature groups will not have much impact on the original search system configuration, because only important and similar feature groups will affect the search results.
In a system that provides positive and negative feedback, the retrieval unit emphasizes the commonality between positive samples and filters out the characteristics of negative samples. Because F <sub>D</sub> (m <sub>1</sub> ,σ <sub>1</sub> ,m <sub>2</sub> ,σ <sub>2</sub> ) Is used to estimate the degree of difference between positive and negative features. Usually the eigenvalues of negative samples will not converge, so we set F <sub>c</sub> ( <img file="TW511010B_D0017.tif" /> , <img file="TW511010B_D0018.tif" /> )=1. In addition, in order to make obvious common features stand out in the retrieval section, when the ratio in equation (6) is greater than a certain value, set w <sub>j</sub> =1。
The control process of the present invention can be summarized as follows: (1) The feature vector of each sample in the database ( <img file="TW511010B_D0019.tif" /><sub>1</sub> , <img file="TW511010B_D0020.tif" /><sub>2</sub> ,..., <img file="TW511010B_D0021.tif" /><sub>N</sub> ) Extract and place it in the feature database; (2) Measure and calculate the characteristics of each feature value among all samples in the database, such as the mean and standard deviation {(m <sub>i</sub> , Σ <sub>i</sub> )﹜ <sub>i</sub> =1,..., <sub>N</sub> Or {p(f <sub>ij</sub> =1,...,M} <sub>j</sub> = <sub>1,...,N</sub> (3) In each search, measure and calculate the mean value (σ) and standard deviation (m) of the average value of each feature value among the enquiry samples; (4) Evaluate whether each feature has obvious commonality, as follows Column method evaluation: F <sub>c</sub> (m,σ)=P(mtenσ)-P(m-σ) and use this value to generate a new feature weight: if(F <sub>c</sub> <P <sub>T</sub> )ω <sub>i</sub> =1 else ω <sub>i</sub> =0; (5) Perform a new similarity search with a new search vector and a new feature value.
In the preferred embodiment of the present invention, the statistical characteristics of the feature values in the development database are used for similarity retrieval. If the feature values are very different, but are very similar in the statistical space, the retrieved images may not The similarity is too in line with subjective cognition. This problem may occur when there are too few samples similar to the retrieved image in the database. Therefore, when the variation value of a certain feature value between the inquiry samples is too large, the feature weight is set to 0, so that it will not have an effect and affect the search results. Therefore, if it is <img file="TW511010B_D0022.tif" /> > <img file="TW511010B_D0023.tif" /> ,then w <sub>i=</sub> 0, <img file="TW511010B_D0024.tif" /> and <img file="TW511010B_D0025.tif" /> It is the standard deviation of the i-th feature between the enquiry samples and all the samples in the database.
Figures 4 and 5 show the results of some similarity searches. There are 30,000 trademark images in the database, which include text totems, animals, general geometric drawings, and combined images and texts. These images are first manually cut in the file. To find out where the image is, the pre-processing system first calculates the minimum circumscribed circle of each trademark graphic. For this type of graphic, the ZM and PZM moments are used as the characteristic value to measure. In the case of a level of 10. The number of eigenvalues is ZM=36 and PZM=66, respectively.
Fig. 4 is a search result of a search query example image 42. The left side is the query example image and the right side is the searched similar images. The similarity is from left to right, and from top to bottom. Select image 43 as a positively correlated image for a new search. If the mean value of the feature vectors of the two samples is used as the new search vector, the search result is shown in Figure 4B. It can be seen from the figure that there are two visually similar New samples from the searched images appear in the search results, such as images 44 and 45. FIG. 4C is a similar image searched by the retrieval method in the preferred embodiment of the present invention. It can be seen that there are three more samples 46, 47, and 48 that are visually similar to the query image. The control method in the preferred embodiment of the present invention can indeed reflect the users definition of similarity in the search results. Another search result is shown in Figure 5 with three sample query samples. Figure 5A uses the present invention. Compared with FIG. 5B, the search results of the control method in the preferred embodiment have more samples similar to the inquiry image, such as 50 and 51.
However, the above are only preferred embodiments of the present invention, and should not be used to limit the scope of implementation of the present invention. That is to say, all equal changes and modifications made in accordance with the scope of the patent application of the present invention should still fall within the scope of the patent of the present invention.
Schematic description
FIG. 1 is a schematic diagram of a database system method for searching based on the similarity of image information content in a preferred embodiment of the present invention.
Fig. 2 is a schematic diagram of the general retrieval method of the present invention related to the retrieving feature.
FIG. 3 is a schematic diagram illustrating whether a certain feature is significantly similar between the sample inquiries in a preferred embodiment of the present invention.
Figures 4 and 5 show the search results used in the image database in the preferred embodiment of the present invention.
Symbol description of main components
10. . . Multimedia Database
12. . . Feature extraction unit
14. . . Feature extraction unit
16. . . Feature database
18. . . Feature database
20. . . Online editing interface
twenty two. . . Universal Search Interface Unit
twenty four. . . General search interface sub-unit
26. . . General search interface sub-unit
42. . . Retrieve the query image
43. . . Positive correlation image
44. . . Similar to the sample of the inquiry image
45. . . Similar to the sample of the inquiry image
46~48. . . Similar to the sample of the inquiry image
49. . . Similar to the sample of the inquiry image
50~51. . . Similar to the sample of the inquiry image
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| TWI396101B | Cited by | Taiwan Province of China | Examiner |
3 members in 2 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 09833818 | United States of America | – | |
| 83381801 | United States of America | A | |
| 20010833818 | – | – | – |
| US20010833818 | – | – | – |
Members3
| Document | Office | Kind | |
|---|---|---|---|
| TW511010BThis record | Taiwan Province of China | B | |
| US2002178149A1 | United States of America | A1 | |
| US6834288B2 | United States of America | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Annulment or lapse of patent due to non-payment of feesLapsedMM4A | MM4A | |
| Issue of patent certificate for granted invention patentGrantedGD4A | GD4A |
Numbers
- Publication
- 511010
- Publication, DOCDB
- 511010
- Publication, EPODOC
- TW511010B
- Application
- 90123293
- Application, DOCDB
- 90123293
- Application, EPODOC
- TW20010123293
Titles5
- Chinese
- 一種影像資訊相似性比對之資料庫檢索系統與方法
- English
- Database retrieval system and method for image information similarity comparison
- English
- A database search system and method for resemblance comparison of image data
- Unlabeled
- 一種影像資訊相似性比對之資料庫檢索系統與方法
- Unlabeled
- Database retrieval system and method for image information similarity comparison
Classification
- CPC, 6
- G06F17/30259
- G06F16/5854
- Y10S707/99932
- Y10S707/99933
- Y10S707/99945
- Y10S707/99948
- IPC, 1
- G06F17 30