Method, device and program for acquiring synonymous vocabulary
Abstract
Problem to be solved.To automatically acquire a synonym of a vocabulary of a specific class from hypertext. According to the present invention, when a keyword and a higher-level conceptual word of the keyword are input, a highly relevant document including the keyword is acquired, and the link destination of the link in which the keyword is included in the anchor text in the document. Identify the document, identify the link that is the link to the document and have the anchor text that contains the keyword, extract the reference string contained in the link to the identified specific document, and put the keyword and the keyword of the same class The anchor text of the link to the referenced document set is analyzed by the link as the anchor text, the frequency of occurrence of the substring contained in the anchor text is calculated, and the frequency of occurrence is used to obtain the anchor text. Analyze and remove common substrings in the anchor text. [Selection diagram] Fig. 1

Term
Term ended
Projected expiry passed 15 December 2025, 0.8 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
7 claims: 3 independent, 4 dependent
- 1HTMLやXMLを含む電子テキストを解析し、固有名詞の別名や略称を含む同義語彙を獲得する同義語彙獲得方法であって、 キーワード検索手段が、キーワードと該キーワードの上位概念語が入力されると、該上位概念語とそれぞれのキーワードを含み関連性が高い文書を取得するキーワード検索ステップと、 文書特定手段が、前記キーワード検索ステップで取得した文書中で、前記キーワードがアンカーテキスト中に含まれるリンクのリンク先の文書を特定する文書特定ステップと、 リンク検索手段が、前記キーワード検索取得ステップで取得した前記文書に対するリンクであり、前記キーワードをアンカーテキストに含むリンクを特定するリンク検索ステップと、 アンカーテキスト特定手段が、前記リンク検索ステップで特定された文書へのリンクに含まれる参照文字列を抽出するアンカーテキスト特定ステップと、 クラス別アンカーテキスト文字列統計解析手段が、前記キーワードと同じクラスのキーワードをアンカーテキストとするリンクによって、参照されている文書集合へのリンクのアンカーテキストを解析し、該アンカーテキスト中に含まれる部分文字列の出現頻度を算出し、アンカーテキスト解析結果として記憶手段に格納するクラス別アンカーテキスト文字列統計解析ステップと、 アンカー文字列クリーニング手段が、前記記憶手段に格納された前記部分文字列の出現頻度に基づいて、前記アンカーテキストを解析し、該アンカーテキストの中で一般的な部分文字列を除去するアンカー文字列クリーニングステップと、を行うことを特徴とする同義語彙獲得方法。
- 2前記キーワード検索ステップにおいて、 入力された前記キーワードと前記上位概念語に基づき、該上位概念語が表す概念を含む、つまり、該上位概念語そのものに限らず、該上位概念語の同義語や同様の意味を表現する語彙の集合が含まれている文書で、かつ、キーワードを含み関連性が高い文書を取得する、請求項1記載の同義語彙獲得方法。
- 3アンカー文字列クリーニングステップにおいて、 前記アンカーテキストの解析をサイト毎に解析を行う、請求項1記載の同義語彙獲得方法。
- 4HTMLやXMLを含む電子テキストを解析し、固有名詞の別名や略称を含む同義語彙を獲得する同義語彙獲得装置であって、 キーワードと該キーワードの上位概念語が入力されると、該上位概念語とそれぞれのキーワードを含み関連性が高い文書を取得するキーワード検索手段と、 前記キーワード検索手段で取得した文書中で、前記キーワードがアンカーテキスト中に含まれるリンクのリンク先の文書を特定する文書特定手段と、 前記キーワード検索手段で取得した前記文書に対するリンクであり、前記キーワードをアンカーテキストに含むリンクを特定するリンク検索手段と、 前記リンク検索手段で特定された文書へのリンクに含まれる参照文字列を抽出するアンカーテキスト特定手段と、 前記キーワードと同じクラスのキーワードをアンカーテキストとするリンクによって、参照されている文書集合へのリンクのアンカーテキストを解析し、該アンカーテキスト中に含まれる部分文字列の出現頻度を算出し、アンカーテキスト解析結果として記憶手段に格納するクラス別アンカーテキスト文字列統計解析手段と、 前記記憶手段に格納された前記部分文字列の出現頻度に基づいて、前記アンカーテキストを解析し、該アンカーテキストの中で一般的な部分文字列を除去するアンカー文字列クリーニング手段と、を有することを特徴とする同義語彙獲得装置。
- 5前記キーワード検索手段は、 入力された前記キーワードと前記上位概念語に基づき、該上位概念語が表す概念を含む、つまり、該上位概念語そのものに限らず、該上位概念語の同義語や同様の意味を表現する語彙の集合が含まれている文書で、かつ、キーワードを含み関連性が高い文書を取得する手段を含む請求項4記載の同義語彙獲得装置。
- 6アンカー文字列クリーニング手段は、 前記アンカーテキストの解析をサイト毎に解析を行う手段を含む請求項4記載の同義語彙獲得装置。
- 7コンピュータを、 請求項4乃至5記載の同義語彙獲得装置として機能させることを特徴とする同義語彙獲得プログラム。
Independent claims7
40 paragraphs, as filed
The present invention relates to a synonymous vocabulary acquisition method, device, and program, and particularly, a synonymous vocabulary acquisition method, device, and device for acquiring vocabulary from tagged texts such as HTML, XML, and SGML in a computer network represented by the Internet. Regarding the program.
In information retrieval in computer networks, the number of search results is often large, and users of search systems are forced to perform a search by keyword and then obtain the information they really want from the obtained search results. Has been done.
For such problems, we extract the proper nouns that point to real-world instances from the text information of the search results from the document, select the proper nouns that are considered to be important in the search results, and search the search results. There is a method of easily realizing an efficient document search by presenting with (for example, see Patent Document 1).
This allows users to view real-world instances as landmarks and narrow down the desired information without having to look at the search results one by one to find the desired information, think of additional keywords, and perform a re-search. it can.
As a basic technique to achieve this, a method for identifying proper nouns in texts is required.
The simplest method is to manually create a dictionary and extract words that match the dictionary words from the text.
Furthermore, we do not have a specific dictionary, but create extraction rules as morpheme (part of speech information) level patterns from learning data in which proper nouns existing in documents are manually specified in advance, and only words included in the learning data in advance. However, there is also a method that enables the extraction of new words (see, for example, Patent Document 2).<patcit num="1"><text>Japanese Unexamined Patent Publication No. 2005-208838</text></patcit><patcit num="2"><text>Japanese Unexamined Patent Publication No. 2003-331254</text></patcit>
<p> However, the above-mentioned conventional technique has the following problems.</p><p> The method of manually creating a dictionary can reliably identify the relevant part in the text, but the cost of updating the dictionary is very high, so it is realistic to collect dictionary words of a wide range of fields and attributes. Difficult to.</p><p> The method of automatically generating rules using learning data and analyzing proper nouns reduces the cost of dictionary construction and rule generation, and makes it possible to extract proper nouns, but it is frequent on the Web. The abbreviations and aliases found in are unaware.</p><p> In other words, it is possible to extract both when using a formal expression and when using an alias such as an abbreviation as a proper noun, but it is judged that the notation refers to the same instance. However, it is judged that the document describes different instances.</p><p> As a method to solve this, there is a method to solve it by using co-occurrence in a document (for example, Mukai et al., Examination of classification label integration method in label-oriented information retrieval, FIT2004), but this is a news article etc. It uses the method that the official name and abbreviation are listed in the formal document of, and is not effective for the informal document that exists a lot on the Web of Blog and bulletin board.</p><p> The present invention has been made in view of the above points, and provides a synonymous vocabulary acquisition method, device, and program capable of automatically acquiring synonyms of a specific class of vocabulary from a text in a computer network. With the goal.</p>
<p> FIG. 1 is a diagram for explaining the principle of the present invention.</p><p> The present invention (claim 1) is a synonymous vocabulary acquisition method for acquiring a synonymous vocabulary including an alias or abbreviation of a proper nomenclature by analyzing an electronic text including HTML or XML, and a keyword search means is a keyword and the keyword. When the higher-level conceptual word of is input (step 1), the keyword search step (step 2) for acquiring a highly relevant document containing the higher-level conceptual word and each keyword, and the document identification means are performed in the keyword search step. In the acquired document, the document identification step (step 3) that identifies the document to which the keyword is included in the anchor text, and the link to the document acquired by the link search method in the keyword search step (step 2). The link search step (step 4), which identifies the link that includes the keyword in the anchor text, and the anchor text identification means, the reference string contained in the link to the document identified in the link search step (step 4). Anchor text identification step to extract (step 5), Anchor text string by class Statistical analysis means analyzes the anchor text of the link to the referenced document set by the link whose anchor text is the keyword of the same class as the keyword, and the subcharacters contained in the anchor text. The class-specific anchor text character string statistical analysis step (step 6) that calculates the frequency of occurrence of columns and stores it in the storage means as the anchor text analysis result, and the anchor character string cleaning means are the substrings stored in the storage means. Anchor string cleaning step (step 7), which analyzes the anchor text based on the frequency of occurrence and removes a general substring in the anchor text, is performed.</p><p> Further, the present invention (claim 2) includes the concept represented by the superordinate concept word based on the input keyword and the superordinate concept word in the keyword search step of the claim 1, that is, is limited to the superordinate conceptual word itself. Instead, a document that includes a set of vocabulary that expresses a synonym for the superordinate conceptual word or a similar meaning, and that includes a keyword and has a high relevance is acquired.</p><p> Further, in the present invention (claim 3), in the anchor character string cleaning step of claim 1, the analysis of the anchor text is performed for each site.</p><p> FIG. 2 is a structural diagram of the principle of the present invention.</p><p> The present invention (claim 4) is a synonymous vocabulary acquisition device that analyzes an electronic text including HTML and XML and acquires a synonymous vocabulary including an alias or abbreviation of a proper nomenclature. When input, the keyword search means 20 for acquiring a highly relevant document containing the higher-level conceptual word and each keyword, and the link for which the keyword is included in the anchor text in the document acquired by the keyword search means 20. Included in the document identification means 25 that identifies the linked document, the link search means 30 that identifies the link that is a link to the document and includes the keyword in the anchor text, and the link to the document identified by the link search means 30. The anchor text of the link to the referenced document set is analyzed by the anchor text identification means 40 for extracting the reference character string and the link using the keyword of the same class as the keyword as the anchor text, and the anchor text is included in the anchor text. Anchor text character string statistical analysis means 50 by class, which calculates the frequency of occurrence of substrings and stores them in storage means 70 as anchor text analysis results, It has an anchor character string cleaning means 60 that analyzes the anchor text based on the appearance frequency of the substring stored in the storage means 70 and removes a general substring in the anchor text.</p><p> Further, the present invention (claim 5) includes the concept represented by the superordinate concept word based on the input keyword and the superordinate concept word in the keyword search means 20 of the claim 4, that is, the superordinate concept word itself. Not limited to this, it includes a means for acquiring a document that includes a set of vocabulary that expresses synonyms of the superordinate concept word or a similar meaning, and that includes a keyword and has a high relevance.</p><p> Further, the present invention (claim 6) includes means for analyzing the anchor text for each site in the anchor character string cleaning means 60 of claim 4.</p><p> The present invention (claim 7) is a synonymous vocabulary acquisition program that causes a computer to function as the synonymous vocabulary acquisition device according to claims 4 to 5.</p>
<p> As described above, according to the present invention, based on a keyword having few specific attributes, the position and pattern in which the keyword of the corresponding attribute appears are automatically extracted, and these two features are used to obtain a keyword candidate with high accuracy. Identify the rules that appear, extract the keywords that match the above specified rule in the text containing multiple keywords specified in advance, and finally extract based on the appearance frequency and distribution of each keyword extracted here. By identifying candidate keywords, it is possible to acquire vocabulary with high accuracy.</p><p> By using the dictionary obtained by acquiring this vocabulary, it is possible to extract keywords with specific attributes from the text.</p>
Hereinafter, embodiments of the present invention will be described with reference to the drawings.
The present invention accepts vocabularies of the same class (hyponymy and hyperconceptual words and keywords) as input, identifies Web page groups referenced by each vocabulary, and specifies anchor texts that refer to those Web page groups. By specifying the named entity part from this anchor text, the synonyms for each input vocabulary are specified. Since the appearance information of the character string in the anchor text, which is the basic data for estimating this last named entity part, is biased depending on the field, field-dependent analysis is performed by inputting multiple vocabularies of the same class. It makes it possible.
FIG. 3 shows the configuration of a synonymous vocabulary acquisition device according to an embodiment of the present invention.
The synonymous vocabulary acquisition device shown in the figure is a keyword input unit 10 in which a keyword and a higher-level conceptual word of the keyword are input, and a search that searches for documents highly related to them based on the input keyword and the higher-level conceptual word. Department (keyword search means) 20, document identification part (document identification means) 25 that finds a URL that matches the keyword, URL extraction part (link search means) 30 that asks for the URL of anchor text that matches the keyword, document identification part 25 Anchor text extraction unit (anchor text identification means) 40 that extracts anchor text from the URL document of URL extraction unit 30, and an anchor text analysis unit (anchor text by class) that obtains statistics (appearance frequency) of substrings of anchor text. Character string statistical analysis means) 50, Anchor text cleaning unit (anchor character string cleaning means) 60 that determines proper nomenclature by deleting unnecessary character strings, etc., Anchor text statistical information DB (storage means) that stores analysis results It is composed of 70 and a synonym output unit 80 that outputs the proper nomenclature specified by the anchor text cleaning unit 60 as a synonym vocabulary.
The description in parentheses above indicates the correspondence with each means in the scope of claims.
FIG. 4 is a flowchart of the operation according to the embodiment of the present invention. Hereinafter, the operation of the above configuration will be described with reference to FIG.
Step 101) The keyword input unit 10 accepts a keyword input from the user. For input, multiple keywords belonging to the same class shall be accepted, and one or more vocabularies belonging to the concept shall be input together with the superordinate concept word. An example of input data is shown in Fig. 5.
Step 102) The search unit 20 performs a keyword search for each of the input keywords, and acquires a document set highly related to the keyword from electronic text such as HTML or XML.
Step 103) The document identification unit 25 identifies a particularly relevant document from the document set acquired by the search unit 20 above for each of the input keywords. The following are examples of means for determining relevance.
As a result of searching in the search unit 20, the document with the strongest relevance to the keyword; The document whose title of the document exactly matches the keyword; Stores the document URL that is judged to be strong. After performing the above processing for all the keywords input to the device, the data contents are passed to the component to perform the next processing. Figure 6 shows an example of data storage in memory.
Step 104) For each of the input keywords, the URL extraction unit 30 extracts a link containing the keyword in the anchor text from the document set acquired by the search unit 20 above, and regards the document as having a high relevance to the keyword. In identifying this document, not all documents are necessarily acquired, and the threshold value is determined based on the frequency of appearance, and it is conceivable that only the documents that exceed the threshold value are regarded as documents highly related to the keyword. .. The URL extraction unit 30 has a memory and stores a keyword and the URL of a document strongly related to the keyword. After performing the above processing for all the input keywords, the data contents are passed to the component that performs the next processing. Figure 7 shows an example of data storage in the memory of the URL extraction unit 30.
Step 105) When all the keywords have been processed, the process proceeds to step 106, and if not, the process returns to step 102.
Step 106) The anchor text extraction unit 40 receives the URLs related to each of the input keywords from the document identification unit 25 and the URL extraction unit 30, and merges the URLs for each keyword. After that, use a search engine or identify the document containing each specified URL from the information in the link database (Fig. 8) provided in the anchor text extraction unit 40, and specify the anchor text of the link to the document. Extract. Figure 8 shows the contents of the data expected in the linked database. In addition, when using a search engine (goo (registered trademark), etc.), for example, if the URL you are looking for is A, you can obtain the desired data by specifying a search request such as "link: A". it can. The anchor text extraction unit 40 has a memory and stores keywords and anchor text. After performing the above processing for all the keywords input to the anchor text extraction unit 40, the data contents are passed to the component (anchor text analysis unit 50) that performs the next processing. Figure 9 shows an example of data storage in the memory of the anchor text extraction unit 40.
Step 107) The anchor text extraction unit 40 determines whether or not the processing of all documents has been completed. If so, the process proceeds to step 108, and if not, the process proceeds to step 106.
Step 108) The anchor text analysis unit 50 calculates the appearance frequency and the appearance frequency of each substring for the anchor text extracted by the anchor text extraction unit 40. As a substring, the frequency of appearance is calculated for all delimiters with n-gram. It is also conceivable to limit the analysis target to the character string that becomes the prefix or suffix of the anchor text. Further, in the anchor text analysis unit 50, it is conceivable that the field-dependent vocabulary remains as it is by analyzing the corpus and using the result of analyzing the anchor text that does not depend on the class. The analysis result in the anchor text analysis unit 50 is stored in the anchor text statistical information DB 70. Figure 10 shows an example of the data registered in the anchor text statistical information DB70.
The anchor text statistical information DB 70 stores the data analyzed by the anchor text analysis unit 50. An example of data is shown in FIG.
Step 109) When the processing of all anchor texts is completed in the anchor text analysis unit 50, the process proceeds to step 110, and if not, the process returns to step 108.
Step 110) For each of the entered keywords, the anchor text cleaning unit 60 removes unnecessary character strings contained in the anchor text based on the data stored in the anchor text statistical information DB 70, and sets it as a synonym candidate. To do. First, based on the information in Fig. 10 (C), the vocabulary used frequently (when the threshold value α is exceeded) in all classes is extracted, considered to be a general vocabulary, and the substring to be removed is selected. Register in the specified stop word list.
Step 111) Next, based on the information in Fig. 10 (B), the vocabulary frequently used in the class is extracted. This is considered to be a general vocabulary, but it is not necessarily a string other than a proper expression (for example, the string "aviation" when creating a list of airlines is frequent but unnecessary. If the information in Fig. 10 (C) is checked and the threshold β is exceeded, or if they are clearly different from the proper expression, such as when they contain consecutive strings of symbols, register them in the stop word list. To do. In addition to simply considering the frequency, it is also conceivable to consider the dispersion and consider that the widely distributed substring is unlikely to be a part of the named entity, and use it as a criterion for the stop word judgment together with the frequency.
Here, the probability that the substring x appears with the keyword i is P (x).<sub>i</sub>), The variance can be evaluated by the following formula for calculating entropy. If this value is close to log | I | (I is a set of keywords), the variance is considered to be large, and it can be determined that there is a high possibility that it is a general word.
<maths num="1"><img file="JP2007164635A_D0001.tif" /></maths> Step 112) If the anchor text cleaning unit 60 completes the processing for all the keywords, the process proceeds to step 113, and if not, the process proceeds to step 110.
Step 113) After that, the anchor text cleaning unit 60 removes the stop word from the data registered in FIG. 10 (A) and reconstructs the data. Then, a vocabulary that appears frequently exceeding the threshold value γ is used as a synonym.
Step 114) The synonym output unit 80 outputs the data determined to be synonymous by the anchor text cleaning unit 60 together with the input keyword.
In the present invention, it is possible to construct the function of each component of the synonymous vocabulary acquisition device as a program, install it on a computer and execute it, or distribute it via a network.
In addition, the constructed program can be stored in a portable storage medium such as a hard disk, a flexible disk, or a CD-ROM, installed on a computer, executed, or distributed.
The present invention is not limited to the above-described embodiment, and various modifications and applications can be made within the scope of the claims.
The present invention is applicable to a technique for extracting aliases and abbreviations of proper nouns from hypertext.
<figref num="1">It is a figure for demonstrating the principle of this invention.</figref><figref num="2">It is a principle block diagram of this invention.</figref><figref num="3">It is a block diagram of the synonymous vocabulary acquisition device in one Embodiment of this invention.</figref><figref num="4">It is a flowchart of operation in one Embodiment of this invention.</figref><figref num="5">This is an example of data received by the keyword input unit according to the embodiment of the present invention.</figref><figref num="6">This is an example of data storage in the memory of the document identification unit according to the embodiment of the present invention.</figref><figref num="7">This is an example of data storage in the memory of the URL extraction unit according to the embodiment of the present invention.</figref><figref num="8">This is an example of the contents of the link database used in the anchor text extraction unit according to the embodiment of the present invention.</figref><figref num="9">This is an example of data storage in the memory of the anchor text extraction unit according to the embodiment of the present invention.</figref><figref num="10">This is a data example of the anchor text statistical information DB according to the embodiment of the present invention.</figref>
Code description
10 Keyword input unit 20 Keyword search method, Search unit 25 Document identification method, Document identification unit 30 Link search method, URL extraction unit 40 Anchor text identification method, Anchor text extraction unit 50 Anchor text by class Statistical analysis method, anchor text Analysis unit 60 Anchor character string cleaning means, anchor text cleaning unit 70 Storage means, anchor text statistical information DB80 Synonym output unit
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9946812B2 | Cited by | United States of America | Applicant |
| US10176274B2 | Cited by | United States of America | Applicant |
| JP2011170790A | Cited by | Japan | Search report |
| US9916397B2 | Cited by | United States of America | Applicant |
| US10007740B2 | Cited by | United States of America | Applicant |
| US9916397B2 | Cited by | United States of America | Applicant |
| US9785726B2 | Cited by | United States of America | Applicant |
| US9785726B2 | Cited by | United States of America | Applicant |
| JP2015176511A | Cited by | Japan | Examiner |
| JP2009086979A | Cited by | Japan | Examiner |
| US9576054B2 | Cited by | United States of America | Applicant |
| JP2012527028A | Cited by | Japan | Examiner |
2 members in 1 office
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2005362386 | Japan | A | |
| JP20050362386 | – | – | – |
Members2
| Document | Office | Kind | |
|---|---|---|---|
| JP2007164635AThis record | Japan | A | |
| JP4143085B2 | Japan | B2 |
15 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 2007164635
- Publication, DOCDB
- 2007164635
- Publication, EPODOC
- JP2007164635
- Application
- 362386
- Application, DOCDB
- 2005362386
- Application, EPODOC
- JP20050362386
Titles2
- Japanese
- 同義語彙獲得方法及び装置及びプログラム
- English
- Synonymous vocabulary acquisition methods, devices and programs
Classification
- IPC, 2
- G06F17 30
- G06F17 28