Nova Patents
EP0802492A1

Document search system

Abstract

A unique character string is extracted from an input document 907, and a similarity search is performed by using the unique character string. The extraction of the unique character string is performed by calculating and evaluating the amount of feature of a character string through comparison between appearance frequency appearing in the input document 907 and appearance frequency in a set of documents 909 to be searched. Then, the extracted unique character string is used for the search. Documents found by the search are evaluated and arranged in the order of evaluation. The similarity factor of document is evaluated by using the appearance frequency of each unique character string in the input document so that higher evaluation is provided to a document in which unique character strings with higher weight appear many times. Such a system and method do not require vocabulary information or grammatical information, which run into difficulties when meeting new words or phrases, and allow a document search to be performed against a vague request of a user for document search.

EP0802492A1, drawing sheet 1
Sheet 1 of 70

Term

Term ended

Projected expiry passed 16 April 2017, 9.4 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

8 claims: 8 independent, 0 dependent

  1. 1
    A method for identifying a unique character string contained in an input document in a computer system, said computer system being able to search at least one stored comparison document, the method comprising the steps of:(a) associating and managing position information for a position in said comparison document where a partial comparison document character string extracted from said comparison document exists with said partial comparison document character string;(b) extracting a partial input character string from said input document, and determining such as a candidate character string;(c) identifying a partial comparison document character sting which matches part of said candidate character string with a predetermined similarity factor or higher;(d) identifying position information associated with said partial comparison document character string which matches with said predetermined similarity factor or higher;and (e) recognizing said candidate character string as the unique character string by comparing appearance frequency information on a part of said candidate character string appearing in said input document with the position information, and evaluating the amount of feature of said candidate character string.
  2. 2
    A method for searching a document which has a character string similar to a partial input character string existing in an input document in a computer from a plurality of documents to be searched stored in the computer, the method comprising the steps of:(a) extracting a partial character string from said input document, and determining such as a candidate character string;(b) evaluating the amount of feature of said candidate character string through comparison between appearance frequency information on a part of said candidate character string appearing in said input document and appearance frequency information on a part of said candidate character string appearing in a comparison document to recognize said candidate character string as a unique character string;and (c) searching the comparison document having a character string similar to said unique character string from said plurality of document to be searched.
  3. 3
    A method for identifying a unique character string contained in an input document in a computer system, said computer system being able to search at least one stored comparison document, the method comprising the steps of:(a) extracting a partial input character string from said input document, and determining such as a candidate character string;and (b) evaluating the amount of feature of said candidate character string through comparison between appearance frequency information on a part of said candidate character string appearing in said input document and appearance frequency information on a part of said candidate character string appearing in said comparison document to recognize said candidate character string as a unique character string.
  4. 4
    A method for evaluating similarity between a comparison document and an input document which contains a first unique character string and a second unique cnaracter string input in a computer, said computer system being able to search a stored comparison document, the method comprising the steps of:(a) calculating a first weight value corresponding to said first unique character string from appearance frequency information on a part of said first unique character string appearing in said input document;(b) calculating a second weight value corresponding to said second unique character string from appearance frequency information on a part of said second unique character string appearing in the input document;(c) calculating a first appearance frequency value on a part of said first unique character string appearing in said comparison document;(d) calculating a second appearance frequency value on a part of said second unique character string appearing in said comparison document;and (e) calculating the similarity factor of said comparison document from the first appearance frequency value taking said first weight value into account and the second appearance frequency value taking said second weight value into account.
  5. 5
    An apparatus for identifying a unique character string contained in an input document in a computer system, said computer system containing at least one stored comparison document, the apparatus comprising:(a) a storage device for storing a position information file which associates and manages position information for a position in said comparison document where a partial comparison document character string extracted from said comparison document exists with said partial comparison document character string;(b) means for extracting a candidate character string from said input document;(c) means for identifying a partial comparison document character string which matches part of said candidate character string with a predetermined similarity factor or higher;(d) means for identifying position information which is associated with said partial comparison document character string with the predetermined similarity factor or higher in said position information file;and (e) means for recognizing said candidate character string as tne unique character string by comparing appearance frequency information on a part of said candidate character string appearing in said input document with said position information, and evaluating the amount of feature of said candidate character string.
  6. 6
    An apparatus for searching a document having a character string similar to a partial input character string which exists in an input document in a computer from a plurality of documents to be searched which are stored in the computer, the apparatus comprising:(a) an input device for identifying said input document and instructing execution of search;(b) means for detecting from said input device the fact that said input document is identified and that said instruction of search is input;(c) means for extracting a candidate character string from said input document in response to the detection of the fact that said input document is identified and that said instruction of search is input;(d) means for calculating the amount of feature through comparison between appearance frequency information on a part of said candidate character string appearing in said input document and appearance frequency information on a part of said candidate character string appearing in a comparison document;(e) means for determining said candidate character string as a unique character string by evaluating said amount of feature;(f) means for searching the document to be searched having a character string similar to said unique character string from a plurality of documents to be searched;and (g) a display device for displaying the document to be searched having a character string similar to said unique character string.
  7. 7
    An apparatus for identifying a unique character string contained in an input document in a computer system, said computer system containing at least one stored comparison document, the apparatus comprising:(a) means for extracting a candidate character string from said input document;and (b) means for determining said candidate character string as a unique cnaracter string by evaluating tne amount or feature of said candidate character string through comparison between appearance frequency information on a part of said candidate character string appearing in said input document and appearance frequency information on a part of said candidate character string appearing in said comparison document.
  8. 8
    An apparatus for evaluating similarity between a comparison document and an input document containing a unique character string in a computer system, said computer system containing a stored comparison document, the apparatus comprising:(a) means for calculating a weight value corresponding to said unique character string from appearance frequency information on a part of said unique character string appearing in said input document;and (b) means for calculating the similarity factor of said comparison document from the appearance frequency information on a part of said unique character string appearing in said comparison document and said weight value.