JP2004192546A

Information retrieval method, device, program, and recording medium

Abstract

Problem to be solved.To perform ranking retrieval targeting a large-scale document set at high speed and low costs.

Solution.A retrieval processing part 7 inputs an appearance location list of an index word from a high-frequency word transposition index 4 in which a list of index words and appearance locations of documents of which the frequencies of the index word are equal to or more than a threshold F is stored when a search word is constituted of only one index word, inputs the appearance location list of the respective index words from the high-frequency word transposition index 4 when the retrieval word is a line of a plurality of index words, calculates a list of locations where all the index words adjacently appear, transfers the list to an adaptation calculation part 6 and receives an adaptation document list. When the obtained adaptation document list refers to T or more documents required for display of retrieval results and when documents with top T-th adaptation have adaptation larger than the maximum adaptation which can be taken by low-frequency documents not to be stored in the high-frequency word transposition index 4, they are outputted.

Copyright (C)2004,JPO&NCIPI

JP2004192546A, drawing sheet 1
Sheet 1 of 2

Term

Term ended

Projected expiry passed 13 December 2022, 3.8 years ago.

  1. Priority and filed
  2. Published
  3. Projected expiry
  4. Today

4 claims: 3 independent, 1 dependent

  1. 1
    From the index creation step that extracts index terms from the documents to be searched and stores them and a list of pairs of their occurrence positions in the transposed index, and from the transposed index, the index term frequency becomes equal to or higher than a predetermined threshold. A high-frequency word extraction step that extracts a list of index terms and occurrence position pairs and stores them in the high-frequency word transposition index and a search term occurrence position list are received, and each document is obtained from the frequency and document size of the search terms in each document. The conformity calculation step that calculates the conformity and outputs it as a conformity document list, and when the search term is received from the user and the search term consists of only one index term, the high-frequency word transposition index is used. Enter a list of index term occurrence positions, and if the search term is a sequence of multiple index terms, enter a list of index term occurrence positions from the high-frequency word transposition index, and enter all index terms. Find a list of positions where are adjacent to each other, and refer to a document in which the number of occurrence positions obtained from the high-frequency word transposition index is T (T is an integer of 1 or more) required to display search results. If so, the list of occurrence positions is passed to the conformity calculation step, a conformance document list is received, and the T-th highest conformity document is a low-frequency document that is not stored in the high-frequency word transposition index. If it has a fit greater than the maximum possible fit, it outputs them, and the high-frequency word transposition index does not provide a list of occurrence positions that refer to T or more documents, or fits. If there is a possibility that the top T documents are not searched correctly, the transposition index is used to obtain a list of the occurrence positions of the search terms, which is output to the conformity calculation step, and the conformity is calculated. An information search method having a search processing step that inputs a list of conforming documents from a calculation step and outputs a maximum of T conforming documents in descending order of conformity. 検索対象の文書から索引語を抽出し、それらと、それらの出現位置の組のリストを転置索引に格納する索引作成ステップと、前記転置索引から、索引語頻度があらかじめ定められた閾値以上になる索引語と出現位置の組のリストを抽出し、高頻度語転置索引に格納する高頻度語抽出ステップと、検索語の出現位置リストを受け取り、各文書における検索語の頻度と文書サイズから文書毎に適合度を計算し、適合文書リストとして出力する適合度計算ステップと、利用者から検索語を受け取り、検索語が1つの索引語のみから構成される場合には前記高頻度語転置索引から該索引語の出現位置のリストを入力し、検索語が複数の索引語の並びである場合には、それぞれの索引語の出現位置のリストを前記高頻度語転置索引から入力し、すべての索引語が隣接して出現する位置のリストを求め、前記高頻度語転置索引から求められた出現位置のリストが、検索結果の表示に必要なT個(Tは1以上の整数)以上の文書を参照している場合には該出現位置のリストを前記適合度計算ステップに渡して、適合文書リストを受け取り、適合度上位T番目の文書が、前記高頻度語転置索引には格納されない低頻度の文書がとり得る最大の適合度よりも大きい適合度を持つ場合は、それらを出力し、前記高頻度語転置索引からT個以上の文書を参照する出現位置のリストが得られなかった場合、あるいは適合度上位T個の文書が正しく検索されていない可能性がある場合には、前記転置索引を用いて検索語の出現位置のリストを求め、それを前記適合度計算ステップに出力し、前記適合度計算ステップから適合文書のリストを入力し、適合文書を、適合度の降順で上位最大T個出力する検索処理ステップを有する情報検索方法。
  2. 2
    A transposed index that stores a set of index terms and occurrence positions, an index creation means that extracts index terms from the document to be searched and stores those index terms and their occurrence position sets in the transposed index, and a high frequency. A list of index terms and occurrence position pairs whose index term frequency is equal to or higher than a predetermined threshold is extracted from the high-frequency word transposition index that stores only the index terms and their occurrence position pairs. Receives the high-frequency word extraction means stored in the high-frequency word index and the list of occurrence positions of search terms, calculates the degree of conformity for each document from the frequency of search terms and the document size in each document, and outputs it as a conforming document list. When the suitability calculation means and the search term are received from the user and the search term is composed of only one index term, a list of appearance positions of the index term is input from the high-frequency word transposition index, and the search term is entered. When is a sequence of a plurality of index terms, a list of occurrence positions of each index term is input from the high-frequency word transposition index, and a list of positions where all index terms appear adjacent to each other is obtained. If the list of occurrence positions obtained from the high-frequency word transposition index refers to T or more documents (T is an integer of 1 or more) required for displaying search results, the list of occurrence positions is described above. The conformance calculation means is passed to receive a conformance document list, and the T-th highest conformity is greater than the maximum conformance that a low-frequency document that is not stored in the high-frequency word transposition index can have. If so, it is possible that they are output and a list of occurrence positions that refer to T or more documents cannot be obtained from the high-frequency word transposition index, or the T documents with the highest degree of conformity are not searched correctly. If there is a possibility, the transposed index is used to obtain a list of appearance positions of search terms, which is output to the conformity calculation means, a list of conform documents is input from the conformity calculation means, and conform documents are entered. An information search device having a search processing means that outputs a maximum of T items in descending order of conformity. 索引語と出現位置の組を格納する転置索引と、検索対象の文書から索引語を抽出し、それら索引語と、それらの出現位置の組を前記転置索引に格納する索引作成手段と、高頻度の索引語とその出現位置の組のみを格納する高頻度語転置索引と、前記転置索引から、索引語頻度があらかじめ定められた閾値以上になる索引語と出現位置の組のリストを抽出し、前記高頻度語索引に格納する高頻度語抽出手段と、検索語の出現位置リストを受け取り、各文書における検索語の頻度と文書サイズから文書毎に適合度を計算し、適合文書リストとして出力する適合度計算手段と、利用者から検索語を受け取り、検索語が1つの索引語のみから構成される場合には前記高頻度語転置索引から該索引語の出現位置のリストを入力し、検索語が複数の索引語の並びである場合には、それぞれの索引語の出現位置のリストを前記高頻度語転置索引から入力し、すべての索引語が隣接して出現する位置のリストを求め、前記高頻度語転置索引から求められた出現位置のリストが、検索結果の表示に必要なT個以上(Tは1以上の整数)の文書を参照している場合には該出現位置のリストを前記適合度計算手段に渡して、適合文書リストを受け取り、適合度上位T番目の文書が、前記高頻度語転置索引には格納されない低頻度の文書がとり得る最大の適合度よりも大きい適合度を持つ場合は、それらを出力し、前記高頻度語転置索引からT個以上の文書を参照する出現位置のリストが得られなかった場合、あるいは適合度上位T個の文書が正しく検索されていない可能性がある場合には、前記転置索引を用いて検索語の出現位置のリストを求め、それを前記適合度計算手段に出力し、前記適合度計算手段から適合文書のリストを入力し、適合文書を、適合度の降順で上位最大T個出力する検索処理手段を有する情報検索装置。
  3. 3
    An information retrieval program for causing a computer to execute the information retrieval method described in claim 1. 請求項1記載の情報検索方法をコンピュータに実行させるための情報検索プログラム。