EP0980043A2

A cluster-based method and system for browsing large document collections

Abstract

Scatter-Gather is a computer based document browsing method which operates in time proportional to a number of documents in a target corpus. The Scatter-Gather method includes: preparing an initial ordering of the corpus using, for example, an off-line computational method; determining a summary of the initial ordering of the corpus for interactive utility; and providing a further ordering of the corpus using, for example, an on-line non-deterministic method. The step of an off-line preparation of an initial ordering of a corpus is non-time-dependent, thus an accurate initial ordering is prepared. The step of determining a summary includes determining a summary for presentation to a user without scrolling on a CRT. The step of providing a further ordering includes truncated group average agglomerate clustering, merging disjointed document sets, center finding, assign-to-nearest and other refinement methods.

EP0980043A2, drawing sheet 1
Sheet 1 of 57

Term

Term ended

Projected expiry passed 15 October 2012, 13.9 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

9 claims: 6 independent, 3 dependent

  1. 1
    A document partitioning Fractionation method in a digital computer for non-hierarchical, linear-time partitioning of a corpus of documents, said Fractionation method comprising the steps of:preparing an ordering of the corpus;determining a partitioning of a desired size from the ordering;andrefining the partitioning.
  2. 4
    The Fractionation method of claims 1, 2 or 3, wherein the determining step comprises truncated group averaging agglomerative clustering which includes limiting a growth of an agglomeration by terminating a group averaging agglomerative clustering before a single over-arching agglomeration is formed.
  3. 5
    The Fractionation method of claims 1, 2 or 3, wherein the refining step includes refining with an assign-to-nearest method for assigning a document to a nearest bucket.
  4. 7
    The Fractionation method of claims 5 or 6, wherein the refining step includes splitting non-similar buckets.
  5. 8
    The Fractionation method of claims 5, 6 or 7, wherein the refining step includes detecting weak-small and incoherent buckets by applying size and two-way link thresholds.
  6. 9
    The Fractionation method of claims 1, 2 or 3, wherein:the determining step includes determining a partitioning of a desired size from the ordering to form a set of buckets, each document of the corpus of documents assigned to only one bucket of the set of buckets;andthe refining step includes performing a predetermined number of iterations of: creating a set of modified buckets from the set of buckets based on content and size of each bucket;andreassigning each document of the corpus of documents to the set of modified buckets.