EP2012240A1

Systems, methods, and software for classifying documents

Abstract

To reduce cost and improve accuracy, the inventors devised systems, methods, and software to aid classification of text, such as headnotes and other documents, to target classes in a target classification system. For example, one system computes composite scores based on: similarity of input text to text assigned to each of the target classes; similarly of non-target classes assigned to the input text and target classes; probability of a target class given a set of one or more non-target classes assigned to the input text; and/or probability of the input text given text assigned to the target to the target classes. The exemplary system then evaluates the composite scores using class-specific decision criteria, such as thresholds, ultimately assigning or recommending assignment of the input text to one or more of the target classes. The exemplary system is particularly suitable for classification systems having thousands of classes.

EP2012240A1, drawing sheet 1
Sheet 1 of 57

Term

Term ended

Projected expiry passed 1 November 2022, 3.9 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

8 claims: 2 independent, 6 dependent

  1. 1
    A computer-implemented method of classifying text to one or more target classes in a target classification system, the method comprising:• identifying one or more noun-word pairs in a portion of text.
  2. 8
    A computer-implemented method of classifying input text to one or more target classes in a target classification system, the method comprising:• identifying a first set of noun-word pairs in the input text, with the first set including at least one noun-word pair formed from a noun and non-adjacent word in the input text;• identifying two or more second sets of noun-word pairs, with each second set including at least one noun-word pair formed from a noun and non-adjacent word in text associated with a respective one of the target classes;• determining a set of scores based on the first and second sets of noun-word pairs;and • classifying or recommending classification of the input text to one or more of the target classes based on the set of scores.