EP0280866A2

Computer method for automatic extraction of commonly specified information from business correspondence.

Abstract

A Parametric Information Extraction (PIE) system has been developed to identify automatically commonly specified information such as author, date, recipient, address, subject statement, etc. from documents in free format. The program-generated data can be used directly or can be supplemented manually to provide automatic indexing or indexing aid, respectively. The PIE system uses structural, syntactic, and semantic knowledge to accomplish its objective. The structural analysis identifies the major components of a document which are the document heading, body, and ending. The heading and the ending which usually contain commonly specified information, are then analyzed by a battery of morphological, syntactic, and semantic pattern-matching procedures that provide specified information in standardized forms that can be easily manipulated by computer.

EP0280866A2, drawing sheet 1
Sheet 1 of 15

Term

Term ended

Projected expiry passed 22 January 2008, 18.7 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

4 claims: 3 independent, 1 dependent

  1. 1
    An information extraction method, for automatically identifying commonly specified information from documents in free format, comprising the steps of:reading in a document to be abstracted;reading in structural, syntactic and semantic know­ledge data base identifying the major information components of the document using said structural knowledge data base;;analyzing the information components so identified, to obtain commonly specified information with said syntactic and said semantic knowledge data bases, employing pattern-matching procedures which provide commonly specified information in standardized form;outputting a formatted frame containing slots corresponding to the specified information that occur in the major information components of said document.
  2. 2
    An automatic method of locating commonly specified information in a document in free format using pat­tern-matching techniques, comprising the steps of:identifying the date information in any location of a document using said pattern-matching techniques and a representation of the general syntactic pat­tern of dates;identifying personal names in any location of a doc­ument using said pattern-matching techniques and a specification of the constituents of personal names;identifying the subject of the document within the heading of the document by using said pattern-match­ing techniques and a specification of the formats used for titles, and other fields that contain the subject information.
  3. 4
    An automatic method of mapping commonly specified information fields from a document in free format into standardized form comprising the steps of:identification of the semantic constituents of each field;application of pre-specified transformations to each constituent of said field and reformatting the re­sulting constituents in a prescribed order;placement of the standardized field in an output data structure which unambiguously identifies the type of data either by position or by explicit tagging.