Nova Patents
EP0809219A2

Document analysis systems and processes

Abstract

A knowledge-based document analysis system and process for identifying and decomposing constrained and unconstrained images of scanned documents (100) is disclosed. Low level features are extracted by a document feature extractor (105) within bitonal and grayscale images. Low level features are passed to a document classification means (115) which forms initial hypotheses about the document class. For constrained documents, a document analysis means (125) sorts through various models to determine the exact type of document and then extracts the relevant fields for character recognition. For unconstrained documents, through the use of a blackboard architecture which includes a knowledge database and knowledge sources, the document analysis means (125) creates information and hypotheses to identify and locate relevant fields within the document. These fields are then sent for optical character recognition.

EP0809219A2, drawing sheet 1
Sheet 1 of 26

Term

Term ended

Projected expiry passed 6 May 2017, 9.4 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

8 claims: 2 independent, 6 dependent

  1. 1
    A system for analyzing a target document including at least one informational element, the system comprising:(a) means (100) for storing a digitized image of the target document;(b) document feature extractor means (105) for extracting low level features from the digitized image;(c) document classification means (115) for classifying the document based upon the extracted low level features;and (d) document analysis means (125) for analyzing the target document in order to extract informational data associated with the at least one informational element, responsive to the document classifying means (115).
  2. 8
    A process for analyzing a target document including at least one informational element, comprising the steps of:(a) storing a digitized image of the target document on a storage device;(b) extracting low level features from the digitized image;(c) classifying the document based upon the extracted low level features;and (d) analyzing the target document in order to extract informational data associated with the at least one informational element.