EP0411231A2

Method for compressing and decompressing forms by means of very large symbol matching.

Abstract

This method relates to the compression of information contained in filled-in forms (0) by separate handling of the corresponding empty forms (CP) and of the information written into them (VP). Samples of the empty forms are pre-scanned, the data obtained digitized and stored in a computer memory to create a forms library. The original, filled-in form (0) to be compressed is then scanned, the data obtained digitized and the retrieved representation of the empty form (CP) is then subtracted, the difference being the digital representation of the filled-in information (VP), which may now be compressed by conventional methods or, preferably, by an adaptive compression scheme using at least two compression ratios depending on the relative content of black pixels in the data to be compressed.

EP0411231A2, drawing sheet 1
Sheet 1 of 69

Term

Term ended

Projected expiry passed 10 October 2009, 17 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

11 claims: 11 independent, 0 dependent

  1. 1
    Method for compressing, for storage or transmission, the information contained in filled-in forms (O) by separate handling of the corresponding empty forms (CP) and of the information written into them (VP), characterized by the steps of:- pre-scanning the empty forms (CP), digitizing the data obtained, and storing the digitized representations relating to each of the empty forms (CP) in a computer memory to create a forms library,- scanning the original, filled-in form (0) to be compressed, digitizing the data obtained,- identifying the particular one of said empty forms (CP) in said forms library and retrieving the digital representation thereof,- subtracting said retrieved representation of the empty form (CP) from said digital representation of the scanned filled-in form (O), the difference being the digital representation of the filled-in information (VP), and- compressing the digital representation of the filled-in information (VP) by appropriate methods.
  2. 2
    Method in accordance with claim 1, characterized in that the scanning parameters, such as brightness and threshold level, are determined separately for the empty form (CP) and for said original filled-in form (O).
  3. 3
    Method in accordance with claim 1, characterized in that prior to the subtraction step, registration information relating to the relative position of the filled-in information (VP) with respect to the completed form (0) is determined and registration of said original filled-in form (O) with said empty form (CP) is performed.
  4. 4
    Method in accordance with claim 3, characterized in that the said registration information is determined through dimensionality reduction.
  5. 5
    Method in accordance with claim 3, characterized in that said registration is performed in the following sequence of steps:- partitioning the the scanned filled-in form (0) into small segments,- estimating, for each segment, the optimum shifts to be performed in x- and y directions,- placing each segment of said original filled-in form (0) at the appropriate area of the output image array using the shift information previously established, so that a complete, shifted image is obtained when the placements for all segments have been completed.
  6. 6
    Method in accordance with claim 5, characterized in that said original filled-in form (O) is partitioned in a proportion corresponding at least approximately to 16 segments per page of 210 x 297 mm size.
  7. 7
    Method in accordance with claim 5, characterized in that said partitioning is performed such that the segments overlap at their margins by a distance corresponding to about two pixels.
  8. 8
    Method in accordance with claim 5, characterized in that said estimated optimum shifts for each pair of neighbouring segments are checked for consistency by assuring that their difference does not exceed a predetermined threshold and, if it does, automatically calling for operator intervention.
  9. 9
    Method in accordance with claim 1, characterized in that said subtraction step involves removing, from the original filled-in form (0), all black pixels which belong to said corresponding empty form (CP) or which are located close to black pixels belonging to said corresponding empty form (CP), and retaining all black pixels which belong to said filled-in information (VP).
  10. 10
    Method in accordance with claim 1, characterized in that said compression step involves the application of at least two different compression ratios depending on the content, in the data to be compressed, of portions containing a comparatively large number of black pixels and portions containing a comparatively small number of black pixels, with the data to be compressed being first subjected to a pre-filtering step to determined the relative compressibility thereof.
  11. 11
    Method in accordance with any one of the preceding claims, characterized in that all steps are performed with the said binary representations being maintained in a byte format without unpacking any of the bytes to their component pixels.