EP1598770A3

Low resolution optical character recognition for camera acquired documents

Abstract

A global optimization framework for optical character recognition (OCR) of low-resolution photographed documents that combines a binarization-type process, segmentation, and recognition into a single process. The framework includes a machine learning approach trained on a large amount of data. A convolutional neural network can be employed to compute a classification function at multiple positions and take grey-level input which eliminates binarization. The framework utilizes preprocessing, layout analysis, character recognition, and word recognition to output high recognition rates. The framework also employs dynamic programming and language models to arrive at the desired output.

EP1598770A3, drawing sheet 1
Sheet 1 of 1

Term

Term ended

Projected expiry passed 19 May 2025, 1.3 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

1 sheet

  1. Sheet 1