Outlier detection during scanning
Abstract
This record has no abstract on file.
Term
Term ended
Expired 9 August 2025, 1.1 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
16 claims: 5 independent, 11 dependent
- 1A step of receiving a scanned image by optically scanning a series of pages of a multi-page document and processing the scanned image to generate a page image corresponding to the original page of the multi-page document. A method of processing a multi-page document, which is a method of processing a multi-page document obtained during the processing of the scanned image.Of the page image corresponding to the original pageIt includes, and if so, a step of automatically determining a target criterion for an image quality parameter based on page characteristics and a step of checking whether the image quality parameter of the page complies with the target criterion. Automatically accepts the page image and, if not,Display the page image andoperatorManuallyAdjust the page image to accept or accept the page imageShi, The page characteristics are the paper size of the original page, the text areaplace,Or, Text areaThe image quality parameters, including dimensions, are the paper size of the scanned image, the text area.place,Or, Text areaThe target criteria, including dimensions, are for the multi-page document.Of the page characteristicsSet based on statistical characteristics, The target criteria include expectations and tolerances for image quality parameters based on the statistical characteristics.The method, which is characterized by that. マルチページ文書の一連のページの光学的な走査により走査画像を受信するステップと、 前記マルチページ文書の元のページに対応するページ画像を生成するために前記走査画像を処理するステップとを含む、 マルチページ文書を処理する方法であって、 前記走査画像の処理中に得られるマルチページ文書の元のページに対応する前記ページ画像のページ特性に基づいて画像品質パラメータのための目標基準を自動的に決定するステップと、 ページの前記画像品質パラメータが前記目標基準に従っているか否かをチェックするステップとを含み、 従っている場合には、前記ページ画像を自動的に受け入れ、 従っていない場合には、前記ページ画像を表示し、操作者は手動で前記ページ画像を受け入れ或いは受入れのために前記ページ画像を調整し、 前記ページ特性は、前記元のページの紙サイズ、テキスト領域場所、又は、テキスト領域寸法を含み、 前記画像品質パラメータは、前記走査画像の紙サイズ、テキスト領域場所、又は、テキスト領域寸法を含み、 前記目標基準は、前記マルチページ文書の前記ページ特性の統計的特性に基づき設定され、前記目標基準は、前記統計的特性に基づく画像品質パラメータに対する期待値及び公差を含むことを特徴とする、 方法。
- 7The step of processing the scanned image comprises detecting the spine of a book in a multi-page document and generating two page images from a single scanned image according to any one of claims 1-6. The method described. 前記走査画像を処理するステップは、マルチページ文書における本の背を検出し、単一の走査画像から2つのページ画像を生成することを含む、請求項1乃至6のうちの何れか1項に記載の方法。
- 12A scanner unit that generates a scanned image by optically scanning a series of pages of a multi-page document, a user interface unit, and processing the scanned image to generate a page image corresponding to the original page of the multi-page document. A document processing system including an image processor unit that includes an outline detection means, and the outline detection means includes the multi-page document obtained during the processing of the scanned image.Of the page image corresponding to the original pageGoal criteria for image quality parameters based on page characteristicsAutomaticallyDetermine and check if the image quality parameters on the page comply with the target criteria, and if so, automatically accept the page image, if not, if not.Display the page image andoperatorManuallyAdjust the page image to accept or accept the page image through the user interface unitShi, The page characteristics are the paper size of the original page, the text areaplace,Or, Text areaThe image quality parameters, including dimensions, are the paper size of the scanned image, the text area.place,Or, Text areaThe target criteria, including dimensions, are for the multi-page document.Of the page characteristicsSet based on statistical characteristics, The target criteria include expectations and tolerances for image quality parameters based on the statistical characteristics.A document processing system characterized by this. マルチページ文書の一連のページの光学的な走査により走査画像を生成するスキャナユニットと、 ユーザインターフェースユニットと、 前記マルチページ文書の元のページに対応するページ画像を生成するために前記走査画像を処理する画像プロセッサユニットとを含む、 文書処理システムであって、 当該文書処理システムは、アウトライア検出手段を含み、 該アウトライア検出手段は、 前記走査画像の処理中に得られる前記マルチページ文書の元のページに対応する前記ページ画像のページ特性に基づいて画像品質パラメータのための目標基準を自動的に決定し、 ページの前記画像品質パラメータが前記目標基準に従っているか否かをチェックし、 従っている場合には、前記ページ画像を自動的に受け入れ、 従っていない場合には、前記ページ画像を表示し、操作者は手動でユーザインターフェースユニットを介して前記ページ画像を受け入れ或いは受入れのために前記ページ画像を調整し、 前記ページ特性は、前記元のページの紙サイズ、テキスト領域場所、又は、テキスト領域寸法を含み、 前記画像品質パラメータは、前記走査画像の紙サイズ、テキスト領域場所、又は、テキスト領域寸法を含み、 前記目標基準は、前記マルチページ文書の前記ページ特性の統計的特性に基づき設定され、前記目標基準は、前記統計的特性に基づく画像品質パラメータに対する期待値及び公差を含むことを特徴とする、 文書処理システム。
- 1312. The outline detection means comprises statistically determining a target range for at least one image quality parameter based on the page characteristics obtained during the processing of the scanned image. System. 前記アウトライア検出手段は、前記走査画像の処理の間に得られる前記ページ特性に基づいて、少なくとも1つの画像品質パラメータのための目標範囲を統計的に決定することを含む、請求項12に記載のシステム。
Independent claims5
67 paragraphs, as filed
The present invention includes a step of receiving a scanned image by optically scanning a series of pages of a multi-page document and a step of processing the scanned image to generate a page image corresponding to the original page of the multi-page document. The present invention relates to a multi-page document processing method.
The present invention further relates to a computer program for processing a multi-page document.
The present invention further relates to a scanner unit that generates a scanned image by optically scanning a series of pages of a multi-page document, a local interface unit, and a page image corresponding to the original page of the multi-page document. The present invention relates to a document processing system including an image processor unit for processing the scanned image.
When a large document needs to be scanned for archiving, it is extremely important that all pages of the document be scanned flawlessly, because later when a scanning error is detected. The original document may no longer be available. Therefore, it is important to check the quality of each scanned image. However, the quality of each scanned image requires a great deal of time and effort, and imposes a great burden on the individual performing the scanning operation. Moreover, checking a large number of images is boring and error prone.
One way to avoid human checking is to use an automated system that automatically checks each new scanned image and, if possible, corrects the defective image with the corresponding image processing technique. Scanned images that do not comply with a given quality standard are referred to herein as "outliers".
Such methods are known from patent application WO98 / 09427 and disclose configurations and methods that guarantee quality during scanning or copying. This method involves feeding in the page (s) to be scanned and checking the quality of the scanned image in a series of steps, with respect to skew, double feed / overlap, shape deviation, and geometric deformation. It includes checking the external characteristics performed, checking the so-called internal quality of the page content, and checking the quality of the information content. If the measured quality meets or is better than the limit, automatic adjustment of the scanned image is performed as needed, after which the scanned image is added to the scanned file. If the measured quality is below the limit, the operator is required to resupply the page for e-scanning.
<p> In known systems, quality checks are based on fixed, pre-programmed quality limits that do not always match the actual situation. If the check result is negative, it has to be rescanned, which forces the operator to resupply the document. However, it can happen that the rejected scanned image is actually still acceptable or can be made acceptable with small adjustments and does not require rescanning.</p><p> The present invention processes a scanned image to generate a series of page images that faithfully correspond to the original page of a multi-page document, while producing a limited number of page images based on automatic detection of quality. The purpose is to provide a method and system that provide the operator with flexible choices to make corrections.</p>
<p> According to the first aspect of the present invention, the above object is a method as described in the opening paragraph, based on image parameters based on the page characteristics of the multi-page document obtained during the processing of the scanned image. It has a step to automatically determine the target criteria for the page and a step to check whether the page characteristics of the page comply with the target criteria, and if it does, it automatically accepts the page image, otherwise it operates. It is achieved by a method characterized by displaying or accepting a page image for a corrective action against a person.</p><p> According to the second aspect of the present invention, the above object is achieved by a computer program for carrying out the method.</p><p> According to the third aspect of the present invention, the above object is a document processing system as described in the opening paragraph, based on the page characteristics of a multi-page document obtained during the processing of the scanned image. Determine target criteria for image parameters, check if the page characteristics of the page comply with the target criteria, and if so, automatically accept the page image, otherwise corrective action for the operator It is achieved by a document processing system comprising an outliner detecting means for displaying a page image or accepting the page image via a user interface unit. The above measures have the following effects. The processing of the scanned image produces a page image based on the detected page characteristics. Various effects of the scanning process on the page image may be compensated or corrected in the scanning image processing step. Page characteristics and corrections are compared to target criteria set based on the statistical characteristics of the multi-page document itself. Therefore, the quality of the page image is measured relative to the characteristics of the document, and then the page image is considered an outlier if the proposed processed page image deviates significantly from the target criteria. The proposed page image is displayed and the operator may accept it or adjust or correct, rescan, or reject the page image with corrective action. This has the effect that while most of the scanned image is processed automatically, the operator is required to see only a limited number of outliers. In addition, the operator can prevent rejection of images that are acceptable or still adjustable.</p><p> In particular, the target criteria are adjusted to the global characteristics of the document, taking into account the characteristics of the actual scanned multi-page document. This effectively improves outlier detection and effectively reduces the number of accurate pages that are mistakenly classified as outliers. Outliers are detected when significant deviations are detected from the target criteria, and only then are the proposed images presented to the operator for authentication or correction. Therefore, errors in the final set of page images are effectively prevented by selective inspection and correction.</p><p> In one embodiment of the method, the method of determining the target criteria comprises statistically determining the target range for at least one image parameter based on the page characteristics obtained during the processing of the scanned image. This has the effect that the target range of expected values is statistically determined or adjusted based on the characteristics of the scanned image of the multi-page document.</p><p> In one embodiment, the image parameters include page size, text area position or dimensions. Such parameters are usually constant throughout multi-page documents such as books and magazines. Outliers are detected and displayed to the operator based on the detection of whether the detected image parameters, such as paper size, are outside the target range of expected values.</p><p> In one embodiment of the method, the step of detecting whether the image parameter complies with the target criteria includes calculating a reliability factor that dictates the reliability of the adjustments made to generate the page image. Confidence is calculated for processing steps such as descue or rotation, for example during processing it is detected that the result of the proposed page image is unreliable due to unclear data. Therefore, the target criteria may include the lowest level of reliability.</p><p> In one embodiment, the expected value includes a priori knowledge of the document for a given parameter of a page of a multi-page document. A priori knowledge may be validated or combined with statistical data from multi-page documents. A general characteristic of a document is that it is generally expected, such as the characters being arranged in horizontal lines. Also, a given set of parameters, such as vertical lines for a Japanese document, may be applied or selected for the appropriate document. By using a priori knowledge, outliers deviate from normal document characteristics and can be easily detected.</p><p> In one embodiment, the predetermined parameter includes the orientation of the text line, and the step of processing the scanned image detects the orientation of the text line and corrects the skew of the scanned image according to the orientation of the detected text line. Including. This has the effect that normal errors during scanning, i.e. the tilted position of the original image on the scanner unit, can be easily corrected.</p><p> In one embodiment, the method comprises constructing a composite set of page images, the composite set having a logical portion corresponding to a range of pages in a multi-page document. This has the effect that logical subdivisions of the original document, such as chapters in a book, can be maintained within a composite set of page images. In certain embodiments, the method comprises receiving a command from the user instructing a subset of scanned images to form a logical portion of a composite set of page images. This allows the operator to easily enter commands during scanning when a consistent range of pages of the original document is started and / or completed.</p><p> In addition, preferred embodiments of the device according to the invention are set forth in the appended claims, the disclosure thereof of which is incorporated herein by reference.</p>
These and other aspects of the invention will become apparent and further taught by reference to the accompanying drawings and by reference to the examples described by the examples in the following description. The figure is not drawn exactly as Scase. In the figure, the same reference code is given to the element corresponding to the already-technical element.
FIG. 1 shows a device for processing a document, where the different parts are shown generally separately. The document is usually a paper document, but may include any type of sheet carrying information, such as transparencies, books, drawings, and the like. The document processing device 1 may be only a scanner, but is preferably a multifunctional device including a printing, copying or fax function, such as a versatile copier. The scanner unit 120 includes a flat bed scanner comprising a glass platen on which the original document can be placed, a CCD array, and an imaging unit having a movable mirror and a lens for imaging the document on the CCD array. In this situation, the CCD array produces an electrical signal that is converted into digital image data in a particularly well known manner. The document feeder 110 includes an input tray 111 for introducing a stack of documents, a transport mechanism (not shown) for transporting documents one by one along the scanner unit 120, and a sending tray in which documents are arranged after scanning. May be equipped with 112.
The multi-page document to be scanned can be entered via the document feeder at the appropriate time. For example, documents and magazines are manually entered on the platen. Additional scanning aids, such as automatic page rotation of the book, may be provided.
The apparatus may have a printer unit 130, including, for example, a known electrophotographic processing section, in which the photoconducting medium is exposed and charged by digital image data through the LED array and then. The toner image is transferred and fixed on the image support, usually on a sheet of paper. Stocks of image supports in different formats and orientations are available in supply section 140. Image supports with toner images are transferred to finishing and delivery section 150, which sets them together, staples them, and deposits them in delivery tray 151, if necessary.
The control unit of the device is schematically indicated by reference numeral 170. The scanning image processing function according to the present invention will be described in detail below with reference to FIG. Cable 171 may connect the control unit 170 to the local network. The network may be preferred, but may be partially or completely wireless.
The device has a user interface 160, including, for example, an operator control panel provided on the device for its operation. The user interface includes a display 161 and keys. The operation of the display for controlling document processing is described below.
In the document processing system according to the invention, scanning may be performed on another device, and image processing described below may be performed on a processor unit having a display and operator interface, such as a user workstation. The processor may be built as a dedicated hardware unit or may include standard processing units and software programs that implement the image processing and correction functions described below.
FIG. 2 shows how to scan a multi-page document. In the first step, the method is initiated at START 21 by optically scanning the multi-page document in step SCAN22. The scan may be performed on a complete document or on a partial document. For each scanning operation, the scanned image is generated by placing a new original page on the scanner or, for the book, both pages on the platen of the scanner unit. A scanned image containing two pages is referred to as a dual scan image. In special cases, the scanned image may include a larger number of sub-images that are automatically processed to separate the page images. The scanned image may be processed directly or stored as an intermediate file, or may be included in the final file, for example to preserve the raw source material.
In the next step PROC23, the scanned image is received and processed from a scan of a series of pages of the multi-page document to generate a page image corresponding to the original page of the multi-page document. For each scanned image, a number of processing steps are performed to retrieve the page image, or in the case of a double-scanned image, two page images. The page image is a representation of the original page, i.e., a processed version of the scanning data provided by the scanning image. Some examples of processing a scanned image into a page image are described below.
According to the present invention, statistical information STAT20 is collected during the processing of step PROC23, and the statistical information is targeted for image parameters based on the page characteristics of the multi-page document obtained during the processing of the scanned image. Used to determine criteria. Expected values for properties such as paper size, text area dimensions and contrast may be determined. Target criteria are expected values and tolerances for given image parameters based on statistically determined characteristics, such as mean, median, and variance-based tolerances.
It should be noted that at least a large number of pages need to be processed first so that the statistical parameters are determined in a reliable manner. Therefore, the initial part of a multi-page document (at least a few pages, but preferably an essential part of the document or even the entire document) needs to be available (ie, scanned and stored). Yes, it is processed until the outlier detection described later is started. As a result, if the initial portion of the multi-page document contains 10 pages and, for example, the first or second scanned image is found to be outliers, a 10 page delay is presented. After the initial portion has been processed, additional scanned images may be tested against the outliers without delay. Initially determined statistical parameters may be refined during the processing of the rest of the document. Alternatively, a complete multi-page document may be scanned and stored, a complete set of scanned images may be processed as an initial run to determine statistical parameters, and an outlier in the second run. Detected based on complete document statistics.
In step 23, the proposed page image is generated by the enhancement or correction processing function based on the detected image parameters in step PROP24. Some examples of such processing functions are described below.
The next step, OUTLIER 25, determines whether the page image is outlier, i.e., whether the image page has properties or image parameters that significantly deviate from the target criteria based on statistical information 20. .. Page images are evaluated by determining if the image parameters are outside the target criteria.
If the page image is not an outlier, processing continues in step STORE 26 by automatically accepting the page image. However, if the page image is detected as an outlier, the process is continued by interacting with the operator as follows. In step DISPLAY30, the page image is displayed to the operator. Therefore, the operator can visually inspect the proposed page image. If the results are unacceptable, the processing of the page image can be manually adjusted in step ADJUST32. For example, if the wrong text area is proposed, the portion of the scanned image that includes the edges of the original text is truncated and the operator may adjust the boundaries of the text area of the proposed page image. Subsequently, the adjusted or accepted page image is saved in step STORE 26.
Finally, if the available scanned image is processed to be detected in step NEXTIMAGE27 and no further pages or document parts need to be scanned to be detected in step NEXTSCAN28, the resulting page The image is combined with the multi-page digital output document of the original multi-page document in step COMBINE 29 and saved, for example, in a file. Note that step NEXTSCAN28 may be omitted if the process is batch-seeking and requires the complete document to be scanned before the process begins. The process is terminated by sending the output document file at END33.
In one embodiment of the method, the step of determining statistical values for image parameters in step PROC23 includes: The page edge detector detects the strongest edge in the scan and selects the four edges that form the bounding box closest to the page area. The paper size related parameters and characteristics detected for individual pages are statistically analyzed and, for example, averaged to estimate the final paper size characteristics of the original page. Expected values for parameters or characteristics are then stored as paper size criteria to be compared with additional pages of the multi-page document.
In one embodiment, the target criteria include a target range for at least one of the image parameters based on the page characteristics obtained during the processing of the scanned image. For example, the image parameter may include the text area position. The text area parameter is usually constant within the range for a multi-page document. Further, more detailed features such as page header, footer position, page sequence or chapter number may be separately detected and stored as expected values. Therefore, lost pages may be detected.
In addition to the expected value based on the statistically analyzed characteristics of the scanned image, the processing in step PROC23 may include a priori knowledge of the document, i.e., certain parameters or characteristics of the pages of a multi-page document. It is assumed to exist. For example, many scanned images Has standard paper sizes such as A4 and Letter. A practical example of a given parameter is the orientation of the text line, that is, the text is hypothetical to be aligned in a straight line in a direction parallel to the edge of the paper, with upright characters arranged in a horizontal line. Conceivable. Therefore, it is assumed that from the detected angle of the text line, the original page needs to be scanned at an oblique position and rotated until the text line is horizontal, that is, until the so-called skew is zero. To. Subsequently, the processing of the scanned image includes detecting the orientation of the text line and correcting the skew of the scanned image according to the orientation of the detected text line.
Further embodiments of the method make it possible to handle a large number of original pages in a scanned image, particularly two pages in a double-scanned image of a book or magazine. Therefore, the double-scanned image includes the spine of the multi-page document and the pages located on both sides of the spine.
FIG. 3 shows an example of a scanned image of a book. The scanned image is a double-scanned image of pages 35, 36 of the book. The page is scanned at 256 gray levels at 300 DPI (dots per inch), but with reduced resolution in the figure. Lines between pages, the so-called spine 37 of a multi-page document, are detected to separate the two pages. The text area 39 is available for each page, but may be used to include a picture, or may be unused on some pages, i.e. white. Paper size may be detected from the border of the white area on pages 35,36. The amount of paper hidden behind the book may vary depending on how the book is positioned. Such changes may be compensated for by reconstructing the original page with appropriate page margins on both the left and right edges, for example by centering the text area 39 independently of the spine of the book. Good.
In the method as shown in FIG. 2, the processing of the scanned image in step PROC23 may include detecting the spine of a multi-page document and generating two page images from a single scanned image. Detection of the original image may be automatic, for example, based on the detected paper size in combination with the orientation of the text line and / or the presence of spine, parallel text areas. Alternatively, the operator may enter a command indicating that the spine type multi-page document should be scanned and processed.
In one embodiment of step PROC23, the processing of the double-scanned image comprises detecting the orientation of the text line independently for each of the two pages. The tilt angle of both pages can vary depending on how the book is positioned. Therefore, by detecting and correcting the skew of both parts of the double-scanned image depending on the orientation of the detected text lines, both images are individually processed to have zero skew.
A further property of multi-page documents is that each page is accurate and has a so-called upright orientation. However, during scanning, the multi-page document may be facing down or sideways. During processing, the orientation of the original page with respect to the scanned image may be detected, and the page image with the top up is generated from the scanned images oriented in different ways by proper rotation. Page orientation may be detected by page layout characteristics such as top or bottom margins, page numbers, and so on. In a particular example, detecting page orientation is based on detecting characters and determining character characteristics.
FIG. 4 shows the characteristics of characters for detecting orientation, especially upside down. The text fragment 40 is parsed and the characters that extend below the bottom baseline 42 are called the descender 44, and the characters that extend above the top baseline 44 are called the ascender 43. Generally, there is a ratio for descenders and ascenders, for example, for Latin languages, a certain ratio is expected. Such a priori knowledge may be used as an initial value. The ratio to a particular document may be statistically determined or adjusted during the processing of the scanned image. The target ratio is applied, it is detected whether the document bitmap (page area in the scanned image) is upside down or straight, and the reliability of the detected orientation is calculated. Each character is classified as ascender, descender or nothing. For example, if the ratio of ascenders to descenders to a complete page is close to the target ratio, then the page is straight. If the ratio is close to the reciprocal of the target ratio, the page is upside down and a 180 degree rotation is performed to correct the page. Outliers are detected when the ratio of (corrected) pages deviates significantly from the target ratio.
Other characteristics of the character may also be used to detect the orientation of the character. Determining the orientation of selected characters, for example letter i, provides the character orientation parameters.
As shown in FIG. 2, the step OUTLIER 25 in this method is for detecting whether or not the image parameter is out of the target range. The target range may include reliability criteria such as: During the various correction and adjustment processes in step PROC23, a reliability factor is calculated that dictates the reliability of the adjustments made to generate the page image. For example, the amount of text lines on a page can be very small. Therefore, the orientation and character features of the text line to be detected can be unreliable and the reliability factor is low. Also, paper edges such as those detected may show gray areas, for example due to the original multi-page document paper that is not completely flattened on the platen. Therefore, from the presence of the gray picture element (pixel) near the detected paper edge, low reliability of the paper edge or orientation is assumed, and a low reliability factor is calculated.
FIG. 5 is a diagram of a component of the document processing system. The document processing system 50 includes a scanner unit 51 that generates a scanned image from an optical scan of a series of pages of a multi-page document 58. The scanner unit may be a part of a scanning and processing device, or may be another scanner device. The document processing system 50 includes a control unit 52 connected to the scanner unit 51, a local memory 57, a display 55, a user command element 56 such as a key, a menu control via a cursor, a user interface 54 including a touch section, and the like. including. The memory may be a solid-state storage device, a magnetic disk, or the like. The processor unit shall be within the expected target range of the scanning control unit 60 that controls the capture of the scanned image, the user interface control unit 63 that communicates with the operator via the user interface 54, and the processed image features. Includes an outline detection unit 62 and an image processor unit 61 that generate an output document 59 containing the checked page image and the page image selectively adjusted by the operator.
The image processor unit 61 processes the scanned image to generate a page image corresponding to the original page of the multi-page document. The outlier detection unit 62 detects whether or not the page image is outlier by determining whether or not the image parameter is outside the target range of the expected value. The outlier detection unit then automatically accepts the page image if it is not an outlier. When the page image is outlier, the page image is displayed on the display 55, allowing the operator to accept or adjust the page image via the user command element 56 on the user interface unit 54.
In system 50, the image processing unit may use a priori knowledge or may be configured to calculate expected values for image parameters as described above. In particular, the expected value may be based on page characteristics obtained during the processing of the scanned image, such as average paper size. System 50 may be configured to build a composite set of page images in output document 59. A composite set has the structure of logical parts that correspond to the page range of a multi-page document, such as chapters and appendices. The local user interface unit 54 has a controllable element, such as an import button, to receive commands from the user to indicate that a subset of scanned images constitutes a logical portion of the composite set of page images. In one embodiment, the system includes a printer unit (not shown) for printing page images or other requested print jobs.
In a practical example, the workflow for generating the output document is as follows. Multi-page documents such as books start in a directory containing tables and scanned images in the form of comma-separated text files. The table provides control data for a series of scanned images, where each line contains several fields: image type, color or black and white, left page page number, right page page number, and both pages. It has the file name of the scanned image including. The next iterative step is performed to process the scanned image. That is, each image is descued (eliminates tilt), both locally adaptive threshold (bi-level) and gray-level images are stored, looking for paper edges and book spines on the scanned image, both left and right. Find the text area of the page, detect out-of-range parameters such as possible errors in orientation and text area selection, pop up the user interface to correct the error, and the paper area of the original multi-page document Cut out the page image from the scanned image by deleting the non-black / gray area, and PDF (Portable Document) The final generation of the file is in a well-known publishing format such as Format) or HTML (Hypertext Markup Language).
For the descued process, for example, Digital Image Processing, page 115, W.Niblack, Prentice Hall The Niblack method described in 1986 generates bi-level images. The Niblack binarization algorithm is a local adaptation method. Mean and standard deviation (stdev) for windows of size (n * n). The window size (n) may be set to, for example, 31. If the text is darker than the background, the following formula is used to calculate the threshold for the center pixel. Threshold = mean (window) )-0.18 * stdev (window). This algorithm is very useful when the image should not be dized or should have low contrast due to bad lighting or due to the original aging. If the background is darker than the text, the -0.18 coefficient must be +0.18. Binary images are used to detect the angle of elements in a scanned image, such as text lines or paper edges. Binary images may also be used for OCR and text area positions.
Various methods for descuing relatively small angles (eg, up to 30 degrees) are commonly used in image processing to calculate angle histograms. The quality of the histogram, such as the lack of obvious peaks, can indicate when the proposed angles are unreliable. Reliability parameters may be derived or used for outlier detection. Note that the dessked page may be upside down, which is not known by the initial skew detection algorithm. Improvements can be achieved by using the previous scanned image, for example, by executing additional rules. That is, if the quality of the histogram is too low, the page is rotated the same as the previous page. If the quality remains low, outliers are detected and the scan results are displayed to the operator for judgment.
As part of the descuing process, page orientation is detected as described above with reference to FIG. Rotation of individual pages or scanned images beyond 90 degrees or 180 degrees with respect to the scanned image is required to achieve an upright orientation of the page image to compensate for positioning the original multi-page document in different ways. To. In some cases, the pages need to be individually descued or may have different orientations. However, the descuing process can fail if the scanned image contains dispersed elements such as numeric pen strokes. The reliability factor to be used for outlier detection may be calculated from the spectrum of angle indications generated during the descubing process.
Detecting the edges of paper and the spine of a book may be performed as follows. The first step is to use a 9x9 kernel to remove text by morphological filtering, closing behavior (basic image processing that fills small openings with subsequent expansion by erosion). A Sobel filter is then used for the image (derivative calculation based on pixel-to-pixel differences in the nxn kernel) to produce an approximate derivative of the image with a strong component at the boundary between the white and black regions. A fixed threshold is then applied to the image to produce a binary candidate book edge. Morphological filters remove most letters and therefore false edges, but the numbers present in the book will still produce false book edge candidates. These candidates are removed by applying cleanup rules, such as the following rules for eight connecting components. If the area of interest is greater than one-fifth of the total area and the aspect ratio is less than 10, the control is an error candidate for the paper edge or book spine and is eliminated. Such a rule removes'blobby'objects and maintains elongated shapes or parts of the outline.
In general, the final image contains only a large number of lines near the book edge and near the spine of the book. To position these lines, the Hough transform is calculated and represents the image in the angular domain as follows: A straight line is used to construct the spectrum of angles and is parameterized in the following form: ρ = xsin (θ) + ycos (θ). Here, ρ is the vertical distance from the origin, and θ is the angle with the normal. Collinear point (x<sub>i</sub>, y<sub>j</sub>), I = 1, ... N, are transformed into N sinusoidal curves in the (ρ, θ) planes that intersect at the points (ρ, θ). On the Hough plane, for example, 10 maximum values near θ = 0 ° give 10 vertical edge candidates, and 20 maximum values near θ = 90 ° give 20 horizontal edge candidates. Since the edge of the spine of the book must also be detected, more horizontal candidates are selected. From this line set, four are selected that form an outer diameter that is close to the expected size. From the set of horizontal lines, candidates that are closer to the middle of the selected top and bottom book edges are selected as book spines. Based on that, two pages are cut out and further processed. If the appropriate line is not selected, or if the proposed cutout page deviates significantly from the target range of page size, outliers are detected and presented to the operator for further processing.
A new way to look for paper areas and middle areas is based on the whiteness of the paper. A gray descued image is used as input, and first, the character object is removed by a closing operation. The result is thresholded by the isodata algorithm to generate, for example, the following binary image.
"Picture Thresholding Using an Iterativ [sic] Selection Method, IEEE Transactions on Systems, Man and Cybernetics, Volume SMC-8, No.8, pp.630-" 2, August 1978 . Histogram is half of maximum dynamic range, t = 2<sup>B-1</sup>Initially segmented into two parts using an initial threshold t such as. The sample average of the gray values related to the foreground pixel (mf) and the sample average of the gray values related to the background pixel (mbkg) are calculated. The new threshold t is now calculated as the average of the averages of the two samples. The process is repeated based on the new threshold until the threshold no longer changes.
The objects present in the binary image are then labeled and the largest one is selected. If the largest object is less than a certain threshold, the second largest object is also selected. Then all the selected objects are copied to the new image. The spar at the edge of the binary object is removed by the morphological opening and the closing fills the hole. The resulting image generally contains only one object, the bounding box of which is measured and used as the paper edge of the book. Gaps can appear in the final paper area mask (eg, due to pictures that are not removed) but do not adversely affect the bounding box of interest.
In the next step, the center of the book (the position of the spine of the book) is determined. Two images, an input image that has been thresholded as described above, and a closing processed image that has been closed using, for example, an 11x11 pattern and filtered by a 3x3 Sobel filter, are for this purpose. Used. Then, some candidates for the center of the book are calculated by finding columns with less than 25 pixel transitions (black to white, or vice versa) in the isodata thresholded image. For these candidates, the center is selected by finding the maximum integrated value of pixels in one column of the Sobel-filtered image.
Note that this step requires a descued source, so this step (especially the central search of a book) will fail if the descued fails. Therefore, the reliability factor for this step may depend on the parameters of the deskew step. The reliability factor is applied to detect if the processed image is outlier and requires operator approval or adjustment.
During processing, the position of the text area on the scanned image may be determined. This paragraph describes the steps required to position the text area of a book page. The calculated value is used to cut out a page image from a black borderless scanned image, or to guide OCR or page number recognition. The basic algorithm only uses the number of column and row transitions. Quality improvements can be made by using layout analysis algorithms. The input image is a black and white descued image and is generated in the descued step described above.
FIG. 6 shows the result of detecting the position of the text area. The figure shows a scanned image 67 and six image parameters including the text area position, where the respective parameters X1, X2, X3, X4, Y1, Y2 indicate the proposed boundaries of the text area. The two text areas 65 and 66 are defined by six coordinate values, in which case the y coordinates Y1 and Y2 of the left text area 65 and the right text area 66 are equal as shown in the figure. Starting from the top of the scanned image, the number of transitions in each row is calculated. When this number exceeds 15, Y1 in the text area is found. The same method is used for Y2. Finding the left and right boundaries of the text area on both pages can be more difficult. This is caused by the fact that a page can contain only a few lines of text. The first occurrence of more than 5 transitions from the left defines X1. Position X4 is located by the first occurrence of more than 25 transitions in the column, starting from the right and more than 30 pixels to the left. Position X2 is positioned with less than the first 5 transitions, with less than 5 transitions at positions more than 10 pixels to the right. The search begins to the right at 1/4 the width of the scanned image. Position X3 is located by the first occurrence of less than 5 transitions, with less than 5 transitions at positions more than 10 pixels to the left. The search begins to the left at 3/4 of the width of the scanned image. Due to a straightforward approach, this step can be erroneous. At least descued images and some text are needed on the page. Some common mistakes are X1, X2, X3 or X4 is the action caused by cutting through some text and cutting out the page number between X2 and X3, or by an empty page. The error detection step detects whether the result should be considered an outlier by calculating a parameter that indicates the likelihood of such an error. For example, an empty page can be detected. If an object in the space between X2 and X3 is detected, the X2 or X3 line may be moved correspondingly to solve the problem, or the result should be displayed to the operator. It may be classified as an outlier.
In the outlier detection unit, the error detection step calculates the error likelihood of the preceding processing step, for example, based on a reliability factor based on histogram quality. Various parameters are used to detect outliers, such as the width of the paper, the spectral characteristics of an area (text or numbers), the quality of different domains such as angles, the size of the object, the color, the whiteness or the contrast. You can. The outlier detection unit may determine additional characteristics of the page image, for example, outliers of the text area position with respect to the paper, the width of the text area. All parameters or characteristics may be compared to the statistical knowledge of the multi-page document collected during processing or the a priori knowledge entered or automatically virtualized by the operator.
In one embodiment, the outlier detector can detect outliers of text area width by imagining regularity with respect to the width of the text area of the page. It should be noted that the top and / or bottom of the left and right text areas may be assumed to be the same, or they may be hypothesized to be different and processed separately. The method used to detect outliers of text area width is expressed by the formula for detecting outliers in the following cases.
<maths num="1"><img file="JP4963809B2_D0001.tif" /></maths>
Where p = text area width and Mp is the median of the text area width. The median is used because there can be very large outliers (eg, zero-width empty pages) that have an excessive adverse effect on the mean. The threshold may be adjusted by the operator to find a working value that substantially detects all errors in the width or position of the text area. In a practical example of a typical book, threshold 14 has been demonstrated to give good results.
Possible errors detected by the outlier function are later displayed to the operator for manual adjustment via the user interface.
The user interface may have the following options for accepting or adjusting the proposed image, such as menus and toolbars on the display screen. The user performs various functions by selecting one of the buttons at the bottom of the image page that is currently being inspected. That is, Orientation: Manual descue by drawing a line that should be horizontal in the pages of the book, or instructions for rotation over 90 or 180 degrees -Paper area: A rectangular section containing the paper area -Paper center: Select the position of the spine of the book · Left page: Select a rectangle that contains the text area on the left page · Right page: Select a rectangle that contains the text area on the right page Going to the next scan suspected of being incorrect, or simply going to the next scan on the line if error detection is not used.
All functions may be performed multiple times as needed for the same image. In the output document, the actions performed by the operator may be recorded and can be viewed again. Therefore, skipping scans marked as outliers is not permanent and can be provided again on restart. The new gray image may be recorded separately if the user performs a manual descuing process.
Finally, the image page is cropped from the scanned image and any descued processing is considered. The double-scanned image is split into two page images at the boundaries of the text area, with an increase of 35 pixels on all sides. This is done to prevent narrow scraping, and there is almost no loss of text in the image. This does not adversely affect the positioning of the image. Finally, a file containing a page image is constructed as a cutout of a series of scanned images. The file may include the number of pages and bookmarks corresponding to the original multi-page document and may be used for optical processing such as Optical Character Recognition (OCR) and adaptive background correction.
A further method of processing a multi-page document involves generating a logical structure in the generated set of page images. Basically, a set of scanned images of a single multi-page document is converted into a composite set of page images, which may be, for example, a single document file. However, the original multi-page document typically has a logical structure, such as a chapter or paragraph, or may include multiple appendices. The original structure is transformed into a similar structure in the composite set, i.e. points to the logical part of the set that corresponds to the page range of the multi-page document. Note that the original structure may be automatically detected, for example, from page numbering or from graphic layout features such as chapter titles in thick and large fonts.
To direct the logical structure, the scanner or processing unit may provide the operator with options for generating the logical structure in the document being scanned. Therefore, the method includes receiving a command from the operator to indicate the structure to be assigned to the series of page images. The command may be given during the operation of a multi-page document. Special buttons may be pressed before and / or after scanning a portion of a multi-page document to indicate that each subset of scanned images constitutes a logical portion of a composite set of page images. Bookmark names, such as import set numbers, may be automatically generated, followed by start and end page numbers. Therefore, it is possible to efficiently scan a structured multi-page document while generating a logical structure in a series of page images to be converted.
FIG. 7 has a number of buttons or keys 72 and a display screen 71 for providing visible data to the operator. The display 71 is of sufficient size and resolution to display the page image obtained from the scan, or at least a portion of the image large enough to determine and adjust the quality of the proposed page image as described above. Is.
In particular, the user interface 70 has an IMPORT button 74 and a START button 73. The START button may be named the OPEN / CLOSE button. The START button opens and closes the scanning operation, and the IMPORT button 74 adds a scanned image of the multi-page portion to the end of the existing set. The origin of each part may be placed as a loose sheet on an automatic document feeder (ADF) or may be manually supplied by the operator by placing a contiguous page of books or documents on the platen. .. After scanning such a portion, the operator presses the IMPORT button to add a further portion to the tail and at the same time define a bookmark in the document file for the just closed portion. Finally, the process may be terminated by pressing the START button and the logical structure of the digital version of the document is completed.
The user interface may also include a special button or menu function that directs the logical portion to start or end on the left or right page of the double-scanned image.
Although the present invention has been described by examples of scanning a book, it should be noted that the present invention is suitable for any multi-page document processing. Moreover, next to the corporate environment, document processing may be of any scale, such as in consumer homes or public commercial services. Moreover, the method of generating the logical structure in the generated series of page images may be applied separately. In this document, the verb'including'and its variants do not preclude the existence of other elements or steps, the singular representation does not preclude the existence of multiple elements, and any reference code claims. The present invention and each unit or means may be realized by appropriate hardware and / or software, and some'units' or'means' are represented by the same item. You may. Furthermore, the scope of the present invention is not limited to the examples, and the present invention lies in each of the above-mentioned novel features and combinations thereof.
<figref num="1">It is a figure which shows the apparatus for processing a document.</figref><figref num="2">It is a figure which shows the scanning method of a multi-page document.</figref><figref num="3">It is a figure which shows an example of the scanned image of a book.</figref><figref num="4">It is a figure which shows the characteristic of the text for detecting the orientation.</figref><figref num="5">It is a figure of the component part of a document processing system.</figref><figref num="6">It is a figure which shows the result of detecting the character area position.</figref><figref num="7">It is a figure which shows the user interface.</figref>
Code description
Page 65,66 image 67 Scanned image
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP09027888A | Cites | Japan |
| JP2004104435A | Cites | Japan |
| JP2002290637A | Cites | Japan |
14 members in 6 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 04077285 | European Patent Office (EPO) | A | |
| 04077285 | European Patent Office (EPO) | A | |
| 040772857 | European Patent Office (EPO) | – | |
| 200404077285 | – | – | – |
| EP20040077285 | – | – | – |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| CN1734469A | China | A | |
| US2006033967A1 | United States of America | A1 | |
| EP1628240A2 | European Patent Office (EPO) | A2 | |
| JP2006054885A | Japan | A | |
| EP1628240A3 | European Patent Office (EPO) | A3 | |
| EP1843275A2 | European Patent Office (EPO) | A2 | |
| EP1843275A3 | European Patent Office (EPO) | A3 | |
| EP1628240B1 | European Patent Office (EPO) | B1 | |
| AT388449T | Austria | T | |
| DE602005005117D1 | Germany | D1 | |
| DE602005005117T2 | Germany | T2 | |
| CN1734469B | China | B | |
| JP4963809B2This record | Japan | B2 | |
| US8564844B2 | United States of America | B2 |
15 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4963809
- Publication, DOCDB
- 4963809
- Publication, EPODOC
- JP4963809B
- Application
- 230964
- Application, DOCDB
- 2005230964
- Application, EPODOC
- JP20050230964
Titles2
- English
- Outlier detection during scanning
- Japanese
- 走査中のアウトライア検出
Classification
- CPC, 6
- G06K9/00469
- G06V30/416
- G06K9/033
- G06V10/987
- G06K9/3208
- G06V10/242
- IPC, 3
- H04N1 00
- G06T1 00
- H04N1 387