Three steps, three results
Scanning creates a digital image of paper. OCR recognizes the text in that image. Data extraction identifies values such as dates, partner names or case identifiers. Classification determines the document type and its possible relationship to a case. Correctly recognized text alone therefore does not prove that the document has been assigned to the right case.
What affects recognition?
The sharpness of the image, the turning of the pages, poor contrast and missing parts can also make processing difficult. In the case of different languages, tables or handwriting, separate samples should be examined. The printed PDF and the scanned image are a different starting point: first, it must be clarified whether extractable text is already available.
What does an identification proposal mean?
The system suggests a possible value or document type. In NOPA's document sorting process, the clerk's control helps ensure that the document goes to the right place. Uncertain identifiers, several possible partners or missing links should be treated as a separate verification task, not filled in by guesswork.
Example of an ambiguous partner name
An incoming letter contains both the name of the sender and the name of the recipient. OCR can recognize both correctly, while the partner to be assigned to the case is still unclear. The administrator compares the proposal with the content of the document and the available master data. In the example, the object of the check is the meaning, not just the character recognition.
How should you test the process?
Compile a representative sample of common documents, less common layouts, and examples of known errors. Record in advance the correct type, the data to be searched and the expected business relationship. Separately measure field accuracy, classification correctness, and time spent on human correction. A single aggregate accuracy number can mask errors in critical identifiers.
What should happen to the exceptions?
Be responsible for illegible, incomplete or unidentifiable documents. After the correction, the relationship between the original document and the final classification should be checked. During the survey, we record the necessary data fields, the control points and the automation that can be used in the given introduction.
Digitization: what should a quote include?
Contact