Skip to main content
Unpublished Paper
A Fast Alignment Scheme for Automatic OCR Evaluation of Books
(2011)
  • Ismet Zeki Yalniz
  • R. Manmatha, University of Massachusetts - Amherst
Abstract

This paper aims to evaluate the accuracy of optical character recognition (OCR) systems on real scanned books. The ground truth e-texts are obtained from the Project Gutenberg website and aligned with their corresponding OCR output using a fast recursive text alignment scheme (RETAS). First, unique words in the vocabulary of the book are aligned with unique words in the OCR output. This process is recursively applied to each text segment in between matching unique words until the text segments become very small. In the final stage, an edit distance based alignment algorithm is used to align these short chunks of texts to generate the final alignment. The proposed approach effectively segments the alignment problem into small subproblems which in turn yields dramatic time savings even when there are large pieces of inserted or deleted text and the OCR accuracy is poor. This approach is used to evaluate the OCR accuracy of real scanned books in English, French, German and Spanish.

Keywords
  • OCR evaluation,
  • sequence alignment,
  • digital libraries
Disciplines
Publication Date
2011
Comments
This is the pre-published version harvested from CIIR.
Citation Information
Ismet Zeki Yalniz and R. Manmatha. "A Fast Alignment Scheme for Automatic OCR Evaluation of Books" (2011)
Available at: http://works.bepress.com/r_manmatha/42/