Skip to main content
Article
Digitization Decisions: Comparing OCR Software for Librarian and Archivist Use
Code4Lib (2021)
  • Leanne Olson, Western University
  • Veronica Berry, Western University
Abstract
This paper is intended to help librarians and archivists who are involved in digitization work choose optical character recognition (OCR) software. The paper provides an introduction to OCR software for digitization projects, and shares the method we developed for easily evaluating the effectiveness of OCR software on resources we are digitizing.

We tested three major OCR programs (Adobe Acrobat, ABBYY FineReader, Tesseract) for accuracy on three different digitized texts from our archives and special collections at the University of Western Ontario. Our test was divided into two parts: a word accuracy test (to determine how searchable the final documents were), and a test with a screen reader (to determine how accessible the final documents were). We share our findings from the tests and make recommendations for OCR work on digitized documents from archives and special collections.
Keywords
  • digitization,
  • ocr,
  • optical character recognition,
  • archives,
  • special collections
Publication Date
September 22, 2021
Citation Information
Leanne Olson and Veronica Berry. "Digitization Decisions: Comparing OCR Software for Librarian and Archivist Use" Code4Lib Iss. 52 (2021) ISSN: 1940-5758
Available at: http://works.bepress.com/leanne-olson/15/