Digitization Decisions: Comparing OCR Software for Librarian and Archivist UseCode4Lib (2021)
This paper is intended to help librarians and archivists who are involved in digitization work choose optical character recognition (OCR) software. The paper provides an introduction to OCR software for digitization projects, and shares the method we developed for easily evaluating the effectiveness of OCR software on resources we are digitizing.
We tested three major OCR programs (Adobe Acrobat, ABBYY FineReader, Tesseract) for accuracy on three different digitized texts from our archives and special collections at the University of Western Ontario. Our test was divided into two parts: a word accuracy test (to determine how searchable the final documents were), and a test with a screen reader (to determine how accessible the final documents were). We share our findings from the tests and make recommendations for OCR work on digitized documents from archives and special collections.
- optical character recognition,
- special collections
Publication DateSeptember 22, 2021
Citation InformationLeanne Olson and Veronica Berry. "Digitization Decisions: Comparing OCR Software for Librarian and Archivist Use" Code4Lib Iss. 52 (2021) ISSN: 1940-5758
Available at: http://works.bepress.com/leanne-olson/15/