Navigating Long S’s and Garbled Full Texts: Searching Strategies for Online 18th-Century Texts

Saturday, January 9, 2027
Grand Ballroom (New Orleans Marriott)
Valerie Sallis, Georgetown University
Searchable full-text digitizations of periodicals and texts have revolutionized historical research in the twenty first century. However, despite the progress made in the last few decades, the Optical Character Recognition (OCR) software used to generate full-text scans of works printed in the hand press era still often struggles to interpret older texts correctly. Issues like the use of now nonstandard characters (the long s, ligatures, etc.), worn type, or blurry images scanned from microfilm versions can cause digital full text transcriptions to be riddled with errors. Problems with full-text versions of eighteenth-century texts due to OCR are common in all major subscription and free databases. The quality of full-texts can vary widely even within the same database depending on when items were processed as digitization projects now span multiple decades.

While the seriousness of this issue for historical research is not often highlighted, inaccuracies within searchable texts can cause historians to miss relevant texts for their research due to a database’s failure to read and search older texts correctly. This poster draws the presenter’s background in library and information science to survey the current state of OCR and eighteenth-century texts and common problems found in multiple databases. The bulk of the poster will center on a digital humanities searching strategy for working around issues with OCR developed by the presenter for her dissertation research tracing the trans-Atlantic republication of eighteenth-century ephemeral sources.

This poster will explore several examples of how this method works applied to eighteenth-century newspaper and almanac articles. Along with explanatory text, the poster will include photos of common OCR problems with eighteenth-century texts and examples searches as well as graphs of statistical information comparing the numbers of search results returned when using the presenter’s method compared to conducting traditional shorter keyword searches. Images will show how the method involves selecting several search phrases of three to five short words from varying points within the text, avoiding common expressions, and looking for unusual combinations of words unlikely to appear in other texts in order to limit false matches. Searches are then made for three to five such phrases (with at least one taken from the beginning, middle, and end of each passage) to compensate for possible confounding results due to errors in OCR. While errors are very commonly found sporadically in OCR full text versions of eighteenth-century sources, they only rarely affect the entirety of the text, so searching for multiple phrases throughout a passage compensates for the high error rate expected on single searches. Frequently, the first several searches for each text returned different results within the same databases, publications, and even at times within the same digitized items—all of which indicate the persistence of major gaps in OCR quality. However, by the fourth or fifth search, typically searches returned only results found in previous searches. Graphs will be presented showing that the multiple searching technique compensated for gaps in OCR quality and returned both higher numbers of results and better-quality matches than traditional keyword searches.

See more of: Poster Session #2
See more of: AHA Sessions