While the seriousness of this issue for historical research is not often highlighted, inaccuracies within searchable texts can cause historians to miss relevant texts for their research due to a database’s failure to read and search older texts correctly. This poster draws the presenter’s background in library and information science to survey the current state of OCR and eighteenth-century texts and common problems found in multiple databases. The bulk of the poster will center on a digital humanities searching strategy for working around issues with OCR developed by the presenter for her dissertation research tracing the trans-Atlantic republication of eighteenth-century ephemeral sources.
This poster will explore several examples of how this method works applied to eighteenth-century newspaper and almanac articles. Along with explanatory text, the poster will include photos of common OCR problems with eighteenth-century texts and examples searches as well as graphs of statistical information comparing the numbers of search results returned when using the presenter’s method compared to conducting traditional shorter keyword searches. Images will show how the method involves selecting several search phrases of three to five short words from varying points within the text, avoiding common expressions, and looking for unusual combinations of words unlikely to appear in other texts in order to limit false matches. Searches are then made for three to five such phrases (with at least one taken from the beginning, middle, and end of each passage) to compensate for possible confounding results due to errors in OCR. While errors are very commonly found sporadically in OCR full text versions of eighteenth-century sources, they only rarely affect the entirety of the text, so searching for multiple phrases throughout a passage compensates for the high error rate expected on single searches. Frequently, the first several searches for each text returned different results within the same databases, publications, and even at times within the same digitized items—all of which indicate the persistence of major gaps in OCR quality. However, by the fourth or fifth search, typically searches returned only results found in previous searches. Graphs will be presented showing that the multiple searching technique compensated for gaps in OCR quality and returned both higher numbers of results and better-quality matches than traditional keyword searches.