Reading the Past Through the AI Lens: A Comparative Analysis of Large Language Models Interpreting 19th-Century American Travel Narratives

Saturday, January 9, 2027
Grand Ballroom (New Orleans Marriott)
Tyler Gillis, University of Central Florida
As artificial intelligence tools become increasingly accessible, historians face urgent questions about their utility and limitations for primary source analysis. This poster presents findings from an empirical study comparing how four major Large Language Models (ChatGPT, Claude, Gemini, and Microsoft Copilot) interpret a late nineteenth-century American travel narrative. Rather than theorizing about AI's potential, this research provides concrete, comparative data on what happens when these tools encounter historical primary sources.

The study employs a structured experimental design testing each LLM under three conditions: minimal context (source text only), partial context (basic provenance information), and full historical framing (author background, genre conventions, and historiographical relevance). This methodology reveals not only what each AI system "sees" in historical documents, but how their interpretations shift, or fail to shift, when provided with the contextual knowledge historians bring to source analysis.

The poster's visual design facilitates direct comparison across models and conditions. The central panel displays excerpts from the primary source alongside color-coded AI responses, allowing viewers to immediately identify patterns in interpretation, emphasis, and omission. Side panels present systematic analysis using a coding framework that tracks: accuracy of factual claims, recognition of genre conventions, handling of period-specific language and assumptions, imposition of anachronistic moral frameworks, and acknowledgment of interpretive uncertainty. Heat maps visualize performance across these categories, making cross-model patterns immediately legible.

A dedicated methodology section illustrates the experimental protocol with flowcharts and sample prompts, enabling other researchers to replicate or extend this work. The poster also includes a "failure analysis" section highlighting specific instances where AI systems confidently produced misleading or historically problematic interpretations, crucial data for understanding the risks of uncritical AI adoption.

Intellectually, this poster argues that the question "Can historians use AI?" is inadequately framed. The more productive questions are: Under what conditions? For what purposes? With what safeguards? The comparative data demonstrates that AI systems vary significantly in their handling of historical sources, that contextual framing substantially affects output quality, and that even sophisticated LLMs exhibit systematic blind spots that align with broader concerns about algorithmic bias and the flattening of historical complexity.

See more of: Poster Session #1
See more of: AHA Sessions