PaperScope
LIVE · 2026-09-22 05:40 UTC

Chronologic: Measuring Language Models' Ability to Represent the Past

Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Edwin Roland, Wenyi Shang, Matthew Wilkens

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23178 v1
Category
Submitted
2026-09-19

Abstract

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct answers. We use historical texts to develop a benchmark for a model's representation of English-language contexts 1831-1930, relying on pairwise comparisons to multiple ground truths and strong distractors to score the hardest questions in an appropriately graduated way. We find that generative tasks are harder than discriminative ones; in fact, reasoning models can typically discern the weakness of their own generated answers. While models pretrained exclusively on historical text lead the pack when evaluated by answer likelihood, they cannot compete with commercial models in free generation. None of the models we tested represent historical contexts in a fully persuasive way yet, but progress toward that goal is evident.

Comment: 22 pages, 3 figures, 8 tables

arXiv abs page · PDF · same-day batch