AI Models Explain Early Word Recognition but Not Complex Reading, New Study Finds
Published on August 10, 2026 in the Proceedings of the National Academy of Sciences, the research tracked the eye movements of 368 adults as they read carefully constructed sentences, including garden‑path sentences that are grammatically correct yet misleading, forcing readers to backtrack and re‑parse.
The team compared the eye‑tracking data with predictions from more than 400 LLMs trained to anticipate the next word in a sequence. "Language models develop their remarkable language understanding capabilities by being trained to predict the next word in a sentence," said William Timkey, a doctoral student at NYU and the study’s lead author. "That led us to ask whether the same predictive processes that drive these AI systems could also explain how humans comprehend sentences."
The results were clear: the time it takes readers to recognize a word from its letters aligns well with the models’ next‑word predictions. In other words, LLMs capture the first stage of reading—identifying the word.
However, when the sentence requires integrating that word into a broader context—especially in garden‑path sentences that prompt a reread—the models’ predictions fall far short. "The predictability of a word really doesn’t even come close to explaining just how much time we spend on difficult words and garden‑path sentences," Timkey said. "LLMs were drastically underpredicting the type of difficulty that we experience when reading."
Tal Linzen, an associate professor at NYU and co‑author, added, "It’s in that second stage of processing—recognizing a word and then integrating it with other words in passages—where we find big gaps between what word predictability can explain and what we need cognitive models to explain."
Brian Dillon, a professor of linguistics at UMass Amherst and senior author, emphasized the broader implications: "We now know a little bit better how humans and models are different," he said. "That is the first step in understanding how we can close that gap, which we want to do because that could have enormous advantages down the road."
The authors caution that while LLMs are valuable tools for cognitive science, they are not sufficient to model the full complexity of human reading. Human readers exhibit backward eye movements and rereading behaviors that current models—reliant solely on next‑word prediction—cannot capture.
The research was funded by several National Science Foundation grants (BCS‑2020914, BCS‑2020945, IIS‑2504953, and IIS‑2504954). Contact for the study is Signe Dugger at UMass Amherst (sdugger@umass.edu, 413‑545‑0146).
Looking ahead, the study suggests that future models of human language processing will need mechanisms beyond simple next‑word prediction, possibly incorporating hierarchical parsing or memory‑based components. Such advances could improve language learning tools and aid in diagnosing and treating reading disorders, while the partial alignment between human and AI reading processes offers a roadmap for researchers seeking to bridge the gap between artificial and biological language comprehension.