PaperScope
LIVE · 2026-09-29 05:40 UTC

How Well Can LLMs Simulate Real Learner Evaluations of Educational Feedback?

Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara, Kyosuke Takami

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35376 v1
Category
Submitted
2026-09-28

Abstract

While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback for high-school biology questions at both the group and individual levels. We compare performance with and without learner-specific information, such as personality traits and evaluation examples, across six models. Our results show that LLMs still have a limited ability to simulate learner evaluations. Providing learner profiles and examples improves score calibration and individual-level simulation, but more often fails to improve group-level consistency. These findings highlight the need to investigate which learner information and adaptation strategies are effective for learner preference simulation.

Comment: Accepted to the EMNLP 2026 Main Conference

arXiv abs page · PDF · same-day batch