PaperScope
LIVE · 2026-09-17 05:40 UTC

Challenges of Auditing: Variability in Outputs of Large Language Models for Health

Yuan Pu, Yewon Chang, Furong Jia, Xunjian Yin, Jessica Ma, Ayman Ali, Monica Agrawal

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.16590 v1
Category
Submitted
2026-09-15

Abstract

People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.

arXiv abs page · PDF · same-day batch