PaperScope
LIVE · 2026-09-03 05:40 UTC

Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration

Sunwhi Kim, Sunyul Kim, Meounggun Jo, Jini Tae

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.30210 v1
Category
Submitted
2026-08-31

Abstract

AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.

Comment: 35 pages, 8 figures, 5 supplementary figures. Human reference data reused (not newly collected) from arXiv:2603.24048. Data and code: https://doi.org/10.5281/zenodo.22148304

arXiv abs page · PDF · same-day batch