PaperScope
LIVE · 2026-09-15 05:40 UTC

From Advertised Improvements to Measured Capabilities: Evaluating ChatGPT Images 2.5 on Forgery Tasks

Ankit Raj, Yuxin Zhang, Kidus Zewde, Tommy Duong, Jiaqi Gan, Xingyu Shen, Yuchen Zhou, Huaiyu Guo, Siyu Zhang, Simiao Ren

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.13617 v1
Category
Submitted
2026-09-12

Abstract

We evaluate whether the improvements advertised for ChatGPT Images 2.5 translate into better performance on forgery tasks with predetermined answers. We compare its Flare and Sunburst API models with GPT-Image-2 re-run in the same week, using receipt-field edits, repeated editing, product placement and fine-print rendering. After image registration, Flare and Sunburst show fewer OCR-detected changes to surrounding receipt text (31.7% and 31.2% versus 44.2% for both GPT-Image-2 baselines), mainly on CORD receipts, without a detectable improvement in target-field correctness. Flare retains fewer earlier edits on CORD receipts, while photo-edit sequences provide little separation between models. Product codes are more often legible with Images 2.5, alongside larger product placement; the analyses do not establish a fidelity gain independent of size. Fine-print improvements remain unresolved below the OCR reliability limit. Refusals are rare and localisation is weak in both generations. At a fixed detection threshold, Community Forensics flags 68.6% of controlled Images 2.5 images averaged across cells, versus 35.9% of self-reported images posted online. These results motivate task-specific evaluation of advertised capabilities and defences, with explicit limits on what automatic checks can establish.

Comment: 26 pages, 6 figures. Ancillary files include the scored manifest for the in-the-wild sample and an arena snapshot

arXiv abs page · PDF · same-day batch