PaperScope
LIVE · 2026-09-22 05:40 UTC

Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control

Simiao Ren, Ankit Raj, Tommy Duong, Yuxin Zhang, Dennis Ng, Xingyu Shen, Kidus Zewde, Yuchen Zhou, Neo Tiangratanakul

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.23953 v1
Category
Submitted
2026-09-20

Abstract

AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what that means for a relying party -- an insurer, a lender, an auditor -- whose evidence is a filed PDF. AgentForge-Bench measures how reliably an off-the-shelf coding agent, driving one of seven open-weight models with a shell and the stock Python PDF stack, alters one dollar amount, date or address in a real filed financial document from a single sentence of intent, graded by rules rather than by a model. Across 1,750 cells, 1,419 (81.1%) satisfy the verifier, and 808 (46.2%) also survive every stricter filter: visible, localized, typeface-matched, original value gone document-wide. A deterministic script with no model in it solves 98 of the 125 documents; the agents solve 124, and none the script solves alone. Agents misreport 41% of their wrong edits as done, no model refused, and the cheapest verified forgery costs 2.4 cents. The raw rate overstates the threat by about a factor of two; the strict rate is still large.

Comment: 15 pages, 12 figures, 2 tables

arXiv abs page · PDF · same-day batch