PaperScope
LIVE · 2026-10-06 05:40 UTC

When Verifiable Counts Depend on Wording: Auditing Wording Robustness in Instruction Following

Qishi Zhan, Seoyeon Jang, Zihan Dong, Minxuan Hu, Ziheng Chen, Tonghui Qu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05278 v1
Category
Submitted
2026-10-04

Abstract

Verifiable instruction-following benchmarks often express each constraint through one fixed template. We test whether scores remain stable when the operational requirement is unchanged but its wording varies. We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range. Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts. In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially: in the nine-model control panel, compliance is 54.9% for the original positive form, 48.2% for a longer positive form, 36.7% when the target appears later, and 33.8% for AVOID1. A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect. Effect sizes, failure directions, weakest forms, and model rankings vary across realizations. Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse. Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior. WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.

arXiv abs page · PDF · same-day batch