Does CoT-Pass@k Really Check the CoT? A Multilingual Mathematical Audit
Tarık Tuna Taşaltı, Burcu Hüdaverdi, David Semedo
Abstract
Pass@k measures whether a model reaches a correct answer under repeated sampling, but never how: a lucky guess counts the same as sound reasoning. CoT-Pass@k was proposed to close that gap, adding an LLM-as-judge that must assess a solution's reasoning chain before it counts. Its value rests entirely on one assumption: that the judge catches flawed reasoning. That assumption has never been tested inside the metric that depends on it, and never outside English, though the metric's claims concern models used in many languages. We report the first audit of that verification step, run under the metric's own protocol on a multilingual suite of five mathematical benchmarks in English, Turkish and Portuguese, two of them natively written. We corrupt correct solutions with deterministic edits that damage the chain and the final answer separately. We observe that all three judges accept corrupted chains almost as often as clean ones. V4-Flash and Qwen3.6 reject a solution sharply only when its final answer is wrong and accept a wrong answer more readily when the chain agrees with it; the metric's own judge accepts most wrong answers as well. Our study shows that chain-answer agreement dominates the two larger judges' verdicts and that all three fail to reliably detect the tested reasoning errors. Consequently the difference Pass@k - CoT-Pass@k averages 19.7 points on an earlier solver generation but only 4.1 on the current one. What little remains depends on the token budgets on both sides and on the generation mode; raising the generation budget moves Pass@64 by more than fifty points while the difference stays at zero. We close with two checks any judged reasoning metric should pass before its numbers are read as evidence about reasoning.