PaperScope
LIVE · 2026-09-15 05:40 UTC

DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

Hongye Yang, Zhihao Xie, Shengjun Xiong, Boxiao Huang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.15122 v1
Category
Submitted
2026-09-14

Abstract

Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as a direct route to more reliable evaluation. Under a fixed budget, however, auditing each program more thoroughly reduces the number of tasks and independent generations that can be evaluated, which can ultimately make model-level estimates less accurate. We study this phenomenon and the conditions under which it arises. We decompose behavioral evaluation into three evidence levels: task templates, stochastic generations, and within-program edits. We define an average failure risk that is invariant to audit depth, and combine three-level variance with measured execution costs to analyze the tradeoff between deeper edit auditing and broader independent coverage. Experiments across two CAD environments and five generation systems show that the value of deeper auditing depends on where evaluation uncertainty originates. When template heterogeneity or generation stochasticity dominates, additional edit checks can increase total estimation error; when within-program state variation is large and generation is expensive, deeper auditing is more valuable. Variance and cost estimates from calibration predict the direction of this change and provide a diagnostic basis for allocating evidence on held-out tasks. These results show that the thoroughness of program inspection can diverge from the reliability of model evaluation, and they help determine whether the next unit of budget should be spent on a new task, a new generation, or additional edit checks.

Comment: 29 pages, 3 figures, 20 tables. Preprint

arXiv abs page · PDF · same-day batch