Prioritizing Repeated LLM Evaluation for Hidden Failure Discovery
Keita Broadwater, Akin Broadwater
Abstract
Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but operationally important failures. A prompt that produces no failures in a small sample may therefore appear reliable despite having a nonzero latent probability of failure under repeated inference. We formulate LLM reliability evaluation as a budget-constrained discovery problem in which each prompt is associated with an unknown per-generation failure probability. We propose a budgeted discovery framework that first performs shallow evaluation across the prompt set and then uses trial-level failure outcomes together with prompt-derived representations to learn a feature-based ranking of failure propensity. The resulting scores prioritize prompts with zero observed shallow failures for deeper evaluation, concentrating the deep-evaluation budget where hidden failures are more likely to be discovered. We evaluate this approach on AIRBench and StrongREJECT across multiple model and system-prompt conditions. The central empirical test asks whether models fit without access to deep-evaluation outcomes can rank prompts with zero observed shallow failures according to their likelihood of producing failures under deeper evaluation. On AIRBench, the highest-ranked 10\% of unresolved prompts achieves 2.54x hidden-failure lift for Qwen 2.5 7B and 1.87x for Gemma 3n E4B, recovering 25.4\% and 18.7\% of subsequently observed hidden failures, respectively, compared with 10\% expected under random allocation. Semantic-neighborhood and feature-ablation analyses further show that this predictive signal can be recovered from multiple representations of prompt content and relationships.