What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review
Shuyang Zhang
Abstract
Large language model (LLM) agents increasingly draw on external tools and reusable skills selected at run time from libraries that hold thousands of entries. Reports that a retriever, router, or skill library "improves" an agent may refer to retrieval recall, the success change from enabling a library, a paired contrast restricted to triggered tasks, or a gain under an approximate budget constraint. This critical review asks what each design compares and under which assumptions. Building on estimand-based approaches to agent evaluation, we describe tool and skill designs along six axes: treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Thirteen core empirical studies anchor the evidence synthesis, supplemented by related methodological work and design-level reading of the wider literature. Our contribution is to make explicit distinctions that some source authors already acknowledge through a decomposition of trigger-conditioned pairing, analytic counterexamples, and comparisons across studies. Pairing on the task does not by itself identify an invocation effect; paired gain and regression counts measure protocol-specific discordance rather than the share of tasks whose expected outcomes worsen; and total effects of module deployment answer a different question from budget-constrained efficiency. We compare curated skill provision with retriever replacement, triggered subsets with all-task outcomes, and observed cost reductions with budget-constrained comparisons. A reporting checklist and worked examples connect these distinctions to information that studies can report. The review runs no new experiments; empirical results come from the cited studies, and numerical toy examples are analytical illustrations.