PaperScope
LIVE · 2026-09-29 05:40 UTC

What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review

Shuyang Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33153 v1
Category
Submitted
2026-09-27

Abstract

Large language model (LLM) agents increasingly draw on external tools and reusable skills selected at run time from libraries that hold thousands of entries. Reports that a retriever, router, or skill library "improves" an agent may refer to retrieval recall, the success change from enabling a library, a paired contrast restricted to triggered tasks, or a gain under an approximate budget constraint. This critical review asks what each design compares and under which assumptions. Building on estimand-based approaches to agent evaluation, we describe tool and skill designs along six axes: treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Thirteen core empirical studies anchor the evidence synthesis, supplemented by related methodological work and design-level reading of the wider literature. Our contribution is to make explicit distinctions that some source authors already acknowledge through a decomposition of trigger-conditioned pairing, analytic counterexamples, and comparisons across studies. Pairing on the task does not by itself identify an invocation effect; paired gain and regression counts measure protocol-specific discordance rather than the share of tasks whose expected outcomes worsen; and total effects of module deployment answer a different question from budget-constrained efficiency. We compare curated skill provision with retriever replacement, triggered subsets with all-task outcomes, and observed cost reductions with budget-constrained comparisons. A reporting checklist and worked examples connect these distinctions to information that studies can report. The review runs no new experiments; empirical results come from the cited studies, and numerical toy examples are analytical illustrations.

Comment: 30 pages, 1 figure. Critical narrative review

arXiv abs page · PDF · same-day batch