PaperScope
LIVE · 2026-10-08 05:40 UTC

Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness

Yihan Li, Hanyi Zhang, Xiaoxi Jiang, Man Guo

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.09489 v1
Category
Submitted
2026-10-07

Abstract

Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.

Comment: 20 pages, 4 figures, 11 tables. Accepted to the main conference of EMNLP 2026

arXiv abs page · PDF · same-day batch