PaperScope
LIVE · 2026-09-30 05:40 UTC

CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data

Philipp E. Glass, Alina Miron

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.37807 v1
Category
Submitted
2026-09-29

Abstract

Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fulfil the requested task. But existing annotation covers evaluation sets of a few thousand prompts at most. We present CompOrca, a compliance labelling over the entirety of the 4,233,923-example OpenOrca corpus. Every example was classified as compliant or noncompliant by five independent passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters), and the corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%) along with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, allowing for filtering the most ambiguous samples. Against 450 human-annotated examples, 150 of them annotated twice (human-human $κ= 0.93$), the unanimous compliance and noncompliance labels are 97.3% and 86.7% precise, the latter a high-precision subset, not a complete enumeration, of noncompliance. Published refusal-detection methods recall only between 0.4% and 94.1% of the noncompliance class. We release the full corpus with its per-row labels and vote counts at https://huggingface.co/datasets/cemiu/CompOrca

Comment: Accepted to PlurVA-LLM Workshop @ AACL-IJCNLP 2026. Dataset available on HuggingFace

arXiv abs page · PDF · same-day batch