PaperScope
LIVE · 2026-09-24 05:40 UTC

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

Felix Ringe

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.26945 v1
Category
Submitted
2026-09-22

Abstract

Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA). Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.

Comment: accepted at the ICML 2026 AI4Law Workshop; 32 pages (main text 9 pages + appendices 23 pages)

arXiv abs page · PDF · same-day batch