PaperScope
LIVE · 2026-10-06 05:40 UTC

MERCI Cards: An LLM Evaluation and Deployment Framework for High-Stakes Domains

Aparna Komarla, Annalisa Szymanski

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04430 v1
Category
Submitted
2026-10-03

Abstract

As LLMs are increasingly deployed in high-stakes professional workflows, engineers and researchers require principled protocols to systematically track, monitor, and improve model performance across deployment cycles. We present a mathematical framework for iterative LLM evaluation and deployment, and demonstrate its application to AI systems used in criminal justice. Our framework formalizes LLM integration in high-stakes, high-risk, and resource-constrained domains across model selection, rubric design, evaluations and deployment via a weighted multi-objective optimization. We demonstrate that MERCI Cards can guide improvements of the system across deployment iterations, direct developer attention toward under-performing areas, and focus user attention on validation and error-correction in the LLM's outputs.

arXiv abs page · PDF · same-day batch