PaperScope
LIVE · 2026-10-07 05:40 UTC

Symphony for Text Generation: Benchmarking Clinical Note Generation

Daniel Varab, Victor Petrén Bach Hansen, Asbjørn W. Helge, Kevin Pelgrims, Mathias Baltzersen, Adrian Young-San Roessler, Vanessa Klungtvedt, Maximilian Brand, Lasse Krogsbøll, Henrik Cullen, Lars Maaløe

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.08161 v1
Category
Submitted
2026-10-06

Abstract

Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.

arXiv abs page · PDF · same-day batch