PaperScope
LIVE · 2026-09-03 05:40 UTC

Generating Clinical Vignettes that Preserve Cognitive Formulations

Amit Oren, Nimrod Hertz-Palmor, Dean Ariel, Guy Laban

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.29995 v1
Category
Submitted
2026-08-30

Abstract

Large language models can generate fluent clinical case vignettes, but fluency alone does not ensure fidelity to a specifiable clinical structure. We introduce FORMA, a theory-grounded framework that compiles a cognitive model of a disorder into a directed weighted graph, samples a person-specific configuration of that graph, and validates whether the generated vignette preserves the specified components and causal links. We instantiate FORMA on Posttraumatic Stress Disorder using the Ehlers and Clark cognitive model, generating 16,500 vignettes across 500 personas, 11 generation models, and three ablation conditions. Evaluation combines an external edge-recovery probe, two clinical experts, a scaled LLM judge, and a clinician user study with 100 licensed practitioners. The cognitive graph is recoverable from full-condition vignettes (MCC = +0.41, AUC = 0.70) but not from zero-shot generation (MCC = +0.01, AUC = 0.50). Experts rate full vignettes substantially higher than zero-shot alternatives, and clinicians perceive them to be human-written 85% of the time, compared with 22% for zero-shot. FORMA also reduces demographic disparity in perceived quality by 1.5-7x. These results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation. A repository with the data and code is available online: https://github.com/Amit-Oren/FORMA.

Comment: Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026. Code and data: https://github.com/Amit-Oren/FORMA

arXiv abs page · PDF · same-day batch