PaperScope
LIVE · 2026-09-04 05:40 UTC

Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

Divyesh Bommana, Mohammad Saim, Tianyu Jiang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.03394 v1
Category
Submitted
2026-09-03

Abstract

Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.

Comment: Accepted to EMNLP 2026 (Main Conference) Dataset and code: https://github.com/cincynlp/Chiaro

arXiv abs page · PDF · same-day batch