PaperScope
LIVE · 2026-09-17 05:40 UTC

Predicting Human Disagreement for Calibrated Dynamic Facial Expression Recognition

Yiming Wang, Frederick W. B. Li, Jingyun Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.17130 v1
Category
Submitted
2026-09-15

Abstract

Dynamic facial expression recognition (DFER) benchmarks such as DFEW provide multiple annotator votes per clip, yet most models collapse them to a majority label and cannot represent human disagreement at inference time. We propose a disagreement-aware DFER framework that trains directly on the raw annotator count vector using a Dirichlet-Multinomial likelihood. Unlike mean-only soft-label objectives, the proposed likelihood provides scale-sensitive supervision for the Dirichlet concentration while preserving the predictive mean. A separate ambiguity head predicts annotation entropy for unseen clips, and a monotone Chow-style reject rule combines predicted ambiguity, vacuity, temporal instability, and input quality for selective prediction. On DFEW, the method preserves recognition accuracy while reducing ECE by 30% and AURC by 15%, and predicted ambiguity reaches a Spearman correlation of 0.52 with the annotation entropy of test clips. The calibration and selective-prediction gains transfer to FERV39k and remain under identity- and movie-disjoint DFEW splits.

Comment: 5 pages, 3 figures, 4 tables. Submitted to ICASSP 2027

arXiv abs page · PDF · same-day batch