PaperScope
LIVE · 2026-09-18 05:40 UTC

E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews

Haoshen Wang, Dongbo Che, Zeyi Xie, Yuanjie Du, Shicheng Hua, Xingyu Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.20001 v1
Category
Submitted
2026-09-17

Abstract

Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.

arXiv abs page · PDF · same-day batch