PaperScope
LIVE · 2026-09-11 05:40 UTC

Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion Recognition

Shirong Lyu, Silu Quan, Yixuan Ding, Chengpeng Wang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.11680 v1
Category
Submitted
2026-09-10

Abstract

3D skeleton-based gait emotion recognition faces high annotation costs, data scarcity, and poor generalization on heterogeneous data. This paper proposes SV-GCN, a single-stream multi-feature fusion framework with temporal invariance. We introduce intra-frame relative motion features to eliminate frame-rate sensitivity and embed heterogeneous cues at shallow layers, enabling early fusion without multi-stream complexity. For variable-length sequences, we design a global mask-guided valid-frame spatio-temporal graph convolution module, introducing frame-rate insensitivity for the first time in this domain. On the E-Gait dataset, our method achieves performance comparable to state-of-the-art while demonstrating strong generalization across varying sequence lengths and frame rates, offering a viable pathway for pre-training on large-scale skeleton-based action recognition datasets.

Comment: Accepted at the 35th International Conference on Artificial Neural Networks (ICANN 2026)

arXiv abs page · PDF · same-day batch