PaperScope
LIVE · 2026-09-29 05:40 UTC

Native Association: Confidence-Aware Human Perception in the Wild with a Foundation VLM

Igal Dmitriev, Ofir Liba

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33450 v1
Category
Submitted
2026-09-27

Abstract

Extracting who is where, on which team, wearing which number from a broadcast frame is typically done by stitching a detector, an OCR engine, and classifiers together -- and the stitching step swaps identities under occlusion. We make association native instead: a 0.77B vision-language model (Florence-2) is fine-tuned to emit all per-person attributes as one grammar-constrained sequence, with each attribute generated inside its owner's block. Output is therefore schema-valid on every frame by construction, and no post-hoc binding step exists to attach a correctly read number to the wrong player: residual misassociation is pure perception error, $\approx4\times$ rarer than zero-shot-prompted frontier APIs' (0.057 vs. 0.21-0.24). On a frozen multi-sport test set, this single pass reaches 0.95 detection F1 (APIs: 0.65-0.75). A single extra forward pass yields a per-field confidence that supports a reject option (jersey precision $0.71\rightarrow0.96$ at half coverage) and routes a training-free zoom-and-re-read for small players. Surprisingly, once the grammar is learned, further parameter-efficient tuning yields no measurable gain under the adaptation configurations we test; the identical recipe on WIDER-Attribute reaches 93.1 mAP given-box, yields the first detection-coupled end-to-end results under its standard test protocol (84.5 mAP), and reproduces the same tuning result. In this regime, the gains live in the structure, not in added weights.

Comment: Accepted at the ECCV 2026 Workshop on Human-Centered Multimodal Intelligence in the Wild (HCMIW). 16 pages, 1 figure, 2 tables; supplementary material included as ancillary file

arXiv abs page · PDF · same-day batch