PaperScope
LIVE · 2026-09-29 05:40 UTC

When Should the Count Change? Learning State Maintenance for Causal Video Counting

Pengyiang Liu, Dongyue Lyu, Junbo Niu, Zhongyue Shi, Jiahao Xie, Si Liu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35416 v1
Category
Submitted
2026-09-28

Abstract

Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain visibility, persistent identities, and completed-event records. A differentiable recurrence trains event transitions over legal paths constrained by count endpoints; visibility and association objectives train the object branch. A multi-source pipeline organizes 39.8K spatial queries and complementary event annotations into counting trajectories. On SVCBench, we evaluate counting adaptation with partial video overlap and held-out groups of linked annotations. Under prefix replay (Full) and persistent streaming (Stream), 4B and 8B models reach 41.9/36.4 and 44.9/38.2 Gaussian Precision Accuracy, respectively. The 8B model gains 10.9/3.2 points over Counting-SFT on the same queries. Matched-graph comparisons isolate phase conditioning and trajectory supervision, assessing training objectives alongside hard decisions. Online video benchmarks and count-conditioned decisions assess online understanding and task eligibility. Project Page: https://PLACEHOLDER.github.io/StaMina/

Comment: 28 pages, 7 figures. Project Page: https://PLACEHOLDER.github.io/StaMina/

arXiv abs page · PDF · same-day batch