PaperScope
LIVE · 2026-10-05 05:40 UTC

When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity

Mayand Gulati, Kerong Wang, WeiChen Au

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.03646 v1
Category
Submitted
2026-10-02

Abstract

Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's probability of ever departing from OTS is at most the chosen $α_E$, without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by $38.4\%$ on the registered suite but reduced it by $7.5\%$ on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.

Comment: 25 pages, 3 figures. Accepted to the E-Values Workshop at NeurIPS 2026 (poster)

arXiv abs page · PDF · same-day batch