PaperScope
LIVE · 2026-09-30 05:40 UTC

State Trace Rationale As Auxiliary Task in Reinforcement Learning

Muhammad U. Nasir, Alex Vogt, Steven D. James, Julian Togelius

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.36867 v1
Category
Submitted
2026-09-29

Abstract

We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.

Comment: Under review as a conference paper at ICLR 2027

arXiv abs page · PDF · same-day batch