PaperScope
LIVE · 2026-10-01 05:40 UTC

LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation

Xiaoqiang Wang, Suyuchen Wang, Bang Liu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39740 v1
Category
Submitted
2026-09-30

Abstract

Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect on the emitted token. These gains teach the policy when memory is more useful than further reasoning, while gradients through counterfactual recall teach which intermediate states should be retained in memory for future use. Across six general and long-context reasoning benchmarks, LatentHarness at 1.4B improves on the strongest baselines by 2.8% and 10.0% relative, respectively, and runs 5.9x faster than the strongest long-context baseline.

Comment: Work in progress

arXiv abs page · PDF · same-day batch