PaperScope
LIVE · 2026-09-09 05:40 UTC

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

Jan Corazza, Daniil Kaminskyi, Simon Lutz, Patrick Nossol, Hadi Partovi Aria, Zhe Xu, Daniel Neider

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.07461 v1
Category
Submitted
2026-09-07

Abstract

We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a framework that jointly learns a control policy and infers the hidden temporal cause of the phase transition. TCIRL maintains a hypothesis deterministic finite automaton (DFA) to track what phase is active and refines it via counterexample-driven SAT-based synthesis. We prove that the hypothesis converges almost surely to a DFA recognizing the true cause language on all attainable label sequences, yielding an optimal policy for the original non-Markovian decision process. Experiments on a genetic therapy gridworld and a traffic signal environment show that TCIRL recovers the correct cause DFA and matches the full-information baseline in both domains.

Comment: 9 pages, 8 figures. Accepted for publication in the Proceedings of the 65th IEEE Conference on Decision and Control (CDC), Honolulu, Hawaii, USA, December 2026. Extended version

arXiv abs page · PDF · same-day batch