PaperScope
LIVE · 2026-10-06 05:40 UTC

Structured Representation Learning for Behavior Cloning: How can we learn to safely control a nuclear power plant?

Perceval Beja-Battais, Alain Grosset{ê}te, Nicolas Vayatis

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.06211 v1
Category
Submitted
2026-10-05

Abstract

Learned models for industrial control are usually judged by aggregate accuracy, but accuracy at the component level does not guarantee safety once it is embedded in the system it is meant to serve. We study this gap on a behavior-cloning task: imitating an expert Nonlinear Model Predictive Control (NMPC) policy for load-following of a Pressurized Water Reactor (PWR), an industrial system with tight safety constraints. We propose a structured architecture encoding variables from each timescale into separate latent spaces, reflecting the physical decomposition of the system, before training a controller to imitate the expert on the product latent space. On long-horizon rollouts, separated embeddings improve both accuracy and feasibility compared with a shared-embedding baseline. Sensitivity analysis further shows that our model yields interpretable representations aligned with the system's physics. However, standalone deployment still leaves several percent of trajectories infeasible regardless of the architecture. Using our method to warmstart the NMPC optimizer rather than acting standalone, we recover full feasibility and near-optimal cost while still cutting computation time by $\sim$15% relative to the expert controller, and even more for abrupt operating changes.

Journal: NeurIPS 2026 Workshop AI Foundations for Power Grids, Dec 2026, Sydney (Australia), Australia

arXiv abs page · PDF · same-day batch