PaperScope
LIVE · 2026-10-06 05:40 UTC

Safe Context Switching for Agents in the Wild: Mitigating Subspace Interference via Orthogonal Adaptation

Akash Das, Ishan Roy

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05219 v1
Category
Submitted
2026-10-04

Abstract

Most Large Language Models exhibit a fundamental tension between two sequential tasks, such as logical reasoning and safety alignment. The high-variance internal states required for sophisticated Chain-of-Thought (CoT) deduction can geometrically interfere with latent representations encoding safety constraints. We identify this phenomenon as Sequential Subspace Interference, showing that standard fine-tuning on logical tasks such as multi-step mathematics and code generation can result in a 23.3% interference penalty on alignment benchmarks, substantially weakening the model's safety priors. This Reasoning Drift is not adequately captured by current adaptation methods because gradients for logical tasks are rarely orthogonal to safety objectives. To address this issue, we propose AURA (Adaptive Unique Residual Allocation), a spectral regularization framework that enforces Spectral Independence between reasoning and safety. By explicitly estimating the null space of the alignment manifold and constraining reasoning updates to its orthogonal complement, AURA enables models to improve logical reasoning without compromising safety. Empirically, AURA recovers 23.0% of the lost performance while preserving greater than 0.98 cosine fidelity to the safe state, demonstrating that reasoning and alignment can be effectively decoupled through geometric regularization.

Journal: ICLR 2026 the 2nd Workshop on World Models: Understanding, Modelling and Scaling

arXiv abs page · PDF · same-day batch