PaperScope
LIVE · 2026-09-09 05:40 UTC

Behavioral Cloning Outperforms Entropy-Regularized RL: Critic-Driven Failure of Actor-Critic Methods on Adaptive Tumor Treatment

Aleksandar Dimitrov, Giacomo Spigler

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.06667 v1
Category
Submitted
2026-09-06

Abstract

Adaptive dosing requires policies that reduce tumor burden without excessive toxicity. Learned dosing policies are typically judged against historical or heuristic comparators, which cannot show whether a policy has found the best behavior available. We instead study a three-population tumor-control ODE in which optimal-control analysis fixes the form of a good schedule -- bang-bang dosing punctuated by a singular arc -- and construct a numerical controller of that form as a proxy for near-optimal behavior. Judged against this reference under a sustained-cure criterion -- 200 consecutive days below 5% carrying capacity -- Soft Actor-Critic (SAC) trained from scratch never reaches cure. Behavioral cloning (BC) of the reference reproduces it (100% sustained cure, 30/30 seeds), but SAC fine-tuning of the cloned policy destroys it across five entropy coefficients, and TD3 and BC-regularized SAC fail identically; the pattern persists under multiplicative pharmacokinetic action noise. Along curative trajectories the post-collapse critic ranks the collapsed-policy action above the reference action in 96% of states, concentrated in the maintenance phase, and the policy settles into a non-curative adaptive-therapy equilibrium. The reference is what makes this legible: against a heuristic comparator the fine-tuned policy would read as a competent controller rather than a failure.

Comment: 9 pages, 3 figures

arXiv abs page · PDF · same-day batch