Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
Xinhan Yang, Lu Lu, Shancong Mou
Abstract
Derivative-informed training improves neural operators by directly supervising their input-output sensitivities, which is crucial when neural operators are used as differentiable surrogates for inverse problems, PDE-constrained optimization, design, and control. However, existing methods rely on offline-generated derivative labels, making data generation slow, storage-intensive, and difficult to adapt across datasets, resolutions, or perturbation bases. We propose sketched tangent consistency loss (sTCL), an on-the-fly derivative-informed training objective that, for PDEs with a known and differentiable residual, enforces sensitivity consistency directly from the governing equation without offline tangent labels or neural-operator architecture changes. sTCL uses randomly sketched input perturbations to provide a lightweight derivative-level physics constraint during training. However, raw forward-sensitivity residual penalties can fail for stiff, ill-conditioned, indefinite, or coupled saddle-point tangent operators. To address this, we introduce lightweight operator-aware loss-conditioning mechanisms selected by a simple tangent-operator decision rule. Across Helmholtz, nonlinear diffusion-reaction, Burgers, Allen-Cahn, and Navier-Stokes, with the same neural-operator backbone for all methods, the PDE-specific sTCL losses achieve solution and Jacobian accuracy comparable to offline derivative-informed training (DIFNO) while eliminating the offline derivative-data generation and storage pipeline. These results show that on-the-fly derivative-informed training need not merely amortize offline tangent-solve cost into training; with appropriate sketching and loss design, sTCL provides an effective drop-in path to derivative-informed neural operators. Code is available at https://github.com/yang9579/Derivative-informed-traning-on-the-fly.