PaperScope
LIVE · 2026-10-07 05:40 UTC

Weight Oracles: Reading Neural Network Weights with Language Models

Krishna Kabra, Constantin Venhoff, Christian Schroeder de Witt

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.07334 v1
Category
Submitted
2026-10-05

Abstract

Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.

Comment: Spotlight at the NeurIPS 2026 Workshop on Neural Network Artifacts as a New Data Modality (NeuralArtifacts), Paris. 14 pages, 10 figures

arXiv abs page · PDF · same-day batch