PaperScope
LIVE · 2026-10-08 05:40 UTC

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models

Domenic Rosati, Alessa Carbo, Ali Dadsetan, Hong Huang, Matthew Young, Subhabrata Majumdar, Frank Rudzicz, Hassan Sajjad

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.09004 v1
Category
Submitted
2026-10-06

Abstract

Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.

Comment: Under submission AISTATS 2026

arXiv abs page · PDF · same-day batch