PaperScope
LIVE · 2026-09-28 05:40 UTC

Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights

Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.31564 v1
Category
Submitted
2026-09-25

Abstract

We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.

arXiv abs page · PDF · same-day batch