PaperScope
LIVE · 2026-09-29 05:40 UTC

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

Carlos Garrido-Munoz, Jorge Calvo-Zaragoza

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.35473 v1
Category
Submitted
2026-09-28

Abstract

In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.

Comment: Accepted at 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

arXiv abs page · PDF · same-day batch