PaperScope
LIVE · 2026-09-15 05:40 UTC

SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction

Daniel Woortmann, Tanguy Magne, Olga Sorkine-Hornung

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.15503 v1
Category
Submitted
2026-09-14

Abstract

Deep learning models have become the standard tool for document rectification and illumination correction, yet their performance is fundamentally bound by their training data. For nearly a decade, the community has heavily relied on Doc3D, a pioneering but increasingly limited document unwarping dataset in terms of scale and quality. To address this bottleneck, we introduce SyntheticDoc, a massive, high-quality dataset designed to push the boundaries of document unwarping. SyntheticDoc is composed of 1,000,000 high-resolution procedurally generated training samples, alongside extensive validation and test sets. Each sample is paired with rich, pixel-perfect annotations, including UV maps, normal maps, albedo and shading. To ensure physical accuracy and photorealism, the paper geometries are generated via a physics-based simulator and rendered using a path tracer. To demonstrate the benefit of our dataset, we train a simple baseline model on SyntheticDoc and report on its performance in comparison to state-of-the-art methods on both document unwarping and illumination correction tasks. Our dataset is available at https://igl.ethz.ch/projects/SyntheticDoc/ and the code used to generate it at https://github.com/tanguymagne/SyntheticDoc .

Comment: D. Woortmann and T. Magne -- Equal contribution. Accepted at ECCV 2026 (Spotlight). 20 pages

Journal: Lecture Notes in Computer Science, vol. 17057, pp. 131-150, Springer (2026)

arXiv abs page · PDF · same-day batch