PaperScope
LIVE · 2026-09-29 05:40 UTC

A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations

Yoav Evron, Michal Bar-Asher Siegal, Michael Fire

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33449 v1
Category
Submitted
2026-09-27

Abstract

Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extraction systems can find illustrations on manuscript pages, but without meaningful categories, large collections remain hard to search and explore. We address this gap by introducing a manually labeled dataset of 15,000 illustrations from manuscripts dating back hundreds of years across 22 categories, and evaluating modern vision models for image classification on this task. The problem is challenging due to stylistic diversity, degradation, and semantic ambiguity, with many images that fit more than one category. We compare fine-tuned CNN and Transformer-based classifiers, zero-shot CLIP, embedding-based classifiers, and direct vision-language models. Results show that fine-tuned image classifiers perform best overall, with ConvNeXt reaching 88.9% accuracy and 81.3% macro-F1. Using CLIP embeddings with XGBoost provides a strong alternative. In contrast, zero-shot CLIP and direct vision-language classification perform substantially worse, highlighting the limits of general-purpose models in this domain. Beyond overall performance, the analysis reveals which categories are visually separable and where errors reflect genuine semantic overlap, suggesting that some limitations arise from the taxonomy itself.

Comment: 17 pages, 5 figures

arXiv abs page · PDF · same-day batch