PaperScope
LIVE · 2026-10-01 05:40 UTC

STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding

Nicolas Thiebaut, Nameer Hirschkind, Xiao Yu, Kyle Spence

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.38636 v1
Category
Submitted
2026-09-29

Abstract

We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.

Comment: 9 pages, 5 figures, 4 tables. A version of this paper appeared in PAKDD 2026 (LNCS vol. 16618)

Journal: Data Science: Foundations and Applications (PAKDD 2026), Lecture Notes in Computer Science, vol. 16618, pp. 473-484, Springer (2027)

arXiv abs page · PDF · same-day batch