PaperScope
LIVE · 2026-09-30 05:40 UTC

Engineering Efficient Self-Play Chess: Search, Replay, and Throughput Under Limited Compute

Bertil Braun

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.37447 v1
Category
Submitted
2026-09-27

Abstract

How strong can an AlphaZero-style chess system become under limited training compute when its entire learning loop is engineered for efficiency? We train from random initialization through searched self-play on a single eight-GPU node for 2.5 days. The resulting 6.32-million-parameter model reaches 3,251 benchmark Elo [3,206, 3,297] at 100,000 searches per move (estimated at under five seconds of thinking time) against a fixed-node Stockfish 13 ladder. The run ingests 3.25 million completed games, involves an estimated 100 billion search simulations, and makes 836.6 million training presentations. We investigate search allocation, replay and restart-state selection, policy representation, progressive model sizing, quantized inference, and throughput engineering. Alongside the retained design, we document plausible alternatives that failed to improve the complete learning loop or did not justify their cost. The reported strength is a result of the integrated system, not an isolated Elo gain attributable to any single choice.

Comment: 26 pages, 17 figures, including appendices. Code and experimental evidence: https://github.com/BertilBraun/Advanced-Techniques-in-Chess-Engines ; model artifacts: https://huggingface.co/BertilBraun/alphazero-chess ; interactive demo: https://chess.bertil-braun.de/

arXiv abs page · PDF · same-day batch