PaperScope
LIVE · 2026-09-29 05:40 UTC

ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation

Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.34717 v1
Category
Submitted
2026-09-28

Abstract

Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search framework. It organizes program candidates as tree states, retains branch-local debugging context, retrieves failure experience across branches, and distinguishes failed checks from unavailable evidence. On HumanEval and MBPP-Sanitized, visible-test ReMCTS improves over direct generation in 8 of 10 model-dataset pairs under held-out evaluation, whereas proxy-only search is less stable. Controlled tree-search, sampling, repair, and memory ablations characterize the source and limits of these gains. A 30-task HumanEval-X C++ pilot further demonstrates compatibility with compiler-backed execution, but does not constitute a broad multilingual evaluation.

Comment: 21 pages, 2 figures. To appear in the Proceedings of EMNLP 2026

arXiv abs page · PDF · same-day batch