PaperScope
LIVE · 2026-09-03 05:40 UTC

MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft

Sewoong Lee, Risham Sidhu, Julia Hockenmaier, Yoonhwa Jung

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2608.28884 v1
Category
Submitted
2026-08-28

Abstract

We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.

Journal: EMNLP 2026 Findings

arXiv abs page · PDF · same-day batch