PaperScope
LIVE · 2026-10-09 05:40 UTC

AI4Fire: Evaluating Large Language Models on Wildfire Tasks

Yue Zhao, Xiyang Hu, Zuobin Xiong, Zhangyu Wang, Ruolin Li

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.10946 v1
Category
Submitted
2026-10-07

Abstract

Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a smoke-free reference frame from the same camera. AI4Fire runs six core models bare and grounded on five wildfire tasks, zero-shot; a sweep adds 29 more. Our literature search on fire tasks found 138 works; none combines this roster, task coverage, and paired bare and grounded runs. We report three findings. (1) Grounding helped most where the addition carried the answer: a read-only SQL tool lifted every core model's database accuracy from at most 16 to at least 88 percent. (2) Simple rules were hard to beat: no core model outperformed repeating today's staffing count, and two open-weight models mostly copied the median of similar earlier fire-days, a worse forecast. (3) Public releases carry hazards: a fire-danger column separates the holdout perfectly, and 67 aerial fire frames carry smoldering or fire-free labels read from a clipped thermal maximum. We release prompts, responses, scores, code, and the survey record.

Comment: 51 pages. Code: https://github.com/yzhao062/AI4Fire

arXiv abs page · PDF · same-day batch