PaperScope
LIVE · 2026-10-06 05:40 UTC

SEIS: Self-Evolving Inference Systems

Zhen Xu, Jingyu Liu, Zongze Li, Tahseen Rabbani, Ce Zhang

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04646 v1
Category
Submitted
2026-10-03

Abstract

Inference systems determine how fast and how cheaply language models can be served, so making them faster has direct practical value. However, prior work focuses mostly on optimizing certain parts such as kernels or memory within the large system. In this work, we take a holistic approach and apply agentic self-evolution to optimize the whole system end-to-end. Our SEIS (Self-Evolving Inference Systems) autonomously optimizes the entire mini-sglang engine without human intervention through iterative sessions with inherited experiences and code changes. Serving Qwen3-0.6B on H100, the resulting engine reaches 3.27X the throughput of the original mini-sglang implementation and beats SOTA engines like vLLM, TensorRT-LLM, and SGLang in the single-request workload. The correctness of the optimized inference engine by SEIS is tested in terms of numerical difference and downstream accuracy on math and long-context retrieval tasks. The code and session histories show that the speedup comes from redesigning the whole engine and that building on earlier sessions beats independent attempts. These results suggest that agentic self-evolution can optimize a complex system end-to-end. The evaluation also has to evolve with the engine, and letting agents evolve it is a natural next step.

Comment: 21 pages

arXiv abs page · PDF · same-day batch