PaperScope
LIVE · 2026-10-06 05:40 UTC

From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deployment

Zeshen Zhang, Han Zhao, Weihao Cui, Quan Chen, Yu Liu, Yongjun Deng, Jing Yang, Jiuchen Shi, Chen Chen, Youmin Chen, Yu Feng, Minyi Guo

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.04956 v1
Category
Submitted
2026-10-04

Abstract

As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services. However, these resource-constrained servers struggle to guarantee heterogeneous Service Level Objectives (SLOs) when serving multiple LoRA-adapted services simultaneously. Existing serving frameworks suffer from severe SLO violations due to the computational overhead of LoRA layers and the rigid nature of batch scheduling. To address this, we propose HALO, a scheduling method tailored for LoRA-assisted on-premise LLM deployment. HALO introduces two key innovations: a spatial multiplexing strategy that overlaps Base and LoRA computations by partitioning GPU Streaming Multiprocessors (SMs), and an SLO-aware scheduler that decouples request execution based on "request-level slack." By prioritizing urgent tasks and utilizing idle budget for traffic shaping, HALO significantly mitigates resource contention. Our evaluation demonstrates that HALO minimizes SLO violations while improving throughput compared to state-of-the-art baselines.

Comment: 22 pages, 16 figures, 4 tables

arXiv abs page · PDF · same-day batch