PaperScope
LIVE · 2026-10-08 05:40 UTC

QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing

Yueying Li, Zhongle Xie, Ke Chen, Lidan Shou

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.09894 v1
Category
Submitted
2026-10-07

Abstract

In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.

arXiv abs page · PDF · same-day batch