PaperScope
LIVE · 2026-09-17 05:40 UTC

Distilling Foundation Models for Agentic What-If Reasoning:Cost, Latency, and Governance in a Hybrid LLM+SLM Architecture

Sourish Dey, Aditya Kumar

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.16091 v1
Category
Submitted
2026-09-14

Abstract

Tabular foundation models deliver strong zero-training predictive performance via in-context learning, but their high inference latency makes them impractical as hot-path decision backends in interactive agentic loops. We distill a TabPFN teacher into a compact feed-forward student across a business-decision simulation on UCI Adult and five OpenML benchmarks: the classification head compresses 53.2M parameters to 8,546 (6,220x); the deployed two-head loan pipeline compresses 111.4M parameters to 17,059 (6,532x). The student retains 95.4-100.5% accuracy and 96.8-100.0% AUC, with the lowest accuracy retention on credit-g at 95.4%; an alpha = 0 hard-label control shows that the teacher's soft targets provide a 2.1-7.0 AUC point gain.

Comment: 13 pages, 1 figure, 7 tables

arXiv abs page · PDF · same-day batch