PaperScope
LIVE · 2026-09-29 05:40 UTC

On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation

Moghis Fereidouni, Anthony Arnold, Sumit Gulwani, Mark Marron, A. B. Siddique

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33141 v1
Category
Submitted
2026-09-27

Abstract

Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resources run and depend on cloud hosted inference to handle the task of transforming natural language inputs into actionable software operations. This reliance on cloud-hosted inference introduces substantial network latency on top of LLM inference times, creates data privacy concerns, and, given the costs of running these models, can rapidly escalate expenses associated with supporting agentic features. This paper introduces a novel means of converting the NL-to-Action problem from a generative one into a classification-centric formulation via on-device operation caches. These caches allow an agentic system to handle frequently occurring classes of actions completely on-device -- reducing latency, enhancing privacy, and lowering operational costs. We show that for a classic NL-to-Formula task, generating Excel Formula in response to user requests, this approach reduces total inference cost by 56% when compared to cloud-only model-routing based inference and, on cache hits, reduces the latency to response latency by 5x.

arXiv abs page · PDF · same-day batch