| 1 | KV-streams for Efficient Compaction in Agentic Reinforcement Learning | Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda +15 | cs.LG | 2026-09-28 |
| 2 | SANTA++: Sampling Attention through Representative Keys | Kyle Lee, Christian Z. Pratt, Ruoyu Fang +4 | cs.LG | 2026-09-28 |
| 3 | Cartridges++: KV Cache Compression without Off-Context Derailment | Sonia Laguna, Joao Monteiro, Marco Cuturi +2 | cs.LG | 2026-09-28 |
| #4 | Language Models Act on Hidden Valence | Cameron Berg, Caspar Kaiser | cs.CL | 2026-09-28 |
| #5 | WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse | Aaryam Sharma | cs.DC | 2026-09-28 |
| #6 | CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion | Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer +3 | cs.LG | 2026-09-28 |
| #7 | TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash | Jay H. Park, Hyungjun Kim, Dong Kim | cs.DC | 2026-09-28 |
| #8 | Dual-Stream Simultaneous Translation via 2D Grid Attention | Yu Pu, Wei-Qiang Zhang | cs.AI | 2026-09-28 |
| #9 | ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport | Zhuchenyang Liu, Ziyi Wang, Yao Zhang +1 | cs.IR | 2026-09-28 |
| #10 | D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation | Zijian Ye, Chengqi Wei, Wei Huang +9 | cs.CV | 2026-09-28 |
| #11 | Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs | Zhengxiang Huang, Shengheng Chen, Chaoyue Niu +6 | cs.OS | 2026-09-28 |
| #12 | OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming | Zongshang Shen, Wangsong Yin, Daliang Xu +2 | cs.AI | 2026-09-28 |
| #13 | WorldAttention: An Efficient Attention Architecture for Interactive Video World Models | Zeyu Zhang, Jinyuan Mao, Dakai An +6 | cs.CV | 2026-09-28 |
| #14 | PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding | Qiuyang Zhang, Kai Zhou, Kai Lu +6 | cs.LG | 2026-09-28 |
| #15 | DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory | Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu +2 | cs.DC | 2026-09-28 |
| #16 | Spexis: Speculative Lookahead Scheduling for LLM Inference | Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi +3 | cs.LG | 2026-09-28 |
| #17 | Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference | Qianer Li, Chengjie Zhang, Jingwen Chen +3 | cs.CV | 2026-09-28 |
| #18 | SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving | Gunho Park, Kyoungho Jeun, Juntaek Oh +3 | cs.LG | 2026-09-28 |
| #19 | MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference | Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3 | cs.LG | 2026-09-28 |
| #20 | KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems | Hyesung Jeon, Hyeongju Ha, Seoyoung Lee +2 | cs.LG | 2026-09-28 |
| #21 | PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction | Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim | cs.LG | 2026-09-28 |
| #22 | Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference | Timothy DeLise, Seth Cromelin | cs.AI | 2026-09-28 |
| #23 | 3D Point Tracking with State Space Models | Masahiro Ogawa, Qi An, Atsushi Yamashita | cs.CV | 2026-09-27 |
| #24 | Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding | Jungseob Lee, Seungyoon Lee, Seongtae Hong +2 | cs.LG | 2026-09-27 |
| #25 | JET: Justification Evaluation in Transformer | Shenghao Ding | cs.LG | 2026-09-27 |
| #26 | Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models | Eduardo Ariño de la Rubia, Szilard Pafka | cs.SE | 2026-09-27 |
| #27 | EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents? | Kunming Shao, Jierun Chen, Jiangnan Yu +7 | cs.DC | 2026-09-27 |
| #28 | Positions Are Not Facts: The Mismatch Between KV Caches and Memory | Changhai Zhou, Yuhua Zhou, Shiyang Zhang +5 | cs.CL | 2026-09-27 |
| #29 | PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention | Kunming Shao, Jierun Chen, Yanli Wang +4 | cs.LG | 2026-09-27 |
| #30 | Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving | Jiantong Jiang, Yue Yang, Peiyu Yang +1 | cs.LG | 2026-09-27 |
| #31 | RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse | Ruoling Qi, Yirui Liu, Xuaner Wu +5 | cs.AI | 2026-09-27 |
| #32 | Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs | Yirui Liu, Ruoling Qi, Xuaner Wu +6 | cs.AI | 2026-09-27 |
| #33 | GSM: Efficient Language Modeling with Shared Global State | Yunao Zheng, Bin Wen, Xiaojie Wang +12 | cs.CL | 2026-09-27 |
| #34 | Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning | Junhao Xiao, Haoxiang Zhao, Menghao Fang +8 | cs.CV | 2026-09-27 |
| #35 | FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward | Sriman Achanta | cs.LG | 2026-09-27 |
| #36 | OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading | Jingyuan Xiao, Jiayue Wang, Yitao Hu +7 | cs.DC | 2026-09-27 |
| #37 | When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression | Haeyong Kang, Chang D. Yoo | cs.LG | 2026-09-27 |
| #38 | When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability | Xingjian Li, Yi Han, Jianhua Z. Huang | cs.LG | 2026-09-27 |
| #39 | FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation | Yiyi Cai, Yuhan Wu, Kunhang Li +6 | cs.CV | 2026-09-27 |
| #40 | On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation | Moghis Fereidouni, Anthony Arnold, Sumit Gulwani +2 | cs.AI | 2026-09-27 |
| #41 | How Linear Attention Remembers | Kichang Lee, JaeYeon Park, Songkuk Kim +1 | cs.LG | 2026-09-27 |
| #42 | SketchSSM: Write to the Full State, Read from a Compact Sketch | Omin Kwon, JoongWon Shin, Minseo Kim +3 | cs.LG | 2026-09-27 |
| #43 | Refreshing Less, Selecting Better: Reusing Stale Gradient Features for Efficient Influence-Based Data Selection | Jianchang Su, Yifan Zhang, Wei Zhang | cs.LG | 2026-09-26 |
| #44 | UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models | Wanqi Yang, Yuexiao Ma, Mei Xie +2 | cs.LG | 2026-09-26 |
| #45 | Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers | Ilya Koziev, Ivan Oseledets | cs.LG | 2026-09-26 |
| #46 | Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference | Xianpeng Shang, Canbin Huang, Jiang Li +4 | cs.LG | 2026-09-26 |
| #47 | KV-Lingo: Learning KV-Cache Translators with Distillation | Valérie Castin, Keitaro Sakamoto, Anastasiia Filippova +3 | cs.CL | 2026-09-26 |
| #48 | In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion | Yikai Wang, Xiao Han, Mengmeng Xu +9 | cs.CV | 2026-09-26 |
| #49 | UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation | Youssef Mansour, Enis Simsar, Fadime Sener +4 | cs.CV | 2026-09-26 |
| #50 | Carnator: Fast Text-to-Video Generation with Generation-Native Compatibility-Guided Cross-Request Reuse | Xingkun Yin, Xuebin Tang, Mingkun Xu +1 | cs.AI | 2026-09-26 |