| 1 | StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training | Bao Tang, Jiahao Guo, Haoxiang Cao +4 | cs.CV | 2026-09-22 |
| 2 | Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning | Yuanteng Chen, Zhilei Liu, Peisong Wang +7 | cs.LG | 2026-09-22 |
| 3 | PP-Net: A Hybrid Physical-Prior Neural Network for Scattered Light Removal in Biomedical Images on Embedded Devices | Yongfei Guo, Tingjin Chu, Mengzhuo Liu +2 | cs.CV | 2026-09-22 |
| #4 | Latent Dataset Distillation for Human Motion Prediction | Ge Tian, Guang Li, Takahiro Ogawa +1 | cs.CV | 2026-09-22 |
| #5 | QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation | Jiaqi Zhao, Xiaobin Hu, Bo Yin +3 | cs.CV | 2026-09-22 |
| #6 | Geometry-Aware Hyperbolic Residual Quantization | Alessio Colombo, Melika Ayoughi | cs.LG | 2026-09-22 |
| #7 | Disaggregated Quantization: Specializing LLM Prefill and Decode | Andrei Panferov, Maximilian Kleinegger, Sweta Priyadarshi +2 | cs.LG | 2026-09-22 |
| #8 | Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement | Akihiro Yoshida, Yuma Ichikawa | cs.LG | 2026-09-22 |
| #9 | RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models | Zhiping Wu, Dongdong Ren, Yangchengyu Zhou +4 | cs.CV | 2026-09-21 |
| #10 | VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models | Jiuyi Xu, Qing Jin, Meida Chen +3 | cs.RO | 2026-09-21 |
| #11 | SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models | Kewei Zhang, Zheng Chen, Haotong Qin +1 | cs.CV | 2026-09-21 |
| #12 | When Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs | Yeji Kim, Mi-Young Kim, Randy Goebel | cs.CL | 2026-09-21 |
| #13 | NPU Accelerator: Quantized Real-Time Vehicle Detection on PYNQ-Z1 Using FINN | Daniel Gutierrez, Antonio Cuesta, Jorge Fe +2 | cs.AR | 2026-09-21 |
| #14 | What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization | Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa +4 | cs.CV | 2026-09-21 |
| #15 | QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation | Demian Pavlyshenko, Bohdan Pavlyshenko | cs.CL | 2026-09-21 |
| #16 | ME-VLM:A Unified VLM for Embodied Cognition and Agent Coordination | Foundation Model, Li Auto Inc | cs.CV | 2026-09-21 |
| #17 | FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding | Hung T. Ho, Khanh D. Nguyen, Quang D. Nguyen +5 | cs.RO | 2026-09-21 |
| #18 | NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware | Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis +2 | cs.LG | 2026-09-21 |
| #19 | The Undetected Damage of Quantization on Retrieval and How to Fix It | Luca Zhou, Alessandro Zirilli, Daniele Solombrino +2 | cs.LG | 2026-09-21 |
| #20 | KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation | Sihyeon Ha, Jaeho Lee, Yo-Seb Jeon | cs.LG | 2026-09-21 |
| #21 | From Bits to Beliefs: Recoverable Semantic Fingerprints for Black-Box Verification of Large Language Models | Jiaxin Hong, Yuxin Peng, Hongyao Yu +3 | cs.CR | 2026-09-21 |
| #22 | Q-DEQ: Discrete Solving and Quantization for Deep Equilibrium Models in Time Series Forecasting under Edge Deployment Coding Constraints | Ruotong Yang, Hongdong Zhu, Qi Gao +3 | cs.LG | 2026-09-21 |
| #23 | On the Efficiency-Safety Dilemma in Large Reasoning Models | Yifei Yang, Zouying Cao, Xingrui Wang +4 | cs.CL | 2026-09-20 |
| #24 | Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization | Kuanlin Chen, Chen-Wei Kuo, Cheng-En Ou | cs.LG | 2026-09-20 |
| #25 | WaveletECO: A Closed-Loop Physical ECO Platform and a Specialized Local Language Model | Guoxiang Xu, Guozhen Ji, Zijian Luo +3 | cs.AR | 2026-09-20 |
| #26 | Neural Residual Modeling for Scientific Data Compression under Guaranteed Error Bounds | Surya Majumder, Liangji Zhu, Sanjay Ranka +1 | cs.LG | 2026-09-19 |
| #27 | SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation | Yuxi Liu, Haoyu Li, Zekun Zhang +12 | cs.CV | 2026-09-19 |
| #28 | Real-Time Plasma State Prediction via FPGA-Accelerated Quantized Recurrent Probabilistic Neural Networks | Daniel Gaytan-Villarreal, Aiken Xie, Tu Pham +5 | physics.plasm-ph | 2026-09-19 |
| #29 | From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving | Twinkll Sisodia | cs.AI | 2026-09-19 |
| #30 | Dual-Locking Learned AI Models: A PIN-Based Sparse QIM Watermarking and Adaptive Index Permutation Approach | Iva Vasic, Jesús Muñoz-Cádiz, Bata Vasic | cs.CR | 2026-09-19 |
| #31 | Towards Full Pipeline FP8 Reinforcement Learning for LLMs | Fanchao Chen, Ziheng Jiang, Ziyun Wei +5 | cs.LG | 2026-09-19 |
| #32 | PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection | Xuanming Shang, Weijia Zhang, Chao Ma | cs.CV | 2026-09-18 |
| #33 | ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation | Boni Hu, Xiong Wei, Haoming Huang +19 | cs.CV | 2026-09-18 |
| #34 | SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference | Harish KB, Jagadeeswaran M, Pradheep P +2 | cs.LG | 2026-09-18 |
| #35 | Multi-Domain Clustering via Measure Quantization | Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante | cs.LG | 2026-09-18 |
| #36 | Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals | Yamato Narita, Issei Sato | cs.LG | 2026-09-18 |
| #37 | Quantization-Aware Kalman Estimation for Diffusion Sampling | Qitan Shi, Cheng Jin, Jiawei Zhang +1 | cs.CV | 2026-09-18 |
| #38 | Score Centering Stabilizes Off-policy Reinforcement Learning | Martin Marek, Max Ryabinin | cs.LG | 2026-09-17 |
| #39 | Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation | Fabian Schmalstieg, Karsten Mueller, Wojciech Samek | cs.CV | 2026-09-17 |
| #40 | QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization | Yujie Li, Zezhi Shao, Chengqing Yu +7 | cs.LG | 2026-09-17 |
| #41 | D-Quant: Driftable Entropy Coding for KV Cache Quantization | Yi Su, Hong Liu, Guanghua Yu +1 | cs.CL | 2026-09-17 |
| #42 | MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration | Yuan Liao, Jae-sun Seo | cs.AR | 2026-09-17 |
| #43 | Predict Before You Deploy: Offline Prediction of Quantization-Induced Task Degradation for World Action Models | Jiuyi Xu, Jinjia Guo, Meida Chen +2 | cs.RO | 2026-09-16 |
| #44 | ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models | Shijie Lian, Bin Yu, Zhaolong Shen +5 | cs.RO | 2026-09-16 |
| #45 | REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception | Geoffroy Keime, Nicolas Cuperlier, Benoit R. Cottereau | cs.RO | 2026-09-16 |
| #46 | ${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models | Chunpu Xu, Zhixuan Liang, Yuhao Zhang +6 | cs.RO | 2026-09-16 |
| #47 | CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling | Chong Jing, Junan Zhang, Zhizheng Wu | cs.SD | 2026-09-16 |
| #48 | Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing | Eunju Shin, Jongbin Ryu | cs.LG | 2026-09-16 |
| #49 | The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction | Yu Lin, Yiming Wang, Runyuan Cai +2 | cs.AI | 2026-09-16 |
| #50 | A Calibrated Instrument for Measuring How Inference Optimizations Affect Output Quality | Jerry Kaplan | cs.CL | 2026-09-16 |