| 1 | PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models | Erik Deinzer, Naya Baslan, Luca Paparusso +3 | cs.CV | 2026-09-18 |
| 2 | GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments | Yichen Liu, Puzhen Yuan, Xiang Zhu +2 | cs.RO | 2026-09-18 |
| 3 | ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation | Boni Hu, Xiong Wei, Haoming Huang +19 | cs.CV | 2026-09-18 |
| #4 | Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies | Xingyu Lin, Zhuang Li, Zhongrun Wu +2 | cs.RO | 2026-09-18 |
| #5 | SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations | Hiroaki Kingetsu, Hiroaki Kurihara, Kaoru Yokoo +2 | cs.RO | 2026-09-18 |
| #6 | VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models | Kaiwen Zhu, Dongfang Liu, Liangkai Liu | cs.RO | 2026-09-18 |
| #7 | FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models | Zhiyuan Gao, Di Wen, Yanxiang Zhan +4 | cs.RO | 2026-09-18 |
| #8 | Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies | Zhipeng Tang, Xinda Chen, Weining Rao +5 | cs.RO | 2026-09-18 |
| #9 | GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies | Xin Chen, Sen Chen, Yujuan Ding +5 | cs.RO | 2026-09-17 |
| #10 | HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface | Zimu Han, Yiming Zeng, Jiyao Zhang +9 | cs.RO | 2026-09-17 |
| #11 | Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning | Haoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang +2 | cs.LG | 2026-09-17 |
| #12 | Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies | Enhao Wu, Fusen Guo, Yuxin Cao +3 | cs.CV | 2026-09-17 |
| #13 | rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference | Kaijun Zhou, Zhiyang Li, Le Chen +1 | cs.RO | 2026-09-16 |
| #14 | FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory | Kemal Oksuz, Alexandru Buburuzan, Yuhan Yao +1 | cs.CV | 2026-09-16 |
| #15 | ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware | Shuai Zhou, Kaisheng Pang, Wenxuan Song +3 | cs.RO | 2026-09-16 |
| #16 | ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models | Shijie Lian, Bin Yu, Zhaolong Shen +5 | cs.RO | 2026-09-16 |
| #17 | WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories | Yuna Oikawa, Kei Endo, Takanori Uzawa +5 | cs.AI | 2026-09-16 |
| #18 | ${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models | Chunpu Xu, Zhixuan Liang, Yuhao Zhang +6 | cs.RO | 2026-09-16 |
| #19 | Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation | Lijie Wang, Zheng Lu, Yiming Wang +12 | cs.RO | 2026-09-16 |
| #20 | Reinforcement Learning for Real-Time Vision-Language-Action Policies | Perry Dong, Kuo-Han Hung, Dorsa Sadigh +1 | cs.RO | 2026-09-16 |
| #21 | Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models | Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev +1 | cs.RO | 2026-09-16 |
| #22 | FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence | Yinhao Li, Weixin Mao, Zihan Lan +21 | cs.RO | 2026-09-15 |
| #23 | Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement | Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas +2 | cs.RO | 2026-09-15 |
| #24 | sensVLA: Spatially-Grounded Vision-Language-Action Model for Autonomous Wheel Loader | Gopi Krishna Erabati, Bjarne Johannsen, Angus Stewart +1 | cs.CV | 2026-09-15 |
| #25 | TEMPO: Learning Temporal Context for Dynamic Robot Manipulation | Zhenyang Feng, Jimin Heo, Erik B. Sudderth +1 | cs.RO | 2026-09-15 |
| #26 | GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving | Xiao Liu, Haoyu Li, Jianghao Leng +2 | cs.CV | 2026-09-14 |
| #27 | IMPACT-VLA: Interaction-aware Multimodal Propagation Attribution via Counterfactual Trajectories for Vision-Language-Action Policies | Jinwoong Kim, Sangjin Park | cs.RO | 2026-09-14 |
| #28 | When Faster VLA Deployment Changes Closed-Loop Behavior: Task Success-Latency Analysis of SmolVLA Across PyTorch and ONNX Variants | Rafiqul Islam | cs.RO | 2026-09-12 |
| #29 | What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency | Luoyang Sun, Guoyang Xia, Fengfa Li +9 | cs.RO | 2026-09-12 |
| #30 | ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting | Chenwei Wang, Dianye Huang, Match W. L. Ko +2 | cs.RO | 2026-09-12 |
| #31 | ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies | Jianming Ma, Rongjun Jin, Xiaxi Si +3 | cs.RO | 2026-09-10 |
| #32 | Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models | Gautam Rajendrakumar Gare, Siyi Li, Hewei Wang +5 | cs.CV | 2026-09-10 |
| #33 | IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies | Kian Hosseinkhani, Qinhe Peng, George Shramko +6 | cs.RO | 2026-09-10 |
| #34 | HuRo: Robotizing Human Videos for Scalable VLA Pretraining | Jinho Jeong, Se June Joo, Jaehyun Kang +4 | cs.RO | 2026-09-09 |
| #35 | Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models | Shengye Dong, Haochen Niu, Hao Liu +3 | cs.AI | 2026-09-09 |
| #36 | Modality-Decoupled Federated Learning for Privacy-Preserving Embodied Intelligence in 6G | Zhuodong Liu, Xiangyu Li, Chunhong Yuan +5 | eess.SP | 2026-09-09 |
| #37 | TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model | Anqi Li, Yuxin Chen, Zhaobo Li +4 | cs.RO | 2026-09-08 |
| #38 | DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination | Yankai Fu, Ning Chen, Junkai Zhao +5 | cs.RO | 2026-09-08 |
| #39 | No Free Checker: A Survey of Verifiers for Robot Policies | Yang Wan, Xihang Yue, Zhirui Liu +7 | cs.RO | 2026-09-08 |
| #40 | WorldAgen: Unified State-Action Prediction with Test-Time World Model Training | Chi Wan, Kangrui Wang, Yuan Si +2 | cs.AI | 2026-09-08 |
| #41 | Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy | Ayoub Kirouane, Georgios Giaples, Christos Petrocheilos | cs.RO | 2026-09-07 |
| #42 | Large Discrete Policy: Advancing Explicit Behavior Modeling with Stochastic Iterative Scoring | Zhenxin Li, Nadine Chang, Xinglong Sun +8 | cs.RO | 2026-09-07 |
| #43 | MEMOBench: A Process Level Memory Benchmark for Robotic Manipulation | Haiyang Sun, Haoxiao Wang, Junming Chen +6 | cs.RO | 2026-09-07 |
| #44 | GIFT: Goal-Injected Fine-Tuning for Efficient Manipulation Policy Adaptation | Xiaoyuan Fang, Shuo Feng, Yuxuan Wang +3 | cs.CV | 2026-09-07 |
| #45 | Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models | Yijie Zhu, Zitong Yu, Wei Li +4 | cs.CV | 2026-09-06 |
| #46 | What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies | Vivek Chavan, Pengtao Xie, Yahuan Shi +3 | cs.RO | 2026-09-04 |
| #47 | Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation | Vivek Chavan, Yahuan Shi, Oliver Heimann +2 | cs.RO | 2026-09-04 |
| #48 | RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? | Zhenxuan Fan, Bo Zhang, Yutong Lin +9 | cs.RO | 2026-09-04 |
| #49 | Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving | Ruoyu Yao, Yusen Xie, Qingzhao Liu +5 | cs.CV | 2026-09-03 |
| #50 | FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation | Yutian Zhang, Siyuan Ma, Liwen Yang +6 | cs.RO | 2026-09-03 |