| 1 | Touvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments | George Xi Wang, Xiangyu Li, Shaoyue Wen +8 | cs.HC | 2026-09-18 |
| 2 | GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression | Pinxin Liu, Haiyang Liu, Jiahao Luo +3 | cs.CV | 2026-09-18 |
| 3 | MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation | Yiguang Yang, Jiankun Peng, Xiaoming Wang +2 | cs.CV | 2026-09-18 |
| #4 | AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining | Di Wu, Dongchen Zheng, Junhe Sheng +8 | cs.RO | 2026-09-18 |
| #5 | ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling | Xuancheng Zhang, Xuetao Liu, Qianying Tang +7 | cs.CV | 2026-09-18 |
| #6 | RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy | Xinyi Che, Zheng Lian, Kuofei Fang +13 | cs.CV | 2026-09-18 |
| #7 | CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding | Zhikun Zhou, Kunyu Peng, Runyi Yang +7 | cs.RO | 2026-09-17 |
| #8 | A Mathematical Model of Motivated Emotional Mind - Cognitive Embodied System | Wiesław L. Galus, Janusz A. Starzyk | q-bio.NC | 2026-09-17 |
| #9 | Navi-Agent: Unlocalized Monocular Navigation Agent | Wenyuan Xie, Mengyang Hong, Yongzhong Wang +9 | cs.RO | 2026-09-17 |
| #10 | Astronex-World 1.0: Real-Time Interactive World Model Foundation | Xin Zhou, Cong Miao | cs.CV | 2026-09-17 |
| #11 | CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding | Shuai Zhang, Hongye Hou, Qinghe Liu +5 | cs.CV | 2026-09-17 |
| #12 | BinoGen: Scaling egocentric binocular data for embodied visual perception and learning | Chunpeng Li, Ya-tang Li | cs.CV | 2026-09-17 |
| #13 | DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum | Haoqiang Kang, Yiming Zhang, Yiyang Guo +5 | cs.LG | 2026-09-17 |
| #14 | AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization | Xu Yuan, Yi Wang, Zhuohang Jiang +8 | cs.CV | 2026-09-17 |
| #15 | EmbodiedMind: Adaptive Data Curation and Prefix-Tree Reinforcement Learning for Efficient Embodied Intelligence | Feifan Wang, Zongbing Zhang, Yu Zhang +8 | cs.RO | 2026-09-17 |
| #16 | SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership | Run Peng, Zinnia Nie, Jing Ding +7 | cs.AI | 2026-09-17 |
| #17 | VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control | Zhongbo Zhang, Jiayi Jin, Yifan Wang +4 | cs.RO | 2026-09-17 |
| #18 | GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning | Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta +3 | cs.RO | 2026-09-16 |
| #19 | In-Context Robot Learning with VLM Agents | Dongzhou Cheng, Taoran Yi, Ye Fang +12 | cs.CV | 2026-09-16 |
| #20 | Track, Articulate, Act: Generating Articulation from Casual Human Videos | Jiaming Zhang, Homanga Bharadhwaj | cs.CV | 2026-09-16 |
| #21 | rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference | Kaijun Zhou, Zhiyang Li, Le Chen +1 | cs.RO | 2026-09-16 |
| #22 | AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution | Jiabin Lou, Yirong Yang, Haopeng Wang +6 | cs.AI | 2026-09-16 |
| #23 | CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models | Tianbin Liu, Jian Zhu, Taiyi Su +4 | cs.CV | 2026-09-16 |
| #24 | StrucPhysVideo: Learning Physical Dynamics from Structured Captions and Robot Actions | Awomo-WM Team, :, Enhui Ma +7 | cs.CV | 2026-09-16 |
| #25 | A Comprehensive Review of Generative Physical Artificial Intelligence | Satyam Gaba, Krutiksinh Rana, Siva Sai +2 | cs.RO | 2026-09-16 |
| #26 | Finder: Agentic Closed-Loop Object Finding for Embodied Grounding | Shixiong Xu, Zhiyuan Chen, Song Ding +4 | cs.CV | 2026-09-16 |
| #27 | Exploring 2D backbone effects for indoor semantic occupancy prediction | Shizhang Fanga, Wanling Yea, Qi Zheng | cs.CV | 2026-09-15 |
| #28 | FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence | Yinhao Li, Weixin Mao, Zihan Lan +21 | cs.RO | 2026-09-15 |
| #29 | NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving | Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin +3 | cs.CV | 2026-09-15 |
| #30 | World Models for Embodied Intelligence: From Plausible to Controllable to Actionable | Nanjie Yao, Hao Wang, Chong Cheng +10 | cs.RO | 2026-09-15 |
| #31 | EgoPathBench: Evaluating Zero-Shot Egocentric Waypoint Decision-Making in Vision-Language Models | Yang Zhao, Zhuo Chen, Xubo Yang | cs.CV | 2026-09-15 |
| #32 | PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models | DeepCybo Team, Yu Bin, Haipeng Cao +51 | cs.CV | 2026-09-14 |
| #33 | Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World | Guocun Wang, Kenkun Liu, Guorui Song +7 | cs.CV | 2026-09-13 |
| #34 | PuzzleMate: Benchmarking MLLMs for Egocentric Puzzle Assistance | Avijit Dasgupta, Shayon Dasgupta, Zakaria Laskar +2 | cs.CV | 2026-09-13 |
| #35 | The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement | Yi Duan, Ying Liu, Zirui Tang +30 | cs.LG | 2026-09-10 |
| #36 | ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI | Zhengran Ji, Jonathan Hyun, Boyuan Chen | cs.MA | 2026-09-10 |
| #37 | ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies | Jianming Ma, Rongjun Jin, Xiaxi Si +3 | cs.RO | 2026-09-10 |
| #38 | Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents | Marica Notte, Ludovica Marinucci, Vieri Giuliano Santucci | cs.AI | 2026-09-10 |
| #39 | BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration | Qianliang Wu, Haobo Jiang, Guangwei Gao +4 | cs.CV | 2026-09-10 |
| #40 | Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval | Yilin Long, Chenming Zhu, Zitang Gou +2 | cs.CV | 2026-09-10 |
| #41 | A Mathematical Theory of Pragmatic Information | Kai Niu, Ping Zhang | cs.IT | 2026-09-10 |
| #42 | ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs | Yizhan Li, Jianxin You, Mengyang Xiong +5 | cs.RO | 2026-09-09 |
| #43 | When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents | Qinzhen Ma, Ruihai Wu | cs.AI | 2026-09-09 |
| #44 | Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge | Yuanchen Bai, Zijian Ding, Angelique Taylor | cs.AI | 2026-09-09 |
| #45 | Show-Harness: Just a VLM Agent Can Play Robots | Yanzhe Chen, Zechen Bai, Zhijun Cao +7 | cs.RO | 2026-09-09 |
| #46 | Seven Sources of Physical AI Capability Formation | Gang Chen | cs.AI | 2026-09-09 |
| #47 | Modality-Decoupled Federated Learning for Privacy-Preserving Embodied Intelligence in 6G | Zhuodong Liu, Xiangyu Li, Chunhong Yuan +5 | eess.SP | 2026-09-09 |
| #48 | Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration | Yiran Qiao, Feng Wang, Jing Ma | cs.AI | 2026-09-08 |
| #49 | VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models | Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6 | cs.CV | 2026-09-08 |
| #50 | The Living Library: Transforming Archival Collections into Conversational Knowledge Systems -- Lessons from the Theodore Roosevelt Presidential Library | Pengce Wang, Lucia Ronchi Darre, Matt Briney +8 | cs.CV | 2026-09-08 |