| 1 | MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education | Luyao Zhu, Xun Wei Yee, Wei Li +2 | cs.AI | 2026-09-16 |
| 2 | ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts | Liyang Fan, Chi Wei, Yitai Li +8 | cs.CL | 2026-09-16 |
| 3 | RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection | Liyang Fan, Xinping Bi, Yitai Li +3 | cs.CV | 2026-09-16 |
| #4 | The Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses | Lydia Manikonda, Mei Si, Sirajam Munira +2 | cs.AI | 2026-09-16 |
| #5 | HAP: A Hand-Driven Active Perception Framework for Egocentric Head Motion Prediction | Yunji Feng, Junyi Ma, Guanzhong Sun +2 | cs.CV | 2026-09-16 |
| #6 | Accuracy- and Real-Time-Aware 4D Radar Preprocessing for Autonomous Driving Perception Systems | Woo-Jin Jung, Dong-Hee Paek, Jeong-Su Park +1 | cs.CV | 2026-09-16 |
| #7 | AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution | Jiabin Lou, Yirong Yang, Haopeng Wang +6 | cs.AI | 2026-09-16 |
| #8 | ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware | Shuai Zhou, Kaisheng Pang, Wenxuan Song +3 | cs.RO | 2026-09-16 |
| #9 | Learning from Distributed Eyes: Leveraging Collaborative Perception for Automated Model Adaptation | Yanan Ma, Yihang Tao, Zhengru Fang +4 | cs.CV | 2026-09-16 |
| #10 | Semantic-ITC: A Frame-wise Indoor Mobile Laser Scanning Dataset and Benchmark for Semantic Segmentation | Haiyang Wu, Muhammad Affan, George Vosselman +1 | cs.CV | 2026-09-16 |
| #11 | Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts | Christopher Bagdon, Carina Silberer, Roman Klinger | cs.CL | 2026-09-16 |
| #12 | Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs | Yuhang Zhu, Meiyi Zhu, Yunkai Dang +4 | cs.CV | 2026-09-16 |
| #13 | Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization | Xuyu Fan, Qi Ming, Zhu Han +6 | cs.CV | 2026-09-16 |
| #14 | Stealthy in Semantics, Antagonistic in Space: Attacking Visible-Infrared Object Detectors via Object-Level Misalignment | Yueqi Zhu, Qi Ming, Guo Cheng +6 | cs.CV | 2026-09-16 |
| #15 | A Comprehensive Review of Generative Physical Artificial Intelligence | Satyam Gaba, Krutiksinh Rana, Siva Sai +2 | cs.RO | 2026-09-16 |
| #16 | Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception | Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns | cs.RO | 2026-09-16 |
| #17 | Finder: Agentic Closed-Loop Object Finding for Embodied Grounding | Shixiong Xu, Zhiyuan Chen, Song Ding +4 | cs.CV | 2026-09-16 |
| #18 | Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning | Xinxin Song, Siyuan Li, Tingxiong Xiao +1 | cs.AI | 2026-09-16 |
| #19 | Collaborative Memory for Multi-Agent VLM Systems | Huixin Zhang, Shao-Jun Xia, Di Wang +3 | cs.AI | 2026-09-15 |
| #20 | Investigating Adversarial Robustness of Heterogeneous Cooperative Perception | Chenyi Wang, Yutong Liu, Qingzhao Zhang +1 | cs.CV | 2026-09-15 |
| #21 | Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models | Ziheng Ren, Qian Gao, Jun Fan +3 | cs.CV | 2026-09-15 |
| #22 | HuMemSLAM: Efficient Human-Inspired Semantic Place Recognition for Robust Visual SLAM | Mayowa Adebambo, Sebastian Donnelly, Armand Amaritei +2 | cs.RO | 2026-09-15 |
| #23 | Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement | Tobias Schaffer, Mohab Elkhayat, Daniela Nicklas +2 | cs.RO | 2026-09-15 |
| #24 | High-Fidelity Video Quality Assessment with VQA-Specific Saliency | Hakan Emre Gedik, Shashank Gupta, Alan Bovik | cs.CV | 2026-09-15 |
| #25 | NeuroSymbEAD: A Large Scale Neuro-Symbolic Caption Dataset for Omni-Directional Embodied Autonomous Driving | Muhammad Ahmed Ullah Khan, Mohammed Elamine, Sheikh Talha Uddin +3 | cs.CV | 2026-09-15 |
| #26 | Bridging Learned Visual Perception and Symbolic Belief-Space Planning | Guy Azran, Michael Navat, Sarah Keren | cs.AI | 2026-09-15 |
| #27 | VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal | Haonan Huang, Tianrui Qiu, Xianghao Zang +10 | cs.CV | 2026-09-15 |
| #28 | TEMPO: Learning Temporal Context for Dynamic Robot Manipulation | Zhenyang Feng, Jimin Heo, Erik B. Sudderth +1 | cs.RO | 2026-09-15 |
| #29 | CoAdapt: An LLM-based Framework for Adaptive Collaborative Perception in IIoT Robotic Swarms | Houssam Hajj Hassan, Antonia Maria Masucci, Lynda Zitoune +1 | cs.AI | 2026-09-15 |
| #30 | Hyper-RED: Scalable Event Pre-training via Semantic Hypergraph Distillation | Meisen Wang, Zhiqiang Tian, Wei Bao +3 | cs.CV | 2026-09-15 |
| #31 | LEAP: Learning Emergent Active Perception for Quadruped Navigation | Ü. Bora Gökbakan, Stéphane Caron, Philippe Souères | cs.RO | 2026-09-15 |
| #32 | World Models for Embodied Intelligence: From Plausible to Controllable to Actionable | Nanjie Yao, Hao Wang, Chong Cheng +10 | cs.RO | 2026-09-15 |
| #33 | Bridging the Perceptual Gap: Residual-Enhanced Downscaling and Manifold-Aware Perception Alignment Adaptation for NR-IQA | Yu Li, Zhengran Shen, Yachun Mi +2 | cs.CV | 2026-09-15 |
| #34 | Can Knowledge Transfer Parameters Be Learned? LePoKet for Efficient Robotic Vision | Yanick C. Tchenko, Felix Mohr, Hicham Hadj-Abdelkader +1 | cs.CV | 2026-09-15 |
| #35 | OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation | Chenhao Qiu, Dawei Li, Yechao Zhang +2 | cs.LG | 2026-09-15 |
| #36 | Can a Neural Encoding Model Replicate an fMRI Visualization Study? | Erfan Nasirzadeh Orang, Zack While | cs.HC | 2026-09-14 |
| #37 | Diversified and Perceptible Counterfactual Examples Leveraging Expert Knowledge | Akram Bensalem, Fahima Djelil, Marie-Jeanne Lesot +1 | cs.AI | 2026-09-14 |
| #38 | Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku | Livia Oddi, Simone Scardapane, Toru Sugimoto +1 | cs.CL | 2026-09-14 |
| #39 | Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception | Yanfeng Shi, Yan Song, Junhui Li +4 | cs.SD | 2026-09-14 |
| #40 | Legislating World-Model-Based Planning with Legal Reasoning | Dylan Waldner, Yiannis Kantaros, Guido Governatori +2 | cs.RO | 2026-09-14 |
| #41 | LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration | Jianhe Zhao, Yanhua Qiu, Zhiyu Zhang +2 | cs.CV | 2026-09-14 |
| #42 | ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models | Hongyu Jin, Wenda Zhang, Runqiu Fei +3 | cs.AI | 2026-09-14 |
| #43 | LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions | Myra Cheng, Lujain Ibrahim, Grace Liu +5 | cs.CY | 2026-09-13 |
| #44 | Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models | Mingze Yin, Xiaohan Wang, Dian Li +8 | cs.CL | 2026-09-13 |
| #45 | Perceive, Refine, Reason: A Calibrated Pipeline for Measuring Indicators in Strategic Visual Communication on Social Media | Weihong Qi, Chen Ling | cs.CV | 2026-09-13 |
| #46 | Investigating the Impacts of Generative AI on Information Seeking | Alexi Orchard, Shannon Lodoen | cs.HC | 2026-09-13 |
| #47 | CGGT: Curve-Grounded Geometry Transformer for 3D Parametric Curve Reconstruction | Zhirui Gao, Renjiao Yi, Yunfan Ye +4 | cs.CV | 2026-09-13 |
| #48 | OptoAgent: A Trustworthy Multi-Agent Framework for Opportunistic Vision Micro-Screening in Classroom Environments | Toqeer Ali Syed, Ali Akarma, Adeel Ahmad +1 | cs.AI | 2026-09-13 |
| #49 | Beyond Scene Description: Multi-Agent Orchestration for Non-visual Access to Virtual Worlds | Toqeer Ali Syed, Ali Akarma, Adeel Ahmad +1 | cs.AI | 2026-09-13 |
| #50 | AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video | Jiaming Tan, Mingliang Zhai, Zhen Li +3 | cs.CV | 2026-09-13 |