| 1 | SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models | Junchao Huang, Guian Fang, Shengju Qian +15 | cs.CV | 2026-09-02 |
| 2 | PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation | Yu Tian, Xintong Jiang, Jan Franklin Adamowski +2 | cs.CV | 2026-09-02 |
| 3 | User Feedback Provides a Unique Signal that LLMs Can not Detect | Shachar Don-Yehiya, Leshem Choshen, Omri Abend | cs.CL | 2026-09-02 |
| #4 | MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion | Aidan Bradshaw, Marco Giordano, David Rode +8 | cs.CV | 2026-09-02 |
| #5 | Benchmarking RAW and RGB Restoration in Image Signal Processors | Zihao Lu, Radu Timofte, Marcos V. Conde | cs.CV | 2026-09-02 |
| #6 | Cliff: Learning Process Rewards from the First Mistake | Peixuan Han, Runhui Wang, Ketan Ramaneti +3 | cs.LG | 2026-09-02 |
| #7 | GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic Design | Adrienne Deganutti, Purvanshi Mehta, Simon Hadfield +1 | cs.CV | 2026-09-02 |
| #8 | AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels | Javier Tirado-Garín, Alan Savio Paul, Shuai Chen +5 | cs.CV | 2026-09-02 |
| #9 | From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution | Yuzhang Luo, Chenpeng Wang, Jianhui Chen +1 | cs.CL | 2026-09-02 |
| #10 | Balancing Frequencies and Pixels in Flow Matching | Lucas Degeorge, Paul Couairon, Arijit Ghosh +3 | cs.CV | 2026-09-02 |
| #11 | RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models | Canjie Liu, Jiawen Kang, Jinbo Wen +1 | cs.CV | 2026-09-02 |
| #12 | Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling | Ziquan Liu, Zhewei Zhu, Xuyang Shi | cs.CV | 2026-09-02 |
| #13 | Dimension Dependent Correlation Gap Bounds under Restricted Independence | Arjun Ramachandra | math.PR | 2026-09-02 |
| #14 | Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting | Ron Begleiter, Katya Egert Berg, Gilad Saban +1 | cs.AI | 2026-09-02 |
| #15 | RGB-to-IR image translation for infrared vehicle detection in unseen UAV domains | Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden +4 | cs.CV | 2026-09-02 |
| #16 | Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs | Xixiang He, Xingming Li, Baiqi Wu +4 | cs.LG | 2026-09-02 |
| #17 | TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis | Jiahao Wu, Zhenqun Yang, Chen Jason Zhang +1 | cs.LG | 2026-09-02 |
| #18 | Doppio: A Dataset for Contactless Weight Estimation of Falling Particles | Simon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5 | cs.CV | 2026-09-02 |
| #19 | ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering | Adrien Mialland, Marc Plantevit, Julien Gallois +1 | cs.IR | 2026-09-02 |
| #20 | Adapting a Foundation Model for Lunar Surface Height Estimation | Patrick Bauer, Marius Schwinning, Melanie Siegel +2 | cs.CV | 2026-09-02 |
| #21 | CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction | Menghao Li, Linjie Mu, Yin Wang +4 | cs.CV | 2026-09-02 |
| #22 | MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts | Matteo Greco, Anudeex Shetty, Andrea Tagarelli +1 | cs.CL | 2026-09-02 |
| #23 | TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models | Leqian Ding, Junning Qiu, Manwen Yang +2 | cs.CV | 2026-09-02 |
| #24 | LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory | Kun-Yang Yu, Yingzhe Li, Hongyu Xu +8 | cs.CV | 2026-09-02 |
| #25 | Towards Zero-Shot Transfer Across Embodiments For Driving VLAs | Caio Azevedo, Stefano Sabatini, Sascha Hornauer +1 | cs.CV | 2026-09-02 |
| #26 | YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification | Quansheng Hu, Qin Sun, Qiansen Dai +4 | cs.CV | 2026-09-02 |
| #27 | If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object Detection | Yinghao Sun, Shuguang Li, Jinliang Shao +1 | cs.CV | 2026-09-02 |
| #28 | Diffusion-Encoding Gaussian Field for Joint k-q dMRI Reconstruction | Zhibo Chen, Yajuan Huang, Yu Guan +3 | cs.CV | 2026-09-02 |
| #29 | RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification | Jierui Li, Zhiyuan Qi, Hao Zhu +7 | cs.CV | 2026-09-02 |
| #30 | InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models | Chengyin Hu, Dingyi Lu, Jiaju Han +5 | cs.CV | 2026-09-02 |
| #31 | Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging | Khawaja Murad ul Hassan, Mehran Ebrahimi | cs.CV | 2026-09-02 |
| #32 | Asymmetric Paired-Annotation Learning for Multi-Structure ULF Pediatric Brain MRI Segmentation | Ha-Hieu Pham, Dang P. M. Cao, Minh Hoang Pham +4 | cs.CV | 2026-09-02 |
| #33 | LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images | Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov | cs.CV | 2026-09-02 |
| #34 | TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval | Uicheol Jung, Juyoung Hong, Hojung Kwon +1 | cs.CV | 2026-09-02 |
| #35 | Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding | Shanji Liu, Kelu Yao, Junxiao Xue +5 | cs.CV | 2026-09-02 |
| #36 | GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation | Qianqian Wang, Yunshan Li, Jiawen Zeng +2 | cs.IR | 2026-09-02 |
| #37 | World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models | Chuhan Zhang, Seiji Ito, Kenta Hoshino +2 | cs.CV | 2026-09-02 |
| #38 | EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision | Ziyuan Jin, Yuxuan Ge, Zheng Tian | cs.AI | 2026-09-02 |
| #39 | Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap | Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh +2 | cs.CV | 2026-09-02 |
| #40 | MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs | Youssef Ennouri, Soonhoi Ha | cs.DC | 2026-09-02 |
| #41 | Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts | Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal +2 | cs.LG | 2026-09-02 |
| #42 | Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision | Sitong Pan, Yipeng Shen, Yilin Lu +3 | cs.AI | 2026-09-02 |
| #43 | Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents | Yanting Yang, Can Jin, Jinman Zhao +6 | cs.LG | 2026-09-02 |
| #44 | Test-Time Logit Prompting for Source-Free Missing Modality Adaptation | Taixi Chen, Nancy Guo | cs.CV | 2026-09-02 |
| #45 | SelfLift: Accelerating Few-Step Diffusion via Self-Recovering Resolution Transition | Tingyan Wen, Chenqian Yan, Xurui Peng +4 | cs.CV | 2026-09-02 |
| #46 | Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification | Xuanbing Wen, Boxu Chen, Le Yang +4 | cs.CV | 2026-09-02 |
| #47 | Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework | Shuyao Xiao, Shengling Wang, Haoyu Niu +4 | cs.CV | 2026-09-02 |
| #48 | On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers | Vignesh Prabhakar, Jialing Pan, Anil Babu Ankisettipalli | cs.LG | 2026-09-01 |
| #49 | Learning with Volterra Neural Networks: A System Theoretic Perspective | Haoyu Yun, Hamid Krim, Yufang Bao | cs.CV | 2026-09-01 |
| #50 | Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence? | Wenlong Wang, Fergal Reid | cs.AI | 2026-09-01 |