| 1 | SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models | Junchao Huang, Guian Fang, Shengju Qian +15 | cs.CV | 2026-09-02 |
| 2 | Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit | Wassim Tenachi, Yashar Hezaveh, Laurence Perreault Levasseur +1 | cs.LG | 2026-09-02 |
| 3 | MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception | Guido Caccianiga, Sergey Prokudin, Yutong Chen +9 | cs.CV | 2026-09-02 |
| #4 | GaLe: memory-efficient Global Approximate and Local Exact features | Alberto Ancilotto, Elisabetta Farella | cs.CV | 2026-09-02 |
| #5 | LoFi RADIO: A Distilled In-Domain Backbone Applied for Artifact-Severity Grading of Ultra-Low-Field Neonatal Brain MR | Jonathan B. Martin, Yashwant Kurmi, Charlotte R. Sappo | eess.IV | 2026-09-02 |
| #6 | RGB-to-IR image translation for infrared vehicle detection in unseen UAV domains | Thijs A. Eker, Ella P. Fokkinga, Jan Erik van Woerden +4 | cs.CV | 2026-09-02 |
| #7 | Doppio: A Dataset for Contactless Weight Estimation of Falling Particles | Simon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5 | cs.CV | 2026-09-02 |
| #8 | Adapting a Foundation Model for Lunar Surface Height Estimation | Patrick Bauer, Marius Schwinning, Melanie Siegel +2 | cs.CV | 2026-09-02 |
| #9 | Seeing Beyond the Lesion: Disease Recognition from Reactive CNS Tissue | Jan Schnorrenberg, Jan Ernsting, Enrico Küllenberg +3 | eess.IV | 2026-09-02 |
| #10 | Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot Interactions | Maitreyee Tewari, Michele Persiani | cs.HC | 2026-09-02 |
| #11 | TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models | Leqian Ding, Junning Qiu, Manwen Yang +2 | cs.CV | 2026-09-02 |
| #12 | VoRTeC: Taming Foundation Flow for One-step Real time Video Compression | Yichong Xia, Qinhong Wu, Qinhong Wu +3 | cs.CV | 2026-09-02 |
| #13 | Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding | Shanji Liu, Kelu Yao, Junxiao Xue +5 | cs.CV | 2026-09-02 |
| #14 | Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap | Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh +2 | cs.CV | 2026-09-02 |
| #15 | A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization | Xianghong Fang, Wenlong Mou, Yuan Yuan +2 | cs.LG | 2026-09-02 |
| #16 | Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development | Yilong Guo, Hanqi Chen, Zixiao Ye +3 | cs.CV | 2026-09-02 |
| #17 | TC-Next: Zero-Shot Multimodal Cyclone Forecasting | Zhe Wang, Sijie Chen, Yiming Luo +2 | cs.LG | 2026-09-02 |
| #18 | Morphology signal in whole slide image foundation models can automatically triage slides | Ayushi Sinha, Shashank Yadav, Benjamin Holmes +9 | cs.CV | 2026-09-02 |
| #19 | Network-Aware Forecasting on Wireless Access Points | Niloo Bahadori, Swadhin Pradhan, Peiman Amini | cs.NI | 2026-09-02 |
| #20 | OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation | Yunqin Zhu, Feng Qiu, Yao Xie | cs.LG | 2026-09-01 |
| #21 | Cross-Model Distillation of a Human-Pose Foundation Model from Unannotated Infant Video for Markerless 3D Pose Estimation | R. James Cotton, Divya Joshi, Colleen Peyton | cs.CV | 2026-09-01 |
| #22 | Interpretable Symptom Vectors for Depression in a Large Language Model | Fangyi Zhu, Ajay Subramanian, Allison Constant +3 | cs.CL | 2026-09-01 |
| #23 | Kirin: Animal Motion Generation from In-the-Wild Video | Brian Nlong Zhao, Zhuoyang Pan, James M. Rehg +2 | cs.CV | 2026-09-01 |
| #24 | Video2Reaction: Training Foundation Video Models to Predict Audience Reaction | Sidong Zhang, Trang Nguyen, Shiv Shankar +4 | cs.CV | 2026-09-01 |
| #25 | D-FROST: Decentralized Federated pRompt-tuning via Optimal tranSporT for Non-IID and Imbalanced Data | Quan Minh Nguyen, Hoang M. Ngo, Trong Nghia Hoang +1 | cs.LG | 2026-09-01 |
| #26 | Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks | Osvaldo M Velarde, Lucas C Parra, Alireza Hashemi +1 | cs.LG | 2026-09-01 |
| #27 | Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models | Kunlin Cai, Kaiyuan Zhang, Zihang Xiang +4 | cs.CR | 2026-09-01 |
| #28 | Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation | Haoyuan Deng, Haichao Liu, Wenkai Guo +6 | cs.RO | 2026-09-01 |
| #29 | What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models | Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3 | cs.CV | 2026-09-01 |
| #30 | MegaStyle++: Scaling Image Style Space through Hierarchical Style Definition | Junyao Gao, Sibo Liu, Jiaxing Li +4 | cs.CV | 2026-09-01 |
| #31 | On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study | Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein | stat.ML | 2026-09-01 |
| #32 | Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations | Qingde Li, Qingqi Hong, Zihan Li +1 | cs.AI | 2026-09-01 |
| #33 | Bandits in Prod: Hyperparameter Optimization at Inference Time | Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine | cs.LG | 2026-09-01 |
| #34 | A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation | Hodong Lee, Sanghee Park, Dohoon Ryu +4 | cs.AI | 2026-09-01 |
| #35 | CMRVision: A Foundation Model for Cardiac MR Image Analysis | Athira J. Jacob, Puneet Sharma, Daniel Rueckert | cs.CV | 2026-09-01 |
| #36 | Seeing the World and the Self from Egocentric Video | Kai Guan, Minchao Jiang, Ruichen WangLi +2 | cs.CV | 2026-09-01 |
| #37 | One Prompt Is Enough: Watermark Laundering Through Foundation Image Models | Jidong Yang, Qi Li, Wei Zong +5 | cs.CV | 2026-09-01 |
| #38 | Monocular Depth Estimation from a Single Image: Progress and Opportunities | Muxin Liu, Xiaoyang Lyu, Yang-Tian Sun +4 | cs.CV | 2026-09-01 |
| #39 | A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference | Shuaicheng Niu, Guohao Chen, Yaofo Chen +14 | cs.LG | 2026-09-01 |
| #40 | Modelpedia: A Catalog of Model Findings for the Meta-Science of AI | Franciszek Bernat, Dawid Płudowski, Michał Jan Włodarczyk +6 | cs.LG | 2026-09-01 |
| #41 | AgentFactory: Towards Automated Agentic System Design and Optimization | Enci Zhang, Haofeng Wang, Yuesheng Zhu +2 | cs.AI | 2026-09-01 |
| #42 | MultiGait: A Multi-Sensor Multi-Perspective Multi-Session Biometric Inference Benchmark and its Dataset | Julian Todt, Felix Morsbach, Philip Dissert +1 | cs.CR | 2026-09-01 |
| #43 | ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation | Jeonghyeok Do, Seungchul Lee, Munchurl Kim | cs.CV | 2026-09-01 |
| #44 | Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting | Udo Schlegel, Shubhangi, Gabriel Dax +3 | cs.CV | 2026-09-01 |
| #45 | MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries | Utsab Ghosh, Roshni Chakraborty | cs.SD | 2026-09-01 |
| #46 | Instella-MoE Technical Report | Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra +10 | cs.CL | 2026-09-01 |
| #47 | Agentic Empirical Asset Pricing: Methodological Foundations | Yingjian Pan, Xiaowei Ding, Kay Giesecke | cs.AI | 2026-09-01 |
| #48 | FTU-Seek: Foundation Model-Guided Hard-Negative Learning for Sparse Functional Tissue Unit Segmentation | Zonghao Liu, Lei Su, Jiguang Yu +4 | cs.CV | 2026-09-01 |
| #49 | Do Satellites See Commuters? A Critical Benchmark of Vision Foundation Models | Ashiq Shukoor Iqbal, Wilson Wongso, Flora D. Salim | cs.CV | 2026-09-01 |
| #50 | EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction | Yunzhen Zhang, Ruoxi Piao, Hasan Onur Keles +1 | cs.LG | 2026-09-01 |