| 1 | SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models | Junchao Huang, Guian Fang, Shengju Qian +15 | cs.CV | 2026-09-02 |
| 2 | Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework | Cagri Temel | cs.RO | 2026-09-02 |
| 3 | RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation | Xiaolei Lang, Ze Kang, Zehao Huang +1 | cs.CV | 2026-09-02 |
| #4 | Video-Based Palm-Vein Authentication under Challenging Conditions | Xiaofeng Yan, Kechen Liu, Abhilash Venkatesh +3 | cs.CV | 2026-09-02 |
| #5 | Query Rewriting for Complex Object Segmentation in 4D Gaussian Representations | Thanh-Khoi Nguyen, Thien-Phuc Tran, Minh-Triet Tran | cs.CV | 2026-09-02 |
| #6 | Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment | Zihao Wang, Xi Xiang, Yuwen Sun +5 | cs.CV | 2026-09-02 |
| #7 | MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval? | Uicheol Jung, Juyoung Hong, Geuntaek Lim +1 | cs.CV | 2026-09-02 |
| #8 | Doppio: A Dataset for Contactless Weight Estimation of Falling Particles | Simon Kiefhaber, Jan-Martin O. Steitz, Julia Grabinski +5 | cs.CV | 2026-09-02 |
| #9 | Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition | Naoto Nishida, Yoshio Ishiguro | cs.CV | 2026-09-02 |
| #10 | DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models | Yotam Eshel, Guy Hadad, Guy Feigenblat +3 | cs.LG | 2026-09-02 |
| #11 | The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation | Yichen Liu, Quanwei Zhang, Haozhe Wang +7 | cs.MM | 2026-09-02 |
| #12 | TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models | Leqian Ding, Junning Qiu, Manwen Yang +2 | cs.CV | 2026-09-02 |
| #13 | YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification | Quansheng Hu, Qin Sun, Qiansen Dai +4 | cs.CV | 2026-09-02 |
| #14 | VoRTeC: Taming Foundation Flow for One-step Real time Video Compression | Yichong Xia, Qinhong Wu, Qinhong Wu +3 | cs.CV | 2026-09-02 |
| #15 | Handwriting Trajectory Recovery via Autoregressive Ordered Stroke Instance Prediction | En-Guang Wang, Yan-Ming Zhang, Fei Yin +1 | cs.CV | 2026-09-02 |
| #16 | Prototype-guided transfer of sparse literature knowledge for electrolyte additive discovery | Weixiang Hong, Hongting Du, Jiayue Tang +4 | physics.chem-ph | 2026-09-02 |
| #17 | TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval | Uicheol Jung, Juyoung Hong, Hojung Kwon +1 | cs.CV | 2026-09-02 |
| #18 | Learning the Constitutive Behavior of Materials via Neural Operators and Causal Attention: Case Studies in Plasticity and Damage | Rishabh Arora, Lisa Scheunemann, Tim Brepols +1 | cs.LG | 2026-09-02 |
| #19 | Progressive Pseudo-Label Optimization for Point-Supervised Change Detection | Hailong Ning, Hao Wang, Yimeng Wang +3 | cs.CV | 2026-09-02 |
| #20 | FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs | Zhengyi Jin, Ru Zhang, Xiao Chen +5 | cs.AI | 2026-09-02 |
| #21 | C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees | S M Rafiuddin, Atriya Sen | cs.CL | 2026-09-02 |
| #22 | A Computational Comparison of Fourier Spectral Differentiation and Spatial Automatic Differentiation in Periodic Physics-Informed Neural Networks | Xilai Liang, Zhao Zhang | cs.LG | 2026-09-02 |
| #23 | Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents | Yanting Yang, Can Jin, Jinman Zhao +6 | cs.LG | 2026-09-02 |
| #24 | SelfLift: Accelerating Few-Step Diffusion via Self-Recovering Resolution Transition | Tingyan Wen, Chenqian Yan, Xurui Peng +4 | cs.CV | 2026-09-02 |
| #25 | Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework | Shuyao Xiao, Shengling Wang, Haoyu Niu +4 | cs.CV | 2026-09-02 |
| #26 | A Unified Particle Filter LSTM for Data-Driven Process Simulation | Parvin Malekzadeh, Opher Baron, Dmitry Krass | cs.LG | 2026-09-02 |
| #27 | Basin Geometry and Reliable Recall of Dynamical Memories in Reservoir Computing | Ling-Wei Kong, Ying-Cheng Lai | nlin.CD | 2026-09-01 |
| #28 | Accurate in space, unreliable in time: how LLMs represent national cultural change | Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp | cs.CY | 2026-09-01 |
| #29 | OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation | Yunqin Zhu, Feng Qiu, Yao Xie | cs.LG | 2026-09-01 |
| #30 | DESA-TTA: Dynamic EMA and Source Anchoring for Test-Time Adaptation | Atif Belal, Lilian Hollard, Marco Pedersoli +1 | cs.CV | 2026-09-01 |
| #31 | Allocate Before You Embed: Adaptive Visual Input Allocation for Video Embeddings | Song Jin, Zhongtao Jiang, Chenglei Shen +5 | cs.CV | 2026-09-01 |
| #32 | A Study of Conditional Diffusion Models for Open-Loop Control under Dry Friction and Stiction | Eric Aislan Antonelo | cs.LG | 2026-09-01 |
| #33 | Evidential Deep Learning for Multi-Modal Anti-UAV Detection | Dmitry Golovchits, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag | cs.CV | 2026-09-01 |
| #34 | From Visual Cues to Spoken Narration: Rethinking Audio Description | Akshita Gupta, Aditya Arora, Federico Tombari +2 | cs.CV | 2026-09-01 |
| #35 | Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models | Kunlin Cai, Kaiyuan Zhang, Zihang Xiang +4 | cs.CR | 2026-09-01 |
| #36 | H3-World: Turning Language Understanding into World Control | Danze Chen, Zeqing Wang, Ziyue Lin +2 | cs.CV | 2026-09-01 |
| #37 | What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models | Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3 | cs.CV | 2026-09-01 |
| #38 | TempCloze: Can Video-LLMs Identify the Missing Middle? | Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu +4 | cs.CV | 2026-09-01 |
| #39 | CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling | Xin Shen, Chengyou Jia, Keshuo Xing +6 | cs.CV | 2026-09-01 |
| #40 | Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds | Clinton Enwerem, John S. Baras, Calin Belta | cs.RO | 2026-09-01 |
| #41 | Predicting Subsurface Abnormalities Growth using Physics-Informed Neural Networks | Mehrdad Shafiei Dizaji, Hoda Azari | cs.LG | 2026-09-01 |
| #42 | CHARM: Character Hallucination for Multicultural Role Play Benchmark | Sunkyung Han, Nahyeon Park, Gaeun Seo +2 | cs.CL | 2026-09-01 |
| #43 | TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models | Chao Zhou, Yiling Chen, Qi Chu +3 | cs.CV | 2026-09-01 |
| #44 | Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations | Yi Fei Cheng, Fan Yang, Iremsu Bas +3 | cs.AI | 2026-09-01 |
| #45 | From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs | Jie Chen, Xiangqian Yu, Yanchao Lian +9 | cs.IR | 2026-09-01 |
| #46 | FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue | Hangyeul Lee, Juyoung Oh, Jaeyong Ko +5 | cs.AI | 2026-09-01 |
| #47 | On Synthesis of Metric Interval Temporal Logics | Hsi-Ming Ho, Shankaranarayanan Krishna, Khushraj Madnani | cs.LO | 2026-09-01 |
| #48 | Does This Moment Justify the Recommendation? Counterfactual Behavior-Grounded Evidence Retrieval for Personalized Video Recommendation | Xin Liu | cs.CV | 2026-09-01 |
| #49 | EvoGS: Modeling Deformation Evolution for Dynamic Gaussian Splatting | Wei Dong, Shahram Shirani, Jun Chen +1 | cs.CV | 2026-09-01 |
| #50 | The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research | Yuni Susanti, Moritz Schubotz | cs.DL | 2026-09-01 |