| 1 | SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators | Yuncong Yang, Zhengtao Han, Furkan Ozyurt +6 | cs.CV | 2026-09-08 |
| 2 | DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination | Yankai Fu, Ning Chen, Junkai Zhao +5 | cs.RO | 2026-09-08 |
| 3 | Do Reasoning Representations Help Humans Evaluate LLM Outputs? | Jaewoo Lim, Sungbok Shin, Sanghyun Hong | cs.LG | 2026-09-08 |
| #4 | Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild | Fei Teng, Sheng Wu, Mengfei Duan +7 | cs.CV | 2026-09-08 |
| #5 | SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL Evaluation | Mohammadhossein Malekpour, Mohamed Riahi, Maxime Lamothe +1 | cs.DB | 2026-09-08 |
| #6 | Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language Representation | Abdullah Al Shafi | cs.CL | 2026-09-08 |
| #7 | Length Generalization for Transformers via Compression | Georg Zetzsche, Hongjian Jiang, Andy Yang +4 | cs.LG | 2026-09-08 |
| #8 | Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics | Ruibo Ming, Lei Sun, Deheng Zhang +8 | cs.CV | 2026-09-08 |
| #9 | CoordFormer: Give Me Any Coordinates and I Will Give You Labels | Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli +1 | cs.CV | 2026-09-08 |
| #10 | AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems | Jaewon Chu, Jinwoo Seo, Jaewon Cho +4 | cs.AI | 2026-09-08 |
| #11 | CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning | Imane Chraki, Pierre Marza, Stergios Christodoulidis +1 | cs.CV | 2026-09-08 |
| #12 | From Glance to Scrutiny: Progressive Distortion Reasoning for Fine-Grained Image Quality Assessment | Aoting Zhang, Mingze Gao, Dongbao Yang +5 | cs.CV | 2026-09-08 |
| #13 | Human-Centric Image Captioning with Subject-Centered Spatial Understanding | Bozhou Li, Jiahang Zhang, Yue Ding +11 | cs.CV | 2026-09-08 |
| #14 | HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting | Namwoo Kim, Hyungryul Baik, Yoonjin Yoon | cs.LG | 2026-09-08 |
| #15 | MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation | Nagito Saito, Shintaro Ito, Koichi Ito +1 | cs.CV | 2026-09-08 |
| #16 | SoftRerank: Hierarchical Soft Fusion with Candidate-Label Reranking for Long-Tailed Micro-Action Recognition | Yichi Zhang, Zhichao Xia, Yanjun Chi +6 | cs.CV | 2026-09-08 |
| #17 | A Better Spur Should Start From Each Objective | Shanwen Mao, Hao Zhang, Guangtao nie +4 | cs.AI | 2026-09-08 |
| #18 | SIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection | Kehan Yan, Yue Tan, Qingfeng Chen +3 | cs.LG | 2026-09-08 |
| #19 | MamMA: A Mamba-Based Pedestrian Trajectory Prediction Algorithm Considering Occupancy Map and Pedestrian Awareness States | Juncen Long, Xiaofeng Jin, Gianluca Bardaro +2 | cs.CV | 2026-09-07 |
| #20 | VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities | Jiahao Shi, Edward Tsien, Yifeng Di +10 | cs.CR | 2026-09-07 |
| #21 | Flexible Motion Generation from Language and Style References | Kai Weixian Lan, Bodie Criswell, Briana Fedkiw +3 | cs.CV | 2026-09-07 |
| #22 | Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning | Vishwas Sathish, Viresh Ranjan, Xinliang Zhu +2 | cs.AI | 2026-09-07 |
| #23 | TDDN: Text-aligned Diffused DINO Network for Puzzle Understanding | Harsha Patnala, Debopriyo Banerjee, Ayush Sunil Munot +1 | cs.CV | 2026-09-07 |
| #24 | SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs | Pengfei Li, Naufal Suryanto, Sicheng Zhang +2 | cs.CV | 2026-09-07 |
| #25 | What Does an LLM-Agent Leaderboard Rank Actually Compare? | Wei-Jung Huang | cs.AI | 2026-09-07 |
| #26 | xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems | Yongchang Peng, Qingshui Gu, Liya Zhu +31 | cs.AI | 2026-09-07 |
| #27 | Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs | Shahzeb Qamar, Lorenz Sparrenberg, Christian Bauckhage +5 | cs.LG | 2026-09-07 |
| #28 | TeMo: Temperature Modulation for Multimodal Contrastive Learning | Dhimitrios Duka, Bernt Schiele, Hilde Kuehne +1 | cs.CV | 2026-09-07 |
| #29 | CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements | Hongxiang Zhao, Mutian Xu, Zeyu Jin +3 | cs.RO | 2026-09-07 |
| #30 | MEMO: Multimodal Evidence Memory Organization for Long-Horizon LLM Agents | Xian Gao, Jinpeng Wang, Jiacheng Ruan +3 | cs.CL | 2026-09-07 |
| #31 | Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval | PS Rishi, Rajeev Ranjan Dwivedi, Vinod K Kurmi | cs.IR | 2026-09-07 |
| #32 | REFINE: Trajectory Representation Learning via Closed-Loop Transcription -- Extended Version | Sean Bin Yang, Ying Sun, Jilin Hu +5 | cs.LG | 2026-09-07 |
| #33 | Fine-grained Distributed Backdoor Attacks in Federated Learning | Jian Wang, Hong Shen, Wei Ke +1 | cs.LG | 2026-09-07 |
| #34 | Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection | Jiangning Zhu, Bowen Li, Shenyu Qiao +4 | cs.CV | 2026-09-07 |
| #35 | Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval | JungMin Yun, YoungBin Kim | cs.AI | 2026-09-07 |
| #36 | Large Discrete Policy: Advancing Explicit Behavior Modeling with Stochastic Iterative Scoring | Zhenxin Li, Nadine Chang, Xinglong Sun +8 | cs.RO | 2026-09-07 |
| #37 | Fine-Grained Visual Preprocessing and Dual-Stream Temporal Modeling for Multimodal Sentiment Analysis on Social Media | Su Li, Yigong Zhang, Lei Xiong +1 | cs.CV | 2026-09-07 |
| #38 | AnomalyCraft-700K: Component-Level Controllable and Verifiable Synthetic Anomalies for Fine-Grained Video Anomaly Understanding | Yuzhou Long, Haodong Zhang, Yunpeng Yang +2 | cs.CV | 2026-09-07 |
| #39 | Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras | Jiaqi Chen, Qinfu Xu, Hao Zhuang +1 | cs.CV | 2026-09-07 |
| #40 | MedGSSR: Generalizable Medical Image Super-Resolution 3D Reconstruction via Hierarchical Feed-forward Gaussian Splatting | Chengkai Wang, Luoyu Hong, Yiting Zhao +7 | eess.IV | 2026-09-06 |
| #41 | AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths | MaoXun Huang, Zhenxing Zhang, Claire Cardie | cs.CL | 2026-09-06 |
| #42 | ADELE - Adaptive Delaunay Grids for High-Fidelity Mesh-Native Reconstruction | Johannes Weidenfeller, Shaofei Wang, Philipp Fürnstahl +1 | cs.CV | 2026-09-06 |
| #43 | DianShi-RxnDB: A Large-Scale, Fine-Grained Organic Reaction Data Platform Built via a Fully Automated Pipeline for Researchers and AI Agents | Yubin Wang, Xingjian Wei, Jiang Wu +33 | cs.CL | 2026-09-06 |
| #44 | Data Efficient Sample Selection for In-Context Learning | V Venktesh, Cem levi, Avishek Anand | cs.LG | 2026-09-06 |
| #45 | EviMap: Evidence-Grounded Hierarchical Topic Maps for Exploring Unlabeled Corpora | Zhiyin Tan, Changxu Duan | cs.IR | 2026-09-06 |
| #46 | FSAN: Flow State Attention Network for Aerodynamic Prediction | Wenxuan Jin, Jianguo Yao, Haibing Guan +1 | cs.CV | 2026-09-06 |
| #47 | VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes | Yan Ma, Jiadi Su, Zhulin Hu +4 | cs.CV | 2026-09-06 |
| #48 | TD-STGT: A Spatio-Temporal Graph Transformer for Mobile Traffic Demand Forecasting | Mohamad Alkadamani, Halim Yanikomeroglu | eess.SY | 2026-09-06 |
| #49 | Multi-History-Step SDE Inversion for Image Editing with Superior Regional Awareness | Haiyan Wei, Yunlong Wang, Huaibo Huang +2 | cs.CV | 2026-09-06 |
| #50 | CAM: Question Answering on Entity-Centric Videos with Continuous Extraction and Adaptive Querying | Yizhou Tian, Zizhe Chen, Shiyuan Deng +7 | cs.CV | 2026-09-06 |