| 1 | Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision | Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali +1 | cs.CV | 2026-09-03 |
| 2 | Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views | Joseph Lee, Yidi Huang, Dokyoon Kim +2 | cs.CL | 2026-09-03 |
| 3 | Last Translation Benchmark | Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241 | cs.CL | 2026-09-03 |
| #4 | Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs | Yujie Zhang, Huiying Lan, Ehsan Aghapour +5 | cs.DC | 2026-09-03 |
| #5 | From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research | Yakov Pyotr Shkolnikov | cs.AI | 2026-09-03 |
| #6 | The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs | Yumeng Shi, Quanyu Long, Yin Wu +1 | cs.CV | 2026-09-03 |
| #7 | Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR | Boyan Li, Bingsen Chen, Chenghao Yang +3 | cs.CL | 2026-09-03 |
| #8 | Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis | Sixu Yan, Shikang Wang, Binhua Huang +11 | cs.RO | 2026-09-03 |
| #9 | Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving | Ruoyu Yao, Yusen Xie, Qingzhao Liu +5 | cs.CV | 2026-09-03 |
| #10 | The Blind Spot in 2D Infants' Pose Estimation:Robust Learning from Noisy Annotations | Emanuele Cardinale, Marco Proietti, Alessandro Cacciatore +3 | cs.CV | 2026-09-03 |
| #11 | Cooperative Multi-Task Semantic Communication for Joint Classification and Regression Tasks | Ahmad Halimi Razlighi, Mohammad Siddiqur Rahman, Maximilian H. V. Tillmann +2 | eess.SP | 2026-09-03 |
| #12 | OSR: Output Space Redistribution for Adaptive Label Removal in Classification Models | Minyi Peng, Darian Gunamardi, Ivan Tjuawinata +2 | cs.LG | 2026-09-03 |
| #13 | Sparse auto-regressive modeling for scene generation from multi-view images | Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel +4 | cs.CV | 2026-09-03 |
| #14 | OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping | Zelong Lv, Sicheng Xu, Jianfeng Xiang +5 | cs.CV | 2026-09-03 |
| #15 | Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated Learning | Michael Khavkin, Kichang Lee, Jaeho Jin +2 | cs.LG | 2026-09-03 |
| #16 | EF1-Constrained Nash Social Welfare with Identical Additive Valuations: Complexity, Guarantees, and Experiments | Zih-Sian Yang, Yi-Hao Chen, Yu-Te Kuan +3 | cs.GT | 2026-09-03 |
| #17 | Semantic Bayesian World Models | Tommaso Soru | cs.AI | 2026-09-03 |
| #18 | VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues | Ernesto Lozano, Alberto Jaenal, Javier Civera | cs.CV | 2026-09-03 |
| #19 | A Peer-Relative Representation Learning Framework for Energy Inefficiency Identification in Mobile Network Sites | Eliud Nyakweba Koto, Jaco du Toit, Adham Stoltz +1 | cs.LG | 2026-09-03 |
| #20 | LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes | Chuyan Chen, Haoxing Chen, Kun Chen +27 | cs.CV | 2026-09-03 |
| #21 | ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation | Javier del Pino, Salvador Rodríguez, Alejandro Garabito +2 | cs.CV | 2026-09-03 |
| #22 | KnowVis: Knowledge-Centric Visual Summarization for Video Lectures | Yi Xu, Yifan Hou, Xiaoyu Zhang | cs.CV | 2026-09-03 |
| #23 | Fill My Mirror: Geometry-Constrained Mirror Inpainting | Ofek Basson, Shimon Vainer, Yacov Hel-Or +1 | cs.CV | 2026-09-03 |
| #24 | Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks | Oline Ranum, Edward Fish, Simon Hadfield +1 | cs.CL | 2026-09-03 |
| #25 | Genetic Algorithms for Tractable Bayesian Network Fusion via Pre-Fusion Edge Pruning | Pablo Torrijos, José A. Gámez, José M. Puerta +1 | cs.NE | 2026-09-03 |
| #26 | SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval | Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden | cs.CV | 2026-09-03 |
| #27 | Observation-Conditioned Latent Energy Priors for Sparse Implicit Neural Shape Completion | Paul Büschl, Ezequiel de la Rosa, Julia Wolleb +3 | cs.CV | 2026-09-03 |
| #28 | Semantic-Aware Subgraph State Space Model for WSI Classification in Histopathology | Feixing Chen, Hao Lu, Lin Luo +1 | cs.CV | 2026-09-03 |
| #29 | CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding | Jing Jiang, Yiran Ling, Ruonan Li +2 | cs.CV | 2026-09-03 |
| #30 | Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning | Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob +14 | cs.LG | 2026-09-03 |
| #31 | Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations | Onat Şahin, Mohammad Altillawi, George Eskandar +2 | cs.CV | 2026-09-03 |
| #32 | Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements | Arianna Miola, Bruno Spaccavento, Lorenzo Silotto +2 | cs.CL | 2026-09-03 |
| #33 | Test-time adaptation for speech enhancement with an autoregressive speech prior | Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda +1 | cs.SD | 2026-09-03 |
| #34 | Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation | Xuanfa Jin, Zhijian Ma, Yongcheng Zeng +3 | cs.CL | 2026-09-03 |
| #35 | From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control | Vincenzo Norman Vitale, Mohammad Solki, Antonia Maria Tulino +2 | cs.NI | 2026-09-03 |
| #36 | The Attention Triangle in Audio-Video Models | Sagi Polaczek, Noa Kraicer, Gal Metzer +4 | cs.AI | 2026-09-03 |
| #37 | Text2Thermal: Physics-Aware Thermal Image Synthesis from Textual Priors | Tayeba Qazi, Brejesh Lall, Prerana Mukherjee | cs.CV | 2026-09-03 |
| #38 | Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation | Yinan Liu, Jiankang Hong, Zhen Gao +1 | cs.AI | 2026-09-03 |
| #39 | Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws | Jie Wang | cs.LG | 2026-09-03 |
| #40 | NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis | Yinan Liu, Hongtai Xia, Haoran Xu +3 | cs.AI | 2026-09-03 |
| #41 | Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection | Yue Zhao, Hua Yu, Yukun Zhao +6 | cs.CV | 2026-09-03 |
| #42 | BMCTrack-d: Pig re-identification and tracking via back marks in challenging camera settings | David Brunner, Maciej Oczak, Marie Bordes +3 | cs.CV | 2026-09-03 |
| #43 | Beyond Straightness: Non-Crossing Flow Matching via Quantile AlignTree Coupling | Junyi Lin, Mengyu Li, Jingxuan Hu +2 | cs.LG | 2026-09-03 |
| #44 | Guide, Not Bind: Why Defeasible Priors Fail in Augmented Lagrangian Causal Discovery | Sairam Sundararaman, Sara Girdhar, Manit Narasimha Murthy +2 | cs.LG | 2026-09-03 |
| #45 | Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning | Heng Wang, Jielin Qiu, Wenting Zhao +7 | cs.CL | 2026-09-03 |
| #46 | Mudragen: Geometrically Supervised Generation of Interacting Two-Hand Mudras for Preserving Indian Classical Dance Heritage | Jagadish Kashinath Kamble, Jayanta Mukhopadhyay, Debaditya Roy +1 | cs.CV | 2026-09-03 |
| #47 | Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation | Yuhe Wu, Guangyu Wang, Yujie Chen +7 | cs.AI | 2026-09-03 |
| #48 | Neural-Collapse-guided Task-Free Continual Anomaly Detection | Xiaotong Kong, Chaoyang Song, Ziai Zhou +3 | cs.CV | 2026-09-03 |
| #49 | Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data | Xiangyang Miao, Kelu Yao, Yekai Huang +7 | cs.CV | 2026-09-03 |
| #50 | TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents | Jinwei Gan | cs.LG | 2026-09-03 |