| 1 | EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning | Jingpu Yang, Fengxian Ji, Mingxuan Cui +4 | cs.CV | 2026-09-08 |
| 2 | Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics | Aleš Kučera, Karel Zimmermann | cs.RO | 2026-09-08 |
| 3 | The Rater Ising-Potts Model with LLM-Derived Weights: An Application to Multi-Category Scoring Reliability | Matthias von Davier | stat.AP | 2026-09-08 |
| #4 | Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks | Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah +1 | cs.CR | 2026-09-08 |
| #5 | The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution? | Boyang Wang, Yunhan Wang, Yalun Wu | cs.SE | 2026-09-08 |
| #6 | Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting | Jinwoo Park, Hyeongwon Kang, Pilsung Kang | cs.LG | 2026-09-08 |
| #7 | CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning | Imane Chraki, Pierre Marza, Stergios Christodoulidis +1 | cs.CV | 2026-09-08 |
| #8 | Miles v0.1: Production-Level Post-Training | RadixArk, :, Tom Chen +11 | cs.LG | 2026-09-08 |
| #9 | VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent | Kevin Chuanpu Fu, Yongsen Zheng, Zee Kin Yeong +1 | cs.CV | 2026-09-08 |
| #10 | Agentic ML Exploration (A-MLE) for Ads Ranking | Erwin Gao, Vinodh Kumar Sunkara, Jingyi Guan +36 | cs.AI | 2026-09-08 |
| #11 | Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation | Runzhong Zhang, Yueqi Duan, Yang Chen +4 | cs.CV | 2026-09-08 |
| #12 | Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation | Jianqiang Xiao, Xiang Deng, Yuexuan Sun +3 | cs.RO | 2026-09-08 |
| #13 | SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents | Pujun Zheng, Zixin Shang, Shufan Jiang +5 | cs.AI | 2026-09-08 |
| #14 | Semi-Supervised Learning under Spatially Biased Sampling | Bright Wiredu Nuakoh, Francky Fouedjio, Stephen Bradshaw +4 | cs.LG | 2026-09-07 |
| #15 | Explainable Temporal Attention-based Defect Detection For Fillet Joints in Real-Time Gas Metal Arc Welding Based on Multi-modal Data | Mobina Mobaraki, Mahyar Asadi, Klaske Van Heusden +1 | cs.AI | 2026-09-07 |
| #16 | SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs | Pengfei Li, Naufal Suryanto, Sicheng Zhang +2 | cs.CV | 2026-09-07 |
| #17 | Kalman Delta Networks: Uncertainty-aware Associative Memory | Ngoc Bui, Tinglin Huang, Rex Ying | cs.LG | 2026-09-07 |
| #18 | Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging | Nazish Khalid, Tausifa Jan Saleem, Amal Saqib +2 | cs.AI | 2026-09-07 |
| #19 | When Superpixels Fail on Documents: A Study of Segmentation for LIME Explanations | Quentin Telnoff, Emanuela Boros, Mickaël Coustaty +3 | cs.CV | 2026-09-07 |
| #20 | Unified Vision-Centric Pedestrian Crossing Action Prediction via Adaptive Patch Projection and Proactive Spatial Rectification | Yao Tian, Le Yang, Binglu Wang | cs.CV | 2026-09-07 |
| #21 | Beyond Fluent Generation: A CPU Reliability Benchmark for MCP-Style Tool Calling in Sub-2B Small Language Models for Edge Deployment | Abrar Shahriar Qurat-Ul-Ain Mastoi | cs.CL | 2026-09-07 |
| #22 | D3ARC: Time-Critical Distributed Disaster Detection for Asynchronous Cooperative Multi-Robot Systems | Nikolaos Koursioumpas, Lina Magoula, Nancy Alonistioti +1 | cs.RO | 2026-09-07 |
| #23 | TAD: Token-Adaptive Contrastive Decoding with Confidence-Guided Gating for Hallucination Mitigation in Large Audio-Language Models | Heyu Chang, Nianwen Si, Hao Zhang +2 | cs.SD | 2026-09-07 |
| #24 | Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration | Xiao Ma, Hong Shen, Hui Tian +2 | cs.LG | 2026-09-07 |
| #25 | EmoMed: An Emotionally-Aware Agent for Multimodal Medical Support with Real-Time Information Retrieval | Ivan Nasonov, Nikita Glazkov, Ivan Makovetskiy +4 | cs.AI | 2026-09-07 |
| #26 | CHILD: Human-in-the-Loop OOD Detection for Safe Clinical Deployment | Jinlun Ye, Kaiyue Lu, Runhe Lai +3 | cs.CV | 2026-09-07 |
| #27 | Beyond Task Success: Stage-Wise Reliability of World Model Planning under Sensing Degradation | Geonmyeong Lee, Byoung-Tak Zhang | cs.RO | 2026-09-07 |
| #28 | PCSDiff: Diffusion-Based Bias Correction and Super Resolution Toward Practical Operational Medium-Term Precipitation Forecast | Yuze Sun, Shiyi Wang, Jiancheng Pan +7 | cs.LG | 2026-09-07 |
| #29 | Characterizing Contention-Induced Reliability Collapse in KV-Cache Timing Side Channels for Multi-Tenant LLM Serving | Rana Abu Bakar | cs.CR | 2026-09-06 |
| #30 | SerenAI: State-transition system inspired by text-based world AI models | Elvin Babayev, Artem Sinitsa, Arash Hajisharifi +1 | cs.AI | 2026-09-06 |
| #31 | SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives | Tianyu Wang, Nianjun Zhou | cs.CL | 2026-09-06 |
| #32 | Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability for Small-Object Detection | Zhaoning Shi, Bo Ma | cs.CV | 2026-09-06 |
| #33 | Separating Capability from Confidence: Grounded Dual-State Calibration for GRPO-Trained Medical Vision-Language Models | Yangyang Xie, Ke Hao, Jiaqi Liu +2 | cs.CV | 2026-09-06 |
| #34 | Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction | Lihui Liu | cs.LG | 2026-09-06 |
| #35 | ChildGaze: A Benchmark Dataset for Collaborative Behavior Understanding in Children | Sindhuja Penchala, Saketh Reddy Kontham, Prachi Bhattacharjee +7 | cs.CV | 2026-09-06 |
| #36 | Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring | Chunyi Zhao, Chao Li | cs.CL | 2026-09-06 |
| #37 | Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence | Urja Pawar, Rajitha Ramanayake, Nabeel Kemal +4 | cs.AI | 2026-09-04 |
| #38 | PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting | Maryam Fakhari, Mehran Safayani | cs.LG | 2026-09-04 |
| #39 | A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR | Thi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo +2 | cs.CL | 2026-09-04 |
| #40 | A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment | Balaji Venkateswaran | cs.AI | 2026-09-04 |
| #41 | VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps | Sunesh Praveen Raja Sundarasami, Taehyoung Kim, Johannes Scherer +5 | cs.CV | 2026-09-04 |
| #42 | MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning | Guanglong Sun, Kanglei Zhou, Liyuan Wang +6 | cs.AI | 2026-09-04 |
| #43 | TreeFI: Value-Aware Statistical Fault Injection for Deep Neural Networks | Noam Bires, Marcello Traiola, Angeliki Kritikakou +1 | cs.AR | 2026-09-04 |
| #44 | From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments | Linsen Zhu, Mengqing Cai | cs.AI | 2026-09-04 |
| #45 | MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models | Changming Xiao, Zhenliang Ni, Jinhui He +2 | cs.AI | 2026-09-04 |
| #46 | LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus | Chi Zhang, Qi Song, Feifei Li +2 | cs.CV | 2026-09-04 |
| #47 | CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving | Zhe Huang, Zhaoxin Fan, Shuo Wang +3 | cs.RO | 2026-09-04 |
| #48 | An Attention-Guided Global and Local Fusion Framework for Lesion-Focused Image Classification | Mst Shafia Tasnima, Md Samaun Elaheea, Tanjim Taharat Aurpab +1 | cs.CV | 2026-09-04 |
| #49 | Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges | Chenqi Li, Minghui Min, Dusit Niyato +1 | cs.AI | 2026-09-04 |
| #50 | Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle | Happy Bhati | cs.SE | 2026-09-04 |