| 1 | Learning Length-Extrapolatable Recurrent Models | Hanwen Jiang | cs.LG | 2026-09-08 |
| 2 | Silver Rate Is (Almost) Optimal for Gradient Descent Acceleration | Yuhan Ye, Kaizhao Liu | math.OC | 2026-09-08 |
| 3 | ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR | Tommy Sha, Skylar Zhai, Siqi Zhao | cs.LG | 2026-09-08 |
| #4 | Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths | Qihao Yuan | cs.AI | 2026-09-08 |
| #5 | SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation | Linnan Zhao, Xu Liu, Lingling Li +3 | cs.CV | 2026-09-08 |
| #6 | Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics | Aleš Kučera, Karel Zimmermann | cs.RO | 2026-09-08 |
| #7 | Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation | Youngrok Park, Sangmin Bae, Hojung Jung +6 | cs.LG | 2026-09-08 |
| #8 | AXS-Net: Interpretable Deep Unfolding for Hyperspectral Image Denoising via Spectral Basis Unmixing and Structured Noise Refinement | Ziyi Guan, Jianping Zhang, Zheng Yang | cs.CV | 2026-09-08 |
| #9 | Suan: Rectifying Direct Preference Safety Alignment in Large Language Models | Oleksandr Cherednichenko, Roman Klypa | cs.LG | 2026-09-08 |
| #10 | Why shared attention vectors fail: a case for outcome-indexed tuning | Lenard Dome | cs.LG | 2026-09-08 |
| #11 | AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems | Jaewon Chu, Jinwoo Seo, Jaewon Cho +4 | cs.AI | 2026-09-08 |
| #12 | Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting | Jinwoo Park, Hyeongwon Kang, Pilsung Kang | cs.LG | 2026-09-08 |
| #13 | The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives | Ruijie Li, Kang Chen, Tianyu Wang | math.OC | 2026-09-08 |
| #14 | Equivariance Breaks the Learning Rate | Andrei Manolache, Mathias Niepert | cs.LG | 2026-09-08 |
| #15 | How to Make the Gradient Mapping Small for Constrained Stochastic Min-Max Problems and Beyond | Ahmet Alacaoglu | math.OC | 2026-09-08 |
| #16 | Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent | Quang-Duy Tran, Trung Le, Bao Duong +2 | cs.LG | 2026-09-08 |
| #17 | Adaptively Incorporating Directional Hints into Zeroth-Order Optimization | Alexander Ryabchenko, Jian Qian, Wenlong Mou | cs.LG | 2026-09-08 |
| #18 | zScore-N: A Neural Network for On-Chain Wallet Reputation Scoring | Girish G N, Ashutosh Sahoo, Akshay SP +2 | cs.AI | 2026-09-08 |
| #19 | Speed Limit for Information Acquisition in Stochastic Learning Dynamics | Shuta Kobayashi, Andreas Dechant | cond-mat.stat-mech | 2026-09-08 |
| #20 | DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory | Rui Bao, Zheng Gao, Xiaoyu Li +3 | cs.CR | 2026-09-08 |
| #21 | A Better Spur Should Start From Each Objective | Shanwen Mao, Hao Zhang, Guangtao nie +4 | cs.AI | 2026-09-08 |
| #22 | A Transformer-Based Delta Expression Encoder for Psilocybin Transcriptional Response: Architecture, Representations, and Biological Validation | Sai Jayakumar | q-bio.GN | 2026-09-08 |
| #23 | When Metrics Reward the Worst Translations: Internalizing Cultural Reasoning for Social Media Translation Evaluation | Yiwen Qiu, Linjuan Wu, Dingming Li +7 | cs.CL | 2026-09-08 |
| #24 | Topology-induced Operators Reveal Complementary Graph Representations without Training | Meng Qin, Jinqiang Cui, Hongwei Zheng +2 | cs.LG | 2026-09-08 |
| #25 | GPU-Enabled Large-Scale Optimization Using Randomized Linear Algebra | Pratik Rathore, Zachary Frangella, Parth Nobel +2 | cs.LG | 2026-09-08 |
| #26 | Sparse Data Augmentation for Optimization with Provable Guarantees | Behrooz Tahmasebi, Melanie Weber | cs.LG | 2026-09-08 |
| #27 | A Machine Learning Framework for Predicting Restaurant Food Waste to Support Sustainable Food Management | Md Mehedi Hasan Naeem, Md Ashraful Islam, Moumita Barua +2 | cs.LG | 2026-09-08 |
| #28 | A Gradient-based yet Spike-Timing-Dependent Solution to the Feedback Learning Problem in Neural Microcircuits | Xiangnan Zhang, Jingxin Liu, Ranqi Lu +6 | cs.NE | 2026-09-08 |
| #29 | TaskGuard: Task-Conditioned Restoration Utility for Risk-Aware Object Detection | Vung Pham | cs.CV | 2026-09-07 |
| #30 | Support Topology and Gradient Mixing in Sinkhorn Layers | Dylan Forde | cs.AI | 2026-09-07 |
| #31 | SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws | Konstantinos Christopher Tsiolis, Denny Wu, Christos Thrampoulidis +1 | stat.ML | 2026-09-07 |
| #32 | Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics | Hanwen Wang, Paris Perdikaris | cs.LG | 2026-09-07 |
| #33 | Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging | Nazish Khalid, Tausifa Jan Saleem, Amal Saqib +2 | cs.AI | 2026-09-07 |
| #34 | A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay | Yuqing Wang, Ioannis G. Kevrekidis, Mikhail Belkin | cs.LG | 2026-09-07 |
| #35 | Local gradient neural operator | Baiming Zhang, Jinsong Tang, Ying Xu +2 | cs.LG | 2026-09-07 |
| #36 | Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems | Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso | cs.DC | 2026-09-07 |
| #37 | Emergent Charging Coordination in Electric Delivery Fleets | Javier Vales-Alonso, Juan J. Alcaraz | cs.LG | 2026-09-07 |
| #38 | MpSub: A Momentum $p$-Dimensional Subspace Trust-Region Method for Derivative-Free Fine-Tuning of Large Language Models | Yuyang Wang, Haoyu Yao, Pengcheng Xie | cs.LG | 2026-09-07 |
| #39 | Translation of Black-Box Clinical Prediction Models into Standalone Transparent Nomograms: Temporal External Validation in Heart Transplantation | Henry Pigot, Paulo J. G. Lisboa, Sandra Ortega-Martorell +3 | cs.LG | 2026-09-07 |
| #40 | Beyond the Matrix Sign: Quadratic Spectral Descent | Qiaozhe Zhang, Jun Sun, Yingzhuang Liu | cs.LG | 2026-09-07 |
| #41 | Human mutation field reveals an equilibrium-like structure with irreversible circulation | Isabella Caranzano, Daniel Maria Busiello, Stefano Priorelli +2 | q-bio.GN | 2026-09-07 |
| #42 | TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning | Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann | cs.LG | 2026-09-07 |
| #43 | A Systematic Analysis of Automatic Differentiation versus Discretization-based Constraints for Physics-Informed PDE Solvers | Xing Guo, Hongwei Tang, Zewei Meng +3 | math.NA | 2026-09-07 |
| #44 | Impact of canny edge detection preprocessing on performance of machine learning models for Parkinson's disease classification | Sameer Bhat, Piotr Szczuko | cs.LG | 2026-09-07 |
| #45 | Parallelism Strategy Chaining for Fast Training Convergence | Minchul Kang, Changyong Shin, Younghun Go +4 | cs.LG | 2026-09-07 |
| #46 | Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration | Xiao Ma, Hong Shen, Hui Tian +2 | cs.LG | 2026-09-07 |
| #47 | PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians | Jiang Qin, Chunji Lv, Yangguang Wei +6 | cs.AI | 2026-09-07 |
| #48 | Mind the Approximation: Fisher-Weighted SVD Compression for ViTs | Moritz Thoma, Maximilian Groezinger, Maximilian Forstenhäusler +7 | cs.CV | 2026-09-07 |
| #49 | Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification | Yimeng Ye, Shuang Chen, Wenxuan Huang +8 | cs.LG | 2026-09-07 |
| #50 | Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer | Zhiwei Ning, Zhen Zhou, Puhua Jiang +7 | cs.CV | 2026-09-07 |