| 1 | AdamX: Cosine similarity meets gradient descent | Francisco Caldas, Ruben Belo, Cláudia Soares | cs.LG | 2026-09-10 |
| 2 | Generalization Analysis of Distributed Kernel-based Robust Gradient Descent Algorithms | Jun-Yi Meng, Zheng-Chu Guo, Yuan Mao | stat.ML | 2026-09-10 |
| 3 | Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry | Mo Zhou, Weihang Xu, Simon S. Du +1 | cs.LG | 2026-09-09 |
| #4 | Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending | Gijs A. F. Niewzwaag, Marijn G. S. Veth, Manuele Massei +1 | cs.LG | 2026-09-09 |
| #5 | Online Inverse Integer Linear Optimization via Small-Gradient Skipping: Constant Regret and Finite Mistakes | Akira Kitaoka | cs.LG | 2026-09-09 |
| #6 | Settling: Equilibrium Inference for Non-Convex Validity Sets | Lyes Saad Saoud | cs.LG | 2026-09-09 |
| #7 | Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader | Ashkan Soleymani, Gabriele Farina, Patrick Jaillet | cs.LG | 2026-09-08 |
| #8 | Silver Rate Is (Almost) Optimal for Gradient Descent Acceleration | Yuhan Ye, Kaizhao Liu | math.OC | 2026-09-08 |
| #9 | The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives | Ruijie Li, Kang Chen, Tianyu Wang | math.OC | 2026-09-08 |
| #10 | Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent | Quang-Duy Tran, Trung Le, Bao Duong +2 | cs.LG | 2026-09-08 |
| #11 | Speed Limit for Information Acquisition in Stochastic Learning Dynamics | Shuta Kobayashi, Andreas Dechant | cond-mat.stat-mech | 2026-09-08 |
| #12 | Sparse Data Augmentation for Optimization with Provable Guarantees | Behrooz Tahmasebi, Melanie Weber | cs.LG | 2026-09-08 |
| #13 | A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay | Yuqing Wang, Ioannis G. Kevrekidis, Mikhail Belkin | cs.LG | 2026-09-07 |
| #14 | Novel Methods for Catheter and Guidewire Segmentation in X-ray Fluoroscopy under a Federated Learning Setting | Chayun Kongtongvattana | cs.CV | 2026-09-06 |
| #15 | Stochastic Nonconvex Bilevel Optimization: Improved Rates Without Rare-Visit Assumption | Daniel Cortild, Mathias Staudigl, Juan Peypouquet +1 | math.OC | 2026-09-06 |
| #16 | Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks | Yiming Ying | cs.LG | 2026-09-06 |
| #17 | Fast Gauss Sums via Flash Attention | Nicolaj Rux, Sebastian Neumayer | cs.LG | 2026-09-04 |
| #18 | Centered Permutation Prefixes for SGD with Random Reshuffling: Sharp Rates, Hölder Geometry, and Composite Proximal Extensions | Jiaxiang Li | math.OC | 2026-09-04 |
| #19 | High-Dimensional Learning Dynamics of Attention-Indexed Models | Yizhou Xu, Margarita Sagitova, Lenka Zdeborová +1 | cs.LG | 2026-09-03 |
| #20 | Projected Riemannian Gradient Descent for the Bures-Wasserstein Barycenter: Dimension-Independent Linear Convergence at Unit Step Size | A. Afham | cs.LG | 2026-09-03 |
| #21 | Improved Gradient Descent Lower Bounds Beyond Nesterov | Yuhan Ye, Kaizhao Liu | math.OC | 2026-09-02 |
| #22 | Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance | Sai Niranjan Ramachandran, Suvrit Sra | cs.LG | 2026-09-02 |
| #23 | Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks | Osvaldo M Velarde, Lucas C Parra, Alireza Hashemi +1 | cs.LG | 2026-09-01 |
| #24 | Rethinking Learnability in Offline Data-driven Optimization | Chao Qian, Chen-Guang Wang, Rong-Xi Tan +1 | cs.LG | 2026-09-01 |
| #25 | The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow | Raphaël Berthier | cs.LG | 2026-09-01 |
| #26 | Subspace Levenberg Marquardt Algorithms in Training Neural Networks | M. Duc Hoang | cs.LG | 2026-09-01 |
| #27 | Operational Regimes in Non-Convex Optimization: A Multiplier-Based Taxonomy | Seyed Mohsen Kazemi, Ali Movaghar, Shaahin hessabi | math.OC | 2026-08-31 |
| #28 | Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not Commute | Xiaoyang Li, Runni Zhou | cs.LG | 2026-08-31 |
| #29 | Reciprocity Separates Gradient Flow from Rotation in Conservative Physical Learning | Ruiwu Niu, Xiaowen Bi, Michaël Antonie van Wyk | cs.LG | 2026-08-31 |
| #30 | Generalization as a robust performance property of learning-enabled dynamical systems | Filippo Fabiani | eess.SY | 2026-08-31 |
| #31 | Convergence rates for the RMSprop optimizer with full control of the hyperparameters | Steffen Dereich, Arnulf Jentzen | cs.LG | 2026-08-31 |
| #32 | REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent | Qian Zhang, Yaoming Li, Zhewen Tan +9 | cs.LG | 2026-08-30 |
| #33 | Quantitative Target Convergence and Uniform-in-Time Propagation of Chaos for Langevin-Regularized SVGD | Sayan Banerjee, Dohyeon Kim | stat.ML | 2026-08-28 |
| #34 | Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations | Yuehao Song, Zhong Chen, Lihui Cen +2 | math.NA | 2026-08-27 |
| #35 | Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules | Liviu Aolaritei, Lucas Lévy, Francis Bach +1 | cs.LG | 2026-08-26 |
| #36 | A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm | Duy Hoang, Bastien Berret, Olivier Bruneau +1 | cs.LG | 2026-08-25 |
| #37 | Dimensionless Controls of Plasticity Under Alternating Tasks: From Evolutionary Biology to Continual Learning | Owen Skriloff | math.OC | 2026-08-24 |
| #38 | Machine Learning Assisted Inverse Design of Pixelated mmWave Patch Antennas | Nadeem Rather, Holger Claussen, Lester Ho | eess.SP | 2026-08-24 |
| #39 | SGHA: A Single-Loop Fully First-Order Algorithm for Nonconvex-Strongly-Convex Bilevel Optimization | Zhihao Gu, Qilong Wu, Junchi Yang | math.OC | 2026-08-24 |
| #40 | One Inverse Step is a Convex Program: Bayes-Limit Calibration of Diffusion Inversion | Gordei Verbii | stat.ML | 2026-08-24 |
| #41 | Stochastic gradient descent with initial regularization | Nabil Kahalé | cs.LG | 2026-08-24 |
| #42 | Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture | Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde +1 | cs.AI | 2026-08-23 |
| #43 | Variational Structure at the Edge of Stability | Eric Regis | cs.LG | 2026-08-21 |
| #44 | Kähler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifold | Andrew Gracyk | cs.LG | 2026-08-20 |
| #45 | Quantum Tensor Network Learning with DMRG | Gustav J L Jäger, Martin B Plenio, Hans-Martin Rieser | quant-ph | 2026-08-19 |
| #46 | The Road Taken: The Role of Optimizers at the Edge of Stability | Jaerin Lee, Kyoung Mu Lee | cs.LG | 2026-08-19 |
| #47 | Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression | Sambit Mishra, Urbashi Mitra | cs.LG | 2026-08-17 |
| #48 | Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System | Alam Noor, Luis Almeida, Kai Li +3 | cs.CV | 2026-08-17 |
| #49 | Differentiable Voxelization of Surface Representations | Tobias Djuren, Ugo Finnendahl, Markus Worchel +2 | cs.GR | 2026-08-16 |
| #50 | Spectral Saliency for Machine Unlearning | Cedar Site Bai, Amber Yijia Zheng, Raymond A. Yeh +1 | cs.LG | 2026-08-16 |