Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1,612 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Worth Their Weight: Randomized and Regularized Block Kaczmarz Algorithms without Preprocessing
Gil Goldshlager, Jiang Hu, Lin Lin · 18 December 2025
Due to the ever growing amounts of data leveraged for machine learning and scientific computing, it is increasingly important to develop algorithms that sample only a small portion of the data at a time. In the case of linear least-squares, the randomized block Kaczmarz method (RBK) is an appealing …
- A Teacher-Student Perspective on the Dynamics of Learning Near the Optimal Point
Carlos Couto, Jos\'e Mour\~ao, M\'ario A. T. Figueiredo, Pedro Ribeiro · 18 December 2025
Near an optimal learning point of a neural network, the learning performance of gradient descent dynamics is dictated by the Hessian matrix of the loss function with respect to the network parameters. We characterize the Hessian eigenspectrum for some classes of teacher-student problems, when the te…
- Dynamic Learning Rate Scheduling based on Loss Changes Leads to Faster Convergence
Shreyas Subramanian, Bala Krishnamoorthy, Pranav Murthy · 17 December 2025
Despite significant advances in optimizers for training, most research works use common scheduler choices like Cosine or exponential decay. In this paper, we study \emph{GreedyLR}, a novel scheduler that adaptively adjusts the learning rate during training based on the current loss. To validate the …
- CurvaDion: Curvature-Adaptive Distributed Orthonormalization
Bhavesh Kumar, Roger Jin, Jeffrey Quesnelle · 17 December 2025
As language models scale to trillions of parameters, distributed training across many GPUs becomes essential, yet gradient synchronization over high-bandwidth, low-latency networks remains a critical bottleneck. While recent methods like Dion reduce per-step communication through low-rank updates, t…
- Generalization performance of narrow one-hidden layer networks in the teacher-student setting
Jean Barbier, Federica Gerace, Alessandro Ingrosso, Clarissa Lauditi, Enrico M. Malatesta, Gibbs Nwemadji, Rodrigo P\'erez Ortiz · 17 December 2025
Understanding the generalization abilities of neural networks for simple input-output distributions is crucial to account for their learning performance on real datasets. The classical teacher-student setting, where a network is trained from data obtained thanks to a label-generating teacher model, …
- Bias-Variance Trade-off for Clipped Stochastic First-Order Methods: From Bounded Variance to Infinite Mean
Chuan He · 17 December 2025
Stochastic optimization is fundamental to modern machine learning. Recent research has extended the study of stochastic first-order methods (SFOMs) from light-tailed to heavy-tailed noise, which frequently arises in practice, with clipping emerging as a key technique for controlling heavy-tailed gra…
- Dropout Neural Network Training Viewed from a Percolation Perspective
Finley Devlin, Jaron Sanders · 17 December 2025
In this work, we investigate the existence and effect of percolation in training deep Neural Networks (NNs) with dropout. Dropout methods are regularisation techniques for training NNs, first introduced by G. Hinton et al. (2012). These methods temporarily remove connections in the NN, randomly at e…
- Evolving Deep Learning Optimizers
Mitchell Marfinetz · 16 December 2025
We present a genetic algorithm framework for automatically discovering deep learning optimization algorithms. Our approach encodes optimizers as genomes that specify combinations of primitive update terms (gradient, momentum, RMS normalization, Adam-style adaptive terms, and sign-based updates) alon…
- Better LMO-based Momentum Methods with Second-Order Information
Sarit Khirirat, Abdurakhmon Sadiev, Yury Demidovich, Peter Richt\'arik · 16 December 2025
The use of momentum in stochastic optimization algorithms has shown empirical success across a range of machine learning tasks. Recently, a new class of stochastic momentum algorithms has emerged within the Linear Minimization Oracle (LMO) framework--leading to state-of-the-art methods, such as Muon…
- Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences
Liviu Aolaritei, Michael I. Jordan · 16 December 2025
We study stopping rules for stochastic gradient descent (SGD) for convex optimization from the perspective of anytime-valid confidence sequences. Classical analyses of SGD provide convergence guarantees in expectation or at a fixed horizon, but offer no statistically valid way to assess, at an arbit…
- A Nonparametric Statistics Approach to Feature Selection in Deep Neural Networks with Theoretical Guarantees
Junye Du, Zhenghao Li, Zhutong Gu, Long Feng · 16 December 2025
This paper tackles the problem of feature selection in a highly challenging setting: $\mathbb{E}(y | \boldsymbol{x}) = G(\boldsymbol{x}_{\mathcal{S}_0})$, where $\mathcal{S}_0$ is the set of relevant features and $G$ is an unknown, potentially nonlinear function subject to mild smoothness conditions…
- Universality of high-dimensional scaling limits of stochastic gradient descent
Reza Gheissari, Aukosh Jagannath · 16 December 2025
We consider statistical tasks in high dimensions whose loss depends on the data only through its projection into a fixed-dimensional subspace spanned by the parameter vectors and certain ground truth vectors. This includes classifying mixture distributions with cross-entropy loss with one and two-la…
- Stochastic Bilevel Optimization with Heavy-Tailed Noise
Zhuanghua Liu, Luo Luo · 16 December 2025
This paper considers the smooth bilevel optimization in which the lower-level problem is strongly convex and the upper-level problem is possibly nonconvex. We focus on the stochastic setting where the algorithm can access the unbiased stochastic gradient evaluation with heavy-tailed noise, which is …
- Alada: Alternating Adaptation of Momentum Method for Memory-Efficient Matrix Optimization
Xiaoyu He, Yu Cai, Jin Jia, Canxi Huang, Wenqing Chen, Zibin Zheng · 16 December 2025
This work proposes Alada, an adaptive momentum method for stochastic optimization over large-scale matrices. Alada employs a rank-one factorization approach to estimate the second moment of gradients, where factors are updated alternatively to minimize the estimation error. Alada achieves sublinear …
- BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
Zhengyang Wang, Ziyue Liu, Ruijie Zhang, Avinash Maurya, Paul Hovland, Bogdan Nicolae, Franck Cappello, Zheng Zhang · 16 December 2025
The scale of transformer model pre-training is constrained by the increasing computation and communication cost. Low-rank bottleneck architectures offer a promising solution to significantly reduce the training time and memory footprint with minimum impact on accuracy. Despite algorithmic efficiency…
- A PyTorch Framework for Scalable Non-Crossing Quantile Regression
Kaihua Chang · 16 December 2025
Quantile regression is fundamental to distributional modeling, yet independent estimation of multiple quantiles frequently produces crossing -- where estimated quantile functions violate monotonicity, implying impossible negative probability densities. While Constrained Joint Quantile Regression (CJ…
- On the Approximation Power of SiLU Networks: Exponential Rates and Depth Efficiency
Koffi O. Ayena · 16 December 2025
This article establishes a comprehensive theoretical framework demonstrating that SiLU (Sigmoid Linear Unit) activation networks achieve exponential approximation rates for smooth functions with explicit and improved complexity control compared to classical ReLU-based constructions. We develop a nov…
- A Variance-Based Analysis of Sample Complexity for Grid Coverage
Lyu Yuhuan · 15 December 2025
Verifying uniform conditions over continuous spaces through random sampling is fundamental in machine learning and control theory, yet classical coverage analyses often yield conservative bounds, particularly at small failure probabilities. We study uniform random sampling on the $d$-dimensional uni…
- Gradient Descent as a Perceptron Algorithm: Understanding Dynamics and Implicit Acceleration
Alexander Tyurin · 15 December 2025
Even for the gradient descent (GD) method applied to neural network training, understanding its optimization dynamics, including convergence rate, iterate trajectories, function value oscillations, and especially its implicit acceleration, remains a challenging problem. We analyze nonlinear models w…
- Proof of a perfect platonic representation hypothesis
Liu Ziyin, Isaac Chuang · 12 December 2025
In this note, we elaborate on and explain in detail the proof given by Ziyin et al. (2025) of the ``perfect" Platonic Representation Hypothesis (PRH) for the embedded deep linear network model (EDLN). We show that if trained with the stochastic gradient descent (SGD), two EDLNs with different widths…
- Robust Gradient Descent via Heavy-Ball Momentum with Predictive Extrapolation
Sarwan Ali · 12 December 2025
Accelerated gradient methods like Nesterov's Accelerated Gradient (NAG) achieve faster convergence on well-conditioned problems but often diverge on ill-conditioned or non-convex landscapes due to aggressive momentum accumulation. We propose Heavy-Ball Synthetic Gradient Extrapolation (HB-SGE), a ro…
- The Interplay of Statistics and Noisy Optimization: Learning Linear Predictors with Random Data Weights
Gabriel Clara, Yazan Mash'al · 12 December 2025
We analyze gradient descent with randomly weighted data points in a linear regression model, under a generic weighting distribution. This includes various forms of stochastic gradient descent, importance sampling, but also extends to weighting distributions with arbitrary continuous values, thereby …
- Supervised Learning of Random Neural Architectures Structured by Latent Random Fields on Compact Boundaryless Multiply-Connected Manifolds
Christian Soize · 12 December 2025
This paper introduces a new probabilistic framework for supervised learning in neural systems. It is designed to model complex, uncertain systems whose random outputs are strongly non-Gaussian given deterministic inputs. The architecture itself is a random object stochastically generated by a latent…
- The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
Alexey Kravatskiy, Ivan Kozyrev, Nikolai Kozlov, Alexander Vinogradov, Daniil Merkulov, Ivan Oseledets · 11 December 2025
In this article, we explore the use of various matrix norms for optimizing functions of weight matrices, a crucial problem in training large language models. Moving beyond the spectral norm underlying the Muon update, we leverage duals of the Ky Fan $k$-norms to introduce a family of Muon-like algor…
- Online Inference of Constrained Optimization: Primal-Dual Optimality and Sequential Quadratic Programming
Yihang Gao, Michael K. Ng, Michael W. Mahoney, Sen Na · 11 December 2025
We study online statistical inference for the solutions of stochastic optimization problems with equality and inequality constraints. Such problems are prevalent in statistics and machine learning, encompassing constrained $M$-estimation, physics-informed models, safe reinforcement learning, and alg…
