Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1,612 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway
Hee-Sung Kim, Sungyoon Lee · 5 June 2026
Recent analyses of multi-pathway Deep Linear Networks use Gradient Flow to predict a "winner-takes-all" specialization in which path symmetry breaks and each feature concentrates in a single pathway. In this work, we show that discrete Gradient Descent (GD) with a large step size tells a different s…
- Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models
Hancheol Park, Geonho Lee, Tairen Piao, Tae-Ho Kim · 5 June 2026
Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of expert parameters still makes quantization essential for practical deployment. Unlike dense models, however, MoE models are sensitive to routing instab…
- PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training
Senmiao Wang, Tiantian Fang, Haoran Zhang, Yushun Zhang, Kunxiang Zhao, Alex Schwing, Ruoyu Sun · 5 June 2026
We propose a preconditioning (PC) layer, a weight parameterization via polynomial preconditioner that ensures stable weight conditioning throughout LLM training. The PC module reshapes the singular-value spectrum of weight matrices via low-degree polynomial preconditioning. After training, the preco…
- Near-Optimal Decentralized Stochastic Convex Optimization over Networks
Nitai Kluger, Amit Attia, Tomer Koren · 4 June 2026
We study decentralized stochastic smooth convex optimization, where $M$ workers minimize an average objective using local stochastic gradients and neighbor-only communication over a fixed gossip network. A central question in this setting is to determine the largest number of workers that can be use…
- Adalina: Adaptive Linear Approximation for the Shapley Value and Beyond
Weida Li, Yaoliang Yu, Bryan Kian Hsiang Low · 4 June 2026
The Shapley value, and its broader family of semi-values, has received much attention in various attribution problems. A fundamental and long-standing challenge is their efficient approximation, since exact computation generally requires an exponential number of utility queries in the number of play…
- When Do Fewer Coordinates Suffice in DP-SGD?
Huiqi Zhang, Fang Xie · 4 June 2026
Differentially private stochastic gradient descent (DP-SGD) injects noise into every updated coordinate, making the injected noise energy scale with the ambient parameter dimension \(d\). We ask when private training can update fewer coordinates without losing the signal needed for optimization. We …
- When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks
Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi · 4 June 2026
In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function. Concretely, we consider a realizable setting where inputs are drawn i.i.d. from a Gaussian distribution and labels follow a planted linear model.…
- Low-Rank Decay for Grokking in Scale-Invariant Transformers: A Spectral-Geometric View
Mingyu Li · 4 June 2026
Modern Transformer architectures frequently employ normalization mechanisms such as RMSNorm and Query-Key Normalization, making parts of the model approximately scale-invariant with respect to weight magnitudes. In this regime, standard Frobenius-norm weight decay acts purely along the radial direct…
- Pseudospectral Bounds for Transient Amplification in Coupled Gradient Descent
Ahanaf Hasan Ariq · 4 June 2026
Coupled gradient descent--where the update of one parameter block depends on another--underlies bilevel optimization, two-time-scale stochastic approximation, and adversarial training. When the coupled Jacobian is block-triangular, asymptotic stability is governed by the spectral radii of the diagon…
- Flatness and Generalization: Learning Multi-Index Models with Homogeneous Neural Networks
Harsh Vardhan, Hossein Taheri, Arya Mazumdar · 4 June 2026
A common heuristic used to explain the generalization of first-order gradient methods on non-convex neural networks is that "flat interpolators generalize well" (Hochreiter and Schmidhuber, 1994; Keskar et al., 2017), where flatness can be measured by the trace of the Hessian of the empirical loss. …
- A Geometric Characterization of the Stationary Plateau for Two-Layer Neural Networks
Tian Ding, Dawei Li, Ruoyu Sun · 4 June 2026
We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions. We focus on the phenomenon of "neuron splitting" where duplicating a hidden neuron yields an affine set of stationary points in a wider networ…
- Edge of Stability Selectively Shapes Learning Across the Data Distribution
Shauna Kwag, Anakha Ganesh, Tomaso Poggio, Pierfrancesco Beneventano · 4 June 2026
Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constraint redistributes learning across subsets of the training distribution, amplifying progress on some groups while suppressing progress on others. Usi…
- Demystifying Pipeline Parallelism: First Theory for PipeDream
Ivan Ilin, Peter Richt\'arik · 3 June 2026
Training modern machine learning models increasingly requires computation to be distributed across many accelerators. Data parallelism remains the default choice and is often paired with tensor-parallel sharding, but model parallelism becomes unavoidable once parameters, activations, or optimizer st…
- Spectral Asymptotics of Neural Network Loss Landscapes: An Exact Decomposition of the Curvature Exponent
Anherutowa Calvo · 3 June 2026
The curvature exponent $\alpha$ in $h_k \propto \sigma_k^\alpha$ -- governing how Hessian eigenvalues scale with gradient singular values -- varies systematically across layer types ($\alpha \approx 2$ for convolutions, $\approx 1$ for transformer attention, $< 1$ for MLP up-projections). Why? We pr…
- Compress then Merge: From Multiple LoRAs into One Low-Rank Adapter
Zhengbao He, Ruiqi Ding, Zhehao Huang, Ruikai Yang, Tao Li, Xiaolin Huang · 3 June 2026
Low-rank adaptation (LoRA) enables parameter-efficient specialization of foundation models, but the proliferation of task-specific adapters fragments capabilities across many adapters, complicating reuse and deployment. We study the problem of merging $T$ LoRAs into a single rank-$r$ LoRA, thereby p…
- Hierarchical RBF-KAN and RBF-SKAN Architectures for Multidimensional Function Approximation and Random Field Learning
Mingtao Xia, Qijing Shen · 3 June 2026
In this manuscript, we propose and analyze hierarchical Kolmogorov--Arnold neural network architectures employing radial basis functions as activation functions for approximating deterministic functions and random field models. Specifically, we develop a hierarchical radial-basis-function Kolmogorov…
- Pruning Deep Neural Networks via the Marchenko--Pastur Distribution
Leonid Berlyand, Theo Bourdais, Houman Owhad, Yitzchak Shmalo · 3 June 2026
We study a Marchenko--Pastur (MP) random-matrix approach to pruning deep neural networks with very small post-pruning fine-tuning budgets. The main practical contribution is accuracy retention under short calibration and fine-tuning schedules, rather than a long post-pruning reoptimization pipeline.…
- Analytical Evaluation of DCA Convergence Properties for Minimizing Prediction Functions of Gaussian RBF Support Vector Regression
Yohei Kakimoto, Yuto Omae, Hirotaka Takahashi · 3 June 2026
For nonconvex optimization problems whose objective is the prediction function of a trained Support Vector Regression (SVR) model with the Gaussian radial basis function (RBF) kernel (RBF-SVR), we present a framework that applies the difference of convex functions (DC) algorithm (DCA) by exploiting …
- Near-Optimal Pure Machine Unlearning for Smooth Strongly Convex Losses
Matthew Regehr, Gautam Kamath, Andrew Lowy · 2 June 2026
Machine unlearning is motivated by legal and user-facing requirements to remove the influence of individuals' data from trained models, such as the right to be forgotten. Prior work has developed algorithms and error bounds for unlearning in smooth strongly convex stochastic optimization, but the fu…
- Fast Generalization after Interpolation via Critically Damped Momentum Optimization
Luca Muscarnera, Silas Ruhrberg Est\'evez, Yuanzhang Xiao, Mihaela Van der Schaar · 2 June 2026
A central problem in machine learning is that models can achieve near-perfect training performance while generalizing substantially less well to unseen examples. This gap is especially acute in high-dimensional, low-sample regimes, where many interpolating solutions exist and optimization must impli…
- Neural Network Compression by Approximate Differential Equivalence
Ravi Dhiman, Andrea Passarella, Mirco Tribastone, Lorenzo Valerio · 2 June 2026
Neural network compression is commonly achieved by pruning parameters based on local importance scores, e.g., magnitude-based pruning. We propose a complementary approach that compresses models by aggregating neurons with similar functional behavior rather than removing weights independently. Our me…
- GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation
Shihao Zhang, Rayan Saab · 2 June 2026
Post-training quantization is widely used for compressing large neural networks, but aggressive low-bit quantization can significantly degrade model quality. A common remedy is to augment the quantized weights with a low-rank correction, leading to approximations of the form $W\approx Q+LR$. In this…
- How Accurately Can a Gaussian Approximate Stochastic Approximation Iterates?
Shaan Ul Haque, Zedong Wang, Zixuan Zhang, Siva Theja Maguluri · 2 June 2026
Stochastic approximation (SA) is a method for finding the root of an operator perturbed by noise. The focus of this paper is studying the distribution of SA iterates in finite time. In general, it is not possible to characterize the exact distribution, and therefore our goal is to find an approximat…
- Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded Scheduler
Dimitris Oikonomou, Nicolas Loizou · 2 June 2026
Sharpness-Aware Minimization (SAM) has established itself as a powerful and widely adopted optimizer for training machine learning models. By explicitly minimizing the sharpness of the loss landscape, SAM often improves generalization while delivering strong empirical performance. However, SAM and i…
- DAGGER: Gradient-Free Construction of Transiently Amplifying Networks under Hard Connectivity Constraints
James C. Ferguson · 2 June 2026
Many networks not only support but also rely on transient non-normal amplification, an orders-of-magnitude increase in the activity of an otherwise stable system. Constructing such networks under hard sign/sparsity/diagonal constraints -- the regime relevant for biological connectomes and structured…
