Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1.612 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Mathematics of Data Science
Afonso S. Bandeira, Amit Singer, Thomas Strohmer · 15. Juli 2026
This book is about the mathematical foundations of data science. 1. Introduction 2. Curses, Blessings, and Surprises in High Dimensions 3. Singular Value Decomposition and Principal Component Analysis 4. Linear Regression and Regularization 5. Graphs, Networks, and Clustering 6. Nonlinea…
- Same Loss, Same Noise, Opposite Schedules: Noise Structure and Optimizer Normalization Jointly Determine Whether Learning-Rate Cooldown Helps
Subham Singh, Ashutosh Mishra, Subha Raut · 15. Juli 2026
The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others. We give a provable account of which case obtains, and it turns on two properties together: the structure of …
- Continuous Policy and Value Iteration for Stochastic Control Problems and Its Convergence
Qi Feng, Gu Wang · 15. Juli 2026
We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework applies to both the entropy-regularized relaxed control problems…
- Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization
Jiajie Zhao, Jianxing Wang, Junjie Yang, Zhiwei Bai, Yaoyu Zhang · 15. Juli 2026
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as …
- Memory Savings at What Cost? A Study of Alternatives to Backpropagation
Kunjal Panchal, Sunav Choudhary, Yuriy Brun, Hui Guan · 14. Juli 2026
Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting …
- M+Adam: Low-Precision Training via Additive-Multiplicative Optimization
Xiaoyuan Liang, Sebastian Loeschcke, Mads Toftrup, Anima Anandkumar · 14. Juli 2026
Training with quantized weights can reduce costs but often results in degraded accuracy, especially when optimization is carried out in low precision, without storing high-precision copies. We identify a key failure mode: under low precision, standard optimizers can get stuck and not make progress, …
- RSLoRA: Training-free Rank Allocation for LoRA via Representational Sensitivity Probing
Jiaqi Liu, Haidong Kang, Qihui Zhao, Guo Yu · 14. Juli 2026
Low-Rank Adaptation (LoRA) has become a cornerstone of parameter-efficient fine-tuning (PEFT); however, the conventional practice of uniform rank assignment ignores the functional heterogeneity of neural layers. Existing rank allocation methods typically struggle with a trade-off between computation…
- Exact Dynamics of Multi-class Stochastic Gradient Descent
Elizabeth Collins-Woodfin, Inbar Seroussi · 14. Juli 2026
We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes. Our main theorem provides exact expressions for quantities of interest, including the risk and the overlap wit…
- Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization
Ethan Smith · 14. Juli 2026
Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps. Adaptive optimizers such as Adam normalize updates per coordinate, but update steps remain addit…
- How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm
Igor Itkin · 14. Juli 2026
A cheap swarm of unreliable agents can be steered to a correct consensus by a few strong, expensive "oracle" correctors. We ask how much one must spend, and where to place the oracles. We model the swarm as a consensus on a graph in which each oracle pins one node toward the truth at a cost-coupled,…
- Sharper Analysis of Single-Loop Methods for Bilevel Optimization
Yubo Zhou, Jun Shu, Luo Luo, Junmin Liu, Deyu Meng, Guang Dai, Haishan Ye · 14. Juli 2026
Bilevel optimization underpins many machine learning applications, including hyperparameter optimization, meta-learning, neural architecture search, and reinforcement learning. While hypergradient-based methods have advanced significantly, a gap persists between theoretical guarantees and practical …
- Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks
Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes · 14. Juli 2026
We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised benchmark coreset selection (our approach), the selection algorithm uses…
- Training Diagonal Linear Networks with Stochastic Sharpness-Aware Minimization
Gabriel Clara, Sophie Langer, Johannes Schmidt-Hieber · 14. Juli 2026
We analyze the landscape and training dynamics of diagonal linear networks in a linear regression task, with the network parameters being perturbed by isotropic normal noise during training. The addition of such noise may be interpreted as a stochastic form of sharpness-aware minimization (SAM) and …
- Spectrally Deconfounded Gradient Boosting
Andrea Nava, Peter B\"uhlmann, Fabio Sigrist · 13. Juli 2026
Flexible machine-learning methods can be sensitive to hidden confounding: they may learn associations induced by unobserved confounders rather than stable signals. Spectral deconfounding mitigates this problem by shrinking high-variance directions of the covariate matrix that, under dense confoundin…
- Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
Qianli Liu, Kaibin Guo, Zicong Hong, Peng Li, Fahao Chen, Haodong Wang, Jian Lin, Song Guo · 13. Juli 2026
Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs. Existing works for optimizing expert placement focus on leverag…
- Nonconvex Composite Functional Constraints via First-Order Augmented Lagrangian Methods under Local Regularity
Linglingzhi Zhu, Jiajin Li · 13. Juli 2026
We study nonasymptotic convergence of primal-dual methods for a class of nonconvex constrained optimization problems with a convex-composite structure. In this class, both the objective and the functional inequality constraints are given by convex Lipschitz outer functions composed with smooth nonli…
- Solving Stochastic Fixed-Point Equations with High Probability
Jelena Diakonikolas · 13. Juli 2026
We study stochastic fixed-point equations $\mathbf{T}(\mathbf{x}) = \mathbf{x}$ over normed spaces $(\mathcal{E}, \|\cdot\|)$, where the operator $\mathbf{T}$ is nonexpansive or contractive and is accessed only through unbiased stochastic evaluations with bounded second central moment. Given $\epsil…
- Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles
Jiseok Chae, Donghwan Kim · 13. Juli 2026
Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimiza…
- Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara, Andrew Saxe · 10. Juli 2026
In this short note we consider the gradient descent dynamics of deep scalar linear networks, $f(x) = \prod_{l=1}^L w_l x$, which enjoy exact time-course solutions for any integer depth. We show that even in this minimal model, the optimal depth-wise learning rate scaling depends on data, whereas dat…
- Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization
Ryusei Yamada, Naoki Sato, Hideaki Iiduka · 10. Juli 2026
Stochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla SGD, particularly wi…
- Distributionally Faithful Imputation via Positive Semi-Definite Kernel Density Estimation
Andrea Basteri, Carlo Ciliberto, Alessandro Rudi · 10. Juli 2026
Missing values undermine statistical inference and machine learning pipelines, yet most imputation methods rely on heuristics or restrictive parametric assumptions that ignore the joint data distribution. We recast imputation under missing completely at random (MCAR) as density estimation from maske…
- Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability in Diffusion Sampling
Yiwei Zhou · 10. Juli 2026
Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily …
- Tubular Neighbourhoods of Pfaffian Sets and Applications to Neural Networks
Paul Lezeau, Martin Lotz · 10. Juli 2026
We derive bounds for the volume of tubular neighbourhoods of smooth Pfaffian hypersurfaces, generalising known results for algebraic varieties. The bounds are given in terms of the Pfaffian format of the defining functions. As an application, we obtain tail bounds on the probability distribution of …
- Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima
Lachlan Ewen MacDonald, Ren\'e Vidal · 10. Juli 2026
An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently vi…
- A law of robustness for two-layer neural networks with arbitrary weights
Yitzchak Shmalo · 10. Juli 2026
Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with $m$ neurons that fits $n$ noisy labels must have Lipschitz constant at least of order $\sqrt{n/m}$, with no restriction on the size of the weights. Bubeck and Sellke proved a universal version of this law fo…
