Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1.612 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Information Hidden in Gradients of Regression with Target Noise
Arash Jamshidi, Katsiaryna Haitsiukevich, Kai Puolam\"aki · 8. April 2026
Second-order information -- such as curvature or data covariance -- is critical for optimisation, diagnostics, and robustness. However, in many modern settings, only the gradients are observable. We show that the gradients alone can reveal the Hessian, equalling the data covariance $\Sigma$ for the …
- Expectation Maximization (EM) Converges for General Agnostic Mixtures
Avishek Ghosh · 8. April 2026
Mixture of linear regression is well studied in statistics and machine learning, where the data points are generated probabilistically using $k$ linear models. Algorithms like Expectation Maximization (EM) may be used to recover the ground truth regressors for this problem. Recently, in \cite{pal202…
- Training Without Orthogonalization, Inference With SVD: A Gradient Analysis of Rotation Representations
Chris Choy · 8. April 2026
Recent work has shown that removing orthogonalization during training and applying it only at inference improves rotation estimation in deep learning, with empirical evidence favoring 9D representations with SVD projection. However, the theoretical understanding of why SVD orthogonalization specific…
- Primal-Dual Methods for Nonsmooth Nonconvex Optimization with Orthogonality Constraints
Linglingzhi Zhu, Wentao Ding, Shangyuan Liu, Anthony Man-Cho So · 7. April 2026
Recent advancements in data science have significantly elevated the importance of orthogonally constrained optimization problems. The Riemannian approach has become a popular technique for addressing these problems due to the advantageous computational and analytical properties of the Stiefel manifo…
- Preservation Is Not Enough for Width Growth: Regime-Sensitive Selection of Dense LM Warm Starts
Eren Unlu · 7. April 2026
Width expansion offers a practical route to reuse smaller causal-language-model checkpoints, but selecting a widened warm start is not solved by zero-step preservation alone. We study dense width growth as a candidate-selection problem over full training states, including copied weights, optimizer m…
- Muon Dynamics as a Spectral Wasserstein Flow
Gabriel Peyr\'e · 7. April 2026
Gradient normalization is central in deep-learning optimization because it stabilizes training and reduces sensitivity to scale. For deep architectures, parameters are naturally grouped into matrices or blocks, so spectral normalizations are often more faithful than coordinatewise Euclidean ones; Mu…
- The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks
El Mehdi Achour, Kathl\'en Kohn, Holger Rauhut · 7. April 2026
We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space can be written as a Riemannian gradient flow on function space (i.e., on the pro…
- Neural Global Optimization via Iterative Refinement from Noisy Samples
Qusay Muzaffar, David Levin, Michael Werman · 7. April 2026
Global optimization of black-box functions from noisy samples is a fundamental challenge in machine learning and scientific computing. Traditional methods such as Bayesian Optimization often converge to local minima on multi-modal functions, while gradient-free methods require many function evaluati…
- The Generalised Kernel Covariance Measure
Luca Bergen, Dino Sejdinovic, Vanessa Didelez · 7. April 2026
We consider the problem of conditional independence (CI) testing and adopt a kernel-based approach. Kernel-based CI tests embed variables in reproducing kernel Hilbert spaces, regress their embeddings on the conditioning variables, and test the resulting residuals for marginal independence. This app…
- Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective
Nicola Aladrah, Emanuele Ballarin, Matteo Biagetti, Alessio Ansuini, Alberto d'Onofrio, Fabio Anselmi · 7. April 2026
A key challenge in machine learning is to explain how learning dynamics select among the many solutions that achieve identical loss values in overparameterized models - a phenomenon known as implicit bias. Controlling this bias provides a direct mechanism on learned representations, which are centra…
- Geometric Limits of Knowledge Distillation: A Minimum-Width Theorem via Superposition Theory
Dawar Jyoti Deka, Nilesh Sarkar · 7. April 2026
Knowledge distillation compresses large teachers into smaller students, but performance saturates at a loss floor that persists across training methods and objectives. We argue this floor is geometric: neural networks represent far more features than dimensions through superposition, and a student o…
- Adaptive Threshold-Driven Continuous Greedy Method for Scalable Submodular Optimization
Mohammadreza Rostami, Solmaz S. Kia · 7. April 2026
Submodular maximization under matroid constraints is a fundamental problem in combinatorial optimization with applications in sensing, data summarization, active learning, and resource allocation. While the Sequential Greedy (SG) algorithm achieves only a $\frac{1}{2}$-approximation due to irrevocab…
- An Improved Last-Iterate Convergence Rate for Anchored Gradient Descent Ascent
Anja Surina, Arun Suggala, George Tsoukalas, Anton Kovsharov, Sergey Shirobokov, Francisco J. R. Ruiz, Pushmeet Kohli, Swarat Chaudhuri · 7. April 2026
We analyze the last-iterate convergence of the Anchored Gradient Descent Ascent algorithm for smooth convex-concave min-max problems. While previous work established a last-iterate rate of $\mathcal{O}(1/t^{2-2p})$ for the squared gradient norm, where $p \in (1/2, 1)$, it remained an open problem wh…
- Grokking as Dimensional Phase Transition in Neural Networks
Ping Wang · 7. April 2026
Neural network grokking -- the abrupt memorization-to-generalization transition -- challenges our understanding of learning dynamics. Through finite-size scaling of gradient avalanche dynamics across eight model scales, we find that grokking is a \textit{dimensional phase transition}: effective dime…
- On the Efficiency of Sinkhorn-Knopp for Entropically Regularized Optimal Transport
Kun He · 7. April 2026
The Sinkhorn--Knopp (SK) algorithm is a cornerstone method for matrix scaling and entropically regularized optimal transport (EOT). Despite its empirical efficiency, existing theoretical guarantees to achieve a target marginal accuracy $\varepsilon$ deteriorate severely in the presence of outliers, …
- Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization
Chiheb Yaakoubi, Cosme Louart, Malik Tiomoko, Zhenyu Liao · 6. April 2026
We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min-Max Theorem (CGMT) to non-Gaussian settings, we derive an asymptotic min-max characterization of key statistics, enabling approximation of th…
- Optimal Projection-Free Adaptive SGD for Matrix Optimization
Dmitry Kovalev · 6. April 2026
Recently, Jiang et al. [2026] developed Leon, a practical variant of One-sided Shampoo [Xie et al., 2025a, An et al., 2025] algorithm for online convex optimization, which does not require computing a costly quadratic projection at each iteration. Unfortunately, according to the existing analysis, L…
- Low-Rank Compression of Pretrained Models via Randomized Subspace Iteration
Farhad Pourkamali-Anaraki · 6. April 2026
The massive scale of pretrained models has made efficient compression essential for practical deployment. Low-rank decomposition based on the singular value decomposition (SVD) provides a principled approach for model reduction, but its exact computation is expensive for large weight matrices. Rando…
- Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability
Eric Gan · 6. April 2026
Empirically, modern deep learning training often occurs at the Edge of Stability (EoS), where the sharpness of the loss exceeds the threshold below which classical convergence analysis applies. Despite recent progress, existing theoretical explanations of EoS either rely on restrictive assumptions o…
- Scalable Mean-Variance Portfolio Optimization via Subspace Embeddings and GPU-Friendly Nesterov-Accelerated Projected Gradient
Yi-Shuai Niu, Yajuan Wang · 6. April 2026
We develop a sketch-based factor reduction and a Nesterov-accelerated projected gradient algorithm (NPGA) with GPU acceleration, yielding a doubly accelerated solver for large-scale constrained mean-variance portfolio optimization. Starting from the sample covariance factor $L$, the method combines …
- Adaptive randomized pivoting and volume sampling
Ethan N. Epperly · 6. April 2026
Adaptive randomized pivoting (ARP) is a recently proposed and highly effective algorithm for column subset selection. This paper reinterprets the ARP algorithm by drawing connections to the volume sampling distribution and active learning algorithms for linear regression. As consequences, this paper…
- Dynamical structure of vanishing gradient and overfitting in multi-layer perceptrons
Alex Al\`i Maleknia, Yuzuru Sato · 6. April 2026
Vanishing gradient and overfitting are two of the most extensively studied problems in the literature about machine learning. However, they are frequently considered in some asymptotic setting, which obscure the underlying dynamical mechanisms responsible for their emergence. In this paper, we aim t…
- Inversion-Free Natural Gradient Descent on Riemannian Manifolds
Dario Draca, Takuo Matsubara, Minh-Ngoc Tran · 6. April 2026
The natural gradient method is widely used in statistical optimization, but its standard formulation assumes a Euclidean parameter space. This paper proposes an inversion-free stochastic natural gradient method for probability distributions whose parameters lie on a Riemannian manifold. The manifold…
- CeRA: Overcoming the Linear Ceiling of Low-Rank Adaptation via Capacity Expansion
Hung-Hsuan Chen · 6. April 2026
Low-Rank Adaptation (LoRA) dominates parameter-efficient fine-tuning (PEFT). However, it faces a ``linear ceiling'': increasing the rank yields diminishing returns in expressive capacity due to intrinsic linear constraints. We introduce CeRA (Capacity-enhanced Rank Adaptation), a weight-level parall…
- The Geometric Anatomy of Capability Acquisition in Transformers
Jayadev Billa · 3. April 2026
Neural networks gain capabilities during training, but the internal changes that precede capability acquisition are not well understood. In particular, the relationship between geometric change and behavioral change, and the effect of task difficulty and model scale on that relationship, is unclear.…
