Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1.612 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- ZeroS: Zero-Sum Linear Attention for Efficient Transformers
Jiecheng Lu, Xu Han, Yan Sun, Viresh Pati, Yubin Kim, Siddhartha Somani, Shihao Yang · 6. Februar 2026
Linear attention methods offer Transformers $O(N)$ complexity but typically underperform standard softmax attention. We identify two fundamental limitations affecting these approaches: the restriction to convex combinations that only permits additive information blending, and uniform accumulated wei…
- Finite-Particle Rates for Regularized Stein Variational Gradient Descent
Ye He, Krishnakumar Balasubramanian, Sayan Banerjee, Promit Ghosal · 6. Februar 2026
We derive finite-particle rates for the regularized Stein variational gradient descent (R-SVGD) algorithm introduced by He et al. (2024) that corrects the constant-order bias of the SVGD by applying a resolvent-type preconditioner to the kernelized Wasserstein gradient. For the resulting interacting…
- Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
Artem Riabinin, Andrey Veprikov, Arman Bolatov, Martin Tak\'a\v{c}, Aleksandr Beznosikov · 6. Februar 2026
We study adaptive learning rate scheduling for norm-constrained optimizers (e.g., Muon and Lion). We introduce a generalized smoothness assumption under which local curvature decreases with the suboptimality gap and empirically verify that this behavior holds along optimization trajectories. Under t…
- A Sketch-and-Project Analysis of Subsampled Natural Gradient Algorithms
Gil Goldshlager, Jiang Hu, Lin Lin · 6. Februar 2026
Subsampled natural gradient descent (SNG) has been used to enable high-precision scientific machine learning, but standard analyses based on stochastic preconditioning fail to provide insight into realistic small-sample settings. We overcome this limitation by instead analyzing SNG as a sketch-and-p…
- Sharpness-Aware Minimization Can Hallucinate Minimizers
Chanwoong Park, Uijeong Jang, Ernest K. Ryu, Insoon Yang · 6. Februar 2026
Sharpness-Aware Minimization (SAM) is widely used to seek flatter minima -- often linked to better generalization. In its standard implementation, SAM updates the current iterate using the loss gradient evaluated at a point perturbed by distance $\rho$ along the normalized gradient direction. We sho…
- Solving Stochastic Variational Inequalities without the Bounded Variance Assumption
Ahmet Alacaoglu, Jun-Hyun Kim · 6. Februar 2026
We analyze algorithms for solving stochastic variational inequalities (VI) without the bounded variance or bounded domain assumptions, where our main focus is min-max optimization with possibly unbounded constraint sets. We focus on two classes of problems: monotone VIs; and structured nonmonotone V…
- Decoupled Orthogonal Dynamics: Regularization for Deep Network Optimizers
Hao Chen, Jinghui Yuan, Hanmin Zhang · 6. Februar 2026
Is the standard weight decay in AdamW truly optimal? Although AdamW decouples weight decay from adaptive gradient scaling, a fundamental conflict remains: the Radial Tug-of-War. In deep learning, gradients tend to increase parameter norms to expand effective capacity while steering directions to lea…
- Limitations of SGD for Multi-Index Models Beyond Statistical Queries
Daniel Barzilai, Ohad Shamir · 6. Februar 2026
Understanding the limitations of gradient methods, and stochastic gradient descent (SGD) in particular, is a central challenge in learning theory. To that end, a commonly used tool is the Statistical Queries (SQ) framework, which studies performance limits of algorithms based on noisy interaction wi…
- Convergence Rate of the Last Iterate of Stochastic Proximal Algorithms
Kevin Kurian Thomas Vaidyan, Michael P. Friedlander, Ahmet Alacaoglu · 6. Februar 2026
We analyze two classical algorithms for solving additively composite convex optimization problems where the objective is the sum of a smooth term and a nonsmooth regularizer: proximal stochastic gradient method for a single regularizer; and the randomized incremental proximal method, which uses the …
- Optimal scaling laws in learning hierarchical multi-index models
Leonardo Defilippis, Florent Krzakala, Bruno Loureiro, Antoine Maillard · 6. Februar 2026
In this work, we provide a sharp theory of scaling laws for two-layer neural networks trained on a class of hierarchical multi-index targets, in a genuinely representation-limited regime. We derive exact information-theoretic scaling laws for subspace recovery and prediction error, revealing how the…
- A Short and Unified Convergence Analysis of the SAG, SAGA, and IAG Algorithms
Feng Zhu, Robert W. Heath Jr., Aritra Mitra · 6. Februar 2026
Stochastic variance-reduced algorithms such as Stochastic Average Gradient (SAG) and SAGA, and their deterministic counterparts like the Incremental Aggregated Gradient (IAG) method, have been extensively studied in large-scale machine learning. Despite their popularity, existing analyses for these …
- Shiva-DiT: Residual-Based Differentiable Top-$k$ Selection for Efficient Diffusion Transformers
Jiaji Zhang, Hailiang Zhao, Guoxuan Zhu, Ruichao Sun, Jiaju Wu, Xinkui Zhao, Hanlin Tang, Weiyi Lu, Kan Liu, Tao Lan, Lin Qu, Shuiguang Deng · 6. Februar 2026
Diffusion Transformers (DiTs) incur prohibitive computational costs due to the quadratic scaling of self-attention. Existing pruning methods fail to simultaneously satisfy differentiability, efficiency, and the strict static budgets required for hardware overhead. To address this, we propose Shiva-D…
- SLAY: Geometry-Aware Spherical Linearized Attention with Yat-Kernel
Jose Miguel Luna, Taha Bouhsine, Krzysztof Choromanski · 6. Februar 2026
We propose a new class of linear-time attention mechanisms based on a relaxed and computationally efficient formulation of the recently introduced E-Product, often referred to as the Yat-kernel (Bouhsine, 2025). The resulting interactions are geometry-aware and inspired by inverse-square interaction…
- Unbiased Single-Queried Gradient for Combinatorial Objective
Thanawat Sornwanee · 6. Februar 2026
In a probabilistic reformulation of a combinatorial problem, we often face an optimization over a hypercube, which corresponds to the Bernoulli probability parameter for each binary variable in the primal problem. The combinatorial nature suggests that an exact gradient computation requires multiple…
- Dual Perspectives on Non-Contrastive Self-Supervised Learning
Jean Ponce (ENS-PSL, NYU), Basile Terver (FAIR, WILLOW), Martial Hebert (CMU), Michael Arbel (Thoth) · 6. Februar 2026
The {\em stop gradient} and {\em exponential moving average} iterative procedures are commonly used in non-contrastive approaches to self-supervised learning to avoid representation collapse, with excellent performance in downstream applications in practice. This presentation investigates these proc…
- Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
Yizhou Xu, Pierfrancesco Beneventano, Isaac Chuang, Liu Ziyin · 6. Februar 2026
A large body of theory and empirical work hypothesizes a connection between the flatness of a neural network's loss landscape during training and its performance. However, there have been conceptually opposite pieces of evidence regarding when SGD prefers flatter or sharper solutions during training…
- On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature
Yikuan Zhang, Ning Yang, Yuhai Tu · 6. Februar 2026
Stochastic Gradient Descent (SGD) introduces anisotropic noise that is correlated with the local curvature of the loss landscape, thereby biasing optimization toward flat minima. Prior work often assumes an equivalence between the Fisher Information Matrix and the Hessian for negative log-likelihood…
- Tight Long-Term Tail Decay of (Clipped) SGD in Non-Convex Optimization
Aleksandar Armacki, Dragana Bajovi\'c, Du\v{s}an Jakoveti\'c, Soummya Kar, Ali H. Sayed · 6. Februar 2026
The study of tail behaviour of SGD-induced processes has been attracting a lot of interest, due to offering strong guarantees with respect to individual runs of an algorithm. While many works provide high-probability guarantees, quantifying the error rate for a fixed probability threshold, there is …
- Radon--Wasserstein Gradient Flows for Interacting-Particle Sampling in High Dimensions
Elias Hess-Childs, Dejan Slep\v{c}ev, Lantian Xu · 6. Februar 2026
Gradient flows of the Kullback--Leibler (KL) divergence, such as the Fokker--Planck equation and Stein Variational Gradient Descent, evolve a distribution toward a target density known only up to a normalizing constant. We introduce new gradient flows of the KL divergence with a remarkable combinati…
- Rate-Optimal Noise Annealing in Semi-Dual Neural Optimal Transport: Tangential Identifiability, Off-Manifold Ambiguity, and Guaranteed Recovery
Raymond Chu, Jaewoong Choi, Dohyun Kwon · 5. Februar 2026
Semi-dual neural optimal transport learns a transport map via a max-min objective, yet training can converge to incorrect or degenerate maps. We fully characterize these spurious solutions in the common regime where data concentrate on low-dimensional manifold: the objective is underconstrained off …
- Geometry-Aware Optimal Transport: Fast Intrinsic Dimension and Wasserstein Distance Estimation
Ferdinand Genans (SU, LPSM), Olivier Wintenberger (SU, LPSM) · 5. Februar 2026
Solving large scale Optimal Transport (OT) in machine learning typically relies on sampling measures to obtain a tractable discrete problem. While the discrete solver's accuracy is controllable, the rate of convergence of the discretization error is governed by the intrinsic dimension of our data. T…
- It's not a Lottery, it's a Race: Understanding How Gradient Descent Adapts the Network's Capacity to the Task
Hannah Pinson · 5. Februar 2026
Our theoretical understanding of neural networks is lagging behind their empirical success. One of the important unexplained phenomena is why and how, during the process of training with gradient descent, the theoretical capacity of neural networks is reduced to an effective capacity that fits the t…
- Principles of Lipschitz continuity in neural networks
R\'ois\'in Luo · 5. Februar 2026
Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable in artificial intelligence. Yet, despite these advances, critical challenges remain -- most notably, ensuring robustness to small input perturbations and generali…
- LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
Andrej Jovanovi\'c, Alex Iacob, Mher Safaryan, Ionut-Vlad Modoranu, Lorenzo Sani, William F. Shen, Xinchi Qiu, Dan Alistarh, Nicholas D. Lane · 5. Februar 2026
Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth. While infrequent communication strategies reduce synchronization frequency, they remain bottlenecked by the memory and communication requirements of optimizer states. Low-rank optimizers can alleviate …
- Supervised Learning as Lossy Compression: Characterizing Generalization and Sample Complexity via Finite Blocklength Analysis
Kosuke Sugiyama, Masato Uchida · 5. Februar 2026
This paper presents a novel information-theoretic perspective on generalization in machine learning by framing the learning problem within the context of lossy compression and applying finite blocklength analysis. In our approach, the sampling of training data formally corresponds to an encoding pro…
