Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1,612 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Born Discrete, Made Smooth: Variational Formulation of Shallow Neural Networks
Matej Benko, Pierre Bousquet, Iwona Chlebicka, B{\l}a\.zej Miasojedow · 3 July 2026
Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-convex landscapes and stochastic heuristics. In this work, we propose a paradigm shift by replacing the discrete training problem of shallow neural net…
- Beyond the Performance Illusion: Structure-Aware Stratified Partitioning and Curriculum Distributionally Robust Optimization for Spatially Correlated Domains
Prathamesh Patil, Arpit Jain, Aswanth Krishnan · 3 July 2026
Performance evaluation in AI systems commonly assumes that random dataset splits produce independent and identically distributed (i.i.d.) subsets. We show that this assumption often breaks down in spatiotemporally correlated domains such as aerial surveillance, precision agriculture, and medical ima…
- Revisiting Decentralized Online Convex Optimization with Compressed Communication
Hao Zhou, Xiaoyu Wang, Chang Yao, Mingli Song, Yuanyu Wan · 3 July 2026
Decentralized online convex optimization (D-OCO) is a popular framework for distributed applications with streaming data. To tackle the communication bottleneck, previous studies have investigated D-OCO with compressed communication and proposed several algorithms that are variants of online gradien…
- Regularized Variational and Spectral Log-Density-Ratio Estimation in the Gaussian Location Model
Francis Bach (SIERRA) · 3 July 2026
We study ridge-regularized log-density-ratio estimation in the Gaussian location model with a common covariance matrix. By affine invariance, the model is written as q $\sim$ N(0, I), p $\sim$ N($\Delta$, I), with linear features, where $\Delta$ is a mean vector. The variational estimator is the emp…
- Gradient-Flow Optimization as Dynamic Random-Effects Inference: Testing and Early Stopping with Applications to Deep Learning
Minhao Yao, Ruoyu Wang, Xihong Lin, Lin Liu, Zhonghua Liu · 3 July 2026
Gradient-flow optimization is usually viewed as an algorithmic procedure for minimizing empirical loss, with training duration selected by validation or heuristic early stopping rules. We develop a statistical inference framework for gradient-flow training. We show that whenever fitted values evolve…
- Local exponential stability of mean-field Langevin descent-ascent and associated particle system
Geuntaek Seo, Minseop Shin, Pierre Monmarch\'e, Beomjun Choi · 3 July 2026
We study the mean-field Langevin descent-ascent (MFL-DA), a coupled optimization dynamics on the space of probability measures for entropically regularized two-player zero-sum games, together with its associated interacting particle system. For general nonconvex-nonconcave payoffs, Wang and Chizat (…
- From Approximation to Emergence: A Theory of Deep Learning
Zhilin Zhao · 3 July 2026
Deep learning has outgrown any single mathematical explanation. From Approximation to Emergence develops a unified, proof-oriented account of modern deep learning theory, tracing a path from the classical foundations of approximation, optimization, and generalization to the contemporary mechanisms o…
- A Unified Lyapunov-IQC Framework for Uniform Stability of Smooth Quadratic First-Order Accelerated Optimizers
Don Li, Dacian Daescu · 3 July 2026
We develop a unified Lyapunov-integral quadratic constraint (IQC) framework for establishing uniform stability of first-order accelerated optimization algorithms in the $\beta$-smooth and $\gamma$-strongly convex regime. Classical analyses of uniform stability, such as the work of Hardt, Recht, and …
- Decision-focused Sparse Tangent Portfolio Optimization
Haeun Jeon, Seunghoon Choi, Hyunglip Bae, Yongjae Lee, Woo Chang Kim · 2 July 2026
Sparse tangent portfolio optimization aims to learn an interpretable, low-cardinality portfolio in the tangency direction of the mean-variance frontier. However, the associated cardinality-constrained formulation is NP-hard, and standard predict-then-optimize pipelines often misalign forecasting acc…
- Function-Counting Theory for Low-Dimensional Data Structures
Konstantin H\"aberle, Helmut B\"olcskei · 2 July 2026
The success of deep learning models in classification and regression is widely attributed to the low-dimensional structure that real-world data tend to exhibit, despite their high-dimensional representation. This work attempts to provide a mathematical framework for binary classification on low-dime…
- Neural Certificate Pricing for Combinatorial Optimization Problems
Jingyi Chen, Xinyuan Zhang, Xinwu Qian · 2 July 2026
Combinatorial optimization (CO) problems are difficult because certifiable discrete structure induces exponential search. One needs to search over the set exponentially many candidates to certify optimality, however, the structural feasibility of a path, packing, or cover can be verified in polynomi…
- Homogenization of $\ell_2$-Adversarial Training in High-Dimensions: Exact Dynamics under Stochastic Gradient Descent
Fabrizzio Sabelli · 2 July 2026
We develop a framework for analyzing the learning dynamics of $\ell_2$-adversarial training of single-index models on Gaussian mixtures in the high-dimensional limit under streaming stochastic gradient descent (SGD). We derive deterministic equivalents for a broad class of statistics of the SGD iter…
- Measuring Dead Directions: Decomposing and Classifying Singular Structure off Canonical Alignment
Tejas Pradeep Shirodkar · 2 July 2026
We give a descent-free, alignment-free measurement of singular structure on trained networks. At a single frozen checkpoint the read recovers the order $k$ of each dead direction from the directional-Fisher rate, the master invariant from which the per-direction learning coefficient $1/(2k)$ follows…
- Generative Refinement for Low-Budget Black-Box Optimization
Edouard R. Dufour, Pascal Fua · 2 July 2026
Black-box optimization is a fundamental science and engineering tool that makes it possible to optimize objectives without gradient information. Unfortunately, as it often requires many function evaluations, it can be challenging when each one is costly. This is especially true when the evaluation f…
- SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates
Konstantinos Emmanouilidis, Lachlan MacDonald, Salma Tarmoun, Rene Vidal · 1 July 2026
Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory. Most prior analyses of the edge of stability phenomenon focus on deterministic gradient descent, leaving the stochastic setting la…
- Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization
Srijan Tiwari, Aditya Chauhan, Manjot Singh · 1 July 2026
Why do neural networks memorize algorithmic training data long before they generalize? We present a geometric case study demonstrating that, on tasks where generalization requires discovering structured low-dimensional circuits, the memorization-generalization delay is driven by radial inflation of …
- Mean-Field Model for Two-Layer Neural Networks Trained with Consensus-Based Optimization
William De Deyn, Michael Herty, Giovanni Samaey · 1 July 2026
We study Consensus-Based Optimization (CBO) for two-layer neural network training. We compare the performance of CBO against Adam on two test cases and demonstrate how a hybrid approach, combining CBO with Adam, provides faster convergence than CBO. Additionally, in the context of multi-task learnin…
- Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
Haoming Meng, Anton Sugolov, Vardan Papyan · 1 July 2026
Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that emerge during training. Motivated by this observation, we introduce \emph{Depth-wise Gradient Augmentation}, a general optimization paradigm in which the update ap…
- On the Convergence of Self-Improving Online LLM Alignment
Xudong Wu, Pangpang Liu, Vaneet Aggarwal, Jiayu Chen · 1 July 2026
The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lack…
- Random Reshuffling Dominates Stochastic Gradient Descent
Zijian Liu · 1 July 2026
Stochastic Gradient Descent ($\textsf{SGD}$) is one of the most classical optimization algorithms with favorable theoretical guarantees, yet the practical implementation of $\textsf{SGD}$ differs subtly from its well-known form and is often referred to as Shuffling Stochastic Gradient Descent ($\tex…
- Revisiting the Volume Hypothesis
Ari Pakman, Lior Kreimer, Yakir Berchenko · 1 July 2026
Modern deep neural networks often contain far more parameters than needed to fit their training data, yet they achieve impressive generalization. A common explanation for this success is the implicit bias of stochastic gradient descent (SGD). An alternative volume hypothesis posits that, within low …
- Curvature-Weighted Gradient Diversity: A Noise Measure for Geometry-Adaptive SGD Schedules
Muhammad Hamza (Indian Institute of Technology Kharagpur), Ayush Goel (Indian Institute of Technology Kharagpur) · 30 June 2026
The standard convergence analysis of mini-batch stochastic gradient descent (SGD) models gradient noise using a single variance term that treats all parameter directions equally, ignoring the fact that noise in high-curvature directions has less impact because learning rates are already constrained …
- ITSPACE: Monotone Gaussian Optimal Transport Updates
Woojoo Na, Jennifer Dy · 30 June 2026
Covariance matrices serve as compact descriptors of feature distributions in many machine-learning pipelines, including domain adaptation and Gaussian embeddings. Under a centered Gaussian approximation, the unregularized Wasserstein-2 optimal-transport (OT) discrepancy admits a closed form on covar…
- Why Do We Need Warm-up? A Theoretical Perspective
Foivos Alimisis, Rustem Islamov, Aurelien Lucchi · 30 June 2026
Learning rate warm-up -- increasing the learning rate at the beginning of training -- has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations remain poorly understood. In this work, we provide a principled explanation for why warm-up improves training. We rely on a…
- IG-Lens: Exact Additive Probability Attribution Across Transformer Layers via Telescoping Integrated Gradients
Duc Anh Nguyen · 30 June 2026
We ask a simple question about decoder-only transformers: \emph{between which two layers is the probability of a predicted token actually produced?} Existing layer-wise readout tools answer only approximately. The logit lens and its trained variant report a per-layer \emph{level} of probability but …
