Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1 414 papiers indexés
Les techniques d’optimisation par gradient stochastique constituent un domaine central de l’intelligence artificielle, où l’on étudie comment ajuster efficacement les paramètres des modèles, notamment les réseaux de neurones, en présence de données bruitées ou incomplètes. Ces méthodes explorent des variantes du gradient descent, comme l’intégration de momentum, la reparamétrisation aléatoire ou des conditions de convergence adaptées à des fonctions non convexes et non lisses. Les travaux récents analysent aussi les garanties théoriques de stabilité, les bornes inférieures de performance ou les comportements des algorithmes face à des discontinuités ou des structures géométriques complexes.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis43 % · 417 articles
- Chine20 % · 189 articles
- Royaume-Uni6,8 % · 65 articles
- Allemagne6,2 % · 60 articles
- France5,8 % · 56 articles
- Inde4,4 % · 42 articles
- Italie4,4 % · 42 articles
- Canada4 % · 38 articles
Sur 962 articles de ce sujet dont au moins un laboratoire est situé. 63 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- Broken scale symmetries in undercomplete linear autoencoders
Farhad Pashakhanloo, Jacob A. Zavatone-Veth · 5 octobre 2026
Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next withou…
- Differential Privacy of Gradient Descent on Perturbed Objectives
Austin Watkins, Raman Arora · 5 octobre 2026
Objective perturbation adds a random linear term to a regularized empirical risk and releases the exact perturbed minimizer. We study the finite computation obtained by releasing the $N$-th iterate of deterministic gradient descent on $w\mapsto F(w;S)+\langle z,w\rangle$, where $z\sim\mathcal N(0,\s…
- Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance
Arda Fazla, Antesh Upadhyay, Ege C. Kaya, M. Berk Sahin, Abolfazl Hashemi · 5 octobre 2026
Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded stochastic gradient variance, yet recent evidence …
- Removing spurious minima for planar features by skip connections
Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho · 2 octobre 2026
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for stu…
- Conflicting Supervision Moves Commitment, Not Capability: A 12.29{\sigma} arrangement effect that is exactly zero under a convention-agnostic score
Wenhui Chen · 2 octobre 2026
"Train a model on the same problems written under two incompatible conventions, both correct, and ask what the ordering of that data writes into the parameters. The learning-rate schedule is not a background condition for that question. It is the averaging operator, and it decides the answer. We pro…
- Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization
Yiyang Lu, Mohammad Pedramfar, Vaneet Aggarwal · 1 octobre 2026
We study decentralized online optimization of upper-linearizable payoffs over an action set under efficient separation access, with applications to online continuous diminishing-return (DR) submodular maximization. We propose Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evalu…
- How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan · 1 octobre 2026
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase…
- Second-Moment Stochastic Approximation Methods
Tao Jiang, Lin Xiao · 30 septembre 2026
Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive …
- Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
Ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou, Youzhi Zhang · 30 septembre 2026
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whet…
- AYLA: Amplifying Gradient Sensitivity via Loss Transformation in Non-Convex Optimization
Behnam Gheshlaghi, Shahin Atakishiyev · 28 septembre 2026
Stochastic Gradient Descent (SGD) and its variants, such as ADAM, are foundational to deep learning optimization, adjusting model parameters through fixed or adaptive learning rates based on loss function gradients. However, these methods often struggle to balance adaptability and efficiency in high…
- Geometric Moment Contraction for Stochastic Nesterov Acceleration
Wei Biao Wu · 28 septembre 2026
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison p…
- A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine
Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta · 28 septembre 2026
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and t…
- Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds
Wei Biao Wu · 28 septembre 2026
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a dir…
- Low-Rank Friction for Memory-Efficient Transformer Pretraining
Rajit Rajpal, Benedict Leimkuhler · 28 septembre 2026
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as A…
- Distribution of hitting times for dissipative random dynamical systems on $\mathbb{R}^d$, with application to stochastic gradient descent
St\'ephane Galatolo, St\'ephane Chr\'etien · 28 septembre 2026
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most ru…
- Generalization behavior of OPTQ and the role of regularization
Erin George, Rayan Saab · 28 septembre 2026
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dat…
- Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li · 28 septembre 2026
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features …
- Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness
Alokendu Mazumder, Ayaan Mohd, Harshit Rawat, Arnab Roy, Mayank Baranwal, Punit Rathore · 28 septembre 2026
We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moment…
- When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
Gongyue Zhang, Honghai Liu · 28 septembre 2026
Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and the global learning rate is less underst…
- Improving the Last-Iterate Guarantees of Anytime Algorithms for Stochastic Monotone Variational Inequalities
Jun-Hyun Kim, Ahmet Alacaoglu · 25 septembre 2026
We analyze a stochastic algorithm with Halpern-type anchoring for constrained convex-concave problems and monotone variational inequalities. This single-loop and single-call algorithm uses one unbiased sample of the gradient operator at every iteration, to be applicable to monotone games with noisy …
- A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning
Ids van der Werf, Sergio Rozada, Antonio G. Marques · 25 septembre 2026
Many iterative algorithms rely on bootstrapping. A variable is updated using a second, frozen copy as a target, which is periodically replaced with the updated variable. Majorize-minimize and inexact proximal-point methods share this structure, as does temporal-difference (TD) learning. However, exi…
- Precise Convergence Speed of Clipped SGD
David A. R. Robin · 25 septembre 2026
We present a tightened convergence analysis of clipped gradient descent on $(L_0, L_1)$-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental p…
- Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections
Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c · 24 septembre 2026
We study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-lay…
- Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences
Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou · 24 septembre 2026
Empirical scaling laws for autoregressive language models relate prediction loss to model size, data size, and optimization compute, but their theoretical origin is still poorly understood in sequential pretraining settings. We study this question in a tractable teacher--student model where a stable…
- The Drift Contract: Spectral Updates for Depth-Robust Local Learning
Fabien Polly · 24 septembre 2026
Local learning trains each layer with its own auxiliary loss and no global backward pass, which makes layer updates structurally parallel. Two problems have kept it marginal: accuracy degrades as depth grows, and hyperparameters are fragile. We apply Muon-style spectral update geometry (momentum ort…
Autres sujets du thème Intelligence artificielle
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Large Language Models7 407 papiers / 12 mois+247 %
- Adversarial Robustness in Machine Learning3 552 papiers / 12 mois+118 %
- Reinforcement Learning in Robotics2 519 papiers / 12 mois+117 %
- Explainable Artificial Intelligence (XAI)2 319 papiers / 12 mois+200 %
- Domain Adaptation and Few-Shot Learning2 059 papiers / 12 mois+67 %
- Advanced Graph Neural Networks1 926 papiers / 12 mois+38 %
