Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1.414 indexierte Paper
Die stochastischen Gradientenoptimierungstechniken bilden ein zentrales Gebiet der künstlichen Intelligenz, in dem untersucht wird, wie die Parameter von Modellen - insbesondere neuronalen Netzen - effizient angepasst werden können, wenn verrauschte oder unvollständige Daten vorliegen. Diese Methoden erforschen Varianten des gradient descent, wie die Integration von momentum, zufällige Reparametrisierung oder Konvergenzbedingungen, die an nichtkonvexe und nichtglatte Funktionen angepasst sind. Aktuelle Arbeiten analysieren zudem theoretische Stabilitätsgarantien, untere Leistungsgrenzen oder das Verhalten der Algorithmen bei Diskontinuitäten oder komplexen geometrischen Strukturen.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten43 % · 417 Artikel
- China20 % · 189 Artikel
- Vereinigtes Königreich6,8 % · 65 Artikel
- Deutschland6,2 % · 60 Artikel
- Frankreich5,8 % · 56 Artikel
- Indien4,4 % · 42 Artikel
- Italien4,4 % · 42 Artikel
- Kanada4 % · 38 Artikel
Über 962 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 63 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Broken scale symmetries in undercomplete linear autoencoders
Farhad Pashakhanloo, Jacob A. Zavatone-Veth · 5. Oktober 2026
Neural network loss landscapes have many symmetries, which are preserved by gradient flow but broken by finite-stepsize stochastic gradient descent (SGD). A canonical example of such a symmetry is scale in homogeneous networks: one can scale up the parameters in one layer and down in the next withou…
- Differential Privacy of Gradient Descent on Perturbed Objectives
Austin Watkins, Raman Arora · 5. Oktober 2026
Objective perturbation adds a random linear term to a regularized empirical risk and releases the exact perturbed minimizer. We study the finite computation obtained by releasing the $N$-th iterate of deterministic gradient descent on $w\mapsto F(w;S)+\langle z,w\rangle$, where $z\sim\mathcal N(0,\s…
- Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance
Arda Fazla, Antesh Upadhyay, Ege C. Kaya, M. Berk Sahin, Abolfazl Hashemi · 5. Oktober 2026
Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded stochastic gradient variance, yet recent evidence …
- Removing spurious minima for planar features by skip connections
Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho · 2. Oktober 2026
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for stu…
- Conflicting Supervision Moves Commitment, Not Capability: A 12.29{\sigma} arrangement effect that is exactly zero under a convention-agnostic score
Wenhui Chen · 2. Oktober 2026
"Train a model on the same problems written under two incompatible conventions, both correct, and ask what the ordering of that data writes into the parameters. The learning-rate schedule is not a background condition for that question. It is the averaging operator, and it decides the answer. We pro…
- Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization
Yiyang Lu, Mohammad Pedramfar, Vaneet Aggarwal · 1. Oktober 2026
We study decentralized online optimization of upper-linearizable payoffs over an action set under efficient separation access, with applications to online continuous diminishing-return (DR) submodular maximization. We propose Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evalu…
- How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan · 1. Oktober 2026
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase…
- Second-Moment Stochastic Approximation Methods
Tao Jiang, Lin Xiao · 30. September 2026
Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive …
- Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
Ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou, Youzhi Zhang · 30. September 2026
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whet…
- AYLA: Amplifying Gradient Sensitivity via Loss Transformation in Non-Convex Optimization
Behnam Gheshlaghi, Shahin Atakishiyev · 28. September 2026
Stochastic Gradient Descent (SGD) and its variants, such as ADAM, are foundational to deep learning optimization, adjusting model parameters through fixed or adaptive learning rates based on loss function gradients. However, these methods often struggle to balance adaptability and efficiency in high…
- Geometric Moment Contraction for Stochastic Nesterov Acceleration
Wei Biao Wu · 28. September 2026
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison p…
- A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine
Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta · 28. September 2026
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and t…
- Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds
Wei Biao Wu · 28. September 2026
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a dir…
- Low-Rank Friction for Memory-Efficient Transformer Pretraining
Rajit Rajpal, Benedict Leimkuhler · 28. September 2026
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as A…
- Distribution of hitting times for dissipative random dynamical systems on $\mathbb{R}^d$, with application to stochastic gradient descent
St\'ephane Galatolo, St\'ephane Chr\'etien · 28. September 2026
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most ru…
- Generalization behavior of OPTQ and the role of regularization
Erin George, Rayan Saab · 28. September 2026
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dat…
- Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li · 28. September 2026
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features …
- Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness
Alokendu Mazumder, Ayaan Mohd, Harshit Rawat, Arnab Roy, Mayank Baranwal, Punit Rathore · 28. September 2026
We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moment…
- When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
Gongyue Zhang, Honghai Liu · 28. September 2026
Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and the global learning rate is less underst…
- Improving the Last-Iterate Guarantees of Anytime Algorithms for Stochastic Monotone Variational Inequalities
Jun-Hyun Kim, Ahmet Alacaoglu · 25. September 2026
We analyze a stochastic algorithm with Halpern-type anchoring for constrained convex-concave problems and monotone variational inequalities. This single-loop and single-call algorithm uses one unbiased sample of the gradient operator at every iteration, to be applicable to monotone games with noisy …
- A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning
Ids van der Werf, Sergio Rozada, Antonio G. Marques · 25. September 2026
Many iterative algorithms rely on bootstrapping. A variable is updated using a second, frozen copy as a target, which is periodically replaced with the updated variable. Majorize-minimize and inexact proximal-point methods share this structure, as does temporal-difference (TD) learning. However, exi…
- Precise Convergence Speed of Clipped SGD
David A. R. Robin · 25. September 2026
We present a tightened convergence analysis of clipped gradient descent on $(L_0, L_1)$-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental p…
- Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections
Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c · 24. September 2026
We study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-lay…
- Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences
Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou · 24. September 2026
Empirical scaling laws for autoregressive language models relate prediction loss to model size, data size, and optimization compute, but their theoretical origin is still poorly understood in sequential pretraining settings. We study this question in a tractable teacher--student model where a stable…
- The Drift Contract: Spectral Updates for Depth-Robust Local Learning
Fabien Polly · 24. September 2026
Local learning trains each layer with its own auxiliary loss and no global backward pass, which makes layer updates structurally parallel. Two problems have kept it marginal: accuracy degrades as depth grows, and hyperparameters are fragile. We apply Muon-style spectral update geometry (momentum ort…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
