Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1414 artículos indexados
Las técnicas de optimización por gradiente estocástico constituyen un ámbito central de la inteligencia artificial, donde se estudia cómo ajustar eficientemente los parámetros de los modelos, en particular las redes neuronales, en presencia de datos ruidosos o incompletos. Estos métodos exploran variantes del gradient descent, como la integración de momentum, la reparametrización aleatoria o condiciones de convergencia adaptadas a funciones no convexas y no suaves. Los trabajos recientes analizan también las garantías teóricas de estabilidad, los límites inferiores de rendimiento o los comportamientos de los algoritmos ante discontinuidades o estructuras geométricas complejas.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos43 % · 417 artículos
- China20 % · 189 artículos
- Reino Unido6,8 % · 65 artículos
- Alemania6,2 % · 60 artículos
- Francia5,8 % · 56 artículos
- India4,4 % · 42 artículos
- Italia4,4 % · 42 artículos
- Canadá4 % · 38 artículos
Sobre 962 artículos de este tema con al menos un laboratorio localizado. 63 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Removing spurious minima for planar features by skip connections
Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho · 2 de octubre de 2026
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for stu…
- Conflicting Supervision Moves Commitment, Not Capability: A 12.29{\sigma} arrangement effect that is exactly zero under a convention-agnostic score
Wenhui Chen · 2 de octubre de 2026
"Train a model on the same problems written under two incompatible conventions, both correct, and ask what the ordering of that data writes into the parameters. The learning-rate schedule is not a background condition for that question. It is the averaging operator, and it decides the answer. We pro…
- Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization
Yiyang Lu, Mohammad Pedramfar, Vaneet Aggarwal · 1 de octubre de 2026
We study decentralized online optimization of upper-linearizable payoffs over an action set under efficient separation access, with applications to online continuous diminishing-return (DR) submodular maximization. We propose Decentralized Barrier Follow-the-Regularized-Leader (Dec-BFTRL), and evalu…
- How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan · 1 de octubre de 2026
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase…
- Second-Moment Stochastic Approximation Methods
Tao Jiang, Lin Xiao · 30 de septiembre de 2026
Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive …
- Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
Ruinan Jin, Difei Cheng, Ling Chen, Jun Luo, Hao Zhou, Youzhi Zhang · 30 de septiembre de 2026
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whet…
- AYLA: Amplifying Gradient Sensitivity via Loss Transformation in Non-Convex Optimization
Behnam Gheshlaghi, Shahin Atakishiyev · 28 de septiembre de 2026
Stochastic Gradient Descent (SGD) and its variants, such as ADAM, are foundational to deep learning optimization, adjusting model parameters through fixed or adaptive learning rates based on loss function gradients. However, these methods often struggle to balance adaptability and efficiency in high…
- Geometric Moment Contraction for Stochastic Nesterov Acceleration
Wei Biao Wu · 28 de septiembre de 2026
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=\Theta_k+\beta(\Theta_k-\Theta_{k-1}),\qquad \Theta_{k+1}=Y_k-\gamma G(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison p…
- A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine
Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta · 28 de septiembre de 2026
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and t…
- Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds
Wei Biao Wu · 28 de septiembre de 2026
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a dir…
- Low-Rank Friction for Memory-Efficient Transformer Pretraining
Rajit Rajpal, Benedict Leimkuhler · 28 de septiembre de 2026
iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $\xi\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as A…
- Distribution of hitting times for dissipative random dynamical systems on $\mathbb{R}^d$, with application to stochastic gradient descent
St\'ephane Galatolo, St\'ephane Chr\'etien · 28 de septiembre de 2026
Machine Learning and more specifically Deep Learning involves solving large scale nonconvex optimization problems. Several algorithms have been proposed in the literature, that seem to achieve satisfactory practical efficiency for difficult instances, the Stochastic Gradient Method being the most ru…
- Generalization behavior of OPTQ and the role of regularization
Erin George, Rayan Saab · 28 de septiembre de 2026
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dat…
- Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li · 28 de septiembre de 2026
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features …
- Why Clipping Matters in AdaGrad? Toward a High-Probability Theory under Generalized Smoothness
Alokendu Mazumder, Ayaan Mohd, Harshit Rawat, Arnab Roy, Mayank Baranwal, Punit Rathore · 28 de septiembre de 2026
We analyze the original same-step coordinate-wise AdaGrad under generalized smoothness and heavy-tailed noise with bounded variance. In this setting, local curvature may grow sub-quadratically with the gradient norm, and stochastic gradients are assumed to have only bounded conditional second moment…
- When the Preconditioning Exponent Turns Negative: Learning-Rate Coupling and Cross-Environment Generalization
Gongyue Zhang, Honghai Liu · 28 de septiembre de 2026
Adaptive optimizers are commonly parameterized by a fixed power of the second-moment estimate. Existing partially adaptive methods study exponents between momentum-like updates and the standard Adam square root, while the interaction between this exponent and the global learning rate is less underst…
- Improving the Last-Iterate Guarantees of Anytime Algorithms for Stochastic Monotone Variational Inequalities
Jun-Hyun Kim, Ahmet Alacaoglu · 25 de septiembre de 2026
We analyze a stochastic algorithm with Halpern-type anchoring for constrained convex-concave problems and monotone variational inequalities. This single-loop and single-call algorithm uses one unbiased sample of the gradient operator at every iteration, to be applicable to monotone games with noisy …
- A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning
Ids van der Werf, Sergio Rozada, Antonio G. Marques · 25 de septiembre de 2026
Many iterative algorithms rely on bootstrapping. A variable is updated using a second, frozen copy as a target, which is periodically replaced with the updated variable. Majorize-minimize and inexact proximal-point methods share this structure, as does temporal-difference (TD) learning. However, exi…
- Precise Convergence Speed of Clipped SGD
David A. R. Robin · 25 de septiembre de 2026
We present a tightened convergence analysis of clipped gradient descent on $(L_0, L_1)$-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental p…
- Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections
Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazi\'c · 24 de septiembre de 2026
We study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-lay…
- Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences
Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou · 24 de septiembre de 2026
Empirical scaling laws for autoregressive language models relate prediction loss to model size, data size, and optimization compute, but their theoretical origin is still poorly understood in sequential pretraining settings. We study this question in a tractable teacher--student model where a stable…
- The Drift Contract: Spectral Updates for Depth-Robust Local Learning
Fabien Polly · 24 de septiembre de 2026
Local learning trains each layer with its own auxiliary loss and no global backward pass, which makes layer updates structurally parallel. Two problems have kept it marginal: accuracy degrades as depth grows, and hyperparameters are fragile. We apply Muon-style spectral update geometry (momentum ort…
- Sharp First-Order Lower Bounds for Higher-Order Smooth Nonconvex Optimization
Dongruo Zhou · 23 de septiembre de 2026
We study the deterministic first-order oracle complexity of finding \(\epsilon\)-stationary points in smooth nonconvex optimization when the objective satisfies higher-order smoothness assumptions. While the classical \(\epsilon^{-2}\) rate is optimal under only Lipschitz gradients, higher-order smo…
- COREM: Cosine-Relation Momentum Reshaping with Stateful Writeback
Yan Wang, Xiaochuan Wang, Yuxiang Sun · 22 de septiembre de 2026
Matrix-valued optimizer states may contain relational structure that is not captured by treating their entries independently. We study whether relations within matrix-valued optimizer states can be exploited to improve optimization. To this end, we introduce a unit-relation-transform abstraction and…
- Brownian Heads for Deep ReLU Representations: Activation Mass and the Cost of Same-Sample Selection
Mahdi Mohammadigohari, Nicole M\"ucke · 21 de septiembre de 2026
Deep representation learning often selects hidden features and fits the final predictor on the same sample, so fixed-feature analysis performed after selection can omit selection cost. We study the conditional empirical Rademacher complexity of deep ReLU representations followed by bounded-norm pred…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
