Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1 612 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Thinned Mean Field Langevin Dynamics
Zonghao Chen, Heishiro Kanagawa, Fran\c{c}ois-Xavier Briol, Chris J. Oates, Lester Mackey · 28 mai 2026
Several important learning tasks can be formulated as minimizing an entropy-regularized objective over an appropriate space of probability distributions. Mean-field Langevin dynamics (MFLD) facilitate computation in this general context, casting the minimizer as the invariant distribution of a McKea…
- Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization
Kristi Topollai, Allan Ma, Tolga Dimlioglu, Sui Jiet Tay, Anna Choromanska · 28 mai 2026
Communication-efficient distributed optimizers such as DiLoCo reduce synchronization costs by letting workers perform many local updates before aggregating their progress with an outer momentum optimizer. Recent theory suggests that the outer optimizer acts on an effective spectrum induced by the in…
- Efficient Pre-Training of LLMs through Truncated SVD Layers
Kaivan Kamali, Kajetan Schweighofer, Hormoz Shahrzad, Olivier Francon, Babak Hodjat, Risto Miikkulainen · 28 mai 2026
The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal weight matrices could in principle reduce parameter counts and computational overhead, most existing methods rely on static rank selection and do not …
- Stochastic Gradient Descent with Momentum is Algorithmically Stable
Yunwen Lei, Zimeng Wang, Xiaoming Yuan · 28 mai 2026
Stochastic gradient descent with momentum (SGDM) is one of the most widely used optimization algorithms in machine learning. While optimization properties of SGDM have been extensively studied in the literature, it remains insufficiently understood whether and when SGDM can generalize well to unseen…
- Deep Neural Network Training as Random Effects: An Optimization-Inference Duality
Minhao Yao, Ruoyu Wang, Xihong Lin, Lin Liu, Zhonghua Liu · 28 mai 2026
Deep neural networks (DNNs) have achieved remarkable empirical success, yet their training dynamics remain understood mainly from optimization rather than statistical principles. Here we develop a statistical framework for DNN training in the over-parameterized regime by showing that the prediction …
- Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency
Yibo Jacky Zhang, Zeyu Tang, Sanmi Koyejo · 28 mai 2026
Backpropagation is the default learning rule for artificial neural networks and is often treated as the settled approach whenever differentiability is available. In this work, we revisit this convention through a theoretical lens of sample efficiency. We introduce a unified vectorized feedback frame…
- Can Entry-Wise Clipping Give Spectral Control of Stochastic Gradients?
Zitao Song, Cedar Site Bai, Zhe Zhang, Brian Bullins, David F. Gleich · 28 mai 2026
Training instabilities such as loss spikes are frequently the result of stochastic gradient noise. Because of rare expressions in language training data, and multiple layer composition, the noise impact is heavy-tailed and survives mini-batch averaging. Existing remedies trade off structure against …
- On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note
Guangyi Zou, Roman Vershynin · 28 mai 2026
This short note presents a dimension-independent subgaussian concentration bound for Gaussian vectors under coordinate-wise nonlinear mappings. Discovered by Gemini 3.5 Flash, this result applies to any bounded function under a well-conditioned covariance. We apply this tool to answer a question of …
- Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings
Stanislav Kirdey, Clark Labs Inc · 28 mai 2026
Clark Hash is a small method for storing neural embeddings in less space. It normalizes each database vector, applies a deterministic sparse signed Johnson-Lindenstrauss projection, clips the result, and stores a fixed-width scalar-quantized code. Queries stay in floating point and are scored agains…
- Learning Theory of the SVRG: Generalization and Convergence Analysis
Yunwen Lei, Zimeng Wang, Xiaoming Yuan · 28 mai 2026
Variance reduction (VR) methods employ stochastic gradients with decreasing variance, and they have been widely applied to solve large-scale optimization problems in machine learning because of their efficiency. Existing theoretical studies of VR methods are mainly focused on the convergence analysi…
- Smoothed Score Queries and the Complexity of Sampling
Jingbo Liu · 28 mai 2026
We study the query complexity of sampling from high-dimensional Gaussian distributions using gradient information. In the standard oracle model, exact gradients expose only matrix-vector products with the precision matrix, leading to polynomial approximation barriers and a characteristic \(\sqrt{\ka…
- Optimal ridge regularization revisited
Jack Timmermans, Sergio A. Alvarez · 28 mai 2026
We consider $L^2$-regularized linear (ridge) regression over a finite data sample $X$ with bounded covariance and linear prediction targets $y$ with additive isotropic noise of finite variance. We present an iterative procedure to compute the optimal regularization strength numerically from the gene…
- Worker Disagreement Reveals Sharp Directions in Local SGD
Tolga Dimlioglu, Kristi Topollai, Anna Choromanska · 28 mai 2026
Deep neural network training often exhibits highly anisotropic loss geometry, where a few sharp dominant Hessian directions coexist with a large flatter bulk. Gradients tend to align disproportionately with these dominant directions, although stable progress often requires movement through flatter b…
- Stochastic global optimization of continuous functions via random walks on Grassmannians
Kartik Gupta, Stephen D. Miller, Pradeep Ravikumar, Ramarathnam Venkatesan · 27 mai 2026
We introduce a stochastic global optimization method based on random walks on Grassmannian manifolds. To minimize a continuous objective $\ell:\mathbb{R}^d\rightarrow\mathbb{R}$, the method repeatedly samples random $k$-dimensional linear subspaces (with $k\ll d$), solves the resulting low-dimension…
- Convergence of Spectral Descent for Non-smooth Optimization
Yixuan Yang, Yuqing He, Song Li · 27 mai 2026
The Muon optimizer has recently demonstrated remarkable empirical success in training large language models. However, the theoretical understanding of its mechanisms remains limited. Current convergence guarantees for Muon rely heavily on smoothness assumptions, leaving its non-smooth convergence be…
- Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks
Ali Hussaini Umar, Alessandro Laio · 27 mai 2026
Neural networks are known to develop latent representations that are $aligned$, namely structurally similar across networks trained with different architectures, training protocols, or training datasets. We study this phenomenon in a controlled setting, where we train an ensemble of networks on regr…
- Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias
James Town, Etienne Boursier, Ben Lewis, Matthias Englert, Ranko Lazic · 27 mai 2026
The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods remains incomplete. This is especially true in settings with mild overparameterization. In this work, we study the gradient flow dynamics of two-lay…
- PILOT: Policy-Informed Learned Optimization for Adaptive Deep Network Training
Sattam Altuuaim, Lama Ayash, Muhammad Mubashar, Naeemullah Khan · 26 mai 2026
Despite the central role of optimization in deep learning, most optimizers rely on update structures whose functional form is fixed before training begins. This static design can limit their ability to respond to changing gradient behavior across the loss landscape, where training may shift between …
- Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
Huangyu Xu, Jingqin Yang, Qianqian Xu, Jiaye Teng · 26 mai 2026
Sparse optimization is a fundamental challenge in various practical applications. A popular approach to sparse optimization is $\ell_p$ regularization. However, it may encounter optimization instability due to the unbounded gradients when $0<p<1$. In this paper, we introduce a novel approach to spar…
- Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding
Rustem Takhanov, Zhenisbek Assylbekov · 26 mai 2026
Conditionally positive definite (CPD) kernels are defined with respect to a function class $\mathcal{F}$. It is well known that such a kernel $K$ is associated with its native space (defined analogously to an RKHS), which in turn gives rise to a learning method -- called conditional kernel ridge reg…
- FLoRIST: Singular Value Thresholding for Efficient and Accurate Federated Fine-Tuning of Large Language Models
Hariharan Ramesh, Jyotikrishna Dass · 26 mai 2026
Integrating Low-Rank Adaptation (LoRA) into federated learning offers a promising solution for parameter-efficient fine-tuning of Large Language Models (LLMs) without sharing local data. However, several methods designed for federated LoRA present significant challenges in balancing communication ef…
- EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
Chung-Yiu Yau, Dawei Li, Athanasios Glentis, Valentyn Boreiko, Hoi-To Wai, Mingyi Hong · 26 mai 2026
Lookahead-based acceleration methods, such as Nesterov's momentum, are widely used in optimization, but they often become unreliable in deep learning training mainly due to stochastic gradient noise and non-convex loss landscapes. In particular, standard lookahead relies on short-horizon update sign…
- LAPLEX: The FFT of Learnable Laplace Kernels
{\L}ukasz Struski, Hanna Blazhko, Piotr Kubaty, Jacek Tabor · 26 mai 2026
Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geometry paid for by dense parameters, random features, or low-rank surrogates. To move beyond this trade-off, we introduce LAPLEX, a class of exact, train…
- Algorithms with Polynomially-Improved Approximation Factors for the $2 \rightarrow q$ Norm, and Applications
Samuel B. Hopkins, Stefan Tiegel · 26 mai 2026
The $2 \rightarrow q$ norm of a matrix $X \in \mathbb{R}^{n \times d}$ is defined as $\lVert X \rVert_{2 \rightarrow q} = \sup_{\lVert v \rVert_2 = 1} \lVert Xv \rVert_q$. We give polynomial-time multiplicative approximation algorithms for this norm when $q > 2$ (i.e. in the hypercontractive setting…
- How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
Rei Higuchi, Ryotaro Kawata, Akifumi Wachi, Shokichi Takakura, Kohei Miyaguchi, Taiji Suzuki · 26 mai 2026
Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depends on errors in reward-tilted regions. We study this feedback in a Gaussian single-index model with $r^*(x) = \sigma^*(…
