Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1.612 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Quantifying Error Propagation and Model Collapse in Diffusion Models
Nail B. Khelifa, Richard E. Turner, Ramji Venkataramanan · 1. Juni 2026
Machine learning models are increasingly trained or fine-tuned on synthetic data. Recursively training on such data has been observed to significantly degrade performance in a wide range of tasks, often characterized by a progressive drift away from the target distribution. In this work, we theoreti…
- Zeroth-Order Non-Log-Concave Sampling with Variance Reduction and Applications to Inverse Problems
M. Berk Sahin, Behzad Sharif, Abolfazl Hashemi · 1. Juni 2026
Sampling from high-dimensional, non-log-concave distributions with unnormalized densities remains a fundamental challenge in machine learning, particularly in black-box settings where gradient information is inaccessible or computationally prohibitive. While Langevin dynamics provides a principled f…
- Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail
Konstantin Nikolaou, Jonas Scheunemann, Sven Krippendorf, Samuel Tovey, Christian Holm · 1. Juni 2026
Neural scaling laws describe predictable power-law relationships between model size, dataset size, compute, and performance. While these laws guide the development of modern foundation models, the mechanisms underpinning them remain poorly understood, in part due to the absence of scalable analysis …
- Destruction is a General Strategy to Learn Generation; Diffusion's Strength is to Take it Seriously; Exploration is the Future
Pierre-Andr\'e No\"el · 1. Juni 2026
I present diffusion models as part of a family of machine learning techniques that withhold information from a model's input and train it to guess the withheld information. I argue that diffusion's destroying approach to withholding is more flexible than typical hand-crafted information withholding …
- Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling
Dmitrii Feoktistov, Timofey Belinsky, Andrey Veprikov, Amir Zainullin, Aleksandr Beznosikov · 1. Juni 2026
Sign-based and LMO-inspired optimizers have recently attracted substantial attention in deep learning due to their strong performance and low memory footprint. However, their fixed-magnitude updates can hurt terminal convergence: they decouple update mechanisms from gradient magnitudes and fail to a…
- S$^3$LDBO: A Snapshot Single-Loop Algorithm for Decentralized Bilevel Optimization
Chao Yin, Youran Dong, Shiqian Ma, Bofan Wang, Junfeng Yang · 1. Juni 2026
Networked AI systems increasingly rely on multiple agents that collaboratively learn and adapt models over communication networks. In such systems, bilevel formulations naturally arise in hyperparameter optimization, data cleaning, and meta-learning, but the repeated evaluation of gradients, Jacobia…
- Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
Luca Benfenati, Matteo Risso, Andrea Vannozzi, Ahmet Caner Y\"uz\"ug\"uler, Lukas Cavigelli, Enrico Macii, Daniele Jahier Pagliari, Alessio Burrello · 1. Juni 2026
Key-value (KV) caching enables fast autoregressive decoding but at long contexts becomes a dominant bottleneck in High Bandwidth Memory (HBM) capacity and bandwidth. A common mitigation is to compress cached keys and values by projecting per-head matrices to a lower rank, storing only the projection…
- Inference of Online Newton Methods with Nesterov's Accelerated Sketching
Haoxuan Wang, Xinchen Du, Sen Na · 1. Juni 2026
Reliable decision-making with streaming data requires principled uncertainty quantification of online methods. While first-order methods enable efficient iterate updates, their inference procedures still require updating proper (covariance) matrices, incurring $O(d^2)$ time and memory complexity, an…
- Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness, and Safety
Seok-Jin Kim · 1. Juni 2026
We study the multi-task linear regression problem in the presence of contaminated tasks. We address the setting where the unknown parameters of a majority of tasks are close in the $\ell_2$-norm, while a fraction of tasks are arbitrary outliers. Existing theoretical frameworks for this problem rely …
- Wall-Clock Complexity for Zeroth-Order Optimization with Tunable Oracle Fidelity
Alexandra Suvorikova, Igor Pavlov, Artem Vasin, Georgii Bychkov, Anastasia Antsiferova, Darina Dvinskikh, Alexander Gasnikov · 1. Juni 2026
Zeroth-order (black-box) optimization is applied when gradients are unavailable and objective evaluations rely on expensive simulations. In many such applications, the oracle fidelity is tunable: higher-accuracy queries reduce noise but incur higher computational costs. To capture this trade-off, we…
- Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
Zaiwei Chen, Siva Theja Maguluri · 1. Juni 2026
We survey Lyapunov-based techniques for the finite-time analysis of stochastic iterative algorithms, also known as stochastic approximation (SA) algorithms, for solving fixed-point equations $\bar{F}(x)=x$, where the operator $\bar{F}(\cdot)$ can only be accessed through a noisy oracle. We first foc…
- Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization Lens
Junbin Qiu, Zhaowei Hong, Renzhe Xu, Yao Shu · 1. Juni 2026
Accurate Zeroth-Order (ZO) Hessian estimation is a cornerstone of derivative-free methods, essential for tasks such as bilevel optimization, Bayesian inference, and uncertainty quantification. However, obtaining a complete suite of low-variance estimators for the Hessian and its inverse in high-dime…
- Stochastic Gradients under Nuisances
Facheng Yu, Ronak Mehta, Alex Luedtke, Zaid Harchaoui · 1. Juni 2026
Stochastic gradient optimization is the dominant learning paradigm for a variety of scenarios, from classical supervised learning to modern self-supervised learning. We consider stochastic gradient algorithms for learning problems whose objectives rely on unknown nuisance parameters, and establish n…
- Bayesian Inference with Shaped Deep Non-linear MLPs
Boris Hanin, Tianze Jiang · 1. Juni 2026
A central aim of deep learning theory is to characterize how neural networks make predictions in the regime of simultaneously large model and training set size. Since the limits of diverging number of model parameters and dataset size do not commute it is not clear a priori what limits exist. In thi…
- Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
Sharan Vaswani, Yifan Sun, Reza Babanezhad · 1. Juni 2026
Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generalize this assumption to objectives whose curvature is an affine function of the objective value. This property is satisfie…
- Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning
Bing Liu, Boao Kong, Limin Lu, Kun Yuan, Chengcheng Zhao · 1. Juni 2026
Decentralized learning often involves a weighted global loss with heterogeneous node weights $\lambda$. We revisit two natural strategies for incorporating these weights: (i) embedding them into the local losses to retain a uniform weight (and thus a doubly stochastic matrix), and (ii) keeping the o…
- From Sublinear to Linear: Local Convergence in Finite-Width Networks via Locally Polyak-Lojasiewicz Regions
Agnideep Aich, Ashit Baran Aich, Bruce Wade · 29. Mai 2026
We study local linear convergence of gradient descent for finite-width feedforward networks under the squared empirical loss. Prior work shows that GD can remain confined to a Locally Quasi-Convex Region (LQCR) around initialization, but only gives a sublinear rate. We show that if the empirical Neu…
- Diffusion Models Are Statistically Optimal for Learning Low-Dimensional Multi-Modal Distributions
Jingda Wu, Changxiao Cai · 29. Mai 2026
Score-based diffusion models have demonstrated remarkable empirical success in learning high-dimensional distributions, particularly those exhibiting low-dimensional and multi-modal structures. However, theoretical understanding of their statistical efficiency remains limited. Existing theories typi…
- How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Jeff A. Bilmes, Gantavya Bhatt, Arnav M. Das · 29. Mai 2026
Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both that common neural-scaling-law objectives and the Vendi Score are submodular. We further show that the Vendi Score is a special case of a broader class of submodu…
- Matching Rates and Optimal Allocation for Federated Probe-Logit Distillation under Heterogeneous Bandwidth Budgets
Prasanjit Dubey, Xiaoming Huo · 29. Mai 2026
In federated language modeling, $K$ nodes each hold $n$ samples but cannot pool data or exchange full-precision gradients or weights. We study the minimax rate at which a conditional distribution over $V$ tokens can be estimated when each node may upload at most $B$ bits per query in a public probe …
- Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime
Paolo Baglioni, Christian Keup, Vincenzo Zimbardo, Rosalba Pacelli, Alessandro Vezzani, Raffaella Burioni, Pietro Rotondo · 29. Mai 2026
The scaling limit where both the size of the training set $P$ and the width $N$ of a deep neural network grow at the same rate, the so-called proportional-width regime, has been intensely studied for shallow, single-hidden-layer networks. However, extending these non-perturbative results from shallo…
- MoSSP: A Momentum-Based Single-Loop Stochastic Penalty Method for Nonconvex Constrained DC-Regularized Optimization
Luxuan Li, Chunfeng Cui, Xiao Wang · 29. Mai 2026
In this paper, we study a structured class of nonconvex constrained stochastic problems with difference-of-convex (DC) regularization, where the feasible set is possibly nonconvex and the concave part of the DC regularizer is allowed to be nonsmooth. The fundamental challenge lies in maintaining fea…
- On Distributional Reinforcement Learning in Chaotic Dynamical Systems
James Rudd-Jones, Mirco Musolesi, Mar\'ia P\'erez-Ortiz · 29. Mai 2026
Chaotic dynamical systems pose a fundamental challenge for Reinforcement Learning (RL): exponential sensitivity to initial conditions induces high-variance bootstrap targets and poorly conditioned gradient updates. Chaotic dynamics arise across scientific and engineering domains, from fluid flows an…
- Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
Mohua Das, Pierfrancesco Beneventano, Shibshankar Dey, Gareth H. McKinkey, Tomaso Poggio · 29. Mai 2026
Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survives the training pipeline. To make the question measurable, we introduce initialization memory: the dependence of the vali…
- Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang · 29. Mai 2026
Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators. In such non-smooth regimes, adaptive optimizers such as Adam suffer fro…
