Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1612 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- All ERMs Can Fail in Stochastic Convex Optimization Lower Bounds in Linear Dimension
Tal Burla, Roi Livni · 10 de febrero de 2026
We study the sample complexity of the best-case Empirical Risk Minimizer in the setting of stochastic convex optimization. We show that there exists an instance in which the sample size is linear in the dimension, learning is possible, but the Empirical Risk Minimizer is likely to be unique and to o…
- ODELoRA: Training Low-Rank Adaptation by Solving Ordinary Differential Equations
Yihang Gao, Vincent Y. F. Tan · 10 de febrero de 2026
Low-rank adaptation (LoRA) has emerged as a widely adopted parameter-efficient fine-tuning method in deep transfer learning, due to its reduced number of trainable parameters and lower memory requirements enabled by Burer-Monteiro factorization on adaptation matrices. However, classical LoRA trainin…
- Improving Online-to-Nonconvex Conversion for Smooth Optimization via Double Optimism
Francisco Patitucci, Ruichen Jiang, Aryan Mokhtari · 10 de febrero de 2026
A recent breakthrough in nonconvex optimization is the online-to-nonconvex conversion framework of [Cutkosky et al., 2023], which reformulates the task of finding an $\varepsilon$-first-order stationary point as an online learning problem. When both the gradient and the Hessian are Lipschitz continu…
- Scalable LinUCB: Low-Rank Design Matrix Updates for Recommenders with Large Action Spaces
Evgenia Shustova, Marina Sheshukova, Sergey Samsonov, Evgeny Frolov · 10 de febrero de 2026
In this paper, we introduce PSI-LinUCB, a scalable variant of LinUCB that enables efficient training, inference, and memory usage by representing the inverse regularized design matrix as a sum of a diagonal matrix and low-rank correction. We derive numerically stable rank-1 and batched updates that …
- Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate
Zhiqi Bu, Shiyun Xu, Jialin Mao · 10 de febrero de 2026
Deep learning has non-convex loss landscape and its optimization dynamics is hard to analyze or control. Nevertheless, the dynamics can be empirically convex-like across various tasks, models, optimizers, hyperparameters, etc. In this work, we examine the applicability of convexity and Lipschitz con…
- ARO: A New Lens On Matrix Optimization For Large Models
Wenbo Gong, Javier Zazo, Qijun Luo, Puqian Wang, James Hensman, Chao Ma · 10 de febrero de 2026
Matrix-based optimizers have attracted growing interest for improving LLM training efficiency, with significant progress centered on orthogonalization/whitening based methods. While yielding substantial performance gains, a fundamental question arises: can we develop new paradigms beyond orthogonali…
- Tighter Information-Theoretic Generalization Bounds via a Novel Class of Change of Measure Inequalities
Yanxiao Liu, Yijun Fan an Deniz G\"und\"uz · 10 de febrero de 2026
In this paper, we propose a novel class of change of measure inequalities via a unified framework based on the data processing inequality for $f$-divergences, which is surprisingly elementary yet powerful enough to yield tighter inequalities. We provide change of measure inequalities in terms of a b…
- Enhancing Newton-Kaczmarz training of Kolmogorov-Arnold networks through concurrency
Andrew Polar, Michael Poluektov · 10 de febrero de 2026
The present paper introduces concurrency-driven enhancements to the training algorithm for the Kolmogorov-Arnold networks (KANs) that is based on the Newton-Kaczmarz (NK) method. As indicated by prior research, the NK-based training for KANs offers state-of-the-art performance in terms of accuracy a…
- The Median is Easier than it Looks: Approximation with a Constant-Depth, Linear-Width ReLU Network
Abhigyan Dutta, Itay Safran, Paul Valiant · 10 de febrero de 2026
We study the approximation of the median of $d$ inputs using ReLU neural networks. We present depth-width tradeoffs under several settings, culminating in a constant-depth, linear-width construction that achieves exponentially small approximation error with respect to the uniform distribution over t…
- Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
Dingzhi Yu, Hongyi Tao, Yuanyu Wan, Luo Luo, Lijun Zhang · 10 de febrero de 2026
While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have recently demonstrated superior empirical performance over AdamW in training large language models (LLM). However, a theoretical understanding of why sign-based …
- Near-optimal Swap Regret Minimization for Convex Losses
Lunjia Hu, Jon Schneider, Yifan Wu · 10 de febrero de 2026
We give a randomized online algorithm that guarantees near-optimal $\widetilde O(\sqrt T)$ expected swap regret against any sequence of $T$ adaptively chosen Lipschitz convex losses on the unit interval. This improves the previous best bound of $\widetilde O(T^{2/3})$ and answers an open question of…
- Efficient Distribution Learning with Error Bounds in Wasserstein Distance
Eduardo Figueiredo, Steven Adams, Luca Laurenti · 10 de febrero de 2026
The Wasserstein distance has emerged as a key metric to quantify distances between probability distributions, with applications in various fields, including machine learning, control theory, decision theory, and biological systems. Consequently, learning an unknown distribution with non-asymptotic a…
- Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
Shuang Liang, Guido Mont\'ufar · 10 de febrero de 2026
We examine gradient descent in matrix factorization and show that under large step sizes the parameter space develops a fractal structure. We derive the exact critical step size for convergence in scalar-vector factorization and show that near criticality the selected minimizer depends sensitively o…
- Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent
Shota Imai, Sota Nishiyama, Masaaki Imaizumi · 10 de febrero de 2026
The dynamics of gradient-based training in neural networks often exhibit nontrivial structures; hence, understanding them remains a central challenge in theoretical machine learning. In particular, a concept of feature unlearning, in which a neural network progressively loses previously learned feat…
- Mutual information and task-relevant latent dimensionality
Paarth Gulati, Eslam Abdelaleem, Audrey Sederberg, Ilya Nemenman · 10 de febrero de 2026
Estimating the dimensionality of the latent representation needed for prediction -- the task-relevant dimension -- is a difficult, largely unsolved problem with broad scientific applications. We cast it as an Information Bottleneck question: what embedding bottleneck dimension is sufficient to compr…
- Scalable Mean-Field Variational Inference via Preconditioned Primal-Dual Optimization
Jinhua Lyu, Tianmin Yu, Ying Ma, Naichen Shi · 10 de febrero de 2026
In this work, we investigate the large-scale mean-field variational inference (MFVI) problem from a mini-batch primal-dual perspective. By reformulating MFVI as a constrained finite-sum problem, we develop a novel primal-dual algorithm based on an augmented Lagrangian formulation, termed primal-dual…
- Provably robust learning of regression neural networks using $\beta$-divergences
Abhik Ghosh, Suryasis Jana · 10 de febrero de 2026
Regression neural networks (NNs) are most commonly trained by minimizing the mean squared prediction error, which is highly sensitive to outliers and data contamination. Existing robust training methods for regression NNs are often limited in scope and rely primarily on empirical validation, with on…
- Parallel Layer Normalization for Universal Approximation
Yunhao Ni, Yuxin Guo, Yuhe Liu, Wenxin Sun, Jie Luo, Wenjun Wu, Lei Huang · 10 de febrero de 2026
This paper studies the approximation capabilities of neural networks that combine layer normalization (LN) with linear layers. We prove that networks consisting of two linear layers with parallel layer normalizations (PLNs) inserted between them (referred to as PLN-Nets) achieve universal approximat…
- From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency
Sizhe Dang, Jiaqi Shao, Xiaodong Zheng, Guang Dai, Yan Song, Haishan Ye · 10 de febrero de 2026
As foundation models continue to scale, pretraining increasingly relies on data-parallel distributed optimization, making bandwidth-limited gradient synchronization a key bottleneck. Orthogonally, projection-based low-rank optimizers were mainly designed for memory efficiency, but remain suboptimal …
- Target noise: A pre-training based neural network initialization for efficient high resolution learning
Shaowen Wang, Tariq Alkhalifah · 9 de febrero de 2026
Weight initialization plays a crucial role in the optimization behavior and convergence efficiency of neural networks. Most existing initialization methods, such as Xavier and Kaiming initializations, rely on random sampling and do not exploit information from the optimization process itself. We pro…
- Reparameterization Proximal Policy Optimization
Hai Zhong, Xun Wang, Zhuoran Li, Longbo Huang · 9 de febrero de 2026
By leveraging differentiable dynamics, Reparameterization Policy Gradient (RPG) achieves high sample efficiency. However, current approaches are hindered by two critical limitations: the under-utilization of computationally expensive dynamics Jacobians and inherent training instability. While sample…
- Trust Regions Sell, But Who's Buying? Overlap Geometry as an Alternative Trust Region for Policy Optimization
Gaurish Trivedi, Alakh Sharma, Kartikey Singh Bhandari, Yash Sinha, Pratik Narang, Dhruv Kumar, Jagat Sesh Challa · 9 de febrero de 2026
Standard trust-region methods constrain policy updates via Kullback-Leibler (KL) divergence. However, KL controls only an average divergence and does not directly prevent rare, large likelihood-ratio excursions that destabilize training--precisely the failure mode that motivates heuristics such as P…
- Optimistic Training and Convergence of Q-Learning -- Extended Version
Prashant Mehta, Sean Meyn · 9 de febrero de 2026
In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the $(\varepsilon,\kappa)$-tamed Gibbs policy; $\kappa$ is inverse temperature, and $\varepsilon>0$ is introduced for additional exploration. Under these assump…
- RanSOM: Second-Order Momentum with Randomized Scaling for Constrained and Unconstrained Optimization
El Mahdi Chayti · 9 de febrero de 2026
Momentum methods, such as Polyak's Heavy Ball, are the standard for training deep networks but suffer from curvature-induced bias in stochastic settings, limiting convergence to suboptimal $\mathcal{O}(\epsilon^{-4})$ rates. Existing corrections typically require expensive auxiliary sampling or rest…
- Missing At Random as Covariate Shift: Correcting Bias in Iterative Imputation
Luke Shannon, Song Liu, Katarzyna Reluga · 9 de febrero de 2026
Accurate imputation of missing data is critical to downstream machine learning performance. We formulate missing data imputation as a risk minimisation problem, which highlights a covariate shift between the observed and unobserved data distributions. This covariate shift induced bias is not account…
