Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1 612 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- On the Intrinsic Dimensions of Data in Kernel Learning
Rustem Takhanov · 23 janvier 2026
The manifold hypothesis suggests that the generalization performance of machine learning methods improves significantly when the intrinsic dimension of the input distribution's support is low. In the context of KRR, we investigate two alternative notions of intrinsic dimension. The first, denoted $d…
- Progressive Power Homotopy for Non-convex Optimization
Chen Xu · 23 janvier 2026
We propose a novel first-order method for non-convex optimization of the form $\max_{\bm{w}\in\mathbb{R}^d}\mathbb{E}_{\bm{x}\sim\mathcal{D}}[f_{\bm{w}}(\bm{x})]$, termed Progressive Power Homotopy (Prog-PowerHP). The method applies stochastic gradient ascent to a surrogate objective obtained by fir…
- Non-Stationary Functional Bilevel Optimization
Jason Bohne, Ieva Petrulionyte, Michael Arbel, Julien Mairal, Pawe{\l} Polak · 23 janvier 2026
Functional bilevel optimization (FBO) provides a powerful framework for hierarchical learning in function spaces, yet current methods are limited to static offline settings and perform suboptimally in online, non-stationary scenarios. We propose SmoothFBO, the first algorithm for non-stationary FBO …
- Panther: Faster and Cheaper Computations with Randomized Numerical Linear Algebra
Fahd Seddik, Abdulrahman Elbedewy, Gaser Sami, Mohamed Abdelmoniem, Yahia Zakaria · 23 janvier 2026
Training modern deep learning models is increasingly constrained by GPU memory and compute limits. While Randomized Numerical Linear Algebra (RandNLA) offers proven techniques to compress these models, the lack of a unified, production-grade library prevents widely adopting these methods. We present…
- Stability, Complexity and Data-Dependent Worst-Case Generalization Bounds
Mario Tuci, Lennart Bastian, Benjamin Dupuis, Nassir Navab, Tolga Birdal, Umut \c{S}im\c{s}ekli · 23 janvier 2026
Providing generalization guarantees for stochastic optimization algorithms remains a key challenge in learning theory. Recently, numerous works demonstrated the impact of the geometric properties of optimization trajectories on generalization performance. These works propose worst-case generalizatio…
- TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction
Yuchen Fang, Xinshou Zheng, Javad Lavaei · 22 janvier 2026
We propose a stochastic trust-region method for unconstrained nonconvex optimization that incorporates stochastic variance-reduced gradients (SVRG) to accelerate convergence. Unlike classical trust-region methods, the proposed algorithm relies solely on stochastic gradient information and does not r…
- Efficient and Minimax-optimal In-context Nonparametric Regression with Transformers
Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma, William G. Underwood, Richard J. Samworth · 22 janvier 2026
We study in-context learning for nonparametric regression with $\alpha$-H\"older smooth regression functions, for some $\alpha>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained transformer with $\Theta(\log n)$ parameters and $\Omega\bigl(n^{2\al…
- One-Sided Matrix Completion from Ultra-Sparse Samples
Hongyang R. Zhang, Zhenshuo Zhang, Huy L. Nguyen, Guanghui Lan · 21 janvier 2026
Matrix completion is a classical problem that has received recurring interest across a wide range of fields. In this paper, we revisit this problem in an ultra-sparse sampling regime, where each entry of an unknown, $n\times d$ matrix $M$ (with $n \ge d$) is observed independently with probability $…
- Beyond Softmax and Entropy: Improving Convergence Guarantees of Policy Gradients by f-SoftArgmax Parameterization with Coupled Regularization
Safwan Labbi, Daniil Tiapkin, Paul Mangold, Eric Moulines · 21 janvier 2026
Policy gradient methods are known to be highly sensitive to the choice of policy parameterization. In particular, the widely used softmax parameterization can induce ill-conditioned optimization landscapes and lead to exponentially slow convergence. Although this can be mitigated by preconditioning,…
- Ordered Local Momentum for Asynchronous Distributed Learning under Arbitrary Delays
Chang-Wei Shi, Shi-Shang Wang, Wu-Jun Li · 21 janvier 2026
Momentum SGD (MSGD) serves as a foundational optimizer in training deep models due to momentum's key role in accelerating convergence and enhancing generalization. Meanwhile, asynchronous distributed learning is crucial for training large-scale deep models, especially when the computing capabilities…
- Empirical Risk Minimization with $f$-Divergence Regularization
Francisco Daunas, I\~naki Esnaola, Samir M. Perlaza, H. Vincent Poor · 21 janvier 2026
In this paper, the solution to the empirical risk minimization problem with $f$-divergence regularization (ERM-$f$DR) is presented and conditions under which the solution also serves as the solution to the minimization of the expected empirical risk subject to an $f$-divergence constraint are establ…
- BiCoLoR: Communication-Efficient Optimization with Bidirectional Compression and Local Training
Laurent Condat, Artavazd Maranjyan, Peter Richt\'arik · 21 janvier 2026
Slow and costly communication is often the main bottleneck in distributed optimization, especially in federated learning where it occurs over wireless networks. We introduce BiCoLoR, a communication-efficient optimization algorithm that combines two widely used and effective strategies: local traini…
- Concatenated Matrix SVD: Compression Bounds, Incremental Approximation, and Error-Constrained Clustering
Maksym Shamrai · 21 janvier 2026
Large collections of matrices arise throughout modern machine learning, signal processing, and scientific computing, where they are commonly compressed by concatenation followed by truncated singular value decomposition (SVD). This strategy enables parameter sharing and efficient reconstruction and …
- Suspicious Alignment of SGD: A Fine-Grained Step Size Condition Analysis
Shenyang Deng, Boyao Liao, Zhuoli Ouyang, Tianyu Pang, Minhak Song, Yaoqing Yang · 21 janvier 2026
This paper explores the suspicious alignment phenomenon in stochastic gradient descent (SGD) under ill-conditioned optimization, where the Hessian spectrum splits into dominant and bulk subspaces. This phenomenon describes the behavior of gradient alignment in SGD updates. Specifically, during the i…
- ButterflyMoE: Sub-Linear Ternary Experts via Structured Butterfly Orbits
Aryan Karmore · 21 janvier 2026
Linear memory scaling stores $N$ independent expert weight matrices requiring $\mathcal{O}(N \cdot d^2)$ memory, which exceeds edge devices memory budget. Current compression methods like quantization, pruning and low-rank factorization reduce constant factors but leave the scaling bottleneck unreso…
- Asymmetric regularization mechanism for GAN training with Variational Inequalities
Spyridon C. Giagtzoglou, Mark H. M. Winands, Barbara Franci · 21 janvier 2026
We formulate the training of generative adversarial networks (GANs) as a Nash equilibrium seeking problem. To stabilize the training process and find a Nash equilibrium, we propose an asymmetric regularization mechanism based on the classic Tikhonov step and on a novel zero-centered gradient penalty…
- Optimistic Gradient Learning with Hessian Corrections for High-Dimensional Black-Box Optimization
Yedidya Kfir, Elad Sarafian, Sarit Kraus, Yoram Louzoun · 21 janvier 2026
Black-box algorithms are designed to optimize functions without relying on their underlying analytical structure or gradient information, making them essential when gradients are inaccessible or difficult to compute. Traditional methods for solving black-box optimization (BBO) problems predominantly…
- Mixture-of-Experts as Soft Clustering: A Dual Jacobian-PCA Spectral Geometry Perspective
Feilong Liu · 21 janvier 2026
Mixture-of-Experts (MoE) architectures are commonly motivated by efficiency and conditional computation, but their effect on the geometry of learned functions and representations remains poorly characterized. In this work, we study MoEs through a geometric lens, interpreting routing as a form of sof…
- On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
Sharan Sahu, Cameron J. Hogan, Martin T. Wells · 21 janvier 2026
While momentum-based acceleration has been studied extensively in deterministic optimization problems, its behavior in nonstationary environments -- where the data distribution and optimal parameters drift over time -- remains underexplored. We analyze the tracking performance of Stochastic Gradient…
- Unit-Consistent (UC) Adjoint for GSD and Backprop in Deep Learning Applications
Jeffrey Uhlmann · 19 janvier 2026
Deep neural networks constructed from linear maps and positively homogeneous nonlinearities (e.g., ReLU) possess a fundamental gauge symmetry: the network function is invariant to node-wise diagonal rescalings. However, standard gradient descent is not equivariant to this symmetry, causing optimizat…
- Transient learning dynamics drive escape from sharp valleys in Stochastic Gradient Descent
Ning Yang, Yikuan Zhang, Qi Ouyang, Chao Tang, Yuhai Tu · 19 janvier 2026
Stochastic gradient descent (SGD) is central to deep learning, yet the dynamical origin of its preference for flatter, more generalizable solutions remains unclear. Here, by analyzing SGD learning dynamics, we identify a nonequilibrium mechanism governing solution selection. Numerical experiments re…
- Near-Optimal Decentralized Stochastic Nonconvex Optimization with Heavy-Tailed Noise
Menglian Wang, Zhuanghua Liu, Luo Luo · 19 janvier 2026
This paper studies decentralized stochastic nonconvex optimization problem over row-stochastic networks. We consider the heavy-tailed gradient noise which is empirically observed in many popular real-world applications. Specifically, we propose a decentralized normalized stochastic gradient descent …
- A New Convergence Analysis of Plug-and-Play Proximal Gradient Descent Under Prior Mismatch
Guixian Xu, Jinglai Li, Junqi Tang · 16 janvier 2026
In this work, we provide a new convergence theory for plug-and-play proximal gradient descent (PnP-PGD) under prior mismatch where the denoiser is trained on a different data distribution to the inference task at hand. To the best of our knowledge, this is the first convergence proof of PnP-PGD unde…
- COALA: Numerically Stable and Efficient Framework for Context-Aware Low-Rank Approximation
Uliana Parkina, Maxim Rakhuba · 16 janvier 2026
Recent studies suggest that context-aware low-rank approximation is a useful tool for compression and fine-tuning of modern large-scale neural networks. In this type of approximation, a norm is weighted by a matrix of input activations, significantly improving metrics over the unweighted case. Never…
- Sobolev Approximation of Deep ReLU Networks in Log-Barron Space
Changhoon Song, Seungchan Ko, Youngjoon Hong · 15 janvier 2026
Universal approximation theorems show that neural networks can approximate any continuous function; however, the number of parameters may grow exponentially with the ambient dimension, so these results do not fully explain the practical success of deep models on high-dimensional data. Barron space t…
