Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1 612 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Relative Wasserstein Angle and the Problem of the $W_2$-Nearest Gaussian Distribution
Binshuai Wang, Peng Wei · 2 février 2026
We study the problem of quantifying how far an empirical distribution deviates from Gaussianity under the framework of optimal transport. By exploiting the cone geometry of the relative translation invariant quadratic Wasserstein space, we introduce two novel geometric quantities, the relative Wasse…
- FlexLoRA: Entropy-Guided Flexible Low-Rank Adaptation
Muqing Liu, Chongjie Si, Yuheng Jia · 2 février 2026
Large pre-trained models achieve remarkable success across diverse domains, yet fully fine-tuning incurs prohibitive computational and memory costs. Parameter-efficient fine-tuning (PEFT) has thus become a mainstream paradigm. Among them, Low-Rank Adaptation (LoRA) introduces trainable low-rank matr…
- Bias-Optimal Bounds for SGD: A Computer-Aided Lyapunov Analysis
Daniel Cortild, Lucas Ketels, Juan Peypouquet, Guillaume Garrigos · 2 février 2026
The non-asymptotic analysis of Stochastic Gradient Descent (SGD) typically yields bounds that decompose into a bias term and a variance term. In this work, we focus on the bias component and study the extent to which SGD can match the optimal convergence behavior of deterministic gradient descent. A…
- Riemannian Lyapunov Optimizer: A Unified Framework for Optimization
Yixuan Wang, Omkar Sudhir Patil, Warren E. Dixon · 2 février 2026
We introduce Riemannian Lyapunov Optimizers (RLOs), a family of optimization algorithms that unifies classic optimizers within one geometric framework. Unlike heuristic improvements to existing optimizers, RLOs are systematically derived from a novel control-theoretic framework that reinterprets opt…
- Approximating $f$-Divergences with Rank Statistics
Viktor Stein, Jos\'e Manuel de Frutos · 2 février 2026
We introduce a rank-statistic approximation of $f$-divergences that avoids explicit density-ratio estimation by working directly with the distribution of ranks. For a resolution parameter $K$, we map the mismatch between two univariate distributions $\mu$ and $\nu$ to a rank histogram on $\{ 0, \ldo…
- Inexact Moreau Envelope Lagrangian Method for Non-Convex Constrained Optimization under Local Error Bound Conditions on Constraint Functions
Yankun Huang, Qihang Lin, Yangyang Xu · 2 février 2026
In this paper, we investigate how structural properties of the constraint system impact the oracle complexity of smooth non-convex optimization problems with convex inequality constraints over a simple polytope. In particular, we show that, under a local error bound condition with exponent $d\in[1,2…
- AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
Thalaiyasingam Ajanthan, Sameera Ramasinghe, Gil Avraham, Hadi Mohaghegh Dolatabadi, Chamin P Hewa Koneputugodage, Violetta Shevchenko, Yan Zuo, Alexander Long · 2 février 2026
Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their scalability. We address this communication bottleneck by introducing a…
- Local Intrinsic Dimension of Representations Predicts Alignment and Generalization in AI Models and Human Brain
Junjie Yu, Wenxiao Ma, Chen Wei, Jianyu Zhang, Haotian Deng, Zihan Deng, Quanying Liu · 2 février 2026
Recent work has found that neural networks with stronger generalization tend to exhibit higher representational alignment with one another across architectures and training paradigms. In this work, we show that models with stronger generalization also align more strongly with human neural activity. …
- On Approximate Computation of Critical Points
Amir Ali Ahmadi, Georgina Hall · 30 janvier 2026
We show that computing even very coarse approximations of critical points is intractable for simple classes of nonconvex functions. More concretely, we prove that if there exists a polynomial-time algorithm that takes as input a polynomial in $n$ variables of constant degree (as low as three) and ou…
- A block-coordinate descent framework for non-convex composite optimization. Application to sparse precision matrix estimation
Guillaume Lauga (LJAD) · 30 janvier 2026
Block-coordinate descent (BCD) is the method of choice to solve numerous large scale optimization problems, however their theoretical study for non-convex optimization, has received less attention. In this paper, we present a new block-coordinate descent (BCD) framework to tackle non-convex composit…
- Efficient Stochastic Optimisation via Sequential Monte Carlo
James Cuin, Davide Carbone, Yanbo Tang, O. Deniz Akyildiz · 30 janvier 2026
The problem of optimising functions with intractable gradients frequently arise in machine learning and statistics, ranging from maximum marginal likelihood estimation procedures to fine-tuning of generative models. Stochastic approximation methods for this class of problems typically require inner …
- Divergence Results and Convergence of a Variance Reduced Version of ADAM
Ruiqi Wang, Diego Klabjan · 30 janvier 2026
Stochastic optimization algorithms using exponential moving averages of the past gradients, such as ADAM, RMSProp and AdaGrad, have been having great successes in many applications, especially in training deep neural networks. ADAM in particular stands out as efficient and robust. Despite of its out…
- Training Memory in Deep Neural Networks: Mechanisms, Evidence, and Measurement Gaps
Vasileios Sevetlidis, George Pavlidis · 30 janvier 2026
Modern deep-learning training is not memoryless. Updates depend on optimizer moments and averaging, data-order policies (random reshuffling vs with-replacement, staged augmentations and replay), the nonconvex path, and auxiliary state (teacher EMA/SWA, contrastive queues, BatchNorm statistics). This…
- KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices
Wuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Danilo Mandic · 30 janvier 2026
The success of Hyper-Connections (HC) in neural networks (NN) has also highlighted issues related to its training instability and restricted scalability. The Manifold-Constrained Hyper-Connections (mHC) mitigate these challenges by projecting the residual connection space onto a Birkhoff polytope, h…
- FISMO: Fisher-Structured Momentum-Orthogonalized Optimizer
Chenrui Xu, Wenjing Yan, Ying-Jun Angela Zhang · 30 janvier 2026
Training large-scale neural networks requires solving nonconvex optimization where the choice of optimizer fundamentally determines both convergence behavior and computational efficiency. While adaptive methods like Adam have long dominated practice, the recently proposed Muon optimizer achieves sup…
- A Trainable Optimizer
Ruiqi Wang, Diego Klabjan · 30 janvier 2026
The concept of learning to optimize involves utilizing a trainable optimization strategy rather than relying on manually defined full gradient estimations such as ADAM. We present a framework that jointly trains the full gradient estimator and the trainable weights of the model. Specifically, we pro…
- Differentiable Knapsack and Top-k Operators via Dynamic Programming
Germain Vivier-Ardisson, Micha\"el E. Sander, Axel Parmentier, Mathieu Blondel · 30 janvier 2026
Knapsack and Top-k operators are useful for selecting discrete subsets of variables. However, their integration into neural networks is challenging as they are piecewise constant, yielding gradients that are zero almost everywhere. In this paper, we propose a unified framework casting these operator…
- High-dimensional learning dynamics of multi-pass Stochastic Gradient Descent in multi-index models
Zhou Fan, Leda Wang · 30 janvier 2026
We study the learning dynamics of a multi-pass, mini-batch Stochastic Gradient Descent (SGD) procedure for empirical risk minimization in high-dimensional multi-index models with isotropic random data. In an asymptotic regime where the sample size $n$ and data dimension $d$ increase proportionally, …
- Solving the Offline and Online Min-Max Problem of Non-smooth Submodular-Concave Functions: A Zeroth-Order Approach
Amir Ali Farzin, Yuen-Man Pun, Philipp Braun, Tyler Summers, Iman Shames · 30 janvier 2026
We consider max-min and min-max problems with objective functions that are possibly non-smooth, submodular with respect to the minimiser and concave with respect to the maximiser. We investigate the performance of a zeroth-order method applied to this problem. The method is based on the subgradient …
- Don't be so Stief! Learning KV Cache low-rank approximation over the Stiefel manifold
Luca Benfenati, Matteo Risso, Andrea Vannozzi, Ahmet Caner Y\"uz\"ug\"uler, Lukas Cavigelli, Enrico Macii, Daniele Jahier Pagliari, Alessio Burrello · 30 janvier 2026
Key--value (KV) caching enables fast autoregressive decoding but at long contexts becomes a dominant bottleneck in High Bandwidth Memory (HBM) capacity and bandwidth. A common mitigation is to compress cached keys and values by projecting per-head matrixes to a lower rank, storing only the projectio…
- Manifold constrained steepest descent
Kaiwei Yang, Lexiao Lai · 30 janvier 2026
Norm-constrained linear minimization oracle (LMO)-based optimizers such as spectral gradient descent and Muon are attractive in large-scale learning, but extending them to manifold-constrained problems is nontrivial and often leads to nested-loop schemes that solve tangent-space subproblems iterativ…
- PPI-SVRG: Unifying Prediction-Powered Inference and Variance Reduction for Semi-Supervised Optimization
Ruicheng Ao, Hongyu Chen, Haoyang Liu, David Simchi-Levi, Will Wei Sun · 30 janvier 2026
We study semi-supervised stochastic optimization when labeled data is scarce but predictions from pre-trained models are available. PPI and SVRG both reduce variance through control variates -- PPI uses predictions, SVRG uses reference gradients. We show they are mathematically equivalent and develo…
- Let the Optimizers Optimize Themselves
Jaerin Lee, Kyoung Mu Lee · 30 janvier 2026
We lay the theoretical foundation for automating optimizer design in gradient-based learning. Based on the greedy principle, we formulate the problem of designing optimizers and their hyperparameters as maximizing the instantaneous decrease in loss. By treating an optimizer as a function that transl…
- Missing-Data-Induced Phase Transitions in Spectral PLS for Multimodal Learning
Anders Gj{\o}lbye, Ida Kargaard, Emma Kargaard, Lars Kai Hansen · 30 janvier 2026
Partial Least Squares (PLS) learns shared structure from paired data via the top singular vectors of the empirical cross-covariance (PLS-SVD), but multimodal datasets often have missing entries in both views. We study PLS-SVD under independent entry-wise missing-completely-at-random masking in a pro…
- Differentiable Integer Linear Programming is not Differentiable & it's not a mere technical problem
Thanawat Sornwanee · 30 janvier 2026
We show how the differentiability method employed in the paper ``Differentiable Integer Linear Programming'', Geng, et al., 2025 as shown in its theorem 5 is incorrect. Moreover, there already exists some downstream work that inherits the same error. The underlying reason comes from that, though bei…
