Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1.612 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Beyond LoRA: Is Sparsity-Induced Adaptation Better?
Elijah Cadenhead, Cristian McGee, Xin Li, El Houcine Bergou, Aritra Dutta · 15. Juni 2026
Low-rank adaptation (LoRA) and its variants provide a memory- and compute-efficient alternative to full fine-tuning of pre-trained models. However, questions remain about the comparative generalizability of these approaches and how the structural restrictions on low-rank updates preserve effective a…
- Compressed Computation is (probably) not Computation in Superposition
Jai Bhagat, Sara Molas-Medina, Giorgi Giglemiani, Stefan Heimersheim · 15. Juni 2026
We study whether the Compressed Computation (CC) toy model (Braun et al., 2025) is an instance of computation in superposition. The CC model appears to compute 100 ReLU functions with just 50 neurons, achieving a better loss than expected from only representing 50 ReLU functions. We show that the mo…
- Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
Guo Yu, Wenlin Liu, Yulan Hu, Hao-Xuan Ma, Jun-Peng Jiang, Han-Jia Ye · 15. Juni 2026
On-policy distillation (\textsc{OPD}) has recently become a prominent post-training recipe by combining two desirable ingredients: on-policy student trajectories and dense teacher supervision. However, how this hybrid changes a model's parameters remains unclear. Across several language and vision-l…
- Free Heavy-Tailed Lunch for Muon: A Theoretical Justification of Empirical Success
Florian H\"ubler, Thomas Pethick, Suvrit Sra · 15. Juni 2026
Non-Euclidean optimisation methods with matrix-valued updates, such as Muon and Scion, have recently shown strong empirical performance for training Transformer models, yet their theoretical advantages over Euclidean methods remain poorly understood. We address this gap in the heavy-tailed non-conve…
- Neural Slack Variables for Shape Constraints
Ruben Wiedemann, Antoine Jacquier, Lukas Gonon · 15. Juni 2026
Enforcing functional inequality constraints such as monotonicity and convexity in neural networks is a fundamental challenge in many industrial and scientific applications. Classical one-sided penalty methods, along with primal-dual methods gated by complementary slackness, provide constraint gradie…
- The Weight Norm Sets the Grokking Timescale: A Causal Delay Law
Truong Xuan Khanh, Doan Hoang Viet, Luu Duc Trung, Phan Thanh Duc · 15. Juni 2026
Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data. Whether the weight norm causes this delay is disputed: some studies report a critical norm at the transition, others observe grokking with no fixed norm at all. We settle this by interv…
- Direct Fisher Score Estimation for Likelihood Maximization
Sherman Khoo, Yakun Wang, Song Liu, Mark Beaumont · 15. Juni 2026
We study the problem of likelihood maximization when the likelihood function is intractable but model simulations are readily available. We propose a sequential, gradient-based optimization method that directly models the Fisher score based on a local score matching technique which uses simulations …
- Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning
Kaiwen Chen, Shuhai Zhang, Qiuwu Chen, Zimo Liu, Linxiao Li, Ying Sun, Yuchen Li, Yifan Zhang, Bo Han, Mingkui Tan · 15. Juni 2026
Large-scale neural network training increasingly relies on matrix-aware optimizers that exploit the structure of weight parameters beyond element-wise adaptation. However, existing matrix-aware methods such as Muon have an underappreciated vulnerability: their core operation, Newton-Schulz iteration…
- Scalable Deep Unfolding of Conic Optimizers
Alex Oshin, Rahul Vodeb Ghosh, Evangelos A. Theodorou · 15. Juni 2026
Deep unfolding (DU) accelerates iterative optimizers by introducing learnable components and training them through unrolled iterations, but extending DU to the large-scale semidefinite programs (SDPs) common in robotics has remained limited. Unrolling a full-update conic solver such as COSMO exposes…
- Simplex-Constrained Sparse Bagging: Transitioning from Uniform Priors to Sparse Posteriors in Ensemble Learning
Meher Sai Preetam, Meher Bhaskar · 15. Juni 2026
We present Simplex-Constrained Sparse Bagging (SCSB), a mathematically rigorous framework for post-training compression and probability calibration of bootstrap-based bagging ensembles. Standard bagging ensembles (such as Random Forests, Bagged SVMs, and Bagged Neural Networks) assign uniform voting…
- Operator Calculus for Population-Based Optimization: A Mean-Field Convergence Theory
Pekka Malo, Lauri Viitasaari, Patrik Nummi, Antti Suominen, Ankur Sinha, Olli Tahvonen · 15. Juni 2026
Population-based and distributional optimization methods, from evolution strategies and consensus-based optimization to covariance-matrix adaptation and stochastic gradient methods viewed as distributional dynamics, are widely used for nonconvex or black-box problems, yet their convergence analyses …
- Nonlinear Two-Time-Scale Stochastic Approximation: A Sharp Phase Transition and How to Beat It
Dhruv Sarkar, Vaneet Aggarwal · 15. Juni 2026
Recent finite-time analyses of nonlinear two-time-scale stochastic approximation show that under contractive assumptions the slow iterate $Y_k$ with stepsizes $\beta_k=\Theta(k^{-1})$ and $\alpha_k=\Theta(k^{-a})$, $a\in(1/2,1)$, generally satisfies a mean-square rate of order $k^{-a}$; decoupled $k…
- LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold
Franz Louis Cesista, Katherine Crowson, C\'edric Simal, Stella Biderman · 12. Juni 2026
Low-Rank Adaptation (LoRA) significantly reduces compute and memory costs for finetuning Deep Learning models but is often harder to tune than dense training: when using factor-wise optimizers such as AdamW, it is sensitive to initialization choices, its optimal learning rates transfer poorly across…
- Different Layers, Different Manifolds: Module-Wise Weight-Space Geometry in Transformer Optimization
Kirato Yoshihara · 12. Juni 2026
Weight-space geometry plays a central role in neural network optimization, yet manifold constraints are often applied uniformly across all weight matrices. In this work, we ask whether different transformer modules prefer different manifold geometries. We study Manifold Muon for GPT-2 pretraining an…
- Last-Iterate Convergence of Optimistic Multiplicative Weight Update
Francesco Orabona · 11. Juni 2026
Optimistic Gradient Descent Ascent (OGDA) and Optimistic Multiplicative-Weights Update (OMWU) are two very popular algorithms to solve convex/concave saddle-point problems, where OMWU is the non-Euclidean, entropic version of OGDA. It is known since the '80s that the last iterate of OGDA asymptotica…
- A Riemannian Approach to Low-Rank Optimal Transport
Pratik Jawanpuria, Bamdev Mishra · 11. Juni 2026
Low-rank optimal transport (OT) mitigates the quadratic scaling of classical solvers, yet existing approaches rely heavily on first-order mirror-descent updates that require careful hyperparameter tuning and ignore the optimization landscape's curvature. To address these limitations, we propose a un…
- Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity
Shira Vansover-Hager, Matan Schliserman, Ofir Schlisselberg, Tomer Koren · 11. Juni 2026
Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training. This raises a basic robustness question, crucial to reproducibility and reliability: how sensitive…
- Simplicity Suffices for Parameter Noise Injection in Stochastic Gradient Descent
Benjamin Leblanc, Louis-Jacob Lebel, Teddy Kana, Richard Kamel · 11. Juni 2026
Injecting noise into the optimization process is a well-established technique for improving the training and generalization of deep neural networks. Yet, despite the breadth of existing approaches, it remains unclear which design choices truly matter in practice. In this work, we investigate paramet…
- Quantized Stochastic Primal-Dual Methods for Distributed Optimization under Relaxed Global Geometry
Susmit Sarkar, Abhinav Raghuvanshi, Kushal Chakrabarti, Mayank Baranwal · 11. Juni 2026
We study distributed optimization with stochastic gradients and finite-bit communication modeled by random (unbiased) quantization. We propose q-PDGD, a quantized stochastic primal-dual method, and analyze it under relaxed global geometry. Under restricted secant inequality (RSI), a constant step-si…
- Conservation Laws from Data Symmetry in Neural Networks
Jakob Galley, Vahid Shahverdi, Axel Flinth · 10. Juni 2026
We explore whether intrinsic symmetries of the training data lead to conserved quantities during gradient-flow training of neural networks. Under the assumption that the loss function is analytic and non-polynomial, we prove that data symmetries generically do not induce any additional integrals of …
- Recoverable but Not Stationary:Local Linear Structures in Weights and Activations
Irina Piontkovskaia, Sergey Nikolenko · 10. Juni 2026
Task vectors, LoRA, activation steering, and random search around pretrained weights all suggest that learned behaviour can be controlled by linear directions. We ask which linear structures actually exist and on what scale. In a synthetic multitask transformer and LoRA adapters on DistilGPT-2 / GPT…
- Limitations of Learning Tanh Neural Networks with Finite Precision
Philipp Grohs, Mat\v{e}j Tr\"odler · 10. Juni 2026
We investigate limitations of learning $\tanh$ neural networks from point evaluations under finite-precision computations and $L^p$ accuracy guarantees, building on Berner, Grohs, and Voigtl\"ander (2023). Our approach is based on a novel construction of sharply localized bump functions via iterated…
- Unifying Local Communications and Local Updates for LLM Pretraining
Pietro Cagnasso, Eugene Belilovsky, Edouard Oyallon · 10. Juni 2026
Communication-efficient pre-training of LLMs is increasingly important as training draws on compute distributed across clusters, data centers, and lower-bandwidth links. Many practical methods reduce communication frequency but still rely on synchronous All-Reduce operations that maintain identical …
- Accelerating SAV-based optimization via randomized low-rank Hessian approximation
Ryo Sagawa, Daisuke Furihata, Yuto Miyatake · 10. Juni 2026
We propose a new optimization method, the Nystr\"om-enhanced relaxed scalar auxiliary variable method (N-RSAV), which incorporates curvature information into the RSAV framework to accelerate convergence while preserving an unconditional modified energy dissipation law. Existing RSAV-based methods re…
- Learning-Guided Integration Contours Construction for Fast Large-Scale Generalized Eigensolvers
Yeqiu Chen, Ziyan Liu, Hong Wang, Lei Liu · 10. Juni 2026
Solving large-scale Generalized Eigenvalue Problems (GEPs) is a fundamental yet computationally prohibitive task in science and engineering. As a promising direction, contour integral (CI) methods offer an efficient and parallelizable framework. However, their performance is critically dependent on …
