Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1612 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Bayesian Fine-tuning in Projected Subspaces
Viktar Dubovik, Patryk Marsza{\l}ek, Jacek Tabor, Tomasz Ku\'smierczyk · 11 de mayo de 2026
Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of large models by decomposing weight updates into low-rank matrices, significantly reducing storage and computational overhead. While effective, standard LoRA lacks mechanisms for uncertainty quantification, leading to overconfident…
- OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling
Yuxuan Lou, Yang You · 11 de mayo de 2026
Muon improves neural-network training by orthogonalizing matrix-valued updates, but it leaves each layer's update magnitude controlled mostly by a global learning rate. We introduce OrScale, a trust-ratio extension of Muon built on a simple rule: the denominator of a layer-wise ratio should measure …
- Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
Clarissa Lauditi, Cengiz Pehlevan, Blake Bordelon · 11 de mayo de 2026
We study the evolution of hidden-weight spectra in wide neural networks trained by (stochastic) gradient descent. We develop a two-level dynamical mean-field theory (DMFT) that jointly tracks bulk and outlier spectral dynamics for spiked ensembles whose spike directions remain statistically dependen…
- Sample Complexity of Stochastic Optimization with Integer Variables
Hongyu Cheng, Yinghao Zheng, Marco Molinaro, Amitabh Basu · 11 de mayo de 2026
We establish sample complexity results for stochastic optimization over the integers, especially with a view to understand the complexity with respect to the corresponding continuous optimization problem. We show that integer optimization can sometimes require strictly more samples and sometimes str…
- Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
Sayantan Choudhury, Xiaoran Cheng, Martin Tak\'a\v{c}, Sen Na, Mladen Kolar · 11 de mayo de 2026
Most first-order optimizers treat matrix-valued parameters as vectors, ignoring the intrinsic geometry of hidden-layer weights in neural networks. Muon addresses this mismatch by updating along the polar factor of a momentum matrix, but its theoretical understanding has lagged behind practice. In pa…
- Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers
Mana Sakai, Masaaki Imaizumi · 11 de mayo de 2026
Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled way to study this question. While existing norm-based bounds for Transformers remove the explicit polynomial dependence o…
- A Rod Flow Model for Adam at the Edge of Stability
Eric Regis, Sinho Chewi · 11 de mayo de 2026
Cohen et al. (arXiv:2207.14484) observed that adaptive gradient methods such as Adam operate at the edge of stability. While there has been significant work on continuous-time modeling of gradient descent at the edge of stability, extending these models to momentum methods remains underdeveloped. In…
- MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
Ionut-Vlad Modoranu, Mher Safaryan, Dan Alistarh · 11 de mayo de 2026
With the rise in scale for deep learning models to billions of parameters, the computational cost of fine-tuning remains a significant barrier to deployment. While Low-Rank Adaptation (LoRA) has become the standard for parameter-efficient fine-tuning, the need to set a predefined, static rank $r$ re…
- When Descent Is Too Stable: Event-Triggered Hamiltonian Learning to Optimize
Yi Wang, Chandrajit Bajaj · 11 de mayo de 2026
Fixed-budget nonconvex optimization can fail not because local descent is unstable, but because it is too stable: after reaching a nearby stationary point, an optimizer may spend the remaining evaluations refining an uninformative local minimum. We formulate this failure mode as a control problem ov…
- Solving Max-Cut to Global Optimality via Feasibility-Preserving Graph Neural Networks
Hao Chen, Chendi Qian, Christopher Morris, Andrea Lodi, Can Li · 11 de mayo de 2026
Exact solution of hard combinatorial optimization problems often relies on strong convex relaxations, but solving these relaxations repeatedly inside a branch-and-bound algorithm can be prohibitively expensive. Hence, we consider this challenge for Max-Cut, where branch and bound commonly uses semid…
- Adaptive Regularization for Sparsity Control in Bregman-Based Optimizers
Ahmad Aloradi, Tim Roith, Emanu\"el A. P. Habets, Daniel Tenbrinck · 11 de mayo de 2026
Sparse training reduces the memory and computational costs of deep neural networks. However, sparse optimization methods, e.g., those adding an $\ell_1$ penalty, often control sparsity only indirectly through a regularization parameter $\lambda$, whose mapping to the final sparsity rate is non-trivi…
- Accelerated Relax-and-Round for Concave Coverage Problems
Matthew Fahrbach, Mehraneh Liaee, Morteza Zadimoghaddam · 11 de mayo de 2026
We present an accelerated relax-and-round algorithm for concave coverage problems, which generalize the classic maximum coverage problem. Building on the relax-and-round framework of Barman et al. [STACS 2021], we propose two significant improvements. First, we replace the linear programming (LP) re…
- Locally Near Optimal Piecewise Linear Regression in High Dimensions via Difference of Max-Affine Functions
Haitham Kanj, Kiryung Lee · 11 de mayo de 2026
This paper presents a parametric solution to piecewise linear regression through the Adaptive Block Gradient Descent (ABGD) algorithm. The heart of the method is the parametrization of piecewise linear functions as the difference of max-affine (DoMA) functions. A non-asymptotic local convergence ana…
- Approximation Error Upper and Lower Bounds for H\"{o}lder Class with Transformers
Xin He, Yuling Jiao, Xiliang Lu, Jerry Zhijian Yang · 11 de mayo de 2026
We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for H\"{o}lder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and …
- Robust stochastic first order methods in heavy-tailed noise via medoid mini-batch gradient sampling
Manojlo Vukovic, Dusan Jakovetic · 11 de mayo de 2026
We consider a first order stochastic optimization framework where, at each iteration, $K$ independent identically distributed (i.i.d.) data point samples are drawn, based on which stochastic gradients can be queried. We allow gradient noise to be heavy-tailed, with possibly infinite variances. For t…
- Closed-Form Last Layer Optimization
Alexandre Galashov, Natha\"el Da Costa, Liyuan Xu, Philipp Hennig, Arthur Gretton · 11 de mayo de 2026
Neural networks are typically optimized with variants of stochastic gradient descent. Under a squared loss, however, the optimal solution to the linear last layer weights is known in closed-form. We propose to leverage this during optimization, treating the last layer as a function of the backbone p…
- Convex Optimization with Nested Evolving Feasible Sets
Karthick Krishna M., Haricharan Balasundaram, Rahul Vaze · 11 de mayo de 2026
Convex Optimization with Nested Evolving Feasible Sets (CONES)} is considered where the objective function $f$ remains fixed but the feasible region evolves over time as a nested sequence $S_1 \supseteq S_2 \supseteq \cdots \supseteq S_T$. The goal of an online algorithm is to simultaneously minimiz…
- Decentralized Time-Varying Optimization for Streaming Data via Temporal Weighting
Muhammad Faraz Ul Abrar, Nicol\`o Michelusi, Erik G. Larsson · 11 de mayo de 2026
Classical optimization theory largely focuses on fixed objective functions, whereas many modern learning systems operate in dynamic environments where data arrive sequentially and decisions must be updated continuously. In this work, we study optimization with streaming data over a distributed netwo…
- Convergent Stochastic Training of Attention and Understanding LoRA
Zhengkai Sun, Dibyakanti Kumar, Alejandro F Frangi, Anirbit Mukherjee, Mingfei Sun · 11 de mayo de 2026
Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achiev…
- Exploring the non-convexity in machine learning using quantum-inspired optimization
Kandula Eswara Sai Kumar, Parth Dhananjay Danve, Abhishek Chopra, Rut Lineswala · 11 de mayo de 2026
The escalating complexity of modern machine learning necessitates solving challenging non-convex optimization problems, particularly in high-dimensional regimes and scenarios contaminated by gross outliers. Traditional approaches, relying on convex relaxations or specialized local search heuristics,…
- FANoS-v2: Feedback-Controlled Momentum with Thermostat Damping for Lightweight Neural Optimization
Nalin Dhiman · 11 de mayo de 2026
\FANOS{} is a PyTorch optimizer that augments RMS-preconditioned momentum with a scalar feedback controller over update energy. The public reference implementation stores momentum in parameter-update units, applies a non-negative thermostat damping coefficient, supports diagonal, factored, and raw-g…
- PolarAdamW: Disentangling Spectral Control and Schur Gauge-Equivariance in Matrix Optimisation
Haozhou Zhang · 11 de mayo de 2026
Muon's matrix-level update couples two distinct effects: spectral control via a polar map, and equivariance under orthogonal changes of multiplicity-space basis (Schur gauge-equivariance). We separate them with PolarAdamW, a controlled hybrid that preserves Muon's polar spectral-norm control but bre…
- QuadNorm: Resolution-Robust Normalization for Neural Operators
Bum Jun Kim, Makoto Kawano, Yusuke Iwasawa, Yutaka Matsuo · 11 de mayo de 2026
Normalization layers in neural operators usually compute statistics by uniformly averaging discrete grid values, making the normalization itself discretization-dependent and thereby a source of transfer error across different resolutions or meshes. To enable discretization robustness, we introduce a…
- Greedy Alignment Principle for Optimizer Selection
Jaerin Lee, Kyoung Mu Lee · 8 de mayo de 2026
Recent works have shown that gradient-update alignment is a powerful signal for modulating optimizer updates, often leading to faster training. We promote this update-wise heuristic as a mathematically grounded principle for selecting and tuning optimizer hyperparameters. By treating gradients and u…
- MARBLE: Multi-Aspect Reward Balance for Diffusion RL
Canyu Zhao, Hao Chen, Yunze Tong, Yu Qiao, Jiacheng Li, Chunhua Shen · 8 de mayo de 2026
Reinforcement learning fine-tuning has become the dominant approach for aligning diffusion models with human preferences. However, assessing images is intrinsically a multi-dimensional task, and multiple evaluation criteria need to be optimized simultaneously. Existing practice deal with multiple re…
