Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1 612 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Stochastic Gradient Descent for Nonparametric Additive Regression
Xin Chen, Jason M. Klusowski · 1 janvier 2026
This paper introduces an iterative algorithm for training nonparametric additive models that enjoys favorable memory storage and computational requirements. The algorithm can be viewed as the functional counterpart of stochastic gradient descent, applied to the coefficients of a truncated basis expa…
- Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
Zijian Liu · 1 janvier 2026
In Online Convex Optimization (OCO), when the stochastic gradient has a finite variance, many algorithms provably work and guarantee a sublinear regret. However, limited results are known if the gradient estimate has a heavy tail, i.e., the stochastic gradient only admits a finite $\mathsf{p}$-th ce…
- Colorful Pinball: Density-Weighted Quantile Regression for Conditional Guarantee of Conformal Prediction
Qianyi Chen, Bo Li · 1 janvier 2026
While conformal prediction provides robust marginal coverage guarantees, achieving reliable conditional coverage for specific inputs remains challenging. Although exact distribution-free conditional coverage is impossible with finite samples, recent work has focused on improving the conditional cove…
- Basic Inequalities for First-Order Optimization with Applications to Statistical Risk Analysis
Seunghoon Paik, Kangjie Zhou, Matus Telgarsky, Ryan J. Tibshirani · 1 janvier 2026
We introduce \textit{basic inequalities} for first-order iterative optimization algorithms, forming a simple and versatile framework that connects implicit and explicit regularization. While related inequalities appear in the literature, we isolate and highlight a specific form and develop it as a w…
- Rethinking Dense Linear Transformations: Stagewise Pairwise Mixing (SPM) for Near-Linear Training in Neural Networks
Peter Farag · 1 janvier 2026
Dense linear layers are a dominant source of computational and parametric cost in modern machine learning models, despite their quadratic complexity and often being misaligned with the compositional structure of learned representations. We introduce Stagewise Pairwise Mixers (SPM), a structured line…
- Federated Multi-Task Clustering
S. Dai, G. Sun, F. Li, X. Tang, Q. Wang, Y. Cong · 30 décembre 2025
Spectral clustering has emerged as one of the most effective clustering algorithms due to its superior performance. However, most existing models are designed for centralized settings, rendering them inapplicable in modern decentralized environments. Moreover, current federated learning approaches o…
- Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods
Zijian Liu, Zhengyuan Zhou · 30 décembre 2025
In the past several years, the last-iterate convergence of the Stochastic Gradient Descent (SGD) algorithm has triggered people's interest due to its good performance in practice but lack of theoretical understanding. For Lipschitz convex functions, different works have established the optimal $O(\l…
- A Simple, Optimal and Efficient Algorithm for Online Exp-Concave Optimization
Yi-Han Wang, Peng Zhao, Zhi-Hua Zhou · 30 décembre 2025
Online eXp-concave Optimization (OXO) is a fundamental problem in online learning. The standard algorithm, Online Newton Step (ONS), balances statistical optimality and computational practicality, guaranteeing an optimal regret of $O(d \log T)$, where $d$ is the dimension and $T$ is the time horizon…
- On Fibonacci Ensembles: An Alternative Approach to Ensemble Learning Inspired by the Timeless Architecture of the Golden Ratio
Ernest Fokou\'e · 30 décembre 2025
Nature rarely reveals her secrets bluntly, yet in the Fibonacci sequence she grants us a glimpse of her quiet architecture of growth, harmony, and recursive stability \citep{Koshy2001Fibonacci, Livio2002GoldenRatio}. From spiral galaxies to the unfolding of leaves, this humble sequence reflects a un…
- Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
G\'erard Ben Arous, Murat A. Erdogdu, Nuri Mert Vural, Denny Wu · 30 décembre 2025
We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where the data is generated as $y \propto \sum_{j=1}^{r}\lambda_j \sigma\left(\langle \boldsymbol{\theta_j}, \boldsymbol{x}\rang…
- A first-order method for nonconvex-strongly-concave constrained minimax optimization
Zhaosong Lu, Sanyou Mei · 30 décembre 2025
In this paper we study a nonconvex-strongly-concave constrained minimax problem. Specifically, we propose a first-order augmented Lagrangian method for solving it, whose subproblems are nonconvex-strongly-concave unconstrained minimax problems and suitably solved by a first-order method developed in…
- Diffusion-based Decentralized Federated Multi-Task Representation Learning
Donghwa Kang, Shana Moothedath · 30 décembre 2025
Representation learning is a widely adopted framework for learning in data-scarce environments to obtain a feature extractor or representation from various different yet related tasks. Despite extensive research on representation learning, decentralized approaches remain relatively underexplored. Th…
- Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis
Zijian Liu · 30 décembre 2025
Optimization under heavy-tailed noise has become popular recently, since it better fits many modern machine learning tasks, as captured by empirical observations. Concretely, instead of a finite second moment on gradient noise, a bounded ${\frak p}$-th moment where ${\frak p}\in(1,2]$ has been recog…
- Beyond Centralization: Provable Communication Efficient Decentralized Multi-Task Learning
Donghwa Kang, Shana Moothedath · 30 décembre 2025
Representation learning is a widely adopted framework for learning in data-scarce environments, aiming to extract common features from related tasks. While centralized approaches have been extensively studied, decentralized methods remain largely underexplored. We study decentralized multi-task repr…
- Directly Constructing Low-Dimensional Solution Subspaces in Deep Neural Networks
Yusuf Kalyoncuoglu · 30 décembre 2025
While it is well-established that the weight matrices and feature manifolds of deep neural networks exhibit a low Intrinsic Dimension (ID), current state-of-the-art models still rely on massive high-dimensional widths. This redundancy is not required for representation, but is strictly necessary to …
- The Affine Divergence: Aligning Activation Updates Beyond Normalisation
George Bird · 30 décembre 2025
A systematic mismatch exists between mathematically ideal and effective activation updates during gradient descent. As intended, parameters update in their direction of steepest descent. However, activations are argued to constitute a more directly impactful quantity to prioritise in optimisation, a…
- Simultaneous Approximation of the Score Function and Its Derivatives by Deep Neural Networks
Konstantin Yakovlev, Nikita Puchkin · 30 décembre 2025
We present a theory for simultaneous approximation of the score function and its derivatives, enabling the handling of data distributions with low-dimensional structure and unbounded support. Our approximation error bounds match those in the literature while relying on assumptions that relax the usu…
- APO: Alpha-Divergence Preference Optimization
Wang Zixian · 30 décembre 2025
Two divergence regimes dominate modern alignment practice. Supervised fine-tuning and many distillation-style objectives implicitly minimize the forward KL divergence KL(q || pi_theta), yielding stable mode-covering updates but often under-exploiting high-reward modes. In contrast, PPO-style online …
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
Chiwun Yang · 29 décembre 2025
The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources. Yet, while empirically validated, its theoretical underpinnings remain poorly understood. This work formalizes the learning dynamics of transf…
- An Equivariance Toolbox for Learning Dynamics
Yongyi Yang, Liu Ziyin · 29 décembre 2025
Many theoretical results in deep learning can be traced to symmetry or equivariance of neural networks under parameter transformations. However, existing analyses are typically problem-specific and focus on first-order consequences such as conservation laws, while the implications for second-order s…
- Pruning as a Game: Equilibrium-Driven Sparsification of Neural Networks
Zubair Shah, Noaman Khan · 29 décembre 2025
Neural network pruning is widely used to reduce model size and computational cost. Yet, most existing methods treat sparsity as an externally imposed constraint, enforced through heuristic importance scores or training-time regularization. In this work, we propose a fundamentally different perspecti…
- Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
Pierre Abillama, Changwoo Lee, Juechu Dong, David Blaauw, Dennis Sylvester, Hun-Seok Kim · 25 décembre 2025
Recent advances in transformer-based foundation models have made them the default choice for many tasks, but their rapidly growing size makes fitting a full model on a single GPU increasingly difficult and their computational cost prohibitive. Block low-rank (BLR) compression techniques address this…
- Relu and softplus neural nets as zero-sum turn-based games
Stephane Gaubert, Yiannis Vlassopoulos · 24 décembre 2025
We show that the output of a ReLU neural network can be interpreted as the value of a zero-sum, turn-based, stopping game, which we call the ReLU net game. The game runs in the direction opposite to that of the network, and the input of the network serves as the terminal reward of the game. In fact,…
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
Yedi Zhang, Andrew Saxe, Peter E. Latham · 24 décembre 2025
Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across architectures, existing theoretical treatments lack a unifying framework. We present a theoretical framework that explai…
- Shallow Neural Networks Learn Low-Degree Spherical Polynomials with Learnable Channel Attention
Yingzhen Yang · 24 décembre 2025
We study the problem of learning a low-degree spherical polynomial of degree $\ell_0 = \Theta(1) \ge 1$ defined on the unit sphere in $\RR^d$ by training an over-parameterized two-layer neural network (NN) with channel attention in this paper. Our main result is the significantly improved sample com…
