Physical Sciences › Computer Science › Artificial Intelligence
Stochastic Gradient Optimization Techniques
1612 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Accelerated Gradient Methods with Biased Gradient Estimates: Risk Sensitivity, High-Probability Guarantees, and Large Deviation Bounds
Mert G\"urb\"uzbalaban, Yasa Syed, Necdet Serhat Aybat · 14 de enero de 2026
We study trade-offs between convergence rate and robustness to gradient errors in the context of first-order methods. Our focus is on generalized momentum methods (GMMs)--a broad class that includes Nesterov's accelerated gradient, heavy-ball, and gradient descent methods--for minimizing smooth stro…
- Convergence of gradient flow for learning convolutional neural networks
Jona-Maria Diederen, Holger Rauhut, Ulrich Terstiege · 14 de enero de 2026
Convolutional neural networks are widely used in imaging and image recognition. Learning such networks from training data leads to the minimization of a non-convex function. This makes the analysis of standard optimization methods such as variants of (stochastic) gradient descent challenging. In thi…
- LDLT L-Lipschitz Network Weight Parameterization Initialization
Marius F. R. Juston, Ramavarapu S. Sreenivas, Dustin Nottage, Ahmet Soylemezoglu · 14 de enero de 2026
We analyze initialization dynamics for LDLT-based $\mathcal{L}$-Lipschitz layers by deriving the exact marginal output variance when the underlying parameter matrix $W_0\in \mathbb{R}^{m\times n}$ is initialized with IID Gaussian entries $\mathcal{N}(0,\sigma^2)$. The Wishart distribution, $S=W_0W_0…
- NOVAK: Unified adaptive optimizer for deep neural networks
Sergii Kavun · 14 de enero de 2026
This work introduces NOVAK, a modular gradient-based optimization algorithm that integrates adaptive moment estimation, rectified learning-rate scheduling, decoupled weight regularization, multiple variants of Nesterov momentum, and lookahead synchronization into a unified, performance-oriented fram…
- Towards A Unified PAC-Bayesian Framework for Norm-based Generalization Bounds
Xinping Yi, Gaojie Jin, Xiaowei Huang, Shi Jin · 14 de enero de 2026
Understanding the generalization behavior of deep neural networks remains a fundamental challenge in modern statistical learning theory. Among existing approaches, PAC-Bayesian norm-based bounds have demonstrated particular promise due to their data-dependent nature and their ability to capture algo…
- Riemannian Zeroth-Order Gradient Estimation with Structure-Preserving Metrics for Geodesically Incomplete Manifolds
Shaocong Ma, Heng Huang · 14 de enero de 2026
In this paper, we study Riemannian zeroth-order optimization in settings where the underlying Riemannian metric $g$ is geodesically incomplete, and the goal is to approximate stationary points with respect to this incomplete metric. To address this challenge, we construct structure-preserving metric…
- The Impact of Anisotropic Covariance Structure on the Training Dynamics and Generalization Error of Linear Networks
Taishi Watanabe, Ryo Karakida, Jun-nosuke Teramae · 13 de enero de 2026
The success of deep neural networks largely depends on the statistical structure of the training data. While learning dynamics and generalization on isotropic data are well-established, the impact of pronounced anisotropy on these crucial aspects is not yet fully understood. We examine the impact of…
- Gradient descent for deep equilibrium single-index models
Sanjit Dandapanthula, Aaditya Ramdas · 13 de enero de 2026
Deep equilibrium models (DEQs) have recently emerged as a powerful paradigm for training infinitely deep weight-tied neural networks that achieve state of the art performance across many modern machine learning tasks. Despite their practical success, theoretically understanding the gradient descent …
- Effectively Leveraging Momentum Terms in Stochastic Line Search Frameworks for Fast Optimization of Finite-Sum Problems
Matteo Lapucci, Davide Pucci · 13 de enero de 2026
In this work, we address unconstrained finite-sum optimization problems, with particular focus on instances originating in large scale deep learning scenarios. Our main interest lies in the exploration of the relationship between recent line search approaches for stochastic optimization in the overp…
- From Sublinear to Linear: Fast Convergence in Deep Networks via Locally Polyak-Lojasiewicz Regions
Agnideep Aich, Ashit Baran Aich, Bruce Wade · 13 de enero de 2026
Gradient descent (GD) on deep neural network loss landscapes is non-convex, yet often converges far faster in practice than classical guarantees suggest. Prior work shows that within locally quasi-convex regions (LQCRs), GD converges to stationary points at sublinear rates, leaving the commonly obse…
- On a Gradient Approach to Chebyshev Center Problems with Applications to Function Learning
Abhinav Raghuvanshi, Mayank Baranwal, Debasish Chatterjee · 13 de enero de 2026
We introduce $\textsf{gradOL}$, the first gradient-based optimization framework for solving Chebyshev center problems, a fundamental challenge in optimal function learning and geometric optimization. $\textsf{gradOL}$ hinges on reformulating the semi-infinite problem as a finitary max-min optimizati…
- A Kernel-based Stochastic Approximation Framework for Nonlinear Operator Learning
Jia-Qi Yang, Lei Shi · 13 de enero de 2026
We develop a stochastic approximation framework for learning nonlinear operators between infinite-dimensional spaces utilizing general Mercer operator-valued kernels. Our framework encompasses two key classes: (i) compact kernels, which admit discrete spectral decompositions, and (ii) diagonal kerne…
- A Fast and Effective Method for Euclidean Anticlustering: The Assignment-Based-Anticlustering Algorithm
Philipp Baumann, Olivier Goldschmidt, Dorit S. Hochbaum, Jason Yang · 13 de enero de 2026
The anticlustering problem is to partition a set of objects into K equal-sized anticlusters such that the sum of distances within anticlusters is maximized. The anticlustering problem is NP-hard. We focus on anticlustering in Euclidean spaces, where the input data is tabular and each object is repre…
- Tight Analysis of Decentralized SGD: A Markov Chain Perspective
Lucas Versini, Paul Mangold, Aymeric Dieuleveut · 13 de enero de 2026
We propose a novel analysis of the Decentralized Stochastic Gradient Descent (DSGD) algorithm with constant step size, interpreting the iterates of the algorithm as a Markov chain. We show that DSGD converges to a stationary distribution, with its bias, to first order, decomposable into two componen…
- Deriving Decoder-Free Sparse Autoencoders from First Principles
Alan Oursland · 13 de enero de 2026
Gradient descent on log-sum-exp (LSE) objectives performs implicit expectation--maximization (EM): the gradient with respect to each component output equals its responsibility. The same theory predicts collapse without volume control analogous to the log-determinant in Gaussian mixture models. We in…
- Implicit bias as a Gauge correction: Theory and Inverse Design
Nicola Aladrah, Emanuele Ballarin, Matteo Biagetti, Alessio Ansuini, Alberto d'Onofrio, Fabio Anselmi · 13 de enero de 2026
A central problem in machine learning theory is to characterize how learning dynamics select particular solutions among the many compatible with the training objective, a phenomenon, called implicit bias, which remains only partially characterized. In the present work, we identify a general mechanis…
- Asymptotics of Non-Convex Generalized Linear Models in High-Dimensions: A proof of the replica formula
Matteo Vilucchio, Yatin Dandi, Mat\'eo Pirio Rossignol, Cedric Gerbelot, Florent Krzakala · 13 de enero de 2026
The analytic characterization of the high-dimensional behavior of optimization for Generalized Linear Models (GLMs) with Gaussian data has been a central focus in statistics and probability in recent years. While convex cases, such as the LASSO, ridge regression, and logistic regression, have been e…
- The Hessian of tall-skinny networks is easy to invert
Ali Rahimi · 13 de enero de 2026
We describe an exact algorithm for solving linear systems $Hx=b$ where $H$ is the Hessian of a deep net. The method computes Hessian-inverse-vector products without storing the Hessian or its inverse in time and storage that scale linearly in the number of layers. Compared to the naive approach of f…
- Wide Neural Networks as a Baseline for the Computational No-Coincidence Conjecture
John Dunbar, Scott Aaronson · 13 de enero de 2026
We establish that randomly initialized neural networks, with large width and a natural choice of hyperparameters, have nearly independent outputs exactly when their activation function is nonlinear with zero mean under the Gaussian measure: $\mathbb{E}_{z \sim \mathcal{N}(0,1)}[\sigma(z)]=0$. For ex…
- Tree-Preconditioned Differentiable Optimization and Axioms as Layers
Yuexin Liao · 13 de enero de 2026
This paper introduces a differentiable framework that embeds the axiomatic structure of Random Utility Models (RUM) directly into deep neural networks. Although projecting empirical choice data onto the RUM polytope is NP-hard in general, we uncover an isomorphism between RUM consistency and flow co…
- Accumulation of Sub-Sampling Matrices with Applications to Statistical Computation
Yifan Chen, Yun Yang · 13 de enero de 2026
With appropriately chosen sampling probabilities, sampling-based random projection can be used to implement large-scale statistical methods, substantially reducing computational cost while maintaining low statistical error. However, computing optimal sampling probabilities is often itself expensive,…
- mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
Yongyi Yang, Jianyang Gao · 12 de enero de 2026
Hyper-Connections (HC) generalizes residual connections by introducing dynamic residual matrices that mix information across multiple residual streams, accelerating convergence in deep neural networks. However, unconstrained residual matrices can compromise training stability. To address this, DeepS…
- mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
Yongyi Yang, Jianyang Gao · 12 de enero de 2026
Hyper-Connections (HC) generalizes residual connections by introducing dynamic residual matrices that mix information across multiple residual streams, accelerating convergence in deep neural networks. However, unconstrained residual matrices can compromise training stability. To address this, DeepS…
- What Functions Does XGBoost Learn?
Dohyeong Ki, Adityanand Guntuboyina · 12 de enero de 2026
This paper establishes a rigorous theoretical foundation for the function class implicitly learned by XGBoost, bridging the gap between its empirical success and our theoretical understanding. We introduce an infinite-dimensional function class $\mathcal{F}^{d, s}_{\infty-\text{ST}}$ that extends fi…
- Communication-Efficient Stochastic Distributed Learning
Xiaoxing Ren, Nicola Bastianello, Karl H. Johansson, Thomas Parisini · 12 de enero de 2026
We address distributed learning problems, both nonconvex and convex, over undirected networks. In particular, we design a novel algorithm based on the distributed Alternating Direction Method of Multipliers (ADMM) to address the challenges of high communication costs, and large datasets. Our design …
