Physical Sciences › Engineering › Mechanics of Materials
Muon and positron interactions and applications
52 papiers indexés
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel — 12 derniers mois
Derniers papiers
- Muon on the Stiefel Manifold Admits an Exact Closed-Form Update
Mikhail Solonko, Molozhavenko Alexander, Maxim Rakhuba · 7 août 2026
We study Muon, a recently proposed matrix-aware optimization method, in the context of the Stiefel manifold. This manifold consists of matrices with orthonormal columns and is ubiquitous in machine learning and scientific computing. Existing extensions of Muon to this manifold rely on heuristic, app…
- MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning
Tongle Wu, Huanyu Dong, Ying Sun, Ziye Ma · 6 août 2026
Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon mitigates gradient anisotropy, it does not explicitly account for the curvature geometry of the loss landscape and may there…
- On MUON optimization: From non-convergence to an error analysis with Polar Express and the Newton-Schulz polynomial from implementations
Thang Do, Steffen Dereich, Arnulf Jentzen · 6 août 2026
Stochastic gradient descent (SGD) optimization methods are the standard instruments for the training of deep neural networks (DNNs). In many relevant artificial intelligence (AI) systems - such as popular large language models (LLMs)-not the standard SGD scheme is used as the optimization method but…
- Muon Meets Mamba: Spectral Optimization for State Space Models
Arslan Battalov, Karim Kramin, Alexander Markotenko, Sofia Sinitsina · 5 août 2026
Muon is a recent optimizer that orthogonalizes the update to each weight matrix with a Newton-Schulz iteration, which performs steepest descent under the spectral norm. Almost all the evidence for it comes from Transformer models, and its behavior on state-space models is largely unreported. We comp…
- CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
Chuyan Chen, Peng Sun, Kun Yuan · 4 août 2026
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct applicati…
- Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback
Maria Smirnova, Alexey Kravatskiy · 3 août 2026
SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget. It outperforms SignSGD in practice, yet it can ascend even on a linear function. Signing the gradien…
- Sharpness-Aware Minimization and Muon: Robustness under the Spectral Norm
Wenzhi Zhong, Edward Milsom, Michael Murray · 29 juillet 2026
Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear pers…
- Muse: Representation Geometry of Muon Beyond Normalized Momentum
Da Chang, Qiankun Shi, Lvgang Zhang, Di He, Yaoshuai Ma, Ganzhao Yuan, Yongxiang Liu · 17 juillet 2026
Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization. We study this representation choice as a form of optimizer geometry and introduce {\method}, a family of Muon-style optimizers that shares…
- Reassessing Muon for Matrix Factorization
Ali Parviz, Gal Mishne, Alex Cloninger · 16 juillet 2026
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretic…
- Muon learns balanced solutions in matrix factorization without slow saddle-to-saddle dynamics
Mark Rhee, Jamie Simon, Dhruva Karkada · 30 juin 2026
Matrix factorization (i.e., problems of the form $\min_{\mathbf{P},\mathbf{Q}} \|\mathbf{M}^\star - \mathbf{P}^\top\mathbf{Q}\|_\mathrm{F}^2$) is a minimal learning problem that exhibits both nonlinear parameter dynamics and representation learning. In this setting, we study how parameter trajectori…
- Aurora: A Leverage-Aware Spectral Optimizer
Alec Dewulf, Dhruv Pai, Li Yang, Ashley Zhang, Ben Keigwin · 29 juin 2026
We show that for tall matrix parameters, like projection matrices in the MLP layers, the Muon update can have row norms that are arbitrarily non-uniform. This can lead to a self-reinforcing feedback loop whereby neurons receive persistently small updates and eventually do not contribute meaningfully…
- Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization
Ziyuan Tang, Tianshi Xu, Yousef Saad, Yuanzhe Xi · 26 juin 2026
Muon-type optimizers construct update directions for dense neural-network weights by applying a finite Newton-Schulz map to momentum-gradient matrices. For an $H \times W$ matrix, with $r=\min\{H,W\}$ and $s=\max\{H,W\}$, $K$ steps of the full-matrix Newton-Schulz update require $O(r^2 s K)$ work an…
- DMuon: Efficient Distributed Muon Training with Near-Adam Overhead
Vincent Chen, Starrick Liu, Regis Cheng, Dance Yang, Shalfun Li, Ryan Yu, Lucy Liang, Hang Su, Roy Gan, Hao Wang, Qian Wang · 26 juin 2026
Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads. The matrix-aware updates offer a compelling alternative to conventional element-wise optimization, particularly as model architectures …
- Towards Understanding the Power and Limits of the Muon Optimizer: A River-Valley Perspective
Tianqi Shen, Jinji Yang, Runze Shi, Jianhao Ma, Jiaye Teng, Ziye Ma · 23 juin 2026
Recently, Muon has gained substantial attention as an appealing alternative to Adam-like optimizers, with many works highlighting its advantages through spectral normalization and improved conditioning. Yet this positive theoretical narrative contrasts with its empirical performance in large languag…
- CacheMuon: Using Temporal Preconditioning To Approximate Polar Factor
Bishnu Dev (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Sushil Bohara (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Martin Tak\'a\v{c} (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Samuel Horv\'ath (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE) · 16 juin 2026
Muon is an optimizer that computes updates using the polar factor of the momentum matrix and has shown strong empirical performance across a range of training settings. A key component of Muon is the Newton-Schulz iteration used to compute this polar factor. Although this avoids the cost of an exact…
- The Spectral Dynamics and Noise Geometry of Muon
Pierfrancesco Beneventano, Mahmoud Abdelmoneum, Tomaso Poggio · 9 juin 2026
Muon replaces a matrix gradient $G=U\Sigma V^\top$ by its polar factor $UV^\top$. This keeps the singular directions selected by the gradient, but makes the update spectrum flat. We study the optimization bias created by this operation. Under explicit alignment assumptions, we prove that the polar u…
- Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering
Xianliang Li, Zihan Zhang, Weiyang Liu, Han Bao · 3 juin 2026
Muon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains unclear. Existing analyses of Muon either remove momentum to study spectral updates in isolation, or retain momentum without explaining why it improves e…
- How Much Orthogonalization Does Muon Need?
Hua Huang · 2 juin 2026
Muon optimizers improve neural-network training by replacing ill-conditioned momentum updates with approximately semi-orthogonal updates. This motivates a practical question: how much orthogonalization does Muon actually require? We study this question using a relaxed cubic Newton--Schulz schedule d…
- LiMuon: Light and Fast Muon Optimizer for Large Models
Feihu Huang, Yuning Luo, Songcan Chen · 1 juin 2026
Large models recently are widely applied in machine learning, so efficient training of large models has received widespread attention. More recently, the useful Muon optimizer is specifically designed for matrix-structured parameters of large models. Although some works have begun to study the Muon …
- MuCon: Clipped Muon Updates for LLM Training
Albert Yi · 27 mai 2026
Muon-style optimizers take a matrix-valued momentum or preconditioned update $B = U \operatorname{diag}(\sigma_1,\ldots,\sigma_r) V^\top$ and replace it with its canonical partial polar factor $\operatorname{Pol}(B) = U V^\top$. This maps every nonzero singular value to one. MuCon is the clipped-Muo…
- MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
Jiacheng Li, Jianchao Tan, Hongtao Xu, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai · 27 mai 2026
The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to produce geometry-aware updates. However, like all first-order methods, Muon can become trapped in sharp local minima. In this work, we present MONA, an o…
- Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer
Aratrika Mustafi, Soumya Mukherjee, Bharath K. Sriperumbudur · 25 mai 2026
We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dua…
- MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
Feihu Huang, Yuning Luo, Songcan Chen · 20 mai 2026
Matrix-structured parameters frequently appear in many artificial intelligence models such as large language models. More recently, an efficient Muon optimizer is designed for matrix parameters of large-scale models, and shows markedly faster convergence than the vector-wise algorithms. Although som…
- Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization
Yury Demidovich, Abhishek Chakraborty, Grigory Malinovsky, Angelia Nedi\'c, Peter Richt\'arik · 20 mai 2026
Muon and related normalized optimizers decouple the choice of update direction from the choice of step scale, but their practical performance remains sensitive to the scale of the normalized step. We study adaptive scaling rules for Muon in general norm geometries and develop three complementary alg…
- AMO: Adaptive Muon Orthogonalization
Xinlin Zhuang, Panyi Ouyang, Yichen Li, Jiangming Shi, Yizhang Chen, Shuman Liu, Ying Qian, Weiyang Liu, Haibo Zhang, Imran Razzak · 19 mai 2026
Muon has recently emerged as a competitive alternative to AdamW for large-scale pre-training, with orthogonalization via Newton-Schulz (NS) iterations as its core operation. Existing Muon variants apply a uniform NS schedule to all parameter matrices, overlooking possible differences in orthogonaliz…
