Physical Sciences › Computer Science › Computational Theory and Mathematics
Numerical Methods and Algorithms
88 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Últimos artículos
- Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2
Yohan Chatelain (Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada), Pablo de Oliveira Castro (Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France) · 2 de octubre de 2026
Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments …
- The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization
Chao Li, Shigeng Wang, Anbang Yao · 2 de octubre de 2026
Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization …
- ShatterQuant: Breaking Uniform Precision with Block-Wise Mixed-Precision on a Systolic Transformer Hardware Accelerator
Mikolaj Walczak, Edward Humes, Chao Fang, Marian Verhelst, Tinoosh Mohsenin · 2 de octubre de 2026
Due to limited support for intra-tensor heterogeneous precision in conventional accelerators, neural network quantization remains largely restricted to per-tensor precision assignment. We present ShatterQuant, a hardware-software co-designed framework enabling mixed-precision quantization within eac…
- ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights
Jonathan Mei, Sang Hyub Kim, Oliver Knitter, Chi Chen, Martin Roetteler · 1 de octubre de 2026
We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant by replacing each its diagonal reconstruction geometry with a tractable dense curvature metric, using a general paradigm popularized by the Shampoo optimizer. For each linear weight, ShamAN-Q fits a Kronecker…
- STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei · 30 de septiembre de 2026
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors prop…
- ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs
Mehdi Makni, Ryan Lucas, Rahul Mazumder · 30 de septiembre de 2026
Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free …
- Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
I Kennedy, T Kennedy · 29 de septiembre de 2026
A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96…
- Zero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet Validation
Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i} · 29 de septiembre de 2026
Contemporary neural inference architectures rely on dense floating-point weight matrices stored in high-bandwidth memory (VRAM), incurring severe memory-wall bottlenecks and preventing native execution inside deterministic virtual machines like the Ethereum Virtual Machine (EVM). Verifying terminati…
- JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
Kaicheng Yang, Kaisen Yang, Chunyu Liu, Xianglong Yan, Haotong Qin, Junyi Wu, Tianao Zhang, Xun Zhang, Shaoqiu Zhang, Youbang Sun, Yulun Zhang · 29 de septiembre de 2026
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization.…
- Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer
Haoran Jin, Kangqi Zhang, Jirong Yang, Barry Lyu, Qiuyi Ding, Ruijie Gao, Nathan Bleier · 29 de septiembre de 2026
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conv…
- Low-Bit Recurrent States in Hybrid Language Models
Hongren Chen, Jiayang He · 28 de septiembre de 2026
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-pre…
- Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
Nux Li · 28 de septiembre de 2026
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the r…
- Softmax Reparameterization for Output-Head Quantization
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King · 28 de septiembre de 2026
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from eve…
- G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou · 28 de septiembre de 2026
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise obj…
- Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters
Gautam Veldanda · 25 de septiembre de 2026
Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block (convolution, diagonal SSM and sparse attention mixed by a per-token route…
- RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models
David Poblaci\'on-Criado, Dario Garcia-Gasulla, Eduardo Quinones · 24 de septiembre de 2026
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. However, quantization affects different layer types in inconsistent ways, so identifying where accuracy lo…
- Predicting Quantization Price for Selecting PTQ Configurations Before Deployment
Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu · 24 de septiembre de 2026
Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstructi…
- Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity
Kasun Dewage, Marianna Pensky, Suranadi De Silva · 23 de septiembre de 2026
Many post-training quantization (PTQ) methods use layer-wise reconstruction, second-order proxy objectives, or activation-aware transformations to reduce quantization-induced error. Whether that error signal predicts the downstream functional impact of quantizing an individual attention projection h…
- Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error
Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie · 23 de septiembre de 2026
A Strassen-type algorithm has many realizations with the same exact product and multiplication count yet different fp8 error because basis changes reshape coefficient geometry, posing the question of which to run. No current account settles this: classical stability controls worst-case $\ell_1$ grow…
- Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement
Akihiro Yoshida, Yuma Ichikawa · 23 de septiembre de 2026
Mixed-precision weight quantization is commonly formulated as a Multiple-Choice Knapsack Problem (MCKP), yet existing solvers rely on scalar sensitivity proxies that collapse each weight matrix's Hessian into a single number and treat every module independently. We prove that even the optimal scalar…
- Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning
Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng · 23 de septiembre de 2026
Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without comp…
- Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference
Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li · 23 de septiembre de 2026
Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four f…
- Perplexity Cost Understates What Activation Quantisation Breaks
Anish Sathyanarayanan · 22 de septiembre de 2026
Activation quantisation is usually evaluated with an aggregate metric, perplexity, averaged over every token a model predicts. We ask whether that average identifies which computations a quantiser damages. Perplexity turns out to be a reliable aggregate signal: across 12 models from four families an…
- Towards Full Pipeline FP8 Reinforcement Learning for LLMs
Fanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong, Du Li, Chi Zhang, Haibin Lin, Shivaram Venkataraman · 22 de septiembre de 2026
Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused o…
- PRQuant: Permutation Residual Quantization for Low-Overhead Inference
Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang · 22 de septiembre de 2026
Accuracy of Low-bit quantization of linear layers is often dominated by a small number of outliers. Although existing methods, such as smoothing, rotation, or residual-based approaches, may mitigate this problem, they often introduce new accuracy bottlenecks to weights. Besides, most of these techni…
