Physical Sciences › Computer Science › Computational Theory and Mathematics
Numerical Methods and Algorithms
88 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Latest papers
- 16-bit Precision of Convolutional Neural Networks on Microcontroller Units for 8-bit Costs
Rui Liu, Benjamin Paa{\ss}en · 5 October 2026
To deploy deep neural networks on edge hardware, highly efficient inference schemes are necessary that retain high accuracy. This work presents W16A16, a high precision (16-bit), fast speed, low energy quantization method. On a widely applied microcontroller architecture Armv7E-M, our proposed appro…
- VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models
Elad Dror Cohen, Ofir Gordon, Lior Dikstein, Idan Achituve, Hai Victor Habi · 5 October 2026
Microscaling (MX) formats are emerging as a hardware-supported approach to efficient training and inference. They combine low-precision elements with shared block scales, but their impact on vision models remains underexplored. We systematically investigate post-training MX quantization across visio…
- Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2
Yohan Chatelain (Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada), Pablo de Oliveira Castro (Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France) · 2 October 2026
Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments …
- The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization
Chao Li, Shigeng Wang, Anbang Yao · 2 October 2026
Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization …
- ShatterQuant: Breaking Uniform Precision with Block-Wise Mixed-Precision on a Systolic Transformer Hardware Accelerator
Mikolaj Walczak, Edward Humes, Chao Fang, Marian Verhelst, Tinoosh Mohsenin · 2 October 2026
Due to limited support for intra-tensor heterogeneous precision in conventional accelerators, neural network quantization remains largely restricted to per-tensor precision assignment. We present ShatterQuant, a hardware-software co-designed framework enabling mixed-precision quantization within eac…
- ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights
Jonathan Mei, Sang Hyub Kim, Oliver Knitter, Chi Chen, Martin Roetteler · 1 October 2026
We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant by replacing each its diagonal reconstruction geometry with a tractable dense curvature metric, using a general paradigm popularized by the Shampoo optimizer. For each linear weight, ShamAN-Q fits a Kronecker…
- STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei · 30 September 2026
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors prop…
- ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs
Mehdi Makni, Ryan Lucas, Rahul Mazumder · 30 September 2026
Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free …
- Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
I Kennedy, T Kennedy · 29 September 2026
A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96…
- Zero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet Validation
Volkan Da\u{g}l{\i}, Zerrin Da\u{g}l{\i}, Da\u{g}han Da\u{g}l{\i} · 29 September 2026
Contemporary neural inference architectures rely on dense floating-point weight matrices stored in high-bandwidth memory (VRAM), incurring severe memory-wall bottlenecks and preventing native execution inside deterministic virtual machines like the Ethereum Virtual Machine (EVM). Verifying terminati…
- JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
Kaicheng Yang, Kaisen Yang, Chunyu Liu, Xianglong Yan, Haotong Qin, Junyi Wu, Tianao Zhang, Xun Zhang, Shaoqiu Zhang, Youbang Sun, Yulun Zhang · 29 September 2026
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization.…
- Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer
Haoran Jin, Kangqi Zhang, Jirong Yang, Barry Lyu, Qiuyi Ding, Ruijie Gao, Nathan Bleier · 29 September 2026
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conv…
- Low-Bit Recurrent States in Hybrid Language Models
Hongren Chen, Jiayang He · 28 September 2026
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-pre…
- Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
Nux Li · 28 September 2026
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the r…
- Softmax Reparameterization for Output-Head Quantization
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King · 28 September 2026
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from eve…
- G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou · 28 September 2026
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise obj…
- Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters
Gautam Veldanda · 25 September 2026
Ternary (1.58-bit) weights are attractive for microcontroller-class language models, but the sub-1M-parameter regime rests mainly on isolated, single-seed comparisons. One prominent example reports that a routed ternary block (convolution, diagonal SSM and sparse attention mixed by a per-token route…
- RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models
David Poblaci\'on-Criado, Dario Garcia-Gasulla, Eduardo Quinones · 24 September 2026
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. However, quantization affects different layer types in inconsistent ways, so identifying where accuracy lo…
- Predicting Quantization Price for Selecting PTQ Configurations Before Deployment
Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu · 24 September 2026
Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstructi…
- Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity
Kasun Dewage, Marianna Pensky, Suranadi De Silva · 23 September 2026
Many post-training quantization (PTQ) methods use layer-wise reconstruction, second-order proxy objectives, or activation-aware transformations to reduce quantization-induced error. Whether that error signal predicts the downstream functional impact of quantizing an individual attention projection h…
- Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error
Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie · 23 September 2026
A Strassen-type algorithm has many realizations with the same exact product and multiplication count yet different fp8 error because basis changes reshape coefficient geometry, posing the question of which to run. No current account settles this: classical stability controls worst-case $\ell_1$ grow…
- Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement
Akihiro Yoshida, Yuma Ichikawa · 23 September 2026
Mixed-precision weight quantization is commonly formulated as a Multiple-Choice Knapsack Problem (MCKP), yet existing solvers rely on scalar sensitivity proxies that collapse each weight matrix's Hessian into a single number and treat every module independently. We prove that even the optimal scalar…
- Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning
Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng · 23 September 2026
Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without comp…
- Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference
Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li · 23 September 2026
Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four f…
- Perplexity Cost Understates What Activation Quantisation Breaks
Anish Sathyanarayanan · 22 September 2026
Activation quantisation is usually evaluated with an aggregate metric, perplexity, averaged over every token a model predicts. We ask whether that average identifies which computations a quantiser damages. Perplexity turns out to be a reliable aggregate signal: across 12 models from four families an…
