Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Advanced Data Compression Techniques
281 papers indexed
Advanced data compression techniques explore methods to reduce the size of artificial intelligence models and the data streams they process, without significantly altering their performance. This research focuses in particular on quantization, which involves representing numerical values with fewer bits, as well as approaches like vector quantization or cache compression, suited to modern architectures such as transformers. Strategies such as post-training quantization, matrix rotation optimization, or dynamic precision allocation aim to reconcile computational efficiency with the preservation of output quality.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China45% · 63 papers
- United States28% · 39 papers
- Hong Kong SAR China7.1% · 10 papers
- South Korea7.1% · 10 papers
- Canada6.4% · 9 papers
- United Kingdom5% · 7 papers
- Japan5% · 7 papers
- Russia4.3% · 6 papers
Across 141 papers on this subject with at least one lab located. 35 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- XOR-Trellis: Ultra-Low-Complexity Dequantization and Curvature-Aware Hadamard-Free LLM Quantization
Xiaofan Que, Nir Elkayam, Spandan Pyakurel, Shuokai Pan, Dibakar Gope · 2 October 2026
Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deployment, however, presents two challenges: reconstructing compressed weig…
- ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models
Lianjun Liu, Shipeng Li, You Huang, Weiqi Yan, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong · 2 October 2026
Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fi…
- ResARC: Residual-Aware AutoRegressive Coding for Ultra-Low Bitrate Image Compression
Qin Yan, Ruixiao Dong, Yutao Xie, Li Li, Ying Chen, Kai Li, Daowen Li, Houqiang Li · 1 October 2026
Progressive autoregressive image codecs provide an appealing paradigm for generative compression by quantizing continuous latents into discrete tokens, transmitting coarse-to-fine prefix tokens and generating the remaining suffix tokens at the decoder. However, their reconstruction quality is fundam…
- Rethinking Generative Image Compression at Extremely Low Bitrates
Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu · 1 October 2026
Generative image compression produces visually plausible reconstructions at low bitrates, yet their behavior as the rate approaches zero remains largely unexplored. When pushed below normal operating rates, representative codecs undergo semantic collapse: rather than gracefully losing source-specifi…
- Security-Enhanced Seed-Based Weight Quantization for Large Language Models
Qiuyu Ren, Sudipta Paria, Aritra Dasgupta, Swarup Bhunia · 1 October 2026
Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random representations but do not explicitly account for the non-uniform sensitivity o…
- Scaling Influence Functions in LLMs through Eigenbasis-Corrected One-Bit Gradient Projection
Jaeseung Heo, J Rosser, Dongwoo Kim · 30 September 2026
Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storin…
- TORQUE: Optimizing What (not) to Quantize Before and After Rotation
Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik · 30 September 2026
Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous quantization works that use random rotation…
- EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates
Hong Zhang, Zhongjie Duan, Yingda Chen · 29 September 2026
Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility req…
- Chameleon: Dynamic Format Adapter for Efficient Diffusion
Arnab Sanyal, Sandeep Chinchali · 29 September 2026
Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the $\mathit{number\ format}$ in advance and only tunes the scale, zero point, or per-layer bit-width. At a fixed bit-width the best f…
- One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion
Yulin Yuan, Ying Zhang, Xiangming Meng · 29 September 2026
Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fix…
- CoViST: Visual Token Compression via Composable States
Qi Zhang, Xiandong Meng, Ronggang Wang, Siwei Ma · 29 September 2026
Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. T…
- PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers
Yanlong Chen, Yining Chen, Song Zhang, Amirhossein Habibian, Yawei Li · 29 September 2026
Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The a…
- Training-Free Bottleneck Width Planning for Convolutional Autoencoders
Guannan Guo · 28 September 2026
Multiscale Spectral Rate-Distortion (MS-SRD) estimates the bottleneck channels required at user-supplied spatial cuts from training images and a normalized mean-squared error (NMSE) bound, without fitting a neural network. Its covariance-tail rule is exact for shared linear block-convolutional autoe…
- Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication
Qinglei Qi, Zhihe Liang, Fengzhan Jing, Shenao Zhu, Lei Zhang, Chenyang Zhang, Shuqing He, Jia Guo · 28 September 2026
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated r…
- Exact Minimax One-Bit Unbiased Compression: Heavy-Tail Necessity and Finite-Randomness Approximation
Tao Jiang, Minbo Gao, Shaowei Cai · 24 September 2026
A pointwise-unbiased one-bit compressor reconstructs every real input in expectation while transmitting one bit. For a scalar source $P$ with CDF $F$, mean $m$, and $\mathcal J(P)=\int_{\mathbb R}\sqrt{F(r)(1-F(r))}\,dr$, we prove that the infimum of the source-averaged reconstruction second moment …
- GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals
Bo Cui, Yaowen Zhang · 24 September 2026
Residual vector quantization (RVQ) turns physiological waveforms into compact token sequences, but conventional masked modeling treats every incorrect token as equally costly. We propose GeoRVQ, a coarse-to-fine masked token model whose objective reflects the local response of a frozen waveform deco…
- Lifelong Learning of Video Diffusion Models From a Single Video Stream
Jason Yoo, Yingchen He, Saeid Naderiparizi, Dylan Green, Gido M. van de Ven, Geoff Pleiss, Frank Wood · 23 September 2026
Video diffusion models can enable embodied agents to anticipate plausible futures from the recent past, but they are typically trained offline on curated datasets--a mismatch with the agents' learning setup at deployment: online, from a single video stream that sequentially outputs one frame at a ti…
- FuncCode: Compressing Kolmogorov--Arnold Networks in Function Space with Hardware-Aware Quantization
Kazi Ahmed Asif Fuad, Lizhong Chen · 23 September 2026
Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions, increasing flexibility but also parameter memory because each edge stores multiple coefficients, often together with a separate base branch. We introduce FuncCode, a basis-agnostic compression approac…
- GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
Baher Mohammad, Ammar Ali, Stamatios Lefkimmiatis · 23 September 2026
Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentiall…
- Damage Predicts Recovery: When Calibration Data Matters in Compressing Financial LLMs
Junyi Ye, Mengjia Yu, Debapriya Hazra, Guiling Wang · 23 September 2026
Post-training quantization and pruning rely on a small calibration corpus. Whether specialized domains such as finance require domain-matched calibration data remains unsettled. We argue that the answer depends on the task-level damage caused by compression rather than on domain mismatch. If compres…
- You've Seen Enough: Quality-Constrained Image Coding for Machines
Khoa Pham-Dinh, Sanaz Nami, Hamed Rezazadegan Tavakoli, Moncef Gabbouj, Farhad Pakdaman · 23 September 2026
Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validate the decisions. Inspired by just-notice…
- KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
Sihyeon Ha, Jaeho Lee, Yo-Seb Jeon · 22 September 2026
What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and qua…
- Neural Residual Modeling for Scientific Data Compression under Guaranteed Error Bounds
Surya Majumder, Liangji Zhu, Sanjay Ranka, Anand Rangarajan · 22 September 2026
Lossy compression of scientific simulation data increasingly relies on learned, latent-space architectures such as Residual Vector Quantization (RVQ), which iteratively quantize a base representation and its residuals to progressively reduce reconstruction error. While effective, RVQ performs this r…
- GPLQ: A General, Practical, and Lightning QAT Method for Vision Transformers
Guang Liang, Xinyao Liu, Jianxin Wu · 22 September 2026
Vision Transformers (ViTs) are essential in computer vision but are computationally intensive, too. Model quantization, particularly to low bit-widths like 4-bit, aims to alleviate this difficulty, yet existing Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) methods exhibit si…
- State-Change Learning for Prediction of Future Events in Endoscopic Videos
Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter, Nicolas Padoy · 18 September 2026
Surgical future prediction, driven by real-time AI analysis of surgical video, is critical for operating room safety and efficiency. It provides actionable insights into upcoming events, their timing, and risks-enabling better resource allocation, timely instrument readiness, and early warnings for …
Other topics in Computer vision and pattern recognition
The topics the OpenAlex classification attaches to the same theme, most active first.
- Multimodal Machine Learning Applications8,069 papers / 12 months+191%
- Generative Adversarial Networks and Image Synthesis4,992 papers / 12 months+39%
- Advanced Neural Network Applications2,354 papers / 12 months+48%
- Advanced Vision and Imaging841 papers / 12 months+78%
- Human Pose and Action Recognition836 papers / 12 months+457%
- Face recognition and analysis482 papers / 12 months+88%
