Physical Sciences › Computer Science › Artificial Intelligence
Algorithms and Data Compression
82 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Neueste Paper
- Capability Scaling-Down Laws for LLM Compression
Xueqi Cheng, Liang Wu, Kelly Wan, Liangjie Hong, Yushun Dong · 5. Oktober 2026
LLM compression reduces inference costs and memory requirements, but selecting a method and configuration remains largely empirical because comparable resource reductions can produce different capability losses. We systematically investigate capability scaling-down laws for LLM compression across pr…
- Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
Hyunsik Kim, Youngmoon Jung · 2. Oktober 2026
Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character re…
- Counting and Min-Cost Encoding for Tokenization in Large Language Models
Shuming Shi, Xiang Zhang, Hao Yu, Wenbo Fei, Changjian Wang, Zhan Wang, Guoqing Pang, Guangye Yu, Quan Lu, Ning Jiang · 2. Oktober 2026
Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a to…
- Sequential Functional Structured Tucker Compression for Large Language Model Attentions
Jiangfeng Chen, Xinyu Wang, Tianshuo Yan, Hanwei Wu, Xiao-Wen Chang, Yang Zhang, Lei Ding · 2. Oktober 2026
Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts…
- How Divergence Becomes Decision Flips in Compressed Language Models
Beatriz Almeida Felicio · 2. Oktober 2026
Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compr…
- Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings
Zifeng Cheng, Jie Zheng, Zhiwei Jiang, Shuwen Wang, Fei Shen, Shiping Ge, Qing Gu · 1. Oktober 2026
Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long do…
- MDL-Calibrated Significance-Gain Pair Encoding: Replication-Aware Automatic Stopping for Subword Tokenization
Azam Nouri · 30. September 2026
Byte-Pair Encoding (BPE) constructs subword vocabularies through greedy pair merging, but conventional BPE requires the number of merges or target vocabulary size to be specified externally. Significance-Gain Pair Encoding (SG-BPE) replaces frequency-only selection with a statistical criterion based…
- FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
Bowen Yang, Jingbo Zhou, Qinghong Miao, Hua Wu · 30. September 2026
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat…
- PC-SubMax: Efficient Prompt Compression via Regularized Submodular Maximization
Ziyi Zhang, Shuang Cui, Haotian Zhang, Xiaoyu Wang · 30. September 2026
While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle'' phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, m…
- Aperture: Merge-Consistent Rotary States for Compressed Tokens
Yuhao Du, Shunian Chen · 30. September 2026
Token compression combines content from several positions, yet rotary position embeddings usually assign the merged token one coordinate. We ask what positional information must survive later merges. Aperture stores Fourier moments of the token's weighted support at the model's rotary frequencies. W…
- HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks
Thai Nguyen, Khang Tran, NhatHai Phan · 30. September 2026
Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages …
- The Price of Token Boundaries: Compression Certificates and Prediction
Yuhao Du, Shunian Chen · 30. September 2026
Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule…
- Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
Ritul Satish, Prasoon Sinha, Akiho Kawada, Neeraja J. Yadwadkar · 29. September 2026
As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to co…
- ROTE: Benchmarking Neural Memorization on Complexity-Controlled Symbolic Sequences
Xinye Chen, Stefan G\"uttel, Mohammad Mozaffari · 29. September 2026
We introduce ROTE (RollOut Testing of Exact memorization), a benchmarking protocol for evaluating symbolic memorization of neural architectures. We study memorization and the extension of symbolic rules in neural sequence models by using sequences whose complexity is regulated by Lempel--Ziv--Welch …
- Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodol\`a · 28. September 2026
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patte…
- Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models
Kaifeng Tan, Yudong Li, Linlin Shen · 24. September 2026
Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following…
- MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression
Ke Wan, Yifan Wang, Liheng Lai, Chen Chen · 24. September 2026
Likelihood-based context compression can account for cross-context redundancy through sequential scoring, but this makes compression outcomes sensitive to context order. We show that different permutations of the same context collection can produce markedly different evidence-retention outcomes unde…
- Distilling Sequential Computation in Transformer Language Models
Zixuan Lan, Jessica Yang, Yanhong Li, Karen Livescu, Jiawei Zhou · 24. September 2026
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or frequently occur as stable units, suggesting that their representations may be compressible. We introduce …
- CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers · 23. September 2026
Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improvin…
- Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure
Shuyang Xiang · 22. September 2026
Standard positional encodings represent position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate chann…
- An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents
Luzhuo Chen, Jiayu Shi · 22. September 2026
Context compression is widely proposed as a way to cut the token bill of LLM coding agents, and public benchmarks report that aggressive compression preserves task-solving quality. These two facts do not imply the third one commonly assumed: that compressing file reads saves money in a real multi-tu…
- TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding
Huapeng Zhou, Huayu Wang, Xinyu Wang · 22. September 2026
Speculative decoding accelerates language-model inference by letting a cheap drafter propose tokens that the target model verifies in parallel. Recent block drafters make drafting nearly free: a single backbone pass emits an entire block of draft tokens. Draft trees promise a further gain -- several…
- An Introduction to Compression-Based Machine Learning
John Hurwitz, Edward Raff, Charles K. Nicholas · 21. September 2026
Any lossless compression algorithm (like gzip) may be converted into a machine learning method, via either Normalized Compression Distance or the Minimum Description Length principle. Any auto-regressive model may be converted into a lossless compression method via entropy coding. This seemingly cir…
- Stringological sequence prediction III: layered ziplines and a tradeoff between efficiency and expressivity
Vanessa Kosoy · 18. September 2026
In previous papers, we began the study of sequence prediction algorithms adapted to stringological word complexity measures. In particular, we defined a complexity measure called Arithmetic Repetition Complexity (ARC) which admits a polynomial-time prediction algorithm with a mistake bound quasiline…
- Z-Loss Backward Geometry in Dense Output Heads and Sparse Routers
Bum Jun Kim · 16. September 2026
Z-loss has been widely applied to the logits of language-model output heads and sparse mixture-of-experts routers. Z-loss constrains the softmax log-normalizers of these output heads and routers, thereby limiting large-logit excursions, reducing finite-precision roundoff exposure, and avoiding train…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
