Physical Sciences › Computer Science › Computer Networks and Communications
Caching and Content Delivery
139 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China52% · 23 papers
- United States39% · 17 papers
- Hong Kong SAR China9.1% · 4 papers
- United Kingdom6.8% · 3 papers
- United Arab Emirates4.5% · 2 papers
- India4.5% · 2 papers
- South Korea4.5% · 2 papers
- Germany4.5% · 2 papers
Across 44 papers on this subject with at least one lab located. 21 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention
Inbasekaran S · 5 October 2026
Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV bu…
- SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention
Zihan Teng, Jiayu Zhao, Wentao Ren, Minhao Fan, Tianrui Ma, Song Chen, Weichen Liu · 5 October 2026
Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compression strategies, token-wise methods reduce cached states but risk information loss through eviction or condensation, while feature-wise methods reduc…
- Exact Memory-Time Optimization for Prefix-Cached Language Model Serving
Shivam Gupta · 5 October 2026
Retaining language-model prefix states trades recomputation against storage time. Optimizing each cached block independently can overcount savings: a resident block is usable only when the required preceding prefix is also available. We introduce Prefix-Certificate Retention (PCR), an exact finite-t…
- SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching
Zhendong Mi, Pu Zhao, Ziyu Hu, Xiaodong Yu, Yanzhi Wang, Grace Li Zhang, Shaoyi Huang · 5 October 2026
Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlyi…
- Budgeted Cache Repair for Cross-Context KV-Cache Reuse
Haeyong Kang, Chang D. Yoo · 5 October 2026
Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A de…
- WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds
Utkarsh Ranjan · 5 October 2026
Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on f…
- Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection
Shuo Yang, Lihao Fang, Yi Zhang, Haixiang Wang, Xincheng Ye, Shufan Chen, Jipeng Guo, Youqing Wang · 5 October 2026
Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluatio…
- Tailoring the Quantization Space for 1-Bit KV Cache Compression
Minsoo Cheong, Donghyun Son, Sungjoo Yoo · 5 October 2026
The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ…
- AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning
Han Yu, Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Hejian Sang, Han Shi, Menglin Zhou, Xuanzhao Dong, Minzhou Huang, Rui Cai, Hao Wang, Alborz Geramifard · 5 October 2026
Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overl…
- CORE: COverage CAlibration and Evicted-Mass REdistribution for KV Cache
Shuxin Liu, Qing Liu, Yi Du, Ou Wu · 5 October 2026
Long-context decoding is increasingly constrained by key--value (KV) cache memory and bandwidth. Existing fixed-budget compression methods typically separate retention from compensation, while a retention ranking specifies neither discarded attention mass nor the direction of induced output error. W…
- KV$^2$: A Self-Refining KV Cache
Johannes Wesch, Danni Liu, Jan Niehues · 5 October 2026
The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are…
- DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou, Zihan Ding, Xingang Pan · 1 October 2026
Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows with the generated history. Existing compression strategies discard history using fixed windows or select tokens through local attention and similarity signals, without directly measuring whe…
- ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation
Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong, Haocheng Sun, Keyu Chen, Siyu An, Ruizhi Qiao, Xing Sun, Wenjie Li, Xiaoyu Shen · 1 October 2026
Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and r…
- Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure
Matias Parij, Pawan Paudel, Tate Berenbaum, Muthaiah Venkatachalam · 1 October 2026
We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh us…
- The Golden Path Hypothesis: Reusable Schedules in Diffusion Caching
Dong Wang, Wenwu Tang, Francesco Corti, Yun Cheng, Lothar Thiele, Olga Saukh · 1 October 2026
Diffusion caching accelerates generation by replacing transformer computation with cached or predicted features at selected denoising steps. We introduce the Golden Path Hypothesis (GPH): under fixed inference conditions, prompt-independent cache schedules can achieve final-output quality comparable…
- Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
Xianpeng Shang, Canbin Huang, Jiang Li, Tian Lan, Qianyi Cai, Xiaojun Quan, Xiangdong Su · 30 September 2026
The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval c…
- Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier
Linye Wei, Shutian Zheng, Haoyu Zeng, Meng Li · 30 September 2026
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared bac…
- Positions Are Not Facts: The Mismatch Between KV Caches and Memory
Changhai Zhou, Yuhua Zhou, Shiyang Zhang, Jun Gao, Zhen Li, Hua Wu, Hanchao Yu, Haifeng Wang · 30 September 2026
When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only replaced values, and deleting old text and …
- KV-Lingo: Learning KV-Cache Translators with Distillation
Val\'erie Castin, Keitaro Sakamoto, Anastasiia Filippova, Jo\~ao Monteiro, Marco Cuturi, Pierre Ablin · 30 September 2026
Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context: although the context has already been p…
- Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning
Zihan Zhang, Shuangjie Yao, Zesen Liu, Zhixiang Zhang, Wai Ip Lai, Dung Hiu Hilton Yeung, Chun Kit Zhang, Fuchen Ma, Yuanyuan Yuan, Yu Jiang, Dongdong She · 30 September 2026
Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response…
- KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi · 30 September 2026
As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by dis…
- CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents
Junjie Yao, Zhangchen Zhou, Zhi-Qin John Xu · 30 September 2026
For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editin…
- ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction
Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li, Adnan Aziz, Chunqiang Tang, Ang Li · 30 September 2026
Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve stro…
- BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
Suyoung Kim, Jahyun Koo, Hyeonjin Kim, Inhyeok Bang, Seunghyun Lee, Hyunjae Oh, Baeseong Park, Dongsoo Lee · 29 September 2026
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatc…
- Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
Timothy DeLise, Seth Cromelin · 29 September 2026
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens …
Other topics in Computer networks and communications
The topics the OpenAlex classification attaches to the same theme, most active first.
- Software System Performance and Reliability395 papers / 12 months+400%
- Constraint Satisfaction and Optimization254 papers / 12 months+220%
- Software-Defined Networks and 5G205 papers / 12 months+400%
- Network Security and Intrusion Detection186 papers / 12 months+260%
- IoT and Edge/Fog Computing150 papers / 12 months+175%
- Advanced Database Systems and Queries130 papers / 12 months+220%
