Physical Sciences › Computer Science › Computer Networks and Communications
Caching and Content Delivery
139 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China52 % · 23 Artikel
- Vereinigte Staaten39 % · 17 Artikel
- Sonderverwaltungsregion Hongkong9,1 % · 4 Artikel
- Vereinigtes Königreich6,8 % · 3 Artikel
- Vereinigte Arabische Emirate4,5 % · 2 Artikel
- Indien4,5 % · 2 Artikel
- Südkorea4,5 % · 2 Artikel
- Deutschland4,5 % · 2 Artikel
Über 44 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 21 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou, Zihan Ding, Xingang Pan · 1. Oktober 2026
Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows with the generated history. Existing compression strategies discard history using fixed windows or select tokens through local attention and similarity signals, without directly measuring whe…
- ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation
Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong, Haocheng Sun, Keyu Chen, Siyu An, Ruizhi Qiao, Xing Sun, Wenjie Li, Xiaoyu Shen · 1. Oktober 2026
Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and r…
- Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure
Matias Parij, Pawan Paudel, Tate Berenbaum, Muthaiah Venkatachalam · 1. Oktober 2026
We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh us…
- The Golden Path Hypothesis: Reusable Schedules in Diffusion Caching
Dong Wang, Wenwu Tang, Francesco Corti, Yun Cheng, Lothar Thiele, Olga Saukh · 1. Oktober 2026
Diffusion caching accelerates generation by replacing transformer computation with cached or predicted features at selected denoising steps. We introduce the Golden Path Hypothesis (GPH): under fixed inference conditions, prompt-independent cache schedules can achieve final-output quality comparable…
- Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
Xianpeng Shang, Canbin Huang, Jiang Li, Tian Lan, Qianyi Cai, Xiaojun Quan, Xiangdong Su · 30. September 2026
The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval c…
- Reciprocal Guidance: Orchestrating Draft and Verify Budgets for Advancing the Diffusion-AR Self-Speculation Frontier
Linye Wei, Shutian Zheng, Haoyu Zeng, Meng Li · 30. September 2026
Diffusion drafting with autoregressive (AR) verification has emerged as a promising paradigm for efficient speculative decoding. Recent self-speculation models, represented by Nemotron-Labs-Diffusion, further simplify the speculative pipeline by unifying drafting and verification within a shared bac…
- Positions Are Not Facts: The Mismatch Between KV Caches and Memory
Changhai Zhou, Yuhua Zhou, Shiyang Zhang, Jun Gao, Zhen Li, Hua Wu, Hanchao Yu, Haifeng Wang · 30. September 2026
When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only replaced values, and deleting old text and …
- KV-Lingo: Learning KV-Cache Translators with Distillation
Val\'erie Castin, Keitaro Sakamoto, Anastasiia Filippova, Jo\~ao Monteiro, Marco Cuturi, Pierre Ablin · 30. September 2026
Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context: although the context has already been p…
- Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning
Zihan Zhang, Shuangjie Yao, Zesen Liu, Zhixiang Zhang, Wai Ip Lai, Dung Hiu Hilton Yeung, Chun Kit Zhang, Fuchen Ma, Yuanyuan Yuan, Yu Jiang, Dongdong She · 30. September 2026
Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response…
- KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
Joao Monteiro, Louis B\'ethune, Anastasiia Filippova, Sonia Laguna, David Grangier, Marco Cuturi · 30. September 2026
As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by dis…
- CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents
Junjie Yao, Zhangchen Zhou, Zhi-Qin John Xu · 30. September 2026
For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editin…
- ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction
Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang, Yanjia Li, Adnan Aziz, Chunqiang Tang, Ang Li · 30. September 2026
Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve stro…
- BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
Suyoung Kim, Jahyun Koo, Hyeonjin Kim, Inhyeok Bang, Seunghyun Lee, Hyunjae Oh, Baeseong Park, Dongsoo Lee · 29. September 2026
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatc…
- Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
Timothy DeLise, Seth Cromelin · 29. September 2026
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens …
- RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse
Ruoling Qi, Yirui Liu, Xuaner Wu, Yuxin Jin, Jian Chen, Jiayu Qin, Yin Chen, Jiawei Shao · 29. September 2026
Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk inter…
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
Yirui Liu, Ruoling Qi, Xuaner Wu, Yuxin Jin, Jian Chen, Penghang Liu, Yafei Huang, Jiawei Shao, Xuelong Li · 29. September 2026
Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbi…
- Memory as a cache: Exact context reuse and deletion by construction
Shengyao Wang, Jiang Liu · 29. September 2026
The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose…
- Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Sai Praneeth Karimireddy, Murali Annavaram · 29. September 2026
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free…
- Receiver-Conditioned Latent Communication gives 94% CacheBack
Maximillian Rossi, Prajwal Raghunath, Haoqing Xuan, Yusen Zhang, Eugene Wu · 29. September 2026
Multi-agent systems distribute large contexts across agents that communicate to solve a task. Text messages are compact but require decoding and may omit evidence the receiving agent needs. Recent latent communication instead transfers KV caches. This avoids text generation and can improve accuracy …
- Evaluating the accuracy of KV cache reuse techniques
Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona · 28. September 2026
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse,…
- CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters
Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen · 28. September 2026
Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores f…
- When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang · 25. September 2026
Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrain…
- Visual Representation and History Modeling for Navigation World Models
Guangfu Guo, Xiaoqian Lu, Rui Liu, Yutong Chen, Kunpeng Liu, Long Cheng · 25. September 2026
Navigation World Models (NWMs) predict action-conditioned visual futures for planning. Two practical challenges are central to their design: selecting a suitable visual representation and efficiently modeling observation history for repeated candidate queries. Standard Global-Softmax attention provi…
- Shared Global KV with Layer-Specific Local History
Xinglang Xian · 24. September 2026
Decoder-only Transformer language models cache keys and values (KV) to reuse past computation during generation. Sharing KV across layers saves storage but reduces the diversity of representations available across depth. We study what local memory should retain alongside shared global KV, separating…
- Marginally Correct Tool Caches Can Reverse Group-Normalized Policy Updates
Shivam Gupta · 24. September 2026
Tool-result caching reduces repeated execution in agent training, but also couples rollout randomness. We study a two-action model in which independent and shared execution preserve every rollout's conditional reward distribution. Despite this marginal agreement, sharing one stochastic result per gr…
Weitere Unterthemen aus Rechnernetze und Kommunikation
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Software System Performance and Reliability395 Papiere / 12 Monate+400 %
- Constraint Satisfaction and Optimization254 Papiere / 12 Monate+220 %
- Software-Defined Networks and 5G205 Papiere / 12 Monate+400 %
- Network Security and Intrusion Detection186 Papiere / 12 Monate+260 %
- IoT and Edge/Fog Computing150 Papiere / 12 Monate+175 %
- Advanced Database Systems and Queries130 Papiere / 12 Monate+220 %
