Physical Sciences › Computer Science › Computer Networks and Communications
Advanced Data Storage Technologies
87 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States55% · 23 papers
- China48% · 20 papers
- Hong Kong SAR China9.5% · 4 papers
- Singapore4.8% · 2 papers
- Japan4.8% · 2 papers
- India4.8% · 2 papers
- United Kingdom4.8% · 2 papers
- South Korea4.8% · 2 papers
Across 42 papers on this subject with at least one lab located. 19 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- KV-streams for Efficient Compaction in Agentic Reinforcement Learning
Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre C\^ot\'e, Laurent Charlin, Guillaume Lajoie · 30 September 2026
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling…
- Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation
Xiaoyu Wu, Weihang Guo, Yifei Wang, Xinze Feng, Lydia E. Kavraki, Zhiwei Steven Wu · 30 September 2026
Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge …
- Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong · 30 September 2026
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compressio…
- How Linear Attention Remembers
Kichang Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko · 29 September 2026
Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memor…
- What does FFN compression change downstream? Same-state causal restoration in diffusion language models
Shaurya Omar · 29 September 2026
Diffusion language models (DLMs) enable flexible, parallel generation, but their iterative denoising remains computationally expensive, motivating increasingly aggressive compression. Existing compression objectives largely measure how well compressed computation approximates the original locally, b…
- LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models
Xianglong Shi, Ruijie Yang, Sirui Zhao, Shukang Yin, Zihao Bian, Tinghao Yi, Enhong Chen · 29 September 2026
Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to a…
- LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound
Baixi Sun, Le Chen, Anjir Ahmed Chowdhury, Xiaolong Ma, Chih-Hsuan Yang, Mingze Xia, Syed Zawad, Sheng Di, Rajkumar Kettimuthu, Huihuo Zheng, Rajeev Thakur, Venkatram Vishwanath, Feng Yan · 29 September 2026
Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory …
- ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou · 28 September 2026
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task …
- HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents
Zihong He, Junxiao Shen, Chen Liang, Hai-Ning Liang · 28 September 2026
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-prompt embeddings provide a verifiable initia…
- PatchKV: Efficient KV Cache Recovery for Dynamically Edited LLM Contexts
Guotao Yang, Rui Guo, Siwei He, Sheng Chen, Yitao Hu, Keqiu Li · 23 September 2026
Long-running LLM agent workflows often revise interior context spans while retaining long suffixes. Although suffix tokens remain unchanged, altered causal histories and rotary positions prevent exact reuse of their offloaded key-value (KV) states. Full suffix recomputation wastes prefill work, whil…
- LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay
Simon P. Villani · 23 September 2026
Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent …
- RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents
Fanyu Zhao, Ruike Cao, Liang Dong, Fugen Yao, Jian Xu, Guanjun Jiang, Han Zhang, Yifei Zhao, Yinsheng Li · 22 September 2026
Long-running LLM agents require memory that persists and evolves across sessions. Text-based memory retrieves and reconstructs past interactions at every query, making long-horizon performance increasingly dependent on retrieval quality and contextual reasoning as histories grow. Parametric memory e…
- Layer-wise Curriculum Learning for Efficient LLM Compression
Donggeon Lee, Dooyeon Na, Seungmin Oh, Jongbin Ryu · 18 September 2026
In this paper, we introduce layer-wise curriculum learning for efficient LLM compression. The proposed method facilitates the knowledge transfer from the teacher model to the student model, utilizing a curriculum learning approach that begins with easier optimization tasks and progressively tackles …
- Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache
Frank Li · 15 September 2026
External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.3-Flash-NVFP4 quantized checkpoint with vLLM and LMCache under four-way tensor parallelism. A complete-hit recovery misma…
- Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
Sara Rizwan, Samaanah Abdus Salam · 9 September 2026
Long context language models now advertise windows of one million tokens, but two habits limit how much of that window is used. Attention heads with nothing useful to read still spend their budget on the first token, which is called the attention sink, and where a fact sits in the context changes wh…
- ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics
Chin Ting Hsu, Yu-Syuan Xu, Ling Zou, Hsien-Kai Kuo, Wen-Huang Cheng · 9 September 2026
Although multimodal Large Language Models (MLLMs) excel in diverse tasks, their scalability remains limited by the memory and computational overhead of KV cache storage. Recent KV cache eviction approaches incorporate a cosine similarity-based diversity metric with importance metrics to selectively …
- KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU
Di Chai, Leye Wang, Zeshen Su, Zhiguo Xia, Zhihang Yu · 7 September 2026
Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity and the model's native context window. Existing systems typically compact older context into summaries or retrieve it later as text, either losing fine-grained execution evidence or repeatedl…
- SGD-KV: Summarization Guided KV Cache Compression
Zeyu Liu, Woomin Song, Xuandi Fu, Sai Muralidhar Jayanthi, Vivek Govindan, Aram Galstyan, Sravan Babu Bodapati, Srikanth Ronanki · 4 September 2026
Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the distinct functional roles of different attention heads. We pr…
- InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing
Shuaiyi Li, Zhisong Zhang, Yang Deng, Chenlong Deng, Tianqing Fang, Hongming Zhang, Haitao Mi, Dong Yu, Wai Lam · 2 September 2026
Although existing model editing methods perform well in recalling exact edit facts, they often struggle in complex scenarios that require deeper semantic understanding rather than mere knowledge regurgitation. Leveraging the strong contextual reasoning abilities of large language models (LLMs), in-c…
- MemoryWalker: Stop Training Agents on Contexts They Never Saw
Zinco J, Xunjie Zhu, Shen Huang, Zhenyi Wang, Pengjun Xie, Jieping Ye · 2 September 2026
Production agent harnesses such as Claude Code and Qwen-Agent compress context during rollout, but training under compression creates a conditioning problem: every eviction branches the effective history, so the learning object is a tree rather than a sequence. Existing linearizations either retain …
- Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
Shmuel Berman, Jia Deng · 2 September 2026
Memory is widely viewed as an important unsolved problem for LLMs and VLMs, and current benchmarks typically evaluate it by testing accuracy over long text or video. However, accuracy alone misses properties that matter for real long-horizon tasks. We introduce ECCBench, a benchmark and evaluation p…
- Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer Evidence
Mohammadali Khodabandehlou, Bhaskar Krishnamachari · 1 September 2026
KV-cache compression reduces LLM inference memory by evicting context tokens, but when the evicted tokens contain answer-bearing evidence, the model may hallucinate instead of recognizing that the compressed context is insufficient. We address this failure from a behavioral perspective: to our knowl…
- Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation
Linze Wu, Xinrui Chen · 31 August 2026
Structured generation underpins large language model (LLM) agents that produce JSON, SQL, and function calls, where a single wrong field can cause the downstream action to fail. Constrained decoding already tracks parser transitions to enforce formal validity, and these transitions expose how genera…
- Fast Weight Attention for Continual Learning
Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao · 31 August 2026
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered her…
- TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy
Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Junyan Zhang, Xuming Hu · 28 August 2026
Long-context inference is bottlenecked by the memory footprint of the key-value (KV) cache, especially for small models under tight resource budgets. Existing KV cache eviction methods score tokens using the model's attention distribution or, in attention-free variants, each key's distance from a gl…
Other topics in Computer networks and communications
The topics the OpenAlex classification attaches to the same theme, most active first.
- Software System Performance and Reliability395 papers / 12 months+400%
- Constraint Satisfaction and Optimization254 papers / 12 months+220%
- Software-Defined Networks and 5G205 papers / 12 months+400%
- Network Security and Intrusion Detection186 papers / 12 months+260%
- IoT and Edge/Fog Computing150 papers / 12 months+175%
- Caching and Content Delivery139 papers / 12 months+1500%
