Physical Sciences › Computer Science › Computer Networks and Communications
Distributed systems and fault tolerance
47 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Latest papers
- Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation
Genliang Zhu, Chu Wang · 2 October 2026
Resource limits are becoming an authorization boundary for AI agents that delegate work across concurrent and failure-prone workers. Parent-child allocation constraints, affine objects, and distributed escrow do not by themselves prevent overspend when replies are lost, effects complete after timeou…
- Exact Record Omission in Delta Attention: A Transport Criterion, Its Cost, and a Replay Certificate
Vishwajith Ramesh · 30 September 2026
An assistant can stop repeating a deleted statement while its recurrent memory still carries that statement's influence. We examine this distinction by saving the state difference immediately after a record, carrying this saved difference, or receipt, through later updates, and comparing the correct…
- When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression
Kang Chen, Xiuze Zhou, Hong Chen, Yuanguo Lin · 30 September 2026
KV cache compression is widely used for long context LLM inference under memory constraints, while deployed systems typically score refusals after generation with keyword filters or learned classifiers. Such monitors are intended to indicate whether a model declined a harmful request under the servi…
- Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures
Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu, Jiangqiang Li, Jun Zhang, Junkai Ji · 29 September 2026
Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purel…
- Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Jiapeng Li · 25 September 2026
When a tool-using agent's write times out or returns a server error, the action may already have taken effect. Retrying blindly duplicates it -- a second charge, a second announcement, a second deployment -- while giving up skips required work. We ask where exactly-once behaviour should be enforced:…
- ERRAND: Budgeted Maintenance of Agent Memory
Beining Wu, Zihao Ding, Jun Huang · 25 September 2026
Deployed agents run on handed-over knowledge: a frozen policy consults a briefing of consolidated items written before the stream begins. The world then moves while the store stands still: paths close, flags change, price bands move; every item was true at handover, and the failure is staleness, not…
- Authorization Revocation for Long-Running AI Agents: Root-Scoped Quiescence under Delegation and Asynchronous Execution
Genliang Zhu, Chu Wang · 21 September 2026
Long-running AI agents outlive initiating processes through credentials, delegated tasks, queues, callbacks, reservations, and provider-side operations. Cancellation, process exit, and credential revocation neither close every pre-cut carrier nor distinguish independently authorized shared work. We …
- When AI Agents Commit: Cognitive Serializability Across Data, Evidence, Policy, and Authority
Jun He, Deying Yu · 18 September 2026
Autonomous agents derive concrete mutations from database reads, retrieved evidence, policy, beliefs, and delegated authority. Those inputs may change while reasoning is in progress. Database isolation orders the submitted transaction; agentic transaction processing determines whether a proposal sat…
- Recoverability as a System Primitive for Long-Horizon AI Agents
Zhihui Zhang, Wei Liu · 15 September 2026
AI agents can be interrupted while editing files, calling tools, or carrying out multi-step tasks. Restarting repeats completed work, but continuing from unverified or outdated progress can carry earlier errors forward. A saved state is not necessarily a suitable place to resume. We introduce recove…
- From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents
Yongjian Lyu, Yang Ren, Ruofei Lai, Wenting Liu · 9 September 2026
Long-running AI agents may read state, reason, wait for tools or human approval, and perform an external action much later. The state that justified the action can change in the meantime. For example, after an agent proposes an 80 GBP refund under a limit of 100, a customer-name change affects only …
- FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
Abhishek Sharma · 7 September 2026
A merchant's payment processor, ledger, ERP and bank feed are updated by messages that get delayed, duplicated, dropped and reordered, so for minutes at a time the four hold contradictory beliefs about the same order. An agent resolving the exception must decide whether to ship goods, re-submit a ca…
- The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems
Bardia Mohammadi, Laurent Bindschaedler · 2 September 2026
Fleets of LLM agents now externalize effects that cannot be fully undone: they move money, deploy code, delete data, and disclose information. Current controls check one effect at a time, so a fleet of individually authorized agents can overdraw its principal's risk under a shared trigger while ever…
- When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
Kazuki Nakayashiki · 27 August 2026
An agent that inherits a consolidated memory may inherit a constraint that was true when written and has since been withdrawn by a newer authoritative record. Under a scarce verification budget, does the agent recover the withdrawal, and if not, is the error avoidable without spending more? We model…
- MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories
Lauri Lovén, Jaakko Sauvola, Jukka Riekki, Sasu Tarkoma · 18 August 2026
Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' memories reconcile a fact phrased two ways, link related facts held apart, or reconcile contradictory knowledge without silently discarding either claim. We present…
- Aborted but Not Forgotten: KV-Cache Retention Breaks Rollback Consistency in Language Agents
Guijia Zhang, Harry Yang · 18 August 2026
Stateful language agents assume a rejected branch can be taken back by clearing it from the application transcript. We show this breaks when the serving session retains key/value (KV) state across the logical abort: the model can continue attending to content the application believes it discarded. W…
- Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State
Fanzhe Wei, Li Liu · 18 August 2026
Runtime compression of serving state trades quality for capacity with no priced guarantee: systems adapt precision on load signals with no soundness statement, and certified approaches budget request-level risk by a union bound over a pre-declared event count. We show the union budget exhausts on ev…
- Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
Jun He, Deying Yu · 13 August 2026
Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Without an explicit control plane, unmediated updates by models, tools, and background workers risk stale overwrites, un-audited exposures, and self-authorizing pr…
- TEPA: Revoking Stale Memories for Conflict-Robust Language Agents
Yan Zhou, Yue Ouyang, Kaiyang Zheng, Suncheng Xiang · 10 August 2026
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can remain retrievable and pollute the prompt. We characterize this failure mode as memory pollution: degr…
- SafeCommit: Certifying When Memory-Grounded Agents May Safely Act
Mayur Akewar, Ravi Ranjan · 6 August 2026
Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is premature commitment: an agent acts before resolving whether its memory grounding is stale, conflicting, incomplete, or corrupted. We formalize this problem as safe …
- Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers
Sajjad Khan · 5 August 2026
A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired. Five widely deployed agent workflow frameworks answer differently, none exposes a machine-checkable contract, and behavior violates eve…
- ALIVE: Warnings Before Exclusion in Budgeted Multi-Source Learning
Xiyang Zhang, Hongzhi Wang, Yuanhe Tian · 3 August 2026
A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions. We ask what evidence should authorize these unequal-persistence actions when finite-population auditing and learning share a budget. ALIVE (Action-Layered Intervention via Evide…
- MemTX: Transactional Belief Commit for Stateful Agent Memory
Xiaoyang Li, Yiqi Wang, Haohui Lu, Zhi Chen, Mo Li, Pingan Song, Mingkai Zheng, Taotao Cai · 29 July 2026
LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a tool call with real side effects. Current agent memory systems treat every accepted write as immediately actionable truth, so a polluted tool result, a stale updat…
- Error Certificates for KV-Cache Eviction via Randomized Design
Peng Xie · 24 July 2026
Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot know what it destroyed: evicted values can be altered so that everything the serving system retains is unchanged while the true attention-output error grows arbit…
- When JSON Is Not Enough: Semantic Reliability of Schema-Constrained LLM Ordering Agents
Yin Li · 22 July 2026
LLM agents are increasingly used as transaction compilers: a user states an intent in natural language, and the model emits a structured object that an API can execute. JSON Schema and provider-level structured-output modes are useful because they remove a large class of parse failures, but they do …
- Teach it to stop, not just to click
Barada Sahu (Cabal AI), Shivesh Pandey (Para AI) · 21 July 2026
Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CUA) across five oracle-graded environments, we show a repaired policy's success rate is dominated by upstream variance: a variance-components decomposition across…
Other topics in Computer networks and communications
The topics the OpenAlex classification attaches to the same theme, most active first.
- Software System Performance and Reliability395 papers / 12 months+400%
- Constraint Satisfaction and Optimization254 papers / 12 months+220%
- Software-Defined Networks and 5G205 papers / 12 months+400%
- Network Security and Intrusion Detection186 papers / 12 months+260%
- IoT and Edge/Fog Computing150 papers / 12 months+175%
- Caching and Content Delivery139 papers / 12 months+1500%
