Physical Sciences › Computer Science › Computer Networks and Communications
Software System Performance and Reliability
395 papers indexed
Research under this theme explores the challenges related to the performance and reliability of software systems, particularly in contexts where artificial intelligence is deployed at scale. It addresses issues such as anomaly detection, fault localization, infrastructure optimization, and workload management, especially for large language models (LLM) and autonomous agents. The work also examines methods to enhance error diagnosis, resource allocation, and system robustness in the face of rare or complex failures.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States43% · 117 papers
- China35% · 94 papers
- India8.6% · 23 papers
- Canada6.3% · 17 papers
- Germany5.6% · 15 papers
- United Kingdom4.1% · 11 papers
- Singapore3.3% · 9 papers
- France2.6% · 7 papers
Across 269 papers on this subject with at least one lab located. 49 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- DeFA: Dependency-Guided Failure Attribution for LLM Agents
Bo Deng, Xinlei Zheng, Yi Wei, Kang Zhou, Chongyang Tao, Renzhao Liang, Xuanren Chen, Lifan Guo, Chi Zhang · 2 October 2026
Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guided framework for agent failure attribution. DeFA first combines prot…
- Incident-Arena: Getting agents to the last nine of reliability
Andre Fu, Malik Drabla, Leon Liu, Meji Abidoye, Marek Suppa, Lata Mishra, Adnan El Assadi, Yiyuan Li · 2 October 2026
AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) co…
- Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
Yasuto Hoshi, Daisuke Miyashita, Jun Deguchi · 30 September 2026
Residual completion augments query-aware sparse attention by estimating the contribution of tokens omitted from the exact sparse computation. We ask whether improving a layer's attention-output reconstruction on the same incoming Q/K/V and selected support necessarily improves the fidelity of the fi…
- REFINE: A Resilient Evolution Framework for Intelligent Enterprise Alert Triage in Security Operations Centers
Huimin Chen, Quan Long, Yanhao Wang · 29 September 2026
Security Operations Centers (SOCs) process large volumes of alerts daily. Alert triage prioritizes high-risk threats while reducing manual review of benign alerts. LLM agents can reason over logs and threat intelligence, but struggle to keep aligned with organization-specific, rapidly evolving SOC o…
- STAR: Adaptive Spatial-Temporal Normalization for Unified Microservice Incident Management
Xinhua Miao, Linyu Zhu, Bowei Yang, Zhengong Cai · 29 September 2026
Automated incident management in large-scale microservice systems relies on learning robust representations from multimodal observability data, including metrics, logs, and traces. Although recent self-supervised frameworks enable unified modeling for anomaly detection (AD), failure triage (FT), and…
- Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries
Haiquan Hu, Yuzhu Liang, Weicheng Tang, Yanzeng Li, Yao Shi, Tian Wang · 29 September 2026
Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boun…
- SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents
Yifang Tian, Yingjian Bai, Yifeng He, Zichun Chong, Yuanchen Gao, Yiran Li, Hans-Arno Jacobsen · 29 September 2026
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuratio…
- From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis
Shu-Xun Yang, Yidong Wang, Zhuoer Feng, Bosi Wen, Jiayi Gui, Dayong Yang, Wenbo Yu, Haoke Zhang, Jie Tang, Cunxiang Wang · 29 September 2026
LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally…
- DAAF: From Failure Localization to Editable System Assets in LLM Agents
Xiaoyang Yuan, Qi Liu, Yubin Ruan, Xinyi Mou, Zhuomeng Zhang, Wenjin Wang, Hanying Jiao, Di Wu, Mingye Xu, Yi Bin, Ke Feng, Zixun Sun · 29 September 2026
Deployed LLM agents increasingly rely on persistent, versioned system assets such as routing rules, knowledge segments, prompt instructions, and reusable skills. Failure-localization methods can identify where an error manifests in an agent or execution trace, but repair requires a different decisio…
- Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness
Bin Li, Dongdong Wang, Siyang Lu · 28 September 2026
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic emp…
- Learning to Remember: Attentive Reinforcement Learning for Edge Serverless Autoscaling
Faraz Shaikh, Gianluca Reali, Mauro Femminella · 24 September 2026
In edge computing, the stochastic and bursty nature of serverless workloads challenges autonomous resource orchestration. Traditional reactive controllers, such as the Kubernetes Horizontal Pod Autoscaler (HPA), suffer from reaction latency, leading to Service Level Objective (SLO) violations during…
- Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring
Imad Bulji\'c · 24 September 2026
Graph neural networks dominate recent work on microservice root-cause analysis, yet recent results question whether the graph contributes. Those results compare whole pipelines, so when a flat model wins one cannot tell whether structure is useless or redundant. We run the comparison they imply on R…
- Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
Weihang Ding, Junfei Zhan · 23 September 2026
Post-training is becoming a service (PTaaS): a customer hands an operator data and a goal, and a forward-deployed engineer (FDE) returns a fine-tuned, evaluated, and deployed model under a budget, a human-approval gate, and reproducibility requirements. Seating an LLM agent in the FDE seat raises a …
- When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention
Peiying Zhu, Sidi Chang · 23 September 2026
Runtime traces can appear transparent, but a closed-loop policy determines which states are visited and which failures become visible. We study a simulated hotel-pricing agent mapping time, inventory, and market state to discrete price actions under varying demand regimes. A fault may leave no aggre…
- ARID: A Deployable Edge AI System for Structured Information Extraction from Industrial Maintenance Work Orders
Kuanlin Chen, Chen-Wei Kuo · 22 September 2026
Maintenance work orders must often be processed offline on embedded hardware, yet downstream software requires predictable structured output. We present ARID (Aviation-inspired Routing for Industrial Deployment), which extracts component, failure mode, symptom, and maintenance action into fixed-sche…
- Towards Secure Cloud-Native Computing: Unveiling Kubernetes Misconfigurations with Large Language Models
Mostafa Anouar Ghorab, Mohamed Aymen Saied · 21 September 2026
In the rapidly evolving landscape of cloud-native computing, Organizations are increasingly adopting infrastructure models that emphasize scalability, flexibility, and efficiency. Kubernetes has become the de facto standard for orchestrating containerized applications in these environments. However,…
- BurnRiSc: Toward Non-Invasive Burnout Screening in Open Source from Public Repository Signals
Timofey Sanko, Yuan Tian, Mariam Guizani · 18 September 2026
Burnout is a chronic occupational syndrome, and open source is close to a worst case for it: maintainers absorb unbounded demand with no manager to reallocate work and no organization to notice decline. The cost is not only personal. Burnout precedes withdrawal, and in projects sustained by a handfu…
- TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary · 15 September 2026
Cloud-hosted large language models (LLMs) are increasingly used for root cause analysis (RCA) in AIOps pipelines, but they introduce data privacy risk, network latency, and per-query cost that scale poorly with production log volumes. We present TriCalRAG, a benchmark evaluating open-weight LLMs ser…
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta · 15 September 2026
The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving…
- DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems
Zehao Wang, Lanjun Wang, Shilong Jin, Junjie Chen, Yanghua Xiao · 7 September 2026
Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracin…
- SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center
Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild · 4 September 2026
Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-form generation offers no guarantee that a recommended contai…
- Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis
Hao Zhou (Jianzhong), Mandar Kulkarni (Jianzhong), Hao Chen (Jianzhong), Yan Xin (Jianzhong), Charlie (Jianzhong), Zhang · 3 September 2026
Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations in modern 5G and emerging 6G networks remains challenging due to complex cross-layer dependencies. While large language models (LLMs) offer promising capabilities for reasoning and kno…
- Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
Ron Begleiter, Katya Egert Berg, Gilad Saban, Gil Shabat · 3 September 2026
Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge when deploying NLP systems in real-world industrial settings. While monolithic Large Language Model (LLM) agents offer unbounded expressivity for tasks like Root Cause Analysis (RCA), they suffer f…
- Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics
Jiani He, Dingyan Shang, Yihua Xu, Shiqi Huang, Yan Lyu, Jize Li, Shangjing Tang · 3 September 2026
Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Signal-Assisted Decision Support converts return notes into a condition factor and a signal-quality score that guide inspe…
- EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems
Jun Hou, Priya Pitre, Yi Fang, Xuan Wang · 2 September 2026
Large language model (LLM) agent failures often contain multiple related errors rather than a single mistake. Existing attribution methods usually identify a responsible agent, step, or root cause, but do not explicitly model dependency between errors. We introduce EDGE, an Error Dependency Graph-gu…
Other topics in Computer networks and communications
The topics the OpenAlex classification attaches to the same theme, most active first.
- Constraint Satisfaction and Optimization254 papers / 12 months+220%
- Software-Defined Networks and 5G205 papers / 12 months+400%
- Network Security and Intrusion Detection186 papers / 12 months+260%
- IoT and Edge/Fog Computing150 papers / 12 months+175%
- Caching and Content Delivery139 papers / 12 months+1500%
- Advanced Database Systems and Queries130 papers / 12 months+220%
