Physical Sciences › Computer Science › Artificial Intelligence
Security and Verification in Computing
418 papers indexed
The research grouped under this theme explores the vulnerabilities and verification mechanisms of artificial intelligence systems, particularly those based on language models and autonomous agents. It analyzes attacks such as prompt injection, skill hijacking, or reasoning exfiltration, as well as methods to detect and counter these threats, whether intentional or emergent. Emphasis is placed on the robustness of architectures, secure skill composition, and the assessment of risks associated with interactions between tools, agents, and digital environments.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States48% · 96 papers
- China29% · 58 papers
- Germany8% · 16 papers
- United Kingdom7% · 14 papers
- Switzerland4.5% · 9 papers
- Singapore4% · 8 papers
- South Korea4% · 8 papers
- Israel3.5% · 7 papers
Across 199 papers on this subject with at least one lab located. 39 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents
Taolin Zhang, Jiuheng Wan, Hanyu Wang, Tingyuan Hu, Chengyu Wang · 2 October 2026
LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from…
- SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses
Xingyu Li, Juefei Pu, Haonan Li, Arrdya Srivastav, Kareem Shehada, Srikanth V. Krishnamurthy, Zhiyun Qian · 2 October 2026
Automated kernel vulnerability reproduction is essential for bug triage, patch validation, and regression testing, but still lacks an effective and efficient solution. The core challenge is twofold: a reproducer must first recover the trigger scaffold needed to reach the vulnerable state and determi…
- Chaining Skills to Hijack LLM Agents
Tian Dong, Zixuan Ma, Haodong Zhao, Huaien Zhang, Shaofeng Li, Hao Chen · 2 October 2026
LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-…
- PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents
Fengpeng Li, Qizhou Wang, Yuke Hu, Kemou Li, Jun Liu, Haiwei Wu, Jiantao Zhou, Di Wang · 2 October 2026
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the sa…
- Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
Birk Torpmann-Hagen, Finn Schwall, Leon Moonen · 2 October 2026
Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce \emph{memetic troja…
- When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models
Pakhapoom Sarapat, Saksorn Ruangtanusak, Kunat Pipatanakul, Pittawat Taveekitworachai · 1 October 2026
We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a…
- Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Tobias Kaisar, Aritra Dhar · 1 October 2026
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging…
- Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
Zezhong Wang, Xueyang Tang, Rui Lian, Yang Lou, Heqing Huang · 1 October 2026
As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious…
- ActionGuard: Tool Call Authorization under Poisoned Skills
Jihun Han, Yejin Jang, Byung Il Kwak, Mee Lan Han · 1 October 2026
LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exf…
- Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents
Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao · 1 October 2026
LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or dis…
- Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun · 1 October 2026
LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority …
- Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses
Yang Wang · 1 October 2026
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval auth…
- Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks
Alexandra Souly, Kai Fronsdal, Abby D'Cruz, Xander Davies, Robert Kirk · 1 October 2026
This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of…
- When Does Randomized Oversight Align AI Agents That Can Conceal?
Joshua S. Gans, Richard Holden · 1 October 2026
Oversight changes the evidence it relies on. We ask when randomized audits and scoring align AI agents that can conceal misconduct and alter records. Stronger auditing makes undeterred violations better hidden. Because the provider writes the agent's objective, sanctions need not stop at forfeiture,…
- Checking Leakage Witnesses versus Certifying Bounded Non-Leakage
Chao Feng, Burkhard Stiller · 30 September 2026
When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, whil…
- Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation
Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah · 30 September 2026
Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, isolating distinct fail…
- Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring
Mohammadali Mohammadkhani, Madhava Krishna, Yash Sarrof, Michael Hahn · 30 September 2026
Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold de…
- ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
Yanjie Li, Xiangyu He, Xuelong Dai, Bin Xiao · 30 September 2026
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because the…
- SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs
Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo, Tianwei Zhang, Kwok-Yan Lam · 30 September 2026
As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growi…
- CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
Mark Russinovich · 30 September 2026
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whethe…
- Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao · 30 September 2026
Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode…
- MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du, Sibo Yi, Yuqi Qing, Zhenguang Liu, Qinming He, Shiwen Cui, Changhua Men · 30 September 2026
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidan…
- SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents
Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye, Yuyuan Li, Haibo Hu · 30 September 2026
Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign…
- Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety
Charlie Summers, Prajwal Raghunath, Aaditya Pai, Mayur Kulkarni, Zhuo Zhang, Oliver Kennedy, Eugene Wu · 30 September 2026
LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that …
- Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift
Wesley Shu · 30 September 2026
Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mechanism that binds a single-use commit authorization to the exact actio…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
