Physical Sciences › Computer Science › Artificial Intelligence
Security and Verification in Computing
418 artículos indexados
Las investigaciones agrupadas bajo este tema exploran las vulnerabilidades y los mecanismos de verificación de los sistemas de inteligencia artificial, en particular aquellos basados en modelos de lenguaje y agentes autónomos. Analizan ataques como el prompt injection, el secuestro de habilidades o la exfiltración de razonamientos, así como los métodos para detectar y contrarrestar estas amenazas, ya sean intencionales o emergentes. Se pone énfasis en la robustez de las arquitecturas, la composición segura de habilidades y la evaluación de riesgos asociados a las interacciones entre herramientas, agentes y entornos digitales.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos48 % · 96 artículos
- China29 % · 58 artículos
- Alemania8 % · 16 artículos
- Reino Unido7 % · 14 artículos
- Suiza4,5 % · 9 artículos
- Singapur4 % · 8 artículos
- Corea del Sur4 % · 8 artículos
- Israel3,5 % · 7 artículos
Sobre 199 artículos de este tema con al menos un laboratorio localizado. 39 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
Zhuowen Liu · 5 de octubre de 2026
LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDo…
- Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu · 5 de octubre de 2026
Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences th…
- EvoRiskBench: An Evolving Benchmark for Runtime Security Risks in Workspace Agents
Shiyi Kuang, Xuemei Luo, Kun Liu, Junhai Li, Rui Tian, Feng Shi, Bo Shen, Nianyu Li, Dehui Li, Ping Chen · 5 de octubre de 2026
Workspace agents combine large language models with execution harnesses to perform stateful, multi-step tasks that access or modify external resources. Existing benchmarks leave gaps in executable coverage of their runtime security risks, while evolving model capabilities, harnesses, tools, and thre…
- Securing Computer-Use Agents Against Branch Steering Attacks
Giulio Zingrillo, Hanna Foerster, Ilia Shumailov, Yiren Zhao, Robert Mullins · 5 de octubre de 2026
Modern Computer Use Agents (CUAs) directly interact with graphical user interfaces and execute third-party web tools, exposing them to indirect prompt injection across every rendered page and tool response. While the Dual-LLM pattern is the primary system-level architecture offering formal security …
- Beyond Predefined Sinks: Security-Aware Dependency Analysis for LLM Agents
Hang Cui · 5 de octubre de 2026
Large language model (LLM)-based agents increasingly connect model-generated decisions to security-sensitive software capabilities such as command execution, filesystem access, network communication, browser control, and external tools. Existing analyses often use predefined sensitive operations as …
- Pincer: Resource Authorization for Agents using a Digital Twin
Mayank Rathee, Alexander Stepanov, Shalin Madabhavi, Jinhao Zhu, Raluca Ada Popa, Ion Stoica · 5 de octubre de 2026
Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses th…
- SideKernel: A Usable microVM Sandbox for AI Coding Agents on macOS
Dimitrios Prasakis · 5 de octubre de 2026
AI coding agents are untrusted system components, yet they require autonomy on the developer machines they run on. This contradiction is a security problem. Sandboxes provide an isolated environment, but for local macOS development, the existing local, open-source options for AI coding agents are fe…
- MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication
Ryuichi Yamafuji Lun, Jingzhen Wang, Shreyas Kolte, Ruiteng Li · 5 de octubre de 2026
Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (AS…
- HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence
Xiqiao Xiong, Moxin Li, Zhixin Ma, Ouxiang Li, Wenjie Wang, Fuli Feng, Xiangnan He · 5 de octubre de 2026
Agent harnesses play a critical role in defenses by enforcing safety constraints to prevent unsafe actions. However, rapidly emerging attacks outpace manual harness adaptation, motivating automated harness evolution. Yet the signals available for harness evolution are often sparse, such as brief des…
- A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns
XinPeng Shen, Lan Zhang, Yixiao Huang, Haoran Cheng, Jiewei Lai, Leilei Chen, Haoxiang Deng · 5 de octubre de 2026
Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety concern. Under benign interaction conditions, an agent may execute an act…
- OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents
Taolin Zhang, Jiuheng Wan, Hanyu Wang, Tingyuan Hu, Chengyu Wang · 2 de octubre de 2026
LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from…
- SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses
Xingyu Li, Juefei Pu, Haonan Li, Arrdya Srivastav, Kareem Shehada, Srikanth V. Krishnamurthy, Zhiyun Qian · 2 de octubre de 2026
Automated kernel vulnerability reproduction is essential for bug triage, patch validation, and regression testing, but still lacks an effective and efficient solution. The core challenge is twofold: a reproducer must first recover the trigger scaffold needed to reach the vulnerable state and determi…
- Chaining Skills to Hijack LLM Agents
Tian Dong, Zixuan Ma, Haodong Zhao, Huaien Zhang, Shaofeng Li, Hao Chen · 2 de octubre de 2026
LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-…
- PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents
Fengpeng Li, Qizhou Wang, Yuke Hu, Kemou Li, Jun Liu, Haiwei Wu, Jiantao Zhou, Di Wang · 2 de octubre de 2026
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the sa…
- Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
Birk Torpmann-Hagen, Finn Schwall, Leon Moonen · 2 de octubre de 2026
Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce \emph{memetic troja…
- When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models
Pakhapoom Sarapat, Saksorn Ruangtanusak, Kunat Pipatanakul, Pittawat Taveekitworachai · 1 de octubre de 2026
We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a…
- Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Tobias Kaisar, Aritra Dhar · 1 de octubre de 2026
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging…
- Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
Zezhong Wang, Xueyang Tang, Rui Lian, Yang Lou, Heqing Huang · 1 de octubre de 2026
As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious…
- ActionGuard: Tool Call Authorization under Poisoned Skills
Jihun Han, Yejin Jang, Byung Il Kwak, Mee Lan Han · 1 de octubre de 2026
LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exf…
- Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents
Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao · 1 de octubre de 2026
LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or dis…
- Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun · 1 de octubre de 2026
LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority …
- Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses
Yang Wang · 1 de octubre de 2026
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval auth…
- Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks
Alexandra Souly, Kai Fronsdal, Abby D'Cruz, Xander Davies, Robert Kirk · 1 de octubre de 2026
This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of…
- When Does Randomized Oversight Align AI Agents That Can Conceal?
Joshua S. Gans, Richard Holden · 1 de octubre de 2026
Oversight changes the evidence it relies on. We ask when randomized audits and scoring align AI agents that can conceal misconduct and alter records. Stronger auditing makes undeterred violations better hidden. Because the provider writes the agent's objective, sanctions need not stop at forfeiture,…
- Checking Leakage Witnesses versus Certifying Bounded Non-Leakage
Chao Feng, Burkhard Stiller · 30 de septiembre de 2026
When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, whil…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
