Physical Sciences › Computer Science › Artificial Intelligence
Security and Verification in Computing
418 indexierte Paper
Die unter diesem Thema zusammengefassten Forschungen untersuchen die Schwachstellen und Verifizierungsmechanismen von Systemen der künstlichen Intelligenz, insbesondere solcher, die auf Sprachmodellen und autonomen Agenten basieren. Sie analysieren Angriffe wie prompt injection, skill hijacking oder reasoning exfiltration sowie Methoden zur Erkennung und Abwehr dieser Bedrohungen, ob sie nun absichtlich oder emergent sind. Der Fokus liegt auf der Robustheit von Architekturen, der sicheren Komposition von Fähigkeiten und der Risikobewertung im Zusammenhang mit Interaktionen zwischen Tools, Agenten und digitalen Umgebungen.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten48 % · 96 Artikel
- China29 % · 58 Artikel
- Deutschland8 % · 16 Artikel
- Vereinigtes Königreich7 % · 14 Artikel
- Schweiz4,5 % · 9 Artikel
- Singapur4 % · 8 Artikel
- Südkorea4 % · 8 Artikel
- Israel3,5 % · 7 Artikel
Über 199 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 39 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- When a Kindergartener Solves Calculus: Measuring Capability Leakage in Role-Prompted Reasoning Models
Pakhapoom Sarapat, Saksorn Ruangtanusak, Kunat Pipatanakul, Pittawat Taveekitworachai · 1. Oktober 2026
We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a…
- Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Tobias Kaisar, Aritra Dhar · 1. Oktober 2026
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging…
- Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks
Zezhong Wang, Xueyang Tang, Rui Lian, Yang Lou, Heqing Huang · 1. Oktober 2026
As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious…
- ActionGuard: Tool Call Authorization under Poisoned Skills
Jihun Han, Yejin Jang, Byung Il Kwak, Mee Lan Han · 1. Oktober 2026
LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exf…
- Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents
Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao · 1. Oktober 2026
LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or dis…
- Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
Yan Wang, Zhihao Zhang, Ke Chen, Kai Chen, Yaqin Zhang, Duohe Ma, Jun Dai, Xiaoyan Sun · 1. Oktober 2026
LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority …
- Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses
Yang Wang · 1. Oktober 2026
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval auth…
- Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks
Alexandra Souly, Kai Fronsdal, Abby D'Cruz, Xander Davies, Robert Kirk · 1. Oktober 2026
This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of…
- When Does Randomized Oversight Align AI Agents That Can Conceal?
Joshua S. Gans, Richard Holden · 1. Oktober 2026
Oversight changes the evidence it relies on. We ask when randomized audits and scoring align AI agents that can conceal misconduct and alter records. Stronger auditing makes undeterred violations better hidden. Because the provider writes the agent's objective, sanctions need not stop at forfeiture,…
- Checking Leakage Witnesses versus Certifying Bounded Non-Leakage
Chao Feng, Burkhard Stiller · 30. September 2026
When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, whil…
- Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation
Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah · 30. September 2026
Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, isolating distinct fail…
- Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring
Mohammadali Mohammadkhani, Madhava Krishna, Yash Sarrof, Michael Hahn · 30. September 2026
Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold de…
- ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
Yanjie Li, Xiangyu He, Xuelong Dai, Bin Xiao · 30. September 2026
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because the…
- SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs
Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo, Tianwei Zhang, Kwok-Yan Lam · 30. September 2026
As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growi…
- CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
Mark Russinovich · 30. September 2026
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whethe…
- Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao · 30. September 2026
Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode…
- MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du, Sibo Yi, Yuqi Qing, Zhenguang Liu, Qinming He, Shiwen Cui, Changhua Men · 30. September 2026
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidan…
- SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents
Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie, Qingqing Ye, Yuyuan Li, Haibo Hu · 30. September 2026
Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign…
- Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety
Charlie Summers, Prajwal Raghunath, Aaditya Pai, Mayur Kulkarni, Zhuo Zhang, Oliver Kennedy, Eugene Wu · 30. September 2026
LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that …
- Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift
Wesley Shu · 30. September 2026
Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mechanism that binds a single-use commit authorization to the exact actio…
- Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?
Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson · 30. September 2026
Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a compl…
- Certified Multi-Source Integrity for Structured Agent Actions
Anmol Pandey, Aditya Jain, Liang Chen, Carsten Maple, Christo Panchev · 29. September 2026
LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-c…
- HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents
Zhuowen Liu, Zhixuan Wang · 29. September 2026
Security operations centers receive far more alerts than analysts can investigate, and organizations that cannot send their telemetry to hosted models must automate triage with small open-weight LLMs on their own hardware. Current LLM agents leave the investigation procedure to the model, and small …
- API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary
Patrick Kenney, Hadi Ahmadi, Denis Lusson, Donald Nguyen, Gurbinder Gill · 29. September 2026
Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in conv…
- AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents
Asif Shahriar, Md Nafiu Rahman, Sadif Ahmed, Farig Sadeque, Md Rizwan Parvez · 29. September 2026
Browser-use agents often carry information in their context as they move between websites. While it may be necessary for task completion, it also creates a privacy risk, especially when the information contains a private fact regarding the user. For example, an agent may learn a user's affiliation a…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
