Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3,552 papers indexed
The study of adversarial robustness in machine learning explores how models respond to deliberate perturbations or data designed to deceive them. Recent work focuses on verifying the stability of predictions, detecting security vulnerabilities, or adapting classifiers to targeted attacks. It also addresses questions such as the calibration of large language models, the optimization of malicious transferability, or unlearning mechanisms to correct biases or undesirable behaviors.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States38% · 939 papers
- China34% · 841 papers
- United Kingdom8.4% · 205 papers
- Germany6.4% · 158 papers
- India5.2% · 128 papers
- Canada5.1% · 125 papers
- Australia4.3% · 106 papers
- South Korea3.7% · 91 papers
Across 2,454 papers on this subject with at least one lab located. 91 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers
Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University) · 2 October 2026
This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at $\epsilo…
- Towards Robust Numerical Claim Verification
Peter R{\o}ysland Aarnes, Vinay Setty · 2 October 2026
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically pert…
- Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim · 2 October 2026
Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether succ…
- False Prophets: On the Security of World Models in Agentic Systems
Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck · 2 October 2026
Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent research proposes to enhance predictive capabilities via specially trained…
- UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao · 2 October 2026
Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship a…
- Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints
Melika Shirian, Kianoosh Vadaei · 2 October 2026
Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, wh…
- Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun · 2 October 2026
Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an ac…
- TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang · 2 October 2026
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient co…
- Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
Jianwei Li, Jung-Eun Kim · 2 October 2026
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, acces…
- Backdoor Containment via Expert Quarantine and Shutdown in LLMs
Jianwei Li, Min-Seon Kim, Jung-Eun Kim · 2 October 2026
Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppre…
- Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole
Cris Huynh · 2 October 2026
When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool,…
- Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation
Ziqi Zhou, Yifan Hu, Yufei Song, Haowen Jiang, Xianlong Wang, Shengshan Hu, Dezhong Yao, Leo Yu Zhang · 1 October 2026
The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial exam…
- Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models
Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su · 1 October 2026
Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image…
- BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
Zhihang Wu, Zhongqi Wang, Jie Zhang, Fengming Gu, Shiguang Shan, Xilin Chen · 1 October 2026
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the inte…
- FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance
Rikhiya Ghosh, Himanshu Kumar, Sriram Venkatapathy, Sahil Wadhwa, Alexandre G. R. Day, Pranab Mohanty · 1 October 2026
In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety failures and pushing responses dangerously close to policy limits. Existing automated red-teaming methods trade off attack effectiveness against generati…
- Probabilistic Adversarial Training
Andi Zhang, Xingyu Zhao, Siddartha Khastgir · 1 October 2026
Building on a probabilistic perspective in which adversarial examples arise from the overlap between a distance-based distribution $p_{\mathrm{dis}}$ and a victim-classifier-induced distribution $p_{\mathrm{vic}}$, we start from a simple intuition: adversarial examples become harder to generate when…
- Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation
Linfeng Jiang, Steven McDonagh, Yuhang Chen, Xingyu Zhao, Siddartha Khastgir, Andi Zhang · 1 October 2026
Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the atta…
- When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic
Muhammad Zawish, Steven Davy · 1 October 2026
This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly …
- ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara · 1 October 2026
Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content …
- Faithful Dual-constrained Erasure for Robust LLM Safety Alignment
Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang · 1 October 2026
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed maliciou…
- Alignment via Training Against Probes Without Losing Monitorability
Lena Libon, Alexander Panfilov, Ben Rank, Xin Chen, Jonas Geiping, Maksym Andriushchenko · 1 October 2026
Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance durin…
- CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition
Zhiyi Mou, Yao Lu, Wangze Ni, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou, Kui Ren · 1 October 2026
Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain inconspicu…
- Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard
Julian Schulz, Lukas F\"ulle, Rieke Fruengel · 1 October 2026
Chain-of-thought monitoring as an approach for AI oversight and control is threatened by the possibility of steganographic reasoning, where LLMs conceal their reasoning inside innocuous-looking text. Two neighbouring capabilities, steganographic messaging (passing a concealed message) and encoded re…
- Safety of Latent Communication in Multi-Agent Systems
Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer · 1 October 2026
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the …
- Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models
Haoran Gu, Handing Wang, Yi Mei, Mengjie Zhang, Yaochu Jin · 30 September 2026
Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating harmful responses despite safety alignme…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
- Natural Language Processing Techniques1,595 papers / 12 months+69%
