Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3 552 papiers indexés
L’étude de la robustesse adversariale en apprentissage automatique explore comment les modèles réagissent à des perturbations délibérées ou à des données conçues pour les tromper. Les travaux récents s’intéressent à la vérification de la stabilité des prédictions, à la détection de failles de sécurité ou à l’adaptation des classificateurs face à des attaques ciblées. Ils abordent aussi des questions comme la calibration des grands modèles de langage, l’optimisation de transferts malveillants ou les mécanismes de désapprentissage pour corriger des biais ou des comportements indésirables.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis38 % · 939 articles
- Chine34 % · 841 articles
- Royaume-Uni8,4 % · 205 articles
- Allemagne6,4 % · 158 articles
- Inde5,2 % · 128 articles
- Canada5,1 % · 125 articles
- Australie4,3 % · 106 articles
- Corée du Sud3,7 % · 91 articles
Sur 2 454 articles de ce sujet dont au moins un laboratoire est situé. 91 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild
Elad David, Max Fomin · 5 octobre 2026
LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to appen…
- CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
Shouli Wang, Yanfeng Jia, Zhihao Ou, Zitao Su, Ruize He, Haotong Xie, Hao Peng, Juanzi Li, Xiaozhi Wang · 5 octobre 2026
During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigat…
- Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
Narek Maloyan · 5 octobre 2026
Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(…
- Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation
Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha · 5 octobre 2026
Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prio…
- Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
Yukiya Horiba, Koshiro Aoki, Shunsuke Yasuki, Bum Jun Kim, Taiki Miyanishi · 5 octobre 2026
Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we me…
- Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama · 5 octobre 2026
A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the …
- A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control
Zhe Zhou, Tianhua Tao · 5 octobre 2026
Post-training with verifiable rewards can induce reward hacking, motivating the use of monitors within the training objective rather than solely for offline auditing. We show that a low monitor readout does not identify whether such an intervention controls behavior. In a code-generation environment…
- Positive-Unlabeled Learning for Agent Safety False Alarm Auditing
Xichen Yan, Chongyang Gao, Kezhen Chen, Guangyi Zhang, Jiaqi Wu, Lixu Wang · 5 octobre 2026
Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native m…
- Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models
Omanshu Thapliyal · 5 octobre 2026
Safety heads are lightweight classifiers attached to pretrained language models for flagging harmful inputs before generation. Their empirical detection performance has been studied, but their formal robustness properties remain largely unexplored. We ask when a State Space Model (SSM)-based safety …
- Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
NaHyeon Park, Minhyun Lee, Hyunjung Shim · 5 octobre 2026
Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivi…
- The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs
Toluwani Aremu, Manit Baser, Mohan Gurusamy, Nils Lukas, Dinil Mon Divakaran · 5 octobre 2026
Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target c…
- Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers
Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University) · 2 octobre 2026
This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at $\epsilo…
- Towards Robust Numerical Claim Verification
Peter R{\o}ysland Aarnes, Vinay Setty · 2 octobre 2026
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically pert…
- Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim · 2 octobre 2026
Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether succ…
- False Prophets: On the Security of World Models in Agentic Systems
Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck · 2 octobre 2026
Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent research proposes to enhance predictive capabilities via specially trained…
- UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao · 2 octobre 2026
Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship a…
- Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints
Melika Shirian, Kianoosh Vadaei · 2 octobre 2026
Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, wh…
- Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun · 2 octobre 2026
Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an ac…
- TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang · 2 octobre 2026
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient co…
- Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
Jianwei Li, Jung-Eun Kim · 2 octobre 2026
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, acces…
- Backdoor Containment via Expert Quarantine and Shutdown in LLMs
Jianwei Li, Min-Seon Kim, Jung-Eun Kim · 2 octobre 2026
Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppre…
- Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole
Cris Huynh · 2 octobre 2026
When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool,…
- Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation
Ziqi Zhou, Yifan Hu, Yufei Song, Haowen Jiang, Xianlong Wang, Shengshan Hu, Dezhong Yao, Leo Yu Zhang · 1 octobre 2026
The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial exam…
- Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models
Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su · 1 octobre 2026
Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image…
- BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
Zhihang Wu, Zhongqi Wang, Jie Zhang, Fengming Gu, Shiguang Shan, Xilin Chen · 1 octobre 2026
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the inte…
Autres sujets du thème Intelligence artificielle
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Large Language Models7 407 papiers / 12 mois+247 %
- Reinforcement Learning in Robotics2 519 papiers / 12 mois+117 %
- Explainable Artificial Intelligence (XAI)2 319 papiers / 12 mois+200 %
- Domain Adaptation and Few-Shot Learning2 059 papiers / 12 mois+67 %
- Advanced Graph Neural Networks1 926 papiers / 12 mois+38 %
- Natural Language Processing Techniques1 595 papiers / 12 mois+69 %
