Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3.164 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
Chengshuai Zhao, Pingchuan Ma, Dawei Li, Bohan Jiang, Zhiyuan Yu, Zhen Tan, Huan Liu · 3. August 2026
The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage. Unlearnable examples (UEs) offer a promising defense by introducing car…
- Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
Weiyi He, Yuping Lin, Jiliang Tang, Yue Xing · 3. August 2026
Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs). While existing mitigation strategies, e.g., latent adversarial training (LAT), have been developed, th…
- Lilith: Backdoor Generalization under Training-Inference Trigger Shift
Zhou Feng, Jiahao Chen, Chunyi Zhou, Yuan Su, Tianyu Du, Yuwen Pu, Jianhai Chen, Jinbao Li, Shouling Ji · 30. Juli 2026
Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger …
- GPT-Red: Automated Red Teaming via Self-Play at Scale
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cer\'on Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen · 30. Juli 2026
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most …
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
Keegan Harris, Brian W. Lee, Ian Waudby-Smith, Philip Amortila, Nika Haghtalab, Michael I. Jordan · 30. Juli 2026
Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself pr…
- ToxScreen: Detecting Whether an LLM Has Been Poisoned
Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov · 30. Juli 2026
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-bo…
- RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning
Pushkal Kumar, Tucker Nielson, Tanish Kolhe, Shubham Zala, Vincent Li · 30. Juli 2026
Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence. We introduce RAGuard, a layered defense against \emph{factual} corpus-poisoning …
- Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita · 30. Juli 2026
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain l…
- On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Yongjian Guo, Wanlun Ma, Lingyu Shen, Xi Xiao, Sheng Wen · 30. Juli 2026
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing …
- Anti-Backdoor Coreset Selection via Cumulative Entropy
Qi Zhao, Christian Wressnegger · 29. Juli 2026
Recent training-time defenses against neural backdoors isolate a benign subset from poisoned training data, to learn a backdoor-free model from it. In this paper, we formulate this defense strategy as a coreset selection problem, giving rise to so-called "Anti-Backdoor Coreset Selection." Since pois…
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cin\`a, Luca Oneto, Iacopo Masi, Fabio Roli · 29. Juli 2026
Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text encoders, and exported computation graphs are distributed by third parties and reused across downstream services. This reuse model creates a security-…
- Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato · 29. Juli 2026
Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly…
- Directional Influence Function: Estimating Training Data Influence in Constrained Learning
Xin Wang (Jeff), R. Tyrrell Rockafellar (Jeff), Xuegang (Jeff), Ban · 28. Juli 2026
As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluence the model solution (e.g., learned parameters) is…
- Not All LLM Reasoning is Visible in the Chain-of-Thought
Vatsal Baherwani, Tom Goldstein, Ashwinee Panda · 28. Juli 2026
A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tas…
- Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety
Domenic Rosati, Ali Dadsetan, Hong Huang, Xijie Zeng, Hassan Chowdhry, Subhabrata Majumdar, Hassan Sajjad, Frank Rudzicz · 28. Juli 2026
A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign adaptability remains difficult: the only prior method with an explicit cur…
- Covariance Last-Layer Ensembles: Function-Space Diversity for Efficient Uncertainty Quantification
H. Martin Gillis, Isaac Xu, Gabriel Spadon, Thomas Trappenberg · 28. Juli 2026
A Last-Layer Ensemble (LLE), $K$ linear units on one shared frozen feature map, is an efficient single-pass approach to the disagreement-based epistemic uncertainty for out-of-distribution (OOD) detection. Its weakness is that members share the backbone gradient and can converge toward the same func…
- DECAF: De-Clustering for Adaptive Representational Unlearning
Anjie Le, Can Peng, Hongcheng Guo, J. Alison Noble · 28. Juli 2026
Machine unlearning, which aims to remove the influence of specific training data from a trained model, is a key requirement for privacy, accountability, and adaptive deployment. We argue that many unlearning methods are vulnerable to a simple clustering attack, which can recover class structure in a…
- Visual Token Compression Enhances Robustness of MLLMs
Shishen Gu, Jiequan Cui, Wenbo Hu, Zenglin Shi, Zhenzhen Hu, Richang Hong · 28. Juli 2026
In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision and language modalities cannot be perfectly aligned, the misaligned visu…
- When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
Tong Zhang, Zexin Li, Simin Chen, Yun Peng · 28. Juli 2026
Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost…
- A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection
Ali Borji · 28. Juli 2026
Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer recognize the image. Prior work showed such examp…
- Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation
Mihai Suteu, Ovidiu Serban · 28. Juli 2026
Deep ensembles provide the most reliable uncertainty estimates in deep learning, but their cost grows linearly with the number of members. Implicit ensembles lower this cost by sharing a single backbone across members. Member diversity is a primary determinant of ensemble quality, yet no implicit en…
- Explainable Reinforcement Learning via Physics-Aware Policy Distillation
Shaker Al-Tamari, Waled Kadour · 28. Juli 2026
In safety-critical sectors such as robotics and automotive engineering, the deployment of Deep Reinforcement Learning (DRL) is often hindered by the black-box nature of deep neural networks. This lack of transparency poses significant challenges for regulatory compliance and human-agent trust. This …
- What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents
Shawn Ray · 28. Juli 2026
Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whether intervention changes future behavior. We separate three questions. First, relative to fixed oracle predicates, a deterministic gate enforces ex…
- Sampling Decisions: Exact Path-Space Correction, Prior Cancellation and Local-Boltzmann Guidance
Michael Chertkov, Sungsoo Ahn, Hamidreza Behjoo · 28. Juli 2026
How can a cheap but biased sequential, finite-horizon sampler over a discrete space be corrected so that its terminal output follows a prescribed Gibbs distribution? We formulate Sampling Decisions as a path-space relative-entropy projection on a growing autoregressive state graph. The unique prior-…
- Securing Multimodal AI through Internal Information Decomposition
Jehyeok Yeon, Hyeonjeong Ha, Qiusi Zhan, Heng Ji · 27. Juli 2026
Multimodal large language models introduce attack surfaces absent in unimodal systems: adversaries can distribute malicious intent across modalities to evade unimodal safeguards. This motivates using cross-modal consistency as a detection signal rather than inspecting each modality in isolation. Our…
