Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3164 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- State Contamination in Memory-Augmented LLM Agents
Yian Wang, Agam Goyal, Yuen Chen, Hari Sundaram · 19 de mayo de 2026
LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon interaction. This makes safety depend not only on individual model outputs, but also on what an agent stores and later reuses. We study a failure mode we…
- Quantitative Linear Logic for Neuro-Symbolic Learning and Verification
Thomas Flinkow, Ekaterina Komendantskaya, Matteo Capucci, Rosemary Monahan · 19 de mayo de 2026
Differentiable Logics are deployed in neuro-symbolic learning tasks as a way of embedding logical constraints in the training objective of neural networks. A differentiable logic consists of a syntax to write logical properties and a semantics to interpret them as real-valued functions to be folded …
- Membership Inference Attacks on Discrete Diffusion Language Models
Shailesh Kasivelrajan · 19 de mayo de 2026
Masked Diffusion Language Models MDLMs replace autoregressive generation with iterative demasking and their privacy properties are largely unstudied. We study membership inference attacks MIA on fine tuned MDLMs and show they are significantly more vulnerable than current grey box baselines suggest.…
- Compositional Adversarial Training for Robust Visual Watermarking
Anirudh Satheesh, Michael-Andrei Panaitescu-Liess, Andrew Xu, Georgios Milis, Heng Huang, Zikui Cai, Furong Huang · 19 de mayo de 2026
Robust watermarking is typically trained with random post-processing augmentation, but random sampling under-covers the combinatorial space of realistic attack pipelines and rarely encounters the rare compositions that actually break detection. This leads to unstable training and poor sample efficie…
- Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
Miles Q. Li, Benjamin C. M. Fung, Boyang Li, Heba Ismail, Farkhund Iqbal · 19 de mayo de 2026
The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a proliferation of safety benchmarks since late 2023. However, these benchmarks have developed independently, with inconsistent threat models, incompatible metri…
- Stress-Testing Neural Network Verifiers with Provably Robust Instances
David Troxell, Yulia Alexandr, Sofia Hunt, Stephanie Lei, Guido Mont\'ufar · 19 de mayo de 2026
Neural network verifiers aim to provide formal guarantees on model behavior, but existing verification benchmarks are fundamentally limited by their lack of ground-truth labels. As a result, verifier evaluation relies on indirect heuristics, which prevents exact scoring and systematic study of verif…
- AI Agents May Always Fall for Prompt Injections
Sahar Abdelnabi, Eugene Bagdasarian · 19 de mayo de 2026
Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data-instruction separation) both fails to detect attacks that operate through contextual manipulation and degrades contextually appropriate behavior. We …
- When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning
Arahan Kujur · 19 de mayo de 2026
We study adversarial action masking in self-play reinforcement learning: an attacker selectively removes legal actions from a victim's action set. Unlike observation or action perturbations, removal eliminates decision options before the agent acts. Across poker games scaling from 6 to 5,531 informa…
- Testable and Actionable Calibration for Full Swap Regret
Konstantina Bairaktari, Lunjia Hu, Huy L. Nguyen, Jonathan Ullman · 19 de mayo de 2026
AI generated predictions increasingly inform decision making in critical tasks, and therefore must be trustworthy. One widely used measure of trustworthiness is calibration, which requires that the predictions match the true frequencies and can be treated like real probabilities of a given outcome. …
- Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
Qinwu Xu · 19 de mayo de 2026
Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling. We propose a stage-wi…
- Catastrophic Overfitting, Entropy Gap and Participation Ratio: A Noiseless $l^p$ Norm Solution for Fast Adversarial Training
Fares B. Mehouachi, Saif Eddin Jabari · 19 de mayo de 2026
Adversarial training is a cornerstone of robust deep learning, but fast methods like the Fast Gradient Sign Method (FGSM) often suffer from Catastrophic Overfitting (CO), where models become robust to single-step attacks but fail against multi-step variants. While existing solutions rely on noise in…
- Universal Adversarial Triggers
Benedict Florance Arockiaraj, Alexander Feng, Jianxiong Cai, Xiaoyu Cheng · 19 de mayo de 2026
Recent works have illustrated that modern NLP models trained for diverse tasks ranging from sentiment analysis to language generation succumb to universal adversarial attacks, a class of input-agnostic attacks where a common trigger sequence is used to attack the model. Although these attacks are su…
- A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
Mohamed elShehaby, Ashraf Matrawy · 19 de mayo de 2026
Gradient-based adversarial attacks subtly manipulate inputs of Machine Learning (ML) models to induce incorrect predictions. This paper investigates whether careful architectural choices alone can yield an inherently robust Deep Neural Network (DNN)-based Network Intrusion Detection Systems (NIDS), …
- CADS: Conformal Adaptive Decision System for Cost-Efficient Image Classification
Turkoglu Mikael, Bary Tim, Thielens Vincent, Dausort Manon, Macq Beno\^it · 19 de mayo de 2026
While high-capacity AI models have advanced state-of-the-art performance, their practical deployment is often hindered by high inference costs, environmental impact, and a "one-size-fits-all" approach that ignores varying sample complexity. In clinical settings for instance, the waste of computation…
- Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings
Anthonio Oladimeji Gabriel, Ahmad Rufai Yusuf · 19 de mayo de 2026
Current clinical artificial intelligence (AI) systems are evaluated almost exclusively on clean, standardised, English-language inputs, conditions that do not reflect the realities of healthcare delivery in low-resource settings. This study presents the first systematic dual audit of two orthogonal …
- Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning
Sunwoo Lee, Mingu Kang, Yonghyeon Jo, Seungyul Han · 19 de mayo de 2026
Cooperation is central to multi-agent reinforcement learning (MARL), yet learned coordination can be fragile when external perturbations disrupt inter-agent interactions. Prior robust MARL methods have primarily considered value-oriented attacks, leaving a gap in robustness when interaction structur…
- Bug or Feature$^2$: Weight Drift, Activation Sparsity, and Spikes
Egor Shvetsov, Aleksandr Serkov, Shokorov Viacheslav, Redko Dmitry, Vladislav Goloshchapov, Evgeny Burnaev · 19 de mayo de 2026
The design of modern neural architectures has converged through incremental empirical choices, yet the mechanisms governing their training dynamics remain only partially understood. We identify and analyze a negative weight drift induced by the interaction between standard losses and positively bias…
- No Free Swap: Protocol-Dependent Layer Redundancy in Transformers
Gabriel Garcia · 19 de mayo de 2026
When researchers ask whether two transformer layers are "equivalent" for compression, they often conflate distinct tests. Replacement asks whether one layer's map can substitute for another's in place; interchange asks whether two layers approximately commute when their positions are swapped. Both a…
- Enabling Adversarial Robustness in AI Models through Kubeflow MLOps
Stavros Bouras, Ioannis Korontanis, Antonios Makris, Konstantinos Tserpes · 18 de mayo de 2026
AI models are increasingly deployed in cloud-native environments to support scalable and automated services. However, while platforms such as Kubernetes provide strong infrastructure orchestration, security mechanisms specifically designed to protect deployed AI models remain limited. This paper pre…
- Learning Context-conditioned Gaussian Overbounds for Convolution-Based Uncertainty Propagation
Ruirui Liu, Xuejie Hou, Yiping Jiang, Hui Ren · 18 de mayo de 2026
Uncertainty quantification is essential in safety-critical settings--from autonomous driving to aviation, finance, and health--where decisions must rely on conservative bounds rather than point estimates. Predictor-level intervals (e.g., from quantile regression, conformal prediction, variance netwo…
- The Geometric Structure of Models Learning Sparse Data
Thomas Walker, T. Mitchell Roddenberry, Ahmed Imtiaz Humayun, Randall Balestriero, Richard Baraniuk · 18 de mayo de 2026
The manifold hypothesis (MH) is often used to explain how machine learning can overcome the curse of dimensionality. However, the MH is only applicable in regimes where the training data provides a sufficiently dense sample of the underlying low-dimensional data manifold, or where such a low-dimensi…
- From Weight Perturbation to Feature Attribution for Explaining Fully Connected Neural Networks
Thodoris Lymperopoulos, Denia Kanellopoulou · 18 de mayo de 2026
Fully Connected Neural Networks (FCNNs) are often regarded as simple and intuitive architectures, yet they serve as the foundation for more complex models. Nonetheless, the lack of consensus on their interpretability continues to pose challenges, underscoring the enduring relevance of simpler, attri…
- Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
Yu Fu, Longxuan Yu, Haz Sameen Shahgir, Zhipeng Wei, Hui Liu, N. Benjamin Erichson, Yue Dong · 18 de mayo de 2026
Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is distributional mismatch: supervised fine-tuning trains the target model on safety demonstrations produced by humans, external models, or fixed self-ge…
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
Jehyeok Yeon, Federico Cinus, Yifan Wu, Luca Luceri · 18 de mayo de 2026
Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity objective treats latent features as independent. This prior can be poorly matched to high-level safety behaviors, where refusal and harmful compliance appear to …
- Training on Documents About Monitoring Leads to CoT Obfuscation
Reilly Haskins, Bilal Chughtai, Joshua Engels · 18 de mayo de 2026
Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning. Motivated by this vulnerability, we study whether monitor-aware models are capable of obfuscating their reas…
