Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3164 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models
Tao Qi, Huili Wang, Yuanhong Huang, Wendan Wang, Lianchao Zhao, Jinrui Wang, Zichen Qin, Shangguang Wang, Yongfeng Huang · 27 de mayo de 2026
The rapid advancement of diffusion-based image generation models has raised serious concerns regarding potential copyright and privacy infringements involving human-created data. Membership inference attacks (MIAs) have emerged as a promising tool for identifying unauthorized data usage during model…
- Probabilistic Smoothing with Ratio-Monotone Transforms for Global Optimization
Kukyoung Jang, Taehyun Cho, Junrui Zhang, Ping Xu, Kyungjae Lee · 27 de mayo de 2026
Probabilistic smoothing is a standard tool for global optimization, but existing methods rely on Gaussian kernels and specific transforms, often resulting in strong hyperparameter sensitivity and limited robustness. We propose a general smoothing framework that combines flexible symmetric unimodal k…
- Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models
Arian Komaei Koma, Seyed Amir Kasaei, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban · 27 de mayo de 2026
Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been introduced to make the model generate such unlearned concepts. These attacks, nevertheless, do not assume a realistic threat model, i.e. they either …
- Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
Yedidia Agnimo, Anna Korba, Annabelle Blangero, Nicolas Chesneau, Karteek Alahari · 27 de mayo de 2026
Large language models (LLMs) are prone to hallucinations, i.e., statements unsupported by the input or training data, hindering reliable deployment. In parallel, numerous uncertainty estimation (UE) methods have been proposed to quantify model confidence and are often implicitly treated as proxies f…
- When Muon Optimizer Meets Adversarial Training: A Theoretical and Empirical Study
Jun Yan, Weiquan Huang, Jiankai Zuo, Yujian Mo, Xi Fang, Chengliang Wu, Zeming Wei · 27 de mayo de 2026
Adversarial training (AT) remains one of the most reliable empirical defenses against adversarial attacks. Its robustness critically depends on how the underlying min-max objective is optimized. In practice, Stochastic Gradient Descent (SGD) optimizer remains the default optimization choice for AT, …
- Detectability in Diversity: Improved Canary Crafting for Privacy Auditing in One Run
Mathieu Dagr\'eou, Aur\'elien Bellet · 27 de mayo de 2026
Privacy auditing aims to empirically assess privacy leakage in machine learning models using membership inference attacks (MIAs), and to derive lower bounds on differential privacy (DP) parameters. Recent one-run auditing methods address the high cost of standard approaches by relying on a single tr…
- Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
Kia-J\"ung Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp · 27 de mayo de 2026
Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal. Unlike instruction-tuned LLMs, where refusal is mediated by a single directional subspace, refusal in larg…
- Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
Tongxi Wu, Jian Zhang, Yang Gao · 27 de mayo de 2026
Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshold mechanism. We challenge this assumption by revealing that safety behavior is governed by an instability region where small perturbations induce stoc…
- Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial
Haoyu Li, Xiangru Zhong, Hao Cheng, Bin Hu, Huan Zhang · 27 de mayo de 2026
Learning-based methods for synthesizing controllers have gained popularity due to their high expressiveness and strong empirical performance. However, in safety-critical scenarios such as autonomous driving, robotics, and power systems, empirical performance alone is insufficient, and formal verific…
- ReasonOps: A Unified Operational Paradigm for Trustworthy Verified LLM Reasoning
Adnan Rashid · 27 de mayo de 2026
Large Language Models (LLMs) have transformed artificial intelligence from primarily generative systems into increasingly capable reasoning agents. Recent advances in theorem proving, autoformalization, symbolic reasoning, and tool-augmented language models demonstrate substantial progress toward ma…
- Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
Kevin Kuo, Chhavi Yadav, Virginia Smith · 27 de mayo de 2026
Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode…
- Robust Koopman Control Barrier Filters for Safe Actor-Critic Reinforcement Learning
Dhruv S. Kushwaha, Zoleikha A. Biron · 27 de mayo de 2026
Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints during both training and deployment. Control barrier functions (CBFs) provide a principled mechanism for enforcing forward invariance through minimally in…
- Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting Attacks
Lin Mu, Guowei Chu, Li Ni, Lei Sang, Yiwen Zhang · 27 de mayo de 2026
Large Language Models (LLMs) have demonstrated remarkable performance across various tasks by effectively utilizing a prompting strategy. However, they are highly sensitive to input perturbations, such as typographical errors or slight character order errors, which can significantly impair their per…
- PRBench: A Standardized Probabilistic Robustness Benchmark
Yi Zhang, Zheng Wang, Zhen Chen, Wenjie Ruan, Qing Guo, Siddartha Khastgir, Carsten Maple, Xingyu Zhao · 27 de mayo de 2026
Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which evaluates models under worst-case scenarios by examining the existence of deterministic adversarial examples (AEs). In contrast, probabilistic robustne…
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
Changyue Jiang, Wenqi Zhang, Xudong Pan, Geng Hong, Min Yang · 27 de mayo de 2026
LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly shapes subsequent actions. Small deviations in these thoughts can therefore propagate into unsafe behaviors, yet existing guardrails typically operate onl…
- Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection
Nico Steckhan, Krutarth Prajapati, Weija Shao, Silvia Vock · 27 de mayo de 2026
Testing object detectors in safety-critical domains requires semantically meaningful probes beyond pixel-level corruptions. We present SemProbe, a tool for semantic robustness probing: users upload deployment images, create masks manually or automatically, select operational design domain-derived fa…
- Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)
Paul Sigloch, Christoph Benzm\"uller · 27 de mayo de 2026
LLMs deployed in high-stakes domains face fundamental reliability challenges: hallucinations, inconsistencies, and privacy vulnerabilities introduce unacceptable risks where errors carry legal, financial, or safety consequences. This paper presents a hybrid verification architecture combining formal…
- Position: AI Safety Requires Effective Controllability
Yige Li, Yunhao Feng, Jun Sun · 27 de mayo de 2026
AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing has improved the behavior of modern language models, but aligned behavior does not by itself guarantee that a deployed agent can be stopped, overridde…
- Cordon-MAS: Defending RAG against Knowledge Poisoning via Information-Flow Control
Zhe Yu, Wenpeng Xing, Gaolei Li, Shuguang Xiong, Hongzhi Wang, Xuyang Teng, Meng Han · 27 de mayo de 2026
Retrieval-augmented generation (RAG) increasingly underpins high-stakes applications, yet remains vulnerable to Confundo-style poisoning where adversarially optimized documents manipulate generated outputs. Existing defenses assume that detecting poisoned evidence prevents harm. We show this assumpt…
- Structure-Adaptive Conformal Inference for Large-Scale Out-of-Distribution Testing
Rongyi Sun, Wenguang Sun, Zinan Zhao · 27 de mayo de 2026
This paper addresses structured out-of-distribution (OOD) testing in high-stakes machine learning applications. Traditional conformal methods rely on joint exchangeability, making it difficult to incorporate auxiliary information such as spatiotemporal or grouping structures. To overcome this limita…
- Curriculum Learning for Safety Alignment
Sandeep Kumar, Virginia Smith, Chhavi Yadav · 27 de mayo de 2026
Direct Preference Optimisation (DPO) is widely used for safety alignment in large language models. However, prior work shows it is brittle and exhibits poor out-of-distribution (OOD) generalisation. In this paper, we investigate whether Curriculum Learning can improve the robustness of DPO-based saf…
- Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
Zedian Shao, Charles Fleming, Teodora Baluta · 27 de mayo de 2026
Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely on fixed trigger phrases that defenses such as outlier detection, clean-data regularization, or online monitoring can neutralize. In this paper, we prop…
- Measuring Prediction Uncertainty in Neural Cellular Automata
Ario Sadafi, Michael Deutges, Nassir Navab, Carsten Marr · 27 de mayo de 2026
Neural cellular automata (NCA) provide a lightweight alternative to encoder-decoder segmentation networks. However, it can be difficult to decide when a prediction should be trusted. Here, we study uncertainty estimation for NCA-based medical image segmentation without modifying the underlying archi…
- Variational Inference for Evidential Deep Learning
Jiawei Tang, Xinyan Du, Hui Liu, Junhui Hou, Yuheng Jia · 27 de mayo de 2026
While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mitigates this by formulating predictions as a Dirichlet distribution over class probabilities to explicitly quantify epistemic uncertainty. However, …
- VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
Qilin Liao, Anamika Lochab, Ruqi Zhang · 27 de mayo de 2026
Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplored vulnerabilities. Existing multimodal red-teaming methods largely rely on brittle templates, focus on single-attack settings, and expose only a narrow subse…
