Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3169 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Perturbing the Phase: Analyzing Adversarial Robustness of Complex-Valued Neural Networks
Florian Eilers, Christof Duhme, Xiaoyi Jiang · 9 de febrero de 2026
Complex-valued neural networks (CVNNs) are rising in popularity for all kinds of applications. To safely use CVNNs in practice, analyzing their robustness against outliers is crucial. One well known technique to understand the behavior of deep neural networks is to investigate their behavior under a…
- Don't Break the Boundary: Continual Unlearning for OOD Detection Based on Free Energy Repulsion
Ningkang Peng, Kun Shao, Jingyang Mao, Linjing Qian, Xiaoqian Peng, Xichen Yang, Yanhui Gu · 9 de febrero de 2026
Deploying trustworthy AI in open-world environments faces a dual challenge: the necessity for robust Out-of-Distribution (OOD) detection to ensure system safety, and the demand for flexible machine unlearning to satisfy privacy compliance and model rectification. However, this objective encounters a…
- Endogenous Resistance to Activation Steering in Language Models
Alex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab, Murat Cubuktepe, Mike Vaiana, Diogo de Lucena, Judd Rosenblatt, Michael S. A. Graziano · 9 de febrero de 2026
Large language models can resist task-misaligned activation steering during inference, sometimes recovering mid-generation to produce improved responses even when steering remains active. We term this Endogenous Steering Resistance (ESR). Using sparse autoencoder (SAE) latents to steer model activat…
- Time-uniform conformal and PAC prediction
Kayla E. Scharfstein, Arun Kumar Kuchibhotla · 9 de febrero de 2026
Given that machine learning algorithms are increasingly being deployed to aid in high stakes decision-making, uncertainty quantification methods that wrap around these black box models such as conformal prediction have received much attention in recent years. In sequential settings, where data are o…
- TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
Sung-Hoon Yoon, Ruizhi Qian, Minda Zhao, Weiyue Li, Mengyu Wang · 9 de febrero de 2026
Large Language Models (LLMs) have become integral to many domains, making their safety a critical priority. Prior jailbreaking research has explored diverse approaches, including prompt optimization, automated red teaming, obfuscation, and reinforcement learning (RL) based methods. However, most exi…
- REBEL: Hidden Knowledge Recovery via Evolutionary-Based Evaluation Loop
Patryk Rybak, Pawe{\l} Batorski, Paul Swoboda, Przemys{\l}aw Spurek · 9 de febrero de 2026
Machine unlearning for LLMs aims to remove sensitive or copyrighted data from trained models. However, the true efficacy of current unlearning methods remains uncertain. Standard evaluation metrics rely on benign queries that often mistake superficial information suppression for genuine knowledge re…
- Temperature Scaling Attack Disrupting Model Confidence in Federated Learning
Kichang Lee, Jaeho Jin, JaeYeon Park, Songkuk Kim, JeongGil Ko · 9 de febrero de 2026
Predictive confidence serves as a foundational control signal in mission-critical systems, directly governing risk-aware logic such as escalation, abstention, and conservative fallback. While prior federated learning attacks predominantly target accuracy or implant backdoors, we identify confidence …
- TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, Matthew Kowal, Nayeema Nonta, Samuel Simko, Stephen Casper, Zhijing Jin, Kellin Pelrine, Sirisha Rambhatla · 9 de febrero de 2026
As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks. However, there is no standard approach to evaluate tamper resistance. Varied data sets…
- Is Gradient Ascent Really Necessary? Memorize to Forget for Machine Unlearning
Zhuo Huang, Qizhou Wang, Ziming Hong, Shanshan Ye, Bo Han, Tongliang Liu · 9 de febrero de 2026
For ethical and safe AI, machine unlearning rises as a critical topic aiming to protect sensitive, private, and copyrighted knowledge from misuse. To achieve this goal, it is common to conduct gradient ascent (GA) to reverse the training on undesired data. However, such a reversal is prone to catast…
- MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs
Junhyeok Lee, Han Jang, Kyu Sung Choi · 9 de febrero de 2026
Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly integrated into clinical workflows; however, prompt injection attacks can steer these systems toward clinically unsafe or misleading outputs. We introduce the Medical Prompt Injection Benchmark (MPIB), a d…
- Action Hallucination in Generative Visual-Language-Action Models
Harold Soh, Eugene Lim · 9 de febrero de 2026
Robot Foundation Models such as Vision-Language-Action models are rapidly reshaping how robot policies are trained and deployed, replacing hand-designed planners with end-to-end generative action models. While these systems demonstrate impressive generalization, it remains unclear whether they funda…
- GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Mark Russinovich, Yanan Cai, Keegan Hines, Giorgio Severi, Blake Bullwinkel, Ahmed Salem · 9 de febrero de 2026
Safety alignment is only as robust as its weakest failure mode. Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-deployment fine-tuning. However, these methods often require extensive data curation and degrade model utility. In thi…
- Empirical Analysis of Adversarial Robustness and Explainability Drift in Cybersecurity Classifiers
Mona Rajhans, Vishal Khawarey · 9 de febrero de 2026
Machine learning (ML) models are increasingly deployed in cybersecurity applications such as phishing detection and network intrusion prevention. However, these models remain vulnerable to adversarial perturbations small, deliberate input modifications that can degrade detection accuracy and comprom…
- PurSAMERE: Reliable Adversarial Purification via Sharpness-Aware Minimization of Expected Reconstruction Error
Vinh Hoang, Sebastian Krumscheid, Holger Rauhut, Ra\'ul Tempone · 9 de febrero de 2026
We propose a novel deterministic purification method to improve adversarial robustness by mapping a potentially adversarial sample toward a nearby sample that lies close to a mode of the data distribution, where classifiers are more reliable. We design the method to be deterministic to ensure reliab…
- Agentic Uncertainty Reveals Agentic Overconfidence
Jean Kaddour, Srijan Patel, Gb\`etondji Dovonon, Leo Richter, Pasquale Minervini, Matt J. Kusner · 9 de febrero de 2026
Can AI agents predict whether they will succeed at a task? We study agentic uncertainty by eliciting success probability estimates before, during, and after task execution. All results exhibit agentic overconfidence: some agents that succeed only 22% of the time predict 77% success. Counterintuitive…
- Exploring Sparsity and Smoothness of Arbitrary $\ell_p$ Norms in Adversarial Attacks
Christof Duhme, Florian Eilers, Xiaoyi Jiang · 9 de febrero de 2026
Adversarial attacks against deep neural networks are commonly constructed under $\ell_p$ norm constraints, most often using $p=1$, $p=2$ or $p=\infty$, and potentially regularized for specific demands such as sparsity or smoothness. These choices are typically made without a systematic investigation…
- Optimal Abstractions for Verifying Properties of Kolmogorov-Arnold Networks (KANs)
Noah Schwartz, Chandra Kanth Nagesh, Sriram Sankaranarayanan, Ramneet Kaur, Tuhin Sahai, Susmit Jha · 9 de febrero de 2026
We present a novel approach for verifying properties of Kolmogorov-Arnold Networks (KANs), a class of neural networks characterized by nonlinear, univariate activation functions typically implemented as piecewise polynomial splines or Gaussian processes. Our method creates mathematical ``abstraction…
- AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
Fengpeng Li, Kemou Li, Qizhou Wang, Bo Han, Jiantao Zhou · 9 de febrero de 2026
Concept erasure helps stop diffusion models (DMs) from generating harmful content; but current methods face robustness retention trade off. Robustness means the model fine-tuned by concept erasure methods resists reactivation of erased concepts, even under semantically related prompts. Retention mea…
- SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory
Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu · 9 de febrero de 2026
Reliable uncertainty quantification (UQ) is essential for deploying large language models (LLMs) in safety-critical scenarios, as it enables them to abstain from responding when uncertain, thereby avoiding hallucinations, i.e., plausible yet factually incorrect responses. However, while semantic UQ …
- Private and interpretable clinical prediction with quantum-inspired tensor train models
Jos\'e Ram\'on Pareja Monturiol, Juliette Sinnott, Roger G. Melko, Mohammad Kohandel · 9 de febrero de 2026
Machine learning in clinical settings must balance predictive accuracy, interpretability, and privacy. Models such as logistic regression (LR) offer transparency, while neural networks (NNs) provide greater predictive power; yet both remain vulnerable to privacy attacks. We empirically assess these …
- Refining the Information Bottleneck via Adversarial Information Separation
Shuai Ning, Zhenpeng Wang, Lin Wang, Bing Chen, Shuangrong Liu, Xu Wu, Jin Zhou, Bo Yang · 9 de febrero de 2026
Generalizing from limited data is particularly critical for models in domains such as material science, where task-relevant features in experimental datasets are often heavily confounded by measurement noise and experimental artifacts. Standard regularization techniques fail to precisely separate me…
- Algebraic Robustness Verification of Neural Networks
Yulia Alexandr, Hao Duan, Guido Mont\'ufar · 9 de febrero de 2026
We formulate formal robustness verification of neural networks as an algebraic optimization problem. We leverage the Euclidean Distance (ED) degree, which is the generic number of complex critical points of the distance minimization problem to a classifier's decision boundary, as an architecture-dep…
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink
Guozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo, Qi Mu, Xiumin Wang, Li Shen · 6 de febrero de 2026
Harmful fine-tuning can invalidate safety alignment of large language models, exposing significant safety risks. In this paper, we utilize the attention sink mechanism to mitigate harmful fine-tuning. Specifically, we first measure a statistic named \emph{sink divergence} for each attention head and…
- Agent2Agent Threats in Safety-Critical LLM Assistants: A Human-Centric Taxonomy
Lukas Stappen, Ahmet Erkan Turan, Johann Hagerer, Georg Groh · 6 de febrero de 2026
The integration of Large Language Model (LLM)-based conversational agents into vehicles creates novel security challenges at the intersection of agentic AI, automotive safety, and inter-agent communication. As these intelligent assistants coordinate with external services via protocols such as Googl…
- LeakBoost: Perceptual-Loss-Based Membership Inference Attack
Amit Kravchik Taub, Fred M. Grabovski, Guy Amit, Yisroel Mirsky · 6 de febrero de 2026
Membership inference attacks (MIAs) aim to determine whether a sample was part of a model's training set, posing serious privacy risks for modern machine-learning systems. Existing MIAs primarily rely on static indicators, such as loss or confidence, and do not fully leverage the dynamic behavior of…
