Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3164 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation
Dan Jacobellis, Neeraja J. Yadwadkar · 8 de mayo de 2026
Modern sensors generate rich, high-fidelity data, yet applications operating on wearable or remote sensing devices remain constrained by bandwidth and power budgets. Standardized codecs such as JPEG and MPEG achieve efficient trade-offs between bitrate and perceptual quality but are designed for hum…
- Mapping Human Anti-collusion Mechanisms to Multi-agent AI Systems
Jamiu Idowu, Ahmed Almasoud, Ayman Alfahid · 8 de mayo de 2026
As multi-agent AI systems become increasingly autonomous, evidence shows they can develop collusive strategies similar to those long observed in human markets and institutions. While human domains have accumulated centuries of anti-collusion mechanisms, it remains unclear how these can be adapted to…
- Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors
Richard Bergna, Stefan Depeweg, Jos\'e Miguel Hern\'andez-Lobato · 8 de mayo de 2026
Prior-Fitted Networks (PFNs) amortize Bayesian prediction by meta-learning over a synthetic task prior, but their standard output is a posterior predictive distribution over noisy observations. For sequential decision-making, such as active learning and Bayesian optimization, acquisition should prio…
- Uncertainty Estimation via Hyperspherical Confidence Mapping
Eunseo Choi, Ho-Yeon Kim, Jaewon Lee, Taeyong jo, Myungjun lee, Heejin Ahn · 8 de mayo de 2026
Quantifying uncertainty in neural network predictions is essential for high-stakes domains such as autonomous driving, healthcare, and manufacturing. While existing approaches often depend on costly sampling or restrictive distributional assumptions, we propose Hyperspherical Confidence Mapping (HCM…
- Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
Qirun Zeng, Eric He, Richard Hoffmann, Xuchuang Wang, Jinhang Zuo · 8 de mayo de 2026
Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting their relevance to real-world systems. We propose a more practical threat model, Fake Data Injection, which reflects realis…
- A Testable Certificate for Constant Collapse in Teacher-Guided VAEs
Zegu Zhang, Jianhua Peng, Jian Zhang · 8 de mayo de 2026
Posterior collapse in variational autoencoders is often diagnosed by its symptoms: a small KL term, a strong decoder, or weak use of the latent code. These signals are useful, but they do not define a collapse boundary. We study a concrete failure mode, input-independent constant collapse, and show …
- Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations
Yuan Du, Mitchel Hill, HanQin Cai · 8 de mayo de 2026
This work studies the robust evaluation of iterative stochastic purification defenses under white-box adversarial attacks. Our key technical insight is that gradient checkpointing makes exact end-to-end gradient computation through long purification trajectories practical by trading additional recom…
- Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text
Tom Kempton, Viktor Drobnyi, Maeve Madigan, Stuart Burrell · 8 de mayo de 2026
The ability to reliably distinguish human-written text from that generated by large language models is of profound societal importance. The dominant approach to this problem exploits the likelihood hypothesis: that machine-generated text should appear more probable to a detector language model than …
- Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
Charlie Griffin, Louis Thomson, Buck Shlegeris, Alessandro Abate · 8 de mayo de 2026
To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This paper introduces AI-Control Games, a formal decision-making model of the red-teaming exercise as a multi-objective, partia…
- FIT to Forget: Robust Continual Unlearning for Large Language Models
Xiaoyu Xu, Minxin Du, Kun Fang, Yaxin Xiao, Zhicong Huang, Cheng Hong, Qingqing Ye, Haibo Hu · 8 de mayo de 2026
While large language models (LLMs) exhibit remarkable capabilities, they increasingly face demands to unlearn memorized privacy-sensitive, copyrighted, or harmful content. Existing unlearning methods primarily focus on \emph{single-shot} scenarios, whereas real-world deletion requests arrive \emph{c…
- Position: Adopt Constraints Over Fixed Penalties in Deep Learning
Juan Ramirez, Meraj Hashemizadeh, Simon Lacoste-Julien · 8 de mayo de 2026
Recent efforts to develop trustworthy AI systems have increased interest in learning problems with explicit requirements, or constraints. In deep learning, however, such problems are often handled through fixed weighted-sum penalization: the constraints are added to the task loss with fixed coeffici…
- SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing
Jialong Sun, Zeming Wei, Jiaxuan Zou, Jiacheng Gong, Jie Fu, Chengyang Dong, Heng Xu, Jialong Li, Bo Liu · 8 de mayo de 2026
Machine unlearning (MU) is essential for enforcing the right to be forgotten in machine learning systems. A key challenge of MU is how to reliably audit whether a model has truly forgotten specified training data. Membership Inference Attacks (MIAs) are widely used for unlearned model auditing, wher…
- Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
Bo Wang, Jia Ni, Mengnan Zhao, Zhan Qin, Kui Ren · 8 de mayo de 2026
The unauthorized use of personal data in model training has emerged as a growing privacy threat. Unlearnable examples (UEs) address this issue by embedding imperceptible perturbations into benign examples to obstruct feature learning. However, existing studies mainly evaluate UEs under from-scratch …
- Information Theoretic Adversarial Training of Large Language Models
Yiwei Zhang, Jeremiah Birrell, Reza Ebrahimi, Rouzbeh Behnia, Jason Pacheco, Elisa Bertino · 8 de mayo de 2026
Large language models (LLMs) remain vulnerable to adversarial prompting despite advances in alignment and safety, often exhibiting harmful behaviors under novel attack strategies. While adversarial training can improve robustness, existing approaches are computationally expensive and difficult to sc…
- DeTrigger: A Gradient-Centric Approach to Backdoor Attack Mitigation in Federated Learning
Kichang Lee, Yujin Shin, Jonghyuk Yun, Songkuk Kim, Jun Han, JeongGil Ko · 8 de mayo de 2026
Federated Learning (FL) enables collaborative model training across distributed devices while preserving local data privacy, making it ideal for mobile and embedded systems. However, the decentralized nature of FL also opens vulnerabilities to model poisoning attacks, particularly backdoor attacks, …
- Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
Yuntai Bao, Qinfeng Li, Xinyan Yu, Xuhong Zhang, Ge Su, Wenqi Zhang, Liu Yan, Haiqin Weng, Jianwei Yin · 8 de mayo de 2026
Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones. However, current approaches to fine-tuned SVs suffer from two limitations. First, they…
- Efficient Techniques for Data Reconstruction, with Finite-Width Recovery Guarantees
Edward Tansley, Roy Makhlouf, Estelle Massart, Coralia Cartis · 8 de mayo de 2026
Data reconstruction attacks on trained neural networks aim to recover the data on which the network has been trained and pose a significant threat to privacy, especially if the training dataset contains sensitive information. Here, we propose a unified optimization formulation of the data reconstruc…
- Laundering AI Authority with Adversarial Examples
Jie Zhang, Pura Peetathawatchai, Florian Tram\`er, Avital Shafran · 7 de mayo de 2026
Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial examples break this assu…
- Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges
Ali Al-Kaswan, Maksim Plotnikov, Maxim H\'ajek, Roland V\'izner, Arie van Deursen, Maliheh Izadi · 7 de mayo de 2026
Large Language Model (LLM) agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings remain poorly understood. We present DeepRed, an open-source benchmark for evaluating LLM-based agents on realistic Capture The Flag (CTF) challenges…
- On the Hardness of Junking LLMs
Marco Rando, Samuel Vaiter · 7 de mayo de 2026
Large language models (LLMs) are known to be vulnerable to jailbreak attacks, which typically rely on carefully designed prompts containing explicit semantic structure. These attacks generally operate by fixing an adversarial instruction and optimizing small adversarial components (e.g., suffixes or…
- Gray-Box Poisoning of Continuous Malware Ingestion Pipelines
Jan Dolej\v{s}, Martin Jure\v{c}ek, R\'obert L\'orencz · 7 de mayo de 2026
Modern malware detection pipelines rely on continuous data ingestion and machine learning to counter the high volume of novel threats. This work investigates a realistic gray-box poisoning threat model targeting these pipelines. Using the secml_malware framework, we generate problem-space adversaria…
- From Beats to Breaches:How Offensive AI Infers Sensitive User Information from Playlists
Stefano Cecconello, Mauro Conti, Luca Pajola, Luca Pasa, Pier Paolo Tricomi · 7 de mayo de 2026
The pervasive integration of AI has enabled Offensive AI: the exploitation of AI for malicious ends across the cyber-kill chain. A critical manifestation is the user attribute inference attack, where AI infers sensitive Personally Identifiable Information (PII) from innocuous public data. We explore…
- Gradients with Respect to Semantics Preserving Embeddings Tell the Uncertainty of Large Language Models
Mingda Li, Rundong Lv, Xinyu Li, Weinan Zhang, Ting Liu · 7 de mayo de 2026
Uncertainty quantification (UQ) is an important technique for ensuring the trustworthiness of LLMs, given their tendency to hallucinate. Existing state-of-the-art UQ approaches for free-form generation rely heavily on sampling, which incurs high computational cost and variance. In this work, we prop…
- Coward: Collision-based OOD Watermarking for Practical Proactive Federated Backdoor Detection
Wenjie Li, Siying Gu, Yiming Li, Shuxin Li, Zhili Chen, Tianwei Zhang, Shu-Tao Xia · 7 de mayo de 2026
Backdoor detection is currently the mainstream defense against backdoor attacks in federated learning (FL), where a small number of malicious clients can upload poisoned updates to compromise the federated global model. Existing backdoor detection techniques fall into two categories, passive and pro…
- Dissociating spatial frequency reliance from adversarial robustness advantages in neurally guided deep convolutional neural networks
Zhenan Shao, Tianyu Ren, Chengxiao Wang, Leyla Isik, Diane M. Beck · 7 de mayo de 2026
Deep convolutional neural networks (DCNNs) have rivaled humans on many visual tasks, yet they remain vulnerable to near-imperceptible perturbations generated by adversarial attacks. Recent work shows that aligning DCNN representations with human visual cortex activity improves adversarial robustness…
