Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3552 artículos indexados
El estudio de la robustez adversarial en aprendizaje automático explora cómo los modelos reaccionan a perturbaciones deliberadas o a datos diseñados para engañarlos. Los trabajos recientes se interesan en la verificación de la estabilidad de las predicciones, la detección de fallos de seguridad o la adaptación de los clasificadores frente a ataques dirigidos. También abordan cuestiones como la calibración de los grandes modelos de lenguaje, la optimización de transferencias maliciosas o los mecanismos de desaprendizaje para corregir sesgos o comportamientos no deseados.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos38 % · 939 artículos
- China34 % · 841 artículos
- Reino Unido8,4 % · 205 artículos
- Alemania6,4 % · 158 artículos
- India5,2 % · 128 artículos
- Canadá5,1 % · 125 artículos
- Australia4,3 % · 106 artículos
- Corea del Sur3,7 % · 91 artículos
Sobre 2454 artículos de este tema con al menos un laboratorio localizado. 91 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers
Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University) · 2 de octubre de 2026
This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at $\epsilo…
- Towards Robust Numerical Claim Verification
Peter R{\o}ysland Aarnes, Vinay Setty · 2 de octubre de 2026
Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically pert…
- Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim · 2 de octubre de 2026
Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether succ…
- False Prophets: On the Security of World Models in Agentic Systems
Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck · 2 de octubre de 2026
Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent research proposes to enhance predictive capabilities via specially trained…
- UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao · 2 de octubre de 2026
Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship a…
- Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints
Melika Shirian, Kianoosh Vadaei · 2 de octubre de 2026
Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, wh…
- Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun · 2 de octubre de 2026
Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an ac…
- TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang · 2 de octubre de 2026
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient co…
- Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
Jianwei Li, Jung-Eun Kim · 2 de octubre de 2026
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, acces…
- Backdoor Containment via Expert Quarantine and Shutdown in LLMs
Jianwei Li, Min-Seon Kim, Jung-Eun Kim · 2 de octubre de 2026
Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppre…
- Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole
Cris Huynh · 2 de octubre de 2026
When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the signal that best describes it. On 108 confirmatory items, the adversary-robust optimum aligns exactly with the salience pole from prior work. Across a 200,000-item pool,…
- Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation
Ziqi Zhou, Yifan Hu, Yufei Song, Haowen Jiang, Xianlong Wang, Shengshan Hu, Dezhong Yao, Leo Yu Zhang · 1 de octubre de 2026
The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial exam…
- Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models
Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su · 1 de octubre de 2026
Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study adversarial images that preserve full-token correctness yet induce errors after compression, even when both inference paths succeed on the clean image…
- BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
Zhihang Wu, Zhongqi Wang, Jie Zhang, Fengming Gu, Shiguang Shan, Xilin Chen · 1 de octubre de 2026
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the inte…
- FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance
Rikhiya Ghosh, Himanshu Kumar, Sriram Venkatapathy, Sahil Wadhwa, Alexandre G. R. Day, Pranab Mohanty · 1 de octubre de 2026
In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety failures and pushing responses dangerously close to policy limits. Existing automated red-teaming methods trade off attack effectiveness against generati…
- Probabilistic Adversarial Training
Andi Zhang, Xingyu Zhao, Siddartha Khastgir · 1 de octubre de 2026
Building on a probabilistic perspective in which adversarial examples arise from the overlap between a distance-based distribution $p_{\mathrm{dis}}$ and a victim-classifier-induced distribution $p_{\mathrm{vic}}$, we start from a simple intuition: adversarial examples become harder to generate when…
- Let the Carrier Carry the Attack: Preserving the Subject in Adversarial Image Generation
Linfeng Jiang, Steven McDonagh, Yuhang Chen, Xingyu Zhao, Siddartha Khastgir, Andi Zhang · 1 de octubre de 2026
Strong unrestricted adversarial attacks can distort the primary object of an image, hereafter referred to as the subject. To preserve subject integrity without compromising attack magnitude, we introduce the carrier: a secondary visual element that provides an auxiliary region to facilitate the atta…
- When Masking Helps or Hurts Robustness in Compressed CLIP: A Pre-Deployment Diagnostic
Muhammad Zawish, Steven Davy · 1 de octubre de 2026
This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly …
- ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models
Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara · 1 de octubre de 2026
Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content …
- Faithful Dual-constrained Erasure for Robust LLM Safety Alignment
Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang · 1 de octubre de 2026
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed maliciou…
- Alignment via Training Against Probes Without Losing Monitorability
Lena Libon, Alexander Panfilov, Ben Rank, Xin Chen, Jonas Geiping, Maksym Andriushchenko · 1 de octubre de 2026
Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance durin…
- CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition
Zhiyi Mou, Yao Lu, Wangze Ni, Di Hong, Dakun Shen, Haoyang Li, Chen Jason Zhang, Alexander Zhou, Kui Ren · 1 de octubre de 2026
Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain inconspicu…
- Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard
Julian Schulz, Lukas F\"ulle, Rieke Fruengel · 1 de octubre de 2026
Chain-of-thought monitoring as an approach for AI oversight and control is threatened by the possibility of steganographic reasoning, where LLMs conceal their reasoning inside innocuous-looking text. Two neighbouring capabilities, steganographic messaging (passing a concealed message) and encoded re…
- Safety of Latent Communication in Multi-Agent Systems
Muhammad Huzaifa, Sina Mavali, Thorsten Eisenhofer · 1 de octubre de 2026
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the …
- Decoding One Safety Trigger Token for Balancing Safety and Usability in Large Language Models
Haoran Gu, Handing Wang, Yi Mei, Mengjie Zhang, Yaochu Jin · 30 de septiembre de 2026
Large Language Models (LLMs) have been extensively used across diverse domains, including virtual assistants, automated code generation, and scientific research. However, they remain vulnerable to jailbreak attacks, which manipulate the models into generating harmful responses despite safety alignme…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Large Language Models7407 artículos / 12 meses+247 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
- Natural Language Processing Techniques1595 artículos / 12 meses+69 %
