Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3164 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Multi-Task Bayesian In-Context Learning
Qingyang Zhu, Eric Karl Oermann, Kyunghyun Cho · 19 de junio de 2026
Bayesian predictive inference provides a principled framework for uncertainty quantification, data efficiency, and robust generalization. However, exact inference is often intractable, and scalable approximations may remain computationally expensive or require restrictive modeling assumptions that d…
- Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
Reza Soosahabi, Vivek Namsani · 19 de junio de 2026
Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequential, especially as attackers adopt model-guided automation…
- Adversarial Dependence Minimization
Pierre-Fran\c{c}ois De Plaen, Tinne Tuytelaars, Marc Proesmans, Luc Van Gool · 19 de junio de 2026
Minimally redundant representations are typically learned by minimizing feature covariance. However, covariance-based methods fail to eliminate all dependencies/redundancies, as linearly uncorrelated variables can still exhibit nonlinear relationships. To address this, we introduce ADM, a differenti…
- LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems
Hanwool Lee, Dasol Choi, Bokyeong Kim, Seung Geun Kim, Haon Park · 19 de junio de 2026
Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a benchmark for multi-turn red-teaming of LLM agents acting as op…
- Shifting-based Optimizable Linear Relaxations for General Activation Functions
Philipp Kern, L\'aszl\'o Antal, Erika \'Abr\'aham, Carsten Sinz · 19 de junio de 2026
The use of neural networks (NNs) is rapidly increasing, including in safety- and security-critical domains. To provide formal guarantees about NN behavior, many verification methods rely on optimizable linear relaxations of activation functions. However, existing techniques depend on hand-crafted re…
- Efficient and Sound Probabilistic Verification for AI Agents
Alaia Solko-Breslin, Pramod Kaushik Mudrakarta, Mihai Christodorescu, Somesh Jha, Krishnamurthy Dj Dvijotham · 19 de junio de 2026
Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulate and enforce policies expressed in a formal language like Datalog offer a promising solution. However, existing approaches are restricted to deterministic polic…
- SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling
Haotian Xu, Zeyang Zhang, Linbao Li, Huadi Zheng, Yu Li, Cheng Zhuo · 19 de junio de 2026
Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses are largely incompatible with speculative inference: they either introduce additional computation or disrupt the draft-verify mechanism, negating acceleration be…
- AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Faqiang Qian, Kang An, Weikun Zhang, Ziliang Wang, Xuhui Zheng, Liangjian Wen, Yong Dai, Mengya Gao, Yichao Wu · 19 de junio de 2026
Post-training alignment of large language models often combines supervised fine-tuning (SFT) on expert demonstrations with reinforcement learning (RL) from preference or verifiable feedback. SFT provides a useful behavioral anchor but can overfit to static demonstrations, whereas RL encourages explo…
- Toward Calibrated Mixture-of-Experts Under Distribution Shift
Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa, Anqi Liu · 19 de junio de 2026
Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve ensemble accuracy and calibration,…
- On the QUEST for Uncertainty Quantification via Highest Density Regions
Sam Goring, Tom Kuipers, Nicola Paoletti, David S. Watson · 19 de junio de 2026
Uncertainty quantification (UQ) is essential for reliable decision-making in safety-critical applications in probabilistic machine learning. For regression problems, dominant scalar UQ approaches - notably, those based on proper scoring rules - measure uncertainty via pointwise predictive risk. This…
- Convex training of Lipschitz-regularized shallow neural networks
Chao Yin, Antoine Lesage-Landry · 19 de junio de 2026
In this work, we introduce a training procedure for shallow neural networks that promotes robustness against adversarial attacks. We solve a non-convex Lipschitz-regularized training program by introducing a convex restriction that can be efficiently solved to global optimality. Our approach can be …
- Open Weight AI Models Require Proportional Evaluation Approaches
Patricia Paskov, Christopher Rodriguez, Sunishchal Dev, Stephen Casper · 19 de junio de 2026
Open-weight AI models (OWMs), or models released with publicly-available weights, are distributing rapidly and approaching the performance levels of leading closed-weight AI models (CWMs). While OWMs offer substantial scientific and economic benefits, their release introduces distinct risk factors f…
- Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation
Ahmad Farooq, Kamran Iqbal · 19 de junio de 2026
Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets. We present the first end-to…
- Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
Jinhan Li, Kexian Tang, Yihan Xu, Zhuorui Ye, Kaifeng Lyu · 18 de junio de 2026
To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how to push safety interventions earlier into the pretraining stage, primarily by filtering unsafe data or rewriting it into safer forms. We argue that pretraining-stage alignment should go beyond making…
- Lifecycle-Aware Dynamic Analysis for Secure ML Model Execution
Gabriele Digregorio, Marco Di Gennaro, Francesco Pastore, Stefano Zanero, Stefano Longari, Michele Carminati · 18 de junio de 2026
The growing reliance on pre-trained Machine Learning (ML) models has introduced new attack surfaces. Recent vulnerabilities demonstrate that malicious behavior can be embedded within model artifacts, often bypassing existing defenses. Current model-scanning solutions primarily rely on static, format…
- Generalised Eigenvalue Geometry of Semantic Adversarial Attacks
Martin Anthony, Kaveh Salehzadeh Nobari · 18 de junio de 2026
Recent empirical work shows that semantically equivalent paraphrases can fool financial sentiment classifiers: although a paraphrase remains close to the original under a strong reference embedding, it may shift the target model's representation enough to change the predicted class. Existing robustn…
- EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
Minseo Kim, Minjae Lee, Seunghyuk Oh, Kevin Galim, Donghoon Kim, Coleman Hooper, Harman Singh, Amir Gholami, Hyung Il Koo, Wonjun Kang · 18 de junio de 2026
Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling strong reasoning and agentic capabilities. However, rollout generation remains a dominant latency bottleneck because autoregressive sampling decodes responses sequentially and a small number of long-tai…
- Quantification of Uncertainty with Adversarial Models in Medical Image Segmentation
Hana Jebril, Thomas Pinetz, G\"unter Klambauer, Hrvoje Bogunovi\'c · 18 de junio de 2026
Reliable pixel-level uncertainty quantification holds the potential to transform clinical workflows by enabling high-fidelity longitudinal monitoring and distinguishing true pathological changes from artifacts. Ideally, these models provide the stability required for critical treatment planning and …
- TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning
Zijie Meng, Ziwei Li, Yufei Liu, Zhiyu Li, Jiyuan Liu, Wenhua Nie, Bingcai Wei, Miao Zhang · 18 de junio de 2026
Safe coordination in networked cyber-physical systems forces learning algorithms to simultaneously handle hybrid discrete-continuous actions, hard training-time safety constraints, and physics-governed dynamics. We show that these three features form a directed cycle of biases that defeats any naive…
- Machine Unlearning for the XGBoost Model with Network Intrusion Datasets
Diana Magalh\~aes, Eva Maia, Jo\~ao Vitorino, Isabel Pra\c{c}a · 18 de junio de 2026
Machine Unlearning (MU) has emerged as an important technique for removing specific data points from trained models without requiring full retraining. However, most existing MU research focuses on deep learning and image data, leaving a gap in the domain of network intrusion detection, which relies …
- Stealthy World Model Manipulation via Data Poisoning
Yibin Hu, Xiaolin Sun, Zizhan Zheng · 18 de junio de 2026
Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments. However, the process of updating world models from collected experience creates a training-time attack surface: adversarially poisoned fine-tuning trajectories can manipulate t…
- Some Complexity Results for Robustness Verification for Binarized Neural Networks
Harshit Goyal, Sudakshina Dutta · 18 de junio de 2026
This paper studies the computational complexity of verification problems for Binarized Neural Networks (BNNs), where activations (and sometimes weights) are binary. We analyze two problems: satisfiability and robustness under uniform image occlusion. We show that BNN satisfiability is NP-complete vi…
- Detecting Hidden ML Training With Zero-Overhead Telemetry
Robi Rahman, Sabiha Tajdari · 18 de junio de 2026
Hardware-enabled monitoring of GPU workloads underpins many proposals for AI compute governance, but if developers can defeat monitoring mechanisms, such schemes are unworkable. We evaluate the adversarial robustness of GPU workload classification using only zero-overhead, privacy-preserving NVML te…
- SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Mingyue Cui, Linghui Shen, Xingyi Yang · 18 de junio de 2026
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that identified "unsafe" SAE features serve as actionable handles for monitoring and intervention. In this paradigm, clamping…
- Signature filtering: a lightweight enhancement for statistical watermark detection in large language models
Chih-Duo Hong, Yen-Pang Chen, Fang Yu · 18 de junio de 2026
Statistical watermarks help organizations attribute large language model (LLM) outputs, yet existing detectors often struggle when watermark signals are weak, texts are repetitive, or watermarks are edited. We propose signature filtering, a detection-time module that enhances watermark detection wit…
