Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3.164 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian, Wei He, Hang Zhou, Jianyuan Guo, Hailin Hu, Lin Ma, Chao Xu, Guohao Dai, Lixue Xia, Yunchao Wei, Yunhe Wang, Yu Wang · 11. Juni 2026
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring. We introduce Claw-SWE-…
- Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers
Jun Wen Leong · 11. Juni 2026
We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a classifier has moved out of distribution. Upon detection, a conformal abstention layer adapts decision thresholds to recover a target error rate eps…
- From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data
Md. Rejaul Korim Sadi, Toufiqur Rahman Tasin, Golam Mostofa Naeem · 11. Juni 2026
Large language models hallucinate--producing fluent, confident, factually wrong outputs--with a consistency that persists across generations and scales. Existing taxonomies classify hallucination by output type, distinguishing intrinsic from extrinsic failures and faithfulness from factuality diverg…
- When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis
Orion Reblitz-Richardson · 11. Juni 2026
Standard linear probing declares a property "encoded" when a classifier on hidden states achieves high accuracy. The protocol works well on a snapshot but breaks across pre-training: probe accuracy saturates within the first few thousand steps, leaving most of training invisible to the instrument. W…
- T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking
Jian-Ping Mei, Weibin Zhang, Ao Yao, Tiantian Zhu, Jie Xiao · 11. Juni 2026
Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against various post-processing attacks on the watermarked model. Model extraction attacks …
- ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing
Chirag Chawla, Pratinav Seth, Vinay Kumar Sankarapu · 11. Juni 2026
Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language. Existing inference-time defenses that mix logits from a safe anchor model require both models to share a vocabulary, which rules them out for the cro…
- Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents
Frank Xiao, Mary Phuong · 11. Juni 2026
Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and untrusted models may render trusted models unreliable monitors. We introduce \emph{bootstrapped monitoring}, a protocol that addresses this by inserti…
- Auditing CoT Answer-Hijack Patches: Source-Control Certificates with Type-I Guarantees
Jianwei Tai · 11. Juni 2026
Chain-of-thought (CoT) answer-hijack templates can flip the final numeric answer of a 7B-8B language model on GSM8K or MATH-500 even when the visible reasoning trace looks fluent. Activation patching is the standard probe for locating where this hijack can be undone, and a successful clean-source pa…
- A prior-free blind detection of information leakage from model predictions
Laurence A. Jacobs · 11. Juni 2026
Data leakage -- contamination of a model with information unavailable at baseline -- is the dominant reproducibility failure in machine-learning-based science, yet detection tools require training code, external data, or domain expertise. None operates on the artifact an auditor most often holds: th…
- Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuous Data Summarization
Yuefang Lian, Longkun Guo, Zhongrui Zhao, Zhigang Lu, Yanan Cai, Shuchao Pang, Dachuan Xu, Jason Xue · 11. Juni 2026
Trustworthy AI requires reliable data-processing pipelines, not only robust downstream predictive models. As an upstream component, data summarization determines which information is retained and passed to subsequent learning or decision modules. Therefore, adversarial perturbations to the summariza…
- Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models
Malikeh Ehghaghi, Bogl\'arka Ecsedi, Marsha Chechik, Colin Raffel · 11. Juni 2026
Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly. In practice, the computational expense of different attack strategies can vary by orders of magnitude. Consequen…
- JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization
Ge Shi, Jun Yin, Donglin Xie, Fangyi Liu, Yucan Li, Menglin Liu · 11. Juni 2026
Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can adapt but often relies on low-level mutations that require many …
- VIA-SD: Verification via Intra-Model Routing for Speculative Decoding
Yuchen Xian, Yang He, Yunqiu Xu, Yi Yang · 11. Juni 2026
Speculative decoding (SD) addresses the high inference costs of LLMs by having lightweight drafters generate candidates for large verifiers to validate in parallel. Existing draft-verify methods use binary decisions: accept or fully recompute. Yet we find that many rejected tokens can be verified co…
- Reinforcement Learning Disrupts Gradient-Based Adversarial Optimization
Xinhai Zou, Chang Zhao, Alireza Aghabagherloo, Dave Singel\'ee, Robin Degraeve, Bart Preneel · 11. Juni 2026
Gradient-based adversarial attacks remain a dominant threat to deep neural networks (DNNs), as they exploit gradient information to efficiently optimize adversarial perturbations. To address this, we investigate whether reinforcement learning (RL) training can disrupt the gradient structure used by …
- Beyond Dark Knowledge: Mixup-Based Distillation for Reliable Predictions
Jos\'e Medina, Paul Honeine, Abdelaziz Bensrhair, Amnir Hadachi · 11. Juni 2026
Knowledge Distillation (KD) and mixup have proven effective at inducing smoothness in class boundaries; KD captures inherent class relationships in probability distributions, and mixup enforces them through convex combinations of inputs. Their interaction, however, remains poorly understood, particu…
- Diffusion-based Cumulative Adversarial Purification for Vision Language Models
Jia Fu, Yongtao Wu, Yihang Chen, Kunyu Peng, Xiao Zhang, Volkan Cevher, Sepideh Pashami, Anders Holst · 11. Juni 2026
Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to adversarial perturbations poses a significant threat to their reliability in real-world applications. Despite often being imperceptible to humans, these perturbations can drastic…
- Adv-TGD: Adversarial Text-Guided Diffusion for Face Recognition Impersonation Attacks
Omid Ahmadieh, Nima Karimian · 11. Juni 2026
The widespread adoption of face recognition (FR) technologies raises serious privacy concerns, as facial data can be exploited without consent. To address this challenge, we propose Adv-TGD, a generative adversarial attack framework that synthesizes photorealistic faces capable of impersonating targ…
- Categorical Robustness Assessment for Machine Learning based Network Intrusion Detection Systems
Mayank Raj, Nathaniel D. Bastian, Lance Fiondella, Gokhan Kul · 11. Juni 2026
Network Intrusion Detection Systems (NIDS) heavily utlize Machine Learning (ML) but ML models can be manipulated via adversarial attacks. These attacks add carefully crafted perturbations to network traffic data that leads to misclassifications. While prior work has demonstrated adversarial vulnerab…
- MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation
Yukuan Zhang, Mengxin Zheng, Qian Lou · 11. Juni 2026
Repository-level benchmarks for evaluating Large Language Model (LLM) code repair on Secure Multi-Party Computation (MPC) software do not yet exist, and directly transplanting general-purpose benchmarks such as SWE-bench fails on three structural fronts: (i) MPC repositories are dominated by generic…
- Signed Compression Progress on a Sealed Audit is Goodhart-Resistant
Ayush Mittal, Dhruv Gupta · 11. Juni 2026
Compression progress is a long-standing proposal for intrinsic motivation: reward an agent when its world model becomes better at predicting or compressing experience. The folk claim is that this reward is "credible" because it is paid only for learning. We make this precise and prove it. If intrins…
- Can we trust our models? Epistemic calibration in second-order classification
Arthur Hoarau · 10. Juni 2026
Uncertainty estimation is critical for deploying machine learning models in high-stakes settings. However, classical calibration only assesses the reliability of predicted probabilities and does not evaluate whether epistemic uncertainty estimates are themselves trustworthy. This limitation is parti…
- SPACR: Single-Pass Adaptive Training of Uncertainty-Aware Conformal Regressors
Soundouss Messoudi, Sylvain Rousseau, S\'ebastien Destercke · 10. Juni 2026
Conformal Prediction (CP) provides robust uncertainty guarantees for predictive models, but is typically applied post hoc, which misaligns model training with the conformal goal of producing efficient (i.e, narrow) intervals. We propose SPACR (Single-Pass Adaptive Conformal Regressor), a novel metho…
- Gradient-Guided Reward Optimization for Inference-time Alignment
Hankun Lin, Ruqi Zhang · 10. Juni 2026
Ensuring the reliability of Large Language Models (LLMs) under distribution drift requires inference-time adaptation. While inference-time alignment methods such as Best-of-$N$ and rejection sampling are widely used, they frame the task as a sampling-intensive, reward-guided search, leading to two k…
- Robust Deep Reinforcement Learning Through Adversarial Attacks and Training : A Survey
Lucas Schott, Josephine Delas, Hatem Hajri, Elies Gherbi, Reda Yaich, Nora Boulahia-Cuppens, Frederic Cuppens, Sylvain Lamprier · 10. Juni 2026
Deep Reinforcement Learning (DRL) is a subfield of machine learning for training autonomous agents that take sequential actions across complex environments. Despite its significant performance in well-known environments, it remains susceptible to minor condition variations, raising concerns about it…
- Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction
Lijia Yu, Jiuxin Cao, Yuchen Qiang, Changhao Chen, Yifei Huang, Bo Liu · 10. Juni 2026
Adversarial examples reveal vulnerabilities in Vision-Language Pre-training (VLP) models and provide insights for improving robustness. A key property is cross-model transferability, which enables transfer-based black-box attacks. However, existing attacks often rely heavily on the surrogate model, …
