Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3169 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
Changyue Li, Jiaying Li, Youliang Yuan, Jiaming He, Zhicong Huang, Pinjia He · 30 de enero de 2026
Multimodal Large Language Models (MLLMs) are increasingly deployed in stateless systems, such as autonomous driving and robotics. This paper investigates a novel threat: Semantic-Aware Hijacking. We explore the feasibility of hijacking multiple stateless decisions simultaneously using a single uni…
- Questioning the Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better
Yizhou Min, Yizhou Lu, Lanqi Li, Zhen Zhang, Jiaye Teng · 30 de enero de 2026
Conformal prediction (CP) has become a cornerstone of distribution-free uncertainty quantification, conventionally evaluated by its coverage and interval length. This work critically examines the sufficiency of these standard metrics. We demonstrate that the interval length might be deceptively impr…
- XFACTORS: Disentangled Information Bottleneck via Contrastive Supervision
Alexandre Myara, Nicolas Bourriez, Thomas Boyer, Thomas Lemercier, Ihab Bendidi, Auguste Genovesio · 30 de enero de 2026
Disentangled representation learning aims to map independent factors of variation to independent representation components. On one hand, purely unsupervised approaches have proven successful on fully disentangled synthetic data, but fail to recover semantic factors from real data without strong indu…
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
Shahriar Kabir Nahin, Hadi Askari, Muhao Chen, Anshuman Chhabra · 30 de enero de 2026
Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best output. A tacit premise behind TTS is that sufficiently diverse candidate pools enhance reliability. In this work, we show that this assumption in TTS introduces…
- Making Foundation Models Probabilistic via Singular Value Ensembles
Mehmet Ozgur Turkoglu, Dominik J. M\"uhlematter, Alexander Becker, Konrad Schindler, Helge Aasen · 30 de enero de 2026
Foundation models have become a dominant paradigm in machine learning, achieving remarkable performance across diverse tasks through large-scale pretraining. However, these models often yield overconfident, uncalibrated predictions. The standard approach to quantifying epistemic uncertainty, trainin…
- Optimal Transport-Induced Samples against Out-of-Distribution Overconfidence
Keke Tang, Ziyong Du, Xiaofei Wang, Weilong Peng, Peican Zhu, Zhihong Tian · 30 de enero de 2026
Deep neural networks (DNNs) often produce overconfident predictions on out-of-distribution (OOD) inputs, undermining their reliability in open-world environments. Singularities in semi-discrete optimal transport (OT) mark regions of semantic ambiguity, where classifiers are particularly prone to unw…
- Position: Certifiable State Integrity in Cyber-Physical Systems -- Why Modular Sovereignty Solves the Plasticity-Stability Paradox
Enzo Nicol\'as Spotorno, Ant\^onio Augusto Medeiros Fr\"ohlich · 30 de enero de 2026
The machine learning community has achieved remarkable success with universal foundation models for time-series and physical dynamics, largely overcoming earlier approximation barriers in smooth or slowly varying regimes through scale and specialized architectures. However, deploying these monolithi…
- FIT: Defying Catastrophic Forgetting in Continual LLM Unlearning
Xiaoyu Xu, Minxin Du, Kun Fang, Zi Liang, Yaxin Xiao, Zhicong Huang, Cheng Hong, Qingqing Ye, Haibo Hu · 30 de enero de 2026
Large language models (LLMs) demonstrate impressive capabilities across diverse tasks but raise concerns about privacy, copyright, and harmful materials. Existing LLM unlearning methods rarely consider the continual and high-volume nature of real-world deletion requests, which can cause utility degr…
- StepShield: When, Not Whether to Intervene on Rogue Agents
Gloria Felicia (University of Virginia), Michael Eniolade (University of the Cumberlands), Jinfeng He (Cornell University), Zitha Sasindran (Indian Institute of Science Bangalore), Hemant Kumar (University of Arizona), Milan Hussain Angati (California State University Northridge), Sandeep Bandarupalli (University of Cincinnati) · 30 de enero de 2026
Existing agent safety benchmarks report binary accuracy, conflating early intervention with post-mortem analysis. A detector that flags a violation at step 8 enables intervention; one that reports it at step 48 provides only forensic value. This distinction is critical, yet current benchmarks cannot…
- Untargeted Jailbreak Attack
Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu, Xiaojun Jia, Di Wang, Zhan Qin, Kui Ren · 30 de enero de 2026
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However, restricting the objective as inducing fixed targets inherently constrains the adversarial search space, limiting the ov…
- Investigation into using stochastic embedding representations for evaluating the trustworthiness of the Fr\'{e}chet Inception Distance
Ciaran Bench, Vivek Desai, Carlijn Roozemond, Ruben van Engen, Spencer A. Thomas · 30 de enero de 2026
Feature embeddings acquired from pretrained models are widely used in medical applications of deep learning to assess the characteristics of datasets; e.g. to determine the quality of synthetic, generated medical images. The Fr\'{e}chet Inception Distance (FID) is one popular synthetic image quality…
- Breaking the Overscaling Curse: Thinking Parallelism Before Parallel Thinking
Yiming Wang, Zhuosheng Zhang, Rui Wang · 30 de enero de 2026
Parallel thinking enhances LLM reasoning by multi-path sampling and aggregation. In system-level evaluations, a global parallelism level N is allocated to all samples, typically set large to maximize overall dataset accuracy. However, due to sample heterogeneity, some samples can achieve comparable …
- Adversarial Vulnerability Transcends Computational Paradigms: Feature Engineering Provides No Defense Against Neural Adversarial Transfer
Achraf Hsain, Ahmed Abdelkader, Emmanuel Baldwin Mbaya, Hamoud Aljamaan · 30 de enero de 2026
Deep neural networks are vulnerable to adversarial examples--inputs with imperceptible perturbations causing misclassification. While adversarial transfer within neural networks is well-documented, whether classical ML pipelines using handcrafted features inherit this vulnerability when attacked via…
- ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
Xingwei Lin, Wenhao Lin, Sicong Cao, Jiahao Yu, Renke Huang, Lei Xue, Chunming Wu · 30 de enero de 2026
Multi-turn jailbreak attacks have emerged as a critical threat to Large Language Models (LLMs), bypassing safety mechanisms by progressively constructing adversarial contexts from scratch and incrementally refining prompts. However, existing methods suffer from the inefficiency of incremental contex…
- Hardware-Triggered Backdoors
Jonas M\"oller, Erik Imgrund, Thorsten Eisenhofer, Konrad Rieck · 30 de enero de 2026
Machine learning models are routinely deployed on a wide range of computing hardware. Although such hardware is typically expected to produce identical results, differences in its design can lead to small numerical variations during inference. In this work, we show that these variations can be explo…
- Evaluating Prediction Uncertainty Estimates from BatchEnsemble
Morten Bl{\o}rstad, Herman Jangsett Mostein, Nello Blaser, Pekka Parviainen · 30 de enero de 2026
Deep learning models struggle with uncertainty estimation. Many approaches are either computationally infeasible or underestimate uncertainty. We investigate \textit{BatchEnsemble} as a general and scalable method for uncertainty estimation across both tabular and time series tasks. To extend BatchE…
- Are We Truly Forgetting? A Critical Re-examination of Machine Unlearning Evaluation Protocols
Yongwoo Kim, Sungmin Cha, Donghyun Kim · 30 de enero de 2026
Machine unlearning is a process to remove specific data points from a trained model while maintaining the performance on the retain data, addressing privacy or legal requirements. Despite its importance, existing unlearning evaluations tend to focus on logit-based metrics under small-scale scenarios…
- On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
Xinwei Zhang, Hangcheng Liu, Li Bai, Hao Wang, Qingqing Ye, Tianwei Zhang, Haibo Hu · 30 de enero de 2026
Visual token compression is widely used to accelerate large vision-language models (LVLMs) by pruning or merging visual tokens, yet its adversarial robustness remains unexplored. We show that existing encoder-based attacks can substantially overestimate the robustness of compressed LVLMs, due to an …
- OpenSec: Measuring Incident Response Agent Calibration Under Adversarial Evidence
Jarrod Barnes · 30 de enero de 2026
As large language models improve, so do their offensive applications: frontier agents now generate working exploits for under $50 in compute (Heelan, 2026). Defensive incident response (IR) agents must keep pace, but existing benchmarks conflate action execution with correct execution, hiding calibr…
- Lossy Common Information in a Learnable Gray-Wyner Network
Anderson de Andrade, Alon Harell, Ivan V. Baji\'c · 30 de enero de 2026
Many computer vision tasks share substantial overlapping information, yet conventional codecs tend to ignore this, leading to redundant and inefficient representations. The Gray-Wyner network, a classical concept from information theory, offers a principled framework for separating common and task-s…
- TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
Chuancheng Shi, Shangze Li, Wenjun Lu, Wenhua Wu, Cong Wang, Zifeng Cheng, Fei Shen, Tat-Seng Chua · 30 de enero de 2026
Despite their capabilities, large foundation models (LFMs) remain susceptible to adversarial manipulation. Current defenses predominantly rely on the "locality hypothesis", suppressing isolated neurons or features. However, harmful semantics act as distributed, cross-layer circuits, rendering such l…
- Small models, big threats: Characterizing safety challenges from low-compute AI models
Prateek Puri · 30 de enero de 2026
Artificial intelligence (AI) systems are revolutionizing fields such as medicine, drug discovery, and materials science; however, many technologists and policymakers are also concerned about the technology's risks. To date, most concrete policies around AI governance have focused on managing AI risk…
- Defining Operational Conditions for Safety-Critical AI-Based Systems from Data
Johann Christensen, Elena Hoemann, Frank K\"oster, Sven Hallerbach · 30 de enero de 2026
Artificial Intelligence (AI) has been on the rise in many domains, including numerous safety-critical applications. However, for complex systems found in the real world, or when data already exist, defining the underlying environmental conditions is extremely challenging. This often results in an in…
- DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
Yisheng Zhong, Zhengbang Yang, Zhuangdi Zhu · 30 de enero de 2026
LLM unlearning is a technique to remove the impacts of undesirable knowledge from the model without retraining from scratch, which is indispensable towards trustworthy AI. Existing unlearning methods face significant limitations: conventional tuning-based unlearning is computationally heavy and pron…
- Shaping capabilities with token-level data filtering
Neil Rathi, Alec Radford · 30 de enero de 2026
Current approaches to reducing undesired capabilities in language models are largely post hoc, and can thus be easily bypassed by adversaries. A natural alternative is to shape capabilities during pretraining itself. On the proxy task of removing medical capabilities, we show that the simple interve…
