Physical Sciences › Computer Science › Artificial Intelligence
Adversarial Robustness in Machine Learning
3164 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Credibility-Weighted Pricing of Autonomous Vehicle Liability Under Operational Design Domain Shift
Doyeon Jang · 17 de junio de 2026
Automated Driving System deployments create a foundational ratemaking challenge: sparse experience, shifting operational design domains, and non-stationary risk across software releases. We propose a hierarchical Bayesian credibility framework pooling across cities, software versions, and territorie…
- LineageMark: Multi-user White-box Watermarking for Contribution Tracing in Model Derivation Chains
Bingxue Zhang, Xiaofeng Xu, Feida Zhu · 17 de junio de 2026
In open large language model (LLM) ecosystems, models are frequently adapted across multiple domains and applications, forming multi-stage derivation chains. Consequently, tracking and verifying historical contributions is essential for model provenance and intellectual property protection. However,…
- Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models
Ramprasath Ganesaraja, Sahil Dilip Panse, Swathika N · 17 de junio de 2026
State Space Models (SSMs) such as Mamba-2 offer linear-time inference but their memory footprint limits edge deployment. Prior ternary SSM work (Slender-Mamba) trains from scratch on 150B tokens; we show a pretrained checkpoint suffices, reducing the marginal token budget by 1,000x. Using grouped qu…
- Adversarial Attacks Leverage Interference Between Features in Superposition
Edward Stevinson, Lucas Prieto, Melih Barsbey, Tolga Birdal · 17 de junio de 2026
Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decision boundary structure, but none provides a representation-level mechanism that explains why specific perturbations succee…
- Theoretical Grounding of Out-Of-Distribution Detection With Reinforcement Learning Optimizer
Salimeh Sekeh, Xin Zhang · 17 de junio de 2026
Out-of-distribution (OOD) detection in dynamic open-world environments requires a model to continually adapt to evolving data distributions while generalizing to covariate-shifted inputs and rejecting semantic-shifted OOD examples. Most existing OOD detection methods optimize only the current-step o…
- Structured Adversarial Camouflage via Voronoi Diagrams
Jens Bayer, Stefan Becker, David M\"unch, Michael Arens, J\"urgen Beyerer · 17 de junio de 2026
Pixel-wise adversarial patches are computationally heavy and often visually detectable, limiting utility in security-critical systems. We present adversarial Voronoi camouflage that optimizes only seed-point locations under fixed, printable palettes using a soft assignment, producing structured, spl…
- Trust-Aware Multi-Agent Traceability: Confidence-Calibrated Knowledge Graphs for Consistent Software Artifact Management
Mohamed Essam, Kareem Wael, Azza Hassan, Ahmed Haitham, Mahmoud Soliman, Samer Saber, Ibrahim Habib · 17 de junio de 2026
Multi-agent AI systems are increasingly used to automate software engineering tasks including requirements analysis, architecture design, test generation, and traceability linking. When these agents operate as a sequential pipeline over shared software artifacts, errors and low-confidence decisions …
- S4oP: Operator-level Pruning of Structured State Space Models for Resource-Constrained Devices
Marco Deano, Filippo Ziche, Nicola Bombieri · 17 de junio de 2026
Structured State Space Models (SSMs), including the S4 and S4D architectures, have recently emerged as powerful alternatives to attention-based models for capturing long-range dependencies in sequential data. Despite their strong empirical performance, deploying these models in time- and resource-co…
- How Inference Compute Shapes Frontier LLM Evaluation
Jessica McFadyen, Ole Jorgensen, Harry Coppock, Kevin Wei, Cozmin Ududec · 17 de junio de 2026
AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still …
- MorphStrata: Layer-Specific Perturbations for Generating Morphence Students in Time-Series Moving Target Defense
Abhishek Bhardwaj, Arnav Doshi, Anusri Nagarajan, Thanh Quynh Nhu Ta, Mohammad Masum, Robert Chun, Jaydip Sen, Saptarshi Sengupta · 17 de junio de 2026
Time-series forecasting models remain vulnerable to gradient-based adversarial attacks while existing defense mechanisms typically incur a trade-off in robustness for bounded response and compute cost. The problem is pronounced in Moving Target Defense where maintaining multiple randomized model ins…
- A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models
Nicola Franco · 17 de junio de 2026
We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundred…
- TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations
Rutger Hendrix, Leonardo G. Russo, Concetto Spampinato, Matteo Pennisi, Giovanni Bellitto · 17 de junio de 2026
The demand for privacy-compliant AI has amplified the need for machine unlearning; yet, existing retraining or distillation-based methods remain unverifiable and computationally costly. We introduce TrustErase, a verifiable, data-free unlearning framework leveraging passport-embedded representations…
- Loss Landscape Poisoning: Targeted Extraction of Unseen Training Data from LLMs
Md Abdullah Al Mamun, Ngoc Phu Doan, Pedram Zaree, Ihsen Alouani, Nael Abu-Ghazaleh · 17 de junio de 2026
Large Language Models are increasingly trained on proprietary or sensitive data, from private healthcare and financial records to user conversations containing secrets. Ensuring the privacy of such data against extraction attacks has become a central concern. In this paper, we ask whether an attacke…
- AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor
Ning Ni, Yingjie Lao · 17 de junio de 2026
Large language models (LLMs) outperform earlier architectures on generative inference and long-context tasks, but their large size introduces significant challenges in memory usage, energy cost, and on-device deployment. Since scaling pre-trained language models improves downstream capability \cite{…
- DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
Siyi Li, Chunyu Sun, Jiahao Zhang, Yuchen Kang, Wuliang Wang, Yu Qiu, Rui Jiang, Haitao Cui, Jie Chen · 17 de junio de 2026
Evaluating a Physical AI stack spans operators that differ by more than three orders of magnitude -- from a single foundation-model decoding step to thousands of physics ticks of whole-body control -- varying orthogonally in modality, reward semantics, and resource profile. No existing framework spa…
- TNODEV: Toolbox for Neural ODE Verification
Abdelrahman Sayed Sayed, Pierre-Jean Meyer, Mohamed Ghazel · 16 de junio de 2026
Neural ordinary differential equations (neural ODE) have started to appear in safety critical settings such as continuous-time controllers for cyber-physical systems and classifiers integrated into automated decision pipelines, raising the question of whether their behavior can be formally verified.…
- Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains
Jinlong Yang · 16 de junio de 2026
Large language model (LLM) systems increasingly use uncertainty signals to allocate limited computation across verification, test-time scaling, tool execution, and other selective-compute decisions. Such policies rely on a \emph{global signal comparability assumption}: equal scores should carry comp…
- Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models
Ke Miao, Jiaxin Li, Hongliang Chen, Yuke Hu, Zhan Qin · 16 de junio de 2026
While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data annotation for safety alignment. However, we observe that LRMs can inher…
- Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades
Zhongye Liu, Yaopei Zeng, Yurui Chang, Lu Lin · 16 de junio de 2026
While multimodal large language models (MLLMs) have shown strong visual reasoning abilities, serving a large model for every query is computationally expensive. MLLM cascades mitigate this cost by first querying a weak but cheaper model and deferring to a strong model when the weak model's output is…
- TriAdReview: Triangular Adversarial Review Architecture for Multi-Model Technical Document Generation
Zhiqiang Zhou, Junliang Dai, Xu Ling · 16 de junio de 2026
Large language models (LLMs) are increasingly used for technical document generation, yet single-model outputs often suffer from over-engineering, security blind spots, and incomplete coverage. We propose TriAdReview, a triangular adversarial review architecture that employs two independent reviewer…
- Model Stealing Through the Lens of Model Multiplicity
Eliott Baltz, Satoshi Hara, Ulrich A\"ivodji · 16 de junio de 2026
Model stealing attacks, where adversaries create high-fidelity surrogate models, are a significant threat to the intellectual property of machine learning services. Conventional wisdom suggests these surrogates could provide adversaries with economic leverage comparable to the original service provi…
- When Confidence Lacks Concepts: Interpretable OOD Detection via Representation Perturbations
Anju Chhetri, Pratik Shrestha, Ramesh Rana, Prashnna Gyawali, Binod Bhattarai · 16 de junio de 2026
Deep neural networks have achieved remarkable performance across medical imaging tasks, yet their tendency to overgeneralize under distributional shifts poses a major obstacle to safe clinical deployment. Out-of-Distribution (OOD) detection methods aim to mitigate this risk, but most existing approa…
- Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
\"Omer Veysel \c{C}a\u{g}atan, Xuandong Zhao · 16 de junio de 2026
Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a central challenge in AI safety. Yet most known instances have been discovered post hoc in frontier systems where controlled study is impractical. We adapt the AI Safet…
- GRAPE: Guided Parameter-Space Evolution for Compact Adversarial Robustness
Zhiyuan Ye (University of Science and Technology of China), Xiangyu Zhou (China Mobile), Ji Qi (China Mobile), Hao Zhang (University of Science and Technology of China), Yi Zhou (China Mobile) · 16 de junio de 2026
Adversarial Training (AT) improves neural network robustness, but most methods train a fixed parameter space from the start. This paper asks whether the order in which parameters become optimizable can affect the final robust solution, even when the final architecture or computation budget is contro…
- AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
Faouzi El Yagoubi, Godwin Badu-Marfo, Ranwa Al Mallah · 16 de junio de 2026
Multi-agent Large Language Model (LLM) systems create privacy risks that current output-only benchmarks cannot measure. When agents coordinate on tasks, sensitive data may pass through inter-agent messages, shared memory, and tool arguments, all pathways that final-output audits typically do not ins…
